Time-motion vector prediction, sub-candidate search.

JP2026074004A5Pending Publication Date: 2026-07-17TENCENT AMERICA LLC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2026-01-08
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing video encoding techniques face inefficiencies in determining temporal motion vector predictor candidates, leading to suboptimal compression ratios and increased data requirements.

Method used

A method for processing video frames to identify and limit the number of temporal motion vector predictor candidates, utilizing a search mechanism to enhance diversity and improve coding efficiency by constructing an MVP list that includes both spatial and temporal candidates.

Benefits of technology

Enhances video encoding efficiency by reducing the data required for motion vector prediction, thereby improving compression ratios and reducing storage and bandwidth needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This invention provides a method and system for determining candidate time-motion vector predictors (TMVPs) for interpretation in video coding. [Solution] The disclosed method includes limiting the number of TMVP candidates in the motion vector predictor (MVP) list and provides various search mechanisms to promote diversity among MVP candidates between TMVP and other types of MVP candidates and to improve coding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Incorporation by Reference

[0001] This application claims the benefit of priority based on U.S. Non-Provisional Application No. 18 / 051,397, filed on October 31, 2022, entitled "Temporal Motion Vector Predictor Candidates Search", and U.S. Provisional Patent Application No. 63 / 402,386, filed on August 30, 2022, entitled "Temporal Motion Vector Predictor Candidates Improvement", and these are incorporated herein by reference in their entirety.

[0002]

[0002] The present disclosure generally relates to video encoding, and more particularly, to methods and systems for determining temporal motion vector predictor candidates for inter prediction in video encoding.

Background Art

[0003]

[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the inventors, as currently identified, is not to be regarded as prior art at the time of filing of this application, either explicitly or implicitly recognized as prior art to the present disclosure, in the same way as aspects of the description that may not be regarded as prior art at the time of filing of this application are set forth in this background section.

[0004]

[0004] Video encoding and decoding can be performed using interpicture prediction with motion compensation. Uncompressed digital video can consist of a series of pictures, each picture having spatial dimensions such as luminance samples of, for example, 1920 x 1080, and associated full or subsampled chromaticity samples. The series of pictures can have a fixed or variable picture rate (alternatively called frame rate), such as 60 pictures per second, i.e., 60 frames per second. Uncompressed video has bitrate requirements specific to streaming or data processing. For example, video with a pixel resolution of 1920 x 1080, a frame rate of 60 frames / second, and chroma subsampling 4:2:0 with 8 bits per pixel per color channel would require a bandwidth of nearly 1.5 Gbit / s. One hour of such video would require more than 600 GByte of storage space.

[0005]

[0005] One objective of video encoding and decoding may be to reduce redundancy in uncompressed input video signals through compression. Compression can, in some cases, help reduce the aforementioned bandwidth and / or storage space requirements by more than two orders of magnitude. Both lossless and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully preserved during encoding and therefore cannot be fully restored during decoding. When using lossy compression, the reconstructed signal may not be identical to the original signal, but even with some information loss, the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for the intended application. In the case of video, lossy compression is widely adopted in many applications. The amount of distortion that can be tolerated depends on the application. For example, users of certain consumer video streaming applications can tolerate more distortion than users of movie or television broadcasting applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances; generally, the greater the distortion tolerance, the more encoding algorithms that produce greater loss and a higher compression ratio become possible.

[0006]

[0006] Video encoders and decoders can utilize techniques from several broad categories and processes, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.

[0007]

[0007] Video codec techniques may include a technique known as intra coding. In intra coding, sample values ​​are represented without referencing samples or other data from a previously reconstructed reference picture. In some video codecs, the picture is spatially subdivided into blocks of samples. When all blocks of samples are coded in intra mode, this picture can be called an intra picture. Intra pictures, and derivatives of intra pictures such as independent decoder refresh pictures, can be used to reset the decoder state and are therefore available as the first picture in the coded video bitstream and video session, or as a still image. The samples of the intra-predicted blocks can then be converted to the frequency domain, and the conversion coefficients thus produced can be quantized before entropy coding. Intra prediction represents a technique that minimizes the sample values ​​in the domain before conversion. In some cases, the smaller the DC value after conversion and the smaller the AC coefficients, the fewer bits are required in a given quantization process size to represent the block after entropy coding.

[0008]

[0008] Traditional intra-encoding, such as that known from the MPEG-2 generation of encoding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to encode / decode blocks based on surrounding sample data and / or metadata acquired during the encoding and / or decoding of spatially adjacent blocks and that precede the data block being intra-encoded or decoded in the decoding order. Such techniques are hereafter referred to as “intra-prediction” techniques. Note that in at least some cases, intra-prediction uses only reference data from the current picture being reconstructed and not reference data from other reference pictures.

[0009]

[0009] There can be many different forms of intra-prediction. When two or more of these techniques are available in a given video coding technique, the technique in use can be called an intra-prediction mode. One or more intra-prediction modes may be provided in a particular codec. In a particular case, a mode may have submodes and / or be associated with various parameters, and the mode / submode information and intra-coding parameters of a video block can be coded individually or included together in a mode codeword. Which codeword should be used for a given combination of mode, submode, and / or parameters may affect the improvement of coding efficiency through intra-prediction, and therefore it is possible to use entropy coding techniques to convert the codeword into a bitstream.

[0010]

[0010] Certain intra-prediction modes were introduced in H.264, improved in H.265, and further improved in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and benchmark set (BMS). Generally, in the case of intra-prediction, it is possible to form a predictor block using the adjacent sample values ​​that become available. For example, the available values ​​of a particular set of adjacent samples along a particular direction and / or line may be copied into the predictor block. The criterion for the direction in use may be encodeable within the bitstream or may be predicted itself.

[0011]

[0011] Referring to Figure 1A, a subset of nine specified predictor directions from the 33 possible intra-predictor directions of H.265 (corresponding to 33 angular modes from the 35 intra-modes specified in H.265) is shown in the lower right. The point where the arrows converge (101) represents the predicted sample. The arrows indicate the direction in which adjacent samples are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from one or more adjacent samples located to the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more adjacent samples located to the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012]

[0012] Referring further to Figure 1A, a square block (104) of 4x4 samples (indicated by a thick dotted line) is shown in the upper left. The square block (104) contains 16 samples, each labeled with "S", the block's position in the Y dimension (e.g., row index), and the block's position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in block (104) in both the Y and X dimensions. Since the block is a 4x4 sample, S44 is in the lower right. Further examples of reference samples following a similar numbering system are shown. Reference samples are labeled with R, the sample's Y position (e.g., row index) relative to block (104), and its X position (column index). In both H.264 and H.265, the predicted samples that are close to and adjacent to the block being reconstructed are used.

[0013]

[0013] Intra-picture prediction of block 104 may begin by copying a reference sample value from an adjacent sample, depending on the signaled prediction direction. For example, suppose the encoded video bitstream includes signaling that indicates the prediction direction of arrow (102) for this block 104—that is, a sample is predicted from one or more prediction samples located to the upper right at a 45-degree angle from the horizontal. In such a case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Thus, sample S44 is predicted from reference sample R08.

[0014]

[0014] In certain cases, especially when the direction cannot be uniformly divided at 45 degrees, the values ​​of multiple reference samples may be combined, for example, through interpolation, in order to calculate a reference sample.

[0015]

[0015] The number of possible directions has increased as video coding techniques continue to develop. In H.264 (2003), for example, nine different directions are available for intra-prediction. This has increased to 33 in H.265 (2013), and JEM / VVC / BMS can support up to 65 directions as of the present disclosure. Demonstration studies have been conducted to help identify the optimal intra-prediction direction, and certain techniques in entropy coding may be used to encode these optimal directions with a small number of bits that accept a specific bit penalty for the direction. Furthermore, the direction itself can sometimes be predicted from similar directions used in the intra-prediction of adjacent decoded blocks.

[0016]

[0016] Figure 1B shows a schematic diagram (180) illustrating the 65 intra prediction directions by JEM to illustrate the increase in the number of prediction directions in various encoding techniques developed over time.

[0017]

[0017] The way in which bits representing intra-prediction directions are mapped to prediction directions in the encoded video bitstream may vary from video coding technique to video coding technique, and can range from, for example, a simple direct mapping of prediction directions to intra-prediction modes to codewords, to complex adaptable schemes with maximum likelihood modes, and similar techniques. However, in all cases, there may be specific directions for intro prediction that are statistically less likely to occur in the video content than certain other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technique, these less likely directions may be represented by more bits than the more likely directions.

[0018]

[0018] Interpicture prediction or interpretation may be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) may be spatially shifted in the direction indicated by a motion vector (hereinafter, MV: motion vector) and then used for predicting the newly reconstructed picture or a portion of a picture (e.g., a block). In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV may have two dimensions X and Y, or three dimensions, the third dimension indicating the reference picture in use (similar to the time dimension).

[0019]

[0019] In some video compression techniques, the current MV applicable to a particular area of ​​sample data is predictable from other MVs, such as other MVs relating to other areas of the sample data that are spatially adjacent to the area being reconstructed and precede the current MV in decoding order. Doing so can substantially reduce the total amount of data required to encode the MV by relying on redundancy removal in the correlated MV, thereby improving compression efficiency. For example, when encoding an input video signal derived from a camera (known as raw video), there is a statistical possibility that an area wider than the area to which a single MV is applicable moves in a similar direction in the video sequence, and therefore in some cases is predictable using similar motion vectors derived from the MVs of adjacent areas, so MV prediction can work effectively. This makes the actual MV of a given area similar to or identical to the MV predicted from the surrounding MVs. Such an MV may then be represented by fewer bits than would be used if the MV were encoded directly rather than being predicted from adjacent MVs after entropy coding. In some cases, MV prediction can be an example of lossless compression of the signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself may be irreversible, for example, due to rounding errors when calculating predictors from several surrounding MVs.

[0020]

[0020] H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016) describes various MV prediction mechanisms. Among the many MV prediction mechanisms specified in H.265, the technique hereafter referred to as "spatial merging" will be described below.

[0021]

[0021] Specifically, referring to Figure 2, the current block (201) contains samples found by the encoder during the motion search process that are predictable from previous blocks of the same size but spatially shifted. Instead of directly encoding this MV, the MV can be derived from metadata associated with one or more reference pictures, such as the most recent reference picture (in decoding order), using the MV associated with one of the five surrounding samples represented as A0, A1, and B0, B1, B2 (202 through 206, respectively). In H.265, the MV prediction can use predictors from the same reference pictures used by adjacent blocks. [Overview of the project]

[0022]

[0022] This disclosure relates generally to video coding, and more particularly to methods and systems for determining candidates for temporal motion vector predictor (TMVP) in interpretation in video coding. For example, the methods of the disclosure include limiting the number of TMVP candidates in a motion vector predictor (MVP) list and providing various search mechanisms to promote diversity of MVP candidates between TMVP and other types of MVP candidates and to improve coding efficiency.

[0023]

[0023] An implementation example discloses a method for processing the current predicted block of the current frame in a video stream. The method may include determining that the current prediction block should be interpredicted by a reference block in a reference frame and that the motion vector of the current prediction block should be predicted by a reference motion vector; identifying a set of candidate prediction blocks in the current frame as a search pool for time motion vector predictor (TMVP) candidates for the current prediction block; searching the set of candidate prediction blocks in search order to identify up to N TMVP candidates for the current prediction block, wherein the search ends in response that N TMVP candidates have been identified, and N is a positive integer; constructing an MVP list indicating a set of motion vector predictor (MVP) candidates, wherein the set of MVP candidates includes a set of spatial MVP (SMVP) candidates and one or more of up to N TMVP candidates; extracting the MVP index of the current prediction block from the video stream; and identifying a reference motion vector for interpreting the current prediction block based on the extracted MVP index and MVP list.

[0024]

[0024] In the above implementation, the current prediction block belongs to the current superblock, the current superblock includes multiple prediction blocks, and the set of candidate prediction blocks includes at least a subset of the multiple prediction blocks of the current superblock.

[0025] In any one of the above implementation forms, a plurality of prediction blocks of the current super block form a prediction block array having a column dimension and a row dimension, and the reconstructed adjacent prediction blocks of the current super block are above the current super block in the row dimension and / or to the left of the current super block in the column dimension. A search order for searching a set of candidate prediction blocks for up to N TMVP candidates starts away from the reconstructed adjacent prediction blocks of the current super block along at least one of the column dimension and the row dimension of the prediction block array and then moves closer.

[0026] In any one of the above implementation forms, the search order for searching a set of candidate prediction blocks for up to N TMVP candidates may start from the bottom row of the prediction block array and proceed to the top row, and for each row, start from the rightmost prediction block and proceed to the leftmost prediction block, or start from the rightmost column of the prediction block array and proceed to the leftmost column, and for each column, start from the bottommost prediction block and proceed to the topmost prediction block, or search the prediction block array along a diagonal direction based on the column dimension and the row dimension, starting from the rightmost and bottommost prediction block of the prediction block array and proceeding towards the leftmost and topmost prediction block.

[0027] In any one of the above implementation forms, the set of candidate prediction blocks includes a subset of a plurality of prediction blocks having M prediction blocks skipped from the plurality of prediction blocks, where M is a non-negative integer.

[0028] In any one of the above implementation forms, the M prediction blocks skipped include at least the leftmost column of the prediction block array.

[0029]

[0029] In any one of the above implementation forms, the skipped M prediction blocks include at least the top row of the prediction block array.

[0030]

[0030] In any one of the above implementations, the skipped M prediction blocks include at least the top row and the leftmost column of the prediction block array.

[0031]

[0031] In any one of the above implementations, the set of candidate prediction blocks excludes any additional prediction blocks outside the current superblock.

[0032]

[0032] In any one of the above implementations, the set of candidate prediction blocks further includes at least one additional prediction block outside the current superblock.

[0033]

[0033] In any one of the above implementations, the method may further include searching a plurality of prediction blocks in the current superblock to identify up to N1 TMVP candidates, wherein N1 is a positive integer less than or equal to N, and ceasing to search the plurality of prediction blocks in the current superblock in response that N1 TMVP candidates have been identified from the plurality of prediction blocks.

[0034]

[0034] In any one of the above implementations, up to one TMVP candidate is identified from at least one additional prediction block outside the current superblock, and the method further includes stopping searching for at least one additional prediction block outside the current superblock in response to one TMVP candidate being identified from at least one additional prediction block outside the current superblock.

[0035]

[0035] In any one of the above implementation forms, the set of candidate prediction blocks includes a maximum of L prediction blocks from multiple prediction blocks, where L is a positive integer less than the total number of prediction blocks in the current superblock.

[0036]

[0036] In the above implementation, L=1, and the set of candidate prediction blocks includes only the rightmost and bottommost prediction block of the prediction block array. In the above implementation, a maximum of L prediction blocks are uniformly distributed among multiple prediction blocks of the current superblock.

[0037]

[0037] In any one of the above implementations, the position of the candidate prediction block set is determined by encoded information including the motion vectors of blocks spatially adjacent to the current superblock, or the number of SMVP candidates already included in the MVP list.

[0038]

[0038] In any one of the above implementation forms, the method further includes, in response to there being N1 SMVP candidates in the MVP list, further limiting the number of TMVP candidates in the MVP list to N2, where N1 and N2 are positive integers and N2 is less than or equal to N.

[0039]

[0039] In any one of the above implementation forms, the context for signaling the index in the MVP list associated with the current prediction block and the inter-prediction mode depends on whether at least N TMVPs are included in the MVP list.

[0040]

[0040] In any one of the above implementations, limiting the number of TMVP candidates in the MVP list to a maximum of N is in response to the current prediction block being encoded under bidirectional interprediction mode.

[0041]

[0041] Furthermore, aspects of the present disclosure provide electronic devices or apparatus including circuit equipment or processors configured to perform any of the above-described implementations.

[0042]

[0042] Furthermore, aspects of the present disclosure provide a non-temporary computer-readable medium that stores instructions causing an electronic device to implement any one of the above-described implementations when executed by the electronic device.

[0043]

[0043] Further features, properties, and various advantages of the subject matter of disclosure will become clearer from the following detailed description and accompanying drawings. [Brief explanation of the drawing]

[0044] [Figure 1A]

[0044] This is a schematic diagram of an exemplary subset of intra-predictive direction modes. [Figure 1B]

[0045] This is an illustrative diagram of the intra-prediction direction. [Figure 2]

[0046] This is a schematic diagram of the current block and its surrounding spatial merge candidates for motion vector prediction in one example. [Figure 3]

[0047] This is a schematic diagram of a simplified block diagram of a communication system (300) according to an embodiment. [Figure 4]

[0048] This is a schematic diagram of a simplified block diagram of a communication system (400) according to an embodiment. [Figure 5]

[0049] This is a schematic diagram of a simplified block diagram of a video decoder according to an embodiment. [Figure 6]

[0050] This is a schematic diagram of a simplified block diagram of a video encoder according to an embodiment. [Figure 7]

[0051] This is a block diagram of a video encoder according to another embodiment. [Figure 8]

[0052] This is a block diagram of a video decoder according to another embodiment. [Figure 9]

[0053] This is a diagram illustrating the coding block partitioning scheme according to an embodiment of the present disclosure. [Figure 10]

[0054] This is a diagram of another scheme for coded block partitioning according to an embodiment of the present disclosure. [Figure 11]

[0055] This is a diagram of another scheme for coded block partitioning according to an embodiment of the present disclosure. [Figure 12]

[0056] This is a diagram illustrating the partitioning of an example implementation of a base block into an encoded block using the example partitioning scheme. [Figure 13]

[0057] This is a diagram illustrating an example of a ternary partitioning scheme. [Figure 14]

[0058] This is a diagram illustrating an example of a quad-tree binary-tree coded block partitioning scheme. [Figure 15]

[0059] This diagram shows a scheme for partitioning an encoded block into a plurality of transform blocks according to an embodiment of the present disclosure, and the encoding order of the transform blocks. [Figure 16]

[0060] This diagram shows another scheme for partitioning an encoded block into a plurality of transform blocks according to an embodiment of the present disclosure, and the encoding order of the transform blocks. [Figure 17]

[0061] This is a diagram of another scheme for partitioning an encoded block into multiple transform blocks according to an embodiment of the present disclosure. [Figure 18]

[0062] This is a diagram illustrating the search process for candidate spatial motion vector predictors in the superblock. [Figure 19]

[0063] This is a diagram illustrating the process for determining the time-motion vector predictor of the current prediction block by linear projection. [Figure 20]

[0064] This is a diagram illustrating the internal and external predicted blocks of an example superblock. [Figure 21]

[0065] This is a diagram illustrating the generation of new motion vector predictor candidates, using a single interpretation block as an example. [Figure 22]

[0066] This is a diagram illustrating the generation of new motion vector predictor candidates, an example of a composite interpretation block. [Figure 23]

[0067] This is a diagram illustrating the example of the reference motion vector candidate bank update process. [Figure 24]

[0068] This is an example diagram showing the sequence for constructing a list of motion vector predictors. [Figure 25]

[0069] This diagram shows an example of a search sequence for identifying candidate time-motion vector predictors for the current superblock. [Figure 26] This diagram shows an example of a search sequence for identifying candidate time-motion vector predictors for the current superblock. [Figure 27] This diagram shows an example of a search sequence for identifying candidate time-motion vector predictors for the current superblock. [Figure 28] This diagram shows an example of a search sequence for identifying candidate time-motion vector predictors for the current superblock. [Figure 29] This diagram shows an example of a search sequence for identifying candidate time-motion vector predictors for the current superblock. [Figure 30] This diagram shows an example of a search sequence for identifying candidate time-motion vector predictors for the current superblock. [Figure 31]

[0070] This is a flowchart of the method according to the embodiments of this disclosure. [Figure 32]

[0071] This is a schematic diagram of a computer system according to an embodiment of the present disclosure. [Modes for carrying out the invention]

[0045]

[0072] Throughout this specification and the claims, terms may have subtle differences in meaning beyond their expressly stated meaning, or meaning implied or suggested in context. The phrases “in one embodiment” or “in some embodiments,” as used herein, do not necessarily refer to the same embodiment, and the phrases “in another embodiment” or “in other embodiments,” as used herein, do not necessarily refer to different embodiments. Similarly, the phrases “in one implementation” or “in some implementations,” as used herein, do not necessarily refer to the same implementation, and the phrases “in another implementation” or “in other implementations,” as used herein, do not necessarily refer to different implementations. For example, the claimed subject matter is intended to include, in whole or in part, combinations of exemplary embodiments / implementations.

[0046]

[0073] In general, technical terms may be understood at least partially from their usage in context. For example, terms such as “and,” “or,” or “and / or,” as used herein, may have various meanings that may at least partially depend on the context in which such terms are used. Typically, when “or” is used to relate a list such as A, B, or C, it is intended to mean A, B, and C in an inclusive sense, as well as A, B, or C in an exclusive sense. Furthermore, terms such as “one or more” or “at least one,” as used herein, may be used at least partially, depending on the context, to describe any feature, structure, or characteristic in a singular sense, or to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms such as “a,” “an,” or “the” may also be understood, at least partially, depending on the context, to convey a singular usage or a plural usage. Furthermore, the terms “based on” or “determined by” may be understood not necessarily as intended to convey an exclusive set of elements, but rather, again, at least in part depending on the context, allow for the presence of additional elements that are not necessarily explicitly stated. Figure 3 illustrates a simplified block diagram of a communication system (300) according to an embodiment of the present disclosure. The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, over a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected over the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) may perform one-way transmission of data. For example, terminal device (310) may encode video data (for example, a stream of video pictures captured by terminal device (310)) for transmission to other terminal devices (320) over the network (350).The encoded video data can be transmitted in the form of one or more encoded video bitstreams. A terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to restore a video picture, and display the video picture according to the restored video data. One-way data transmission may be performed in media serving applications, etc.

[0047]

[0074] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which may be performed, for example, during a video conferencing application. In the case of bidirectional data transmission, in one example, each terminal device of terminal devices (330) and (340) may encode video data (for example, a stream of video pictures captured by the terminal device) for transmission to other terminal devices of terminal devices (330) and (340) over the network (350). Each terminal device of terminal devices (330) and (340) may also receive encoded video data transmitted by other terminal devices of terminal devices (330) and (340), decode the encoded video data to restore video pictures, and display the video pictures on an accessible display device, depending on the restored video data.

[0048]

[0075] In the example in Figure 3, terminal devices (310), (320), (330), and (340) may be implemented as servers, personal computers, and smartphones, respectively, but the applicability of the fundamental principles of this disclosure is not limited thereto. Embodiments of this disclosure may be implemented as desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, etc. Network (350) represents any number or type of network that transmits encoded video data between terminal devices (310), (320), (330), and (340), including, for example, wireline and / or wireless communication networks. Communication network (350)9 may exchange data on circuit-switched, packet-switched, and / or other types of channels. Typical networks include telecommunication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of network (350) may not be important to the operation of this disclosure unless expressly described herein.

[0049]

[0076] Figure 4 illustrates the arrangement of a video encoder and video decoder in a video streaming environment as an example of application to the subject matter of disclosure. The subject matter of disclosure may be similarly applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, and storage of compressed video on digital media such as CDs, DVDs, and memory sticks.

[0050]

[0077] The video streaming system may include a video capture subsystem (413) which may include a video source (401), such as a digital camera, for producing a stream (402) of uncompressed video pictures or images. In one example, the stream (402) of video pictures includes samples recorded by the digital camera of the video source 401. The stream (402) of video pictures is shown as a thick line to emphasize its large data volume compared to encoded video data (404) (or encoded video bitstream), but is processable by an electronic device (420) which includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof for enabling or implementing aspects of the subject matter of the disclosure as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is drawn as a thin line to highlight its smaller data volume compared to the stream of uncompressed video picture (402), but can be stored in a streaming server (405) for future use or directly in a downstream video device (not shown). One or more streaming client subsystems, such as client subsystems (406) and (408) in Figure 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include a video decoder (410) within, for example, an electronic device (430). The video decoder (410) decodes the incoming copy of the encoded video data (407) to produce an output stream (411) of a video picture that is uncompressed and can be rendered to a display (412) (e.g., a display screen) or other rendering device (not shown). The video decoder 410 may be configured to perform some or all of the various functions described in this disclosure.In some streaming systems, encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to specific video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, a video encoding standard under development is informally known as Versatile Video Coding (VVC). The subject of disclosure may also be used in the context of VVC and other video encoding standards.

[0051]

[0078] It should be noted that electronic devices (420) and (430) may include other components (not shown). For example, electronic device (420) may include a video decoder (not shown), and electronic device (430) may similarly include a video encoder (not shown).

[0052]

[0079] Figure 5 shows a block diagram of a video decoder (510) according to one of the embodiments of the present disclosure described below. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit device). The video decoder (510) may be used in place of the video decoder (410) in the example of Figure 4.

[0053]

[0080] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510). In the same or different embodiments, one encoded video sequence may be decoded at a time, and the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with multiple video frames or images. The encoded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device storing encoded video data or to a streaming source that transmitted the encoded video data. The receiver (531) may receive encoded video data with other data such as encoded audio data and / or auxiliary data streams, and the encoded video data may be forwarded to its respective processing circuit equipment (not shown). The receiver (531) may isolate the encoded video sequences from other data. To prevent network jitter, a buffer memory (515) may be placed between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In certain applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be outside the video decoder (510) (not shown) and isolated from the video decoder (510). In yet another application, for example to prevent network jitter, there may be a buffer memory (not shown) outside the video decoder (510), and for example to handle playback timing, there may be another additional buffer memory (515) inside the video decoder (510). When the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (515) may not be necessary or can be made smaller.For use on best-effort packet networks such as the Internet, a sufficiently large buffer memory (515) may be required, and its size may be quite large. Such buffer memory may be provided in an adaptable size and may be at least partially implemented in the operating system or a similar element (not shown) outside the video decoder (510).

[0054]

[0081] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the encoded video sequence. These categories of symbols may include information used to manage the operation of the video decoder (510), and possibly information for controlling a rendering device such as a display (512) (e.g., a display screen), which may or may not be an integrated part of an electronic device (530), but can be coupled to an electronic device (530) as shown in Figure 5. The control information for the rendering device may be in the form of supplemental enhancement information (SEI messages) or video usability information (VUI) parameter set fragments (not shown). The parser (520) may parse / entropy decode the encoded video sequence received by the parser (520). The entropy coding of the encoded video sequence can conform to video coding techniques or standards and can follow a variety of principles, including variable-length coding, Huffman coding, and context-independent or inconsistent arithmetic coding. The parser(520) may extract from the encoded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder, based on at least one parameter corresponding to a subgroup. Subgroups can include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), and so on. Furthermore, the parser(520) may extract encoded video sequence information such as transform coefficients (e.g., Fourier transform coefficients), quantization parameter values, and motion vectors.

[0055]

[0082] The parser (520) may perform an entropy decoding / parse operation on the video sequence received from the buffer memory (515) in order to produce a symbol (521).

[0056]

[0083] The reconstruction of the symbol (521) may require multiple different processing or functional units, depending on the type or part of the encoded video picture (inter and intra picture, inter and intra block, etc.) and other factors. The units required and how they are required may be controlled by subgroup control information parsed from the encoded video sequence by the parser (520). The flow of such subgroup control information between the parser (520) and the multiple processing or functional units described below is not depicted for brevity.

[0057]

[0084] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical implementations operating under commercial constraints, many of these functional units can interact closely with each other and be at least partially integrated. However, in order to clearly illustrate the various functions of the subject matter of the disclosure, the conceptual subdivision into functional units is adopted in this disclosure below.

[0058]

[0085] The first unit may include a scaler / inverse unit (551). The scaler / inverse unit (551) may receive quantized transformation coefficients and control information from the parser (520), the control information including information indicating which type of inverse transform to use, the block size, quantization factors / parameters, the quantization scaling matrix, and lie as symbol (521). The scaler / inverse unit (551) can output a block having sample values ​​that can be input to the aggregator (555).

[0059]

[0086] In some cases, the samples output from the scaler / inverse transform (551) may relate to intra-encoded blocks, i.e., blocks that can use prediction information from previously reconstructed portions of the current picture, rather than prediction information from previously reconstructed pictures. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate blocks of the same size and shape as the block being reconstructed, using surrounding block information that has already been reconstructed and stored in the current picture buffer (558). The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) may add the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a sample-by-sample basis.

[0060]

[0087] In other cases, the output samples of the scaler / inverse unit (551) can relate to intercoded and potentially motion-compensated blocks. In such cases, the motion-compensated prediction unit (553) can access the reference picture memory (557) to fetch samples to be used for interpicture prediction. After motion-compensating the fetched samples according to the symbols (521) relating to the blocks, these samples can be added by the aggregator (555) to the output of the scaler / inverse unit (551) to generate output sample information (the output of unit 551 is sometimes called residual samples or residual signals). The addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches the predicted samples are controllable by motion vectors and are available to the motion-compensated prediction unit (553) in the form of symbols (521) which can have, for example, X, Y components (shift) and a reference picture component (time). Motion compensation may further include interpolation for sample values ​​fetched from reference picture memory (557) when the accurate motion vector of a subsample is used, and may be further associated with a motion vector prediction mechanism, etc.

[0061]

[0088] The output samples of the aggregator (555) can undergo various loop filtering techniques in the loop filter unit (556). The video compression technique may include in-loop filtering techniques, which are controlled by parameters contained in the encoded video sequence (also called the encoded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), but can also be in response to metadata obtained during decoding of earlier portions (in decoding order) of the encoded picture or encoded video sequence, and in response to previously reconstructed and loop-filtered sample values. Several types of loop filters may be included as part of the loop filter unit 556 in various orders, as will be described in more detail below.

[0062]

[0089] The output of the loop filter unit (556) can be a sample stream that can be output to the rendering device (512) and stored in the reference picture memory (557) for use in future interpicture prediction.

[0063]

[0090] A particular encoded picture can be used as a reference picture for future interpicture prediction as soon as it is fully reconstructed. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (520)), the buffer (558) of the current picture can become part of the reference picture memory (557), and any unused current picture buffer can be reallocated before starting subsequent reconstructions of encoded pictures.

[0064]

[0091] The video decoder (510) may perform decoding operations using a specified video compression technique adopted in a standard such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence adheres to both the syntax of the video compression technique or standard and a profile as documented in the video compression technique or standard. Specifically, the profile allows for the selection of a particular tool from all the tools available in the video compression technique or standard, as the only tool available for use under that profile. To be standards compliant, the complexity of the encoded video sequence may be within the range defined at the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted through the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled within the encoded video sequence.

[0065]

[0092] In some embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, time, space, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, or transmission error correction codes.

[0066]

[0093] Figure 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmission circuit device). The video encoder (603) can be used in place of the video encoder (403) in the example of Figure 4.

[0067]

[0094] The video encoder (603) may receive video samples from a video source (601) (not part of the electronic device (620) in the example in Figure 6), and the video source (601) may capture a video image that will be encoded by the video encoder (603). In another example, the video source (601) may be provided as part of the electronic device (620).

[0068]

[0095] The video source (601) may provide a source video sequence to be encoded by a video encoder (603) in the form of a digital video sample stream, the digital video sample stream can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCb, RGB, XYZ, ...), and any suitable sampling structure (e.g., YCrCb4:2:0, YCrCb4:4:4). In a media serving system, the video source (601) may be a storage device capable of storing pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures or images that convey motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, each pixel may have one or more samples depending on the sampling structure, color space, etc., at the time of use. Those skilled in the art will readily understand the relationship between pixels and samples. The following explanation focuses on the sample.

[0069]

[0096] In some embodiments, the video encoder (603) may encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other arbitrary time constraints, as required by the application. Enforcing an appropriate encoding speed is a component of one function of the controller (650). In some embodiments, the controller (650) may be functionally coupled to and control other functional units, such as those described below. The couplings are not depicted for brevity. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other appropriate functions of the video encoder (603) optimized for a particular system design.

[0070]

[0097] In some embodiments, the video encoder (603) may be configured to operate within an encoding loop. In an overly simplified explanation, in one example, the encoding loop may include a source coder (630) (responsible for producing symbols, such as a symbol stream, based on, for example, an input picture to be encoded and a reference picture), as well as a (local) decoder (633) built into the video encoder (603). The decoder (633) reconstructs the symbols in a manner similar to that produced by a (remote) decoder to produce sample data, even if the built-in decoder 633 processes the encoded video stream by the source coder 630 without performing entropy encoding (since any compression between symbols and the encoded video bitstream in entropy encoding may be reversible in the video compression techniques considered in the subject of disclosure). The reconstructed sample stream (sample data) is input to a reference picture memory (634). The decoding of the symbol stream results in a bit-exact outcome independent of the decoder location (local or remote), so the contents of the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder "sees" the exact same sample values ​​as the reference picture samples that the decoder would "see" when using predictions during decoding. This fundamental principle of reference picture simultaneity (and the resulting drift if simultaneity cannot be maintained, for example, due to channel errors) is used to improve encoding quality.

[0071]

[0098] The operation of the "local" decoder (633) can be the same as that of a "remote" decoder, such as the video decoder (510) already described in detail above with reference to Figure 5. However, as briefly also referred to in Figure 5, since symbols are available and the encoding / decoding of symbols to the encoded video sequence by the entropicorder (645) and parser (520) can be reversible, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), does not have to be performed entirely by the local decoder (633) in the encoder.

[0072]

[0099] At this point, it can be said that any decoder technique other than parse / entropy decoding that exists only in the decoder must necessarily exist in substantially the same functional form within the corresponding encoder. For this reason, the subject matter of the disclosure sometimes focuses on the decoder operation in cooperation with the decoding portion of the encoder. Therefore, the description of the encoder technique is the reverse of the more comprehensive description of the decoder technique and can therefore be omitted. A more detailed description of the encoder is provided below only in specific areas or embodiments.

[0073]

[0100] In operation in some implementations, the source coder (630) may perform motion-compensated predictive coding, which predictively codes the input picture by referencing one or more previously coded pictures from a video sequence designated as a “reference picture”. Thus, the coding engine (632) codes the difference (or residue) of color channels between a pixel block of the input picture and a pixel block of the reference picture which may be selected as a predictive reference to the input picture. The terms “residue” and its adjective form “residual” may be used interchangeably.

[0074]

[0101] The local video decoder (633) may decode the encoded video data of a picture that may be designated as a reference picture based on symbols generated by the source coder (630). The operation of the encoding engine (632) may conveniently be an irreversible process. When the encoded video data is decoded in a video decoder (not shown in Figure 6), the reconstructed video sequence may typically be a copy of the source video sequence containing some errors. The local video decoder (633) may replicate the decoding process that the video decoder may perform on the reference picture and store the reconstructed reference picture in a reference picture cache (634). Thus, the video encoder (603) may locally store a copy of the reconstructed reference picture that has content in common with the reconstructed reference picture that will be obtained by a far-end (remote) video decoder (without transmission errors).

[0075]

[0102] The predictor (635) may perform a predictive search for the coding engine (632). That is, in order to encode a new picture, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or for specific metadata such as reference picture motion vectors, block shapes, etc., which may serve as appropriate predictive criteria for the new picture. The predictor (635) may operate on a sample block-by-pixel-block basis to find appropriate predictive criteria. In some cases, the input picture may have predictive criteria created from multiple reference pictures stored in the reference picture memory (634), as determined by the search results obtained by the predictor (635).

[0076]

[0103] The controller (650) may manage the encoding operation of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode video data.

[0077]

[0104] The outputs of all the aforementioned functional units may undergo entropy coding in the entropicorder (645). The entropicorder (645) converts the symbols generated by the various functional units into an encoded video sequence by lossless compression of the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0078]

[0105] The transmitter (640) may buffer the encoded video sequence produced by the entropicorder (645) in preparation for transmission over the communication channel (660), and the communication channel (660) may be a hardware / software link to a storage device that is to store the encoded video data. The transmitter (640) may merge the encoded video data from the videocoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0079]

[0106] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a specific encoded picture type to each encoded picture, and a specific encoded picture type may affect the encoding technique that can be applied to each picture. For example, a picture may often be assigned one of the following picture types:

[0080]

[0107] An intra-picture (I-picture) may be encoded and decoded without using any other pictures in the sequence as a source for prediction. Some video codecs allow different types of intra-pictures, including, for example, independent decoder refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures, as well as their respective uses and characteristics.

[0081]

[0108] The prediction picture (P-picture) may be encoded and decoded using intra-prediction or inter-prediction with at most one motion vector and reference index to predict the sample values ​​of each block.

[0082]

[0109] A bidirectional prediction picture (B-picture) may be encoded and decoded using intra-prediction or inter-prediction with at most two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple prediction pictures may use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0083]

[0110] A source picture may generally be spatially subdivided into multiple sample-encoded blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block may be encoded. Blocks may be predictively encoded, referencing other (already encoded) blocks determined by the encoding assignment applied to each picture in the block. For example, blocks of picture I may be non-predictively encoded, or they may be predictively encoded (spatial or intra-predictive), referencing already encoded blocks of the same picture. Pixel blocks of picture P may be predictively encoded via spatial or temporal prediction, referencing one previously encoded reference picture. Blocks of picture B may be predictively encoded via spatial or temporal prediction, referencing one or two previously encoded reference pictures. The source picture or an intermediate picture may be subdivided into other types of blocks for other purposes. The subdivision into encoded blocks and other types of blocks may or may not follow the same manner as described in further detail below.

[0084]

[0111] The video encoder (603) may perform encoding operations in accordance with a specified video encoding technique or standard, such as ITU-T Recommendation H.265. During such operations, the video encoder (603) may perform various compression operations, including predictive encoding operations that leverage temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0085]

[0112] In some embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source coder (630) may include such data as part of the encoded video sequence. The additional data may include time / space / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0086]

[0113] Video may be captured as multiple time-series source pictures (video pictures). Intra-picture prediction (often abbreviated as intra-prediction) utilizes the spatial correlation of a given picture, while inter-picture prediction utilizes the temporal or other correlation between pictures. For example, a particular picture being encoded / decoded, called the current picture, may be partitioned into blocks. Blocks within the current picture may be encoded by a vector called a motion vector when they are analogous to a reference block in a previously encoded and still-buffered reference picture in the video. The motion vector points to a reference block in the reference picture and, in cases where multiple reference pictures are in use, may have a third dimension that identifies the reference picture.

[0087]

[0114] In some embodiments, a bi-prediction technique can be used for interpicture prediction. According to such a bi-prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, which precede the current picture in the video in decoding order (but may be past or future in display order, respectively). Blocks in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture, and a second motion vector pointing to a second reference block in the second reference picture. The blocks can be predicted collectively by combining the first and second reference blocks.

[0088]

[0115] Furthermore, to improve coding efficiency, merge mode techniques may be used during interpicture prediction.

[0089]

[0116] According to some embodiments of this disclosure, predictions such as interpicture prediction and intrapicture prediction are performed in units of blocks. For example, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs in a picture may have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU may contain three parallel coding tree blocks (CTBs), such as one lumen CTB and two chromen CTBs. Each CTU can be recursively quadtree-partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one 64x64 pixel CU or four 32x32 pixel CUs. Each of the one or more 32x32 blocks may be further divided into four 16x16 pixel CUs. In some embodiments, each CU may be analyzed during encoding to determine the prediction type for the CU from among various prediction types, such as inter-prediction type or intra-prediction type. A CU may be divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation during encoding (encoding / decoding) is performed in units of prediction blocks. Dividing a CU into PUs (or PBs for different color channels) may be performed in various spatial patterns. For example, a luma or chroma PB may include a matrix of sample values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, or 16x8 samples.

[0090]

[0117] Figure 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values ​​in the current video picture in a sequence of video pictures, and to encode the processing block into an encoded picture which is part of an encoded video sequence. The exemplary video encoder (703) may be used instead of the video encoder (403) in the example of Figure 4.

[0091]

[0118] For example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as an 8x8 sample prediction block. The video encoder (703) then determines, for example using rate-distortion optimization (RDO), whether the processing block is best encoded using intra-mode, inter-mode, or bi-prediction mode. When it is determined that the processing block is encoded in intra-mode, the video encoder (703) may encode the processing block into an encoded picture using the intra-prediction technique, and when it is determined that the processing block is encoded in inter-mode or bi-prediction mode, the video encoder (703) may encode the processing block into an encoded picture using the inter-prediction or bi-prediction technique, respectively. In some embodiments, merge mode may be used as a submode of inter-picture prediction, in which case the motion vector is derived from one or more motion vector predictors without the benefit of an external encoded motion vector component of the predictor. In some other embodiments, there may be a motion vector component applicable to the subject block. Therefore, the video encoder (703) may include components not explicitly shown in Figure 7, such as a mode determination module, to determine the prediction mode of the processing block.

[0092]

[0119] In the example in Figure 7, the video encoder (703) includes an interencoder (730), an intraencoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725), all linked together as shown in the example arrangement in Figure 7.

[0093]

[0120] The interencoder (730) is configured to receive a sample of the current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in the previous and subsequent pictures in display order), generate interprediction information (e.g., descriptions of redundant information by intercoding techniques, motion vectors, merge mode information), and compute interprediction results (e.g., predicted blocks) based on the interprediction information using any appropriate technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information using a decoding unit 633 incorporated into the illustrative encoder 620 in Figure 6 (shown as a residual decoder 728 in Figure 7, as described in more detail below).

[0094]

[0121] The intra encoder (722) is configured to receive a sample of the current block (e.g., a processing block), compare the block to an already encoded block in the same picture, and generate a converted quantization coefficient. In some cases, it is also configured to generate intra prediction information (e.g., intra prediction direction information using one or more intra encoding techniques). The intra encoder (722) may compute an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.

[0095]

[0122] The general controller (721) may be configured to determine general control data and to control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines the prediction mode of a block and sends control signals to the switch (726) based on the prediction mode. For example, when the prediction mode is intra-mode, the general controller (721) controls the switch (726) to select the intra-mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select the intra-prediction information and include the intra-prediction information in the bitstream. When the prediction mode of a block is inter-mode, the general controller (721) controls the switch (726) to select the inter-prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select the inter-prediction information and include the inter-prediction information in the bitstream.

[0096]

[0123] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and the predicted result of a block selected from the intra encoder (722) or interencoder (730). The residual encoder (724) may be configured to encode the residual data to generate conversion coefficients. For example, the residual encoder (724) may be configured to convert the residual data from the spatial domain to the frequency domain to generate conversion coefficients. The conversion coefficients are then quantized to obtain quantized conversion coefficients. In various embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is configured to perform the inverse transform to generate decoded residual data. The decoded residual data is then available for use by the intra encoder (722) and interencoder (730). For example, an interencoder (730) can generate decoded blocks based on decoded residual data and interprediction information, and an intraencoder (722) can generate decoded blocks based on decoded residual data and intraprediction information. The decoded blocks are appropriately processed to generate decoded pictures, which are buffered in a memory circuit (not shown) and can be used as reference pictures.

[0097]

[0124] The entropy encoder (725) may be configured to format the bitstream to include encoded blocks and perform entropy coding. The entropy encoder (725) may be configured to include various types of information in the bitstream. For example, the entropy encoder (725) may be configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. Residual information may be omitted when encoding blocks in inter-mode or bi-prediction mode merge submode.

[0098]

[0125] Figure 8 shows a diagram of an illustrative video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive an encoded picture, which is part of an encoded video sequence, and to decode the encoded picture to produce a reconstructed picture. In one example, the video decoder (810) may be used instead of the video decoder (410) in the example of Figure 4.

[0099]

[0126] In the example in Figure 8, the video decoder (810) includes an entropy decoder (871), an interdecoder (880), a residual decoder (873), a reconstruction module (874), and an intradecoder (872), all coupled together as shown in the illustrative arrangement in Figure 8.

[0100]

[0127] The entropy decoder (871) can be configured to reconstruct specific symbols from the encoded picture that represent the syntax elements that make up the encoded picture. Such symbols may include, for example, the mode in which the block is encoded (e.g., intra-mode, inter-mode, bi-prediction mode, merge sub-mode, or another sub-mode), prediction information (e.g., intra-prediction information or inter-prediction information) that can identify specific samples or metadata used for prediction by the intra decoder (872) or inter-decoder (880), and residual information in the form of, for example, quantized transformation coefficients. In one example, when the prediction mode is inter or bi-prediction mode, the inter-prediction information is provided to the inter-decoder (880), and when the prediction type is intra-prediction type, the intra-prediction information is provided to the intra decoder (872). The residual information can be dequantized and provided to the residual decoder (873).

[0101]

[0128] The interdecoder (880) may be configured to receive interprediction information and generate interprediction results based on the interprediction information.

[0102]

[0129] The intra decoder (872) may be configured to receive intra prediction information and generate prediction results based on the intra prediction information.

[0103]

[0130] The residual decoder (873) may be configured to perform inverse quantization to extract dequantized conversion coefficients, process the dequantized conversion coefficients, and convert the residual from the frequency domain to the spatial domain. Furthermore, the residual decoder (873) may utilize certain control information (to include quantization parameters (QP)) that may be provided by the entropy decoder (871) (this may only be control information of a small amount of data, so no data path is drawn).

[0104]

[0131] The reconstruction module (874) may be configured to combine the residual output by the residual decoder (873) and the prediction results (optionally output by the inter or intra prediction module) in the spatial domain to form a reconstructed block that forms part of the reconstructed picture as part of the reconstructed video. Note that other appropriate operations, such as deblocking operations, may be performed to improve visual quality.

[0105]

[0132] It should be noted that the video encoders (403), (603), and (703), as well as the video decoders (410), (510), and (810), can be implemented using any suitable technique. In some embodiments, the video encoders (403), (603), and (703), as well as the video decoders (410), (510), and (810), can be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), and (603), as well as the video decoders (410), (510), and (810), can be implemented using one or more processors that execute software instructions.

[0106]

[0133] Returning to the explanation of block partitioning for encoding and decoding, general partitioning may begin with a base block and follow a predefined set of rules, a specific pattern, a partition tree, or any partition structure or scheme. Partitioning may be hierarchical and recursive. After dividing or partitioning the base block according to one or a combination of the example partitioning procedures described below or other procedures, a final set of partitions or encoding blocks may be obtained. Each of these partitions may be at one of the various partitioning levels in the partitioning hierarchy and may be of various shapes. Each partition may be called an encoding block (CB). In the various example partitioning implementations further described below, each resulting CB may be of either an acceptable size or a partitioning level. Such partitions are called encoding blocks because they may form units in which some basic encoding / decoding decisions may be made, encoding / decoding parameters may be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the encoding block partitioning structure of the tree. The coded blocks may be luma coded blocks or chroma coded blocks. The CB tree structure for each color is sometimes called a coded block tree (CBT).

[0107]

[0134] The coding blocks for all color channels are sometimes collectively called coding units (CUs). The hierarchical structure of all color channels is sometimes collectively called coding tree units (CTUs). The partitioning patterns or structures of the various color channels in a CTU may be the same or different.

[0108]

[0135] In some implementations, the partition tree scheme or structure used for luma and chroma channels may not need to be the same. In other words, luma and chroma channels may have separate coding tree structures or patterns. Furthermore, whether luma and chroma channels use the same coding partition tree structure, different coding partition tree structures, and whether they should use the same coding partition tree structure at all may depend on whether the encoded slice is a P slice, a B slice, or an I slice. For example, in the case of an I slice, chroma and luma channels may have separate coding partition tree structures or coding partition tree structure modes, while in the case of a P or B slice, luma and chroma channels may share the same coding partition tree scheme. When separate coding partition tree structures or modes are applied, luma channels may be partitioned into CBs by one coding partition tree structure, and chroma channels may be partitioned into chroma CBs by another coding partition tree structure.

[0109]

[0136] In some implementations, a predetermined partitioning pattern may be applied to the base block. As shown in Figure 9, the exemplary 4-way partition tree may begin at a first predefined level (e.g., a 64x64 block level or other size as the base block size), and the base block may be partitioned hierarchically down to a predefined minimum level (e.g., a 4x4 level). For example, the base block may follow four predefined partitioning options or patterns indicated by 902, 904, 906, and 908, and partitions designated as R may repeat the same partitioning options indicated in Figure 9 at a lower scale down to the minimum level (e.g., a 4x4 level), thus allowing recursive partitioning. In some implementations, additional restrictions may apply to the partitioning scheme in Figure 9. In the implementation of Figure 9, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) may be allowed, but they are not allowed to be recursive, while square partitions are allowed to be recursive. Partitioning according to Figure 9 with regression generates a final set of coded blocks, if necessary. A coded tree depth may be further defined to indicate the partitioning depth from the root node or root block. For example, the coded tree depth of the root node or root block may be set to 0, such as a 64x64 block, and after the root block is further partitioned once according to Figure 9, the coded tree depth increases by 1. The maximum or deepest level from the 64x64 base block to the smallest 4x4 partition should be 4 (starting from level 0) in the scheme described above. Such a partitioning scheme may be applied to one or more of the color channels. Each color channel may be partitioned independently according to the scheme in Figure 9 (for example, a partitioning pattern or option from a predefined pattern may be determined independently for each color channel at each hierarchical level).Alternatively, two or more of the color channels may share the same hierarchical pattern tree as shown in Figure 9 (for example, the same partitioning pattern or option from a predefined set of patterns may be selected for two or more color channels at each hierarchical level).

[0110]

[0137] Figure 10 shows another example of a predefined partitioning pattern that allows recursive partitioning to form a partitioning tree. The example 10-way partitioning structure or pattern may be predefined, as shown in Figure 10. The root block may start at a predefined level (e.g., from a base block at the 128x128 or 64x64 level). The example partitioning structure in Figure 10 includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. The partition type with three subpartitions indicated as 1002, 1004, 1006, and 1008 in the second row of Figure 10 may be called a “T-type” partition. The “T-type” partitions 1002, 1004, 1006, and 1008 may be called the left T-type, upper T-type, right T-type, and lower T-type. In some implementation examples, none of the rectangular partitions in Figure 10 are allowed to be further subdivided. A coding tree depth may be further defined to indicate the partitioning depth from the root node or root block. For example, the coding tree depth of the root node or root block may be set to 0, and after the root block is further partitioned once according to Figure 10, the coding tree depth increases by 1. In some implementations, only all square partitions of 1010 may be allowed to undergo recursive partitioning to the next level of the partitioning tree according to the pattern in Figure 10. In other words, recursive partitioning may not be allowed for square partitions in T-type patterns 1002, 1004, 1006, and 1008. The partitioning procedure according to Figure 10 with regression generates a final set of coding blocks if necessary. Such a scheme may be applied to one or more of the color channels. In some implementations, more flexibility may be added to the use of partitions below the 8x8 level. For example, 2x2 chroma interpretation may be used in certain cases.

[0111]

[0138] In some other implementations of coded block partitioning, a quadtree structure may be used to partition a base block or intermediate block into quadtree partitions. Such quadtree partitioning may be applied hierarchically and recursively to partitions of any square shape. Whether the base block or intermediate block or partition is further quadtree partitioned may be adapted to various local characteristics of the base block or intermediate block / partition. Quadtree partitioning at picture boundaries may be further adapted. For example, implicit quadtree partitioning may be performed at picture boundaries so that a block continues to quadtree partition until its size fits the picture boundary.

[0112]

[0139] In some other implementations, hierarchical binary partitioning from the base block may be used. In such a scheme, the base block or intermediate level block may be partitioned into two binaries. The partitioning may be horizontal or vertical. For example, horizontal partitioning may divide the base block or intermediate block into equal left and right binaries. Similarly, vertical partitioning may divide the base block or intermediate block into equal upper and lower binaries. Such partitioning may be hierarchical and recursive. In each of the base block or intermediate block, a determination may be made as to whether the binary partitioning scheme should be continued, and if so, whether horizontal or vertical partitioning should be used. In some implementations, further partitioning may stop at a predefined minimum binary size (in one or both dimensions). Alternatively, further partitioning may stop as soon as a predefined binary level or depth from the base block is reached. In some implementations, the aspect ratio of the binaries may be restricted. For example, the aspect ratio of a section does not have to be less than 1:4 (or greater than 4:1). Therefore, a vertically elongated section with a vertical-to-horizontal aspect ratio of 4:1 can simply be further divided vertically into two more sections, an upper section and a lower section, each having a vertical-to-horizontal aspect ratio of 2:1.

[0113]

[0140] In some further other examples, a three-partitioning scheme may be used to partition a base block or any intermediate block, as shown in Figure 13. Three-partitioning may be performed, vertically as shown in 1302 of Figure 13, or horizontally as shown in 1304 of Figure 13. The partitioning ratios in the example in Figure 13 are shown as 1:2:1 vertically or horizontally, but other ratios may be predefined. In some implementations, two or more different ratios may be predefined. Such a three-partitioning scheme may be used to complement a quadtree or binary partitioning structure, as such a ternary tree partitioning has the ability to capture objects at the block center within one contiguous partition, while quadtrees and binary trees always partition along the block center and therefore should partition objects into separate partitions. In some implementations, the width and height of the partitions in the example ternary tree are always powers of 2 to avoid further transformations.

[0114]

[0141] The partitioning schemes described above may be combined in any manner at different partitioning levels. As one example, the quadtree and bipartitioning schemes described above may be combined to partition a base block into a quadtree-binary-tree (QTBT) structure. In such a scheme, the base block or intermediate block / partition may be quadtree partitioned or bipartitioned, and a set of predefined conditions may be applied, if specified. A specific example is illustrated in Figure 14. In the example in Figure 14, the base block is first quadtree partitioned into four partitions, as shown by 1402, 1404, 1406, and 1408. Each of the resulting partitions is then quadtree partitioned into four further partitions (e.g., 1408), or bipartitioned into two further partitions at the next level (e.g., 1402 and 1406, which are symmetrically partitioned horizontally or vertically), or not partitioned at all (e.g., 1404). Binary or quadtree partitioning may be recursively permitted for square-shaped partitions, as shown by the partitioning patterns in the overall example of 1410 and the corresponding tree structures / representations in 1420, where solid lines represent quadtree partitioning and discontinuous lines represent binary partitioning. A flag may be used for each binary node (a leafless partition consisting of two parts) to indicate whether the binary partitioning is horizontal or vertical. For example, as shown in 1420 in accordance with the partitioning structure of 1410, a flag "0" may represent horizontal binary partitioning and a flag "1" may represent vertical binary partitioning. In the case of quadtree partitioned partitions, there is no need to indicate the type of partitioning, as quadtree partitioning always causes a block or partition to be divided both horizontally and vertically to produce four subblocks / partitions of equal size. In some implementations, a flag "1" may represent horizontal binary partitioning and a flag "0" may represent vertical binary partitioning.

[0115]

[0142] In some implementations of QTBT, the quadtree and the binary rule set may be represented by the following predefined parameters and their corresponding functions. - CTU size: The size of the root node (base block) of a quad tree. - MinQTSize: Minimum allowable 4-minute leaf node size - MaxBTSize: Maximum allowable binary tree root node size - MaxBTDepth: Maximum allowed binary tree depth - MinBTSize: Minimum allowable 2-minute leaf node size In some implementation examples of the QTBT partitioning structure, the CTU size may be set as a 128x128 chroma sample with two corresponding 64x64 blocks of chroma samples (when illustrative chroma subsampling is considered and used), MinQTSize may be set as 16x16, MaxBTSize may be set as 64x64, MinBTSize (both width and height) may be set as 4x4, and MaxBTDepth may be set as 4. Quadtree partitioning may be applied to the CTU first to generate a quadtree leaf node. A quadtree leaf node may have a size from its minimum allowable size of 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If a node is 128x128, the node will not be initially partitioned by a binary tree because its size exceeds MaxBTSize (i.e., 64x64). Otherwise, nodes that do not exceed MaxBTSize may be partitioned by binary trees. In the example in Figure 14, the base block is 128x128. According to a predefined set of rules, the base block may only be quadruparted. The base block has a partitioning depth of 0. Each of the four resulting partitions is 64x64 and does not exceed MaxBTSize, so they may be further quadruparted or binary at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning may not be considered. When a binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal partitioning may not be considered. Similarly, when a binary tree node has a height equal to MinBTSize, further vertical partitioning is not considered.

[0116]

[0143] In some implementations, the above QTBT scheme may be configured to support the flexibility for lumens and chromians to have the same QTBT structure or separate QTBT structures. For example, in the case of P and B slices, the lumens and chromians CTBs in one CTU may share the same QTBT structure. However, in the case of I slices, the lumens CTB may be partitioned into CBs by a QTBT structure, and the chromians CTB may be partitioned into chromians CBs by a different QTBT structure. This means that CUs may be used to refer to different color channels in an I slice; for example, an I slice may consist of an encoded block for the lumens component or an encoded block for two chromians, and a CU in a P or B slice may consist of an encoded block for all three color components.

[0117]

[0144] In some other implementations, the QTBT scheme may be supplemented by the three-part scheme described above. Such implementations are sometimes called multi-type tree (MTT) structures. For example, in addition to the bipartiteization of nodes, one of the three-partition patterns in Figure 13 may be selected. In some implementations, only square nodes may undergo tripartiteization. Additional flags may be used to indicate whether the tripartitioning is horizontal or vertical.

[0118]

[0145] Two-level or multi-level tree designs, such as QTBT implementations, and QTBT implementations supplemented by triplication, may be primarily motivated by complexity reduction. Logically, the complexity of traversing a tree is T D Here, T represents the number of partition types and D is the depth of the tree. A trade-off may be made by using multiple types (T) while reducing the depth (D).

[0119]

[0146] In some implementations, the CB may be further partitioned. For example, the CB may be further partitioned into multiple prediction blocks (PBs) for intra or interframe predictions during the encoding and decoding process. In other words, the CB may be further divided into different subpartitions, in which individual prediction decisions / configurations may be performed. Simultaneously, the CB may be further partitioned into multiple transformation blocks (TBs) to accurately represent the level at which transformations or inverse transformations of the video data are performed. The partitioning schemes of the CB into PBs and TBs may be the same or different. For example, each partitioning scheme may be implemented using its own procedure based on various characteristics of the video data, for example. The PB and TB partitioning schemes may be irrelevant in some implementations. The PB and TB partitioning schemes and boundaries may be correlated in some other implementations. In some implementations, for example, the TBs may be partitioned after the PB partitioning, and in particular, each PB may be determined following the partitioning of the encoding block and then further partitioned into one or more TBs. For example, in some implementations, a PB may be divided into one, two, four, or other numbers of TBs.

[0120]

[0147] In some implementations, lumar channels and chroma channels may be treated separately for the partitioning of base blocks into coded blocks, and further into prediction and / or transform blocks. For example, in some implementations, partitioning of coded blocks into prediction and / or transform blocks may be permitted for lumar channels, while such partitioning of coded blocks into prediction and / or transform blocks may not be permitted for chroma channels. In such implementations, transformation and / or prediction of lumar blocks may therefore only be performed at the coded block level. As another example, the minimum transform block size of lumar channels and chroma channels may differ; for example, coded blocks of lumar channels may be partitioned into smaller transform and / or predict blocks than those of chroma channels. As yet another example, the maximum depth of partitioning coded blocks into transform and / or predict blocks may differ between lumar channels and chroma channels; for example, coded blocks of lumar channels may be partitioned into deeper transform and / or predict blocks than those of chroma channels. For example, a luma-encoded block may be partitioned into transformation blocks of multiple sizes, which can be represented by recursive partitions that descend by up to two levels. Transformation block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, as well as transformation block sizes from 4x4 to 64x64, may be permitted. However, in the case of a chroma block, only the maximum possible transformation block specified for the luma block may be permitted.

[0121]

[0148] In some implementations of partitioning coded blocks into PB, the depth, shape, and / or other properties of the PB partitioning may depend on whether the PB is intra-coded or inter-coded.

[0122]

[0149] The partitioning of the coding blocks (or prediction blocks) into transformation blocks may be carried out recursively or non-recursively in various illustrative schemes, including but not limited to quadtree partitioning and predefined pattern partitioning, and further considering the transformation blocks at the boundaries of the coding or prediction blocks. In general, the resulting transformation blocks may be of different partitioning levels, may not be the same size, and may not have to be square in shape (for example, the resulting transformation blocks may be rectangles with several acceptable sizes and aspect ratios). Further examples are described in more detail below in relation to Figures 15, 16, and 17.

[0123]

[0150] However, in some other implementations, CBs obtained through any of the above partitioning schemes may be used as the basic or minimal coding block for prediction and / or transformation. In other words, no further partitioning is performed for inter-prediction / intra-prediction and / or transformation. For example, CBs obtained from the above QTBT scheme may be used directly as units for performing prediction. Specifically, such a QTBT structure eliminates the concept of multiple partition types, i.e., such a QTBT structure eliminates the distinction between CU, PU, ​​and TU, and supports more flexibility for the CU / CB partition shapes as described above. In such a QTBT block structure, CU / CB can have a square or rectangular shape. The leaf nodes of such a QTBT are used as units for prediction and transformation processing with no further partitioning at all. This means that in such an exemplary QTBT coding block structure, CU, PU, ​​and TU have the same block size.

[0124]

[0151] The various CB partitioning schemes described above, as well as further partitioning of CBs into PBs and / or TBs (excluding PB / TB partitioning), may be combined in any manner. The following specific implementations are provided as non-limiting examples.

[0125]

[0152] Specific implementation examples of encoding block and transform block partitioning are described below. In such implementation examples, the base block may be partitioned into encoding blocks using recursive quadtree partitioning or the predefined partitioning patterns described above (such as those in Figures 9 and 10). At each level, whether further quadtree partitioning of a particular partition should be continued may be determined by local video data characteristics. The resulting CBs may be at various quadtree partitioning levels and of various sizes. The decision of whether to encode a picture area using interpicture (time) prediction or intrapicture (spatial) prediction may be made at the CB level (or at the CU level, in the case of three color channels in total). Each CB may be further partitioned into one, two, four, or other number of PBs, depending on the predefined PB partitioning type. The same prediction process may be applied within a single PB, and relational information may be transmitted to the decoder for each PB. After obtaining the residual blocks by applying a prediction process based on the PB partitioning type, the CB can be partitioned into TBs according to another quadtree structure similar to the coding tree of the CB. In this particular implementation, the CB or TB may, but do not need to be, limited to a square shape. Furthermore, in this particular example, the PB may be square or rectangular for interpretation, or square only for intrapretation. The coding block may be partitioned into, for example, four square-shaped TBs. Each TB may be recursively further partitioned (using quadtree partitioning) into smaller TBs called residual quadtrees (RQTs).

[0126]

[0153] Further examples of implementations for partitioning base blocks into CBs, PBs, and / or TBs are described below. For example, instead of using multiple partition unit types such as those shown in Figure 9 or Figure 10, a quadtree with nested multitype trees using 2- and 3-partition segmentation structures (e.g., QTBT or QTBT with 3 partitions as described above) may be used. The separation of CBs, PBs, and TBs (i.e., partitioning CBs into PBs and / or TBs, and partitioning PBs into TBs) may be abandoned except when necessary for CBs that are too large for the maximum transformation length (such CBs may require further partitioning). The partitioning scheme in this example may be designed to support more flexibility in CB partition shapes so that both prediction and transformation are feasible for CB levels without further partitioning. In such coding tree structures, CBs may have a square or rectangular shape. Specifically, a coding tree block (CTB) may first be partitioned by a quadtree structure. The quadtree leaf nodes may then be further partitioned by a nested multitype tree structure. An example of a nested multitype tree structure using 2 or 3 partitions is shown in Figure 11. Specifically, the example multitype tree structure in Figure 11 includes four partition types, called vertical 2-partition (SPLIT_BT_VER)(1102), horizontal 2-partition (SPLIT_BT_HOR)(1104), vertical 3-partition (SPLIT_TT_VER)(1106), and horizontal 3-partition (SPLIT_TT_HOR)(1108). Thus, CB corresponds to a leaf of the multitype tree. In this implementation example, this segmentation is used for both predictive and transformive processing where no further partitioning is needed, as long as CB is not too large relative to the maximum transform length. This means that in most cases, in a quadtree with a nested multitype tree coding block structure, CB, PB, and TB have the same block size. The exception occurs when the maximum supported transform length is smaller than the width or height of the color components of CB. In some implementations, in addition to two or three partitions, the nested pattern in Figure 11 may further include a quadtree partition.

[0127]

[0154] Figure 12 shows one specific example of a quadtree having a nested multitype tree coding block structure of block partitions (including quadtree, 2, and 3 partition options) for a single base block. More specifically, Figure 12 shows that the base block 1200 has been quadtree-partitioned into four square partitions 1202, 1204, 1206, and 1208. The multitype tree structure in Figure 11, and the decision to use further quadtrees for further partitioning, are made for each of the quadtree-partitioned partitions. In the example in Figure 12, partition 1204 is not further partitioned. Parts 1202 and 1208 each adopt a different quadtree partition. In the case of partition 1202, the top-left, top-right, bottom-left, and bottom-right partitions, which are quadtree-partitioned at the second level, adopt a third level partition of the quadtree: horizontal 2 partition 1104 in Figure 11, no partition, and horizontal 3 partition 1108 in Figure 11, respectively. Section 1208 employs a different quadtree partitioning pattern, with the top-left, top-right, bottom-left, and bottom-right sections of the second-level quadtree partitioning employing the third-level partitioning, no partitioning, no partitioning, and horizontal 2 partitioning 1104 of Figure 11, respectively, of the vertical 3 partitioning 1106 in Figure 11. Two of the sub-sections of the top-left section of the third-level partitioning 1208 are further partitioned according to the horizontal 2 partitioning 1104 and horizontal 3 partitioning 1108 of Figure 11, respectively. Section 1206 employs a second-level partitioning pattern that divides into two sections according to the vertical 2 partitioning 1102 of Figure 11, and each of the divided sections is further partitioned at the third level according to the horizontal 3 partitioning 1108 and vertical 2 partitioning 1102 of Figure 11. A fourth-level partitioning pattern is further applied to one of these according to the horizontal 2 partitioning 1104 of Figure 11.

[0128]

[0155] As a specific example, the maximum luma conversion size may be 64x64, and the maximum supported chroma conversion size may be, for example, 32x32, which may differ from that of luma. Even if the CB in the above example in Figure 12 is not further divided into smaller PB and / or TB as a whole, if the width or height of the luma-encoded block or chroma-encoded block is greater than the maximum conversion width or height, the luma-encoded block or chroma-encoded block may be automatically divided horizontally and / or vertically to satisfy the conversion size limit in that direction.

[0129]

[0156] In the specific examples of base block partitioning to CB as described above, the coding tree scheme may support the ability of lumens and chromens to have separate block tree structures. For example, in the case of P and B slices, lumens and chromens CTBs in a single CTU may share the same coding tree structure. In the case of I slices, for example, lumens and chromens may have separate coding block tree structures. Where separate block tree structures are applied, a lumens CTB may be partitioned into lumens CBs by one coding tree structure, and a chromens CTB may be partitioned into chromens CBs by another coding tree structure. This means that a CU in an I slice may consist of coding blocks for the lumens component or coding blocks for the two chromens component, and that a CU in a P or B slice will always consist of coding blocks for all three color components, unless the video is monochrome.

[0130]

[0157] When an encoded block is further partitioned into multiple transform blocks, these transform blocks may be arranged in the bitstream according to various orders or scanning patterns. Implementation examples of partitioning encoded or predictive blocks into transform blocks, and the encoding order of the transform blocks, are described in more detail below. In some implementations, as described above, the transform partitioning may support multiple shapes of transform blocks, such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, with transform block sizes ranging from, for example, 4x4 to 64x64. In some implementations, if the encoded block is smaller than or equal to 64x64, the transform block partitioning may be applied only to the luma component, in which case the transform block size for the chroma block is the same as the encoded block size. Otherwise, if the encoding block width or height is greater than 64, both the luma and chroma encoding blocks may be implicitly divided into transformation blocks that are multiples of min(W,64) × min(H,64) and min(W,32) × min(H,32), respectively.

[0131]

[0158] In some implementations of transform block partitioning, for both intra and inter encoded blocks, the encoded block may be further partitioned into multiple transform blocks having a predefined number of partitioning depths (e.g., 2 levels). The transform block partitioning depth and size may be relevant. For some implementations, the mapping from the transform size at the current depth to the transform size at the next depth is shown below in Table 1.

[0132] [Table 1]

[0133]

[0159] Based on the example mapping in Table 1, for a 1:1 square block, the next level of transformation partitioning may produce four 1:1 square subtransformation blocks. The transformation partitioning may stop at, for example, 4x4. Thus, a transformation size of 4x4 at the current depth corresponds to the same size 4x4 at the next depth. In the example in Table 1, for a 1:2 / 2:1 non-square block, the next level of transformation partitioning may produce two 1:1 square subtransformation blocks, while for a 1:4 / 4:1 non-square block, the next level of transformation partitioning may produce two 1:2 / 2:1 subtransformation blocks.

[0134]

[0160] In some implementations, additional restrictions on transform block partitioning may be applied to the Luma components of an intra-encoded block. For example, at each level of transform partitioning, all sub-transform blocks may be restricted to being of equal size. For instance, for a 32x16 encoded block, a level 1 transform partition produces two 16x16 sub-transform blocks, and a level 2 transform partition produces eight 8x8 sub-transform blocks. In other words, the second level partitioning must be applied to all first-level sub-blocks to ensure that the transform units are of equal size. An example of transform block partitioning for an intra-encoded square block according to Table 1 is shown in Figure 15, along with the coding order illustrated by the arrows. Specifically, 1502 shows a square encoded block. The first-level partitioning into four equally sized transform blocks according to Table 1 is shown in 1504, along with the coding order indicated by the arrows. The second-level partitioning of all first-level equally sized blocks into 16 equally sized transform blocks according to Table 1 is shown in 1506, along with the coding order indicated by the arrows.

[0135]

[0161] In some implementations, the above restrictions on intra-coding may not apply to the Luma component of an inter-encoded block. For example, after the first level of transform partitioning, one of the sub-transform blocks may be further independently partitioned at another level. Thus, the resulting transform blocks may or may not be the same size. An example partitioning of an inter-encoded block into transform blocks in that encoding order is shown in Figure 16. In the example in Figure 16, the inter-encoded block 1602 is partitioned into transform blocks at two levels according to Table 1. At the first level, the inter-encoded block is partitioned into four transform blocks of equal size. Then, just one of the four transform blocks (but not all of them) is further partitioned into four sub-transform blocks, resulting in a total of seven transform blocks of two different sizes, as shown by 1604. The example encoding order of these seven transform blocks is indicated by the arrow in 1604 in Figure 16.

[0136]

[0162] In some implementations, several additional restrictions may be applied to the chroma component compared to the transform block. For example, the transform block size for the chroma component can be about the same size as the encoded block size, but it cannot be smaller than a predefined size, such as 8x8.

[0137]

[0163] In some other implementations, for coding blocks with a width (W) or height (H) greater than 64, both the luma and chroma coding blocks may be implicitly divided into transformation units that are multiples of min(W,64) × min(H,64) and min(W,32) × min(H,32), respectively. Herein, "min(a,b)" may return the smaller value of a and b.

[0138]

[0164] Figure 17 further illustrates another alternative illustrative scheme for partitioning a coded block or prediction block into transform blocks. As shown in Figure 17, instead of using recursive transform partitioning, a predefined set of partitioning types may be applied to the coded block depending on the transform type of the coded block. In the particular example shown in Figure 17, one of the six illustrative partitioning types may be applied to divide the coded block into a varying number of transform blocks. Such schemes for generating transform block partitioning may be applied to either coded blocks or prediction blocks.

[0139]

[0165] More specifically, the partitioning scheme in Figure 17 provides up to six exemplary partition types for any given transformation type (where "transformation type" refers to a major transformation type, such as ADST). In this scheme, a transformation partition type may be assigned to each coded block or predictive block, for example, based on rate distortion cost. In the example, the transformation partition type assigned to a coded block or predictive block may be determined based on the transformation type of the coded block or predictive block. A particular transformation partition type may correspond to a transformation block partition size and pattern, as shown by the six transformation partition types illustrated in Figure 17. The correspondence between various transformation types and various transformation partition types may be predefined. An example is shown below, where capitalized labels indicate transformation partition types that may be assigned to coded blocks or predictive blocks based on rate distortion cost.

[0140]

[0166] • PARTITION_NONE: Allocates a conversion size equal to the block size.

[0141]

[0167] • PARTITION_SPLIT: Assigns a conversion size that is half the width of the block size and half the height of the block size.

[0142]

[0168] • PARTITION_HORZ: Allocates a conversion size that has the same width as the block size and half the height of the block size.

[0143]

[0169] • PARTITION_VERT: Allocates a conversion size that is half the width of the block size and the same height as the block size.

[0144]

[0170] • PARTITION_HORZ4: Allocates a conversion size that has the same width as the block size and 1 / 4 of the block size's height.

[0145]

[0171] • PARTITION_VERT4: Allocates a conversion size that is 1 / 4 the width of the block size and the same height as the block size.

[0146]

[0172] In the example above, all conversion partition types, as shown in Figure 17, include a uniform conversion size for partitioned conversion blocks. This is merely an example, not an limitation. Some other implementations may use mixed conversion block sizes for partitioned conversion blocks of a particular partition type (or pattern).

[0147]

[0173] Next, the PB (or CB, also called PB when not further partitioned into prediction blocks) obtained from any of the above partitioning schemes can become individual blocks for coding via intra or interprediction. In the case of interprediction for the current PB, the residue between the current block and the prediction block may be generated, coded, and included in the coded bitstream.

[0148]

[0174] Interpretation may be performed, for example, in single-reference mode or composite-reference mode. In some implementations, a skip flag may be included first in the bitstream for the current block (or at a higher level) to indicate whether the current block is intercoded and should not be skipped. If the current block is intercoded, another flag may be included further in the bitstream as a signal to indicate whether single-reference mode or composite-reference mode is used for predicting the current block. In single-reference mode, one reference block may be used to generate the predict block for the current block. In composite-reference mode, two or more reference blocks may be used to generate the predict block, for example, by a weighted average. Composite-reference mode is sometimes called two or more-reference mode, two-reference mode, or multiple-reference mode. One or more reference blocks may be identified using one or more reference frame indices and additionally using one or more corresponding motion vectors, the corresponding one or more motion vectors indicating a shift between the reference block and the current block in location relative to the frame, for example, horizontal and vertical pixels. For example, the interpretation block for the current block may be generated as a prediction block in single-reference mode from a single reference block identified by a single motion vector in a reference frame, while in the case of composite-reference mode, the prediction block may be generated by a weighted average of two reference blocks in two reference frames indicated by two reference frame indices and two corresponding motion vectors. The motion vectors may be encoded and included in the bitstream in various ways.

[0149]

[0175] In some implementations, the encoding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures may be maintained in the DPB, waiting to be displayed (in the decoding system), and some images / pictures in the DPB may be used as reference frames to enable inter prediction (in the decoding or encoding system). In some implementations, reference frames in the DPB may be tagged as short-term or long-term references for the current image being encoded or decoded. For example, short-term reference frames may include frames used for inter prediction of blocks in the current frame, or in a predefined number (e.g., two) of subsequent video frames closest to the current frame in decoding order. Long-term reference frames may include frames in the DPB that can be used to predict image blocks in frames that are more than a predefined number of frames away from the current frame in decoding order. Information about such tags for short-term and long-term reference frames may be called a Reference Picture Set (RPS) and may be added to the header of each frame in the encoded bitstream. Each frame in an encoded video stream may be identified by a Picture Order Counter (POC), which is either absolutely numbered according to the playback sequence or relates to a group of pictures starting, for example, with frame I.

[0150]

[0176] In some implementations, one or more reference picture lists, including the identification of short-term and long-term reference frames for inter-prediction, may be formed based on information in the RPS. For example, a single picture reference list may be formed for one-way inter-prediction and represented as L0 reference (or reference list 0), while two picture reference lists may be formed for bidirectional inter-prediction and represented as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be arranged in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. One-way inter-prediction may be in single-reference mode, or in composite-reference mode, where multiple criteria for generating prediction blocks by weighted averaging in composite prediction mode are present when the block to be predicted is on the same side of the frame. Bidirectional inter-prediction may be in composite mode only, since bidirectional inter-prediction involves at least two reference blocks.

[0151]

[0177] In some implementations, a merge mode (MM) may be used for interpretation. Generally, in merge mode, one or more motion vectors in a single reference prediction, or in a composite reference prediction for the current PB, may be derived from other motion vectors rather than being computed and signaled independently. For example, in an encoding system, the current motion vector for the current PB may be represented by the difference between the current motion vector and one or more other already encoded motion vectors (called reference motion vectors). Such a difference of motion vectors, rather than the entire current motion vector, may be encoded and included in the bitstream and linked to the reference motion vectors. Correspondingly, in a decoding system, the motion vector corresponding to the current PB may be derived based on the decoded motion vector difference and the decoded reference motion vector linked to it. As a unique aspect of general merge mode (MM) interpretation, such interpretation based on motion vector differences is sometimes called merge mode with motion vector differences (MMVD). Thus, MM in general, or MMVD in particular, may be implemented to improve encoding efficiency by leveraging correlations between motion vectors associated with different PBs. For example, adjacent PBs may have similar motion vectors, and therefore the MVD can be small and efficiently encodeable. As another example, motion vectors may be temporally correlated (between frames) with respect to blocks that are similarly in / located in space.

[0152]

[0178] In some implementations, the MM flag may be included in the bitstream during the encoding process to indicate whether the current PB is in merge mode. Alternatively, the MMVD flag may be included and signaled in the bitstream during the encoding process to indicate whether the current PB is in MMVD mode. The MM and / or MMVD flags or indicators may be provided at PB level, CB level, CU level, CTB level, CTU level, slice level, frame level, picture level, sequence level, etc. In a specific example, both the MM and MMVD flags may be included for the current CU, and the MMVD flag may be signaled immediately after the skip flag and MM flag to specify whether MMVD mode is used for the current CU.

[0153]

[0179] In some implementations of MMVD, a list of reference motion vectors (RMVs), or MV predictor candidates for motion vector prediction, may be formed for the block being predicted. The list of RMV candidates may contain a predetermined number (e.g., two) of MV predictor candidate blocks, the motion vectors of which may be used to predict the current motion vector. RMV candidate blocks may include blocks selected from adjacent blocks within the same frame and / or from time blocks (e.g., blocks placed equivalently in frames preceding or following the current frame). These options represent blocks in spatial or temporal locations relative to the current block that are likely to have similar or identical motion vectors to the current block. The size of the list of MV predictor candidates may be predetermined. For example, the list may contain two or more candidates. To be on the list of RMV candidates, candidate blocks may be required to have, for example, the same reference frame (or multiple reference frames) as the current block, and must exist (e.g., boundary checking must be performed when the current block is near the edge of a frame), must have been encoded during the encoding process, and / or must have been decoded during the decoding process. In some implementations, the list of merge candidates may first contain spatially adjacent blocks (scanned in a specific predefined order), if available and meeting the above conditions, followed by time blocks if space is still available in the list. Adjacent RMV candidate blocks may be selected, for example, from the leftmost and topmost blocks of the current block. The list of RMV predictor candidates may be dynamically formed as a Dynamic Reference List (DRL) at various levels (sequence, picture, frame, slice, superblock, etc.). The DRL may be signaled with a bitstream.

[0154]

[0180] In some implementations, the actual MV predictor candidates used as reference motion vectors for predicting the motion vector of the current block may be signaled. In cases where the RMV candidate list contains two candidates, a one-bit flag called a merge candidate flag may be used to indicate the selection of the reference merge candidate. If the current block is predicted in composite mode, each of the multiple motion vectors predicted using the MV predictors may be associated with a reference motion vector from the merge candidate list. The encoder may determine which of the RMV candidates more precisely predicts the MV of the current encoded block and signal the selection as an index to the DRL.

[0155]

[0181] In some implementations of MMVD, after an RMV candidate is selected and used as a base motion vector predictor to predict the motion vector, a motion vector difference (MVD or deltaMV representing the difference between the predicted motion vector and the base candidate motion vector) may be calculated in the encoding system. Such an MVD may include information representing the magnitude and direction of the MV difference, and both the magnitude and direction of the MV difference may be signaled in the bitstream. The magnitude and direction of the motion difference may be signaled in various ways.

[0156]

[0182] In some implementations of MMVD, a distance index may be used to specify the magnitude information of the motion vector difference and to indicate one of a set of predefined offsets that represent a predefined motion vector difference from the starting point (reference motion vector). The MV offset by the signaled index may then be added to the horizontal or vertical component of the start (reference) motion vector. Whether the horizontal or vertical component of the reference motion vector should be offset may be determined by the directional information of the MVD. Exemplary predefined relationships between distance indices and predefined offsets are specified in Table 2.

[0157] [Table 2]

[0158]

[0183] In some implementations of MMVD, a direction index may be further signaled and used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction may be limited to one of the horizontal or vertical directions. Exemplary 2-bit direction indices are shown in Table 3. In the examples in Table 3, the interpretation of the MVD may differ depending on the information of the start / reference MV. For example, when the start / reference MV corresponds to a uni-prediction block or to a bi-prediction block where both reference frame lists point to the same side of the current picture (i.e., when the POCs of both reference pictures are both greater than the POC of the current picture, or both are less than the POC of the current picture), the sign in Table 3 may specify the sign (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a biprediction block having two reference pictures on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is less than the POC of the current picture), and when the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the sign in Table 3 may specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 may have the opposite value (the opposite sign of the offset). Otherwise, when the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the sign in Table 3 may specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign of the offset to the reference MV associated with picture reference list 0 has the opposite value.

[0159] [Table 3]

[0160]

[0184] In some implementations, MVD may be scaled according to the difference in POC for each direction. If the difference in POC in both lists is the same, scaling is not necessary. Otherwise, if the difference in POC in reference list 0 is greater than the difference in POC in reference list 1, the MVD for reference list 1 is scaled. If the difference in POC in reference list 1 is greater than that of list 0, the MVD for list 0 may be scaled in the same manner. If the starting MV is single-predicted, the MVD is added to the available MV or reference MV.

[0161]

[0185] In some implementations of MVD coding and signaling for bidirectional composite prediction, symmetric MVD coding may be performed such that, in addition to coding and signaling two MVDs separately, or as an alternative, only one MVD may require signaling, and the other MVD may be derived from the signaled MVD. In such an implementation, motion information containing the reference picture indices of both List 0 and List 1 is signaled. However, only the MVD associated with reference list 0 is signaled, and the MVD associated with reference list 1 is derived without signaling. Specifically, at the slice level, a flag called "mvd_l1_zero_flag" may be included in the bitstream to indicate whether reference list 1 is not signaled in the bitstream. If this flag is 1, indicating that reference list 1 is equal to zero (and therefore not signaled), a bidirectional prediction flag called "BiDirPredFlag" may be set to 0, meaning there is no bidirectional prediction. Otherwise, if mvd_l1_zero_flag is zero, BiDirPredFlag may be set to 1 if the nearest reference picture in List 0 and the nearest reference picture in List 1 form a forward-backward pair of reference pictures, or a backward-forward pair of reference pictures, and the reference pictures in both List 0 and List 1 are short-circuit reference pictures. Otherwise, BiDirPredFlag is set to 0. A BiDirPredFlag of 1 may indicate that a symmetric mode flag is signaled in the bitstream as an additional. The decoder may extract the symmetric mode flag from the bitstream when BiDirPredFlag is 1. The symmetric mode flag may be signaled, for example, at the CU level (if necessary) to indicate whether a symmetric MVD coding mode is being used for the corresponding CU.When the symmetric mode flag is 1, it indicates the use of the symmetric MVD coding mode, and that only the reference picture indices in both List 0 and List 1 (called "mvp_l0_flag" and "mvp_l1_flag") are signaled along with the MVD associated with List 0 (called "MVD0"), and that the other motion vector difference "MVD1" is derived rather than signaled. For example, MVD1 may be derived as -MVD0. Thus, in the exemplary symmetric MVD mode, only one MVD is signaled.

[0162]

[0186] In some other implementations of MV prediction, harmonized schemes may be used to perform general merge-mode MMVD and several other types of MV prediction for both single-reference-mode and composite-reference-mode MV prediction. Various syntactic elements may be used to signal the manner in which the MV of the current block is predicted.

[0163]

[0187] For example, in the case of a single reference mode, the following MV prediction modes may be signaled.

[0164]

[0188] NEARMV - Uses one of the motion vector predictors (MVPs) in a list that is directly indicated by a DRL (Dynamic Reference List) index that has no MVD at all.

[0165]

[0189] NEWMV - Uses one of the motion vector predictors (MVPs) in a list signaled by the DRL index as a criterion, and applies delta to the MVP (e.g., using MVD).

[0166]

[0190] GLOBALMV - Uses motion vectors based on frame-level global motion parameters.

[0167]

[0191] Similarly, if the composite reference interpretation mode uses two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes may be signaled:

[0168]

[0192] NEAR_NEARMV - Uses one of the motion vector predictors (MVPs) in a list, signaled by a DRL index without MVD for each of the two MVs that will be predicted.

[0169]

[0193] NEAR_NEWMV - To predict the first of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the base MV without MVD, and to predict the second of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the base MV, along with the Delta MV (MVD) signaled as an additional.

[0170]

[0194] NEW_NEARMV - To predict the second of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the base MV without MVD, and to predict the first of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the base MV, along with the Delta MV (MVD) signaled as an additional.

[0171]

[0195] NEW_NEWMV - Uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the base MV, and uses this together with the additionally signaled delta MV to make predictions for each of the two MVs.

[0172]

[0196] GLOBAL_GLOBALMV - Uses MV from each criterion based on the global motion parameters at that frame level.

[0173]

[0197] Therefore, the term "NEAR" above refers to MV prediction using a reference MV with no MVD at all as a general merge mode, while the term "NEW" refers to MV prediction that involves using a reference MV and offsetting the reference MV with a signaled or derived MVD, similar to the MMVD mode. In the case of composite interpretation, for example, two MVDs may be correlated, and such correlation may be leveraged to reduce the amount of information required to signal two motion vector deltas, although both the reference-based motion vector and motion vector delta above may generally be different or unrelated between the two references or two MVDs. To leverage such correlations, joint signaling of the two MVDs may be performed and indicated in a bitstream, as will be described in more detail below.

[0174]

[0198] In some implementations of MVD, a predefined pixel resolution for MVD may be permitted. For example, 1 / 8 pixel motion vector precision (or accuracy) may be permitted. The above-described MVDs for various MV prediction modes may be constructed and signaled in various ways. In some implementations, various syntax elements may be used to signal the above-described motion vector difference in reference frame list 0 or list 1.

[0175]

[0199] For example, a syntax element called "mv_joint" may specify which component of the associated motion vector difference is non-zero. In the case of MVD, this signals all non-zero components together. For example, mv_joint can have the following values: A value of 0 may indicate that there are no non-zero MVDs along the horizontal or vertical axis. A value of 1 may indicate that there is a non-zero MVD only along the horizontal direction. 2 may indicate that there is a non-zero MVD only along the vertical direction. 3 may indicate that there is a non-zero MVD along both the horizontal and vertical directions.

[0176]

[0200] When the "mv_joint" syntax element for MVD signals that there are no non-zero MVD components, further MVD information does not need to be signaled. However, if the "mv_joint" syntax signals that there are one or two non-zero components, additional syntax elements may be signaled for each of the non-zero MVD components, as described below.

[0177]

[0201] For example, a syntax element called "mv_sign" may be used to additionally specify whether the corresponding motion vector difference component is positive or negative.

[0178]

[0202] As another example, a syntax element called "mv_class" may be used to specify the class of a motion vector difference from a predefined set of classes for corresponding non-zero MVD components. These predefined classes for motion vector differences may be used, for example, to divide a continuous magnitude space of motion vector differences into non-overlapping ranges, each range corresponding to an MVD class. Thus, the signaled MVD class indicates the magnitude range of the corresponding MVD component. In the implementation example shown in Table 4 below, higher classes correspond to motion vector differences with larger magnitude ranges. In Table 4, the symbol (n,m) is used to represent a range of motion vector differences greater than n pixels and less than or equal to m pixels.

[0179] [Table 4]

[0180]

[0203] In some other examples, an additional syntax element called "mv_bit" may be used to specify the integer part of the offset between the non-zero motion vector difference component and the start magnitude of the MV class magnitude range that is signaled accordingly. Thus, mv_bit may indicate the magnitude or size of the MVD. The number of bits required in "my_bit" to signal the full range for each MVD class may vary depending on the MV class. For example, MV_CLASS0 and MV_CLASS1 in the implementations of Table 4 may require only a single bit to indicate an integer pixel offset of 1 or 2 from 0 of the start MVD, while each higher MV_CLASS in the implementation examples of Table 4 may progressively require one more bit for "mv_bit" than the previous MV_CLASS.

[0181]

[0204] In some other implementations, a syntax element called "mv_fr" may be used to specify the first two fractional bits of the motion vector difference for the corresponding non-zero MVD component, while a syntax element called "mv_hp" may be used to specify the third fractional bit (high-resolution bit) of the motion vector difference for the corresponding non-zero MVD component. Two bits of "mv_fr" essentially provide a 1 / 4 pixel MVD resolution, while the "mv_hp" bits may further provide an 1 / 8 pixel resolution. In some other implementations, two or more "mv_hp" bits may be used to provide MVD pixel resolutions finer than 1 / 8 pixels. In some implementations, additional flags may be signaled at one or more of various levels to indicate whether 1 / 8 pixel or higher MVD resolutions are supported. If an MVD resolution does not apply to a particular coding unit, the above syntax elements for the corresponding unsupported MVD resolutions may not be signaled.

[0182]

[0205] In some of the implementation examples above, fractional resolution may be independent of different MVD classes. In other words, a similar option for motion vector resolution may be provided using a predefined number of "mv_fr" and "mv_hp" bits to signal fractional MVDs of non-zero MVD components, regardless of the magnitude of the motion vector difference.

[0183]

[0206] However, in some other implementations, the resolution of the motion vector difference across different MVD magnitude classes may be distinct or adaptable. Specifically, higher resolution MVD for larger MVD magnitudes in higher MVD classes may not result in a statistically significant improvement in compression efficiency or coding gain. Therefore, MVD may be coded by reducing the resolution (integer pixel resolution or fractional pixel resolution) for larger MVD magnitude ranges corresponding to higher MVD magnitude classes. Similarly, MVD may generally be coded by reducing the resolution (integer pixel resolution or fractional pixel resolution) for larger MVD values. Such MVD class-dependent or MVD magnitude-dependent MVD resolutions are sometimes referred to as adaptable MVD resolution, magnitude-dependent adaptable MVD resolution, or magnitude-dependent MVD resolution. The term “resolution” may also be referred to as “pixel resolution.” Adaptable MVD resolution may be implemented in various ways, as illustrated by the implementation examples below, to achieve better overall compression efficiency. In particular, statistical observations suggest that treating the MVD resolution for large magnitude or high-class MVD at a level similar to that of small magnitude or low-class MVD in non-conforming modes may not significantly improve the inter-predictive residual coding efficiency of blocks with large magnitude or high-class MVD. This suggests that the reduction in the number of signaling bits by aiming for a less precise MVD may outweigh the additional bits required to encode the inter-predictive residual resulting from such a less precise MVD. In other words, using a higher MVD resolution for large magnitude or high-class MVD may not yield more coding gain than using a lower MVD resolution.

[0184]

[0207] In some common implementations, the pixel resolution or precision of an MVD may decrease or cease to increase as the MVD class increases. Decreasing the pixel resolution of an MVD corresponds to a coarser MVD (or a larger step from one MVD level to the next). In some implementations, the correspondence between MVD pixel resolution and MVD class may be specified, predefined, or preconfigured, and therefore may not need to be signaled in the encoded bitstream.

[0185]

[0208] In some implementation examples, the MV classes in Table 3 may each be associated with a different MVD pixel resolution.

[0186]

[0209] In some implementations, each MVD class may be associated with a single acceptable resolution. In some other implementations, one or more MVD classes may each be associated with two or more optional MVD pixel resolutions. The signal in the bitstream for a current MVD component having such MVD classes may be followed by additional signaling to indicate the optional pixel resolution selected for the current MVD component. In some implementations, adaptively acceptable MVD pixel resolutions may include, but are not limited to, 1 / 64-pel (pixel), 1 / 32-pel, 1 / 16-pel, 1 / 8-pel, 1-4-pel, 1 / 2-pel, 1-pel, 2-pel, 4-pel, etc. (in descending order of resolution). Thus, each of the ascending MVD classes may be associated with one of these MVD pixel resolutions in non-ascending order. In some implementations, an MVD class may be associated with two or more of the above resolutions, where the higher resolution may be lower than or equal to the lower resolution of the preceding MVD class. For example, if MV_CLASS_3 in Table 4 is associated with optional 1-pel and 2-pel resolutions, then the highest resolution that MV_CLASS_4 in Table 4 may be associated with should be 2-pel. In some other implementations, the highest acceptable resolution of an MV class may be higher than the lowest acceptable resolution of a preceding (lower) MV class. However, the average of the acceptable resolutions of ascending MV classes may simply be non-ascending.

[0187]

[0210] In some implementations, when fractional pixel resolutions higher than 1 / 8pel are permitted, the "mv_fr" and "mv_hp" signaling may be expanded accordingly to a total of more than 3 fractional bits.

[0188]

[0211] In some implementations, fractional pixel resolution may only be permitted for MVD classes that are below or equal to a threshold MVD class. For example, fractional pixel resolution may only be permitted for MVD_class_0 and not for all other MV classes in Table 4. Similarly, fractional pixel resolution may only be permitted for MVD classes that are below or equal to any one of the other MV classes in Table 4. For other MVD classes that are above the threshold MVD class, only integer pixel resolutions for the MVD are permitted. In this manner, fractional resolution signaling, such as one or more of the "mv-fr" and / or "mv-hp" bits, may not need to be signaled for MVDs that are signaled with MVD classes that are higher than or equal to the threshold MVD class. For MVD classes with resolutions lower than 1 pixel, the number of bits in the "mv-bit" signaling may be further reduced. For example, in the case of MV_CLASS_5 in Table 4, the range of the MVD pixel offset is (32, 64), and therefore 5 bits are needed to signal the entire range with 1-pel resolution. However, if MV_CLASS_5 is associated with a 2-pel MVD resolution (a resolution lower than 1-pixel resolution), then 4 bits may be needed for "mv-bit" instead of 5 bits, and neither "mv-fr" nor "mv-hp" needs to be signaled as MV-CLASS_5 following the signaling of "mv_class".

[0189]

[0212] In some implementations, fractional pixel resolution may only be permitted for MVDs with integer values ​​below a threshold integer pixel value. For example, fractional pixel resolution may only be permitted for MVDs smaller than 5 pixels. Corresponding to this example, fractional resolution may be permitted for MV_CLASS_0 and MV_CLASS_1 in Table 4, but not for all other MV classes. As another example, fractional pixel resolution may only be permitted for MVDs smaller than 7 pixels. Corresponding to this example, fractional resolution may be permitted for MV_CLASS_0 and MV_CLASS_1 in Table 4 (with ranges less than 5 pixels), but not for MV_CLASS_3 and above (with ranges greater than 5 pixels). For MVDs belonging to MV_CLASS_2 whose pixel range encompasses 5 pixels, fractional pixel resolution of the MVD may or may not be permitted depending on the "mv-bit" value. Fractional pixel resolution may be allowed if the "m-bit" value is signaled as 1 or 2 (and therefore the integer part of the signaled MVD is 5 or 6 calculated as the beginning of the pixel range of MV_CLASS_2 with an offset of 1 or 2 as indicated by "m-bit"). Otherwise, fractional pixel resolution may not be allowed if the "mv-bit" value is signaled as 3 or 4 (such that the integer part of the signaled MVD is 7 or 8).

[0190]

[0213] In some other implementations, for MV classes above a threshold MV class, only a single MVD value may be allowed. For example, such a threshold MV class may be MV_CLASS2. Therefore, MV_CLASS_2 and above are only allowed to have a single MVD value and do not require fractional pixel resolution. The single allowed MVD values ​​for these MV classes may be predefined. In some examples, the allowed single value may be the higher end value of the respective range for these MV classes in Table 4. For example, MV_CLASS_2 through MV_CLASS_10 may be above or equal to the threshold class of MV_CLASS2, and the single allowed MVD values ​​for these classes may be predefined as 8, 16, 32, 64, 128, 256, 512, 1024, and 2048, respectively, as shown in Table 5. In some other examples, the allowed single value may be the midpoint value of the respective range for these MV classes in Table 4. For example, MV_CLASS_2 through MV_CLASS_10 may exceed the class threshold, and a single acceptable MVD value for these classes may be predefined as 3, 6, 12, 24, 48, 96, 192, 384, 768, and 1536, respectively. Any other value within the range may also be defined as a single acceptable resolution for each MVD class.

[0191] [Table 5]

[0192]

[0214] In the above implementation, when the signaled "mv_class" is equal to or exceeds a predefined MVD class threshold, the "mv_class" signaling alone is sufficient to determine the MVD value. The magnitude and direction of the MVD should then be determined using "mv_class" and "mv_sign".

[0193]

[0215] Therefore, when the MVD is signaled for only one reference frame (from reference frame list 0 or list 1, but not from both), or when it is signaled together for two reference frames, the precision (or resolution) of the MVD may depend on the associated motion vector difference class and / or the magnitude of the MVD in Table 3. Various other adaptable MVD resolution schemes depending on the MVD magnitude or class are assumed.

[0194]

[0216] Returning to the various composite interprediction modes in which each MV may be predicted by a reference motion vector and encoded by an MVD, the two MVDs may be signaled separately or together in the bitstream, as described above. Therefore, in some implementation examples, in addition to the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes described above, another interprediction mode called JOINT_NEWMV may be introduced for a mode in which the MVDs of reference list 0 and reference list 1 are signaled together. Specifically, when the interprediction mode is indicated as NEW_NEWMV, the MVDs of reference list 0 and reference list 1 are signaled separately, while when the interprediction mode is indicated as JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are signaled together. In particular, in the case of joint MVD, it may be necessary for only one MVD called joint_delta_mv to be signaled and transmitted in the bitstream, and the MVDs of reference list 0 and reference list 1 may be derived from joint_delta_mv. The derived MVD may then be combined with the reference motion vectors in reference list 0 or reference list 1 to generate two motion vectors for identifying the position of the reference block for composite inter prediction.

[0195]

[0217] In some implementations of composite interpretation, the JOINT_NEWMV mode may be signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. In such implementations, the bitstream may contain syntax for instructing any one of these alternative composite interpretation modes at any of the various signaling levels (e.g., sequence level, picture level, frame level, slice level, tile level, superblock level, etc.). Alternatively, the JOINT_NEWMV mode may be run as a submode of the NEW_NEWMV mode. In other words, during the NEW_NEWMV mode, the two MVDs of two reference blocks are either signaled together (and thus a JOINT_NEWMV submode) or not signaled at all (another submode of the NEW_NEWMV mode). In such an implementation, a first syntax element may be included in the bitstream to indicate one of the following modes: NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV. When the first syntax element indicates that the NEW_NEWMV mode has been selected for a coded block, a second syntax element may be further included in the bitstream and extractable by the decoder to indicate whether the MVD of the coded block is signaled separately or together.

[0196]

[0218] In the case of joint MVD implementations in composite interpretation, the MVD associated with the reference MV may be derived from a signaled joint MVD from a bitstream, such as the joint_delta_mv described above. Such a derivation may involve scaling the signaled joint MVD to obtain, for example, one or both of the two MVDs. In other words, the signaled joint MVD may be scaled before being added to the motion vector predictor (MVP) or reference MV. As a result of scaling, the precision or pixel resolution of the scaled MVD may differ from the acceptable precision of the motion vector difference. In some implementations, such an MVD scaled from a jointly signaled MVD may first be quantized to the acceptable precision of the MVD of the current picture or slice or tile or superblock or encoded block before being added to the reference MVD for generating the motion vector.

[0197]

[0219] In some implementations, the frame index of the reference frame in the composite interprediction mode may be signaled in the bitstream. The frame index may correspond to a picture order counter (POC) associated with the reference frame. The distance between the reference frame and the current frame may be defined and expressed as the difference between the corresponding POCs. The direction of the reference frame (before or after the current frame) may be expressed by its sign. Thus, a signed distance may be used to represent the position of the reference frame relative to the current frame. The reference frames for the composite interprediction mode are sometimes referred to as the first reference frame and the second reference frame.

[0198]

[0220] In some implementations, when the JOINT_NEWMV mode is signaled and the POC distances between two reference frames and the current frame are different, the MVD may be scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, the distance between reference frame list 0 and the current frame may be represented as td0, and the distance between reference frame list 1 and the current frame may be represented as td1. When td0 is equal to or greater than td1, joint_mvd may be used directly for reference list 0, and the MVD for reference list 1 may be derived from joint_mvd based on equation (1).

[0199]

number

[0200]

[0221] Otherwise, if td1 is equal to or greater than td0, joint_mvd is used directly for base list 1, and the MVD for base list 0 is derived from joint_mvd based on equation (2).

[0201]

number

[0202]

[0222] Returning to the motion vector predictors (MVPs) and MVP lists described above, in some implementations, such motion vector predictor lists may be established for a group of coding blocks (e.g., a superblock) such that each prediction block in the superblock contains motion vector candidates for predicting motion vectors. The maximum size of the MVP list may be predefined or configured. Candidates in the MVP list may be established using a predefined set of rules. For example, these candidates may be selected from already reconstructed motion vectors belonging to blocks spatially close to the current coding block or superblock (called spatial motion vector predictors, or SMVPs). Alternatively or additionally, these candidates may belong to blocks in the base frame of the current coding block or superblock (called temporal motion vector predictors, or TMVPs). Spatial motion vector predictors may be adjacent or non-adjacent SMVPs. Adjacent SMVPs may refer to motion vector predictors belonging to prediction blocks adjacent to the current coding block or superblock. Non-adjacent SMVPs may refer to motion vector predictors belonging to prediction blocks that are not directly adjacent to the current coding block or superblock. Other types of MVP candidates may be further derived from the reconstructed motion vector. As another example, one or more additional MVP banks may be maintained as one of the sources for establishing the MVP list, as will be described in more detail below.

[0203]

[0223] The MVP list may be constructed to hold a predetermined number of reconstructed MVP candidates (SMVP, TMVP, or other derived MVPs, or other types of MVP candidates) on both the encoder and decoder sides of the current coded block or superblock. When encoding the current predicted block in interprediction mode, the encoder should select an MVP from the candidates in the MVP candidate list that provides the best coding efficiency as a motion vector predictor for the current predicted block. The index of the selected MVP in the MVP list may be signaled in the bitstream. Correspondingly, the decoder should update the MVP list for the current coded block or superblock as the bitstream is reconstructed, extract the MVP index of the current interpredicted predicted block, retrieve an MVP from the MVP candidate list according to the extracted MVP index in the MVP list, and use the MVP as a motion vector predictor for the current predicted block to reconstruct the motion vector of the current predicted block (for example, by combining the motion vector predictor extracted from the MVP list with the corresponding MVD). For example, the MVP list may represent a stack with a predetermined fixed size.

[0204]

[0224] For example, SMVP may be derived from spatially close prediction blocks, which include spatially close and adjacent prediction blocks and spatially close but not adjacent blocks, where spatially close and adjacent prediction blocks are directly adjacent to the current block or superblock above or to the left of the block or superblock (assuming these are preceding blocks that have already been reconstructed), and spatially close but not adjacent blocks are close to the current block or superblock but not directly adjacent. An example of a set of spatially close prediction blocks for a Luma block or superblock is illustrated in Figure 18.

[0205]

[0225] Figure 18 shows superblock 1802, which in this example contains 16 prediction blocks (each prediction block is interpreted). For example, each prediction block may be an 8x8 block. Superblock 1802 may be associated with various adjacent and non-adjacent prediction blocks, indicated by various smaller squares around superblock 1802 in Figure 18. Only adjacent prediction blocks are shown at the top and left, as they represent prediction blocks that have already been reconstructed (e.g., in a decoder or in the loop decoding unit of an encoder). Adjacent prediction blocks shaded with diagonal lines in Figure 18 represent adjacent and adjacent prediction blocks, while other prediction blocks represent adjacent and non-adjacent prediction blocks of the current superblock 1802.

[0206]

[0226] In some implementations, the spatially adjacent prediction blocks in Figure 18 may be examined or searched against a predefined set of rules to determine whether any of the motion vectors of these blocks should be considered as entries in the MVP candidate list. For example, these adjacent prediction blocks may be examined or searched to determine whether they are associated with the same reference frame index (for interpretation) as the current prediction block in superblock 1802. If they do not share the same reference frame index as the current prediction block, their motion vectors may not be eligible for entry in the MVP candidate list. Since the size of the MVP candidate list is limited (e.g., 4 or another number of candidates), these adjacent prediction blocks are examined / searched and ranked in a predetermined order. The search order may be predefined. An example of a predefined search order is illustrated in Figure 18, where adjacent and contiguous prediction blocks on the upper side are examined first from left to right, as indicated by arrow 1, and then adjacent and contiguous prediction blocks on the left side are examined from top to bottom, as indicated by arrow 2. As indicated by arrow 3, the adjacent prediction block to the upper right of superblock 1801 is then examined. Following this, as indicated by arrow 4, the adjacent prediction block to the upper left corner of the current superblock 1802 is examined, and then, in the search order of this example, the adjacent prediction blocks in the second row from the top, the second column from the left, the third row from the top, and the third column from the left are examined. The search order within adjacent rows of prediction blocks may be from left to right, while the search order within adjacent columns of prediction blocks may be from top to bottom. During the check and search process, adjacent prediction blocks associated with the same reference frame as the current prediction block may be flagged as eligible for entry in the MVP candidate list if there is still space available in the list. The terms MP candidate list and MP list are used interchangeably.

[0207]

[0227] In some implementations, adjacent SMVP candidates are placed in the MVP list first, before TMVPs, if their corresponding adjacent prediction block shares the same reference frame as the current block. Non-adjacent SMVP candidates are placed in the MVP list after TMVPs, provided there is still space and their corresponding adjacent prediction block shares the same reference frame as the current block. Therefore, in some implementations, all SMVP candidates have the same reference picture as the current block. For example, if the current prediction block has a single reference picture (single-reference interpretation), an MVP candidate having the same single reference picture as the current block's reference picture is eligible to be placed in the MVP candidate list, or an MVP candidate having a composite reference picture, where one of the two reference pictures is the same as the reference picture of the current block which has a single reference picture, is eligible to be placed in the MVP candidate list. As another example, in the case of the current block having a composite criterion picture (e.g., two criterion pictures), a candidate block may be added to the MVP candidate list (if there is still space) only if the candidate block is also predicted under the composite criterion and the two criterion pictures of the candidate block are the same as the two criterion pictures of the current block.

[0208]

[0228] In some other implementations, the SMVP candidate search may include more or fewer adjacent and non-adjacent block rows or columns than those illustrated in Figure 18.

[0209]

[0229] In some implementations, as described above and in more detail below, the TMVP may also be derived from the reference picture or frame of the current prediction block and included in the MVP candidate list. To generate the time MV predictor, the MVs of multiple frames are initially stored along with the frame index associated with each frame. Then, for each prediction block of the current frame (e.g., every 8x8 block), the MVs of multiple frames through which the spatial trajectory passes in the current block are identified and stored in the time MV buffer along with the frame indices of the corresponding multiple frames. For interpretation using a single reference frame, the MVs are stored in the 8x8 unit to perform time motion vector predictions for future frames, regardless of whether the reference frame is a forward or backward reference frame. For composite interpretation, only the forward MVs are stored in the 8x8 unit to perform time motion vector predictions for future frames.

[0210]

[0230] An example is shown in Figure 19. In this example, MV ref The MV of frame 1906 (R1, right-hand side of Figure 19) for the block indicated within frame 1906, called the MV, points from R1 to its reference frame 1908 (left-hand side of Figure 19). In this example, the MV ref This passes through the 8x8 block 1910 of the current frame 1902. MV ref This is stored in the time MV buffer associated with this 8x8 block 1910. In some implementations, the motion vector MV passing through a specific current prediction block is stored. ref To identify MVs, multiple frames, such as R1, are scanned in a predefined order, for example, LAST_FRAME, BWDREF_FRAME, ALTREF_FRAME, ALTREF2_FRAME, and LAST2_FRAME in reference frame list 0 and reference frame list 1. In some implementations, MVs from higher indexed reference frames (in scan order) do not replace previously identified MVs assigned by lower indexed reference frames (in scan order).

[0211]

[0231] Next, given the predefined block coordinates of the current block within the current frame, the associated MV stored in the time MV buffer is used to derive a TMVP that points from the current block in Figure 19 to the reference frame 1904, such as MV0 in Figure 19. ref It may be identified and projected onto the current block. Such a TMVP may be considered an MVP candidate in the MVP candidate list after an adjacent SMVP search, for example, as described above or in other orders, if there is space in the list.

[0212]

[0232] In some implementations, the TMVP may be searched or determined at the granularity of a group of prediction blocks (e.g., a superblock). In other words, an entire superblock may be associated with the same one or more TMVP candidates. A superblock may contain multiple prediction blocks. For example, a superblock may be 16x16 or 32x32, while the prediction blocks may be 8x8. Thus, a superblock may contain 4 or 16 prediction blocks accordingly. For example, motion vectors may be derived and reconstructed at the prediction block level. Therefore, searching for the TMVP of a superblock may involve scanning multiple prediction blocks to identify their motion vectors and determining which of these motion vectors, or scaled versions thereof, should be considered TMVP candidates, and in what order should the MVP candidate list for the superblock be considered? The multiple prediction blocks scanned in the TMVP candidate search may include blocks both inside the superblock (referred to as internal TMVP blocks) and outside the superblock (referred to as external TMVP blocks). In some implementations, the set of prediction blocks to be scanned, and the order in which the set of prediction blocks is scanned, may be predefined, preconfigured, or adaptively configured. The block positions of the set of prediction blocks relative to the superblock may, accordingly, be redefined, preconfigured, or adaptively configured.

[0213]

[0233] An example is shown in Figure 20. In Figure 20, a 16x16 superblock contains four prediction blocks B0, B1, B2, and B3. Thus, these prediction blocks form an inner block. Additionally, a set of outer prediction blocks, such as B4 through B6, may be predefined. These prediction blocks, which are both inner and outer TMVP blocks, may be searched / tested to determine and identify valid TMVP candidates for the superblock containing B0 through B3. Thus, in this example, up to seven blocks are checked for valid time MV predictors. The search order may be predefined from these inner and outer prediction blocks. For example, the scan may proceed from the inner blocks to the outer blocks. During the scan process, the TMVP for each block scanned, if identified, is placed in the MV predictor list if there is still space after other scans (including, but not limited to, higher-ranked SMVP scans). For example, and as described above, time MV predictors may be checked after adjacent spatial MV predictors but before non-adjacent spatial MV predictors. In one example, the TMVP search order may be B0→B1→B2→B3→B4→B5→B6. Various other exemplary search implementations under different configurations are provided in further detail below in relation to Figures 25-30.

[0214]

[0234] In some implementations, all spatial and temporal MV candidates may be pooled for the derivation of MV predictors, and each predictor may be assigned a weight determined during a scan of spatially and temporally adjacent predictor blocks. Based on the associated weights, the MVP candidates are sorted, ranked, and up to a predetermined number of candidates (e.g., four) are identified and added to the MVP list. This list of MV predictors is also called the Dynamic Criteria List (DRL) and is used in dynamic MV prediction mode as described above. The DRL may be implemented as a list of MVP indices.

[0215]

[0235] In some implementations, if the MVP list is not yet full after pooling TMVPs and SMVPs, an additional search for MVP candidates may be performed, and such additional MVP candidates may be used to complete the MVP list. These additional MVP candidates may include, for example, whole MVs, zero MVs, and composite MVs combined without scaling. In some implementations, only adjacent TMVPs and SMVPs may be pooled before the additional search. In other words, in these implementations, non-adjacent TMVPs may be excluded.

[0216]

[0236] In some implementations, adjacent SMVP candidates, TMVP candidates, and non-adjacent SMVP candidates (if permitted) added to the MVP list may be further sorted. Examples of such sorting processes may be based on candidate weights. Candidate weights may be predefined, for example, based on the overlap area between the current block and the candidate block in space.

[0217]

[0237] In some implementations, certain types of MVPs may not need to be considered and therefore may not be affected during the sorting process. For example, external / non-adjacent and TMVP candidates may not be considered during the sorting process, meaning that the sorting process only affects spatially adjacent candidates.

[0218]

[0238] In some other implementations, further MVP candidates may be derived. MVP candidates may be derived for a single reference picture and in composite modes. For example, the above restriction on SMVPs, which requires that adjacent blocks share a reference frame with the current block in order to be considered an SMVP candidate, may be relaxed. For example, in the case of a single interpretation, if the reference frame of an adjacent prediction block of the current block is not the same as the reference frame of the current block, but they are in the same direction (forward or backward), a time scaling algorithm may be used to scale the MV of the adjacent prediction block to the reference frame of the current block in order to form the MVP of the motion vector of the current block.

[0219]

[0239] An example is shown in Figure 21. Figure 21 shows a motion vector MV12110 from the adjacent prediction block 2102 of the current superblock 2104 in the current frame 2105 to the reference prediction block 2106 of the adjacent prediction block 2102 in the reference frame 2108, such motion vector MV12110 may be scaled according to the frame position of the current frame 2105, the reference frame 2106 of the adjacent prediction block 2102, and the reference frame 2120 of the current superblock 2104 to generate an MV02130 as shown in Figure 21 as an MVP candidate for the current superblock 2104.

[0220]

[0240] As another example, in a situation where the current frame 2204 has a current superblock 2202 under composite interpretation, as shown in Figure 22, MVs composed of different adjacent prediction blocks of the current superblock 2202 are utilized to derive the MVP of the current block, for example, when the base frame of the configured MV is the same as the current block 2202. In Figure 22, the configured MVs (MV2, MV3) have the same base frames 2210 and 2212 as the current block 2202, but these may be from different adjacent blocks 2206 and 2208.

[0221]

[0241] The above implementation example may be partially extended to implement a motion vector candidate bank mechanism. For example, multiple MV bank buffers may be implemented, each bank buffer may be associated with a unique reference frame type, such unique reference frame type corresponding to a single or pair of reference frames covering single and composite intermodes, respectively. All bank buffers may be implemented with the same size. When a new MV is added to a full bank buffer, existing MVs are moved out to make space for the new MV.

[0222]

[0242] The encoded block (e.g., the superblock) can refer to the MV candidate bank to collect reference MV candidates for the MV list, in addition to those obtained for the MV list in the other embodiments described above. After encoding the superblock, the MV bank is updated with the MVs used by the prediction block of the superblock.

[0223]

[0243] In some implementations, encoding may be performed within a tile. Each tile may be associated with an independent MV reference bank, which is used by all superblocks within the tile. At the beginning of encoding each tile, the corresponding bank is emptied. Then, while encoding each superblock within this tile, MVs from the bank may be used as MV reference candidates. At the end of encoding the superblocks, the bank is updated.

[0224]

[0244] An illustrative bank update process based on superblocks within a tile is illustrated in Figure 23. In Figure 23, tile 2302 contains multiple superblocks (SBs). The current superblock being encoded is shown as 2304. Superblocks above or to the left of 2304 in the same row are encoded (or reconstructed by the decoder), while superblocks below or to the right of 2304 in the same row are to be encoded (or reconstructed by the decoder), as indicated in Figure 23. After superblock 2304 is encoded, as indicated by the arrows in Figure 23, the first (e.g., up to 64) candidate MVs used by each encoded block inside superblock 2304 are added to bank 2306. The update may further include a removal process.

[0225]

[0245] In some implementations, to use the MV candidate bank, after the aforementioned MV candidate scan of the superblock has been performed, if there are still empty slots in the MV candidate list, the codec may refer to the MV candidate bank (in the buffer matched with the base frame type) for additional MV candidates. For example, if an MV in the bank does not already exist in the list, it may be added to the MVP candidate list by proceeding from the end of the bank to the beginning until the MVP list is full.

[0226]

[0246] According to the various implementations described above, the construction of an example of an MVP list may follow the search / processing order illustrated in Figure 24, which includes adjacent SMVPs (2402), a sorting process for existing candidates (2404), TMVPs (2406), non-adjacent SMVPs (2408), derived candidates (2410), reserve or additional MVP candidates (2412), and candidates from the base MV candidate bank (2414), until a predefined number of MVP candidates are identified.

[0227]

[0247] As mentioned above, the space or number of MVP candidates in the MVP list may be limited. For example, the MVP list may be limited to four MVP candidates. Often, there may be a large number of identifiable TMVPs according to the implementation form described above or other embodiments, and therefore, after searching for adjacent SMVPs following the procedure in Figure 24, for example, the MVP list should almost always be filled. Thus, other MVP candidates (e.g., MVP candidates belonging to non-adjacent SMVPs, other MVPs, additional MVPs, or MVPs from the MVP bank) may almost always be deprived of the opportunity to be evaluated for inclusion in the MVP list. Therefore, in some implementation examples described below, a limit may be imposed on the number of TMVP candidates that may be included in the MVP candidate list. Such limits may be predefined or configured at various coding levels. Such predetermined or configured limits on TMVPs may be based on video type and statistics to facilitate the achievement of improved coding gain.

[0228]

[0248] In some implementations, the number of TMVP candidates that may be inserted into the MVP candidate list may be limited to N, where N is a positive integer. For example, N may be 1 or 2. In such embodiments, the resulting MVP list may not be excessively dominated by TMVPs and may be more balanced by including other MVP types.

[0229]

[0249] In some further implementations, the process for constructing the MVP list may be designed so that various types of MVPs, including the aforementioned SMVP, TMVP, and other types of derived and / or additional MVPs, complement each other to bring together improved coding efficiency.

[0230]

[0250] For example, considering that adjacent and contiguous blocks or adjacent and non-adjacent blocks scanned to find SMVP candidates are located above or to the left of the current superblock, complementary TMVPs may be more likely if the scan of internal predictor blocks within the current superblock for TMVP begins with internal predictor blocks within the current superblock that are further away from those above and to the left of the current superblock. In such cases, TMVPs of predictor blocks that are further away from those scanned for SMVP, and therefore more complementary to them, are provided with an earlier opportunity to create an MVP list.

[0231]

[0251] Such internal prediction block scan orders may be clearly distinguishable from those described above with respect to Figure 20. There can be various concrete implementations of such scans. For example, some orders may be better than others in terms of improving overall encoding efficiency depending on the type and / or characteristics of the video frame being encoded. Scan orders for this purpose between internal prediction blocks may be predefined or configured and signaled in the bitstream. For example, there may be predefined orders with predefined order indices. The selection of a scan order from among several predefined scan orders may be performed by the encoder, and the index of the selected scan order may be signaled to the superblock in the bitstream used.

[0232]

[0252] In some implementations, internal prediction blocks may be scanned against the superblock to identify candidate TMVPs in the order illustrated in Figures 25 and 26, where the internal blocks are scanned from right to left and bottom to top. Specifically, Figure 25 shows an example superblock of size 16x16, each containing four 8x8 prediction blocks. The example scan order of the internal blocks proceeds from right to left and bottom to top, for example, B3→B2→B1→B0. Similarly, Figure 26 shows an example superblock of size 32x32, each containing sixteen 8x8 prediction blocks. The example scan order of the internal blocks proceeds from right to left and top to bottom, for example, B15→B14→B13→B12→B11→B10→B9→B8→B7→B6→B5→B4→B3→B2→B1→B0. These implementations may, for example, provide SMVP candidates by scanning for TMVP candidates, starting with adjacent and neighboring blocks and adjacent and non-adjacent blocks, and further away (at least horizontally), and thus complementing them, starting with internal blocks.

[0233]

[0253] In some alternative implementations, internal prediction blocks may be scanned against the superblock to identify candidate TMVPs in the order illustrated in Figures 27 and 28. Specifically, Figure 27 shows an illustrative 16x16 superblock, each containing four 8x8 prediction blocks. The illustrative scan order of the internal blocks proceeds from bottom to top and left to right, for example, B3→B1→B2→B0. Similarly, Figure 28 shows an illustrative 32x32 superblock, each containing sixteen 8x8 prediction blocks. The illustrative scan order of the internal blocks proceeds from right to left and top to bottom, for example, B15→B11→B7→B3→B14→B10→B6→B2→B13→B9→B5→B1→B12→B8→B4→B0. These implementations may, for example, provide SMVP candidates by scanning for TMVP candidates, starting with adjacent and neighboring blocks and adjacent and non-adjacent blocks, and then further away (at least vertically) and thus complementing them, starting with internal blocks.

[0234]

[0254] In some other alternative implementations, the internal prediction blocks may be scanned against the superblock to identify candidate TMVPs in the order illustrated in Figures 29 and 30. Specifically, Figure 29 shows an illustrative 16x16 superblock, each containing four 8x8 prediction blocks. The scan order of the illustrative internal blocks proceeds from bottom to top and left to right, for example, B3→B2→B1→B0. Similarly, Figure 30 shows an illustrative 32x32 superblock, each containing sixteen 8x8 prediction blocks. The scan order of the illustrative internal blocks proceeds from right to left and top to bottom, for example, B15→B14→B11→B13→B10→B7→B12→B9→B6→B3→B8→B5→B2→B4→B1→B0. These implementations may, for example, provide SMVP candidates by scanning for TMVP candidates, starting with adjacent and neighboring blocks and adjacent and non-adjacent blocks, and further away (both vertically and horizontally), and thus complementing them, starting with internal blocks.

[0235]

[0255] As long as there is still space in the MVP list when an internal block is scanned, up to a limited number N TMVPs may be identified and included in the MVP list, following the scan order of the example internal predictive block above. If fewer than N TMVPs are identified when an internal block is scanned, and there is still space in the MVP list, and the next MVP candidate to be scanned is an external block, then such external block is scanned as long as the total number of TMVPs is N or less (if N as a whole is configured for all TMVPs) and there is still space in the MVP list.

[0236]

[0256] In some other variations of the above implementation examples shown in Figures 25 to 30, M candidates may be skipped during the scanning of the internal prediction block, where M can be a non-negative value. When M is equal to 0, this is the same as the implementation in Figures 25 to 30.

[0237]

[0257] The internal prediction blocks of candidates that will be skipped during the TMVP scan may be, for example, near the top and / or left side of the superblock, in order to facilitate the generation of TMVPs in MVP lists that are less complementary to other MVP candidates such as SMVPs, as described above. The internal blocks to be skipped may be in units of rows or columns of the prediction block; for example, one or more rows or columns of an internal block may be skipped during the TMVP scan.

[0238]

[0258] For example, if there are two or more internal TMVP block rows, the first internal TMVP block row (the top row) may be skipped. Specifically, as shown in Figures 25, 27, and 29, B0 and B1 may be skipped, and only the internal TMVP blocks B2 and B3 are checked in the respective order shown in Figures 25, 27, and 29. For example, as shown in Figures 26, 28, and 30, B0 through B3 may be skipped, and the remaining internal blocks are checked in the respective order shown in Figures 26, 28, and 30.

[0239]

[0259] As another example, if there are two or more internal TMVP columns, the first internal TMVP block column may be skipped. For example, as shown in Figures 25, 27, and 29, B0 and B2 may be skipped, and only the internal TMVP blocks B1 and B3 are checked in the respective order shown in Figures 25, 27, and 29. For example, as shown in Figures 26, 28, and 30, B0, B4, B8, and B12 may be skipped, and the remaining internal blocks are checked in the respective order shown in Figures 26, 28, and 30.

[0240]

[0260] As another example, if there are two or more internal TMVP columns and two or more internal TMVP rows, both the first internal TMVP block column and the first internal TMVP row may be skipped. For example, as shown in Figures 25, 27, and 29, B0, B1, and B2 may be skipped, and only the internal TMVP block B3 is checked. For example, as shown in Figures 26, 28, and 30, B0 through B3, B4, B8, and B12 may be skipped, and the remaining internal blocks B5, B6, B7, B9, B10, B11, B13, B14, and B15 are checked in the respective order shown in Figures 26, 28, and 30.

[0241]

[0261] In some implementations, during a TMVP scan, some of the internal prediction blocks of the current superblock may be skipped, as well as some of the outer TMVP blocks. These outer TMVP blocks are not checked for insertion into the MVP list, regardless of the size of the current superblock (e.g., 16x16 or 32x32). For example, in Figure 20, outer blocks such as B4, B5, and B6 may be skipped, and they do not need to be checked or inserted into the MVP list.

[0242]

[0262] The upper limit for N motion vector predictors of TMVP candidates in the MVP list may apply to inner blocks only, or outer blocks only, or to both inner and outer blocks as a whole. In some implementations, up to one TMVP candidate from an inner TMVP block may be inserted into the MVP list. In some other implementations, up to one TMVP candidate from an outer TMVP block may be inserted into the MVP list. In some other implementations, up to one TMVP candidate from both inner and outer TMVP blocks may be inserted into the MVP list.

[0243]

[0263] In some implementations, the number of prediction blocks to be scanned to identify a TMVP candidate may be limited to K, where K is a positive integer. For example, K may be limited to 4. Thus, only K blocks may be scanned for TMVP candidates. Such a restriction may apply to scanning internal TMVP blocks, or external TMVP blocks, or internal and external TMVP blocks as a whole. In some implementations, separate K values ​​may be assigned to internal and external blocks. These restrictions may be predefined or fixed. They may be independent of the superblock size. Alternatively, these restrictions may be adaptively configured and indicated in the bitstream. For example, K internal blocks may be checked to identify a TMVP for insertion into the MVP list. The K internal prediction blocks to be checked may be uniformly distributed within a region where TMVP blocks are aggregated, for example, according to some predefined or configured pattern. For example, the internal prediction blocks to be checked may be evenly distributed between B0 and B3 in Figures 25, 27, and 29, or between B0 and B15 in Figures 26, 28, and 30. As another example, K may be applied to an internal prediction block and set to 1. In other words, only one internal block is checked. For example, the internal block to be checked may be block B2 in Figures 25, 27, and 29, or block B15 in Figures 16, 28, and 30.

[0244]

[0264] The upper limit on the number of TMVPs to be inserted into the MVP list (the integer N above) and the upper limit on the number of prediction blocks to check for TMVP candidates (the integer K above) may both be specified. These may be predefined or configured. They may be interdependent. For example, N and P may be specified as N and the difference between P and N. In some constraints, P and N may be specified as the same integer.

[0245]

[0265] In some implementations, the upper limit for the N motion vector predictors of a TMVP candidate in the MVP list may be conditionally applied. For example, the upper limit N may be applied to the current superblock if the current superblock is encoded in bidirectional interprediction mode and the two reference frames point to the current frame in different directions (the POC of one reference frame is smaller than the POC of the current frame, and the POC of the other reference frame is larger than the POC of the current frame), and may not be applied otherwise.

[0246]

[0266] In some implementations, the conditional application of the upper limit N to TMVPs in the MVP list may refer to the number of SMVP candidates already added to the MVP list. For example, if there are already N2 SMVP candidates added to the MVP list, the upper limit N2 for TMVP candidates that can be added to the MVP list may be effective; otherwise, the limit of N2 will not apply. For example, N1 may be set to 2 and N2 may be set to 1. N1 and N2 may be predefined or configured. N1 and N2 may be superblock size dependent or independent of the superblock size.

[0247]

[0267] Moving on to some aspects of signaling MVPs in a bitstream, in some implementations, the selection of an MVP from an MVP list or DRL index may correlate with the number of TMVPs inserted into the MVP list. Therefore, the context for signaling the DRL index and inter-prediction mode may depend on whether at least L TMVPs have been inserted into the MVP list, where L is any positive integer. In one example, L may be set to 1. In other words, if there are at least L TMVPs in the MVP list, a first context may be used to signal the DRL index and inter-prediction mode for a particular prediction block; otherwise, a second context may be used to signal the DRL index.

[0248]

[0268] In some implementations, the position of a candidate TMVP block in the MVP list may be implicitly signaled, for example, through other encoded information, which may include, but not be limited to, the motion vectors of spatially adjacent blocks or the number of spatial MVPs already inserted in the MVP list.

[0249]

[0269] Figure 31 shows a flowchart 3100 of an exemplary method that follows the principles underlying the above implementation for searching for TMVP candidates and constructing an MVP list. The exemplary method flow begins at S3101. In S3110, after receiving the video bitstream, it is determined that the current prediction block should be interpreted by a reference block in a reference frame, and that the motion vector of the current prediction block should be predicted by a reference motion vector. In S3120, the set of candidate prediction blocks in the current frame is identified as a search pool of TMVP candidates for the current prediction block. In S3130, the set of candidate prediction blocks is searched in search order to identify up to N TMVP candidates for the current prediction block, and the search ends in response that N TMVP candidates have been identified, where N is a positive integer. In S3140, an MVP list is constructed that points to a set of MVP candidates, which includes a set of SMVP candidates and one or more of up to N TMVP candidates. In S3150, the MVP index of the current prediction block is extracted from the video stream. In S3160, a reference motion vector for inter-predicting the current prediction block is identified according to the extracted MVP index and MVP list. The exemplary method stops at S3199.

[0250]

[0270] In the embodiments and implementations of this disclosure, any of the processes and / or operations may be combined or arranged in any quantity or order as desired. Two or more of the processes and / or operations may be performed in parallel. The embodiments and implementations in this disclosure may be used separately or in combination in any order. Furthermore, each of the methods (or embodiments), encoders, and decoders may be implemented by processing circuit equipment (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-temporary computer-readable medium. The embodiments in this disclosure may apply to luma blocks or chroma blocks. The term block may be interpreted as a prediction block, coding block, or coding unit, or CU. The term block may also be used to refer to a transformation block. In the following sections, when block size is described, block size may refer to block width or height, or the maximum width and height, or the minimum width and height, or the area (width * height), or the aspect ratio of the block (width:height, or height:width).

[0251]

[0271] The techniques described above can be implemented as computer software physically stored on one or more computer-readable media, using computer-readable instructions. For example, Figure 32 shows a computer system (3200) suitable for implementing a particular embodiment of the subject matter of the disclosure.

[0252]

[0272] Computer software can be encoded using any suitable machine language or computer language, and these languages ​​may undergo mechanisms such as assembly, compilation, and linking to produce code with executable instructions, either directly or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0253]

[0273] The command is executable on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0254]

[0274] The components shown in FIG. 32 for the computer system (3200) are exemplary in nature and are not intended to suggest any limitation as to the usage or functional scope of the computer software implementing embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency relationships and requirements regarding any one or combination of the components exemplified in the exemplary embodiments of the computer system (3200).

[0255]

[0275] The computer system (3200) may include a specific human interface input device. Such a human interface input device can respond to input by one or more human users, for example, through tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown). Also, the human interface device can be used to capture certain media that is not necessarily directly related to conscious input by humans, such as audio (speech, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).

[0256]

[0276] The input human interface device may include one or more of the following (only one of each is depicted): keyboard (3201), mouse (3202), trackpad (3203), touchscreen (3210), data glove (not shown), joystick (3205), microphone (3206), scanner (3207), and camera (3208).

[0257]

[0277] The computer system (3200) may further include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (3210), data glove (not shown), or joystick (3205), but which may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (3209), headphones (not shown)), visual output devices (e.g., screens (3210) for including CRT screens, LCD screens, plasma screens, OLED screens, etc., each with or without touchscreen input capability, each with or without tactile feedback capability - some of these may have the ability to output two-dimensional visual output or three-dimensional or more output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0258]

[0278] Furthermore, the computer system (3200) may include human-accessible storage devices and their associated media, such as CD / DVD ROM / RW (3220) media including CD / DVD media (3221), thumb drives (3222), removable hard drives or solid-state drives (3223), older magnetic media such as tapes and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0259]

[0279] Furthermore, those skilled in the art will understand that the term “computer-readable medium” as used in conjunction with the subject matter of this disclosure does not encompass a transmission medium, carrier wave, or other transient signal.

[0260]

[0280] Furthermore, the computer system (3200) may include an interface (3254) to one or more communication networks (3255). For example, the network may be wireless, wireline, or optical. In addition, the network may be local, wide-area, metropolitan, automotive and industrial, real-time, or latency-tolerant. Examples of networks include local area networks such as Ethernet®, cellular networks to include wireless LAN, GSM, 3G, 4G, 5G, LTE®, etc., TV wireline or wireless wide-area digital networks to include cable TV, satellite TV, and terrestrial television broadcasting, and automotive and industrial networks to include CAN buses, etc. Certain networks generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (3249) (e.g., a USB port on the computer system (3200)), while others are generally integrated into the core of the computer system (3200) by attachments to system buses as described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (3200) can communicate with other entities. Such communication can be one-way and receive-only (e.g., television broadcasting), one-way and transmit-only (e.g., CANbus to a specific CANbus device), or bidirectional, for example, to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks are available for each of these networks and network interfaces as described above.

[0261]

[0281] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be attached to the core (3240) of the computer system (3200).

[0262]

[0282] The core (3240) may include one or more central processing units (CPUs) (3241), graphics processing units (GPUs) (3242), special programmable processing units in the form of field-programmable gate areas (FPGAs) (3243), hardware accelerators for specific tasks (3244), graphics adapters (3250), etc. These devices may be connected via a system bus (3248) along with read-only memory (ROM) (3245), random access memory (3246), internal mass storage such as internal user-inaccessible hard drives (3247), SSDs, etc. In some computer systems, the system bus (3248) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core's system bus (3248) or via a peripheral bus (3249). In one example, a screen (3210) can be connected to a graphics adapter (3250). Peripheral bus architectures include PCI, USB, etc.

[0263]

[0283] The CPU (3241), GPU (3242), FPGA (3243), and accelerator (3244) can be combined to execute specific instructions that can create the aforementioned computer code. This computer code can be stored in ROM (3245) or RAM (3246). Temporary data can also be stored in RAM (3246), while permanent data can be stored, for example, in internal mass storage (3247). High-speed storage and retrieval to and from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more CPUs (3241), GPUs (3242), mass storage (3247), ROMs (3245), RAM (3246), etc.

[0264]

[0284] Computer-readable media can have computer code to perform various computer execution operations. The media and computer code can be specifically designed and constructed for this disclosure, or they can be of a type that is well known and available to those skilled in the computer software technology.

[0265]

[0285] As a non-limiting example, a computer system having an architecture (3200), and specifically a core (3240), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) running software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage as described above, and media associated with specific storage of the core (3240) of a non-transient nature, such as core internal mass storage (3247) or ROM (3245). Software performing various embodiments of the present disclosure can be stored in such devices and executed by the core (3240). The computer-readable media can include one or more memory devices or chips, depending on the specific needs. The software can cause the core (3240) and specifically the processor of the core (including a CPU, GPU, FPGA, etc.) to execute specific processes, or specific parts of specific processes, as described herein, including defining data structures stored in RAM (3246) and modifying such data structures according to processes defined by the software. As an addition or alternative, a computer system may provide functionality resulting from logic, which is wired together or otherwise contained in circuitry (e.g., an accelerator (3244)), capable of operating in place of or in conjunction with software, to perform a particular process or a particular part of a particular process described herein. References to software may, where appropriate, encompass logic, and vice versa. References to computer-readable media may, where appropriate, encompass circuitry containing software for execution (such as an integrated circuit (IC)), circuitry containing logic for execution, or both. This disclosure encompasses any appropriate combination of hardware and software.

[0266]

[0286] While this disclosure has described several exemplary embodiments, there are many equivalents with modifications, rearrangements, and various substitutions, which are within the scope of this disclosure. Therefore, those skilled in the art are capable of devising numerous systems and methods, which, although not expressly illustrated or described herein, will be recognized as embodying the principles of this disclosure and thus remaining within its spirit and scope.

[0267] Addendum A: Acronyms JEM: Joint Exploration Model VVC: Versatile Video Coding BMS: Benchmark Set MV: Motion Vector HEVC: High Efficiency Video Coding SEI:Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Conversion Unit PU: Prediction Unit CTU: Encoding Tree Unit CTB: Encoded Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal-to-noise ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Solid State Drive IC: Integrated Circuit HDR: High Dynamic Range SDR: Standard Dynamic Range JVET: Joint Video Exploration Team MPM: Most Probable Mode WAIP: Wide Angle Intra Prediction CU: Coding Unit PU: Prediction Unit TU: Transform Unit CTU: Coding Tree Unit PDPC: Position Dependent Prediction Combination ISP: Intra Sub Part SPS: Sequence Parameter Setting PPS: Picture Parameter Set APS: Adaptation Parameter Set VPS: Video Parameter Set DPS: Decoding Parameter Set ALF: Adaptive Loop Filter SAO: Sample Adaptive Offset CC-ALF: Cross Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter CCSO: Cross Component Sample Offset LSO: Local Sample Offset LR: Loop Restriction Filter AV1: AOMedia Video 1 AV2: AOMedia Video 2 MVD: Motion Vector Difference CfL: Chroma from Luma SDT: Semi-decoupled tree SDP: Semi-decoupled partitioning SST: Semi-separated tree SB: Superblock IBC (or IntraBC): Intrablock Copy CDF: Cumulative Density Function SCC: Screen Content Coding GBI: Generalized Bidirectional Prediction BCW: Biprediction with CU-level weights CIIP: Combined Intra-Interface Prediction POC: Picture Order Count RPS: Reference Picture Set DPB: Decoded Picture Buffer MMVD: Merge mode with motion vector difference