Methods, apparatus and storage media for processing video blocks in a video stream

By adopting a joint coding signaling scheme with adaptive motion vector difference resolution, the problem of low coding efficiency in composite reference frame prediction is solved, and more efficient video compression is achieved.

CN116941243BActive Publication Date: 2026-05-26TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2022-07-27
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing video coding techniques struggle to effectively utilize motion vector differences with adaptive resolution for joint coding in composite reference frame prediction, resulting in low coding efficiency.

Method used

A joint coding signaling scheme with adaptive motion vector difference resolution is adopted. By receiving the video stream, it is determined whether the current video block applies joint motion vector difference coding and adaptive pixel resolution, and then decoded based on syntax elements.

Benefits of technology

It improves the efficiency of video encoding, reduces the amount of encoded data, and increases the video compression rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116941243B_ABST
    Figure CN116941243B_ABST
Patent Text Reader

Abstract

This disclosure generally relates to video coding, and more specifically, to a method and system for providing a signaling scheme for joint coding of motion vector difference with adaptive resolution in composite reference inter-frame prediction. An exemplary method for processing a current video block in a video stream is disclosed. The method includes: receiving the video stream; determining from the video stream whether joint motion vector difference (MVD) coding is applied to the current video block; determining from the video stream whether adaptive MVD pixel resolution is applied to the current video block; and decoding the current video block based on the joint MVD coding and whether the adaptive MVD pixel resolution is applied to the current video block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references

[0002] This application is based on and claims priority to U.S. non-provisional patent application No. 17 / 869,232, filed July 20, 2022, which claims priority to U.S. provisional patent application No. 63 / 307,413, filed February 7, 2022, entitled “Joint Coding for Adaptive Motion Vector Difference Resolution.” The entire contents of these earlier applications are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to video coding, and more specifically, to methods and systems for providing a signaling scheme for joint coding of motion vector differences with adaptive resolution in composite reference inter-frame prediction. Background Technology

[0004] The background description provided herein is for the purpose of presenting the general content of this disclosure. The extent of the work of the currently named inventors described in this background section and in various aspects of this specification does not indicate that it was prior art at the time of filing of this application, nor is it expressly or implied that it was acknowledged as prior art to this disclosure.

[0005] Video encoding and decoding can be performed using inter-frame picture prediction with motion compensation. Uncompressed digital video can comprise a series of pictures, each with a spatial size of, for example, a 1920x1080 luminance sample and an associated full or subsampled chrominance sample. This series of pictures can have a fixed or variable picture rate (or frame rate), for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, a 4:2:0 video with a 1920x1080 pixel resolution, a frame rate of 60 frames per second, and chrominance subsampling of 8 bits per pixel per color channel requires close to 1.5 Gbit / s of bandwidth. One hour of such video would require more than 600 GB of storage space.

[0006] One goal of video encoding and decoding is to reduce redundancy in the uncompressed input video signal through compression. Compression can help reduce bandwidth or storage requirements by two or more orders of magnitude in some cases. Lossless compression, lossy compression, and combinations thereof can be used. Lossless compression refers to a technique that reconstructs an exact copy of the original signal from the compressed original signal through a decoding process. Lossy compression refers to an encoding / decoding process in which the original video information cannot be fully preserved during encoding and cannot be fully recovered during decoding. When using lossy compression, the reconstructed signal may differ from the original signal, but despite some information loss, the distortion between the original and reconstructed signals is small enough that the reconstructed signal can be used for the intended application. In the case of video, lossy compression is widely used in many applications. The tolerable amount of distortion depends on the application. For example, users of some consumer video streaming applications may tolerate higher distortion than users of film or television broadcasting applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows the encoding algorithm to produce higher loss and a higher compression ratio.

[0007] Video encoders and decoders can utilize techniques from a wide range of categories and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.

[0008] Video codec techniques can include techniques called intra-frame coding. In intra-frame coding, sample values ​​are represented without reference to samples or other data from a previously reconstructed reference picture. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are encoded in intra-frame mode, this picture can be called an intra-frame picture. Intra-frame pictures and their derivatives (e.g., standalone decoder refresh pictures) can be used to reset the decoder state and are therefore used as the first picture in the encoded video bitstream and video session, or as a still image. The samples of the block after intra-frame prediction can then be transformed to the domain, and the transform coefficients can be quantized before entropy coding. Intra-frame prediction represents a technique for minimizing the sample values ​​in the pre-transformed domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are needed to represent the entropy-coded block at a given quantization step size.

[0009] Traditional intra-frame coding (e.g., intra-frame coding known from, for example, MPEG-2 generation coding techniques) does not use intra-frame prediction. However, some newer video compression techniques include attempts at block coding / decoding based on, for example, surrounding sample data and / or metadata, which is acquired during spatially adjacent coding / decoding and decoded before the data blocks being intra-frame coded or decoded. This technique is hereinafter referred to as "intra-frame prediction." It is important to note that, in at least some cases, intra-frame prediction uses only reference data from the current frame being reconstructed, and not reference data from other reference frames.

[0010] There can be many different forms of intra-prediction. When more than one such technique is used in a given video coding technique, the technique used is called an intra-prediction mode. One or more intra-prediction modes can be provided in a particular codec. In some cases, a mode can have sub-modes and / or can be associated with various parameters, and mode / sub-mode information and intra-coding parameters for video blocks can be encoded separately or together in the mode codeword. Which codeword is used for a given mode / sub-mode / parameter combination can have an impact on the coding efficiency gain through intra-prediction, and entropy coding techniques can also be used to convert codewords into bitstreams.

[0011] Some intra-frame prediction modes were introduced with H.264, improved in H.265, and further refined in newer coding techniques such as Joint Probe Model (JEM), Next-Generation Video Coding (VVC), and Baseline Matrix (BMS). Typically, for intra-frame prediction, already available neighboring sample values ​​can be used to form predictor blocks. For example, available values ​​for specific neighboring sample groups along certain directions and / or lines can be copied into a predictor block. References to the directions in use can be encoded in the bitstream or can be predicted themselves.

[0012] refer to Figure 1A The lower right corner depicts a subset of nine predictor directions defined by the 33 possible intra-frame predictor directions of H.265 (corresponding to the 33 angular modes of the 35 intra-frame modes specified in H.265). The point (101) where the arrows converge represents the sample being predicted. The arrows indicate the direction in which neighboring samples are used to predict the sample at 101. For example, arrow (102) indicates predicting sample (101) from one or more neighboring samples to the upper right at a 45-degree angle to the horizontal. Similarly, arrow (103) indicates predicting sample (101) from one or more neighboring samples to the lower left of sample (101) at a 22.5-degree angle to the horizontal.

[0013] Still referencing Figure 1AA 4×4 sample square block (104) is depicted in the upper left (represented by a bold dashed line). The square block (104) comprises 16 samples, each labeled with an "S" indicating its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from top to bottom) and the first sample in the X dimension (from left to right). Similarly, sample S44 is the fourth sample in both the Y and X dimensions of block (104). Since the block size is 4×4 samples, S44 is located in the lower right. An example reference sample following a similar numbering scheme is further shown. The reference sample is labeled with an "R" indicating its Y position (e.g., row index) and X position (column index) relative to block (104). In H.264 and H.265, predicted samples adjacent to the block being reconstructed are used.

[0014] Intra-frame picture prediction for block (104) can begin by copying reference sample values ​​from neighboring samples, based on the prediction direction represented by a signal. For example, assuming the encoded video bitstream includes signaling, for block 104, the signaling indicates the prediction direction of the arrow (102)—that is, the sample direction is predicted at a 45-degree angle to the horizontal from one or more prediction samples to the upper right. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Sample S44 is then predicted from reference sample R08.

[0015] In some cases, the values ​​of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample; especially when the orientation cannot be uniformly divided by 45 degrees.

[0016] As video coding technology continues to evolve, the number of possible directions has increased. For example, in H.264 (2003), nine different directions were available for intra-frame prediction. This increased to 33 directions in H.265 (2013), and at the time of this application, JEM / VVC / BMS supports up to 65 directions. Experimental studies have been conducted to help identify the most suitable intra-frame prediction directions, and certain techniques in entropy coding can be used to encode those most suitable directions with a small number of bits, thus accepting a certain bit disadvantage for the direction. Furthermore, sometimes the direction itself can be predicted from adjacent directions used in intra-frame prediction of already decoded adjacent blocks.

[0017] Figure 1B A schematic diagram (180) depicting 65 intra-frame prediction directions according to JEM is shown to illustrate the increase in the number of prediction directions in various coding techniques over time.

[0018] The mapping of bits representing intra-prediction directions to prediction directions in an encoded video bitstream can vary depending on the video coding technique; and the range can be, for example, from a simple direct mapping of prediction directions to intra-prediction modes, to codewords, to complex adaptive schemes involving the most probable modes, and similar techniques. However, in all cases, there may be some intra-prediction directions that are statistically less likely to occur in the video content compared to certain other directions. Since the goal of video compression is to reduce redundancy, in well-designed video coding techniques, those less likely directions can be represented by more bits than the more likely directions.

[0019] Inter-frame image prediction, or inter-prediction, can be based on motion compensation. In motion compensation, sample data from a previously reconstructed image or a portion thereof (the reference image), after being spatially offset along a direction indicated by a motion vector (hereafter referred to as MV), can be used to predict a newly reconstructed image or image portion (e.g., a patch). In some cases, the reference image can be the same as the image currently being reconstructed. MV can have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference image being used (similar to a temporal dimension).

[0020] In some video compression techniques, the current MV applicable to a region of sample data can be predicted based on other MVs, such as MVs related to other regions of sample data that are spatially adjacent to the region being reconstructed and whose decoding order precedes the current MV. Doing so significantly reduces the overall amount of data required to encode the MV by relying on eliminating redundancy in related MVs, thus improving compression ratio. MV prediction can work effectively, for example, because when encoding the input video signal obtained from the camera (called natural video), there is a statistical probability that a region larger than the region applicable to a single MV moves in similar directions within the video sequence. Therefore, in some cases, this larger region can be predicted using similar motion vectors derived from the MVs of neighboring regions. This results in the actual MV of a given region being similar to or the same as the MV predicted from the surrounding MVs. After entropy coding, such an MV can be represented with fewer bits than would be needed if the MV were encoded directly instead of predicted from neighboring MVs. In some cases, MV prediction can be an example of lossless compression of the signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating predictions based on multiple surrounding MVs.

[0021] H.265 / HEVC (ITU-T H.265 Recommendation, “High Efficiency Video Coding”, December 2016) describes various MV prediction mechanisms. In addition to the various MV prediction mechanisms specified in H.265, the following describes a technique hereinafter referred to as “spatial combining”.

[0022] Specifically, still refer to Figure 2 In spatial merging, the current block (201) includes samples that have been discovered by the encoder during the motion search process and are predictable based on the previous block of the same size that has been spatially shifted. MVs can be derived from metadata associated with one or more reference images, for example, from the most recent (in decoding order) reference image, using the MV associated with any of the five surrounding samples (denoted as A0, A1, and B0, B1, B2 (from 202 to 206) respectively), instead of directly encoding the MV. In H.265, MV prediction can utilize predictors of the same reference images used by adjacent blocks. Summary of the Invention

[0023] This disclosure generally relates to video coding, and more specifically, to methods and systems for providing a signaling scheme for joint coding of motion vector differences with adaptive resolution in composite reference inter-frame prediction.

[0024] In an exemplary embodiment, a method for processing a current video block of a video stream is disclosed. The method may include: receiving the video stream; determining from the video stream whether joint motion vector difference (MVD) coding is applied to the current video block; determining from the video stream whether adaptive MVD pixel resolution is applied to the current video block; and decoding the current video block based on the joint MVD coding and whether the adaptive MVD pixel resolution is applied to the current video block.

[0025] In the above embodiments, determining whether joint MVD coding and adaptive MVD pixel resolution are applied to the current video block may include: determining that the current video block is inter-coded in a composite reference mode based on at least two reference blocks associated with at least two corresponding motion vectors; and in response to determining that the current video block is inter-coded in the composite reference mode, extracting at least one syntax element from the video stream, the at least one syntax element indicating whether the motion vector difference associated with the at least two corresponding motion vectors is jointly represented or decoded and / or whether adaptive MVD pixel resolution is applied to the encoding of the motion vector difference.

[0026] In any of the above embodiments, when the at least one syntax element indicates that the motion vector difference is not jointly represented or decoded and no adaptive MVD pixel resolution is applied, the method may further include: extracting the MVD category or magnitude of the joint MVD from the video stream; determining the current MVD pixel resolution of the joint MVD based on the MVD category or magnitude; extracting the joint MVD from the video stream based on the current MVD pixel resolution; and deriving the at least two corresponding motion vectors based on the joint MVD.

[0027] In any of the above embodiments, when the at least one syntax element indicates that the motion vector difference is not jointly represented or decoded and no adaptive MVD pixel resolution is applied, the method may further include: extracting a joint MVD from the video stream based on a fixed MVD pixel resolution; and deriving the at least two corresponding motion vectors based on the joint MVD.

[0028] In any of the above embodiments, when the at least one syntax element indicates that the motion vector difference is not jointly represented or decoded by a signal and no adaptive MVD pixel resolution is applied, the method may further include: extracting MVD categories or magnitudes associated with the at least two corresponding motion vectors from the video stream; determining the current MVD pixel resolution of the at least two corresponding motion vectors based on the MVD categories or magnitudes; extracting individual MVDs from the video stream based on the current MVD pixel resolutions; and deriving the at least two corresponding motion vectors based at least on the individual MVDs.

[0029] In any of the above embodiments, when the at least one syntax element indicates that the motion vector difference is not jointly represented or decoded by a signal and no adaptive MVD pixel resolution is applied, the method may further include: extracting individual MVDs from the video stream based on a fixed MVD pixel resolution; and deriving the at least two corresponding motion vectors based at least on the individual MVDs.

[0030] In any of the above embodiments, the at least one syntax element may include a first flag and a second flag, the first flag indicating whether the motion vector difference of the current video block is jointly represented or decoded using signals, and the second flag indicating whether adaptive MVD pixel resolution is applied to the current video block.

[0031] In any of the above embodiments, in the video stream, the first flag is indicated by a signal before the second flag.

[0032] In any of the above embodiments, a single context can be used to signal the second flag in the video stream.

[0033] In any of the above embodiments, the context for signaling the second flag depends on the encoded information of the current video block and / or the adjacent video blocks of the current video block.

[0034] In any of the above embodiments, the encoded information may include at least one of the following: the value of the second flag, the reference frame index, or the MVD candidate index of the adjacent video blocks of the current video block.

[0035] In any of the above embodiments, in the video stream, the second flag is indicated by a signal preceding the first flag.

[0036] In any of the above embodiments, the video stream may include inter-frame prediction syntax elements indicating at least one of the following composite inter-frame prediction coding modes for the current video block: a first mode for composite reference inter-frame prediction with joint MVD encoded at adaptive MVD pixel resolution; a second mode for composite reference inter-frame prediction with joint MVD encoded at fixed MVD pixel resolution; a third mode for composite reference inter-frame prediction with independent MVD encoded at adaptive MVD pixel resolution; or a fourth mode for composite reference inter-frame prediction with independent MVD encoded at fixed MVD pixel resolution. Determining whether joint MVD encoding and adaptive MVD pixel resolution are applied to the current video block may include extracting and determining the value of the inter-frame prediction syntax element from the video stream.

[0037] In any of the above embodiments, the method may further include: when joint MVD encoding and adaptive MVD pixel resolution are applied to the current video block, optical flow refinement is always applied to the current video block if a predefined set of conditions is met.

[0038] In any of the above embodiments, the predefined condition set may include video block size constraints.

[0039] In any of the above embodiments, the method may further include: disallowing position-dependent composite prediction when joint MVD coding is applied to the current video block or when both joint MVD coding and adaptive MVD pixel resolution are applied to the current video block.

[0040] In any of the above embodiments, the location-related composite prediction may include a composite wedge-based prediction.

[0041] In any of the above embodiments, the method may further include: when joint MVD coding is applied to the current video block or when both joint MVD coding and adaptive MVD pixel resolution are applied to the current video block, interpolation filters other than regular, smoothing, or sharpening filters are not allowed.

[0042] In any of the above embodiments, when determining whether joint MVD coding or adaptive MVD pixel resolution is applied to the current video block, frame or sequence-level syntax elements may override lower-level signaling.

[0043] This disclosure also provides a video encoding or decoding device or apparatus, including circuitry configured to perform any of the above method embodiments.

[0044] The invention also provides a non-transient computer-readable medium for storing instructions that, when executed by a computer to perform video decoding and / or encoding of the machine, cause the computer to perform a method of video decoding and / or encoding. Attached Figure Description

[0045] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0046] Figure 1A A schematic diagram of an exemplary subset of intra-frame predicted direction patterns is shown;

[0047] Figure 1B A schematic diagram of an exemplary intra-frame prediction direction is shown.

[0048] Figure 2 A schematic diagram of the current block and its surrounding spatial merging candidates for motion vector prediction is shown in one example;

[0049] Figure 3 A simplified block diagram of a communication system (300) according to an example embodiment is shown;

[0050] Figure 4 A simplified block diagram of a communication system (400) according to an example embodiment is shown;

[0051] Figure 5 A simplified block diagram of a video decoder according to an example embodiment is shown in the schematic diagram;

[0052] Figure 6 A simplified block diagram of a video decoder according to an example embodiment is shown in the schematic diagram;

[0053] Figure 7 A block diagram of a video encoder according to another example embodiment is shown;

[0054] Figure 8 A block diagram of a video decoder according to another example embodiment is shown;

[0055] Figure 9 A scheme for dividing coded blocks according to an example embodiment of the present disclosure is shown;

[0056] Figure 10 Another scheme for dividing coded blocks according to an example embodiment of this disclosure is shown;

[0057] Figure 11 Another scheme for dividing coded blocks according to an example embodiment of this disclosure is shown;

[0058] Figure 12 An example of dividing a basic block into coded blocks according to a sample segmentation scheme is shown;

[0059] Figure 13 An example ternary partitioning scheme is shown;

[0060] Figure 14 An example quadtree / binary tree coding block partitioning scheme is shown;

[0061] Figure 15 A scheme for dividing a coded block into multiple transform blocks and the encoding order of the transform blocks, according to an example embodiment of the present disclosure, is shown.

[0062] Figure 16 Another scheme for dividing a coded block into multiple transform blocks and the encoding order of the transform blocks, according to an example embodiment of the present disclosure, is shown;

[0063] Figure 17 A scheme for dividing a coded block into multiple transform blocks according to an example embodiment of the present disclosure is shown;

[0064] Figure 18 A flowchart of a method according to an example embodiment of the present disclosure is shown;

[0065] Figure 19 A schematic diagram of a computer system according to an example embodiment of the present disclosure is shown. Detailed Implementation

[0066] Throughout the specification and claims, terms may have subtle meanings implied or implied in the context that go beyond their expressly stated meanings. The phrases “in one embodiment” or “in some embodiments” as used herein do not necessarily refer to the same embodiment, and the phrases “in another embodiment” or “in other embodiments” as used herein do not necessarily refer to different embodiments. Similarly, the phrases “in one implementation” or “in some implementations” as used herein do not necessarily refer to the same implementation, and the phrases “in another implementation” or “in other implementations” as used herein do not necessarily refer to different implementations. For example, the claimed subject matter includes combinations of all or some exemplary embodiments / implementations.

[0067] Generally, terms can be understood, at least in part, from their usage in context. For example, terms such as “and,” “or,” or “and / or” as used herein can include a variety of meanings, which can depend at least in part on the context in which these terms are used. Typically, “or,” when used to relate a list such as A, B, or C, is intended to mean A, B, and C (used here in an inclusive sense) and A, B, or C (used here in an exclusive sense). Furthermore, the terms “one or more” or “at least one,” as used herein, can be used, at least in part, to describe any feature, structure, or characteristic in a singular sense, or to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms such as “a,” “an,” or “the” can be understood to convey either a singular or a plural usage, at least in part on the context. Moreover, the terms “based on” or “determined by” can be understood to not necessarily convey an exclusive set of factors, but rather to allow for the presence of other factors that are not necessarily explicitly described, which also depends at least in part on the context. Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present disclosure is shown. The communication system (300) includes a plurality of terminal devices capable of communicating with each other, for example, via a network (350). For example, the communication system (300) includes pairs (310) and (320) of terminal devices interconnected via the network (350). Figure 12 In the example, the first terminal device (310) and (320) can perform unidirectional data transmission. For example, terminal device (310) can encode video data (e.g., a video image stream captured by terminal device (310)) for transmission over network (350) to another terminal device (320). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. Terminal device (320) can receive the encoded video data from network (350), decode the encoded video data to recover the video images, and display the video images based on the recovered video data. Unidirectional data transmission can be implemented in media service applications, etc.

[0068] In another example, the communication system (300) includes a pair of second terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which may be implemented, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) may encode video data (e.g., a video image stream captured by the terminal device) for transmission over a network (350) to the other terminal device (330) and (340). Each of the terminal devices (330) and (340) may also receive encoded video data transmitted by the other terminal device (330) and (340), and may decode the encoded video data to recover video images, and may display the video images on an accessible display device based on the recovered video data.

[0069] exist Figure 3 In the examples, terminal devices (310), (320), (330), and (340) may be implemented as servers, personal computers, and smartphones, but the applicability of the basic principles of this disclosure is not limited thereto. Embodiments of this disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, and / or similar devices. Network (350) refers to any number or type of network that transmits encoded video data between terminal devices (310), (320), (330), and (340), including, for example, wired (connected) and / or wireless communication networks. Communication network (350) may exchange data in circuit-switched, packet-switched, and / or other types of channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this discussion, unless explicitly explained herein, the architecture and topology of network (350) may be of little importance to the operation of this disclosure.

[0070] As an example of the application of the disclosed topic, Figure 4 The placement of a video encoder and video decoder in a streaming environment is illustrated. The disclosed subject matter is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0071] The video streaming system may include a video acquisition subsystem (413) that may include a video source (401), such as a digital camera, for creating uncompressed video pictures or image streams (402). In the example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. The video picture stream (402), depicted as a thick line to emphasize its high data volume, may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. Compared to the uncompressed video image stream (402), the encoded video data (404) (or encoded video bitstream (404)) depicted as thin lines to emphasize its lower data volume can be stored on a streaming server (405) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems, for example, Figure 4 Client subsystems (406) and (408) can access a streaming server (405) to retrieve copies (407) and (409) of encoded video data (404). Client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and produces an uncompressed output video picture stream (411) that can be displayed on a display (412) (e.g., a screen) or another presentation device (not depicted). The video decoder 410 may be configured to perform some or all of the various functions described in this disclosure. In some streaming systems, the encoded video data (404), encoded video data (407), and encoded video data (409) (e.g., video bitstreams) may be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The topics presented can be used in the context of VVC and other video coding standards.

[0072] It is understood that electronic devices (420) and (430) may include other components (not shown). For example, electronic device (420) may include a video decoder (not shown), and electronic device (430) may include a video encoder (not shown).

[0073] Figure 5A block diagram of a video decoder (510) according to any embodiment of the present disclosure is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may be used in place of... Figure 4 The example video decoder (410).

[0074] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one encoded video sequence may be decoded at a time, wherein the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with multiple video frames or images. Encoded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive encoded video data that can be forwarded to their respective processing circuitry (not depicted), as well as other data, such as encoded audio data and / or auxiliary data streams. The receiver (531) may separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (515) may be provided between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be located outside and separate from the video decoder (510) (not depicted). Still in other applications, a buffer memory (not depicted) may be located outside the video decoder (510) for purposes such as preventing network jitter, and another buffer memory (515) may be located inside the video decoder (510) for purposes such as handling playback timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be necessary, or it may be made smaller. For use on best-effort packet networks (e.g., the Internet), a buffer memory (515) of sufficient size may be required, and its size may be relatively large. Such a buffer memory can be implemented with adaptive sizing and can be implemented, at least partially, in an operating system or similar component (not shown) outside the video decoder (510).

[0075] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. These symbols may or may not be integral to the electronic device (530), but may be coupled to the electronic device (530). Figure 5 As shown in the diagram. Control information for one (or more) display devices may be in the form of a parameter set fragment (not depicted) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (520) may perform parsing / entropy decoding on the encoded video sequence received by the parser (520). The entropy coding of the encoded video sequence may be performed according to video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set of at least one subgroup of pixels in the subgroup from the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a subgroup. Subgroups may include Group of Pictures (GOP), pictures, tiles, stripes, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser (520) can also extract information from the encoded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc.

[0076] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0077] Depending on the type of encoded video frames or portions thereof (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (521) may involve multiple different processing or functional units. The units involved and the manner in which they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (520). For the sake of brevity, the flow of such subgroup control information between the parser (520) and the various processing or functional units described below is not depicted.

[0078] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical implementations operating under commercial constraints, many of these functional units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, the conceptual subdivision into functional units is adopted in the following disclosure.

[0079] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive quantization transform coefficients as symbols (521) from the parser (520) and control information, including information indicating which inverse transform type, block size, quantization factor / parameter, quantization scaling matrix, etc., are used as symbols (521). The scaler / inverse transform unit (551) may output a block containing sample values, which may be input into the aggregator (555).

[0080] In some cases, the output samples of the scaler / inverse transform (551) may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate blocks of the same size and shape as the blocks being reconstructed, using information from the reconstructed surrounding blocks and block information stored in the current picture buffer (558). For example, the current picture buffer (558) buffers partially reconstructed current images and / or fully reconstructed current images. In some implementations, the aggregator (555) may add the predictive information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) based on each sample.

[0081] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to inter-frame coded and potential motion-compensated blocks. In this case, the motion-compensated prediction unit (553) can access the reference image memory (557) to extract samples for inter-frame prediction. After motion compensation of the extracted samples according to the symbols (521) belonging to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of unit 551 may be referred to as residual samples or residual signals) to generate output sample information. The motion-compensated prediction unit (553) can obtain the predicted samples from the address in the reference image memory (557) under the control of a motion vector, and the motion vector is available to the motion-compensated prediction unit (553) in the form of symbols (521), which may have, for example, X and Y components (shift) and a reference image component (time). Motion compensation may also include interpolation of sample values ​​extracted from a reference image memory (557) when using subsampled precise motion vectors, and may also be associated with motion vector prediction mechanisms, etc.

[0082] The output samples of the aggregator (555) can be subjected to various loop filtering techniques in the loop filter unit (556). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and available to the loop filter unit (556) as symbols (521) from the parser (520). However, video compression techniques may also respond to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. Several types of loop filters may be included in various orders as part of the loop filter unit 556, which will be described in further detail below.

[0083] The output of the loop filter unit (556) can be a sample stream, which can be output to the rendering device (512) and stored in the reference image memory (557) for subsequent inter-frame image prediction.

[0084] Once fully reconstructed, certain encoded images can be used as reference images for future inter-frame prediction. For example, once the encoded image corresponding to the current image has been fully reconstructed and the encoded image (by, for example, the parser (520)) is identified as the reference image, the current image buffer (558) can become part of the reference image memory (557), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0085] The video decoder (510) can perform decoding operations according to a predetermined video compression technique, such as that used in the ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under that configuration file. For standard conformance, the complexity of the encoded video sequence may be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.

[0086] In some example embodiments, the receiver (531) may receive additional (redundant) data along with the encoded video. This additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant stripes, redundant pictures, forward error correction codes, etc.

[0087] Figure 6 A block diagram of a video encoder (603) according to an example embodiment of the present disclosure is shown. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., transmission circuitry). The video encoder (603) may be used in place of Figure 4 The example video encoder (403).

[0088] The video encoder (603) can obtain data from the video source (601) (not) Figure 6 In one example, the electronic device (620) receives video samples, the video source being capable of capturing video images to be encoded by a video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620).

[0089] A video source (601) can provide a sequence of source video samples encoded by a video encoder (603) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601YCrCB, RGB, XYZ, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device capable of storing previously prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures or images, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc., used. The relationship between pixels and samples will be readily understood by those skilled in the art. The following focuses on describing samples.

[0090] According to some example embodiments, the video encoder (603) can encode and compress images of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate constitutes a function of the controller (650). In some embodiments, as described below, the controller (650) can be functionally coupled to and control other functional units. For simplicity, coupling is not depicted in the figures. Parameters set by the controller (650) may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions related to the video encoder (603) optimized for a particular system design.

[0091] In some example embodiments, the video encoder (603) may be configured to operate within an encoding loop. As a simplified description, in this example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and one(or more) reference images) and a (local) decoder (633) embedded within the video encoder (603). Even if the embedded decoder 633 processes the video stream encoded by the source encoder 630 without entropy encoding, the decoder (633) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder would create them (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols in entropy encoding and the encoded video bitstream can be lossless). The reconstructed sample stream (sample data) is input to a reference image memory (634). Since the decoding of the symbol stream produces bit-accurate results independent of the decoder location (local or remote), the contents of the reference image memory (634) also correspond bit-accurately between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the encoder's prediction section are exactly the same as the sample values ​​that the decoder will "see" during the prediction phase. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is used to improve coding quality.

[0092] The operation of the “local” decoder (633) can be combined with, for example, the above-mentioned Figure 5 The operation of the video decoder (510) described in detail is the same as that of the "remote" decoder. However, a further brief reference is provided. Figure 5 When symbols are available and the entropy encoder (645) and parser (520) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding portion of the video decoder (510), including the buffer (515) and parser (520), may not be fully implemented in the encoder’s local decoder (633).

[0093] It can then be observed that any decoder technique, except for parsing / entropy decoding which may exist solely in the decoder, must also exist in the corresponding encoder in essentially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operation, which relates to the decoding portion of the encoder. Since encoder techniques are inverses of fully described decoder techniques, the description of encoder techniques can be simplified. A more detailed description is provided below, focusing only on certain areas or aspects.

[0094] During operation, in some example implementations, the source encoder (630) may perform motion-compensated predictive coding, referencing one or more previously encoded images from the video sequence designated as "reference images," which predictively encodes the input image. In this manner, the encoding engine (632) encodes the color channel differences (or residuals) between pixel blocks of the input image and pixel blocks of one (or more) reference images, which may be selected as prediction references for the input image. The term "residual" and its adjective form "residual" are used interchangeably.

[0095] The local video decoder (633) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (630). The operation of the encoding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded by the video decoder (633), Figure 6 When the source video sequence (not shown) is decoded, the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in a reference image cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0096] The predictor (635) can perform a prediction search against the encoding engine (632). That is, for a new image to be encoded, the predictor (635) can search in the reference image memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (635) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (635), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (634).

[0097] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0098] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, thereby transforming the symbols into an encoded video sequence.

[0099] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0100] The controller (650) can manage the operation of the video encoder (603). During encoding, the controller (650) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types:

[0101] An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Variations of I-pictures and their corresponding applications and characteristics are well understood by those skilled in the art.

[0102] A predictive picture (P-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block.

[0103] A bidirectional predictive picture (B-picture) can be a picture that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0104] Source images are typically spatially subdivided into multiple sample coding blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and coded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignment of the corresponding images applied to the block. For example, a block of an I image can be non-predictively coded, or it can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. A pixel block of a P image can be predictively coded with reference to a previously coded reference image via spatial prediction or temporal prediction. A block of a B image can be predictively coded with reference to one or two previously coded reference images via spatial prediction or temporal prediction. For other purposes, source images or intermediate images can be subdivided into other types of blocks. As described further in detail below, the partitioning of coding blocks and other types of blocks may or may not follow the same pattern.

[0105] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0106] In some example embodiments, the transmitter (640) may transmit additional data while transmitting encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and stripes, SEI messages, VUI parameter set fragments, etc.

[0107] The captured video can be presented as multiple source images (video images) in a time-series format. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes temporal or other correlations between images. For example, a specific image being encoded / decoded can be segmented into blocks; this specific image being encoded / decoded is called the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference image.

[0108] In some example embodiments, bidirectional prediction techniques can be used for inter-frame image prediction. According to this bidirectional prediction technique, two reference images are used, such as a first reference image and a second reference image that precede the current image in the video in decoding order (but may be past and future in display order, respectively). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be jointly predicted using a combination of the first and second reference blocks.

[0109] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0110] According to some example embodiments of this disclosure, predictions, such as inter-frame picture prediction and intra-frame picture prediction, are performed on a block-by-block basis. For example, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression, with CTUs in the pictures having the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU may include three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU may be recursively split into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU may be split into one 64×64 pixel CU, or four 32×32 pixel CUs. Each of one or more of the 32×32 blocks may be further subdivided into four 16×16 pixel CUs. In some example embodiments, each CU may be analyzed during encoding to determine the prediction type used for the CU in its respective prediction type, such as inter-frame prediction type or intra-frame prediction type. Depending on temporal and / or spatial predictability, a CU can be divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In embodiments, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. The division of the CU into PUs (or PBs for different color channels) can be performed using various spatial patterns. For example, a luma or chroma PB may comprise a matrix of sample values ​​(e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 samples, etc.

[0111] Figure 7 A diagram of a video encoder (703) according to another example embodiment of the present disclosure is shown. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values ​​within a current video image in a video image sequence, and to encode the processing block into an encoded image that is part of an encoded video sequence. The example video encoder (703) can be used instead of Figure 4 The example video encoder (403).

[0112] For example, a video encoder (703) receives a matrix of sample values ​​for a processing block, such as an 8×8 sample prediction block. The video encoder (703) then uses, for example, rate-distortion optimization (RDO) to determine whether to use intra-frame mode, inter-frame mode, or bidirectional prediction mode to optimally encode the processing block. When the processing block to be encoded is determined in intra-frame mode, the video encoder (703) can use intra-frame prediction techniques to encode the processing block into an encoded picture; and when the processing block to be encoded is determined in inter-frame mode or bidirectional prediction mode, the video encoder (703) can use inter-frame prediction or bidirectional prediction techniques to encode the processing block into an encoded picture, respectively. In some example embodiments, a merge mode can be used as a sub-mode for inter-frame picture prediction, wherein motion vectors are derived from one or more motion vector predictors without the aid of encoded motion vector components outside the predictor. In some other example embodiments, motion vector components applicable to the subject block may exist. Therefore, the video encoder (703) may include... Figure 7 Components not explicitly shown (such as the pattern determination module) are used to determine the prediction pattern of the processing block.

[0113] exist Figure 7 In the example, the video encoder (703) includes, for example, Figure 7 The example setup shows an inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together.

[0114] The inter-frame encoder (730) is configured to receive samples of the current block (e.g., the processing block), compare the block with one or more reference blocks in a reference image (e.g., to display blocks in previous and subsequent images in sequence), generate inter-frame prediction information (e.g., a description of redundancy information based on the inter-frame coding technique, motion vectors, merging mode information), and compute inter-frame prediction results (e.g., predicted blocks) based on the inter-frame prediction information using any suitable technique. In some examples, the reference image is used using embedded... Figure 6 The decoding unit 633 in the example encoder 620 (described in detail below, shown as) Figure 7 The residual decoder 728 decodes a decoded reference image based on encoded video information.

[0115] The intra encoder (722) is configured to receive samples of the current block (e.g., the processed block), compare the block with previously encoded blocks in the same image, generate quantization coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information based on one or more intra coding techniques). The intra encoder (722) can calculate intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same image.

[0116] A general controller (721) can be configured to determine general control data and control other components of the video encoder (703) based on that general control data. In an example, the general controller (721) determines the prediction mode of a block and provides control signals to a switch (726) based on that prediction mode. For example, when the prediction mode is an intra-frame mode, the general controller (721) controls the switch (726) to select an intra-frame mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select intra-frame prediction information and include that intra-frame prediction information in the bitstream; and when the prediction mode of a block is an inter-frame mode, the general controller (721) controls the switch (726) to select an inter-frame prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select inter-frame prediction information and include that inter-frame prediction information in the bitstream.

[0117] A residual calculator (723) can be configured to calculate the difference (residual data) between the received block and the prediction result of a block selected from the intra encoder (722) or the inter encoder (730). A residual encoder (724) can be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) can be configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to quantization to obtain quantized transform coefficients. In various example embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter-frame prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra-frame prediction information. The decoded blocks are processed appropriately to generate a decoded image, which can be buffered in a memory circuit (not shown) and used as a reference image.

[0118] An entropy encoder (725) can be configured to format a bitstream to produce encoded blocks. The entropy encoder (725) is configured to include various types of information in the bitstream. For example, the entropy encoder (725) can be configured to include general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. Residual information may not be present when encoding blocks in inter-frame mode or a merged sub-mode of bidirectional prediction mode.

[0119] Figure 8 A diagram of an example video decoder (810) according to another embodiment of the present disclosure is shown. The video decoder (810) is configured to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed image. In the example, the video decoder (810) can be used instead of Figure 4 The example video decoder (410).

[0120] exist Figure 8 In the example, the video decoder (810) includes, for example, Figure 8 The example setup shows an entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) coupled together.

[0121] The entropy decoder (871) can be configured to reconstruct certain symbols from an encoded picture, representing the syntax elements constituting the encoded picture. Such symbols may include, for example, the mode used to encode the block (e.g., intra-frame mode, inter-frame mode, bidirectional prediction mode, merged sub-mode, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information) that can be identified for prediction by the intra-frame decoder (872) or inter-frame decoder (880), residual information in the form of, for example, quantized transform coefficients, and so on. In the example, when the prediction mode is inter-frame or bidirectional prediction mode, inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is intra-frame prediction type, intra-frame prediction information is provided to the intra-frame decoder (872). The residual information may be inversely quantized and provided to the residual decoder (873).

[0122] The inter-frame decoder (880) can be configured to receive inter-frame prediction information and generate inter-frame prediction results based on the inter-frame prediction information.

[0123] The intra-frame decoder (872) can be configured to receive intra-frame prediction information and generate prediction results based on the intra-frame prediction information.

[0124] The residual decoder (873) can be configured to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) can also utilize certain control information (to include quantizer parameters (QP)) provided by the entropy decoder (871) (the data path is not depicted because this is only low-data-volume control information).

[0125] The reconstruction module (874) can be configured to combine the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module, depending on the specific situation) in the spatial domain to form a reconstructed block, which forms part of the reconstructed image as part of the reconstructed video. It can be noted that other suitable operations, such as deblocking, can also be performed to improve visual quality.

[0126] It is noteworthy that any suitable technology can be used to implement the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810). In some example embodiments, one or more integrated circuits may be used to implement the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810). In another embodiment, one or more processors executing software instructions may be used to implement the video encoder (403), video encoder (603), and video encoder (603), as well as the video decoder (410), video decoder (510), and video decoder (810).

[0127] Returning to block partitioning for encoding and decoding, partitioning can generally begin with a base block and can follow a predefined set of rules, a specific pattern, a partition tree, or any partitioning structure or scheme. Partitioning can be hierarchical and recursive. After partitioning or segmenting the base block according to any of the example partitioning processes described below, or other processes or combinations thereof, the final partitions or groups of coded blocks are obtained. Each of these partitions can be at one of various partitioning levels in the partitioning hierarchy and can have various shapes. Each partition can be called a coded block (CB). For the various example partitioning implementations further described below, each resulting CB can be of any allowed size and partitioning level. Because such partitions can form units for which some basic encoding / decoding decisions can be made, and for which encoding / decoding parameters can be optimized, determined, and represented as signals in the encoded video bitstream, such partitions are called coded blocks. The highest or deepest level in the final partition represents the depth of the coded block partitioning structure of the tree. A coded block can be a luma coded block or a chroma coded block. The CB tree structure for each color can be called a coded block tree (CBT).

[0128] The coding blocks for all color channels can be collectively referred to as coding units (CUs). The hierarchical structure of all color channels can be collectively referred to as coding tree units (CTUs). The partitioning patterns or structures of various color channels within a CTU can be the same or different.

[0129] In some implementations, the partition tree scheme or structure used for the luma and chroma channels does not need to be the same. In other words, the luma and chroma channels can have separate coding tree structures or patterns. Furthermore, whether the luma and chroma channels use the same or different coding partition tree structures, and the actual coding partition tree structure to be used, can depend on whether the encoded stripe is a P, B, or I stripe. For example, for an I stripe, the chroma and luma channels can have separate coding partition tree structures or coding partition tree structure patterns, while for P or B stripes, the luma and chroma channels can share the same coding partition tree scheme. When separate coding partition tree structures or patterns are applied, the luma channel can be partitioned into CBs using one coding partition tree structure, and the chroma channel can be partitioned into chroma CBs using another coding partition tree structure.

[0130] In some example implementations, a predetermined segmentation pattern can be applied to the base block. For example... Figure 9 As shown, an exemplary 4-way partitioning tree can begin at a first predefined level (e.g., a 64×64 block level or other size, as the base block size), and the base block can be hierarchically partitioned down to a predefined lowest level (e.g., a 4×4 level). For example, the base block can be subject to four predefined partitioning options or patterns indicated by 902, 904, 906, and 908, where the partition specified as R is allowed to be recursively partitioned, i.e., as... Figure 9 The same partitioning options indicated in the code can be repeated at lower scales up to the lowest level (e.g., a 4x4 level). In some implementations, additional restrictions may be applied. Figure 9 The segmentation scheme. In Figure 9 In this implementation, rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) are allowed, but they may not be allowed to be recursive, while square partitions are allowed to be recursive. If necessary, according to... Figure 9 The code tree is segmented and recursively generated into the final coded block group. The code tree depth can be further defined to indicate the segmentation depth from the root node or root block. For example, the code tree depth of the root node or root block (e.g., a 64×64 block) can be set to 0, and the depth of the root block can be determined based on... Figure 9 After being further divided once, the coding tree depth increases by 1. For the above scheme, the maximum or deepest level from the 64x64 base block to the 4x4 minimum partition will be 4 (starting from level 0). This partitioning scheme can be applied to one or more color channels. Each color channel can be divided according to... Figure 9 The segmentation scheme is performed independently (e.g., for each color channel at each layer level, the segmentation pattern or option in a predefined pattern can be determined independently). Optionally, two or more color channels can share... Figure 9 The same hierarchical pattern tree (for example, the same segmentation pattern or option in a predefined pattern can be selected for two or more color channels at each hierarchical level).

[0131] Figure 10 Another example is shown that enables recursive partitioning to form a partition tree. For example... Figure 10 As shown, an exemplary 10-way partitioning structure or pattern can be predefined. The root block can start at a predefined level (e.g., a base block at a 128×128 level or a 64×64 level). Figure 10 The example partitioning structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. Figure 10 The partition type with three sub-partitions indicated by 1002, 1004, 1006, and 1008 in the second line can be called a "T-type" partition. The "T-type" partitions 1002, 1004, 1006, and 1008 can be referred to as left T-type, top T-type, right T-type, and bottom T-type. In some exemplary embodiments, further subdivision is not permitted. Figure 10 Any rectangular partition within the rectangular partition. The coding tree depth can be further defined to indicate the segmentation depth from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., a 128×128 block) can be set to 0, and the depth of the root block can be determined according to... Figure 10After being further segmented once, the coding tree depth increases by 1. In some implementations, only full-square partitions in 1010 can be recursively segmented according to... Figure 10 The next level of the partitioning tree for the pattern. In other words, for the square partitions within the T-patterns 1002, 1004, 1006, and 1008, recursive partitioning is not allowed. If necessary, according to Figure 10 The process involves partitioning the blocks and recursively generating the final coded block groups. This partitioning scheme can be applied to one or more color channels. In some implementations, more flexibility can be added when using partitions at levels lower than 8×8. For example, 2×2 chroma inter-frame prediction can be used in certain situations.

[0132] In some other example implementations of coded block segmentation, a quadtree structure can be used to segment a base block or intermediate block into quadtree partitions. This quadtree segmentation can be applied hierarchically and recursively to any square partition. Whether to further quadtree segment the base block, intermediate block, or partition can be adjusted based on various local features of the base block or intermediate block / partition. Quadtree segmentation at image boundaries can be further adjusted. For example, implicit quadtree segmentation can be performed at image boundaries such that the block will remain quadtree segmented until its size fits the image boundaries.

[0133] In some other example implementations, hierarchical binary partitioning from the base block can be used. For such a scheme, the base block or intermediate block can be partitioned into two partitions. The binary partitioning can be horizontal or vertical. For example, horizontal binary partitioning can divide the base block or intermediate block into equal right and left partitions. Similarly, vertical binary partitioning can divide the base block or intermediate block into equal upper and lower partitions. This binary partitioning can be hierarchical and recursive. A decision can be made at each point in the base block or intermediate block regarding whether to continue the binary partitioning scheme, and if so, whether to use horizontal or vertical binary partitioning. In some implementations, further partitioning can stop at a predefined minimum partition size (one or two dimensions). Alternatively, further partitioning can stop once a predefined partition level or depth from the base block is reached. In some implementations, the aspect ratio of the partitions can be limited. For example, the aspect ratio of the partitions can be no less than 1:4 (or greater than 4:1). Therefore, a vertical strip partition with a vertical-to-horizontal aspect ratio of 4:1 can be further divided vertically into an upper partition and a lower partition with a vertical-to-horizontal aspect ratio of 2:1.

[0134] In some other examples, such as Figure 13 As shown, the ternary partitioning scheme can be used to partition base blocks or any intermediate blocks. Ternary patterns can be implemented vertically, such as... Figure 13As shown in 1302, or horizontally implemented, as... Figure 13 As shown in 1304. Figure 13 The example segmentation ratio in the example is although Figure 13 The example partition ratio is shown vertically or horizontally as 1:2:1, but other ratios can be predefined. In some implementations, two or more different ratios can be predefined. Since ternary tree partitioning can capture objects located at the center of a block within a contiguous partition, while quadtrees and binary trees always partition along the block center, thus dividing objects into different partitions, this ternary partitioning scheme can be used to complement quadtree or binary partitioning structures. In some implementations, the width and height of the partitions in the example ternary tree are always powers of 2 to avoid additional transformations.

[0135] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the quadtree and binary partitioning schemes described above can be combined to partition the base block into a quadtree-binary-tree (QTBT) structure. In such a scheme, the base block or intermediate block / partition can be a quadtree partition or a binary partition, if specified, conforming to a predefined set of conditions. Specific examples include... Figure 14 As shown, in Figure 14 In the examples, such as 1402, 1404, 1406, and 1408, the base block is a first quadtree divided into four partitions. Subsequently, each partition in the resulting partitions is either divided into four further partitions by the quadtree (e.g., 1408), or binary-coded into two further partitions at the next level (horizontally or vertically, e.g., 1402 or 1406, both symmetrical), or is non-partitioned (e.g., 1404). For square partitions, it is possible to recursively enable binary or quadtree partitioning, as shown in the overall example partitioning pattern in 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree partitions and dashed lines represent binary tree partitions. Flags can be used for each binary partition node (non-leaf binary partition) to indicate whether the binary partition is horizontal or vertical. For example, as shown in 1420, consistent with the partitioning structure of 1410, the flag "0" can indicate a horizontal binary partition, and the flag "1" can indicate a vertical binary partition. Since quadtree partitioning always divides blocks or partitions horizontally and vertically to produce four sub-blocks / partitions of equal size, it is not necessary to indicate the partitioning type for quadtree partitioning. In some implementations, the flag "1" can indicate a horizontal binary partition, and the flag "0" can indicate a vertical binary partition.

[0136] In some example implementations of QTBT, quadtrees and binary splitting rule sets can be represented by the following predefined parameters and their associated corresponding functions:

[0137] -CTU size: The size of the root node (base block) of the quadtree.

[0138] -MinQTSize: The minimum allowed size of a quadtree leaf node.

[0139] -MaxBTSize: The maximum allowed size of the root node of the binary tree.

[0140] -MaxBTDepth: Maximum allowed binary tree depth

[0141] -MinBTSize: The minimum allowed size of a binary leaf node.

[0142] In some example implementations of the QTBT segmentation structure, the CTU size can be set (when considering and using example chroma subsamples) to have 128×128 luminance samples with two corresponding 64×64 chroma sample blocks, the MinQTSize can be set to 16×16, the MaxBTSize can be set to 64×64, and the MinBTSize (for width and height) can be set to 4×4, and the MaxBTDepth can be set to 4. Quadtree segmentation can be applied to the CTU first to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from their minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTUsize). If a node is 128×128, it will not be segmented by the binary tree first because its size exceeds the MaxBTSize (i.e., 64×64). Otherwise, nodes not exceeding the MaxBTSize can be segmented by the binary tree. Figure 12 In the example, the base block is 128×128. According to the predefined rule set, the base block can only be quadtree-partitioned. The partitioning depth of the base block is 0. Each of the four resulting partitions is 64x64, not exceeding MaxBTSize, and can be further partitioned into quadtrees or binary trees at level 1. Continue the process. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning can be disregarded. When the width of a binary tree node equals MinBTSize (i.e., 4), further horizontal partitioning can be disregarded. Similarly, when the height of a binary tree node equals MinBTSize, further vertical partitioning is disregarded.

[0143] In some example implementations, the QTBT scheme described above can be configured to support the flexibility of having the same QTBT structure for luma and chroma, or having separate QTBT structures. For example, for P and B stripes, the luma and chroma CTBs in a CTU can have the same QTBT structure. However, for I stripes, the luma CTB can be segmented into CBs using a QTBT structure, and the chroma CTB can be segmented into chroma CBs using another QTBT structure. This means that CUs can be used to refer to different color channels in an I stripe; for example, an I stripe may include a block of coded luma components or a block of coded two chroma components, and a CU in a P or B stripe may include a block of coded all three color components.

[0144] In some other implementations, the QTBT scheme can be supplemented by the aforementioned ternary branch scheme. This implementation can be called a multi-type-tree (MTT) structure. For example, in addition to binary branching of nodes, other options can be selected... Figure 13 One of the three-way splitting patterns. In some implementations, only square nodes can be split into three parts. Additional flags can be used to indicate whether the split is horizontal or vertical.

[0145] The design of two- or multi-level trees, such as QTBT implementations and QTBT implementations supplemented by ternary partitioning, can be primarily driven by the reduction of complexity. Theoretically, the complexity of traversing a tree is T1. D Here, T represents the number of partition types, and D is the depth of the tree. A trade-off can be achieved by using multiple types (T) to reduce the depth (D).

[0146] In some implementations, the CB can be further segmented. For example, to perform intra-frame or inter-frame prediction during the encoding and decoding process, the CB can be further segmented into multiple prediction blocks. In other words, the CB can be further divided into different sub-partitions, in which separate prediction decisions / configurations can be made. In parallel, to describe the level at which the transform or inverse transform of the video data is performed, the CB can be further segmented into multiple transform blocks (TBs). The schemes for segmenting the CB into PBs and TBs can be the same or different. For example, each segmentation scheme can be performed using its own process based on, for example, various features of the video data. In some example implementations, the PB and TB segmentation schemes can be independent. In some other example implementations, the PB and TB segmentation schemes and boundaries can be related. In some implementations, for example, TBs can be segmented after PB segmentation, and specifically, each PB is determined after segmenting the coded block and can then be further segmented into one or more TBs. For example, in some implementations, a PB can be segmented into one, two, four, or other numbers of TBs.

[0147] In some implementations, the luma and chroma channels can be treated differently to segment a base block into coding blocks and further into prediction blocks and / or transform blocks. For example, in some implementations, for the luma channel, segmenting a coding block into prediction blocks and / or transform blocks may be permitted, while for the chroma channel, segmenting a coding block into prediction blocks and / or transform blocks may not be permitted. In such implementations, transform and / or prediction of the luma block can therefore be performed only at the coding block level. As another example, the minimum transform block size for the luma channel and one(s) of the chroma channels may be different; for example, the coding block for the luma channel may be segmented into transform and / or prediction blocks smaller than those for the chroma channel. As yet another example, the maximum depth to which a coding block is segmented into transform and / or prediction blocks may differ between the luma and chroma channels; for example, the coding block for the luma channel may be segmented into transform and / or prediction blocks deeper than those for one(s) of the chroma channels. For a specific example, a luma-coded block can be divided into multiple transform blocks of various sizes, which can be represented by recursive partitioning and can have transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, and transform block sizes ranging from 4×4 to 64×64. However, for a chroma block, only the largest possible transform block is allowed to be specified for the luma block.

[0148] In some example implementations, for the segmentation of a coded block into a PB, the depth, shape, and / or other characteristics of the PB segmentation may depend on whether the PB is intra-coded or inter-coded.

[0149] Segmenting a coded block (or prediction block) into transform blocks can be implemented in various example schemes, including but not limited to recursive or non-recursive quadtree segmentation and predetermined pattern segmentation, with additional consideration given to transform blocks at the boundaries of the coded or prediction blocks. Typically, the resulting transform blocks can be at different segmentation levels, can have different sizes, and do not need to be square (e.g., they can be rectangles of various sizes and aspect ratios). The following combines... Figure 15 , 16 17. Further examples are described in more detail.

[0150] However, in some other implementations, the CB obtained through any of the above-described segmentation schemes can be used as the basic or minimum coding block for prediction and / or transform. In other words, no further segmentation is performed for the purpose of performing inter-frame prediction / intra-frame prediction and / or transform. For example, the CB obtained from the above-described QTBT scheme can be directly used as the unit for performing prediction. Specifically, such a QTBT structure eliminates the concept of multiple partition types, that is, it eliminates the separation of CU, PU, ​​and TU, and supports greater flexibility in the CU / CB partition shape as described above. In this QTBT block structure, the CU / CB can be square or rectangular in shape. The leaf nodes of this QTBT are used as units for prediction and transform processing without any further segmentation. This means that in this exemplary QTBT coding block structure, the CU, PU, ​​and TU have the same block size.

[0151] The various CB partitioning schemes described above, and the further partitioning of CB into PB and / or TB (including no PB / TB partitioning), can be combined in any way. The following specific implementations are provided as non-limiting examples.

[0152] The following describes a specific example implementation of the coding block and transform block segmentation. In such an example implementation, the recursive quadtree segmentation described above or (e.g.) can be used. Figure 9 and Figure 10The predefined segmentation patterns (those in the code) divide the basic block into coded blocks. At each level, whether further quadtree segmentation of a particular partition should continue can be determined by local video data features. The resulting CBs can be at various quadtree segmentation levels and have various sizes. The decision on whether to use inter-frame picture (temporal) or intra-frame picture (spatial) prediction to encode picture regions can be made at the CB level (or CU level, for all three color channels). Each CB can be further segmented into one, two, four, or other numbers of PBs according to predefined PB segmentation types. Within a PB, the same prediction process can be applied, and relevant information can be sent to the decoder based on the PB. After obtaining the residual block by applying the prediction process based on the PB segmentation type, the CB can be segmented into TBs according to another quadtree structure similar to the CB. In this particular implementation, the CB or TB can be, but is not limited to, squares. Furthermore, in this particular example, for inter-frame prediction, the PB can be a square or a rectangle, and for intra-frame prediction, the PB can be only a square. The coded block can be segmented into, for example, four square TBs. Each TB can be further recursively partitioned (using quadtrees) into smaller TBs, called a Residual Quadtree (RQT).

[0153] Another exemplary implementation of dividing the base block into CB, PB, and / or TB is further described below. For example, without using... Figure 9 or Figure 10 Instead of the multiple partitioning unit types shown, a quadtree with nested multi-type trees (using binary and ternary partitioning structures, such as QTBTs as described above or QTBTs with ternary partitioning) can be used. The separation of CB, PB, and TB can be omitted (i.e., partitioning CB into PB and / or TB, and partitioning PB into TB) unless the size of the CB is too large for the maximum transform length, in which case further partitioning of the CB may be necessary. This example partitioning scheme can be designed to support greater flexibility in the shape of CB partitions so that both prediction and transform can be performed at the CB level without further partitioning. In this coding tree structure, the CB can be square or rectangular. Specifically, the coding tree block (CTB) can first be partitioned using a quadtree structure. Then, the leaf nodes of the quadtree can be further partitioned using nested multi-type tree structures. Examples of nested multi-type tree structures using binary or ternary partitioning are shown below. Figure 11 shown. Specifically, Figure 11The example multi-type tree structure includes four partition types, referred to as vertical binary partition (SPLIT_BT_VER) (1102), horizontal binary partition (SPLIT_BT_HOR) (1104), vertical ternary partition (SPLIT_TT_VER) (1106), and horizontal ternary partition (SPLIT_TT_HOR) (1108). The CB then corresponds to the leaf of the multi-type tree. In this example implementation, unless the CB is too large for the maximum transform length, the partition is used for prediction and transform processing without any further partitioning. This means that, in most cases, in a quadtree with a nested multi-type tree coding block structure, the CB, PB, and TB have the same block size. An anomaly occurs when the maximum supported transform length is less than the width or height of the color component of the CB. In some implementations, in addition to binary or ternary partitions, Figure 11 Nested patterns can also include quadtree partitioning.

[0154] Figure 12 A concrete example of a quadtree with a nested, multi-type tree-coded block structure featuring block partitioning options (including quadtree, binary, and ternary partitioning) for a base block is shown. More detailed... Figure 12 The diagram shows that base block 1200 is divided into four square partitions 1202, 1204, 1206, and 1208 by a quadtree. Further usage is described for each quadtree partition. Figure 11 The determination of multiple tree types and quadtrees used for further segmentation. Figure 12 In the example, partition 1204 is not further divided. Partitions 1202 and 1208 are each divided using another quadtree. For partition 1202, the second-level quadtree is used to divide the upper left, upper right, lower left, and lower right partitions respectively. Figure 11 Third-level partitioning of a quadtree, horizontal binary partitioning 1104, no partitioning and Figure 11 The horizontal ternary partition is 1108. Partition 1208 uses another quadtree for partitioning. The upper-left, upper-right, lower-left, and lower-right partitions of the second-level quadtree are respectively... Figure 11 The vertical ternary partition 1106, the third-order partition, no partition, no partition and Figure 11 Horizontal binary partition 1104. Based on respectively Figure 11 The horizontal binary partition 1104 and the horizontal ternary partition 1108 further divide the third-level upper-left partition of 1208 into two sub-partitions. Following... Figure 11 After the vertical binary partition 1102, partition 1206 is divided into two partitions using the second-level partitioning mode. These two partitions are based on... Figure 11 The horizontal ternary partition 1108 and the vertical binary partition 1102 are further divided into a third level. According to... Figure 11The horizontal binary partition 1104, the fourth level partition is further applied to one of the two partitions.

[0155] For the specific example above, the maximum luminance transformation size can be 64×64, and the supported maximum chrominance transformation size can be different from the luminance, for example, 32×32. Even above Figure 12 Example CB in the example even Figure 12 The example CB above is usually not further divided into smaller PB and / or TB. When the width or height of the luma-coded block or chroma-coded block is greater than the maximum transform width or height, the luma-coded block or chroma-coded block can be automatically divided in the horizontal and / or vertical directions to conform to the transform size limit in that direction.

[0156] As described above, in a specific example of segmenting a base block into CBs, the coding tree scheme can support the ability for luma and chroma to have separate block tree structures. For example, for P and B slices, the luma and chroma CTBs in a CTU can have the same coding tree structure. For example, for I slices, luma and chroma can have separate coding block tree structures. When separate block tree structures are applied, the luma CTB is segmented into luma CBs using one coding tree structure, and the chroma CTB is segmented into chroma CBs using another coding tree structure. This means that a CU in an I slice can include coding blocks for the luma component or coding blocks for both chroma components, and a CU in a P or B slice always includes coding blocks for all three color components, unless the video is monochrome.

[0157] When a coded block is further divided into multiple transform blocks, these transform blocks can be ordered in the bitstream in various orders or scanning methods. The following describes in further detail exemplary implementations of dividing a coded block or prediction block into transform blocks and the encoding order of these transform blocks. In some exemplary implementations, as described above, transform segmentation can support transform blocks of multiple shapes, such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, where the transform block size ranges from, for example, 4×4 to 64×64. In some implementations, if the coded block is less than or equal to 64×64, then transform block segmentation can be applied only to the luma component, so that for the chroma block, the transform block size is the same as the coded block size. Otherwise, if the width or height of the coded block is greater than 64, then both the luma and chroma coded blocks can be implicitly segmented into multiples of min(W, 64)×min(H, 64) and min(W, 32)×min(H, 32) transform blocks, respectively.

[0158] In some example implementations of transform block partitioning, for intra-frame and inter-frame coded blocks, the coded blocks can be further partitioned into multiple transform blocks, with a partitioning depth of up to a predetermined number of levels (e.g., 2 levels). The transform block partitioning depth and size can be correlated. For some example implementations, the mapping from the transform size at the current depth to the transform size at the next depth is shown in Table 1 below.

[0159] Table 1. Settings for changing the segment size

[0160]

[0161] Based on the example mapping in Table 1, for a 1:1 square block, the next-level transform segmentation can create four 1:1 square sub-transform blocks. For example, the transform segmentation can stop at 4x4. Thus, the transform size at the current depth of 4×4 corresponds to the same size 4×4 at the next depth. In the examples in Table 1, for a 1:2 / 2:1 non-square block, the next-level transform segmentation can create two 1:1 square sub-transform blocks, while for a 1:4 / 4:1 non-square block, the next-level transform segmentation can create two 1:2 / 2:1 sub-transform blocks.

[0162] In some example implementations, additional constraints may be applied relative to the transform block segmentation for the luminance component of an intra-coded block. For instance, for each level of transform segmentation, all sub-transform blocks may be constrained to have equal sizes. For example, for a 32x16 coded block, a level 1 transform segmentation creates two 16x16 sub-transform blocks, and a level 2 transform segmentation creates eight 8x8 sub-transform blocks. In other words, the second-level segmentation must be applied to all first-level sub-blocks to maintain equal transform unit sizes. Figure 15 Examples of transform block segmentation based on intra-frame coded square blocks from Table 1 are shown, along with the encoding order indicated by arrows. Specifically, 1502 shows a square coded block. 1504 shows a first-level segmentation, dividing the block into four equal-sized transform blocks according to Table 1, with the encoding order indicated by arrows. 1506 shows a second-level segmentation, dividing all first-level equal-sized blocks into 16 equal-sized transform blocks according to Table 1, with the encoding order indicated by arrows.

[0163] In some example implementations, the aforementioned intra-frame coding restrictions may not apply to the luminance components of inter-frame coded blocks. For instance, after the first-level transform segmentation, any sub-transform block within a sub-transform block can be further independently segmented at a level higher than one. Therefore, the resulting transform blocks may or may not have the same size. Figure 16 An example of segmenting inter-frame coded blocks into transform blocks and their coding order is shown. Figure 16In the example, the inter-frame coded block 1602 is divided into two levels of transform blocks according to Table 1. At the first level, the inter-frame coded block is divided into four transform blocks of equal size. Then, as shown in 1604, only one (not all) of the four transform blocks is further divided into four sub-transform blocks, resulting in a total of seven transform blocks with two different sizes. The example coding order of these seven transform blocks is... Figure 16 The arrow in 1604 is shown.

[0164] In some example implementations, additional constraints may be applied to the transform block for one (or more) chroma components. For example, the transform block size may be the same as the coding block size for one (or more) chroma components, but not smaller than a predefined size, such as 8×8.

[0165] In some other example implementations, for a coding block with a width (W) or height (H) greater than 64, the luma and chroma coding blocks can be implicitly divided into transform units that are multiples of min(W, 64) × min(H, 64) and min(W, 32) × min(H, 32), respectively. Here, in this disclosure, "min(a, b)" can return the smaller value between a and b.

[0166] Figure 17 Another alternative example scheme for segmenting a coded block or prediction block into transform blocks is further illustrated. For example... Figure 17 As shown, a predefined set of segmentation types can be applied to the coded block based on its transformation type, without using recursive transformation segmentation. Figure 17 In the specific example shown, one of the six example segmentation types can be applied to segment the coded block into various numbers of transform blocks. This scheme for generating transform block segments can be applied to either the coded block or the prediction block.

[0167] More in detail, Figure 17 The segmentation scheme provides up to six example segmentation types for any given transform type (transformation type refers to, for example, the type of the main transform, such as ADST, etc.). In this scheme, a transform segmentation type can be assigned to each coding block or prediction block based on factors such as rate-distortion cost. In the examples, the transform segmentation type assigned to a coding block or prediction block can be determined based on the transform type of the coding block or prediction block. A specific transform segmentation type can correspond to a transform block segmentation size and mode, such as... Figure 17 The six transform partition types are shown. The correspondence between various transform types and transform partition types can be predefined. In the examples shown below, uppercase labels indicate the transform partition type that can be assigned to a coding block or prediction block based on rate-distortion cost:

[0168] • PARTITION_NONE: Allocates a transform size equal to the block size.

[0169] • PARTITION_SPLIT: Allocates a transform size that is half the width and half the height of the block size.

[0170] • PARTITION_HORZ: Allocates a transform size with the same width as the block size and half the height of the block size.

[0171] • PARTITION_VERT: Allocates a transform size that is half the block size in width and the same height as the block size.

[0172] • PARTITION_HORZ4: Allocate a transform size with the same width as the block size and a height that is 1 / 4 of the block size.

[0173] • PARTITION_VERT4: Allocates a transform size of 1 / 4 the block size in width and the same height as the block size.

[0174] In the example above, such as Figure 17 The transformation segmentation types shown all include a uniform transform size for the segmented transform blocks. This is merely an example and not a limitation. In some other implementations, mixed transform block sizes may be used for the transform blocks of segments within a specific segmentation type (or pattern).

[0175] The PB (or CB, also called PB when it is not further divided into prediction blocks) obtained from any of the above partitioning schemes can then become a single block for encoding via intra-frame or inter-frame prediction. For inter-frame prediction of the current PB, the residual between the current block and the prediction block can be generated, encoded, and included in the encoded bitstream.

[0176] Inter-frame prediction can be implemented, for example, in single-reference mode or composite reference mode. In some implementations, a skip flag may be included first in the bitstream of the current block (or at a higher level) to indicate whether the current block is inter-coded and not skipped. If the current block is inter-coded, another flag may also be included in the bitstream as a signal to indicate whether single-reference mode or composite reference mode is used for the prediction of the current block. For single-reference mode, a single reference block can be used to generate the prediction block for the current block. For composite reference mode, two or more reference blocks can be used to generate the prediction block by, for example, a weighted average. Composite reference mode may be referred to as more than one reference mode, two reference modes, or multiple reference modes. One or more reference blocks can be identified using one or more reference frame indices and additionally using one or more corresponding motion vectors, which indicate one or more offsets in position (e.g., in horizontal and vertical pixels) between the reference block (or more) and the current block. For example, the inter-frame prediction block for the current block can be generated from a single reference block identified by a motion vector in a reference frame as a prediction block in single-reference mode, while for composite reference mode, the prediction block can be generated by a weighted average of two reference blocks in two reference frames indicated by two reference frame indices and two corresponding motion vectors. One (or more) motion vectors can be encoded and included in the bitstream in various ways.

[0177] In some implementations, the encoding or decoding system may maintain a Decoded Picture Buffer (DPB). Some images / pictures may be maintained in the DPB awaiting display (in the decoding system), and some images / pictures in the DPB may be used as reference frames for inter-frame prediction (in the decoding or encoding system). In some implementations, the reference frames in the DPB may be marked as short-term or long-term references for the current image being encoded or decoded. For example, a short-term reference frame may include a frame for inter-frame prediction of blocks in the current frame or in subsequent video frames that are closest to the current frame in the decoding sequence (e.g., 2). A long-term reference frame may include a frame in the DPB that can be used to predict image blocks in frames that are more than a predetermined number of frames away from the current frame in the decoding sequence. This tagging information for short-term and long-term reference frames may be referred to as a Reference Picture Set (RPS) and may be added to the header of each frame in the encoded bitstream. Each frame in the encoded video bitstream may be identified by a Picture Order Counter (POC), which is numbered in an absolute manner according to playback order or associated with a group of pictures, for example, starting from an I-frame.

[0178] In some example implementations, one or more reference picture lists containing identifiers of short-term and long-term reference frames for inter-frame prediction can be formed based on information in the RPS. For example, a single picture reference list, denoted as L0 reference (or reference list 0), can be formed for unidirectional inter-frame prediction, while two picture reference lists can be formed for bidirectional inter-frame prediction, denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 and L1 lists can be ordered in various predetermined ways. The lengths of the L0 and L1 lists can be signaled in the video bitstream. Unidirectional inter-frame prediction can be a single-reference mode, or a composite reference mode when multiple references used to generate a prediction block by weighted averaging in a composite prediction mode are on the same side of the block to be predicted. Since bidirectional inter-frame prediction involves at least two reference blocks, bidirectional inter-frame prediction can be only a composite mode.

[0179] In some implementations, a merge mode (MM) can be implemented for inter-frame prediction. Typically, for a merge mode, for the current PB, one or more motion vectors in a single-reference prediction or a composite-reference prediction can be derived from one (or more) other motion vectors, rather than being independently computed and represented by signals. For example, in a coding system, one (or more) current motion vectors of the current PB can be represented by one (or more) differences between one (or more) current motion vectors and one or more other already encoded motion vectors (called reference motion vectors). This one (or more) difference in one (or more) motion vectors, rather than the entirety of one (or more) current motion vectors, can be encoded and included in the bitstream and can be linked to one (or more) reference motion vectors. Accordingly, in a decoding system, one (or more) motion vectors corresponding to the current PB can be derived based on one (or more) decoded motion vector differences and one (or more) decoded reference motion vectors linked to them. As a specific form of inter-frame prediction in general merge mode (MM), this inter-frame prediction based on one (or more) motion vector differences can be called merge mode with motion vector differences (MMVD). Therefore, a typical MM or a specific MMVD can be implemented to leverage the correlation between motion vectors associated with different PBs to improve coding efficiency. For example, adjacent PBs may have similar motion vectors, so the MVD can be small and can be encoded efficiently. As another example, for spatially similar co-occurring / co-occurring blocks, motion vectors can be temporally (between frames).

[0180] In some example implementations, the MM flag may be included in the bitstream during the encoding process to indicate whether the current PB is in merge mode. Alternatively or additionally, an MMVD flag may be included during the encoding process and signaled in the bitstream to indicate whether the current PB is in MMVD mode. MM and / or MMVD flags or indicators may be provided at the PB level, CB level, CU level, CTB level, CTU level, stripe level, picture level, etc. For a particular example, the current CU may include both the MM flag and the MMVD flag, and the MMVD flag may be signaled immediately after the skip flag and the MM flag to specify whether MMVD mode is used for the current CU.

[0181] In some example implementations of MMVD, a list of reference motion vectors (RMVs) or MV predictor candidates for motion vector prediction can be formed for the block being predicted. The RMV candidate list can contain a predetermined number (e.g., 2) of MV predictor candidate blocks whose motion vectors can be used to predict the current motion vector. RMV candidate blocks can include blocks selected from adjacent blocks and / or temporal blocks within the same frame (e.g., identical co-occurring blocks in the previous or next frame of the current frame). These options represent blocks with spatial or temporal locations relative to the current block, which may have similar or identical motion vectors to the current block. The size of the MV predictor candidate list can be predetermined. For example, the list can contain two or more candidates. For a candidate block to be included in the RMV candidate list, it must, for example, have the same reference frame (or multiple frames) as the current block (e.g., boundary checks are required when the current block is near a frame edge), and must have been encoded during encoding and / or decoded during decoding. In some implementations, if the merge candidate list is available and the above conditions are met, the merge candidate list can be first filled with spatially adjacent blocks (scanned in a specific predefined order), and then time blocks can be filled if space is still available in the list. For example, adjacent RMV candidate blocks can be selected from the left and top blocks of the current block. The RMV predictor candidate list can be dynamically formed as a dynamic reference list (DRL) at different levels (sequence, picture, frame, strip, superblock, etc.). It can be sent in the bitstream by signaling.

[0182] In some implementations, actual MV predictor candidates can be signaled as reference motion vectors for predicting the motion vectors of the current block. When the RMV candidate list contains two candidates, a one-bit flag called the merge candidate flag can be used to indicate the selection of a reference merge candidate. For the current block predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor can be associated with a reference motion vector from the merge candidate list. The encoder can determine which RMV candidate more closely predicts the current coded block and send that selection as an index signal to the DRL.

[0183] In some example implementations of MMVD, after selecting RMV candidates and using them as the base motion vector predictor for the motion vector to be predicted, the motion vector difference (MVD, or delta MV, representing the difference between the motion vector to be predicted and the reference candidate motion vector) can be calculated in the encoding system. This MVD may include information representing the magnitude and direction of the MV difference, both of which can be signaled in the bitstream. The motion difference magnitude and direction can be signaled in various ways.

[0184] In some example implementations of MMVD, a distance index can be used to specify the magnitude information of the motion vector difference and indicate one of a predefined set of offsets representing a predefined motion vector difference relative to a starting point (reference motion vector). The MV offset, based on the index sent by the signal, can then be added to the horizontal or vertical component of the starting (reference) motion vector. Whether the horizontal or vertical component of the reference motion vector should be offset can be determined using the direction information of the MVD. Table 2 specifies examples of predefined relationships between distance indices and predefined offsets.

[0185] Table 2 shows an example relationship between distance index and predefined MV offset.

[0186]

[0187] In some example implementations of MMVD, the direction index can also be signaled and used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction can be restricted to either the horizontal or vertical direction. Examples of 2-bit direction indices are shown in Table 3. In the examples in Table 3, the interpretation of the MVD can vary depending on the information of the start / reference MV. For example, when the start / reference MV corresponds to a single prediction block or a dual prediction block, where both reference frame lists point to the same side of the current image (i.e., the POC of both reference images is greater than or less than the POC of the current image), the symbols in Table 3 can specify the sign (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a double prediction block, where the two reference images are located on different sides of the current image (i.e., the POC of one reference image is greater than the POC of the current image, and the POC of the other reference image is less than the POC of the current image), and the difference between the reference POC in image reference list 0 and the current frame is greater than the difference between the reference POC in image reference list 1 and the current frame, the symbols in Table 3 can specify the sign of the MV offset added to the reference MV corresponding to the reference image in image reference list 0, and the sign of the offset of the MV corresponding to the reference image in image reference list 1 can have the opposite value (opposite sign of the offset). Otherwise, if the difference between the reference POC in image reference list 1 and the current frame is greater than the difference between the reference POC in image reference list 0 and the current frame, then the symbols in Table 3 can specify the sign of the MV offset added to the reference MV associated with image reference list 1, and the sign of the offset of the reference MV associated with image reference list 0 has the opposite value.

[0188] Table 3 shows an example implementation of the sign of the MV offset specified by the direction index.

[0189] Directional IDX 00 01 10 11 x-axis (horizontal) + - N N y-axis (perpendicular) N N + -

[0190] In some example implementations, the MVD can be scaled based on the difference in POCs in each direction. If the differences in POCs in the two lists are the same, no scaling is required. Otherwise, if the difference in POCs in reference list 0 is greater than the difference in POCs in reference list 1, the MVD of reference list 1 is scaled. If the difference in POCs in reference list 1 is greater than that in list 0, the MVD of list 0 can be scaled in the same way. If the initial MV is a single prediction, the MVD is added to the available or reference MV.

[0191] In some example implementations of bidirectional composite prediction MVD encoding and signaling, in addition to separately encoding and signaling both MVDs, or alternatively, symmetric MVD encoding can be implemented such that only one MVD needs to be signaled, while the other MVD can be derived from the signaled MVD. In such an implementation, motion information including reference picture indices of both list-0 and list-1 is signaled. However, only the MVD associated with, for example, reference list-0 is signaled, and the MVD associated with reference list-1 is derived instead of signaled. Specifically, at the stripe level, a flag may be included in the bitstream, called “mvd_11_zero_flag”, to indicate whether reference list-1 is not signaled in the bitstream. If this flag is 1, indicating that reference list-1 is equal to zero (and therefore not signaled), then the bidirectional prediction flag called “BiDirPredFlag” can be set to 0, meaning there is no bidirectional prediction. Otherwise, if `mvd_11_zero_flag` is zero, and if the most recent reference image in `list-0` and the most recent reference image in `list-1` form a forward and backward reference image pair or a backward and forward reference image pair, then `BiDirPredFlag` can be set to 1, and both `list-0` and `list-1` reference images are short-lived reference images. Otherwise, `BiDirPredFlag` is set to 0. `BiDirPredFlag` being 1 indicates that the symmetric mode flag is additionally signaled in the bitstream. When `BiDirPredFlag` is 1, the decoder can extract the symmetric mode flag from the bitstream. For example, (if needed) the symmetric mode flag can be signaled at the CU level, and it can indicate whether the symmetric MVD coding mode is being used by the corresponding CU. When the symmetric mode flag is 1, it indicates the use of the symmetric MVD encoding mode, and only the reference picture indices of both list-0 and list-1 (referred to as "mvp_10_flag" and "mvp_11_flag") and their associated MVD (referred to as "MVD0") are signaled, while another motion vector difference "MVD1" is derived instead of being signaled. For example, MVD1 can be derived as -MVD0. Therefore, in the exemplary symmetric MVD mode, only one MVD is signaled. In some other example implementations for MV prediction, coordination schemes can be used to implement generally merged mode, MMVD, and some other types of MV prediction for single-reference mode and composite reference mode MV prediction. Various syntax elements can be used to signal the prediction of the MV of the current block.

[0192] For example, for a single-reference mode, the following MV prediction modes can be represented by signals:

[0193] NEARMV directly uses one of the motion vector predictors (MVPs) in the list indicated by the DRL (Dynamic Reference List) index, without using any MVD.

[0194] NEWMV - Uses one of the motion vector predictors (MVPs) in a list represented by signals via DRL as a reference and applies an increment to the MVP (e.g., using MVD).

[0195] GLOBALMV - Uses motion vectors based on frame-level global motion parameters.

[0196] Similarly, for a composite reference frame prediction mode using two reference frames corresponding to the two MVs to be predicted, the following MV prediction mode can be represented by a signal:

[0197] NEAR_NEARMV - Uses one of the motion vector predictors (MVPs) in a list represented by signals via DRL index as a reference, instead of using MVD for each of the two MVs to be predicted.

[0198] NEAR_NEWMV - To predict the first of two motion vectors, one of the motion vector predictors (MVPs) in the list represented by signals via DRL index is used as the reference MV without MVD; to predict the second of two motion vectors, combined with the additional incremental MV (MVD) represented by signals, one of the motion vector predictors (MVPs) in the list represented by signals via DRL index is used as the reference MV.

[0199] NEW_NEARMV - To predict the second of two motion vectors, one of the motion vector predictors (MVPs) in the list represented by signals through the DRL index is used as a reference MV without MVD; to predict the first of two motion vectors, in combination with an additional incremental MV (MVD) represented by signals through the DRL index, one of the motion vector predictors (MVPs) in the list represented by signals through the DRL index is used as a reference MV.

[0200] NEW_NEWMV - Uses one of the motion vector predictors (MVPs) in a list represented by signals via DRL as a reference MV, and combines it with another incremental MV represented by signals to predict each of the two MVs.

[0201] GLOBAL_GLOBALMV - Uses MV from each reference based on frame-level global motion parameters.

[0202] Therefore, the term "NEAR" above refers to MV prediction using a reference MV without MVD as the general merging mode, while the term "NEW" refers to MV prediction involving cancellation using a reference MV and signaling the MVD as in MMVD mode. For composite inter-frame prediction, the aforementioned reference base motion vector and motion vector increments can generally be different or independent between the two references, even if they can be correlated, and this correlation can be used to reduce the amount of information required to signal the two motion vector increments. In this case, joint signaling of the two MVDs can be implemented and indicated in the bitstream.

[0203] The Dynamic Reference List (DRL) above can be used to store the indexed motion vector sets that are dynamically maintained and considered as candidate motion vector predictors.

[0204] In some example implementations, a predefined resolution for MVD can be allowed. For example, a motion vector precision (or accuracy) of 1 / 8 pixel can be allowed. MVD in the various MV prediction modes described above can be constructed and signaled in various ways. In some implementations, various syntax elements can be used to signal one (or more) of the motion vector differences described above in reference frame list 0 or list 1.

[0205] For example, a syntax element called "mv_joint" can specify which components of the motion vector difference it is associated with are non-zero. For MVD, this is the joint signal sent for all non-zero components. For example, the value of mv_joint is:

[0206] 0 can indicate that there is no non-zero MVD along the horizontal or vertical direction;

[0207] 1 can indicate that non-zero MVD exists only in the horizontal direction;

[0208] 2 can indicate that non-zero MVD exists only in the vertical direction;

[0209] 0 can indicate that there is no non-zero MVD along both the horizontal and vertical directions.

[0210] If the "mv_joint" syntax element of MVD signals that no non-zero MVD component exists, then no further MVD information needs to be signaled. However, if the "mv_joint" syntax signals that one or two non-zero components exist, then additional syntax elements for each non-zero MVD component can be signaled, as described below.

[0211] For example, a syntax element called "mv_sign" can be used to additionally specify whether the corresponding motion vector difference component is positive or negative.

[0212] For example, a syntax element called "mv_class" can be used to specify the class of motion vector differences for a predefined set of classes corresponding to non-zero MVD components. For instance, predefined classes of motion vector differences can be used to divide the continuous amplitude space of motion vector differences into non-overlapping ranges, each range corresponding to an MVD class. Therefore, the MVD class transmitted by the signal indicates the amplitude range of the corresponding MVD component. In the exemplary implementation shown in Table 4 below, higher classes correspond to motion vector differences with larger amplitude ranges. In Table 4, the symbol (n, m) is used to represent the range of motion vector differences greater than n pixels and less than or equal to m pixels.

[0213] Table 4. Amplitude Categories of Motion Vector Difference

[0214] MV category MVD amplitude MV_CLASS_0 (0、2] MV_CLASS_1 (2、4] MV_CLASS_2 (4、8] MV_CLASS_3 (8、16] MV_CLASS_4 (16、32] MV_CLASS_5 (32、64] MV_CLASS_6 (64、128] MV_CLASS_7 (128、256] MV_CLASS_8 (256、512] MV_CLASS_9 (512、1024] MV_CLASS_10 (1024、2048]

[0215] In some other examples, the syntax element called "mv_bit" can be further used to specify the integer portion of the offset between the non-zero motion vector difference component and the starting amplitude of the corresponding MV class amplitude range signaled. Thus, mv_bit can indicate the amplitude or oscillation of the MVD. The number of bits required in "my_bit" to signal the full range for each MVD class can vary as a function of the MV class. For example, in the implementation of Table 4, MV_CLASS 0 and MV_CLASS 1 may require only a single bit to indicate an integer pixel offset of 1 or 2 from the starting MVD 0; each higher MV_CLASS in the exemplary implementation of Table 4 may require an "mv_bit" that progressively increases by one bit compared to the previous MV_CLASS.

[0216] In some other examples, the syntax element called "mv_fr" can be further used to specify the first two decimal places of the motion vector difference corresponding to the non-zero MVD components, while the syntax element called "mv_hp" can be used to specify the third decimal place (high-resolution bit) of the motion vector difference corresponding to the non-zero MVD components. Two bits of "mv_fr" essentially provide 1 / 4 pixel MVD resolution, while "mv_hp" bits can further provide 1 / 8 pixel resolution. In some other implementations, more than one "mv_hp" bit can be used to provide a finer MVD pixel resolution than 1 / 8 pixel. In some example implementations, additional flags can be signaled at one or more levels to indicate whether 1 / 8 pixel or higher MVD resolutions are supported. If the MVD resolution is not applied to a particular coding unit, then the syntax element for the corresponding unsupported MVD resolution described above may not be signaled.

[0217] In some of the example implementations above, fractional resolution can be independent of the different classes of MVD. In other words, regardless of the magnitude of the motion vector difference, a similar option for motion vector resolution can be provided using a predefined number of "mv_fr" and "mv_hp" bits to signal fractional MVDs of non-zero MVD components.

[0218] However, in some other example implementations, the resolution of motion vector differences in various MVD amplitude categories can be differentiated. Specifically, a high-resolution MVD with a large MVD amplitude for a higher MVD category may not provide a statistically significant improvement in compression efficiency. Therefore, for a larger range of MVD amplitudes, corresponding to higher MVD amplitude categories, the MVD can be encoded with a decreasing resolution (integer pixel resolution or fractional pixel resolution). Similarly, for generally large MVD values, the MVD can be encoded with a decreasing resolution (integer pixel resolution or fractional pixel resolution). This MVD resolution, which depends on the MVD category or depends on the MVD amplitude, is generally referred to as adaptive MVD resolution, amplitude-dependent adaptive MVD resolution, or amplitude-dependent MVD resolution. The term "resolution" can also be referred to as "pixel resolution," and adaptive MVD resolution can be implemented in various ways as described in the example implementations below to achieve better overall compression efficiency. In particular, due to statistical observations, processing MVD resolutions for large-amplitude or high-class MVDs in a non-adaptive manner at a level similar to that for low-amplitude or low-class MVDs may not significantly increase the inter-frame prediction residual coding efficiency for blocks with large-amplitude or high-class MVDs. Therefore, the number of signal bits reduced by targeting lower-precision MVDs may outweigh the additional bits required to code the inter-frame prediction residuals due to such lower-precision MVDs. In other words, using higher MVD resolutions for large-amplitude or high-class MVDs may not yield much coding gain compared to using lower MVD resolutions.

[0219] In some general example implementations, the pixel resolution or precision of the MVD may or may not increase with the increase of the MVD category. Reducing the pixel resolution of the MVD corresponds to a coarser MVD (or a larger step size from one MVD level to the next). In some implementations, the correspondence between MVD pixel resolution and MVD category can be specified, predefined, or preconfigured, and therefore does not need to be signaled in the encoded bitstream.

[0220] In some example implementations, the MV categories in Table 3 can be associated with different MVD pixel resolutions.

[0221] In some example implementations, each MVD category may be associated with a single allowed resolution. In some other implementations, one or more MVD categories may be associated with two or more optional MVD pixel resolutions. Therefore, an additional signal indicating the optional pixel resolution selected for the current MVD component may follow the signal in the bitstream of the current MVD component having such an MVD category.

[0222] In some example implementations, the adaptively allowed MVD pixel resolutions may include, but are not limited to, 1 / 64 pixel (1 / 64-pel), 1 / 32-pel, 1 / 16-pel, 1 / 8-pel, 1-4-pel, 1 / 2-pel, 1-pel, 2-pel, 4-pel… (in descending order of resolution). Therefore, each ascending MVD category may be associated with one of these MVD pixel resolutions in a non-ascending manner. In some implementations, an MVD category may be associated with two or more resolutions, and a higher resolution may be lower than or equal to a lower resolution of a previous MVD category. For example, if MV_CLASS_3 in Table 4 is associated with optional 1-pel and 2-pel resolutions, then the highest resolution that MV_CLASS_4 in Table 4 can be associated with is 2-pel. In some other implementations, the highest allowed resolution of an MV category may be higher than the lowest allowed resolution of a preceding (lower) MV category. However, the average allowed resolution of ascending MV categories may simply be non-ascending.

[0223] In some implementations, when fractional pixel resolutions higher than 1 / 8-pel are allowed, the “mv_fr” and “mv_hp” signaling can be extended accordingly to a total of more than 3 fractional bits.

[0224] In some example implementations, fractional pixel resolution may be allowed only for MVD categories that are lower than or equal to a threshold MVD category. For example, in Table 4, fractional pixel resolution may be allowed only for MVD-CLASS 0, and not for any of the other MV categories. Similarly, fractional pixel resolution may be allowed only for MVD categories that are lower than or equal to any of the other MV categories in Table 4. For other MVD categories that are higher than the threshold MVD category, only integer pixel resolution of the MVD is allowed. In this way, for MVDs signaled with MVDs of MVD categories higher than or equal to the threshold MVD category, it may not be necessary to signal fractional resolution signaling, such as one or more of the “mv-fr” and / or “mv-hp” bits. For MVD categories with a resolution lower than 1 pixel, the number of bits in the “mv-bit” signaling may be further reduced. For example, for MV_CLASS_5 in Table 4, the MVD pixel offset range is (32, 64], so 5 bits are needed to signal the entire range at 1-pel resolution. However, if MV_CLASS_5 is associated with a 2-pel MVD resolution (below 1-pel resolution), then "mv-bit" may need 4 bits instead of 5 bits, and neither "mv-fr" nor "mv-hp" needs to be signaled as MV-CLASS_5 after "MV_CLASS" is signaled.

[0225] In some example implementations, fractional pixel resolution may be allowed only for MVDs with integer values ​​below a threshold integer pixel value. For example, fractional pixel resolution may be allowed only for MVDs with values ​​less than 5 pixels. Corresponding to this example, fractional resolution is allowed for MV_CLASS_0 and MV_CLASS_1 in Table 4, but not for all other MV categories. As another example, fractional pixel resolution may be allowed only for MVDs with values ​​less than 7 pixels. Corresponding to this example, fractional resolution is allowed for MV_CLASS_0 and MV_CLASS_1 in Table 4 (with a range of less than 5 pixels), but not for MV_CLASS_3 and higher (with a range of more than 5 pixels). For MVDs belonging to MV_CLASS_2, whose pixel range includes 5 pixels, fractional pixel resolution is allowed for the MVD based on the "mv-bit" value. If the "m-bit" value is signaled as 1 or 2 (making the integer part of the signaled MVD 5 or 6, calculated as starting from the pixel range of MV_CLASS_2 with an offset of 1 or 2 indicated by "m-bit"), then fractional pixel resolution is allowed. Otherwise, if the "mv-bit" value is signaled as 3 or 4 (making the integer part of the signaled MVD 7 or 8), then fractional pixel resolution is not allowed.

[0226] In some other implementations, for MV categories equal to or higher than a threshold MV category, only a single MVD value may be allowed. For example, such a threshold MV category could be MV_CLASS2. Therefore, MV_CLASS_2 and above may be allowed only with a single MVD value and without fractional pixel resolution. The single allowed MVD value for these MV categories may be predefined. In some examples, the allowed single value may be the higher end of the corresponding range for these MV categories in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be threshold categories higher than or equal to MV_CLASS2, and the single allowed MVD values ​​for these categories may be predefined as 8, 16, 32, 64, 128, 256, 512, 1024, and 2048, respectively. In some other examples, the allowed single value may be the middle value of the corresponding range for these MV categories in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 can be higher than the category thresholds, and the individual allowed MVD values ​​for these categories can be predefined as 3, 6, 12, 24, 48, 96, 192, 384, 768, and 1536, respectively. Any other value within the range can also be defined as an individual allowed resolution for each MVD category.

[0227] In the above implementation, only the "mv_class" signaling is sufficient to determine the MVD value when the "mv_class" of the transmitted signal is equal to or higher than a predefined MVD class threshold. Then, "mv_class" and "mv_sign" are used to determine the magnitude and direction of the MVD.

[0228] Therefore, when MVD is signaled for only one reference frame (or from reference frame list 0 or list 1, but not from both), or when MVD is signaled for two reference frames in combination, the accuracy (or resolution) of MVD may depend on the category of motion vector difference associated with Table 3 and / or the magnitude of MVD.

[0229] In some other example implementations, the pixel resolution or precision of the MVD may decrease or not increase as the MVD amplitude increases. For example, the pixel resolution may depend on the integer portion of the MVD amplitude. In some implementations, fractional pixel resolution may be allowed only for MVD amplitudes less than or equal to an amplitude threshold. For the decoder, the integer portion of the MVD amplitude may first be extracted from the bitstream. The pixel resolution may then be determined, and a determination may then be made regarding the presence of any fractional MVDs in the bitstream that need to be parsed (e.g., if fractional pixel resolution is not allowed for a particular extracted integer MVD amplitude, then no fractional MVD bits are included in the bitstream that needs to be extracted). The example implementations described above concerning adaptive MVD pixel resolution related to MVD categories are applied to adaptive MVD pixel resolution related to MVD amplitude. For a particular example, MVD categories that are above or include an amplitude threshold may be allowed to have only one predefined value.

[0230] As described in detail above, for the composite reference frame prediction mode, the current coded block can be predicted by two or more reference blocks. Each reference block can be associated with a reference frame (one-way or bi-way relative to the current frame associated with the current coded block). Each reference block can be associated with a motion vector relative to the current coded block. The motion vectors of the reference blocks of the current coded block can typically be different (although they can be the same in some cases). As mentioned above, each of these motion vectors can be predicted by, for example, combining a reference motion vector selected from candidate motion vectors in the DRL with the corresponding MVD.

[0231] In some implementations, for the prediction of any two or more motion vectors associated with two or more reference blocks corresponding to the current coded block, a single MVD (instead of two or more MVDs) can be jointly signaled, and the actual two or more MVDs associated with each of the two or more predicted motion vectors can be derived from the signaled MVD. In other words, for the encoder, two or more reference blocks corresponding to two or more reference frames can be identified first for predicting the current coded block in the current frame. The two or more motion vectors corresponding to the two or more reference blocks can be determined by the encoder. The predictor / reference motion vectors used to predict the two or more motion vectors can be selected from the DRL (e.g., the predictor / reference motion vector can be identified as the motion vector most similar to at least two motion vectors to be predicted). In some other implementations, two or more predictor / reference motion vectors can be identified (in other words, different motion vectors can be associated with different candidate predictor / reference motion vectors in the DRL). The encoder can then obtain two or more MVDs corresponding to the two or more motion vectors by taking the difference between the two or more motion vectors to be predicted and one(or more) corresponding predictor / reference motion vectors (a common predictor / reference motion vector or a separate predictor / reference motion vector). In some example implementations, the two or more MVDs can be jointly encoded into a single MVD in various ways and have information items that can be used to decode the jointly encoded MVD, so that a single MVD of the two or more MVDs can be derived at the decoder.

[0232] Accordingly, the decoder can first extract the jointly encoded MVD and other information items from the bitstream. Then, the decoder can derive a single MVD of two or more MVDs based on the jointly encoded / signaled MVD and other information items. Then, the decoder can derive two or more motion vectors based on the derived single MVD and a single predictor / reference motion vector or a separate predictor / reference motion vector from the DRL written to the bitstream.

[0233] Jointly encoding multiple MVDs of the current coded block in composite reference prediction can be referred to as a joint MVD coding mode. The jointly encoded MVDs can also be based on a fixed or adaptive pixel resolution. Therefore, during the encoding process, the encoder can flexibly determine whether to use joint MVDs to encode the current block in composite reference mode and whether to encode one (or more) MVDs at an adaptive resolution, and represent the selection of these different coding modes with signals in the bitstream. Signaling regarding whether to apply joint MVDs and / or adaptive MVD pixel resolutions can be explicitly or implicitly included in the bitstream in various ways using one or more syntax elements.

[0234] In some example implementations, when encoding the current block in composite reference mode, the signal in the bitstream can be explicitly or implicitly used to indicate whether the motion vector of the current coded block in composite reference mode is predicted using joint MVD. When predicting the motion vector of the current coded block in composite reference mode using joint MVD, one (or more) flags can be signaled in the bitstream to indicate whether adaptive MVD resolution is applied to the jointly encoded MVD. This signaling can be provided at various coding levels, including but not limited to sequence level, frame level, image level, slice level, superblock level, coding unit level, and coding block level.

[0235] In some example implementations, when adaptive MVD resolution is applied to a jointly encoded MVD coding mode, as signaled in various ways in the bitstream, the joint MVD of two (or more) reference frames can be further signaled, and the precision of the MVD can be implicitly determined based on the MVD category and / or MVD magnitude (e.g., the adaptive MVD resolution in Table 4 or any other adaptive pixel resolution related to the aforementioned MVD category or magnitude). Therefore, the decoder can determine the MVD pixel resolution accordingly by extracting the MVD category or magnitude signaled in the bitstream. This adaptive resolution scheme can be based on a predefined or dynamically configured correspondence between the MVD category or magnitude and the MVD pixel resolution. Thus, the decoder can first determine from the bitstream whether to apply the joint MVD resolution, and if so, whether to apply the adaptive MVD pixel resolution.

[0236] In some other example implementations, when applying joint MVD by signal representation or from the bitstream without applying adaptive MVD resolution to the current coded block, the MVD of two (or more) reference frames is determined as joint signal representation; however, the precision of the joint signal representation MVD is fixed and independent of the category and / or magnitude of the joint signal representation MVD, rather than being adaptive.

[0237] In some specific example implementations, an explicit or implicit flag in the bitstream may first be signaled to indicate whether the current coded block in the composite reference prediction mode is predicted based on joint MVD. When indicating that the current block in the composite reference prediction mode is predicted using joint MVD via such implicit or explicit signaling, another explicit flag, called amvd_flag, may be further included in / signed in the bitstream to indicate whether adaptive MVD resolution is applied to the joint signaled MVD.

[0238] In some implementations, the context can be derived and used for entropy coding of the adaptive MVD pixel resolution flag (e.g., amvd_flag). In some alternative implementations, the context for entropy coding of the adaptive MVD pixel resolution flag can depend on the encoded information of the current coding block and / or its neighboring blocks. Such encoded information may include, for example, but is not limited to: the adaptive MVD pixel resolution flags (e.g., amvd_flag values) of the neighboring blocks of the current coding block, the reference frame index of the current coding block, and / or one (or more) MVD candidate indices of the current coding block in the DRL. The context used to represent the adaptive MVD pixel resolution flag with a signal can be determined by one or more such encoded information. The basis for this context design may be a probabilistic model of whether the adaptive MVD pixel resolution is more similar to neighboring blocks, and / or more similar to coding blocks using similar MVD candidate indices and / or frame indices.

[0239] In some specific implementation examples, an explicit or implicit flag in the bitstream (e.g., an explicit amvd_flag) may first be signaled to indicate whether the current coded block in the composite reference prediction mode is predicted based on the adaptive MVD pixel resolution. When the current block in the composite reference prediction mode is indicated by such an explicit or implicit flag to be predicted based on the adaptive MVD pixel resolution, another explicit flag, called jmvd_flag, may be further included in / signed in the bitstream to indicate whether joint MVD prediction of the motion vectors of the current coded block is applied. Therefore, the decoder can first determine from the bitstream whether the composite reference inter-frame prediction mode is applied, and if so, whether the adaptive MVD pixel resolution is applied to the current coding. If the current coded block is inter-coded in the composite reference inter-frame prediction mode and the adaptive MVD pixel resolution is applied, then the decoder can then determine from the bitstream whether joint MVD is applied to the current block. If the current coded block is not inter-coded in the composite reference inter-frame prediction mode, then the encoder can determine that joint MVD is not applied regardless of whether the adaptive MVD pixel resolution is applied.

[0240] In some example implementations, the joint MVD and adaptive MVD resolution modes can be implemented as sub-prediction modes of the normal composite reference inter-frame prediction mode (e.g., the NEW_NEWMV composite reference inter-frame prediction mode described above) using predicted motion vectors. Specifically, as mentioned above, in the normal composite reference inter-frame prediction mode, for example, with two reference frames (the basic principle applies to composite reference inter-frame prediction cases with more than two reference frames), MVs can be predicted or constructed in various ways, and may or may not involve MVD. For example, in the NEAR_NEARMV mode, both MVs are predicted directly without using any MVD; in the NEAR_NEWMV or NEW_NEARMV mode, one of the two MVs is predicted using MVD, while the other MV is not; in the NEW_NEWMV mode, both MVs are predicted using MVD; in the GLOBAL_GLOBALMV mode, both MVDs are based on frame-level global motion parameters. By adding one or more sub-modes to the joint MVD with adaptive resolution, sub-modes in the composite reference inter-frame prediction mode with the exemplary two reference frames can, for example, include:

[0241] ADAPTIVE_JOINTMV - To predict two motion vectors (MVs), a joint MVD represented by a signal for the two reference frames with adaptive MVD pixel resolution is used, along with one (or more) motion vector predictors represented by a DRL list through one or more DRL indices.

[0242] (MVP).

[0243] FIXED_JOINTMV - To predict two MVs, use pixel resolution without adaptive MVD.

[0244] The MVD (or a joint signal representation for two reference frames with a fixed resolution) and one (or more) DRLs in a list represented by one or more DRL indices.

[0245] MVP.

[0246] FIXED-FIXEDMV - To predict two MVs, MVDs are represented by signals individually or independently for two reference frames. The two MVDs do not have an adaptive MVD pixel resolution (or have a fixed resolution), and MVPs are in the DRL list via one or more DRL index signaling.

[0247] FIXED-ADAPTIVEMV; ADAPTIVE-FIXEDMV - To predict two MVs, MVDs are represented by signals individually or independently, one MVD having an adaptive MVD pixel resolution, but the other MVD not having an adaptive MVD pixel resolution (fixed MVD resolution).

[0248] And the MVP in the DRL list represented by one or more DRL indices using signals.

[0249] ADAPTIVE-ADAPTIVEMV- To predict two MVs, MVDs are represented individually or independently with signals for two reference frames. The two MVDs have adaptive MVD pixel resolution, and MVPs are represented in a DRL list with signals through one or more DRL indices.

[0250] NEAR_NEARMV - For each of the two MVs to be predicted, use one of the MVPs in the DRL list represented by the DRL index, instead of using MVD.

[0251] NEAR_NEWMV; FIXED_NEARMV - To directly predict one of two MVs, use one of the MVPs in the DRL list represented by the DRL index as the reference MV, without using any MVD; To predict the other MV of the two motion vectors, combine it with another incremental (delta) MV represented by the signal with a fixed MVD pixel resolution.

[0252] (MVD) uses one of the MVPs in the DRL list, which is represented by a signal via the DRL index, as the reference MV.

[0253] NEAR_NEWMV or ADAPTIVE_NEARMV - To predict one of two MVs, use one of the MVPs in the DRL list represented by signals via DRL indexing.

[0254] As a reference MV, without using any MVD; to predict the other of the two motion vectors, an additional incremental MV represented by the signal with adaptive MVD pixel resolution is combined.

[0255] (MVD) uses one of the MVPs in the DRL list, which is represented by a signal via the DRL index, as the reference MV.

[0256] GLOBAL_GLOBALMV - Uses MV from each reference based on frame-level global motion parameters.

[0257] ADAPTIVE_JOINTMV, FIXED_JOINTMV, FIXED–FIXEDMV, FIXED–ADAPTIVEMV, and ADAPTIVE–FIXEDMV can constitute sub-modes of the normal NEW-NEWMV mode, while NEAR_FIXEDMV, FIXED_NEARMV, NEAR_ADAPTIVEMV, and ADAPTIVE_NEARMV can constitute various variations of the normal NEW_NEARMV and NEAR_NEWMV modes. The actual sub-modes of the composite reference inter-frame prediction mode can be implemented as subsets of the above sub-modes, and it is not necessary to include the entire list of sub-modes. For example, when predicting two MVs, the adaptive mode can be applied to both MVDs, or neither MVD can apply the adaptive mode. Therefore, it is not necessary to implement FIXED–ADAPTIVEMV and ADAPTIVE–FIXEDMV. When implementing the various sub-modes of the composite reference inter-frame prediction mode listed above, they can be represented by signals using a syntax in the bitstream and other MVD and MV-related syntaxes.

[0258] In the exemplary simplification of the above scheme, the normal NEW_NEWMV mode can therefore be replaced by two sub-modes. In the first sub-mode: the MVD is jointly encoded with adaptive resolution. In the second sub-mode, the MVD is encoded separately / independently, and both have adaptive resolution.

[0259] In this implementation, syntax elements can be used in the bitstream to represent composite reference inter-predictive sub-modes as signals. By extracting these syntax elements from the bitstream, the decoder can determine whether the joint MVD is associated. Based on the same syntax elements, it can be further determined whether adaptive MVD pixel resolution is applied, and if so, which one(s) MVD(s) applies the adaptive MVD pixel resolution.

[0260] In some example implementations, when adaptive MVD resolution is applied to a joint MVD coding mode, optical flow thinning can always be applied to the current block if certain conditions are met. Such conditions may include, but are not limited to, coding block size constraints. Optical flow thinning can be used to perform pixel-level compensation for fine motion missed by typical block-based motion compensation schemes. When adaptive MVD resolution is applied to a joint MVD coding mode, optical flow thinning can always be achieved to obtain a non-negligible coding gain. When adaptive MVD resolution or a joint MVD coding mode is not applied, optical flow thinning can be skipped because the coding gain of optical flow thinning is negligible.

[0261] In some example implementations, position-dependent composite prediction may be disabled when the current coding block is encoded in the joint MVD coding mode. For example, in this case, composite wedge-based prediction may not be allowed. When joint MVD is applied, position-dependent composite prediction may not yield significant coding gain, so it can be disabled to conserve coding resources. When joint MVD is not applied, position-dependent composite prediction, such as composite wedge-based prediction, can be enabled and allowed.

[0262] In some example implementations, position-dependent composite prediction may be disabled when adaptive MVD resolution is applied to the joint MVD coding mode. For example, composite wedge-based prediction may not be allowed in this case. When adaptive MVD pixel resolution is applied to joint MVD, position-dependent composite prediction may not yield significant coding gain and can therefore be disabled to conserve coding resources. When adaptive MVD pixel resolution is not applied or joint MVD is not applied, position-dependent composite prediction, such as composite wedge-based prediction, can be enabled and allowed.

[0263] In some example implementations, when a coding block is encoded in a joint MVD coding mode, only one type of interpolation filter may be used. For example, in this case, only one of a regular filter, a smooth filter, or a sharpening filter may be used to reduce encoder complexity. For instance, when joint MVD is applied to the current coding block, the coding gain from using more complex interpolation filters may not provide substantial coding gain. In this implementation, when the coding block is not encoded with joint MVD, the type of interpolation filter is not limited to the filter types mentioned above.

[0264] Similarly, in some other implementations, when adaptive MVD pixel resolution is applied to the joint MVD coding mode, only one type of interpolation filter can be used. For example, in this case, only one of the regular filter, smoothing filter, or sharpening filter can be used to reduce encoder complexity. For example, when joint MVD or adaptive MVD pixel resolution is applied to the current coding block, the coding gain from using more complex interpolation filtering may not provide substantial coding gain. In this implementation, when the coding block is not encoded with joint MVD or adaptive MVD pixel resolution, the type of interpolation filter is not limited to the filter types mentioned above.

[0265] In some other implementations, adaptive MVD resolution for a joint MVD coding mode is not permitted when adaptive MVD resolution or joint MVD coding for the current video sequence or frame is not allowed (e.g., via sequence or frame-level signaling). In other words, for example, if adaptive MVD is signaled as not being applied at a higher level (e.g., sequence-level or frame-level), then adaptive MVD resolution for joint MVD coding is not permitted regardless of lower-level signaling. Similarly, if joint MVD is signaled as not being applied at a higher level, then adaptive MVD resolution for joint MVD coding is not permitted regardless of lower-level signaling.

[0266] Figure 18 A flowchart 1800 illustrates an example method following the basic principles of the aforementioned joint adaptive MVD resolution implementation. The example decoding method flow begins at 1801. In S1810, a video stream is received. In S1820, it is determined from the video stream whether joint motion vector difference (MVD) encoding is applied to the current video block. In S1830, it is determined from the video stream whether adaptive MVD pixel resolution is applied to the current video block. In S1840, the current video block is decoded based on the joint MVD encoding and whether the adaptive MVD pixel resolution is applied to the current video block. The example method stops at S1899.

[0267] In the embodiments and implementations of this disclosure, any number or order of steps and / or operations can be combined or arranged as needed. Two or more of the steps and / or operations can be performed in parallel. The embodiments and implementations of this disclosure can be used individually or in combination in any order. Furthermore, each method (or implementation), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transient computer-readable medium. The embodiments in this disclosure can be applied to luma blocks or chroma blocks. The term block can be interpreted as a prediction block, coding block, or coding unit, i.e., CU. The term block in this document can also be used to refer to a transform block. In the following items, when referring to block size, it can refer to the width or height of the block, or the maximum value of the width and height, or the minimum value of the width and height, or the area size of the block (width * height), or the aspect ratio (width:height, or height:width).

[0268] The above techniques can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 19 A computer system (1900) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0269] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode, etc.

[0270] The instructions can be executed on various types of computers or their components, including personal computers, tablets, servers, smartphones, gaming devices, and Internet of Things devices.

[0271] Figure 19 The components of the computer system (1900) shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of computer software implementing embodiments of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any component or combination of components shown in the exemplary embodiments of the computer system (1900).

[0272] Computer systems (1900) may include certain human-computer interface input devices. Such human-computer interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). Human-computer interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images acquired from still image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0273] Input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard (1901), mouse (1902), touchpad (1903), touch screen (1910), data glove (not shown), joystick (1905), microphone (1906), scanner (1907), camera (1908).

[0274] The computer system (1900) may include certain human-machine interface output devices. Such human-machine interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touchscreen (1910), data gloves (not shown), or joysticks (1905), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (1909), headphones (not depicted)), and visual output devices (e.g., screens including CRT screens, LCD screens, plasma screens, OLED screens (1910), each with or without touchscreen input functionality, each with or without tactile feedback functionality—some of which are capable of outputting two-dimensional or more three-dimensional visual outputs through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).

[0275] Computer systems (1900) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1920) with media such as CD / DVD (1921), finger drives (1922), removable hard disk drives or solid-state drives (1923), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0276] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0277] The computer system (1900) may also include an interface (1954) to one or more communication networks (1955). The network may be, for example, a wireless network, a wired network, or an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CAN buses, etc. Some networks typically require external network interface adapters (e.g., a USB port on the computer system (1900)) to connect to certain general-purpose data ports or peripheral buses (1949); other network interfaces are typically integrated into the core of the computer system (1900) by connecting to the system bus (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system), as described below. The computer system (1900) can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., CANbus connected to certain CANbus devices), or bidirectional, such as connecting to other computer systems using a local area network (LAN) or wide area network (WAN) digital network. As mentioned above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.

[0278] The aforementioned human-machine interface device, human-machine accessible storage device, and network interface can be attached to the kernel (1940) of the computer system (1900).

[0279] The core (1940) may include one or more central processing units (CPU) (1941), graphics processing units (GPUs) (1942), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1943), hardware accelerators for certain tasks (1944), graphics adapters (1950), etc. These devices, along with read-only memory (ROM) (1945), random access memory (1946), and internal mass storage such as internal non-user-accessible hard disk drives, SSDs, etc. (1947), can be connected via a system bus (1948). In some computer systems, the system bus (1948) can be accessed in the form of one or more physical plugs to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the core's system bus (1948) or connected to the core's system bus (1848) via a peripheral bus (1949). In one example, a screen (1910) may be connected to a graphics adapter (1950). Peripheral bus architectures include PCI, USB, etc.

[0280] The CPU (1941), GPU (1942), FPGA (1943), and accelerator (1944) can execute certain instructions, which can be combined to form the aforementioned computer code. This computer code can be stored in ROM (1945) or RAM (1946). Transient data can also be stored in RAM (1946), while permanent data can be stored, for example, in internal mass storage (1947). Fast storage and retrieval to any storage device can be achieved by using a cache, which can be closely associated with one or more CPUs (1941), GPUs (1942), mass storage (1947), ROM (1945), RAM (1946), etc.

[0281] Computer-readable media may have computer code thereon for performing various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.

[0282] As a non-limiting example, a computer system having an architecture (1900), particularly a kernel (1940), can be made functional by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, and some non-transitory memory of the kernel (1940), such as internal kernel mass storage (1947) or ROM (1945). Software implementing the various embodiments of this disclosure can be stored in such means and executed by the kernel (1940). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can cause the kernel (1940), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform a particular process or a particular portion of a particular process described herein, including defining data structures stored in RAM (1946) and modifying such data structures according to a software-defined process. Additionally or alternatively, a computer system may be made functional by hard-wired or otherwise embodied logic in circuitry (e.g., an accelerator (1944)) that may replace or operate with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., an integrated circuit (IC)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.

[0283] Although several exemplary embodiments have been described in this disclosure, modifications, substitutions, and various equivalent alternatives that fall within the scope of this disclosure exist. Therefore, it should be understood that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.

[0284] Appendix A: Acronyms

[0285] JEM: Joint Exploration Model

[0286] VVC: Universal Video Coding

[0287] BMS: Benchmark Set

[0288] MV: Motion Vector

[0289] HEVC: High-Efficiency Video Coding

[0290] SEI: Supplemental Enhancement Information

[0291] VUI: Video Availability Information

[0292] GOP: Image Group

[0293] TU: Transformation Unit

[0294] PU: Prediction Unit

[0295] CTU: Coding Tree Unit

[0296] CTB: Coded Tree Block

[0297] PB: Prediction Block

[0298] HRD: Hypothetical Reference Decoder

[0299] SNR: Signal-to-noise ratio

[0300] CPU: Central Processing Unit

[0301] GPU: Graphics Processing Unit

[0302] CRT: Cathode Ray Tube

[0303] LCD: Liquid Crystal Display

[0304] OLED: Organic Light Emitting Diode

[0305] CD: CD-ROM

[0306] DVD: Digital Video Disc

[0307] ROM: Read-Only Memory

[0308] RAM: Random Access Memory

[0309] ASIC: Application-Specific Integrated Circuit

[0310] PLD: Programmable Logic Device

[0311] LAN: Local Area Network

[0312] GSM: Global System for Mobile Communications

[0313] LTE: Long Term Evolution

[0314] CANBus: Controller Area Network Bus

[0315] USB: Universal Serial Bus

[0316] PCI: Interconnect Peripheral Devices

[0317] FPGA: Field Programmable Gate Domain

[0318] SSD: Solid State Drive

[0319] IC: Integrated Circuit

[0320] HDR: High Dynamic Range

[0321] SDR: Standard Dynamic Range

[0322] JVET: Joint Video Exploration Committee

[0323] MPM: Most Likely Pattern

[0324] WAIP: Wide-angle Intra-frame Prediction

[0325] CU: Encoding Unit

[0326] PU: Prediction Unit

[0327] TU: Transformation Unit

[0328] CTU: Coding Tree Unit

[0329] PDPC: Location-Related Prediction Combination

[0330] ISP: Intra-Frame Sub-Partition

[0331] SPS: Sequence Parameter Settings

[0332] PPS: Image Parameter Set

[0333] APS: Adaptive Parameter Set

[0334] VPS: Video Parameter Set

[0335] DPS: Decoding Parameter Set

[0336] ALF: Adaptive Loop Filter

[0337] SAO: Sampling Adaptive Offset

[0338] CC-ALF: Cross-component adaptive loop filter

[0339] CDEF: Constrained Direction Enhancement Filter

[0340] CCSO: Cross Component Sample Offset

[0341] LSO: Local Sampling Offset

[0342] LR: Loop Recovery Filter

[0343] AV1 (AOMedia Video 2): Open Media Alliance Video 1

[0344] AV2 (AOMedia Video 2): Open Media Alliance Video 2

[0345] MVD: Motion Vector Difference

[0346] CfL: Predicting chromaticity from luminance

[0347] SDT: Semi-decoupled tree

[0348] SDP: Semi-decoupled partitioning

[0349] SST (Semi Separate Tree): A tree structured in a semi-separate manner.

[0350] SB: Superblock

[0351] IBC (or IntraBC): Intra-block copy

[0352] CDF: Cumulative Density Function

[0353] SCC: Screen Content Encoding

[0354] GBI: Generalized Dual Prediction

[0355] BCW: Dual Prediction Based on CU-Level Weights

[0356] CIIP: Combined Intra-Inter-Frame Prediction

[0357] POC: Image Sequential Counting

[0358] RPS: Reference Image Gallery

[0359] DPB: Decoding Image Buffer

[0360] MMVD: Merging Mode with Motion Vector Difference

Claims

1. A method for processing a current video block in a video stream, characterized in that, include: Receive the video stream; Based on at least two reference blocks associated with at least two corresponding motion vectors, the current video block is determined to be inter-frame coded in composite reference mode; In response to determining that the current video block is inter-coded in the composite reference mode, a first flag and a second flag are extracted from the video stream, wherein the first flag indicates whether the motion vector difference (MVD) of the current video block is jointly represented or decoded using a signal, and the second flag indicates whether an adaptive MVD pixel resolution is applied to the current video block, wherein the adaptive MVD pixel resolution is determined based on the correspondence between MVD attributes and MVD pixel resolution, and the MVD attributes include at least one of the following: MVD category, MVD magnitude; The current video block is decoded based on the first flag and the second flag, wherein when the first flag indicates the joint MVD encoding and the second flag indicates that the adaptive MVD pixel resolution is applied to the current video block, optical flow refinement is applied to the current video block if the video block size constraint is satisfied.

2. The method according to claim 1, characterized in that, The method further includes the following when the first flag indicates that the motion vector difference is jointly represented or decoded using a signal, and the second flag indicates that adaptive MVD pixel resolution is applied: Extract the MVD category or magnitude of the joint MVD from the video stream; The current MVD pixel resolution of the joint MVD is determined based on the MVD category or magnitude; Extract the joint MVD from the video stream based on the current MVD pixel resolution; and The at least two corresponding motion vectors are derived based on the joint MVD.

3. The method according to claim 1, characterized in that, The method further includes the following steps when the first flag indicates that the motion vector difference is jointly represented or decoded using a signal, and the second flag indicates that no adaptive MVD pixel resolution is applied: Extract joint MVD from the video stream based on a fixed MVD pixel resolution; and The at least two corresponding motion vectors are derived based on the joint MVD.

4. The method according to claim 1, characterized in that, The method further includes the following when the first flag indicates that the motion vector difference has not been jointly represented or decoded by a signal, and the second flag indicates that adaptive MVD pixel resolution is applied: Extract the MVD category or magnitude associated with the at least two corresponding motion vectors from the video stream; The current MVD pixel resolution of the at least two corresponding motion vectors is determined based on the MVD category or magnitude, respectively. Extract individual MVDs from the video stream based on the current MVD pixel resolution; and At least two corresponding motion vectors are derived based on the individual MVDs.

5. The method according to claim 1, characterized in that, The method further includes the following steps when the first flag indicates that the motion vector difference has not been jointly represented or decoded by a signal, and the second flag indicates that adaptive MVD pixel resolution has not been applied: Extract individual MVDs from the video stream based on a fixed MVD pixel resolution; and At least two corresponding motion vectors are derived based on the individual MVDs.

6. The method according to claim 1, characterized in that, In the video stream, the first flag is indicated by a signal before the second flag.

7. The method according to claim 6, characterized in that, The context is used to represent the entropy encoding of the second flag, wherein the second flag includes amvd_flag.

8. The method according to claim 7, characterized in that, The context used to represent the second flag with a signal depends on the encoded information of the current video block and / or the adjacent video blocks of the current video block.

9. The method according to claim 8, characterized in that, The encoded information includes at least one of the following: the value of the second flag, the reference frame index, or the MVD candidate index of the adjacent video blocks of the current video block.

10. The method according to claim 1, characterized in that, In the video stream, the second flag is indicated by a signal preceding the first flag.

11. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Position-dependent composite prediction is not allowed when the joint MVD coding is applied to the current video block or when both the joint MVD coding and the adaptive MVD pixel resolution are applied to the current video block.

12. The method according to claim 11, characterized in that, The location-related composite predictions include predictions based on composite wedges.

13. The method according to any one of claims 1 to 8, characterized in that, The method further includes: When the joint MVD encoding is applied to the current video block, or when both the joint MVD encoding and the adaptive MVD pixel resolution are applied to the current video block, only one type of interpolation filter is used, and interpolation filters other than regular filters, smoothing filters, or sharpening filters are not allowed.

14. An apparatus for processing a current video block in a video stream, characterized in that, It includes a memory for storing computer instructions and a processor for executing the computer instructions to perform the method according to any one of claims 1 to 13.

15. A computer-readable storage medium, characterized in that, Includes instructions that, when run on a computer, cause the computer to perform the method according to any one of claims 1 to 13.

16. A method for processing video streams, characterized in that, The video stream is decoded based on the method described in any one of claims 1 to 13.