Method and electronic device for decoding inter prediction video block of video stream

By receiving and processing the motion vector difference (MVD) and reference motion vector information in the video stream, the problem of inefficient decoding of inter-frame prediction video blocks in the prior art is solved, and a more efficient video decoding process is realized.

CN120050422APending Publication Date: 2025-05-27TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510202134.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-03-30
Filing Date
2022-04-13
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the inter-prediction video block decoding process of video streams, it is difficult to efficiently process the information of motion vector difference (MVD) and reference motion vectors, resulting in low decoding efficiency.

Method used

By receiving the MVD information in the video stream, determining its pixel resolution, and extracting additional MVD information to identify the selected optional pixel resolution, ultimately decoding the inter-prediction video block based on this information.

Benefits of technology

The decoding efficiency of the inter-frame prediction video block of video streams is improved, and the MVD and reference motion vector information can be processed more accurately, thereby optimizing the video decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050422A_ABST
    Figure CN120050422A_ABST
Patent Text Reader

Abstract

The invention provides a method for decoding an inter-frame prediction video block of a video stream and electronic equipment. The method comprises the following steps: receiving the video stream; determining that a motion vector difference (MVD) between a motion vector associated with the inter prediction video block and a reference motion vector is written to the video stream, the reference motion vector corresponding to one reference picture in only one of a reference frame list 0 and a reference frame list 1 unless the MVD is jointly written for two reference pictures; obtaining, from the video stream, an indication of a size range of the MVD in a plurality of predefined motion vector difference size ranges; determining the pixel resolution of the MVD according to the size range of the MVD; identifying additional MVD information in the video stream based on the pixel resolution of the MVD, the additional MVD information indicating a selectable pixel resolution selected by the MVD; extracting additional MVD information from the video stream; and decoding the inter prediction video block based on the pixel resolution of the MVD, the additional MVD information, the reference motion vector, and a reference frame associated with the motion vector.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application with the application date of April 13, 2022, the Chinese patent application number of 202280007737.9, and the invention title of "Method and Electronic Device for Decoding Inter-Frame Predicted Video Blocks of a Video Stream". Technical Field

[0002] This application generally relates to video encoding / decoding, and more particularly to a method and an electronic device for decoding inter-frame predicted video blocks of a video stream. Background Art

[0003] The background art provided herein is for the purpose of generally describing the content of the embodiments of this application. Certain work of the inventors (i.e., the work already described in this background art section) and the content in the specification regarding certain prior art that has not become prior art before the filing date are not regarded as prior art with respect to the embodiments of this application, whether explicitly or implicitly.

[0004] Inter-frame picture prediction with motion compensation can be used to perform video encoding and decoding. Uncompressed digital video can include a series of pictures, each picture having a spatial size of, for example, 1920×1080 luminance samples and associated full-sampled or subsampled chrominance samples. The series of pictures can have a fixed or variable picture rate (alternatively, called frame rate) of, for example, 60 pictures per second or 60 frames per second. Uncompressed video has specific bitrate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and a 4:2:0 chrominance subsampling of 8 bits per pixel per color channel requires a bandwidth of nearly 1.5 Gbit / s. An hour of such video requires more than 600 GB of storage space.

[0005] One purpose of video encoding and decoding can be to reduce redundancy in an uncompressed input video signal through compression. Compression can help reduce the above-mentioned bandwidth and / or storage space requirements, in some cases by two orders of magnitude or more than two orders of magnitude. Lossless compression and lossy compression, as well as combinations thereof, can be employed. Lossless compression refers to techniques by which an exact copy of the original signal can be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to an encoding / decoding process in which the original video information is not fully retained during encoding and not fully restored during decoding. When lossy compression is used, the reconstructed signal may be different from the original signal. Although some information is lost, the distortion between the original signal and the reconstructed signal is small enough for the reconstructed signal to be useful for the intended application. In the case of video, lossy compression is widely used in many applications. The amount of distortion that can be tolerated depends on the application. For example, users of some consumer video streaming applications can tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable through a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows the encoding algorithm to produce higher losses and higher compression ratios.

[0006] Video encoders and decoders can utilize techniques from multiple categories and steps, which include, for example, motion compensation, Fourier transform, quantization, and entropy coding.

[0007] Video codec technology can include techniques known as intra-frame encoding. In intra-frame encoding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, a picture is spatially subdivided into sample blocks. When all sample blocks are encoded in the intra-frame mode, the picture can be referred to as an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) can be used to reset the decoder state and thus can be used as the first picture in an encoded video bitstream and a video session, or as a still image. Then, the samples of the blocks after intra-frame prediction can be transformed into the frequency domain, and the transformation coefficients so generated can be quantized before entropy coding. Intra-frame prediction represents a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are required to represent the block after entropy coding for a given quantization step size.

[0008] For example, traditional intra-frame coding known from, e.g., MPEG-2 generation coding techniques does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to encode / decode a block based on, for example, surrounding sample data and / or metadata obtained during spatial neighboring encoding and / or decoding and preceding the data block to be intra-frame encoded or decoded in decoding order. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. It should be noted that, at least in some cases, intra-frame prediction uses only reference data from the currently being reconstructed picture and not reference data from other reference pictures.

[0009] Intra-frame prediction can have many different forms. When more than one such technique is available in a given video coding technique, the technique in use can be referred to as an intra-frame prediction mode. One or more intra-frame prediction modes can be provided in a particular codec. In some cases, a mode can have sub-modes, and / or can be associated with various parameters, and the mode / sub-mode information and intra-frame coding parameters for a video block can be encoded separately or jointly included in a mode codeword. Which codeword is used for a given mode, sub-mode, and / or parameter combination can affect the coding efficiency gain through intra-frame prediction, and the entropy coding technique used to convert the codeword into a bitstream can also have an impact on it.

[0010] H.264 introduced certain intra-frame prediction modes, which were improved in H.265 and further improved in new coding techniques such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), Benchmark Set (BMS), etc. Generally, for intra-frame prediction, available neighboring sample values that have become available can be used to form a prediction block. For example, the available values of a particular set of neighboring samples can be copied into the prediction block along some directions and / or lines. The reference to the direction used can be encoded in the bitstream, or it can be predicted itself. The reference to the direction used can be encoded in the bitstream, or it can be predicted itself.

[0011] Reference Figure 1A, a subset of 9 prediction directions specified among 33 possible intra prediction directions of H.265 (corresponding to 33 angular modes out of 35 intra modes specified in H.265) is depicted in the lower right. The point (101) where the arrows converge indicates the sample being predicted. The arrows indicate the directions along which adjacent samples are used to predict the sample at 101. For example, arrow (102) indicates that sample (101) is predicted from one or more adjacent samples at the upper right, at a 45-degree angle to the horizontal direction. Similarly, arrow (103) indicates that sample (101) is predicted from one or more adjacent samples at the lower left of sample (101), at a 22.5-degree angle to the horizontal direction.

[0012] Still referring to Figure 1A , a square block (104) of 4×4 samples is depicted in the upper left (represented by a bold dashed line). Square block (104) contains 16 samples, and each sample is labeled using "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (starting from the top) and the first sample in the X dimension (starting from the left). Similarly, sample S44 is the fourth sample in both the Y dimension and the X dimension within block (104). Since the block is 4×4 samples in size, S44 is in the lower right corner. Example reference samples following a similar numbering scheme are also shown. The reference samples are labeled using "R" and their Y position (e.g., row index) and X position (e.g., column index) relative to block (104). In H.264 and H.265, prediction samples adjacent to the block being reconstructed are used.

[0013] Intra picture prediction of block 104 can start by copying reference sample values from adjacent samples according to the signaled prediction direction. For example, assume that the encoded video bitstream includes signaling that indicates, for block 104, the prediction direction of arrow (102), i.e., predicting samples from one or more prediction samples at the upper right, at a 45-degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted according to reference sample R08.

[0014] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, in order to calculate the reference sample, especially when the direction is not divisible by 45 degrees.

[0015] As video coding technology continues to evolve, the number of possible directions increases. For example, in H.264 (in 2003), nine different directions were available for intra prediction. In H.265 (in 2013), this increased to 33 directions, and at the time of the embodiments of the present application, JEM / VVC / BMS can support up to 65 directions. Experimental studies have been conducted to help identify the most suitable intra prediction directions, and some techniques in entropy coding can be used to encode those most suitable directions with a small number of bits, accepting a certain bit cost for the directions. Additionally, sometimes the direction itself can be predicted from neighboring directions used in the intra prediction of already decoded neighboring blocks.

[0016] Figure 1B A schematic diagram (180) is shown, which depicts 65 intra prediction directions according to JEM, to illustrate the increase in the number of prediction directions in various coding techniques developed over time.

[0017] The way the bits representing the intra prediction direction in the encoded video bitstream are mapped to the prediction direction may vary depending on the video coding technology; for example, it can range from a simple direct mapping of the prediction direction to the intra prediction mode, mapping to codewords, mapping to complex adaptive schemes involving the most likely modes, and similar techniques. However, in all cases, for intra prediction, there may be certain directions that are statistically less likely to occur in the video content compared to some other directions. Since the goal of video compression is to reduce redundancy, in a well-designed video coding technology, those less likely directions will be represented by more bits compared to the likely directions.

[0018] Inter-picture prediction or inter prediction can be based on motion compensation. In motion compensation, sample data from a previously reconstructed picture or a portion thereof (reference picture) can be used to predict a newly reconstructed picture or picture portion (e.g., block) after being spatially offset along the direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference picture being used (similar to the temporal dimension).

[0019] In some video compression techniques, a current motion vector (MV) applicable to a certain region of sample data may be predicted based on other MVs, for example, other MVs that are related to other regions of sample data adjacent in space to the region being reconstructed and that are in the decoding order before the current MV. Doing so can significantly reduce the total amount of data required to encode the MVs by eliminating redundancy in the related MVs, thereby increasing the compression efficiency. MV prediction can work effectively, for example, because when encoding an input video signal obtained from a camera (referred to as natural video), there is the following statistical possibility: in a video sequence, regions larger than a region applicable to a single MV move along similar directions. Therefore, in some cases, similar motion vectors derived from the MVs of adjacent regions can be used to predict the larger region. This makes the actual MV for a given region similar to or the same as the MV predicted based on the surrounding MVs. Further, after entropy coding, the MV can be represented using fewer bits than when directly encoding the MV (instead of predicting the MV based on adjacent MVs). In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, for example, due to rounding errors that occur when calculating the predicted value based on multiple surrounding MVs, the MV prediction itself can be lossy.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms specified in H.265, the technique described herein is a technique hereinafter referred to as "spatial merge".

[0021] Specifically, referring to Figure 2 , the current block (201) includes samples that have been discovered by the encoder during the motion search process, and the samples can be predicted based on a previous block of the same size that has had a spatial offset. Additionally, the MV can be derived from metadata associated with one or more reference pictures instead of directly encoding the MV. For example, using the MV associated with any one of five surrounding samples A0, A1 and B0, B1, B2 (corresponding to 202 to 206 respectively), the MV is derived from the metadata of the nearest reference picture (in the decoding order). In H.265, MV prediction can use the predicted value from the same reference picture that the adjacent blocks are using. Summary of the Invention

[0022] Embodiments of the present application provide a method for decoding an inter-frame predicted video block of a video stream, an electronic device, and a method for storing or transmitting a video stream.

[0023] The technical solution of the embodiments of the present application is implemented as follows:

[0024] An embodiment of the present application provides a method for decoding an inter - predicted video block of a video stream. The method includes:

[0025] Receiving the video stream;

[0026] Determining that a motion vector difference (MVD) between a motion vector and a reference motion vector associated with the inter - predicted video block is written into the video stream, where the reference motion vector corresponds to a reference picture in only one of reference frame list 0 and reference frame list 1, unless the MVD is written jointly for two reference pictures;

[0027] Obtaining an indication of the size range of the MVD in a plurality of predefined motion vector difference size ranges from the video stream;

[0028] Determining a pixel resolution of the MVD according to the size range of the MVD;

[0029] Identifying additional MVD information in the video stream based on the pixel resolution of the MVD, where the additional MVD information indicates an optional pixel resolution selected for the MVD;

[0030] Extracting the additional MVD information from the video stream; and

[0031] Decoding the inter - predicted video block based on the pixel resolution of the MVD, the additional MVD information, the reference motion vector, and a reference frame associated with the motion vector.

[0032] In some embodiments, an embodiment of the present application further provides an electronic device. The electronic device includes a memory for storing computer instructions and a processor in communication with the memory. Wherein, when the processor executes the computer instructions, the electronic device is configured to execute the decoding method of the embodiment of the present application.

[0033] In some embodiments, an embodiment of the present application further provides a method for storing or transmitting a video stream, where the video stream is decoded based on the decoding method of the embodiment of the present application. Description of the Drawings

[0034] The further features, properties, and various advantages of the disclosed subject matter will become more apparent through the following detailed description and the drawings, in which:

[0035] Figure 1A A schematic diagram showing an exemplary subset of intra - prediction direction modes.

[0036] Figure 1B An illustration showing an exemplary intra - prediction direction.

[0037] Figure 2 Shows a schematic diagram of a current block for motion vector prediction and candidate blocks for spatial merging around it in an example.

[0038] Figure 3 Shows a schematic diagram of a simplified block diagram of a communication system according to an example embodiment.

[0039] Figure 4 Shows a schematic diagram of a simplified block diagram of a video streaming system according to an example embodiment.

[0040] Figure 5 Shows a schematic diagram of a simplified block diagram of a video decoder according to an example embodiment.

[0041] Figure 6 Shows a schematic diagram of a simplified block diagram of a video encoder according to an example embodiment.

[0042] Figure 7 Shows a block diagram of a video encoder according to another example embodiment.

[0043] Figure 8 Shows a block diagram of a video decoder according to another example embodiment.

[0044] Figure 9 Shows a scheme for coding block partitioning according to an example embodiment of the present application.

[0045] Figure 10 Shows another scheme for coding block partitioning according to an example embodiment of the present application.

[0046] Figure 11 Shows another scheme for coding block partitioning according to an example embodiment of the present application.

[0047] Figure 12 Shows an example of partitioning a basic block into coding blocks according to an example partitioning scheme.

[0048] Figure 13 Shows an exemplary ternary partitioning scheme.

[0049] Figure 14 Shows an exemplary quadtree binary tree coding block partitioning scheme.

[0050] Figure 15 Shows a scheme for partitioning a coding block into multiple transform blocks and the coding order of these transform blocks according to an example embodiment of the present application.

[0051] Figure 16Shows another scheme for dividing an encoding block into multiple transform blocks according to an exemplary embodiment of the present application and the encoding order of these transform blocks.

[0052] Figure 17 Shows another scheme for dividing an encoding block into multiple transform blocks according to an exemplary embodiment of the present application.

[0053] Figure 18 Shows a flowchart of a method according to an exemplary embodiment of the present application.

[0054] Figure 19 Shows another flowchart of a method according to an exemplary embodiment of the present application.

[0055] Figure 20 Shows a schematic diagram of a computer system according to an exemplary embodiment of the present application. Detailed implementation manners

[0056] Throughout the specification and claims, terms may have nuanced meanings that are implied or implicit in the context beyond their explicitly stated meanings. The phrases "in one embodiment" or "in some embodiments" used herein do not necessarily refer to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" used herein do not necessarily refer to different embodiments. Similarly, the phrases "in one implementation" or "in some implementations" used herein do not necessarily refer to the same implementation, and the phrases "in another implementation" or "in other implementations" used herein do not necessarily refer to different implementations. For example, the claimed subject matter includes combinations of all or part of the exemplary embodiments / implementations.

[0057] Generally speaking, terms can be understood at least in part from their usage in context. For example, terms such as "and", "or", or "and / or" used herein can include various meanings that can depend at least in part on the context in which these terms are used. Generally, "or" if used to relate a list such as A, B, or C is intended to mean A, B, and C (used herein in the inclusive sense) as well as A, B, or C (used herein in the exclusive sense). In addition, the terms "one or more" or "at least one" used herein, depending at least in part on the context, can be used to describe any feature, structure, or property in the singular sense or can be used to describe a combination of features, structures, or properties in the plural sense. Similarly, terms such as "a", "an", or "the" can also be understood to convey singular usage or convey plural usage, which depends at least in part on the context. In addition, terms belonging to "based on" or "determined by" can be understood not necessarily to convey a set of exclusive factors but may allow for the existence of other factors that are not necessarily explicitly described, which also depends at least in part on the context.Figure 3 A simplified block diagram of a communication system (300) according to an embodiment of the present application is shown. The communication system (300) includes a plurality of terminal devices, which can communicate with each other through, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected by a network (350). In Figure 3 the example of, the first pair of terminal devices (310) and (320) can perform unidirectional data transmission. For example, the terminal device (310) can encode video data (e.g., video data of a video picture stream collected by the terminal device (310)) for transmission through the network (350) to another terminal device (320). The encoded video data is transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video picture, and display the video picture according to the recovered video data. Unidirectional data transmission can be implemented in applications such as media services.

[0058] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, which can be implemented, for example, during a video conferencing application. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., video data of a video picture stream collected by the terminal device) for transmission through the network (350) to the other of the terminal devices (330) and (340). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to recover the video picture, and display the video picture on an accessible display device according to the recovered video data.

[0059] In Figure 3In the example, the terminal devices (310), (320), (330), and (340) can be implemented as servers, personal computers, and smart phones. However, the applicability of the basic principles of the embodiments of this application is not limited thereto. The embodiments of the embodiments of this application can be implemented on desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, and / or the like. The network (350) represents any number or type of network that transmits the encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit-switched channels, packet-switched channels, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless explicitly stated herein, the architecture and topology of the network (350) may be irrelevant to the operation of the embodiments of this application.

[0060] As an example of an application for the disclosed subject matter, Figure 4 illustrates the placement of a video encoder and a video decoder in a video streaming environment. The disclosed subject matter is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0061] The video streaming system 400 may include a video capture subsystem (413), and the video capture subsystem (413) may include, for example, a video source (401) of a digital camera. The video source (401) creates an uncompressed video picture stream or image (402). In one example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. Compared with the encoded video data (404) (or encoded video bitstream), the video picture stream (402) is depicted as a thick line to emphasize the high data volume of the video picture stream. The video picture stream (302) can be processed by an electronic device (420), and the electronic device (320) includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter described in more detail below. Compared with the uncompressed video picture stream (402), the encoded video data (404) (or encoded video bitstream (404)), which is depicted as a thin line to emphasize the lower data volume, can be stored on the streaming server (405) for future use, or directly stored in a downstream video device (not shown). One or more streaming client subsystems, such as Figure 4The client subsystems (406) and client subsystem (408) therein can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an uncompressed output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or other presentation device (not depicted). The video decoder 410 can be configured to perform some or all of the various functions described in embodiments of the present application. In some streaming systems, the encoded video data (404), video data (407), and video data (409) (e.g., video bitstreams) can be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In an embodiment, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC and can also be used in other video coding standards.

[0062] It should be noted that the electronic device (420) and the electronic device (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).

[0063] Hereinafter, Figure 5 A block diagram of a video decoder (510) according to any embodiment of an embodiment of the present application is shown. The video decoder (510) can be provided in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used to replace Figure 4 the video decoder (410) in the example of

[0064] A receiver (531) may receive one or more encoded video sequences to be decoded by a video decoder (510). In the same or another embodiment, one encoded video sequence may be decoded at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with a plurality of video frames or pictures. The encoded video sequences may be received from a channel (501), which may be a hardware / software link leading to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective processing circuits (not depicted). The receiver (531) may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (515) may be provided between the receiver (531) and an entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be located external to and separate from the video decoder (510) (not depicted). In still other applications, a buffer memory (not depicted) may be provided external to the video decoder (510) to, for example, prevent network jitter, and another additional buffer memory (515) may be provided inside the video decoder (510) to, for example, handle the playback timing. When the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be possible not to configure the buffer memory (515), or the buffer memory may be made smaller. For use on a service packet network such as the Internet, a buffer memory (515) of sufficient size may be required, and the size of the buffer memory (515) may be relatively large. Such a buffer memory may be implemented with an adaptive size and may be implemented at least partially in an operating system or a similar element (not depicted) external to the video decoder (510).

[0065] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequences. The categories of these symbols include information for managing the operation of the video decoder (510), and potential information for controlling a rendering device such as a display (512) (e.g., a display screen), which may or may not be an integral part of the electronic device (530), but may be coupled to the electronic device (530), such as Figure 5As shown. The control information for presenting the device may be in the form of Supplemental Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set segment (not depicted). The parser (520) may perform parsing / entropy decoding on the encoded video sequence received by the parser (520). The entropy coding of the encoded video sequence may be performed according to video coding techniques or standards and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc.

[0066] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0067] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) may involve multiple different processing or functional units. Which units are involved and the way they are involved may be controlled by the parser (520) through subgroup control information parsed from the encoded video sequence. For simplicity, such subgroup control information flows between the parser (520) and multiple processing or functional units below are not depicted.

[0068] In addition to the functional blocks already mentioned, the video decoder (510) may be conceptually divided into several functional units as described below. In practical implementations operating under commercial constraints, many of these functional units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, in the following embodiments of the present application, it is conceptually divided into multiple functional units.

[0069] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive, from the parser (520), quantized transform coefficients as symbols (521) and control information including information indicating which type of inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output a block including sample values, and the sample values may be input into an aggregator (555).

[0070] In some cases, the output samples of the scaler / inverse transform (551) may belong to an intra-coded block; that is, a block that does not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may use surrounding block information that has been reconstructed and stored in the current picture buffer (558) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (555) may add, on a per-sample basis, the prediction information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0071] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (553) may access a reference picture memory (557) to extract samples for inter-picture prediction. After motion-compensating the extracted samples according to the symbols (521) belonging to the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of unit 551 may be referred to as residual samples or a residual signal), thereby generating output sample information. The extraction of prediction samples by the motion compensation prediction unit (553) from an address within the reference picture memory (557) may be controlled by a motion vector, and the motion vector may be provided to the motion compensation prediction unit (553) for use in the form of symbols (521), and the symbols (521) may have, for example, an X component, a Y component (offset), and a reference picture component (temporal). Motion compensation may also include interpolation of the sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, and may also be associated with a motion vector prediction mechanism, etc.

[0072] The output samples of the aggregator (555) can be adopted by various loop filtering techniques in the loop filter unit (556). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and the parameters can be used as symbols (521) from the parser (520) for the loop filter unit (556). However, in other embodiments, video compression techniques can also respond to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values. Multiple types of loop filters can be included in various orders as part of the loop filter unit 556, as will be described in further detail below.

[0073] The output of the loop filter unit (556) can be a sample stream that can be output to the rendering device (512) and stored in the reference picture memory (557) for future inter-picture prediction.

[0074] Once fully reconstructed, some encoded pictures can be used as reference pictures for future inter-picture prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.

[0075] The video decoder (510) can perform decoding operations according to a predetermined video compression technique adopted in a standard such as the ITU-T H.265 recommendation. In the sense that an encoded video sequence conforms to the syntax specified by the video compression technique or standard used, the encoded video sequence can conform to the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, a profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under that profile. To conform to the standard, the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0076] In some exemplary embodiments, a receiver (531) may receive additional (redundant) data along with the encoded video. This additional data may be part of the encoded video sequence. The additional data may be used by a video decoder (510) to decode the data appropriately and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0077] Figure 6 is a block diagram of a video encoder (603) according to an exemplary embodiment disclosed in the present application. The video encoder (603) may be included in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 the video encoder (403) in the example of

[0078] The video encoder (603) may receive video samples from a video source (601) (not Figure 6 part of the electronic device (620) in the example of

[0079] The video source (601) may collect video images to be encoded by the video encoder (603). In another example, the video source (601) may be implemented as part of the electronic device (620). The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603). The digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 YCrCb, RGB, XYZ,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) may be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera that collects local image information as a video sequence. The video data may be provided as a plurality of individual pictures or images that, when viewed in sequence, are given motion. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0080] According to some exemplary embodiments, a video encoder (603) may encode and compress pictures of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed constitutes a function of a controller (650). In some embodiments, the controller (650) may be functionally coupled to other functional units as described below and control the other functional units. For simplicity, couplings are not depicted in the figures. Parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be used for other suitable functions that relate to the video encoder (603) optimized for a certain system design.

[0081] In some exemplary embodiments, the video encoder (603) may be configured to operate in an encoding loop. As an over-simplified description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and an (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to how a (remote) decoder creates sample data, even though the embedded decoder 633 processes the encoded video stream through the source encoder 630 without entropy coding (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols in the entropy coding and the encoded video bitstream can be lossless). The reconstructed sample stream (sample data) is input into a reference picture memory (634). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is used to improve the encoding quality.

[0082] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder that has been described in detail above in connection with Figure 5 the video decoder (510). However, briefly referring additionally to Figure 5, since the symbols are available and the entropy encoder (645) and the parser (520) are capable of losslessly encoding / decoding the symbols into the encoded video sequence, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633) and in the encoder.

[0083] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding that may only exist in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter sometimes focuses on decoder operations, which cooperate with the decoding part of the encoder. Thus, the description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. Only some areas or aspects of the encoder are described in more detail below.

[0084] During operation, in some exemplary implementations, the source encoder (630) may perform motion-compensated predictive coding that predictively encodes an input picture by referring to one or more previously encoded pictures designated as "reference pictures" in the video sequence. In this way, the encoding engine (632) encodes the differences (or residuals) in the color channels between the pixel blocks of the input picture and the pixel blocks of the reference picture, which can be selected as the prediction reference for the input picture. The term "residual" and its adjective form "residual" can be used interchangeably.

[0085] The local video decoder (633) may decode the encoded video data of the picture that can be designated as a reference picture based on the symbols created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the encoded video data can be decoded on a video decoder ( Figure 6 not shown), the reconstructed video sequence may generally be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that can be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture that will be obtained by a remote video decoder.

[0086] Predictor (635) can perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, predictor (635) can search in the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata that can be used as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. Predictor (635) can operate on a per-pixel-block basis for sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by predictor (635), it can be determined that the input picture may have a prediction reference taken from multiple reference pictures stored in reference picture memory (634).

[0087] Controller (650) can manage the encoding operations of source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0088] The outputs of all the above functional units can be entropy encoded in entropy encoder (645). Entropy encoder (645) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0089] Transmitter (640) can buffer the encoded video sequence created by entropy encoder (645) to prepare for transmission over communication channel (660), which can be a hardware / software link to a storage device that will store the encoded video data. Transmitter (640) can merge the encoded video data from video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0090] Controller (650) can manage the operations of video encoder (603). During encoding, controller (650) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can generally be assigned to any of the following picture types:

[0091] Intra picture (I picture), which can be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are aware of the variants of I pictures and their corresponding applications and characteristics.

[0092] A predictive picture (P picture) that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and reference index to predict the sample values of each block.

[0093] A bi-predictive picture (B picture) that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0094] A source picture can generally be spatially subdivided into multiple sample coding blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the coding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be non-predictively encoded, or the block can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be predictively encoded with reference to a previously encoded reference picture through spatial prediction or through temporal prediction. Blocks of a B picture can be predictively encoded with reference to one or two previously encoded reference pictures through spatial prediction or through temporal prediction. For other purposes, a source picture or an intermediate processed picture can be subdivided into other types of blocks. The partitioning of coding blocks and other types of blocks may or may not follow the same manner, as described in further detail below.

[0095] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0096] In one exemplary embodiment, the transmitter (640) can transmit additional data when transmitting the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, SEI messages, VUI parameter set fragments, etc.

[0097] The captured video can be used as a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (usually abbreviated as intra prediction) utilizes the spatial correlation in a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. For example, a specific picture being encoded / decoded can be divided into blocks, and the specific picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0098] In some exemplary embodiments, bidirectional prediction techniques can be used for inter-picture prediction. According to such bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture that are before the current picture in the video in decoding order (but may be past and future respectively in display order). The block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be jointly predicted by a combination of the first reference block and the second reference block.

[0099] In addition, merge mode techniques can be used for inter-picture prediction to improve encoding efficiency.

[0100] According to some exemplary embodiments of the embodiments of the present application, predictions such as inter-picture prediction and intra-picture prediction are performed on a per-block basis. For example, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression. The CTUs in a picture can have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU can include three parallel coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or four 32×32 pixel CUs. Each of one or more 32×32 blocks can be further divided into four 16×16 pixel CUs. In some exemplary embodiments, each CU can be analyzed during encoding to determine the prediction type for the CU among various prediction types, such as inter-prediction type or intra-prediction type. Depending on the temporal and / or spatial predictability, a CU can be divided into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed on a per-prediction block basis. The division of a CU into PUs (or PBs of different color channels) can be performed in various spatial patterns. For example, a luminance or chrominance PB can include a matrix of values for samples (e.g., luminance values), and the samples can be, for example, 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 samples, etc.

[0101] Figure 7 FIG. shows a video encoder (703) according to another exemplary embodiment of the embodiments of the present application. The video encoder (703) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence, and encode the processing block into an encoded picture that is part of an encoded video sequence. The exemplary video encoder (703) can be used to replace Figure 4 the video encoder (403) in the example of

[0102] For example, a video encoder (703) receives a matrix of sample values for processing a block, such as a prediction block of 8×8 samples. Then, the video encoder (703) uses, for example, rate-distortion optimization (RDO) to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to best encode the processing block. When it is determined to encode the processing block in the intra mode, the video encoder (703) may use an intra prediction technique to encode the processing block into an encoded picture; and when it is determined to encode the processing block in the inter mode or the bi-prediction mode, the video encoder (703) may use an inter prediction or a bi-prediction technique, respectively, to encode the processing block into an encoded picture. In some exemplary embodiments, a merge mode may be used as an inter-picture prediction sub-mode, where a motion vector is derived from one or more motion vector predictors without relying on encoded motion vector components external to the predictor. In some other exemplary embodiments, there may be motion vector components applicable to the subject block. Thus, the video encoder (703) may include components not explicitly shown in Figure 7 such as a mode decision module for determining the prediction mode of the processing block.

[0103] In Figure 7 the example of Figure 7 the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in the exemplary arrangement of

[0104] The inter encoder (730) is configured to receive samples of a current block (e.g., the processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter prediction information (e.g., a description of redundancy information, a motion vector, merge mode information according to inter coding techniques), and calculate an inter prediction result (e.g., a predicted block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded using a decoding unit 633 based on encoded video information, and the decoding unit 633 is embedded in Figure 6 the exemplary encoder 620 of Figure 7 which is shown as

[0105] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture, generate quantized coefficients after transformation, and in some cases also generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). The intra encoder (722) may calculate an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.

[0106] The general controller (721) may be configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general controller (721) determines a prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is the intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream; and when the prediction mode of the block is the inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0107] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and a block prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transform coefficients. For example, the residual encoder (724) may be configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various exemplary embodiments, the video encoder (703) further includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transformation and generate decoded residual data. The decoded residual data may be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) may generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) may generate a decoded block based on the decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture may be buffered in a memory circuit (not shown) and used as a reference picture.

[0108] The entropy encoder (725) can be configured to format the bitstream to include the encoded blocks and perform entropy encoding. The entropy encoder (725) is configured to include various information in the bitstream. For example, the entropy encoder (725) can be configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. When encoding a block in the merge sub-mode of the inter mode or the bi-prediction mode, there may be no residual information.

[0109] Figure 8 FIG. shows an exemplary video decoder (810) according to another embodiment of an embodiment of the present application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In one example, the video decoder (810) can be used instead of Figure 4 the video decoder (410) in the example of

[0110] In Figure 8 the example of Figure 8 the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as shown in the exemplary arrangement of

[0111] The entropy decoder (871) can be used to reconstruct certain symbols from the encoded picture, where these symbols represent the syntax elements that make up the encoded picture. Such symbols can include, for example, the mode for encoding a block (e.g., intra mode, inter mode, bi-prediction mode, merge sub-mode, or another sub-mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify certain samples or metadata for use by the intra decoder (872) or the inter decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, etc. In one example, when the prediction mode is the inter or bi-prediction mode, the inter prediction information is provided to the inter decoder (880); and when the prediction type is the intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).

[0112] The inter decoder (880) can be configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.

[0113] The intra decoder (872) can be configured to receive the intra prediction information and generate a prediction result based on the intra prediction information.

[0114] The residual decoder (873) can be configured to perform inverse quantization to extract the dequantized transform coefficients, and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) can also use certain control information (used to include Quantizer Parameter (QP)), which can be provided by the entropy decoder (871) (the data path is not depicted as this is merely low-volume control information).

[0115] The reconstruction module (874) can be configured to combine, in the spatial domain, the residual output by the residual decoder (873) with the prediction result (which can be output by the inter-frame prediction module or the intra-frame prediction module, as the case may be) to form a reconstructed block, and the reconstructed block forms part of a reconstructed picture, and the reconstructed picture is part of a reconstructed video. It should be noted that other suitable operations such as deblocking operations can also be performed to improve the visual quality.

[0116] It should be noted that any suitable technology can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In some exemplary embodiments, one or more integrated circuits can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (403), (603), and (603) and the video decoders (410), (510), and (810).

[0117] Turning to block partitioning for encoding and decoding, the general partitioning can start from a basic block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After splitting or partitioning the basic block according to any of the example partitioning processes described below or other processes or a combination thereof, a final set of partitions or coding blocks can be obtained. Each of these partitions can be at one of the different partitioning levels in the partitioning hierarchy and can have various shapes. Each partition can be referred to as a coding block (CB). For the various example partitioning implementations described further below, each generated CB can be of any allowed size and partitioning level. Such partitions are called coding blocks because they can form units for which some basic encoding / decoding can be performed and for which encoding / decoding parameters can be optimized, determined, and written in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure of the tree. Coding blocks can be luminance coding blocks or chrominance coding blocks. The CB tree structure for each color can be referred to as a coding block tree (CBT).

[0118] The coding blocks of all color channels can be collectively referred to as coding units (CUs). The hierarchical structures of all color channels can be collectively referred to as coding tree units (CTUs). The partitioning patterns or structures of different color channels in a CTU may be the same or different.

[0119] In some implementations, the partitioning tree schemes or structures for the luminance and chrominance channels may not need to be the same. In other words, the luminance and chrominance channels can have separate coding tree structures or patterns. Additionally, whether the luminance and chrominance channels use the same or different coding partitioning tree structures and the actual coding partitioning tree structure to be used can depend on whether the slice being encoded is a P slice, a B slice, or an I slice. For example, for an I slice, the chrominance channel and the luminance channel can have separate coding partitioning tree structures or coding partitioning tree structure patterns, while for a P slice or a B slice, the luminance and chrominance channels can share the same coding partitioning tree scheme. When applying separate coding partitioning tree structures or patterns, the luminance channel can be divided into CBs by one coding partitioning tree structure and the chrominance channel can be divided into chrominance CBs by another coding partitioning tree structure.

[0120] In some exemplary implementations, a predefined partitioning pattern can be applied to the basic block. As Figure 9As shown, an exemplary 4-way partitioning tree can start from a first predefined level (e.g., the 64×64 block level or other size, as the basic block size), and the basic block can be hierarchically partitioned down to a predefined lowest level (e.g., the 4×4 level). For example, the basic block can have four predefined partitioning options or patterns indicated by 902, 904, 906, and 908, where the partition designated as R is allowed for recursive partitioning, such that the same partitioning option as indicated in Figure 9 can be repeated at a lower scale until the lowest level (e.g., the 4×4 level). In some implementations, additional restrictions can be applied to the Figure 9 partitioning scheme. In the Figure 9 implementation, rectangular partitioning (e.g., 1:2 / 2:1 rectangular partitioning) can be allowed, but these rectangular partitions are not allowed to be recursive, while square partitions are allowed to be recursive. If needed, the recursive partitioning according to Figure 9 generates a final set of coded blocks. The coding tree depth can be further defined to indicate the depth of the split from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 64×64 blocks) can be set to 0, and after the root block is further split once according to Figure 9 , the coding tree depth increases by 1. For the above scheme, the maximum or deepest level from the 64×64 basic block to the 4×4 minimum partition will be 4 (starting from level 0). This partitioning scheme can be applied to one or more color channels. Each color channel can be partitioned independently according to the Figure 9 scheme (e.g., the partitioning pattern or option in the predefined pattern can be independently determined for each color channel at each hierarchical level). Optionally, two or more color channels can share the Figure 9 same hierarchical pattern tree (e.g., the same partitioning pattern or option in the predefined pattern can be selected for two or more color channels at each hierarchical level).

[0121] Figure 10 shows another exemplary predefined partitioning pattern that allows recursive partitioning to form a partitioning tree. As shown in Figure 10 , an exemplary 10-way partitioning structure or pattern can be predefined. The root block can start from a predefined level (e.g., starting from the 128×128 level or the 64×64 level basic block). Figure 10 The example partitioning structure includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. The partitioning types of the three sub-partitions 1002, 1004, 1006, and 1008 indicated in the second row of Figure 10 can be referred to as "T-shaped" partitions. The "T-shaped" partitions 1002, 1004, 1006, and 1008 can be respectively referred to as left T-shaped, upper T-shaped, right T-shaped, and lower T-shaped. In some example implementations, further subdivision is not allowedFigure 10 any one of the rectangular partitions. The coding tree depth can be further defined to indicate the depth of the split from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 128×128 block) can be set to 0, and after the root block is further split once according to Figure 10 it, the coding tree depth increases by 1. In some implementations, only the full square partitions in 1010 are allowed to be recursively partitioned to the next level of the partition tree according to Figure 10 the pattern. In other words, for the square partitions within the T-shaped patterns 1002, 1004, 1006, and 1008, recursive partitioning may not be allowed. If needed, the recursive partitioning process according to Figure 10 will generate a final set of coded blocks. This scheme can be applied to one or more color channels. In some implementations, more flexibility can be added when using partitions below the 8×8 level. For example, 2×2 chrominance inter prediction can be used in some cases.

[0122] In some other exemplary implementations for coded block partitioning, a quadtree structure can be used to split a basic block or an intermediate block into quadtree partitions. This quadtree splitting can be applied hierarchically and recursively to any square partition. Whether the basic block or intermediate block or partition is further quadtree split can be adapted to various local characteristics of the basic block or intermediate block or partition. The quadtree partitioning at the picture boundary can be further adjusted. For example, an implicit quadtree split can be performed at the picture boundary so that the block will remain quadtree split until its size fits the picture boundary.

[0123] In some other example implementations, hierarchical binary partitioning from a basic block can be used. For such a scheme, a basic block or an intermediate-level block can be divided into two partitions. The binary partitioning can be horizontal or vertical. For example, a horizontal binary partitioning can split a basic block or an intermediate block into equal right and left partitions. Similarly, a vertical binary partitioning can split a basic block or an intermediate block into equal upper and lower partitions. This binary partitioning can be hierarchical and recursive. It can be decided on each basic block or intermediate block whether the binary partitioning scheme should continue, and if the scheme continues, whether to use horizontal or vertical binary partitioning. In some implementations, further partitioning may stop at a predefined minimum partition size (in one or two dimensions). Optionally, once a predefined partition level or depth from the basic block is reached, further partitioning can be stopped. In some implementations, the aspect ratio of the partition may be restricted. For example, the aspect ratio of the partition can be not less than 1:4 (or greater than 4:1). Thus, a vertical bar partition with an aspect ratio of 4:1 can only be further vertically binary partitioned into upper and lower partitions, each with an aspect ratio of 2:1.

[0124] In still some other examples, such as Figure 13 shown, a ternary partitioning scheme can be used to partition a basic block or any intermediate block. The ternary pattern can be implemented vertically as shown in 1302 of Figure 13 , or horizontally as shown in 1304 of Figure 13 . Although the exemplary split ratios (vertically or horizontally) in Figure 13 are shown as 1:2:1, other ratios can also be predefined. In some implementations, two or more different ratios can be predefined. Such a ternary partitioning scheme can be used to complement a quadtree or binary partitioning structure because this ternary tree partitioning can capture an object located at the center of a block in a continuous partition, while a quadtree and a binary tree always partition along the block center, so the object is divided into separate partitions. In some implementations, the width and height of the example ternary tree partitioning are always powers of 2 to avoid additional transformations.

[0125] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree and binary partitioning schemes can be combined to partition a basic block into a quadtree - binary - tree (QTBT) structure. In such a scheme, the basic block or intermediate block / partition can be either quadtree - split or binary - split, if specified, depending on a set of predefined conditions. Figure 14 A specific example is shown. In the example of Figure 14 , a basic block is first quadtree - divided into four partitions as shown in 1402, 1404, 1406, and 1408. Thereafter, each generated partition is either quadtree - split into four further partitions (e.g., 1408) at the next level, or binary - split into two further partitions (e.g., horizontally or vertically, e.g., 1402 or 1406, both are symmetric), or not split (e.g., 1404). For square partitions, binary or quadtree splitting can be recursively allowed, as shown in the overall partitioning pattern example of 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree splitting and dashed lines represent binary tree splitting. A flag can be used for each binary - split node (non - leaf binary partition) to indicate whether the binary split is horizontal or vertical. For example, as shown in 1420 and consistent with the partitioning structure of 1410, the flag "0" can represent a horizontal binary split, and the flag "1" can represent a vertical binary split. For quadtree - split partitions, no indication of the split type is needed because quadtree splitting always splits a block or partition horizontally and vertically to produce 4 sub - blocks / partitions of equal size. In some implementations, the flag "1" can represent a horizontal binary split, and the flag "0" can represent a vertical binary split.

[0126] In some example implementations of QTBT, the quadtree and binary splitting rule sets can be represented by the following predefined parameters and their associated corresponding functions:

[0127] CTU size: The size of the root node of the quadtree (the size of the basic block)

[0128] MinQTSize: The minimum allowed size of the quadtree leaf nodes

[0129] MaxBTSize: The maximum allowed size of the binary tree root node

[0130] MaxBTDepth: The maximum allowed depth of the binary tree

[0131] MaxBTSize: The minimum allowed size of the binary tree leaf nodes

[0132] In some example implementations of the QTBT partitioning structure, the CTU size can be set to 128×128 luminance samples and two corresponding 64×64 chrominance sample blocks (when considering and using the example chrominance subsampling), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. The quadtree partitioning can be applied to the CTU first to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from its minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If a node is 128×128, it will not be split by the binary tree first because its size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes that do not exceed MaxBTSize can be partitioned by the binary tree. In Figure 14 the example, the basic block is 128×128. According to the predefined rule set, the basic block can only be split by the quadtree. The partitioning depth of the basic block is 0. Each of the four resulting partitions is 64×64, which does not exceed MaxBTSize and can be further split by the quadtree or binary tree at level 1. This process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When the width of the binary tree node equals MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when the height of the binary tree node equals MinBTSize, further vertical splitting is no longer considered.

[0133] In some example implementations, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure or separate QTBT structures for luminance and chrominance. For example, for P slices and B slices, the luminance and chrominance CTBs in a CTU can share the same QTBT structure. However, for I slices, the luminance CTB can be divided into CBs by a QTBT structure, and the chrominance CTB can be divided into chrominance CBs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice. For example, an I slice can consist of coding blocks of the luminance component or coding blocks of the two chrominance components, and a CU in a P slice or B slice can consist of coding blocks of all three color components.

[0134] In some other implementations, the QTBT scheme can be supplemented with the above-mentioned ternary scheme. Such an implementation can be referred to as a multi-type-tree (MTT) structure. For example, in addition to the binary splitting of nodes, one of the ternary partitioning patterns can be selected. Figure 13 In some implementations, only square nodes can be ternary split. An additional flag can be used to indicate whether the ternary partitioning is horizontal or vertical.

[0135] The design of two-level or multi-level trees, such as the QTBT implementation and the QTBT implementation supplemented by ternary splitting, may be mainly for complexity reduction. Theoretically, the complexity of traversing the tree is T D , where T represents the number of splitting types and D is the depth of the tree. A trade-off can be made by using multiple types (T) to reduce the depth (D) simultaneously.

[0136] In some implementations, the CB can be further divided. For example, for the purpose of intra-frame or inter-frame prediction during the encoding and decoding processes, the CB can be further divided into multiple prediction blocks. In other words, the CB can be further divided into different sub-partitions where separate prediction decisions / configurations can be made. In parallel, for the purpose of depicting the level of performing the transform or inverse transform of video data, the CB can be further divided into multiple transform blocks (TBs). The scheme for dividing the CB into PBs and TBs can be the same or different. For example, each partitioning scheme can use its own process to perform, e.g., based on various characteristics of the video data. In some example implementations, the PB and TB partitioning schemes may be independent. In some other example implementations, the partitioning schemes and boundaries of the PBs and TBs may be interrelated. For example, in some implementations, the TB can be divided after the PB partitioning. In particular, each PB is determined after the coding block partitioning and can then be further divided into one or more TBs. For example, in some implementations, a PB can be divided into one, two, four, or other numbers of TBs.

[0137] In some implementations, in order to divide the basic block into coding blocks and further into prediction blocks and / or transform blocks, the luma channel and the chroma channel may be processed differently. For example, in some implementations, for the luma channel, it may be allowed to divide the coding block into prediction blocks and / or transform blocks, while for the chroma channel, it may not be allowed to divide the coding block into prediction blocks and / or transform blocks. Therefore, in such an implementation, the transformation and / or prediction of the luma block may be performed only at the coding block level. For another example, the minimum transform block size of the luma channel and the chroma channel may be different, for example, the coding block of the luma channel may be allowed to be divided into smaller transform and / or prediction blocks than the chroma channel. For yet another example, the maximum depth of dividing the coding block into transform blocks and / or prediction blocks may be different between the luma channel and the chroma channel, for example, it may be allowed to divide the coding block of the luma channel into deeper transform blocks and / or prediction blocks than the chroma channel. For a specific example, the luma coding block may be partitioned into transform blocks of multiple sizes, which may be represented by recursive partitioning with up to 2 levels, and may allow transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4, and transform block sizes from 4×4 to 64×64. However, for chroma blocks, only the largest possible transform block specified for the luma block may be allowed.

[0138] In some example implementations of partitioning a coding block into PBs, the depth, shape, and / or other characteristics of the PB partition may depend on whether the PB is intra-coded or inter-coded.

[0139] The partitioning of the coding block (or prediction block) into transform blocks can be implemented in various example schemes, including but not limited to recursively or non-recursively quadtree partitioning and predetermined pattern partitioning, and additionally considering transform blocks at the boundaries of the coding block or prediction block. In general, the resulting transform blocks may be at different partitioning levels, may not have the same size, and may not need to be square in shape (e.g., they may be rectangular with some allowed size and aspect ratio). Figure 15 , Figure 16 and Figure 17 Other examples are described in more detail.

[0140] However, in some other implementations, the CBs obtained via any of the above partitioning schemes can be used as the basic or minimum coding blocks for prediction and / or transformation. In other words, for the purpose of performing inter-frame prediction / intra-frame prediction and / or transformation, no further splitting is performed. For example, the CBs obtained according to the above QTBT scheme can be directly used as the units for performing prediction. Specifically, this QTBT structure eliminates the concept of multiple partitioning types, i.e., it eliminates the separation of CUs, PUs, and TUs, and provides greater flexibility for the CU / CB partitioning shapes as described above. In this QTBT block structure, the CU / CB can have a square or rectangular shape. The leaf nodes of this QTBT are used as the units for prediction and transformation processing without any further partitioning. This means that in this exemplary QTBT coding block structure, the CUs, PUs, and TUs have the same block size.

[0141] The above various CB partitioning schemes and the further partitioning of CBs into PBs and / or TBs (excluding PB / TB partitioning) can be combined in any way. The following specific implementations are provided as non-limiting examples.

[0142] Specific example implementations of coding block and transform block partitioning are described below. In such an example implementation, a recursive quadtree split or the above predefined split patterns (e.g., Figure 9 and Figure 10 the patterns in) can be used to split a basic block into coding blocks. At each level, whether a particular partition should continue to be further split by a quadtree can be determined by local video data characteristics. The resulting CBs can be at various quadtree split levels and have various sizes. The decision on whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode a picture region can be made at the CB level (or the CU level, for all three color channels). Each CB can be further split into one, two, four, or other numbers of PBs according to a predefined PB split type. Within a PB, the same prediction processing can be applied, and relevant information can be sent to the decoder based on the PB. After obtaining the residual blocks by applying the prediction process based on the PB split type, the CB can be partitioned into TBs according to another quadtree structure similar to the coding tree of the CB. In this specific implementation, the CB or TB can be, but is not limited to, square. Additionally, in this specific example, for inter-frame prediction, the PB can be square or rectangular, while for intra-frame prediction, the PB can be only square. The coding block can be divided into, for example, four square TBs. Each TB can be further recursively split (using quadtree split) into smaller TBs, called Residual Quadtree (RQT).

[0143] Another example implementation for dividing a basic block into CBs, PBs, and / or TBs is further described below. For example, instead of using multiple partitioning unit types such as Figure 9 or Figure 10 shown in, a quadtree with an embedded multi-type tree having binary and ternary split structures (e.g., QTBT or QTBT with ternary split as described above) can be used. The separation of CBs, PBs, and TBs (i.e., dividing CBs into PBs and / or TBs, and dividing PBs into TBs) may be waived unless the size of the CBs is too large for the maximum transform length, and such CBs may need to be further split. This example partitioning scheme can be designed to support greater flexibility in the CB partitioning shape, such that both prediction and transformation can be performed at the CB level without further partitioning. In such an encoding tree structure, a CB can have a square or rectangular shape. Specifically, a Coding Tree Block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by an embedded multi-type tree structure. Figure 11 An example of using an embedded multi-type tree structure with binary or ternary split is shown in. Specifically, Figure 11 the exemplary multi-type tree structure includes four split types, namely vertical binary split (SPLIT_BT_VER) (1102), horizontal binary split (SPLIT_BT_HOR) (1104), vertical ternary split (SPLIT_TT_VER) (1106), and horizontal ternary split (SPLIT_TT_HOR) (1108). Then, the CBs correspond to the leaves of the multi-type tree. In this example implementation, unless the CB is too large for the maximum transform length, this split is used for both prediction and transform processing without any further partitioning. This means that in most cases, in a quadtree with an embedded multi-type tree coding block structure, the CBs, PBs, and TBs have the same block size. Exceptions occur when the maximum supported transform length is less than the width or height of the color components of the CB. In some implementations, in addition to binary or ternary split, Figure 11 the embedded mode can also include quadtree split.

[0144] Figure 12 A specific example of a quadtree with an embedded multi-type tree coding block structure for block partitioning (including quadtree, binary tree, and ternary split options) of a basic block is shown in. More specifically, Figure 12 shows that the basic block 1200 is split into four square partitions 1202, 1204, 1206, and 1208 by a quadtree. For each quadtree split partition, it is decided to further use Figure 11 the multi-type tree structure and quadtree for further splitting. In Figure 12In the example, partition 1204 is not further divided. Partitions 1202 and 1208 each adopt another quadtree division. For partition 1202, the upper-left, upper-right, lower-left, and lower-right partitions of the secondary quadtree division respectively adopt a quadtree, Figure 11 the horizontal binary division 1104, no division, and Figure 11 the three-level division of the horizontal ternary division 1108. Partition 1208 adopts another quadtree division. The upper-left, upper-right, lower-left, and lower-right partitions of the secondary quadtree division respectively adopt Figure 11 the vertical ternary division 1106, no division, no division, Figure 11 the three-level division of the horizontal binary division 1104. The two sub-partitions of the third-level upper-left division of 1208 are further divided respectively according to Figure 11 the horizontal binary division 1104 and the horizontal ternary division 1108. Partition 1206 adopts the second-level division mode following Figure 11 the vertical binary division 1102, is divided into two partitions, and then the third-level division is performed according to Figure 11 the horizontal ternary division 1108 and the vertical binary division 1102. According to Figure 11 the horizontal binary division 1104, the fourth-level division is further applied to one of them.

[0145] For the above specific example, the maximum luminance transform size can be 64×64, and the supported maximum chrominance transform size can be different from the luminance, such as 32×32. Even if Figure 12 the above example CBs in

[0146] are generally not further divided into smaller PBs and / or TBs, when the width or height of the luminance coding block or chrominance coding block is greater than the maximum transform width or height, the luminance coding block or chrominance coding block can be automatically divided in the horizontal and / or vertical directions to meet the transform size limit in that direction.

[0147] When a coding block is further divided into multiple transform blocks, the transform blocks therein can be sorted in the bitstream in various orders or scan patterns. Example implementations of dividing a coding block or a prediction block into transform blocks and the coding order of the transform blocks will be described in further detail below. In some example implementations, as described above, the transform partitioning can support transform blocks of multiple shapes, such as 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1, and the transform block sizes range from 4×4 to 64×64. In some implementations, if the coding block is less than or equal to 64×64, the transform block partitioning can be applied only to the luminance component, such that for the chrominance blocks, the transform block size is the same as the coding block size. Otherwise, if the width or height of the coding block is greater than 64, both the luminance and chrominance coding blocks can be implicitly divided into multiples of min(W, 64)×min(H, 64) and min(W, 32)×min(H, 32) transform blocks, respectively.

[0148] In some example implementations of the transform block partitioning, for intra-coded blocks and inter-coded blocks, the coding block can be further divided into multiple transform blocks, and the partitioning depth can reach a predetermined number of levels (e.g., 2 levels). The transform block partitioning depth and size can be related. For some example implementations, the mapping from the transform block size at the current depth to the transform block size at the next depth is shown in Table 1 below.

[0149] Table 1 Transform Block Partitioning Size Settings

[0150]

[0151]

[0152] Based on the example mapping in Table 1, for a 1:1 square block, the next-level transform partitioning can create four 1:1 square sub-transform blocks. The transform block partitioning can stop, for example, at 4×4. Thus, the transform block size at the current depth of 4×4 corresponds to the same size of 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next-level transform partitioning can create two 1:1 square sub-transform blocks, and for a 1:4 / 4:1 non-square block, the next-level transform partitioning can create two 1:2 / 2:1 sub-transform blocks.

[0153] In some example implementations, for the luminance component of an intra-coded block, additional restrictions can be imposed with respect to the transform block partitioning. For example, for each level of the transform partitioning, all sub-transform blocks can be restricted to have equal sizes. For example, for a 32×16 coding block, the first-level transform partitioning creates two 16×16 sub-transform blocks, and the second-level transform partitioning creates eight 8×8 sub-transform blocks. In other words, the second-level partitioning must be applied to all first-level sub-blocks to keep the transform unit sizes equal. Figure 15An example of dividing an intra-coded square block into transform blocks according to Table 1 is shown, as well as the coding order indicated by the arrows. Specifically, 1502 shows the square coding block. In 1504, the coding block is divided into 4 transform blocks of equal size according to Table 1 at the first level of division, and the coding order of these 4 transform blocks is indicated by the arrows. In 1506, all the blocks of equal size at the first level are divided into 16 transform blocks of equal size according to Table 1 at the second level of division, and the coding order of these 16 transform blocks is indicated by the arrows.

[0154] In some example implementations, the above restrictions on intra-coding may not apply to the luminance component of an inter-coded block. For example, after the first-level transform split, any one of the sub-transform blocks can be further independently split into more than one level. Thus, the resulting transform blocks can be of the same size or of different sizes. Figure 16 An example of dividing an inter-coded block into multiple transform blocks with their coding order is shown. In Figure 16 the example of, according to Table 1, the inter-coded block 1602 is divided into transform blocks of two levels. At the first level, the inter-coded block is split into four transform blocks of equal size. Then, only one (not all) of the four transform blocks is further split into four sub-transform blocks, resulting in a total of 7 transform blocks with two different sizes, as shown in 1604. The example coding order of these 7 transform blocks is indicated by the Figure 16 arrows in 1604 of.

[0155] In some example implementations, some additional restrictions on transform blocks can be applied to one or more chrominance components. For example, for one or more chrominance components, the transform block size can be as large as the coding block size, but not less than a predefined size, such as 8×8.

[0156] In some other example implementations, for coding blocks with a width (W) or height (H) greater than 64, the luminance and chrominance coding blocks can be implicitly split into multiples of min(W, 64)×min(H, 64) and min(W, 32)×min(H, 32) transform units, respectively. Here, in the embodiments of the present application, "min(a, b)" can return the smaller value between a and b.

[0157] Figure 17 A further alternative example scheme for dividing a coding block or a prediction block into multiple transform blocks is shown. As Figure 17 shown, instead of using recursive transform division, a set of predefined division types can be applied to the coding block according to the transform type of the coding block. In Figure 17In the specific examples shown, one of six example partition types can be applied to split coding blocks into various numbers of transform blocks. This scheme for generating transform block partitions can be applied to coding blocks or prediction blocks.

[0158] More specifically, Figure 17 the partitioning scheme provides up to six example partition types for any given transform type (the transform type refers to, for example, the type of the main transform, such as ADST, etc.). In this scheme, a transform partition type can be assigned to each coding block or prediction block based on, for example, rate-distortion cost. In an example, the transform partition type assigned to a coding block or prediction block can be determined based on the transform type of the coding block or prediction block. A specific transform partition type can correspond to a transform block split size and pattern, as Figure 17 shown by the six transform partition types shown in

[0159] ·PARTITION_NONE (no partition): Assign a transform size equal to the block size.

[0160] ·PARTITION_SPLIT (split partition): Assign a transform size whose width is one-half of the block size width and height is one-half of the block size height.

[0161] ·PARTITION_HORZ (horizontal partition): Assign a transform size whose width is the same as the block size width and height is one-half of the block size height.

[0162] ·PARTITION_VERT (vertical partition): Assign a transform size whose width is one-half of the block size width and height is the same as the block size height.

[0163] ·PARTITION_HORZ4 (horizontal 4 partition): Assign a transform size whose width is the same as the block size width and height is one-fourth of the block size height.

[0164] ·PARTITION_VERT4 (vertical 4 partition): Assign a transform size whose width is one-fourth of the block size width and height is the same as the block size height.

[0165] In the above example, the transform partition types shown in Figure 17 all include a uniform transform size for the partitioned transform blocks. This is just an example and is not restrictive. In some other implementations, a mixed transform block size can be used for the partitioned transform blocks in a specific partition type (or pattern).

[0166] The PBs (or CBs, also referred to as PBs when not further divided into prediction blocks) obtained according to any of the above partitioning schemes can become a single block for encoding via intra-frame or inter-frame prediction. For inter-frame prediction of the current PB, a residual between the current block and the prediction block can be generated, encoded, and included in the encoded bitstream.

[0167] Inter-frame prediction can be implemented, for example, in a single-reference mode or a composite-reference mode. In some implementations, a skip flag can be included first in the bitstream of the current block (or a higher level) to indicate whether the current block is inter-frame encoded and not skipped. If the current block is inter-frame encoded, another flag can be further included in the bitstream as a signal to indicate whether the current block uses the single-reference mode or the composite-reference mode. For the single-reference mode, one reference block can be used to generate the prediction block of the current block. For the composite-reference mode, two or more reference blocks can be used, for example, to generate the prediction block by weighted averaging. The composite-reference mode can be referred to as more than one reference mode, two-reference mode, or multi-reference mode. One or more reference frame indices can be used and, in addition, one or more corresponding motion vectors indicating the position (e.g., horizontal and vertical pixels) offset between the reference block and the current block can be used to identify one or more reference blocks. For example, the inter-frame prediction block of the current block can be generated from a single reference block identified by a motion vector in a reference frame as the prediction block in the single-reference mode, while for the composite-reference mode, the prediction block can be generated by weighted averaging of two reference blocks in two reference frames indicated by two motion vectors. The motion vectors can be encoded and included in the bitstream in various ways.

[0168] In some implementations, an encoding or decoding system may have a decoded picture buffer (DPB). Some images / pictures can be stored in the DPB waiting to be displayed (in the decoding system), and some images / pictures in the DPB can be used as reference frames for inter-frame prediction. In some implementations, the reference frames in the DPB can be marked as short-term references or long-term references for the currently encoded or decoded image. For example, short-term reference frames may include frames used to perform operations such as inter-frame prediction of blocks in the current frame or inter-frame prediction of blocks in a predetermined number (e.g., 2) of subsequent video frames closest to the current frame in decoding order. Long-term reference frames may include frames in the DPB that can be used to predict image blocks in frames that are more than a predefined number away from the current frame in decoding order. Information about such tags for short-term and long-term reference frames can be referred to as a Reference Picture Set (RPS), and this information can be added to the header of each frame in the encoded bitstream. Each frame in the encoded video stream can be identified by a Picture Order Counter (POC), which is numbered in an absolute manner according to the playback sequence or in an order related to a group of pictures starting from, for example, an I-frame.

[0169] In some example implementations, one or more reference picture lists containing the identities of short-term and long-term reference frames for inter-frame prediction can be formed based on the information in the RPS. For example, a single picture reference list can be formed for uni-directional inter-frame prediction, denoted as the L0 reference (or reference list 0), and two picture reference lists can be formed for bi-directional inter-frame prediction, with each prediction direction denoted as L0 (or reference list 0) and L1 (or reference list 1). The reference frames included in the L0 and L1 lists can be sorted in various predetermined ways. The lengths of the L0 and L1 lists can be written into the video bitstream. When multiple reference frames used to generate a predicted block by weighted average in a composite prediction mode are on the same side of the block to be predicted, uni-directional inter-frame prediction can be in a single-reference mode or in a composite-reference mode. Bi-directional inter-frame prediction can only be in a composite mode because bi-directional inter-frame prediction involves at least two reference blocks.

[0170] In some implementations, a merge mode (MM) for inter - frame prediction can be implemented. Generally, for the merge mode, the motion vector in the single - reference prediction of the current PB or one or more motion vectors in the composite - reference prediction can be derived from other motion vectors instead of being independently calculated and written. For example, in an encoding system, the current motion vector of the current PB can be reduced to the difference between the current motion vector and one or more other already - encoded motion vectors (referred to as reference motion vectors). Only this difference of the motion vectors, rather than the entire current motion vector, can be encoded and included in the bitstream, and this difference can be linked to the reference motion vector. Correspondingly, in a decoding system, the motion vector corresponding to the current PB can be derived based on the decoded motion - vector difference and the decoded reference motion vector linked to it. As a specific form of the general merge mode (MM) inter - frame prediction, this inter - frame prediction based on the motion - vector difference can be referred to as a merge mode with motion - vector difference (MMVD). Thus, the general MM or the specific MMVD can be implemented to utilize the correlation between the motion vectors associated with different PBs to improve the encoding efficiency. For example, adjacent PBs can have similar motion vectors. For another example, for blocks with similar positioning / locations in space, the motion vectors can be temporally (between frames) correlated.

[0171] In some example implementations, during the encoding process, an MM flag can be included in the bitstream to indicate whether the current PB is in the merge mode. Additionally or alternatively, an MMVD flag can be included in and written to the bitstream during the encoding process to indicate whether the current PB is in the MMVD mode. The MM and / or MMVD flags or indicators can be provided at the PB level, CB level, CU level, CTB level, CTU level, slice level, picture level, etc. For a specific example, for the current CU, an MM flag and an MMVD flag can be included, and the MMVD flag can be written immediately after the skip flag and the MM flag to specify whether the MMVD mode is used for the current CU.

[0172] In some example implementations of MMVD, a merge candidate list for motion vector prediction may be formed for the predicted block. The merge candidate list may contain a predetermined number (e.g., 2) of MV predictor candidate blocks, whose motion vectors may be used to predict the current motion vector. The MVD candidate blocks may include blocks selected from adjacent blocks and / or temporal blocks (e.g., blocks at the same position in the previous or next frame of the current frame) of the same frame. These options represent blocks that are in spatial or temporal positions relative to the current block, and these blocks may have similar or identical motion vectors to the current block. The size of the MV predictor candidate list may be predetermined. For example, the list may contain two candidate blocks. To be on the merge candidate list, a candidate block may, for example, need to have the same one (or more) reference frames as the current block, must exist (e.g., when the current block is near the edge of the frame, boundary checks need to be performed), and must have been encoded during the encoding process and / or decoded during the decoding process. In some implementations, if the merge candidate list is available and meets the above conditions, it may first be filled with spatially adjacent blocks (scanned in a specific predefined order), and then, if there is still space available in the list, it may be filled with temporal blocks. For example, these adjacent candidate blocks may be selected from the left and top blocks of the current block. The merged MV predictor candidate list may be written to the bitstream.

[0173] In some implementations, the actual merge candidates that are used as reference motion vectors for predicting the motion vector of the current block may be written to the bitstream. In the case where the merge candidate list contains two candidates, a one-bit flag called the merge candidate flag may be used to indicate the selection of the reference merge candidate. For a current block predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor may be associated with a reference motion vector from the merge candidate list.

[0174] In some example implementations of MMVD, after selecting a merge candidate and using it as the base motion vector predictor for the motion vector to be predicted, a motion vector difference (MVD or delta MV, representing the difference between the motion vector to be predicted and the reference candidate motion vector) may be calculated in the encoding system. Such an MVD may include information representing the magnitude and direction of the MV difference, and this information may be written to the bitstream. The motion difference magnitude and the motion difference direction may be written to the bitstream in various ways.

[0175] In some example implementations of MMVD, a distance index can be used to specify the magnitude information of the motion vector difference and indicate one of a set of predefined offsets that represent predefined motion vector differences relative to a starting point (reference motion vector). Then, the MV offset according to the index indicated by the signal can be added to the horizontal or vertical component of the starting (reference) motion vector. The offset of the horizontal or vertical component of the reference motion vector should be determined by the exemplary direction information of the MVD. An example of the predefined relationship between the distance index and the predefined offsets is specified in Table 2.

[0176] Table 2 Example relationships between distance index and predefined MV offsets

[0177] Distance Index 0 1 2 3 4 5 6 7 Offset (in luminance samples) 1 / 4 1 / 2 1 2 4 8 16 32

[0178] In some exemplary implementations of MMVD, a direction index can be further written into the bitstream and used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction can be defined as either the horizontal direction or the vertical direction. An example of a 2-bit direction index is shown in Table 3. In the example of Table 3, the interpretation of the MVD can vary according to the information of the starting / reference MVs. For example, when the starting / reference MV corresponds to a single-prediction block or to a dual-prediction block where both reference frame lists point to the same side of the current picture (i.e., the POCs of both reference pictures are greater than the POC of the current picture, or both are less than the POC of the current picture), the signs in Table 3 can specify the signs (directions) of the MV offsets added to the starting / reference MV. When the starting / reference MV corresponds to a dual-prediction block of two reference pictures on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture while the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the signs in Table 3 can specify the signs of the MV offsets added to the reference MV corresponding to the reference picture in picture reference list 0, and the signs of the offsets of the MV corresponding to the reference picture in picture reference list 1 can have opposite values (opposite signs for the offsets). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the signs in Table 3 can specify the signs of the MV offsets added to the reference MV associated with picture reference list 1, and the signs of the offsets of the reference MV associated with picture reference list 0 have opposite values.

[0179] Table 3 Example implementations of signs of MV offsets specified by the direction index

[0180] Direction IDX 00 01 10 11 x - axis (horizontal) + - N / A N / A y - axis (vertical) N / A N / A + -

[0181] In some example implementations, the MVD can be scaled based on the difference in POCs in each direction. If the difference in POCs in the two lists is the same, no adjustment is required. Otherwise, if the POC difference in reference list 0 is greater than the POC difference in reference list 1, the MVD of reference list 1 is adjusted. If the POC difference in reference list 1 is greater than the POC difference in list 0, the MVD of list 0 can be adjusted in the same way. If the starting MV is single predicted, the MVD is added to the available MV or the reference MV.

[0182] In some example implementations for MVD coding and writing for bidirectional composite prediction, in addition to separately coding and writing two MVDs, symmetric MVD coding can be implemented such that only one MVD needs to be written and the other MVD can be derived based on the written MVD. In such an implementation, the motion information including the reference picture indices of both list-0 and list-1 is written into the bitstream. However, only the MVD associated with, for example, reference list-0 is written, and the MVD associated with reference list-1 is not written but derived. Specifically, at the slice level, a flag, called "mvd_l1_zero_flag", can be included in the bitstream to indicate whether reference list-1 is not written into the bitstream. If this flag is 1, indicating that reference list-1 is equal to zero (thus not written), the bidirectional prediction flag, called "BiDirPredFlag", can be set to 0, which means there is no bidirectional prediction. Otherwise, if mvd_l1_zero_flag is zero, if the nearest reference picture in list-0 and the nearest reference picture in list-1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag can be set to 1, and the reference pictures of list-0 and list-1 are both short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. BiDirPredFlag being 1 means that the symmetric mode flag is additionally written into the bitstream. When BiDirPredFlag is 1, the decoder can extract the symmetric mode flag from the bitstream. For example, the symmetric mode flag can be written at the CU level (if needed), and it indicates whether the symmetric MVD coding mode is being used for the corresponding CU. When the symmetric mode flag is 1, it means that the symmetric MVD coding mode is used, and only the reference picture indices of both list-0 and list-1 (called "mvp_l0_flag" and "mvp_l1_flag") and the MVD associated with list-0 (called "MVD0") are written, and the other motion vector difference "MVD1" will be derived instead of being written. For example, MVD1 can be derived as -MVD0. Thus, only one MVD is written into the bitstream in the exemplary symmetric MVD mode. In some other exemplary implementations of MV prediction, for single-reference mode and composite-reference mode MV prediction, a coordination scheme can be used to implement the general merge mode MMVD and some other types of MV prediction. Various syntax elements can be used to represent the way to predict the MV of the current block.

[0183] For example, for the single-reference mode, the following MV prediction modes can be written into the bitstream:

[0184] NEARMV - directly use one of the motion vector predictors (MVPs) in the list indicated by the Dynamic Reference List (DRL) index without using any MVD.

[0185] NEWMV – use one of the motion vector predictors (MVPs) in the list written by the DRL index as a reference and apply an increment to the MVP (e.g., use MVD).

[0186] GLOBALMV – use a motion vector based on frame-level global motion parameters.

[0187] Similarly, for the composite reference inter-prediction mode that uses two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes can be written into the bitstream:

[0188] NEAR_NEARMV – for each of the two MVs to be predicted, use one of the motion vector predictors (MVPs) in the list written by the DRL index without using MVD.

[0189] NEAR_NEWMV – to predict the first of the two motion vectors, use one of the motion vector predictors (MVPs) in the list written by the DRL index as a reference MV without using MVD; to predict the second of the two motion vectors, use one of the motion vector predictors (MVPs) in the list written by the DRL index as a reference MV and combine it with an additionally written incremental MV (MVD).

[0190] NEW_NEARMV – to predict the second of the two motion vectors, use one of the motion vector predictors (MVPs) in the list written by the DRL index as a reference MV without using MVD; to predict the first of the two motion vectors, use one of the motion vector predictors (MVPs) in the list written by the DRL index as a reference MV and combine it with an additionally written incremental MV (MVD).

[0191] NEW_NEWMV – use one of the motion vector predictors (MVPs) in the list written by the DRL index as a reference MV and use it in combination with an additionally written incremental MV to predict each of the two MVs.

[0192] GLOBAL_GLOBALMV – use each reference MV according to frame-level global motion parameters.

[0193] Therefore, the above term "NEAR" refers to MV prediction that uses the reference MV without using the MVD as a general merge mode, while the term "NEW" refers to MV prediction that involves using the reference MV and offsetting it with the written MVD as the MMVD mode. For composite inter prediction, both the above reference basic motion vector and motion vector difference can generally be different or independent between the two references, even if they can be correlated, and this correlation can be utilized to reduce the amount of information required to write the two motion vector differences. In this case, joint writing of the two MVDs can be achieved and indicated in the bitstream.

[0194] The above dynamic reference list (DRL) can be used to save a set of indexed motion vectors, which are dynamically saved and considered candidate motion vector predictors.

[0195] In some example implementations, a predefined resolution of the MVD can be allowed. For example, a motion vector precision (or accuracy) of 1 / 8 pixel can be allowed. The MVDs described above in various MV prediction modes can be constructed and written into the bitstream in various ways. In some implementations, various syntax elements can be used to represent the above motion vector differences in reference frame list 0 or list 1.

[0196] For example, a syntax element called "mv_joint" can specify which components of the associated motion vector difference are non-zero. For the MVD, this is a joint write for all non-zero components. For example, mv_joint has the following values.

[0197] 0 can indicate that there is no non-zero MVD in the horizontal or vertical direction;

[0198] 1 can indicate that there is a non-zero MVD only in the horizontal direction;

[0199] 2 can indicate that there is a non-zero MVD only in the vertical direction;

[0200] 3 can indicate that there are non-zero MVDs in both the horizontal and vertical directions.

[0201] When the "mv_joint" syntax element of the MVD signal indicates that there are no non-zero MVD components, then no further MVD information can be written. If the "mv_joint" syntax indicates the presence of one or two non-zero components, additional syntax elements can be written for each non-zero MVD component, as described below.

[0202] For example, a syntax element called "mv_sign" can be used to additionally specify whether the corresponding motion vector difference component is positive or negative.

[0203] For another example, a syntax element called "mv_class" can be used to specify the class of the motion vector difference in a predefined set of classes for the corresponding non-zero MVD component. For example, the predefined classes of the motion vector difference can be used to partition the continuous magnitude space of the motion vector difference into non-overlapping ranges, each range corresponding to an MVD class. Thus, the written MVD class indicates the magnitude range of the corresponding MVD component. In the example implementation shown in Table 4, higher classes correspond to motion vector differences with larger magnitude ranges. In Table 4, the notation (n, m] is used to represent the range of motion vector differences greater than n pixels and less than or equal to m pixels.

[0204] Table 4 Magnitude classes of motion vector differences

[0205]

[0206]

[0207] In some other examples, a syntax element called "mv_bit" can be further used to specify the integer part of the offset between the non-zero motion vector difference component and the starting magnitude of the corresponding written MV class. The number of bits required in "mv_bit" to write the entire range for each MVD class can vary as a function of the MV class. For example, in the implementation of Table 4, MV_CLASS 0 and MV_CLASS1 may only require a single bit to represent an integer pixel offset of 1 or 2 starting from an MVD of 0. Each higher MV_CLASS may require one more "mv_bit" than the previous MV_CLASS.

[0208] In some other examples, a syntax element called "mv_fr" can be further used to specify the first 2 fractional bits of the motion vector difference for the corresponding non-zero MVD component, while a syntax element called "mv_hp" can be used to specify the third fractional bit (high-resolution bit) of the motion vector difference for the corresponding non-zero MVD component. The two-bit "mv_fr" actually provides an MVD resolution of 1 / 4 pixel, and the "mv_hp" bit can further provide a resolution of 1 / 8 pixel. In some other implementations, multiple "mv_hp" bits can be used to provide a finer MVD pixel resolution than 1 / 8 pixel. In some example implementations, additional flags can be written to the bitstream at one or more different levels to indicate whether 1 / 8 pixel or higher MVD resolution is supported. If the MVD resolution is not applied to a particular coding unit, then the syntax elements for the corresponding unsupported MVD resolution may not be written to the bitstream.

[0209] In the example implementation above, the fractional resolution may be independent of different classes of MVDs. In other words, regardless of the magnitude of the motion vector difference, a predefined number of "mv_fr" and "mv_hp" bits can be used to write the fractional MVD of the non-zero MVD component, thereby providing similar options for motion vector resolution.

[0210] However, in some other example implementations, the resolution of the motion vector difference can be differentiated among various MVD size classes. Specifically, for larger MVDs in higher MVD classes, a high-resolution MVD may not provide a statistically significant improvement in compression efficiency. Thus, for a larger range of MVD sizes, which correspond to higher MVD size classes, the MVD can be encoded with decreasing resolution (integer pixel resolution or fractional pixel resolution). Similarly, for generally larger MVD values, the MVD can be encoded with decreasing resolution (integer pixel resolution or fractional pixel resolution). This MVD resolution that depends on the MVD class or on the MVD size is generally referred to as adaptive MVD resolution. Adaptive MVD resolution can be implemented in various scenarios described in the example implementations below to achieve overall better compression efficiency. In particular, since it has been statistically observed that processing the MVD resolution of larger or higher-level MVDs in a non-adaptive manner similar to that of smaller or lower-level MVDs may not significantly increase the inter-prediction residual coding efficiency, the number of bits of signaling reduced by focusing on lower-precision MVDs may be greater than the additional bits required for the inter-frame prediction residual as a result of such lower-precision MVDs.

[0211] In some general example implementations, the pixel resolution or precision of the MVD can decrease or not increase as the MVD class increases. Decreasing the pixel resolution of the MVD corresponds to a coarser MVD (or a larger step from one MVD level to the next). In some implementations, the correspondence between the MVD pixel resolution and the MVD class can be specified, predefined, or pre-configured, and thus may not need to be written into the encoded bitstream.

[0212] In some example implementations, the MV classes in Table 4 can each be associated with a different MVD pixel resolution.

[0213] In some example implementations, each MVD class can be associated with a single allowed resolution. In some other implementations, one or more MVD classes can be associated with two or more optional MVD pixel resolutions. Thus, the signaling in the bitstream for such an MVD class is followed by additional signaling for indicating which of the optional pixel resolutions is selected for the current MVD component.

[0214] In some example implementations, the adaptively allowed MVD pixel resolutions can include, but are not limited to, 1 / 64 pixel, 1 / 32 pixel, 1 / 16 pixel, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels... (in descending order of resolution). Thus, each ascending MVD class can be associated with one of these resolutions in a non-ascending manner. In some implementations, an MVD class can be associated with two or more resolutions, and the higher resolution can be lower than or equal to the lower resolution of the previous MVD class. For example, if MV_CLASS_3 in Table 4 can be associated with optional 1 pixel and 2 pixel resolutions, the highest resolution that MV_CLASS_4 in Table 4 can be associated with will be 2 pixels. In some other implementations, the highest allowed resolution of an MV class may be higher than the lowest allowed resolution of the previous (lower) MV class. However, the allowed average resolution of ascending MV classes may only be non-ascending.

[0215] In some implementations, when fractional pixel resolutions higher than 1 / 8 pixel are allowed, the "mv_fr" and "mv_hp" signaling can be extended accordingly to a total of more than 3 fractional bits.

[0216] In some example implementations, fractional pixel resolutions can only be allowed for MVD classes that are lower than or equal to a threshold MVD class. For example, fractional pixel resolutions can only be allowed for MVD-CLASS 0 and not for all other MV classes in Table 4. Similarly, fractional pixel resolutions can only be allowed for MVD classes that are lower than or equal to any one of the other MV classes in Table 4. For other MVD classes higher than the threshold MVD class, only integer pixel resolutions of the MVD are allowed. In this way, for MVDs written with MVD classes higher than or equal to the threshold MVD class, it may not be necessary to write fractional resolution signaling such as one or more of the "mv-fr" and / or "mv-hp" bits. For MVD classes with resolutions lower than 1 pixel, the number of bits in the "mv-bit" signaling can be further reduced. For example, for MV_CLASS_5 in Table 4, the range of the MVD pixel offset is (32, 64], so 5 bits are required to write the entire range of 1 pixel resolution. However, if MV_CLASS_5 is associated with a 2 pixel MVD resolution, "mv-bit" may require 4 bits instead of 5 bits, and neither "mv-fr" nor "mv-hp" needs to be written after writing "mv_class" as "MV_CLASS_5".

[0217] In some example implementations, an MVD having an integer value below a threshold integer pixel value may only allow fractional pixel resolution. For example, for an MVD less than 5 pixels, fractional pixel resolution may be allowed only. Corresponding to this example, fractional resolution may be allowed for MV_CLASS_0 and MV_CLASS_1 of Table 4, while fractional resolution may not be allowed for all other MV classes. As another example, for an MVD less than 7 pixels, fractional pixel resolution may be allowed only. Corresponding to this example, fractional resolution may be allowed for MV_CLASS_0 and MV_CLASS_1 (range less than 5 pixels) of Table 4, while fractional resolution may not be allowed for MV_CLASS_3 and higher (range greater than 5 pixels). For an MVD belonging to MV_CLASS_2 with a pixel range of 5 pixels, the fractional pixel resolution of the MVD may be allowed or not allowed according to the "m-bit" value. If the "m-bit" value is written as 1 or 2 (such that the integer part of the written MVD is 5 or 6, calculated as the start of the pixel range of MV_CLASS_2 with an offset of 1 or 2 represented by the "m-bit"), then fractional pixel resolution may be allowed. Otherwise, if the "m-bit" value is written as 3 or 4 (such that the integer part of the written MVD is 7 or 8), then fractional pixel resolution may not be allowed.

[0218] In some other implementations, for MV classes equal to or higher than a threshold MV class, only a single MVD value may be allowed. For example, such a threshold MV class may be MV_CLASS2. Thus, MV_CLASS_2 and above may only be allowed to have a single MVD value and no fractional pixel resolution. The single allowed MVD values for these MV classes may be predefined. In some examples, the allowed single value may be the higher end value of the respective ranges of these MV classes in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be higher than or equal to the threshold class_2, and the single allowed MVD values for these classes may be predefined as 8, 16, 32, 64, 128, 256, 512, 1024, and 2048 respectively. In some other examples, the allowed single value may be the middle value of the respective ranges of these MV classes in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be higher than the class threshold, and the single allowed MVD values for these classes may be predefined as 3, 6, 12, 24, 48, 96, 192, 384, 768, and 1536 respectively. Any other value within the range may also be defined as the single allowed resolution for the respective MVD classes.

[0219] In the above implementation, when the written "mv_class" is equal to or higher than a predefined MVD level threshold, only the "mv_class" signaling is sufficient to determine the MVD value. Then, "mv_class" and "mv_sign" are used to determine the magnitude and direction of the MVD.

[0220] When the MVD is written only for one reference frame (from reference frame list 0 or list 1, but not both), or jointly for two reference frames, the accuracy (or resolution) of the MVD can depend on the category of the associated motion vector difference in Table 3 and / or the magnitude of the MVD.

[0221] In some other implementations, the pixel resolution or accuracy of the MVD can decrease or not increase as the MVD magnitude increases. For example, the pixel resolution can depend on the integer part of the MVD magnitude. In some implementations, a fractional pixel resolution can be allowed only for MVD magnitudes less than or equal to an amplitude threshold. For the decoder, the integer part of the MVD magnitude can first be extracted from the bitstream. Then the pixel resolution can be determined, and then it can be decided whether there is any fractional MVD in the bitstream and needs to be parsed (e.g., if fractional pixel resolution is not allowed for a particular extracted MVD integer size, then the bitstream to be extracted may not contain fractional MVD bits). The above example implementations related to adaptive MVD pixel resolution depending on the MVD category apply to adaptive MVD pixel resolution depending on the MVD magnitude. For a specific example, MVD categories above or including a size threshold may be allowed only a predefined value.

[0222] The above various example implementations are for the single reference mode. These implementations also apply to the NEW_NEARMV, NEAR_NEWMV, and / or NEW_NEWMV mode examples in composite prediction under MMVD. These implementations generally apply to the writing of the MVD.

[0223] Figure 18FIG. 1800 is a flowchart of an example method that follows the principles of the above-described implementation for adaptive MVD resolution. The example decoding method flow starts at S1801. In S1810, a video stream is received, and a motion vector difference (MVD) between a motion vector associated with an inter-predicted video block and a reference motion vector is written to the video stream, where the reference motion vector corresponds to a reference picture in only one of reference frame list 0 and reference frame list 1, unless the MVD is written jointly for two reference pictures. In S1820, an indication of the size range of the MVD in a plurality of predefined motion vector difference size ranges is obtained from the video stream. In S1830, the pixel resolution of the MVD is determined based on the size range. In S1840, additional MVD information in the video stream is identified based on the pixel resolution. In S1850, the additional MVD information is extracted from the video stream. In S1860, the inter-predicted video block is decoded based on the pixel resolution, the additional MVD information, the reference motion vector, and the reference frame associated with the motion vector. The example method flow ends at S1899.

[0224] Figure 19 FIG. 1900 is another flowchart of an example method that follows the principles of the above-described implementation for adaptive MVD resolution. The example method flow starts at S1901. In S1910, a video stream is received, and a motion vector difference (MVD) between a motion vector associated with an inter-predicted video block and a reference motion vector is written to the video stream, where the reference motion vector corresponds to a reference picture in only one of reference frame list 0 and reference frame list 1, unless the MVD is written jointly for two reference pictures. In S1920, the integer part of the size of the MVD is extracted from the video stream. In S1930, the pixel resolution of the MVD is determined based on the integer part of the size of the MVD. In S1940, additional MVD information in the video stream is identified based on the pixel resolution. In S1950, the inter-predicted video block is decoded based on the pixel resolution, the integer part of the size of the MVD, the additional MVD information, the reference motion vector, and the reference frame associated with the motion vector. The example method flow ends at S1999.

[0225] In the embodiments and implementations of the embodiments of the present application, any steps and / or operations can be combined or arranged in any quantity or order as needed. Two or more of the steps and / or operations can be executed in parallel. The embodiments and implementations in the embodiments of the present application can be used alone or in any combination in any order. Additionally, each method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments in the embodiments of the present application can be applied to luminance blocks or chrominance blocks. The term block can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term block here can also be used to refer to a transform block. In the following, when referring to the block size, it can refer to the width or height of the block, or the maximum of the width and height, or the minimum of the width and height, or the area size (width * height), or the aspect ratio of the block (width:height, or height:width).

[0226] The above techniques can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 20 FIG. shows a computer system (2000) suitable for implementing certain embodiments of the disclosed subject matter.

[0227] The computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be subject to mechanisms such as assembly, compilation, linking, or the like to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretive code, microcode, etc.

[0228] The instructions can be executed on various types of computers or their components, such as including personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0229] Figure 20 The components shown for the computer system (2000) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. The configuration of the components should also not be interpreted as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (2000).

[0230] A computer system (2000) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users, such as the following: tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0231] The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard (2001), mouse (2002), touchpad (2003), touch screen (2010), data glove (not shown), joystick (2005), microphone (2006), scanner (2007), camera (2008).

[0232] The computer system (2000) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of the touch screen (2010), data glove (not shown), or joystick (2005), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (2009), headphones (not shown)), visual output devices (e.g., screens (2010) including CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input function, each screen having or not having tactile feedback function - some of which can output two-dimensional visual output or more than three-dimensional output through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).

[0233] The computer system (2000) may also include human-accessible storage devices and their associated media: for example, optical media including CD / DVD ROM / RW (2020) with media such as CD / DVD (2021), thumb drives (2022), removable hard disk drives or solid state drives (2023), traditional magnetic media such as tapes and floppy disks (not shown), devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown), etc.

[0234] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover a transmission medium, a carrier wave, or other transient signals.

[0235] The computer system (2000) may also include an interface (2054) to one or more communication networks (2055). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial television including CAN Bus, and so on. Certain networks typically require an external network interface adapter (such as the USB port of the computer system (2000)) connected to certain common data ports or peripheral buses (2049); as described below, other network interfaces are typically integrated into the kernel of the computer system (2000) by connecting to the system bus (for example, connecting the Ethernet interface in a PC computer system or connecting the cellular network interface in a smartphone computer system). The computer system (2000) may communicate with other entities using any of these networks. Such communication may be only one-way reception (such as broadcast television), only one-way transmission (such as CANbus connected to certain CANbus devices), or two-way, for example, using a local area network or a wide area digital network to connect to other computer systems. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.

[0236] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel (2040) of the computer system (2000).

[0237] The kernel (2040) may include one or more central processing units (CPUs) (2041), a graphics processing unit (GPU) (2042), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2043), a hardware accelerator (2044) for certain tasks, a graphics adapter (2050), etc. These devices, as well as a read-only memory (ROM) (2045), a random access memory (2046), and an internal mass storage such as an internal hard disk drive, SSD, etc. that is not accessible to users (2047) can be connected via a system bus (2048). In some computer systems, the system bus (2048) can be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices can be directly connected to the system bus of the kernel (2048) or connected to the system bus of the kernel (1848) via a peripheral bus (2049). In one example, a screen (2010) can be connected to the graphics adapter (2050). The architecture of the peripheral bus includes PCI, USB, etc.

[0238] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) can execute certain instructions, which can be combined to form the aforementioned computer code. This computer code can be stored in the ROM (2045) or the RAM (2046). Transitional data can also be stored in the RAM (2046), while permanent data can be stored in, for example, the internal mass storage (2047). Fast storage and retrieval of any storage device can be achieved through the use of a cache, which can be closely associated with one or more CPUs (2041), GPUs (2042), mass storage (2047), ROM (2045), RAM (2046), etc.

[0239] A computer-readable medium can have computer code for performing various computer-implemented operations thereon. The medium and the computer code can be media and computer code that are specifically designed and constructed for the purposes of the embodiments of the present application, or the medium and the computer code can be of the type that is well-known and available to those skilled in the field of computer software.

[0240] As a non-limiting example, a computer system having an architecture (2000), particularly a kernel (2040), can provide functionality due to software executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-transitory memories of the kernel (2040), such as the on-kernel mass memory (2047) or ROM (2045). Software implementing various embodiments of the present application can be stored in such devices and executed by the kernel (2040). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the kernel (2040), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific parts of specific processes described herein, including defining data structures (2046) stored in the RAM and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system can provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., an accelerator (2044)), which can replace the software or operate together with the software to execute specific processes or specific parts of specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include circuits (e.g., integrated circuits (ICs)) storing software for execution, logic circuits embodying the logic for execution, or both. Embodiments of the present application include any suitable combination of hardware and software.

[0241] Although some exemplary embodiments have been described in embodiments of the present application, there are changes, permutations, and various alternative equivalents that fall within the scope of embodiments of the present application. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of embodiments of the present application and thus fall within the spirit and scope of embodiments of the present application.

[0242] Appendix A: Abbreviations

[0243] JEM: Joint Exploration Model

[0244] VVC: Next Generation Video Coding

[0245] BMS: Benchmark Set

[0246] MV: Motion Vector

[0247] HEVC: High Efficiency Video Coding

[0248] SEI: Supplementary Enhancement Information

[0249] VUI: Video Usability Information

[0250] GOPs: Groups of Pictures

[0251] TUs: Transform Units

[0252] PUs: Prediction Units

[0253] CTUs: Coding Tree Units

[0254] CTBs: Coding Tree Blocks

[0255] PBs: Prediction Blocks

[0256] HRD: Hypothetical Reference Decoder

[0257] SNR: Signal-to-Noise Ratio

[0258] CPUs: Central Processing Units

[0259] GPUs: Graphics Processing Units

[0260] CRT: Cathode Ray Tube

[0261] LCD: Liquid Crystal Display

[0262] OLED: Organic Light-Emitting Diode

[0263] CD: Compact Disc

[0264] DVD: Digital Versatile Disc

[0265] ROM: Read-Only Memory

[0266] RAM: Random Access Memory

[0267] ASIC: Application-Specific Integrated Circuit

[0268] PLD: Programmable Logic Device

[0269] LAN: Local Area Network

[0270] GSM: Global System for Mobile Communications

[0271] LTE: Long-Term Evolution

[0272] CANBus: Controller Area Network Bus

[0273] USB: Universal Serial Bus

[0274] PCI: Peripheral Component Interconnect

[0275] FPGA: Field-Programmable Gate Array

[0276] SSD: Solid State Drive

[0277] IC: Integrated Circuit

[0278] HDR: High Dynamic Range

[0279] SDR: Standard Dynamic Range

[0280] JVET: Joint Video Exploration Team

[0281] MPM: Most Probable Mode

[0282] WAIP: Wide Angle Intra Prediction

[0283] CU: Coding Unit

[0284] PU: Prediction Unit

[0285] TU: Transform Unit

[0286] CTU: Coding Tree Unit

[0287] PDPC: Position-Dependent Prediction Combination

[0288] ISP: Intra Sub-Block Partitioning

[0289] SPS: Sequence Parameter Set

[0290] PPS: Picture Parameter Set

[0291] APS: Adaptive Parameter Set

[0292] VPS: Video Parameter Set

[0293] DPS: Decoding Parameter Set

[0294] ALF: Adaptive Loop Filter

[0295] SAO: Sample Adaptive Offset CC-ALF: Cross-Component Adaptive Loop Filter CDEF: Constrained Directional Enhancement Filter

[0296] CCSO: Cross-Component Sample Offset

[0297] LSO: Local Sample Offset

[0298] LR: Loop Restoration Filter

[0299] AV1: Alliance for Open Media Video 1

[0300] AV2: Alliance for Open Media Video 2

[0301] MVD: Motion Vector Difference

[0302] CfL: Chrominance from Luminance Prediction

[0303] SDT: Semi-Decoupled Tree

[0304] SDP: Semi-Decoupled Partitioning

[0305] SST: Semi-Separation Tree

[0306] SB: Super Block

[0307] IBC (or IntraBC): Intra-Block Copy

[0308] CDF: Cumulative Distribution Function

[0309] SCC: Screen Content Coding

[0310] GBI: Generalized Bi-Prediction

[0311] BCW: CU-Level Weighted Bi-Prediction

[0312] CIIP: Combined Intra-Inter Prediction

[0313] POC: Picture Order Count

[0314] RPS: Reference Picture Set

[0315] DPB: Decoded Picture Buffer

[0316] MMVD: Motion Vector Difference Merge Mode

Claims

1. A method for decoding an inter - predicted video block of a video stream, characterized in that, the method comprises: receiving the video stream; determining that a motion vector difference (MVD) between a motion vector associated with the inter - predicted video block and a reference motion vector is written into the video stream, wherein the reference motion vector corresponds to a reference picture in only one of reference frame list 0 and reference frame list 1, unless the MVD is jointly written for two reference pictures; obtaining an indication of a size range of the MVD in a plurality of predefined motion vector difference size ranges from the video stream; determining a pixel resolution of the MVD according to the size range of the MVD; identifying additional MVD information in the video stream based on the pixel resolution of the MVD, wherein the additional MVD information indicates an optional pixel resolution selected by the MVD; extracting the additional MVD information from the video stream; and decoding the inter - predicted video block based on the pixel resolution of the MVD, the additional MVD information, the reference motion vector, and a reference frame associated with the motion vector.

2. The method according to claim 1, characterized in that, The pixel resolution of the MVD is 2 n pixels, where n is an integer with a value range between -6 and 11, inclusive of -6 and 11.

3. The method according to claim 1, characterized in that, the plurality of predefined size ranges of the MVD are associated with pixel resolutions in a predefined manner in descending order, wherein a higher pixel resolution is associated with a smaller pixel resolution value.

4. The method according to claim 1, characterized in that, the indication of the size range of the MVD comprises: extracting a first predefined syntax element from the video stream, wherein the first predefined syntax element indicates an MVD category of the MVD in a predefined set of MVD categories, and a lower MVD category corresponds to a smaller MVD size range; and determining the size range of the MVD according to the MVD category.

5. The method according to claim 1, characterized in that, the determining the pixel resolution of the MVD according to the size range of the MVD comprises: determining whether the size range of the MVD is higher than a preset MVD range threshold level; when determining that the size range of the MVD is higher than the preset MVD range threshold level, determining that the pixel resolution of the MVD is an integer number of pixels; and when determining that the size range of the MVD is not higher than the preset MVD range threshold level, determining that the pixel resolution of the MVD is a fractional number of pixels.

6. The method according to claim 5, characterized in that, the identifying additional MVD information in the video stream based on the pixel resolution of the MVD comprises: parsing the video stream according to a second predefined syntax element to obtain an integer - pixel part of the MVD; and when determining that the pixel resolution of the MVD is a fractional number of pixels, the method further comprises: parsing the video stream according to at least a third predefined syntax element to obtain a fractional - pixel part of the MVD.

7. The method according to claim 5, characterized in that, The preset MVD range threshold level includes the lowest MVD or the second lowest MVD in the set of predefined MVD categories.

8. The method according to claim 5, wherein, each MVD in the set of predefined MVD categories with a size range higher than the preset MVD range threshold level is associated with a single allowed integer MVD pixel value.

9. The method according to claim 8, wherein, the single allowed integer pixel value includes the pixel value corresponding to the higher value in the corresponding size range.

10. The method according to claim 8, wherein, the single allowed integer pixel value includes the pixel value corresponding to the midpoint value in the corresponding size range.

11. The method according to claim 6, wherein, determining the pixel resolution of the MVD according to the size range of the MVD includes: determining whether the size range of the MVD is lower than, includes, or higher than a preset MVD threshold size value; when determining that the size range of the MVD is higher than the preset MVD threshold size value, determining that the pixel resolution of the MVD is an integer number of pixels; and when determining that the size range of the MVD is not higher than the preset MVD threshold size value, determining that the pixel resolution of the MVD is a fractional number of pixels.

12. The method according to claim 11, wherein, further comprising: when determining that the size range of the MVD includes the preset MVD threshold size value: extracting a second predefined syntax element from the video stream, the second predefined syntax element indicating the MVD size offset of the MVD relative to the starting size of the size range of the MVD; obtaining the integer size of the MVD based on the size range of the MVD and the MVD size offset; when the integer size of the MVD is not higher than the preset MVD threshold size value, determining that the pixel resolution of the MVD is a fraction; and when the integer size of the MVD is higher than the preset MVD threshold size value, determining that the pixel resolution of the MVD is not a fraction.

13. The method according to claim 12, wherein, when determining that the pixel resolution of the MVD is a fraction, identifying additional MVD information in the video stream based on the pixel resolution of the MVD includes: parsing the video stream according to a third predefined syntax element to obtain the fractional part of the MVD.

14. The method according to claim 11, wherein, the preset MVD threshold size value is less than 4 pixels.

15. The method according to claim 1, wherein, the MVD pixel resolution associated with the multiple predefined motion vector difference size ranges is different from one size range to another.

16. An electronic device, wherein, The electronic device includes a memory for storing computer instructions and a processor communicating with the memory. Wherein, when executing the computer instructions, the processor is configured to cause the electronic device to execute the decoding method described in any one of claims 1 to 15.

17. A method for storing or transmitting a video stream, characterized in that, the video stream is decoded based on the decoding method described in any one of claims 1 to 15.