Method, device, and readable storage medium for processing a current video block of a video stream
By introducing an encoding method with adaptive MVD pixel resolution in the video stream, the problem of poor MVD processing efficiency and quality in the prior art is solved, and more efficient and high-quality video encoding is achieved.
Patent Information
- Application Number
- CN202280012272.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-05-25
- Filing Date
- 2022-06-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-06-03
AI Technical Summary
When existing video encoding technologies deal with motion vector differences (MVD), it is difficult to achieve adaptive resolution adjustment, resulting in poor encoding efficiency and quality.
By introducing an encoding method of adaptive MVD pixel resolution in the video stream, the resolution is dynamically adjusted according to the magnitude and category of the MVD to ensure the best encoding performance in different scenarios.
Improves the efficiency and quality of video encoding, reduces bandwidth and storage requirements, and enhances adaptability to different scenarios.
Smart Images

Figure CN116830572B_ABST
Abstract
Description
[0001] Incorporation by Reference
[0002] This application claims priority to U.S. Non - Provisional Patent Application No. 17 / 824,193, filed on May 25, 2022, with the title "Schemes for Adjusting Adaptive Resolution for Motion Vector Difference", which in turn claims priority to U.S. Provisional Patent Application No. 63 / 302,518, filed on January 24, 2022, with the title "Further Improvement for Adaptive MVD Resolution". These prior applications are hereby incorporated by reference in their entirety into this application. Technical Field
[0003] This application generally relates to video coding and decoding, and more particularly to a method, an apparatus, and a readable storage medium for processing a current video block of a video stream. Background Art
[0004] The background art provided herein is for the purpose of generally presenting the context of this application. To the extent that the work of the currently named inventors is described in this background art section, and aspects that may not qualify as prior art at the time of filing of this application, are neither expressly nor implicitly admitted as prior art of this application.
[0005] Video encoding and decoding can use inter - picture prediction with motion compensation. An uncompressed digital video can include a series of pictures, each picture having a certain spatial dimension. For example, it has 1920×1080 luma samples and associated full - chroma samples or subsampled chroma samples. The series of pictures can have a fixed or variable picture rate (alternatively referred to as frame rate), e.g., 60 pictures per second or 60 frames per second. Uncompressed video has specific bit - rate requirements for streaming or data processing. For example, a video with a pixel resolution of 1920×1080, a frame rate of 60 frames per second, and 4:2:0 chroma subsampling with 8 bits per pixel per color channel requires a bandwidth of nearly 1.5 Gbit / s. Such a video requires more than 600 GB of storage space per hour.
[0006] One purpose of video encoding and decoding can be to reduce redundancy in an uncompressed input video signal through compression. Compression can help reduce the above-mentioned bandwidth and / or storage space requirements, and in some cases, can reduce them by two or more orders of magnitude. Both lossless compression and lossy compression, as well as combinations thereof, can be used for video encoding and decoding. Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal through the decoding process. Lossy compression refers to an encoding / decoding process in which the original video signal is not fully preserved during the encoding process and not fully restored during the decoding process. When lossy compression is used, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application, although there is some information loss. For video, lossy compression is widely used in many applications. The amount of distortion that lossy compression can tolerate depends on the application. For example, consumer users of certain video streaming applications can tolerate higher distortion compared to users of movie or television broadcast applications. The compression ratio that a particular encoding algorithm can achieve can be selected or adjusted to reflect various distortion tolerances: the higher the distortion that can be tolerated, the more lossy and higher compression ratio encoding algorithms are typically allowed to be used.
[0007] Video encoders and decoders can use several major categories of techniques and steps, including, for example, motion compensation, Fourier transform, quantization, and entropy coding.
[0008] Video coding and decoding techniques can include intra-frame coding techniques. In intra-frame coding techniques, the representation of sample values does not refer to samples or other data in previously reconstructed reference pictures. In some video coding and decoding techniques, a picture is spatially divided into sample blocks. When all sample blocks are encoded in the intra-frame mode, the picture can be called an intra-frame picture. Intra-frame pictures and their derivative pictures, for example, pictures refreshed by an independent decoder, can be used to reset the state of the decoder, and thus can be used as the first picture in an encoded video bitstream and a video session, or as a still picture. Then, the samples of the blocks predicted intra-frame can be transformed into the frequency domain, and the transformed coefficients thus generated can be quantized before entropy coding. Intra-frame prediction represents a technique for minimizing sample values in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are required to represent the block after entropy coding for a given quantization step size.
[0009] Traditional intra - frame coding techniques, for example, the well - known MPEG - 2 coding technique, do not use intra - frame prediction. However, some newer video compression techniques include techniques that attempt to encode / decoder blocks based on, for example, neighboring sample data and / or metadata obtained during the encoding and / or decoding of data blocks that are spatially adjacent and prior in decoding order to the data block being intra - frame encoded or decoded. Thus, such techniques are called "intra - frame prediction" techniques. Note that, at least in some cases, intra - frame prediction only uses reference data in the currently reconstructed picture and does not use reference data in other reference pictures.
[0010] Intra - frame prediction can have many different forms. When more than one such technique is available in a given video coding technique, the technique in use can be called an intra - frame prediction mode. One or more intra - frame prediction modes can be provided in a particular encoding / decoding. In some cases, some modes have sub - modes and / or are associated with various parameters. The mode / sub - mode information and intra - frame coding parameters of a video block can be encoded separately or can be collectively included in a mode codeword. Which codeword is used for a given combination of mode / sub - mode and / or parameter affects the coding efficiency gain through intra - frame prediction, and the entropy coding technique used to translate the codeword into the bitstream also has an impact on it.
[0011] The H.264 standard introduced intra - frame prediction for a certain mode, and the H.265 standard improved it. In newer coding techniques, such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), Benchmark Set (BMS), etc., it has been further improved. Generally, for intra - frame prediction, a predictor block can be formed using available neighboring sample values. For example, the available values of neighboring samples along a specific direction and / or a specific set of rows can be copied into the predictor block. The reference to the direction used can be encoded into the bitstream, or it can be predicted itself.
[0012] Reference Figure 1A , depicted in the lower right corner of which is a known subset of 9 predictor directions out of 33 possible intra - frame predictor directions (corresponding to 33 angular modes of 35 intra - frame modes defined in the H.265 standard) of the H.265 standard. Among them, the convergence point (101) of each arrow represents the sample being predicted. The arrows represent the directions of predicting the sample at 101 using neighboring samples. For example, arrow (102) represents predicting sample (101) based on one or more neighboring samples in the upper right corner at an angle of 45 degrees with respect to the horizontal axis. Similarly, arrow (103) represents predicting sample (101) based on one or more neighboring samples in the lower left corner at an angle of 22.5 degrees with respect to the horizontal direction.
[0013] Still referring to Figure 1A as shown Figure 1A Depicted in the upper left corner is a square block (104) with 4×4 samples (represented by bold dashed lines). The square block (104) includes 16 samples, each sample being labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second (counting from top to bottom) in the Y dimension and the first (counting from left to right) in the X dimension. Similarly, sample S44 is the fourth sample in both the X dimension and the Y dimension in the block (104). Since the size of the block is 4×4 samples, S44 is in its lower right corner. Figure 1A Example reference samples are further shown, and the reference samples follow a similar numbering method. The reference samples are labeled with R, their Y position (e.g., row index) and X position (e.g., column index) relative to the block (104). In the H.264 standard and the H.265 standard, predictive samples adjacent to the block in reconstruction are used.
[0014] Intra-picture prediction of block 104 can start by copying the reference sample values of adjacent samples according to the prediction direction represented by the signal. For example, assume that there is signaling in the encoded video bitstream, and for block 104, this signaling represents the prediction direction of arrow (102), that is, samples in the block are predicted based on one or more reference samples in the upper right corner at a 45-degree angle to the horizontal direction. In this case, samples S41, S32, S23, S14 are predicted based on the same reference sample R05. Sample S44 is predicted based on reference sample R08.
[0015] In some cases, the values of multiple reference samples can be combined, for example, by interpolation, to calculate a reference sample; especially when the direction cannot be evenly divisible by 45 degrees.
[0016] With the continuous development of video coding technology, the number of possible directions is also increasing. In the H.264 standard (in 2003), for example, 9 different directions can be used for intra prediction. In the H.265 standard (in 2013), it increased to 33 directions. By the time of the invention of this application, JEM / VVC / BMS can support up to 65 directions. Currently, some experimental studies have been conducted to help identify the most suitable intra prediction directions, and some entropy coding techniques encode these most suitable directions with very few bits, accepting a certain bit cost for the directions. In addition, sometimes these directions themselves can be predicted based on the adjacent directions used in intra prediction of adjacent decoded blocks.
[0017] Figure 1BFIG. 180 shows a schematic diagram depicting 65 intra prediction directions according to JEM, for illustrating that the number of prediction directions in various coding techniques increases over time.
[0018] In an encoded video bitstream, the mapping of the bits representing the intra prediction direction to the prediction direction varies with the video coding technique; for example, the range can vary from simply directly mapping the prediction direction of the intra prediction mode to the codeword, to complex adaptive schemes involving the most probable mode and similar techniques. However, in all these cases, statistically, certain directions for intra prediction are less likely to occur in video content compared to other directions. Since the purpose of video compression is to reduce redundancy, in better performing video coding techniques, these less likely directions are represented with more bits compared to the more likely directions.
[0019] Intra prediction or inter prediction can be based on motion compensation. In motion compensation, a block of sample data from a previously reconstructed picture or a part thereof (reference picture) can be spatially shifted in the direction indicated by a motion vector (hereinafter referred to as MV) and then used to predict a newly reconstructed picture or picture part (e.g., block). In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions, with the third dimension indicating the reference picture in use (i.e., the temporal dimension).
[0020] In some video compression techniques, the current MV applicable to a certain sample data region can be predicted based on other MVs, for example, based on other MVs associated with another sample data region that is spatially adjacent to the region being reconstructed and has a decoding order prior to the said MV. By doing so, the total amount of data required for encoding the MV can be significantly reduced by eliminating the redundancy of the associated MVs, thereby improving the compression efficiency. For example, MV prediction can operate effectively because when encoding an input video signal from a camera (referred to as natural video), there is a statistical likelihood that regions larger than the region applicable to a single MV move in a similar direction in the video sequence, and thus, in some cases, similar motion vectors derived from the MVs of adjacent regions can be used for prediction. This makes the actual MV in a given region similar or identical to the MV predicted based on the surrounding MVs. Such an MV, after entropy coding, can be represented with fewer bits compared to the number of bits used when directly encoding the MV instead of predicting it based on adjacent MVs. In some cases, MV prediction can be an instance of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself may be lossy, for example, due to rounding errors when calculating the predictor based on several surrounding MVs.
[0021] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, "High Efficiency Video Coding", December 2016). Among the various MV prediction mechanisms specified in H.265, the technique called "spatial merge" is described below.
[0022] Specifically, referring to Figure 2 As shown, the current block (201) includes samples that the encoder found during the motion search and that can be predicted from a previous block of the same size as the current block (which has been spatially shifted). The MV is not directly encoded, but rather can be derived from metadata associated with one or more reference pictures (e.g., the most recent (in decoding order) reference picture) using the MV associated with any one of five surrounding samples (denoted A0, A1, B0, B1, B2 (202 to 206 respectively)). In H.265, MV prediction can use predictors of the same reference picture as that used by its neighboring blocks. SUMMARY OF THE INVENTION
[0023] The present disclosure generally relates to video coding, and more particularly to methods and systems for signaling various motion vectors or syntax related to motion vector differences based on a magnitude-dependent adaptive resolution depending on whether a motion vector difference in inter prediction is applied.
[0024] In an example embodiment, a method for processing a current video block of a video stream is disclosed. The method includes: receiving a video stream; determining that the current video block is inter-coded based on a predicted block and a motion vector (MV), where the MV is to be derived from a reference motion vector (RMV) and a motion vector difference (MVD) for the current video block. The method further includes: in response to determining that the MVD is encoded with an adaptive MVD pixel resolution: determining a reference MVD pixel precision for the current video block; identifying a maximum allowed MVD pixel precision; determining an allowed MVD level set for the current video block based on the reference MVD pixel precision and the maximum allowed MVD pixel precision; and deriving the MVD from the video stream according to at least one MVD parameter for the current video block signaled in the video stream and the allowed MVD level set.
[0025] In the above embodiment, the reference MVD pixel precision for the current video block is specified / signaled / derived at the sequence level, picture level, frame level, superblock level, or coding block level.
[0026] In any of the above embodiments, the reference MVD pixel precision for the current video block depends on the MVD category associated with the MVD of the current video block.
[0027] In any of the above embodiments, the reference MVD pixel precision for the current video block depends on the MVD magnitude of the MVD of the current video block. In any of the above embodiments, the maximum allowed MVD pixel precision is predefined.
[0028] In any of the above embodiments, the method may further include: determining a current MVD category from a predefined set of MVD categories. Determining a set of allowed MVD levels for the MVD based on the reference MVD pixel precision and the maximum allowed MVD pixel precision may include: excluding, from a reference MVD level set determined based on the reference MVD pixel precision and the current MVD category, MVD levels associated with an MVD pixel precision equal to or higher than the maximum allowed MVD pixel precision, to determine a set of allowed MVD levels for the current video block.
[0029] In any of the above embodiments, the maximum allowed MVD pixel precision is 1 / 4 pixel.
[0030] In any of the above embodiments, MVD levels associated with a 1 / 8 pixel or higher precision are excluded from the set of allowed MVD levels for the current video block.
[0031] In any of the above embodiments, the method may further include: determining a current MVD category from a predefined set of MVD categories. When the current MVD category is equal to or lower than a threshold MVD category, MVD levels associated with a fractional MVD precision may be included in the set of allowed MVD levels regardless of the reference MVD precision.
[0032] In any of the above embodiments, the threshold MVD category may be the lowest MVD category in the predefined set of MVD categories.
[0033] In any of the above embodiments, the method may further include: determining the magnitude of the MVD, wherein an MVD level associated with an MVD precision higher than a threshold MVD precision is allowed to be used in the set of allowed MVD levels only when the magnitude of the MVD is equal to or lower than a threshold MVD magnitude.
[0034] In any of the above embodiments, the threshold MVD magnitude is 2 pixels or less.
[0035] In any of the above embodiments, the threshold MVD precision is 1 pixel.
[0036] In any of the above embodiments, an MVD level associated with a 1 / 4 pixel or higher MVD precision is allowed to be used only when the magnitude of the MVD is equal to or lower than 1 / 2 pixel. In any of the above embodiments, the maximum allowed MVD pixel precision is not greater than the reference MVD pixel precision.
[0037] In another embodiment, a method for processing a current video block of a video stream is provided. The method includes: receiving a video stream; determining that the current video block is inter-coded and associated with a plurality of reference frames; and determining, based on signaling in the video stream, whether an adaptive motion vector difference (MVD) pixel resolution is applied to at least one of the plurality of reference frames.
[0038] In the above embodiment, the signaling may include a single-bit flag to indicate whether the adaptive MVD pixel resolution is applied to all of the plurality of reference frames or not applied to any of the plurality of reference frames.
[0039] In any of the above embodiments, the signaling includes separate flags, each corresponding to one of the plurality of reference frames, to indicate whether the adaptive MVD pixel resolution is applied.
[0040] In any of the above embodiments, for each of the plurality of reference frames, the signaling includes: an implicit indication for indicating that the adaptive MVD pixel resolution is not applied when the MVD corresponding to each of the plurality of reference frames is zero; and a single-bit flag for indicating whether the adaptive MVD pixel resolution is applied when the MVD corresponding to each of the plurality of reference frames is non-zero.
[0041] In another embodiment, a method for processing a current video block of a video stream is provided. The method includes: receiving a video stream; determining, based on a prediction block and a motion vector (MV), that the current video block is inter-coded, where the MV is to be derived from a reference motion vector (RMV) and a motion vector difference (MVD) for the current video block; determining a current MVD category of the MVD from a predefined set of MVD categories; deriving at least one context for entropy decoding at least one explicit signaling in the video stream, the at least one explicit signaling included in the video stream to specify an MVD pixel resolution for at least one component of the MVD; and using the at least one context to entropy decode the at least one explicit signaling in the video stream to determine the MVD pixel resolution for at least one component of the MVD.
[0042] In the above embodiment, at least one component of the MVD may include a horizontal component and a vertical component of the MVD, and the at least one context may include two separate contexts, each associated with one of the horizontal component and the vertical component of the MVD, and the horizontal component and the vertical component are associated with separate MVD pixel resolutions.
[0043] Aspects of the present disclosure also provide a video encoding or decoding device or apparatus including circuitry configured to perform any of the above method embodiments.
[0044] Aspects of the present disclosure also provide a non - volatile computer - readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform a method of video decoding and / or encoding. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0046] Figure 1A A schematic diagram showing an exemplary subset of intra - prediction direction modes;
[0047] Figure 1B An illustration showing an exemplary intra - prediction direction;
[0048] Figure 2 A schematic diagram showing a current block for motion - vector prediction and its surrounding spatial merge candidates in one example;
[0049] Figure 3 A schematic diagram showing a simplified block diagram of a communication system (300) according to an example embodiment;
[0050] Figure 4 A schematic diagram showing a simplified block diagram of a communication system (400) according to an example embodiment;
[0051] Figure 5 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment;
[0052] Figure 6 A schematic diagram showing a simplified block diagram of a video encoder according to an example embodiment;
[0053] Figure 7 A block diagram of a video encoder according to another example embodiment;
[0054] Figure 8 A block diagram of a video decoder according to another example embodiment;
[0055] Figure 9 A scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0056] Figure 10 Another scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0057] Figure 11 Another scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0058] Figure 12 An example of partitioning a basic block into coding blocks according to an example partitioning scheme;
[0059] Figure 13 shows an example ternary partitioning scheme;
[0060] Figure 14 shows an example quadtree binary tree coded block partitioning scheme;
[0061] Figure 15 shows a scheme for partitioning a coded block into a plurality of transform blocks and the coding order of the transform blocks according to an example embodiment of the present disclosure;
[0062] Figure 16 shows another scheme for partitioning a coded block into a plurality of transform blocks and the coding order of the transform blocks according to an example embodiment of the present disclosure;
[0063] Figure 17 shows another scheme for partitioning a coded block into a plurality of transform blocks according to an example embodiment of the present disclosure;
[0064] Figure 18 shows a flowchart of a method according to an example embodiment of the present disclosure;
[0065] Figure 19 shows another flowchart of a method according to an example embodiment of the present disclosure;
[0066] Figure 20 shows another flowchart of a method according to an example embodiment of the present disclosure;
[0067] Figure 21 shows a schematic illustration of a computer system according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0068] Throughout the specification and claims, terms may have meanings implied or implicit in the context in addition to the explicitly stated meanings. The phrases "in one embodiment" or "in some embodiments" used in this application are not necessarily referring to the same embodiment, and the phrases "in another embodiment" or "in other embodiments" used in this application are not necessarily referring to different embodiments. Similarly, the phrases "in one implementation" or "in some implementations" used in this application are not necessarily referring to the same implementation, and the phrases "in another implementation" or "in other implementations" used in this application are not necessarily referring to different implementations. For example, it is intended that the claimed subject matter includes combinations of all or part of the example embodiments / implementations.
[0069] Generally speaking, terms can be understood at least in part according to their usage in context. For example, terms such as "and", "or", or "and / or" used in this application can include various meanings, which may depend at least in part on the context in which these terms are used. Generally, "or", if used in connection with a list, such as A, B, or C, is intended to mean A, B, and C (used here in an inclusive sense) as well as A, B, or C (used here in an exclusive sense). Additionally, the terms "one or more" or "at least one" used in this application, depending at least on the context, can be used to describe any feature, structure, or property in the singular sense, or can be used to describe a combination of features, structures, or properties in the plural sense. Similarly, terms such as "a" ("a" or "an") or "the" can be understood to convey singular usage or to convey plural usage, at least in part depending on the context. Additionally, the terms "based on" or "determined by" can be understood not necessarily to be intended to represent an exclusive set of factors, but rather, other factors that are not necessarily explicitly described can be allowed, again, at least in part depending on the context. Figure 3 FIG. shows a simplified block diagram of a communication system (300) according to an embodiment disclosed in this application. The communication system (300) includes a plurality of terminal devices, and the terminal devices can communicate with each other through, for example, a network (350). By way of example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected by a network (350). In Figure 3 the example of, the first pair of terminal devices (310) and (320) can perform unidirectional data transmission. For example, the terminal device (310) can encode video data (e.g., a video picture stream collected by the terminal device (310)) for transmission to another terminal device (320) through the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video pictures, and display the video pictures based on the recovered video data. Unidirectional data transmission can be implemented in applications such as media services.
[0070] In another example, a communication system (300) includes a second pair of terminal devices (330) and (340) that perform a two-way transmission of encoded video data, which can be implemented, for example, during a video conference. For two-way data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., of a video picture stream captured by the terminal device) for transmission over a network (350) to the other of the terminal devices (330) and (340). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), can decode the encoded video data to recover the video picture, and can display the video picture on an accessible display device based on the recovered video data.
[0071] In Figure 3 the example, the terminal devices (310), (320), (330), and (340) may be implemented as servers, personal computers, and smart phones, but the applicability of the underlying principles disclosed in this application is not limited thereto. The embodiments disclosed in this application can be implemented in laptop computers, notebook computers, tablet computers, media players, wearable computers, dedicated video conferencing devices, and the like. The network (350) represents any number or type of network that conveys encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, wired (wired) and / or wireless communication networks. The communication network (350) can exchange data in circuit-switched, packet-switched channels, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless explicitly explained hereinafter, the architecture and topology of the network (350) may be immaterial to the operation disclosed in this application.
[0072] As an example of the application of the subject matter disclosed in this application, Figure 4 shows how video encoders and video decoders are placed in a video streaming environment. The subject matter disclosed in this application is equally applicable to other video applications, including, for example, video conferencing, digital TV, broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, and the like.
[0073] A video streaming system may include a video acquisition subsystem (413), and the acquisition subsystem may include a video source (401) such as a digital camera, which is used to create an uncompressed video picture or image stream (402). In one example, the video picture stream (402) includes samples recorded by the digital camera of the video source 401. The video picture stream (402) is depicted as a thick line to emphasize that it has a higher data volume compared to the encoded video data (404) (or encoded video bitstream). The video picture stream (402) may be processed by an electronic device (420), and the electronic device (420) includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (404) (or encoded video bitstream (404)) is depicted as a thin line to emphasize that it has a lower data volume compared to the uncompressed video picture stream (402), and it may be stored on a streaming server (405) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems, for example, Figure 4 the client subsystem (406) and the client subsystem (408) in
[0074]
[0075] Figure 5 may access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and generates an uncompressed output video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or another presentation device (not depicted). The video decoder 410 may be configured to perform some or all of the various functions described in this disclosure. In some streaming systems, the encoded video data (404), video data (407), and video data (409) (e.g., video bitstream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the subject matter disclosed in this application may be used in the context of the VVC standard and other video coding standards.
[0074]
[0075] Figure 5 It should be noted that the electronic device (420) and the electronic device (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may further include a video encoder (not shown).
[0075] Figure 5is a block diagram of a video decoder (510) according to any embodiment disclosed herein. The video decoder (510) may be provided in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used in place of Figure 4 the video decoder (410) in the example.
[0076] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510). In the same or another embodiment, one encoded video sequence is decoded at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. Each video sequence may be associated with a plurality of video frames or images. The encoded video sequence may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data or a streaming source transmitting the encoded video data. The receiver (531) may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective processing circuits (not depicted). The receiver (531) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) may be provided between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) may be implemented as part of the video decoder (510). In other applications, the buffer memory (515) may be provided external to and separate from the video decoder (510) (not depicted). In still other applications, a buffer memory (not depicted) is provided external to the video decoder (510) to, for example, prevent network jitter, and another additional buffer memory (515) may be configured inside the video decoder (510) to, for example, handle playback timing. And when the receiver (531) receives data from a storage / forward device with sufficient bandwidth and controllability, or from an isochronous synchronous network, it may also not be necessary to configure the buffer memory (515), or the buffer memory may be made smaller. For use on a best-effort service packet network such as the Internet, it may also be necessary to have a buffer memory (515) of sufficient size, and the size of this buffer memory may be relatively large. Such a buffer memory may be implemented with an adaptive size and may be at least partially implemented in an operating system or a similar element (not depicted) external to the video decoder (510).
[0077] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (510), and potential information for controlling a display device such as a display (512) (e.g., a display screen), which may or may not be part of an electronic device (530), but may be coupled to the electronic device (530), as Figure 5 shown. The control information for the display device may be a Supplemental Enhancement Information (SEI) message or a parameter set segment of Video Usability Information (VUI) (not depicted). The parser (520) may perform parsing / entropy decoding on the encoded video sequence it receives. The entropy coding of the encoded video sequence may be performed according to a video coding technology or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (520) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The parser (520) may also extract information from the encoded video sequence, such as transform coefficients (e.g., Fourier transform), quantizer parameter values, motion vectors, and so on.
[0078] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515), thereby creating symbols (521).
[0079] Depending on the type of the encoded video picture or the encoded video picture part (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (521) may involve multiple different processing units or functional units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by the parser (520) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (520) and the multiple processing units or functional units below are not described.
[0080] In addition to the functional blocks already mentioned, the video decoder (510) can conceptually be subdivided into several functional units as described below. In actual embodiments operating under commercial constraints, many of these functional units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of clearly describing the various functions of the disclosed subject matter, the present disclosure conceptually subdivides the functional units below.
[0081] The first unit may include a scaler / inverse transform unit (551). The scaler / inverse transform unit (551) may receive quantized transform coefficients as symbols (521) and control information from the parser (520), including indication of which type of inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (551) may output a block including sample values, which may be input into an aggregator (555).
[0082] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) may generate a block having the same size and shape as the block being reconstructed by using surrounding block information that has been reconstructed and stored in the current picture buffer (558). For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some embodiments, the aggregator (555) may add the prediction information generated by the intra-prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0083] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-coded and potentially motion-compensated block. In such cases, the motion compensation prediction unit (553) may access the reference picture memory (557) to extract samples for inter-picture prediction. After motion compensating the extracted samples according to the symbols (521) belonging to the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (the output of unit 551 may be referred to as residual samples or a residual signal), thereby generating output sample information. The motion compensation prediction unit (553) obtaining the prediction samples from an address within the reference picture memory (557) may be controlled by a motion vector, and the motion vector is in the form of the symbols (521) for use by the motion compensation prediction unit (553), where the symbols (521) include, for example, X, Y components (displacements) and a reference picture component (temporal). Motion compensation may also include interpolation of sample values extracted from the reference picture memory (557) when using sub-sample accurate motion vectors, and may also be associated with a motion vector prediction mechanism, etc.
[0084] The output samples of the aggregator (555) may be employed in the loop filter unit (556) by various loop filtering techniques. The video compression technique may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and the parameters are available to the loop filter unit (556) as symbols (521) from the parser (520), but may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and in response to previously reconstructed and loop-filtered sample values. Several types of loop filters may be included as part of the loop filter unit 556 in various orders, which will be described in detail below.
[0085] The output of the loop filter unit (556) may be a sample stream that may be output to the display device (512) and stored in the reference picture memory (557) for subsequent inter-picture prediction.
[0086] Once fully reconstructed, certain encoded pictures may be used as reference pictures for future inter-picture prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (520)), the current picture buffer (558) may become part of the reference picture memory (557), and a new current picture buffer may be reallocated before starting to reconstruct subsequent encoded pictures.
[0087] The video decoder (510) may perform decoding operations according to a predetermined video compression technique adopted in a standard such as the ITU-T H.265 recommendation. In the sense that the encoded video sequence conforms to the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence may comply with the syntax specified by the video compression technique or standard used. Specifically, the profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. To comply with the standard, it is also required that the complexity of the encoded video sequence can be within the range defined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the restrictions set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0088] In some example embodiments, the receiver (531) may receive additional (redundant) data together with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant strips, redundant pictures, forward error correction codes, etc.
[0089] Figure 6 is a block diagram of a video encoder (603) according to an example embodiment disclosed in the present application. The video encoder (603) may be provided in an electronic device (620). The electronic device (620) may further include a transmitter (640) (e.g., a transmission circuit). The video encoder (603) may be used to replace Figure 4 the video encoder (403) in the example.
[0090] The video encoder (603) may receive video samples from a video source (601) (not Figure 6 part of the electronic device (620) in the example), and the video source may capture video images to be encoded by the video encoder (603). In another embodiment, the video source (601) may be implemented as part of the electronic device (620).
[0091] A video source (601) can provide a source video sequence in the form of a digital video sample stream to be encoded by a video encoder (603). The digital video sample stream can have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 YCrCb, RGB, XYZ, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) can be a storage device capable of storing previously prepared videos. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. The video data can be provided as a plurality of individual pictures or images that are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0092] According to some example embodiments, the video encoder (603) can encode and compress pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed constitutes a function of a controller (650). In some embodiments, the controller (650) can be functionally coupled to and control other functional units as described below. For simplicity, couplings are not labeled in the figure. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be used for other suitable functions that relate to optimizing the video encoder (603) for a certain system design.
[0093] In some example embodiments, a video encoder (603) may operate in an encoding loop. As a simple description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols based on an input picture to be encoded and reference pictures, e.g., a symbol stream) and an embedded (local) decoder (633) within the video encoder (603). Even though the embedded decoder 633 processes the non-entropy-coded encoded video stream of the source encoder 630, the decoder (633) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (since in the video compression techniques contemplated by the subject matter disclosed in this application, any compression between symbols and the encoded video bitstream in entropy coding can be lossless). The reconstructed sample stream (sample data) is input into a reference picture memory (634). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (634) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is used to improve the encoding quality.
[0094] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder that has been described in detail above in connection with Figure 5 the video decoder (510). However, briefly referring additionally to Figure 5 , when the symbols are available and the entropy encoder (645) and the parser (520) can encode / decode the symbols losslessly into the encoded / decoded video sequence, the entropy decoding part of the video decoder (510), including the buffer memory (515) and the parser (520), may not be fully implemented in the local decoder (633) in the encoder.
[0095] At this point, it can be observed that any decoder technology other than parsing / entropy decoding that may only exist in the decoder must also exist in the corresponding encoder in a substantially identical functional form. For this reason, this application sometimes focuses on decoder operations, which are similar to the decoding part of the encoder. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described comprehensively. A more detailed description of the encoder is only provided in certain areas or aspects below.
[0096] During operation, in some example embodiments, the source encoder (630) may perform motion-compensated predictive coding, predicting and encoding an input picture by referring to one or more previously encoded pictures designated as "reference pictures" in a video sequence. In this way, the encoding engine (632) encodes the difference (or residual) in the color channels between a pixel block of the input picture and a pixel block of the reference picture, which can be selected as a prediction reference for the input picture. The terms "residual" and its adjective "residual" are used interchangeably.
[0097] The local video decoder (633) may decode the encoded video data of a picture that can be designated as a reference picture, based on the symbols created by the source encoder (630). The operation of the encoding engine (632) may be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 6 not shown), the reconstructed video sequence is typically a copy of the source video sequence with some errors. The local video decoder (633) duplicates the decoding process that can be performed by the video decoder on the reference picture, and may store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture, which has the same content (in the absence of transmission errors) as the reconstructed reference picture to be obtained by a remote video decoder.
[0098] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or some metadata that can serve as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor (635) may operate on a per-pixel block basis of sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (635), it may be determined that the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0099] The controller (650) may manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0100] The outputs of all the above functional units may be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0101] The transmitter (640) may buffer the encoded video sequence created by the entropy encoder (645) to prepare for transmission over the communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) may combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0102] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign a certain encoded picture type to each encoded picture, which may affect the encoding techniques applicable to the corresponding picture. For example, pictures may typically be assigned to any of the following picture types:
[0103] An intra picture (I picture), which may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are aware of the variants of I pictures and their corresponding applications and characteristics.
[0104] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.
[0105] A bi - predictive picture (B picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.
[0106] The source picture can typically be spatially subdivided into multiple sample blocks (e.g., each block having 8×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the encoding assignment of the corresponding picture applied to the 'block'. For example, the blocks of an I picture can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to the already encoded blocks of the same picture (spatial prediction or intra-frame prediction). The pixel blocks of a P picture can be prediction-encoded with reference to a previously encoded reference picture by spatial prediction or by temporal prediction. The blocks of a B picture can be prediction-encoded with reference to one or two previously encoded reference pictures by spatial prediction or by temporal prediction. The source picture or the intermediate processed picture can be subdivided into other types of blocks for other purposes. The division of the encoded blocks and other types of blocks can follow or not follow the same way, which is further described below.
[0107] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (603) can perform various compression operations, including prediction coding operations that utilize the temporal and spatial redundancies in the input video sequence. Accordingly, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0108] In some example embodiments, the transmitter (640) can transmit additional data and the encoded video. The source encoder (630) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and strips, SEI messages, VUI parameter set fragments, etc.
[0109] The captured video can be multiple source pictures (video pictures) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) utilizes the spatial correlation in a given picture, while inter-picture prediction utilizes the temporal or other correlations between pictures. For example, a specific picture being encoded / decoded can be segmented into blocks, and the specific picture being encoded / decoded is called the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, it can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0110] In some example embodiments, bidirectional prediction techniques can be used for inter - picture prediction. According to such bidirectional prediction techniques, two reference pictures are used. For example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past or future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block can be jointly predicted by a combination of the first reference block and the second reference block.
[0111] In addition, merge mode techniques can be used for inter - picture prediction to improve coding efficiency.
[0112] According to some example embodiments disclosed in the present application, predictions such as inter - picture prediction and intra - picture prediction are performed in units of blocks. For example, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression. CTUs in a picture have the same size, for example, 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU can include three parallel coding tree blocks (CTBs): one luminance CTB and two chrominance CTBs. Further, each CTU can be split into one or more coding units (CUs) in a quadtree manner. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. Each of one or more of the 32×32 blocks can be further split into 4 CUs with 16×16 pixels. In some example implementations, each CU can be analyzed during the encoding process to determine a prediction type for the CU among various prediction types, such as an inter - frame prediction type or an intra - frame prediction type. In addition, depending on temporal and / or spatial predictability, a CU can be split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, prediction operations in encoding (encoding / decoding) are performed in units of prediction blocks. Splitting a CU into PUs (or PBs with different color channels) can be performed in various spatial patterns. A luminance or chrominance PB, for example, can include a matrix of sample values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and so on.
[0113] Figure 7A diagram of a video encoder (703) according to another exemplary embodiment disclosed in the present application is shown. The video encoder (703) is used to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures, and encode the processing block into an encoded picture that is part of an encoded video sequence. An exemplary video encoder (703) can be used to replace Figure 4 the video encoder (403) in the example.
[0114] For example, the video encoder (703) receives a matrix of sample values for a processing block, which is, for example, a prediction block of 8×8 samples, etc. The video encoder (703) uses, for example, rate-distortion optimization (RDO) to determine whether to use an intra mode, an inter mode, or a bi-prediction mode to encode the processing block. When it is determined to encode the processing block in the intra mode, the video encoder (703) can use intra prediction techniques to encode the processing block into the encoded picture; and when it is determined to encode the processing block in the inter mode or the bi-prediction mode, the video encoder (703) can use inter prediction or bi-prediction techniques respectively to encode the processing block into the encoded picture. In some exemplary embodiments, the merge mode can be used as a sub-mode of inter-picture prediction, where a motion vector is derived from one or more motion vector prediction values without relying on encoded motion vector components external to the prediction values. In some other exemplary embodiments, there may be motion vector components applicable to the subject block. Accordingly, the video encoder (703) may include Figure 7 components not explicitly shown in, for example, a mode decision module for determining the prediction mode of the processing block.
[0115] In Figure 7 the example of, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in Figure 7 .
[0116] The inter encoder (730) is used to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter prediction information (e.g., a description of redundant information according to inter coding techniques, a motion vector, merge mode information), and calculate an inter prediction result (e.g., a predicted block) based on the inter prediction information using any suitable technique. In some examples, the reference picture is decoded using a decoding unit 633 (e.g., Figure 6 in the example encoder 620 embedded in Figure 7a residual decoder 728, which decodes a decoded reference picture based on encoded video information, as described in further detail below.
[0117] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture, generate quantized coefficients after transformation, and also generate intra prediction information in some cases (e.g., based on intra prediction direction information of one or more intra coding techniques). The intra encoder (722) further calculates an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same picture.
[0118] The general controller (721) may be configured to determine general control data and control other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines a prediction mode of a block and provides a control signal to the switch (726) based on the prediction mode. For example, when the prediction mode is the intra mode, the general controller (721) controls the switch (726) to select an intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select intra prediction information and add the intra prediction information to the bitstream; and when the prediction mode for the block is the inter mode, the general controller (721) controls the switch (726) to select an inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select inter prediction information and add the inter prediction information to the bitstream.
[0119] The residual calculator (723) may be configured to calculate the difference (residual data) between the received block and a prediction result of a block selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) may be configured to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (724) is configured to convert the residual data from the time domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. For example, the video encoder (703) may further include a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform to generate decoded residual data. The decoded residual data may be appropriately used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) may generate a decoded block based on the decoded residual data and inter prediction information, and the intra encoder (722) may generate a decoded block based on the decoded residual data and intra prediction information. The decoded block is appropriately processed to generate a decoded picture, which may be buffered in a memory circuit (not shown) and used as a reference picture.
[0120] An entropy encoder (725) can be used to format a bitstream to produce an encoded block and perform entropy encoding. The entropy encoder (725) is used to include various information in the bitstream. For example, the entropy encoder (725) is used to obtain general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. When encoding a block in the merge submode of the inter mode or the bi - directional prediction mode, there is no residual information.
[0121] Figure 8 FIG. shows an example video decoder (810) according to another embodiment disclosed in the present application. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In one example, the video decoder (810) can be used to replace Figure 4 the video decoder (410) in the example.
[0122] In Figure 8 the example of, the video decoder (810) includes an entropy decoder (871), an inter - frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra - frame decoder (872) coupled together as shown in the example arrangement in Figure 8 ...
[0123] The entropy decoder (871) can be used to reconstruct certain symbols from the encoded picture, and these symbols represent the syntax elements that make up the encoded picture. Such symbols can include, for example, the mode of block coding (e.g., intra mode, inter mode, bi - directional prediction mode, merge submode, or another submode), prediction information (e.g., intra prediction information or inter prediction information) that can separately identify certain samples or metadata for use by the intra - frame decoder (872) or the inter - frame decoder (880) for prediction, residual information in the form of, for example, quantized transform coefficients, and so on. In one example, when the prediction mode is the inter - frame prediction mode or the bi - directional prediction mode, the inter - frame prediction information is provided to the inter - frame decoder (880); and when the prediction type is the intra - frame prediction type, the intra - frame prediction information is provided to the intra - frame decoder (872). The residual information can be inverse - quantized and provided to the residual decoder (873).
[0124] The inter - frame decoder (880) can be used to receive inter - frame prediction information and generate an inter - frame prediction result based on the inter - frame prediction information.
[0125] The intra - frame decoder (872) can be used to receive intra - frame prediction information and generate a prediction result based on the intra - frame prediction information.
[0126] The residual decoder (873) can be used to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also utilize certain control information (for obtaining the quantizer parameter QP), which may be provided by the entropy decoder (871) (the data path is not labeled as this may be just low volume control information).
[0127] The reconstruction module (874) can be used to combine, in the spatial domain, the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) to form a reconstructed block, which forms part of a reconstructed picture, which in turn may be part of a reconstructed video. It should be noted that other suitable operations such as deblocking operations may be performed to improve the visual quality.
[0128] It should be noted that any suitable technology can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In an embodiment, one or more integrated circuits can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810).
[0129] Turning to block partitioning for encoding and decoding, a general partition can start from a basic block and can follow a predefined set of rules, a specific pattern, a partition tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After partitioning or dividing the basic block following any example partitioning process or other processes described below or a combination thereof, a final set of partitions or encoded blocks can be obtained. Each of these partitions can be at one of various partition levels in the partition hierarchy and can have various shapes. Each of the partitions can be referred to as an encoded block (CB). For the various example partitioning embodiments described further below, each resulting CB can have any allowed size and partition level. Such a partition is called an encoded block because they can form units for which some basic encoding / decoding decisions can be made and the encoding / decoding parameters of these units can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the encoded block partitioning structure of the tree. The encoded block can be a luminance encoded block or a chrominance encoded block. The CB tree structure for each color can be referred to as an encoded block tree (CBT).
[0130] The encoded blocks for all color channels can be collectively referred to as a coding unit (CU). The hierarchy for all color channels can be collectively referred to as a coding tree unit (CTU). The partitioning pattern or structure for the various color channels in a CTU can be the same or different.
[0131] In some embodiments, the partitioning tree scheme or structure for the luminance channel and the chrominance channel may not have to be the same. In other words, the luminance channel and the chrominance channel can have separate coding tree structures or patterns. Further, whether the luminance channel and the chrominance channel use the same or different coding partition tree structures and the actual coding partition tree structure to be used can depend on whether the strip being encoded is a P-strip, a B-strip, or an I-strip. For example, for an I-strip, the chrominance channel and the luminance channel can have separate coding partition tree structures or coding partition tree structure patterns, while for a P-strip or a B-strip, the luminance channel and the chrominance channel can share the same coding partition tree scheme. When applying separate coding partition tree structures or patterns, the luminance channel can be partitioned into CBs by one coding partition tree structure and the chrominance channel can be partitioned into chrominance CBs by another coding partition tree structure.
[0132] In some example embodiments, a predefined partitioning pattern can be applied to the basic block. As Figure 9As shown, an example 4-way partition tree can start from a first predefined level (e.g., the 64×64 block level or other size as the basic block size), and the basic block can be hierarchically partitioned down to a predefined lowest level (e.g., the 4×4 level). For example, the basic block can go through four predefined partition options or patterns indicated by 902, 904, 906, and 908, where the partition designated as R allows for recursive partitioning because the same partition option as indicated in Figure 9 can be repeated at a lower ratio until the lowest level (e.g., the 4×4 level). In some embodiments, additional restrictions can be applied to the Figure 9 partitioning scheme. In the Figure 9 embodiment, the use of rectangular partitions (e.g., 1:2 / 2:1 rectangular partitions) can be allowed, but they may not be allowed to be recursive, while square partitions are allowed to be recursive. If needed, following the Figure 9 recursive partitioning generates the final set of coded blocks. The coding tree depth can be further defined to indicate the depth of the split from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., the 64×64 block) can be set to 0, and after the root block is further split once following Figure 9 , the coding tree depth increases by 1. For the above scheme, the maximum level or deepest level from the 64×64 basic block to the 4×4 minimum partition will be 4 (starting from level 0). This partitioning scheme can be applied to one or more color channels. The Figure 9 scheme can be followed to partition each color channel independently (e.g., the partitioning pattern or option in the predefined pattern can be independently determined for each color channel at each hierarchical level). Optionally, two or more color channels can share the Figure 9 same hierarchical pattern tree (e.g., the same partitioning pattern or option in the predefined pattern can be selected for two or more color channels at each hierarchical level).
[0133] Figure 10 shows another example predefined partitioning pattern that allows the use of recursive partitioning to form a partition tree. As shown in Figure 10 , an example 10-way partition structure or pattern can be predefined. The root block can start from a predefined level (e.g., starting from a basic block at the 128×128 level or the 64×64 level). Figure 10 The example partitioning structure includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. Figure 10 The partitioning types with 3 sub-partitions indicated as 1002, 1004, 1006, and 1008 in the second row ofFigure 10 None of the rectangular partitions are allowed to be further subdivided. The coding tree depth can be further defined to indicate the depth of the split from the root node or root block. For example, the coding tree depth of the root node or root block (e.g., 128×128 block) can be set to 0, and after the root block is further split once, the coding tree depth is increased by 1. In some embodiments, only all square partitions in 1010 are allowed to follow Figure 10 the pattern of recursive partitioning to the next level of the partition tree. In other words, for the square partitions 1002, 1004, 1006, and 1008 within the T-shaped pattern, recursive partitioning may not be allowed. If needed, the recursive partitioning process following Figure 10 generates the final set of coded blocks. Such a scheme can be applied to one or more color channels. In some embodiments, more flexibility can be added for the use of partitions below the 8×8 level. For example, 2×2 chrominance inter prediction can be used in some cases. Figure 10 In some other exemplary embodiments for coding block partitioning, a quadtree structure can be used to split a basic block or an intermediate block into quadtree partitions. Such quadtree splitting can be applied hierarchically and recursively to any square partition. Whether it is the basic block, the intermediate block, or the partition that undergoes further quadtree splitting, it can be adapted to various local characteristics of the basic block or the intermediate block / partition. The quadtree partitioning of the picture boundary can be further adjusted. For example, implicit quadtree splitting can be performed at the picture boundary such that a block will continue to be quadtree split until its size fits the picture boundary.
[0134] In some other exemplary embodiments for coding block partitioning, a quadtree structure can be used to split a basic block or an intermediate block into quadtree partitions. Such quadtree splitting can be applied hierarchically and recursively to any square partition. Whether it is the basic block, the intermediate block, or the partition that undergoes further quadtree splitting, it can be adapted to various local characteristics of the basic block or the intermediate block / partition. The quadtree partitioning of the picture boundary can be further adjusted. For example, implicit quadtree splitting can be performed at the picture boundary such that a block will continue to be quadtree split until its size fits the picture boundary.
[0135] In some other example embodiments, a hierarchical binary partitioning from a basic block can be used. For such a scheme, the basic block or an intermediate level block can be partitioned into two partitions. The binary partitioning can be horizontal or vertical. For example, a horizontal binary partitioning can split the basic block or intermediate block into equal left and right partitions. Similarly, a vertical binary partitioning can split the basic block or intermediate block into equal upper and lower partitions. Such binary partitioning can be hierarchical and recursive. A decision can be made at each of the basic blocks or intermediate blocks as to whether the binary partitioning scheme should continue, and if the scheme continues further, a decision can be made as to whether a horizontal binary partitioning or a vertical binary partitioning should be used. In some embodiments, the further partitioning can stop at a predefined minimum partitioning size (in one or two dimensions). Optionally, once a predefined partitioning level or depth starting from the basic block is reached, the further partitioning can stop. In some embodiments, the aspect ratio of the partitioning can be restricted. For example, the aspect ratio of the partitioning can be not less than 1:4 (or greater than 4:1). Thus, a vertical bar partitioning having a vertical to horizontal aspect ratio of 4:1 can be further vertically binary partitioned only into upper and lower partitions, each having a vertical to horizontal aspect ratio of 2:1.
[0136] In still some other examples, as Figure 13 shown, a ternary partitioning scheme can be used to partition a basic block or any intermediate block. The ternary pattern can be implemented vertically as shown by 1302 in Figure 13 or horizontally as shown by 1304 in Figure 13 . Although Figure 13 the example vertical or horizontal split ratios in Figure 13 are shown as 1:2:1, other ratios can be predefined. In some embodiments, two or more different ratios can be predefined. Such a ternary partitioning scheme can be used to compensate for a quadtree or binary partitioning structure because such a ternary tree partitioning can collect an object located at the center of the block in a continuous partition, while a quadtree and a binary tree always split along the center of the block and thus split the object into separate partitions. In some embodiments, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid having additional transformations.
[0137] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree and binary partitioning schemes can be combined to partition a basic block into a quadtree - binary tree (QTBT) structure. In such a scheme, the basic block or intermediate block / partition can be either quadtree split or binary split, subject to a predefined set of conditions if specified. A specific example is illustrated in Figure 14 . In Figure 14In the example of [description], the basic block is first divided into four partitions by a quadtree, as shown in 1402, 1404, 1406, and 1408. Thereafter, each of the resulting partitions is either divided into four additional partitions by a quadtree (such as 1408), or divided into two additional partitions by a binary split at the next level (horizontally or vertically, such as 1402 or 1406, for example both are symmetric), or not divided (such as 1404). For square partitions, binary or quadtree splitting can be recursively allowed, as shown in the overall example partitioning pattern of 1410 and the corresponding tree structure / representation in 1420, where solid lines represent quadtree splitting and dashed lines represent binary splitting. Flags can be used for each binary split node (non-leaf binary partition) to indicate whether the binary split is horizontal or vertical. For example, as shown in 1420, consistent with the partitioning structure of 1410, the flag "0" can represent a horizontal binary split, and the flag "1" can represent a vertical binary split. For quadtree split partitions, there is no need to indicate the split type because quadtree splitting always splits the block or partition horizontally and vertically to produce 4 sub-blocks / partitions of equal size. In some embodiments, the flag "1" can represent a horizontal binary split, and the flag "0" can represent a vertical binary split.
[0138] In some example embodiments of QTBT, the quadtree and binary splitting rule sets can be represented by the following predefined parameters and the corresponding functions associated with them:
[0139] - CTU size: the size of the root node of the quadtree (the size of the basic block)
[0140] - MinQTSize: the minimum allowed quadtree leaf node size
[0141] - MaxBTSize: the maximum allowed binary tree root node size
[0142] - MaxBTDepth: the maximum allowed binary tree depth
[0143] - MinBTSize: the minimum allowed binary tree leaf node size
[0144] In some example embodiments of the QTBT partition structure, the CTU size can be set to 128×128 luma samples having two corresponding 64×64 chroma sample blocks (when considering and using example chroma subsampling), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. Quadtree partitioning can be first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from its minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If a node is 128×128, it will not be first split by the binary tree because the size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes not exceeding MaxBTSize can be partitioned by the binary tree. In Figure 14 the example of Figure 14 , the base block is 128×128. According to a predefined set of rules, the base block can only be split by the quadtree. The partition depth of the base block is 0. Each of the four resulting partitions is 64×64, which does not exceed MaxBTSize, and can be further split by the quadtree or binary tree at level 1. This process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When the width of a binary tree node equals MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when the height of a binary tree node equals MinBTSize, further vertical splitting is not considered.
[0145] In some example embodiments, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure or separate QTBT structures for luma and chroma. For example, for P slices and B slices, the luma CTB and chroma CTB in a CTU can share the same QTBT structure. However, for I slices, the luma CTB can be partitioned into CUs by the QTBT structure, and the chroma CTB can be partitioned into chroma CUs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice. For example, an I slice can consist of coded blocks of the luma component or coded blocks of two chroma components, and a CU in a P slice or B slice can consist of coded blocks of all three color components.
[0146] In some other embodiments, the QTBT scheme can be supplemented with the above-mentioned ternary scheme. This embodiment can be referred to as a multi-type-tree (MTT) structure. For example, in addition to the binary splitting of nodes, Figure 13One of the three-way partitioning patterns. In some embodiments, only square nodes can be three-way split. An additional flag can be used to indicate whether the three-way partitioning is horizontal or vertical.
[0147] The design of two-level or multi-level trees such as QTBT embodiments and QTBT embodiments supplemented by three-way splitting can be mainly driven by reducing complexity. Theoretically, the complexity of traversing the tree is T D , where T represents the number of splitting types and D is the depth of the tree. A trade-off can be made by using multiple types (T) while reducing the depth (D) simultaneously.
[0148] In some embodiments, a CB can be further partitioned. For example, for the purpose of intra prediction or inter prediction during the encoding and decoding processes, a CB can be further partitioned into multiple prediction blocks (PBs). In other words, a CB can be further divided into different sub-partitions in which separate prediction decisions / configurations can be made. At the same time, for the purpose of depicting the level at which the transformation or inverse transformation of video data is performed, a CB can be further partitioned into multiple transform blocks (TBs). The partitioning schemes of CBs into PBs and TBs can be the same or different. For example, each partitioning scheme can be performed using its own process based on various characteristics of the video data, for example. In some example embodiments, the PB and TB partitioning schemes can be independent. In some other example embodiments, the PB and TB partitioning schemes and boundaries can be related. In some embodiments, for example, TBs can be partitioned after PB partitioning. Specifically, each PB (after being determined after partitioning of the coding block) can then be further partitioned into one or more TBs. For example, in some embodiments, a PB can be split into one, two, four, or other numbers of TBs.
[0149] In some embodiments, to partition a basic block into coding blocks and further into prediction blocks and / or transform blocks, the luminance channel and the chrominance channels may be processed in different ways. For example, in some embodiments, for the luminance channel, partitioning of coding blocks into prediction blocks and / or transform blocks may be allowed, while for one or more chrominance channels, partitioning of coding blocks into prediction blocks and / or transform blocks may not be allowed. In such an embodiment, the transformation and / or prediction of luminance blocks may thus be performed only at the coding block level. As another example, the minimum transform block size for the luminance channel and one or more chrominance channels may be different. For example, coding blocks for the luminance channel may be allowed to be partitioned into smaller transform blocks and / or prediction blocks than for the chrominance channels. As yet another example, the maximum depth of partitioning coding blocks into transform blocks and / or prediction blocks may be different between the luminance channel and the chrominance channels. For example, coding blocks for the luminance channel may be allowed to be partitioned into deeper transform blocks and / or prediction blocks than for one or more chrominance channels. As a specific example, a luminance coding block may be partitioned into transform blocks of multiple sizes, which may be represented by a partitioning that recurs down up to 2 levels, and transform block shapes such as square, 2:1 / 1:2, and 4:1 / 1:4 and transform block sizes from 4×4 to 64×64 may be allowed. However, for chrominance blocks, only the maximum possible transform blocks specified for luminance blocks may be allowed.
[0150] In some example embodiments for partitioning coding blocks into PBs, the depth, shape, and / or other characteristics of the PB partitioning may depend on whether the PB is intra-coded or inter-coded.
[0151] Partitioning of coding blocks (or prediction blocks) into transform blocks may be implemented in various example schemes, including but not limited to recursively or non-recursively quadtree splitting and predefined pattern splitting, with additional consideration for transform blocks at the boundaries of coding blocks or prediction blocks. Generally, the resulting transform blocks may be at different splitting levels, may not have the same size, and may not need to be square-shaped (e.g., they may be rectangles with some allowed sizes and aspect ratios). Other examples are described in more detail below with respect to Figure 15 、 Figure 16 and Figure 17 More examples are described in more detail.
[0152] However, in some other embodiments, the CBs obtained via any of the above partitioning schemes can be used as basic blocks or minimum coding blocks for prediction and / or transformation. In other words, for the purpose of performing inter-frame prediction / intra-frame prediction and / or for the purpose of transformation, no further splitting is performed. For example, the CBs obtained from the above QTBT scheme can be directly used as units for performing prediction. Specifically, such a QTBT structure removes the concept of multiple partitioning types, i.e., it removes the distinction between CUs, PUs, and TUs, providing greater flexibility for the CU / CB partitioning shapes as described above. In such a QTBT block structure, the CU / CB can have a square or rectangular shape. The leaf nodes of such a QTBT are used as units for prediction and transformation processing without any further partitioning. This means that the CUs, PUs, and TUs have the same block size in this example QTBT coding block structure.
[0153] The above various CB partitioning schemes and the further partitioning of CBs into PBs and / or TBs (excluding PB / TB partitioning) can be combined in any way. The following specific embodiments are provided as non-limiting examples.
[0154] The following describes specific example embodiments of coding block and transform block partitioning. In such an example embodiment, a recursive quadtree split or a predefined split pattern described above (e.g., Figure 9 and Figure 10 the patterns therein) can be used to split a basic block into coding blocks. At each level, whether the further quadtree split of a particular partition should continue can be determined by local video data characteristics. The resulting CBs can be at various quadtree split levels and have various sizes. A decision can be made at the CB level (or CU level, for all three color channels) as to whether to encode a picture region using inter-frame picture (temporal) prediction or intra-frame picture (spatial) prediction. Each CB can be further split into one, two, four, or other numbers of PBs according to a predefined PB split type. Within a PB, the same prediction process can be applied, and relevant information can be transmitted to the decoder on the basis of the PB. After obtaining a residual block by applying a prediction process based on the PB split type, the CB can be partitioned into TBs according to another quadtree structure similar to the coding tree used for the CB. In this specific embodiment, the CB or TB can be, but is not limited to, a square shape. Further, in this specific example, the PB can be square or rectangular for inter-frame prediction and can be only square for intra-frame prediction. The coding block can be split into, for example, four square TBs. Each TB can be further recursively split (using quadtree split) into smaller TBs, called residual quadtree (RQT).
[0155] Another example embodiment for partitioning a basic block into CBs, PBs, and / or TBs is further described below. For example, instead of using a multi-partition unit type such as Figure 9 or Figure 10 as shown, a quadtree with a nested multi-type tree using a binary and ternary split segmentation structure (e.g., QTBT or QTBT with ternary split as described above) can be used. The distinction between CBs, PBs, and TBs can be dispensed with (i.e., partition CBs into PBs and / or TBs, and partition PBs into TBs) unless a CB with a size too large for the maximum transform length is required, in which case such a CB may need to be further split. This example partitioning scheme can be designed to provide more flexibility in the CB partition shape such that both prediction and transformation can be performed at the CB level without further partitioning. In such a coding tree structure, a CB can have a square or rectangular shape. Specifically, a coding tree block (CTB) can first be partitioned by a quadtree structure. Then, the quadtree leaf nodes can be further partitioned by a nested multi-type tree structure. Figure 11 An example of a nested multi-type tree structure using binary or ternary split is shown in Figure 11 . Specifically, the example multi-type tree structure of Figure 11 includes four split types, called vertical binary split (SPLIT_BT_VER) (1102), horizontal binary split (SPLIT_BT_HOR) (1104), vertical ternary split (SPLIT_TT_VER) (1106), and horizontal ternary split (SPLIT_TT_HOR) (1108). Then, a CB corresponds to a leaf of the multi-type tree. In this example embodiment, this segmentation is used for prediction and transformation processing without any further partitioning unless the CB is too large for the maximum transform length. This means that in most cases, CBs, PBs, and TBs have the same block size in a quadtree with a nested multi-type tree coding block structure. Exceptions occur when the maximum supported transform length is less than the width or height of the color components of a CB. In some embodiments, in addition to binary or ternary split,
[0156] Figure 12 A specific example of a quadtree with a nested multi-type tree coding block structure (including quadtree, binary, and ternary split options) with block partitioning for one basic block is shown in Figure 12 . More specifically, Figure 11 shows that the basic block 1200 is split by a quadtree into four square partitions 1202, 1204, 1206, and 1208. A decision is made for each quadtree split partition to further use Figure 12In the example, partition 1204 is not further divided. Each of partitions 1202 and 1208 undergoes another quadtree division. For partition 1202, the upper-left, upper-right, lower-left, and lower-right partitions of the second-level quadtree division respectively undergo the third-level division of the quadtree, Figure 11 the horizontal binary division 1104, non-division, and Figure 11 the horizontal ternary division 1108. Partition 1208 undergoes another quadtree division, and the upper-left, upper-right, lower-left, and lower-right partitions of the second-level quadtree division respectively undergo Figure 11 the third-level division of the vertical ternary division 1106, non-division, non-division, and Figure 11 the horizontal binary division 1104. The two sub-partitions of the third-level upper-left partition of 1208 are further divided according to Figure 11 the horizontal binary division 1104 and the horizontal ternary division 1108 respectively. Partition 1206 undergoes a second-level division pattern following Figure 11 the vertical binary division 1102, and is divided into two partitions, which are further divided at the third level according to Figure 11 the horizontal ternary division 1108 and the vertical binary division 1102 respectively. According to Figure 11 the horizontal binary division 1104, a fourth-level division is further applied to one of them.
[0157] For the above specific example, the maximum luminance transform size can be 64×64, and the maximum supported chrominance transform size can be different from that of luminance. For example, it can be 32×32. Even though the example CB in the above Figure 12 is generally not further divided into smaller PBs and / or TBs, when the width or height of a luminance coding block or a chrominance coding block is greater than the maximum transform width or height, the luminance coding block or the chrominance coding block can be automatically divided in the horizontal and / or vertical directions to meet the transform size limit in that direction.
[0158] In the above specific example for partitioning a basic block into CBs, as described above, the coding tree scheme can support the ability for luminance and chrominance to have separate block tree structures. For example, for P slices and B slices, the luminance CTB and the chrominance CTB in a CTU can share the same coding tree structure. For example, for I slices, luminance and chrominance can have separate coding block tree structures. When applying separate block tree structures, the luminance CTB can be partitioned into luminance CBs through one coding tree structure, and the chrominance CTB can be partitioned into chrominance CBs through another coding tree structure. This means that a CU in an I slice can be composed of coding blocks of the luminance component or coding blocks of the two chrominance components, and a CU in a P slice or a B slice is always composed of coding blocks of all three color components, unless the video is monochromatic.
[0159] When a coding block is further partitioned into multiple transform blocks, the transform blocks therein can be sorted in the bitstream in various orders or scan patterns. Example embodiments for partitioning a coding block or a prediction block into transform blocks and the coding order of the transform blocks are described in further detail below. In some example embodiments, as described above, transform partitioning can support transform blocks of multiple shapes (e.g., 1:1 (square), 1:2 / 2:1, and 1:4 / 4:1), where the transform block size ranges from, for example, 4×4 to 64×64. In some embodiments, if the coding block is less than or equal to 64×64, transform block partitioning can be applied only to the luminance component, such that for chrominance blocks, the transform block size is the same as the coding block size. Otherwise, if the width or height of the coding block is greater than 64, both the luminance and chrominance coding blocks can be implicitly segmented into multiples of min(W,64)x min(H,64) and min(W,32)x min(H,32) transform blocks, respectively.
[0160] In some example embodiments of transform block partitioning, for both intra-coded blocks and inter-coded blocks, the coding block can be further partitioned into multiple transform blocks with a partitioning depth of up to a predefined number of levels (e.g., 2 levels). The transform block partitioning depth and the partitioning size can be related. For some example embodiments, the mapping from the transform size at the current depth to the transform size at the next depth is as shown in Table 1 below.
[0161] Table 1: Transform Partitioning Size Settings
[0162]
[0163] Based on the example mapping in Table 1, for a 1:1 square block, the next-level transform split can create four 1:1 square sub-transform blocks. The transform partitioning can stop, for example, at 4×4. Thus, the transform size of 4×4 at the current depth corresponds to the same size of 4×4 at the next depth. In the example of Table 1, for a 1:2 / 2:1 non-square block, the next-level transform split can create two 1:1 square sub-transform blocks, while for a 1:4 / 4:1 non-square block, the next-level transform split can create two 1:2 / 2:1 sub-transform blocks.
[0164] In some example embodiments, for the luminance component of an intra-coded block, additional restrictions can be applied to transform block partitioning. For example, for each level of transform partitioning, all its sub-transform blocks can be restricted to have equal sizes. For example, for a 32×16 coding block, the level 1 transform split creates two 16×16 sub-transform blocks, and the level 2 transform split creates eight 8×8 sub-transform blocks. In other words, the second-level split must be applied to all first-level sub-blocks to keep the transform unit sizes equal. Figure 15An example of transform block partitioning for an intra-coded square block according to Table 1 is shown, along with the coding order indicated by the arrow diagram. Specifically, 1502 shows the square coded block. A first-level partitioning into 4 equally-sized transform blocks according to Table 1 is shown in 1504, with the coding order indicated by the arrow. A second-level partitioning of all the first-level equally-sized blocks into 16 equally-sized transform blocks according to Table 1 is shown in 1506, with the coding order indicated by the arrow.
[0165] In some example embodiments, for the luminance component of an inter-coded block, the above restrictions on intra-coding may not be applied. For example, after the first-level transform partitioning, any one of the sub-transform blocks may be further independently re-partitioned one level. Thus, the resulting transform blocks may or may not have the same size. Figure 16 An example of partitioning an inter-coded block into transforms with their coding order is shown. In Figure 16 the example, according to Table 1, the inter-coded block 1602 is partitioned into transform blocks in two levels. At the first level, the inter-coded block is partitioned into four equally-sized transform blocks. Then, as shown in 1604, only one (not all) of the four transform blocks is further partitioned into four sub-transform blocks, resulting in a total of 7 transform blocks with two different sizes. The example coding order of these 7 transform blocks is shown by the arrow in Figure 16 1604.
[0166] In some example embodiments, some additional restrictions on transform blocks may be applied for one or more chrominance components. For example, for one or more chrominance components, the transform block size may be as large as the coded block size, but not less than a predefined size, such as 8×8.
[0167] In some other example embodiments, for a coded block with a width (W) or height (H) greater than 64, both the luminance coded block and the chrominance coded block may be implicitly partitioned into multiples of min(W,64)×min(H,64) and min(W,32)×min(H,32) transform units, respectively. Here, in the present disclosure, "min(a,b)" may return the smaller value between a and b.
[0168] Figure 17 Another alternative example scheme for partitioning a coded block or a prediction block into transform blocks is further shown. As Figure 17 shown, instead of using recursive transform partitioning, a predefined set of partitioning types may be applied to the coded block according to the transform type of the coded block. In the specific example shown in Figure 17 , one of 6 example partitioning types may be applied to partition the coded block into various numbers of transform blocks. This scheme for generating transform block partitioning may be applied to a coded block or a prediction block.
[0169] More specifically, Figure 17 the partitioning scheme of provides up to 6 example partitioning types for any given transform type (the transform type refers to the type of the main transform, such as ADST and others). In this scheme, a transform partitioning type can be assigned to each coding block or prediction block based on (for example) the rate-distortion cost. In one example, the transform partitioning type assigned to a coding block or prediction block can be determined based on the transform type of the coding block or prediction block. As Figure 17 shown by the 6 transform partitioning types shown, a specific transform partitioning type can correspond to a transform block split size and a split pattern. The correspondence between various transform types and various transform partitioning types can be predefined. An example is shown below, where the capital markings indicate the transform partitioning types that can be assigned to a coding block or prediction block based on the rate-distortion cost:
[0170] ·PARTITION_NONE: Assign a transform size equal to the block size.
[0171] ·PARTITION_SPLIT: Assign a transform size whose width is 1 / 2 of the block size width and height is 1 / 2 of the block size height.
[0172] ·PARTITION_HORZ: Assign a transform size whose width is the same as the block size width and height is 1 / 2 of the block size height.
[0173] ·PARTITION_VERT: Assign a transform size whose width is 1 / 2 of the block size width and height is the same as the block size height.
[0174] ·PARTITION_HORZ4: Assign a transform size whose width is the same as the block size width and height is 1 / 4 of the block size height.
[0175] ·PARTITION_VERT4: Assign a transform size whose width is 1 / 4 of the block size width and height is the same as the block size height.
[0176] In the above example, all the transform partitioning types shown as Figure 17 include a unified transform size for the partitioned transform blocks. This is only an example and not a limitation. In some other embodiments, mixed transform block sizes can be used for the partitioned transform blocks of a specific partitioning type (or mode).
[0177] The PB (or CB, also referred to as PB when not further partitioned into prediction blocks) obtained from any of the above partitioning schemes can then become individual blocks for encoding via intra prediction or inter prediction. For inter prediction of the current PB, the residual between the current block and the prediction block can be generated, encoded, and included in the encoded bitstream.
[0178] Inter prediction can be implemented, for example, in a single reference mode or a composite reference mode. In some embodiments, a skip flag can first be included in the bitstream (or at a higher level) for the current block to indicate whether the current block is inter-coded and not skipped. If the current block is inter-coded, another flag can be further included in the bitstream as a signal to indicate whether a single reference mode or a composite reference mode is used for the prediction of the current block. For the single reference mode, one reference block can be used to generate the prediction block for the current block. For the composite reference mode, two or more reference blocks (for example) can be used to generate the prediction block by weighted averaging. The composite reference mode can refer to a mode with more than one reference, a two-reference mode, or a multi-reference mode. One or more reference blocks can be identified using one reference frame index or multiple reference frame indexes additionally used with corresponding one motion vector or multiple motion vectors, where the one motion vector or multiple motion vectors indicate one or more shifts in position (e.g., in horizontal and vertical pixels) between the one or more reference blocks and the current block. For example, in the single reference mode, the inter prediction block for the current block can be generated from a single reference block identified by one motion vector in a reference frame, while for the composite reference mode, the prediction block can be generated by weighted averaging of two reference blocks in two reference frames indicated by two reference frame indexes and two corresponding motion vectors. One or more motion vectors can be encoded and included in the bitstream in various ways.
[0179] In some embodiments, an encoding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures may be kept in the DPB for display (in a decoding system), and some images / pictures in the DPB may be used as reference frames for inter-frame prediction (in a decoding system or an encoding system). In some embodiments, the reference frames in the DPB may be marked as short-term references or long-term references for the current image being encoded or decoded. For example, a short-term reference frame may include a frame used for inter-frame prediction of blocks in a predefined number (e.g., 2) of subsequent video frames in or closest to the current frame in decoding order. A long-term reference frame may include frames in the DPB that may be used to predict image blocks in frames that are more than a predefined number away from the current frame in decoding order. Information about such tags for short-term and long-term reference frames may be referred to as a reference picture set (RPS), and may be added to the header of each frame in the encoded bitstream. Each frame in the encoded video stream may be identified by a picture order count (POC), which is numbered in an absolute manner according to the playback sequence or is related to a group of pictures starting from an I-frame, for example.
[0180] In some example embodiments, one or more reference picture lists may be formed based on the information in the RPS, which contain the identities of short-term and long-term reference frames for inter-frame prediction. For example, a single picture reference list may be formed for uni-directional inter-frame prediction, denoted as the L0 reference (or reference list 0), while two picture reference lists may be formed for bi-directional inter-frame prediction, denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be sorted in various predefined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. Uni-directional inter-frame prediction may be in a single-reference mode or in a composite-reference mode when multiple references used for generating a prediction block by weighted averaging in a composite prediction mode are on the same side of the block to be predicted. Bi-directional inter-frame prediction may only be in a composite mode since bi-directional inter-frame prediction involves at least two reference blocks.
[0181] In some embodiments, a merge mode (MM) for inter prediction may be implemented. Generally, for the merge mode, the motion vector in a single reference prediction for the current PB or one or more motion vectors in a composite reference prediction may be derived from one or more other motion vectors, rather than being independently computed and signaled. For example, in an encoding system, one or more current motion vectors for the current PB may be represented by one or more differences between one or more current motion vectors and one or more previously encoded motion vectors (referred to as reference motion vectors). One or more such differences in one or more motion vectors rather than the entire one or more current motion vectors may be encoded and included in the bitstream, and may be linked to one or more reference motion vectors. Accordingly, in a decoding system, one or more motion vectors corresponding to the current PB may be derived based on one or more decoded motion vector differences and one or more decoded reference motion vectors linked to them. As a specific form of general merge mode (MM) inter prediction, this inter prediction based on one or more motion vector differences may be referred to as merge mode with motion vector difference (MMVD). Thus, generally MM may be implemented, or specifically MMVD may be implemented, to exploit the correlation between motion vectors associated with different PBs to improve encoding efficiency. For example, adjacent PBs may have similar motion vectors, and thus the MVD may be small and may be efficiently encoded. For another example, motion vectors may be associated with blocks that are similar / located at similar positions in time (between frames) and in space.
[0182] In some example embodiments, an MM flag may be included in the bitstream during the encoding process to indicate whether the current PB is in the merge mode. Additionally, or alternatively, an MMVD flag may be included in the bitstream and signaled during the encoding process to indicate whether the current PB is in the MMVD mode. The MM and / or MMVD flag or indicator may be provided at the PB level, CB level, CU level, CTB level, CTU level, slice level, picture level, etc. For a specific example, both an MM flag and an MMVD flag may be included for the current CU, and the MMVD flag may be signaled immediately after the skip flag and the MM flag to specify whether the MMVD mode is used for the current CU.
[0183] In some example embodiments of MMVD, a list of reference motion vectors (RMVs) or MV predictor candidates for motion vector prediction can be formed for the block being predicted. The list of RMV candidates can contain a predetermined number (e.g., 2) of MV predictor candidate blocks, whose motion vectors can be used to predict the current motion vector. The RMV candidate blocks can include blocks selected from adjacent blocks and / or temporal blocks in the same frame (e.g., blocks at the same position in the previous or next frame of the current frame). These options represent blocks that may have a similar or identical motion vector to the current block at a spatial or temporal position relative to the current block. The size of the list of MV predictor candidates can be predetermined. For example, the list can contain two or more candidates. For a candidate block to be on the list of RMV candidates, e.g., the candidate block may need to have the same reference frame (or multiple frames) as the current block, must exist (e.g., a boundary check needs to be performed when the current block is close to the edge of the frame), and must have been encoded during the encoding process and / or decoded during the decoding process. In some embodiments, the list of merge candidates can be first filled with spatially adjacent blocks (scanned in a specific predefined order) (if available and meeting the above conditions), and then filled with temporal blocks (if space is still available in the list). For example, adjacent RMV candidate blocks can be selected from the left and top blocks of the current block. The list of RMV predictor candidates can be dynamically formed as a dynamic reference list (DRL) at various levels (sequence, picture, frame, slice, superblock, etc.). The DRL can be signaled in the bitstream.
[0184] In some embodiments, the actual MV predictor candidate that is used as the reference motion vector for predicting the motion vector of the current block can be signaled. In the case where the RMV candidate list contains two candidates, a 1-bit flag called the merge candidate flag can be used to indicate the selection of the reference merge candidate. For a current block predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor can be associated with a reference motion vector from the merge candidate list. The encoder can determine which RMV candidate more closely predicts the current encoded block and signal this selection as an index to the DRL.
[0185] In some example embodiments of MMVD, after selecting an RMV candidate and using it as the base motion vector prediction value for the motion vector to be predicted, a motion vector difference (MVD or delta MV, representing the difference between the motion vector to be predicted and the reference candidate motion vector) can be calculated in the encoding system. Such an MVD can include information representing both the magnitude of the MV difference and the direction of the MV difference, both of which can be signaled in the bitstream. The motion difference magnitude and the motion difference direction can be signaled in various ways.
[0186] In some example embodiments of MMVD, a distance index may be used to specify the magnitude information of the motion vector difference and to indicate one of a set of predefined offsets representing a predefined motion vector difference from a starting point (reference motion vector). The MV offset according to the signaled index may then be added to the horizontal or vertical component of the starting (reference) motion vector. Whether the horizontal or vertical component of the reference motion vector should be offset may be determined by the direction information of the MVD. An example predefined relationship between the distance index and the predefined offsets is specified in Table 2.
[0187] Table 2 - Example relationship between distance index and predefined MV offsets
[0188]
[0189] In some example embodiments of MMVD, a direction index may be further signaled and used to represent the direction of the MVD relative to the reference motion vector. In some embodiments, the direction may be restricted to either the horizontal or vertical direction. Example 2-bit direction indexes are shown in Table 3. In the example of Table 3, the interpretation of the MVD may vary according to the information of the starting / reference MV. For example, when the starting / reference MV corresponds to a single-prediction block or corresponds to a bi-prediction block and both reference frame lists point to the same side of the current picture (i.e., the POCs of both reference pictures are greater than the POC of the current picture, or both are less than the POC of the current picture), the signs in Table 3 may specify the sign (direction) of the MV offset added to the starting / reference MV. When the starting / reference MV corresponds to a bi-prediction block and the reference pictures are on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the signs in Table 3 may specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset for the MV corresponding to the reference picture in picture reference list 1 may have an opposite value (opposite sign for the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the signs in Table 3 may specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign of the offset for the reference MV associated with picture reference list 0 has an opposite value.
[0190] Table 3 - Example embodiments of signs for MV offsets specified by direction index
[0191] Direction IDX 00 01 10 11 X-axis (horizontal) + - Not applicable Not applicable Y-axis (vertical) Not applicable Not applicable + -
[0192] In some example embodiments, the MVD may be scaled according to the difference in POC in each direction. If the difference in POC in the two lists is the same, no scaling is required. Otherwise, if the difference in POC in reference list 0 is greater than the difference in POC in reference list 1, the MVD for reference list 1 is scaled. If the POC difference in reference list 1 is greater than the POC difference in reference list 0, the MVD for list 0 may be scaled in the same manner. If the starting MV is single predicted, the MVD is added to the available or reference MV.
[0193] In some example embodiments of MVD coding and signaling for bidirectional composite prediction, in addition to or alternatively separately coding and signaling two MVDs, symmetric MVD coding may be implemented such that only one MVD needs to be signaled and the other MVD can be derived from the signaled MVD. In such an embodiment, the motion information that includes the reference picture indices of list-0 and list-1 is signaled. However, only the MVD associated with, for example, reference list-0 is signaled, and the MVD associated with reference list-1 is derived without being signaled. Specifically, at the slice level, a flag may be included in the bitstream, called "mvd_l1_zero_flag", for indicating whether reference list-1 is not signaled in the bitstream. If this flag is 1, indicating that reference list-1 is equal to 0 (and thus not signaled), then the bidirectional prediction flag, called "BiDirPredFlag", may be set to 0, meaning that there is no bidirectional prediction. Otherwise, if mvd_l1_zero_flag is zero, then if the nearest reference picture in list-0 and the nearest reference picture in list-1 form a forward and backward reference picture pair or a backward and forward reference picture pair, BiDirPredFlag may be set to 1, and both the list-0 and list-1 reference pictures are short-term reference pictures. Otherwise BiDirPredFlag is set to 0. A BiDirPredFlag of 1 may indicate that a symmetric mode flag is additionally signaled in the bitstream. When BiDirPredFlag is 1, the decoder may extract the symmetric mode flag from the bitstream. For example, the symmetric mode flag may be signaled (if needed) at the CU level, and it may indicate whether the symmetric MVD coding mode is being used for the corresponding CU. When the symmetric mode flag is 1, indicating the use of the symmetric MVD coding mode, only the reference picture indices of both list-0 and list-1 (called "mvp_l0_flag" and "mvp_l1_flag") are signaled with the MVD associated with list-0 (called "MVD0"), and the other motion vector difference "MVD1" will be derived rather than being signaled. For example, MVD1 may be derived as -MVD0. Thus, only one MVD is signaled in the example symmetric MVD mode. In some other example embodiments for MV prediction, a coordinated scheme may be used to implement general merge mode, MMVD, and some other types of MV prediction for both single reference mode and composite reference mode MV prediction. Various syntax elements may be used to signal the way of predicting the MV for the current block.
[0194] For example, for a single reference mode, the following MV prediction modes may be signaled:
[0195] NEARMV - Directly use one of the motion vector predictors (MVPs) in the list indicated by the dynamic reference list (DRL) index without any MVD.
[0196] NEWMV - Use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and apply an increment to the MVP (e.g., use MVD).
[0197] GLOBALMV - Use a motion vector based on frame-level global motion parameters.
[0198] Similarly, for the composite reference inter-prediction mode using two reference frames corresponding to two MVs to be predicted, the following MV prediction modes can be signaled:
[0199] NEAR_NEARMV - For each of the two MVs to be predicted, directly use one of the motion vector predictors (MVPs) in the list signaled by the DRL index without MVD.
[0200] NEAR_NEWMV - To predict the first of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference MV without MVD; to predict the second of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference MV and combine it with an additionally signaled incremental MV (MVD).
[0201] NEW_NEARMV - To predict the second of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference MV without MVD; to predict the first of two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference MV and combine it with an additionally signaled incremental MV (MVD).
[0202] NEW_NEWMV - Use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference MV and combine it with an additionally signaled incremental MV to predict each of the two MVs.
[0203] GLOBAL_GLOBALMV - Use the MVs from each reference based on their frame-level global motion parameters.
[0204] Accordingly, the above term "NEAR" refers to MV prediction using a reference MV without MVD as a general merge mode, while the term "NEW" refers to MV prediction involving using a reference MV and offsetting it with a signaled MVD as in the MMVD mode. For composite inter prediction, the above reference basic motion vector and motion vector difference can typically be different or independent between two references, even if they can be correlated, and this correlation can be exploited to reduce the amount of information needed to signal the two motion vector differences. In such cases, joint signaling of the two MVDs can be implemented and indicated in the bitstream.
[0205] The above dynamic reference list (DRL) can be used to hold a set of indexed motion vectors that are dynamically maintained and considered as candidate motion vector predictors.
[0206] In some example embodiments, a predefined resolution of the MVD can be allowed. For example, a motion vector precision (or accuracy) of 1 / 8 pixel can be allowed. The MVD in the above various MV prediction modes can be constructed and signaled in various ways. In some embodiments, various syntax elements can be used to signal one or more of the above motion vector differences in reference frame list 0 or list 1.
[0207] For example, a syntax element called "mv_joint" can specify which components of the associated motion vector difference are non-zero. For the MVD, this is for all non-zero components to be jointly signaled. For example, mv_joint has values
[0208] 0 can indicate that there is no non-zero MVD along the horizontal or vertical direction;
[0209] 1 can indicate that there is only a non-zero MVD along the horizontal direction;
[0210] 2 can indicate that there is only a non-zero MVD along the vertical direction;
[0211] 3 can indicate that there are non-zero MVDs along both the horizontal and vertical directions.
[0212] When the "mv_joint" syntax element for the MVD signals the absence of non-zero MVD components, then the MVD information is no longer signaled. However, if the "mv_joint" syntax signals the presence of one or two non-zero components, then additional syntax elements can be signaled for each of the non-zero MVD components as described below.
[0213] For example, a syntax element called "mv_sign" can be used to additionally specify whether the corresponding motion vector difference component is positive or negative.
[0214] For another example, a syntax element called "mv_class" can be used to specify a class of the motion vector difference in a predefined set of classes for the corresponding non-zero MVD component. For example, a predefined class for the motion vector difference can be used to divide the continuous magnitude space of the motion vector difference into non-overlapping ranges, where each range corresponds to an MVD class. Thus, the signaled MVD class indicates the magnitude range of the corresponding MVD component. In the example implementation shown in Table 4 below, higher classes correspond to motion vector differences with larger magnitude ranges. In Table 4, the notation (n,m] is used to denote a range of motion vector differences greater than n pixels and less than or equal to m pixels.
[0215] Table 4: Magnitude Classes for Motion Vector Differences
[0216]
[0217]
[0218] In some other examples, a syntax element called "mv_bit" can be further used to specify the integer part of the offset between the non-zero motion vector difference component and the starting magnitude of the correspondingly signaled MV class magnitude range. In this way, mv_bit can indicate the magnitude or amplitude of the MVD. The number of bits required to signal the full range of each MVD class in "mv_bit" can vary as a function of the MV class. For example, MV_CLASS 0 and MV_CLASS1 in the implementation of Table 4 may only require a single bit to indicate an integer pixel offset of 1 or 2 from the starting MVD of 0; each higher MV_CLASS in the example implementation of Table 4 may require one more bit than the previous MV_CLASS for "mv_bit".
[0219] In some other examples, a syntax element called "mv_fr" can be further used to specify the first 2 fractional bits of the motion vector difference for the corresponding non-zero MVD component, and a syntax element called "mv_hp" can be used to specify the third fractional bit (high-resolution bit) of the motion vector difference for the corresponding non-zero MVD component. The 2-bit "mv_fr" substantially provides 1 / 4 pixel MVD resolution, and the "mv_hp" bit can further provide 1 / 8 pixel resolution. In some other implementations, more than one "mv_hp" bit can be used to provide a finer MVD pixel resolution than 1 / 8 pixel. In some example implementations, an additional flag can be signaled at one or more levels in each level to indicate whether 1 / 8 pixel or higher MVD resolution is supported. If the MVD resolution is not applied to a particular coding unit, the above syntax elements for the corresponding non-supported MVD resolution may not be signaled.
[0220] In some of the above example embodiments, the fractional resolution can be independent of different classes of MVDs. In other words, regardless of the magnitude of the motion vector difference, a predefined number of "mv_fr" and "mv_hp" bits can be used to provide similar options for motion vector resolution for signaling fractional MVDs of non-zero MVD components.
[0221] However, in some other example embodiments, the resolution of the motion vector difference can be distinguished among various MVD magnitude classes. Specifically, high-resolution MVDs for large MVD magnitudes in higher MVD classes may not provide a statistically significant improvement in compression efficiency. Thus, for larger MVD magnitude ranges corresponding to higher MVD magnitude classes, the MVD can be encoded at a reduced resolution (either integer pixel resolution or fractional pixel resolution). Similarly, generally for larger MVD values, the MVD can be encoded at a reduced resolution (either integer pixel resolution or fractional pixel resolution). This MVD resolution that depends on the MVD class or on the MVD magnitude is generally referred to as adaptive MVD resolution, amplitude-dependent adaptive MVD resolution, or magnitude-dependent MVD resolution. The term "resolution" can be further referred to as "pixel resolution". Adaptive MVD resolution can be implemented in various ways as described in the following example embodiments to achieve overall better compression efficiency. In particular, since processing the MVD resolution of large magnitude or high class MVDs in a non-adaptive manner at a level similar to that of low magnitude or low class MVDs may not significantly increase the statistically observed inter-frame prediction residual coding efficiency of blocks with large magnitude or high class MVDs, reducing the number of signaling bits by targeting a less precise MVD may be greater than the additional bits required to encode the inter-frame prediction residual due to such a less precise MVD. In other words, using a higher MVD resolution for large magnitude or high class MVDs may not result in more coding gain than using a lower MVD resolution.
[0222] In some general example embodiments, the pixel resolution or precision of the MVD can decrease or not increase as the MVD class increases. Decreasing the pixel resolution of the MVD corresponds to a coarser MVD (or a larger step from one MVD level to the next). In some embodiments, the correspondence between the MVD pixel resolution and the MVD class can be specified, predefined, or preconfigured and thus may not need to be signaled in the encoded bitstream.
[0223] In some example embodiments, the MV classes of Table 3 can each be associated with a different MVD pixel resolution.
[0224] In some example embodiments, each MVD class may be associated with a single allowed resolution. In some other embodiments, one or more MVD classes may be associated with two or more optional MVD pixel resolutions. Thus, the signal in the bitstream of the current MVD component having such an MVD class may follow additional signaling for indicating the optional pixel resolution selected for the current MVD component.
[0225] In some example embodiments, the adaptively allowed MVD pixel resolutions may include but are not limited to 1 / 64-pel (pixel), 1 / 32 pixel, 1 / 16 pixel, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels... (in descending order of resolution). Thus, each ascending MVD class may be associated with one of these MVD pixel resolutions in a non-ascending manner. In some embodiments, an MVD class may be associated with two or more of the above resolutions, and the higher resolution may be lower than or equal to the lower resolution of the previous MVD class. For example, if MV_CLASS_3 in Table 4 is associated with the optional 1 pixel and 2 pixel resolutions, the highest resolution that MV_CLASS_4 in Table 4 can be associated with will be 2 pixels. In some other embodiments, the highest allowed resolution of an MV class may be higher than the lowest allowed resolution of the previous (lower) MV class. However, the average of the allowed resolutions of ascending MV classes may only be non-ascending.
[0226] In some embodiments, when allowing fractional pixel resolutions higher than 1 / 8 pixel, the "mv_fr" and "mv_hp" signaling may be correspondingly extended to a total of more than 3 fractional bits.
[0227] In some example embodiments, fractional pixel resolution may only allow MVD classes that are less than or equal to a threshold MVD class. For example, fractional pixel resolution may only allow MVD-CLASS 0 and not all of the other MV classes in Table 4. Similarly, fractional pixel resolution may only allow MVD classes that are less than or equal to any one of the other MV classes in Table 4. For other MVD classes that are higher than the threshold MVD class, only integer pixel resolution of the MVD is allowed. In this way, fractional resolution signaling such as one or more bits in "mv-fr" and / or "mv-hp" may not be required for MVD signaling of MVD classes signaled at a MVD class that is higher than or equal to the threshold MVD class. For MVD classes with a resolution less than 1 pixel, the number of bits in the "mv-bit" signaling may be further reduced. For example, for MV_CLASS_5 in Table 4, the range of MVD pixel offsets is (32,64], so 5 bits are required to signal the entire range at 1 pixel resolution. However, if MV_CLASS_5 is associated with a 2 pixel MVD resolution (a lower resolution than 1 pixel resolution), "mv_bit" may require 4 bits instead of 5 bits, and none of "mv-fr" and "mv-hp" need to be signaled after the signaling of "mv_class" as MV-CLASS_5.
[0228] In some example embodiments, fractional pixel resolution may only allow MVDs with integer values less than a threshold integer pixel value. For example, fractional pixel resolution may only allow MVDs less than 5 pixels. Corresponding to this example, fractional resolution may allow MV_CLASS_0 and MV_CLASS_1 of Table 4 and not allow all other MV classes. For another example, fractional pixel resolution may only allow MVDs less than 7 pixels.. Corresponding to this example, fractional resolution may allow MV_CLASS_0 and MV_CLASS_1 (range less than 5 pixels) of Table 4 and not allow MV_CLASS_3 and higher (range greater than 5 pixels). For MVDs belonging to MV_CLASS_2, whose pixel range includes 5 pixels, depending on the "mv-bit" value, fractional pixel resolution of the MVD may or may not be allowed. If the "m-bit" value is signaled as 1 or 2 (such that the integer part of the signaled MVD is 5 or 6, calculated as the start of the pixel range of MV_CLASS_2 with an offset of 1 or 2 indicated by "m-bit"), then fractional pixel resolution may be allowed. Otherwise, if the "mv-bit" value is signaled as 3 or 4 (such that the integer part of the signaled MVD is 7 or 8), then fractional pixel resolution may not be allowed.
[0229] In some other embodiments, for MV classes equal to or higher than a threshold MV class, only a single MVD value may be allowed. For example, such a threshold MV class may be MV_CLASS2. Thus, for MV_CLASS_2 and above, only a single MVD value with no fractional pixel resolution may be allowed. The single allowed MVD values for these MV classes may be predefined. In some examples, the single allowed value may be the higher end value of the corresponding ranges for these MV classes in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be threshold classes higher than or equal to MV_CLASS2, and the single allowed MVD values for these classes may be predefined as 8, 16, 32, 64, 128, 256, 512, 1024, and 2048, respectively. In some other examples, the single allowed value may be the midpoint value of the corresponding ranges for these MV classes in Table 4. For example, MV_CLASS_2 to MV_CLASS_10 may be higher than the class threshold, and the single allowed MVD values for these classes may be predefined as 3, 6, 12, 24, 48, 96, 192, 384, 768, and 1536, respectively. Any other value within the range may also be defined as the single allowed resolution for the corresponding MVD class.
[0230] In the above embodiments, when the signaling "mv_class" is equal to or higher than a predefined MVD class threshold, only the "mv_class" signaling is sufficient to determine the MVD value. Then the "mv_class" and "mv_sign" will be used to determine the magnitude and direction of the MVD.
[0231] Thus, when signaling the MVD for only one reference frame (from reference frame list 0 or list 1, but not both) or signaling the MVD jointly for two reference frames, the accuracy (or resolution) of the MVD may depend on the associated class of the motion vector difference in Table 3 and / or the magnitude of the MVD.
[0232] In some other embodiments, the pixel resolution or precision of the MVD may decrease or not increase as the MVD magnitude increases. For example, the pixel resolution may depend on the integer part of the MVD magnitude. In some embodiments, the fractional pixel resolution may only allow MVD magnitudes less than or equal to the magnitude threshold. For the decoder, the integer part of the MVD magnitude may first be extracted from the bitstream. Subsequently, the pixel resolution may be determined, and then a decision may be made as to whether any fractional MVD is present in the bitstream and needs to be parsed (e.g., if the fractional pixel resolution does not allow a particular extracted MVD integer magnitude, then the fractional MVD bits may not be included in the bitstream to be extracted). The above example embodiments related to adaptive MVD pixel resolution depending on the MVD class apply to adaptive MVD pixel resolution depending on the MVD magnitude. For a specific example, MVD classes above or including the magnitude threshold may only have one predefined value. The above various example embodiments apply to the single-reference mode. These embodiments also apply to the example NEW_NEARMV, NEAR_NEWMV, and / or NEW_NEWMV modes in composite prediction under MMVD. These embodiments generally apply to adaptive resolution for any MVD.
[0233] In a specific example embodiment of adaptive MVD pixel resolution, the MVD pixel resolution for MVD magnitudes below 1 may be fractional, and for MV_CLASS_1 and above MV classes, only a single MVD magnitude equal to the end value of the corresponding MVD magnitude range in Table 4 may be allowed. In such an example, the allowed MVD values indicating the allowed fractional pixel resolutions for 1 / 8, 1 / 4, or 1 / 2 pixels are indicated in Table 4.
[0234] Table 5
[0235]
[0236]
[0237] For coding blocks, it can be signaled (exported) explicitly or implicitly whether adaptive MVD pixel resolution is used. When it is signaled that adaptive MVD pixel resolution is not used, it means that different MVD categories can follow the MVD ranges indicated in Table 4, and a non - adaptive MVD pixel resolution can be defined or signaled. This non - adaptive resolution can be fractional (such as 1 / 8, 1 / 4, or 1 / 2 pixel) or non - fractional (such as 1, 2, 4... pixels), and it will apply to all MVD categories. The non - adaptive resolution basically determines the number of bits required to signal the above - mentioned mv_bit, mv_fr, and mv_hp. When the non - adaptive resolution is fractional, it can only determine the number of bits required to signal mv_fr and mv_hp for all MVD categories (regardless of the MVD category), and the number of bits to signal Mv_bit can depend on the MVD category.
[0238] When it is signaled that adaptive MVD pixel resolution is used, MVD levels or values allowed by an adaptive mode such as those shown in Table 5 can be predefined or signaled. For example, according to a specific scheme of adaptive MVD resolution, they can be signaled in the bitstream in various ways. In the example of Table 5, a signaling syntax set can be used to indicate a fractional resolution (e.g., 1 / 8 pixel), a magnitude threshold below which the signaled fractional resolution applies (e.g., an MVD magnitude of 1 pixel). Another syntax set (possibly more complex) can be used to signal other adaptive MVD resolution schemes. This indication of the adaptive MVD pixel resolution scheme can be signaled at one of various coding levels such as sequence level, picture level, frame level, slice level, super - block level, or coding - block level.
[0239] In some example embodiments, an overall adaptive MVD pixel resolution scheme, including but not limited to those shown in Table 5, can be defined or signaled at a specific coding level (e.g., sequence level, picture level, frame level, slice level, super - block level). This adaptive MVD pixel resolution scheme can be further modified at the same coding level or another coding level, such that the allowed MVD pixel resolution values for various MVD categories can be adjusted or modified at the same coding level or another coding level. If no adjustment is made at a specific coding level, the signaled or predefined adaptive MVD pixel resolution scheme is applied without modification. For example, an overall adaptive MVD pixel resolution scheme can be defined or signaled at the frame level, while adjustments can be made at one or more super - block levels or coding - block levels, and vice versa.
[0240] Such adjustment can be implemented as a limit on MVD precision or an extension of MVD precision. Information associated with such adjustment can be predefined or signaled. The predefined adjustment can be applicable to all coding blocks. Optionally, the predefined adjustment can be activated at various coding levels by signaling.
[0241] In some embodiments, such adjustment can be embodied as the maximum allowable MVD precision. For a particular coding block, when applying adaptive MVD resolution, such maximum allowable MVD precision can be different from the MVD pixel precision of the adaptive MVD pixel resolution scheme specified / signaled / derived at the picture level, or super-block level, or coded block level, as described above. In such a case, the allowable MVD resolution values for various MVD categories can be determined by adopting both the allowable values specified or derived by the adaptive MVD pixel resolution scheme and the maximum allowable MVD precision.
[0242] For example, assume that for a certain coding level, the adaptive MVD pixel resolution scheme of Table 5 is predefined / signaled / derived. Further assume that the maximum allowable precision is 1 / 4 pixel, which means that regardless of the adaptive MVD resolution associated with Table 5, no MVD category is allowed to use a precision equal to or higher than 1 / 8 pixel. Then, by applying the maximum allowable pixel precision as a limit indifferently to Table 5, the allowable MVD pixel levels or values for various MVD categories can be modified to:
[0243] Table 6
[0244]
[0245]
[0246] By defining / signaling the maximum allowable MVD pixel precision as 1 / 4 pixel and not allowing all MVD categories to use 1 / 8 or higher pixel precision is just an example. In another example, the maximum allowable pixel precision can be defined / signaled as 1 / 2 pixel. The corresponding allowable MVD values for MV_CLASS_0 above can become: (1 / 2, 1, 2) for the adaptive resolution scheme with fractional pixel resolutions of 1 / 8 pixel, 1 / 4 pixel, and 1 / 2 pixel, and (1, 2) for the adaptive resolution scheme with a pixel resolution of 1 pixel.
[0247] As illustrated by the examples in the above-described embodiments regarding Table 6, when the application of adaptive MVD resolution uses a defined / signaled / derived adaptive resolution scheme at a specific coding level and an additional defined / signaled maximum allowable MVD accuracy at the same or a different coding level, it may be necessary / to limit such maximum allowable MVD accuracy to be no greater than the MVD resolution in the adaptive resolution scheme. In other words, the MVD accuracy of the actual application derived by considering both the adaptive resolution scheme and the maximum allowable accuracy is truncated by the MVD resolution in the adaptive resolution scheme (i.e., when the maximum allowable accuracy is greater than the defined / signaled / derived resolution by the adaptive resolution scheme, the maximum allowable accuracy will be invalid).
[0248] However, in some other embodiments, such truncation may not be required, and the defined / signaled maximum allowable MVD accuracy can control the actual MVD resolution for at least some MVD classes. In these embodiments, when applying adaptive MVD resolution (as indicated by definition / signaling / derivation at various coding levels, as described above), for the adjustment of the MVD level for at least some MVD classes, it may involve increasing rather than limiting the defined / signaled / derived adaptive MVD resolution in the adaptive MVD pixel resolution scheme associated with, for example, Table 5. For example, an adjustment can be made to allow the use of an accuracy higher than the specified / signaled / derived from the adaptive resolution scheme of the MVD class, where the MV class is equal to or lower than the defined or signaled threshold MVD class level. As described above, such higher accuracy can be defined / signaled as the maximum allowable MVD accuracy. Regardless of the specified / signaled / derived MVD resolution in the adaptive resolution scheme, this maximum allowable accuracy can be applied to MVD classes equal to or lower than the threshold MVD class level. Specifically, this threshold MVD class level can be (but does not have to be) MV_CLASS_0 (or the lowest MVD class level of the MVD class set, such as the MVD class set in Table 5). The maximum allowable pixel accuracy can be predefined / signaled. The maximum allowable pixel accuracy can be a fraction. For a specific example, in the case where the threshold MVD class in the adaptive resolution scheme of Table 5 is MV_CLASS_0, when the MVD pixel resolution for MV_CLASS_0 is a non-fractional 1 pixel and the maximum allowable fractional pixel accuracy for adjustment is 1 / 8, 1 / 4, or 1 / 2 pixel, the adjusted allowable MVD values will be:
[0249] Table 7
[0250]
[0251] In some example embodiments of alternative Table 7, a threshold MVD magnitude may be used instead of the threshold MVD class level. In these embodiments, for an MVD whose magnitude is equal to or below the threshold MVD magnitude rather than the threshold MVD class level, a higher precision may be imposed on the specified / signaled maximum allowable MVD precision. In such an embodiment, in addition to the mv_class information, the mv_bit information may be signaled early enough in the video stream such that the magnitude of the MVD can be determined in time to determine the allowable MVD value. For example, by replacing the threshold MVD class with a threshold MVD magnitude of 1 / 2 pixel and still assuming that the adaptive MVD resolution for MV_CLASS_0 is 1 pixel in the adaptive resolution scheme, Table 7 would become Table 8 below:
[0252] Table 8
[0253]
[0254]
[0255] In some other exemplary embodiments, the above adjustment may involve allowing only a specific precision and lower precisions (e.g., fractional precisions 1 / 8, 1 / 4, or 1 / 2 and lower) when the magnitude of the MVD is equal to or below the threshold MVD magnitude. In such an embodiment, again, in addition to the mv_class information, the mv_bit information may be signaled early enough in the video stream such that the magnitude of the MVD can be determined in time to determine the allowable MVD value.
[0256] In such an embodiment, no additional resolution is imposed on the MVD values derived from the adaptive resolution scheme (such as Table 5). Instead, when the magnitude of the MVD is higher than the threshold MVD magnitude, MVD values associated with a resolution equal to or higher than the defined / signaled precision level may not be allowed. Again assuming, in the example of Table 5, and further assuming that for MVD magnitudes higher than the threshold MVD magnitude of 1 / 2 pixel, MVD values associated with a resolution equal to or higher than the defined / signaled precision of 1 / 8 pixel precision are not allowed. Then Table 5 would be adjusted to:
[0257] Table 9
[0258]
[0259] Specifically, as shown above, the allowed MVD values (1 / 8, 2 / 8, 3 / 8, 1 / 2, 5 / 8, 6 / 8, 7 / 8, 1, 2) for the fractional resolution of MV_CLASS_0 and 1 / 8 pixel are adjusted to (1 / 8, 2 / 8, 3 / 8, 1 / 2, 6 / 8, 1, 2), where the MVD values associated with 1 / 8 precision are only allowed to be used and maintained when the threshold MVD magnitude is equal to or less than 1 / 2 pixel. Beyond the 1 / 2 pixel magnitude, MVD values associated with 1 / 8 precision, such as 5 / 8 pixel value and 7 / 8 pixel value, are not allowed to be used.
[0260] Similarly, in the example of Table 5, it is assumed that when the MVD magnitude is higher than the threshold magnitude of 1 / 2 pixel, MVD values associated with the resolution of the defined / signaled precision equal to or higher than 1 / 4 pixel precision are not allowed to be used. Then Table 5 will be adjusted to:
[0261] Table 9
[0262]
[0263] In some of the above embodiments, the threshold MVD magnitude can be 2 pixels or less, such as the 1 / 2 pixel magnitude threshold given in the above example.
[0264] The above example embodiments are descriptions regarding specific MVDs, regardless of whether the inter-frame prediction mode is a single reference mode or a composite reference mode. In some other example embodiments of the composite reference mode (where the MV is predicted by multiple reference frames), a defined / signaling set can be used to indicate whether adaptive MVD resolution is applied and to which reference frame or reference frames among the multiple reference frames it is applied.
[0265] In some example embodiments, when signaling the MVD for multiple reference frames, one (or more) flags / indexes can be signaled to indicate whether adaptive MVD resolution is applied.
[0266] For example, when signaling the MVD for multiple reference frames (e.g., in the above NEW_NEWMV mode or other composite reference inter-frame prediction modes), one flag / index can be signaled in the video stream to indicate whether adaptive MVD resolution is applied to the signaling of the MVD for all multiple reference frames. If the flag / index is 1 (or 0), it indicates that adaptive MVD resolution is applied to the signaling of the MVD for all multiple reference frames. Otherwise, if the flag / index is 0 (or 1), adaptive MVD coding is not applied to the signaling of the MVD for any of the multiple reference frames. In this embodiment, for multiple inter-frame prediction reference frames, adaptive MVD resolution is applied in an all-or-nothing scheme.
[0267] In some other examples, when signaling the MVD for multiple reference frames (e.g., in the NEW_NEWMV mode or other composite inter prediction modes for the dual reference frame composite inter prediction mode described above), a flag / index can be signaled separately for each reference frame to indicate whether adaptive MVD resolution is applied to each reference frame. In such an implementation, it can be determined separately for each of the reference frames whether to apply adaptive MVD resolution. A decision on whether to apply adaptive MVD resolution can be made independently for each of the multiple reference frames at the encoder and signaled separately in the video stream.
[0268] In some example implementations, when signaling the MVD for multiple reference frames, for each of the multiple reference frames, if the MVD of the reference frame is non-zero, a flag / index can be signaled to indicate whether adaptive MVD resolution is applied to the reference frame. Otherwise, there is no need to signal the flag / index. In other words, if it is signaled / indicated that the MVD for a particular reference frame is zero, there is no need to determine whether to apply adaptive MVD resolution, and thus no corresponding signaling is required in the video stream. However, in such an implementation, an indication that the MVD is zero needs to be signaled before making a decision on whether to apply adaptive resolution.
[0269] Turning further to the signaling of MVD resolution, in some example implementations, a flag / index can be signaled to explicitly indicate the MVD resolution for the current coding block, and the context used for entropy coding such a flag / index can depend on the MVD category associated with the MVD. Such a flag / index can be any MVD resolution used to derive an adaptive resolution scheme such as Table 5, or the maximum allowed MVD precision described above.
[0270] In some example implementations regarding the signaling of MVD resolution, the individual components of the MVD can be signaled separately. The MVD can include, for example, a horizontal component and a vertical component. A flag / index can be signaled for each of the horizontal and vertical components of the MVD to indicate the MVD resolution of the horizontal and vertical components, respectively.
[0271] In some example embodiments, an MVD resolution flag / index may be signaled after the MVD class information. Based on the value of the signaled MVD class information such as MV_CLASS_0, MV_CLASS_1, MV_CLASS_2, etc., a context value may be derived and used to signal the MVD resolution flag / index for indicating the MVD resolution. In other words, different contexts for different MVD classes or different groups of MVD classes may be used to entropy code one or more syntaxes of the signaling for the MVD resolution.
[0272] Figure 18 FIG. 1800 is a flow chart of an example method illustrating the principles of an embodiment following the above adaptive MVD resolution. The example decoding method flow starts at S1801. At S1810, a video stream is received. At S1820, based on a prediction block and a motion vector (MV), it is determined that a video block is inter-coded, where the MV is to be derived from a reference motion vector (RMV) and a motion vector difference (MVD) for the video block. At S1830, in response to determining that the MVD is encoded with an adaptive MVD pixel resolution: determine a reference MVD pixel precision for the current video block; identify a maximum allowed MVD pixel precision; based on the reference MVD pixel precision and the maximum allowed MVD pixel precision, determine an allowed MVD level set for the current video block; and derive the MVD from the video stream according to at least one MVD parameter signaled in the video stream for the current video block and the allowed MVD level set. The example method stops at S1899.
[0273] Figure 19 FIG. 1900 is a flow chart of another example method illustrating the principles of an embodiment following the above adaptive MVD resolution. The example decoding method flow starts at S1901. At S1910, a video stream is received. At S1920, it is determined that the current video block is inter-coded and associated with multiple reference frames. At S1930, further based on the signaling in the video stream, it is determined whether an adaptive motion vector difference (MVD) pixel resolution is applied to at least one of the multiple reference frames. The example method stops at S1999.
[0274] Figure 20FIG. 2000 is a flowchart of an example method showing the principle of an embodiment following the above adaptive MVD resolution. The example decoding method flow starts at S2001. At S2010, a video stream is received. At S2020, based on a prediction block and a motion vector (MV), it is determined that a video block is inter-coded, where the MV is to be derived from a reference motion vector (RMV) and a motion vector difference (MVD) for the video block. At S2030, a current MVD category of the MVD is determined from a predefined set of MVD categories. At S2040, based on the current MVD category, at least one context for entropy decoding at least one explicit signaling in the video stream is derived, where the at least one explicit signaling is included in the video stream to specify an MVD pixel resolution for at least one component of the MVD. At S2050, the at least one explicit signaling in the video stream is entropy decoded using the at least one context to determine the MVD pixel resolution for at least one component of the MVD. The example method terminates at S2099.
[0275] In embodiments and implementations of the present disclosure, any steps and / or operations can be combined or arranged in any number or order as needed. Two or more steps and / or operations can be performed in parallel. The embodiments and implementations in the present disclosure can be used alone or in any combination in any order. Further, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium. The embodiments in the present disclosure can be applied to luminance blocks or chrominance blocks. The term "block" can be interpreted as a prediction block, a coding block, or a coding unit (i.e., CU). The term "block" can also be used herein to refer to a transform block. In the following items, when referring to the block size, it can refer to the block width or height, or the maximum of the width and height, or the minimum of the width and height, or the area size (width * height), or the aspect ratio of the block (width: height, or height: width).
[0276] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 21 FIG. 2100 shows a computer system suitable for implementing certain embodiments of the disclosed subject matter.
[0277] The computer software can be encoded using any suitable machine code or computer language, which can be subject to assembly, compilation, linking, or similar mechanisms to create code including instructions that can be executed directly or by interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.
[0278] The instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablets, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0279] Figure 21 The components shown for the computer system (2100) are exemplary in nature and are not intended to imply any limitation regarding the scope of use or functionality of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependence on or requirement for any one component or combination of components shown in the exemplary embodiments of the computer system (2100).
[0280] The computer system (2100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input by one or more human users through, for example, tactile input (e.g., key presses, swipes, data glove movements), audio input (e.g., voice, taps), visual input (e.g., gestures), olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media not directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0281] The input human-machine interface devices may include one or more of the following (each depicted only one): keyboard (2101), mouse (2102), trackpad (2103), touch screen (2110), data glove (not shown), joystick (2105), microphone (2106), scanner (2107), camera (2108).
[0282] The computer system (2100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., the tactile feedback of a touch screen (2110), a data glove (not shown), or a joystick (2105), but there may also be tactile feedback devices that do not act as input devices), audio output devices (e.g., speakers (2109), headphones (not depicted)), visual output devices (e.g., a screen (2110), including a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, an organic light emitting diode (OLED) screen, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or output greater than three dimensions through, for example, stereoscopic flat painting output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted).
[0283] The computer system (2100) may also include human-accessible storage devices and the associated media of the storage devices, such as optical media, including CD / DVD ROM / RW (2120) with media such as CD / DVD (2121), thumb drives (2122), removable hard disk drives, or solid state drives (2123), legacy magnetic media such as tapes and floppy disks (not depicted), dedicated devices based on ROM / application specific integrated circuit (ASIC) / programmable logic device (PLD), such as security protection devices (not depicted), and so on.
[0284] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0285] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of networks include, for example, Ethernet, local area networks of wireless LAN, cellular networks including Global System for Mobile Communications (GSM), Third Generation (3G), Fourth Generation (4G), Fifth Generation (5G), Long Term Evolution (LTE), etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular networks and industrial networks including Controller Area Network Bus (CAN Bus), etc. Some networks typically require an external network interface adapter attached to certain general-purpose data ports or peripheral buses (2149) (e.g., Universal Serial Bus (USB) ports of the computer system (2100)); other networks are typically integrated into the core of the computer system (2100) by attaching to the system bus as described below (e.g., integrated into a PC computer system through an Ethernet interface, or integrated into a smartphone computer system through a cellular network interface). By using any of these networks, the computer system (2100) can communicate with other entities. Such communication can be only one-way reception (e.g., broadcast TV), only one-way transmission (e.g., CANBus connected to certain CANBus devices), or two-way, for example, using a local digital network or a wide area digital network to connect to other computer systems. Certain protocols and protocol stacks can be used on each of the networks and network interfaces as described above.
[0286] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be attached to the core (2140) of the computer system (2100).
[0287] The core (2140) may include one or more central processing units (CPUs) (2141), a graphics processing unit (GPU) (2142), a dedicated programmable processing unit in the form of field programmable gate areas (FPGAs) (2143), a hardware accelerator for certain tasks (2144), a graphics adapter (2150), and so on. These devices, together with a read-only memory (ROM) (2145), a random access memory (2146), and internal mass storage devices such as internal hard disk drives, solid state drives (SSDs), etc., that are not user-accessible (2147), may be connected via a system bus (2148). In some computer systems, the system bus (2148) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (2149) to the system bus (2148) of the core. In one example, a screen (2110) may be connected to the graphics adapter (2150). Architectures for the peripheral bus include Peripheral Component Interconnect (PCI), USB, and so on.
[0288] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) may execute certain instructions, which, when combined, may constitute the aforementioned computer code. The computer code may be stored in the ROM (2145) or the RAM (2146). Transitional data may also be stored in the RAM (2146), while permanent data may be stored, for example, in the internal mass storage device (2147). Fast storage and retrieval of any memory device may be achieved by using a cache memory, which may be closely associated with one or more CPUs (2141), GPUs (2142), mass storage devices (2147), ROM (2145), RAM (2146), etc.
[0289] Computer code for performing various computer-implemented operations may be present on a computer-readable medium. The medium and the computer code may be those designed and constructed specifically for the purposes of this application, or may be of the kind well-known and available to those skilled in the field of computer software.
[0290] As a non - limiting example, a computer system having an architecture (2100) and in particular a core (2140) can provide functions resulting from a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer - readable media. Such computer - readable media can be media associated with the user - accessible mass storage device introduced above and certain non - transient storage devices of the core (2140) (e.g., the on - core mass storage device (2147) or ROM (2145)). The software implementing various embodiments of the present application can be stored in such devices and executed by the core (2140). Depending on specific requirements, the computer - readable media can include one or more memory devices or chips. The software can cause the core (2140) and specifically the processors therein (including CPU, GPU, FPGA, etc.) to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in the RAM (2146) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functions resulting from logic hard - wired or otherwise embodied in a circuit (e.g., accelerator (2144)), which can operate instead of or in conjunction with the software to execute specific processes or specific parts of specific processes described herein. When appropriate, references to software can encompass logic and vice versa. When appropriate, references to computer - readable media can encompass a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both such circuits. The present application encompasses any suitable combination of hardware and software.
[0291] Although the present application describes several exemplary embodiments, within the scope of the present application, there can be various modifications, permutations, and various alternative equivalents. Therefore, it should be understood that within the spirit and scope of the application, those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, can embody the principles of the present application.
[0292] Appendix A: Abbreviations
[0293] JEM: Joint Exploration Model, Joint Development Model
[0294] VVC: Versatile Video Coding, Multifunctional Video Coding
[0295] BMS: Benchmark Set, Benchmark Collection
[0296] MV: Motion Vector, Motion Vector
[0297] HEVC: High Efficiency Video Coding, High-Efficiency Video Coding
[0298] SEI: Supplementary Enhancement Information, Supplementary Enhancement Information
[0299] VUI: Video Usability Information, Video Usability Information
[0300] GOP: Groups of Pictures, Picture Groups
[0301] TU: Transform Unit, Transform Unit
[0302] PU: Prediction Unit, Prediction Unit
[0303] CTU: Coding Tree Unit, Coding Tree Unit
[0304] CTB: Coding Tree Block, Coding Tree Block
[0305] PB: Prediction Block, Prediction Block
[0306] HRD: Hypothetical Reference Decoder, Hypothetical Reference Decoder
[0307] SNR: Signal Noise Ratio, Signal-to-Noise Ratio
[0308] CPU: Central Processing Unit, Central Processing Unit
[0309] GPU: Graphics Processing Unit, Graphics Processing Unit
[0310] CRT: Cathode Ray Tube, Cathode Ray Tube
[0311] LCD: Liquid-Crystal Display, Liquid Crystal Display
[0312] OLED: Organic Light-Emitting Diode, Organic Light-Emitting Diode
[0313] CD: Compact Disc, Compact Disc
[0314] DVD: Digital Video Disc, Digital Video Disc
[0315] ROM: Read-Only Memory, read-only memory
[0316] RAM: Random Access Memory, random access memory
[0317] ASIC: Application-Specific Integrated Circuit, application-specific integrated circuit
[0318] PLD: Programmable Logic Device, programmable logic device
[0319] LAN: Local Area Network, local area network
[0320] GSM: Global System for Mobile communications, global system for mobile communications
[0321] LTE: Long-Term Evolution, long-term evolution
[0322] CANBus: Controller Area Network Bus, controller area network bus
[0323] USB: Universal Serial Bus, universal serial bus
[0324] PCI: Peripheral Component Interconnect, peripheral component interconnect
[0325] FPGA: Field Programmable Gate Array, field programmable gate array
[0326] SSD: Solid-state drive, solid-state drive
[0327] IC: Integrated Circuit, integrated circuit
[0328] HDR: high dynamic range, high dynamic range
[0329] SDR: standard dynamic range, standard dynamic range
[0330] JVET: Joint Video Exploration Team, joint video exploration team
[0331] MPM: Most Probable Mode, the most likely mode
[0332] WAIP: Wide - Angle Intra Prediction
[0333] CU: Coding Unit
[0334] PU: Prediction Unit
[0335] TU: Transform Unit
[0336] CTU: Coding Tree Unit
[0337] PDPC: Position Dependent Prediction Combination, ISP: Intra Sub - Partitions
[0338] SPS: Sequence Parameter Setting
[0339] PPS: Picture Parameter Set
[0340] APS: Adaptation Parameter Set
[0341] VPS: Video Parameter Set
[0342] DPS: Decoding Parameter Set
[0343] ALF: Adaptive Loop Filter
[0344] SAO: Sample Adaptive Offset
[0345] CC-ALF: Cross-Component Adaptive Loop Filter, a cross-component adaptive loop filter; CDEF: Constrained Directional Enhancement Filter, a constrained directional enhancement filter; CCSO: Cross-Component Sample Offset, a cross-component sample offset
[0346] LSO: Local Sample Offset, a local sample offset
[0347] LR: Loop Restoration Filter, a loop restoration filter
[0348] AV1: AOMedia Video 1, an open media alliance video 1
[0349] AV2: AOMedia Video 2, an open media alliance video 2
[0350] MVD: Motion Vector difference, a motion vector difference
[0351] CfL: Chroma from Luma, chroma predicted from luma
[0352] SDT: Semi Decoupled Tree, a semi-decoupled tree
[0353] SDP: Semi Decoupled Partitioning, a semi-decoupled partitioning
[0354] SST: Semi Separate Tree, a semi-separate tree
[0355] SB: Super Block, a super block
[0356] IBC (or IntraBC): Intra Block Copy, an intra block copy
[0357] CDF: Cumulative Density Function, a cumulative density function
[0358] SCC: Screen Content Coding, screen content coding
[0359] GBI: Generalized Bi-prediction, a generalized bi-prediction
[0360] BCW: Bi-prediction with CU-level Weights, CIIP: Combined intra-inter prediction
[0361] POC: Picture Order Count
[0362] RPS: Reference Picture Set
[0363] DPB: Decoded Picture Buffer
[0364] MMVD: Merge Mode with Motion Vector Difference
Claims
1. A video decoding method, characterized in that, comprising: receiving an encoded video stream, the encoded video stream including a bitstream obtained by encoding a current frame composed of a plurality of video blocks, the plurality of blocks including a current video block; determining that the current video block is inter-coded based on a predicted block and a motion vector (MV), wherein the MV is to be derived from a reference motion vector (RMV) and a motion vector difference (MVD) for the current video block; and in response to determining that the MVD is encoded with an adaptive MVD pixel resolution: determining a reference MVD pixel precision for the current frame, the reference MVD pixel precision depending on an MV category associated with the MVD or an MVD magnitude of the MVD; identifying a maximum allowable MVD pixel precision for the current video block, the maximum allowable MVD pixel precision being predefined and the maximum allowable MVD pixel precision being less than the reference MVD pixel precision; determining a set of allowable MVD pixel precisions for the current video block by limiting the reference MVD pixel precision according to the maximum allowable MVD pixel precision; and deriving the MVD from the video stream according to at least one MVD parameter for the current video block signaled in the video stream and the set of allowable MVD pixel precisions.
2. The method according to claim 1, characterized in that, the reference MVD pixel precision for the current frame is specified / signaled / derived at a frame level.
3. The method according to any one of claims 1 to 2, characterized in that, further comprising: determining a current MV category from a predefined set of MV categories, wherein determining the set of allowable MVD pixel precisions for the current video block by limiting the reference MVD pixel precision according to the maximum allowable MVD pixel precision includes: excluding MVD pixel precisions associated with MVD pixel precisions equal to or higher than the maximum allowable MVD pixel precision from a set of reference MVD pixel precisions determined based on the reference MVD pixel precision and the current MV category to determine the set of allowable MVD pixel precisions for the current video block.
4. The method according to claim 3, characterized in that, the maximum allowable MVD pixel precision is 1 / 4 pixel.
5. The method according to any one of claims 1 to 2, characterized in that, excluding MVD pixel precisions associated with 1 / 8 pixel or higher precision from the set of allowable MVD pixel precisions for the current video block.
6. The method according to any one of claims 1 to 2, characterized in that, further comprising: determining a current MV category from a predefined set of MV categories, wherein: when the current MV category is equal to or lower than a threshold MV category, MVD pixel precisions associated with fractional MVD precisions are included in the set of allowable MVD pixel precisions.
7. The method according to claim 6, characterized in that, The threshold MV class is the lowest MV class in the predefined set of MV classes: MV_CLASS_0.
8. The method according to any one of claims 1 to 2, wherein, further comprising: determining a magnitude of the MVD, wherein use of an MVD pixel precision associated with an MVD precision higher than a threshold MVD precision is allowed in the allowed MVD pixel precision set only if the magnitude of the MVD is equal to or lower than a threshold MVD magnitude.
9. The method according to claim 8, wherein, the threshold MVD magnitude is 2 pixels or less.
10. The method according to claim 9, wherein, the threshold MVD magnitude is 1 pixel.
11. The method according to claim 8, wherein, use of an MVD pixel precision associated with an MVD precision of 1 / 4 pixel or higher is allowed only if the magnitude of the MVD is equal to or lower than 1 / 2 pixel.
12. The method according to claim 1, wherein, further comprising: determining that the current video block is inter-coded and associated with a plurality of reference frames; and determining, based on signaling in the video stream, whether an adaptive motion vector difference (MVD) pixel resolution is applied to at least one of the plurality of reference frames.
13. The method according to claim 12, wherein, the signaling includes a single-bit flag to indicate whether the adaptive MVD pixel resolution is applied to all of the plurality of reference frames or not applied to any of the plurality of reference frames.
14. The method according to claim 12, wherein, the signaling includes separate flags, each corresponding to one of the plurality of reference frames, to indicate whether the adaptive MVD pixel resolution is applied.
15. The method according to claim 12, wherein, for each of the plurality of reference frames, the signaling includes: an implicit indication for indicating that the adaptive MVD pixel resolution is not applied when the MVD corresponding to each of the plurality of reference frames is zero; and a single-bit flag for indicating whether the adaptive MVD pixel resolution is applied when the MVD corresponding to each of the plurality of reference frames is non-zero.
16. The method according to claim 1, wherein, further comprising: determining a current MV class of the MVD from a predefined set of MV classes; deriving, based on the current MV class, at least one context for entropy decoding at least one explicit signaling in the video stream, the at least one explicit signaling being included in the video stream to specify an MVD pixel resolution for at least one component of the MVD; and using the at least one context to perform entropy decoding on the at least one explicit signaling in the video stream to determine the MVD pixel resolution for at least one component of the MVD.
17. The method according to claim 16, wherein, At least one component of the MVD includes a horizontal component and a vertical component of the MVD, the at least one context includes two separate contexts, each context being associated with one of the horizontal component and the vertical component of the MVD, and the horizontal component and the vertical component are associated with separate MVD pixel resolutions.
18. A video encoding method, applied to a local decoder included in an encoder, characterized in that it includes: receiving an encoded video stream, the encoded video stream including a bitstream obtained by encoding a current frame composed of a plurality of video blocks, the plurality of blocks including a current video block; determining that the current video block of the video stream is inter-coded based on a prediction block and a motion vector (MV), wherein the MV is to be derived from a reference motion vector (RMV) and a motion vector difference (MVD) for the current video block; and when it is determined that the MVD is encoded with an adaptive MVD pixel resolution: determining a reference MVD pixel precision for the current frame, the reference MVD pixel precision depending on an MV category associated with the MVD or an MVD magnitude of the MVD; identifying a maximum allowed MVD pixel precision for the current video block, the maximum allowed MVD pixel precision being predefined and the maximum allowed MVD pixel precision being less than the reference MVD pixel precision; determining a set of allowed MVD pixel precisions for the current video block by limiting the reference MVD pixel precision according to the maximum allowed MVD pixel precision; and deriving the MVD from the video stream according to at least one MVD parameter for the current video block signaled in the video stream and the set of allowed MVD pixel precisions.
19. A video processing device, characterized in that it includes a processor and a memory for storing computer instructions, and when the computer instructions are executed, the processor is configured to execute the method according to any one of claims 1 to 18.
20. A non-transitory computer-readable storage medium storing instructions, which when executed by a computer, cause the computer to execute the method according to any one of claims 1 to 18.
21. A method for storing a video bitstream, characterized in that storing a video bitstream on a non-transitory computer-readable medium, the video bitstream being decoded according to the video decoding method according to any one of claims 1-17, or being generated based on the video encoding method according to claim 18.
Citation Information
Patent Citations
Methods and apparatus for adaptive coding of motion information
US20120201293A1
Restrictions on motion vector difference
WO2020259681A1