Method and apparatus for processing video data

CN116601959BActive Publication Date: 2026-08-21QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180084523.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-23
Filing Date
2021-11-24
Publication Date
2026-08-21
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

因此,为了满足这些需求所需要的大量视频数据为处理和存储视频数据的通信网络和设备带来了负担

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116601959B_ABST
    Figure CN116601959B_ABST
Patent Text Reader

Abstract

Systems and techniques for overlapping block motion compensation (OBMC) are provided. A method can include determining that an OBMC mode is enabled for a current subblock of video data; determining, for a neighboring subblock that is contiguous with the current subblock, whether a first condition, a second condition, and a third condition are satisfied, the first condition including that all reference picture lists used to predict the current subblock are used to predict the neighboring subblock, the second condition including that a same reference picture is used to determine motion vectors associated with the current subblock and the neighboring subblock, and the third condition including that a difference between motion vectors of the current subblock and the neighboring subblock does not exceed a threshold; and based on determining that the OBMC mode is enabled and that the first condition, the second condition, and the third condition are satisfied, determining not to use motion information of the neighboring subblock for motion compensation of the current subblock.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to video encoding and decoding. For example, aspects of this disclosure relate to systems and techniques for performing overlapping block motion compensation. Background Technology

[0002] Digital video capabilities can be integrated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game consoles, video game control panels, cellular or satellite phones (so-called "smartphones"), video conferencing equipment, video streaming devices, and more. Such devices enable the processing and output of video data for consumption. Digital video data comprises vast amounts of data to meet the needs of both consumers and video providers. For example, consumers of video data expect the highest quality video with high fidelity, high resolution, and high frame rates. Therefore, the large amounts of video data required to meet these needs place a burden on the communication networks and equipment used to process and store this video data.

[0003] Digital video devices can implement video decoding technologies for compressing video data. Video decoding can be performed according to one or more video decoding standards or formats. Examples of video decoding standards or formats include Universal Video Decoding (VVC), High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), MPEG-2 Part 2 Decoding (MPEG stands for Moving Picture Experts Group), and proprietary video codecs / formats (such as AOMediavideo 1 (AV1) developed by the Open Media Alliance). Video decoding typically utilizes predictive methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that take advantage of redundancy present in video images or sequences. The goal of video decoding technology is to compress video data to a form using a lower bitrate while avoiding or minimizing degradation in video quality. As more and more video services become available, there is a need for decoding technologies with better decoding efficiency. Summary of the Invention

[0004] Systems, methods, and computer-readable media for performing Overlapped Block Motion Compensation (OBMC) are disclosed. According to at least one example, a method for performing OBMC is provided. The example method may include: determining that an Overlapped Block Motion Compensation (OBMC) mode is enabled for a current sub-block of a video data block; for at least one adjacent sub-block, determining whether a first, second, and third condition are satisfied, the first condition including a list of all reference images in one or more reference image lists used to predict the current sub-block for predicting the adjacent sub-blocks; the second condition including the same one or more reference images used to determine motion vectors associated with the current sub-block and the adjacent sub-blocks; and the third condition including a first difference between the horizontal motion vectors of the current sub-block and the adjacent sub-blocks and a second difference between the vertical motion vectors of the current sub-block and the adjacent sub-blocks not exceeding a motion vector difference threshold, wherein the motion vector difference threshold is greater than zero; and based on determining that an OBMC mode is used for the current sub-block and determining that the first, second, and third conditions are satisfied, determining that motion information from the adjacent sub-blocks is not used for motion compensation of the current sub-block.

[0005] According to at least one example, a non-transitory computer-readable medium for OBMC is provided. The example non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause one or more processors to: determine that Overlapping Block Motion Compensation (OBMC) mode is enabled for a current sub-block of a video data block; for at least one adjacent sub-block adjacent to the current sub-block, determine whether a first condition, a second condition, and a third condition are satisfied, the first condition including all reference image lists in one or more reference image lists used for predicting the current sub-block for predicting the adjacent sub-blocks; the second condition including the same one or more reference images used to determine motion vectors associated with the current sub-block and the adjacent sub-blocks; and the third condition including a first difference between the horizontal motion vectors of the current sub-block and the adjacent sub-blocks and a second difference between the vertical motion vectors of the current sub-block and the adjacent sub-blocks not exceeding a motion vector difference threshold, wherein the motion vector difference threshold is greater than zero; and based on determining that OBMC mode is used for the current sub-block and determining that the first, second, and third conditions are satisfied, determine that motion information from the adjacent sub-blocks is not used for motion compensation of the current sub-block.

[0006] According to at least one example, an apparatus for Overlap Block Motion Compensation (OBMC) is provided. The example apparatus may include: a memory; and one or more processors coupled to the memory. The one or more processors are configured to: determine that Overlap Block Motion Compensation (OBMC) mode is enabled for a current sub-block of a video data block; for at least one adjacent sub-block adjacent to the current sub-block, determine whether a first condition, a second condition, and a third condition are satisfied, the first condition including all reference image lists in one or more reference image lists used for predicting the current sub-block for predicting adjacent sub-blocks; the second condition including the same one or more reference images used to determine motion vectors associated with the current sub-block and adjacent sub-blocks; and the third condition including a first difference between the horizontal motion vectors of the current sub-block and adjacent sub-blocks and a second difference between the vertical motion vectors of the current sub-block and adjacent sub-blocks not exceeding a motion vector difference threshold, wherein the motion vector difference threshold is greater than zero; and based on determining that OBMC mode is used for the current sub-block and determining that the first, second, and third conditions are satisfied, determine that motion information from adjacent sub-blocks is not used for motion compensation of the current sub-block.

[0007] According to at least one example, another apparatus for OBMC is provided. The example apparatus may include: components for determining that Overlapping Block Motion Compensation (OBMC) mode is enabled for a current sub-block of a video data block; components for determining, for at least one adjacent sub-block, whether a first condition, a second condition, and a third condition are satisfied, wherein the first condition includes all reference image lists in one or more reference image lists used to predict the current sub-block for predicting the adjacent sub-block; the second condition includes the same one or more reference images used to determine motion vectors associated with the current sub-block and the adjacent sub-block; and the third condition includes a first difference between the horizontal motion vectors of the current sub-block and the adjacent sub-block, and a second difference between the vertical motion vectors of the current sub-block and the adjacent sub-block, not exceeding a motion vector difference threshold, wherein the motion vector threshold is greater than zero; and components for determining that motion information from the adjacent sub-blocks is not used for motion compensation of the current sub-block based on determining that OBMC mode is used for the current sub-block and determining that the first, second, and third conditions are satisfied.

[0008] In some aspects, the methods, non-transitory computer-readable media, and apparatus may include: determining to perform a sub-block boundary OBMC mode for the current sub-block based on determining to use a decoder-side motion vector refinement (DMVR) mode, a sub-block-based temporal motion vector prediction (SbTMVP) mode, or an affine motion compensation prediction mode for the current sub-block.

[0009] In some cases, performing a sub-block boundary OBMC pattern for the current sub-block may include: determining a first prediction associated with the current sub-block, a second prediction associated with a first OBMC block adjacent to the top border of the current sub-block, a third prediction associated with a second OBMC block adjacent to the left border of the current sub-block, a fourth prediction associated with a third OBMC block adjacent to the bottom border of the current sub-block, and a fifth prediction associated with a fourth OBMC block adjacent to the right border of the current sub-block; determining a sixth prediction based on the result of applying a first weight to the first prediction, applying a second weight to the second prediction, applying a third weight to the third prediction, applying a fourth weight to the fourth prediction, and applying a fifth weight to the fifth prediction; and generating a hybrid sub-block corresponding to the current sub-block based on the sixth prediction.

[0010] In some examples, each of the second, third, fourth, and fifth weights may include one or more weight values ​​associated with one or more samples from the corresponding sub-block from the current sub-block. In some cases, the sum of the weight values ​​of the corner samples of the current sub-block is greater than the sum of the weight values ​​of the other boundary samples of the current sub-block. In some examples, the sum of the weight values ​​of the other boundary samples of the current sub-block is greater than the sum of the weight values ​​of the non-boundary samples of the current sub-block.

[0011] In some aspects, the method, non-transitory computer-readable medium, and apparatus may include: determining that an additional block of video data uses a local illumination compensation (LIC) mode; and, based on determining that the LIC mode is used for the additional block, skipping signaling associated with information related to the OBMC mode used for the additional block.

[0012] In some cases, signaling that skips information associated with the OBMC mode used for additional blocks may include signaling a syntax flag with a null value that is associated with the OBMC mode.

[0013] In some aspects, the method, non-transitory computer-readable medium, and apparatus may include: receiving a signal including a syntax flag having a null value, the syntax flag being associated with an OBMC mode for an additional block of video data. In some aspects, the method, non-transitory computer-readable medium, and apparatus may include: determining, based on the syntax flag having a null value, not to use the OBMC mode for the additional block.

[0014] In some examples, signaling that skips information associated with the OBMC mode used for the extra block may include: determining whether the OBMC mode is not used or enabled for the extra block based on whether the LIC mode is used for the extra block; and skipping the signal notification of the value associated with the OBMC mode used for the extra block.

[0015] In some aspects, the method, non-transitory computer-readable medium, and apparatus may include: determining whether to enable OBMC mode for an additional block; and based on determining whether to enable OBMC mode for an additional block and determining that LIC mode is used for the additional block, determining to skip signaling information associated with the OBMC mode used for the additional block.

[0016] In some aspects, the method, non-transitory computer-readable medium, and apparatus may include: determining a decoding unit (CU) boundary OBMC mode for a current sub-block of a video data block; and determining a final prediction for the current sub-block based on the sum of a first result of applying weights associated with the current sub-block to a corresponding prediction associated with the current sub-block and a second result of applying one or more corresponding weights to one or more corresponding predictions associated with one or more sub-blocks adjacent to the current sub-block.

[0017] In some examples, determining not to use motion information from adjacent sub-blocks for motion compensation of the current sub-block may include: skipping the use of motion information from adjacent sub-blocks for motion compensation of the current sub-block.

[0018] In some examples, the OBMC mode may include the sub-block boundary OBMC mode.

[0019] In some aspects, one or more of the device are, may be part of, or may include: mobile devices, camera devices, encoders, decoders, Internet of Things (IoT) devices, and / or extended reality (XR) devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices). In some aspects, the device includes a camera device. In some examples, the device may include or be part of: vehicles, mobile devices (e.g., mobile phones or so-called "smartphones" or other mobile devices), wearable devices, personal computers, laptop computers, tablet computers, server computers, robotic devices or systems, aviation systems, or other devices. In some aspects, the device includes an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, the device includes one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device includes one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, the device may include one or more sensors.

[0020] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to define the scope of the claimed subject matter. The subject matter should be understood by referring to the appropriate portions of the entire specification, any or all of the drawings, and each claim.

[0021] The foregoing and other features and embodiments will become more apparent upon reference to the following description, claims and drawings. Attached Figure Description

[0022] In order to describe the various advantages and features that can be obtained from this disclosure, a more specific description of the above principles will be presented by reference to specific embodiments thereof shown in the accompanying drawings. It is understood that these drawings depict only exemplary embodiments of this disclosure and are not intended to limit its scope; the principles herein are described and explained with additional specificity and detail using the drawings, in which:

[0023] Figure 1 This is a block diagram illustrating examples of encoding and decoding devices according to some examples of this disclosure;

[0024] Figure 2A This is a conceptual diagram illustrating example space adjacent motion vector candidates for merging patterns, according to some examples of this disclosure;

[0025] Figure 2B This is a conceptual diagram illustrating example space neighboring motion vector candidates for advanced motion vector prediction (AMVP) patterns according to some examples of this disclosure;

[0026] Figure 3A This is a conceptual diagram illustrating example time motion vector predictor (TMVP) candidates according to some examples of this disclosure;

[0027] Figure 3B This is a conceptual diagram illustrating an example of motion vector scaling based on some examples of this disclosure;

[0028] Figure 4A This is a conceptual diagram illustrating examples of neighboring samples of the current decoding unit for estimating motion compensation parameters for the current decoding unit, according to some examples of this disclosure;

[0029] Figure 4B This is a conceptual diagram illustrating an example of neighboring samples of a reference block used to estimate motion compensation parameters for the current decoding unit, according to some examples of this disclosure;

[0030] Figure 5 This is a diagram illustrating an example of OBMC hybridization for a decoder boundary overlap block motion compensation (OBMC) mode according to some examples of this disclosure;

[0031] Figure 6 This is a diagram illustrating an example of OBMC blending for a sub-block boundary overlap block motion compensation (OBMC) mode according to some examples of this disclosure;

[0032] Figure 7 and Figure 8 This is a table showing examples of the sum of weighted factors from overlapping block motion compensation sub-blocks used for overlapping block motion compensation, according to some examples of this disclosure;

[0033] Figure 9 This is a diagram illustrating example decoding units with sub-blocks in a block of video data according to some examples of this disclosure;

[0034] Figure 10 This is a flowchart illustrating an example process for performing overlapping block motion compensation according to some examples of this disclosure;

[0035] Figure 11 This is a flowchart illustrating another example process for performing overlapping block motion compensation, according to some examples of this disclosure;

[0036] Figure 12 This is a block diagram illustrating an example video encoding device according to some examples of this disclosure; and

[0037] Figure 13 This is a block diagram illustrating an example video decoding device according to some examples of this disclosure. Detailed Implementation

[0038] Certain aspects and embodiments of this disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments can be applied independently, and some can be applied in combination. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments may be practiced without these specific details. The drawings and description are not intended to be limiting.

[0039] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of these exemplary embodiments will provide those skilled in the art with a feasible description for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the scope of this application as set forth in the appended claims.

[0040] Video compression techniques used in video decoding can include applying different prediction modes (including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques) to reduce or remove redundancy inherent in the video sequence. A video encoder can segment each frame of the original video sequence into rectangular regions, referred to as video blocks or decoding units (described in more detail below). These video blocks can be encoded using specific prediction modes.

[0041] Motion compensation is typically used when decoding video data for video compression. In some examples, motion compensation may include and / or implement algorithmic techniques for predicting frames in a video based on previous and / or future frames by taking into account the motion of elements (e.g., objects) in the camera and / or video. Motion compensation can describe an image based on a transformation from a reference image to the current image. The reference image may be a temporally previous image or even an image from the future. In some examples, motion compensation can improve compression efficiency by allowing accurate synthesis of images from previously sent and / or stored images.

[0042] An example of motion compensation techniques includes Block Motion Compensation (BMC) (also known as Motion Compensated Discrete Cosine Transform (MC DCT)), where frames are segmented into non-overlapping pixel blocks, and each block is predicted from one or more blocks in one or more reference frames. In BMC, blocks are shifted to the location of the predicted block. This shift is represented by a motion vector (MV), or motion compensation vector. To take advantage of redundancy between adjacent block vectors, BMC can be used to encode only the difference between the current and previous motion vectors in the video bitstream. In some cases, BMC may introduce discontinuities at block boundaries (e.g., block artifacts). Such artifacts may appear as sharp horizontal and vertical edges, often perceptible to the human eye, and produce false edges and ringing effects (e.g., large coefficients in high-frequency subbands) due to the quantization of the coefficients of the Fourier correlation transform used for transform decoding of the residual frame.

[0043] Typically, in BMC (Block Motion Compensation), the current reconstructed block consists of a predicted block from a previous frame (e.g., referenced by motion vectors) and residual data transmitted in the bitstream used for the current block. Another example of motion compensation techniques is Overlapping Block Motion Compensation (OBMC). OBMC can improve prediction accuracy and avoid block artifacts. In OBMC, predictions can be, or can include, a weighted sum of multiple predictions. In some cases, the block may be larger in every dimension and may overlap with adjacent blocks. In this case, each pixel may belong to multiple blocks. For example, in some illustrative examples, each pixel may belong to four different blocks. In such a scheme, OBMC can implement four predictions for each pixel, which are summed to calculate a weighted average.

[0044] In some cases, specific syntax (e.g., one or more specific syntax elements) can be used at the CU level to turn OBMC on and off. In some examples, there are two orientation modes in the OBMC (e.g., top, left, right, bottom, or below), including the CU boundary OBMC mode and the sub-block boundary OBMC mode. When using the CU boundary OBMC mode, the original prediction block of the current CU MV is mixed with another prediction block (e.g., the "OBMC block") from the adjacent CU MV. In some examples, the top-left sub-block in the CU (e.g., the first or leftmost sub-block in the first / top row of the CU) has both top and left OBMC blocks, while other topmost sub-blocks (e.g., other sub-blocks in the first / top row of the CU) may only have top OBMC blocks. Other leftmost sub-blocks (e.g., sub-blocks to the left of the CU in the first column of the CU) may only have left OBMC blocks.

[0045] When sub-CU decoding tools are enabled in the current CU, sub-block boundary OBMC modes (e.g., affine motion compensation prediction, advanced temporal motion vector prediction (ATMVP), etc.) can be enabled. In sub-block boundary modes, the MV of the current sub-block can be used to sequentially blend individual OBMC blocks using the MVs of connected adjacent sub-blocks with the original prediction block. In some cases, CU boundary OBMC modes can be performed before sub-block boundary OBMC modes, and the predefined blending order for sub-block boundary OBMC modes can include top, left, bottom, and right.

[0046] The prediction of the MV based on the neighboring sub-blocks N (e.g., the sub-blocks above, to the left, below, and to the right of the current sub-block) can be represented as P. N The prediction of the MV based on the current sub-block can be represented as P. CWhen sub-block N contains the same motion information as the current sub-block, the original prediction block may not be mixed with the prediction block based on the MV of sub-block N. In some cases, P can be... N The four rows / columns of the sample and P C The same samples are mixed together. In some examples, weighting factors of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 can be used for P. N Furthermore, the corresponding weighting factors 3 / 4, 7 / 8, 15 / 16, and 31 / 32 can be used for P. C In some cases, if the height or width of the decoded block is equal to four, or if the CU is decoded using a sub-CU mode, then only P is allowed. N Two rows and / or two columns are used for OBMC mixing.

[0047] This document describes systems, apparatuses, methods, and computer-readable media (collectively, “Systems and Techniques”) for performing improved video decoding. In some aspects, the systems and techniques described herein can be used to perform Overlapping Block Motion Compensation (OBMC). For example, Local Illumination Compensation (LIC) is a decoding tool that uses a linear model with scaling factors and offsets to change the illumination of the currently predicted block based on a reference block. In some aspects, since both OBMC and LIC tune the prediction, the systems and techniques described herein can disable OBMC when LIC is enabled, or can disable LIC when OBMC is enabled. Alternatively, in some aspects, the systems and techniques described herein can skip OBMC signaling when LIC is enabled, or skip LIC signaling when OBMC is enabled.

[0048] In some aspects, the systems and techniques described herein can implement multi-hypothesis prediction (MHP) to improve inter-frame prediction modes, such as Advanced Motion Vector Prediction (AMVP) mode, skip and merge modes, and intra-frame modes. In some examples, the systems and techniques described herein can combine prediction modes with additional merge index prediction. Merge index prediction can be performed as in merge mode, where the merge index is signaled to obtain motion information for motion compensation prediction. Since OBMC and MHP typically require access to different reference images for prediction, the decoder can utilize large buffers for processing. To reduce memory buffer usage, the systems and techniques described herein can disable OBMC when MHP is enabled, or disable MHP when OBMC is enabled. In other examples, the systems and techniques described herein can alternatively skip OBMC signaling when MHP is enabled, or skip MHP signaling when OBMC is enabled. In some cases, the systems and techniques described herein can allow MHP and OBMC to be enabled simultaneously when the current slice is an inter-frame B-slice.

[0049] Some video decoding standards (such as VVC) support Geometric Partitioning (GEO) mode for inter-frame prediction. When using this mode, the CU (Computer Unit) can be split into two parts by geometrically positioned lines. The location of the split lines can be mathematically derived from the angle and offset parameters of a specific partition. Since OBMC and GEO typically require access to different reference images for prediction, the decoder can utilize large buffers. In some cases, to reduce memory buffer usage, the systems and techniques described herein can disable OBMC when GEO is enabled, disable GEO when OBMC is enabled, skip OBMC signaling when GEO is enabled, or skip GEO signaling when OBMC is enabled. In some cases, when the current slice is an inter-frame B-slice, it is permissible to enable both GEO and OBMC simultaneously.

[0050] In some video decoding standards (such as VVC), affine motion compensation prediction, sub-block-based temporal motion vector prediction (SbTMVP), and decoder-side motion vector refinement (DMVR) can be supported for inter-frame prediction. These decoding tools generate different motion vector predictions (MVs) for sub-blocks in the CU. The SbTMVP mode can be one of the affine merging candidates. Therefore, in some examples, the systems and techniques described herein can allow the sub-block boundary OBMC mode to be enabled when the current CU uses the affine motion compensation prediction mode, when the current CU enables SbTMVP, or when the current CU enables DMVR. In some cases, the systems and techniques described herein can infer that the sub-block boundary OBMC mode is enabled when the current CU enables DMVR.

[0051] In some cases, different weighting factors can be applied to the CU boundary OBMC model and / or the sub-block boundary OBMC model. In other cases, the CU boundary OBMC model and the sub-block boundary OBMC model can share the same weighting factors. For example, in JEM, the CU boundary OBMC model and the sub-block boundary OBMC model can share the same weighting factors as follows: the final prediction used for mixing can be expressed as P = W C * P C + W N * P N , where P N P represents the prediction of the MV based on the neighboring sub-blocks N (e.g., the sub-blocks above, to the left, below, and to the right). C It is a prediction based on the MV of the current sub-block, and the CU boundary OBMC mode and the sub-block boundary OBMC mode use the same value for W. C and W N For the current sub-block, the nearest sample row / column to the first, second, third, and fourth neighboring sub-blocks N can be weighted by the weighting factor W. NThe values ​​are set to 1 / 4, 1 / 8, 1 / 16, and 1 / 32. Sub-blocks can have a size of 4×4. The first element, 1 / 4, represents the row or column of samples closest to the neighboring sub-block N, while the last element, 1 / 32, represents the row or column of samples farthest from the neighboring sub-block N. The weight W of the current sub-block... C It can be equal to 1 – W N (Weights of adjacent sub-blocks). Since sub-blocks in a CU used for sub-CU patterns may have more connections with neighboring blocks, the weighting factor used for the sub-block boundary OBMC pattern can differ from the weighting factor used for the CU boundary OBMC pattern. Therefore, the system and techniques described in this paper can provide different weighting factors.

[0052] In some examples, the weighting factor can be as follows. In the CU boundary OBMC mode, W N It can be set to {a1, b1, c1, d1}. Otherwise, W N It can be set to {a2, b2, c2, d2}, where {a1, b1, c1, d1} is different from {a2, b2, c2, d2}. In the example, a2 can be less than a1, b2 can be less than b1, c2 can be less than c1, and / or d2 can be less than d1.

[0053] In JEM, the predefined mixing order for sub-block boundary OBMC patterns is top, left, bottom, and right. In some cases, this order can increase computational complexity, degrade performance, lead to unequal weighting, and / or cause inconsistencies. In some examples, this order can cause problems because sequential computation is not friendly to parallel hardware designs. In some cases, this can lead to unequal weighting. For example, during the mixing process, the OBMC blocks of adjacent sub-blocks may contribute more to the final sample prediction in a later sub-block mix than in a previous sub-block mix. The system and techniques described in this paper can mix the prediction of the current sub-block with four OBMC sub-blocks in a single formula, fixing the weighting factors without bias towards any particular adjacent sub-block. For example, the final prediction could be P = w1 * P c + w2 * P top + w3 * P left +w4 * P below + w5 * P right , where P top It is a prediction based on the MV of the top adjacent sub-block, P left It is a prediction based on the MV of the left-adjacent sub-block, P below It is a prediction based on the MV of the adjacent sub-blocks below, P rightThe prediction is based on the MV of the right-hand adjacent sub-block, and w1, w2, w3, w4, and w5 are weighting factors. In some cases, the weight w1 can be equal to 1 – w2 – w3 – w4 – w5. Because the prediction of the MV based on the adjacent sub-block N may add / include / introduce noise to the samples in the row / column farthest from sub-block N, the system and techniques described herein can set the value of each of the weights w2, w3, w4, and w5 to {a, b, c, 0} for the sample rows / columns that are respectively closest to the adjacent sub-block N {first, second, third, fourth} of the current sub-block. For example, the first element a can be the sample row or column that is closest (e.g., adjacent) to the adjacent sub-block N for the current sub-block, and the last element 0 can be the sample row or column that is farthest from the adjacent sub-block N for the current sub-block. Using the positions (0, 0), (0, 1), and (1, 1) relative to the top-left sample of the current sub-block with a size of 4×4 samples as examples, the final prediction P(x, y) can be derived as follows:

[0054] P(0, 0) = w1 * P c (0, 0) + a * P top (0, 0) + a * P left (0, 0)

[0055] P(0, 1) = w1 * P c (0, 1) + b * P top (0, 1) + a * P left (0, 1) + c * P below (0, 1)

[0056] P(1, 1) = w1 * P c (1, 1) + b * P top (1, 1) + b * P left (1, 1) + c * P below (1, 1) + c * P right (1, 1)

[0057] For a 4×4 current sub-block, an example of the sum of weighting factors from adjacent OBMC sub-blocks (e.g., w2 + w3 + w4 + w5) can be shown in Table 1 below. In some cases, the weighting factors can be shifted left to avoid division. For example, {a', b', c', 0} can be set to {a << shift, b << shift, c << shift, 0}, where shift is a positive integer. In this example, the weight w1 can be equal to (1 << shift) – a' – b' – c', and P can be equal to (w1 * P) c + w2 * P top + w3 * P left + w4 * P below + w5 * P right +(1<<(shift-1)))>> shift. An example of setting {a', b', c', 0} is {15, 8, 3, 0}, where the value is the result of shifting the original value 6 left, and w1 equals (1 << 6) – a – b – c. P = (w1 * P) c + w2 * P top + w3 * P left + w4 * P below + w5 * P right +(1<<5))>> 6.

[0058] Table 1. Sum of weighting factors from OBMC sub-blocks for {a, b, c, 0}

[0059] a+b+c 2b+2c 2b+2c a+b+c a+b+c 2b+2c 2b+2c a+b+c 2a a+b+c a+b+c 2a

[0060] In some respects, for the row / column of the sample closest to the neighboring sub-blocks N {first, second, third, fourth} respectively in the current sub-block, the values ​​of w2, w3, w4, and w5 can be set to {a, b, 0, 0}. Using the positions (0, 0), (0, 1), and (1, 1) of the top-left sample relative to the current sub-block with a size of 4×4 samples as an example, the final prediction P(x, y) can be derived as follows:

[0061] P(0, 0) = w1 * P c (0, 0) + a * P top (0, 0) + a * P left (0, 0)

[0062] P(0, 1) = w1 * P c (0, 1) + b * P top(0, 1) + a * P left (0, 1)

[0063] P(1, 1) = w1 * P c (1, 1) + b * P top (1, 1) + b * P left (1, 1)

[0064] For the current 4×4 sub-block, the example sum of weighting factors from adjacent OBMC sub-blocks (e.g., w2 + w3 + w4 + w5) is shown in Table 2 below.

[0065] Table 2. Sum of weighting factors from OBMC sub-blocks for {a, b, 0, 0}

[0066] a+b 2b 2b a+b a+b 2b 2b a+b 2a a+b a+b 2a

[0067] In some examples, weights can be chosen such that the sum of w2 + w3 + w4 + w5 at corner samples (e.g., samples at (0,0), (0,3), (3,0), and (3,3)) is greater than the sum of w2 + w3 + w4 + w5 at other boundary samples (e.g., samples at (0,1), (0,2), (1,0), (2,0), (3,1), (3,2), (1,3), and (2,3)), and / or the sum of w2 + w3 + w4 + w5 at boundary samples is greater than the value at intermediate samples (e.g., samples at (1,1), (1,2), (2,1), and (2,2)).

[0068] In some cases, motion compensation is skipped during OBMC based on the similarity between the current sub-block's MV and the MVs of its spatial neighbors (e.g., top, left, bottom, and right). For example, before each invocation of motion compensation using motion information from a given neighboring block / sub-block, the MV of the neighboring block / sub-block can be compared to the current sub-block's MV based on one or more of the following conditions. One or more conditions may include, for example: a first condition that all prediction lists used by the neighboring blocks / sub-blocks (e.g., list L0 or list L1 in unidirectional prediction, or both L0 and L1 in bidirectional prediction) are also used for the prediction of the current sub-block; a second condition that the MVs of the neighboring sub-blocks and the current sub-block use the same reference picture; and / or a third condition that the absolute value of the horizontal MV difference between the neighboring MV and the current MV is not greater than (or does not exceed) a predefined MV difference threshold T and the absolute value of the vertical MV difference between the neighboring MV and the current MV is not greater than a predefined MV difference threshold T (if bidirectional prediction is used, both L0 and L1 MVs can be checked).

[0069] In some examples, if the first, second, and third conditions are met, motion compensation using a given neighboring block / sub-block is not performed, and the OBMC sub-block using the MV of a given neighboring block / sub-block N is disabled and not mixed with the original sub-block. In some cases, the CU boundary OBMC mode and the sub-block boundary OBMC mode can have different values ​​for the threshold T. If the mode is the CU boundary OBMC mode, T is set to T1; otherwise, T is set to T2, where T1 and T2 are greater than 0. In some cases, when conditions are met, the lossy algorithm for skipping neighboring blocks / sub-blocks can be applied only to the sub-block boundary OBMC mode. The CU boundary OBMC mode can alternatively apply a lossless algorithm that skips neighboring blocks / sub-blocks when one or more of the following conditions are met: a fourth condition that all prediction lists used by neighboring blocks / sub-blocks (e.g., L0 or L1 in one-way prediction, or both L0 and L1 in two-way prediction) are also used for the prediction of the current sub-block; a fifth condition that neighboring MVs and the current MV use the same reference picture; and a sixth condition that neighboring MVs and the current MV are the same (if two-way prediction is used, both L0 and L1 MVs can be checked).

[0070] In some cases, when conditions one, two, and three are met, the lossy algorithm for skipping adjacent blocks / sub-blocks is only applied to the CU boundary OBMC mode. In some cases, when conditions four, five, and six are met, the lossless algorithm for skipping adjacent blocks / sub-blocks can be applied to the sub-block boundary OBMC mode.

[0071] In some aspects, lossy fast algorithms can be implemented in CU boundary OBMC mode to save encoding and decoding time. For example, the first OBMC block and adjacent OBMC blocks can be merged into a larger OBMC block and generated together if one or more conditions are met. One or more conditions may include, for example, the following: a condition that all prediction lists used by the first adjacent block of the current CU (e.g., L0 or L1 in unidirectional prediction, or both L0 and L1 in bidirectional prediction) are also used for the prediction of the second adjacent block of the current CU (in the same direction as the first adjacent block); a condition that the MV of the first adjacent block and the MV of the second adjacent block use the same reference picture; and a condition that the absolute value of the horizontal MV difference between the MV of the first adjacent block and the MV of the second adjacent block is not greater than a predefined MV difference threshold T3 and the absolute value of the vertical MV difference between the MV of the first adjacent block and the MV of the second adjacent block is not greater than a predefined MV difference threshold T3 (if bidirectional prediction is used, both L0 and L1 MVs can be checked).

[0072] In some aspects, lossy fast algorithms can be implemented in subblock boundary OBMC mode to save encoding and decoding time. In some examples, SbTMVP mode and DMVR are performed on an 8×8 basis, and affine motion compensation is performed on a 4×4 basis. The systems and techniques described herein can implement subblock boundary OBMC mode on an 8×8 basis. In some cases, the systems and techniques described herein can perform a similarity check at each 8×8 subblock to determine whether the 8×8 subblock should be split into four 4×4 subblocks, and if so, perform OBMC on a 4×4 basis. In some examples, the algorithm may include: for each 8×8 sub-block, enabling four 4×4 OBMC sub-blocks (e.g., P, Q, R, and S) when at least one of the following conditions is not met: a first condition that the prediction lists used by sub-blocks P, Q, R, and S are the same (e.g., L0 or L1 in one-way prediction, or both L0 and L1 in two-way prediction); a second condition that the MVs of sub-blocks P, Q, R, and S use the same reference picture; and a third condition that the absolute value of the horizontal MV difference between the MVs of any two sub-blocks (e.g., P and Q, P and R, P and S, Q and R, Q and S, and R and S) is not greater than a predefined MV difference threshold T4 and the absolute value of the vertical MV difference between the MVs of any two sub-blocks (e.g., P and Q, P and R, P and S, Q and R, Q and S, and R and S) is not greater than a predefined MV difference threshold T4 (if two-way prediction is used, both L0 and L1 MVs may be checked).

[0073] If all the above conditions are met, the system and techniques described herein can perform 8×8 sub-block OBMC, wherein 8×8 OBMC sub-blocks from the top, left, bottom, and right MVs are generated using OBMC blending for sub-block boundary OBMC modes. Otherwise, when at least one of the above conditions is not met, OBMC is performed on a 4×4 basis within the 8×8 sub-block, and each 4×4 sub-block in the 8×8 sub-block generates four OBMC sub-blocks from the top, left, bottom, and right MVs.

[0074] In some aspects, when the CU is decoding using merge mode, the OBMC flag is copied from the adjacent block in a manner similar to motion information copying in merge mode. Otherwise, when the CU is not decoding using merge mode, the OBMC flag can be signaled to the CU to indicate whether OBMC is applicable.

[0075] The systems and techniques described herein can be applied to any of existing video codecs (e.g., High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), or other suitable existing video codecs), and / or can be efficient decoding tools for any video decoding standard under development and / or future video decoding standards, such as Universal Video Decoding (VVC), Joint Exploratory Model (JEM), VP9, ​​AV1 format / codec, and / or other video decoding standards under development or to be developed.

[0076] Further details about the system and technology will be described using the diagrams.

[0077] Figure 1 This is a block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or receiving device may include electronic devices such as mobile or landline phones (e.g., smartphones, cellular phones, etc.), desktop computers, laptop or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. In some examples, the source device and receiving device may include one or more wireless transceivers for wireless communication. The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcasting or transmission, encoding of digital video for storage on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. As used herein, the term decoding may refer to both encoding and / or decoding. In some examples, system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.

[0078] Encoding device 104 (or encoder) can be used to encode video data using video decoding standards, formats, codecs, or protocols to generate an encoded video bitstream. Examples of video decoding standards and formats / codecs include ITU-T H.261, ISO / IEC MPEG-1 video (visual), ITU-T H.262 or ISO / IEC MPEG-2 video, ITU-T H.263, ISO / IEC MPEG-4 video, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Decoding (SVC) and Multi-View Video Decoding (MVC) extensions), High Efficiency Video Decoding (HEVC) or ITU-T H.265, and Universal Video Decoding (VVC) or ITU-T H.266. Various extensions to HEVC exist for handling multi-layer video decoding, including range and screen content decoding extensions, 3D video decoding (3D-HEVC), multi-view extensions (MV-HEVC), and scalable extensions (SHVC). The ITU-T Video Decoding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) Joint Collaborative Working Group on Video Decoding (JCT-VC), as well as the Joint Collaborative Working Group on 3D Video Decoding Extensions (JCT-3V), have developed HEVC and its extensions. VP9, ​​AOMedia Video 1 (AV1) developed by the Alliance for Open Media (AOMedia), and Basic Video Decoding (EVC) are the technologies described in this paper that can be applied to other video decoding standards.

[0079] The techniques described herein can be applied to any of existing video codecs (e.g., High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), or other suitable existing video codecs), and / or can be efficient decoding tools for any video decoding standard under development and / or future video decoding standards (such as VVC and / or other video decoding standards under development or to be developed). For example, the examples described herein can be performed using video codecs such as VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein can also be applied to other decoding standards, codecs, or formats, such as MPEG, JPEG (or other decoding standards for still images), VP9, ​​AV1, their extensions, or other suitable decoding standards that are already available or not yet available or under development. For example, in some examples, encoding device 104 and / or decoding device 112 can operate according to proprietary video codecs / formats (such as AV1, extensions of AVI, and / or subsequent versions of AV1 (e.g., AV2) or other proprietary formats or industry standards). Therefore, although the techniques and systems described herein may be referenced to specific video decoding standards, it will be understood by those skilled in the art that the description should not be construed as applicable only to that particular standard.

[0080] refer to Figure 1 Video source 102 can provide video data to encoding device 104. Video source 102 can be part of a source device, or it can be part of a device other than a source device. Video source 102 can include video capture devices (e.g., camcorders, mobile phone cameras, video phones, etc.), video storage units containing stored video, video servers or content providers providing video data, video feed interfaces receiving video from video servers or content providers, computer graphics systems for generating computer graphics video data, combinations of such sources, or any other suitable video source.

[0081] Video data from video source 102 may include one or more input pictures or frames. A picture or frame is a still image that, in some cases, is part of the video. In some examples, the data from video source 102 may be a still image that is not part of the video. In HEVC, VVC, and other video decoding specifications, a video sequence may include a series of pictures. A picture may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. Chrominance samples may also be referred to herein as “chroma” samples. A pixel may refer to all three components (luminance and chrominance samples) at a given location in the array of pictures. In other cases, a picture may be monochrome and may include only an array of luminance samples, in which case the terms pixel and sample may be used interchangeably. The same techniques used to refer to the various samples described herein for illustrative purposes can be applied to pixels (e.g., all three sample components at a given location in the array of pictures). The same technique can be applied to individual samples, referring to the example techniques described in this paper for illustrative purposes for the reference pixels (e.g., all three sample components at a given location in an array of images).

[0082] Encoding device 104's encoder engine 106 (or encoder) encodes the video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. A decoded video sequence (CVS) comprises a series of access units (AUs) starting from an AU that has a random access point picture in the base layer and certain attributes, continuing to the next AU that has a random access point picture in the base layer and certain attributes, and excluding that next AU. For example, certain attributes of the random access point picture that starts the CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (where the RASL flag is equal to 0) does not start a CVS. An access unit (AU) comprises one or more decoded pictures and control information corresponding to decoded pictures sharing the same output time. Decoded slices of pictures are encapsulated at the bitstream level as data units called Network Abstraction Layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs, which include NAL units. Each of the NAL units has a NAL unit header. In one example, the header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header use specified bits and are therefore visible to all kinds of systems and transport layers, such as transport streams, Real-Time Transport (RTP) protocols, file formats, and others.

[0083] The HEVC standard contains two types of NAL units: Video Decoding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain decoded picture data that forms the decoded video bitstream. For example, the bit sequence forming the decoded video bitstream resides within a VCL NAL unit. A VCL NAL unit may include a slice or fragment of the decoded picture data (described below), while non-VCL NAL units contain control information relating to one or more decoded pictures. In some cases, NAL units may be referred to as packets. A HEVC AU includes: VCL NAL units containing decoded picture data, and non-VCL NAL units (if any) corresponding to the decoded picture data. Among other information, non-VCL NAL units may also contain a set of parameters with high-level information relating to the encoded video bitstream. For example, the parameter set may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). In some cases, each slice or other portion of the bitstream may reference a single valid PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slices or other portions of the bitstream.

[0084] NAL units can contain bit sequences that form a decoded representation of video data (e.g., an encoded video bitstream, a CVS of a bitstream, etc.), such as a decoded representation of images in a video. Encoder engine 106 generates a decoded representation of an image by dividing each image into multiple slices. Slices are independent of other slices, allowing information within a slice to be decoded without relying on data from other slices within the same image. A slice includes one or more segments, comprising independent segments and (if present) one or more dependent segments that depend on previous segments.

[0085] In HEVC, the slice is then divided into decode tree blocks (CTBs) for luma and chroma samples. One or more CTBs for luma samples and one CTB for chroma samples, along with the syntax used for the samples, are called decode tree units (CTUs). CTUs can also be called "tree blocks" or "maximum decoder units" (LCUs). A CTU is the basic processing unit used for HEVC encoding. A CTU can be subdivided into multiple decoder units (CUs) of different sizes. A CU contains arrays of luma and chroma samples called decoder blocks (CBs).

[0086] Luminance and chrominance CBs can be further subdivided into prediction blocks (PBs). A PB is a sample block of either the luminance or chrominance component that uses the same motion parameters for inter-frame prediction or intra-block copy (IBC) prediction (when available or enabled for use). A luminance PB and one or more chrominance PBs, along with their associated syntax, form a prediction unit (PU). For inter-frame prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and used for inter-frame prediction of the luminance PB and one or more chrominance PBs. Motion parameters can also be referred to as motion information. CBs can also be subdivided into one or more transform blocks (TBs). A TB represents a square block of samples of the chrominance component, to which a residual transform (e.g., in some cases, the same two-dimensional transform) is applied to decode the prediction residual signal. A transform unit (TU) represents a TB of luminance and chrominance samples and the corresponding syntax elements. Transform decoding is described in more detail below.

[0087] The size of a CU corresponds to the size of the decoding mode and can be square. For example, the size of a CU can be 8×8 samples, 16×16 samples, 32×32 samples, 64×64 samples, or any other suitable size up to the corresponding CTU size. The phrase "N×N" is used herein to refer to the pixel size of a video block in both the vertical and horizontal dimensions (e.g., 8 pixels × 8 pixels). Pixels in a block can be arranged in columns and rows. In some implementations, a block may not have the same number of pixels in the horizontal direction as it does in the vertical direction. The syntax data associated with a CU can describe, for example, the segmentation of the CU into one or more PUs. The segmentation mode can differ between CUs encoded using intra-frame prediction mode and inter-frame prediction mode. PUs can be segmented into non-square shapes. The syntax data associated with a CU can also describe, for example, the segmentation of the CU into one or more TUs according to the CTU. TUs can be square or non-square.

[0088] According to the HEVC standard, transform units (TUs) can be used to perform transforms. The TU can be different for different CUs. The size of the TU can be set based on the size of the PU within a given CU. The TU can have the same size as the PU or be smaller than the PU. In some examples, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can correspond to TUs. The pixel differences associated with the TU can be transformed to produce transform coefficients. The transform coefficients can then be quantized by the encoder engine 106.

[0089] Once the video data is segmented into Units (CUs), the encoder engine 106 uses a prediction mode to predict each Unit (PU). The prediction unit or block is then subtracted from the original video data to obtain the residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. Prediction modes can include intra-frame prediction (or intra-picture prediction) or inter-frame prediction (or inter-picture prediction). Intra-frame prediction utilizes the correlation between samples that are spatially adjacent within a picture. For example, using intra-frame prediction, each PU is predicted from neighboring image data in the same picture, using, for example, DC prediction to find an average value for the PU, planar prediction to adapt a planar surface to the PU, orientation prediction to infer from neighboring data, or any other appropriate prediction type. Inter-frame prediction uses temporal correlations between pictures to derive motion-compensated predictions for blocks of image samples. For example, using inter-frame prediction, each PU is predicted from image data in one or more reference pictures (in the output order before or after the current picture) using motion-compensated prediction. For example, a decision can be made at the CU level whether to use inter-picture prediction or intra-picture prediction to decode a picture region.

[0090] Encoder engine 106 and decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, the video decoder (such as encoder engine 106 and / or decoder engine 116) segments the image into multiple decoder tree units (CTUs) (wherein, the CTB of luminance samples and one or more CTBs of chrominance samples, along with the syntax used for the samples, are collectively referred to as CTUs). The video decoder can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple segmentation types, such as the distinction between CU, PU, ​​and TU in HEVC. The QTBT structure includes two levels: a first level segmented according to quadtree segmentation and a second level segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoder units (CUs).

[0091] In the MTT partitioning structure, blocks can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. A ternary tree partition is a partition in which a block is divided into three sub-blocks. In some examples, a ternary tree partition divides a block into three sub-blocks without using a center to partition the original block. The partitioning types in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.

[0092] When operating according to the AV1 codec, encoding device 104 and decoding device 112 can be configured to decode video data in blocks. In AV1, the largest decoded block that can be processed is called a superblock. In AV1, a superblock can be 128×128 luminance samples or 64×64 luminance samples. However, in subsequent video decoding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luminance sample sizes. In some examples, the superblock is the top level of a block quadtree. Encoding device 104 can further divide the superblock into smaller decoded blocks. Encoding device 104 can use square or non-square partitioning to divide the superblock and other decoded blocks into smaller blocks. Non-square blocks can include N / 2×N, N×N / 2, N / 4×N, and N×N / 4 blocks. Encoding device 104 and decoding device 112 can perform separate prediction and transformation processes for each in the decoded block.

[0093] AV1 also defines tiles for video data. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the encoding device 104 and the decoding device 112 can encode and decode the decoding blocks within a tile separately, without using video data from other tiles. However, the encoding device 104 and the decoding device 112 can perform filtering across tile boundaries. Tiles can be uniform or non-uniform in size. Tile-based decoding enables parallel processing and / or multithreading for encoder and decoder implementations.

[0094] In some examples, the video decoder may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video decoder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for the two chrominance components (or two QTBT and / or MTT structures for the respective chrominance components).

[0095] The video decoder can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partitioning structures.

[0096] In some examples, one or more slices of an image are assigned slice types. Slice types include intra-decoded slices (I-slices), inter-decoded P-slices, and inter-decoded B-slices. An I-slice (intra-decoded, independently decodeable) is a slice of an image that is decoded solely by intra-frame prediction and is therefore independently decodeable because an I-slice requires only intra-frame data to predict any prediction unit or prediction block of the slice. A P-slice (one-way prediction frame) is a slice of an image that can be decoded using both intra-frame and one-way inter-frame prediction. Each prediction unit or prediction block within a P-slice is decoded using either intra-frame or inter-frame prediction. When inter-frame prediction is applied, the prediction unit or prediction block is predicted using only one reference image, and therefore the reference sample comes from only one reference region within a frame. A B-slice (two-way prediction frame) is a slice of an image that can be decoded using both intra-frame and inter-frame prediction (e.g., two-way or one-way prediction). Bidirectional prediction of a B-slice's prediction unit or block can be performed from two reference images, where each image contributes a reference region, and the sample sets of the two reference regions are weighted (e.g., using equal weights or different weights) to generate the prediction signal for the bidirectional prediction block. As explained above, a slice of an image is decoded independently. In some cases, an image may be decoded as a single slice.

[0097] As mentioned above, intra-image prediction utilizes the correlation between spatially adjacent samples within the image. Several intra-prediction modes exist (also referred to as "intra-frame modes"). In some examples, intra-frame prediction for luma blocks includes 35 modes, comprising planar modes, DC modes, and 33 angular modes (e.g., diagonal intra-prediction modes and angular modes adjacent to the diagonal intra-prediction modes). The 35 intra-prediction modes are indexed as shown in Table 1 below. In other examples, more intra-frame modes may be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with angular modes may differ from those used in HEVC.

[0098] Table 3 Specification of Intra-Frame Prediction Modes and Associated Names

[0099] 0 INTRA_PLANAR 1 INTRA_DC 2..34 INTRA_ANGULAR2..INTRA_ANGULAR34

[0100] Inter-image prediction utilizes the temporal correlation between images to derive motion-compensated predictions for image sample blocks. A translational motion model is used, where the position of a block in a previously decoded image (reference image) is determined by motion vectors. It means that among them Specifies the horizontal displacement of the reference block relative to the position of the current block, while Specifies the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector ( The motion vector can be integer sample precision (also known as integer precision), in which case the motion vector points to an integer pixel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector ( Motion vectors can have fractional sample precision (also known as fractional pixel precision or non-integer precision) to capture the motion of the underlying object more accurately, without being limited to the integer pixel grid of the reference frame. The precision of a motion vector can be expressed by its quantization level. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., ¼ pixel, ½ pixel, or other values ​​below 1 pixel). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference image to derive the predicted signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values ​​at fractional positions. The previously decoded reference image is indicated by a reference index (refIdx) for a list of reference images. The motion vector and reference index can be referred to as motion parameters. Two types of inter-image prediction can be performed, including one-way prediction and two-way prediction.

[0101] In the case of inter-frame prediction using bidirectional prediction (also known as bidirectional inter-frame prediction), two sets of motion parameters are used ( and The system generates two motion-compensated predictions (from the same reference image or possibly from different reference images). For example, in the case of bidirectional prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. The two motion-compensated predictions are then combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference images that can be used in bidirectional prediction are stored in two separate lists, denoted as List 0 and List 1, respectively. Motion parameters can be derived at the encoder using a motion estimation process.

[0102] In the case of using one-way prediction for inter-frame prediction (also known as one-way inter-frame prediction), a set of motion parameters is used ( This is used to generate motion-compensated predictions from a reference image. For example, in the case of unidirectional prediction, each prediction block uses at most one motion-compensated prediction signal and generates P prediction units.

[0103] The prediction unit (PU) may include data related to the prediction process (e.g., motion parameters or other appropriate data). For example, when the PU is encoded using intra-frame prediction, the PU may include data describing the intra-frame prediction mode used for the PU. As another example, when the PU is encoded using inter-frame prediction, the PU may include data defining the motion vectors used for the PU. The data defining the motion vectors used for the PU may describe, for example, the horizontal components of the motion vectors (…). ), the vertical component of the motion vector ( The resolution used for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference image to which the motion vector points, the reference index, the list of reference images used for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.

[0104] AV1 includes two common techniques for encoding and decoding blocks of video data. These two common techniques are intra-frame prediction (e.g., intra-frame prediction or spatial prediction) and inter-frame prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using intra-frame prediction modes to predict blocks of video data for the current frame, the encoding device 104 and the decoding device 112 do not use video data from other frames of the video data. For most intra-frame prediction modes, the video encoding device 104 encodes the blocks of the current frame based on the difference between sample values ​​in the current block and predicted values ​​generated based on reference samples in the same frame. The video encoding device 104 determines the predicted values ​​generated based on the reference samples according to the intra-frame prediction mode.

[0105] After performing prediction using intra-frame prediction and / or inter-frame prediction, encoding device 104 can perform transform and quantization. For example, after prediction, encoder engine 106 can compute a residual value corresponding to the PU. The residual value can include the pixel difference between the current pixel block (PU) being decoded and the prediction block used to predict the current block (e.g., a predicted version of the current block). For example, after generating a prediction block (e.g., performing inter-frame prediction or intra-frame prediction), encoder engine 106 can generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel differences that quantize the differences between the pixel values ​​of the current block and the pixel values ​​of the prediction block. In some examples, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of the pixel values.

[0106] Block transforms are used to transform any residual data that may remain after prediction is performed. These block transforms can be based on discrete cosine transform, discrete sine transform, integer transform, wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., sizes of 32×32, 16×16, 8×8, 4×4, or other suitable sizes) can be applied to the residual data in each CU. In some embodiments, TUs can be used for transform and quantization processes implemented by encoder engine 106. A given CU having one or more PUs may also include one or more TUs. As described further in detail below, residual values ​​can be transformed into transform coefficients using block transforms, and then quantized and scanned using TUs to produce serialized transform coefficients for entropy decoding.

[0107] In some embodiments, after intra-frame prediction or inter-frame prediction decoding using the PU of the CU, the encoder engine 106 can compute residual data for the TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). The TU may include coefficients in the transform domain after applying the block transform. As previously described, the residual data may correspond to the pixel difference between a pixel in the unencoded image and the predicted value corresponding to the PU. The encoder engine 106 can form a TU including the residual data for the CU, and can then transform the TU to produce transform coefficients for the CU.

[0108] The encoder engine 106 can perform quantization of the transform coefficients. Quantization provides further compression by reducing the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.

[0109] Once quantization is performed, the decoded video bitstream includes quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), segmentation information, and any other suitable data (such as other syntax data). The different elements of the decoded video bitstream can then be entropy-coded by encoder engine 106. In some examples, encoder engine 106 can scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy-coded. In some examples, encoder engine 106 can perform adaptive scanning. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), encoder engine 106 can entropy-code that vector. For example, encoder engine 106 can use context-adaptive variable-length decoding, context-adaptive binary arithmetic decoding, syntax-based context-adaptive binary arithmetic decoding, probabilistic interval segmentation entropy decoding, or another suitable entropy coding technique.

[0110] The output 110 of encoding device 104 can send NAL units constituting the encoded video bitstream data to decoding device 112 of the receiving device on communication link 120. The input 114 of decoding device 112 can receive the NAL units. Communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network may include any wireless interface or combination of wireless interfaces and may include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi). TM Radio frequency (RF), Ultra-wideband (UWB), WiFi Direct, Cellular, Long Term Evolution (LTE), WiMax TM Wired networks can include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, coaxial cable Ethernet, digital signal line (DSL), etc.). Various devices can be used to implement wired and / or wireless networks, such as base stations, routers, access points, bridges, gateways, switches, etc. Encoded video bitstream data can be modulated according to communication standards such as wireless communication protocols and transmitted to receiving devices.

[0111] In some examples, encoding device 104 may store encoded video bitstream data in storage device 108. Output 110 may retrieve the encoded video bitstream data from encoder engine 106 or from storage device 108. Storage device 108 may include any of a variety of distributed or locally accessed data storage media. For example, storage device 108 may include hard disk drives, storage disks, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. Storage device 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In other examples, storage device 108 may correspond to a file server or another intermediate storage device that may store encoded video generated by a source device. In such cases, receiving device including decoding device 112 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing encoded video data and sending such encoded video data to the receiving device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data via any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage device 108 can be streaming, downloading, or a combination thereof.

[0112] Input 114 of decoding device 112 receives encoded video bitstream data and may provide the video bitstream data to decoder engine 116 or to storage device 118 for later use by decoder engine 116. For example, storage device 118 may include a DPB for storing reference pictures used in inter-frame prediction. A receiving device including decoding device 112 may receive the encoded video data to be decoded via storage device 108. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the receiving device. The communication medium used to transmit the encoded video data may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be used to facilitate communication from the source device to the receiving device.

[0113] Decoder engine 116 decodes the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more decoded video sequences that constitute the encoded video data. Decoder engine 116 can then rescale the encoded video bitstream data and perform an inverse transform on it. The residual data is then passed to the prediction stage of decoder engine 116. Decoder engine 116 then predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (the residual data).

[0114] Video decoding device 112 can output decoded video to video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, video destination device 122 may be part of a receiving device that includes decoding device 112. In some aspects, video destination device 122 may be part of a separate device, distinct from the receiving device.

[0115] In some embodiments, the video encoding device 104 and / or the video decoding device 112 may be integrated with the audio encoding device and the audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 may also include other hardware or software necessary for implementing the above-described decoding techniques, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 may be integrated as part of a combined encoder / decoder (codec) in the respective device.

[0116] exist Figure 1 The example system shown is merely an illustrative example that may be used herein. The techniques used to process video data using the techniques described herein can be implemented by any digital video encoding and / or decoding device. Although, in general, the techniques of this disclosure are implemented by video encoding or video decoding devices, the techniques can also be implemented by a combined video encoder-decoder, commonly referred to as a "CODEC". Furthermore, the techniques of this disclosure can also be implemented by a video preprocessor. The source device and receiving device are merely examples of such encoding devices, wherein the source device generates decoded video data for transmission to the receiving device. In some examples, the source device and receiving device may operate in a substantially symmetrical manner, such that each of these devices includes video encoding and decoding components. Thus, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0117] Extensions to the HEVC standard include the Multi-View Video Decoding Extension, known as MV-HEVC, and the Scalable Video Decoding Extension, known as SHVC. MV-HEVC and SHVC extensions share the concept of layered decoding, where different layers are included in the encoded video bitstream. Each layer in the decoded video sequence is addressed by a unique layer identifier (ID). The layer ID can be present in the header of a NAL unit to identify the layer associated with that NAL unit. In MV-HEVC, different layers typically represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided to represent the video bitstream with different spatial resolutions (or picture resolutions) or different reconstruction fidelities. A scalable layer can include a base layer (where layer ID = 0) and one or more enhancement layers (where layer IDs = 1, 2, ..., n). The base layer may conform to the first version of the HEVC profile and represents the lowest available layer in the bitstream. Compared to the base layer, enhancement layers have increased spatial resolution, temporal resolution, or frame rate and / or reconstruction fidelity (or quality). Enhancement layers are organized hierarchically and may depend on (or not depend on) lower layers. In some examples, a single-standard codec can be used to decode different layers (e.g., encoding all layers using HEVC, SHVC, or other decoding standards). In other examples, multi-standard codecs can be used to decode different layers. For example, AVC can be used to decode the base layer, while MV-HEVC extensions of the SHVC and / or HEVC standards can be used to decode one or more enhancement layers.

[0118] Typically, a layer consists of a set of VCL NAL units and a corresponding set of non-VCL NAL units. Specific layer ID values ​​are assigned to NAL units. Layers can be hierarchical in the sense that they can depend on lower layers. A layer set refers to a self-contained set of layers represented within a bitstream, meaning that layers within a layer set can depend on other layers in that set during decoding, but not on any other layers for decoding. Therefore, the layers in a layer set can form independent bitstreams that can represent video content. A set of layers in a layer set can be obtained from another bitstream through sub-bitstream extraction operations. A layer set can correspond to the set of layers that will be decoded when the decoder wants to operate according to certain parameters.

[0119] As previously described, the HEVC bitstream includes a set of NAL units, which includes VCL NAL units and non-VCL NAL units. VCL NAL units contain decoded picture data that forms the decoded video bitstream. For example, a bit sequence forming the decoded video bitstream exists within a VCL NAL unit. Non-VCL NAL units may also contain parameter sets with high-level information relating to the encoded video bitstream, among other information. For example, parameter sets may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). Examples of objectives for parameter sets include bit rate efficiency, error recovery capability, and providing a system-level interface. Each slice references a single valid PPS, SPS, and VPS to access information that the decoding device 112 can use to decode the slice. Decoding identifiers (IDs) may be provided for each parameter set, including a VPS ID, an SPS ID, and a PPS ID. An SPS includes an SPS ID and a VPS ID. A PPS includes a PPS ID and an SPS ID. Each slice header includes a PPS ID. Using these IDs, the valid parameter set for a given slice can be identified.

[0120] The PPS includes information applicable to all slices within a given image. Therefore, all slices in an image reference the same PPS. Slices in different images can also reference the same PPS. The SPS includes information applicable to all images within the same decoded video sequence (CVS) or bitstream. As previously described, a decoded video sequence is a series of access units (AUs) that begins with a random access point image (e.g., an Instant Decoding Reference (IDR) image, a Broken Link Access (BLA) image, or other suitable random access point image) in the base layer and has certain properties (as described above), until the next AU has a random access point image in the base layer and has certain properties (or the end of the bitstream) and does not include that next AU. The information in the SPS may not change between images in the decoded video sequence. Images in a decoded video sequence can use the same SPS. The VPS includes information applicable to all layers within the decoded video sequence or bitstream. The VPS includes a syntax structure with syntax elements applicable to the entire decoded video sequence. In some embodiments, the VPS, SPS, or PPS can be transmitted in-band along with the encoded bitstream. In some embodiments, VPS, SPS, or PPS can be transmitted out of band in a different transmission than the NAL unit containing decoded video data.

[0121] This disclosure may generally relate to "signaling" certain information (such as syntax elements). The term "signaling" can generally refer to the transmission of values ​​for syntax elements and / or other data for decoding encoded video data. For example, video encoding device 104 may signal values ​​for syntax elements in the bitstream. Generally, signaling refers to generating values ​​in the bitstream. As described above, video source 102 may transmit the bitstream to video destination device 122 substantially in real time or not in real time (such as when syntax elements are stored in storage device 108 for later retrieval by video destination device 122).

[0122] Video bitstreams may also include Supplemental Enhancement Information (SEI) messages. For example, an SEI NAL unit can be part of the video bitstream. In some cases, SEI messages may contain information not required by the decoding process. For example, the information in the SEI message may not be necessary for the decoder to decode the video images in the bitstream, but the decoder can use the information to improve the display or processing of the images (e.g., the decoded output). The information in the SEI message can be embedded metadata. In an illustrative example, a decoder-side entity can use the information in the SEI message to improve the visibility of the content. In some cases, certain application standards may mandate the presence of such SEI messages in the bitstream, allowing for quality improvements for all devices conforming to that application standard (e.g., carrying SEI messages for frame-compatible planar stereoscopic 3DTV video formats, where SEI messages are carried for each frame of the video, processing recovery point SEI messages, and using generalized scanning to scan rectangular SEI messages in DVB, among many other examples).

[0123] As described above, for each block, a set of motion information (also referred to herein as motion parameters) may be available. This set of motion information contains motion information for the forward and backward prediction directions. The forward and backward prediction directions can be two prediction directions in a bidirectional prediction mode, in which case the terms "forward" and "backward" do not necessarily have geometric meaning. Instead, "forward" and "backward" correspond to the reference image list 0 (RefPicList0 or L0) and reference image list 1 (RefPicList1 or L1) for the current image. In some examples, when only one reference image list is available for an image or slice, only RefPicList0 is available, and the motion information for each block of the slice is always forward.

[0124] In some examples, motion vectors and their reference indices are used in the decoding process (e.g., motion compensation). Such motion vectors with associated reference indices are represented as a unidirectional set of predictions of motion information. For each prediction direction, the motion information may include a reference index and a motion vector. In some cases, for simplicity, the motion vector itself may be referenced in a manner that assumes it has an associated reference index. The reference index is used to identify a reference image in the current list of reference images (RefPicList0 or RefPicList1). The motion vector has horizontal and vertical components, providing the offset from the coordinate position in the current image to the coordinate position in the reference image identified by the reference index. For example, the reference index may indicate a specific reference image that should be used for a block in the current image, and the motion vector may indicate where the best-matching block in the reference image (the block that best matches the current block) is located in the reference image.

[0125] Picture order counts (POCs) can be used in video decoding standards to identify the display order of pictures. Although it is possible for two pictures within a decoded video sequence to have the same POC value, this is generally not the case within a decoded video sequence. When multiple decoded video sequences exist in a bitstream, pictures with the same POC value are likely closer to each other in terms of decoding order. The POC value of a picture can be used for reference picture list construction (such as reference picture set derivation in HEVC) and motion vector scaling.

[0126] In H.264 / AVC, each inter-frame large block (MB) can be segmented in four different ways: one 16×16 MB partition; two 16×8 MB partitions; two 8×16 MB partitions; and four 8×8 MB partitions. Different MB partitions within a single MB can have different reference index values ​​(RefPicList0 or RefPicList1) for each direction. In some cases, when the MB is not segmented into four 8×8 MB partitions, it may have only one motion vector for each MB partition in each direction. In some cases, when the MB is segmented into four 8×8 MB partitions, each 8×8 MB partition can be further subdivided into sub-blocks, in which case each sub-block can have a different motion vector in each direction. In some examples, there are four different ways to obtain sub-blocks from 8×8 MB partitions, including: one 8×8 sub-block; two 8×4 sub-blocks; two 4×8 sub-blocks; and four 4×4 sub-blocks. Each sub-block can have a different motion vector in each direction. Therefore, motion vectors can exist at a level equal to or higher than the sub-block level.

[0127] In AVC, for skip and / or direct modes in B slices, the direct temporal mode can be implemented at the MB level or MB partition level. For each MB partition, the motion vector is derived using the motion vector of the block that is co-located with the current MB partition in RefPicList1[0] of the current block. Each motion vector in the co-located block can be scaled based on the POC distance.

[0128] In AVC, spatial direct mode can also be executed. For example, in AVC, direct mode can also predict motion information based on spatial neighbors.

[0129] As mentioned above, in HEVC, the largest decoding unit in a slice is called a decode tree block (CTB). A CTB contains a quadtree whose nodes are decoding units. In the HEVC master profile, the size of the CTB can range from 16×16 to 64×64. In some cases, an 8×8 CTB size is supported. Decoding units (CUs) can have the same size as the CTB and can be as small as 8×8. In some cases, a pattern can be used to decode each decoding unit. When a CU is inter-decoded, it can be further divided into 2 or 4 prediction units (PUs), or when further division is not applicable, the CU can be a single PU. When two PUs exist within a CU, they can be rectangles of half the size, or two rectangles of ¼ or ¾ the size of the CU.

[0130] When the CU is decoded inter-frame, there is a set of motion information for each PU. Additionally, a unique inter-frame prediction mode can be used to decode each PU to deduce the set of motion information.

[0131] For example, for motion prediction in HEVC, there are two inter-frame prediction modes for prediction units (PUs): merge mode and Advanced Motion Vector Prediction (AMVP) mode. Special cases considered as merge are skipped. In AMVP or merge mode, a candidate list of motion vectors (MVs) for multiple motion vector predictors can be maintained. In merge mode, the motion vector of the current PU and its reference index are generated by extracting a candidate from the MV candidate list. In some examples, one or more scaling window offsets can be included in the MV candidate list along with the stored motion vectors.

[0132] In an example where the MV candidate list is used for block motion prediction, the MV candidate list can be constructed separately by the encoding and decoding devices. For example, the MV candidate list can be generated by the encoding device when encoding the block, and can be generated by the decoding device when decoding the block. Information related to motion information candidates in the MV candidate list (e.g., information related to one or more motion vectors, information related to one or more LIC flags that may be stored in the MV candidate list in some cases, and / or other information) can be signaled between the encoding and decoding devices. For example, in merge mode, index values ​​for stored motion information candidates (e.g., in syntax structures such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), slice headers, Supplemental Enhancement Information (SEI) messages sent in or separately from the video bitstream, and / or other signaling) can be signaled from the encoding device to the decoding device. The decoding device can construct the MV candidate list and use the signaled references or indexes to obtain one or more motion information candidates from the constructed MV candidate list for motion compensation prediction. For example, decoding device 112 can construct a list of motion vector (MV) candidates and use motion vectors from index positions (and in some cases, the LIC flag) to perform motion prediction on the block. In the case of AMVP mode, in addition to the reference or index, the difference or residual value can be signaled as an increment. For example, in AMVP mode, the decoding device can construct one or more lists of MV candidates and apply the increment value to one or more motion information candidates obtained using the signaled index value when performing motion compensation prediction on the block.

[0133] In some examples, the MV candidate list contains up to five candidates for the merge mode and two candidates for the AMVP mode. In other examples, different numbers of candidates may be included in the MV candidate list for the merge mode and / or AMVP mode. The merge candidate may contain a set of motion information. For example, the set of motion information may include motion vectors corresponding to two lists of reference images (list 0 and list 1) and reference indices. If the merge candidate is identified by a merge index, the reference image is used for the prediction of the current block and to determine the associated motion vector. However, in AMVP mode, for each potential prediction direction from list 0 or list 1, the reference index, along with the MVP index for the MV candidate list, needs to be explicitly signaled, because the AMVP candidate only contains motion vectors. In AMVP mode, the predicted motion vectors can be further refined.

[0134] As seen above, the merged candidate corresponds to the complete set of motion information, while the AMVP candidate contains only a motion vector and a reference index for a specific prediction direction. Candidates for both modes can be similarly derived from the same spatially and temporally adjacent blocks.

[0135] In some examples, the merge mode allows an inter-frame predicted PU to inherit one or more motion vectors, prediction directions, and one or more reference image indices from an inter-frame predicted PU that includes motion data locations selected from a set of spatially adjacent motion data locations and one of two temporally co-located motion data locations. For AMVP mode, one or more motion vectors of the PU can be predicted and decoded relative to one or more motion vector predictors (MVPs) from an AMVP candidate list constructed by the encoder and / or decoder. In some cases, for unidirectional inter-frame prediction of the PU, the encoder and / or decoder can generate a single AMVP candidate list. In some cases, for bidirectional prediction of the PU, the encoder and / or decoder can generate two AMVP candidate lists, one using motion data from spatially and temporally adjacent PUs from the forward prediction direction, and one using motion data from spatially and temporally adjacent PUs from the backward prediction direction.

[0136] Candidates for both modes can be derived from spatially and / or temporally adjacent blocks. For example, Figure 2A and Figure 2B Includes a conceptual diagram showing spatially adjacent candidates. Figure 2A Candidate spatially adjacent motion vectors (MVs) for merging patterns are shown. Figure 2B Spatial neighbor motion vector (MV) candidates for AMVP mode are shown. Spatial MV candidates are derived from neighboring blocks for a specific PU (PU0), but the method of generating candidates based on blocks differs for merge and AMVP modes.

[0137] In merging mode, the encoder can form a list of merging candidates by considering merging candidates from various motion data locations. For example, ... Figure 2A As shown, regarding in Figure 2A The spatially adjacent motion data positions, indicated by numbers 0 to 4, can derive up to five spatial MV candidates. In the merged candidate list, the MV candidates can be ordered according to the sequence indicated by numbers 0 to 4. For example, positions and order can include: left position (0), top position (1), upper right position (2), lower left position (3), and upper left position (4). Figure 2AIn this context, block 200 includes PU0 202 and PU1 204. In some examples, when the video decoder uses a merging mode to decode motion information for PU0 202, the video decoder can add motion information from spatial neighbor blocks 210, 212, 214, 216, and 218 to the candidate list in the order described above.

[0138] exist Figure 2B In the AVMP mode shown, adjacent blocks can be divided into two groups: the left group including blocks 0 and 1, and the upper group including blocks 2, 3, and 4. Figure 2B In this context, blocks 0, 1, 2, 3, and 4 are labeled as blocks 230, 232, 234, 236, and 238, respectively. Here, block 220 includes PU0 222 and PU1 224, and blocks 230, 232, 234, 236, and 238 represent the spatial neighbors of PU0 222. For each group, potential candidates in neighboring blocks that reference the same reference image as the reference image indicated by the reference index signaled have the highest priority to be selected to form the final candidates for that group. All neighboring blocks may not contain motion vectors pointing to the same reference image. Therefore, if no such candidate can be found, the first available candidate is scaled to form the final candidates, thereby compensating for temporal distance differences.

[0139] Figure 3A and Figure 3B Includes a conceptual diagram illustrating the prediction of time motion vectors. Figure 3A An example CU 300 including PU0 302 and PU1 304 is shown. PU0 302 includes a central block 310 for PU0 302 and a lower right block 306 for PU0 302. Figure 3A An external block 308, which can predict motion information for PU0 302 from motion information, is also shown, as discussed below. Figure 3B The current image 342 shows the current block 326, which includes motion information to be predicted for it. Figure 3B Also shown are the co-located image 330 of the current image 342 (including the co-located block 324 of the current block 326), the current reference image 340, and the co-located reference image 332. The co-located block 324 is predicted using the co-located motion vector 320, which serves as a temporal motion vector predictor (TMVP) candidate 322 for motion information of block 326.

[0140] The video decoder can add Temporal Motion Vector Predictor (TMVP) candidates (e.g., TMVP candidate 322) (if enabled and available) to the MV candidate list, following any spatial motion vector candidates. The motion vector derivation process for TMVP candidates is the same for both merge and AMVP modes. However, in some cases, the target reference index for TMVP candidates is always set to zero in merge mode.

[0141] like Figure 3A As shown, the primary block position for TMVP candidate derivation is block 306 to the lower right outside co-located PU 304, to compensate for the bias of the upper and left blocks used to generate spatially adjacent candidates. However, if block 306 is located outside the current CTB (or LCU) row (e.g., as...), Figure 3A (as shown in block 308) or if the motion information for block 306 is unavailable, then the block is replaced by the center block 310 of PU 302.

[0142] refer to Figure 3B The motion vectors used for TMVP candidate 322 can be derived from the co-op block 324 of co-op image 330, indicated at the slice level. Similar to the temporal direct mode in AVC, the motion vectors of the TMVP candidate can undergo motion vector scaling, which is performed to compensate for the distance differences between the current image 342 and the current reference image 340, and between the co-op image 330 and the co-op reference image 332. That is, motion vector 320 can be scaled to produce TMVP candidate 322 based on the distance differences between the current image (e.g., current image 342) and the current reference image (e.g., current reference image 340), and between the co-op image (e.g., co-op image 330) and the co-op reference image (e.g., co-op reference image 332).

[0143] Other aspects of motion prediction are covered in the HEVC standard and / or other standards, formats, or codecs. For example, several other aspects of merge mode and AMVP mode are covered. One aspect includes motion vector scaling. Regarding motion vector scaling, it is assumed that the value of a motion vector is proportional to the distance between the images at rendering time. Motion vectors associate two images (a reference image and an image containing the motion vector, i.e., the containing image). When a motion vector is used to predict other motion vectors, the distance between the containing image and the reference image is calculated based on the Image Order Count (POC) value.

[0144] For a motion vector to be predicted, its associated containing image and reference image can be different. Therefore, a new distance is calculated (based on the Proof-of-Concept). Furthermore, the motion vector is scaled based on these two POC distances. For spatially adjacent candidates, the containing image used for two motion vectors is the same, while the reference image is different. In HEVC, motion vector scaling applies to both TMVP and AMVP for spatially and temporally adjacent candidates.

[0145] Another aspect of motion prediction involves artificial motion vector candidate generation. For example, if the list of motion vector candidates is incomplete, artificial motion vector candidates are generated and inserted at the end of the list until all candidates are obtained. In merge mode, there are two types of artificial motion vector candidates: combined candidates derived only for B-slices; and zero candidates used only for AMVP (if the first type does not provide enough artificial candidates). For each pair of candidates that is already in the candidate list and has the necessary motion information, a bidirectional combined motion vector candidate is derived by referring to a combination of the motion vector of the first candidate in the image in list 0 and the motion vector of the second candidate in the image in list 1.

[0146] In some implementations, a pruning process can be performed when new candidates are added to or inserted into the MV candidate list. For example, in some cases, MV candidates from different blocks may include the same information. In such cases, storing duplicate motion information from multiple MV candidates in the MV candidate list can lead to redundancy and reduced efficiency. In some examples, the pruning process can eliminate or minimize redundancy in the MV candidate list. For example, the pruning process may include comparing potential MV candidates to be added to the MV candidate list with MV candidates already stored in the MV candidate list. In an illustrative example, the horizontal displacement of the stored motion vectors (…) can be used as a comparison. ) and vertical displacement ( (Indicates the position of the reference block relative to the current block) and the horizontal displacement of the potential candidate motion vector ( ) and vertical displacement ( The comparison is performed. If the comparison reveals that the potential candidate's motion vector does not match any of the one or more stored motion vectors, the potential candidate is not considered a candidate to be pruned and can be added to the MV candidate list. If a match is found based on the comparison, the potential MV candidate is not added to the MV candidate list, thus avoiding the insertion of the same candidate. In some cases, to reduce complexity, only a limited number of comparisons are performed during the pruning process, instead of comparing each potential MV candidate with all existing candidates.

[0147] In some decoding schemes such as HEVC, weighted prediction (WP) is supported. In this case, a scaling factor (denoted by a), a shift (denoted by s), and an offset (denoted by b) are used in motion compensation. Assuming the pixel value at position (x, y) in the reference image is p(x, y), then p'(x, y) = ((a*p(x, y) + (1 << (s-1))) >> s) + b replaces p(x, y) as the predicted value in motion compensation.

[0148] When WP is enabled, for each reference image in the current slice, a signaling flag is used to indicate whether WP is applicable to the reference image. If WP is applicable to a reference image, a set of WP parameters (i.e., a, s, and b) is sent to the decoder, and this set of WP parameters is used for motion compensation based on the reference image. In some examples, to flexibly turn WP on / off for the luma and chroma components, WP flags and WP parameters for the luma and chroma components are signaled separately. In WP, the same set of WP parameters can be used for all pixels in a reference image.

[0149] Figure 4A This is a diagram illustrating an example of neighboring reconstructed samples of the current block 402 and neighboring samples of the reference block 404 used for unidirectional inter-frame prediction. The motion vector MV 410 can be decoded for the current block 402, where MV 410 may include a reference index for a list of reference images and / or other motion information for identifying the reference block 404. For example, MV may include horizontal and vertical components, providing coordinate offsets from the coordinate position in the current image to the coordinates in the reference image identified by the reference index. Figure 4B This is a diagram illustrating an example of neighboring reconstructed samples of the current block 422 and neighboring samples of the first reference block 424 and the second reference block 426 used for bidirectional inter-frame prediction. In this case, two motion vectors MV0 and MV1 can be decoded for the current block 422 to identify the first reference block 424 and the second reference block 426, respectively.

[0150] As explained earlier, OBMC is an example motion compensation technique that can be implemented for motion compensation. OBMC can improve prediction accuracy and avoid block artifacts. In OBMC, a prediction can be or include a weighted sum of multiple predictions. In some cases, blocks may be large in each dimension and may overlap with adjacent blocks in quadrants. Therefore, each pixel may belong to multiple blocks. For example, in some illustrative cases, each pixel may belong to four blocks. In such a scheme, OBMC can implement four predictions for each pixel, which are summed into a weighted average.

[0151] In some cases, specific syntax can be used at the CU level to enable and disable OBMC. In some examples, two orientation modes exist within the OBMC (e.g., top, left, right, bottom, or below), including the CU boundary OBMC mode and the sub-block boundary OBMC mode. When using the CU boundary OBMC mode, the original prediction block of the current CU MV is mixed with another prediction block (e.g., an "OBMC block") from the adjacent CU MV. In some examples, the top-left sub-block in the CU (e.g., the first or leftmost sub-block in the first / top row of the CU) has both top and left OBMC blocks, while other topmost sub-blocks (e.g., other sub-blocks in the first / top row of the CU) may only have top OBMC blocks. Other leftmost sub-blocks (e.g., the sub-block to the left of the CU in the first column of the CU) may only have left OBMC blocks.

[0152] When sub-CU decoding tools (e.g., affine motion compensation prediction, advanced temporal motion vector prediction (ATMVP), etc.) are enabled in the current CU, a sub-block boundary OBMC mode can be enabled, which allows the use of different prediction vectors (MVs) on a sub-block basis. In the sub-block boundary OBMC mode, individual OBMC blocks using the MVs of connected adjacent sub-blocks can be blended with the original prediction blocks using the MVs of the current sub-block. In some examples, in the sub-block boundary OBMC mode, individual OBMC blocks using the MVs of connected adjacent sub-blocks can be blended with the original prediction blocks using the MVs of the current sub-block in parallel, as further described herein. In other examples, in the sub-block boundary mode, individual OBMC blocks using the MVs of connected adjacent sub-blocks can be blended sequentially with the original prediction blocks using the MVs of the current sub-block. In some cases, the CU boundary OBMC mode can be executed before the sub-block boundary OBMC mode, and the predefined blending order for the sub-block boundary OBMC mode can include top, left, bottom, and right.

[0153] The prediction of the MV based on the neighboring sub-blocks N (e.g., the sub-blocks above, to the left, below, and to the right of the current sub-block) can be represented as P. N And the prediction based on the MV of the current sub-block can be represented as P C When sub-block N contains the same motion information as the current sub-block, the original prediction block may not be mixed with the prediction block based on the MV of sub-block N. In some cases, P can be... N The 4 rows / columns of the sample and P C The same samples in the same dataset are mixed. In some examples, weighting factors of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 can be used for P. N Furthermore, the corresponding weighting factors 3 / 4, 7 / 8, 15 / 16, and 31 / 32 can be used for P. CIn some cases, if the height or width of the decoded block is equal to 4, or if the CU is decoding using a sub-CU mode, then only P is allowed. N The 2 rows / columns are used for OBMC mixing.

[0154] Figure 5 This is a diagram illustrating an example of OBMC blending for CU boundary OBMC modes. (See diagram for example.) Figure 5 As shown, when using the CU boundary OBMC mode, the original prediction block of the current CU motion vector (MV) will be used (in Figure 5 The original block is referred to as "the original block" and another predicted block using the adjacent CU MV (in Figure 5 The top left sub-block of CU 530 may have top and left OBMC blocks that can be used to generate the blended block as described herein. Other topmost sub-blocks of CU 530 only have top OBMC blocks that can be used to generate the blended block as described herein. For example, sub-block 502 located at the top of CU 530 only has a top OBMC block, in Figure 5 The OBMC sub-block 504 is shown in the diagram. OBMC sub-block 504 can be a sub-block of the top adjacent CU, and it can include one or more sub-blocks. The other leftmost sub-blocks of CU 530 only have left-side OBMC blocks that can be used to generate the hybrid block as described herein. For example, sub-block 506 of CU 530 only has left-side OBMC blocks, in... Figure 5 The OBMC subblock 508 is shown in the diagram. The OBMC subblock 508 can be a subblock of the left adjacent CU, and it can include one or more subblocks.

[0155] exist Figure 5In the example shown, sub-block 502 and OBMC sub-block 504 can be used to generate hybrid block 515. For example, the MV of sub-block 502 can be used to predict the sample of CU 530 at the location of sub-block 502, and then multiplied by weight factor 510 to generate a first prediction result for sub-block 502. Similarly, the MV of OBMC sub-block 504 can be used to predict the sample of CU 530 at the location of sub-block 502, and then multiplied by weight factor 512 to generate a second prediction result for sub-block 502. The first prediction result generated for sub-block 502 can be added to the second prediction result generated for sub-block 502 to derive hybrid block 515. Weight factor 510 can be the same as or different from weight factor 512. In some examples, weight factor 510 can be different from weight factor 512. In some cases, weight factor 510 may depend on the distance from the mixed image data and / or samples from sub-block 502 to the CU boundary (e.g., to the boundary of CU 530), and weight factor 512 may depend on the distance from the mixed image data and / or samples from sub-block 502 to the CU boundary (e.g., to the boundary of CU 530). Weight factors 510 and 512 can add up to 1.

[0156] Sub-block 506 and OBMC sub-block 508 can be used to generate hybrid block 520. For example, the MV of sub-block 506 can be used to predict the sample of CU 530 at the location of sub-block 506, and then multiplied by weight factor 516 to generate a first prediction result for sub-block 506. Similarly, the MV of OBMC sub-block 508 can be used to predict the sample of CU 530 at the location of sub-block 506, and then multiplied by weight factor 518 to generate a second prediction result for sub-block 506. The first prediction result generated for sub-block 506 can be added to the second prediction result generated for sub-block 506 to derive hybrid block 520. Weight factor 516 can be the same as or different from weight factor 518. In some examples, weight factor 516 can be different from weight factor 518. In some cases, weight factor 516 may depend on the distance of the mixed image data and / or samples from sub-block 506 to the CU boundary (e.g., to the boundary of CU 530), and weight factor 518 may depend on the distance of the mixed image data and / or samples from sub-block 506 to the CU boundary (e.g., to the boundary of CU 530).

[0157] Figure 6 This is a diagram illustrating an example of OBMC blending for sub-block boundary OBMC modes. In some examples, sub-block boundary OBMC modes can be enabled when a sub-CU decoding tool (e.g., affine mode or tool, Advanced Time Motion Vector Prediction (ATMVP) mode or tool, etc.) is enabled for the current CU. Figure 6 As shown, four separate OBMC blocks using the MVs of four connected adjacent sub-blocks are blended with the original prediction block using the MV of the current sub-block. In other words, in addition to the original prediction using the MV of the current sub-block, four predictions for the sample of the current sub-block 602 are generated using the MVs from four separate OBMC blocks, and then combined with the original prediction to form a blended block 625. For example, sub-block 602 of CU 630 can be blended with adjacent OBMC blocks 604 to 610. In some cases, sub-block 602 can be blended with OBMC blocks 604 to 610 according to the blending order used for the sub-block boundary OBMC patterns. In some examples, the blending order may include the top OBMC block (e.g., OBMC block 604), the left OBMC block (e.g., OBMC block 606), the bottom OBMC block (e.g., OBMC block 608), and finally the right OBMC block (e.g., OBMC block 610). In some cases, sub-block 602 can be mixed with OBMC blocks 604 to 610 in parallel, as further described herein.

[0158] exist Figure 6 In the example shown, sub-block 602 can be blended with each OBMC block 620 according to formula 622. Formula 622 can be executed once for each of OBMC blocks 604 to 610, and the corresponding results can be summed to generate blended block 625. For example, OBMC block 620 in formula 622 can represent the OBMC blocks from OBMC blocks 604 to 610 used in formula 622. In some examples, the weighting factor 612 can depend on the location of the image data and / or samples within the sub-block 602 being blended. In some examples, the weighting factor 612 can depend on the distance of the image data and / or samples from the corresponding OBMC block being blended (e.g., OBMC block 604, OBMC block 606, OBMC block 608, OBMC block 610).

[0159] For example, when the prediction of the MV using OBMC block 604 is mixed with the prediction of the MV using sub-block 602 according to formula 622, OBMC block 620 can represent OBMC block 604. Here, the original prediction of sub-block 602 can be multiplied by weighting factor 612, and the result can be added to the result of multiplying the prediction of the MV using OBMC block 604 by weighting factor 614. When the prediction of the MV using OBMC block 606 is mixed with the prediction of the MV using sub-block 602 according to formula 622, OBMC block 620 can also represent OBMC block 606. Here, the original prediction of sub-block 602 can be multiplied by weighting factor 612, and the result can be added to the result of multiplying the prediction of the MV using OBMC block 606 by weighting factor 614. When the prediction of the MV using OBMC block 608 is mixed with the prediction of the MV using sub-block 602 according to formula 622, OBMC block 620 can also represent OBMC block 608. The original prediction of sub-block 602 can be multiplied by weighting factor 612, and the result can be added to the result of multiplying the prediction of the MV using OBMC block 608 by weighting factor 614. Finally, when the prediction of the MV using OBMC block 610 is mixed with the prediction of the MV using sub-block 602 according to formula 622, OBMC block 620 can represent OBMC block 610. The original prediction of sub-block 602 can be multiplied by weighting factor 612, and the result can be added to the result of multiplying the prediction of the MV using OBMC block 610 by weighting factor 614. The results from formula 622 for each of OBMC blocks 604 to 610 can be added to derive the mixed block 625.

[0160] Parallel mixing according to Formula 622 can be friendly to parallel hardware computing designs, avoiding or limiting unequal weighting, inconsistencies, etc. For example, in JEM, the predefined sequential mixing order for sub-block boundary OBMC patterns is top, left, bottom, and right. This order may increase computational complexity, reduce performance, lead to unequal weighting, and / or cause inconsistencies. In some examples, this sequential order may cause problems because sequential computation is not friendly to parallel hardware designs. Furthermore, this sequential order may lead to unequal weighting. For example, during the mixing process, the OBMC blocks of adjacent sub-blocks in later sub-block mixing may contribute more to the final sample predictions than in earlier sub-block mixing.

[0161] On the other hand, the systems and techniques described in this paper can achieve, for example Figure 6The formula for parallel mixing shown blends the prediction of the current sub-block with four OBMC sub-blocks, and the weighting factor can be fixed without biasing towards any particular neighboring sub-block. For example, using the formula for implementing parallel mixing, the final prediction P could be P = w1 * P c + w2 *P top + w3 * P left + w4 * P below + w5 * P right , where P top It is a prediction based on the MV of the top adjacent sub-block, P left It is a prediction based on the MV of the left-adjacent sub-block, P below It is a prediction based on the MV of the adjacent sub-blocks below, P right The prediction is based on the MV of the right-hand adjacent sub-block, and w1, w2, w3, w4, and w5 are the corresponding weighting factors. In some cases, the weight w1 can be equal to 1 – w2 – w3 – w4 – w5. Because the prediction of the MV based on the adjacent sub-block N may add / include / introduce noise to the samples in the row / column farthest from sub-block N, the system and technique described in this paper can set the values ​​of each of the weights w2, w3, w4, and w5 to {a, b, c, 0} for the row / column of the sample closest to the adjacent sub-block N {first, second, third, fourth} of the current sub-block.

[0162] For example, the first element 'a' (e.g., the weighting factor 'a') can be used for the sample row or column closest to the corresponding neighboring sub-block N, and the last element '0' can be used for the sample row or column furthest from the corresponding neighboring sub-block N. Using the positions (0,0), (0,1), and (1,1) of the top-left sample relative to the current sub-block of size 4×4 samples as examples, the final prediction P(x, y) can be derived as follows:

[0163] P(0, 0) = w1 * P c (0, 0) + a * P top (0, 0) + a * P left (0, 0)

[0164] P(0, 1) = w1 * P c (0, 1) + b * P top (0, 1) + a * P left (0, 1) + c * P below (0, 1)

[0165] P(1, 1) = w1 * P c(1, 1) + b * P top (1, 1) + b * P left (1, 1) + c * P below (1, 1) + c * P right (1, 1)

[0166] The example sum of weighting factors from the adjacent OBMC sub-blocks used for the 4×4 current sub-block (e.g., w2 + w3 + w4 + w5) can be as follows: Figure 7 As shown in Table 700. In some cases, weighting factors can be shifted left to avoid division operations that may increase computational complexity / burden and / or cause inconsistent results. For example, {a', b', c', 0} can be set to {a << shift, b << shift, c << shift, 0}, where shift is a positive integer. In this example, weight w1 can be equal to (1 << shift) – a' – b' – c', and P can be equal to (w1 * P) c + w2 * P top + w3 * P left + w4* P below + w5 * P right + (1<<(shift-1))) >> shift. An illustrative example of setting {a', b', c', 0} is {15, 8, 3, 0}, where the values ​​are the result of shifting the original values ​​6 left, and w1 equals (1 << 6) – a – b – c. P = (w1 * P c + w2 * P top + w3 * P left + w4 * P below + w5 * P right + (1<<5)) >> 6.

[0167] In some respects, for the row / column of the sample closest to the neighboring sub-blocks N{first, second, third, fourth} respectively in the current sub-block, the values ​​of w2, w3, w4, and w5 can be set to {a, b, 0, 0}. Using the positions (0, 0), (0, 1), and (1, 1) of the top-left sample of the current sub-block with a size of 4×4 samples as examples, the final prediction P(x, y) can be derived as follows:

[0168] P(0, 0) = w1 * P c (0, 0) + a * P top (0, 0) + a * P left(0, 0)

[0169] P(0, 1) = w1 * P c (0, 1) + b * P top (0, 1) + a * P left (0, 1)

[0170] P(1, 1) = w1 * P c (1, 1) + b * P top (1, 1) + b * P left (1, 1)

[0171] Example sums of weighting factors from neighboring OBMC sub-blocks used for the 4×4 current sub-block (e.g., w2 + w3 + w4 + w5) in Figure 8 As shown in Table 800, in some examples, a weighting factor can be selected such that the sum of w2 + w3 + w4 + w5 at corner samples (e.g., samples at (0,0), (0,3), (3,0), and (3,3)) is greater than the sum of w2 + w3 + w4 + w5 at other boundary samples (e.g., samples at (0,1), (0,2), (1,0), (2,0), (3,1), (3,2), (1,3), and (2,3)), and / or the sum of w2 + w3 + w4 + w5 at boundary samples is greater than the value at intermediate samples (e.g., samples at (1,1), (1,2), (2,1), and (2,2)).

[0172] In some cases, motion compensation can be skipped during the OBMC process based on the similarity between the current sub-block's MV and the MVs of its spatial neighbors (e.g., top, left, bottom, and right). For example, before each invocation of motion compensation using motion information from a given neighboring block / sub-block, the MV of the neighboring block / sub-block can be compared with the current sub-block's MV based on one or more of the following conditions. One or more conditions may include, for example: a first condition that all prediction lists used by the neighboring blocks / sub-blocks (e.g., list L0 or list L1 in unidirectional prediction, or both L0 and L1 in bidirectional prediction) are also used for the prediction of the current sub-block; a second condition that the MVs of the neighboring blocks / sub-blocks and the current sub-block use the same reference picture; and / or a third condition that the absolute value of the horizontal MV difference between the neighboring MV and the current MV is not greater than a predefined MV difference threshold T and the absolute value of the vertical MV difference between the neighboring MV and the current MV is not greater than a predefined MV difference threshold T (if bidirectional prediction is used, both L0 and L1 MVs can be checked).

[0173] In some examples, if the first, second, and third conditions are met, motion compensation using a given neighboring block / sub-block is not performed, and the OBMC sub-block using the MV of a given neighboring block / sub-block N is disabled and not mixed with the original sub-block. In some cases, the CU boundary OBMC mode and the sub-block boundary OBMC mode can have different values ​​for the threshold T. If the mode is the CU boundary OBMC mode, T is set to T1; otherwise, T is set to T2, where T1 and T2 are greater than 0. In some cases, when conditions are met, the lossy algorithm for skipping neighboring blocks / sub-blocks can be applied only to the sub-block boundary OBMC mode. The CU boundary OBMC mode can alternatively apply a lossless algorithm that skips neighboring blocks / sub-blocks when one or more of the following conditions are met: a fourth condition that all prediction lists used by neighboring blocks / sub-blocks (e.g., L0 or L1 in one-way prediction, or both L0 and L1 in two-way prediction) are also used for the prediction of the current sub-block; a fifth condition that neighboring MVs and the current MV use the same reference picture; and a sixth condition that neighboring MVs and the current MV are the same (if two-way prediction is used, both L0 and L1 MVs can be checked).

[0174] In some cases, when conditions one, two, and three are met, the lossy algorithm for skipping adjacent blocks / sub-blocks is only applied to the CU boundary OBMC mode. In some cases, when conditions four, five, and six are met, the lossless algorithm for skipping adjacent blocks / sub-blocks can be applied to the sub-block boundary OBMC mode.

[0175] In some aspects, lossy fast algorithms can be implemented in CU boundary OBMC mode to save encoding and decoding time. For example, the first OBMC block and adjacent OBMC blocks can be merged into a larger OBMC block and generated together if one or more conditions are met. One or more conditions may include, for example, the following: a condition that all prediction lists used by the first adjacent block of the current CU (e.g., L0 or L1 in unidirectional prediction, or both L0 and L1 in bidirectional prediction) are also used for the prediction of the second adjacent block of the current CU (in the same direction as the first adjacent block); a condition that the MV of the first adjacent block and the MV of the second adjacent block use the same reference picture; and a condition that the absolute value of the horizontal MV difference between the MV of the first adjacent block and the MV of the second adjacent block is not greater than a predefined MV difference threshold T3 and the absolute value of the vertical MV difference between the MV of the first adjacent block and the MV of the second adjacent block is not greater than a predefined MV difference threshold T3 (if bidirectional prediction is used, both L0 and L1 MVs can be checked).

[0176] In some aspects, lossy fast algorithms can be implemented in subblock boundary OBMC mode to save encoding and decoding time. In some examples, SbTMVP mode and DMVR are performed on an 8×8 basis, and affine motion compensation is performed on a 4×4 basis. The systems and techniques described herein can implement subblock boundary OBMC mode on an 8×8 basis. In some cases, the systems and techniques described herein can perform a similarity check at each 8×8 subblock to determine whether the 8×8 subblock should be split into four 4×4 subblocks, and if so, perform OBMC on a 4×4 basis.

[0177] Figure 9 This is a diagram illustrating an example CU 910 with sub-blocks 902 to 908 in an 8×8 block. In some examples, for each 8×8 sub-block, the lossy fast algorithm in the sub-block boundary OBMC mode may include four 4×4 OBMC sub-blocks (e.g., OBMC sub-block 902 (P), OBMC sub-block 904 (Q), OBMC sub-block 906 (R), and OBMC sub-block 908 (S)). OBMC subblocks 902 through 908 may be enabled for OBMC mixing if at least one of the following conditions is not met: a first condition that the prediction lists used by subblocks 902(P), 904(Q), 906(R), and 908(S) are the same (e.g., L0 or L1 in unidirectional prediction, or both L0 and L1 in bidirectional prediction); a second condition that the MVs of subblocks 902(P), 904(Q), 906(R), and 908(S) use the same reference picture; and any two subblocks (e.g., 902(P) and 904(Q), 902(P) and 906(R), 902(P) and 908(S)) are not used. The absolute value of the horizontal MV difference between the MVs of 904(Q) and 906(R), 904(Q) and 908(S), and 906(R) and 908(S) is not greater than the predefined MV difference threshold T4. The absolute value of the vertical MV difference between the MVs of any two sub-blocks (e.g., 902(P) and 904(Q), 902(P) and 906(R), 902(P) and 908(S), 904(Q) and 906(R), 904(Q) and 908(S), and 906(R) and 908(S)) is not greater than the predefined MV difference threshold T4 (if bidirectional prediction is used, both L0 and L1 MVs can be checked).

[0178] If all the above conditions are met, the system and techniques described herein can perform 8×8 sub-block OBMC, wherein the 8×8 OBMC sub-blocks from the top, left, bottom, and right MVs are generated using OBMC blending for sub-block boundary OBMC modes. Otherwise, when at least one of the above conditions is not met, OBMC is performed on a 4×4 basis within the 8×8 sub-block, and each 4×4 sub-block in the 8×8 sub-block generates four OBMC sub-blocks from the top, left, bottom, and right MVs.

[0179] In some aspects, when the CU is decoding using merge mode, the OBMC flag is copied from the adjacent block in a manner similar to motion information copying in merge mode. Otherwise, when the CU is not decoding using merge mode, the OBMC flag can be signaled to the CU to indicate whether OBMC is applicable.

[0180] Figure 10 This is a flowchart illustrating an example procedure 1000 for performing OBMC. At block 1002, procedure 1000 may include: determining that OBMC mode is enabled for the current sub-block of the video data block. In some examples, OBMC mode may include sub-block boundary OBMC mode.

[0181] At box 1004, process 1000 may include: determining a first prediction associated with the current sub-block, a second prediction associated with a first OBMC block adjacent to the top border of the current sub-block, a third prediction associated with a second OBMC block adjacent to the left border of the current sub-block, a fourth prediction associated with a third OBMC block adjacent to the bottom border of the current sub-block, and a fifth prediction associated with a fourth OBMC block adjacent to the right border of the current sub-block.

[0182] At box 1006, process 1000 may include: determining a sixth prediction based on the results of applying a first weight to a first prediction, applying a second weight to a second prediction, applying a third weight to a third prediction, applying a fourth weight to a fourth prediction, and applying a fifth weight to a fifth prediction. In some cases, the sum of the weights of the corner samples of a corresponding sub-block (e.g., the current sub-block, the first OBMC block, the second OBMC block, the third OBMC block, and the fourth OBMC block) may be greater than the sum of the weights of other boundary samples of the corresponding sub-block. In some cases, the sum of the weights of other boundary samples may be greater than the sum of the weights of non-boundary samples of the corresponding sub-block (e.g., samples not adjacent to the boundary of the sub-block).

[0183] For example, in some cases, each of the first, second, third, and fourth weights may include one or more weight values ​​associated with one or more samples from the corresponding sub-block of the current sub-block, the first OBMC block, the second OBMC block, the third OBMC block, or the fourth OBMC block. Furthermore, the sum of the weight values ​​of the corner samples of the corresponding sub-block may be greater than the sum of the weight values ​​of the other boundary samples of the corresponding sub-block, and the sum of the weight values ​​of the other boundary samples of the corresponding sub-block may be greater than the sum of the weight values ​​of the non-boundary samples of the corresponding sub-block.

[0184] At box 1008, process 1000 may include: generating a hybrid sub-block corresponding to the current sub-block of the video data block based on a sixth prediction.

[0185] Figure 11 This is a flowchart illustrating another example procedure 1100 for performing OBMC. At block 1102, procedure 1100 may include: determining that OBMC mode is enabled for the current sub-block of the video data block. In some examples, OBMC mode may include sub-block boundary OBMC mode.

[0186] At box 1104, process 1100 may include determining whether a first condition, a second condition, and a third condition are satisfied for at least one adjacent sub-block adjacent to the current sub-block. In some examples, the first condition may include a list of all reference images in one or more reference image lists used to predict adjacent sub-blocks for predicting the current sub-block.

[0187] In some examples, the second condition may include one or more of the same reference images used to determine motion vectors associated with the current sub-block and neighboring sub-blocks.

[0188] In some examples, the third condition may include a first difference between the horizontal motion vectors of the current sub-block and its neighboring sub-blocks, and a second difference between the vertical motion vectors of the current sub-block and its neighboring sub-blocks, not exceeding a motion vector difference threshold. In some examples, the motion vector difference threshold is greater than zero.

[0189] At box 1106, process 1100 may include: determining not to use motion information of adjacent sub-blocks for motion compensation of the current sub-block based on determining that OBMC mode is enabled for the current sub-block and determining that the first condition, the second condition, and the third condition are met.

[0190] In some aspects, process 1100 may include: determining to perform a sub-block boundary OBMC mode for the current sub-block based on determining whether to use a decoder-side motion vector refinement (DMVR) mode, a sub-block-based temporal motion vector prediction (SbTMVP) mode, or an affine motion compensation prediction mode for the current sub-block.

[0191] In some aspects, process 1100 may include: performing a sub-block boundary OBMC pattern for a sub-block. In some cases, performing a sub-block boundary OBMC pattern for a sub-block may include determining a first prediction associated with the current sub-block, a second prediction associated with a first OBMC block adjacent to the top border of the current sub-block, a third prediction associated with a second OBMC block adjacent to the left border of the current sub-block, a fourth prediction associated with a third OBMC block adjacent to the bottom border of the current sub-block, and a fifth prediction associated with a fourth OBMC block adjacent to the right border of the current sub-block; determining a sixth prediction based on the result of applying a first weight to the first prediction, applying a second weight to the second prediction, applying a third weight to the third prediction, applying a fourth weight to the fourth prediction, and applying a fifth weight to the fifth prediction; and generating a hybrid sub-block corresponding to the current sub-block based on the sixth prediction.

[0192] In some cases, the sum of the weights of the corner samples of a corresponding sub-block (e.g., the current sub-block, the first OBMC block, the second OBMC block, the third OBMC block, and the fourth OBMC block) can be greater than the sum of the weights of the other boundary samples of the corresponding sub-block. In some cases, the sum of the weights of the other boundary samples can be greater than the sum of the weights of the non-boundary samples of the corresponding sub-block (e.g., samples that are not adjacent to the boundary of the current sub-block).

[0193] For example, in some cases, each of the second, third, fourth, and fifth weights may include one or more weight values ​​associated with one or more samples from the corresponding sub-block of the current sub-block. Furthermore, the sum of the weight values ​​of the corner samples of the current sub-block may be greater than the sum of the weight values ​​of the other boundary samples of the current sub-block, and the sum of the weight values ​​of the other boundary samples of the current sub-block may be greater than the sum of the weight values ​​of the non-boundary samples of the current sub-block.

[0194] In some aspects, process 1100 may include: determining that a Local Illumination Compensation (LIC) mode is used for an additional block of video data; and, based on the determination that the LIC mode is used for the additional block, skipping signaling of information associated with the OBMC mode used for the additional block. In some examples, signaling of skipping information associated with the OBMC mode used for the additional block may include signaling a syntax flag associated with the OBMC mode with a null value (e.g., excluding the value used for the flag). In some aspects, process 1100 may include: receiving a signal including a syntax flag with a null value associated with the OBMC mode used for the additional block of video data. In some aspects, process 1100 may include: determining that the OBMC mode is not used for the additional block based on the syntax flag with a null value.

[0195] In some cases, signaling that skips information associated with the OBMC mode used for the extra block may include determining that the LIC mode is used for the extra block, determining that the OBMC mode is not used or enabled for the extra block, and skipping the signal notification associated with the OBMC mode used for the extra block.

[0196] In some aspects, process 1100 may include: determining whether OBMC mode is enabled for the additional block, and based on determining whether OBMC mode is enabled for the additional block and determining that LIC mode is used for the additional block, determining to skip signaling information associated with the OBMC mode used for the additional block.

[0197] In some aspects, process 1100 may include: determining the current sub-block of the video data block using the Decoding Unit (CU) Boundary OBMC mode; and determining a final prediction for the current sub-block based on the sum of a first result of applying weights associated with the current sub-block to the corresponding prediction associated with the current sub-block and a second result of applying one or more corresponding weights to one or more corresponding predictions associated with one or more sub-blocks adjacent to the current sub-block.

[0198] In some examples, determining not to use motion information from adjacent sub-blocks for motion compensation of the current sub-block may include skipping the use of motion information from adjacent sub-blocks for motion compensation of the current sub-block.

[0199] In some cases, process 1000 and / or process 1100 can be implemented by an encoder and / or decoder.

[0200] In some implementations, the processes (or methods) described herein (including processes 1000 and 1100) can be computed by a computing device or apparatus (such as in...) Figure 1 The process can be executed by system 100 as shown. For example, the process can be performed by... Figure 1 and Figure 12 The encoding device 104 shown, another video source-side device or video transmission device, in Figure 1 and Figure 13 The decoding device 112 shown, and / or another client-side device (such as a player device, display, or any other client-side device) may be used to perform the steps of process 1000 and / or process 1100. In some cases, the computing device or apparatus may include one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to perform the steps of process 1000 and / or process 1100.

[0201] In some examples, a computing device may include a mobile device, a desktop computer, a server computer and / or a server system, or other types of computing devices. Components of the computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented in circuitry. For example, components may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. In some examples, a computing device or apparatus may include a camera configured to capture video data (e.g., video sequences) comprising video frames. In some examples, the camera or other capturing device that captures the video data is separate from the computing device, in which case the computing device receives or acquires the captured video data. The computing device may include a network interface configured to transmit video data. The network interface may be configured to transmit Internet Protocol (IP) based data or other types of data. In some examples, the computing device or apparatus may include a display for showing output video content, such as a sample of a video bitstream image.

[0202] The above process can be described using a logic flowchart, where operations represent a series of operations that can be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents a computer-executable instruction stored on one or more computer-readable storage media that performs the described operation when executed by one or more processors. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement these processes.

[0203] Furthermore, the above processes can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, implemented by hardware or a combination thereof. As mentioned above, the code can be stored, for example, in the form of a computer program comprising multiple instructions executable by one or more processors on a computer-readable or machine-readable storage medium. The computer-readable or machine-readable storage medium can be non-transitory.

[0204] The decoding techniques discussed in this paper can be implemented in an example video encoding and decoding system (e.g., System 100). In some examples, the system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source and destination devices can include any of a variety of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, mobile phones (such as so-called "smartphones"), so-called "smart tablets," televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source and destination devices can be equipped for wireless communication.

[0205] A destination device can receive encoded video data to be decoded via a computer-readable medium. The computer-readable medium can be any type of medium or device capable of moving encoded video data from a source device to a destination device. In one example, the computer-readable medium may include a communication medium enabling the source device to transmit encoded video data directly to the destination device in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network (LAN), a wide area network (WAN), or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device to the destination device.

[0206] In some examples, encoded data can be output from an output interface to a storage device. Similarly, encoded data can be accessed from a storage device via an input interface. The storage device can include any of a variety of distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In other examples, the storage device can correspond to a file server or another intermediate storage device that can store encoded video generated by the source device. The destination device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing and sending encoded video data to the destination device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device can access the encoded video data via any standard data connection, including an Internet connection. This can include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., DSL, cable modems, etc.), or combinations thereof, suitable for accessing encoded video data stored on a file server. Transfer of encoded video data from the storage device can be streaming, downloading, or a combination thereof.

[0207] The techniques disclosed herein are not necessarily limited to wireless applications or setups. The techniques described above can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (e.g., HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding digital video stored on data storage media, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0208] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the techniques disclosed herein. In other examples, the source and destination devices may include other components or arrangements. For example, the source device may receive video data from an external video source such as an external camera. Similarly, the destination device may interface with an external display device, rather than including an integrated display device.

[0209] The example system described above is merely an example. Techniques for processing video data in parallel can be implemented by any digital video encoding and / or decoding device. Although, in general, the techniques disclosed herein are implemented by video encoding devices, they can also be implemented by video encoders / decoders, commonly referred to as "CODECs." Furthermore, the techniques disclosed herein can also be implemented by video preprocessors. The source and destination devices are merely examples of such decoding devices, where the source device generates decoded video data for transmission to the destination device. In some examples, the source and destination devices can operate in a substantially symmetrical manner, such that each of these devices includes both video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0210] Video sources may include video capture devices, such as cameras, video storage units containing previously captured video, and / or video feed interfaces for receiving video from video content providers. Alternatively, a video source may generate computer graphics-based data as source video, or a combination of live video, stored video, and computer-generated video. In some cases, if the video source is a camera, the source and destination devices may form what is known as a mobile phone camera or videophone. However, as described above, the techniques described in this disclosure are generally applicable to video decoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by a video encoder. The encoded video data can then be output to a computer-readable medium via an output interface.

[0211] As mentioned, computer-readable media can include temporary media such as those transmitted via wireless broadcasting or wired networks, or storage media such as hard disks, flash drives, compressed optical discs, digital versatile optical discs, Blu-ray discs (i.e., non-temporary storage media), or other computer-readable media. In some examples, a network server (not shown) can, for example, receive encoded video data from a source device via a network transmission and provide the encoded video data to a destination device. Similarly, a computing device in a media production facility, such as an optical disc stamping facility, can receive encoded video data from a source device and manufacture an optical disc containing the encoded video data. Therefore, in the various examples, computer-readable media can be understood to include one or more computer-readable media of various forms.

[0212] The input interface of the destination device receives information from a computer-readable medium. The information in the computer-readable medium may include syntax information defined by the video encoder (which is also used by the video decoder), including characteristics of description blocks and other decoding units (e.g., groups of pictures (GOPs)) and / or processed syntax elements. The display device displays the decoded video data to the user and may include any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device. Various embodiments of this application have been described.

[0213] exist Figure 12 and Figure 13 The specific details of the encoding device 104 and the decoding device 112 are shown in the figure. Figure 12 This is a block diagram illustrating an example encoding device 104 that can implement one or more of the techniques described in this disclosure. Encoding device 104 can, for example, generate the syntax structures described herein (e.g., syntax structures of VPS, SPS, PPS, or other syntax elements). Encoding device 104 can perform intra-frame prediction and inter-frame prediction decoding of video blocks within a video slice. As previously described, intra-frame decoding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-frame decoding relies at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. Intra-frame mode (I-mode) can refer to any of several spatial-based compression modes. Inter-frame modes such as one-way prediction (P-mode) or two-way prediction (B-mode) can refer to any of several temporal-based compression modes.

[0214] Encoding device 104 includes a segmentation unit 35, a prediction processing unit 41, a filter unit 63, an image memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-frame prediction processing unit 46. For video block reconstruction, encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. Filter unit 63 is intended to represent one or more loop filters, such as deblocking filters, adaptive loop filters (ALF), and sample adaptive offset (SAO) filters. Although in Figure 12The filter unit 63 is shown as an in-loop filter, but in other configurations, the filter unit 63 may be implemented as a post-loop filter. The post-processing device 57 may perform additional processing on the encoded video data generated by the encoding device 104. In some cases, the techniques of this disclosure may be implemented by the encoding device 104. However, in other cases, one or more techniques of this disclosure may be implemented by the post-processing device 57.

[0215] like Figure 12 As shown, encoding device 104 receives video data, and segmentation unit 35 segments the data into video blocks. This segmentation may also include, for example, segmentation into slices, segments, tiles, or other larger units based on the quadtree structure of LCUs and CUs, as well as video block segmentation. Encoding device 104 generally illustrates components for encoding video blocks within a video slice to be encoded. The slice may be divided into multiple video blocks (and possibly into a set of video blocks referred to as tiles). Prediction processing unit 41 may select one of several possible decoding modes for the current video block based on error results (e.g., decoding rate and distortion level), such as one or more of various intra-frame prediction decoding modes or inter-frame prediction decoding modes. Prediction processing unit 41 may provide the obtained intra-frame or inter-frame decoded blocks to summer 50 to generate residual block data, and to summer 62 to reconstruct the encoded blocks for use as reference pictures.

[0216] The intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-prediction decoding of the current video block relative to one or more adjacent blocks in the same frame or slice as the current video block to be decoded, to provide spatial compression. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction decoding of the current video block relative to one or more prediction blocks in one or more reference pictures, to provide temporal compression.

[0217] Motion estimation unit 42 can be configured to determine an inter-frame prediction mode for video slices based on a predetermined pattern for the video sequence. The predetermined pattern can designate video slices in the sequence as P-slices, B-slices, or GPB-slices. Motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are shown separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of the video slices. Motion vectors can, for example, indicate the displacement of a prediction unit (PU) of a video slice within the current video frame or picture relative to a prediction slice within a reference picture.

[0218] A predicted block is a block found to closely match the PU of the video block to be decoded in terms of pixel difference, which can be determined by the sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some examples, encoding device 104 can compute values ​​for pixel positions below integers for a reference image stored in image memory 64. For example, encoding device 104 can interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional-pixel positions and output motion vectors with fractional-pixel precision.

[0219] The motion estimation unit 42 calculates the motion vector for the PU by comparing the position of the PU in the video block within the inter-frame decoded slice with the position of the predicted block in the reference image. Reference images can be selected from either a first list of reference images (list 0) or a second list of reference images (list 1), each of which has an identifier stored in one or more reference images in the image memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.

[0220] Motion compensation performed by motion compensation unit 44 may involve obtaining or generating prediction blocks based on motion vectors determined by motion estimation, possibly performing interpolation with sub-pixel precision. Upon receiving the motion vector of the PU for the current video block, motion compensation unit 44 can locate the prediction block pointed to by the motion vector in a list of reference images. Encoding device 104 forms a residual video block by subtracting the pixel values ​​of the prediction block from the pixel values ​​of the current video block being decoded. The pixel difference forms residual data for that block and may include both luminance difference components and chrominance difference components. Summer 50 represents one or more components performing this subtraction operation. Motion compensation unit 44 may also generate syntax elements associated with video blocks and video slices for use by decoding device 112 when decoding video blocks of video slices.

[0221] As described above, the intra-prediction processing unit 46 can perform intra-prediction on the current block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44. Specifically, the intra-prediction processing unit 46 can determine the intra-prediction mode to be used for encoding the current block. In some examples, the intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during a separate encoding process, and the intra-prediction processing unit 46 can select a suitable intra-prediction mode from the tested modes. For example, the intra-prediction processing unit 46 can use rate-distortion analysis for various tested intra-prediction modes to calculate rate-distortion values, and can select the intra-prediction mode with the best rate-distortion characteristics from the tested modes. Rate-distortion analysis typically determines the amount of distortion (or error) between the encoded block and the original uncoded block encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. Intra-prediction processing unit 46 can calculate the ratio based on the distortion and rate for various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for that block.

[0222] In any case, after selecting an intra-prediction mode for a block, the intra-prediction processing unit 46 can provide information indicating the selected intra-prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode. The encoding device 104 can include in the transmitted bitstream configuration data definitions for various blocks, as well as indications of the most probable intra-prediction mode to be used for each context, an intra-prediction mode index table, and a modified intra-prediction mode index table. The bitstream configuration data may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also referred to as codeword maps).

[0223] After prediction processing unit 41 generates a prediction block for the current video block via inter-frame prediction or intra-frame prediction, encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to transform processing unit 52. Transform processing unit 52 uses a transform (such as Discrete Cosine Transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients. Transform processing unit 52 can transform the residual video data from the pixel domain to the transform domain (such as the frequency domain).

[0224] The transform processing unit 52 can send the obtained transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of these coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can then perform a scan of a matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform this scan.

[0225] After quantization, entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, entropy coding unit 56 can perform context-adaptive variable-length decoding (CAVLC), context-adaptive binary arithmetic decoding (CABAC), syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, or another entropy coding technique. After entropy coding by entropy coding unit 56, the encoded bitstream can be sent to decoding device 112, or stored for later transmission or retrieval by decoding device 112. Entropy coding unit 56 can also perform entropy coding on the motion vectors and other syntax elements of the current video slice being decoded.

[0226] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual blocks in the pixel domain for later use as reference blocks in reference images. Motion compensation unit 44 can calculate the reference block by adding the residual block to a predicted block of one of the reference images in the reference image list. Motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual block to calculate pixel values ​​below the integer level for motion estimation. Summer 62 adds the reconstructed residual block to the motion-compensated predicted block generated by motion compensation unit 44 to produce a reference block for storage in image memory 64. The reference block can be used by motion estimation unit 42 and motion compensation unit 44 for inter-frame prediction of blocks in subsequent video frames or images.

[0227] In this way, Figure 12 The encoding device 104 indicates that it is configured to perform any of the techniques described herein (including those mentioned above). Figure 10 The described process and / or the above regarding Figure 11 Examples of video encoders (described in the process). In some cases, some of the techniques in this disclosure can also be implemented by post-processing device 57.

[0228] Figure 13This is a block diagram illustrating an example decoding device 112. Decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and an image memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-frame prediction processing unit 84. In some examples, decoding device 112 can perform operations generally related to... Figure 12 The encoding stage described by the encoding device 104 is the opposite of the decoding stage.

[0229] During the decoding process, decoding device 112 receives an encoded video bitstream sent by encoding device 104, which represents video blocks of encoded video slices and associated syntax elements. In some embodiments, decoding device 112 may receive the encoded video bitstream from encoding device 104. In some embodiments, decoding device 112 may receive the encoded video bitstream from network entity 79 (such as a server, a media-aware network element (MANE), a video editor / cutter, or other such device configured to implement one or more of the techniques described above). Network entity 79 may or may not include encoding device 104. Before sending the encoded video bitstream to decoding device 112, network entity 79 may implement some of the techniques described in this disclosure. In some video decoding systems, network entity 79 and decoding device 112 may be part of a separate device, while in other cases, the functionality described with respect to network entity 79 may be performed by the same device including decoding device 112.

[0230] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or video block level. The entropy decoding unit 80 can process and parse both fixed-length and variable-length syntax elements in more parameter sets such as VPS, SPS, and PPS.

[0231] When a video slice is decoded into a slice that has undergone intra-frame decoding (I), the intra-frame prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for a video block for the current video slice based on the intra-frame prediction mode notified by a signal and data from previously decoded blocks from the current frame or picture. When a video frame is decoded into a slice that has undergone inter-frame decoding (i.e., B, P, or GPB), the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for a video block for the current video slice based on motion vectors received from the entropy decoding unit 80 and other syntax elements. A prediction block can be generated from one of the reference pictures in the reference picture list. The decoding device 112 can construct the reference frame list, i.e., list 0 and list 1, based on the reference pictures stored in the picture memory 92 using a default construction technique.

[0232] Motion compensation unit 82 determines prediction information for video blocks in the current video slice by parsing motion vectors and other syntax elements, and uses this prediction information to generate prediction blocks for the current video slice being decoded. For example, motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra-frame or inter-frame prediction), the inter-frame prediction slice type (e.g., B-slice, P-slice, or GPB-slice), construction information for one or more reference picture lists for the slice, motion vectors for each inter-frame encoded video block in the slice, inter-frame prediction state for each inter-frame decoded video block in the slice, and other information for decoding video blocks in the current video slice.

[0233] The motion compensation unit 82 can also perform interpolation based on an interpolation filter. The motion compensation unit 82 can use the interpolation filter used by the encoding device 104 during the encoding of the video block to calculate the interpolated values ​​for pixels less than an integer value for the reference block. In this case, the motion compensation unit 82 can determine the interpolation filter used by the encoding device 104 based on the received syntax elements, and can use the interpolation filter to generate a prediction block.

[0234] The inverse quantization unit 86 inverse-quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include determining the degree of quantization using quantization parameters calculated by the encoding device 104 for each video block in the video slice, and similarly determining the degree of inverse quantization to be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients to produce residual blocks in the pixel domain.

[0235] After the motion compensation unit 82 generates a prediction block for the current video block based on motion vectors and other syntax elements, the decoding device 112 forms a decoded video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82. The summer 90 represents one or more components performing this summation operation. Loop filters (in or after the decoding loop) can also be used, if needed, to smooth pixel transitions or otherwise improve video quality. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although in Figure 13 The filter unit 91 is shown as a filter in the loop, but in other configurations, filter unit 91 can be implemented as a post-loop filter. The decoded video block in a given frame or image is then stored in image memory 92, which stores reference images for subsequent motion compensation. Image memory 92 also stores the decoded video for later display on a display device (such as...). Figure 1 The video is displayed on the destination device 122 shown in the figure.

[0236] In this way, Figure 13 The decoding device 112 indicates that it is configured to perform any of the techniques described herein (including those mentioned above). Figure 10 The described process and the above about Figure 11 An example of a video decoder (describing the process).

[0237] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and excludes carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as compact optical discs (CDs) or digital laser discs (DVDs), flash memory, memory, or memory devices. Computer-readable media may have code and / or machine-executable instructions stored thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or sent via any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0238] In some embodiments, computer-readable storage devices, media, and memories may include cable or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0239] Specific details are provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that these embodiments can be practiced without these specific details. For clarity, in some instances, the techniques described herein may be presented as comprising individual functional blocks including devices, device components, steps or routines in a software-embodied method, or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure these embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring these embodiments.

[0240] The various embodiments described above may be presented as processes or methods, depicted as flowcharts, schematic diagrams, data flow diagrams, structural diagrams, or block diagrams. While a flowchart may describe operations as a sequential process, many of these operations may be performed in parallel or simultaneously. Furthermore, the order of operations may be rearranged. A process terminates upon completion of its operations, but may have additional steps not included in the diagram. A process may correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.

[0241] The processes and methods described in the examples above can be implemented using computer-executable instructions, which are stored in or otherwise made available from a computer-readable medium. Such instructions may include, for example, instructions or data that cause a general-purpose computer, special-purpose computer, or processing device to perform or otherwise configure it to perform a particular function or a particular set of functions. Parts of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include hard disks or optical disks, flash memory, USB devices equipped with non-volatile memory, network storage devices, etc.

[0242] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may employ any of a variety of form factors. When implemented using software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored on a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mount devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or plug-in cards. By further example, such functionality may also be implemented on a circuit board between different chips or different processes executed in a single device.

[0243] Instructions, media for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0244] In the foregoing description, aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it should be understood that the inventive concepts may be embodied and employed differently in other ways, and the appended claims are intended to be construed as including such variations, in addition to those limited by the prior art. Various features and aspects of the above-described applications may be used individually or collectively. Furthermore, embodiments may be used in any number of environments and applications other than those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, the methods have been described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in a different order than that described.

[0245] Those skilled in the art will understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“>”) respectively. ") and greater than or equal to (" Replace with the ") symbol.

[0246] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing a circuit or other hardware to perform the operation, designing a programmable circuit (e.g., a microprocessor or other suitable circuit) to perform the operation, or any combination thereof.

[0247] The phrase “coupled to” refers to any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0248] In this disclosure, the language of claims stating "at least one of..." and / or "one or more" in the set indicates that one or more members of the set (in any combination) satisfy the claim. For example, the language of claims stating "at least one of A and B" means A, B, or A and B. In another example, the language of claims stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language of "at least one of..." and / or "one or more" in the set does not limit the set to items listed in the set. For example, the language of claims stating "at least one of A and B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0249] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above regarding their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0250] The techniques described herein can also be implemented using electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, mobile phones with wireless communication devices, or integrated circuit devices with multiple uses (including applications in mobile phones with wireless communication devices and other devices). Any feature described as a module or component can be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code that, when executed, performs one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Alternatively, the technology may be implemented, at least in part, by a computer-readable communication medium (such as a propagating signal or wave) that carries or transmits program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer.

[0251] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).

[0252] Illustrative examples of this disclosure include:

[0253] Aspect 1. An apparatus for processing video data, comprising: a memory; and one or more processors coupled to the memory, the one or more processors being configured to: determine that an Overlapping Block Motion Compensation (OBMC) mode is enabled for a current sub-block of a video data block; for at least one adjacent sub-block adjacent to the current sub-block: determine whether a first condition, a second condition, and a third condition are satisfied, the first condition comprising all reference image lists in one or more reference image lists used for predicting the current sub-block for predicting the adjacent sub-block; the second condition comprising the same one or more reference images used for determining motion vectors associated with the current sub-block and the adjacent sub-block; and the third condition comprising a first difference between the horizontal motion vectors of the current sub-block and the adjacent sub-block and a second difference between the vertical motion vectors of the current sub-block and the adjacent sub-block not exceeding a motion vector difference threshold, wherein the motion vector difference threshold is greater than zero; and based on determining that an OBMC mode is enabled for the current sub-block and determining that the first condition, the second condition, and the third condition are satisfied, determine that motion information of the adjacent sub-blocks is not used for motion compensation of the current sub-block.

[0254] Aspect 2. The apparatus according to Aspect 1, wherein one or more processors are configured to: determine to perform a sub-block boundary OBMC mode for the current sub-block based on determining to use a decoder-side motion vector refinement (DMVR) mode, a sub-block-based temporal motion vector prediction (SbTMVP) mode, or an affine motion compensation prediction mode for the current sub-block.

[0255] Aspect 3. The apparatus according to Aspect 2, wherein, in order to perform a sub-block boundary OBMC mode for the current sub-block, one or more processors are configured to: determine a first prediction associated with the current sub-block, a second prediction associated with a first OBMC block adjacent to the top border of the current sub-block, a third prediction associated with a second OBMC block adjacent to the left border of the current sub-block, a fourth prediction associated with a third OBMC block adjacent to the bottom border of the current sub-block, and a fifth prediction associated with a fourth OBMC block adjacent to the right border of the current sub-block; determine a sixth prediction based on the result of applying a first weight to the first prediction, applying a second weight to the second prediction, applying a third weight to the third prediction, applying a fourth weight to the fourth prediction, and applying a fifth weight to the fifth prediction; and generate a hybrid sub-block corresponding to the current sub-block based on the sixth prediction.

[0256] Aspect 4. The apparatus according to aspect 3, wherein each of the second weight, the third weight, the fourth weight, and the fifth weight includes one or more weight values ​​associated with one or more samples from the corresponding sub-block of the current sub-block, wherein the sum of the weight values ​​of the corner samples of the current sub-block is greater than the sum of the weight values ​​of the other boundary samples of the current sub-block.

[0257] Aspect 5. The apparatus according to aspect 4, wherein the sum of the weight values ​​of the other boundary samples of the current sub-block is greater than the sum of the weight values ​​of the non-boundary samples of the current sub-block.

[0258] Aspect 6. The apparatus according to any one of Aspects 1 to 5, wherein one or more processors are configured to: determine that a Local Illumination Compensation (LIC) mode is used for an additional block of video data; and based on the determination that a LIC mode is used for the additional block, skip signaling of information associated with an OBMC mode for the additional block.

[0259] Aspect 7. The apparatus according to aspect 6, wherein, in order to skip signaling associated with information for an OBMC mode for an additional block, one or more processors are configured to signal a syntax flag having a null value, the syntax flag being associated with the OBMC mode.

[0260] Aspect 8. The apparatus according to any one of Aspects 6 to 7, wherein one or more processors are configured to: receive a signal including a syntax flag having a null value, the syntax flag being associated with an OBMC mode for additional blocks of video data.

[0261] Aspect 9. The apparatus according to any one of Aspects 7 to 8, wherein one or more processors are configured to: determine, based on a syntax flag having a null value, not to use the OBMC mode for an additional block.

[0262] Aspect 10. The apparatus according to any one of Aspects 6 to 9, wherein, in order to skip signaling associated with information for an OBMC mode for an additional block, one or more processors are configured to: determine whether to use or enable an OBMC mode for the additional block based on determining that an LIC mode is used for the additional block; and skip signaling associated with values ​​for an OBMC mode for the additional block.

[0263] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein one or more processors are configured to: determine whether to enable OBMC mode for an additional block; and based on the determination of whether to enable OBMC mode for an additional block and the determination of using LIC mode for an additional block, determine to skip signaling information associated with the OBMC mode for the additional block.

[0264] Aspect 12. The apparatus according to any one of Aspects 1 to 11, wherein one or more processors are configured to: determine a decoding unit (CU) boundary OBMC mode for a current sub-block of a video data block; and determine a final prediction for the current sub-block based on the sum of a first result of applying weights associated with the current sub-block to a corresponding prediction associated with the current sub-block and a second result of applying one or more corresponding weights to one or more corresponding predictions associated with one or more sub-blocks adjacent to the current sub-block.

[0265] Aspect 13. The apparatus according to any one of Aspects 1 to 12, wherein, in order to determine that motion information of adjacent sub-blocks should not be used for motion compensation of the current sub-block, one or more processors are configured to: skip using motion information of adjacent sub-blocks for motion compensation of the current sub-block.

[0266] Aspect 14. The apparatus according to any one of aspects 1 to 13, wherein the apparatus includes a decoder.

[0267] Aspect 15. The apparatus according to any one of aspects 1 to 14 further includes: a display configured to display one or more output images associated with video data.

[0268] Aspect 16. The apparatus according to any one of Aspects 1 to 15, wherein the OBMC mode includes a sub-block boundary OBMC mode.

[0269] Aspect 17. The apparatus according to any one of aspects 1 to 16, wherein the apparatus includes an encoder.

[0270] Aspect 18. The apparatus according to any one of aspects 1 to 17 further includes: a camera configured to capture images associated with video data.

[0271] Aspect 19. The apparatus according to any one of aspects 1 to 18, wherein the apparatus is a mobile device.

[0272] Aspect 20. A method for processing video data, comprising: determining that an Overlapping Block Motion Compensation (OBMC) mode is enabled for a current sub-block of a video data block; for at least one adjacent sub-block adjacent to the current sub-block: determining whether a first condition, a second condition, and a third condition are satisfied, the first condition comprising all reference image lists in one or more reference image lists used for predicting the current sub-block for predicting the adjacent sub-block; the second condition comprising the same one or more reference images used to determine motion vectors associated with the current sub-block and the adjacent sub-block; and the third condition comprising a first difference between the horizontal motion vectors of the current sub-block and the adjacent sub-block and a second difference between the vertical motion vectors of the current sub-block and the adjacent sub-block not exceeding a motion vector difference threshold, wherein the motion vector difference threshold is greater than zero; and based on determining that the OBMC mode is used for the current sub-block and determining that the first condition, the second condition, and the third condition are satisfied, determining not to use motion information of the adjacent sub-blocks for motion compensation of the current sub-block.

[0273] Aspect 21. The method according to aspect 20 further includes: determining to perform a sub-block boundary OBMC mode for the current sub-block based on determining to use a decoder-side motion vector refinement (DMVR) mode, a sub-block-based temporal motion vector prediction (SbTMVP) mode, or an affine motion compensation prediction mode for the current sub-block.

[0274] Aspect 22. The method according to aspect 21, wherein performing a sub-block boundary OBMC pattern for the current sub-block includes: determining a first prediction associated with the current sub-block, a second prediction associated with a first OBMC block adjacent to the top border of the current sub-block, a third prediction associated with a second OBMC block adjacent to the left border of the current sub-block, a fourth prediction associated with a third OBMC block adjacent to the bottom border of the current sub-block, and a fifth prediction associated with a fourth OBMC block adjacent to the right border of the current sub-block; determining a sixth prediction based on the result of applying a first weight to the first prediction, applying a second weight to the second prediction, applying a third weight to the third prediction, applying a fourth weight to the fourth prediction, and applying a fifth weight to the fifth prediction; and generating a hybrid sub-block corresponding to the current sub-block based on the sixth prediction.

[0275] Aspect 23. The method according to aspect 22, wherein each of the second weight, the third weight, the fourth weight, and the fifth weight includes one or more weight values ​​associated with one or more samples from the corresponding sub-block of the current sub-block, wherein the sum of the weight values ​​of the corner samples of the current sub-block is greater than the sum of the weight values ​​of the other boundary samples of the current sub-block.

[0276] Aspect 24. The method according to aspect 23, wherein the sum of the weights of the other boundary samples of the current sub-block is greater than the sum of the weights of the non-boundary samples of the current sub-block.

[0277] Aspect 25. The method according to any one of Aspects 20 to 24 further includes: determining that a Local Illumination Compensation (LIC) mode is used for an additional block of video data; and, based on determining that a LIC mode is used for the additional block, skipping signaling associated with information for an OBMC mode used for the additional block.

[0278] Aspect 26. The method according to aspect 25, wherein signaling that skips information associated with an OBMC mode for an additional block includes: signaling a syntax flag having a null value, the syntax flag being associated with the OBMC mode.

[0279] Aspect 27. The method according to any one of Aspects 25 to 26 further includes: receiving a signal including a syntax flag having a null value, the syntax flag being associated with an OBMC mode for additional blocks of video data.

[0280] Aspect 28. The method according to any one of Aspects 26 to 27 further includes: determining, based on a syntax flag having a null value, not to use the OBMC mode for an additional block.

[0281] Aspect 29. The method according to any one of Aspects 25 to 28, wherein signaling that skips information associated with the OBMC mode for the additional block includes: determining that the LIC mode is not used or enabled for the additional block based on determining that the LIC mode is used for the additional block; and skipping the signal notification of the value associated with the OBMC mode for the additional block.

[0282] Aspect 30. The method according to any one of Aspects 25 to 29 further includes: determining whether OBMC mode is enabled for the additional block; and based on determining whether OBMC mode is enabled for the additional block and determining that LIC mode is used for the additional block, determining to skip signaling information associated with the OBMC mode used for the additional block.

[0283] Aspect 31. The method according to any one of Aspects 20 to 30, further comprising: determining a decoding unit (CU) boundary OBMC mode for a current sub-block of a video data block; and determining a final prediction for the current sub-block based on the sum of a first result of applying weights associated with the current sub-block to a corresponding prediction associated with the current sub-block and a second result of applying one or more corresponding weights to one or more corresponding predictions associated with one or more sub-blocks adjacent to the current sub-block.

[0284] Aspect 32. The method according to any one of Aspects 20 to 31, wherein determining not to use motion information of adjacent sub-blocks for motion compensation of the current sub-block includes: skipping the use of motion information of adjacent sub-blocks for motion compensation of the current sub-block.

[0285] Aspect 33. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause one or more processors to perform the method according to any one of aspects 20 to 32.

[0286] Aspect 34. An apparatus comprising components for performing the method according to any one of aspects 20 to 32.

Claims

1. An apparatus for processing video data, comprising: A memory that stores instructions; as well as One or more processors, the one or more processors being configured to execute the instructions to cause the device to: Determine whether Overlapping Block Motion Compensation (OBMC) mode is enabled for the current sub-block of the video data block; For at least one adjacent sub-block that is adjacent to the current sub-block: It is determined that the first, second, and third conditions are met. The first condition includes a list of reference images used for unidirectional prediction of the current sub-block being used to predict the neighboring sub-blocks; The second condition includes one or more identical reference images used to determine motion vectors associated with the current sub-block and the neighboring sub-blocks; as well as The third condition includes a first difference between the horizontal motion vectors of the current sub-block and the adjacent sub-block and a second difference between the vertical motion vectors of the current sub-block and the adjacent sub-block not exceeding a motion vector difference threshold, wherein the motion vector difference threshold is greater than zero; as well as Based on the determination that the OBMC mode is enabled for the current sub-block and the determination that the first condition, the second condition, and the third condition are satisfied, it is determined that the motion information of the adjacent sub-blocks will not be used for the motion compensation of the current sub-block.

2. The apparatus according to claim 1, wherein, The one or more processors are configured to execute the instructions to cause the device to: Based on determining whether to use the decoder-side motion vector refinement (DMVR) mode, the sub-block-based temporal motion vector prediction (SbTMVP) mode, or the affine motion compensation prediction mode for the current sub-block, it is determined whether to execute the sub-block boundary (OBMC) mode for the current sub-block.

3. The apparatus according to claim 2, wherein, In order to execute the sub-block boundary OBMC mode for the current sub-block, the one or more processors are configured to execute the instructions to cause the device to: Determine a first prediction associated with the current sub-block, a second prediction associated with a first OBMC block adjacent to the top border of the current sub-block, a third prediction associated with a second OBMC block adjacent to the left border of the current sub-block, a fourth prediction associated with a third OBMC block adjacent to the bottom border of the current sub-block, and a fifth prediction associated with a fourth OBMC block adjacent to the right border of the current sub-block. A sixth prediction is determined based on the results of applying a first weight to the first prediction, applying a second weight to the second prediction, applying a third weight to the third prediction, applying a fourth weight to the fourth prediction, and applying a fifth weight to the fifth prediction. as well as Based on the sixth prediction, a hybrid sub-block corresponding to the current sub-block is generated.

4. The apparatus according to claim 3, wherein, Each of the second weight, the third weight, the fourth weight, and the fifth weight includes one or more weight values ​​associated with one or more samples from the corresponding sub-block of the current sub-block, wherein the sum of the weight values ​​of the corner samples of the current sub-block is greater than the sum of the weight values ​​of the other boundary samples of the current sub-block.

5. The apparatus according to claim 4, wherein, The sum of the weight values ​​of the other boundary samples of the current sub-block is greater than the sum of the weight values ​​of the non-boundary samples of the current sub-block.

6. The apparatus of claim 1, wherein the one or more processors are configured to execute the instructions to cause the apparatus to: Determine to use local illumination compensation LIC mode for additional blocks of video data; and Based on the determination that the LIC mode is used for the additional block, signaling associated with information for the OBMC mode used for the additional block is skipped.

7. The apparatus according to claim 6, wherein, In order to skip the signaling associated with the OBMC mode for the additional block, the one or more processors are configured to execute the instructions to cause the device to: The syntax flag with a null value is signaled, and the syntax flag is associated with the OBMC mode.

8. The apparatus of claim 6, wherein the one or more processors are configured to execute the instructions to cause the apparatus to: The signal received includes a syntax flag with a null value, which is associated with the OBMC mode for additional blocks of video data.

9. The apparatus according to claim 8, wherein, The one or more processors are configured to execute the instructions to cause the device to: Based on the syntax flag having the null value, it is determined that the OBMC mode will not be used for the additional block.

10. The apparatus according to claim 6, wherein, In order to skip the signaling associated with the OBMC mode for the additional block, the one or more processors are configured to execute the instructions to cause the device to: Based on determining that the LIC mode is used for the additional block, it is determined whether the OBMC mode is not used or enabled for the additional block; as well as Skip the signal notification associated with the OBMC mode used for the additional block.

11. The apparatus according to claim 6, wherein, The one or more processors are configured to execute the instructions to cause the device to: Determine whether to enable the OBMC mode for the additional block; as well as Based on determining whether to enable the OBMC mode for the additional block and determining that the LIC mode is used for the additional block, it is determined to skip signaling information associated with the OBMC mode used for the additional block.

12. The apparatus according to claim 1, wherein, The one or more processors are configured to execute the instructions to cause the device to: Determine that the current sub-block of the video data block uses the CU boundary OBMC mode; as well as The final prediction for the current sub-block is determined based on the sum of a first result of applying weights associated with the current sub-block to the corresponding prediction associated with the current sub-block and a second result of applying one or more corresponding weights to one or more corresponding predictions associated with one or more sub-blocks adjacent to the current sub-block.

13. The apparatus according to claim 1, wherein, To determine that motion information from the adjacent sub-blocks should not be used for motion compensation of the current sub-block, the one or more processors are configured to execute the instructions to cause the device to: Skip using the motion information of the adjacent sub-blocks for motion compensation of the current sub-block.

14. The apparatus according to claim 1, wherein, The device includes a decoder.

15. The apparatus of claim 14, further comprising: A display configured to display one or more output images associated with the video data.

16. The apparatus according to claim 1, wherein, The OBMC mode includes the sub-block boundary OBMC mode.

17. The apparatus according to claim 1, wherein, The device includes an encoder.

18. The apparatus of claim 17, further comprising a camera configured to capture images associated with the video data.

19. The apparatus according to claim 1, wherein, The device is a mobile device.

20. A method for processing video data, comprising: Determine whether Overlapping Block Motion Compensation (OBMC) mode is enabled for the current sub-block of the video data block; For at least one adjacent sub-block that is adjacent to the current sub-block, determine whether the first condition, the second condition, and the third condition are satisfied. The first condition includes a list of reference images used for unidirectional prediction of the current sub-block being used to predict the neighboring sub-blocks; The second condition includes one or more identical reference images used to determine motion vectors associated with the current sub-block and the neighboring sub-blocks; as well as The third condition includes a first difference between the horizontal motion vectors of the current sub-block and the adjacent sub-block and a second difference between the vertical motion vectors of the current sub-block and the adjacent sub-block not exceeding a motion vector difference threshold, wherein the motion vector difference threshold is greater than zero; as well as Based on the determination that the OBMC mode is used for the current sub-block and the determination that the first condition, the second condition, and the third condition are satisfied, it is determined that the motion information of the adjacent sub-blocks will not be used for the motion compensation of the current sub-block.

21. The method of claim 20, further comprising: Based on determining whether to use the decoder-side motion vector refinement (DMVR) mode, the sub-block-based temporal motion vector prediction (SbTMVP) mode, or the affine motion compensation prediction mode for the current sub-block, it is determined whether to execute the sub-block boundary (OBMC) mode for the current sub-block.

22. The method according to claim 21, wherein, Executing the sub-block boundary OBMC mode for the current sub-block includes: Determine a first prediction associated with the current sub-block, a second prediction associated with a first OBMC block adjacent to the top border of the current sub-block, a third prediction associated with a second OBMC block adjacent to the left border of the current sub-block, a fourth prediction associated with a third OBMC block adjacent to the bottom border of the current sub-block, and a fifth prediction associated with a fourth OBMC block adjacent to the right border of the current sub-block. A sixth prediction is determined based on the results of applying a first weight to the first prediction, applying a second weight to the second prediction, applying a third weight to the third prediction, applying a fourth weight to the fourth prediction, and applying a fifth weight to the fifth prediction; and Based on the sixth prediction, a hybrid sub-block corresponding to the current sub-block is generated.

23. The method according to claim 22, wherein, Each of the second weight, the third weight, the fourth weight, and the fifth weight includes one or more weight values ​​associated with one or more samples from the corresponding sub-block of the current sub-block, wherein the sum of the weight values ​​of the corner samples of the current sub-block is greater than the sum of the weight values ​​of the other boundary samples of the current sub-block.

24. The method according to claim 23, wherein, The sum of the weight values ​​of the other boundary samples of the current sub-block is greater than the sum of the weight values ​​of the non-boundary samples of the current sub-block.

25. The method of claim 20, further comprising: Determine whether to use the local illumination compensation LIC mode for additional blocks of video data; as well as Based on the determination that the LIC mode is used for the additional block, signaling associated with information for the OBMC mode used for the additional block is skipped.

26. The method of claim 25, wherein, The signaling that skips the information associated with the OBMC mode used for the additional block includes: The syntax flag with a null value is signaled, and the syntax flag is associated with the OBMC mode.

27. The method of claim 25, further comprising: The signal received includes a syntax flag with a null value, which is associated with the OBMC mode for additional blocks of video data.

28. The method of claim 27, further comprising: Based on the syntax flag having the null value, it is determined that the OBMC mode will not be used for the additional block.

29. The method according to claim 25, wherein, The signaling that skips the information associated with the OBMC mode used for the additional block includes: Based on determining that the LIC mode is used for the additional block, it is determined that the OBMC mode is either not used or enabled for the additional block; and Skip the signal notification associated with the OBMC mode used for the additional block.

30. The method of claim 25, further comprising: Determine whether to enable the OBMC mode for the additional block; as well as Based on determining whether to enable the OBMC mode for the additional block and determining that the LIC mode is used for the additional block, it is determined to skip signaling information associated with the OBMC mode used for the additional block.

31. The method of claim 20, further comprising: Determine that the current sub-block of the video data block uses the CU boundary OBMC mode; as well as The final prediction for the current sub-block is determined based on the sum of a first result of applying weights associated with the current sub-block to the corresponding prediction associated with the current sub-block and a second result of applying one or more corresponding weights to one or more corresponding predictions associated with one or more sub-blocks adjacent to the current sub-block.

32. The method according to claim 20, wherein, Determining not to use the motion information of the adjacent sub-blocks for the motion compensation of the current sub-block includes: Skip motion compensation for the current sub-block and use motion information from the adjacent sub-blocks.

33. The method according to claim 20, wherein, The OBMC mode includes the sub-block boundary OBMC mode.

Citation Information

Patent Citations

  • Weighted prediction in video coding

    WO2020147747A1