Method and device for performing optical flow prediction correction on affine decoding block

Through the optical flow prediction and correction method of affine decoding block, the problems of high coding complexity and low prediction accuracy in the affine motion compensation method are solved, and better complexity and accuracy balance in video decoding are achieved, and compression performance is improved.

CN118945317BActive Publication Date: 2025-08-26HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410918815.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-28
Filing Date
2020-03-20
Publication Date
2025-08-26
Estimated Expiration
2040-03-20

AI Technical Summary

Technical Problem

The existing sub-block-based affine motion compensation method has the problem of high coding complexity and low prediction accuracy in video decoding, and it is difficult to achieve a good balance between complexity and prediction accuracy.

Method used

Optical flow prediction correction (PROF) method is used to perform optical flow processing on the current sub-block of the affine coded block, and the incremental predicted value and the predicted sample value of the current sub-block are corrected to ensure that calculations are performed only when the prediction accuracy is improved to balance the coding complexity and prediction accuracy.

Benefits of technology

The overall compression performance of the video decoding method is improved, and the compression ratio of the decoding method is improved by using optical flow at the pixel/sample level particle size to correct the affine motion compensation prediction based on sub-blocks to avoid unnecessary increase in computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118945317B_ABST
    Figure CN118945317B_ABST
Patent Text Reader

Abstract

The present invention relates to an apparatus, an encoder, a decoder, and various corresponding methods for performing optical flow prediction refinement (PROF) on an affine decoded block. When the affine decoded block satisfies multiple optical flow decision conditions, a PROF process is performed on a current sub-block of the affine decoded block to obtain a modified predicted sample value of the current sub-block of the affine decoded block. After performing sub-block-based affine motion compensation, the predicted sample value of the current sample of the current sub-block is modified by adding a delta prediction value. Therefore, the present invention can better balance decoding complexity and prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202080023321.7, and the original application date is March 20, 2020. The entire content of the original application is incorporated into this application by reference.

[0002] Cross-reference to related applications

[0003] This patent application claims priority to U.S. Provisional Patent Application No. 62 / 821,440, filed on March 20, 2019, and to U.S. Provisional Patent Application No. 62 / 839,765, filed on April 28, 2019. The entire disclosures of the above patent applications are incorporated herein by reference. Technical Field

[0004] Embodiments of the present invention generally relate to the field of image processing, and more particularly to a method for modifying sub-block based affine motion compensation prediction values ​​using optical flow when one or more constraints are required. Background Art

[0005] Video coding (video encoding and decoding) is widely used in digital video applications such as broadcast digital television, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat and video conferencing, DVD and Blu-ray Disc, video content acquisition and editing systems, and camcorders for security applications.

[0006] Even relatively short videos require a large amount of video data to describe them, which can create difficulties when the data is to be streamed or otherwise transmitted across communication networks with limited bandwidth capacity. Therefore, video data is typically compressed before being transmitted over modern telecommunications networks. Since memory resources may be limited, the size of the video can also be an issue when storing the video on a storage device. Video compression devices typically use software and / or hardware to encode the video data at the source side before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination side by a video decompression device that decodes the video data. With limited network resources and the growing demand for higher video quality, there is a need for improved compression and decompression techniques that can increase the compression ratio with little impact on image quality.

[0007] In recent years, affine tools have been introduced in general video coding. In theory, affine motion model parameters can be used to derive the motion vector for each sample in a coding block. However, due to the high complexity of generating sample-based affine motion compensation prediction values, sub-block-based affine motion compensation methods are used. In this method, the coding block is divided into multiple sub-blocks, each of which is assigned a motion vector (MV) derived using the affine motion model parameters. However, the accuracy of sub-block-based prediction is not high. Therefore, it is necessary to strike a good balance between decoding complexity and prediction accuracy. Summary of the Invention

[0008] Embodiments of the present application provide various methods and apparatus for encoding and decoding according to the independent claims. Embodiments of the present application provide various apparatus and methods for performing optical flow prediction refinement (PROF) on affine decoded blocks, thereby achieving a good balance between the complexity and accuracy of sub-block-based affine prediction.

[0009] Embodiments are defined by the features of the independent claims, while further advantageous implementations of these embodiments are defined by the features of the dependent claims.

[0010] Particular embodiments are outlined in the accompanying independent claims, further embodiments are outlined in the dependent claims.

[0011] The above and other objects are achieved by the subject matter claimed in the independent claims. Other implementations are apparent from the dependent claims, the description and the drawings.

[0012] According to a first aspect, the present invention relates to a method for performing optical flow prediction refinement (PROF) on an affine decoded block (i.e., a block encoded or decoded using an affine tool). The method is applied to a sub-block consisting of samples within the affine decoded block. The method is performed by an encoding device or a decoding device. The method may include:

[0013] performing a PROF process on a current sub-block of the affine decoded block to obtain a modified predicted sample value (i.e., a final predicted sample value) of the current sub-block (e.g., each sub-block) of the affine decoded block, wherein the affine decoded block does not satisfy a plurality of PROF application constraints;

[0014] The performing of the PROF process on the current sub-block of the affine decoding block includes: performing optical flow processing on the current sub-block to obtain a delta prediction value of a current sample of the current sub-block; and obtaining a revised predicted sample value of the current sample based on the delta prediction value of the current sample and the predicted sample value of the current sample of the current sub-block (performing optical flow processing on the current sub-block to obtain a delta prediction value of the current sub-block; and obtaining a revised predicted sample value of the current sub-block based on the delta prediction value of the current sub-block and the predicted sample value of the current sub-block). It will be understood that when the revised predicted sample value of each sub-block of the affine decoding block is generated, the revised predicted sample value of the affine decoding block is naturally generated.

[0015] Therefore, an improved method is provided that can better balance decoding complexity and prediction accuracy. The prediction refinement with optical flow (PROF) process is performed according to conditions, thereby using optical flow to refine sub-block-based affine motion compensation predictions at the pixel / sample level granularity. These conditions ensure that the calculations involved in PROF are only performed when prediction accuracy can be improved, thereby avoiding unnecessary increases in computational complexity. Therefore, the beneficial effects achieved by the disclosed technology improve the overall compression performance of the decoding method.

[0016] It should be noted that the terms "block," "coding block," or "image block" used in this disclosure may include transform units (TUs), prediction units (PUs), coding units (CUs), etc. In VVC, transform units and coding units are aligned in most cases, except in a few scenarios when TU tiling or sub-block transform (SBT) is used. It is understood that the terms "block," "image block," "coding block," and "picture block" are used interchangeably herein. The terms "affine block," "affine picture block," "affine coded block," and "affine motion block" are used interchangeably herein. The terms "sample" and "pixel" are used interchangeably in this disclosure. The terms "predicted sample value" and "predicted pixel value" are used interchangeably in this disclosure. The terms "sample position" and "pixel position" are used interchangeably in this disclosure.

[0017] According to the first aspect, in a possible implementation of the method, before performing the PROF process on the current sub-block of the affine decoding block, the method further includes: determining that the affine decoding block does not satisfy the multiple PROF application constraints.

[0018] According to the first aspect, in one possible implementation of the method, the multiple PROF application constraints include: first indication information indicating that PROF is disabled for an image including the affine decoding block, or first indication information indicating that PROF is disabled for multiple slices associated with the image including the affine decoding block; and second indication information indicating that the affine decoding block is not split, that is, the variable fallbackModeTriggered is set to 1. It is understood that when the variable fallbackModeTriggered is set to 1, the affine decoding block does not need to be split, that is, each sub-block of the affine decoding block has the same motion vector, indicating that the affine decoding block has only translational motion. When the variable fallbackModeTriggered is set to 0, the affine decoding block needs to be split, that is, each sub-block of the affine decoding block has its own motion vector, indicating that the affine decoding block has non-translational motion.

[0019] The present invention may not apply PROF to the affine decoding block in some cases. These cases are determined based on the multiple PROF application constraints. This can better balance decoding complexity and prediction accuracy.

[0020] According to the first aspect or any foregoing implementation of the first aspect, in a possible implementation of the method, performing optical flow processing on the current sub-block to obtain an incremental prediction value of a current sample of the current sub-block includes:

[0021] Obtaining a second prediction matrix (in one example, the second prediction matrix is ​​generated based on a first prediction matrix corresponding to the predicted sample values ​​of the current sub-block. Here, the predicted sample values ​​of the current sub-block can be obtained by performing sub-block-based affine motion compensation on the current sub-block), wherein a size of the second prediction matrix is ​​larger than a size of the first prediction matrix (for example, a size of the first prediction matrix is ​​sbWidth*sbHeight, a size of the second prediction matrix is ​​(sbWidth+2)*(sbHeight+2), and variables sbWidth and sbHeight represent a width and a height of the current sub-block, respectively). That is, obtaining the second prediction matrix includes: generating a first prediction matrix based on motion information of the current sub-block, wherein elements of the first prediction matrix correspond to the predicted sample values ​​of the current sub-block, and generating the second prediction matrix based on the first prediction matrix; or generating the second prediction matrix based on the motion information of the current sub-block;

[0022] Generate a horizontal prediction gradient matrix and a vertical prediction gradient matrix according to the second prediction matrix, wherein a size of the second prediction matrix is ​​greater than or equal to a size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix (for example, a size of the horizontal prediction gradient matrix or the vertical prediction gradient matrix is ​​sbWidth*sbHeight, and a size of the second prediction matrix is ​​(sbWidth+2)*(sbHeight+2));

[0023] An incremental prediction value (ΔI(i, j)) of the current sample of the current subblock is calculated based on the horizontal prediction gradient value of the current sample in the horizontal prediction gradient matrix, the vertical prediction gradient value of the current sample in the vertical prediction gradient matrix, and the difference (MVD) between the motion vector of the current sample of the current subblock and the motion vector of the center sample of the subblock. It can be understood that the MVD has a horizontal component and a vertical component. The horizontal prediction gradient value of the current sample in the horizontal prediction gradient matrix corresponds to the horizontal component of the MVD, and the vertical prediction gradient value of the current sample in the vertical prediction gradient matrix corresponds to the vertical component of the MVD.

[0024] It should be noted that the affine block can be an encoding block or a decoding block of an image of a video signal. The current sub-block of the affine decoding block is a 4×4 block, etc. The luminance position (xCb, yCb) represents the position of the upper left sample of the affine decoding block relative to the upper left sample of the current image. Each sample of the current sub-block can be indexed by the absolute position of the sample relative to (with respect to or relative to) the upper left sample of the image (for example, (x, y)), or by the relative position of the sample relative to the upper left sample of the sub-block (other coordinates) (for example, (xSb+i, ySb+j)). Here, (xSb, ySb) represents the coordinates of the upper left sample of the sub-block relative to the upper left sample of the image.

[0025] The first prediction matrix can be a two-dimensional array including multiple rows and columns, and an element of the array can be indexed by (i, j), where i is the horizontal / row index and j is the vertical / column index. The range of i and j can be i=0..sbWidth–1, j=0..sbHeight–1, and so on. Here, sbWidth represents the width of the sub-block and sbHeight represents the height of the sub-block. In some examples, the size of the first prediction matrix is ​​the same as the size of the current sub-block. For example, the size of the first prediction matrix can be 4×4, and the size of the current sub-block can be 4×4.

[0026] The second prediction matrix can be a two-dimensional array including multiple rows and columns, and an element of the array can be indexed by (i, j), where i is the horizontal / row index and j is the vertical / column index. The range of i and j can be i=–1..sbWidth and j=–1..sbHeight, and so on. Here, sbWidth represents the width of the sub-block and sbHeight represents the height of the sub-block. In some examples, the size of the second prediction matrix is ​​larger than the size of the first prediction matrix. That is, the size of the second prediction matrix can be larger than the size of the current sub-block. For example, the size of the second prediction matrix can be (sbWidth+2)*(sbHeight+2), while the size of the current sub-block is sbWidth*sbHeight. For example, the size of the second prediction matrix can be 6×6, while the size of the current sub-block is 4×4.

[0027] The horizontal and vertical prediction gradient matrices can be any two-dimensional arrays comprising multiple rows and columns, and an element of the array can be indexed by (i, j), where x is the horizontal / row index and y is the vertical / column index. The range of i and j can be i = 0..sbWidth–1, j = 0..sbHeight–1, and so on. sbWidth represents the width of the sub-block, and sbHeight represents the height of the sub-block. In some examples, the size of the horizontal and vertical prediction gradient matrices is the same as the size of the current sub-block. For example, the size of the horizontal and vertical prediction gradient matrices can be 4×4, and the size of the current sub-block is 4×4.

[0028] If the position (x, y) of an element of the horizontal prediction gradient matrix in the horizontal prediction gradient matrix is ​​the same as the position (p, q) of an element of the vertical prediction gradient matrix in the vertical prediction gradient matrix, that is, (x, y) = (p, q), then the two elements correspond.

[0029] Therefore, the PROF process can use optical flow to modify the sub-block based affine motion compensation prediction value at sample level granularity without increasing the memory access bandwidth (because the second prediction matrix is ​​based on the first prediction matrix or the (original) predicted sample value of the current sub-block), thereby achieving more refined motion compensation.

[0030] According to the first aspect or any of the above implementations of the first aspect, in one possible implementation of the method, a motion vector difference between a motion vector of a current sample unit (e.g., a 2×2 sample block) of the current sample and a motion vector of a center sample of the sub-block is used as the difference between the motion vector of the current sample of the current sub-block and the motion vector of the center sample of the sub-block. Here, the motion vector of the center sample of the sub-block can be understood as the MV of the sub-block to which the current sample (i, j) belongs (i.e., the sub-block MV). Calculating the motion vector difference using a sample unit such as a 2×2 sample block can balance processing overhead and prediction accuracy.

[0031] According to the first aspect or any of the above implementations of the first aspect, in one possible implementation of the method, an element of the second prediction matrix is ​​represented as I1(p,q), where the value range of p is [–1, sbW], and the value range of q is [–1, sbH];

[0032] An element of the horizontal prediction gradient matrix is ​​denoted as X(i, j) and corresponds to a sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1];

[0033] An element of the vertical prediction gradient matrix is ​​denoted as Y(i, j) and corresponds to the sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1], where,

[0034] sbW represents the width of the current sub-block in the affine decoding block, and sbH represents the height of the current sub-block in the affine decoding block.

[0035] In another representation, an element of the second prediction matrix is ​​represented as I1(p,q), where the value range of p is [0,subW+1] and the value range of q is [0,subH+1];

[0036] An element of the horizontal prediction gradient matrix is ​​denoted as X(i, j) and corresponds to a sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [1, sbW] and the value range of j is [1, sbH];

[0037] An element of the vertical prediction gradient matrix is ​​denoted as Y(i, j) and corresponds to the sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [1, sbW] and the value range of j is [1, sbH], where,

[0038] sbW represents the width of the current sub-block in the affine decoding block, and sbH represents the height of the current sub-block in the affine decoding block.

[0039] It can be understood that when the value range of p is [0, subW+1] and the value range of q is [0, subH+1], the upper left sample (or coordinate origin) is located at (1, 1); and when the value range of p is [–1, subW] and the value range of q is [–1, subH], the upper left sample (or coordinate origin) is located at (0, 0).

[0040] According to the first aspect or any of the above-mentioned implementations of the first aspect, in one possible implementation of the method, before performing the PROF process on the current sub-block of the affine decoding block, the method further includes: performing sub-block-based affine motion compensation on the current sub-block of the affine decoding block to obtain the (original or to-be-corrected) predicted sample value of the current sub-block.

[0041] According to a second aspect of the present invention, a method for performing optical flow prediction refinement (PROF) on an affine decoded block is provided. The method comprises:

[0042] Performing a PROF process on a current sub-block of the affine decoded block to obtain a modified predicted sample value (i.e., a final predicted sample value) of the current sub-block of the affine decoded block, wherein the affine decoded block satisfies a plurality of optical flow decision conditions (here, satisfying the plurality of optical flow decision conditions means not satisfying all PROF application constraints);

[0043] The PROF process performed on the current sub-block of the affine decoding block includes: performing optical flow processing on the current sub-block to obtain an incremental prediction value of the current sample of the current sub-block; and obtaining a corrected prediction sample value of the current sample based on the incremental prediction value of the current sample and the (original or to-be-corrected) prediction sample value of the current sample of the current sub-block.

[0044] Therefore, an improved method is provided that can better balance decoding complexity and prediction accuracy. The prediction refinement with optical flow (PROF) process is performed according to conditions, thereby using optical flow to modify the sub-block-based affine motion compensation prediction value at the pixel / sample level granularity. These conditions ensure that the calculations involved in PROF are only performed when prediction accuracy can be improved, thereby avoiding unnecessary increases in computational complexity. Therefore, the beneficial effects achieved by the disclosed technology improve the overall compression performance of the decoding method.

[0045] According to the second aspect, in a possible implementation of the method, before performing the PROF process on the current sub-block of the affine decoding block, the method further includes: determining whether the affine decoding block satisfies the multiple optical flow decision conditions.

[0046] According to the second aspect, in one possible implementation of the method, the multiple optical flow decision conditions include: first indication information indicating that PROF is enabled for an image including the affine decoding block, or first indication information indicating that PROF is enabled for multiple slices associated with the image including the affine decoding block; and second indication information indicating that the affine decoding block is segmented, for example, a variable fallbackModeTriggered is set to 0. It is understandable that when the variable fallbackModeTriggered is set to 0, the affine decoding block needs to be segmented, that is, each sub-block of the affine decoding block has its own motion vector, indicating that the affine decoding block has non-translational motion.

[0047] According to the design of the multiple PROF application constraints, PROF can be applied when all PROF application constraints are not satisfied. Therefore, the present invention can balance decoding complexity and prediction accuracy.

[0048] According to the second aspect or any of the foregoing implementations of the second aspect, in a possible implementation of the method, performing optical flow processing on the current sub-block to obtain an incremental prediction value of a current sample of the current sub-block includes:

[0049] Obtaining a second prediction matrix, wherein elements of the second prediction matrix are based on predicted sample values ​​of the current sub-block; in some examples, obtaining the second prediction matrix includes: generating a first prediction matrix based on motion information of the current sub-block, wherein elements of the first prediction matrix correspond to predicted sample values ​​of the current sub-block, and generating the second prediction matrix based on the first prediction matrix; or generating the second prediction matrix based on the motion information of the current sub-block;

[0050] generating a horizontal prediction gradient matrix and a vertical prediction gradient matrix according to the second prediction matrix, wherein a size of the second prediction matrix is ​​greater than or equal to a size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix;

[0051] Calculate an incremental prediction value (ΔI(i, j)) of the current sample of the current subblock according to a horizontal prediction gradient value of the current sample in the horizontal prediction gradient matrix, a vertical prediction gradient value of the current sample in the vertical prediction gradient matrix, and a difference between a motion vector of the current sample of the current subblock and a motion vector of a center sample of the subblock.

[0052] According to the second aspect or any of the above implementations of the second aspect, in one possible implementation of the method, the method further includes: performing sub-block-based affine motion compensation on the current sub-block of the affine decoding block to obtain the (original) predicted sample value of the current sub-block of the affine decoding block.

[0053] According to the second aspect or any of the foregoing implementations of the second aspect, in a possible implementation of the method, a motion vector difference between a motion vector of a current sample unit (for example, a 2×2 sample block) to which the current sample belongs and a motion vector of a center sample of the sub-block is used as a difference between the motion vector of the current sample of the current sub-block and the motion vector of the center sample of the sub-block.

[0054] According to the second aspect or any of the above implementations of the second aspect, in a possible implementation of the method,

[0055] An element of the second prediction matrix is ​​represented by I1(p,q), where the value range of p is [–1, sbW] and the value range of q is [–1, sbH];

[0056] An element of the horizontal prediction gradient matrix is ​​denoted as X(i, j) and corresponds to a sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1];

[0057] An element of the vertical prediction gradient matrix is ​​denoted as Y(i, j) and corresponds to the sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1], where,

[0058] sbW represents the width of the current sub-block in the affine decoding block, and sbH represents the height of the current sub-block in the affine decoding block.

[0059] According to a third aspect, the present invention relates to an apparatus for performing optical flow prediction refinement (PROF) on an affine decoded block (i.e., a block encoded or decoded using an affine tool). The apparatus corresponds to an encoding apparatus or a decoding apparatus. The apparatus may include:

[0060] A determining unit is configured to determine that the affine decoding block does not satisfy a plurality of PROF application constraints.

[0061] a prediction processing unit, configured to perform a PROF process on a current sub-block of the affine decoded block to obtain a modified predicted sample value (i.e., a final predicted sample value) of the current sub-block (e.g., each sub-block) of the affine decoded block, wherein the affine decoded block does not satisfy a plurality of PROF application constraints;

[0062] The prediction processing unit is configured to perform optical flow processing on the current sub-block to obtain an incremental prediction value of a current sample of the current sub-block; and obtain a revised predicted sample value of the current sample based on the incremental prediction value of the current sample and the predicted sample value of the current sample of the current sub-block (performing optical flow processing on the current sub-block to obtain an incremental prediction value of the current sub-block; and obtaining a revised predicted sample value of the current sub-block based on the incremental prediction value of the current sub-block and the predicted sample value of the current sub-block). It will be understood that when the revised predicted sample value of each sub-block of the affine decoding block is generated, the revised predicted sample value of the affine decoding block is naturally generated.

[0063] According to the third aspect, in one possible implementation of the device, the multiple PROF application constraints include: first indication information indicating that PROF is disabled for an image including the affine decoding block, or first indication information indicating that PROF is disabled for multiple slices associated with the image including the affine decoding block; and second indication information indicating that the affine decoding block is not split, that is, the variable fallbackModeTriggered is set to 1. It is understood that when the variable fallbackModeTriggered is set to 1, the affine decoding block does not need to be split, that is, each sub-block of the affine decoding block has the same motion vector, indicating that the affine decoding block has only translational motion. When the variable fallbackModeTriggered is set to 0, the affine decoding block needs to be split, that is, each sub-block of the affine decoding block has its own motion vector, indicating that the affine decoding block has non-translational motion.

[0064] According to the third aspect or any of the above implementations of the third aspect, in a possible implementation of the device, the prediction processing unit is configured to: obtain a second prediction matrix (in one example, the second prediction matrix is ​​generated based on a first prediction matrix corresponding to the predicted sample values ​​of the current sub-block. Here, the predicted sample values ​​of the current sub-block can be obtained by performing sub-block-based affine motion compensation on the current sub-block.), wherein a size of the second prediction matrix is ​​larger than a size of the first prediction matrix (for example, a size of the first prediction matrix is ​​sbWidth*sbHeight, a size of the second prediction matrix is ​​(sbWidth+2)*(sbHeight+2), and variables sbWidth and sbHeight represent a width and a height of the current sub-block, respectively). That is, obtaining the second prediction matrix includes: generating a first prediction matrix based on motion information of the current sub-block, wherein elements of the first prediction matrix correspond to the predicted sample values ​​of the current sub-block, and generating the second prediction matrix based on the first prediction matrix; or generating the second prediction matrix based on the motion information of the current sub-block;

[0065] Generate a horizontal prediction gradient matrix and a vertical prediction gradient matrix according to the second prediction matrix, wherein a size of the second prediction matrix is ​​greater than or equal to a size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix (for example, a size of the horizontal prediction gradient matrix or the vertical prediction gradient matrix is ​​sbWidth*sbHeight, and a size of the second prediction matrix is ​​(sbWidth+2)*(sbHeight+2));

[0066] Calculate an incremental prediction value (ΔI(i, j)) of the current sample of the current subblock according to a horizontal prediction gradient value of the current sample in the horizontal prediction gradient matrix, a vertical prediction gradient value of the current sample in the vertical prediction gradient matrix, and a difference between a motion vector of the current sample of the current subblock and a motion vector of a center sample of the subblock.

[0067] According to the third aspect or any of the above implementations of the third aspect, in a possible implementation of the device, the motion vector difference between the motion vector of the current sample unit (e.g., a 2×2 sample block) of the current sample and the motion vector of the center sample of the sub-block is included as the difference between the motion vector of the current sample of the current sub-block and the motion vector of the center sample of the sub-block. Here, the motion vector of the center sample of the sub-block can be understood as the MV of the sub-block to which the current sample (i, j) belongs (i.e., the sub-block MV). Using sample units such as 2×2 sample blocks to calculate the motion vector difference can balance processing overhead and prediction accuracy. According to the third aspect or any of the above implementations of the third aspect, in a possible implementation of the device, an element of the second prediction matrix is ​​represented as I1(p,q), where the value range of p is [–1,sbW] and the value range of q is [–1,sbH];

[0068] An element of the horizontal prediction gradient matrix is ​​denoted as X(i, j) and corresponds to a sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1];

[0069] An element of the vertical prediction gradient matrix is ​​denoted as Y(i, j) and corresponds to the sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1], where,

[0070] sbW represents the width of the current sub-block in the affine decoding block, and sbH represents the height of the current sub-block in the affine decoding block.

[0071] According to the third aspect or any of the above implementations of the third aspect, in one possible implementation of the device, the prediction processing unit is used to perform sub-block-based affine motion compensation on the current sub-block of the affine decoding block to obtain the (original or to be corrected) predicted sample value of the current sub-block.

[0072] According to a fourth aspect of the present invention, there is provided an apparatus for performing optical flow prediction refinement (PROF) on an affine decoded block. The apparatus may include:

[0073] a determining unit, configured to determine whether the affine decoding block satisfies a plurality of optical flow decision conditions (here, satisfying the plurality of optical flow decision conditions means not satisfying all PROF application constraints);

[0074] A prediction processing unit is configured to perform a PROF process on a current sub-block of the affine decoding block to obtain a modified predicted sample value (i.e., a final predicted sample value) of the current sub-block of the affine decoding block, wherein the affine decoding block satisfies multiple optical flow decision conditions; the prediction processing unit is configured to: perform optical flow processing on the current sub-block to obtain an incremental predicted value of a current sample of the current sub-block; and obtain a modified predicted sample value of the current sample based on the incremental predicted value of the current sample and the (original or to-be-modified) predicted sample value of the current sample of the current sub-block.

[0075] According to the fourth aspect, in one possible implementation of the device, the multiple optical flow decision conditions include: first indication information indicating that PROF is enabled for an image including the affine decoding block, or first indication information indicating that PROF is enabled for multiple slices associated with the image including the affine decoding block; and second indication information indicating that the affine decoding block is segmented, for example, a variable fallbackModeTriggered is set to 0. It is understandable that when the variable fallbackModeTriggered is set to 0, the affine decoding block needs to be segmented, that is, each sub-block of the affine decoding block has its own motion vector, indicating that the affine decoding block has non-translational motion.

[0076] According to the fourth aspect or any of the foregoing implementations of the fourth aspect, in one possible implementation of the device, the prediction processing unit is configured to:

[0077] Obtaining a second prediction matrix, wherein elements of the second prediction matrix are based on predicted sample values ​​of the current sub-block; in some examples, obtaining the second prediction matrix includes: generating a first prediction matrix based on motion information of the current sub-block, wherein elements of the first prediction matrix correspond to predicted sample values ​​of the current sub-block, and generating the second prediction matrix based on the first prediction matrix; or generating the second prediction matrix based on the motion information of the current sub-block;

[0078] generating a horizontal prediction gradient matrix and a vertical prediction gradient matrix according to the second prediction matrix, wherein a size of the second prediction matrix is ​​greater than or equal to a size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix;

[0079] Calculate an incremental prediction value (ΔI(i, j)) of the current sample of the current subblock according to a horizontal prediction gradient value of the current sample in the horizontal prediction gradient matrix, a vertical prediction gradient value of the current sample in the vertical prediction gradient matrix, and a difference between a motion vector of the current sample of the current subblock and a motion vector of a center sample of the subblock.

[0080] According to the fourth aspect or any of the above implementations of the fourth aspect, in one possible implementation of the device, the prediction processing unit is used to perform sub-block-based affine motion compensation on the current sub-block of the affine decoding block to obtain the (original) predicted sample value of the current sub-block of the affine decoding block.

[0081] According to the fourth aspect or any foregoing implementation manner of the fourth aspect, in a possible implementation manner of the device, a motion vector difference between a motion vector of a current sample unit (for example, a 2×2 sample block) to which the current sample belongs and a motion vector of a center sample of the sub-block is used as a difference between the motion vector of the current sample of the current sub-block and the motion vector of the center sample of the sub-block.

[0082] According to the fourth aspect or any of the above implementations of the fourth aspect, in a possible implementation of the device, an element of the second prediction matrix is ​​represented as I1(p,q), where the value range of p is [–1, sbW], and the value range of q is [–1, sbH].

[0083] An element of the horizontal prediction gradient matrix is ​​denoted as X(i, j) and corresponds to a sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1];

[0084] An element of the vertical prediction gradient matrix is ​​denoted as Y(i, j) and corresponds to the sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1], where,

[0085] sbW represents the width of the current sub-block in the affine decoding block, and sbH represents the height of the current sub-block in the affine decoding block.

[0086] The method provided in the first aspect of the present invention can be performed by the apparatus provided in the third aspect of the present invention. Other features and implementations of the apparatus provided in the third aspect of the present invention correspond to the features and implementations of the method provided in the first aspect of the present invention.

[0087] The method provided in the second aspect of the present invention can be performed by the apparatus provided in the fourth aspect of the present invention. Other features and implementations of the apparatus provided in the fourth aspect of the present invention correspond to the features and implementations of the method provided in the second aspect of the present invention.

[0088] According to a fifth aspect, the present invention relates to an encoder (20). The encoder (20) comprises a processing circuit for executing the method provided by the first or second aspect or an implementation thereof.

[0089] According to a sixth aspect, the present invention relates to a decoder (30). The decoder (30) comprises a processing circuit for executing the method provided by the first or second aspect or an implementation thereof.

[0090] According to a seventh aspect, the present invention relates to a decoder. The decoder comprises:

[0091] one or more processors;

[0092] A non-transitory computer-readable storage medium is coupled to the one or more processors and stores a program executed by the one or more processors, wherein when the one or more processors execute the program, the decoder is used to execute the method provided by the first aspect or its implementation.

[0093] According to an eighth aspect, the present invention relates to an encoder. The encoder comprises:

[0094] one or more processors;

[0095] A non-transitory computer-readable storage medium is coupled to the one or more processors and stores a program executed by the one or more processors, wherein when the one or more processors execute the program, the encoder is used to execute the method provided by the first aspect or its implementation.

[0096] According to a ninth aspect, the present invention relates to an apparatus for encoding a video stream, wherein the apparatus comprises a processor and a memory, wherein the memory stores instructions that cause the processor to execute the method provided by the second aspect.

[0097] According to a tenth aspect, the present invention relates to an apparatus for decoding a video stream, wherein the apparatus comprises a processor and a memory, wherein the memory stores instructions that cause the processor to execute the method provided by the first aspect.

[0098] According to an eleventh aspect, the present invention relates to a computer program. The computer program comprises program code. When the program code is executed on a computer, it is used to perform the method provided by the first or second aspect or any possible embodiment of the first or second aspect.

[0099] According to a twelfth aspect, a computer-readable storage medium storing instructions is provided. When executed, the instructions cause one or more processors to decode video data. The instructions cause the one or more processors to perform the method provided by the first or second aspect, or any possible embodiment of the first or second aspect.

[0100] According to another aspect, a video image encoding method is provided. The video image encoding method includes: determining indication information, wherein the indication information is used to indicate whether a to-be-encoded image block is to be encoded according to a target inter-frame prediction method, the target inter-frame prediction method including the inter-frame prediction method provided by the first or second aspect or any possible embodiment of the first or second aspect; and encoding the indication information in a bitstream.

[0101] According to another aspect, a video image decoding method is provided. The video image decoding method includes: parsing a bitstream to obtain indication information, wherein the indication information is used to indicate whether a to-be-decoded image block is to be processed according to a target inter-frame prediction method, the target inter-frame prediction method including the inter-frame prediction method provided by the first or second aspect or any possible embodiment of the first or second aspect; when the indication information indicates that processing is to be performed according to the target inter-frame prediction method, processing the to-be-decoded image block according to the target inter-frame prediction method.

[0102] The following drawings and description set forth in detail one or more embodiments. Other features, objects, and advantages are apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] The embodiments of the present invention are described in more detail below with reference to the accompanying drawings and schematic diagrams, in which:

[0104] Figure 1A is a block diagram of an example of a video decoding system for implementing an embodiment of the present invention;

[0105] Figure 1B is a block diagram of another example of a video decoding system for implementing an embodiment of the present invention;

[0106] Figure 2 is a block diagram of an example of a video encoder for implementing an embodiment of the present invention;

[0107] Figure 3is a block diagram of an exemplary structure of a video decoder for implementing an embodiment of the present invention;

[0108] Figure 4 A block diagram of an example of an encoding device or a decoding device.

[0109] Figure 5 A block diagram of another example of an encoding device or a decoding device.

[0110] Figure 6 Schematic diagram of spatial and temporal candidate motion information of the current block.

[0111] Figure 7 Schematic diagram of the current affine decoding block and the adjacent affine decoding block where A1 is located.

[0112] Figure 8A A schematic diagram illustrating an example of a constructed control point motion vector prediction method.

[0113] Figure 8B A schematic diagram illustrating an example of a constructed control point motion vector prediction method.

[0114] Figure 9A A flowchart of a decoding method according to an embodiment of the present application.

[0115] Figure 9B Schematic diagram of the constructed control point motion vector prediction method.

[0116] Figure 9C Schematic diagram of samples or pixels of the current affine decoding block and motion vectors of the upper left control point and the upper right control point.

[0117] Figure 9D A schematic diagram of a 6×6 prediction signal window and a 4×4 sub-block for calculating and generating a horizontal prediction gradient matrix and a vertical prediction gradient matrix.

[0118] Figure 9E Schematic diagram of an 18×18 prediction signal window and a 16×16 block for calculating and generating a horizontal prediction gradient matrix and a vertical prediction gradient matrix.

[0119] Figure 10 It is the sample MV (denoted as v(i,j)) calculated for the sample position (i,j) and the sub-block MV (V SB ) between the two (red arrow).

[0120] Figure 11AA schematic diagram of a method for performing optical flow prediction refinement (PROF) on an affine decoding block is provided in accordance with an embodiment of the present invention.

[0121] Figure 11B A schematic diagram of another method for performing optical flow prediction and correction (PROF) on an affine decoded block is provided according to an embodiment of the present invention.

[0122] Figure 12 A schematic diagram of the PROF process provided in an embodiment of the present invention.

[0123] Figure 13 A schematic diagram of the surrounding area (area or region) and internal area of ​​an (M+2)*(N+2) prediction block provided in one embodiment of the present invention.

[0124] Figure 14 A schematic diagram of the surrounding area and internal area of ​​a (M+2)*(N+2) prediction block provided by another embodiment of the present invention.

[0125] Figure 15 A block diagram of an exemplary structure of an apparatus for performing optical flow prediction and correction (PROF) on an affine decoded block of a video signal is provided in some aspects of the present invention.

[0126] Figure 16 A block diagram of an exemplary structure of a content provision system for implementing content distribution services.

[0127] Figure 17 A block diagram of an example structure of a terminal device.

[0128] In the following, identical reference numerals denote identical features or at least functionally equivalent features, unless expressly stated otherwise. DETAILED DESCRIPTION

[0129] In the following description, reference is made to the accompanying drawings that form a part hereof and illustrate, by way of illustration, specific aspects of embodiments of the invention or in which embodiments of the invention may be used. It is understood that embodiments of the invention may be used in other aspects and may include structural or logical variations not depicted in the accompanying drawings. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of the invention is defined by the appended claims.

[0130] For example, it should be understood that the disclosure relating to describing a method may also apply to a corresponding device or system for performing the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units (e.g., functional units) to perform the described one or more method steps (e.g., one unit performs one or more steps, or multiple units perform one or more of the multiple steps respectively), even if such one or more units are not explicitly described or illustrated in the accompanying drawings. On the other hand, for example, if a specific device is described in terms of one or more units (e.g., functional units), the corresponding method may include a step to perform the function of one or more units (e.g., one step performs the function of one or more units, or multiple steps perform the function of one or more of the multiple units respectively), even if such one or more steps are not explicitly described or illustrated in the accompanying drawings. In addition, it should be understood that, unless otherwise expressly stated, the features of the various exemplary embodiments and / or aspects described herein may be combined with each other.

[0131] Video decoding generally refers to processing a series of images that make up a video or video sequence. In the field of video decoding, the terms "frame" and "picture / image" can be used as synonyms. Video decoding (or generally referred to as decoding) includes two parts: video encoding and video decoding. Video encoding is performed on the source side and generally includes processing (for example, by compression) the original video image to reduce the amount of data required to represent the video image (thereby making it more efficient to store and / or transmit). Video decoding is performed on the destination side and generally includes inverse processing relative to the encoder to reconstruct the video image. The "decoding" of the video image (or generally referred to as image) involved in the embodiment should be understood as the "encoding" or "decoding" of the video image or the corresponding video sequence. The encoding part and the decoding part are also collectively referred to as codec (CODEC) (encoding and decoding).

[0132] In the case of lossless video decoding, the original video image can be reconstructed, that is, the reconstructed video image has the same quality as the original video image (assuming there is no transmission loss or other data loss during storage or transmission). In the case of lossy video decoding, further compression is performed through quantization, etc. to reduce the amount of data representing the video image, but the decoder side cannot fully reconstruct the video image, that is, the quality of the reconstructed video image is lower or worse than the quality of the original video image.

[0133] Several video coding standards belong to the group of "lossy hybrid video codecs" (i.e., combining spatial and temporal prediction in the sample domain with 2D transform coding using quantization in the transform domain). Each image in a video sequence is typically partitioned into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, on the encoder side, video is typically processed (i.e., encoded) at the block (video block) level. For example, a prediction block is generated through spatial (intra-frame) prediction and / or temporal (inter-frame) prediction; the prediction block is subtracted from the current block (currently processed block / to-be-processed block) to produce a residual block; the residual block is transformed and quantized in the transform domain to reduce the amount of data to be transmitted (compressed). On the decoder side, the coded or compressed block is processed inversely to the encoder to reconstruct the current block for representation. Furthermore, the encoder and decoder share the same processing steps, so that the encoder and decoder generate the same prediction blocks (e.g., intra-frame and inter-frame prediction blocks) and / or reconstructed blocks for processing (i.e., decoding) subsequent blocks.

[0134] In the following embodiment of the video decoding system 10, the video encoder 20 and the video decoder 30 are configured according to FIG. Figure 3 Provide a description.

[0135] Figure 1A FIG1 is a schematic block diagram of an exemplary decoding system 10, such as a video decoding system 10 (or simply decoding system 10) that can utilize the techniques of the present application. The video encoder 20 (or simply encoder 20) and the video decoder 30 (or simply decoder 30) in the video decoding system 10 are two examples of devices that can implement the techniques using various examples described in this application.

[0136] like Figure 1A As shown, the decoding system 10 includes a source device 12 for providing encoded image data 21 to a destination device 14 or the like for decoding the encoded image data 21 .

[0137] The source device 12 includes an encoder 20 and may additionally (ie, optionally) include an image source 16 , a pre-processor (or pre-processing unit) 18 (eg, image pre-processor 18 ), and a communication interface or communication unit 22 .

[0138] Image source 16 may include or may be any type of image capture device for capturing real-world images, etc.; and / or any type of image generation device (e.g., a computer graphics processor for generating computer-animated images); or any type of device for acquiring and / or providing real-world images, computer-animated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory (memory / storage) for storing any of the above images.

[0139] In order to distinguish the pre-processor 18 and the processing performed by the pre-processing unit 18 , the image or image data 17 may also be referred to as a raw image or raw image data 17 .

[0140] The preprocessor 18 is configured to receive (raw) image data 17 and perform preprocessing on the image data 17 to obtain a preprocessed image 19 or preprocessed image data 19. The preprocessing performed by the preprocessor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color grading, or denoising. It will be appreciated that the preprocessing unit 18 may be an optional component.

[0141] The video encoder 20 is configured to receive the pre-processed image data 19 and provide encoded image data 21 (combined with Figure 2 etc. for more details).

[0142] The communication interface 22 in the source device 12 can be used to receive the encoded image data 21 and send the encoded image data 21 (or data obtained after further processing of the encoded image data 21) to another device (such as the destination device 14) or any other device through the communication channel 13 for storage or direct reconstruction.

[0143] Destination device 14 includes a decoder 30 (eg, video decoder 30 ), and may additionally (ie, optionally) include a communication interface or communication unit 28 , a post-processor 32 (or post-processing unit 32 ), and a display device 34 .

[0144] The communication interface 28 in the destination device 14 is used to receive the encoded image data 21 (or data obtained by further processing the encoded image data 21), for example, directly from the source device 12 or from any other source such as a storage device (such as an encoded image data storage device), and provide the encoded image data 21 to the decoder 30.

[0145] Communication interface 22 and communication interface 28 may be used to send or receive encoded image data 21 or encoded data 13 via a direct communication link (e.g., a direct wired or wireless connection) between source device 12 and destination device 14, or via any type of network (e.g., a wired network, a wireless network, or any combination thereof, or any type of private and public network, or any combination thereof).

[0146] For example, the communication interface 22 may be used to encapsulate the encoded image data 21 into a suitable format (eg, data packets) and / or process the encoded image data through any type of transmission encoding or processing for transmission over a communication link or network.

[0147] For example, the communication interface 28 corresponding to the communication interface 22 may be configured to receive transmission data and process the transmission data through any type of corresponding transmission decoding or processing and / or decapsulation to obtain the encoded image data 21 .

[0148] Both the communication interface 22 and the communication interface 28 can be configured as Figure 1A The unidirectional communication interface indicated by the arrow of the communication channel 13 pointing from the source device 12 to the destination device 14, or configured as a bidirectional communication interface, can be used to send and receive messages, etc. to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission (such as encoded image data transmission), etc.

[0149] The decoder 30 is configured to receive the encoded image data 21 and provide decoded image data 31 or decoded image 31 (hereinafter referred to as Figure 3 or Figure 5 etc. for more details).

[0150] The post-processor 32 in the destination device 14 is configured to post-process the decoded image data 31 (also referred to as reconstructed image data) (e.g., the decoded image 31) to obtain post-processed image data 33 (e.g., the post-processed image 33). The post-processing performed by the post-processing unit 32 may include color format conversion (e.g., from YCbCr to RGB), color grading, trimming, resampling, or any other processing to provide the decoded image data 31 for display by a display device 34, etc.

[0151] The display device 34 in the destination device 14 is configured to receive the post-processed image data 33 for displaying the image to a user, viewer, or the like. The display device 34 may be or include any type of display (e.g., an integrated or external display or screen) for displaying the reconstructed image. For example, the display may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display.

[0152] although Figure 1A Source device 12 and destination device 14 are shown as separate devices, but in embodiments, a device may also include both source device 12 and destination device 14 or the functionality of both source device 12 and destination device 14, i.e., source device 12 or the corresponding functionality and destination device 14 or the corresponding functionality. In these embodiments, source device 12 or the corresponding functionality and destination device 14 or the corresponding functionality may be implemented using the same hardware and / or software or through separate hardware and / or software, or any combination thereof.

[0153] According to the description, Figure 1A The presence and (precise) division of the different units or functionalities in the source device 12 and / or destination device 14 shown may vary depending on the actual device and application, as will be apparent to the skilled person.

[0154] The encoder 20 (eg, video encoder 20) or the decoder 30 (eg, video decoder 30) or both the encoder 20 and the decoder 30 may be configured to generate a video signal. Figure 1B The encoder 20 may be implemented by the processing circuit 46, such as one or more microprocessors, one or more digital signal processors (DSPs), one or more application-specific integrated circuits (ASICs), one or more field-programmable gate arrays (FPGAs), one or more discrete logic, one or more video decoding dedicated processors, or any combination thereof. Figure 2 The various modules described in the encoder 20 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented by the processing circuit 46 to include reference to Figure 3 The various modules described in the decoder 30 and / or any other decoder system or subsystem described herein. The processing circuitry may be used to perform the various operations described below. Figure 5 As shown, if the above technology is partially implemented in software, a device can store the instructions of the software in a suitable non-transitory computer-readable medium and use one or more processors to execute these instructions in hardware to perform the technology of the present invention. The video encoder 20 or the video decoder 30 can be integrated into a single device as part of a combined codec (CODEC), such as Figure 1B shown.

[0155] Source device 12 and destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as a notebook / laptop computer, a mobile phone, a smartphone, a tablet or a tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content distribution server), a broadcast receiver device, a broadcast transmitter device, etc., and may not use any type of operating system. In some cases, source device 12 and destination device 14 may be configured for wireless communication. Therefore, source device 12 and destination device 14 may be wireless communication devices.

[0156] In some cases, Figure 1A The video decoding system 10 shown is merely exemplary, and the present technology can be applied to video decoding (e.g., video encoding or video decoding) settings that do not necessarily include any data communication between the encoding device and the decoding device. In other examples, data is retrieved from local storage, streamed over a network, and so on. The video encoding device can encode the data and store the data in the memory, and / or the video decoding device can retrieve the data from the memory and decode the data. In some examples, the encoding and decoding are performed by devices that do not communicate with each other and only encode the data to the memory and / or retrieve the data from the memory and decode the data.

[0157] For ease of description, this document (for example) refers to the High-Efficiency Video Coding (HEVC) or the next-generation video coding standard Versatile Video Coding (VVC) reference software developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG) to describe embodiments of the present invention. Those skilled in the art will appreciate that embodiments of the present invention are not limited to HEVC or VVC.

[0158] Encoders and encoding methods

[0159] Figure 2 FIG. 2 is a schematic block diagram of an exemplary video encoder 20 for implementing the technology of the present application. Figure 2 In the example of FIG, the video encoder 20 includes an input terminal 201 (or input interface 201), a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output terminal 272 (or output interface 272). The mode selection unit 260 may include an inter-frame prediction unit 244, an intra-frame prediction unit 254, and a segmentation unit 262. The inter-frame prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The illustrated video encoder 20 may also be referred to as a hybrid video encoder or a hybrid video codec-based video encoder.

[0160] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 may constitute a forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-frame prediction unit 244, and the intra-frame prediction unit 254 may constitute a backward signal path of the video encoder 20, wherein the backward signal path of the video encoder 20 corresponds to the decoder (see Figure 3The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter-frame prediction unit 244, and the intra-frame prediction unit 254 also constitute the "built-in decoder" of the video encoder 20.

[0161] Images and image segmentation (images and blocks)

[0162] Encoder 20 can be configured to receive an image 17 (or image data 17) via input 201 or the like. Image 17 can be an image from a series of images comprising a video or video sequence. The received image or image data can also be a pre-processed image 19 (or pre-processed image data 19). For simplicity, the following description uses image 17. Image 17 can also be referred to as a current image or a picture to be decoded (particularly in video decoding to distinguish the current image from other images (e.g., previously encoded and / or decoded images) in the same video sequence (i.e., a video sequence that also includes the current image)).

[0163] A (digital) image is, or can be considered to be, a two-dimensional array or matrix consisting of samples with intensity values. The samples in the array are also referred to as pixels (or pels, short for picture elements). The number of samples in the horizontal and vertical directions (or axes) of the array or image defines the image's size and / or resolution. To represent color, three color components are typically used; that is, an image can be represented as or include three arrays of samples. In the RGB format or color space, an image consists of corresponding arrays of red, green, and blue samples. However, in video decoding, each pixel is typically represented in a luminance and chrominance format or color space, such as YCbCr, consisting of a luminance component, represented by Y (sometimes also represented by L), and two chrominance components, represented by Cb and Cr. The luminance (luma) component Y represents brightness or grayscale intensity (for example, in grayscale images, both are the same), while the two chrominance (chroma) components, Cb and Cr, represent chrominance or color information components. Thus, an image in YCbCr format includes a luma sample array consisting of luma sample values ​​(Y) and two chroma sample arrays consisting of chroma values ​​(Cb and Cr). An image in RGB format can be converted or transformed to YCbCr format, and vice versa. This process is also called color conversion or transformation. If the image is black and white, the image can include only a luma sample array. Accordingly, the image can be, for example, a luma sample array in black and white format or a luma sample array and two corresponding chroma sample arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0164] In an embodiment, the video encoder 20 may include an image segmentation unit ( Figure 2 ), is used to partition the image 17 into a plurality of (typically non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTBs) or coding tree units (CTUs) (H.265 / HEVC and VVC). The image partitioning unit may be used to use the same block size for all images of a video sequence and a corresponding grid of defined block sizes, or to vary the block size between images or subsets or groups of images and to partition each image into a plurality of corresponding blocks.

[0165] In other embodiments, the video encoder may be configured to directly receive a block 203 of the image 17, such as one, several or all blocks constituting the image 17. The image block 203 may also be referred to as a current image block or an image block to be decoded.

[0166] Similar to image 17, image block 203 is also or can be considered as a two-dimensional array or matrix consisting of samples having intensity values ​​(sample values), but the size of image block 203 is smaller than that of image 17. In other words, depending on the color format used, block 203 may include, for example, one sample array (e.g., a luma array in the case of black and white image 17, or a luma array or chroma array in the case of a color image), three sample arrays (e.g., one luma array and two chroma arrays in the case of a color image 17), or any other number and / or type of arrays. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Accordingly, a block may be an M×N (M columns×N rows) sample array, or an M×N transform coefficient array, etc.

[0167] In an embodiment, Figure 2 The illustrated video encoder 20 may be configured to encode the image 17 block by block, eg, performing encoding and prediction on each block 203 .

[0168] In an embodiment, Figure 2 The illustrated video encoder 20 may also be configured to partition and / or encode a picture using slices (also referred to as video slices). A picture may be partitioned into or encoded using one or more (usually non-overlapping) slices, each of which may include one or more blocks (e.g., CTUs).

[0169] In an embodiment, Figure 2The video encoder 20 shown can also be used to segment and / or encode an image using tile groups (also called video tile groups) and / or blocks (also called video blocks). An image can be segmented into one or more tile groups (usually non-overlapping) or encoded using one or more tile groups (usually non-overlapping); each tile group can include one or more blocks (e.g., CTUs) or one or more tiles, etc.; each tile can be rectangular, etc., and can include one or more complete or partial blocks (e.g., CTUs), etc.

[0170] Residual calculation

[0171] The residual calculation unit 204 can be used to calculate the residual block 205 (also referred to as residual 205) based on the image block 203 and the prediction block 265 (the prediction block 265 is described in detail later) in the following manner to obtain the residual block 205 in the sample domain: for example, subtracting the sample value of the prediction block 265 from the sample value of the image block 203 sample by sample (pixel by pixel).

[0172] Transform

[0173] The transform processing unit 206 may be configured to perform a transform such as a discrete cosine transform (DCT) or a discrete sine transform (DST) on the sample values ​​of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207 may also be referred to as transform residual coefficients, representing the residual block 205 in the transform domain.

[0174] The transform processing unit 206 can be used to perform an integer approximation of DCT / DST (e.g., a transform specified for H.265 / HEVC). Compared to the orthogonal DCT transform, this integer approximation is usually scaled by a certain factor. In order to maintain the norm of the residual block after the forward transform and inverse transform, other scaling factors are used as part of the transform process. The scaling factor is usually selected based on certain constraints, such as whether the scaling factor is a power of 2 for the shift operation, the bit depth of the transform coefficients, the trade-off between accuracy and implementation cost, etc. For example, a specific scaling factor is specified for the inverse transform by the inverse transform processing unit 212, etc. (and on the video decoder 30 side, the corresponding inverse transform by the inverse transform processing unit 312, etc.); accordingly, on the encoder 20 side, the corresponding scaling factor can be specified for the forward transform by the transform processing unit 206, etc.

[0175] In an embodiment, the video encoder 20 (correspondingly, the transform processing unit 206) can be used to output transform parameters such as one or more transform types, for example, directly output or output after encoding or compression by the entropy coding unit 270, so that (for example) the video decoder 30 can receive and use the transform parameters for decoding.

[0176] Quantification

[0177] The quantization unit 208 may be configured to quantize the transform coefficients 207 by performing scalar quantization or vector quantization to obtain quantized coefficients 209 . The quantized coefficients 209 may also be referred to as quantized transform coefficients 209 or quantized residual coefficients 209 .

[0178] The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, during quantization, an n-bit transform coefficient can be rounded down to an m-bit transform coefficient, where n is greater than m. The degree of quantization can be modified by adjusting a quantization parameter (QP). For example, for scalar quantization, varying degrees of scaling can be used to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. The appropriate quantization step size can be represented by a quantization parameter (QP). For example, a quantization parameter can be an index into a set of predefined applicable quantization step sizes. For example, a smaller quantization parameter can correspond to fine quantization (a smaller quantization step size), while a larger quantization parameter can correspond to coarse quantization (a larger quantization step size), and vice versa. Quantization can include division by the quantization step size, while the corresponding or inverse quantization performed by the inverse quantization unit 210, etc., can include multiplication by the quantization step size. Embodiments according to some standards, such as HEVC, can use the quantization parameter to determine the quantization step size. Generally, the quantization step size can be calculated using a fixed-point approximation of an equation involving division based on the quantization parameter. Other scaling factors can be introduced for quantization and dequantization to restore the norm of the residual block that may be modified by the scaling used in the fixed-point approximation of the equations for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used, which is signaled by the encoder to the decoder via the bitstream or other means. Quantization is a lossy operation, and the larger the quantization step size, the greater the loss.

[0179] In an embodiment, the video encoder 20 (correspondingly, the quantization unit 208) can be used to output a quantization parameter (QP), for example, directly output or output after being encoded by the entropy coding unit 270, so that (for example) the video decoder 30 can receive and use the quantization parameter for decoding.

[0180] Dequantization

[0181] The inverse quantization unit 210 is configured to inversely quantize the quantized coefficients as performed by the quantization unit 208 to obtain dequantized coefficients 211, for example, by performing an inverse quantization scheme of the quantization scheme performed by the quantization unit 208 based on or using the same quantization step size as the quantization unit 208. The dequantized coefficients 211, which may also be referred to as dequantized residual coefficients 211, correspond to the transform coefficients 207, but are generally different from the transform coefficients due to losses caused by quantization.

[0182] Inverse transform

[0183] The inverse transform processing unit 212 is configured to perform an inverse transform of the transform performed by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as a transform block 213.

[0184] reconstruction

[0185] The reconstruction unit 214 (e.g., an adder or summer 214) is used to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the sample domain by, for example, adding the sample values ​​of the reconstructed residual block 213 and the sample values ​​of the prediction block 265 sample by sample.

[0186] Filtering

[0187] The loop filter unit 220 (or simply "loop filter" 220) is used to filter the reconstructed block 215 to obtain a filtered block 221, or generally to filter the reconstructed samples to obtain filtered samples. For example, the loop filter unit is used to smoothly perform pixel transitions or otherwise improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, a collaborative filter, or any combination thereof. Although the loop filter unit 220 is in Figure 2 2. In FIG. 2 , the loop filter unit 220 is shown as an in-loop filter, but in other configurations, the loop filter unit 220 may be implemented as a post-loop filter. The filtered block 221 may also be referred to as a filtered reconstruction block 221.

[0188] In an embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) can be used to output loop filter parameters (e.g., sample adaptive offset information), for example, directly output or output after being encoded by the entropy coding unit 270, so that (for example) the decoder 30 can receive and use the same loop filter parameters or the corresponding loop filter for decoding.

[0189] Decoded image buffer

[0190] The decoded picture buffer (DPB) 230 may be a memory that stores reference images or generally stores reference image data for use by the video encoder 20 when encoding video data. The DPB 230 may be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be used to store one or more filtered blocks 221. The decoded picture buffer 230 may also be used to store other previously filtered blocks (e.g., previously filtered reconstructed blocks 221) from the same current image or a different image (e.g., a previously reconstructed image), and may provide a previously fully reconstructed (i.e., decoded) image (and corresponding reference blocks and samples) and / or a partially reconstructed current image (and corresponding reference blocks and samples) for inter-frame prediction, etc. If the reconstructed block 215 is not filtered by the loop filter unit 220, the decoded picture buffer (DPB) 230 can also be used to store one or more unfiltered reconstructed blocks 215, or generally to store unfiltered reconstructed samples, or reconstructed blocks or reconstructed samples that have not undergone any other processing.

[0191] Mode selection (segmentation and prediction)

[0192] The mode selection unit 260 includes a segmentation unit 262, an inter-frame prediction unit 244, and an intra-frame prediction unit 254, and is configured to receive or obtain original image data such as the original block 203 (current block 203 of the current image 17) and reconstructed image data (e.g., filtered and / or unfiltered reconstructed samples or reconstructed blocks of the same (current) image and / or one or more previously decoded images) from the decoded image buffer 230 or other buffers (e.g., line buffers, not shown). The reconstructed image data is used as reference image data required for prediction such as inter-frame prediction or intra-frame prediction to obtain a prediction block 265 or a prediction value 265.

[0193] The mode selection unit 260 can be used to determine or select a segmentation method for the current block prediction mode (including non-segmentation) and determine or select a prediction mode (such as an intra-frame prediction mode or an inter-frame prediction mode), generate a corresponding prediction block 265, and calculate the residual block 205 and reconstruct the reconstruction block 215.

[0194] In an embodiment, the mode selection unit 260 can be used to select a partitioning method and a prediction mode (for example, from prediction modes supported or available by the mode selection unit 260), wherein the prediction mode provides the best match or minimum residual (minimum residual means better compression in transmission or storage), or provides the minimum indicated overhead (minimum indicated overhead means better compression in transmission or storage), or considers or balances both of the above. The mode selection unit 260 can be used to determine the partitioning method and the prediction mode based on rate distortion optimization (RDO), that is, select the prediction mode that provides the minimum rate distortion. The terms "best", "minimum", "optimal", etc. in this document do not necessarily mean "best", "minimum", "optimal", etc. in general, but can also refer to situations where termination or selection criteria are met. For example, values ​​exceeding or falling below a threshold or other constraints may result in a "suboptimal selection" but reduce complexity and processing time.

[0195] In other words, the partitioning unit 262 can be used to partition the block 203 into smaller partitions or sub-blocks (forming blocks again) in the following manner: for example, by iteratively using quad-tree (QT) partitioning, binary-tree (BT) partitioning or triple-tree (TT) partitioning or any combination thereof, and for performing prediction on each of the partitions or sub-blocks, etc., wherein the mode selection includes selecting a tree structure of the partition block 203 and using a prediction mode for each of the block partitions or sub-blocks.

[0196] The segmentation (eg, performed by segmentation unit 260) and prediction processes (performed by inter-prediction unit 244 and intra-prediction unit 254) performed by exemplary video encoder 20 are described in detail below.

[0197] segmentation

[0198] The segmentation unit 262 can segment (or divide) the current block 203 into smaller segments, such as smaller blocks of square or rectangular size. These smaller blocks (also referred to as sub-blocks) can be further segmented into even smaller segments. This is also known as tree segmentation or hierarchical tree segmentation. A root block at root tree level 0 (hierarchical level 0, depth 0) can be recursively segmented into two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchical level 1, depth 1). These blocks can be further segmented into two or more blocks at the next lower level, such as tree level 2 (hierarchical level 2, depth 2), and so on, until the segmentation is completed (because the end criteria are met, such as reaching the maximum tree depth or minimum block size). Blocks that are not further segmented are also referred to as leaf blocks or leaf nodes of the tree. A tree segmented into two segments is called a binary tree (BT), a tree segmented into three segments is called a ternary tree (TT), and a tree segmented into four segments is called a quad tree (QT).

[0199] As described above, the term "block" used herein may be a portion of an image, in particular a square or rectangular portion. With reference to HEVC and VVC, etc., a block may be or may correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or correspond to multiple pairs of blocks, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).

[0200] For example, a coding tree unit (CTU) can be or include a CTB consisting of luma samples and two corresponding CTBs consisting of chroma samples from an image with three sample arrays, or a CTB consisting of samples from a monochrome image or an image coded using three separate color planes and syntax structures. These syntax structures are used to decode the samples. Accordingly, a coding tree block (CTB) can be an N×N block of samples, where N can be set to a value such that a component is divided into multiple CTBs. This is one type of partitioning. A coding unit (CU) can be or include a coding block consisting of luma samples and two corresponding coding blocks consisting of chroma samples from an image with three sample arrays, or a coding block consisting of samples from a monochrome image or an image coded using three separate color planes and syntax structures. These syntax structures are used to decode the samples. Accordingly, a coding block (CB) can be an M×N block of samples, where M and N can be set to values ​​such that a CTB is divided into multiple coding blocks. This is one type of partitioning.

[0201] In an embodiment, for example according to HEVC, a coding tree unit (CTU) may be divided into multiple CUs using a quadtree structure represented as a coding tree. It is decided at the CU level whether to use inter (temporal) prediction or intra (spatial) prediction to decode an image region. Each CU may be further divided into 1, 2, or 4 PUs according to the PU partition type. The same prediction process is performed within a PU, and relevant information is sent to the decoder in units of PUs. After the residual block is obtained by performing the prediction process according to the PU partition type, the CU may be partitioned into transform units (TUs) according to other quadtree structures similar to the coding tree for the CU.

[0202] In an embodiment, for example, according to the latest video coding standard currently under development called Versatile Video Coding (VVC), a quad-tree combined with a binary tree (quad-tree and binary-tree, QTBT) segmentation is used to segment the coding block. In the QTBT block structure, a CU can be square or rectangular. For example, the coding tree unit (CTU) is first segmented by a quadtree structure. The quadtree leaf nodes are further segmented by a binary tree or a ternary / triple tree structure. The segmented leaf nodes are called coding units (CUs), and this segmentation is used for prediction and transform processing without any further segmentation. This means that in the QTBT coding block structure, the block sizes of CU, PU, ​​and TU are the same. At the same time, multiple segmentations such as ternary tree segmentation can be used with the QTBT block structure.

[0203] In one example, mode select unit 260 in video encoder 20 may be used to perform any combination of the segmentation techniques described herein.

[0204] As described above, the video encoder 20 is configured to determine or select the best or optimal prediction mode from a (predetermined) prediction mode set. The prediction mode set may include intra-frame prediction mode and / or inter-frame prediction mode, etc.

[0205] Intra-frame prediction

[0206] The intra-frame prediction mode set may include 35 different intra-frame prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in HEVC, or may include 67 different intra-frame prediction modes, such as non-directional modes like DC (or mean) mode and planar mode, or directional modes as defined in VVC.

[0207] The intra prediction unit 254 is configured to generate an intra prediction block 265 using reconstructed samples of neighboring blocks of the same current image according to an intra prediction mode in the intra prediction mode set.

[0208] The intra-frame prediction unit 254 (or generally referred to as the mode selection unit 260) is also used to output intra-frame prediction parameters (or generally referred to as information representing the selected intra-frame prediction mode of the block) to the entropy coding unit 270 in the form of syntax elements 266 to be included in the encoded image data 21, so that (for example) the video decoder 30 can receive and use the prediction parameters for decoding.

[0209] Inter-frame prediction

[0210] The set of (possible) inter-frame prediction modes depends on the available reference picture (i.e. (for example) at least part of the decoded picture stored in the DPB 230 as mentioned above) and other inter-frame prediction parameters, for example on whether the entire reference picture or only a part of the reference picture (e.g. a search window area around the area of ​​the current block) is used to search for the best matching reference block, and / or for example on whether pixel interpolation is performed (e.g. half / half pixel interpolation and / or quarter pixel interpolation).

[0211] In addition to the above-mentioned prediction modes, skip mode and / or direct mode may also be used.

[0212] The inter-frame prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both in Figure 2 ). The motion estimation unit may be configured to receive or acquire an image block 203 (the current image block 203 of the current image 17) and a decoded image 231, or at least one or more previously reconstructed blocks (e.g., reconstructed blocks of one or more other / different previously decoded images 231), for motion estimation. For example, a video sequence may include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 may be part of or constitute a series of images, the series of images constituting the video sequence.

[0213] For example, the encoder 20 may be configured to select a reference block from a plurality of reference blocks of the same or different images in a plurality of other images, and provide the reference image (or reference image index) and / or the offset (spatial offset) between the position (x-coordinate, y-coordinate) of the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. This offset is also referred to as a motion vector (MV).

[0214] The motion compensation unit is configured to obtain (e.g., receive) inter-frame prediction parameters and perform inter-frame prediction based on or using the inter-frame prediction parameters to obtain an inter-frame prediction block 265. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction block based on a motion / block vector determined by motion estimation, and may also include performing interpolation to obtain sub-sample accuracy. Interpolation filtering can generate additional pixel samples based on known pixel samples, thereby potentially increasing the number of candidate prediction blocks that can be used to decode the image block. Upon receiving the motion vector corresponding to the PU of the current image block, the motion compensation unit may locate the prediction block pointed to by the motion vector in one of the reference picture lists.

[0215] The motion compensation unit may also generate syntax elements associated with blocks and video slices for use by video decoder 30 when decoding image blocks of the video slices. In addition to or as an alternative to slices and corresponding syntax elements, tile groups and / or tiles and corresponding syntax elements may also be generated or used.

[0216] Entropy Coding

[0217] The entropy coding unit 270 is configured to apply or not apply an entropy coding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC (CAVLC) scheme, an arithmetic coding scheme, binarization, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the (uncompressed) quantized coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements, resulting in coded image data 21 that can be output via an output terminal 272 in the form of a coded bitstream 21, for example, so that the video decoder 30 can receive and use these parameters for decoding. The coded bitstream 21 can be sent to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.

[0218] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform based encoder 20 may directly quantize the residual signal for certain blocks or frames without the transform processing unit 206. In another implementation, the encoder 20 may include the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.

[0219] Decoder and decoding method

[0220] Figure 3An example of a video decoder 30 for implementing the technology of the present application is shown. The video decoder 30 is configured to receive coded image data 21 (e.g., coded codestream 21) coded by an encoder 20, for example, and obtain a decoded image 331. The coded image data or codestream includes information for decoding the coded image data, such as data representing image blocks of coded video slices (and / or partition groups or partitions) and related syntax elements.

[0221] exist Figure 3 In the example of FIG. 3 , the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., a summer 314), a loop filter 320, a decoded picture buffer (DPB) 330, a mode application unit 360, an inter-frame prediction unit 344, and an intra-frame prediction unit 354. The inter-frame prediction unit 344 may be or may include a motion compensation unit. In some examples, the video decoder 30 may perform substantially the same motion compensation as the reference video. Figure 2 The encoding pass described in the video encoder 100 is the inverse of the decoding pass.

[0222] As described with reference to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter-frame prediction unit 344, and intra-frame prediction unit 354 also constitute the "built-in decoder" of video encoder 20. Accordingly, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 212, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Therefore, the explanations of the corresponding units and functions of video encoder 20 apply accordingly to the corresponding units and functions of video decoder 30.

[0223] Entropy decoding

[0224] The entropy decoding unit 304 is used to parse the code stream 21 (or generally referred to as the encoded image data 21) and perform entropy decoding on the encoded image data 21 to obtain quantization coefficients 309 and / or decoded encoding parameters ( Figure 3, such as any or all of inter-frame prediction parameters (e.g., reference picture index and motion vector), intra-frame prediction parameters (e.g., intra-frame prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 can be used to apply a decoding algorithm or scheme corresponding to the encoding scheme described with reference to the entropy encoding unit 270 in the encoder 20. The entropy decoding unit 304 can also be used to provide the inter-frame prediction parameters, intra-frame prediction parameters, and / or other syntax elements to the mode application unit 360 and to provide other parameters to other units in the decoder 30. The video decoder 30 can receive syntax elements at the video slice level and / or the video block level. In addition to or as an alternative to slices and corresponding syntax elements, partition groups and / or partitions and corresponding syntax elements can also be received and / or used.

[0225] Dequantization

[0226] The inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally, information related to inverse quantization) and quantization coefficients from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse quantize the decoded quantization coefficients 309 according to the quantization parameters to obtain dequantized coefficients 311. The dequantized coefficients 311 may also be referred to as transform coefficients 311. The inverse quantization process may include using the quantization parameter determined by the video encoder 20 for each video block in a video slice (or tile or group of tiles) to determine a degree of quantization, and thus a degree of inverse quantization to be performed.

[0227] Inverse transform

[0228] The inverse transform processing unit 312 can be configured to receive the dequantized coefficients 311 (also referred to as transform coefficients 311) and transform the dequantized coefficients 311 to obtain a reconstructed residual block 213 in the sample domain. The reconstructed residual block 213 can also be referred to as a transform block 313. The transform can be an inverse transform, such as an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 can also be configured to receive transform parameters or corresponding information from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304, etc.) to determine the transform to be performed on the dequantized coefficients 311.

[0229] reconstruction

[0230] The reconstruction unit 314 (e.g., adder or summer 314) can be used to add the reconstructed residual block 313 to the prediction block 365 to obtain the reconstructed block 315 in the sample domain by, for example, adding the sample values ​​of the reconstructed residual block 313 and the sample values ​​of the prediction block 365.

[0231] Filtering

[0232] The loop filter unit 320 (in the decoding loop or after) is used to filter the reconstructed block 315 to obtain a filtered block 321, so as to smoothly perform pixel conversion or improve video quality in other ways. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, a collaborative filter, or any combination thereof. Although the loop filter unit 320 is in Figure 3 3. Although shown as an in-loop filter in FIG. 3, in other configurations, the loop filter unit 320 may be implemented as a post-loop filter.

[0233] Decoded image buffer

[0234] The decoded video block 321 of one picture is then stored in a decoded picture buffer 330 , which stores the decoded picture 331 as a reference picture for subsequent motion compensation and / or output or display of other pictures.

[0235] The decoder 30 is configured to output the decoded image 311 through an output terminal 312 or the like, for display to a user or for viewing by the user.

[0236] predict

[0237] The inter-frame prediction unit 344 can be functionally the same as the inter-frame prediction unit 244 (particularly the motion compensation unit), and the intra-frame prediction unit 354 can be functionally the same as the intra-frame prediction unit 254, and can determine the partitioning or segmentation and perform prediction based on the segmentation and / or prediction parameters or corresponding information received from the coded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304, etc.). The mode application unit 360 can be used to perform prediction (intra-frame prediction or inter-frame prediction) on each block based on the reconstructed image, block or corresponding samples (filtered or unfiltered) to obtain a prediction block 365.

[0238] When the video slice is coded as an intra-coded (I) slice, the intra-prediction unit 354 in the mode application unit 360 is configured to generate a prediction block 365 for the image block of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks in the current image. When the video image is coded as an inter-coded (i.e., B or P) slice, the inter-prediction unit 344 (e.g., a motion compensation unit) in the mode application unit 360 is configured to generate a prediction block 365 for the video block of the current video slice based on the motion vector and other syntax elements received from the entropy decoding unit 304. For inter-prediction, these prediction blocks can be generated based on one of the reference pictures in one of the reference picture lists. Video decoder 30 can construct reference frame list 0 and list 1 using a default construction technique based on the reference pictures stored in DPB 330. The same or similar processes may be applied to or by embodiments that use partition groups (e.g., video partition groups) and / or partitions (e.g., video partitions) in addition to or instead of slices (e.g., video slices), e.g., video may be coded using I, P, or B partition groups and / or partitions.

[0239] Mode application unit 360 is configured to determine prediction information for video blocks in the current video slice by parsing motion vectors or related information and other syntax elements, and use the prediction information to generate a prediction block for the current video block being decoded. For example, mode application unit 360 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction), an inter prediction slice type (e.g., B slice, P slice, or GPB slice), construction information for one or more reference picture lists for the slice, motion vectors for each inter-coded video block in the slice, inter prediction status for each inter-coded video block in the slice, and other information to decode the video blocks in the current video slice. In addition to slices (e.g., video slices), the same or similar processes can be applied to partition groups (e.g., video partition groups) and / or partitions (e.g., video partitions). For example, video can be coded using I, P, or B partition groups and / or partitions.

[0240] In an embodiment, Figure 3 The video decoder 30 shown may be configured to partition and / or decode a picture using slices (also referred to as video slices). A picture may be partitioned into or decoded using one or more (usually non-overlapping) slices, each of which may include one or more blocks (e.g., CTUs).

[0241] In an embodiment, Figure 3The video decoder 30 shown can be used to segment and / or decode an image using partition groups (also called video partition groups) and / or blocks (also called video blocks). An image can be segmented into one or more partition groups (usually non-overlapping) or decoded using one or more partition groups (usually non-overlapping); each partition group can include one or more blocks (e.g., CTUs) or one or more partitions, etc.; each partition can be rectangular, etc., and can include one or more complete or partial blocks (e.g., CTUs), etc.

[0242] Other variations of the video decoder 30 may be used to decode the encoded image data 21. For example, the decoder 30 may generate an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 may directly dequantize the residual signal for certain blocks or frames without the inverse transform processing unit 312. In another implementation, the video decoder 30 may include the dequantization unit 310 and the inverse transform processing unit 312 combined into a single unit.

[0243] It should be understood that the processing result of the current step may be further processed in the encoder 20 and the decoder 30 before being output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, the processing result of the interpolation filtering, motion vector derivation, or loop filtering may be further processed, such as clipping or shifting.

[0244] It should be noted that further operations can be performed on the derived motion vector of the current block (including but not limited to control point motion vectors in affine mode, sub-block motion vectors in affine mode, planar mode, and ATMVP mode, and temporal motion vectors). For example, the motion vector value can be constrained to a predefined range based on the representation bits of the motion vector. If the representation bits of the motion vector are bitDepth, the range is –2^(bitDepth–1) to 2^(bitDepth–1)–1, where “^” represents a power. For example, if bitDepth is set to 16, the range is –32768 to 32767; if bitDepth is set to 18, the range is –131072 to 131071. For example, the values ​​of the derived motion vectors (e.g., the MVs of four 4×4 sub-blocks in an 8×8 block) can be constrained so that the maximum difference between the integer parts of the MVs of these four 4×4 sub-blocks does not exceed N pixels, for example, not more than 1 pixel.

[0245] Figure 4 Schematic diagram of a video decoding device 400 provided in accordance with an embodiment of the present invention. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 400 may be a decoder (e.g., Figure 1A) or an encoder (e.g., a video decoder 30 in Figure 1A Video encoder 20 in ).

[0246] Video decoding device 400 includes: an input port 410 (or input port 410) and a receiving unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing the data; a transmitting unit (Tx) 440 and an output port 450 (or output port 450) for transmitting the data; and a memory 460 for storing the data. Video encoding device 400 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to input port 410, receiving unit 420, transmitting unit 440, and output port 450, serving as an outlet or inlet for optical or electrical signals.

[0247] Processor 430 is implemented using both hardware and software. Processor 430 can be implemented as one or more CPU chips, one or more cores (e.g., a multi-core processor), one or more FPGAs, one or more ASICs, and one or more DSPs. Processor 430 communicates with input port 410, receiving unit 420, transmitting unit 440, output port 450, and memory 460. Processor 430 includes a decoding module 470. Decoding module 470 implements the disclosed embodiments described above. For example, decoding module 470 performs, processes, prepares, or provides various decoding operations. Therefore, the inclusion of decoding module 470 provides substantial improvements to the functionality of video decoding device 400 and influences the transition of video decoding device 400 to different states. Alternatively, decoding module 470 can be implemented using instructions stored in memory 460 and executed by processor 430.

[0248] The memory 460 may include one or more magnetic disks, one or more tape drives, and one or more solid-state drives, and may be used as an overflow data storage device to store programs when they are selected for execution, as well as to store instructions and data read during the execution of the programs. For example, the memory 460 may be volatile and / or non-volatile, and may be a read-only memory (ROM), a random access memory (RAM), a ternary content-addressable memory (TCAM), and / or a static random-access memory (SRAM).

[0249] Figure 5 A simplified block diagram of an apparatus 500 is provided for an exemplary embodiment. The apparatus 500 may be used as Figure 1A Either or both of the source device 12 and the destination device 14.

[0250] The processor 502 in the apparatus 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices, now or hereafter developed, capable of manipulating or processing information. While the disclosed implementations may be implemented using a single processor, such as the processor 502 shown, using multiple processors may improve speed and efficiency.

[0251] In one implementation, the memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that are accessed by the processor 502 via a bus 512. The memory 504 may also include an operating system 508 and application programs 510. The application programs 510 include at least one program that causes the processor 502 to perform the methods described herein. For example, the application programs 510 may include application 1 through application N, and may also include a video decoding application that performs the methods described herein.

[0252] The apparatus 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element capable of sensing touch input. The display 518 may be coupled to the processor 502 via the bus 512.

[0253] Although the bus 512 of the device 500 is shown as a single bus here, there may be multiple buses 512. In addition, the auxiliary memory 514 may be directly coupled to other components in the device 500 or may be accessed via a network and may include a single integrated unit (e.g., a memory card) or multiple units (e.g., multiple memory cards). Therefore, the device 500 can have a variety of configurations.

[0254] The concepts in this application are first described below.

[0255] 1. Inter-frame prediction mode

[0256] In HEVC, two inter-frame prediction modes are used, namely advanced motion vector prediction (AMVP) mode and merge mode.

[0257] For the AMVP mode, the coded blocks that are spatially or temporally adjacent to the current block (referred to as adjacent blocks) are first traversed. A candidate motion vector list (also called a motion information candidate list) is constructed based on the motion information of each adjacent block. The optimal motion vector is then determined from the candidate motion vector list using the rate-distortion cost. The candidate motion information with the lowest rate-distortion cost is used as the motion vector predictor (MVP) for the current block. The positions of the adjacent blocks and their traversal order are predefined. The rate-distortion cost is calculated according to formula (1), where J represents the rate-distortion cost (RDcost), SAD is the sum of absolute differences (SAD) between the predicted sample values ​​obtained after motion estimation using the candidate motion vector predictor and the original sample values, R represents the bit rate, and λ represents the Lagrange multiplier. The encoder transmits the index value of the selected motion vector predictor in the candidate motion vector list and the reference frame index value to the decoder. Furthermore, a motion search is performed in the neighborhood centered on the MVP to obtain the actual motion vector of the current block. The encoder transmits the difference between the MVP and the actual motion vector (motion vector difference) to the decoder.

[0258] J=SAD+λR (1)

[0259] For the fusion mode, the motion information of the coded blocks adjacent to the current block in space or time is first used to build a candidate motion vector list. Then, the optimal motion information is determined from the candidate motion vector list by calculating the rate-distortion cost as the motion information of the current block. The index value of the position of the optimal motion information in the candidate motion vector list (hereinafter referred to as the fusion index) is passed to the decoding end. The spatial and temporal candidate motion information of the current block is as follows: Figure 6 As shown. The spatial candidate motion information comes from the five spatially adjacent blocks (A0, A1, B0, B1 and B2). If the adjacent block is unavailable (the adjacent block does not exist or the adjacent block is not encoded or the prediction mode used by the adjacent block is not the inter-frame prediction mode), the motion information of the adjacent block is not added to the candidate motion vector list. The temporal candidate motion information of the current block is obtained by scaling the MV of the corresponding position block in the reference frame according to the picture order count (POC) of the reference frame and the current frame. First, determine whether the block at position T in the reference frame is available. If not, select the block at position C.

[0260] Similar to the AMVP mode, the positions of adjacent blocks and their traversal order in the fusion mode are also predefined. In addition, the positions of adjacent blocks and their traversal order may be different in different modes.

[0261] As you can see, in both AMVP mode and fusion mode, a candidate motion vector list (also called a candidate list, or simply a candidate list) needs to be maintained. Each time new motion information is added to the candidate list, a check is performed to see if the same motion information already exists in the list. If so, the motion information will not be added to the list. This check process is called pruning the candidate motion vector list. List pruning is done to prevent the same motion information from appearing in the list, thus avoiding redundant rate-distortion cost calculations.

[0262] In HEVC inter-frame prediction, all samples within a coding block use the same motion information. Motion compensation is then performed based on this motion information to obtain the predicted values ​​of the samples within the coding block. However, not all samples within a coding block have the same motion characteristics. Using the same motion information across coding blocks can lead to inaccurate motion compensation predictions, which in turn increases residual information.

[0263] Existing video coding standards use block matching motion estimation based on translational motion models and assume that all samples in a block have consistent motion. However, in the real world, motion is diverse, and many objects experience non-translational motion, such as rotating objects, roller coasters spinning in different directions, fireworks, and stunts in movies. This is especially true for moving objects in user-generated content (UGC) scenes. Decoding these objects using block motion compensation techniques based on translational motion models, as used in current coding standards, can significantly impact decoding efficiency. Therefore, non-translational motion models, such as affine motion models, have been developed to further improve decoding efficiency.

[0264] Based on this, according to the different motion models, the AMVP mode can be divided into the AMVP mode based on the translation model and the AMVP mode based on the non-translation model (such as the AMVP mode based on the affine model); the fusion mode can be divided into the fusion mode based on the translation model and the fusion mode based on the non-translation model (such as the fusion mode based on the affine model).

[0265] 2. Non-translational motion model

[0266] Non-translational motion model prediction uses the same motion model at the codec to derive motion information for each sub-block (also called a sub-motion compensation unit or basic motion compensation unit) within the current block. Motion compensation is then performed based on the sub-block's motion information to produce a predicted block, thereby improving prediction efficiency. Common non-translational motion models include the 4-parameter affine motion model and the 6-parameter affine motion model.

[0267] The sub-motion compensation unit (also referred to as a sub-block) in the embodiment of the present application can be a sample or a sample block of size N1×N2 obtained according to a specific segmentation method, where N1×N2 and N1×N2 are both positive integers, and N1×N2 can be equal to N1×N2 or not equal to N1×N2.

[0268] The 4-parameter affine motion model is shown in formula (2):

[0269]

[0270] The 4-parameter affine motion model can be represented by the motion vectors of two samples and their coordinates relative to the upper left sample of the current block. The samples used to represent the motion model parameters are called control points. If the upper left sample (0,0) and the upper right sample (W,0) are used as control points, the corresponding motion vectors (vx0,vy0) and (vx0,vy0) of the upper left and upper right control points of the current block are first determined. Then, the motion information of each sub-motion compensation unit in the current block is obtained according to formula (3), where (x,y) is the coordinate of the sub-motion compensation unit relative to the upper left sample of the current block (for example, the coordinates of the upper left sample), and W is the width of the current block. It should be understood that other control points can also be used, such as samples at positions (2,2) and (W+2,2) or (–2,–2) and (W–2,–2). The choice of control points is not limited here.

[0271]

[0272] The 6-parameter affine motion model is shown in formula (4):

[0273]

[0274] The 6-parameter affine motion model can be represented by the motion vectors of three samples and their coordinates relative to the upper left sample of the current block. If the upper left sample (0,0), upper right sample (W,0), and lower left sample (0,H) of the current block are used as control points, the motion vectors of the upper left control point, upper right control point, and lower left control point of the current block are first determined as (vx0,vy0), (vx0,vy0), and (vx0,vy0), respectively. Then, the motion information of each sub-motion compensation unit in the current block is obtained according to formula (5), where (x,y) is the coordinate of the sub-motion compensation unit relative to the upper left sample of the current block, and W and H are the width and height of the current block, respectively. It should be understood that other control points, such as samples at positions (2,2), (W+2,2), and (2,H+2), or (–2,–2), (W–2,–2), and (–2,H–2), can also be used as control points, and this is not limited here.

[0275]

[0276] The coding block predicted by the affine motion model is called the affine decoding block.

[0277] Typically, an Advanced Motion Vector Prediction (AMVP) mode based on an affine motion model or a merge mode based on an affine motion model may be used to obtain motion information of control points of an affine decoding block.

[0278] The motion information of the control points of the current coding block can be obtained by an inherited control point motion vector prediction method or a constructed control point motion vector prediction method.

[0279] 3. Inherited control point motion vector prediction method

[0280] The inherited control point motion vector prediction method refers to determining the candidate control point motion vectors of the current block by using the motion model of the adjacent encoded affine decoded blocks.

[0281] by Figure 7 Taking the current block shown as an example, the adjacent position blocks around the current block are traversed in a set order, such as A1->B1->B0->A0->B2, and the affine decoding block where the adjacent position block of the current block is located is found and the control point motion information of the affine decoding block is obtained. Then, the control point motion vector (for fusion mode) or the control point motion vector prediction value (for AMVP mode) of the current block is derived through the motion model constructed according to the control point motion information of the affine decoding block. The order A1->B1->B0->A0->B2 mentioned above is only used as an example and should not be interpreted as restrictive. Other orders can also be used. In addition, the adjacent position blocks are not limited to A1, B1, B0, A0 and B2, and various adjacent position blocks can be used.

[0282] The adjacent position block may be a sample, a sample block of a preset size obtained according to a specific segmentation method, for example, a 4×4 sample block, a 4×2 sample block, or a sample block of other sizes. These block sizes are for illustrative purposes only and should not be construed as limiting.

[0283] The following describes the determination process using A1 as an example, and other cases can be deduced in the same way.

[0284] like Figure 7As shown, if the coding block where A1 is located is a 4-parameter affine decoding block, the motion vector (vx4, vy4) of the upper left sample (x4, y4) and the motion vector (vx5, vy5) of the upper right sample (x5, y5) of the affine decoding block are obtained; the motion vector (vx0, vy0) of the upper left sample (x0, y0) of the current affine decoding block is calculated according to formula (6), and the motion vector (vx1, vy1) of the upper right sample (x1, y1) of the current affine decoding block is calculated according to formula (7).

[0285]

[0286] The combination of the motion vector (vx0, vy0) of the upper left sample (x0, y0) of the current block and the motion vector (vx1, vy1) of the upper right sample (x1, y1) obtained based on the affine decoding block where A1 is located is the candidate control point motion vector of the current block.

[0287] If the coding block where A1 is located is a 6-parameter affine decoding block, the motion vector (vx4, vy4) of the upper left sample (x4, y4), the motion vector (vx5, vy5) of the upper right sample (x5, y5), and the motion vector (vx6, vy6) of the lower left sample (x6, y6) of the affine decoding block are obtained; the motion vector (vx0, vy0) of the upper left sample (x0, y0) of the current block is calculated according to formula (8), the motion vector (vx1, vy1) of the upper right sample (x1, y1) of the current block is calculated according to formula (9), and the motion vector (vx2, vy2) of the lower left sample (x2, y2) of the current block is calculated according to formula (10).

[0288]

[0289] The combination of the motion vector (vx0, vy0) of the upper left sample (x0, y0) of the current block obtained based on the affine decoding block where A1 is located, the motion vector (vx1, vy1) of the upper right sample (x1, y1), and the motion vector (vx2, vy2) of the lower left sample (x2, y2) of the current block is the candidate control point motion vector of the current block.

[0290] It should be noted that other motion models, candidate positions, and search traversal orders may also be applicable to this application, and the embodiments of this application will not go into details about this.

[0291] It should be noted that the method of using other control points to represent the motion model of the adjacent and current coding blocks can also be applied to the present application, which will not be described in detail here.

[0292] 4. Constructed control point motion vector prediction method 1

[0293] The constructed control point motion vector prediction method refers to combining the motion vectors of adjacent coded blocks around the control point of the current block as the control point motion vector of the current affine decoded block, without considering whether the adjacent coded blocks are affine decoded blocks.

[0294] The motion information of the adjacent coded blocks around the current coded block is used to determine the motion vectors of the upper left and upper right samples of the current block. Figure 8A The constructed control point motion vector prediction method is described using as an example. It should be noted that Figure 8A This is merely an example and should not be construed as limiting.

[0295] like Figure 8A As shown, the motion vectors of the adjacent coded blocks A2, B2, and B3 of the upper left sample are used as candidate motion vectors for the motion vector of the upper left sample of the current block; the motion vectors of the adjacent coded blocks B1 and B0 of the upper right sample are used as candidate motion vectors for the motion vector of the upper right sample of the current block. The candidate motion vectors of the upper left sample and the upper right sample are combined to form multiple pairs. The motion vectors of the two coded blocks included in the pair can be used as candidate control point motion vectors for the current block, as shown in the following formula (11A):

[0296] {v A2 ,v B1},{v A2 ,v B0},{v B2 ,v B1},{v B2 ,v B0},{v B3 ,v B1},{v B3 ,v B0} (11A)

[0297] Among them, v A2 represents the motion vector of A2, v A2 represents the motion vector of B1, v A2 represents the motion vector of B0, v A2 represents the motion vector of B2, v A2 Indicates the motion vector of B3.

[0298] like Figure 8AAs shown, the motion vectors of the adjacent coded blocks A2, B2, and B3 of the upper left sample are used as candidates for the motion vector of the upper left sample of the current block; the motion vectors of the adjacent coded blocks B1 and B0 of the upper right sample are used as candidates for the motion vector of the upper right sample of the current block, and the motion vectors of the adjacent coded blocks A0 and A1 of the lower left sample are used as candidates for the motion vector of the lower left sample of the current block. The candidate motion vectors of the upper left sample, the upper right sample, and the lower left sample are combined to form a triple. The motion vectors of the three coded blocks included in the triple can be used as candidate control point motion vectors of the current block, as shown in the following formulas (11B) and (11C):

[0299] {v A2 ,v B1 ,v A0},{v A2 ,v B0 ,v A0},{v B2 ,v B1 ,v A0},{v B2 ,v B0 ,v A0},{v B3 ,v B1 ,v A0},{v B3 ,v B0 ,v A0} (11B)

[0300] {v A2 ,v B1 ,v A1},{v A2 ,v B0 ,v A1},{v B2 ,v B1 ,v A1},{v B2 ,v B0 ,v A1},{v B3 ,v B1 ,v A1},{v B3 ,v B0 ,v A1} (11C)

[0301] Among them, v A2 represents the motion vector of A2, v A2 represents the motion vector of B1, v A2 represents the motion vector of B0, v A2 represents the motion vector of B2, v A2represents the motion vector of B3, v A2 represents the motion vector of A0, v A2 Indicates the motion vector of A1.

[0302] It should be noted that other methods of combining control point motion vectors are also applicable to this application and will not be described in detail here.

[0303] It should be noted that the method of using other control points to represent the motion model of the adjacent and current coding blocks can also be applied to the present application, which will not be described in detail here.

[0304] 5. Constructed control point motion vector prediction method 2, such as Figure 8B shown.

[0305] Step 501: Obtain motion information of each control point of the current block.

[0306] by Figure 8A As shown in the example, CP k (k=1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial neighbors of the current block, used to predict CP1, CP2, or CP3; T is the temporal neighbor of the current block, used to predict CP4.

[0307] Assume that the coordinates of CP1, CP2, CP3 and CP4 are (0, 0), (W, 0), (H, 0) and (W, H), respectively, where W and H represent the width and height of the current block.

[0308] For each control point, its motion information is obtained in the following order:

[0309] (1) For CP1, the check order is B2->A2->B3. If B2 is available, CP1 uses the motion information of B2. Otherwise, A2 and B3 are checked in sequence. If the motion information of all three locations is unavailable, the motion information of CP1 cannot be obtained.

[0310] (2) For CP2, the check order is B0->B1. If B0 is available, CP2 uses the motion information of B0. Otherwise, B1 is checked. If the motion information of both locations is unavailable, the motion information of CP2 cannot be obtained.

[0311] (3) For CP3, the check order is A0->A1. If A0 is available, CP3 uses the motion information of A0. Otherwise, A1 is checked. If the motion information of both locations is unavailable, the motion information of CP3 cannot be obtained.

[0312] (4) For CP4, the motion information of T is used.

[0313] Here, X available indicates that block X (eg, A0, A1, A2, B0, B1, B2, B3, or T) has been encoded and uses inter-frame prediction mode. Otherwise, X is not available.

[0314] It should be noted that other methods for obtaining motion information of control points may also be applicable to this application and will not be described in detail here.

[0315] Step 502: Combine the motion information of each control point to obtain constructed control point motion information.

[0316] The motion information of two control points is combined into a binary tuple to construct a four-parameter affine motion model. The combinations of motion information of two control points can be {CP1, CP4}, {CP2, CP3}, {CP1, CP2}, {CP2, CP4}, {CP1, CP3}, and {CP3, CP4}. For example, a four-parameter affine motion model constructed using a binary tuple consisting of control points CP1 and CP2 can be written as Affine(CP1, CP2).

[0317] The motion information of the three control points is combined into a triplet to construct a six-parameter affine motion model. The three control points can be combined in the following ways: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4}. For example, a six-parameter affine motion model constructed using a triplet of the CP1, CP2, and CP3 control points can be written as Affine(CP1, CP2, CP3).

[0318] The motion information of the four control points is combined into a quaternion to construct an 8-parameter bilinear motion model. The 8-parameter bilinear model constructed using the quaternion consisting of the CP1, CP2, CP3, and CP4 control points is denoted as Bilinear(CP1,CP2,CP3,CP4).

[0319] In the embodiment of the present application, for the convenience of description, the combination of motion information of two control points (or 2 encoded blocks) is referred to as a binary group, the combination of motion information of three control points (or 3 encoded blocks) is referred to as a triplet, and the combination of motion information of four control points (or 4 encoded blocks) is referred to as a quadruple.

[0320] The models are traversed in a preset order. If the motion information of a control point corresponding to the combined model is unavailable, the model is considered unavailable. Otherwise, the reference frame index of the model is determined and the motion vector of the control point is scaled. If the motion information of all the control points after scaling is consistent, the model is considered invalid. If the motion information of all the control points controlling the model is available and the model is valid, the motion information of the control points that construct the model is added to the motion information candidate list.

[0321] The method for scaling the motion vector of the control point is shown in formula (12):

[0322]

[0323] Among them, CurPoc represents the POC number of the current frame, CurPoc represents the POC number of the reference frame of the current block, CurPoc represents the POC number of the reference frame of the control point, CurPoc represents the scaled motion vector, and MV represents the motion vector of the control point.

[0324] It should be noted that it is also possible to convert a combination of different control points into control points at the same position.

[0325] For example, the four-parameter affine motion model obtained by combining {CP1, CP4}, {CP2, CP3}, {CP2, CP4}, {CP1, CP3}, or {CP3, CP4} is converted to a model represented by {CP1, CP2} or {CP1, CP2, CP3}. The conversion method is to substitute the motion vectors and coordinate information of the control points {CP1, CP4}, {CP2, CP3}, {CP2, CP4}, {CP1, CP3}, or {CP3, CP4} into formula (2) to obtain the model parameters, and then substitute the coordinate information of the control point {CP1, CP2} into formula (3) to obtain its motion vector.

[0326] More directly, the conversion can be performed according to the following formulas (13) to (21), where W represents the width of the current block and H represents the height of the current block. In formulas (13) to (21), (vx0, vy0) represents the motion vector of CP1, (vx0, vy0) represents the motion vector of CP2, (vx0, vy0) represents the motion vector of CP3, and (vx0, vy0) represents the motion vector of CP4.

[0327] The conversion of {CP1, CP2} to {CP1, CP2, CP3} can be achieved by the following formula (13), that is, the motion vector of CP3 in {CP1, CP2, CP3} can be determined by formula (13):

[0328]

[0329] The conversion of {CP1, CP3} to {CP1, CP2} or {CP1, CP2, CP3} can be achieved by the following formula (14):

[0330]

[0331] The conversion of {CP2, CP3} to {CP1, CP2} or {CP1, CP2, CP3} can be achieved by the following formula (15):

[0332]

[0333] The conversion of {CP1, CP4} to {CP1, CP2} or {CP1, CP2, CP3} can be achieved by the following formula (16) or (17):

[0334]

[0335] The conversion of {CP2, CP4} to {CP1, CP2} can be achieved by the following formula (18), and the conversion of {CP2, CP4} to {CP1, CP2, CP3} can be achieved by the following formulas (18) and (19):

[0336]

[0337] The conversion of {CP3, CP4} to {CP1, CP2} can be achieved by the following formula (20), and the conversion of {CP3, CP4} to {CP1, CP2, CP3} can be achieved by the following formulas (20) and (21):

[0338]

[0339] For example, the 6-parameter affine motion model obtained by combining {CP1, CP2, CP4}, {CP2, CP3, CP4}, or {CP1, CP3, CP4} is converted to a model represented by {CP1, CP2, CP3}. The conversion method is to substitute the motion vectors and coordinates of the control points {CP1, CP2, CP4}, {CP2, CP3, CP4}, or {CP1, CP3, CP4} into formula (4) to obtain the model parameters, and then substitute the coordinates of {CP1, CP2, CP3} into formula (5) to obtain its motion vector.

[0340] More directly, the conversion can be performed according to the following formulas (22) to (24), where W represents the width of the current block and H represents the height of the current block. In formulas (13) to (21), (vx0, vy0) represents the motion vector of CP1, (vx0, vy0) represents the motion vector of CP2, (vx0, vy0) represents the motion vector of CP3, and (vx0, vy0) represents the motion vector of CP4.

[0341] The conversion of {CP1, CP2, CP4} to {CP1, CP2, CP3} can be achieved by the following formula (22):

[0342]

[0343] The conversion of {CP2, CP3, CP4} to {CP1, CP2, CP3} can be achieved by the following formula (23):

[0344]

[0345] The conversion of {CP1, CP3, CP4} to {CP1, CP2, CP3} can be achieved by the following formula (24):

[0346]

[0347] 6. Advanced motion vector prediction mode based on affine motion model (Affine AMVP mode)

[0348] (1) Constructing a list of candidate motion vectors

[0349] The inherited control point motion vector prediction method and / or the constructed control point motion vector prediction method are used to construct a candidate motion vector list for the AMVP mode based on the affine motion model. In an embodiment of the present application, the candidate motion vector list for the AMVP mode based on the affine motion model may be referred to as a control point motion vector predictor candidate list, and the motion vector predictor for each control point includes the motion vectors of two (4-parameter affine motion model) control points or the motion vectors of three (6-parameter affine motion model) control points.

[0350] Optionally, the control point motion vector predictor candidate list is pruned and sorted according to specific rules, and may be truncated or padded to include a specific number of control point motion vector predictor candidates.

[0351] (2) Determine the optimal control point motion vector prediction value candidate

[0352] At the encoder, based on each control point motion vector prediction candidate (e.g., X-tuple candidate) in the control point motion vector prediction candidate list, the motion vector of each sub-motion compensation unit in the current coding block is obtained by formula (3) or (5), and then the sample value of the corresponding position in the reference frame pointed to by the motion vector of each sub-motion compensation unit is obtained. This sample value is used as the prediction value to perform motion compensation using the affine motion model. The average value of the difference between the original value and the prediction value of each sample in the current coding block is calculated, and the control point motion vector prediction candidate corresponding to the minimum average value is selected as the optimal control point motion vector prediction candidate and used as the motion vector prediction value of two or three control points in the current coding block. The index number representing the position of the optimal control point motion vector prediction candidate (e.g., X-tuple candidate) in the control point motion vector prediction candidate list is encoded into the bitstream and sent to the decoder.

[0353] At the decoding end, the index number is parsed, and a control point motion vector predictor (CPMVP) (eg, an X-tuple candidate) is determined from a control point motion vector predictor candidate list according to the index number.

[0354] (3) Determine the control point motion vector

[0355] At the encoding end, a motion search is performed within a certain search range using the control point motion vector prediction value as the search starting point to obtain the control point motion vector (CPMV), and the difference between the control point motion vector and the control point motion vector prediction value (CPMVD) is passed to the decoding end.

[0356] At the decoding end, the control point motion vector differences are parsed from the bitstream and added to the control point motion vector prediction values ​​to obtain the corresponding control point motion vectors.

[0357] 7. Affine Merge mode

[0358] A control point motion vector fusion candidate list is constructed using the inherited control point motion vector prediction method and / or the constructed control point motion vector prediction method.

[0359] Optionally, the control point motion vector fusion candidate list is pruned and sorted according to specific rules, and may be truncated or padded to a specific number.

[0360] At the encoding end, based on each control point motion vector candidate (e.g., X-tuple candidate) in the fusion candidate list, the motion vector of each sub-motion compensation unit (sample or sample block of size N1×N2 obtained according to a specific segmentation method) in the current coding block is obtained by formula (3) or (5). Then, the sample value of the position in the reference frame pointed to by the motion vector of each sub-motion compensation unit is obtained, and these sample values ​​are used as predicted sample values ​​to perform affine motion compensation. The average value of the difference between the original value and the predicted value of each sample in the current coding block is calculated, and the control point motion vector (CPMV) candidate (e.g., a two-tuple candidate or a three-tuple candidate) corresponding to the minimum average value of the difference is selected as the motion vector of the two or three control points of the current coding block. The index number representing the position of the control point motion vector in the candidate list is encoded into the bitstream and sent to the decoder.

[0361] At the decoding end, the index number is parsed, and the control point motion vector (CPMV) is determined from the control point motion vector fusion candidate list according to the index number.

[0362] In addition, it should be noted that, in this application, "at least one" means one or more, and "plurality" means more than two. "And / or" describes the association relationship of associated objects, indicating that there can be three kinds of relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.

[0363] In the present application, when the inter prediction mode is used to decode the current block, a syntax element may be used to signal the inter prediction mode.

[0364] The partial syntax structure currently used to parse the inter-frame prediction mode used by the current block can be seen in Table 1. Table 1 lists partial syntax of the inter-frame prediction mode. Syntax elements in the syntax structure can also be represented by other identifiers.

[0365] Table 1

[0366]

[0367]

[0368] In Table 1, inter_affine_flag[x0][y0] equal to 1 indicates that when decoding a P or B partition group, affine-based motion compensation is used for the current coding unit to generate the prediction samples for the current coding unit. inter_affine_flag[x0][y0] equal to 0 indicates that the coding unit is not predicted using affine-based motion compensation. If inter_affine_flag[x0][y0] is not present, inter_affine_flag[x0][y0] is inferred to be 0.

[0369] inter_pred_idc[x0][y0] indicates whether the current coding unit uses list0, list1 or bidirectional prediction, as shown in Table 2. The array index x0, y0 represents the position (x0, y0) of the top left luma sample of the considered coding block relative to the top left luma sample of the image.

[0370] If inter_pred_idc[x0][y0] does not exist, then inter_pred_idc[x0][y0] is inferred to be equal to PRED_L0.

[0371] Table 2

[0372]

[0373] sps_affine_enabled_flag indicates whether inter prediction uses affine-based motion compensation. If sps_affine_enabled_flag is 0, the syntax is constrained so that affine-based motion compensation is not used in CVS, and inter_affine_flag and cu_affine_type_flag are not present in the CVS coding unit syntax. Otherwise (sps_affine_enabled_flag is 1), affine-based motion compensation can be used in CVS.

[0374] The syntax element inter_affine_flag[x0][y0] (or affine_inter_flag[x0][y0]) can be used to indicate whether the AMVP mode based on the affine motion model is used for the current block when the slice where the current block is located is a P-type slice or a B-type slice. When this syntax element does not appear in the bitstream, it defaults to 0. For example, inter_affine_flag[x0][y0] = 1 indicates that the AMVP mode based on the affine motion model is used for the current block; inter_affine_flag[x0][y0] = 0 indicates that the AMVP mode based on the affine motion model is not used for the current block, and the AMVP mode based on the translational motion model can be used. In other words, inter_affine_flag[x0][y0] equal to 1 indicates that when decoding the P or B partition group, the affine model-based motion compensation is used for the current coding unit to generate the prediction samples of the current coding unit. inter_affine_flag[x0][y0] equal to 0 indicates that the coding unit is not predicted by affine model-based motion compensation. If inter_affine_flag[x0][y0] is not present, then inter_affine_flag[x0][y0] is inferred to be equal to 0.

[0375] The syntax element cu_affine_type_flag[x0][y0] can be used to indicate whether the 6-parameter affine motion model is used for motion compensation for the current block when the slice where the current block is located is a P-type slice or a B-type slice. cu_affine_type_flag[x0][y0]=0 indicates that the 6-parameter affine motion model is not used for motion compensation for the current block, and only the 4-parameter affine motion model can be used for motion compensation; cu_affine_type_flag[x0][y0]=1 indicates that the 6-parameter affine motion model is used for motion compensation for the current block. In other words, cu_affine_type_flag[x0][y0] equal to 1 indicates that when decoding P or B partition groups, motion compensation based on the 6-parameter affine model is used for the current coding unit to generate the prediction samples of the current coding unit. cu_affine_type_flag[x0][y0] equal to 0 indicates that motion compensation based on the 4-parameter affine model is used to generate the prediction samples of the current coding unit.

[0376] As shown in Table 3, MotionModelIdc[x0][y0]=1 indicates the use of a 4-parameter affine motion model, MotionModelIdc[x0][y0]=2 indicates the use of a 6-parameter affine motion model, and MotionModelIdc[x0][y0]=0 indicates the use of a translational motion model.

[0377] Table 3

[0378] Motion model for motion compensation 0 translational motion 1 4-parameter affine motion 2 6-parameter affine motion

[0379] The variables MaxNumMergeCand and MaxAffineNumMrgCand are used to indicate the maximum list length, indicating the maximum length of the constructed candidate motion vector list. inter_pred_idc[x0][y0] is used to indicate the prediction direction. PRED_L1 is used to indicate backward prediction. num_ref_idx_l0_active_minus1 indicates the number of reference frames in the forward reference frame list, and ref_idx_l0[x0][y0] indicates the forward reference frame index value of the current block. mvd_coding(x0,y0,0,0) indicates the first motion vector difference. mvp_l0_flag[x0][y0] indicates the forward MVP candidate list index value. PRED_L0 indicates forward prediction. num_ref_idx_l1_active_minus1 indicates the number of reference frames in the backward reference frame list. ref_idx_l1[x0][y0] represents the backward reference frame index value of the current block, and mvp_l1_flag[x0][y0] represents the backward MVP candidate list index value.

[0380] In Table 1, ae(v) represents a syntax element encoded using context-based adaptive binary arithmetic coding (CABAC).

[0381] Figure 9A Flowchart of the decoding method according to one embodiment of the present invention. The process may be performed by the inter-frame prediction unit 344 of the video decoder 30. The process is described as a series of steps or operations. It should be understood that the process may be performed in various orders and / or may occur simultaneously, not limited to Figure 9A Assume that a video decoder is used to decode a video data stream having a plurality of video frames using a process comprising: Figure 9A The inter-frame prediction process is shown.

[0382] Step 601: Parse the bitstream according to the syntax structure shown in Table 1 to determine the inter-frame prediction mode of the current block.

[0383] If it is determined that the inter prediction mode of the current block is the AMVP mode based on the affine motion model, step 602a is executed.

[0384] For example, the syntax elements merge_flag=0 and inter_affine_flag=1 indicate that the inter prediction mode of the current block is the AMVP mode based on the affine motion model.

[0385] If it is determined that the inter prediction mode of the current block is a merge mode based on an affine motion model, step 602b is executed.

[0386] For example, the syntax elements merge_flag=1 and inter_affine_flag=1 indicate that the inter prediction mode of the current block is a fusion mode based on an affine motion model.

[0387] Step 602a: Construct a candidate motion vector list corresponding to the AMVP mode based on the affine motion model.

[0388] Using the inherited control point motion vector prediction method and / or the constructed control point motion vector prediction method, one or more control point motion vector candidates (e.g., one or more X-tuple candidates) of the current block can be derived, and these control point motion vector candidates can be added to the candidate motion vector list.

[0389] The candidate motion vector list can include a two-tuple list (the current coding block uses a 4-parameter affine motion model) or a three-tuple list. The two-tuple list includes one or more two-tuples for constructing a 4-parameter affine motion model. The three-tuple list includes one or more triplets for constructing a 6-parameter affine motion model. It will be understood that each two-tuple candidate includes two candidate control point motion vectors for the current block.

[0390] Optionally, the list of candidate motion vector pairs / triples is pruned and sorted according to specific rules, and may be truncated or padded to include a specific number of pairs or triplets.

[0391] A1: Describe the process of constructing a candidate motion vector list using the inherited control motion vector prediction method.

[0392] by Figure 7 As shown in the example, the adjacent blocks around the current block are traversed in the order A1->B1->B0->A0->B2, and the affine decoded block containing the adjacent block of the current block is found and the control point motion information of the affine decoded block is obtained. The motion model is then constructed using the control point motion information of the affine decoded block to derive the candidate control point motion information of the current block. For details, please refer to the relevant description of the control point motion vector prediction method inherited from point 3 above, which will not be repeated here.

[0393] In one example, the affine motion model used by the current block is a 4-parameter affine motion model (i.e., MotionModelIdc = 1). In this example, if the adjacent affine decoded block uses a 4-parameter affine motion model, the motion vectors of the two control points of the affine decoded block are obtained: the motion vector (vx4, vy4) of the upper left control point (x4, y4) and the motion vector (vx5, vy5) of the upper right control point (x5, y5). The affine decoded block is an affine decoded block predicted using the affine motion model during the encoding phase.

[0394] Using the 4-parameter affine motion model composed of two control points of adjacent affine decoding blocks, the motion vectors of the two control points on the upper left and upper right of the current block are derived according to the 4-parameter affine motion model formulas (6) and (7).

[0395] If the adjacent affine decoding block uses a 6-parameter affine motion model, the motion vectors of the three control points of the adjacent affine decoding block are obtained, for example Figure 7 In the figure, the motion vector (vx4,vy4) of the upper left control point (x4,y4), the motion vector (vx5,vy5) of the upper right control point (x5,y5), and the motion vector (vx6,vy6) of the lower left control point (x6,y6) are shown.

[0396] Using the 6-parameter affine motion model composed of the 3 control points of the adjacent affine decoding blocks, the motion vectors of the upper left and upper right control points of the current block are derived according to the 6-parameter motion model formulas (8) and (9).

[0397] In another example, the affine motion model used by the current decoding block is a 6-parameter affine motion model (ie, MotionModelIdc=2).

[0398] If the affine motion model used by the adjacent affine decoding block is a 6-parameter affine motion model, the motion vectors of the three control points of the adjacent affine decoding block are obtained, for example Figure 7 In the figure, the motion vector of the upper left control point (x4,y4) is (vx4,vy4), the motion vector of the upper right control point (x5,y6) is (vx5,vy5), and the motion vector of the lower left control point (x6,y6) is (vx6,vy6).

[0399] Using the 6-parameter affine motion model composed of the 3 control points of the adjacent affine decoding blocks, the motion vectors of the 3 control points at the top left, top right, and bottom left of the current block are derived respectively according to the formulas (8), (9), and (10) corresponding to the 6-parameter affine motion model.

[0400] If the affine motion model used by the adjacent affine decoding block is a 4-parameter affine motion model, the motion vectors of the two control points of the affine decoding block are obtained: the motion vector (vx4, vy4) of the upper left control point (x4, y4) and the motion vector (vx5, vy5) of the upper right control point (x5, y5).

[0401] Using the 4-parameter affine motion model composed of 2 control points of adjacent affine decoding blocks, the motion vectors of the top left, top right and bottom left control points of the current block are derived according to the 4-parameter affine motion model formulas (6) and (7).

[0402] It should be noted that other motion models, candidate positions, and search orders may also be applicable to the present application. In addition, other control points may also be used to represent the motion models of adjacent and current coding blocks.

[0403] A2: Describe the process of constructing a candidate motion vector list using the constructed control point motion vector prediction method.

[0404] In one example, the affine motion model used by the current decoding block is a 4-parameter affine motion model (i.e., MotionModelIdc=1). In this example, the motion information of the adjacent coded blocks of the current coding block is used to determine the motion vectors of the upper left sample and the upper right sample of the current coding block. Specifically, the constructed control point motion vector prediction method 1 or the constructed control point motion vector prediction method 2 can be used to construct the candidate motion vector list. For specific methods, refer to the description in points 4 and 5 above.

[0405] In another example, the affine motion model used by the current decoding block is a 6-parameter affine motion model (i.e., MotionModelIdc=2). In this example, the motion information of the adjacent coded blocks of the current coding block is used to determine the motion vectors of the upper left sample, upper right sample, and lower left sample of the current coding block. Specifically, the constructed control point motion vector prediction method 1 or the constructed control point motion vector prediction method 2 can be used to construct the candidate motion vector list. For specific methods, refer to the description in points 4 and 5 above.

[0406] It should be noted that other control point motion information combination methods may also be used.

[0407] Step 603a: parse the code stream and determine the optimal control point motion vector prediction value (ie, the optimal tuple candidate).

[0408] B1: If the affine motion model used by the current decoding block is a 4-parameter affine motion model (MotionModelIdc=1), parse the index number from the bitstream and determine the optimal motion vector prediction values ​​of the two control points from the candidate motion vector list based on the index number.

[0409] For example, the index number is mvp_l0_flag or mvp_l1_flag.

[0410] B2: If the affine motion model used by the current decoding block is a 6-parameter affine motion model (MotionModelIdc=2), parse the index number from the bitstream and determine the optimal motion vector prediction value of the three control points from the candidate motion vector list according to the index number.

[0411] Step 604a: Parse the code stream and determine the control point motion vector.

[0412] C1: If the affine motion model used for the current decoding block is a 4-parameter affine motion model (MotionModelIdc = 1), the motion vector differences of the two control points of the current block are obtained from the decoded bitstream. The motion vector values ​​of the two control points are obtained based on the motion vector differences and the motion vector prediction values. Taking forward prediction as an example, the motion vector differences of the two control points are mvd_coding(x0,y0,0,0) and mvd_coding(x0,y0,0,1), respectively.

[0413] For example, the motion vector differences of the upper left control point and the upper right control point are obtained by decoding the code stream, and are added to the motion vector prediction values ​​of these two points to obtain the motion vectors of the upper left control point and the upper right control point of the current block.

[0414] C2: The affine motion model used by the current decoding block is a 6-parameter affine motion model (MotionModelIdc=2).

[0415] Decode the motion vector differences of the three control points of the current block from the bitstream. Determine the motion vector values ​​of the three control points based on the motion vector differences and the motion vector prediction values. For forward prediction (i.e., list 0), the motion vector differences of the three control points are mvd_coding(x0,y0,0,0), mvd_coding(x0,y0,0,1), and mvd_coding(x0,y0,0,2).

[0416] For example, the motion vector differences of the upper left control point, upper right control point and lower left control point are decoded from the code stream and added to the motion vector prediction values ​​respectively to obtain the motion vectors of the upper left control point, upper right control point and lower left control point of the current block.

[0417] Step 605a: Obtain the motion vector of each sub-block in the current block according to the control point motion information and the affine motion model used by the current decoding block.

[0418] A sub-block of the current affine decoding block can be equivalent to a motion compensation unit, and the width and height of the sub-block are smaller than the width and height of the current block. The motion information of the preset position sample in the sub-block or motion compensation unit can be used to represent the motion information of all samples in the sub-block or motion compensation unit. Assuming that the size of the motion compensation unit is M×N, the preset position sample can be the center sample (M / 2, N / 2), the upper left sample (0, 0), the upper right sample (M–1, 0) or the sample at other positions of the motion compensation unit. The following takes the center sample of the motion compensation unit as an example, see Figure 9C As shown in the figure, V0 represents the motion vector of the upper left control point, and V1 represents the motion vector of the upper right control point. Each small box represents a motion compensation unit.

[0419] The coordinates of the center sample of the motion compensation unit relative to the upper left sample of the current affine decoding block are calculated according to formula (25), where i represents the i-th motion compensation unit in the horizontal direction (from left to right), j represents the j-th motion compensation unit in the vertical direction (from top to bottom), (x ( i,j ) ,x ( i,j ) ) represents the coordinates of the center sample of the (i, j)th motion compensation unit relative to the upper left control point sample of the current affine decoding block.

[0420] If the affine motion model used by the current affine decoding block is a 6-parameter affine motion model, then (x ( i,j ) ,x ( i,j ) ) is substituted into the 6-parameter affine motion model formula (26) to obtain the motion vector (x) of the center sample of each motion compensation unit. ( i,j ) ,x ( i,j ) ), as mentioned above, the motion vector of the center pixel of the motion compensation unit is used as the motion vector of all samples in the motion compensation unit.

[0421] If the affine motion model used by the current affine decoding block is a 4-affine motion model, then (x ( i,j ) ,x ( i,j ) ) is substituted into the 4-parameter affine motion model formula (27) to obtain the motion vector (x( i,j ) ,x ( i,j ) ) as the motion vectors of all samples in the motion compensation unit.

[0422]

[0423] Step 606a: For each sub-block, perform motion compensation according to the determined motion vector of the sub-block to obtain a predicted sample value of the sub-block.

[0424] As described above, if it is determined in step 601 that the inter prediction mode of the current block is the merge mode based on the affine motion model, step 602b is executed.

[0425] Step 602b: Construct a motion information candidate list corresponding to the fusion mode based on the affine motion model.

[0426] Specifically, the inherited control point motion vector prediction method and / or the constructed control point motion vector prediction method may be used to construct a motion information candidate list corresponding to a fusion mode based on an affine motion model.

[0427] Optionally, the motion information candidate list is pruned and sorted according to specific rules, and may be truncated or padded to include a specific number of motion information.

[0428] D1: Describe the process of constructing a candidate motion vector list using the inherited control motion vector prediction method.

[0429] The inherited control point motion vector prediction method is used to derive the candidate control point motion information of the current block and add it to the motion information candidate list.

[0430] according to Figure 8A The adjacent blocks around the current block are traversed in the order of A1->B1->B0->A0->B2, and the affine decoding block where the adjacent block is located is found. The control point motion information of the affine decoding block is obtained, and then the motion model of the current block is used to derive the candidate control point motion information of the current block.

[0431] If the candidate motion vector list is empty at this time, the candidate control point motion information obtained above is added to the candidate list; otherwise, the motion information in the candidate motion vector list is traversed in sequence to check whether there is motion information in the candidate motion vector list that is the same as the candidate control point motion information. If there is no motion information in the candidate motion vector list that is the same as the candidate control point motion information, the candidate control point motion information is added to the candidate motion vector list.

[0432] To determine whether two candidate motion information are identical, the forward (list 0) and backward (list 1) reference frames, as well as the horizontal and vertical components of the forward and backward motion vectors, are checked. The two motion information are considered different only if all of these elements are different.

[0433] If the number of motion information in the candidate motion vector list reaches the maximum list length MaxAffineNumMrgCand (a positive integer, such as 1, 2, 3, 4 or 5, 5 is used as an example below), the candidate list is constructed, otherwise the next adjacent block is traversed.

[0434] D2: Use the constructed control point motion vector prediction method to derive the candidate control point motion information of the current block and add it to the motion information candidate list. Figure 9B An example of a flowchart of a constructed control point motion vector prediction method is shown.

[0435] Step 601c: Obtain motion information of each control point of the current block. This step is similar to step 501 in "5. Constructing Control Point Motion Vector Prediction Method 2" and will not be repeated here.

[0436] Step 602c: Combine the motion information of each control point to obtain the constructed control point motion information. Figure 8B The process is similar to step 501 in , and will not be described again here.

[0437] Step 603c: Add the constructed control point motion information to the candidate motion vector list.

[0438] If the length of the candidate list is less than the maximum list length MaxAffineNumMrgCand, the combinations of control point motion information are traversed in a preset order to obtain a valid combination as the candidate control point motion information. If the candidate motion vector list is empty, the motion information of the candidate control point is added to the candidate motion vector list. Otherwise, the motion information in the candidate motion vector list is traversed in sequence to check whether there is motion information in the candidate motion vector list that is identical to the motion information of the candidate control point. If there is no motion information in the candidate motion vector list that is identical to the motion information of the candidate control point, the motion information of the candidate control point is added to the candidate motion vector list.

[0439] For example, one preset order is as follows: Affine(CP1,CP2,CP3)->Affine(CP1,CP2,CP4)->Affine(CP1,CP3,CP4)->Affine(CP2,CP3,CP4)->Affine(CP1,CP2)->Affine(CP1,CP3)->Affine(CP2,CP3)->Affine(CP1,CP4)->Affine(CP2,CP4)->Affine(CP3,CP4), a total of 10 combinations.

[0440] If the motion information of the control points corresponding to the combination is unavailable, the combination is considered unavailable. If the combination is available, the reference frame index of the combination is determined (when there are two control points, the reference frame index with the smallest index is selected as the reference frame index of the combination; when there are more than two control points, the reference frame index with the largest number of occurrences is selected as the reference frame index of the combination. If there are multiple reference frame indices with the same number of occurrences, the reference frame index with the smallest index is selected as the reference frame index of the combination), and the motion vectors of the control points are scaled. If the motion information of all control points after scaling is consistent, the combination is illegal.

[0441] Optionally, the embodiment of the present application can also fill the candidate motion vector list. For example, after the above-mentioned traversal process, if the length of the candidate motion vector list is less than the maximum list length MaxAffineNumMrgCand, the candidate motion vector list can be filled until the length of the list is equal to MaxAffineNumMrgCand.

[0442] The method of filling with zero motion vectors or combining the candidate motion information already in the existing list (eg weighted average) can be used to fill the list. It should be noted that other methods of obtaining the candidate motion vector list filling method are also applicable to this application.

[0443] Step 603b: Analyze the code stream and determine the optimal control point motion information.

[0444] Parse the index number and determine the optimal control point motion information from the candidate motion vector list based on the index number.

[0445] Step 604b: Obtain the motion vector of each sub-block in the current block according to the optimal control point motion information and the affine motion model used by the current decoding block.

[0446] This step is similar to step 605a.

[0447] Step 605b: For each sub-block, perform motion compensation according to the determined motion vector of the sub-block to obtain a predicted sample value of the sub-block.

[0448] As described above, after the motion vector of each sub-block is obtained through step 605a and step 604b, motion compensation of the sub-block is performed through step 606a and step 605b respectively. That is, sub-block-based affine motion compensation is performed on the current sub-block of the affine decoding block, and the details of the predicted sample value of the current sub-block of the affine decoding block are obtained as described above. In traditional designs, the size of the sub-block is set to 4×4, that is, each 4×4 unit uses a corresponding / different motion vector to perform motion compensation. Generally, the smaller the size of the sub-block, the higher the computational complexity of motion compensation, and the better the prediction effect. In order to take into account both the computational complexity of motion compensation and the prediction accuracy, a process of prediction signal refinement with optical flow (PROF) is proposed after sub-block-level motion compensation. The specific steps of this process are as follows:

[0449] (1) The motion vector of each sub-block is obtained through steps 605a and 604b, and motion compensation of the sub-block is performed through steps 606a and 605b to obtain the prediction signal I(i, j) of the sub-block. It should be noted that the PROF process does not include step (1).

[0450] (2) Calculate the horizontal gradient value g of the prediction signal of the sub-block x (i, j) and vertical gradient value g x (i,j), is calculated as follows:

[0451] g x (i,j)=I(i+1,j)-I(i-1,j)

[0452] g y (i,j)=I(i,j+1)-I(i,j-1)

[0453] According to the formula, to obtain the gradient value of a 4×4 block (4×4 gradient value), a 6×6 prediction signal window 900 is required, as shown in Figure 9D shown.

[0454] This can be achieved in different ways:

[0455] (a) After obtaining the prediction matrix for the sub-block based on the sub-block's motion information (e.g., motion vector), the horizontal and vertical gradient matrices for the sub-block are obtained. In other words, based on the motion vectors of the M×N sub-block, an (M+2)*(N+2) prediction block is obtained through interpolation. For example, based on the sub-block's motion vectors, a 6×6 prediction signal is directly interpolated, and a 4×4 gradient value (i.e., a 4×4 gradient matrix) is calculated.

[0456] (b) According to the motion vector of the sub-block, interpolation is performed to obtain a 4×4 prediction signal (i.e., the first prediction matrix), and then the prediction signal is expanded to obtain a 6×6 prediction signal (i.e., the second prediction matrix), and a 4×4 gradient value is calculated (i.e., the 4×4 gradient matrix).

[0457] (c) Based on the motion vector of each sub-block, interpolation is performed to obtain each 4×4 prediction signal (i.e., the first prediction matrix), and the prediction signal of w*h is obtained by combination. Then, the prediction signal of w*h is expanded to obtain the prediction signal of (w+2)*(h+2), and the gradient value of w*h is calculated (i.e., the gradient matrix of w*h), and then the gradient value of each 4×4 is obtained (i.e., the gradient matrix of 4×4).

[0458] It should be noted that, according to the motion vector of the M×N sub-block, directly interpolating the (M+2)*(N+2) prediction block includes the following implementation methods:

[0459] (a1) For the surrounding area ( Figure 13 (white samples in ), get the integer sample of the upper left sample of the position pointed by the motion vector. For the inner area ( Figure 13 The gray sample in the image is obtained to obtain the sample at the position pointed by the motion vector. If the sample is a fractional sample, the interpolation filter is used to interpolate the sample.

[0460] like Figure 14 As shown, A, B, C, and D are integer samples. The motion vector of the M×N sub-block has 1 / 16 sample precision. dx / 16 is the horizontal distance between the fractional sample and the integer sample to the upper left, and dy / 16 is the vertical distance between the fractional sample and the integer sample to the upper left. For the surrounding area, the sample value of A is used as the predicted sample value for the sample location. For the internal area, the predicted sample value for the sample location is interpolated using an interpolation filter.

[0461] (a2) For the surrounding area ( Figure 13 (white samples in ), get the integer sample closest to the position pointed by the motion vector. For the inner area ( Figure 13 The gray sample in the image is obtained to obtain the sample at the position pointed by the motion vector. If the sample is a fractional sample, the interpolation filter is used to interpolate the sample.

[0462] like Figure 14 As shown, for the surrounding area, the integer sample closest to the position pointed by the motion vector is selected according to dx and dy.

[0463] (a3) For both the surrounding area and the internal area, obtain the sample at the position pointed by the motion vector. If the sample is a fractional sample, use the interpolation filter to interpolate the sample.

[0464] It should be understood that (a), (b) and (c) are three different implementations.

[0465] (3) Calculate the incremental forecast value. The calculation method is as follows:

[0466] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j)

[0467] Where (i, j) represents the current sample of the sub-block, and Δv(i, j) is the difference between the motion vector of the current sample of the current sub-block and the motion vector of the center sample of the sub-block (e.g. Figure 10 As shown), it can be calculated according to the above formula, Δv(i,j) and Δv(i,j) are the horizontal offset value and vertical offset value of the difference between the motion vector of the current sample of the current sub-block and the motion vector of the center sample of the sub-block. Alternatively, in the simplified method, the motion vector difference between the motion vector of each 2×2 sample block to which the current sample belongs and the motion vector of the center sample of the sub-block can be calculated. In contrast, Δv(i,j): the motion vector difference is calculated for each pixel or sample (for example, 16 times in a 4×4 sub-block), while in the simplified method, the motion vector difference is calculated for each 2×2 sub-block (for example, 4 times in a 4×4 sub-block). It should be noted that the sub-block here can be a 4×4 sub-block or an m×n sub-block. For example, m here is greater than or equal to 4, or n here is greater than or equal to 4.

[0468]

[0469] For a 4-parameter affine model:

[0470]

[0471] For a 6-parameter affine model:

[0472]

[0473] Where (v0x, v0y), (v0x, v0y) and (v0x, v0y) are the motion vectors of the upper left, upper right and lower left control points, and w and h are the width and height of the affine decoding block (CU).

[0474] (4) Perform forecast correction:

[0475] I′(i,j)=I(i,j)+ΔI(i,j)

[0476] Among them, I(i,j) represents the predicted value of the sample (i,j) of the sub-block (that is, the predicted sample value at position (i,j) in the sub-block), I(i,j) represents the incremental predicted value of the sample (i,j) of the sub-block, and I(i,j) is the corrected predicted sample value of the sample (i,j) of the sub-block.

[0477] According to an embodiment of the present invention, “the prediction refinement with optical flow (PROF) process is performed according to a condition, thereby refining the sub-block-based affine motion compensation prediction value using the optical flow” is described as follows.

[0478] Example 1

[0479] A method for using optical flow to obtain incremental prediction values ​​for sub-blocks (specifically, incremental prediction values ​​for each sample of the sub-block) can be applied to unidirectional affine decoding blocks and can also be applied to bidirectional affine decoding blocks. If this method is applied to bidirectional affine prediction blocks, the above steps (1) to (4) need to be performed twice, which has a high computational complexity. In order to reduce the complexity of this method, the present invention proposes to constrain it, that is, this method is only used to correct the predicted sample values ​​when the affine decoding block is a unidirectional affine decoding block.

[0480] At the decoding end, the syntax elements parsed from the bitstream indicate unidirectional prediction or bidirectional prediction. This syntax element can be used to determine whether an affine decoded block is a unidirectional affine decoded block.

[0481] On the encoder side, the structure of B-frames and P-frames is user-defined, and the use of unidirectional or bidirectional prediction in B-frames is determined by RDO. In other words, for B-frames, the encoder can determine whether to use unidirectional or bidirectional prediction for the current affine image block based on the RDO cost. For example, the encoder attempts to select a mechanism that minimizes RDO among forward prediction, backward prediction, and bidirectional prediction.

[0482] Example 2

[0483] To reduce the complexity of using optical flow to correct the predicted signal, this method can only be used when the sub-block size is large. That is, optical flow is used to obtain the predicted offset value of the sub-block. For example, the sub-block size of the unidirectional affine decoding block can be set to 4×4, and the sub-block size of the bidirectional affine decoding block can be set to 8×8, 8×4, or 4×8. In this example, this method is only used when the sub-block size is larger than 4×4. For another example, the sub-block size can be adaptively selected based on information such as the motion vector of the control point of the affine decoding block, the width and height of the affine decoding block, etc. This method is only used when the sub-block size is larger than 4×4.

[0484] In addition, in step (2), both methods (a) and (b) can ensure that the predictions of each 4×4 sub-block of the affine decoding block have no dependencies and can be executed in parallel. However, method (a) increases the complexity of the interpolation calculation. Although method (b) does not increase the complexity, the gradient value of the boundary is calculated by expanding the samples, and the accuracy is not high. Method (c) can improve the accuracy of the gradient calculation, but there is a dependency between each 4×4 sub-block, that is, the optical flow correction can only be performed after the entire CU completes the interpolation.

[0485] like Figure 9E As shown, in order to take into account the accuracy of the degree of parallelism and gradient calculation, the present invention proposes to calculate the gradient value according to the granularity of 16×16. Assuming size_w=min(w,16), size_h=min(h,16), for each size_w*size_h in the affine decoding block, the predicted value (predicted sample value) of each 4×4 sub-block in the affine decoding block is calculated, and the prediction signal of size_w*size_h is combined to obtain the edge, and then the prediction signal of size_w*size_h is expanded (for example, expanded outward by 2 samples, such as padding) to obtain a prediction signal of (size_w+2)*(size_h+2), and the gradient value of size_w*size_h is calculated, and then the gradient value of each 4×4 is obtained. It should be understood that the number of samples expanded outward in this application is not limited to 2 samples, which is related to the gradient calculation. If the gradient calculation is 3 taps, 2 samples are expanded outward, which is related to the filter for calculating the gradient. Assuming that the number of filter taps is T, the supported increase area or surrounding area is T / 2 (integer divisible)*2.

[0486] Figure 11A A method for performing optical flow prediction refinement (PROF) on an affine decoded block, provided by an embodiment, is shown. The method can be performed by a decoding device (e.g., a decoding device or a decoder). The method includes the following steps:

[0487] S1101: Determine whether multiple optical flow decision conditions are met.

[0488] Here, the optical flow decision conditions can also be referred to as conditions under which PROF can be applied. If all the optical flow decision conditions are met, PROF is applied to the current sub-block of the affine decoded block. The following examples describe the optical flow decision conditions. In some examples, the optical flow decision conditions can be replaced or rewritten as PROF application constraints. If one PROF application constraint is met, PROF is not applied to the current sub-block of the affine decoded block. In these examples, step S1101 becomes: determining that none of the multiple PROF application constraints are met.

[0489] S1102: Performing a PROF process on the current sub-block of the affine decoded block to obtain a modified predicted sample value of the current sub-block of the affine decoded block, wherein the affine decoded block satisfies all of the multiple optical flow decision conditions. Here, the modified predicted sample value of the current sub-block can be understood as a final predicted sample value obtained after adding the prediction correction value to the current sub-block.

[0490] In step S1102, optical flow (prediction refinement with optical flow, PROF) processing is performed on one or more sub-blocks (e.g., the current sub-block or each sub-block) in the current affine image block to obtain incremental (delta) prediction values ​​(e.g., ΔI(i,j)) of the one or more sub-blocks (e.g., the current sub-block or each sub-block) in the current affine image block.

[0491] Step S1102 includes: obtaining a modified prediction sample value (e.g., prediction signal I′(i,j)) of the sub-block based on the incremental prediction value (e.g., I′(i,j)) of the sub-block and the prediction sample value (e.g., prediction signal I′(i,j)) of the sub-block.

[0492] Specifically, step S1102 includes: obtaining a corrected prediction value (e.g., prediction signal I′(i,j)) of the current sample in the sub-block based on the incremental prediction value (e.g., I′(i,j)) of the current sample in the sub-block and the prediction value (e.g., prediction signal I′(i,j)) of the current sample in the sub-block.

[0493] In one possible design, the multiple optical flow decision conditions include one or more of the following:

[0494] (a) The indication information obtained by parsing or derivation (e.g., sps_prof_enabled_flag or sps_bdof_enabled_flag) indicates that PROF is enabled for the current sequence, image, slice, or tile group, e.g., sps_prof_enabled_flag or sps_bdof_enabled_flag = 1. It will be understood that if a PROF application constraint is used in S1101 instead of an optical flow decision condition, the condition can be converted into a PROF application constraint as follows: (a) The indication information indicates that PROF is disabled for the current sequence, image, slice, or tile group, e.g., sps_prof_disabled_flag or sps_bdof_disabled_flag = 1.

[0495] The indication information obtained by parsing parameter sets such as SPS and PPS, slice header or tile group header indicates whether PROF is enabled for the current sequence, picture, slice or tile group.

[0496] Specifically, use sps_prof_enabled_flag to control whether PROF is enabled, and the syntax and semantics of sps_prof_enabled_flag are as follows:

[0497]

[0498] sps_prof_enabled_flag equal to 0 disables optical flow prediction modification for affine motion compensation. sps_prof_enabled_flag equal to 1 enables optical flow prediction modification for affine motion compensation.

[0499] Optionally, use sps_bdof_enabled_flag to control whether PROF is enabled.

[0500] It should be understood that in this embodiment, other conditions are derived (e.g., the master switch determines whether PROF is enabled) provided that the above conditions are met. In other words, if PROF is enabled for the current sequence, image, strip, or block group, a further determination is made as to whether the current affine image block satisfies the other optical flow decision conditions described below. If PROF is not enabled for the current sequence, image, strip, or block group, there is no need to determine whether the current affine image block satisfies the other optical flow decision conditions provided below.

[0501] (b) The derived indication information (e.g., the variable fallbackModeTriggered) indicates that the current affine decoded block is to be segmented, e.g., fallbackModeTriggered = 0. It is understood that if PROF application constraints are used in S1101 instead of optical flow decision conditions, then condition (b) can be converted into PROF application constraints as follows: (b) The derived indication information indicates that the current affine decoded block is not to be segmented, e.g., fallbackModeTriggered = 1.

[0502] The variable fallbackModeTriggered is derived from the affine parameters, and whether PROF is used depends on the parameter fallbackModeTriggered. When fallbackModeTriggered is 1, it indicates that the current affine decoding block is not being split. When fallbackModeTriggered is 0, it indicates that the affine decoding block is to be split (for example, the affine decoding block is to be split into multiple sub-blocks, such as 4×4 sub-blocks). PROF will be used when the current affine decoding block is to be split.

[0503] Specifically, the variable fallbackModeTriggered can be derived through the following process:

[0504] Initially the variable fallbackModeTriggered is set to 1 and further derivation is as follows:

[0505] –Variables bxWX4, bxHX4, bxWX h 、bxHX h 、bxWX v and bxHX v The derivation is as follows:

[0506] maxW4=Max(0,Max(4*(2048+dHorX),

[0507] Max(4*dHorY,4*(2048+dHorX)+4*dHorY)))(8-775)

[0508] minW4=Min(0,Min(4*(2048+dHorX),

[0509] Min(4*dHorY,4*(2048+dHorX)+4*dHorY)))(8-775)

[0510] maxH4=Max(0,Max(4*dVerX,

[0511] Max(4*(2048+dVerY),4*dVerX+4*(2048+dVerY))))(8-775)

[0512] minH4=Min(0,Min(4*dVerX,

[0513] Min(4*(2048+dVerY),4*dVerX+4*(2048+dVerY))))(8-775)

[0514] bxWX4 = ( ( maxW4 – minW4 ) >> 11 ) + 9 (8-775)

[0515] bxHX4 = ( ( maxH4 – minH4 ) >> 11 ) + 9 (8-775)

[0516] wxya h = ( (Max( 0, 4 * ( 2048 + dHorX ) ) – Min( 0, 4 * ( 2048 + dHorX) ) ) >> 11 ) + 9 (8-775)

[0517] xXd h = ( ( Max( 0, 4 * dVerX ) – Min( 0, 4 * dVerX ) ) >> 11 ) + 9 (8-775)

[0518] wxya v = ( ( Max( 0, 4 * dVerY ) – Min( 0, 4 * dVerY ) ) >> 11 ) + 9 (8-775)

[0519] xXd v = ( ( Max( 0, 4 * ( 2048 + dHorY ) ) – Min( 0, 4 * ( 2048 + dHorY) ) ) >> 11 ) + 9 (8-775)

[0520] – If inter_pred_idc[xCb][yCb] is equal to PRED_BI and bxWX4*bxHX4 is less than or equal to 225, set fallbackModeTriggered to 0.

[0521] – If bxWXh*bxHXh is less than or equal to 165, and bxWXv*bxHXv is less than or equal to 165, then set fallbackModeTriggered to 0.

[0522] (c) The current affine image block is a unidirectionally predicted affine image block.

[0523] (d) The size of the sub-block in the affine image block is larger than N×N, where N=4.

[0524] (e) The current affine image block is a unidirectionally predicted affine image block, and the size of a sub-block in the affine image block is equal to N×N, where N=4.

[0525] (f) The current affine image block is a bidirectionally predicted affine image block, and a size of a sub-block in the affine image block is larger than N×N, where N=4.

[0526] The current affine image block is the current affine coding block. Whether the current affine image block is a unidirectionally predicted affine image block is determined by the following method:

[0527] At the encoding end, the current affine image block is determined to use unidirectional prediction according to the rate-distortion criterion RDO.

[0528] The current affine image block is the current affine decoding block. Whether the current affine image block is a unidirectionally predicted affine image block is determined by the following method:

[0529] At the decoding end, in AMVP mode, the prediction direction indication information is used to indicate a unidirectional prediction direction (for example, only forward prediction or only backward prediction), and the prediction direction indication information is obtained by parsing the bitstream or derivation; or

[0530] At the decoding end, in fusion mode, the candidate motion information corresponding to the candidate index in the candidate list includes first motion information corresponding to the first reference frame list, or the candidate motion information corresponding to the candidate index in the candidate list includes second motion information corresponding to the second reference frame list.

[0531] In one possible design, the prediction direction indication information includes a syntax element inter_pred_idc[x0][y0], where

[0532] inter_pred_idc[x0][y0]=PRED_L0, used to indicate forward prediction;

[0533] inter_pred_idc[x0][y0]=PRED_L1, used to indicate backward prediction; or

[0534] The prediction direction indication information includes predFlagL0 and / or predFlagL1, wherein

[0535] predFlagL0=1 and predFlagL1=0, used to indicate forward prediction;

[0536] predFlagL1=1 and predFlagL0=0 are used to indicate backward prediction.

[0537] It should be noted that the optical flow decision conditions (or PROF application constraints) are not limited to the above examples, but other or different optical flow decision conditions (or PROF application constraints) can be set according to different application scenarios. For example, the above conditions (a) to (f) can be replaced by other conditions, such as the following optical flow decision condition: if all control point MVs of the affine decoding block are different from each other, PROF can be applied to the affine decoding block, or the following PROF application constraint condition: if all control point MVs of the affine decoding block are the same, PROF is not applied to the affine decoding block; for example, the following optical flow decision condition: if the resolution of the current image where the affine decoding block is located and the resolution of the reference image of the affine decoding block are the same, PROF can be applied to the affine decoding block, for example, RprConstraintsActive[X][refIdxLX] is equal to 0, or the following PROF application constraint condition: if the resolution of the current image where the affine decoding block is located and the resolution of the reference image of the affine decoding block are different, PROF is not applied to the affine decoding block, for example, RprConstraintsActive[X][refIdxLX] is equal to 1.

[0538] In one possible design, in step S1102, performing optical flow (prediction refinement with optical flow, PROF) processing on one or more sub-blocks (e.g., each sub-block or the current sub-block) in the current affine image block to obtain incremental prediction values ​​(e.g., ΔI(i, j)) of the one or more sub-blocks (e.g., each sub-block or the current sub-block) in the current affine image block may include the following steps:

[0539] Step 1: Obtain a second prediction matrix according to motion information (eg, motion vector) of a current sub-block in the current affine image block.

[0540] For example, according to the motion vector of the M×N sub-block, a prediction block of (M+2)*(N+2) (ie, the second prediction matrix) is obtained by interpolation. Different implementations are provided above.

[0541] Step 2: Calculate a horizontal prediction gradient matrix and a vertical prediction gradient matrix according to the second prediction matrix, wherein the size of the second prediction matrix is ​​greater than or equal to the size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix.

[0542] Step 3: Calculate an incremental prediction value (ΔI(i, j)) of the current sample in the sub-block according to the horizontal prediction gradient value of the current sample in the horizontal prediction gradient matrix, the vertical prediction gradient value of the current sample in the vertical prediction gradient matrix, and the difference between the motion vector of the current sample of the current sub-block and the motion vector of the center sample of the sub-block.

[0543] Accordingly, in step S1102, obtaining a corrected prediction value (e.g., prediction signal I′(i,j)) of the sub-block according to the incremental prediction value (e.g., I′(i,j)) of the sub-block and the prediction sample value (e.g., prediction signal I′(i,j)) of the sub-block may include:

[0544] According to the incremental prediction value of the current sample in the sub-block (e.g., I′(i,j)) and the predicted sample value of the current sample (e.g., prediction signal I′(i,j)), a corrected predicted sample value of the current sample (e.g., prediction signal I′(i,j)) is obtained.

[0545] It should be understood that the predicted sample value of the sub-block (eg, the prediction signal I(i, j)) may be an M×N prediction block in an (M+2)*(N+2) prediction block.

[0546] For step 3, in one implementation, the motion vector differences between the motion vectors of different samples in the current sub-block and the motion vector of the center sample of the sub-block are different. In another implementation, the motion vector difference between the motion vector of the current sample unit (e.g., a 2×2 sample block) including the current sample and the motion vector of the center sample of the sub-block is used as the motion vector difference between the motion vector of the current sample of the current sub-block and the motion vector of the center sample of the sub-block. In other words, to balance processing overhead and prediction accuracy, assuming that the current sample unit includes sample A and sample B, the motion vector difference between the motion vector of the current sample unit (e.g., a 2×2 sample block) and the motion vector of the center sample of the sub-block can be used as the motion vector difference between the motion vector of sample A in the sub-block and the motion vector of the center sample of the sub-block; in addition, the motion vector difference between the motion vector of the current sample unit and the motion vector of the center sample of the sub-block can be used as the motion vector difference between the motion vector of sample B in the sub-block and the motion vector of the center sample of the sub-block.

[0547] In one implementation, the second prediction matrix in step 1 is expressed as I1(p,q), where the value range of p is [–1, sbW] and the value range of q is [–1, sbH];

[0548] The horizontal prediction gradient matrix is ​​expressed as X(i,j), where the value range of i is [0,sbW–1] and the value range of j is [0,sbH–1];

[0549] The vertical prediction gradient matrix is ​​expressed as Y(i,j), where the value range of i is [0,sbW–1] and the value range of j is [0,sbH–1].

[0550] sbW represents the width of the current sub-block in the current affine image block, sbH represents the height of the current sub-block in the current affine image block, (x, y) represents the position coordinates of each sample (also called sample) in the current sub-block in the current affine image block, and the element located at (x, y) can correspond to the element located at (i, j).

[0551] In another possible design, in step 1102, performing optical flow (prediction refinement with optical flow, PROF) processing on one or more sub-blocks (e.g., each sub-block) in the current affine image block to obtain incremental prediction values ​​(also called prediction value offset values, such as ΔI(i, j)) of the one or more sub-blocks (e.g., each sub-block) in the current affine image block includes the following steps: Figure 12 As shown:

[0552] S1202: Obtain or generate a second prediction matrix based on the first prediction matrix, wherein the first prediction matrix (e.g., the first prediction signal I(i, j) or the 4×4 prediction block) of the sub-block (e.g., each sub-block) corresponds to the predicted sample value of the current sub-block. Perform sub-block-based affine motion compensation on the current sub-block of the affine decoding block to obtain the predicted sample value of the current sub-block of the affine decoding block, such as Figure 9A shown.

[0553] S1203: Calculate a horizontal prediction gradient matrix and a vertical prediction gradient matrix based on the second prediction matrix, wherein a size of the second prediction matrix is ​​greater than or equal to a size of the first prediction matrix, and a size of the second prediction matrix is ​​greater than or equal to a size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix.

[0554] S1204: Calculate an incremental prediction value matrix (e.g., ΔI(i, j) of a prediction signal) for the sub-block based on the horizontal prediction gradient matrix, the vertical prediction gradient matrix, and a motion vector difference between a motion vector of a current sample unit (e.g., a current sample or a current sample block, e.g., a 2×2 sample block) of the sub-block and a motion vector of a center sample of the sub-block.

[0555] The step of obtaining a modified predicted sample value (e.g., predicted signal I′(i,j)) of the sub-block according to the incremental predicted value (e.g., I′(i,j)) of the sub-block and the predicted sample value (e.g., predicted signal I′(i,j)) of the sub-block comprises:

[0556] S1205: Obtain a modified third prediction matrix (eg, prediction signal I′(i,j)) of the sub-block based on the incremental prediction value matrix (eg, I′(i,j)) and the first prediction matrix (eg, prediction signal I′(i,j)).

[0557] It should be understood that, here, I(i, j) represents the predicted sample value of the current sample in the current sub-block (e.g., the original predicted value obtained through motion compensation), I(i, j) represents the incremental predicted value of the current sample in the current sub-block, and I(i, j) represents the revised predicted sample value of the current sample in the current sub-block. For example, original predicted sample value + incremental predicted value = revised predicted sample value. It should be understood that obtaining revised predicted sample values ​​for multiple samples (e.g., all samples) in the current sub-block is equivalent to obtaining the revised predicted sample value of the current sub-block.

[0558] In different possible implementations, the gradient value can be calculated sample by sample, and the incremental prediction value can also be calculated sample by sample. Alternatively, the gradient value matrix can be obtained first, and then the incremental prediction value can be calculated. This application does not impose any restrictions on this. In an optional implementation, the first prediction matrix and the second prediction matrix represent the same prediction matrix.

[0559] When the size of the second prediction matrix is ​​equal to the size of the first prediction matrix and the size of the second prediction matrix is ​​equal to the size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix, in one possible implementation, a (w–2)*(h–2) gradient matrix is ​​calculated using a w*h prediction matrix, and the gradient matrix is ​​padded to obtain a size of w*h, where w*h represents the size of the current sub-block. For example, the size of the first prediction matrix and the second prediction matrix are both w*h, or the first prediction matrix and the second prediction matrix represent the same prediction matrix.

[0560] like Figure 11BAs shown, another embodiment of the present application provides another method for performing optical flow prediction refinement (PROF) on an affine decoding block. The method includes the following steps:

[0561] S1110: Determine whether multiple optical flow decision conditions are met. Here, the optical flow decision conditions refer to conditions under which PROF can be applied.

[0562] S1111: If the multiple optical flow decision conditions are met, the first indicator (e.g., applyProfFlag) is set to true (true), and an optical flow prediction refinement (PROF) process is performed on the current sub-block of the affine decoding block to obtain a modified predicted sample value of the current sub-block of the affine decoding block. In step S1111, an optical flow (PROF) process is performed on one or more sub-blocks (e.g., each sub-block) in the current affine image block to obtain an incremental prediction value (also known as a prediction value offset value, e.g., ΔI(i,j)) of the one or more sub-blocks (e.g., each sub-block) in the current affine image block.

[0563] In step S1111, the corrected prediction sample value (e.g., prediction signal I′(i,j)) of the sub-block is obtained based on the incremental prediction value (e.g., I′(i,j)) and the prediction sample value (e.g., prediction signal I′(i,j)) of the sub-block.

[0564] It should be understood that, here, I(i, j) represents the predicted sample value of the current sample in the current sub-block (e.g., the original predicted sample value obtained through motion compensation), I(i, j) represents the incremental predicted value of the current sample in the current sub-block, and I(i, j) represents the revised predicted sample value of the current sample in the current sub-block. For example, original predicted sample value + incremental predicted value = revised predicted sample value. It should be understood that obtaining revised predicted sample values ​​for multiple samples (e.g., all samples) in the current sub-block is equivalent to obtaining the revised predicted sample value of the current sub-block.

[0565] It can be understood that when generating the modified predicted sample values ​​of each sub-block of the affine decoded block, the modified predicted sample values ​​of the affine decoded block are naturally generated. S1113: When at least one of the multiple optical flow decision conditions is not satisfied, set the first indicator (e.g., applyProfFlag) to false and skip the PROF process.

[0566] It is understood that if the PROF application constraint is used to determine whether to apply PROF, then step S1110 becomes: determining whether multiple PROF application constraints are not satisfied. In this case, step S1111 becomes: if multiple PROF application constraints are not satisfied, setting the first indicator (e.g., applyProfFlag) to true, and performing an optical flow prediction refinement (PROF) process on the current sub-block of the affine decoded block to obtain a modified predicted sample value of the current sub-block of the affine decoded block. Accordingly, step S1113 becomes: if at least one of multiple PROF application constraints is satisfied, setting the first indicator (e.g., applyProfFlag) to false, and skipping the PROF process.

[0567] In one implementation, if the first indicator (e.g., applyProfFlag) is a first value (e.g., 1), performing optical flow (e.g., PROF) processing on the one or more sub-blocks (e.g., each sub-block) in the current affine image block; or

[0568] If the first indicator (eg, applyProfFlag) is a second value (eg, 0), optical flow (eg, PROF) processing is not performed on the one or more sub-blocks (eg, each sub-block) in the current affine image block.

[0569] In one implementation, the value of the first indicator is determined according to whether the optical flow decision condition is satisfied, wherein the optical flow decision condition includes one or more of the following:

[0570] The first indication information (e.g., sps_prof_enabled_flag or sps_bdof_enabled_flag) is used to indicate that PROF is enabled for the current image unit. It should be noted that the current image unit here can be the current sequence, the current image, the current slice, or the current block (slice) group, etc. These examples of the current image unit are not limited.

[0571] The second indication information (e.g., fallbackModeTriggered) is used to indicate that the current affine image block is segmented;

[0572] The current affine image block is a unidirectionally predicted affine image block;

[0573] The size of the sub-block in the affine image block is greater than N×N, where N=4;

[0574] The current affine image block is a unidirectionally predicted affine image block, and a size of a subblock in the affine image block is equal to N×N, where N=4; or

[0575] The current affine image block is a bidirectionally predicted affine image block, and a size of a sub-block in the affine image block is larger than N×N, where N=4.

[0576] It should be noted that the present application includes but is not limited to the above-mentioned optical flow decision conditions, and other or different optical flow decision conditions can be set according to different application scenarios.

[0577] In the embodiment of the present application, for example, when all of the following conditions are met, applyProfFlag is set to 1:

[0578] –sps_prof_enabled_flag==1,

[0579] –fallbackModeTriggered==0,

[0580] –inter_pred_idc[x0][y0]=PRED_L0 or PRED_L1, (or predFlagL0=1, predFlagL1=0; or predFlagL1

[0581] =1, predFlagL0=0),

[0582] –Other conditions.

[0583] In another embodiment, for example, applyProfFlag is set to 1 when all of the following conditions are met:

[0584] –sps_prof_enabled_flag==1,

[0585] –fallbackModeTriggered==0,

[0586] –Other conditions.

[0587] It is understood that if PROF application constraints are used in S1110 instead of optical flow decision conditions, applyProfFlag is set to 1 when all of the following constraints are not met:

[0588] –sps_prof_disabled_flag==1,

[0589] –fallbackModeTriggered==1,

[0590] –Other conditions.

[0591] It should be understood that regarding the execution entities of the various steps in the prediction method provided in the embodiments of the present application and the addition and changes of these steps, reference can be made to the description of the corresponding method above. For the sake of brevity, this specification will not elaborate on them.

[0592] Another embodiment of the present application also provides another PROF process, including:

[0593] According to the motion information (such as motion vector) of multiple sub-blocks in the current affine image block, the first prediction matrix of the M*N block is obtained. For example, the M*N block is 16*16, such as Figure 9E As shown, for example, a 16*16 block (or a 16*16 window) includes 16 4*4 sub-blocks;

[0594] Calculating a horizontal prediction gradient matrix and a vertical prediction gradient matrix according to a second prediction matrix, wherein a size of the second prediction matrix is ​​greater than or equal to a size of the first prediction matrix, and a size of the second prediction matrix is ​​greater than or equal to a size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix;

[0595] Calculating an incremental prediction value matrix (e.g., ΔI(i, j) of a prediction signal) for the M*N block based on the horizontal prediction gradient matrix, the vertical prediction gradient matrix, and a motion vector difference between a motion vector of a current pixel unit (e.g., a current pixel or a current pixel block, e.g., a 2×2 pixel block) in the M*N block and a motion vector of a central pixel of the M*N block;

[0596] According to the incremental prediction value matrix (eg, I′(i,j)) and the first prediction matrix (eg, prediction signal I′(i,j)), a modified third prediction matrix (eg, prediction signal I′(i,j)) of the M*N block is obtained.

[0597] In another possible design, the first prediction matrix is ​​represented as I1(i, j), where the value range of i is [0, size_w–1], and the value range of j is [0, size_h–1];

[0598] The second prediction matrix is ​​represented by I2(i,j), where the value range of i is [–1, size_w], the value range of j is [–1, size_h], size_w=min(W,m), size_h=min(H,m), and m=16;

[0599] The horizontal prediction gradient matrix is ​​represented as X(i,j), where the value range of i is [0, size_w–1] and the value range of j is [0, size_h–1];

[0600] The vertical prediction gradient matrix is ​​expressed as Y(i,j), where the value range of i is [0, size_w–1] and the value range of j is [0, size_h–1].

[0601] W represents the width of the current affine image block, H represents the height of the current affine image block, and (x, y) represents the position coordinates of each sample in the current affine image block.

[0602] It should be understood that if Figure 9E As shown, in some examples, the affine image block is implicitly divided into 16×16 blocks, and the gradient matrix is ​​calculated for each 16×16 block. Accordingly, the second prediction matrix is ​​represented as I2(i,j), where the value range of i is [–1, size_w], the value range of j is [–1, size_h], size_w = min(w,m), size_h = min(h,m), and m = 16. The horizontal prediction gradient matrix is ​​represented as X(i,j), where the value range of i is [0, size_w–1], and the value range of j is [0, size_h–1]. The vertical prediction gradient matrix is ​​represented as Y(i,j), where the value range of i is [0, size_w–1], and the value range of j is [0, size_h–1].

[0603] In another possible design, calculating the horizontal prediction gradient matrix and the vertical prediction gradient matrix (e.g., gradient values ​​of size_w*size_h) according to the second prediction matrix (e.g., a prediction signal of (size_w+2)*(size_h+2)) includes:

[0604] According to the second prediction matrix, the horizontal prediction gradient matrix and the vertical prediction gradient matrix are calculated, wherein the horizontal prediction gradient matrix and the vertical prediction gradient matrix include the horizontal prediction gradient matrix and the vertical prediction gradient matrix of the sub-block, wherein

[0605] The second prediction matrix is ​​represented by I2(i,j), where the value range of i is [–1, size_w], the value range of j is [–1, size_h], size_w=min(W,m), size_h=min(H,m), and m=16;

[0606] The horizontal prediction gradient matrix is ​​represented as X(i,j), where the value range of i is [0, size_w–1] and the value range of j is [0, size_h–1];

[0607] The vertical prediction gradient matrix is ​​expressed as Y(i,j), where the value range of i is [0, size_w–1] and the value range of j is [0, size_h–1], where,

[0608] W represents the width of the current affine image block, H represents the height of the current affine image block, and (i, j) represents the position coordinates of each sample in the current affine image block.

[0609] As can be seen from the above description, the current affine image block is implicitly divided into 16×16 blocks, and the gradient matrix is ​​calculated for each 16×16 block. It should be understood that m=16 is used as an example here and should not be construed as limiting. Various other values ​​of m can be used, such as m=32.

[0610] In one possible design, the method is used for unidirectional prediction, and the motion information includes first motion information corresponding to a first reference frame list or second motion information corresponding to a second reference frame list;

[0611] The first prediction matrix includes (is) a first initial prediction matrix or a second initial prediction matrix, wherein the first initial prediction matrix is ​​obtained according to the first motion information, and the second initial prediction matrix is ​​obtained according to the second motion information;

[0612] The horizontal prediction gradient matrix includes (is) a first horizontal prediction gradient matrix or a second horizontal prediction gradient matrix, wherein the first horizontal prediction gradient matrix is ​​calculated based on an extended first initial prediction matrix, and the second horizontal prediction gradient matrix is ​​calculated based on an extended second initial prediction matrix;

[0613] The vertical prediction gradient matrix includes (is) a first vertical prediction gradient matrix or a second vertical prediction gradient matrix, wherein the first vertical prediction gradient matrix is ​​calculated based on an extended first initial prediction matrix, and the second vertical prediction gradient matrix is ​​calculated based on an extended second initial prediction matrix;

[0614] The incremental prediction value matrix includes (is) a first incremental prediction value matrix corresponding to the first reference frame list or a second incremental prediction value matrix corresponding to the second reference frame list, wherein the first incremental prediction value matrix is ​​calculated based on the first horizontal prediction gradient matrix, the first vertical prediction gradient matrix, and a first motion vector difference (e.g., a forward motion vector difference) of each sample unit in the sub-block relative to a center sample of the sub-block, and the second incremental prediction value matrix is ​​calculated based on the second horizontal prediction gradient matrix, the second vertical prediction gradient matrix, and a second motion vector difference (e.g., a backward motion vector difference) of each sample unit in the sub-block relative to a center sample of the sub-block.

[0615] In one possible design, the method is used for bidirectional prediction, and the motion information includes first motion information corresponding to a first reference frame list and second motion information corresponding to a second reference frame list;

[0616] The first prediction matrix includes a first initial prediction matrix and a second initial prediction matrix, wherein the first initial prediction matrix is ​​obtained according to the first motion information, and the second initial prediction matrix is ​​obtained according to the second motion information;

[0617] The horizontal prediction gradient matrix includes a first horizontal prediction gradient matrix and a second horizontal prediction gradient matrix, wherein the first horizontal prediction gradient matrix is ​​calculated based on an extended first initial prediction matrix, and the second horizontal prediction gradient matrix is ​​calculated based on an extended second initial prediction matrix;

[0618] The vertical prediction gradient matrix includes a first vertical prediction gradient matrix and a second vertical prediction gradient matrix, wherein the first vertical prediction gradient matrix is ​​calculated based on an extended first initial prediction matrix, and the second vertical prediction gradient matrix is ​​calculated based on an extended second initial prediction matrix;

[0619] The incremental prediction value matrix includes a first incremental prediction value matrix corresponding to the first reference frame list and a second incremental prediction value matrix corresponding to the second reference frame list, wherein the first incremental prediction value matrix is ​​calculated based on the first horizontal prediction gradient matrix, the first vertical prediction gradient matrix, and the first motion vector difference (for example, forward motion vector difference) of each sample unit in the sub-block relative to the center sample of the sub-block, and the second incremental prediction value matrix is ​​calculated based on the second horizontal prediction gradient matrix, the second vertical prediction gradient matrix, and the second motion vector difference (for example, backward motion vector difference) of each sample unit in the sub-block relative to the center sample of the sub-block.

[0620] In one possible design, the method is used for unidirectional prediction;

[0621] The motion information includes first motion information corresponding to the first reference frame list or second motion information corresponding to the second reference frame list;

[0622] The first prediction matrix includes (is) a first initial prediction matrix or a second initial prediction matrix, wherein the first initial prediction matrix is ​​obtained according to the first motion information, and the second initial prediction matrix is ​​obtained according to the second motion information.

[0623] In one possible design, the method is used for bidirectional prediction;

[0624] The motion information includes first motion information corresponding to the first reference frame list and second motion information corresponding to the second reference frame list;

[0625] The first prediction matrix includes a first initial prediction matrix and a second initial prediction matrix, wherein the first initial prediction matrix is ​​obtained according to the first motion information, and the second initial prediction matrix is ​​obtained according to the second motion information;

[0626] The acquiring, according to the motion information of the sub-block, a prediction matrix of the sub-block includes:

[0627] Performing a weighted summation on sample values ​​at the same position in the first initial prediction matrix and the second initial prediction matrix to obtain the prediction matrix for the sub-block. It should be understood that before performing the weighted summation, the sample values ​​in the first initial prediction matrix and the second initial prediction matrix may be modified separately.

[0628] In another possible design, the PROF process is described as the following four steps.

[0629] Step (1): Perform affine motion compensation based on sub-blocks to generate sub-block prediction values ​​I(i,j). For example, the value range of i is [0,subW+1] or [–1,subW], and the value range of j is [0,subH+1] or [–1,subH]. It can be understood that when the value range of i is [0,subW+1] and the value range of j is [0,subH+1], the upper left sample (or coordinate origin) is located at (1,1); and when the value range of i is [–1,subW] and the value range of j is [–1,subH], the upper left sample is located at (0,0).

[0630] Step (2): Use a 3-tap filter [–1, 0, 1] to calculate the spatial gradient g of the sub-block prediction value at each sample position x (i,j) and g x (i,j).

[0631] g x (i,j)=I(i+1,j)-I(i-1,j)

[0632] g y (i,j)=I(i,j+1)-I(i,j-1)

[0633] The sub-block predictions are extended by one sample on each side for gradient calculation. To reduce memory bandwidth and complexity, samples on the extended boundaries are copied from the nearest integer sample position in the reference image. This avoids interpolation of the padded area.

[0634] Step (3): Calculate the brightness prediction correction value through the optical flow equation.

[0635] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j)

[0636] Where Δv(i,j) is the difference between the sample MV calculated for the sample position Δv(i,j) (denoted as Δv(i,j)) and the sub-block MV of the sub-block to which the sample Δv(i,j) belongs, as Figure 10 shown.

[0637] In other words, the MV of each 4×4 center sample is calculated, and then the MV of each sample of the sub-block is calculated, and the difference Δv(i, j) between the MV of each sample and the MV of the center sample can be obtained.

[0638] Since the affine model parameters and the sample positions relative to the sub-block center do not change between sub-blocks, Δv(i,j) can be calculated for the first sub-block and then used for other sub-blocks in the same CU. Assuming x and x are the horizontal and vertical offsets of the sample position to the sub-block center, x can be derived by the following equation:

[0639]

[0640] For a 4-parameter affine model,

[0641]

[0642] For the 6-parameter affine model,

[0643]

[0644] Among them, (v0x, v0y), (v0x, v0y) and (v0x, v0y) are the motion vectors of the upper left, upper right and lower left control points, and (v0x, v0y) and (v0x, v0y) are the width and height of the CU.

[0645] Step (4): Finally, the brightness prediction correction value is added to the sub-block prediction value I(i, j). As shown in the following formula, the final prediction value I' is generated.

[0646] I′(i,j)=O(i,j)+ΔI(i,j)

[0647] Figure 15 Another aspect of the present invention provides an apparatus 1500 for performing optical flow prediction refinement (PROF) on an affine decoded block. In one example, the apparatus 1500 includes:

[0648] A determining unit 1501 is configured to determine that multiple PROF application constraints are not satisfied;

[0649] The prediction processing unit 1503 is configured to perform a prediction refinement with optical flow (PROF) process on the current sub-block of the affine decoding block to obtain a modified predicted sample value of the current sub-block of the affine decoding block. It will be understood that when the modified predicted sample value of each sub-block of the affine decoding block is generated, the modified predicted sample value of the affine decoding block is naturally generated.

[0650] In another example, the apparatus 1500 includes:

[0651] A determining unit 1501 is configured to determine whether a plurality of optical flow decision conditions are satisfied (here, the plurality of optical flow decision conditions refer to conditions under which PROF can be applied);

[0652] The prediction processing unit 1503 is configured to perform a PROF process on the current sub-block of the affine decoding block to obtain a modified predicted sample value of the current sub-block of the affine decoding block. It will be understood that when the modified predicted sample value of each sub-block of the affine decoding block is generated, the modified predicted sample value of the affine decoding block is naturally generated.

[0653] Accordingly, in one example, an exemplary structure of the apparatus 1500 may correspond to Figure 2 In another example, an exemplary structure of the apparatus 1500 may correspond to Figure 3 The decoder 30 in .

[0654] In another example, an exemplary structure of the apparatus 1500 may correspond to Figure 2 In another example, an exemplary structure of the apparatus 1500 may correspond to Figure 3 The inter-frame prediction unit 344 in .

[0655] It is understood that the determination unit and prediction processing unit (corresponding to the inter-frame prediction module) in the encoder 20 or decoder 30 provided in the embodiments of the present application are functional entities that implement the various execution steps included in the corresponding methods described above, namely, functional entities that fully implement the steps in the methods of the present application as well as extensions and variations of these steps. For details, please refer to the description of the corresponding methods described above. For the sake of brevity, they will not be repeated here.

[0656] The following explains the application of the encoding method and decoding method shown in the above embodiments and the system using these applications.

[0657] Figure 16 FIG3 is a block diagram of a content delivery system 3100 for implementing a content distribution service. The content delivery system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 via a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 includes, but is not limited to, Wi-Fi, Ethernet, wired, wireless (3G / 4G / 5G), USB, or any combination thereof.

[0658] Capture device 3102 generates data, and can encode data by the coding method shown in the above embodiment.Alternatively, capture device 3102 can distribute data to a streaming media server (not shown), and this server encodes data and sends encoded data to terminal device 3106.Capture device 3102 includes but is not limited to camera, smart phone or tablet computer, computer or notebook computer, video conferencing system, PDA, vehicle-mounted device, or any one combination thereof etc.For example, capture device 3102 can include source device 12 as described above.When data includes video, the video encoder 20 included in capture device 3102 can actually perform video coding process.When data includes audio (i.e. sound), the audio encoder included in capture device 3102 can actually perform audio coding process.For some actual scenes, capture device 3102 distributes encoded video data and encoded audio data by multiplexing encoded video data and encoded audio data together.For other actual scenes, for example, in a video conferencing system, encoded audio data and encoded video data are not multiplexed. The capture device 3102 distributes the encoded audio data and the encoded video data to the terminal device 3106, respectively.

[0659] In content delivery system 3100, terminal device 310 receives and regenerates encoded data. Terminal device 3106 may be a device capable of receiving and recovering data, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a television 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, an in-vehicle device 3124, or a combination of any of the above devices capable of decoding the encoded data. For example, terminal device 3106 may include destination device 14 described above. When the encoded data includes video, the video decoder 30 included in the terminal device prioritizes video decoding. When the encoded data includes audio, the audio decoder included in the terminal device prioritizes audio decoding.

[0660] For terminal devices with a display, such as a smartphone or tablet 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a television 3114, a personal digital assistant (PDA) 3122, or an in-vehicle device 3124, the terminal device can feed the decoded data to its display. For terminal devices without a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is connected thereto to receive and display the decoded data.

[0661] When each device in the system performs encoding or decoding, the image encoding device or the image decoding device as shown in the above-described embodiments can be used.

[0662] Figure 17 FIG3 is a schematic diagram of an exemplary structure of a terminal device 3106. After the terminal device 3106 receives a stream from the capture device 3102, the protocol processing unit 3202 analyzes the transport protocol of the stream. The protocol includes, but is not limited to, Real Time Streaming Protocol (RTSP), Hyper Text Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real-time Transport Protocol (RTP), Real-time Messaging Protocol (RTMP), or any combination thereof.

[0663] After the protocol processing unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As mentioned above, in some practical scenarios, such as in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this case, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.

[0664] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optional subtitles are generated. The video decoder 3206, including the video decoder 30 described in the above embodiment, decodes the video ES using the decoding method shown in the above embodiment to generate video frames and feeds the data to the synchronization unit 3212. The audio decoder 3208 decodes the audio ES to generate audio frames and feeds the data to the synchronization unit 3212. Optionally, the video frames can be stored in a buffer (not shown in FIG. Y) before being fed to the synchronization unit 3212. Similarly, the audio frames can be stored in a buffer (not shown in FIG. Y) before being fed to the synchronization unit 3212.

[0665] Synchronization unit 3212 synchronizes video and audio frames and provides the video / audio to video / audio display 3214. For example, synchronization unit 3212 synchronizes the presentation of video information and audio information. The information can be coded in syntax using timestamps associated with the presentation of the coded audio and visual data as well as timestamps associated with the transmission of the data stream itself.

[0666] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes the subtitles with the video frames and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.

[0667] The present invention is not limited to the above-mentioned system, and the image encoding device or the image decoding device in the above-mentioned embodiments may be included in other systems such as automobile systems.

[0668] The mathematical operators used in this application are similar to those in the C programming language; however, the results of integer division and arithmetic shift operations are more precisely defined, and other operations such as exponentiation and real-valued division are defined.

[0669] The explanation of the relevant contents, implementation methods of the relevant steps, beneficial effects, etc. in this embodiment can refer to the corresponding parts above, and simple modifications can also be made on the basis of the corresponding parts above. No further details will be given here.

[0670] It should be noted that, in the absence of conflict, some features of any two or more of the above embodiments can be combined into a new embodiment. In addition, some features of any of the above embodiments can be independently used as an embodiment.

[0671] The above content mainly introduces the technical solutions provided by the embodiments of the present application from the perspective of methods. In order to realize the above functions, corresponding hardware structures and / or software modules for performing various functions are included. Those skilled in the art should easily appreciate that, in conjunction with the various examples described in the embodiments disclosed in the specification, the present application can implement units and algorithm steps by hardware or a combination of hardware and computer software. Whether a function is performed by hardware or by hardware driven by computer software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0672] In the embodiment of the present application, the encoder / decoder can be divided into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. In actual implementation, there may be other division methods.

[0673] Although embodiments of the present invention are primarily described with respect to video decoding, it should be noted that embodiments of the decoding system 10, encoder 20, and decoder 30 (respectively, system 10), as well as other embodiments described herein, can also be used for still image processing or decoding, i.e., processing or decoding a single image in video decoding that is independent of any previous or subsequent images. Generally speaking, if image processing and decoding are limited to a single image 17, only the inter-frame prediction units 244 (encoder) and 344 (decoder) are unavailable. All other functionalities (also referred to as tools or techniques) of the video encoder 20 and video decoder 30 can also be used for still image processing, such as residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354, and / or loop filtering 220 / 320, entropy coding 270, and entropy decoding 304.

[0674] The embodiments of the encoder 20 and decoder 30, etc. and the functions described herein with reference to the encoder 20 and decoder 30, etc. can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, these functions can be stored as one or more instructions or codes in a computer-readable medium or sent via a communication medium and executed by a hardware-based processing unit. The computer-readable medium can include a computer-readable storage medium, corresponding to a tangible medium (such as a data storage medium), or any communication medium that facilitates the transfer of a computer program from one place to another according to a communication protocol, etc. In this way, the computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium can be any available medium that is accessed by one or more computers or one or more processors to retrieve instructions, codes and / or data structures for implementing the techniques described in the present invention. The computer program product can include a computer-readable medium.

[0675] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. In addition, any connection can be appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wireless technologies such as infrared, radio and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technologies such as infrared, radio and microwave are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals or other transient media, but rather refer to non-transient tangible storage media. Disks and optical disks as used herein include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0676] Instructions can be executed by one or more processors such as one or more digital signal processors (DSPs), one or more general-purpose microprocessors, one or more application-specific integrated circuits (ASICs), one or more field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein may refer to any of the above structures or any other structure suitable for implementing the techniques described herein. In addition, in some aspects, the various functions described herein may be provided within dedicated hardware and / or software modules for encoding and decoding, or incorporated into a combined codec. Moreover, these techniques may be fully implemented in one or more circuits or logic elements.

[0677] The present technology can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a set of ICs (e.g., chipsets). This disclosure describes various components, modules, or units to emphasize the functional aspects of devices for performing the disclosed technology, but they do not necessarily need to be implemented by different hardware units. Instead, as described above, the various units can be combined in a codec hardware unit in conjunction with appropriate software and / or firmware, or provided by a collection of interoperable hardware units including one or more processors as described above.

Claims

1. A method for performing optical flow prediction refinement (PROF) on an affine decoded block, characterized in that: The method comprises: When the first indicator is set to true, performing a PROF process on the current sub-block of the affine decoding block to obtain a modified prediction sample value of the current sub-block of the affine decoding block; If the affine decoding block satisfies one or more PROF application constraints of a plurality of PROF application constraints, the first indicator is set to false; if the affine decoding block does not satisfy the plurality of PROF application constraints, the first indicator is set to true; The multiple PROF application constraints include: The first indication information indicates that PROF is disabled for the image including the affine decoding block; and The second indication information indicates that the affine decoding block is not segmented; The performing of the PROF process on the current sub-block of the affine decoding block comprises: Obtaining a second prediction matrix, wherein the second prediction matrix is ​​generated according to the motion information of the current sub-block; generating a horizontal prediction gradient matrix and a vertical prediction gradient matrix according to the second prediction matrix, wherein the horizontal prediction gradient matrix and the vertical prediction gradient matrix have the same size, and the size of the second prediction matrix is ​​greater than or equal to the size of the horizontal prediction gradient matrix and the vertical prediction gradient matrix; Calculating an incremental prediction value of the current sample of the current subblock according to a horizontal prediction gradient value of the current sample in the horizontal prediction gradient matrix, a vertical prediction gradient value of the current sample in the vertical prediction gradient matrix, and a difference between a motion vector of the current sample of the current subblock and a motion vector of a center sample of the current subblock; Obtain a modified predicted sample value of the current sample of the current sub-block according to the incremental predicted value of the current sample of the current sub-block and the predicted sample value of the current sample of the current sub-block.

2. The method according to claim 1, characterized in that The method further comprises: Determining whether the affine decoding block satisfies the multiple PROF application constraints; Accordingly, If the affine decoded block satisfies one or more PROF application constraints of the multiple PROF application constraints, the first indicator is set to false, and the PROF process for the current sub-block of the affine decoded block is skipped; If the affine decoding block does not satisfy the multiple PROF application constraints, a first indicator is set to true, and a PROF process is performed on the current sub-block of the affine decoding block to obtain a modified prediction sample value of the current sub-block of the affine decoding block.

3. The method according to claim 1, characterized in that One of the multiple PROF application constraints is that the variable fallbackModeTriggered is set to 1.

4. The method according to claim 1, wherein The acquiring the second prediction matrix includes: generating a first prediction matrix according to the motion information of the current sub-block, wherein elements of the first prediction matrix correspond to predicted sample values ​​of the current sub-block, and generating the second prediction matrix according to the first prediction matrix; or, The acquiring the second prediction matrix includes: generating the second prediction matrix according to the motion information of the current sub-block.

5. The method according to claim 4, characterized in that An element of the second prediction matrix is ​​represented by I1(p,q), where the value range of p is [–1, sbW] and the value range of q is [–1, sbH]; An element of the horizontal prediction gradient matrix is ​​denoted as X(i, j) and corresponds to a sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1]; An element of the vertical prediction gradient matrix is ​​denoted as Y(i, j) and corresponds to the sample (i, j) of the current sub-block in the affine decoding block, where the value range of i is [0, sbW–1] and the value range of j is [0, sbH–1], where, sbW represents the width of the current sub-block in the affine decoding block, and sbH represents the height of the current sub-block in the affine decoding block.

6. The method according to any one of claims 1 to 3, characterized in that Before performing the PROF process on the current sub-block of the affine decoding block, the method further includes: Sub-block-based affine motion compensation is performed on the current sub-block of the affine decoding block to obtain a predicted sample value of the current sub-block.

7. An encoder (20), characterized in that The encoder (20) comprises processing circuitry for performing the method according to any one of claims 1 to 6.

8. A decoder (30), characterized in that The decoder (30) comprises processing circuitry for performing the method according to any one of claims 1 to 6.

9. A computer program product, characterized in that The computer program product comprises program instructions for executing the method according to any one of claims 1 to 6.

10. A decoder, characterized in that: The decoder comprises: one or more processors; A non-transitory computer-readable storage medium, coupled to the one or more processors and storing a program executed by the one or more processors, wherein when the one or more processors execute the program, the decoder is configured to perform the method according to any one of claims 1 to 6.

11. An encoder, characterized in that: The encoder comprises: one or more processors; A non-transitory computer-readable storage medium, coupled to the one or more processors and storing a program executed by the one or more processors, wherein when the one or more processors execute the program, the encoder is configured to perform the method according to any one of claims 1 to 6.

12. A non-transitory computer-readable storage medium comprising program instructions, characterized in that: When the program instructions are executed by a computer device, the computer device is caused to perform the method according to any one of claims 1 to 6.