Inter-frame prediction method, encoder, decoder, and computer storage medium
By using two-dimensional filters to perform point-based quadratic prediction in the inter prediction method, the existing method has solved the problem of poor encoding performance under large motion vector deviations, and achieved more efficient inter prediction and encoding and decoding efficiency.
Patent Information
- Application Number
- CN202210644545.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-29
- Filing Date
- 2021-07-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-07-13
AI Technical Summary
When the existing inter prediction method has a large deviation between the motion vector of the pixel position in the sub-block and the motion vector of the sub-block, the effect is poor, resulting in a degradation of encoding performance.
An inter prediction method is proposed, by analyzing the code stream to obtain the prediction mode parameters of the current block, determine the first motion vector of the sub-block, and perform filtering processing of the two-dimensional filter based on this, and correct the first predicted value to obtain the second predicted value.
This method can effectively improve coding performance and codec efficiency in all scenarios, and is suitable for different sub-block sizes and motion vector deviations.
Smart Images

Figure CN114866783B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application No. 202180005841.X, entitled "Inter-frame Prediction Method, Encoder, Decoder, and Computer Storage Medium", which is the national stage entry of PCT International Patent Application PCT / CN2021 / 106081 with a filing date of July 13, 2021.
[0002] Cross-reference to related applications
[0003] This application claims the priority of a Chinese patent application filed with the Chinese Patent Office on July 29, 2020, with an application number of 202010746227.6 and an application title of "Inter-frame Prediction Method, Encoder, Decoder, and Computer Storage Medium", the entire content of which is incorporated herein by reference. Technical field
[0004] This application relates to the field of video coding and decoding technologies, and in particular to an inter-frame prediction method, an encoder, a decoder, and a computer storage medium. Background art
[0005] In the field of video coding and decoding, in general, in order to balance performance and cost, affine prediction in Versatile Video Coding (VVC) and the Audio Videocoding Standard Workgroup of China (AVS) is implemented based on sub-blocks. Currently, prediction refinement with optical flow (PROF) using the optical flow principle has been proposed to correct the sub-block-based affine prediction, thereby improving the compression performance.
[0006] However, the application of PROF is based on the situation where the deviation between the motion vector of the pixel position within the sub-block and the sub-block motion vector is very small. That is to say, the optical flow calculation method of PROF is effective when the deviation between the motion vector of the pixel position within the sub-block and the sub-block motion vector is very small. However, since PROF depends on the gradients in the horizontal and vertical directions of the reference position, when the actual position is far from the reference position, the gradients in the horizontal and vertical directions of the reference position cannot truly reflect the gradients in the horizontal and vertical directions between the reference position and the actual position. Therefore, when the deviation between the motion vector of the pixel position within the sub-block and the sub-block motion vector is large, this method is not particularly effective. Summary of the invention
[0007] The present application provides an inter-frame prediction method, an encoder, a decoder, and a computer storage medium, which can greatly improve the coding performance and thus improve the encoding and decoding efficiency.
[0008] The technical solution of the present application is implemented as follows:
[0009] In a first aspect, an embodiment of the present application provides an inter-frame prediction method applied to a decoder. The method includes:
[0010] Parse the code stream to obtain the prediction mode parameter of the current block;
[0011] When the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, determine the first motion vector of the sub-blocks of the current block; wherein, the current block includes multiple sub-blocks;
[0012] Based on the first motion vector, determine the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block; wherein, the pixel position is the position of the pixel points within the sub-block;
[0013] Determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for secondary prediction processing according to a preset shape;
[0014] Based on the filtering coefficients and the first prediction value, determine the second prediction value of the sub-block, and determine the second prediction value as the inter-frame prediction value of the sub-block.
[0015] In a second aspect, an embodiment of the present application provides an inter-frame prediction method applied to an encoder. The method includes:
[0016] Determine the prediction mode parameter of the current block;
[0017] When the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, determine the first motion vector of the sub-blocks of the current block; wherein, the current block includes multiple sub-blocks;
[0018] Based on the first motion vector, determine the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block; wherein, the pixel position is the position of the pixel points within the sub-block;
[0019] Determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for secondary prediction according to a preset shape;
[0020] Based on the filtering coefficients and the first prediction value, determine the second prediction value of the sub-block, and determine the second prediction value as the inter-frame prediction value of the sub-block.
[0021] In a third aspect, an embodiment of the present application provides a decoder, which includes a parsing part and a first determination part;
[0022] The parsing part is configured to parse a bitstream and obtain a prediction mode parameter of a current block;
[0023] The first determination part is configured to, when the prediction mode parameter indicates using an inter prediction mode to determine an inter prediction value of the current block, determine a first motion vector of a sub-block of the current block; wherein, the current block includes a plurality of sub-blocks; and determine a first prediction value of the sub-block and a motion vector deviation between the pixel position and the sub-block based on the first motion vector; wherein, the pixel position is the position of a pixel point within the sub-block; and determine a filtering coefficient of a two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for performing a secondary prediction process according to a preset shape; and determine a second prediction value of the sub-block based on the filtering coefficient and the first prediction value, and determine the second prediction value as the inter prediction value of the sub-block.
[0024] In a fourth aspect, an embodiment of the present application provides a decoder, which includes a first processor and a first memory storing executable instructions of the first processor. When the instructions are executed, the first processor implements the inter prediction method as described above when executed.
[0025] In a fifth aspect, an embodiment of the present application provides an encoder, which includes a second determination part;
[0026] The second determination part is configured to determine a prediction mode parameter of a current block; and when the prediction mode parameter indicates using an inter prediction mode to determine an inter prediction value of the current block, determine a first motion vector of a sub-block of the current block; wherein, the current block includes a plurality of sub-blocks; and determine a first prediction value of the sub-block and a motion vector deviation between the pixel position and the sub-block based on the first motion vector; wherein, the pixel position is the position of a pixel point within the sub-block; and determine a filtering coefficient of a two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for using secondary prediction according to a preset shape; and determine a second prediction value of the sub-block based on the filtering coefficient and the first prediction value, and determine the second prediction value as the inter prediction value of the sub-block.
[0027] In a sixth aspect, an embodiment of the present application provides an encoder, which includes a second processor and a second memory storing executable instructions of the second processor. When the instructions are executed, the second processor implements the inter prediction method as described above when executed.
[0028] In a seventh aspect, an embodiment of the present application provides a computer storage medium storing a computer program, which, when executed by a first processor and a second processor, implements the inter-frame prediction method as described above.
[0029] An inter-frame prediction method, an encoder, a decoder, and a computer storage medium provided by an embodiment of the present application. The decoder parses a bitstream to obtain a prediction mode parameter of a current block; when the prediction mode parameter indicates using an inter-frame prediction mode to determine an inter-frame prediction value of the current block, a first motion vector of a sub-block of the current block is determined; wherein the current block includes multiple sub-blocks; a first prediction value of the sub-block and a motion vector deviation between a pixel position and the sub-block are determined based on the first motion vector; wherein the pixel position is the position of a pixel point within the sub-block; a filtering coefficient of a two-dimensional filter is determined according to the motion vector deviation; wherein the two-dimensional filter is used for performing a secondary prediction process according to a preset shape; a second prediction value of the sub-block is determined based on the filtering coefficient and the first prediction value, and the second prediction value is determined as the inter-frame prediction value of the sub-block. The encoder determines a prediction mode parameter of the current block; when the prediction mode parameter indicates using an inter-frame prediction mode to determine an inter-frame prediction value of the current block, a first motion vector of a sub-block of the current block is determined; wherein the current block includes multiple sub-blocks; a first prediction value of the sub-block and a motion vector deviation between a pixel position and the sub-block are determined based on the first motion vector; wherein the pixel position is the position of a pixel point within the sub-block; a filtering coefficient of a two-dimensional filter is determined according to the motion vector deviation; wherein the two-dimensional filter is used for using a secondary prediction according to a preset shape; a second prediction value of the sub-block is determined based on the filtering coefficient and the first prediction value, and the second prediction value is determined as the inter-frame prediction value of the sub-block. That is to say, for the inter-frame prediction method proposed in the present application, after prediction based on sub-blocks, for pixel positions with a deviation between the motion vector and the motion vector of the sub-block, point-based secondary prediction can be performed on the basis of the first prediction value of the sub-block to obtain a second prediction value. Among them, for the point-based secondary prediction using a two-dimensional filter, the information of several points constituting a preset shape can be used to correct the first prediction value, and finally the corrected second prediction value is obtained. The inter-frame prediction method proposed in the present application can be well applied to all scenarios, greatly improving the coding performance, thereby improving the encoding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Schematic diagram of an affine model Figure 1 ;
[0031] Figure 2 Schematic diagram of an affine model Figure 2 ;
[0032] Figure 3 Schematic diagram of pixel interpolation;
[0033] Figure 4 Schematic diagram of sub-block interpolation Figure 1 ;
[0034] Figure 5 Schematic diagram of sub-block interpolation Figure 2 ;
[0035] Figure 6 Schematic diagram of motion vectors for each sub-block;
[0036] Figure 7 Schematic diagram of sample positions;
[0037] Figure 8 Schematic diagram of positions of affine-predicted luminance samples;
[0038] Figure 9 Schematic diagram of positions of affine-predicted integer-pixel samples and fractional-pixel samples;
[0039] Figure 10 Schematic diagram of positions of affine-predicted chrominance samples;
[0040] Figure 11 Schematic diagram of positions of affine-predicted integer-pixel samples and fractional-pixel samples;
[0041] Figure 12 Schematic diagram of the process of affine prediction;
[0042] Figure 13 Schematic diagram of the process of PROF-corrected predicted values Figure 1 ;
[0043] Figure 14 Schematic diagram of the process of PROF-corrected predicted values Figure 2 ;
[0044] Figure 15 Schematic diagram of the block diagram of a video coding system provided by an embodiment of the present application;
[0045] Figure 16 Schematic diagram of the block diagram of a video decoding system provided by an embodiment of the present application;
[0046] Figure 17 Schematic diagram of the implementation process of the inter-frame prediction method Figure 1 ;
[0047] Figure 18 Schematic diagram of a two-dimensional filter Figure 1 ;
[0048] Figure 19 Schematic diagram of a two-dimensional filter Figure 2 ;
[0049] Figure 20 Schematic diagram of the implementation process of the inter-frame prediction methodFigure 2 ;
[0050] Figure 21 Schematic diagram of the implementation process of the inter-frame prediction method Figure 3 ;
[0051] Figure 22 Schematic diagram of the implementation process of the inter-frame prediction method Figure 4 ;
[0052] Figure 23 Schematic diagram of filtering the boundary pixel positions;
[0053] Figure 24 Schematic diagram of the expansion of the boundary;
[0054] Figure 25 Schematic diagram of the implementation process of the inter-frame prediction method Figure 5 ;
[0055] Figure 26 Schematic diagram of the composition structure of the decoder Figure 1 ;
[0056] Figure 27 Schematic diagram of the composition structure of the decoder Figure 2 ;
[0057] Figure 28 Schematic diagram of the composition structure of the encoder Figure 1 ;
[0058] Figure 29 Schematic diagram of the composition structure of the encoder Figure 2 . Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the related application, rather than limiting the application. Additionally, it should be noted that for the sake of description, only the parts related to the related application are shown in the accompanying drawings.
[0060] In video images, generally, the first image component, the second image component, and the third image component are used to represent the current block (Coding Block, CB); among them, these three image components are respectively a luminance component, a blue chrominance component, and a red chrominance component. Specifically, the luminance component is usually represented by the symbol Y, the blue chrominance component is usually represented by the symbol Cb or U, and the red chrominance component is usually represented by the symbol Cr or V; thus, the video image can be represented in the YCbCr format or the YUV format.
[0061] Currently, general video coding standards all adopt a block-based hybrid coding framework. Each frame in a video image is divided into square largest coding units (LCUs) of the same size (such as 128×128, 64×64, etc.), and each largest coding unit can also be divided into rectangular coding units (CUs) according to rules; moreover, the coding units may also be divided into smaller prediction units (PUs). Specifically, the hybrid coding framework may include modules such as prediction, transform, quantization, entropy coding, and in-loop filter; among them, the prediction module may include intra prediction and inter prediction, and the inter prediction may include motion estimation and motion compensation. Since there is a strong correlation between adjacent pixels within a frame of a video image, using the intra prediction method in video coding technology can eliminate the spatial redundancy between adjacent pixels; however, since there is also a strong similarity between adjacent frames in a video image, using the inter prediction method in video coding technology can eliminate the temporal redundancy between adjacent frames, thereby improving the coding efficiency. The following description of this application will focus on inter prediction in detail.
[0062] Inter-frame prediction uses already encoded / decoded frames to predict the parts to be encoded / decoded in the current frame. In a block-based encoding / decoding framework, the parts to be encoded / decoded are usually coding units or prediction units. Here, the coding units or prediction units to be encoded / decoded are collectively referred to as the current block. Translational motion is a common and simple motion pattern in videos, so translational prediction is also a traditional prediction method in video encoding / decoding. Translational motion in a video can be understood as a part of the content moving from a certain position on one frame to a certain position on another frame over time. A simple unidirectional prediction of translation can be represented by a motion vector (MV) between a certain frame and the current frame. The certain frame mentioned here is a reference frame of the current frame. Through this motion information containing the reference frame and the motion vector, the current block can find a reference block on the reference frame that has the same size as the current block, and use this reference block as the prediction block for the current block. In ideal translational motion, the content of the current block does not change in terms of deformation, rotation, etc., nor in terms of brightness and color between different frames. However, the content in videos does not always conform to such an ideal situation. Bidirectional prediction can solve the above problems to a certain extent. Usually, bidirectional prediction refers to bidirectional translational prediction. Bidirectional prediction is to use the motion information of two reference frames and motion vectors to find two reference blocks with the same size as the current block from the two reference frames (the two reference frames may be the same reference frame) respectively, and generate the prediction block for the current block using these two reference blocks. The generation methods include averaging, weighted averaging, and some other calculations, etc.
[0063] In this application, prediction can be considered as a part of motion compensation. Some documents will call the prediction in this application motion compensation. For example, the affine prediction in this application is called affine motion compensation in some documents.
[0064] Rotation, zooming in, zooming out, distortion, deformation, etc. are also common changes in videos. However, ordinary translational prediction cannot handle such changes well. Therefore, the affine prediction model is applied to video encoding / decoding, such as the affine in VVC and AVS. Among them, the affine prediction models of VVC and AVS3 are similar. In changes such as rotation, zooming in, zooming out, distortion, deformation, etc., it can be considered that not all points of the current block use the same MV. Therefore, it is necessary to derive the MV for each point. The affine prediction model derives the MV for each point through calculation using a small number of parameters. The affine prediction models of VVC and AVS3 both use the 2-control-point (4-parameter) and 3-control-point (6-parameter) models. The 2 control points are the upper left corner and the upper right corner of the current block, and the 3 control points are the upper left corner, the upper right corner, and the lower left corner of the current block. Exemplarily, Figure 1 is a schematic diagram of the affine model Figure 1 , Figure 2 is a schematic diagram of the affine modelFigure 2 , as shown in Figure 1 and 2 . Since each MV includes an x - component and a y - component, 2 control points have 4 parameters, and 3 control points have 6 parameters.
[0065] According to the affine prediction model, an MV can be derived for each pixel position. Each pixel position can find its corresponding position in the reference frame. If this position is not an integer - pixel position, then the value of this fractional - pixel position needs to be obtained by interpolation. Currently, the interpolation methods used in video coding and decoding standards are usually implemented by finite - impulse - response (FIR) filters, and the complexity (cost) of implementing in this way is very high. For example, in AVS3, an 8 - tap interpolation filter is used for the luminance component, and the fractional - pixel accuracy in the normal mode is 1 / 4 pixel, and the fractional - pixel accuracy in the affine mode is 1 / 16 pixel. For each fractional - pixel point that meets the 1 / 16 - pixel accuracy, 8 integer - pixels in the horizontal direction and 8 integer - pixels in the vertical direction, that is, 64 integer - pixels, are required for interpolation. Figure 3 Schematic diagram of pixel interpolation, as shown in Figure 3 . The circular pixel is the fractional - pixel point to be obtained, the dark - colored square pixel is the position of the integer - pixel corresponding to this fractional - pixel, and the vector between the two is the motion vector of the fractional - pixel. The light - colored square pixels are the pixels required for interpolating the circular fractional - pixel position. To obtain the value of this fractional - pixel position, the pixel values of these 8x8 light - colored square pixel regions need to be interpolated, including the dark - colored pixel positions.
[0066] In traditional translational prediction, the MV of each pixel position in the current block is the same. If the concept of sub - blocks is further introduced, the sizes of sub - blocks are such as 4x4, 8x8, etc. Figure 4 Schematic diagram of sub - block interpolation Figure 1 , the pixel region required for interpolating a 4x4 block is as shown in Figure 4 . Figure 5 Schematic diagram of sub - block interpolation Figure 2 , the pixel region required for interpolating an 8x8 block is as shown in Figure 5 .
[0067] If the MV of each pixel position in a sub - block is the same, then the pixel positions in a sub - block can be interpolated together, thus sharing the bandwidth, using the filter of the same phase, and sharing the intermediate values of the interpolation process. However, if each pixel point uses an MV, then the bandwidth will increase, and different - phase filters may be used and the intermediate values of the interpolation process cannot be shared.
[0068] Point-based affine prediction is very costly. Therefore, to balance performance and cost, affine prediction in VVC and AVS3 is implemented based on sub-blocks. The sub-block sizes in AVS3 are 4x4 and 8x8, while VVC uses a 4x4 sub-block size. Each sub-block has an MV, and the pixel positions within the sub-block share the same MV. Thus, interpolation is uniformly performed for all pixel positions within the sub-block. Through the above method, the motion compensation complexity of sub-block-based affine prediction is similar to that of other sub-block-based prediction methods.
[0069] It can be seen that in the sub-block-based affine prediction method, the pixel positions within the sub-block share the same MV. Among them, the method to determine this shared MV is to take the MV at the center of the current sub-block. For sub-blocks with at least one even number of pixels in the horizontal and vertical directions, such as 4x4 and 8x8, their centers actually fall on non-integer pixel positions. In current standards, an integer pixel position is taken. For example, for a 4x4 sub-block, the pixel position at a distance of (2, 2) from the upper left corner is taken. For an 8x8 sub-block, the pixel position at a distance of (4, 4) from the upper left corner is taken.
[0070] The affine prediction model can derive the MV for each pixel position based on the control points (2 control points or 3 control points) used by the current block. In sub-block-based affine prediction, the MV for this position is calculated according to the pixel position described in the previous paragraph and used as the MV of the sub-block. Figure 6 Schematic diagram of the motion vector for each sub-block, as Figure 6 shown. To derive the motion vector for each sub-block, the motion vector sampled at the center of each sub-block is shown in the figure, rounded to 1 / 16 precision, and then motion compensation is performed.
[0071] With the development of technology, a method called Prediction Refinement using Optical Flow (PROF) has been proposed. This technology can improve the prediction value of block-based affine prediction without increasing the bandwidth. After the sub-block-based affine prediction is completed, the horizontal and vertical gradients of each pixel point that has completed the sub-block-based affine prediction are calculated. When calculating the gradients in PROF in VVC, a 3-tap filter [-1, 0, 1] is used, and the calculation method is the same as that of Bi-directional Optical flow (BDOF). Then, for each pixel position, the motion vector deviation is calculated, which is the difference between the motion vector of the current pixel position and the MV used for the entire sub-block. These motion vector deviations can all be calculated according to the formula of the affine prediction model. Due to the characteristics of the formula, the motion vector deviations at the same position of some sub-blocks are the same. For these sub-blocks, only a set of motion vector deviations needs to be calculated, and other sub-blocks can directly reuse these values. For each pixel position, the correction value of the prediction value at this pixel position is calculated using the horizontal and vertical gradients and the motion vector deviation (including the horizontal deviation and the vertical deviation) of this point. Then, the original prediction value, that is, the prediction value of the sub-block-based affine prediction, is added to the correction value of the prediction value to obtain the corrected prediction value.
[0072] When calculating the horizontal and vertical gradients, a [-1, 0, 1] filter is used, that is, for the current pixel position, the prediction values of the pixel positions with a distance of one to the left and a distance of one to the right in the horizontal direction will be used, and the prediction values of the pixel positions with a distance of one above and a distance of one below in the vertical direction will be used. If the current pixel position is the boundary position of the current block, then some of the above pixel positions will exceed the boundary of the current block by one pixel distance. The prediction values at the boundary of the current block are used to fill the positions with a one-pixel distance outward to meet the gradient calculation, so there is no need to additionally increase the prediction values that exceed the boundary of the current block by one pixel distance. Since the gradient calculation only needs to use the prediction values of the sub-block-based affine prediction, there is no need to increase additional bandwidth.
[0073] In the VVC standard text, the MVs of each sub-block of the current block and the motion vector deviations of each pixel position within the sub-block are derived from the MV of the control point. In VVC, the pixel positions used as the sub-block MV for each sub-block are the same, so only a set of motion vector deviations of the sub-blocks needs to be derived, and other sub-blocks can reuse this sub-block. Further, in the VVC standard text, for the description of the PROF process, the calculation of the motion vector deviation by PROF is included in the above process.
[0074] When calculating the MV of a sub-block in AVS3's affine prediction, the basic principle is the same as that of VVC. However, AVS3 has special processing for the upper-left sub-block A, upper-right sub-block B, and lower-left sub-block C of the current block.
[0075] The following is the description in the AVS3 standard text for the derivation of the motion vector array of the affine motion unit sub-block:
[0076] If there are 3 motion vectors in the affine control point motion vector group, then the motion vector group can be expressed as mvsAffine(mv0, mv1, mv2); if there are 2 motion vectors in the affine control point motion vector group, then the motion vector group can be expressed as mvsAffine(mv0, mv1). Then, the motion vector array of the affine motion unit sub-block can be derived according to the following steps:
[0077] 1. Calculate the variables dHorX, dVerX, dHorY, and dVerY:
[0078] dHorX = (mv1_x - mv0_x) << (7 - Log(width));
[0079] dHorY = (mv1_y - mv0_y) << (7 - Log(width));
[0080] If the motion vector group is mvsAffine(mv0, mv1, mv2), then:
[0081] dVerX = (mv2_x - mv0_x) << (7 - Log(height));
[0082] dVerY = (mv2_y - mv0_y) << (7 - Log(height));
[0083] If the motion vector group is mvsAffine(mv0, mv1), then:
[0084] dVerX = -dHorY;
[0085] dVerY = dHorX;
[0086] It should be noted that Figure 7 is a schematic diagram of the sample position, such as Figure 7As shown, (xE, yE) is the position of the top-left sample of the current prediction unit's luminance prediction block in the luminance sample matrix of the current image. The width and height of the current prediction unit are width and height respectively, the width and height of each sub-block are subwidth and subheight respectively. The sub-block where the top-left sample of the current prediction unit's luminance prediction block is located is A, the sub-block where the top-right sample is located is B, and the sub-block where the bottom-left sample is located is C.
[0087] 2.1. If the prediction reference mode of the current prediction unit is 'Pred_List01' or AffineSubblockSizeFlag is equal to 1 (AffineSubblockSizeFlag is used to indicate the size of the sub-block), then both subwidth and subheight are equal to 8, (x, y) are the coordinates of the top-left position of an 8x8 sub-block, and the motion vector mvE (mvE_x, mvE_y) of each 8x8 luminance sub-block can be calculated as follows:
[0088] If the current sub-block is A, then both xPos and yPos are equal to 0;
[0089] If the current sub-block is B, then xPos is equal to width and yPos is equal to 0;
[0090] If the current sub-block is C and there are 3 motion vectors in mvsAffine, then xPos is equal to 0 and yPos is equal to height;
[0091] Otherwise, xPos is equal to (x - xE) + 4 and yPos is equal to (y - yE) + 4;
[0092] Therefore, the motion vector mvE of the current 8x8 sub-block is:
[0093] mvE_x = Clip3(-131072, 131071, Rounding((mv0_x << 7) + dHorX × xPos + dVerX × yPos, 7));
[0094] mvE_y = Clip3(-131072, 131071, Rounding((mv0_y << 7) + dHorY × xPos + dVerY × yPos, 7));
[0095] 2.2. If the prediction reference mode of the current prediction unit is 'Pred_List0' or 'Pred_List1', and AffineSubblockSizeFlag is equal to 0, then both subwidth and subheight are equal to 4, (x, y) are the coordinates of the upper-left corner position of the sub-block with a size of 4x4, and calculate the motion vector mvE (mvE_x, mvE_y) of each 4x4 luma sub-block:
[0096] If the current sub-block is A, then both xPos and yPos are equal to 0;
[0097] If the current sub-block is B, then xPos is equal to width and yPos is equal to 0;
[0098] If the current sub-block is C and there are 3 motion vectors in mvAffine, then xPos is equal to 0 and yPos is equal to height;
[0099] Otherwise, xPos is equal to (x - xE) + 2 and yPos is equal to (y - yE) + 2;
[0100] Therefore, the motion vector mvE of the current 4x4 sub-block is:
[0101] mvE_x = Clip3(-131072, 131071, Rounding((mv0_x << 7) + dHorX × xPos + dVerX × yPos, 7));
[0102] mvE_y = Clip3(-131072, 131071, Rounding((mv0_y << 7) + dHorY × xPos + dVerY × yPos, 7)).
[0103] The following is the description of the affine prediction sample derivation and luma / chroma sample interpolation in AVS3 text:
[0104] If the position of the top-left sample of the luma prediction block of the current prediction unit in the luma sample matrix of the current image is (xE, yE).
[0105] If the prediction reference mode of the current prediction unit is 'PRED_List0' and the value of AffineSubblockSizeFlag is 0, mv0E0 is the LO motion vector of the 4x4 unit at the position (xE + x, yE + y) in the MvArrayL0 motion vector set. The value of the element predMatrixL0[x][y] in the luminance prediction sample matrix predMatrixL0 is the sample value at the position (((xE + x) << 4) + mv0E0_x, ((yE + y) << 4) + mv0E0_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL0 in the reference image list 0. The value of the element predMatrixL0[x][y] in the chrominance prediction sample matrix predMatrixL0 is the sample value at the position (((xE + 2x) << 4) + MvC_x, ((yE + 2y) << 4) + MvC_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL0 in the reference image list 0. Where x1 = ((xE + 2x) >> 3) << 3, y1 = ((yE + 2y) >> 3) << 3, mv1E0 is the LO motion vector of the 4x4 unit at the position (x1, y1) in the MvArrayL0 motion vector set, mv2E0 is the LO motion vector of the 4x4 unit at the position (x1 + 4, y1) in the MvArrayL0 motion vector set, mv3E0 is the LO motion vector of the 4x4 unit at the position (x1, y1 + 4) in the MvArrayL0 motion vector set, and mv4E0 is the LO motion vector of the 4x4 unit at the position (x1 + 4, y1 + 4) in the MvArrayL0 motion vector set.
[0106] MvC_x = (mv1E0_x + mv2E0_x + mv3E0_x + mv4E0_x + 2) >> 2
[0107] MvC_y = (mv1E0_y + mv2E0_y + mv3E0_y + mv4E0_y + 2) >> 2
[0108] If the prediction reference mode of the current prediction unit is 'PRED_List0' and the value of AffineSubblockSizeFlag is 1, mv0E0 is the LO motion vector of an 8x8 unit at the position (xE + x, yE + y) in the MvArrayL0 motion vector set. The value of the element predMatrixL0[x][y] in the luminance prediction sample matrix predMatrixL0 is the sample value at the position (((xE + x) << 4) + mv0E0_x, ((yE + y) << 4) + mv0E0_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL0 in the reference picture list 0. The value of the element predMatrixL0[x][y] in the chrominance prediction sample matrix predMatrixL0 is the sample value at the position (((xE + 2x) << 4) + MvC_x, ((yE + 2y) << 4) + MvC_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL0 in the reference picture list 0. Wherein, MvC_x is equal to mv0E0_x, and MvC_y is equal to mv0E0.
[0109] If the prediction reference mode of the current prediction unit is 'PRED_List1' and the value of AffineSubblockSizeFlag is 0, mv0E1 is the L1 motion vector of the 4x4 unit at the position (xE + x, yE + y) in the MvArrayL1 motion vector set. The value of the element predMatrixL1[x][y] in the luminance prediction sample matrix predMatrixL1 is the sample value at the position (((xE + x) << 4) + mv0E1_x, ((yE + y) << 4) + mv0E1_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL1 in the reference picture list 1. The value of the element predMatrixL1[x][y] in the chrominance prediction sample matrix predMatrixL1 is the sample value at the position (((xE + 2x) << 4) + MvC_x, ((yE + 2y) << 4) + MvC_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL1 in the reference picture list 1. Where x1 = ((xE + 2x) >> 3) << 3, y1 = ((yE + 2y) >> 3) << 3, mv1E1 is the L1 motion vector of the 4x4 unit at the position (x1, y1) in the MvArrayL1 motion vector set, mv2E1 is the L1 motion vector of the 4x4 unit at the position (x1 + 4, y1) in the MvArrayL1 motion vector set, mv3E1 is the L1 motion vector of the 4x4 unit at the position (x1, y1 + 4) in the MvArrayL1 motion vector set, and mv4E1 is the L1 motion vector of the 4x4 unit at the position (x1 + 4, y1 + 4) in the MvArrayL1 motion vector set.
[0110] MvC_x = (mv1E1_x + mv2E1_x + mv3E1_x + mv4E1_x + 2) >> 2
[0111] MvC_y = (mv1E1_y + mv2E1_y + mv3E1_y + mv4E1_y + 2) >> 2
[0112] If the prediction reference mode of the current prediction unit is 'PRED_List1' and the value of AffineSubblockSizeFlag is 1, mv0E1 is the L1 motion vector of the 8x8 unit at the position (xE + x, yE + y) in the MvArrayL1 motion vector set. The value of the element predMatrixL1[x][y] in the luminance prediction sample matrix predMatrixL1 is the sample value at the position (((xE + x) << 4) + mv0E1_x, ((yE + y) << 4) + mv0E1_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL1 in the reference picture list 1. The value of the element predMatrixL1[x][y] in the chrominance prediction sample matrix predMatrixL1 is the sample value at the position (((xE + 2x) << 4) + MvC_x, ((yE + 2y) << 4) + MvC_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL1 in the reference picture list 1. Where MvC_x is equal to mv0E1_x and MvC_y is equal to mv0E1.
[0113] If the prediction mode of the current prediction unit is 'PRED_List01', mv0E0 is the L0 motion vector of the 8x8 unit at the position (xE + x, yE + y) in the MvArrayL0 motion vector set, and mv0E1 is the L1 motion vector of the 8x8 unit at the position (x, y) in the MvArrayL1 motion vector set. The value of the element predMatrixL0[x][y] in the luminance prediction sample matrix predMatrixL0 is the sample value at the position ((((xE + x) << 4) + mv0E0_x, ((yE + y) << 4) + mv0E0_y)) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL0 in the reference image queue 0. The value of the element predMatrixL0[x][y] in the chrominance prediction sample matrix predMatrixL0 is the sample value at the position ((((xE + 2x) << 4) + MvC0_x, ((yE + 2y) << 4) + MvC0_y)) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL0 in the reference image queue 0. The value of the element predMatrixL1[x][y] in the luminance prediction sample matrix predMatrixL1 is the sample value at the position ((((xE + x) << 4) + mv0E1_x, ((yE + y) << 4) + mv0E1_y)) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL1 in the reference image queue 1. The value of the element predMatrixL1[x][y] in the chrominance prediction sample matrix predMatrixL1 is the sample value at the position ((((xE + 2x) << 4) + MvC1_x, ((yE + 2y) << 4) + MvC1_y)) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL1 in the reference image queue 1. Where MvC0_x is equal to mv0E0_x, MvC0_y is equal to mv0E0_y, MvC1_x is equal to mv0E1_x, and MvC1_y is equal to mv0E1_y.
[0114] Among them, the element values at each position in the 1 / 16 precision luminance sample matrix and the 1 / 32 precision chrominance sample matrix of the reference image are obtained by the interpolation method defined by the following affine luminance sample interpolation process and affine chrominance sample interpolation process. Integer samples outside the reference image should be replaced by the integer sample (edge or corner sample) closest to this sample within the image, that is, the motion vector can point to samples outside the reference image.
[0115] Specifically, the affine luminance sample interpolation process is as follows:
[0116] Figure 8 For the position schematic diagram of the affine predicted luminance sample, as Figure 8As shown, A, B, C, and D are adjacent integer pixel samples. dx and dy are the horizontal and vertical distances between the fractional pixel sample a(dx, dy) and the integer pixel sample A around A. dx is equal to fx & 15, and dy is equal to fy & 15, where (fx, fy) are the coordinates of the fractional pixel sample in the luminance sample matrix with a 1 / 16 precision. Figure 9 It is a schematic diagram of the positions of the affine prediction integer pixel samples and fractional pixel samples. The specific positions of the integer pixel Ax,y and the 255 fractional pixel samples ax,y(dx, dy) around it are as Figure 9 shown.
[0117] Specifically, the sample position ax, 0 (x = 1 to 15) is filtered by 8 integer values that are the closest to the interpolation point in the horizontal direction. The method for obtaining the predicted value is as follows:
[0118] ax, 0 = Clip1((fL[x][0] × A - 3 , 0 + fL[x][1] × A - 2 , 0 + fL[x][2] × A - 1 , 0 + fL[x][3] × A 0 , 0 + fL[x][4] × A 1 , 0 + fL[x][5] × A 2 , 0 + fL[x][6] × A 3 , 0 + fL[x][7] × A 4 , 0 + 32) >> 6).
[0119] Specifically, the sample position a 0 , y (y = 1 to 15) is filtered by 8 integer values that are the closest to the interpolation point in the vertical direction. The method for obtaining the predicted value is as follows:
[0120] a 0 , y = Clip1((fL[y][0] × A 0 , - 3 + fL[y][1] × A - 2 , 0 + fL[y][2] × A - 1 , 0 + fL[y][3] × A 0 , 0 + fL[y][4] × A 1 , 0+fL[y][5]×A 2 , 0 +fL[y][6]×A 3 , 0 +fL[y][7]×A - 4 , 0 +32)>>6)。
[0121] Specifically, the predicted values of the sample positions ax,y (x = 1 to 15, y = 1 to 15) are obtained as follows:
[0122] ax,y = Clip1((fL[y][0]×a'x,y - 3 +fL[y][1]×a'x,y - 2 +fL[y][2]×a'x,y - 1 +fL[y][3]×a'x,y + fL[y][4]×a'x,y +1 +fL[y][5]×a'x,y + 2 + fL[y][6]×a'x,y +3 +fL[y][7]×a'x,y +4 +(1<<(19 - BitDepth)))>>(20 - BitDepth))。
[0123] Where:
[0124] a'x,y = (fL[x][0]×A - 3 ,y + fL[x][1]×A - 2 ,y + fL[x][2]×A - 1 ,y + fL[x][3]×A 0 ,y + fL[x][4]×A 1 ,y + fL[x][5]×A 2 ,y + fL[x][6]×A 3 ,y + fL[x][7]×A 4 ,y + ((1<<(BitDepth - 8))>>1))>>(BitDepth - 8)。
[0125] The luminance interpolation filter coefficients are shown in Table 1:
[0126] Table 1
[0127]
[0128] Specifically, the affine chrominance sample interpolation process is as follows:
[0129] Figure 10Schematic diagram of the position for affine prediction of chrominance samples, as Figure 10 shown, A, B, C, D are adjacent integer-pixel samples, dx and dy are the horizontal and vertical distances between the fractional-pixel sample a(dx, dy) around the integer-pixel sample A and A, dx is equal to fx&31, dy is equal to fy&31, where (fx, fy) are the coordinates of the fractional-pixel sample in the chrominance sample matrix with 1 / 32 precision. Figure 11 Schematic diagram of the positions of affine prediction integer-pixel samples and fractional-pixel samples, integer-pixel A x,y and 1023 fractional-pixel samples a x,y (dx, dy) around it are shown specifically as Figure 11 shown.
[0130] Specifically, for fractional-pixel points where dx is equal to 0 or dy is equal to 0, they can be directly obtained by chrominance integer-pixel interpolation. For points where dx is not equal to 0 and dy is not equal to 0, the fractional-pixels on the integer-pixel row (dy equal to 0) are used for calculation:
[0131] if(dx == 0){
[0132] a x,y (0, dy) = Clip3(0, (1<<BitDepth)-1, (fC[dy][0]×A x,y - 1 + fC[dy][1]×A x,y + fC[dy][2]×A x,y + 1 + fC[dy][3]×A x,y + 2 + 32)>>6)
[0133] }}
[0134] else if(dy == 0){
[0135] a x,y (dx, 0) = Clip3(0, (1<<BitDepth)-1, (fC[dx][0]×A x-1,y + fC[dx][1]×A x,y + fC[dx][2]×A x+1,y + fC[dx][3]×A x+2,y+32 )>>6)
[0136] }}
[0137] else{
[0138] a x,y (dx, dy) = Clip3(0, (1<<BitDepth)-1, (C[dy][0]×a' x,y-1 (dx, 0) + C[dy][1]×a'x,y (dx, 0) + C[dy][2] × a' x,y +1
[0139] (dx, 0) + C[dy][3] × a' x,y +2(dx, 0) + (1 << (19 - BitDepth))) >> (20 - BitDepth))
[0140] }
[0141] where a' x,y (dx, 0) is the temporary value of the sub - pixel on the integer - pixel row, defined as: a' x,y (dx, 0) = (fC[dx][0] × A x-1,y + fC[dx][1] × A x,y + fC[dx][2] × A x+1,y+ fC[dx][3] × A x+2,y + ((1 << (BitDepth - 8)) >> 1)) >> (BitDepth - 8).
[0142] The chrominance interpolation filter coefficients are shown in Table 2:
[0143] Table 2
[0144]
[0145]
[0146] Figure 12 is a schematic diagram of the process of affine prediction, as Figure 12 shown. The common methods of affine prediction may include the following steps:
[0147] Step 101: Determine the motion vector of the control point.
[0148] Step 102: Determine the motion vector of the sub - block according to the motion vector of the control point.
[0149] Step 103: Predict the sub - block according to the motion vector of the sub - block.
[0150] Currently, Figure 13 is a schematic diagram of the process for correcting the prediction value by PROF Figure 1 , as Figure 13 shown. When improving the prediction value of block - based affine prediction through PROF, it may specifically include the following steps:
[0151] Step 101: Determine the motion vector of the control point.
[0152] Step 102: Determine the motion vector of the sub - block according to the motion vector of the control point.
[0153] Step 103: Predict the sub-block according to the motion vector of the sub-block.
[0154] Step 104: Determine the deviation of the motion vector of each position within the sub-block from the motion vector of the sub-block according to the motion vector of the control point and the motion vector of the sub-block.
[0155] Step 105: Determine the motion vector of the sub-block according to the motion vector of the control point.
[0156] Step 106: Use the prediction value based on the sub-block to derive the gradients in the horizontal and vertical directions for each position.
[0157] Step 107: Utilize the optical flow principle to calculate the deviation value of the prediction value for each position according to the motion vector deviation of each position and the gradients in the horizontal and vertical directions.
[0158] Step 108: Add the deviation value of the prediction value to the prediction value based on the sub-block for each position to obtain the corrected prediction value.
[0159] Currently, Figure 14 is the process schematic for PROF to correct the prediction value Figure 2 , as Figure 14 shown, when improving the prediction value of the block-based affine prediction through PROF, the following steps may specifically be included:
[0160] Step 101: Determine the motion vector of the control point.
[0161] Step 109: Determine the motion vector of the sub-block and the deviation of the motion vector of each position within the sub-block from the motion vector of the sub-block according to the motion vector of the control point.
[0162] Step 103: Predict the sub-block according to the motion vector of the sub-block.
[0163] Step 106: Use the prediction value based on the sub-block to derive the gradients in the horizontal and vertical directions for each position.
[0164] Step 107: Utilize the optical flow principle to calculate the deviation value of the prediction value for each position according to the motion vector deviation of each position and the gradients in the horizontal and vertical directions.
[0165] Step 108: Add the deviation value of the prediction value to the prediction value based on the sub-block for each position to obtain the corrected prediction value.
[0166] The predictive correction PROF using the optical flow principle can correct the block-based affine prediction using the optical flow principle, improving the compression performance. However, the application of PROF is based on the situation where the deviation between the motion vector of the pixel positions within the block and the block motion vector is very small. That is to say, the optical flow calculation method of PROF is effective when the deviation between the motion vector of the pixel positions within the block and the block motion vector is very small. However, since PROF depends on the gradients in the horizontal and vertical directions of the reference position, when the actual position is far from the reference position, the gradients in the horizontal and vertical directions of the reference position cannot truly reflect the gradients in the horizontal and vertical directions between the reference position and the actual position. Therefore, when the deviation between the motion vector of the pixel positions within the block and the block motion vector is large, this method is not particularly effective.
[0167] Thus, it can be seen that the existing method for correcting the prediction value of PROF is not rigorous. When improving the affine prediction, it cannot be well applied to all scenarios, and the coding performance needs to be improved.
[0168] To solve the defects existing in the prior art, in the embodiments of the present application, after the block-based prediction, for the pixel positions where the motion vector deviates from the motion vector of the block, a point-based secondary prediction can be performed on the basis of the first prediction value of the block to obtain a second prediction value. Among them, the point-based secondary prediction using a two-dimensional filter can correct the first prediction value using the information of several points forming a preset shape, and finally obtain the corrected second prediction value. The inter-frame prediction method proposed in the present application can be well applied to all scenarios, greatly improving the coding performance, thereby improving the encoding and decoding efficiency.
[0169] It should be understood that the embodiments of the present application provide a video coding system. Figure 15 It is a schematic block diagram of the composition of a video coding system provided by the embodiments of the present application, as Figure 15As shown in the figure, the video encoding system 11 may include: a transformation unit 111, a quantization unit 112, a mode selection and encoding control logic unit 113, an intra prediction unit 114, an inter prediction unit 115 (including motion compensation and motion estimation), an inverse quantization unit 116, an inverse transformation unit 117, a loop filter unit 118, an encoding unit 119, and a decoded image buffer unit 110; for the input original video signal, a video reconstruction block can be obtained through the division of a Coding Tree Unit (CTU). The encoding mode is determined by the mode selection and encoding control logic unit 113. Then, for the residual pixel information obtained after intra or inter prediction, the transformation unit 111 and the quantization unit 112 perform a transformation on the video reconstruction block, including transforming the residual information from the pixel domain to the transform domain and quantizing the obtained transform coefficients to further reduce the bit rate; the intra prediction unit 114 is used to perform intra prediction on the video reconstruction block; among them, the intra prediction unit 114 is used to determine the optimal intra prediction mode (i.e., the target prediction mode) of the video reconstruction block; the inter prediction unit 115 is used to perform inter prediction encoding of the received video reconstruction block relative to one or more blocks in one or more reference frames to provide temporal prediction information; among them, motion estimation is the process of generating a motion vector, and the motion vector can estimate the motion of the video reconstruction block. Then, motion compensation is performed based on the motion vector determined by motion estimation; after determining the inter prediction mode, the inter prediction unit 115 is further used to provide the selected inter prediction data to the encoding unit 119, and also send the calculated and determined motion vector data to the encoding unit 119; in addition, the inverse quantization unit 116 and the inverse transformation unit 117 are used for the reconstruction of the video reconstruction block, reconstructing the residual block in the pixel domain. The reconstructed residual block removes block effect artifacts through the loop filter unit 118. Then, the reconstructed residual block is added to a predictive block in the frame of the decoded image buffer unit 110 to generate a reconstructed video reconstruction block; the encoding unit 119 is used to encode various encoding parameters and the quantized transform coefficients. The decoded image buffer unit 110 is used to store the reconstructed video reconstruction blocks for prediction reference. As the video image encoding progresses, new reconstructed video reconstruction blocks will be continuously generated, and these reconstructed video reconstruction blocks will be stored in the decoded image buffer unit 110.
[0170] An embodiment of this application also provides a video decoding system. Figure 16 It is a schematic block diagram of the composition of a video decoding system provided by an embodiment of this application, as Figure 16As shown, the video decoding system 12 may include: a decoding unit 121, an inverse transformation unit 127, an inverse quantization unit 122, an intra prediction unit 123, a motion compensation unit 124, a loop filter unit 125, and a decoded image buffer unit 126; after the input video signal is encoded by the video encoding system 11, the bitstream of the video signal is output; the bitstream is input into the video decoding system 12, and first passes through the decoding unit 121 to obtain the decoded transform coefficients; the transform coefficients are processed by the inverse transformation unit 127 and the inverse quantization unit 122 to generate a residual block in the pixel domain; the intra prediction unit 123 can be used to generate the prediction data of the current video decoding block based on the determined intra prediction direction and the data of the previously decoded block from the current frame or picture; the motion compensation unit 124 determines the prediction information for the video decoding block by analyzing the motion vector and other associated syntax elements, and uses the prediction information to generate a predictive block of the video decoding block being decoded; by summing the residual block from the inverse transformation unit 127 and the inverse quantization unit 122 and the corresponding predictive block generated by the intra prediction unit 123 or the motion compensation unit 124, a decoded video block is formed; the decoded video signal passes through the loop filter unit 125 to remove block effect artifacts, which can improve the video quality; then the decoded video block is stored in the decoded image buffer unit 126, and the decoded image buffer unit 126 stores the reference image for subsequent intra prediction or motion compensation, and is also used for the output of the video signal to obtain the restored original video signal.
[0171] The inter prediction method provided by the embodiment of the present application mainly acts on the inter prediction unit 215 of the video encoding system 11 and the inter prediction unit of the video decoding system 12, that is, the motion compensation unit 124; that is to say, if a better prediction effect can be obtained by the inter prediction method provided by the embodiment of the present application in the video encoding system 11, then correspondingly, in the video decoding system 12, the video decoding recovery quality can also be improved.
[0172] Based on this, the technical solution of the present application will be further elaborated in detail below in conjunction with the drawings and embodiments. Before the detailed elaboration, it should be noted that the "first", "second", "third", etc. mentioned throughout the specification are only used to distinguish different features and do not have functions such as limiting priority, sequence, or size relationship.
[0173] It should be noted that this embodiment is described by taking the AVS3 standard as an example, and the inter prediction method proposed by the present application can also be applied to other coding standard technologies such as VVC, and the application does not make specific limitations on this.
[0174] An embodiment of the present application provides an inter-frame prediction method, which is applied to a video decoding device, i.e., a decoder. The functions implemented by this method can be achieved by a first processor in the decoder calling a computer program. Of course, the computer program can be stored in the first memory. It can be seen that the decoder at least includes a first processor and a first memory.
[0175] Further, in the embodiment of the present application, Figure 17 Schematic diagram of the implementation process of the inter-frame prediction method Figure 1 , as Figure 17 shown, the method for the decoder to perform inter-frame prediction may include the following steps:
[0176] Step 201, parse the bitstream to obtain the prediction mode parameter of the current block.
[0177] In the embodiment of the present application, the decoder may first parse the binary bitstream to obtain the prediction mode parameter of the current block. Among them, the prediction mode parameter can be used to determine the prediction mode used for the current block.
[0178] It should be noted that the image to be decoded can be divided into multiple image blocks, and the currently to-be-decoded image block can be referred to as the current block (which can be represented by CU), and the image blocks adjacent to the current block can be referred to as adjacent blocks; that is, in the image to be decoded, there is an adjacent relationship between the current block and the adjacent blocks. Here, each current block may include a first image component, a second image component, and a third image component, that is, the current block represents the image block in the image to be decoded for which the first image component, the second image component, or the third image component prediction is to be performed currently.
[0179] Among them, assuming that the current block performs the first image component prediction, and the first image component is a luminance component, that is, the image component to be predicted is a luminance component, then the current block can also be referred to as a luminance block; or, assuming that the current block performs the second image component prediction, and the second image component is a chrominance component, that is, the image component to be predicted is a chrominance component, then the current block can also be referred to as a chrominance block.
[0180] Further, in the embodiment of the present application, the prediction mode parameter can not only indicate the prediction mode adopted by the current block, but also indicate the parameters related to this prediction mode.
[0181] It can be understood that, in the embodiment of the present application, the prediction mode may include an inter-frame prediction mode, a traditional intra-frame prediction mode, a non-traditional intra-frame prediction mode, etc.
[0182] That is to say, on the encoding side, the encoder can select the optimal prediction mode to pre-encode the current block. During this process, the prediction mode of the current block can be determined, and then the prediction mode parameters used to indicate the prediction mode can be determined, so as to write the corresponding prediction mode parameters into the code stream and transmit them from the encoder to the decoder.
[0183] Correspondingly, on the decoder side, the decoder can directly obtain the prediction mode parameters of the current block by parsing the code stream, and determine the prediction mode used by the current block and the relevant parameters corresponding to the prediction mode according to the parsed prediction mode parameters.
[0184] Further, in the embodiments of the present application, after parsing and obtaining the prediction mode parameters, the decoder can determine whether the current block uses the inter-frame prediction mode based on the prediction mode parameters.
[0185] Step 202, when the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, determine the first motion vector of the sub-blocks of the current block; wherein, the current block includes multiple sub-blocks.
[0186] In the embodiments of the present application, after parsing and obtaining the prediction mode parameters, if the parsed prediction mode parameters indicate that the current block uses the inter-frame prediction mode to determine the inter-frame prediction value of the current block, then the decoder can first determine the first motion vector of each sub-block of the current block. Wherein, one sub-block corresponds to one first motion vector.
[0187] It should be noted that, in the embodiments of the present application, the current block is the image block to be decoded in the current frame, and the current frame is decoded sequentially in the form of image blocks. The current block is the image block to be decoded at the next moment in the current frame in this order. The current block can have various specifications and sizes, such as specifications of 16×16, 32×32 or 32×16, etc., where the numbers represent the number of rows and columns of pixel points on the current block.
[0188] Furthermore, in the embodiments of the present application, the current block can be divided into multiple sub-blocks, wherein the size of each sub-block is the same, and the sub-block is a set of pixel points with a smaller specification. The size of the sub-block can be 8×8 or 4×4.
[0189] Exemplarily, in the present application, the size of the current block is 16×16 and can be divided into 4 sub-blocks with a size of 8×8 each.
[0190] It can be understood that, in the embodiments of the present application, when the decoder parses the code stream and obtains that the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, the inter-frame prediction method provided by the embodiments of the present application can be continued to be used.
[0191] In an embodiment of the present application, further, when the prediction mode parameter indicates that the inter - prediction mode is used to determine the inter - prediction value of the current block, the method for the decoder to determine the first motion vector of the sub - blocks of the current block may include the following steps:
[0192] Step 202a, parse the bitstream to obtain the affine mode parameter and the prediction reference mode of the current block.
[0193] Step 202b, when the affine mode parameter indicates the use of the affine mode, determine the control point mode and the sub - block size parameter.
[0194] Step 202c, determine the first motion vector according to the prediction reference mode, the control point mode, and the sub - block size parameter.
[0195] In an embodiment of the present application, after the decoder parses and obtains the prediction mode parameter, if the obtained prediction mode parameter indicates that the current block uses the inter - prediction mode to determine the inter - prediction value of the current block, then the decoder can obtain the affine mode parameter and the prediction reference mode by parsing the bitstream.
[0196] It should be noted that, in an embodiment of the present application, the affine mode parameter is used to indicate whether to use the affine mode. Specifically, the affine mode parameter may be the affine motion compensation enable flag affine_enable_flag. The decoder can further determine whether to use the affine mode by determining the value of the affine mode parameter.
[0197] That is to say, in the present application, the affine mode parameter can be a binary variable. If the value of the affine mode parameter is 1, it indicates the use of the affine mode; if the value of the affine mode parameter is 0, it indicates the non - use of the affine mode.
[0198] It can be understood that, in the present application, if the decoder parses the bitstream and does not obtain the affine mode parameter, it can also be understood as indicating the non - use of the affine mode.
[0199] Exemplarily, in the present application, the value of the affine mode parameter can be equal to the value of the affine motion compensation enable flag affine_enable_flag. If the value of affine_enable_flag is '1', it means that affine motion compensation can be used; if the value of affine_enable_flag is '0', it means that affine motion compensation should not be used.
[0200] Further, in an embodiment of the present application, if the affine mode parameter obtained by the decoder parsing the bitstream indicates the use of the affine mode, then the decoder can obtain the control point mode and the sub - block size parameter.
[0201] It should be noted that in the embodiments of the present application, the control point mode is used to determine the number of control points. In the affine model, a sub-block can have 2 control points or 3 control points. Correspondingly, the control point mode can be the control point mode corresponding to 2 control points, or the control point mode corresponding to 3 control points. That is, the control point mode can include a 4-parameter mode and a 6-parameter mode.
[0202] It can be understood that in the embodiments of the present application, for the AVS3 standard, if the current block uses the affine mode, the decoder also needs to determine the number of control points of the current block in the affine mode, so as to determine whether to use the 4-parameter (2 control points) mode or the 6-parameter (3 control points) mode.
[0203] Furthermore, in the embodiments of the present application, if the affine mode parameter obtained by the decoder parsing the bitstream indicates the use of the affine mode, the decoder can further obtain the sub-block size parameter by parsing the bitstream.
[0204] Specifically, the sub-block size parameter can be determined by the affine prediction sub-block size flag affine_subblock_size_flag. The decoder parses the bitstream to obtain the sub-block size flag, and determines the size of the sub-block of the current block according to the value of the sub-block flag. Among them, the size of the sub-block can be 8×8 or 4×4. Specifically, in the present application, the sub-block size flag can be a binary variable. If the value of the sub-block size flag is 1, it indicates that the sub-block size parameter is 8×8; if the value of the sub-block size flag is 0, it indicates that the sub-block size parameter is 4×4.
[0205] Exemplarily, in the present application, the value of the sub-block size flag can be equal to the value of the affine prediction sub-block size flag affine_subblock_size_flag. If the value of affine_subblock_size_flag is '1', the current block is divided into sub-blocks with a size of 8×8; if the value of affine_subblock_size_flag is '0', the current block is divided into sub-blocks with a size of 4×4.
[0206] It can be understood that in the present application, when the decoder parses the bitstream and does not obtain the sub-block size flag, it can also be understood that the current block is divided into 4×4 sub-blocks. That is to say, if the affine_subblock_size_flag does not exist in the bitstream, the value of the sub-block size flag can be directly set to 0.
[0207] Further, in the embodiments of the present application, after determining the control point mode and the sub-block size parameter, the decoder can further determine the first motion vector of the sub-block in the current block according to the prediction reference mode, the control point mode, and the sub-block size parameter.
[0208] Specifically, in the embodiments of the present application, the decoder can first determine the control point motion vector group according to the prediction reference mode; then, based on the control point motion vector group, the control point mode, and the sub-block size parameter, determine the first motion vector of the sub-block.
[0209] It can be understood that, in the embodiments of the present application, the control point motion vector group can be used to determine the motion vectors of the control points.
[0210] It should be noted that, in the embodiments of the present application, the decoder can traverse each sub-block in the current block according to the above method, and use the control point motion vector group, the control point mode, and the sub-block size parameter of each sub-block to determine the first motion vector of each sub-block, so as to construct a motion vector set according to the first motion vector of each sub-block.
[0211] It can be understood that, in the embodiments of the present application, the motion vector set of the current block can include the first motion vectors of each sub-block of the current block.
[0212] Further, in the embodiments of the present application, when the decoder determines the first motion vector according to the control point motion vector group, the control point mode, and the sub-block size parameter, it can first determine the difference variable according to the control point motion vector group, the control point mode, and the size parameter of the current block; then, based on the prediction mode parameter and the sub-block size parameter, determine the sub-block position; finally, use the difference variable and the sub-block position to determine the first motion vector of the sub-block, and further obtain the motion vector set of multiple sub-blocks of the current block.
[0213] Exemplarily, in the present application, the difference variable can include 4 variables, specifically dHorX, dVerX, dHorY, and dVerY. When calculating the difference variable, the decoder needs to first determine the control point motion vector group, where the control point motion vector group can represent the motion vectors of the control points.
[0214] Specifically, if the control point mode is the 6-parameter mode, that is, there are 3 control points, then the control point motion vector group can be a motion vector group including 3 motion vectors, denoted as mvsAffine(mv0, mv1, mv2); if the control point mode is the 4-parameter mode, that is, there are 2 control points, then the control point motion vector group can be a motion vector group including 2 motion vectors, denoted as mvsAffine(mv0, mv1).
[0215] Next, the decoder can use the control point motion vector group to calculate the difference variables:
[0216] dHorX = (mv1_x - mv0_x) << (7 - Log(width));
[0217] dHorY = (mv1_y - mv0_y) << (7 - Log(width));
[0218] If the motion vector group is mvsAffine(mv0, mv1, mv2), then:
[0219] dVerX = (mv2_x - mv0_x) << (7 - Log(height));
[0220] dVerY = (mv2_y - mv0_y) << (7 - Log(height));
[0221] If the motion vector group is mvsAffine(mv0, mv1), then:
[0222] dVerX = -dHorY;
[0223] dVerY = dHorX.
[0224] Where width and height are the width and height of the current block respectively, that is, the size parameters of the current block. Specifically, the size parameters of the current block can be obtained by the decoder through parsing the bitstream.
[0225] Furthermore, in the embodiments of the present application, after determining the difference variables, the decoder can then determine the sub-block position based on the prediction mode parameter and the sub-block size parameter. Specifically, the decoder can determine the size of the sub-block through the sub-block size flag, and at the same time can determine which prediction mode to specifically use through the prediction mode parameter, and then can determine the sub-block position according to the size of the sub-block and the prediction mode used.
[0226] Exemplarily, in the present application, if the value of the prediction reference mode of the current block is 2, i.e., the third reference mode 'Pred_List01', or if the value of the sub-block size flag is 1, i.e., both the width subwidth and the height subheight of the sub-block are equal to 8, and (x, y) are the coordinates of the upper left corner position of the 8×8 sub-block, then the coordinates xPos and yPos of the sub-block position can be determined in the following manner:
[0227] If the sub-block is the control point at the upper left corner of the current block, then both xPos and yPos are equal to 0;
[0228] If the sub-block is the control point at the upper right corner of the current block, then xPos is equal to width and yPos is equal to 0;
[0229] If the sub-block is the control point at the lower left corner of the current block, and the control point motion vector group can be a motion vector group including 3 motion vectors, then xPos is equal to 0 and yPos is equal to height;
[0230] Otherwise, xPos is equal to (x - xE) + 4 and yPos is equal to (y - yE) + 4.
[0231] Exemplarily, in the present application, if the value of the prediction reference mode of the current block is 0 or 1, i.e., the first reference mode 'Pred_List0' or the second reference mode 'Pred_List1', and the value of the sub-block size flag is 0, i.e., both the width subwidth and the height subheight of the sub-block are equal to 4, and (x, y) are the coordinates of the upper left corner position of the 4×4 sub-block, then the coordinates xPos and yPos of the sub-block position can be determined in the following manner:
[0232] If the sub-block is the control point at the upper left corner of the current block, then both xPos and yPos are equal to 0;
[0233] If the sub-block is the control point at the upper right corner of the current block, then xPos is equal to width and yPos is equal to 0;
[0234] If the sub-block is the control point at the lower left corner of the current block, and the control point motion vector group can be a motion vector group including 3 motion vectors, then xPos is equal to 0 and yPos is equal to height;
[0235] Otherwise, xPos is equal to (x - xE) + 2 and yPos is equal to (y - yE) + 2.
[0236] Further, in the embodiments of the present application, after the decoder calculates and obtains the sub-block position, it can then determine the first motion vector of the sub-block based on the sub-block position and the difference variable. Finally, by traversing each sub-block of the current block and obtaining the first motion vector of each sub-block, a motion vector set of multiple sub-blocks of the current block can be constructed and obtained.
[0237] Exemplarily, in the present application, after the decoder determines the sub-block positions xPos and yPos, the first motion vector mvE (mvE_x, mvE_y) of the sub-block can be determined by the following method
[0238] mvE_x = Clip3(-131072, 131071, Rounding((mv0_x << 7) + dHorX × xPos + dVerX × yPos, 7));
[0239] mvE_y = Clip3(-131072, 131071, Rounding((mv0_y << 7) + dHorY × xPos + dVerY × yPos, 7)).
[0240] It should be noted that, in the present application, when determining the deviation between each position within the sub-block and the motion vector of the sub-block, if the affine prediction model is used for the current block, the motion vector of each position within the sub-block can be calculated according to the formula of the affine prediction model, and the deviation between them can be obtained by subtracting the motion vector of the sub-block. If the motion vectors of the sub-blocks all select the motion vectors of the same position within the sub-block, for example, a 4x4 block uses the position (2, 2) from the upper left corner, and an 8x8 block uses the position (4, 4) from the upper left corner. According to the affine models used in current standards including VVC and AVS3, the motion vector deviations of the same position in each sub-block are the same. However, in the case of the upper left corner, upper right corner, and the lower left corner in the case of 3 control points (the positions A, B, C shown in the above AVS3 text), the positions used are different from those used in other blocks. Correspondingly, when calculating the motion vector deviations of the sub-blocks in the upper left corner, upper right corner, and the lower left corner in the case of 3 control points, they are also different from those of other blocks. Specifically, as in the embodiments. Figure 7 shown, the motion vector deviation of the sub-block at the lower left corner in the case of the upper left corner, upper right corner, and 3 control points is different from that of other blocks. Specifically, as in the embodiments.
[0241] Step 203: Determine the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block based on the first motion vector; wherein, the pixel position is the position of the pixel point within the sub-block.
[0242] In the embodiments of the present application, after the decoder determines the first motion vector of each sub-block of the current block, it can determine the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block respectively based on the first motion vector of the sub-block.
[0243] It can be understood that, in the embodiments of the present application, step 203 may specifically include:
[0244] Step 203a: Determine the first prediction value of the sub-block based on the first motion vector.
[0245] Step 203b: Determine the motion vector deviation between the pixel position and the sub-block based on the first motion vector.
[0246] Among them, the inter-frame prediction method proposed in the embodiments of the present application does not limit the order in which the decoder executes step 203a and step 203b. That is to say, in the present application, after determining the first motion vector of each sub-block of the current block, the decoder may first execute step 203a, then execute step 203b, or may first execute step 203b, and then execute step 203a, or may also execute step 203a and step 203b simultaneously.
[0247] Furthermore, in the embodiments of the present application, when the decoder determines the first prediction value of the sub-block based on the first motion vector, it may first determine a sample matrix; wherein, the sample matrix includes a luminance sample matrix and a chrominance sample matrix; and then may determine the first prediction value according to the prediction reference mode, the sub-block size parameter, the sample matrix, and the motion vector set.
[0248] It should be noted that, in the embodiments of the present application, when the decoder determines the first prediction value according to the prediction reference mode, the sub-block size parameter, the sample matrix, and the motion vector set, it may first determine a target motion vector from the motion vector set according to the prediction reference mode and the sub-block size parameter; and then may use the reference image queue and reference index corresponding to the prediction reference mode, the sample matrix, and the target motion vector to determine a prediction sample matrix; wherein, the prediction sample matrix includes the first prediction values of multiple sub-blocks.
[0249] Specifically, in the embodiments of the present application, the sample matrix may include a luminance sample matrix and a chrominance sample matrix. Correspondingly, the prediction sample matrix determined by the decoder may include a luminance prediction sample matrix and a chrominance prediction sample matrix. Among them, the luminance prediction sample matrix includes the first luminance prediction values of multiple sub-blocks, and the chrominance prediction sample matrix includes the first chrominance prediction values of multiple sub-blocks. The first luminance prediction value and the first chrominance prediction value constitute the first prediction value of the sub-block.
[0250] Exemplarily, in the present application, it is assumed that the position of the top-left sample of the current block in the luminance sample matrix of the current image is (xE, yE). If the value of the prediction reference mode of the current block is 0, that is, the first reference mode 'PRED_List0' is used, and the value of the sub-block size flag is 0, that is, the sub-block size parameter is 4×4, then the target motion vector mv0E0 is the first motion vector of the 4×4 sub-block at the position (xE + x, yE + y) in the motion vector set of the current block. The value of the element predMatrixL0[x][y] in the luminance prediction sample matrix predMatrixL0 is the sample value at the position (((xE + x) << 4) + mv0E0_x, ((yE + y) << 4) + mv0E0_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL0 in the reference image queue 0. The value of the element predMatrixL0[x][y] in the chrominance prediction sample matrix predMatrixL0 is the sample value at the position (((xE + 2×x) << 4) + MvC_x, ((yE + 2×y) << 4) + MvC_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL0 in the reference image queue 0. Wherein, x1 = ((xE + 2×x) >> 3) << 3, y1 = ((yE + 2×y) >> 3) << 3, mv1E0 is the first motion vector of the 4×4 unit at the position (x1, y1) in the motion vector set of the current block, mv2E0 is the first motion vector of the 4×4 unit at the position (x1 + 4, y1) in the motion vector set of the current block, mv3E0 is the first motion vector of the 4×4 unit at the position (x1, y1 + 4) in the motion vector set of the current block, and mv4E0 is the first motion vector of the 4×4 unit at the position (x1 + 4, y1 + 4) in the motion vector set of the current block.
[0251] Specifically, MvC_x and MvC_y can be determined in the following manner:
[0252] MvC_x = (mv1E0_x + mv2E0_x + mv3E0_x + mv4E0_x + 2) >> 2
[0253] MvC_y = (mv1E0_y + mv2E0_y + mv3E0_y + mv4E0_y + 2) >> 2
[0254] Exemplarily, in the present application, it is assumed that the position of the sample in the upper left corner of the current block in the luminance sample matrix of the current image is (xE, yE). If the predicted reference mode value of the current block is 0, that is, the first reference mode 'PRED_List0' is used, and the value of the sub-block size flag is 1, that is, the sub-block size parameter is 8×8, then the target motion vector mv0E0 is the first motion vector of the 8×8 unit at the position (xE + x, yE + y) in the motion vector set of the current block. The value of the element predMatrixL0[x][y] in the luminance prediction sample matrix predMatrixL0 is the sample value at the position (((xE + x) << 4) + mv0E0_x, ((yE + y) << 4) + mv0E0_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL0 in the reference image queue 0. The value of the element predMatrixL0[x][y] in the chrominance prediction sample matrix predMatrixL0 is the sample value at the position (((xE + 2×x) << 4) + MvC_x, ((yE + 2×y) << 4) + MvC_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL0 in the reference image queue 0. Wherein, MvC_x is equal to mv0E0_x, and MvC_y is equal to mv0E0.
[0255] Exemplarily, in the present application, it is assumed that the position of the top-left sample of the current block in the luminance sample matrix of the current image is (xE, yE). If the predicted reference mode value of the current block is 1, that is, the second reference mode 'PRED_List1' is used, and the value of the sub-block size flag is 0, that is, the sub-block size parameter is 4×4, then the target motion vector mv0E1 is the first motion vector of the 4×4 unit at the position (xE + x, yE + y) in the motion vector set of the current block. The value of the element predMatrixL1[x][y] in the luminance prediction sample matrix predMatrixL1 is the sample value at the position (((xE + x) << 4) + mv0E1_x, ((yE + y) << 4) + mv0E1_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL1 in the reference image queue 1. The value of the element predMatrixL1[x][y] in the chrominance prediction sample matrix predMatrixL1 is the sample value at the position (((xE + 2×x) << 4) + MvC_x, ((yE + 2×y) << 4) + MvC_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL1 in the reference image queue 1. Wherein, x1 = ((xE + 2×x) >> 3) << 3, y1 = ((yE + 2×y) >> 3) << 3, mv1E1 is the first motion vector of the 4×4 unit at the position (x1, y1) in the first motion vector set of MvArray, mv2E1 is the first motion vector of the 4×4 unit at the position (x1 + 4, y1) in the first motion vector set of MvArray, mv3E1 is the first motion vector of the 4×4 unit at the position (x1, y1 + 4) in the first motion vector set of MvArray, and mv4E1 is the first motion vector of the 4×4 unit at the position (x1 + 4, y1 + 4) in the first motion vector set of MvArray.
[0256] Specifically, MvC_x and MvC_y can be determined in the following manner:
[0257] MvC_x = (mv1E1_x + mv2E1_x + mv3E1_x + mv4E1_x + 2) >> 2
[0258] MvC_y = (mv1E1_y + mv2E1_y + mv3E1_y + mv4E1_y + 2) >> 2
[0259] Exemplarily, in the present application, it is assumed that the position of the top - left sample of the current block in the luminance sample matrix of the current image is (xE, yE). If the prediction reference mode value of the current block is 1, that is, the second reference mode 'PRED_List1' is used, and the value of the sub - block size flag is 1, that is, the sub - block size parameter is 8×8, then the target motion vector mv0E1 is the first motion vector of the 8×8 unit at the position (xE + x, yE + y) in the motion vector set of the current block. The value of the element predMatrixL1[x][y] in the luminance prediction sample matrix predMatrixL1 is the sample value at the position (((xE + x) << 4)+mv0E1_x, ((yE + y) << 4)+mv0E1_y) in the 1 / 16 - precision luminance sample matrix with the reference index RefIdxL1 in the reference image queue 1. The value of the element predMatrixL1[x][y] in the chrominance prediction sample matrix predMatrixL1 is the sample value at the position (((xE + 2×x) << 4)+MvC_x, ((yE + 2×y) << 4)+MvC_y) in the 1 / 32 - precision chrominance sample matrix with the reference index RefIdxL1 in the reference image queue 1. Where MvC_x is equal to mv0E1_x and MvC_y is equal to mv0E1.
[0260] Exemplarily, in the present application, it is assumed that the position of the top-left sample of the current block in the luminance sample matrix of the current image is (xE, yE). If the predicted reference mode value of the current block is 2, that is, the third reference mode 'PRED_List01' is used, then the target motion vector mv0E0 is the first motion vector of the 8×8 unit at the position (xE + x, yE + y) in the motion vector set of the current block, and the target motion vector mv0E1 is the first motion vector of the 8×8 unit at the position (x, y) in the motion vector set of the current block. The value of the element predMatrixL0[x][y] in the luminance prediction sample matrix predMatrixL0 is the sample value at the position ((((xE + x) << 4) + mv0E0_x, ((yE + y) << 4) + mv0E0_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL0 in the reference image queue 0. The value of the element predMatrixL0[x][y] in the chrominance prediction sample matrix predMatrixL0 is the sample value at the position ((((xE + 2×x) << 4) + MvC0_x, ((yE + 2×y) << 4) + MvC0_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL0 in the reference image queue 0. The value of the element predMatrixL1[x][y] in the luminance prediction sample matrix predMatrixL1 is the sample value at the position ((((xE + x) << 4) + mv0E1_x, ((yE + y) << 4) + mv0E1_y) in the 1 / 16 precision luminance sample matrix with the reference index RefIdxL1 in the reference image queue 1. The value of the element predMatrixL1[x][y] in the chrominance prediction sample matrix predMatrixL1 is the sample value at the position ((((xE + 2×x) << 4) + MvC1_x, ((yE + 2×y) << 4) + MvC1_y) in the 1 / 32 precision chrominance sample matrix with the reference index RefIdxL1 in the reference image queue 1. Where MvC0_x is equal to mv0E0_x, MvC0_y is equal to mv0E0_y, MvC1_x is equal to mv0E1_x, and MvC1_y is equal to mv0E1_y.
[0261] It should be noted that in the embodiments of the application, the luminance sample matrix in the sample matrix may be a 1 / 16 precision luminance sample matrix, and the chrominance sample matrix in the sample matrix may be a 1 / 32 precision chrominance sample matrix.
[0262] It can be understood that in the embodiments of the present application, for different predicted reference modes, the reference image queue and reference index obtained by the decoder by parsing the code stream are different.
[0263] Further, in the embodiments of the present application, when the decoder determines the sample matrix, it may first obtain the luminance interpolation filter coefficients and the chrominance interpolation filter coefficients; then, it may determine the luminance sample matrix based on the luminance interpolation filter coefficients, and at the same time, determine the chrominance sample matrix based on the chrominance interpolation filter coefficients.
[0264] Exemplarily, in the present application, when the decoder determines the luminance sample matrix, the obtained luminance interpolation filter coefficients are as shown in Table 1 above, and then, according to the Figure 8 and Figure 9 shown pixel positions and sample positions, the luminance sample matrix is calculated and obtained.
[0265] Specifically, the sample position a x,0 (x = 1 to 15) is filtered by 8 integer values closest to the interpolation point in the horizontal direction, and the prediction value is obtained as follows:
[0266] a x,0 = Clip1((fL[x][0] × A -3,0 + fL[x][1] × A -2,0 + fL[x][2] × A -1,0 + fL[x][3] × A 0,0 + fL[x][4] × A 1,0 + fL[x][5] × A 2,0 + fL[x][6]
[0267] × A 3,0 + fL[x][7] × A 4,0 + 32) >> 6).
[0268] Specifically, the sample position a 0,y (y = 1 to 15) is filtered by 8 integer values closest to the interpolation point in the vertical direction, and the prediction value is obtained as follows:
[0269] a 0,y = Clip1((fL[y][0] × A 0,-3 + fL[y][1] × A -2,0 + fL[y][2] × A -1,0 + fL[y][3] × A 0,0 + fL[y][4] × A 1,0 + fL[y][5] × A 2,0 + fL[y][6] × A 3,
[0270] 0 + fL[y][7] × A -4,0 + 32) >> 6).
[0271] Specifically, the sample position a x,y (where x = 1 to 15, y = 1 to 15) is obtained as follows:
[0272] a x,y= Clip1((fL[y][0] × a' x,y-3 + fL[y][1] × a' x,y-2 + fL[y][2] × a'x,y-1 + fL[y][3] × a'x,y + fL[y][4] × a'x,y+1 + fL[y][5] × a'x,y + 2 + fL[y][6] × a'x,
[0273] y+3 + fL[y][7] × a'x,y+4 +(1 << (19 - BitDepth))) >> (20 - BitDepth))。
[0274] Where:
[0275] a' x,y = (fL[x][0] × A -3,y + fL[x][1] × A -2,y + fL[x][2] × A -1,y+ fL[x][3] × A 0,y + fL[x][4] × A 1,y + fL[x][5] × A 2,y + fL[x][6] × A 3,y + fL[x][7] × A 4,
[0276] y+ ((1 << (BitDepth - 8)) >> 1)) >> (BitDepth - 8)。
[0277] Exemplarily, in the present application, when the decoder determines the chrominance sample matrix, it can first parse the bitstream to obtain the chrominance interpolation filter coefficients as shown in Table 2 above, and then calculate the chrominance sample matrix according to the pixel positions and sample positions as Figure 10 and Figure 11 shown.
[0278] Specifically, for sub-pixel points where dx is equal to 0 or dy is equal to 0, chrominance integer-pixel interpolation can be directly used. For points where dx is not equal to 0 and dy is not equal to 0, sub-pixels on the integer-pixel row (dy equal to 0) are used for calculation:
[0279] if (dx == 0) {
[0280] a x,y (0, dy) = Clip3(0, (1 << BitDepth) - 1, (fC[dy][0] × A x,y -1 + fC[dy][1] × A x,y + fC[dy][2] × A x,y + 1 + fC[dy][3] × A x,y + 2 + 32) >> 6)
[0281] }
[0282] else if (dy == 0) {
[0283] a x,y (dx, 0) = Clip3(0, (1 << BitDepth) - 1, (fC[dx][0] × A x-1,y + fC[dx][1] × A x,y + fC[dx][2] × A x+1,y + fC[dx][3] × A x+2,y+32 ) >> 6)
[0284] }
[0285] else {
[0286] a x,y (dx, dy) = Clip3(0, (1 << BitDepth) - 1, (C[dy][0] × a' x,y-1 (dx, 0) + C[dy][1] × a' x,y (dx, 0) + C[dy][2] × a' x,y + 1
[0287] (dx, 0) + C[dy][3] × a' x,y + 2(dx, 0) + (1 << (19 - BitDepth))) >> (20 - BitDepth))
[0288] }
[0289] Wherein, a' x,y (dx, 0) is the temporary value of the sub - pixel on the whole - pixel row, defined as: a' x,y (dx, 0) = (fC[dx][0] × A x-1,y + fC[dx][1] × A x,y + fC[dx][2] × A x+1,y+ fC[dx][3] × A x+2,y + ((1 << (BitDepth - 8)) >> 1)) >> (BitDepth - 8).
[0290] Further, in an embodiment of the present application, when the decoder is based on the motion vector deviation between the pixel position and the sub-block, it may first parse the bitstream to obtain the secondary prediction parameters; if the secondary prediction parameters indicate the use of secondary prediction, then the decoder may determine the motion vector deviation between the sub-block and each pixel position based on the difference variable.
[0291] Specifically, in an embodiment of the present application, when the decoder determines the motion vector deviation between the sub-block and each pixel position based on the difference variable, it may, according to the method proposed in step 202 above, determine the four difference variables dHorX, dVerX, dHorY, and dVerY based on the control point motion vector group, the control point mode, and the size parameters of the current block, and then further determine the motion vector deviation corresponding to each pixel position in the sub-block by using the difference variable.
[0292] Exemplarily, in the present application, width and height are respectively the width and height of the current block obtained by the decoder, and the width subwidth and height subheight of the sub-block are determined by using the sub-block size parameters. Assuming that (i, j) is the coordinate of any pixel point inside the sub-block, where the value range of i is 0 to (subwidth - 1), and the value range of j is 0 to (subheight - 1), then the motion vector deviation of each pixel (i, j) position inside the four different types of sub-blocks can be calculated by the following method:
[0293] If the sub-block is the control point A at the upper left corner of the current block, then the motion vector deviation dMvA[i][j] of the (i, j) pixel:
[0294] dMvA[i][j][0] = dHorX × i + dVerX × j
[0295] dMvA[i][j][1] = dHorY × i + dVerY × j;
[0296] If the sub-block is the control point B at the upper right corner of the current block, then the motion vector deviation dMvB[i][j] of the (i, j) pixel:
[0297] dMvB[i][j][0] = dHorX × (i - subwidth) + dVerX × j
[0298] dMvB[i][j][1] = dHorY × (i - subwidth) + dVerY × j;
[0299] If the sub-block is the control point C at the lower left corner of the current block, and the motion vector group of the control point can be a motion vector group including 3 motion vectors, then the motion vector deviation dMvC[i][j] of the pixel at (i, j):
[0300] dMvC[i][j][0] = dHorX × i + dVerX × (j - subheight)
[0301] dMvC[i][j][1] = dHorY × i + dVerY × (j - subheight);
[0302] Otherwise, the motion vector deviation dMvN[i][j] of the pixel at (i, j):
[0303] dMvN[i][j][0] = dHorX × (i - (subwidth >> 1)) + dVerX × (j - (subheight >> 1))
[0304] dMvN[i][j][1] = dHorY × (i - (subwidth >> 1)) + dVerY × (j - (subheight >> 1)).
[0305] Wherein, dMvX[i][j][0] represents the deviation value of the motion vector deviation in the horizontal component, and dMvX[i][j][1] represents the deviation value of the motion vector deviation in the vertical component. X is A, B, C or N.
[0306] It can be understood that in the embodiments of the present application, after the decoder determines the motion vector deviation between the sub-block and each pixel position based on the difference variable, it can use all the motion vector deviations corresponding to all the pixel positions within the sub-block to construct a motion vector deviation matrix corresponding to the sub-block. It can be seen that the motion vector deviation matrix includes the motion vector deviation between the sub-block and any internal pixel point, that is, the motion vector deviation.
[0307] Further, in the embodiments of the present application, if the secondary prediction parameter obtained by the decoder parsing the code stream indicates not to use secondary prediction, then the decoder can directly select the first prediction value of the sub-block of the current block obtained in step 203a above as the second prediction value of the sub-block, without performing the following steps 204 and 205.
[0308] Specifically, in the embodiments of the present application, if the secondary prediction parameter indicates not to use secondary prediction, then the decoder can use the prediction sample matrix to determine the second prediction value. The prediction sample matrix includes the first prediction values of multiple sub-blocks, and the decoder can determine the first prediction value of the sub-block where the pixel position is located as its own second prediction value.
[0309] Exemplarily, in the present application, if the prediction reference mode of the current block takes a value of 0 or 1, that is, the first reference mode 'PRED_List0' is used, or the second reference mode 'PRED_List1' is used, then the first prediction value of the sub-block where the pixel position is located can be directly selected from the prediction sample matrix including one luminance prediction sample matrix and two chrominance prediction sample matrices, and this first prediction value is determined as the inter-frame prediction value of the pixel position, that is, the second prediction value.
[0310] Exemplarily, in the present application, if the prediction reference mode of the current block takes a value of 2, that is, the third reference mode 'PRED_List01' is used, then the mean operation can be first performed on the 2 luminance prediction sample matrices (2 groups of 4 chrominance prediction sample matrices in total) included in the prediction sample matrix to obtain 1 averaged luminance prediction sample (2 averaged chrominance prediction samples), and finally the first prediction value of the sub-block where the pixel position is located is selected from this averaged luminance prediction sample (2 averaged chrominance prediction samples), and this first prediction value is determined as the inter-frame prediction value of the pixel position, that is, the second prediction value.
[0311] Step 204: Determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for performing secondary prediction processing according to a preset shape.
[0312] In the embodiment of the present application, after the decoder respectively determines the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block based on the first motion vector, the filtering coefficients of the two-dimensional filter can be further determined according to the motion vector deviation.
[0313] It should be noted that, in the embodiment of the present application, the filter coefficients of the two-dimensional filter are related to the motion vector deviation corresponding to the pixel position. That is to say, for different pixel positions, if the corresponding motion vector deviations are different, then the filtering coefficients of the two-dimensional filter used are also different.
[0314] It can be understood that, in the embodiment of the present application, the two-dimensional filter is used for performing secondary prediction by using a plurality of adjacent pixel positions that form the preset shape. Wherein, the preset shape is a rectangle, a rhombus or any symmetric shape.
[0315] That is to say, in the present application, the two-dimensional filter for performing secondary prediction is a filter formed by adjacent points that form the preset shape. The adjacent points that form the preset shape can include multiple points, for example, formed by 9 points. The preset shape can be a symmetric shape, for example, the preset shape can include a rectangle, a rhombus or any other symmetric shape.
[0316] Exemplarily, in the present application, the two-dimensional filter is a rectangular filter. Specifically, the two-dimensional filter is a filter composed of 9 adjacent pixel positions forming a rectangle. Among the 9 pixel positions, the pixel position at the center is the pixel position of the pixel that needs to be secondarily predicted currently.
[0317] Further, in the embodiments of the present application, when the decoder determines the filtering coefficients of the two-dimensional filter according to the motion vector deviation, it may first parse the bitstream to obtain the scaling parameter, and then may determine the filter coefficients corresponding to the pixel positions according to the scaling parameter and the motion vector deviation.
[0318] It should be noted that, in the embodiments of the present application, the scaling parameter may include at least one scaling value, and the motion vector deviation includes a horizontal deviation and a vertical deviation; wherein, the at least one scaling value is a non-zero real number.
[0319] Specifically, in the present application, when the two-dimensional filter uses 9 adjacent pixel positions forming a rectangle for secondary prediction, the pixel position at the center of the rectangle is the position to be predicted, and the other 8 pixel positions are successively located at the adjacent positions of the upper left, upper, upper right, right, lower right, lower, lower left, and left of the position to be predicted.
[0320] Correspondingly, in the present application, the decoder may calculate and obtain 9 filter coefficients corresponding to the 9 adjacent pixel positions based on the at least one scaling value and the motion vector deviation of the position to be predicted according to a pre-designed calculation rule.
[0321] It should be noted that, in the present application, the pre-designed calculation rule may include various different calculation methods, such as addition operation, subtraction operation, multiplication operation, etc. Among them, for different pixel positions, different calculation methods may be used to calculate the filter coefficients.
[0322] It can be understood that, in the present application, among the multiple filter coefficients corresponding to multiple pixel positions calculated by the decoder according to different calculation methods in the pre-designed calculation rule, some filter coefficients may be a first-order function of the motion vector deviation, that is, the two are linearly related, and may also be a second-order function or a higher-order function of the motion vector deviation, that is, the two are non-linearly related.
[0323] That is to say, in the present application, any one of the multiple filter coefficients corresponding to multiple adjacent pixel positions may be a first-order function, a second-order function, or a higher-order function of the motion vector deviation.
[0324] Exemplarily, in the present application, it is assumed that the motion vector deviation of the pixel position is (dmv_x, dmv_y). Among them, if the coordinates of the target pixel position are (i, j), then dmv_x can be expressed as dMvX[i][j][0], which represents the deviation value of the motion vector deviation in the horizontal component, and dmv_y can be expressed as dMvX[i][j][1], which represents the deviation value of the motion vector deviation in the vertical component.
[0325] Correspondingly, Table 3 shows the filter coefficients obtained based on the motion vector deviation (dmv_x, dmv_y). As shown in Table 3, for a two-dimensional filter, according to the motion vector deviation of the pixel position (the horizontal deviation is dmv_x, and the vertical deviation is dmv_y) and different scale parameters, such as m and n, nine filter coefficients corresponding to nine adjacent pixel positions can be obtained. Among them, the decoder can directly set the filter coefficient of the central position to be predicted to 1.
[0326] Table 3
[0327]
[0328]
[0329] Among them, the scale parameters m and n are generally decimals or fractions. A possible situation is that both m and n are powers of 2, such as 1 / 2, 1 / 4, 1 / 8, etc. Here, dmv_x and dmv_y are their actual magnitudes, that is, 1 in dmv_x and dmv_y represents the distance of 1 pixel, and dmv_x and dmv_y are decimals or fractions.
[0330] It should be noted that in the embodiments of the present application, compared with the existing 8-tap filter, the motion vectors of the integer pixel positions and fractional pixel positions corresponding to the currently common 8-tap filter are non-negative in both the horizontal and vertical directions, and their magnitudes are all within 0 pixels to 1 pixel, that is, dmv_x and dmv_y cannot be negative. However, in the present application, the motion vectors of the integer pixel positions and fractional pixel positions corresponding to the filter can be negative in both the horizontal and vertical directions, that is, dmv_x and dmv_y can be negative.
[0331] Exemplarily, in the embodiments of the present application, if the scale parameter m is 1 / 16 and n is 1 / 2, then the above Table 3 can be represented as Table 4 below:
[0332] Table 4
[0333] Pixel position Filter coefficient Upper left (-dmv_x - dmv_y) / 16 Left -dmv_x / 2 Lower left (-dmv_x + dmv_y) / 16 Upper -dmv_y / 2 Center 1 Lower dmv_y / 2 Upper right (dmv_x - dmv_y) / 16 Right dmv_x / 2 Lower right (dmv_x + dmv_y) / 16
[0334] It can be understood that in the embodiments of the present application, in video coding and decoding technologies and standards, magnification factors are usually used to avoid decimal and floating-point operations, and then the calculated results are reduced by an appropriate factor to obtain the correct results. Left shift is usually used for magnification, and right shift is usually used for reduction. Therefore, when performing secondary prediction through a two-dimensional filter, it will be written in the following form in practical applications:
[0335] Assume that the motion vector deviation at the pixel position is (dmv_x, dmv_y), and after left shift by shift1, it becomes (dmv_x’, dmv_y’). Based on Table 4 above, the coefficients of the two-dimensional filter can be expressed as Table 5 below:
[0336] Table 5
[0337] Pixel position Filter coefficient Upper left -dmv_x’ - dmv_y’ Left -dmv_x’ × 8 Lower left -dmv_x’ + dmv_y’ Upper -dmv_y’ × 8 Center 16 << shift1 Lower dmv_y’ × 8 Upper right dmv_x’ - dmv_y’ Right dmv_x × 8 Lower right dmv_x’ + dmv_y’
[0338] Figure 18 Schematic of the two-dimensional filter Figure 1 , as Figure 18 shown, taking the result of sub-block-based prediction as the basis for secondary prediction, the light-colored squares are the integer pixel positions of the filter, that is, the positions obtained from sub-block-based prediction. The circles are the fractional pixel positions where secondary prediction needs to be performed, that is, the positions of the pixel positions. The dark-colored squares are the integer pixel positions corresponding to the fractional pixel positions, and 9 integer pixel positions as shown in the figure are required to interpolate this fractional pixel position.
[0339] Figure 19 Schematic of the two-dimensional filter Figure 2 , as Figure 19 shown, taking the result of sub-block-based prediction as the basis for secondary prediction, the light-colored squares are the integer pixel positions of the filter, that is, the positions obtained from sub-block-based prediction. The circles are the fractional pixel positions where secondary prediction needs to be performed, that is, the positions of the pixel positions. The dark-colored squares are the integer pixel positions corresponding to the fractional pixel positions, and 13 integer pixel positions as shown in the figure are required to interpolate this fractional pixel position.
[0340] Step 205: Determine the second prediction value of the sub-block based on the filter coefficients and the first prediction value, and determine the second prediction value as the inter-frame prediction value of the sub-block.
[0341] In the embodiments of the present application, after the decoder determines the filter coefficients of the two-dimensional filter according to the motion vector deviation, it can determine the second prediction value of the sub-block based on the filter coefficients and the first prediction value, so as to realize secondary prediction of the sub-block in combination with the motion vector deviation at the pixel position and complete the correction of the first prediction value.
[0342] It can be understood that in the embodiments of the present application, the decoder determines the filter coefficients by using the motion vector deviation corresponding to the pixel positions, so that the first prediction value can be corrected by a two-dimensional filter according to the filter coefficients, and the corrected second prediction value of the sub-block can be obtained. It can be seen that the second prediction value is a corrected value based on the first prediction value.
[0343] Further, in the embodiments of the present application, when the decoder determines the second prediction value of the sub-block based on the filter coefficients and the first prediction value, it can first perform a multiplication operation on the filter coefficients and the first prediction value of the sub-block to obtain a product result. After traversing all the pixel positions in the sub-block, an addition operation is performed on the product results of all the pixel positions in the sub-block to obtain a summation result. Finally, the addition result can be normalized, and finally the corrected second prediction value of the sub-block can be obtained.
[0344] It should be noted that in the embodiments of the present application, before performing the secondary prediction, generally, the first prediction value of the sub-block where the pixel position is located is used as the prediction value before correction of the pixel position. Therefore, when performing filtering by a two-dimensional filter, the filter coefficients can be multiplied by the prediction value of the corresponding pixel position, that is, the first prediction value, and the product results corresponding to each pixel position are accumulated and then normalized.
[0345] It can be understood that in the present application, the decoder can perform normalization processing in various ways. For example, the result of accumulating the product of the filter coefficients and the prediction value of the corresponding pixel position can be shifted 4 + shift1 bits to the right. Or, the result of accumulating the product of the filter coefficients and the prediction value of the corresponding pixel position can be added with (1 << (3 + shift1)), and then shifted 4 + shift1 bits to the right.
[0346] It can be seen that in the present application, after obtaining the motion vector deviation corresponding to the pixel positions inside the sub-block, for each sub-block and each pixel position in each sub-block, filtering can be performed using a two-dimensional filter based on the first prediction value of the motion compensation of the sub-block according to the motion vector deviation, and the secondary prediction of the sub-block is completed to obtain a new second prediction value.
[0347] Further, in the embodiments of the present application, the two-dimensional filter can be understood as performing secondary prediction by using multiple adjacent pixel positions that form the preset shape. The preset shape can be a rectangle, a rhombus, or any symmetric shape.
[0348] Specifically, in the embodiments of the present application, when the two-dimensional filter performs secondary prediction using 9 adjacent pixel positions forming a rectangle, the prediction sample matrix of the current block and the set of motion vector deviations of the sub-blocks of the current block can be determined first; wherein, the motion vector deviation matrix includes the motion vector deviations corresponding to all pixel positions; then, based on the 9 adjacent pixel positions forming a rectangle, using the prediction sample matrix and the set of motion vector deviations, the sample matrix after secondary prediction of the current block is determined.
[0349] Exemplarily, in the present application, if the width and height of the current block are width and height respectively, and the width and height of each sub-block are subwidth and subheight respectively. As Figure 7 shown, the sub-block where the top-left sample of the luminance prediction sample matrix of the current block is located is A, the sub-block where the top-right sample is located is B, the sub-block where the bottom-left sample is located is C, and the sub-blocks where other positions are located are other sub-blocks.
[0350] For each sub-block in the current block, if the motion vector deviation matrix of the sub-block is dMv, then:
[0351] 1. If the sub-block is A, dMv is equal to dMvA;
[0352] 2. If the sub-block is B, dMv is equal to dMvB;
[0353] 3. If the sub-block is C and there are 3 motion vectors in the control point motion vector group mvAffine of the sub-block, dMv is equal to dMvC;
[0354] 4. If the sub-block is other sub-blocks other than A, B, and C, dMv is equal to dMvN.
[0355] Further, assuming that (x, y) is the coordinate of the top-left position of the current sub-block, and (i, j) are the coordinates of the pixels inside the luminance sub-block, the value range of i is 0 to (subwidth - 1), and the value range of j is 0 to (subheight - 1), based on the prediction sample matrix of the sub-block being PredMatrixSb and the prediction sample matrix after secondary prediction being PredMatrixS, the prediction sample PredMatrixS[x + i][y + j] after secondary prediction of (x + i, y + j) can be calculated according to the following method:
[0356] PredMatrixS[x + i][y + j] =
[0357] (UPLEFT(x + i, y + j) × (-dMv[i][j][0] - dMv[i][j][1]) +
[0358] UP(x + i, y + j) × ((-dMv[i][j][1]) << 3) +
[0359] UPRIGHT(x + i, y + j) × (dMv[i][j][0] - dMv[i][j][1]) +
[0360] LEFT(x + i, y + j) × ((-dMv[i][j][0]) << 3) +
[0361] CENTER(x + i, y + j) × (1 << 15) +
[0362] RIGHT(x + i, y + j) × (dMv[i][j][0] << 3) +
[0363] DOWNLEFT(x + i, y + j) × (-dMv[i][j][0] + dMv[i][j][1]) +
[0364] DOWN(x + i, y + j) × (dMv[i][j][1] << 3) +
[0365] DOWNRIGHT(x + i, y + j) × (dMv[i][j][0] + dMv[i][j][1]) +
[0366] (1 << 10)) >> 11
[0367] PredMatrixS[x + i][y + j] = Clip3(0, (1 << BitDepth) - 1, PredMatrixS[x + i][y + j]).
[0368] where UPLEFT(x + i, y + j) = predMatrix[x + i - 1 < 0? 0 : x + i - 1][y + j - 1 < 0? 0 : y + j - 1]
[0369] UP(x + i, y + j) = predMatrix[x + i][y + j - 1 < 0? 0 : y + j - 1]
[0370] UPRIGHT(x + i, y + j) = predMatrix[x + i + 1 > width - 1? width - 1 : x + i + 1][y + j - 1 < 0? 0 : y + j - 1]
[0371] LEFT(x + i, y + j) = predMatrix[x + i - 1 < 0? 0 : x + i - 1][y + j]
[0372] CENTER(x+i, y+j) = predMatrix[x+i][y+j]
[0373] RIGHT(x+i, y+j) = predMatrix[x+i+1>width-1? width-1 : x+i+1][y+j]
[0374] DOWNLEFT(x+i, y+j) = predMatrix[x+i-1<0? 0 : x+i-1][y+j+1<0? 0 : y+j-1]
[0375] DOWN(x+i, y+j) = predMatrix[x+i][y+j-1<0? 0 : y+j-1]
[0376] DOWNRIGHT(x+i, y+j) = predMatrix[x+i+1>width-1? width-1 : x+i+1][y+j+1>height-1? height-1 : y+j+1].
[0377] It can be understood that in the embodiments of the present application, the CENTER(x+i, y+j) pixel position can be the central position among the above 9 adjacent pixel positions that form a rectangle. Then, based on the (x+i, y+j) pixel position and the other 8 adjacent pixel positions, secondary prediction processing can be performed. Specifically, the other 8 pixel positions are respectively UP (upper), UPRIGHT (upper right), LEFT (left), RIGHT (right), DOWNLEFT (lower left), DOWN (lower), DOWNRIGHT (lower right), and UPLEFT (upper left).
[0378] It should be noted that in the present application, the accuracy of the calculation formula of PredMatrixS[x+i][y+j] can use a lower accuracy. For example, shift each term on the right side of each multiplication. For example, both dMv[i][j][0] and dMv[i][j][1] are shifted right by shift3 bits. Correspondingly, 1<<15 becomes 1<<(15-shift3),...+(1<<10))>>11 becomes...+(1<<(10-shift3)))>>(11-shift3).
[0379] Exemplarily, the magnitude of the motion vector can be restricted within a reasonable range. For example, the positive and negative values of the motion vector in the horizontal and vertical directions used above do not exceed 1 pixel or 1 / 2 pixel or 1 / 4 pixel, etc.
[0380] It can be understood that if the prediction reference mode of the current block is 'Pred_List01', the decoder averages multiple prediction sample matrices of each component to obtain the final prediction sample matrix of the component. For example, two luminance prediction sample matrices are averaged to obtain a new luminance prediction sample matrix.
[0381] Furthermore, in the embodiments of the present application, after obtaining the prediction sample matrix of the current block, if the current block has no transform coefficients, the prediction matrix is used as the decoding result of the current block. If the current block still has transform coefficients, then the transform coefficients can be decoded first, and the residual matrix can be obtained through inverse transform and inverse quantization, and the residual matrix is added to the prediction matrix to obtain the decoding result.
[0382] In summary, through the inter-frame prediction method proposed in steps 201 to 205, after sub-block-based prediction, for pixel positions with motion vectors deviated from the motion vectors of the sub-blocks, point-based secondary prediction is performed on the basis of sub-block-based prediction. Point-based secondary prediction uses the information of multiple points forming a preset shape such as a rectangle or a rhombus. Point-based secondary prediction uses a two-dimensional filter. The two-dimensional filter is a filter formed by adjacent points forming a preset shape. The adjacent points forming a preset shape can be nine points. For a pixel position, the result of filter processing is the new prediction value of this position.
[0383] The inter-frame prediction method proposed in the present application is applicable to the case where the motion vector of the pixel position is deviated from the motion vector of the sub-block. That is to say, in the present application, the decoder performs secondary prediction because the pixel position is not fully predicted after sub-block-based prediction. If the motion vector of the pixel position is not deviated from the motion vector of the sub-block, then the second prediction value after prediction using the inter-frame prediction method proposed in the present application should be the same as the first prediction value before using the inter-frame prediction method proposed in the present application. In the present application, when the motion vector of the pixel position is not deviated from the motion vector of the sub-block, it is also possible to choose not to use the inter-frame prediction method proposed in the present application for this pixel position.
[0384] The inter-frame prediction method proposed in this application can be applicable to the situation where the motion vectors of all pixel positions deviate from the motion vectors of the sub-blocks, including cases such as the affine prediction model where the motion vectors of each pixel position can be calculated, but only sub-block-based prediction is performed, that is, the motion vectors of all pixel positions in the current block can be calculated, and the motion vectors of all pixel positions in the current block are not the same as the motion vectors of the sub-block where it is located. It can also be applicable to the situation where the motion vectors of not all pixel positions can be calculated according to a model, but the motion vectors of some pixel positions inside the current block, such as a sub-block, change and are different from the motion vectors of other positions. Assuming that these contents change continuously at the time point of the current frame, then the motion vectors between these adjacent parts with different motion vectors, such as sub-blocks, may change continuously. At this time, after recalculating the motion vectors of some related pixel positions by means of a calculation model, etc., the situation where these motion vectors deviate from the original motion vectors. That is to say, the inter-frame prediction method proposed in this application can be applicable to affine prediction or other scenarios, and does not limit the application scenarios.
[0385] The two-dimensional filter used in the inter-frame prediction method proposed in this application can be a rectangular filter, a diamond filter, or a filter of other shapes. Exemplarily, the two-dimensional filter is a filter composed of 9 adjacent pixel positions forming a rectangle. The pixel position at the center among the 9 pixel positions is the pixel position to be predicted currently.
[0386] It should be noted that the inter-frame prediction method proposed in this application can be applicable to any image component. In this embodiment, a quadratic prediction scheme is exemplarily used for the luminance component, but it can also be used for the chrominance component or any component in other formats. The inter-frame prediction method proposed in this application can also be applicable to any video format, including but not limited to the YUV format, including but not limited to the luminance component of the YUV format.
[0387] This embodiment provides an inter-frame prediction method. The decoder parses the bitstream to obtain the prediction mode parameter of the current block. When the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, the first motion vector of the sub-blocks of the current block is determined. Wherein, the current block includes multiple sub-blocks. Based on the first motion vector, the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block are determined. Wherein, the pixel position is the position of the pixel point within the sub-block. The filtering coefficient of the two-dimensional filter is determined according to the motion vector deviation. Wherein, the two-dimensional filter is used for secondary prediction processing according to a preset shape. Based on the filtering coefficient and the first prediction value, the second prediction value of the sub-block is determined, and the second prediction value is determined as the inter-frame prediction value of the sub-block. That is to say, the inter-frame prediction method proposed in this application can, after the prediction based on the sub-block, perform point-based secondary prediction on the pixel positions where the motion vector deviates from the motion vector of the sub-block, based on the first prediction value of the sub-block, to obtain the second prediction value. Among them, the point-based secondary prediction using the two-dimensional filter can use the information of several points constituting the preset shape to correct the first prediction value, and finally obtain the corrected second prediction value. The inter-frame prediction method proposed in this application can be well applied to all scenarios, greatly improving the coding performance, thereby improving the encoding and decoding efficiency.
[0388] Based on the above embodiment, in another embodiment of the present application, Figure 20 Schematic diagram of the implementation process of the inter-frame prediction method Figure 2 , as Figure 20 shown, the method for the decoder to perform inter-frame prediction may include the following steps:
[0389] Step 301: Predict the sub-block according to the first motion vector of the sub-block to obtain the first prediction value.
[0390] Step 302: Determine the motion vector deviation between each position within the sub-block and the motion vector of the sub-block.
[0391] Step 303: Filter the first prediction value using the two-dimensional filter according to the motion vector deviation of each position to obtain the second prediction value.
[0392] That is to say, the inter-frame prediction method proposed in this application can, after the prediction based on the sub-block, perform point-based secondary prediction on the pixel positions where the motion vector deviates from the motion vector of the sub-block, based on the prediction of the sub-block, and finally complete the correction of the first prediction value to obtain a new prediction value, that is, the second prediction value.
[0393] Specifically, point-based quadratic prediction uses a two-dimensional filter. The two-dimensional filter is a filter formed by adjacent points that form a preset shape. The adjacent points that form the preset shape can be nine points. For a pixel position, the result of the filter processing is the new predicted value at that position. Among them, the filtering coefficients of the two-dimensional filter are determined by the motion vector deviation at each position. The input of the two-dimensional filter is the first predicted value, and the output is the second predicted value.
[0394] Furthermore, in this application, if the affine mode is used for the current block in this application, then it is necessary to first determine the motion vectors of the control points. Figure 21 Schematic diagram of the implementation process of the inter-frame prediction method Figure 3 , such as Figure 21 shown, the method for the decoder to perform inter-frame prediction may include the following steps:
[0395] Step 304: Determine the second motion vector of the control point.
[0396] Step 305: Determine the first motion vector of the sub-block according to the second motion vector.
[0397] Step 301: Perform prediction on the sub-block according to the first motion vector of the sub-block to obtain the first predicted value.
[0398] Step 302: Determine the motion vector deviation between each position in the sub-block and the motion vector of the sub-block.
[0399] Step 303: Filter the first predicted value using a two-dimensional filter according to the motion vector deviation at each position to obtain the second predicted value.
[0400] It can be seen that in the embodiments of this application, after determining the motion vectors of the control points, the motion vectors of the control points can be used to perform prediction processing on the sub-block, and after determining the motion vector deviation between the pixel positions in the sub-block and the motion vector of the sub-block, for the pixel positions with a deviation between the motion vector and the motion vector of the sub-block, based on the prediction of the sub-block, point-based quadratic prediction is performed, and finally the correction of the first predicted value is completed to obtain a new predicted value, that is, the second predicted value.
[0401] Furthermore, in this application, the two steps of determining the first motion vector of the sub-block and determining the motion vector deviation between each position in the sub-block and the motion vector of the sub-block can also be performed simultaneously. Figure 22 Schematic diagram of the implementation process of the inter-frame prediction method Figure 4 , such as Figure 22 shown, the method for the decoder to perform inter-frame prediction may include the following steps:
[0402] Step 304: Determine the second motion vector of the control point.
[0403] Step 306: Determine the first motion vector of the sub-block and the motion vector deviation of each position within the sub-block according to the second motion vector.
[0404] Step 301: Perform prediction on the sub-block according to the first motion vector of the sub-block to obtain a first prediction value.
[0405] Step 303: Filter the first prediction value by using a two-dimensional filter according to the motion vector deviation of each position to obtain a second prediction value.
[0406] It can be seen that in the embodiment of the present application, after determining the motion vector of the control point, the first motion vector of the sub-block and the motion vector deviation of each position within the sub-block can be determined simultaneously. Then, for the pixel positions with a deviation between the motion vector and the motion vector of the sub-block, a point-based secondary prediction is performed on the basis of the prediction based on the sub-block, and finally the first prediction value is corrected to obtain a new prediction value, that is, the second prediction value.
[0407] It should be noted that in the embodiment of the present application, since the affine model can explicitly calculate the motion vector of each pixel position, or rather the deviation between each pixel position within the sub-block and the motion vector of the sub-block, therefore, the inter-frame prediction method proposed in the present application can be used to improve the affine prediction. Of course, this method can also be applied to the improvement of other sub-block-based predictions.
[0408] Furthermore, the inter-frame prediction method proposed in the present application can be based on the AVS3 standard or can also be applied to the VVC standard, and the present application does not make specific limitations.
[0409] Exemplarily, in the present application, parse the sequence header information of the bitstream to determine whether to use the secondary prediction method proposed in the present application when decoding the sequence. Generally, the control level of this secondary prediction method is at the sequence level. If so, judge according to the above method. If the control level of this secondary prediction method is at the picture level, that is, it is necessary to judge whether to use the secondary prediction method proposed in the present application when decoding each frame of image, then when decoding each frame of image, parse the picture header information of the bitstream to determine whether to use this technology when decoding this frame of image. If this technology is used for the current sequence or the current frame of image, then the coding units that meet the conditions in the current sequence or the current frame of image use the secondary prediction method proposed in the present application. In this embodiment, this technology is used for the coding units decoded in the affine mode, that is, correspondingly, the coding units using the affine mode in the current sequence or the current frame of image use this technology. The affine mode is controlled at the sequence level. Whether this sequence uses the affine mode needs to parse the information of the sequence header. The text description is the sequence header definition shown in Table 6:
[0410] Table 6
[0411]
[0412] Among them, the affine motion compensation enable flag affine_enable_flag is a binary variable. A value of '1' indicates that affine motion compensation can be used; a value of '0' indicates that affine motion compensation should not be used. The value of AffineEnableFlag is equal to the value of affine_enable_flag.
[0413] Furthermore, if sequence-level control is used, then the sequence header definitions shown in Table 7 below can be added to the text:
[0414] Table 7
[0415]
[0416] Among them, the secondary prediction enable flag secondary_pred_enable_flag is a binary variable. A value of '1' indicates that secondary prediction can be used; a value of '0' indicates that secondary prediction should not be used. The value of SencondaryPredEnableFlag is equal to the value of secondary_pred_enable_flag.
[0417] Furthermore, if the secondary prediction method proposed in this application is only used for the affine mode, then the sequence header definitions shown in Table 8 below can also be added to the text:
[0418] Table 8
[0419]
[0420]
[0421] Furthermore, in AVS3, the size of the sub-blocks for affine prediction can be 4x4 or 8x8. Which size to choose is controlled at the picture level, and the text describes the inter-picture prediction picture header definitions shown in Table 9 below:
[0422] Table 9
[0423]
[0424] Among them, the affine prediction sub-block size flag affine_subblock_size_flag is a binary variable. A value of '0' indicates that the minimum size of the affine prediction sub-block of the current image is 4x4; a value of '1' indicates that the minimum size is 8x8. The value of AffineSubblockSizeFlag is equal to the value of affine_subblock_size_flag. If affine_subblock_size_flag does not exist in the bitstream, the value of AffineSubblockSizeFlag is equal to 0.
[0425] When decoding the current block, it is necessary to parse whether it uses the affine mode. In AVS3, if the current block uses the affine mode, it is also necessary to determine whether it uses the 4-parameter (2 control points) mode or the 6-parameter (3 control points) mode, the prediction reference mode, and the motion information of each control point. The prediction reference modes include the prediction reference mode 'Pred_List0' that only refers to reference frame list 0, the prediction reference mode 'Pred_List1' that only refers to reference frame list 1, and the prediction reference mode 'Pred_List01' that refers to both reference frame list 0 and reference frame list 1. Determining whether it uses the 4-parameter (2 control points) mode or the 6-parameter (3 control points) mode, the prediction reference mode, and the motion information of each control point have different determination methods in different modes (direct mode, skip mode, normal mode).
[0426] The affine mode flag affine_flag is a binary variable. A value of '1' indicates that the current block is in the affine mode; a value of '0' indicates that it is not in the affine mode. The value of AffineFlag is equal to the value of affine_flag. If affine_flag does not exist in the bitstream, the value of AffineFlag is 0.
[0427] The affine motion vector index cu_affine_cand_idx, which is the affine mode index value in the skip mode or direct mode. The value of AffineCandIdx is equal to the value of cu_affine_cand_idx. If cu_affine_cand_idx does not exist in the bitstream, the value of AffineCandIdx is equal to 0.
[0428] The affine adaptive motion vector precision index affine_amvr_index is used to determine the affine motion vector precision of the coding unit. The value of AffineAmvrIndex is equal to the value of affine_amvr_index. If affine_amvr_index does not exist in the bitstream, the value of AffineAmvrIndex is equal to 0.
[0429] The absolute value of the horizontal component difference of the motion vector in affine inter prediction mode L0, mv_diff_x_abs_l0_affine, the absolute value of the vertical component difference of the motion vector in affine inter prediction mode L0, mv_diff_y_abs_l0_affine, the absolute value of the motion vector difference from reference picture list 0. MvDiffXAbsL0Affine is equal to the value of mv_diff_x_abs_l0_affine, and MvDiffYAbsAffineAffine is equal to the value of mv_diff_y_abs_l0_affine.
[0430] The sign value of the horizontal component difference of the motion vector in affine inter prediction mode L0, mv_diff_x_sign_l0_affine, the sign value of the vertical component difference of the motion vector in affine inter prediction mode L0, mv_diff_y_sign_l0_affine, the sign bit of the motion vector difference from reference picture list 0. The value of MvDiffXSignL0Affine is equal to the value of mv_diff_x_sign_l0_affine, and the value of MvDiffYSignL0Affine is equal to the value of mv_diff_y_sign_l0_affine. If mv_diff_x_sign_l0_affine or mv_diff_y_sign_l0_affine does not exist in the bitstream, the value of MvDiffXSignL0Affine or MvDiffYSignL0Affine is 0. If the value of MvDiffXSignL0Affine is 0, MvDiffXL0Affine is equal to MvDiffXAbsL0Affine; if the value of MvDiffXSignL0Affine is 1, MvDiffXL0Affine is equal to -MvDiffXAbsL0Affine. If the value of MvDiffYSignL0Affine is 0, MvDiffYL0Affine is equal to MvDiffYAbsL0Affine; if the value of MvDiffYSignL0Affine is 1, MvDiffYL0Affine is equal to -MvDiffYAbsL0Affine. The value range of MvDiffXL0Affine and MvDiffYL0Affine is -32768 to 32767.
[0431] The absolute value of the horizontal component difference of the motion vector in affine inter-frame mode L1, mv_diff_x_abs_l1_affine, the absolute value of the vertical component difference of the motion vector in affine inter-frame mode L1, mv_diff_y_abs_l1_affine, the absolute value of the motion vector difference from reference picture list 1. MvDiffXAbsL1Affine is equal to the value of mv_diff_x_abs_l1_affine, and MvDiffYAbsAffineAffine is equal to the value of mv_diff_y_abs_l1_affine.
[0432] The sign value of the horizontal component difference of the motion vector in affine inter-frame mode L1, mv_diff_x_sign_l1_affine, the sign value of the vertical component difference of the motion vector in affine inter-frame mode L1, mv_diff_y_sign_l1_affine, the sign bit of the motion vector difference from reference picture list 1. The value of MvDiffXSignL1Affine is equal to the value of mv_diff_x_sign_l1_affine, and the value of MvDiffYSignL1Affine is equal to mv_diff_y_sign_l1_affine. If mv_diff_x_sign_l1_affine or mv_diff_y_sign_l1_affine does not exist in the bitstream, the value of MvDiffXSignL1Affine or MvDiffYSignL1Affine is 0. If the value of MvDiffXSignL1Affine is 0, MvDiffXL1Affine is equal to MvDiffXAbsL1Affine; if the value of MvDiffXSignL1Affine is 1, MvDiffXL1Affine is equal to -MvDiffXAbsL1Affine. If the value of MvDiffYSignL1Affine is 0, MvDiffYL1Affine is equal to MvDiffYAbsL1Affine; if the value of MvDiffYSignL1Affine is 1, MvDiffYL0Affine is equal to -MvDiffYAbsL1Affine. The value range of MvDiffXL1Affine and MvDiffYL1Affine is -32768 to 32767.
[0433] Furthermore, in an embodiment of the present application, if the current block uses the affine mode, and the control point mode, prediction reference mode, and sub-block size parameter of the affine mode are also determined, the motion vector of each sub-block can be derived using the affine model.
[0434] If the prediction reference mode of the current prediction unit (where the prediction unit is equal to the coding block) is 'Pred_List0', mvs_L0 (the set of affine control point motion vectors corresponding to reference frame list 0) is used as the set of affine control point motion vectors, and the L0 motion vector set of the current prediction unit (denoted as MvArrayL0) is obtained by the method of deriving the motion vector array of the affine motion unit sub-blocks. The MvArrayL0 set consists of the L0 motion vectors of all luminance prediction sub-blocks.
[0435] If the prediction reference mode of the current prediction unit (where the prediction unit is equal to the coding block) is 'Pred_List1', mvs_L1 (the set of affine control point motion vectors corresponding to reference frame list 1) is used as the set of affine control point motion vectors, and the L1 motion vector set of the current prediction unit (denoted as MvArrayL1) is obtained by the method of deriving the motion vector array of the affine motion unit sub-blocks. MvArrayL1 consists of the L1 motion vectors of all luminance prediction sub-blocks.
[0436] If the prediction reference mode of the current prediction unit (where the prediction unit is equal to the coding block) is 'Pred_List01', mvs_L0 and mvs_L1 are respectively used as the sets of affine control point motion vectors and the L0 motion vector set (denoted as MvArrayL0) and L1 motion vector set (denoted as MvArrayL1) of the current prediction unit are obtained according to the method defined in "deriving the motion vector array of the affine motion unit sub-blocks". MvArrayL0 and MvArrayL1 respectively consist of the L0 motion vectors and L1 motion vectors of all luminance prediction sub-blocks.
[0437] Next, the matrix of chrominance motion vectors can be further obtained, and the affine luminance sample interpolation and affine chrominance sample interpolation are used to obtain the luminance prediction sample matrix in the affine mode, as well as the chrominance prediction sample matrix.
[0438] If the current sequence or the current frame image does not use this technology, if the prediction reference mode of the current prediction unit (where the prediction unit is equal to the coding block) is 'Pred_List0' or 'Pred_List1', then the obtained luminance prediction sample matrix and chrominance prediction sample matrix are the prediction sample matrices of the current prediction unit. If the prediction reference mode of the current prediction unit (where the prediction unit is equal to the coding block) is 'Pred_List01', then the two obtained luminance prediction sample matrices are averaged to obtain the luminance prediction sample matrix of the current prediction unit. For each chrominance component, the two groups (two in each group, a total of four) of chrominance prediction sample matrices obtained in the previous step are averaged to obtain the chrominance prediction sample matrix of the current prediction unit.
[0439] If the current sequence or the current frame image uses this technology, export the internal pixel motion vector deviation matrix of the affine motion unit sub-block:
[0440] If the prediction reference mode of the current prediction unit (here the prediction unit is equal to the coding block) is 'Pred_List0', use mvs_L0 (the affine control point motion vector group corresponding to reference frame list 0) as the affine control point motion vector group, and obtain the 4 sub-block internal motion vector deviation matrix motion vector sets dMvA_L0, dMvB_L0, dMvC_L0, dMvN_L0 of L0 of the current prediction unit through the method of exporting the internal pixel motion vector deviation matrix of the affine motion unit sub-block. Then use dMvA, dMvB, dMvC, dMvN exported from the internal pixel motion vector deviation matrix of the affine motion unit sub-block as dMvA_L0, dMvB_L0, dMvC_L0, dMvN_L0 respectively.
[0441] If the prediction reference mode of the current prediction unit (here the prediction unit is equal to the coding block) is 'Pred_List1', use mvs_L1 (the affine control point motion vector group corresponding to reference frame list 1) as the affine control point motion vector group, and obtain the 4 sub-block internal motion vector deviation matrix motion vector sets dMvA_L1, dMvB_L1, dMvC_L1, dMvN_L1 of L1 of the current prediction unit through the method of exporting the internal pixel motion vector deviation matrix of the affine motion unit sub-block. Then use dMvA, dMvB, dMvC, dMvN exported from the internal pixel motion vector deviation matrix of the affine motion unit sub-block as dMvA_L1, dMvB_L1, dMvC_L1, dMvN_L1 respectively.
[0442] If the prediction reference mode of the current prediction unit (here the prediction unit is equal to the coding block) is 'Pred_List01', use mvs_L0 and mvs_L1 as the affine control point motion vector groups respectively, and obtain the 4 sub-block internal motion vector deviation matrix motion vector sets dMvA_L0, dMvB_L0, dMvC_L0, dMvN_L0 of L0 of the current prediction unit and the 4 sub-block internal motion vector deviation matrix motion vector sets dMvA_L1, dMvB_L1, dMvC_L1, dMvN_L1 of L1 through the method of exporting the internal pixel motion vector deviation matrix of the affine motion unit sub-block.
[0443] Furthermore, in this application, the internal pixel motion vector deviation matrix of the affine motion unit sub-block can be exported according to the following method:
[0444] If there are 3 motion vectors in the affine control point motion vector group, the motion vector group is expressed as mvsAffine(mv0, mv1, mv2); otherwise (if there are 2 motion vectors in the affine control point motion vector group), the motion vector group is expressed as mvsAffine(mv0, mv1).
[0445] 1. Calculate variables dHorX, dVerX, dHorY, and dVerY:
[0446] dHorX = (mv1_x - mv0_x) << (7 - Log(width))
[0447] dHorY = (mv1_y - mv0_y) << (7 - Log(width))
[0448] If there are 3 motion vectors in mvsAffine, then:
[0449] dVerX = (mv2_x - mv0_x) << (7 - Log(height))
[0450] dVerY = (mv2_y - mv0_y) << (7 - Log(height))
[0451] Otherwise (if there are 2 motion vectors in mvsAffine):
[0452] dVerX = -dHorY
[0453] dVerY = dHorX.
[0454] As Figure 7 , assume that the width and height of the previous prediction unit are width and height respectively, and the width and height of each sub-block are subwidth and subheight respectively. The sub-block where the top-left sample of the current prediction unit luminance prediction block is located is A, the sub-block where the top-right sample is located is B, the sub-block where the bottom-left sample is located is C, and the sub-blocks at other positions are other sub-blocks.
[0455] (i, j) are the coordinates of the pixels inside the luminance sub-block, the value range of i is 0 to (subwidth - 1), the value range of j is 0 to (subheight - 1), and calculate the motion vector deviation of each pixel (i, j) position inside the 4 luminance prediction sub-blocks:
[0456] 2.1. If the current sub-block is A, the motion vector deviation dMvA[i][j] of the (i, j) pixel:
[0457] dMvA[i][j][0] = dHorX × i + dVerX × j
[0458] dMvA[i][j][1] = dHorY * i + dVerY * j;
[0459] 2.2. If the current sub - block is B, the motion vector deviation dMvB[i][j] of the pixel at (i, j):
[0460] dMvB[i][j][0] = dHorX * (i - subwidth) + dVerX * j
[0461] dMvB[i][j][1] = dHorY * (i - subwidth) + dVerY * j;
[0462] 2.3. If the current sub - block is C and there are 3 motion vectors in mvAffine, the motion vector deviation dMvC[i][j] of the pixel at (i, j):
[0463] dMvC[i][j][0] = dHorX * i + dVerX * (j - subheight)
[0464] dMvC[i][j][1] = dHorY * i + dVerY * (j - subheight;
[0465] 2.4. The motion vector deviation dMvN[i][j] of the pixel at (i, j):
[0466] dMvN[i][j][0] = dHorX * (i - (subwidth >> 1)) + dVerX * (j - (subheight >> 1))
[0467] dMvN[i][j][1] = dHorY * (i - (subwidth >> 1)) + dVerY * (j - (subheight >> 1)).
[0468] Among them, dMvX[i][j][0] represents the value of the horizontal component, and dMvX[i][j][1] represents the value of the vertical component. X is A, B, C or N.
[0469] Furthermore, in this application, after deriving the motion vector deviation matrix of the internal pixels of the affine motion unit sub - block, for each sub - block and for each pixel position, according to the motion vector deviation, based on the first prediction value of the motion compensation of the sub - block, a two - dimensional filter is used for filtering to obtain a new second prediction value. In this embodiment, this inter - frame prediction technique is used for the luminance component. However, this inter - frame prediction technique can also be used for the chrominance component, or any component in other formats.
[0470] Specifically, in the embodiments of the present application, when the two-dimensional filter performs secondary prediction using 9 adjacent pixel positions forming a rectangle, it can first determine the prediction sample matrix of the current block and the set of motion vector deviations of the sub-blocks of the current block; wherein, the motion vector deviation matrix includes the motion vector deviations corresponding to all pixel positions; then, based on the 9 adjacent pixel positions forming a rectangle, using the prediction sample matrix and the set of motion vector deviations, determine the sample matrix after secondary prediction of the current block.
[0471] Exemplarily, in the present application, the process of the two-dimensional filter performing secondary prediction using 9 adjacent pixel positions forming a rectangle is as follows:
[0472] If the prediction reference mode of the current prediction unit (here the prediction unit is equal to the coding block) is 'Pred_List0', take the luminance prediction sample matrix PredMatrixL0 as PredMatrixSb, and obtain PredMatrixS as the new luminance prediction sample matrix PredMatrixL0 through the method of deriving the prediction matrix of affine secondary prediction.
[0473] If the prediction reference mode of the current prediction unit (here the prediction unit is equal to the coding block) is 'Pred_List1', take the luminance prediction sample matrix PredMatrixL1 as PredMatrixSb, and obtain PredMatrixS as the new luminance prediction sample matrix PredMatrixL1 through the method of deriving the prediction matrix of affine secondary prediction.
[0474] If the prediction reference mode of the current prediction unit (here the prediction unit is equal to the coding block) is 'Pred_List01', take the luminance prediction sample matrices PredMatrixL0 and PredMatrixL1 as PredMatrixSb respectively, and obtain PredMatrixS as the new luminance prediction sample matrices PredMatrixL0 and PredMatrixL1 through the method of deriving the prediction matrix of affine secondary prediction.
[0475] Furthermore, in the present application, the method for deriving the prediction matrix of affine secondary prediction can be as follows:
[0476] As Figure 7 , assume that the width and height of the current block are width and height respectively, and the width and height of each sub-block are subwidth and subheight respectively. The sub-block where the top-left sample of the current block luminance prediction sample matrix is located is A, the sub-block where the top-right sample is located is B, the sub-block where the bottom-left sample is located is C, and the sub-blocks where other positions are located are other sub-blocks.
[0477] For each sub-block in the current block, denote the motion vector deviation matrix of the sub-block as dMv. Perform the following operations:
[0478] 1. If the sub-block is A, dMv is equal to dMvA;
[0479] 2. If the sub-block is B, dMv is equal to dMvB;
[0480] 3. If the sub-block is C and there are 3 motion vectors in mvAffine, dMv is equal to dMvC;
[0481] 4. dMv is equal to dMvN.
[0482] (x, y) are the coordinates of the upper-left corner position of the sub-block, (i, j) are the coordinates of the pixels inside the luminance sub-block, the value range of i is 0 to (subwidth - 1), the value range of j is 0 to (subheight - 1), the prediction sample matrix based on the sub-block is PredMatrixSb, and the prediction sample matrix of the quadratic prediction is PredMatrixS. Calculate the prediction sample PredMatrixS[x + i][y + j] of the quadratic prediction of (x + i, y + j) according to the following method:
[0483] PredMatrixS[x + i][y + j] = (UPLEFT(x + i, y + j) × (-dMv[i][j][0] - dMv[i][j][1]) + UP(x + i, y + j) × ((-dMv[i][j][1]) << 3) +
[0484] UPRIGHT(x + i, y + j) × (dMv[i][j][0] - dMv[i][j][1]) +
[0485] LEFT(x + i, y + j) × ((-dMv[i][j][0]) << 3) +
[0486] CENTER(x + i, y + j) × (1 << 15) +
[0487] RIGHT(x + i, y + j) × (dMv[i][j][0] << 3) +
[0488] DOWNLEFT(x + i, y + j) × (-dMv[i][j][0] + dMv[i][j][1]) +
[0489] DOWN(x + i, y + j) × (dMv[i][j][1] << 3) +
[0490] DOWNRIGHT(x + i, y + j) × (dMv[i][j][0] + dMv[i][j][1]) +
[0491] ((1 << 10)) >> 11
[0492] PredMatrixS[x + i][y + j] = Clip3(0, (1 << BitDepth) - 1, PredMatrixS[x + i][y + j]).
[0493] where UPLEFT(x + i, y + j) = predMatrix[x + i - 1 < 0? 0 : x + i - 1][y + j - 1 < 0? 0 : y + j - 1]
[0494] UP(x + i, y + j) = predMatrix[x + i][y + j - 1 < 0? 0 : y + j - 1]
[0495] UPRIGHT(x + i, y + j) = predMatrix[x + i + 1 > width - 1? width - 1 : x + i + 1][y + j - 1 < 0? 0 : y + j - 1]
[0496] LEFT(x + i, y + j) = predMatrix[x + i - 1 < 0? 0 : x + i - 1][y + j]
[0497] CENTER(x + i, y + j) = predMatrix[x + i][y + j]
[0498] RIGHT(x + i, y + j) = predMatrix[x + i + 1 > width - 1? width - 1 : x + i + 1][y + j]
[0499] DOWNLEFT(x + i, y + j) = predMatrix[x + i - 1 < 0? 0 : x + i - 1][y + j + 1 < 0? 0 : y + j - 1]
[0500] DOWN(x + i, y + j) = predMatrix[x + i][y + j - 1 < 0? 0 : y + j - 1]
[0501] DOWNRIGHT(x + i, y + j) = predMatrix[x + i + 1 > width - 1? width - 1 : x + i + 1][y + j + 1 > height - 1? height - 1 : y + j + 1].
[0502] It can be understood that in the embodiments of the present application, the CENTER(x+i, y+j) pixel position can be the central position among the above-mentioned 9 adjacent pixel positions that form a rectangle. Then, based on the (x+i, y+j) pixel position and the other 8 adjacent pixel positions, a secondary prediction process can be performed. Specifically, the other 8 pixel positions are respectively UP (up), UPRIGHT (upper right), LEFT (left), RIGHT (right), DOWNLEFT (lower left), DOWN (down), DOWNRIGHT (lower right), and UPLEFT (upper left).
[0503] It should be noted that in the embodiments of the present application, when filtering each pixel position in the current block, the value of the corresponding pixel position used by the two-dimensional filter is the predicted value based on the prediction of the sub-block, and the filtered result is the new predicted value obtained by the present technology. Specifically, according to the shape of the two-dimensional filter (such as a rectangle or a rhombus), it can be known that the pixel positions required by the two-dimensional filter may exceed the boundary of the current block. Figure 23 For the schematic diagram of filtering the boundary pixel position, as Figure 23 shown, taking the above-mentioned 9-point rectangular filter as an example, if the pixel position to be filtered is on the boundary of the current block, then when filtering this point, some of the corresponding pixel positions of the filter will exceed the boundary of the current block.
[0504] To address the above problem, the present application proposes a solution, specifically by expanding the boundary of the current block based on the prediction of the sub-block, so that the pixel positions required by the two-dimensional filter do not exceed the boundary of the expanded block. That is to say, if the pixel positions to be used by the two-dimensional filter do not belong to the current block, then the current block can be expanded based on the boundary position of the current block. Figure 24 For the schematic diagram of the boundary expansion, as Figure 24 shown, taking the above-mentioned 9-point rectangular filter as an example, the current block based on the prediction of the sub-block needs to expand one pixel position on each of the upper, lower, left, and right boundaries. A simple expansion method is that the pixels expanded on the left copy the pixel values of the left boundary corresponding to their horizontal direction, the pixels expanded on the right copy the pixel values of the right boundary corresponding to their horizontal direction, the pixels expanded on the upper copy the pixel values of the upper boundary corresponding to their vertical direction, the pixels expanded on the lower copy the pixel values of the lower boundary corresponding to their vertical direction, and the four vertices of the expansion can copy the pixel values of their corresponding vertices.
[0505] It can be understood that the method of ensuring that the pixel positions required by the two-dimensional filter do not exceed the boundary of the expanded block through the expansion process can also be applied to two-dimensional filters of other shapes, and the present application does not make specific limitations.
[0506] In view of the above problems, the present application also proposes a solution. If the pixel position corresponding to a certain position of the two-dimensional filter exceeds the boundary of the current block, then adjust its corresponding pixel position to a position within the current block. That is to say, if the pixel position to be used by the two-dimensional filter does not belong to the current block, then the pixel position within the current block can be used to replace that pixel position. For example, assume that the upper left pixel position of the current block is (0, 0), the width of the current block is width, and the height of the current block is height. Then the horizontal range of the current block is 0 to (width - 1), and the vertical range of the current block is 0 to (height - 1). If the pixel position required by the filter is (x, y), if x is less than 0, then set x to 0. If x is greater than width - 1, then set x to width - 1. If y is less than 0, then set y to 0. If y is greater than height - 1, then set y to height - 1.
[0507] It can be understood that the method of using the pixel position within the current block to replace the pixel position exceeding the current block to ensure that the pixel position required by the two-dimensional filter does not exceed the boundary of the extended block can also be applied to two-dimensional filters of other shapes, and the application does not make specific limitations.
[0508] An embodiment of the present application provides an inter-frame prediction method. After prediction based on sub-blocks, for pixel positions with a deviation between the motion vector and the motion vector of the sub-block, a point-based secondary prediction can be performed on the basis of the first prediction value based on the sub-block to obtain a second prediction value. Among them, the point-based secondary prediction using a two-dimensional filter can use the information of several points forming a preset shape to correct the first prediction value, and finally obtain the corrected second prediction value. The inter-frame prediction method proposed in the present application can be well applied to all scenarios, greatly improving the coding performance, thereby improving the encoding and decoding efficiency.
[0509] An embodiment of the present application provides an inter-frame prediction method, which is applied to a video coding device, that is, an encoder. The functions implemented by this method can be realized by a second processor in the encoder calling a computer program. Of course, the computer program can be stored in a second memory. It can be seen that the encoder at least includes a second processor and a second memory.
[0510] Figure 25 Schematic diagram of the implementation process of the inter-frame prediction method Figure 5 , such as Figure 25 shown, the method for the encoder to perform inter-frame prediction may include the following steps:
[0511] Step 401, determine the prediction mode parameters of the current block.
[0512] In an embodiment of the present application, the encoder may first determine the prediction mode parameters of the current block. Specifically, the encoder may first determine the prediction mode used for the current block, and then determine the corresponding prediction mode parameters based on this prediction mode. Among them, the prediction mode parameters may be used to determine the prediction mode used for the current block.
[0513] It should be noted that, in an embodiment of the present application, the image to be encoded may be divided into multiple image blocks. The currently to-be-encoded image block may be referred to as the current block, and the image blocks adjacent to the current block may be referred to as adjacent blocks; that is, in the image to be encoded, there is an adjacent relationship between the current block and the adjacent blocks. Here, each current block may include a first image component, a second image component, and a third image component; that is, the current block is the image block in the image to be encoded for which the first image component, the second image component, or the third image component prediction is to be performed currently.
[0514] Among them, assuming that the current block performs the first image component prediction, and the first image component is the luminance component, that is, the image component to be predicted is the luminance component, then the current block may also be referred to as a luminance block; or, assuming that the current block performs the second image component prediction, and the second image component is the chrominance component, that is, the image component to be predicted is the chrominance component, then the current block may also be referred to as a chrominance block.
[0515] It should be noted that, in an embodiment of the present application, the prediction mode parameters indicate the prediction mode adopted by the current block and the parameters related to this prediction mode. Here, for the determination of the prediction mode parameters, a simple decision strategy may be adopted, such as determining according to the magnitude of the distortion value; or a complex decision strategy may be adopted, such as determining according to the result of rate distortion optimization (RDO). The embodiments of the present application do not make any limitations. Generally speaking, the RDO method may be used to determine the prediction mode parameters of the current block.
[0516] Specifically, in some embodiments, when determining the prediction mode parameters of the current block, the encoder may first perform pre-encoding processing on the current block using multiple prediction modes to obtain the rate distortion cost values corresponding to each prediction mode; then select the minimum rate distortion cost value from the obtained multiple rate distortion cost values, and determine the prediction mode parameters of the current block according to the prediction mode corresponding to the minimum rate distortion cost value.
[0517] That is to say, on the encoder side, multiple prediction modes can be adopted for the current block to perform pre-encoding processing on the current block respectively. Here, the multiple prediction modes usually include an inter-frame prediction mode, a traditional intra-frame prediction mode, and a non-traditional intra-frame prediction mode; among them, the traditional intra-frame prediction mode can include a direct current (DC) mode, a planar (PLANAR) mode, an angular mode, etc., and the non-traditional intra-frame prediction mode can include a matrix-based intra prediction (MIP) mode, a cross-component linear model prediction (CCLM) mode, an intra block copy (IBC) mode, a PLT (Palette) mode, etc., and the inter-frame prediction mode can include a normal inter-frame prediction mode, a GPM mode, an AWP mode, etc.
[0518] In this way, after pre-encoding the current block using multiple prediction modes respectively, the rate-distortion cost value corresponding to each prediction mode can be obtained; then, the minimum rate-distortion cost value is selected from the obtained multiple rate-distortion cost values, and the prediction mode corresponding to the minimum rate-distortion cost value is determined as the prediction mode parameter of the current block. In addition, after pre-encoding the current block using multiple prediction modes respectively, the distortion value corresponding to each prediction mode can be obtained; then, the minimum distortion value is selected from the obtained multiple distortion values, and then the prediction mode corresponding to the minimum distortion value is determined as the prediction mode used for the current block, and the corresponding prediction mode parameter is set according to this prediction mode. In this way, finally, the current block is encoded using the determined prediction mode parameter, and in this prediction mode, the prediction residual can be made smaller, and the encoding efficiency can be improved.
[0519] That is to say, on the encoding side, the encoder can select the optimal prediction mode to perform pre-encoding on the current block. During this process, the prediction mode of the current block can be determined, and then the prediction mode parameter used to indicate the prediction mode is determined, so as to write the corresponding prediction mode parameter into the code stream and transmit it from the encoder to the decoder.
[0520] Correspondingly, on the decoder side, the decoder can directly obtain the prediction mode parameter of the current block by parsing the code stream, and determine the prediction mode used for the current block and the related parameters corresponding to this prediction mode according to the parsed prediction mode parameter.
[0521] Step 402: When the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, determine the first motion vector of the sub-blocks of the current block; where the current block includes multiple sub-blocks.
[0522] In an embodiment of the present application, if the prediction mode parameter indicates that the current block uses the inter-frame prediction mode to determine the inter-frame prediction value of the current block, the encoder may first determine the first motion vector of each sub-block of the current block. Among them, each sub-block corresponds to a first motion vector.
[0523] It should be noted that, in an embodiment of the present application, the current block is an image block to be encoded in the current frame. The current frame is encoded sequentially in the form of image blocks, and the current block is the next image block to be encoded in the current frame in this order. The current block can have various specifications, such as specifications of 16×16, 32×32, or 32×16, etc., where the numbers represent the number of rows and columns of pixel points on the current block.
[0524] Furthermore, in an embodiment of the present application, the current block can be divided into multiple sub-blocks. Among them, the size of each sub-block is the same, and the sub-block is a set of pixel points with a smaller specification. The size of the sub-block can be 8×8 or 4×4.
[0525] Exemplarily, in the present application, if the size of the current block is 16×16, it can be divided into 4 sub-blocks with a size of 8×8 each.
[0526] It can be understood that, in an embodiment of the present application, when the encoder determines that the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, the inter-frame prediction method provided by the embodiment of the present application can be continued to be adopted.
[0527] In an embodiment of the present application, further, when the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, when the encoder determines the first motion vector of the sub-blocks of the current block, it can determine the affine mode parameter and the prediction reference mode of the current block. When the affine mode parameter indicates using the affine mode, the control point mode and the sub-block size parameter are determined. Finally, the first motion vector can be determined according to the prediction reference mode, the control point mode, and the sub-block size parameter.
[0528] In an embodiment of the present application, after the encoder determines the prediction mode parameter, if the prediction mode parameter indicates that the current block uses the inter-frame prediction mode to determine the inter-frame prediction value of the current block, the encoder can determine the affine mode parameter and the prediction reference mode.
[0529] It should be noted that, in an embodiment of the present application, the affine mode parameter is used to indicate whether to use the affine mode. Specifically, the affine mode parameter can be the affine motion compensation enable flag affine_enable_flag. By determining the value of the affine mode parameter, the encoder can further determine whether to use the affine mode.
[0530] That is to say, in the present application, the affine mode parameter can be a binary variable. If the value of the affine mode parameter is 1, it indicates that the affine mode is used; if the value of the affine mode parameter is 0, it indicates that the affine mode is not used.
[0531] Exemplarily, in the present application, the value of the affine mode parameter can be equal to the value of the affine motion compensation enable flag affine_enable_flag. If the value of affine_enable_flag is '1', it means that affine motion compensation can be used; if the value of affine_enable_flag is '0', it means that affine motion compensation should not be used.
[0532] Furthermore, in the embodiments of the present application, if the affine mode parameter determined by the encoder indicates that the affine mode is used, then the encoder can obtain the control point mode and the sub-block size parameter.
[0533] It should be noted that, in the embodiments of the present application, the control point mode is used to determine the number of control points. In the affine model, a sub-block can have 2 control points or 3 control points. Correspondingly, the control point mode can be the control point mode corresponding to 2 control points, or the control point mode corresponding to 3 control points. That is, the control point mode can include a 4-parameter mode and a 6-parameter mode.
[0534] It can be understood that, in the embodiments of the present application, for the AVS3 standard, if the current block uses the affine mode, then the encoder also needs to determine the number of control points of the current block in the affine mode, so as to determine whether the 4-parameter (2 control points) mode or the 6-parameter (3 control points) mode is used.
[0535] Furthermore, in the embodiments of the present application, if the affine mode parameter determined by the encoder indicates that the affine mode is used, then the encoder can further determine the sub-block size parameter.
[0536] Specifically, the sub-block size parameter can be characterized by the affine prediction sub-block size flag affine_subblock_size_flag. The encoder can set the value of the sub-block size flag to indicate the sub-block size parameter, that is, to indicate the size of the sub-block of the current block. Among them, the size of the sub-block can be 8×8 or 4×4. Specifically, in the present application, the sub-block size flag can be a binary variable. If the value of the sub-block size flag is 1, it indicates that the sub-block size parameter is 8×8; if the value of the sub-block size flag is 0, it indicates that the sub-block size parameter is 4×4.
[0537] Exemplarily, in the present application, the value of the sub-block size flag may be equal to the value of the affine prediction sub-block size flag affine_subblock_size_flag. If the value of affine_subblock_size_flag is '1', the current block is divided into sub-blocks with a size of 8×8; if the value of affine_subblock_size_flag is '0', the current block is divided into sub-blocks with a size of 4×4.
[0538] Further, in the embodiments of the present application, after the encoder determines the control point mode and the sub-block size parameter, it can further determine the first motion vector of the sub-blocks in the current block according to the prediction reference mode, the control point mode, and the sub-block size parameter.
[0539] Specifically, in the embodiments of the present application, the encoder may first determine the control point motion vector group according to the prediction reference mode; then, based on the control point motion vector group, the control point mode, and the sub-block size parameter, it can determine the first motion vector of the sub-block.
[0540] It can be understood that, in the embodiments of the present application, the control point motion vector group can be used to determine the motion vectors of the control points.
[0541] It should be noted that, in the embodiments of the present application, the encoder can traverse each sub-block in the current block according to the above method, and use the control point motion vector group, the control point mode, and the sub-block size parameter of each sub-block to determine the first motion vector of each sub-block, so as to construct a motion vector set according to the first motion vector of each sub-block.
[0542] It can be understood that, in the embodiments of the present application, the motion vector set of the current block may include the first motion vectors of each sub-block of the current block.
[0543] Further, in the embodiments of the present application, when the encoder determines the first motion vector according to the control point motion vector group, the control point mode, and the sub-block size parameter, it may first determine the difference variable according to the control point motion vector group, the control point mode, and the size parameter of the current block; then, based on the prediction mode parameter and the sub-block size parameter, it can determine the sub-block position; finally, it can use the difference variable and the sub-block position to determine the first motion vector of the sub-block, and further obtain the motion vector set of multiple sub-blocks of the current block.
[0544] It should be noted that in this application, when determining the deviation between each position in the sub-block and the motion vector of the sub-block, if the current block uses an affine prediction model, the motion vector of each position in the sub-block can be calculated according to the formula of the affine prediction model, and the deviation between them can be obtained by subtracting the motion vector of the sub-block. If the motion vectors of the sub-blocks all select the motion vectors of the same position within the sub-block, for example, a 4x4 block uses the position (2, 2) from the upper left corner, and an 8x8 block uses the position (4, 4) from the upper left corner, according to the affine models used in current standards including VVC and AVS3, the motion vector deviations of the same position in each sub-block are the same. However, the positions used by AVS at the upper left corner, upper right corner, and the lower left corner in the case of 3 control points (the positions A, B, and C shown in Figure 7 as in the AVS3 text) are different from those used by other blocks. Correspondingly, when calculating the motion vector deviations of the sub-blocks at the upper left corner, upper right corner, and the lower left corner in the case of 3 control points, they are also different from those of other blocks. Specifically, as in the embodiments.
[0545] Step 403: Determine the first prediction value of the sub-block and the deviation of the motion vector between the pixel position and the sub-block based on the first motion vector; wherein, the pixel position is the position of the pixel point within the sub-block.
[0546] In the embodiments of this application, after the encoder determines the first motion vector of each sub-block of the current block, it can determine the first prediction value of the sub-block and the deviation of the motion vector between the pixel position and the sub-block respectively based on the first motion vector of the sub-block.
[0547] It can be understood that in the embodiments of this application, step 403 may specifically include:
[0548] Step 403a: Determine the first prediction value of the sub-block based on the first motion vector.
[0549] Step 403b: Determine the deviation of the motion vector between the pixel position and the sub-block based on the first motion vector.
[0550] Among them, the inter-frame prediction method proposed in the embodiments of this application does not limit the order in which the encoder executes step 403a and step 403b. That is to say, in this application, after determining the first motion vector of each sub-block of the current block, the encoder can first execute step 403a and then execute step 403b, or first execute step 403b and then execute step 403a, or execute step 403a and step 403b simultaneously.
[0551] Further, in the embodiments of the present application, when the encoder determines the first prediction value of the sub-block based on the first motion vector, it may first determine a sample matrix; wherein, the sample matrix includes a luminance sample matrix and a chrominance sample matrix; then, it may determine the first prediction value according to the prediction reference mode, the sub-block size parameter, the sample matrix, and the set of motion vectors.
[0552] It should be noted that, in the embodiments of the present application, when the encoder determines the first prediction value according to the prediction reference mode, the sub-block size parameter, the sample matrix, and the set of motion vectors, it may first determine a target motion vector from the set of motion vectors according to the prediction reference mode and the sub-block size parameter; then, it may use the reference image queue and reference index corresponding to the prediction reference mode, the sample matrix, and the target motion vector to determine a prediction sample matrix; wherein, the prediction sample matrix includes the first prediction values of multiple sub-blocks.
[0553] Specifically, in the embodiments of the present application, the sample matrix may include a luminance sample matrix and a chrominance sample matrix. Correspondingly, the prediction sample matrix determined by the encoder may include a luminance prediction sample matrix and a chrominance prediction sample matrix. Among them, the luminance prediction sample matrix includes the first luminance prediction values of multiple sub-blocks, and the chrominance prediction sample matrix includes the first chrominance prediction values of multiple sub-blocks. The first luminance prediction value and the first chrominance prediction value constitute the first prediction value of the sub-block.
[0554] It should be noted that, in the embodiments of the application, the luminance sample matrix in the sample matrix may be a 1 / 16 precision luminance sample matrix, and the chrominance sample matrix in the sample matrix may be a 1 / 32 precision chrominance sample matrix.
[0555] It can be understood that, in the embodiments of the present application, for different prediction reference modes, the reference image queue and reference index obtained by the encoder are different.
[0556] Further, in the embodiments of the present application, when the encoder determines the sample matrix, it may first obtain luminance interpolation filter coefficients and chrominance interpolation filter coefficients; then, it may determine the luminance sample matrix based on the luminance interpolation filter coefficients, and at the same time, it may determine the chrominance sample matrix based on the chrominance interpolation filter coefficients.
[0557] Further, in the embodiments of the present application, when the encoder is based on the motion vector deviation between the pixel position and the sub-block, it may determine a quadratic prediction parameter; if the quadratic prediction parameter indicates the use of quadratic prediction, then the encoder may determine the motion vector deviation between the sub-block and each pixel position based on the difference variable.
[0558] It can be understood that in the embodiments of the present application, after the encoder determines the motion vector deviation between the sub-block and each pixel position based on the difference variable, it can use all the motion vector deviations corresponding to all the pixel positions within the sub-block to construct a motion vector deviation matrix corresponding to the sub-block. It can be seen that the motion vector deviation matrix includes the motion vector deviation between the sub-block and any internal pixel point, that is, the motion vector deviation.
[0559] Further, in the embodiments of the present application, if the secondary prediction parameter determined by the encoder indicates not to use secondary prediction, then the encoder can directly select the first prediction value of the sub-block of the current block obtained in step 403a above as the second prediction value of the sub-block, without performing the following steps 404 and 405.
[0560] Specifically, in the embodiments of the present application, if the secondary prediction parameter indicates not to use secondary prediction, then the encoder can use the prediction sample matrix to determine the second prediction value. The prediction sample matrix includes the first prediction values of multiple sub-blocks, and the encoder can determine the first prediction value of the sub-block where the pixel position is located as its own second prediction value.
[0561] Exemplarily, in the present application, if the prediction reference mode of the current block takes a value of 0 or 1, that is, the first reference mode 'PRED_List0' is used, or the second reference mode 'PRED_List1' is used, then the first prediction value of the sub-block where the pixel position is located can be directly selected from the prediction sample matrix including 1 luminance prediction sample matrix (2 chrominance prediction sample matrices), and this first prediction value is determined as the inter-frame prediction value of the pixel position, that is, the second prediction value.
[0562] Exemplarily, in the present application, if the prediction reference mode of the current block takes a value of 2, that is, the third reference mode 'PRED_List01' is used, then the mean operation can be first performed on the 2 luminance prediction sample matrices (2 groups of 4 chrominance prediction sample matrices) included in the prediction sample matrix to obtain 1 averaged luminance prediction sample (2 averaged chrominance prediction samples), and finally the first prediction value of the sub-block where the pixel position is located is selected from this averaged luminance prediction sample (2 averaged chrominance prediction samples), and this first prediction value is determined as the inter-frame prediction value of the pixel position, that is, the second prediction value.
[0563] Step 404: Determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for performing secondary prediction processing according to a preset shape.
[0564] In an embodiment of the present application, after the encoder determines the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block based on the first motion vector, it can further determine the filtering coefficient of the two-dimensional filter according to the motion vector deviation.
[0565] It should be noted that, in an embodiment of the present application, the filter coefficient of the two-dimensional filter is related to the motion vector deviation corresponding to the pixel position. That is to say, for different pixel positions, if the corresponding motion vector deviations are different, the filtering coefficients of the two-dimensional filters used are also different.
[0566] Furthermore, in an embodiment of the present application, when the encoder determines the filtering coefficient of the two-dimensional filter according to the motion vector deviation, it can first determine the proportional parameter, and then determine the filter coefficient corresponding to the pixel position according to the proportional parameter and the motion vector deviation.
[0567] It should be noted that, in an embodiment of the present application, the proportional parameter may include at least one proportional value, and the motion vector deviation includes a horizontal deviation and a vertical deviation; wherein, the at least one proportional value is a non-zero real number.
[0568] Specifically, in the present application, when the two-dimensional filter uses 9 adjacent pixel positions forming a rectangle for secondary prediction, the pixel position at the center of the rectangle is the position to be predicted, and the other 8 pixel positions are successively located at the upper left, upper, upper right, right, lower right, lower, lower left, and left adjacent to the position to be predicted. Calculate the 9 filter coefficients corresponding to the 9 adjacent pixel positions.
[0569] It should be noted that, in the present application, the pre-designed calculation rules may include various different calculation methods, such as addition operation, subtraction operation, multiplication operation, etc. Among them, for different pixel positions, different calculation methods can be used to calculate the filter coefficients.
[0570] It can be understood that, in the present application, among the multiple filter coefficients corresponding to multiple pixel positions calculated by the encoder according to different calculation methods in the pre-designed calculation rules, some filter coefficients can be a first-order function of the motion vector deviation, that is, the two are linearly related, and can also be a second-order function or a higher-order function of the motion vector deviation, that is, the two are non-linearly related.
[0571] That is to say, in the present application, any one of the multiple filter coefficients corresponding to multiple adjacent pixel positions can be a first-order function, a second-order function or a higher-order function of the motion vector deviation.
[0572] Exemplarily, in the present application, it is assumed that the motion vector deviation of the pixel position is (dmv_x, dmv_y). Among them, if the coordinates of the target pixel position are (i, j), then dmv_x can be expressed as dMvX[i][j][0], which represents the deviation value of the motion vector in the horizontal component, and dmv_y can be expressed as dMvX[i][j][1], which represents the deviation value of the motion vector in the vertical component.
[0573] Among them, the proportional parameter is generally a decimal or a fraction. A possible situation is that the proportional parameters are all powers of 2, such as 1 / 2, 1 / 4, 1 / 8, etc. Here, both dmv_x and dmv_y are their actual magnitudes, that is, 1 in dmv_x and dmv_y represents the distance of 1 pixel, and dmv_x and dmv_y are decimals or fractions.
[0574] It should be noted that in the embodiments of the present application, compared with the existing 8-tap filter, the motion vectors of the integer pixel positions and sub-pixel positions corresponding to the currently common 8-tap filter are non-negative in both the horizontal and vertical directions, and their magnitudes are all between 0 pixels and 1 pixel, that is, dmv_x and dmv_y cannot be negative. However, in the present application, the motion vectors of the integer pixel positions and sub-pixel positions corresponding to the filter can be negative in both the horizontal and vertical directions, that is, dmv_x and dmv_y can be negative.
[0575] It should be noted that in the embodiments of the present application, the two-dimensional filter for secondary prediction is a filter composed of adjacent points forming a preset shape. The adjacent points forming the preset shape can include multiple points, for example, composed of 9 points.
[0576] It can be understood that in the embodiments of the present application, the preset shape can be a symmetric shape. For example, the preset shape can include a rectangle, a rhombus, or any other symmetric shape.
[0577] Exemplarily, in the present application, the two-dimensional filter is a rectangular filter. Specifically, the two-dimensional filter is a filter composed of 9 adjacent pixel positions forming a rectangle. Among the 9 pixel positions, the pixel position at the center is the pixel position of the pixel currently required for secondary prediction.
[0578] Step 405: Determine the second prediction value of the sub-block based on the filtering coefficient and the first prediction value, and determine the second prediction value as the inter-frame prediction value of the sub-block.
[0579] In an embodiment of the present application, after the encoder determines the filtering coefficients of the two-dimensional filter according to the motion vector deviation, it can determine the second prediction value of the sub-block based on the filtering coefficients and the first prediction value, so that it can achieve the secondary prediction of the sub-block by combining the motion vector deviation with the pixel position, and complete the correction of the first prediction value.
[0580] It can be understood that in an embodiment of the present application, the encoder uses the motion vector deviation corresponding to the pixel position to determine the filter coefficients, so that the first prediction value can be corrected by the two-dimensional filter according to the filter coefficients, and the corrected second prediction value of the sub-block can be obtained. It can be seen that the second prediction value is a corrected value based on the first prediction value.
[0581] Further, in an embodiment of the present application, when the encoder determines the second prediction value of the sub-block based on the filtering coefficients and the first prediction value, it can first perform a multiplication operation on the filter coefficients and the first prediction value of the sub-block to obtain a product result. After traversing all the pixel positions in the sub-block, it then performs an addition operation on the product results of all the pixel positions in the sub-block to obtain a summation result. Finally, it can perform a normalization process on the addition result, and finally the corrected second prediction value of the sub-block can be obtained.
[0582] It should be noted that in an embodiment of the present application, before performing the secondary prediction, generally, the first prediction value of the sub-block where the pixel position is located is used as the prediction value before correction of the pixel position. Therefore, when filtering through the two-dimensional filter, the filter coefficients can be multiplied by the prediction value corresponding to the pixel position, that is, the first prediction value, and the product results corresponding to each pixel position are accumulated and then normalized.
[0583] It can be understood that in the present application, the encoder can perform the normalization process in various ways. For example, the result of accumulating the product of the filter coefficients and the prediction values of the corresponding pixel positions can be shifted 4 + shift1 bits to the right. Or, the result of accumulating the product of the filter coefficients and the prediction values of the corresponding pixel positions can be added with (1 << (3 + shift1)), and then shifted 4 + shift1 bits to the right.
[0584] It can be seen that in the present application, after obtaining the motion vector deviation corresponding to the pixel position inside the sub-block, for each sub-block and each pixel position in each sub-block, according to the motion vector deviation, based on the first prediction value of the motion compensation of the sub-block, filtering can be performed using a two-dimensional filter to complete the secondary prediction of the sub-block and obtain a new second prediction value.
[0585] Furthermore, in the embodiments of the present application, the two-dimensional filter can be understood as performing secondary prediction using a plurality of adjacent pixel positions that form the preset shape. Among them, the preset shape can be a rectangle, a rhombus, or any symmetric shape.
[0586] Specifically, in the embodiments of the present application, when the two-dimensional filter performs secondary prediction using 9 adjacent pixel positions that form a rectangle, the prediction sample matrix of the current block and the set of motion vector deviations of the sub-blocks of the current block can be determined first; among them, the motion vector deviation matrix includes the motion vector deviations corresponding to all pixel positions; then, based on the 9 adjacent pixel positions that form a rectangle, using the prediction sample matrix and the set of motion vector deviations, the sample matrix after secondary prediction of the current block is determined.
[0587] Exemplarily, the magnitude of the motion vector can be restricted within a reasonable range, such as the positive and negative values of the motion vector used above do not exceed 1 pixel or 1 / 2 pixel or 1 / 4 pixel, etc. in the horizontal and vertical directions.
[0588] It can be understood that if the prediction reference mode of the current block is 'Pred_List01', then the encoder averages the multiple prediction sample matrices of each component to obtain the final prediction sample matrix of the component. For example, the two luminance prediction sample matrices are averaged to obtain a new luminance prediction sample matrix.
[0589] Furthermore, in the embodiments of the present application, after obtaining the prediction sample matrix of the current block, if the current block has no transform coefficients, then the prediction matrix is used as the coding result of the current block. If the current block still has transform coefficients, then the transform coefficients can be encoded first, and the residual matrix is obtained through inverse transformation and inverse quantization, and the residual matrix is added to the prediction matrix to obtain the coding result.
[0590] In summary, through the inter-frame prediction method proposed in steps 401 to 405, after the prediction based on sub-blocks, for the pixel positions where the motion vector has a deviation from the motion vector of the sub-block, point-based secondary prediction is performed on the basis of the prediction based on sub-blocks. The point-based secondary prediction uses the information of multiple points that form preset shapes such as rectangles and rhombuses. The point-based secondary prediction uses a two-dimensional filter. The two-dimensional filter is a filter formed by adjacent points that form a preset shape. Adjacent points that form a pre-
[0591] This embodiment provides an inter-frame prediction method. After sub-block based prediction, for pixel positions with a deviation between the motion vector and the sub-block's motion vector, a point-based secondary prediction can be performed on the basis of the first prediction value of the sub-block to obtain a second prediction value. Among them, the point-based secondary prediction using a two-dimensional filter can correct the first prediction value using the information of several points forming a preset shape, and finally obtain the corrected second prediction value. The inter-frame prediction method proposed in this application can be well applied to all scenarios, greatly improving the coding performance, and thus improving the encoding and decoding efficiency.
[0592] Based on the above embodiment, in another embodiment of the present application, Figure 26 Schematic diagram of the composition structure of the decoder Figure 1 , such as Figure 26 As shown, the decoder 300 proposed in the embodiment of the present application may include a parsing part 301 and a first determination part 302;
[0593] The parsing part 301 is configured to parse the code stream and obtain the prediction mode parameter of the current block;
[0594] The first determination part 302 is configured to determine the first motion vector of the sub-blocks of the current block when the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block; wherein, the current block includes multiple sub-blocks; and determine the first prediction value of the sub-blocks and the motion vector deviation between the pixel position and the sub-blocks based on the first motion vector; wherein, the pixel position is the position of the pixel points within the sub-block; and determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for secondary prediction processing according to a preset shape; and determine the second prediction value of the sub-blocks based on the filtering coefficients and the first prediction value, and determine the second prediction value as the inter-frame prediction value of the sub-blocks.
[0595] Figure 27 Schematic diagram of the composition structure of the decoder Figure 2 , such as Figure 27 As shown, the decoder 300 proposed in the embodiment of the present application may further include a first processor 303, a first memory 304 storing executable instructions of the first processor 303, a first communication interface 305, and a first bus 306 for connecting the first processor 303, the first memory 304, and the first communication interface 305.
[0596] Further, in the embodiment of the present application, the above-mentioned first processor 303 is used to parse the code stream and obtain the prediction mode parameter of the current block;
[0597] When the prediction mode parameter indicates that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, determine the first motion vector of the sub-blocks of the current block; wherein, the current block includes a plurality of sub-blocks; determine the first prediction value of the sub-blocks and the motion vector deviation between the pixel position and the sub-block based on the first motion vector; wherein, the pixel position is the position of the pixel points within the sub-block; determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for secondary prediction processing according to a preset shape; determine the second prediction value of the sub-block based on the filtering coefficients and the first prediction value, and determine the second prediction value as the inter-frame prediction value of the sub-block.
[0598] If the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0599] The embodiment of the present application provides a decoder. After prediction based on sub-blocks, the decoder can perform point-based secondary prediction on pixel positions with a deviation between the motion vector and the motion vector of the sub-blocks on the basis of the first prediction value of the sub-blocks to obtain a second prediction value. Among them, the point-based secondary prediction using a two-dimensional filter can use the information of several points constituting a preset shape to correct the first prediction value, and finally obtain the corrected second prediction value. The inter-frame prediction method proposed in the present application can be well applicable to all scenarios, greatly improving the coding performance, and thus improving the encoding and decoding efficiency.
[0600] Figure 28 Schematic diagram of the composition structure of the encoder Figure 1 , such as Figure 28 shown, the encoder 400 proposed in the embodiment of the present application may include a second determination part 401;
[0601] The second determination part 401 is configured to determine the prediction mode parameter of the current block; and when the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, determine the first motion vector of the sub-blocks of the current block; wherein, the current block includes a plurality of sub-blocks; and determine the first prediction value of the sub-blocks and the motion vector deviation between the pixel position and the sub-blocks based on the first motion vector; wherein, the pixel position is the position of the pixel points within the sub-blocks; and determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used to perform quadratic prediction according to a preset shape; and determine the second prediction value of the sub-blocks based on the filtering coefficients and the first prediction value, and determine the second prediction value as the inter-frame prediction value of the sub-blocks.
[0602] Figure 29 Schematic diagram of the composition structure of the encoder Figure 2 , such as Figure 29 As shown, the encoder 400 proposed in the embodiment of the present application may further include a second processor 402, a second memory 403 storing executable instructions of the second processor 402, a second communication interface 404, and a second bus 405 for connecting the second processor 402, the second memory 403, and the second communication interface 404.
[0603] Further, in the embodiment of the present application, the above-mentioned second processor 402 is used to determine the prediction mode parameter of the current block; when the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, determine the first motion vector of the sub-blocks of the current block; wherein, the current block includes a plurality of sub-blocks; determine the first prediction value of the sub-blocks and the motion vector deviation between the pixel position and the sub-blocks based on the first motion vector; wherein, the pixel position is the position of the pixel points within the sub-blocks; determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used to perform quadratic prediction according to a preset shape; determine the second prediction value of the sub-blocks based on the filtering coefficients and the first prediction value, and determine the second prediction value as the inter-frame prediction value of the sub-blocks.
[0604] When the integrated unit is implemented in the form of a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0605] The embodiment of the present application provides an encoder. After sub-block-based prediction, for pixel positions where the motion vector deviates from the motion vector of the sub-block, a point-based secondary prediction can be performed on the basis of the first prediction value of the sub-block to obtain a second prediction value. Among them, the point-based secondary prediction using a two-dimensional filter can use the information of several points forming a preset shape to correct the first prediction value, and finally obtain the corrected second prediction value. The inter-frame prediction method proposed in the present application can be well applied to all scenarios, greatly improving the coding performance, and thus improving the encoding and decoding efficiency.
[0606] The embodiment of the present application provides a computer-readable storage medium and a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the method as described in the above embodiment.
[0607] Specifically, the program instructions corresponding to an inter-frame prediction method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the program instructions corresponding to an inter-frame prediction method in the storage medium are read or executed by an electronic device, the following steps are included:
[0608] Parse the bitstream to obtain the prediction mode parameter of the current block;
[0609] When the prediction mode parameter indicates using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, determine the first motion vector of the sub-blocks of the current block; wherein, the current block includes multiple sub-blocks;
[0610] Based on the first motion vector, determine the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block; wherein, the pixel position is the position of the pixel point within the sub-block;
[0611] Determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for performing secondary prediction processing according to a preset shape;
[0612] Based on the filtering coefficients and the first prediction value, determine the second prediction value of the sub-block, and determine the second prediction value as the inter-frame prediction value of the sub-block.
[0613] Specifically, the program instructions corresponding to an inter-frame prediction method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the program instructions corresponding to an inter-frame prediction method in the storage media are read or executed by an electronic device, the following steps are included:
[0614] Determine the prediction mode parameters of the current block;
[0615] When the prediction mode parameters indicate using the inter-frame prediction mode to determine the inter-frame prediction value of the current block, determine the first motion vector of the sub-blocks of the current block; wherein, the current block includes multiple sub-blocks;
[0616] Based on the first motion vector, determine the first prediction value of the sub-block and the motion vector deviation between the pixel position and the sub-block; wherein, the pixel position is the position of the pixel points within the sub-block;
[0617] Determine the filtering coefficients of the two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for performing secondary prediction according to a preset shape;
[0618] Based on the filtering coefficients and the first prediction value, determine the second prediction value of the sub-block, and determine the second prediction value as the inter-frame prediction value of the sub-block.
[0619] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program codes.
[0620] This application is described with reference to the schematic implementation flow diagrams and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the schematic implementation flow diagrams and / or block diagrams, as well as the combination of processes and / or blocks in the schematic implementation flow diagrams and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the processes or blocks in the schematic implementation flow Figure 1 in one or more of the processes Figure 1 or blocks.
[0621] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more of the processes or blocks in the schematic implementation flow Figure 1 in one or more of the processes Figure 1 or blocks.
[0622] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the processes or blocks in the schematic implementation flow Figure 1 in one or more of the processes Figure 1 or blocks.
[0623] The methods disclosed in several method embodiments provided by this application can be combined arbitrarily without conflict to obtain new method embodiments.
[0624] The features disclosed in several product embodiments provided by this application can be combined arbitrarily without conflict to obtain new product embodiments.
[0625] The features disclosed in several method or device embodiments provided by this application can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.
[0626] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0627] Industrial applicability
[0628] A method for inter-frame prediction, an encoder, a decoder, and a computer storage medium provided by an embodiment of the present application. The decoder parses a bitstream to obtain a prediction mode parameter of a current block; when the prediction mode parameter indicates that an inter-frame prediction mode is used to determine an inter-frame prediction value of the current block, a first motion vector of a sub-block of the current block is determined; wherein, the current block includes a plurality of sub-blocks; a first prediction value of the sub-block and a motion vector deviation between a pixel position and the sub-block are determined based on the first motion vector; wherein, the pixel position is the position of a pixel point within the sub-block; a filtering coefficient of a two-dimensional filter is determined according to the motion vector deviation; wherein, the two-dimensional filter is used for performing a secondary prediction process according to a preset shape; a second prediction value of the sub-block is determined based on the filtering coefficient and the first prediction value, and the second prediction value is determined as the inter-frame prediction value of the sub-block. The encoder determines a prediction mode parameter of the current block; when the prediction mode parameter indicates that an inter-frame prediction mode is used to determine an inter-frame prediction value of the current block, a first motion vector of a sub-block of the current block is determined; wherein, the current block includes a plurality of sub-blocks; a first prediction value of the sub-block and a motion vector deviation between a pixel position and the sub-block are determined based on the first motion vector; wherein, the pixel position is the position of a pixel point within the sub-block; a filtering coefficient of a two-dimensional filter is determined according to the motion vector deviation; wherein, the two-dimensional filter is used for using a secondary prediction according to a preset shape; a second prediction value of the sub-block is determined based on the filtering coefficient and the first prediction value, and the second prediction value is determined as the inter-frame prediction value of the sub-block. That is to say, the inter-frame prediction method proposed in the present application can perform point-based secondary prediction on a pixel position with a deviation between a motion vector and the motion vector of a sub-block after prediction based on the sub-block, and obtain a second prediction value based on the first prediction value of the sub-block. Among them, the point-based secondary prediction using the two-dimensional filter can correct the first prediction value using the information of several points constituting a preset shape, and finally obtain a corrected second prediction value. The inter-frame prediction method proposed in the present application can be well applied to all scenarios, greatly improving the coding performance, and thus improving the encoding and decoding efficiency.
Claims
1. An inter-frame prediction method, characterized in that, applied to a decoder, the method comprises: parsing a bitstream to obtain a prediction mode parameter of a current block; when the prediction mode parameter indicates using an inter-frame prediction mode to determine an inter-frame prediction value of the current block, determining a first motion vector of a sub-block of the current block; wherein, the current block comprises a plurality of sub-blocks; determining a first prediction value of the sub-block based on the first motion vector, and a motion vector deviation between a pixel position of the sub-block and the sub-block; determining filter coefficients of a two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for performing a secondary prediction process according to a preset shape; determining a second prediction value of the sub-block based on the filter coefficients and the first prediction value, and determining the second prediction value as the inter-frame prediction value of the sub-block.
2. The method according to claim 1, characterized in that, the determining the first motion vector of the sub-block of the current block comprises: parsing the bitstream to obtain an affine mode parameter and a prediction reference mode of the current block; when the affine mode parameter indicates using an affine mode, determining a control point mode and a sub-block size parameter; wherein, the control point mode is used for determining the number of control points; determining the first motion vector according to the prediction reference mode, the control point mode and the sub-block size parameter.
3. The method according to claim 2, characterized in that, the determining the first motion vector according to the prediction reference mode, the control point mode and the sub-block size parameter comprises: determining a control point motion vector group according to the prediction reference mode; determining the first motion vector according to the control point motion vector group, the control point mode and the sub-block size parameter.
4. The method according to claim 3, characterized in that, the method further comprises: traversing each sub-block of the current block, and constructing a motion vector set according to the first motion vector of each sub-block.
5. The method according to claim 3, characterized in that, the determining the first motion vector according to the control point motion vector group, the control point mode and the sub-block size parameter comprises: determining a difference variable according to the control point motion vector group, the control point mode and a size parameter of the current block; determining a sub-block position based on the prediction mode parameter and the sub-block size parameter; using the difference variable and the sub-block position to determine the first motion vector of the sub-block.
6. The method according to claim 4, characterized in that, the determining the first prediction value of the sub-block based on the first motion vector, and the motion vector deviation between the pixel position of the sub-block and the sub-block comprises: determining a sample matrix; wherein, the sample matrix comprises a luminance sample matrix and a chrominance sample matrix; determining the first prediction value according to the prediction reference mode, the sub-block size parameter, the sample matrix and the motion vector set.
7. The method according to claim 2, characterized in that, Determining the first prediction value of the sub-block and the motion vector deviation between the pixel position of the sub-block and the sub-block based on the first motion vector includes: Parsing the bitstream to obtain secondary prediction parameters; When the secondary prediction parameters indicate the use of secondary prediction, determining the motion vector deviation between the sub-block and each pixel position based on the difference variable; wherein, the difference variable is determined according to the control point motion vector group, the control point mode, and the size parameters of the current block; the control point motion vector group is determined according to the prediction reference mode of the current block.
8. The method according to claim 2, wherein, the method further includes: Parsing the bitstream to obtain secondary prediction parameters; When the secondary prediction parameters indicate the non-use of secondary prediction, determining the first prediction value in the prediction sample matrix as the second prediction value.
9. The method according to claim 1, wherein, the two-dimensional filter is used to perform secondary prediction using a plurality of adjacent pixel positions that form the preset shape.
10. The method according to claim 9, wherein, the preset shape is a rectangle, a rhombus, or any symmetric shape.
11. The method according to claim 10, wherein, determining the filter coefficients of the two-dimensional filter according to the motion vector deviation includes: Obtaining a scale parameter; Determining the filter coefficients corresponding to the pixel positions according to the scale parameter and the motion vector deviation.
12. The method according to claim 11, wherein, the scale parameter includes at least one scale value, and the motion vector deviation includes a horizontal deviation and a vertical deviation; wherein, the at least one scale value is a non-zero real number.
13. The method according to claim 12, wherein, When the two-dimensional filter performs secondary prediction using 9 adjacent pixel positions that form a rectangle, the pixel position at the center of the rectangle is the position to be predicted, and the other 8 pixel positions are successively located at the adjacent positions of the upper left, upper, upper right, right, lower right, lower, lower left, and left of the position to be predicted.
14. The method according to claim 13, wherein, determining the filter coefficients corresponding to the pixel positions according to the scale parameter and the motion vector deviation includes: Calculating 9 filter coefficients corresponding to the 9 adjacent pixel positions based on the at least one scale value and the motion vector deviation of the position to be predicted according to a preset calculation rule.
15. The method according to claim 14, wherein, Setting the filter coefficient of the position to be predicted to 1.
16. The method according to claim 13, wherein, determining the filter coefficients corresponding to the pixel positions according to the scale parameter and the motion vector deviation includes: Performing a shift process on the motion vector deviation to obtain a shifted motion vector deviation; Calculating the filter coefficients based on the at least one scale value and the shifted motion vector deviation of the position to be predicted according to a preset calculation rule.
17. The method according to claim 9, wherein, any one of the plurality of filter coefficients corresponding to the plurality of adjacent pixel positions is a linear function, quadratic function or higher-order function of the motion vector deviation.
18. The method according to claim 1, wherein, determining the second prediction value of the sub-block based on the filter coefficient and the first prediction value includes: performing a multiplication operation on the filter coefficient and the first prediction value to obtain a product result corresponding to the pixel position; performing an addition operation on the product results of all pixel positions of the sub-block to obtain a summation result; performing a normalization process on the summation result to obtain the second prediction value.
19. An inter-frame prediction method, wherein, applied to an encoder, the method includes: determining a prediction mode parameter of a current block; when the prediction mode parameter indicates using an inter-frame prediction mode to determine an inter-frame prediction value of the current block, determining a first motion vector of a sub-block of the current block; wherein, the current block includes a plurality of sub-blocks; determining a first prediction value of the sub-block and a motion vector deviation between the pixel position of the sub-block and the sub-block based on the first motion vector; determining filter coefficients of a two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used for quadratic prediction according to a preset shape; determining a second prediction value of the sub-block based on the filter coefficients and the first prediction value, and determining the second prediction value as the inter-frame prediction value of the sub-block.
20. The method according to claim 19, wherein, determining the first motion vector of the sub-block of the current block includes: determining an affine mode parameter and a prediction reference mode of the current block; when the affine mode parameter indicates using an affine mode, determining a control point mode and a sub-block size parameter; wherein, the control point mode is used to determine the number of control points; determining the first motion vector according to the prediction reference mode, the control point mode and the sub-block size parameter.
21. The method according to claim 20, wherein, determining the first motion vector according to the prediction reference mode, the control point mode and the sub-block size parameter includes: determining a control point motion vector group according to the prediction reference mode; determining the first motion vector according to the control point motion vector group, the control point mode and the sub-block size parameter.
22. The method according to claim 21, wherein, the method further includes: traversing each sub-block of the current block, and constructing a motion vector set according to the first motion vector of each sub-block.
23. The method according to claim 21, wherein, determining the first motion vector according to the control point motion vector group, the control point mode and the sub-block size parameter includes: determining a difference variable according to the control point motion vector group, the control point mode and the size parameter of the current block; Determine the sub - block position based on the predicted pattern parameter and the sub - block size parameter; Determine the first motion vector of the sub - block by using the difference variable and the sub - block position.
24. The method according to claim 22, wherein, The determining the first prediction value of the sub - block based on the first motion vector, and the motion vector deviation between the pixel position of the sub - block and the sub - block, includes: Determine a sample matrix; wherein, the sample matrix includes a luminance sample matrix and a chrominance sample matrix; Determine the first prediction value according to the prediction reference pattern, the sub - block size parameter, the sample matrix, and the set of motion vectors.
25. The method according to claim 20, wherein, The determining the first prediction value of the sub - block based on the first motion vector, and the motion vector deviation between the pixel position of the sub - block and the sub - block, includes: Determine a quadratic prediction parameter; When the quadratic prediction parameter indicates using quadratic prediction, determine the motion vector deviation between the sub - block and each pixel position based on the difference variable.
26. The method according to claim 20, wherein, The method further includes: Determine a quadratic prediction parameter; When the quadratic prediction parameter indicates not using quadratic prediction, determine the first prediction value in the prediction sample matrix as the second prediction value.
27. The method according to claim 19, wherein, The two - dimensional filter is used to perform quadratic prediction by using a plurality of adjacent pixel positions that form the preset shape.
28. The method according to claim 27, wherein, The preset shape is a rectangle, a rhombus, or any symmetric shape.
29. The method according to claim 28, wherein, The determining the filter coefficients of the two - dimensional filter according to the motion vector deviation includes: Determine a scaling parameter; Determine the filter coefficient corresponding to the pixel position according to the scaling parameter and the motion vector deviation.
30. The method according to claim 29, wherein, The scaling parameter includes at least one scaling value, and the motion vector deviation includes a horizontal deviation and a vertical deviation; wherein, the at least one scaling value is a non - zero real number.
31. The method according to claim 30, wherein, When the two - dimensional filter performs quadratic prediction by using 9 adjacent pixel positions that form a rectangle, the pixel position at the center of the rectangle is the position to be predicted, and the other 8 pixel positions are successively located at the adjacent positions of the position to be predicted at the upper - left, upper, upper - right, right, lower - right, lower, lower - left, and left.
32. The method according to claim 31, wherein, The determining the filter coefficient corresponding to the pixel position according to the scaling parameter and the motion vector deviation includes: Calculate 9 filter coefficients corresponding to the 9 adjacent pixel positions based on the at least one scaling value and the motion vector deviation of the position to be predicted according to a pre - designed calculation rule.
33. The method according to claim 32, wherein, Set the filter coefficient of the to-be-predicted position to 1.
34. The method according to claim 31, wherein, determining the filter coefficient corresponding to the pixel position according to the ratio parameter and the motion vector deviation includes: performing a shift process on the motion vector deviation to obtain a shifted motion vector deviation; calculating the filter coefficient based on the at least one ratio value and the shifted motion vector deviation of the to-be-predicted position according to a pre-designed calculation rule.
35. The method according to claim 27, wherein, any one of the filter coefficients corresponding to the multiple adjacent pixel positions is a linear function, quadratic function or higher-order function of the motion vector deviation.
36. The method according to claim 19, wherein, determining the second prediction value of the sub-block based on the filter coefficient and the first prediction value includes: performing a multiplication operation on the filter coefficient and the first prediction value to obtain a product result corresponding to the pixel position; performing an addition operation on the product results of all pixel positions of the sub-block to obtain a summation result; performing a normalization process on the summation result to obtain the second prediction value.
37. A decoder, wherein, the decoder includes a parsing part and a first determination part; the parsing part is configured to parse a bitstream to obtain a prediction mode parameter of a current block; the first determination part is configured to, when the prediction mode parameter indicates using an inter prediction mode to determine an inter prediction value of the current block, determine a first motion vector of a sub-block of the current block; wherein, the current block includes multiple sub-blocks; and determine a first prediction value of the sub-block, a motion vector deviation between the pixel position of the sub-block and the sub-block based on the first motion vector; and determine a filter coefficient of a two-dimensional filter according to the motion vector deviation; wherein, the two-dimensional filter is used to perform a secondary prediction process according to a preset shape; and determine a second prediction value of the sub-block based on the filter coefficient and the first prediction value, and determine the second prediction value as the inter prediction value of the sub-block.
38. A decoder, the decoder includes a first processor and a first memory storing executable instructions of the first processor. When the instructions are executed, the first processor implements the method according to any one of claims 1-18.
39. An encoder, wherein, the encoder includes a second determination part; The second determination part is configured to determine prediction mode parameters of a current block; and when the prediction mode parameters indicate that an inter-frame prediction mode is used to determine an inter-frame prediction value of the current block, determine a first motion vector of a sub-block of the current block; wherein the current block includes a plurality of sub-blocks; and determine a first prediction value of the sub-block, and a motion vector deviation between a pixel position of the sub-block and the sub-block based on the first motion vector; and determine filter coefficients of a two-dimensional filter according to the motion vector deviation; wherein the two-dimensional filter is used for quadratic prediction according to a preset shape; and determine a second prediction value of the sub-block based on the filter coefficients and the first prediction value, and determine the second prediction value as the inter-frame prediction value of the sub-block.
40. An encoder, the encoder includes a second processor and a second memory storing executable instructions of the second processor. When the instructions are executed, the second processor implements the method according to any one of claims 19-36.
41. A computer storage medium, the computer storage medium stores a computer program, and when the computer program is executed by a first processor, the method according to any one of claims 1-18 is implemented.
42. A computer storage medium, the computer storage medium stores a computer program, and when the computer program is executed by a first processor, the method according to any one of claims 19-36 is implemented.
Citation Information
Patent Citations
Affine prediction method and related device thereof
CN111050168A