Inter-frame prediction method, encoder, decoder and computer storage medium

By combining sub-block-based prediction and point-based prediction in video encoding and decoding technology and performing weighted averaging, the problem of inaccurate affine prediction and increased complexity in the prior art is solved, and more efficient inter-frame prediction is achieved.

CN114982228BActive Publication Date: 2025-05-13GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080093968.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-16
Publication Date
2025-05-13
Estimated Expiration
2040-10-16

AI Technical Summary

Technical Problem

In the existing video encoding and decoding technology, affine prediction based on sub-blocks is not accurate enough, resulting in the need to improve through interleaving prediction, but interleaving prediction increases complexity and bandwidth, reducing encoding and decoding efficiency.

Method used

An inter prediction method is proposed, by using a combination of sub-block-based prediction and point-based prediction in the current block, a set of predicted values ​​is obtained, and the final predicted values ​​are determined through a weighted average, thereby improving the accuracy of inter prediction and reducing the computational complexity and bandwidth.

Benefits of technology

While improving inter prediction accuracy, it reduces the computational complexity and bandwidth, significantly improving encoding performance and encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114982228B_ABST
    Figure CN114982228B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses an inter-frame prediction method, an encoder, a decoder and a computer storage medium. The decoder parses a code stream to obtain prediction mode parameters of a current block; when the prediction mode parameters indicate that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, the first prediction value of the pixel in each sub-block is determined using the first motion vector of each sub-block of the current block, and the second prediction value of each pixel in the current block is determined using the second motion vector of each pixel; for a pixel in the current block, a first weight and a second weight of the pixel are determined; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; based on the first weight, the second weight, the first prediction value and the second prediction value, the prediction value of the pixel is determined; according to the prediction value of the pixel, a third prediction value of the current block is determined; wherein the third prediction value is used to determine the reconstruction value of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video coding and decoding, and in particular to an inter-frame prediction method, an encoder, a decoder and a computer storage medium. Background Art

[0002] In the field of video coding and decoding, in order to balance performance and cost, in general, affine prediction in Versatile Video Coding (VVC) and Audio Videocoding Standard Workgroup of China (AVS) is implemented based on sub-blocks.

[0003] Sub-block based affine prediction is not accurate enough, so it is necessary to improve sub-block based affine prediction through interweaved prediction and other technologies. Although the new sub-block division method introduced by interweaved prediction can improve the prediction accuracy to a certain extent, interweaved prediction also brings obvious complexity and reduces the encoding and decoding efficiency. Summary of the invention

[0004] The present application proposes an inter-frame prediction method, an encoder, a decoder and a computer storage medium, which can greatly improve the encoding performance, thereby improving the encoding and decoding efficiency.

[0005] The technical solution of this application is implemented as follows:

[0006] In a first aspect, an embodiment of the present application provides an inter-frame prediction method, which is applied to a decoder. The method includes:

[0007] Parse the bitstream and obtain the prediction mode parameters of the current block;

[0008] When the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, a first prediction value of a pixel in each sub-block is determined by using a first motion vector of each sub-block of the current block, and a second prediction value of each pixel in the current block is determined by using a second motion vector of each pixel; wherein the current block includes one or more sub-blocks; and the current block includes one or more pixels;

[0009] For a pixel point in the current block, determine a first weight and a second weight of the pixel point; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value;

[0010] Determine a predicted value of the pixel based on the first weight, the second weight, the first predicted value, and the second predicted value;

[0011] A third prediction value of the current block is determined according to the prediction value of the pixel point; wherein the third prediction value is used to determine a reconstruction value of the current block.

[0012] In a second aspect, an embodiment of the present application provides an inter-frame prediction method, which is applied to an encoder, and the method includes:

[0013] Determining prediction mode parameters for the current block;

[0014] When the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, a first prediction value of a pixel in each sub-block is determined by using a first motion vector of each sub-block of the current block, and a second prediction value of each pixel in the current block is determined by using a second motion vector of each pixel; wherein the current block includes one or more sub-blocks; and the current block includes one or more pixels;

[0015] For a pixel point in the current block, determine a first weight and a second weight of the pixel point; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value;

[0016] Determine a predicted value of the pixel based on the first weight, the second weight, the first predicted value, and the second predicted value;

[0017] Determine a third prediction value of the current block according to the prediction value of the pixel point; wherein the third prediction value is used to determine the residual of the current block.

[0018] In a third aspect, an embodiment of the present application provides a decoder, the decoder comprising a parsing part, a first determining part;

[0019] The parsing part is configured to parse the bitstream and obtain the prediction mode parameters of the current block;

[0020] The first determination part is configured to, when the prediction mode parameter indicates that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, determine the first prediction value of the pixel in each sub-block using the first motion vector of each sub-block of the current block, and determine the second prediction value of each pixel in the current block using the second motion vector of each pixel; wherein the current block includes one or more sub-blocks; the current block includes one or more pixels; for a pixel in the current block, determine the first weight and the second weight of the pixel; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; determine the prediction value of the pixel based on the first weight, the second weight, the first prediction value and the second prediction value; determine the third prediction value of the current block based on the prediction value of the pixel; wherein the third prediction value is used to determine the reconstruction value of the current block.

[0021] In a fourth aspect, an embodiment of the present application provides a decoder, comprising a first processor and a first memory storing instructions executable by the first processor, wherein when the instructions are executed, the first processor implements the inter-frame prediction method as described above.

[0022] In a fifth aspect, an embodiment of the present application provides an encoder, the encoder comprising a second determining part;

[0023] The second determination part is configured to determine the prediction mode parameters of the current block; when the prediction mode parameters indicate that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, the first prediction value of the pixel points in each sub-block is determined using the first motion vector of each sub-block of the current block, and the second prediction value of each pixel point of the current block is determined using the second motion vector of each pixel point of the current block; wherein the current block includes one or more sub-blocks; the current block includes one or more pixels; for a pixel point in the current block, a first weight and a second weight of the pixel point are determined; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; based on the first weight, the second weight, the first prediction value and the second prediction value, the prediction value of the pixel point is determined; according to the prediction value of the pixel point, a third prediction value of the current block is determined; wherein the third prediction value is used to determine the residual of the current block.

[0024] In a sixth aspect, an embodiment of the present application provides an encoder, comprising a second processor and a second memory storing instructions executable by the second processor, wherein when the instructions are executed, the second processor implements the inter-frame prediction method as described above.

[0025] In the seventh aspect, an embodiment of the present application provides a computer storage medium, which stores a computer program. When the computer program is executed by a first processor, it implements the inter-frame prediction method as described in the first aspect. When the computer program is executed by a second processor, it implements the inter-frame prediction method as described in the second aspect.

[0026] The embodiments of the present application provide an inter-frame prediction method, an encoder, a decoder and a computer storage medium. The decoder parses a bit stream to obtain prediction mode parameters of a current block. When the prediction mode parameters indicate that an inter-frame prediction value of the current block is determined using an inter-frame prediction mode, a first prediction value of a pixel in each sub-block is determined using a first motion vector of each sub-block of the current block, and a second prediction value of each pixel in the current block is determined using a second motion vector of each pixel. The current block includes one or more sub-blocks. The current block includes one or more pixels. For a pixel in the current block, a first weight and a second weight of the pixel are determined. The first weight corresponds to a first prediction value, and the second weight corresponds to a second prediction value. The prediction value of the pixel is determined based on the first weight, the second weight, the first prediction value and the second prediction value. The third prediction value of the current block is determined based on the prediction value of the pixel. The third prediction value is used to determine a reconstruction value of the current block. That is to say, the inter-frame prediction method proposed in the present application can use sub-block-based prediction for the current block to obtain a set of prediction values, namely, the first prediction value, and use point-based prediction for the current block to obtain another set of prediction values, namely, the second prediction value, and after determining the first weight of the sub-block-based prediction and the second weight of the point-based prediction for the same pixel point in the current block, use the first weight and the second weight to perform weighted averaging on the first prediction value and the second prediction value, and finally obtain a new prediction value for the current block, thereby improving the accuracy of the inter-frame prediction while reducing the complexity of the calculation and reducing the bandwidth, thereby greatly improving the encoding performance and improving the encoding and decoding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Schematic diagram of pixel interpolation;

[0028] Figure 2 Interpolation for sub-blocks Figure 1 ;

[0029] Figure 3 Interpolation for sub-blocks Figure 2 ;

[0030] Figure 4 A schematic diagram of interleaving predictions;

[0031] Figure 5 The weight of the prediction value based on the sub-block is shown as Figure 1 ;

[0032] Figure 6 The weight of the prediction value based on the sub-block is shown as Figure 2 ;

[0033] Figure 7 A schematic block diagram of a video encoding system provided in an embodiment of the present application;

[0034] Figure 8 A schematic block diagram of a video decoding system provided in an embodiment of the present application;

[0035] Fig. 9 Schematic diagram of the implementation process of the inter-frame prediction method Figure 1 ;

[0036] Fig.10 The first weight is shown Figure 1 ;

[0037] Fig.11 The first weight is shown Figure 2 ;

[0038] Fig.12 Schematic diagram of the implementation process of the inter-frame prediction method Figure 2 ;

[0039] Fig.13 It is a schematic diagram of one-way prediction;

[0040] Fig.14 Schematic diagram of the implementation process of the inter-frame prediction method Figure 3 ;

[0041] Fig.15 This is a schematic diagram of bidirectional prediction. Figure 1 ;

[0042] Fig.16 Schematic diagram of the implementation process of the inter-frame prediction method Figure 4 ;

[0043] Fig.17 This is a schematic diagram of bidirectional prediction. Figure 2 ;

[0044] Fig.18 This is an illustration of interpolation filtering. Figure 1 ;

[0045] Fig.19 This is an illustration of interpolation filtering. Figure 2 ;

[0046] Fig. 20 This is an illustration of interpolation filtering. Figure 3 ;

[0047] Fig.21 This is an illustration of interpolation filtering. Figure 4 ;

[0048] Fig. 22 This is an illustration of interpolation filtering. Figure 5 ;

[0049] Fig.23 This is an illustration of interpolation filtering. Figure 6 ;

[0050] Fig.24 Schematic diagram of the implementation process of the inter-frame prediction method Figure 5 ;

[0051] Fig.25 The structure of the decoder is shown in Figure 2. Figure 1 ;

[0052] Fig.26 The structure of the decoder is shown in Figure 2. Figure 2 ;

[0053] Fig. 27 The structure of the encoder is shown in Figure 2. Figure 1 ;

[0054] Fig.28 The structure of the encoder is shown in Figure 2. Figure 2 . DETAILED DESCRIPTION

[0055] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. It is understood that the specific embodiments described herein are only used to explain the related applications, rather than to limit the applications. It should also be noted that, for the convenience of description, only the parts related to the related applications are shown in the accompanying drawings.

[0056] At present, the common video coding standards are based on a block-based hybrid coding framework. Each frame in the video image is divided into square largest coding units (LCU) of the same size (such as 128×128, 64×64, etc.), and each largest coding unit can be divided into rectangular coding units (CU) according to rules; and the coding unit may be divided into smaller prediction units (PU). Specifically, the hybrid coding framework may include modules such as prediction, transform, quantization, entropy coding, and in loop filter; among them, the prediction module may include intra prediction and inter prediction, and inter prediction may include motion estimation and motion compensation. Since there is a strong correlation between adjacent pixels in a frame of a video image, the intra-frame prediction method used in video coding and decoding technology can eliminate the spatial redundancy between adjacent pixels; however, since there is also a strong similarity between adjacent frames in a video image, the inter-frame prediction method is used in video coding and decoding technology to eliminate the temporal redundancy between adjacent frames, thereby improving the coding efficiency. The following application will be described in detail with inter-frame prediction.

[0057] Inter-frame prediction is to use the frames that have been encoded / decoded to predict the part of the current frame that needs to be encoded / decoded. In the block-based coding and decoding framework, the part that needs to be encoded / decoded is usually a coding unit or a prediction unit. Here, the coding unit or prediction unit that needs to be encoded / decoded is collectively referred to as the current block. Translational motion is a common and simple motion mode in video, so the prediction of translation is also a traditional prediction method in video coding and decoding. Translational motion in video can be understood as a part of the content moving from a certain position on one frame to a certain position on another frame over time. A simple unidirectional prediction of translation can be represented by a motion vector (motion vector, MV) between a certain frame and the current frame. The certain frame mentioned here is a reference frame of the current frame. The current block can find a reference block on the reference frame with the same size as the current block through the motion information containing the reference frame and the motion vector, and use this reference block as the prediction block of the current block. In an ideal translational motion, the content of the current block does not have deformation, rotation, brightness, color, etc. between different frames. However, the content in the video does not always meet this ideal situation. Bidirectional prediction can solve the above problems to a certain extent. Usually, bidirectional prediction refers to bidirectional translation prediction. Bidirectional prediction is to use the motion information of two reference frames and motion vectors to find two reference blocks of the same size as the current block from two reference frames (the two reference frames may be the same reference frame), and use these two reference blocks to generate a prediction block of the current block. The generation methods include averaging, weighted averaging, and some other calculations.

[0058] In the present application, prediction can be considered as a part of motion compensation. Some documents refer to the prediction in the present application as motion compensation. For example, the affine prediction mentioned in the present application may be referred to as affine motion compensation in some documents.

[0059] Rotation, enlargement, reduction, distortion, deformation, etc. are also common changes in videos. However, ordinary translation prediction cannot handle such changes well, so affine prediction models are applied to video encoding and decoding, such as affine in VVC and AVS, where the affine prediction models of VVC and AVS3 are similar. In changes such as rotation, enlargement, reduction, distortion, deformation, etc., it can be considered that not all points of the current block use the same MV, so it is necessary to derive the MV of each point. The affine prediction model uses a small number of parameters to derive the MV of each point through calculation. The affine prediction models of VVC and AVS3 both use 2 control point (4 parameters) and 3 control point (6 parameters) models. 2 control points are 2 control points at the upper left and upper right corners of the current block, and 3 control points are 3 control points at the upper left, upper right and lower left corners of the current block. In general, since each MV includes an x ​​component and a y component, 2 control points have 4 parameters and 3 control points have 6 parameters.

[0060] According to the affine prediction model, an MV can be derived for each pixel. Each pixel can find its corresponding position in the reference frame. If this position is not an integer pixel, then the value of this sub-pixel point needs to be obtained through interpolation. The interpolation method used in the current video coding standard is usually implemented by a finite impulse response (FIR) filter, and the complexity (cost) of implementing this method is very high. For example, in AVS3, an 8-tap interpolation filter is used for the brightness component, and the sub-pixel accuracy of the normal mode is 1 / 4 pixel, and the sub-pixel accuracy of the affine mode is 1 / 16 pixel. For each sub-pixel point that meets the 1 / 16 pixel accuracy, 8 integer pixels in the horizontal direction and 8 integer pixels in the vertical direction, that is, 64 integer pixels, need to be interpolated. Figure 1 is a schematic diagram of pixel interpolation, such as Figure 1 As shown, the circular pixel is the desired sub-pixel point, the dark square pixel is the position of the whole pixel corresponding to the sub-pixel, the vector between the two is the motion vector of the sub-pixel, and the light square pixel is the pixel needed for the interpolation of the circular sub-pixel point. To obtain the value of the sub-pixel point, the pixel values ​​of these 8x8 light square pixel areas need to be interpolated, which also includes the dark pixels.

[0061] In the traditional translation prediction, the MV of each pixel in the current block is the same. If the concept of sub-block is further introduced, the size of the sub-block is 4x4, 8x8, etc. Figure 2 Interpolation for sub-blocks Figure 1 , the pixel area required for 4x4 block interpolation is as follows Figure 2 shown. Figure 3 Interpolation for sub-blocks Figure 2 , the pixel area required for 8x8 block interpolation is as follows Figure 3 shown.

[0062] If the MV of each pixel in a sub-block is the same, then the pixels in a sub-block can be interpolated together, thereby sharing the bandwidth, using filters with the same phase, and sharing the intermediate values ​​of the interpolation process. However, if each pixel uses an MV, the bandwidth will increase, and filters with different phases may be used, and the intermediate values ​​of the interpolation process cannot be shared.

[0063] Point-based affine prediction is very expensive. Therefore, in order to balance performance and cost, affine prediction in VVC and AVS3 is implemented based on sub-blocks. The sub-block size in AVS3 is 4x4 and 8x8, and the sub-block size of 4x4 is used in VVC. Each sub-block has an MV, and the pixels inside the sub-block share the same MV. Therefore, all pixels inside the sub-block are uniformly interpolated. Through the above method, the motion compensation complexity of sub-block-based affine prediction is similar to that of other sub-block-based prediction methods.

[0064] It can be seen that in the affine prediction method based on sub-blocks, the pixels inside the sub-block share the same MV, and the method to determine this shared MV is to take the MV of the center of the current sub-block. For sub-blocks such as 4x4 and 8x8 that have at least one even number of pixels in the horizontal and vertical directions, their centers actually fall on non-integer pixels. The current standards all take an integer pixel position. For example, for a 4x4 sub-block, take the pixel point with a distance of (2, 2) from the upper left corner. For an 8x8 sub-block, take the pixel point with a distance of (4, 4) from the upper left corner.

[0065] The affine prediction model can derive the MV of each pixel based on the control points (2 control points or 3 control points) used by the current block. In the sub-block-based affine prediction, the MV of this position is calculated based on the pixel points in the previous segment as the MV of the sub-block. In order to derive the motion vector of each sub-block, the motion vector of the center sample of each sub-block can be rounded to 1 / 16 precision, and then motion compensation is performed.

[0066] With the development of technology, affine prediction can also use bidirectional prediction, which can be understood as using two sets of affine model parameters for the current block. These two sets of affine model parameters may come from different reference frames or from the same reference frame. Affine prediction is performed on the two sets of affine model parameters respectively, and then the results of the two sets of affine prediction are averaged or weighted averaged to finally obtain the prediction block of the current block.

[0067] The affine prediction based on sub-blocks is not accurate enough, so some technologies have been proposed to further improve the affine prediction based on sub-blocks, one of which is called interlaced prediction. Interlaced prediction adds a new sub-block division method, which is different from the original division method. Figure 4 is a schematic diagram of interleaved prediction, such as Figure 4As shown, Pattern 0 is the existing division method, and Pattern 1 is a newly added division method for interlaced prediction. Compared with Pattern 0, the division of all sub-blocks in Pattern 1 is offset by half the length of the sub-block in the horizontal and vertical directions respectively. This is a typical division method, and Pattern 0 and Pattern 1 can also be divided using other methods. Similar to Pattern 0, each sub-block of Pattern 1 can also use an affine model to calculate an MV, and then use this MV for sub-block-based prediction. Therefore, two prediction values ​​can be obtained for each point of the current block, one from Pattern 0 and the other from Pattern 1, thereby forming the prediction block P0 and prediction block P1 corresponding to the current block, and then obtaining the new prediction block P corresponding to the current block based on P0 and P1.

[0068] In general, the closer the distance to the sub-block reference position (determined as the position of the sub-block MV), the higher the prediction accuracy, and the farther the distance to the sub-block reference position, the lower the prediction accuracy. In the current standard, the pixel position at the upper left corner position (2, 2) of the 4x4 sub-block is taken as the sub-block reference position. The pixel position at the upper left corner position (4, 4) of the 8x8 sub-block is taken as the sub-block reference position. Therefore, the weights of the prediction values ​​corresponding to different pixel positions in the sub-block may also be different. Figure 5 The weight of the prediction value based on the sub-block is shown as Figure 1 , Figure 6 The weight of the prediction value based on the sub-block is shown as Figure 2 ,like Figure 5 and 6 As shown, for different pixel positions, the weights of the prediction values ​​based on the sub-blocks are also different. For example, the weights of the prediction values ​​of the pixel positions at positions (2, 2), (2, 3), (3, 2), and (3, 3) in the 4x4 sub-block are relatively large, taking a value of 3, while the weights of the prediction values ​​of other pixel positions are relatively small, taking a value of 1. Correspondingly, the weights of the prediction values ​​of the 16 pixel positions in the middle area of ​​the 8x8 sub-block are relatively large, taking a value of 3, while the weights of the prediction values ​​of other pixel positions are relatively small, taking a value of 1. Further, the prediction values ​​at each pixel position in P0 and P1 can be weighted averaged according to their corresponding weight values. Specifically, if the weights of a pixel position in P0 and P1 are equal, that is, the weight values ​​in both P0 and P1 are 1, or the weight values ​​in both P0 and P1 are 3, then the average is 1:1. Otherwise, the weighted average is performed according to the corresponding weights, such as 3:1 or 1:3.

[0069] It can be understood that the interleaved prediction method can be regarded as an improvement on each unidirectional prediction in the unidirectional prediction or bidirectional prediction. If the interleaved prediction is applied to the bidirectional prediction, then there will be an interleaved prediction for each unidirectional prediction. The prediction value of the original unidirectional prediction and the prediction value of the interleaved prediction are weighted averaged to obtain the prediction value of the new unidirectional prediction, and the prediction values ​​of the two new unidirectional predictions are averaged (or weighted averaged) to obtain the final prediction value.

[0070] Interleaved prediction uses different division methods so that the reference positions of the sub-blocks of the two division methods are also interleaved, so that some points with low prediction accuracy in one division method have higher prediction accuracy in another division method, and reduce the error through weighted averaging, thereby improving compression performance.

[0071] However, the interleaved prediction technology also has certain drawbacks. Specifically, the interleaved prediction uses two different sub-block division methods, and each sub-block division method requires sub-block-based prediction separately. It can be seen that compared with the original prediction method, the complexity of the interleaved prediction method is more than doubled just based on the sub-block prediction.

[0072] On the one hand, the amount of calculation increases. This is because the sub-block division method of the original prediction method is unified, while the interleaved prediction will divide into smaller sub-blocks. When the number of interpolation results is the same, and when both use horizontally and vertically separable filters with the same number of filter taps, smaller sub-blocks mean more calculations. Take a 4x4 sub-block and 4 2x2 sub-blocks, all using 8-tap horizontally and vertically separable filters as an example. A 4x4 sub-block needs to interpolate 4x(4+7) points in the horizontal direction and 4x4 points in the vertical direction, for a total of 60 points. A 2x2 sub-block needs to interpolate 2x(2+7) points in the horizontal direction and 2x2 points in the vertical direction, for a total of 4 sub-blocks and 88 points. The amount of calculation is more than doubled.

[0073] On the other hand, the bandwidth will increase, because each sub-block of Pattern1 has a new MV, and the MV of the sub-block of Pattern1 is not necessarily the same as the MV of the sub-block of Pattern0 around it, and the MV of the sub-block of Pattern1 is misaligned with the sub-block of Pattern0 around it, so the bandwidth will increase. The amount of bandwidth increase depends on the specific hardware implementation, the control of the interpolation order between sub-blocks, and the algorithm's restrictions on MVs.

[0074] On the other hand, the differences between small blocks such as 2x2, 4x2, and 2x4 will also bring additional complexity, as well as the control of the interpolation order between each divided sub-block, the storage of intermediate values, the control of weighted timing, etc., which will bring additional complexity.

[0075] In summary, the performance of interleaving prediction comes from "interleaving", but "interleaving" also brings obvious complexity. Based on this technology, although some simplified methods have been proposed, they are all carried out under the framework of "interleaving" and cannot effectively reduce the complexity brought by "interleaving". In other words, the above problems can only be improved but not eliminated.

[0076] It can be seen that although the new sub-block division method introduced by interleaved prediction can improve the prediction accuracy to a certain extent, interleaved prediction also brings obvious complexity and reduces the encoding and decoding efficiency.

[0077] In order to solve the defects existing in the prior art, in an embodiment of the present application, sub-block-based prediction can be used for the current block to obtain a set of prediction values, namely, the first prediction value, and point-based prediction can be used for the current block to obtain another set of prediction values, namely, the second prediction value, and after determining the first weight of the sub-block-based prediction and the second weight of the point-based prediction for the same pixel point in the current block, the first weight and the second weight are used to perform weighted averaging on the first prediction value and the second prediction value, and finally a new prediction value of the current block is obtained, thereby improving the accuracy of inter-frame prediction while reducing the complexity of calculation and reducing bandwidth, thereby greatly improving the encoding performance and improving the encoding and decoding efficiency.

[0078] It should be understood that the embodiment of the present application provides a video encoding system. Figure 7 A schematic block diagram of a video encoding system provided in an embodiment of the present application is shown in FIG. Figure 7As shown, the video coding system 11 may include: a transform unit 111, a quantization unit 112, a mode selection and coding control logic unit 113, an intra-frame prediction unit 114, an inter-frame prediction unit 115 (including: motion compensation and motion estimation), an inverse quantization unit 116, an inverse transform unit 117, a loop filtering unit 118, a coding unit 119 and a decoded image buffer unit 110; for the input original video signal, a coding tree block (Coding Tree A video reconstruction block can be obtained by dividing the video reconstruction block into a video unit (CTU), and the coding mode is determined by the mode selection and coding control logic unit 113. Then, the residual pixel information obtained after intra-frame or inter-frame prediction is transformed by the transformation unit 111 and the quantization unit 112, including transforming the residual information from the pixel domain to the transform domain, and quantizing the obtained transform coefficients to further reduce the bit rate; the intra-frame prediction unit 114 is used to perform intra-frame prediction on the video reconstruction block; wherein the intra-frame prediction unit 114 is used to determine the optimal intra-frame prediction mode (i.e., the target prediction mode) of the video reconstruction block; the inter-frame prediction unit 115 is used to perform inter-frame prediction coding of the received video reconstruction block relative to one or more blocks in one or more reference frames to provide temporal prediction information; wherein, the motion estimation is used to generate The motion vector process can estimate the motion of the video reconstruction block, and then the motion compensation is performed based on the motion vector determined by the motion estimation; after determining the inter-frame prediction mode, the inter-frame prediction unit 115 is also used to provide the selected inter-frame prediction data to the encoding unit 119, and the calculated motion vector data is also sent to the encoding unit 119; in addition, the inverse quantization unit 116 and the inverse transformation unit 117 are used to reconstruct the video reconstruction block, reconstruct the residual block in the pixel domain, and the reconstructed residual block is removed by the loop filter unit 118. The reconstructed residual block is then added to a predictive block in the frame of the decoded image cache unit 110 to generate a reconstructed video reconstruction block; the encoding unit 119 is used to encode various encoding parameters and quantized transform coefficients. The decoded image cache unit 110 is used to store the reconstructed video reconstruction block for prediction reference. As the video image encoding proceeds, new reconstructed video reconstruction blocks will be continuously generated, and these reconstructed video reconstruction blocks will be stored in the decoded image cache unit 110.

[0079] The embodiment of the present application also provides a video decoding system, Figure 8 A schematic block diagram of a video decoding system provided in an embodiment of the present application is shown in FIG. Figure 8As shown, the video decoding system 12 may include: a decoding unit 121, an inverse transform unit 127, an inverse quantization unit 122, an intra-frame prediction unit 123, a motion compensation unit 124, a loop filter unit 125 and a decoded image cache unit 126; after the input video signal is encoded by the video encoding system 11, a code stream of the video signal is output; the code stream is input into the video decoding system 12, and first passes through the decoding unit 121 to obtain a decoded transform coefficient; the transform coefficient is processed by the inverse transform unit 127 and the inverse quantization unit 122 to generate a residual block in the pixel domain; the intra-frame prediction unit 123 can be used to generate a prediction number of the current video decoding block based on the determined intra-frame prediction direction and the data of the previously decoded block from the current frame or picture. The motion compensation unit 124 determines prediction information for the video decoding block by analyzing the motion vector and other associated syntax elements, and uses the prediction information to generate a predictive block of the video decoding block being decoded; a decoded video block is formed by summing the residual block from the inverse transform unit 127 and the inverse quantization unit 122 with the corresponding predictive block generated by the intra-frame prediction unit 123 or the motion compensation unit 124; the decoded video signal passes through the loop filtering unit 125 to remove the block effect artifacts, which can improve the video quality; the decoded video block is then stored in the decoded image cache unit 126, which stores the reference image used for subsequent intra-frame prediction or motion compensation, and is also used for the output of the video signal to obtain the restored original video signal.

[0080] An inter-frame prediction method provided in an embodiment of the present application mainly acts on the inter-frame prediction unit 215 of the video encoding system 11 and the inter-frame prediction unit of the video decoding system 12, namely, the motion compensation unit 124; that is, if the video encoding system 11 can obtain a better prediction effect through the inter-frame prediction method provided in an embodiment of the present application, then, correspondingly, in the video decoding system 12, the video decoding recovery quality can also be improved.

[0081] Based on this, the technical solution of the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. Before the detailed description, it should be noted that the "first", "second", "third", etc. mentioned throughout the specification are only for distinguishing different features and do not have the function of limiting priority, sequence, size relationship, etc.

[0082] It should be noted that this embodiment is exemplified based on the AVS3 standard. The inter-frame prediction method proposed in this application can also be applied to other coding standard technologies such as VVC, and the application does not make any specific limitation on this.

[0083] The embodiment of the present application provides an inter-frame prediction method, which is applied to a video decoding device, i.e., a decoder. The function implemented by the method can be implemented by calling a computer program by a first processor in the decoder, and of course the computer program can be stored in a first memory. It can be seen that the decoder includes at least a first processor and a first memory.

[0084] Furthermore, in the embodiments of the present application, Fig. 9 Schematic diagram of the implementation process of the inter-frame prediction method Figure 1 ,like Fig. 9 As shown, the method for the decoder to perform inter-frame prediction may include the following steps:

[0085] Step 101: parse the bitstream to obtain prediction mode parameters of the current block.

[0086] In an embodiment of the present application, the decoder may first parse the binary code stream to obtain prediction mode parameters of the current block, wherein the prediction mode parameters may be used to determine the prediction mode used by the current block.

[0087] It should be noted that the image to be decoded can be divided into multiple image blocks, and the current image block to be decoded can be called a current block (which can be represented by CU), and the image blocks adjacent to the current block can be called adjacent blocks; that is, in the image to be decoded, there is an adjacent relationship between the current block and the adjacent blocks. Here, each current block can include a first image component, a second image component, and a third image component, that is, the current block represents an image block in the image to be decoded that is currently to be predicted for the first image component, the second image component, or the third image component.

[0088] Among them, assuming that the current block predicts the first image component, and the first image component is a luminance component, that is, the image component to be predicted is a luminance component, then the current block can also be called a luminance block; or, assuming that the current block predicts the second image component, and the second image component is a chrominance component, that is, the image component to be predicted is a chrominance component, then the current block can also be called a chrominance block.

[0089] Further, in an embodiment of the present application, the prediction mode parameter may not only indicate the prediction mode adopted by the current block, but also indicate parameters related to the prediction mode.

[0090] It can be understood that, in the embodiments of the present application, the prediction mode may include an inter-frame prediction mode, a traditional intra-frame prediction mode, a non-traditional intra-frame prediction mode, and the like.

[0091] That is to say, on the encoding side, the encoder can select the optimal prediction mode to pre-encode the current block. In this process, the prediction mode of the current block can be determined, and then the prediction mode parameters used to indicate the prediction mode can be determined, so that the corresponding prediction mode parameters are written into the bitstream and transmitted from the encoder to the decoder.

[0092] Accordingly, on the decoder side, the decoder can directly obtain the prediction mode parameters of the current block by parsing the bitstream, and determine the prediction mode used by the current block and the relevant parameters corresponding to the prediction mode according to the prediction mode parameters obtained by parsing.

[0093] Further, in an embodiment of the present application, after parsing and obtaining the prediction mode parameters, the decoder may determine whether the current block uses the inter-frame prediction mode based on the prediction mode parameters.

[0094] Step 102: When the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, the first motion vector of each sub-block of the current block is used to determine the first prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point of the current block is used to determine the second prediction value of each pixel point; wherein the current block includes one or more sub-blocks; and the current block includes one or more pixels.

[0095] In the implementation of the present application, after the decoder parses and obtains the prediction mode parameters, if the prediction mode parameters obtained by parsing indicate that the current block uses the inter-frame prediction mode to determine the inter-frame prediction value of the current block, then the decoder can use the first motion vector of each sub-block of the current block to determine the first prediction value of the pixel points in each sub-block. At the same time, the decoder can use the second motion vector of each pixel point of the current block to determine the second prediction value of each pixel point.

[0096] That is to say, in the present application, the decoder can use sub-block-based prediction for the current block to obtain a set of prediction values, namely, a first prediction value corresponding to each pixel in each sub-block. At the same time, the decoder can also use point-based prediction for the current block to obtain another set of prediction values, namely, a second prediction value corresponding to each pixel.

[0097] It can be understood that, in the present application, the first prediction value is the prediction value of the pixel point in the sub-block obtained by prediction based on the sub-block, and the second prediction value is the prediction value of the pixel point in the current block obtained by prediction based on the point. It can be seen that for the same pixel (pixel position) in the current block, the decoder can obtain the corresponding two prediction values ​​in different ways, namely, the first prediction value obtained by prediction based on the sub-block, and the second prediction value obtained by prediction based on the point.

[0098] It can be understood that, in the present application, the current block may include one or more sub-blocks; the current block may include one or more pixel points (samples), and each pixel point corresponds to a pixel position and a pixel value.

[0099] Exemplarily, in an embodiment of the present application, when determining the first prediction value of a pixel point in each sub-block, the decoder may first determine the first motion vector of each sub-block of the current block, and then determine the first prediction value corresponding to the pixel point in the sub-block based on the first motion vector.

[0100] That is to say, in the present application, when predicting based on sub-blocks, the decoder can divide the current block into several sub-blocks and determine a motion vector for each sub-block, that is, each sub-block corresponds to a first motion vector, and then the decoder can use the first motion vector of each sub-block to predict the sub-block and obtain the first prediction value corresponding to the pixel point in the sub-block.

[0101] Exemplarily, in an embodiment of the present application, when determining the second prediction value of each pixel point, the decoder may first determine the second motion vector of each pixel point of the current block, and then determine the second prediction value corresponding to the pixel point based on the second motion vector.

[0102] That is to say, in the present application, the decoder can determine a motion vector for each pixel when predicting based on sub-blocks, that is, each pixel has a corresponding second motion vector, and then the decoder can use the second motion vector of each pixel to predict the pixel to obtain the corresponding second prediction value.

[0103] It is understandable that in the embodiments of the present application, the first motion vector of the sub-block of the current block and the second motion vector of the pixel can be derived using an affine model. Further, the decoder can also use other methods to determine the first motion vector and the second motion vector, such as other sub-block-based prediction techniques, such as decoder-side motion vector refinement (DMVR), BIO (bi-directional optical flow) called BDOF in VVC, sub-block based temporal motion vector prediction (SbTMVP), motion vector angular prediction (MVAP), etc. The decoder can also determine the second motion vector of the pixel in the sub-block according to the first motion vector of the sub-block, such as determining the second motion vector of any pixel in one of the sub-blocks according to the first motion vector of two or more sub-blocks.

[0104] Further, in an embodiment of the present application, the decoder may use different prediction methods for sub-block-based prediction and point-based prediction. For example, the decoder may use different interpolation methods for sub-block-based prediction and point-based prediction. Specifically, in sub-block-based prediction and point-based prediction, the decoder may use different interpolation filters to interpolate pixel values.

[0105] Exemplarily, in the present application, a relatively complex filter may be used for sub-block-based prediction, and a relatively simple filter may be used for point-based prediction.

[0106] Exemplarily, in the present application, sub-block-based prediction may use a filter with a larger number of taps, and point-based prediction may use a filter with a smaller number of taps.

[0107] Exemplarily, in the present application, sub-block-based prediction and point-based prediction may use a separable two-dimensional filter or a non-separable two-dimensional filter, respectively.

[0108] It should be noted that in an embodiment of the present application, the separable filter can be a two-dimensional filter. Specifically, the separable filter can be a horizontally and vertically separable two-dimensional filter or a horizontally and vertically separable two-dimensional filter, and can be composed of a one-dimensional filter in the horizontal direction and a one-dimensional filter in the vertical direction.

[0109] That is to say, in the present application, if a separable filter is used, the separable filter can perform filtering processing in the horizontal direction and the vertical direction respectively.

[0110] It can be understood that in the present application, as a separable filter that is separable horizontally and vertically, the two-dimensional pixel points can be filtered separately in two directions. Specifically, the separable filter can first filter in one direction (such as horizontal direction or vertical direction) to obtain the intermediate value corresponding to the direction, and then filter the intermediate value in another direction (such as vertical direction or horizontal direction) to obtain the final filtering result.

[0111] It should be noted that in this application, separable filters have been used in many common coding and decoding scenarios, such as inter-frame block-based prediction interpolation filtering, inter-frame sub-block-based prediction interpolation filtering, and affine sub-block-based prediction interpolation filtering. Furthermore, separable filters can also be applied to block-based or sub-block-based motion compensation in HEVC, VVC, and AVS3, that is, the motion compensation prediction interpolation filter can include a separable two-dimensional filter. Exemplarily, the brightness interpolation filter of affine motion compensation prediction in AVS3 can be a separable two-dimensional filter, where the brightness interpolation filter coefficients are shown in Table 1:

[0112] Table 1

[0113]

[0114] Furthermore, in an embodiment of the present application, when performing sub-block-based prediction, for a sub-block of the current block, the decoder can determine a first filtering parameter based on a first motion vector; then, based on the first filtering parameter, a first filter can be used to perform filtering processing to obtain a first prediction value.

[0115] Specifically, in the present application, the first filter can be any one of the following filters: an n-tap interpolation filter, a separable two-dimensional filter, and an inseparable two-dimensional filter; wherein n is any one of the following values: 8, 6, 5, 4, 3, 2.

[0116] It should be noted that, in an embodiment of the present application, the first filtering parameter of the first filter may include the filter coefficient of the first filter, and may also include the filter phase of the first filter, which is not specifically limited in the present application.

[0117] Exemplarily, in the present application, if the first filtering parameter is the filter coefficient of the first filter, then when the decoder determines the first filtering parameter based on the first motion vector, it can first determine the first scale parameter; and then determine the first filtering parameter based on the first scale parameter and the first motion vector.

[0118] Exemplarily, in the present application, if the first filtering parameter is the filter phase of the first filter, then when the decoder determines the first filtering parameter based on the first motion vector, it can first determine a mapping table of the first phase and the motion vector; and then determine the first filtering parameter based on the mapping table of the first phase and the motion vector and the first motion vector.

[0119] Furthermore, in an embodiment of the present application, when performing point-based prediction, for a pixel point of the current block, the decoder can determine a second filtering parameter based on a second motion vector; then, based on the second filtering parameter, a second filter can be used to perform filtering processing to obtain a second prediction value.

[0120] Specifically, in the present application, the second filter can be any one of the following filters: an m-tap interpolation filter, a separable two-dimensional filter, and an inseparable two-dimensional filter; wherein m is any one of the following values: 8, 6, 5, 4, 3, 2.

[0121] It can be understood that in the present application, m can be less than or equal to n, that is, sub-block-based prediction can use a filter with a larger number of taps, and point-based prediction can use a filter with a smaller number of taps.

[0122] It should be noted that, in an embodiment of the present application, the second filtering parameter of the second filter may include the filter coefficient of the second filter, and may also include the filter phase of the second filter, which is not specifically limited in the present application.

[0123] Exemplarily, in the present application, if the second filtering parameter is the filter coefficient of the second filter, then when the decoder determines the second filtering parameter based on the second motion vector, it can first determine the second scale parameter; and then determine the second filtering parameter based on the second scale parameter and the second motion vector.

[0124] Exemplarily, in the present application, if the second filtering parameter is the filter phase of the second filter, then when the decoder determines the second filtering parameter based on the second motion vector, it can first determine a mapping table of the second phase and the motion vector; and then determine the second filtering parameter based on the mapping table of the second phase and the motion vector and the second motion vector.

[0125] That is to say, in the embodiment of the present application, the first filter based on the prediction of the sub-block and the second filter based on the prediction of the point can adopt different forms respectively. Specifically, the first filter coefficient of the first filter and the second filter coefficient of the second filter can be determined respectively in different ways. For example, the prediction based on the sub-block uses the form of the coefficient determined from the table according to the motion vector of the sub-pixel as shown in Table 1 above. The prediction based on the point can use the form of determining the coefficient from the table according to the motion vector of the sub-pixel, or the form of calculating the coefficient according to the motion vector of the sub-pixel.

[0126] Exemplarily, in an embodiment of the present application, when determining the first prediction value of a pixel point in each sub-block of the current block, the decoder may also first determine the motion vector deviation between each sub-block and the current block; and then determine the first prediction value of the pixel point in each sub-block based on the motion vector deviation.

[0127] It should be noted that, in the present application, the first proportional parameter or the second proportional parameter may include at least one proportional value, wherein the at least one proportional value is a non-zero real number.

[0128] Furthermore, in an embodiment of the present application, any one of the filter coefficients obtained by the decoder based on the motion vector calculation may be a linear function (polynomial), a quadratic function (polynomial) or a higher-order function (polynomial) of the motion vector, and the present application does not make any specific limitation thereto.

[0129] That is to say, in the present application, when the decoder calculates multiple filter coefficients corresponding to multiple pixel points according to different calculation methods in the preset calculation rules, some filter coefficients can be a linear function (polynomial) of the motion vector, that is, the two are in a linear relationship, or they can be a quadratic function (polynomial) or a higher-order function (polynomial) of the motion vector, that is, the two are in a non-linear relationship.

[0130] In some scenarios, even without calculating the first motion vector first, the first prediction value of the sub-block can be determined. For example, bidirectional optical flow, referred to as BIO in AVS and BDOF in VVC, uses bidirectional optical flow to calculate the motion vector deviation between the sub-block and the current block, and then uses the principle of optical flow to obtain a new prediction value based on the gradient, motion vector deviation and the current prediction value.

[0131] It should be noted that in the present application, in some scenarios, calculations are performed on a block (coding block or prediction block), such as a 4x4 block, as a whole, such as the calculation of motion vector deviation, correction of prediction value, etc., which can also be processed according to the idea of ​​sub-blocks in the present application.

[0132] It should be noted that in the embodiment of the present application, the current block is an image block to be decoded in the current frame, and the current frame is decoded in the form of image blocks in a certain order, and the current block is an image block to be decoded at the next moment in the current frame in the order. The current block can have a variety of specifications, such as 1616, 3232 or 3216, where the numbers represent the number of rows and columns of pixels on the current block.

[0133] Furthermore, in an embodiment of the present application, the current block can be divided into a plurality of sub-blocks, wherein each sub-block has the same size, and the sub-block is a set of pixels of a smaller size. The size of the sub-block can be 8x8 or 4x4.

[0134] Exemplarily, in the present application, the size of the current block is 1616, which can be divided into 4 sub-blocks of 8x8 size each.

[0135] It can be understood that in the embodiment of the present application, when the decoder parses the code stream and obtains the prediction mode parameter indicating the use of the inter-frame prediction mode to determine the inter-frame prediction value of the current block, the inter-frame prediction method provided in the embodiment of the present application can continue to be used.

[0136] Furthermore, in an embodiment of the present application, the decoder may use a sub-block of the current block as a reference block; then use the first motion vector of the reference block to determine the first prediction value of the pixel points in the reference block, and use the second motion vector of each pixel point of the reference block to determine the second prediction value of each pixel point; wherein the reference block includes one or more pixel points.

[0137] It should be noted that, in the embodiment of the present application, since the current block includes one or more sub-blocks, and at the same time, the current block includes one or more pixel points, therefore, if a pixel point belongs to a sub-block, then while prediction is performed based on the sub-block, the pixel point can also be predicted based on the pixel point using the pixel points in the sub-block. That is to say, when the decoder performs sub-block-based prediction and point-based prediction, the point-based prediction and the sub-block-based prediction can share the same reference block, wherein the reference block is a block composed of reference pixels required for interpolation filtering. Specifically, for a sub-block, each pixel point constituting the sub-block can be determined as a reference block, and the reference block used for the sub-block-based prediction of the sub-block and the reference block used for the point-based prediction of each point are all within this determined reference block. Preferably, the sub-block-based prediction and the point-based prediction of the corresponding point in the sub-block can be performed simultaneously, or the sub-block-based prediction and the point-based prediction of the corresponding point in the sub-block can be performed continuously, thereby avoiding repeated reading of reference pixels into the cache.

[0138] It is understandable that in the present application, if the point-based prediction and the sub-block-based prediction can share the same reference block, then the decoder can limit the MV of the point-based prediction so as to limit whether to increase the bandwidth. For example, the sub-block-based prediction uses an 8-tap filter that is separable horizontally and vertically, and the point-based prediction uses a 4-tap filter that is separable horizontally and vertically. Then, if the MV of the point in the upper left corner exceeds the MV of the sub-block by 2 pixels to the left or upward, the reference block required for the point will not exceed the reference block of the sub-block. Then, if the MV of the point in the upper right corner exceeds the MV of the sub-block by 2 pixels to the right or upward, the reference block required for the point will not exceed the reference block of the sub-block. Accordingly, the points inside the sub-block can move in a larger range.

[0139] Furthermore, the restriction on the MV of the point can be appropriately relaxed, and the performance and complexity can be balanced to appropriately increase the size of the reference block.

[0140] Step 103: for a pixel in the current block, determine a first weight and a second weight of the pixel; wherein the first weight corresponds to a first prediction value, and the second weight corresponds to a second prediction value.

[0141] In the implementation of the present application, for a pixel point in the current block, the decoder can respectively determine the first weight and the second weight of the pixel point; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value. Specifically, the first weight and the second weight of the same pixel point can be the same or different.

[0142] It should be noted that in an embodiment of the present application, for a pixel point in the current block, the decoder can obtain a first prediction value and a second prediction value through sub-block-based prediction and point-based prediction, respectively. Correspondingly, the decoder can also set weight values ​​for these two prediction values ​​to determine a first weight corresponding to the first prediction value and a second weight corresponding to the second prediction value.

[0143] It can be understood that in the present application, for a pixel point in the current block, the first weight is the weight value corresponding to the decoder's prediction based on the sub-block, and the second weight is the weight value corresponding to the decoder's prediction based on the point.

[0144] Furthermore, in an embodiment of the present application, the decoder may use a variety of different methods to respectively determine the first weight and the second weight, wherein the decoder may use the same method to determine the first weight and the second weight, or may use different methods to determine the first weight and the second weight.

[0145] Exemplarily, in the present application, when determining the first weight corresponding to a sub-block-based prediction of a pixel in the current block, the decoder can first determine the target sub-block corresponding to the pixel; wherein the target sub-block includes the pixel, that is, determine which sub-block of the multiple sub-blocks of the current block the pixel belongs to; then determine the first distance between the pixel position of the pixel and the reference position of the target sub-block; wherein the reference position is the position used by the first motion vector of the target sub-block; finally, the decoder can determine the first weight according to the first distance; wherein the first distance is inversely proportional to the first weight, that is, the larger the first distance, the farther the pixel position of the pixel is from the reference position, and the smaller the first weight corresponding to the pixel.

[0146] That is, in the present application, based on the prediction of the sub-block, the decoder gives a larger weight to the pixel position that is closer to the pixel position used by the first motion vector of the sub-block, and a smaller weight to the pixel position that is farther from the pixel position used by the first motion vector of the sub-block.

[0147] Exemplarily, in the present application, when determining the first weight corresponding to a sub-block-based prediction of a pixel in the current block, the decoder can first determine the target sub-block corresponding to the pixel; wherein the target sub-block includes the pixel, that is, it is determined to which sub-block of multiple sub-blocks of the current block the pixel belongs; then the deviation value between the motion vector of the pixel and the first motion vector of the target sub-block can be determined; finally, the first weight can be further determined according to the deviation value; wherein the deviation value is inversely proportional to the first weight, that is, the smaller the deviation value, the closer the motion vector of the pixel is to the first motion vector of the target sub-block, and the larger the first weight corresponding to the pixel.

[0148] That is to say, in the present application, based on the prediction of the sub-block, the decoder gives a larger weight to the pixel point whose motion vector is closer to the first motion vector of the sub-block, and a smaller weight to the pixel point whose motion vector is far from the first motion vector of the sub-block.

[0149] Exemplarily, in the present application, when determining the second weight corresponding to the point-based prediction of a pixel in the current block, the decoder can first determine the corresponding sub-pixels and integer pixels of the pixel in the reference image; then determine the second distance between the sub-pixels and the integer pixels; finally, determine the second weight according to the second distance; wherein the second distance is inversely proportional to the second weight, that is, the larger the second distance, the greater the distance between the sub-pixels and the integer pixels, and the smaller the second weight corresponding to the pixel.

[0150] That is to say, in the present application, based on point prediction, the decoder gives a larger weight to the sub-pixel position in the reference image corresponding to the pixel position of the pixel point that is closer to the integer pixel position, and a smaller weight to the sub-pixel position in the reference image corresponding to the pixel position of the pixel point that is farther from the integer pixel.

[0151] Exemplarily, in the present application, when determining the second weight corresponding to the point-based prediction of a pixel in the current block, the decoder can first calculate the absolute value of the motion vector of the pixel; then the second weight can be further determined according to the absolute value; wherein the absolute value is inversely proportional to the second weight, that is, the smaller the absolute value, the larger the corresponding second weight.

[0152] That is to say, in the present application, based on point prediction, the decoder gives a larger weight to the smaller absolute value of the motion vector of the pixel point, and a smaller weight to the larger absolute value of the motion vector of the pixel point. For example, the second weight corresponding to the motion vector of the pixel point is (1 / 4, 1 / 4), which is greater than the second weight corresponding to the motion vector of (1 / 2, 1 / 2).

[0153] It can be understood that, in the present application, the decoder may use different methods or the same method when setting the first weight and the second weight.

[0154] Furthermore, in the present application, the setting of the first weight based on the prediction of the sub-block can use a method similar to the interlaced prediction, that is, based on the prediction of the sub-block, a larger first weight is set for the pixel position that is closer to the pixel position used by the first motion vector of the sub-block, and a smaller first weight is set for the pixel position that is farther from the pixel position used by the first motion vector of the sub-block.

[0155] For example, Fig.10 The first weight is shown Figure 1 ,like Fig.10 As shown, the second weight corresponding to the pixel position closer to the pixel position used by the first motion vector of the sub-block is set to 4, and the second weight corresponding to the pixel position farther from the pixel position used by the first motion vector of the sub-block is set to 1. The present application does not limit the specific value of the second weight.

[0156] For example, Fig.11 The first weight is shown Figure 2 ,like Fig.11 As shown, the second weight corresponding to the pixel position closer to the pixel position used by the first motion vector of the sub-block is set to 3, and the second weight corresponding to the pixel position farther from the pixel position used by the first motion vector of the sub-block is set to 1. The present application does not limit the specific value of the second weight.

[0157] Furthermore, in the present application, the second weight for point-based prediction may be set using a different setting method from the second weight for sub-block-based prediction. For example, the decoder may set a larger second weight for a pixel point whose pixel position corresponds to a sub-pixel position in the reference image that is closer to an integer pixel, and set a smaller second weight for a pixel point whose pixel position corresponds to a sub-pixel position in the reference image that is farther from an integer pixel. Or, a larger second weight is set for a pixel point whose motion vector is closer to an integer, and a smaller second weight is set for a pixel point whose motion vector is farther from an integer.

[0158] Illustratively, in the present application, the second weight of the pixel position where the motion vector of the pixel is (1 / 4, 1 / 4) is greater than the second weight of the pixel position where the motion vector of the pixel is (1 / 2, 1 / 2).

[0159] Further, in the present application, if the decoder first determines the position of the second filter, that is, the decoder first determines the motion vector of the integer pixel and the motion vector of the sub-pixel corresponding to the current pixel, then when setting the second weight based on point prediction, the pixel position with a larger absolute value of the sub-pixel motion vector is set to a smaller second weight, and the pixel position with a smaller absolute value of the sub-pixel motion vector is set to a larger second weight.

[0160] Exemplarily, in the present application, the method of deriving the second weight w from the sub-pixel motion vector (x, y) is:

[0161] w=clip3(0,1,1-abs(x)-abs(y))

[0162] Where abs represents the absolute value, and clip3(a, b, c) means that if c is less than a, the result is a, if c is greater than b, the result is b, otherwise the result is c. Specifically, this is a normalized expression, where 1 in x and y represents 1 pixel. For the convenience of calculation, this formula can also be deformed, such as by shifting to remove fractions / decimals.

[0163] It can be understood that, in the present application, in order to match the first weight based on the prediction value of the sub-block, the second weight w can be multiplied by a multiple, etc.

[0164] It can be understood that the inter-frame prediction method proposed in the embodiment of the present application does not limit the order in which the decoder executes step 102 and step 103. That is to say, in the present application, the decoder can first execute step 102 and then step 103, or first execute step 103 and then step 102, or execute step 102 and step 103 at the same time.

[0165] Step 104: Determine the predicted value of the pixel based on the first weight, the second weight, the first predicted value, and the second predicted value.

[0166] In an embodiment of the present application, after the decoder obtains the first prediction value and the corresponding first weight based on the prediction of the sub-block, and obtains the second prediction value and the corresponding second weight based on the prediction of the point, the decoder can use the first weight, the second weight, the first prediction value and the second prediction value to obtain the prediction value of a pixel point in the current block.

[0167] It can be understood that in an embodiment of the present application, after performing sub-block-based prediction and point-based prediction, for a pixel point in the current block, the decoder can, after determining the first weight and the second weight respectively corresponding to the pixel point, perform weighted averaging on the first prediction value and the second prediction value using the first weight and the second weight, thereby obtaining the prediction value of the pixel point.

[0168] Fig.12 Schematic diagram of the implementation process of the inter-frame prediction method Figure 2 ,like Fig.12 As shown, before determining the third prediction value of the current block according to the prediction value of the pixel point, that is, before step 105, the method for the decoder to perform inter-frame prediction may further include the following steps:

[0169] Step 106: For a pixel in the current block, determine a prediction value of the pixel based on the first prediction value and the second prediction value.

[0170] In an embodiment of the present application, after the decoder obtains a first prediction value based on sub-block prediction and a second prediction value based on point prediction, it can use the first prediction value and the second prediction value to obtain a prediction value of a pixel point in the current block.

[0171] It can be understood that in the embodiment of the present application, compared with step 104, the decoder can also directly average the first prediction value and the second prediction value of the same pixel point after respectively obtaining the first prediction value and the second prediction value to obtain the prediction value of the pixel point.

[0172] That is to say, in the present application, the decoder may not perform a weighted average operation on the first prediction value and the second prediction value, but directly use the two prediction values ​​to determine the final prediction result.

[0173] It should be noted that, in the present application, the method proposed in step 106 can also be used as one of the cases of step 104, that is, the case where the first weight and the second weight are equal.

[0174] Step 105 : Determine a third prediction value of the current block according to the prediction value of the pixel point; wherein the third prediction value is used to determine a reconstruction value of the current block.

[0175] In an embodiment of the present application, after determining the prediction value of the pixel based on the first weight, the second weight, the first prediction value and the second prediction value, the decoder can further determine the third prediction value of the current block according to the prediction value of the pixel.

[0176] It should be noted that in the present application, the decoder can filter all the pixels in the current block according to the above method, and use the prediction values ​​of all the pixels in the current block to determine the third prediction value of the current block, or filter some of the pixels in the current block, and use the prediction values ​​of some of the pixels in the current block to determine the third prediction value of the sub-block.

[0177] It can be understood that, in the embodiment of the present application, the third prediction value is used to determine the reconstruction value of the current block.

[0178] Furthermore, in an embodiment of the present application, the decoder can, after traversing all or part of the pixels in the current block, perform an addition operation on the predicted values ​​of all or part of the pixels in the current block to obtain a sum result, and finally normalize the addition result to finally obtain a third predicted value of the current block.

[0179] In the embodiment of the present application, further, the inter-frame prediction method proposed in the above steps 101 to 106 can be applied to both unidirectional prediction and bidirectional prediction.

[0180] Specifically, Fig.13 A schematic diagram of one-way prediction is shown in Figure 2. Fig.13 As shown, the decoder uses sub-block-based prediction for the current block to obtain a set of prediction values, namely, the first prediction value, and then generates a sub-block-based prediction block. At the same time, the decoder uses point-based prediction for the current block to obtain another set of prediction values, namely, the second prediction value, and then generates a point-based prediction block. The two prediction values ​​for the same pixel are weighted averaged to obtain the prediction value of the pixel, and finally the prediction block of the current block is obtained.

[0181] Furthermore, in the embodiments of the present application, Fig.14 Schematic diagram of the implementation process of the inter-frame prediction method Figure 3 ,like Fig.14 As shown, the method for the decoder to perform inter-frame prediction may also include the following steps:

[0182] Step 201: parse the bitstream and determine the inter-frame prediction direction.

[0183] Step 202: If the inter-frame prediction direction is bidirectional prediction, based on the first prediction direction, the first motion vector of each sub-block is used to determine the fourth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the fifth prediction value of each pixel point; based on the second prediction direction, the first motion vector of each sub-block is used to determine the sixth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the seventh prediction value of each pixel point.

[0184] Step 203: For a pixel of the current block, determine a third weight and a fourth weight of the pixel based on the first prediction direction; and determine a fifth weight and a sixth weight of the pixel based on the second prediction direction.

[0185] Step 204: Based on the third weight, the fourth weight, the fourth prediction value and the fifth prediction value, obtain the prediction value of the pixel point corresponding to the first prediction direction; based on the fifth weight, the sixth weight, the sixth prediction value and the seventh prediction value, obtain the prediction value of the pixel point corresponding to the second prediction direction.

[0186] Step 205: Determine the prediction value of the pixel point according to the prediction value corresponding to the first prediction direction and the prediction value corresponding to the second prediction direction.

[0187] In contrast, Fig.15 This is a schematic diagram of bidirectional prediction. Figure 1 ,like Fig.15 As shown, for the first reference direction and the second reference direction, the decoder can perform sub-block-based prediction and point-based prediction for each unidirectional prediction, and obtain two prediction values ​​corresponding to the prediction direction for the same pixel, such as the fourth prediction value and the fifth prediction value, and finally obtain the sub-block-based prediction block and the point-based prediction block in each reference direction, and then obtain the prediction value of the pixel in the prediction direction by weighted average of the third weight and the fourth weight, so as to obtain the prediction blocks of the first reference direction and the second reference direction respectively. Finally, according to the prediction value of the pixel in each prediction direction, the final prediction result of the pixel is calculated, that is, the prediction block corresponding to the current block is obtained.

[0188] That is to say, in the present application, when the decoder performs bidirectional prediction, for each unidirectional prediction, a set of prediction values ​​is obtained by using sub-block-based prediction for the current block, and another set of prediction values ​​is obtained by using point-based prediction for the current block. The two prediction values ​​for the same pixel are weighted averaged to obtain the prediction value of the pixel. The prediction values ​​of the two unidirectional predictions are then averaged (or weighted averaged) to finally obtain the prediction value of the bidirectional prediction.

[0189] It should be noted that in the present application, the averaging operation or weighted averaging operation in the inter-frame prediction processing process can be performed after the entire prediction block that needs to be averaged or weighted averaged is obtained, or it can be performed after the prediction blocks of the sub-blocks that need to be averaged or weighted averaged are obtained, or it can be performed after the prediction blocks of the points that need to be averaged or weighted averaged are obtained, and the present application does not make specific limitations.

[0190] Furthermore, in the embodiments of the present application, Fig.16 Schematic diagram of the implementation process of the inter-frame prediction method Figure 4 ,like Fig.16 As shown, the method for the decoder to perform inter-frame prediction may also include the following steps:

[0191] Step 201: parse the bitstream and determine the inter-frame prediction direction.

[0192] Step 202: If the inter-frame prediction direction is bidirectional prediction, based on the first prediction direction, the first motion vector of each sub-block is used to determine the fourth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the fifth prediction value of each pixel point; based on the second prediction direction, the first motion vector of each sub-block is used to determine the sixth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the seventh prediction value of each pixel point.

[0193] Step 203: For a pixel of the current block, determine a third weight and a fourth weight of the pixel based on the first prediction direction; and determine a fifth weight and a sixth weight of the pixel based on the second prediction direction.

[0194] Step 206: Determine the predicted value of the pixel based on the third weight, the fourth weight, the fifth weight, the sixth weight, the fourth predicted value, the fifth predicted value, the sixth predicted value and the seventh predicted value.

[0195] In contrast, Fig.17 This is a schematic diagram of bidirectional prediction. Figure 2 ,like Fig.17 As shown, for the first reference direction and the second reference direction, the decoder can perform sub-block-based prediction and point-based prediction for each unidirectional prediction, respectively, and obtain two prediction values ​​corresponding to the prediction direction for the same pixel point, such as the fourth prediction value and the fifth prediction value, and finally obtain the sub-block-based prediction block and the point-based prediction block in each reference direction, and after determining the two prediction values ​​on each unidirectional prediction, a weighted average calculation can be performed using the weight corresponding to each prediction value, and finally the final prediction result of the pixel point is obtained, that is, the prediction block corresponding to the current block is obtained.

[0196] That is to say, in the present application, when the decoder performs bidirectional prediction, for each unidirectional prediction, a set of prediction values ​​is obtained for the current block using sub-block-based prediction, and another set of prediction values ​​is obtained for the current block using point-based prediction. The four prediction values ​​of the two unidirectional predictions are then averaged (or weighted averaged), and finally the prediction value of the bidirectional prediction can be obtained.

[0197] It can be seen that for bidirectional prediction, the decoder can determine the respective prediction values ​​of the two prediction directions and then perform averaging or weighted averaging operations, or it may not calculate the prediction value of each prediction direction separately, but directly obtain it by weighted averaging the prediction block based on sub-blocks in the first prediction direction, the prediction block based on points in the first prediction direction, the prediction block based on sub-blocks in the second prediction direction, and the prediction block based on points in the second prediction direction.

[0198] It is understandable that, in the present application, the averaging operation can be considered as a special weighted averaging operation, and the pixel point pixel position can be understood as a different expression of the same concept in some scenes.

[0199] Furthermore, the inter-frame prediction method proposed in the present application can be used only for the luminance component, or for the luminance component and the chrominance component, or for any other format, such as one, several, or all components of RGB, etc. The embodiment of the present application is described by taking the luminance component as an example, but is not limited to the luminance component.

[0200] That is to say, the inter-frame prediction method proposed in the present application can be applied to any image component. In the present embodiment, the prediction scheme is exemplarily used for the luminance component, but it can also be used for the chrominance component or any component in other formats. The inter-frame prediction method proposed in the present application can also be applied to any video format, including but not limited to the YUV format, including but not limited to the luminance component in the YUV format.

[0201] This embodiment provides an inter-frame prediction method, which can use sub-block-based prediction for the current block to obtain a set of prediction values, namely, first prediction values, and use point-based prediction for the current block to obtain another set of prediction values, namely, second prediction values, and after determining the first weight of sub-block-based prediction and the second weight of point-based prediction for the same pixel point in the current block, use the first weight and the second weight to perform weighted averaging on the first prediction value and the second prediction value, and finally obtain a new prediction value for the current block, thereby improving the accuracy of inter-frame prediction while reducing the complexity of calculation and reducing bandwidth, thereby greatly improving encoding performance and improving encoding and decoding efficiency.

[0202] Based on the above embodiments, in this application, if a horizontally and vertically separable 8-tap filter is used to predict each pixel point separately, each pixel point needs to correspond to interpolation 8+1, a total of 9 points, and each pixel point in a 4x4 sub-block needs to correspond to interpolation 3.75 points on average. It can be seen that the computational complexity will be much higher than the prediction based on sub-blocks, and the bandwidth will also be greatly increased. Therefore, this application proposes to use a simpler prediction method to perform point-based prediction. For example, the filter corresponding to the sub-block-based prediction uses an 8-tap filter, while the filter corresponding to the point-based prediction uses a 4-tap filter.

[0203] Exemplarily, in the present application, when performing point-based prediction, the decoder can use a horizontally and vertically separable 4-tap filter. Furthermore, the decoder can directly reuse the currently common 4-tap filter based on sub-block prediction, which is currently mainly used for chrominance component interpolation processing. The filter coefficients of AVS3 for affine sub-block-based chrominance are shown in Table 2:

[0204] Table 2

[0205]

[0206]

[0207] Exemplarily, in the present application, the decoder may also use a 3x3 filter for point-based prediction, wherein the 3x3 filter may be a horizontally and vertically inseparable filter, or a horizontally and vertically separable 3-tap filter. In comparison, the computational complexity and bandwidth pressure of using a 3x3 filter for point-based prediction are smaller than using a 4-tap filter.

[0208] Specifically, in this application, a 3-tap filter that is separable horizontally and vertically is taken as an example, when performing point-based prediction. Each pixel point needs to interpolate 3 intermediate points in the horizontal direction and 1 point in the vertical direction, a total of 4 points. Since there are only 3 taps, only 12 multiplications are required. On average, each point of a 4x4 sub-block needs to interpolate 3.75 points. If an 8-tap filter is used, an average of 30 multiplications are required. If a 6-tap filter is used, an average of 22.5 multiplications are required.

[0209] Further, in an embodiment of the present application, if a horizontally and vertically separable 3-tap filter is used, since the center of the 3-tap filter is at the position of the middle tap, the motion vector of the sub-pixel can be set to a negative value. Specifically, when setting the motion vector of the sub-pixel of the 3-tap filter, the motion vector of the sub-pixel can be set in the range of -1 / 2 pixel to 1 / 2 pixel.

[0210] For example, if the motion vector of a pixel is 1 / 4 pixel, the decoder can use the pixel at the corresponding position of the pixel in the reference image as the center and divide the pixel into 1 / 4 motion vectors. Fig.18 This is an illustration of interpolation filtering. Figure 1 ,like Fig.18 As shown, pixel 1 is the pixel at the corresponding position of current pixel 2 in the reference image, and three pixel points centered on pixel 1 can be used to perform interpolation filtering on current pixel 2.

[0211] For example, if the motion vector of a pixel is 3 / 4 pixel, the decoder can use the pixel one pixel to the right of the pixel in the reference image as the center and use the sub-pixel motion vector as -1 / 4. Fig.19 This is an illustration of interpolation filtering. Figure 2 ,like Fig.19 As shown, pixel 1 is the pixel at the corresponding position of current pixel 2 in the reference image, pixel 3 is a pixel to the right of current pixel 2, and three pixels centered on pixel 3 can be used to perform interpolation filtering on current pixel 2.

[0212] Further, in an embodiment of the present application, if a horizontally and vertically separable 3-tap filter is used, when setting the sub-pixel motion vector of the 3-tap filter, the sub-pixel motion vector can also be set in the range of -1 pixel to 1 pixel.

[0213] For example, if the motion vector of a pixel is 3 / 4 pixel, the decoder can use the pixel corresponding to the pixel in the reference image as the center and use the sub-pixel motion vector as 3 / 4 pixel. Fig. 20 This is an illustration of interpolation filtering. Figure 3 ,like Fig. 20 As shown, pixel 1 is the pixel at the corresponding position of current pixel 2 in the reference image, and three pixel points centered on pixel 1 can be used to perform interpolation filtering on current pixel 2.

[0214] Furthermore, in an embodiment of the present application, when performing point-based prediction, the position of the filter used for each pixel point can be determined first, and then the pixel-by-pixel motion vector can be directly calculated based on the position of the filter, and finally interpolation filtering can be performed based on the three points used by the filter.

[0215] For example, the decoder can determine the pixel at the same position as the current pixel in the reference image and use the pixel as the center of the filter corresponding to the current pixel. Specifically, if the motion vector of a pixel is 5 / 4 pixels, then the motion vector of the corresponding sub-pixel is also 5 / 4 pixels, and the motion vector of the integer pixel is 0 pixels. Fig.21 This is an illustration of interpolation filtering. Figure 4 ,like Fig.21 As shown in the figure, pixel 1 is the pixel at the corresponding position of current pixel 2 in the reference image, and three pixels centered on pixel 1 can be used to perform interpolation filtering on current pixel 2. In this scenario, the motion vector of the whole pixel can be called motion vector 1, and the motion vector of the sub-pixel can be called motion vector 2, where motion vector 1 is 0 pixels and motion vector 2 is 5 / 4 pixels.

[0216] It is understandable that in the present application, the above interpolation filtering method can set the position of the reference pixels used by the filter for a group of pixels according to a unified rule, so that it is easier to implement. For example, the same motion vector 1 is set for each pixel of the same sub-block, so that the reference pixels used by the filter of each pixel can be arranged regularly, similar to the rule of the reference pixels used for each point in the prediction based on the sub-block, except that the number of reference pixels used for the prediction based on the point is less. Accordingly, the unified reading of reference pixels can also be realized for hardware, and it is also more beneficial for parallel operation of software, such as Single Instruction Multiple Data (SIMD), thereby reducing the control bandwidth.

[0217] For example, Fig. 22 This is an illustration of interpolation filtering. Figure 5 ,like Fig. 22 As shown in the figure, pixel points A, B, C, and D are the reference pixel points corresponding to the same positions of the four adjacent pixel points in the current block in the reference image, and pixel points a, b, c, and d respectively represent their actual positions in the reference image. If the value range of the sub-pixel motion vector is set to -1 / 2 pixel to 1 / 2 pixel, then in the point-based prediction, the horizontal filters of the horizontally and vertically separable 3-tap filter used are the four horizontal boxes in the figure. It can be seen that the filters of pixel point A and pixel point B differ by 1 pixel in the horizontal direction, but the filters of pixel point A and pixel point C differ by 2 pixels in the horizontal direction, and the filters of pixel point D and pixel point C differ by 1 pixel in the horizontal direction.

[0218] For example, Fig.23 This is an illustration of interpolation filtering. Figure 6 ,like Fig.23As shown, pixel points A, B, C, and D are the reference pixel points corresponding to the same positions of the adjacent pixel points in the current block in the reference image, and pixel points a, b, c, and d represent their actual positions in the reference image respectively. If the position of the filter is determined first, then in point-based prediction, the center of the horizontal filter in the horizontally and vertically separable 3-tap filter used is the position corresponding to these pixel points, that is, the motion vector of the integer pixel (motion vector 1) is 0, and it can be determined that the horizontal filters corresponding to the 4 adjacent pixel points also differ by 1 pixel.

[0219] Further, in an embodiment of the present application, when determining the position of a filter for a group of pixels, that is, when determining the motion vector 1 (motion vector of an integer pixel) of a group of pixels, the principle of dispersing the motion vector 2 (motion vector of a sub-pixel) corresponding to the group of pixels relatively evenly around 0 can be followed. Specifically, after determining the motion vector 2 corresponding to the group of pixels, the motion vectors 2 can be summed up, and then the position of the filter for the group of pixels can be determined according to the position corresponding to the minimum summation result, that is, the decoder can determine the motion vector of the integer pixel corresponding to the minimum sum of the motion vectors 2 corresponding to the group of pixels as the motion vector 1 corresponding to the group of pixels.

[0220] It is understandable that in the present application, the group of pixel points proposed by the above method can be the pixel points of the sub-block divided based on the prediction of the sub-block, or can be the group divided based on the prediction of the sub-block, such as dividing a 4x4 sub-block into 4 2x2 groups. It can also be "interlaced" with the sub-block divided based on the prediction of the sub-block, but because there is no data dependency between the points and the intermediate results of the interpolation are not shared, the complexity is lower than the interlaced prediction.

[0221] Further, in the present application, the decoder can further determine the second filter based on the prediction of the point in other ways. For example, the decoder can determine the second filter parameter of the second filter in the form of a function or a polynomial, wherein the function or polynomial can include a spline function, a piecewise function, etc. The decoder can also further determine the second filter parameter of the second filter in the form of a lookup table (a mapping table of phase and motion vector). This application does not make specific limitations.

[0222] Exemplarily, in the present application, if the second filter used in point-based prediction is a 3-tap filter, then, assuming that the sub-pixel motion vector (motion vector 2) corresponding to the position of the current pixel is (mv_x, mv_y), then the second filtering parameter in the horizontal direction of the second filter, that is, the coefficient of the filter can be expressed as shown in Table 3 below:

[0223] Table 3

[0224] Pixels Filter coefficients Left mv_x×mv_x-mv_x / 2 center 1-2×mv_x×mv_x right mv_x×mv_x+mv_x / 2

[0225] The second filtering parameters of the second filter in the vertical direction, that is, the coefficients of the filter can be expressed as shown in Table 4 below:

[0226] Table 4

[0227] Pixels Filter coefficients superior mv_y×mv_y-mv_y / 2 center 1-2×mv_y×mv_y Down mv_y×mv_y+mv_y / 2

[0228] It should be noted that in the present application, the filter coefficients calculated by the decoder based on the motion vector of the sub-pixel corresponding to the current pixel according to the preset calculation rules can be considered as a function (polynomial) about mv_x or mv_y. However, since the filter of the separable filter in the horizontal direction or vertical direction has only 3 taps, when the sub-pixel point deviates from the center by a large amount, the accuracy of the filter coefficients calculated using the same calculation rules, that is, using the same simple function (such as a first-order or second-order function) may be poor.

[0229] In order to improve the accuracy of the filter coefficients, in the present application, when using the motion vector of the sub-pixel corresponding to the current pixel to calculate the filter coefficients, a piecewise function can be used to replace the original simple function, that is, at least one filter coefficient, namely the second filter coefficient, can be calculated by a piecewise function.

[0230] It is understood that in the present application, the piecewise function may be referred to as a spline function. Exemplarily, when the value (absolute value) of mv_x or mv_y is less than (or equal to) a threshold, a function (polynomial) may be used to derive the corresponding filter coefficient, and when the value (absolute value) of mv_x or mv_y is greater than (or equal to) the threshold, another function (polynomial) may be used to derive the corresponding filter coefficient. That is, a two-segment piecewise function is used to calculate the filter coefficient.

[0231] Furthermore, in the embodiments of the present application, a three-segment or multi-segment piecewise function may be used to calculate the filter coefficients.

[0232] Specifically, in the present application, when calculating different filter coefficients among multiple filter coefficients corresponding to pixel points, the same piecewise function or different piecewise functions can be used; when calculating different filter coefficients among multiple filter coefficients corresponding to pixel points, the thresholds used for mv_x or mv_y may or may not be all the same.

[0233] Further, in an embodiment of the present application, the decoder may also use a preset upper limit value and / or a preset lower limit value to limit the final calculation result during the process of calculating the filter coefficient. For example, when the calculated filter coefficient is greater than or equal to the preset upper limit value, the preset upper limit value is directly derived as the corresponding filter coefficient, or when the calculated filter coefficient is less than or equal to the preset lower limit value, the preset lower limit value is directly derived as the corresponding filter coefficient.

[0234] It can be understood that the piecewise function method or size limitation method proposed in the present application can not only be used in separable filters, but can also be applied to common two-dimensional filters, that is, two-dimensional filters that are inseparable horizontally and vertically. Specifically, for at least one coefficient of the filter, it can be calculated according to the piecewise function, or the size of the final calculation result can be limited according to a preset upper limit value and a preset lower limit value.

[0235] It should be noted that in the present application, the method of limiting the size of the filter coefficients can be understood as a more specific piecewise function. Therefore, when calculating the filter coefficients, the decoder can use both the piecewise function method and the size limitation method.

[0236] Furthermore, in the present application, if the second filtering parameter of the second filter is the filter phase, then when determining the second filtering parameter, the decoder can first determine a mapping table of the second phase and the motion vector, and then determine the filter phase corresponding to the pixel point based on the mapping table of the second phase and the motion vector and the second motion vector, that is, the second filtering parameter.

[0237] It should be noted that, in the present application, when the decoder determines the mapping table of the first phase and the motion vector or the mapping table of the second phase and the motion vector, it can be obtained through training or calculation.

[0238] That is to say, in the embodiment of the present application, the mapping table of the first phase and the motion vector or the mapping table of the second phase and the motion vector can be obtained by derivation through a certain calculation formula or by direct training. When calculating the mapping table of the phase and the motion vector, it can also be derived in the form of a piecewise function or by other more complex formulas. Of course, there can also be a filter phase that does not need to be derived from a formula but is trained.

[0239] In an embodiment of the present application, if the accuracy of the motion vector is high, or the accuracy of the motion vector used in filtering is high, or there are many possible values ​​of mv_x or mv_y, or the complexity of calculating the coefficients based on the motion vector is not high, then it is more reasonable to determine the filter parameters by calculating the filter coefficients through the motion vector and the proportional parameter. However, if the accuracy of the motion vector is not high, or the accuracy of the motion vector used in filtering is not high, or there are not many possible values ​​of mv_x or mv_y, or the complexity of calculating the coefficients based on the motion vector is high, then it is preferred to determine the filter parameters as the filter phase, that is, to make the separable filter into multiple phases, each phase having a set of fixed coefficients.

[0240] Exemplarily, if the motion vector accuracy during filtering is 1 / 16 pixel accuracy, the maximum value is 16 / 16 pixel, and the minimum value is -16 / 16 pixel, then there are a total of 33 possible values ​​of mv_x or mv_y. For example, the mapping table of phase and motion vector is shown in Table 5, where for the horizontal direction, mv corresponds to the motion vector difference mv_x of the pixel point in the horizontal direction, coef0 is the coefficient of the adjacent pixel point to the left of the pixel point, coef1 is the coefficient of the pixel point, and coef2 is the coefficient of the adjacent pixel point to the right of the pixel point. For the vertical direction, mv corresponds to the motion vector difference mv_y of the pixel point in the vertical direction, coef0 is the coefficient of the adjacent pixel point above the pixel point, coef1 is the coefficient of the pixel point, and coef2 is the coefficient of the adjacent pixel point below the pixel point.

[0241] Furthermore, in an embodiment of the present application, the decoder may determine the precision parameter of the motion vector, and determine the reduction parameter based on the precision parameter; thus, after filtering using the filter phase, reduction processing and / or right shift processing may be performed according to the reduction parameter.

[0242] It can be understood that in the present application, if the motion vector accuracy during filtering is 1 / 16 pixel accuracy, then the filter coefficients are magnified 256 times. Therefore, after completing the filtering process and obtaining the filtering result, the filtering result needs to be reduced by 256 times, or shifted right by 8 bits.

[0243] Furthermore, in the present application, reduction or right shifting can be performed separately after unidirectional filtering; or it can be performed after both horizontal and vertical filtering are completed; wherein, if reduction or right shifting is performed separately after unidirectional filtering, then it is necessary to adjust the reduction multiple or the number of right shift bits to ensure that the total reduction multiple or the number of right shift bits is consistent with the magnification multiple of the filter coefficient.

[0244] Table 5

[0245] mv coef0 coef1 coef2 -16 256 0 0 -15 240 16 0 -14 224 32 0 -13 208 48 0 -12 192 64 0 -11 176 80 0 -10 160 96 0 -9 144 112 0 -8 128 128 0 -7 105 158 -7 -6 84 184 -12 -5 65 206 -15 -4 48 224 -16 -3 33 238 -15 -2 20 248 -12 -1 9 254 -7 0 0 256 0 1 -7 254 9 2 -12 248 20 3 -15 238 33 4 -16 224 48 5 -15 206 65 6 -12 184 84 7 -7 158 105 8 0 128 128 9 0 112 144 10 0 96 160 11 0 80 176 12 0 64 192 13 0 48 208 14 0 32 224 15 0 16 240 16 0 0 256

[0246] This embodiment provides an inter-frame prediction method, which can use sub-block-based prediction for the current block to obtain a set of prediction values, namely, first prediction values, and use point-based prediction for the current block to obtain another set of prediction values, namely, second prediction values, and after determining the first weight of sub-block-based prediction and the second weight of point-based prediction for the same pixel point in the current block, use the first weight and the second weight to perform weighted averaging on the first prediction value and the second prediction value, and finally obtain a new prediction value for the current block, thereby improving the accuracy of inter-frame prediction while reducing the complexity of calculation and reducing bandwidth, thereby greatly improving encoding performance and improving encoding and decoding efficiency.

[0247] Fig.24 Schematic diagram of the implementation process of the inter-frame prediction method Figure 5 ,like Fig.24 As shown, the method for the encoder to perform inter-frame prediction may include the following steps:

[0248] Step 401: Determine prediction mode parameters of the current block.

[0249] In an embodiment of the present application, the encoder may first determine the prediction mode parameters of the current block. Specifically, the encoder may first determine the prediction mode used by the current block, and then determine the corresponding prediction mode parameters based on the prediction mode. The prediction mode parameters may be used to determine the prediction mode used by the current block.

[0250] It should be noted that, in the embodiments of the present application, the prediction mode parameters indicate the prediction mode used by the current block and the parameters related to the prediction mode. Here, for determining the prediction mode parameters, a simple decision strategy can be adopted, such as determining according to the size of the distortion value; or a complex decision strategy can be adopted, such as determining according to the result of rate distortion optimization (RDO), which is not limited in the embodiments of the present application. Generally speaking, the RDO method can be used to determine the prediction mode parameters of the current block.

[0251] Specifically, in some embodiments, when determining the prediction mode parameters of the current block, the encoder may first perform pre-encoding processing on the current block using multiple prediction modes to obtain rate-distortion cost values ​​corresponding to each prediction mode; then select the minimum rate-distortion cost value from the multiple rate-distortion cost values ​​obtained, and determine the prediction mode parameters of the current block according to the prediction mode corresponding to the minimum rate-distortion cost value.

[0252] That is to say, on the encoder side, multiple prediction modes can be used for the current block to perform pre-coding processing on the current block respectively. Here, the multiple prediction modes generally include inter-frame prediction mode, traditional intra-frame prediction mode and non-traditional intra-frame prediction mode; among which, the traditional intra-frame prediction mode may include direct current (DC) mode, plane (PLANAR) mode and angle mode, etc., the non-traditional intra-frame prediction mode may include matrix-based intra-frame prediction (Matrix-based Intra Prediction, MIP) mode, cross-component linear model prediction (Cross-component Linear Model Prediction, CCLM) mode, intra-frame block copy (Intra Block Copy, IBC) mode and PLT (Palette) mode, etc., and the inter-frame prediction mode may include ordinary inter-frame prediction mode, GPM mode and AWP mode, etc.

[0253] In this way, after pre-encoding the current block using multiple prediction modes, the rate-distortion cost value corresponding to each prediction mode can be obtained; then the minimum rate-distortion cost value is selected from the multiple rate-distortion cost values ​​obtained, and the prediction mode corresponding to the minimum rate-distortion cost value is determined as the prediction mode parameter of the current block. In addition, after pre-encoding the current block using multiple prediction modes, the distortion value corresponding to each prediction mode can be obtained; then the minimum distortion value is selected from the multiple distortion values ​​obtained, and then the prediction mode corresponding to the minimum distortion value is determined as the prediction mode used by the current block, and the corresponding prediction mode parameters are set according to the prediction mode. In this way, the current block is finally encoded using the determined prediction mode parameters, and in this prediction mode, the prediction residual can be made smaller, which can improve the coding efficiency.

[0254] That is to say, on the encoding side, the encoder can select the optimal prediction mode to pre-encode the current block. In this process, the prediction mode of the current block can be determined, and then the prediction mode parameters used to indicate the prediction mode can be determined, so that the corresponding prediction mode parameters are written into the bitstream and transmitted from the encoder to the decoder.

[0255] Accordingly, on the decoder side, the decoder can directly obtain the prediction mode parameters of the current block by parsing the bitstream, and determine the prediction mode used by the current block and the relevant parameters corresponding to the prediction mode according to the prediction mode parameters obtained by parsing.

[0256] Step 402: When the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, the first motion vector of each sub-block of the current block is used to determine the first prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point of the current block is used to determine the second prediction value of each pixel point; wherein the current block includes one or more sub-blocks; and the current block includes one or more pixels.

[0257] In the implementation of the present application, after the encoder determines the prediction mode parameters, if the prediction mode parameters indicate that the current block uses the inter-frame prediction mode to determine the inter-frame prediction value of the current block, then the encoder can use the first motion vector of each sub-block of the current block to determine the first prediction value of the pixel points in each sub-block. At the same time, the encoder can use the second motion vector of each pixel point of the current block to determine the second prediction value of each pixel point.

[0258] That is to say, in the present application, the encoder can use sub-block-based prediction for the current block to obtain a set of prediction values, namely, a first prediction value corresponding to a pixel point in each sub-block. At the same time, the encoder can also use point-based prediction for the current block to obtain another set of prediction values, namely, a second prediction value corresponding to each pixel point.

[0259] It can be understood that, in the present application, the first prediction value is the prediction value of the pixel point in the sub-block obtained by prediction based on the sub-block, and the second prediction value is the prediction value of the pixel point in the current block obtained by prediction based on the point. It can be seen that for the same pixel (pixel position) in the current block, the encoder can obtain the corresponding two prediction values ​​in different ways, namely, the first prediction value obtained by prediction based on the sub-block, and the second prediction value obtained by prediction based on the point.

[0260] It can be understood that, in the present application, the current block may include one or more sub-blocks; the current block may include one or more pixel points (samples), and each pixel point corresponds to a pixel position and a pixel value.

[0261] Exemplarily, in an embodiment of the present application, when determining the first prediction value of a pixel point in each sub-block, the encoder may first determine the first motion vector of each sub-block of the current block, and then determine the first prediction value corresponding to the pixel point in the sub-block based on the first motion vector.

[0262] That is to say, in the present application, when making predictions based on sub-blocks, the encoder can divide the current block into several sub-blocks and determine a motion vector for each sub-block, that is, each sub-block corresponds to a first motion vector, and then the encoder can use the first motion vector of each sub-block to predict the sub-block and obtain the first prediction value corresponding to the pixel point in the sub-block.

[0263] Exemplarily, in an embodiment of the present application, when determining the second prediction value of each pixel point, the encoder may first determine the second motion vector of each pixel point of the current block, and then determine the second prediction value corresponding to the pixel point based on the second motion vector.

[0264] That is to say, in the present application, the encoder can determine a motion vector for each pixel when predicting based on sub-blocks, that is, each pixel has a corresponding second motion vector, and then the encoder can use the second motion vector of each pixel to predict the pixel to obtain the corresponding second prediction value.

[0265] It is understandable that in the embodiment of the present application, the first motion vector of the sub-block and the second motion vector of the pixel of the current block can be derived using an affine model. Further, the encoder can also determine the first motion vector and the second motion vector in other ways, such as other prediction techniques based on sub-blocks. The encoder can also determine the second motion vector of the pixel in the sub-block according to the first motion vector of the sub-block, such as determining the second motion vector of any pixel in one of the sub-blocks according to the first motion vectors of two or more sub-blocks.

[0266] Further, in an embodiment of the present application, the encoder may use different prediction methods for sub-block-based prediction and point-based prediction. For example, the encoder may use different interpolation methods for sub-block-based prediction and point-based prediction. Specifically, in sub-block-based prediction and point-based prediction, the encoder may use different interpolation filters to interpolate pixel values.

[0267] Exemplarily, in the present application, a relatively complex filter may be used for sub-block-based prediction, and a relatively simple filter may be used for point-based prediction.

[0268] Exemplarily, in the present application, sub-block-based prediction may use a filter with a larger number of taps, and point-based prediction may use a filter with a smaller number of taps.

[0269] Exemplarily, in the present application, sub-block-based prediction and point-based prediction may use a separable two-dimensional filter or a non-separable two-dimensional filter, respectively.

[0270] It should be noted that in an embodiment of the present application, the separable filter can be a two-dimensional filter. Specifically, the separable filter can be a horizontally and vertically separable two-dimensional filter or a horizontally and vertically separable two-dimensional filter, and can be composed of a one-dimensional filter in the horizontal direction and a one-dimensional filter in the vertical direction.

[0271] That is to say, in the present application, if a separable filter is used, the separable filter can perform filtering processing in the horizontal direction and the vertical direction respectively.

[0272] It can be understood that in the present application, as a separable filter that is separable horizontally and vertically, the two-dimensional pixel points can be filtered separately in two directions. Specifically, the separable filter can first filter in one direction (such as horizontal direction or vertical direction) to obtain the intermediate value corresponding to the direction, and then filter the intermediate value in another direction (such as vertical direction or horizontal direction) to obtain the final filtering result.

[0273] It should be noted that in the present application, separable filters have been used in many common coding and decoding scenarios, such as inter-frame block-based prediction interpolation filtering, inter-frame sub-block-based prediction interpolation filtering, and affine sub-block-based prediction interpolation filtering. Furthermore, separable filters can also be applied to block- or sub-block-based motion compensation in HEVC, VVC, and AVS3, that is, the motion compensation prediction interpolation filter can include a separable two-dimensional filter. Exemplarily, the brightness interpolation filter of affine motion compensation prediction in AVS3 can be a separable two-dimensional filter, where the brightness interpolation filter coefficients are shown in Table 1 above.

[0274] Furthermore, in an embodiment of the present application, when performing sub-block-based prediction, for a sub-block of the current block, the encoder can determine a first filtering parameter based on a first motion vector; then, based on the first filtering parameter, a first filter can be used to perform filtering processing to obtain a first prediction value.

[0275] Specifically, in the present application, the first filter can be any one of the following filters: an n-tap interpolation filter, a separable two-dimensional filter, and an inseparable two-dimensional filter; wherein n is any one of the following values: 8, 6, 5, 4, 3, 2.

[0276] It should be noted that, in an embodiment of the present application, the first filtering parameter of the first filter may include the filter coefficient of the first filter, and may also include the filter phase of the first filter, which is not specifically limited in the present application.

[0277] Exemplarily, in the present application, if the first filtering parameter is the filter coefficient of the first filter, then when the encoder determines the first filtering parameter based on the first motion vector, it can first determine the first scale parameter; and then determine the first filtering parameter based on the first scale parameter and the first motion vector.

[0278] Exemplarily, in the present application, if the first filtering parameter is the filter phase of the first filter, then when the encoder determines the first filtering parameter based on the first motion vector, it can first determine a mapping table of the first phase and the motion vector; and then determine the first filtering parameter based on the mapping table of the first phase and the motion vector and the first motion vector.

[0279] Furthermore, in an embodiment of the present application, when performing point-based prediction, for a pixel point of the current block, the encoder can determine a second filtering parameter based on a second motion vector; then, based on the second filtering parameter, a second filter can be used to perform filtering processing to obtain a second prediction value.

[0280] Specifically, in the present application, the second filter can be any one of the following filters: an m-tap interpolation filter, a separable two-dimensional filter, and an inseparable two-dimensional filter; wherein m is any one of the following values: 8, 6, 5, 4, 3, 2.

[0281] It can be understood that in the present application, m can be less than or equal to n, that is, sub-block-based prediction can use a filter with a larger number of taps, and point-based prediction can use a filter with a smaller number of taps.

[0282] It should be noted that, in an embodiment of the present application, the second filtering parameter of the second filter may include the filter coefficient of the second filter, and may also include the filter phase of the second filter, which is not specifically limited in the present application.

[0283] Exemplarily, in the present application, if the second filtering parameter is the filter coefficient of the second filter, then when the encoder determines the second filtering parameter based on the second motion vector, it can first determine the second scale parameter; and then determine the second filtering parameter based on the second scale parameter and the second motion vector.

[0284] Exemplarily, in the present application, if the second filtering parameter is the filter phase of the second filter, then when the encoder determines the second filtering parameter based on the second motion vector, it can first determine a mapping table of the second phase and the motion vector; and then determine the second filtering parameter based on the mapping table of the second phase and the motion vector and the second motion vector.

[0285] That is to say, in the embodiment of the present application, the first filter based on the prediction of the sub-block and the second filter based on the prediction of the point can adopt different forms respectively. Specifically, the first filter coefficient of the first filter and the second filter coefficient of the second filter can be determined respectively in different ways. For example, the prediction based on the sub-block uses the form of the coefficient determined from the table according to the motion vector of the sub-pixel as shown in Table 1 above. The prediction based on the point can use the form of determining the coefficient from the table according to the motion vector of the sub-pixel, or the form of calculating the coefficient according to the motion vector of the sub-pixel.

[0286] Exemplarily, in an embodiment of the present application, when determining the first prediction value of a pixel point in each sub-block of the current block, the encoder may also first determine the motion vector deviation between each sub-block and the current block; and then determine the first prediction value of the pixel point in each sub-block based on the motion vector deviation.

[0287] It should be noted that, in the present application, the first proportional parameter or the second proportional parameter may include at least one proportional value, wherein the at least one proportional value is a non-zero real number.

[0288] Furthermore, in an embodiment of the present application, any one of the filter coefficients obtained by the encoder based on the motion vector calculation may be a linear function (polynomial), a quadratic function (polynomial) or a higher-order function (polynomial) of the motion vector, and the present application does not make any specific limitation thereto.

[0289] That is to say, in the present application, when the encoder calculates multiple filter coefficients corresponding to multiple pixel points according to different calculation methods in the preset calculation rules, some filter coefficients can be a linear function (polynomial) of the motion vector, that is, the two are in a linear relationship, or they can be a quadratic function (polynomial) or a higher-order function (polynomial) of the motion vector, that is, the two are in a non-linear relationship.

[0290] In some scenarios, even without calculating the first motion vector first, the first prediction value of the sub-block can be determined. For example, bidirectional optical flow, referred to as BIO in AVS and BDOF in VVC, uses bidirectional optical flow to calculate the motion vector deviation between the sub-block and the current block, and then uses the principle of optical flow to obtain a new prediction value based on the gradient, motion vector deviation and the current prediction value.

[0291] It should be noted that in the present application, in some scenarios, calculations are performed on a block (coding block or prediction block), such as a 4x4 block, as a whole, such as the calculation of motion vector deviation, correction of prediction value, etc., which can also be processed according to the idea of ​​sub-blocks in the present application.

[0292] It should be noted that in the embodiment of the present application, the current block is an image block to be encoded in the current frame, the current frame is encoded in the form of image blocks in a certain order, and the current block is the image block to be encoded at the next moment in the current frame in the order. The current block can have a variety of specifications, such as 1616, 3232 or 3216, where the numbers represent the number of rows and columns of pixels on the current block.

[0293] Furthermore, in an embodiment of the present application, the current block can be divided into a plurality of sub-blocks, wherein each sub-block has the same size, and the sub-block is a set of pixels of a smaller size. The size of the sub-block can be 8x8 or 4x4.

[0294] Exemplarily, in the present application, the size of the current block is 1616, which can be divided into 4 sub-blocks of 8x8 size each.

[0295] It can be understood that in the embodiment of the present application, when the encoder determines that the prediction mode parameter indicates the use of the inter-frame prediction mode to determine the inter-frame prediction value of the current block, the inter-frame prediction method provided in the embodiment of the present application can continue to be used.

[0296] Furthermore, in an embodiment of the present application, the encoder may use a sub-block of the current block as a reference block; then use the first motion vector of the reference block to determine the first prediction value of the pixel points in the reference block, and use the second motion vector of each pixel point of the reference block to determine the second prediction value of each pixel point; wherein the reference block includes one or more pixel points.

[0297] It should be noted that, in the embodiment of the present application, since the current block includes one or more sub-blocks, and at the same time, the current block includes one or more pixel points, therefore, if a pixel point belongs to a sub-block, then while prediction is performed based on the sub-block, the pixel point can also be predicted based on the pixel point using the pixel points in the sub-block. That is to say, when the encoder performs sub-block-based prediction and point-based prediction, the point-based prediction and the sub-block-based prediction can share the same reference block, wherein the reference block is a block composed of reference pixels required for interpolation filtering. Specifically, for a sub-block, each pixel point constituting the sub-block can be determined as a reference block, and the reference block used for the sub-block-based prediction of the sub-block and the reference block used for the point-based prediction of each point are all within this determined reference block. Preferably, the sub-block-based prediction and the point-based prediction of the corresponding point in the sub-block can be performed simultaneously, or the sub-block-based prediction and the point-based prediction of the corresponding point in the sub-block can be performed continuously, thereby avoiding repeated reading of reference pixels into the cache.

[0298] It is understandable that in the present application, if the point-based prediction and the sub-block-based prediction can share the same reference block, then the encoder can limit the MV of the point-based prediction so as to limit whether to increase the bandwidth. For example, the sub-block-based prediction uses an 8-tap filter that is separable horizontally and vertically, and the point-based prediction uses a 4-tap filter that is separable horizontally and vertically. Then, if the MV of the point in the upper left corner exceeds the MV of the sub-block by 2 pixels to the left or upward, the reference block required for the point will not exceed the reference block of the sub-block. Then, if the MV of the point in the upper right corner exceeds the MV of the sub-block by 2 pixels to the right or upward, the reference block required for the point will not exceed the reference block of the sub-block. Accordingly, the points inside the sub-block can move in a larger range.

[0299] Furthermore, the restriction on the MV of the point can be appropriately relaxed, and the performance and complexity can be balanced to appropriately increase the size of the reference block.

[0300] Step 403: for a pixel in the current block, determine a first weight and a second weight of the pixel; wherein the first weight corresponds to a first prediction value, and the second weight corresponds to a second prediction value.

[0301] In the implementation of the present application, for a pixel point in the current block, the encoder can respectively determine a first weight and a second weight of the pixel point; wherein the first weight corresponds to a first prediction value, and the second weight corresponds to a second prediction value. Specifically, the first weight and the second weight of the same pixel point can be the same or different.

[0302] It should be noted that in an embodiment of the present application, for a pixel point in the current block, the encoder can obtain a first prediction value and a second prediction value through sub-block-based prediction and point-based prediction, respectively. Correspondingly, the encoder can also set weight values ​​for these two prediction values ​​to determine a first weight corresponding to the first prediction value and a second weight corresponding to the second prediction value.

[0303] It can be understood that in the present application, for a pixel point in the current block, the first weight is the weight value corresponding to the encoder's prediction based on the sub-block, and the second weight is the weight value corresponding to the encoder's prediction based on the point.

[0304] Furthermore, in an embodiment of the present application, the encoder may use a variety of different methods to respectively determine the first weight and the second weight, wherein the encoder may use the same method to determine the first weight and the second weight, or may use different methods to determine the first weight and the second weight.

[0305] Exemplarily, in the present application, when the encoder determines the first weight corresponding to a sub-block-based prediction of a pixel in the current block, it can first determine the target sub-block corresponding to the pixel; wherein the target sub-block includes the pixel, that is, it is determined to which sub-block of multiple sub-blocks of the current block the pixel belongs; then the first distance between the pixel position and the reference position of the target sub-block can be determined; wherein the reference position is the position used by the first motion vector of the target sub-block; finally, the encoder can determine the first weight according to the first distance; wherein the first distance is inversely proportional to the first weight, that is, the larger the first distance, the farther the pixel position of the pixel is from the reference position, and the smaller the first weight corresponding to the pixel.

[0306] That is to say, in the present application, based on the prediction of the sub-block, the encoder gives a larger weight to the pixel position that is closer to the pixel position used by the first motion vector of the sub-block, and a smaller weight to the pixel position that is farther from the pixel position used by the first motion vector of the sub-block.

[0307] Exemplarily, in the present application, when the encoder determines the first weight corresponding to a sub-block-based prediction of a pixel in the current block, it can first determine the target sub-block corresponding to the pixel; wherein the target sub-block includes the pixel, that is, it is determined to which sub-block of multiple sub-blocks of the current block the pixel belongs; then the deviation value between the motion vector of the pixel and the first motion vector of the target sub-block can be determined; finally, the first weight can be further determined according to the deviation value; wherein the deviation value is inversely proportional to the first weight, that is, the smaller the deviation value, the closer the motion vector of the pixel is to the first motion vector of the target sub-block, and the larger the first weight corresponding to the pixel.

[0308] That is to say, in the present application, based on the prediction of the sub-block, the encoder gives a larger weight to the pixel point whose motion vector is closer to the first motion vector of the sub-block, and a smaller weight to the pixel point whose motion vector is far from the first motion vector of the sub-block.

[0309] Exemplarily, in the present application, when the encoder determines the second weight corresponding to the point-based prediction of a pixel in the current block, it can first determine the corresponding sub-pixel and integer pixel of the pixel point in the reference image; then it can determine the second distance between the sub-pixel and the integer pixel; finally, it can determine the second weight according to the second distance; wherein the second distance is inversely proportional to the second weight, that is, the larger the second distance, the greater the distance between the sub-pixel and the integer pixel, and the smaller the second weight corresponding to the pixel point.

[0310] That is to say, in the present application, based on point prediction, the encoder gives a larger weight to the sub-pixel position in the reference image corresponding to the pixel position of the pixel point that is closer to the integer pixel position, and a smaller weight to the sub-pixel position in the reference image corresponding to the pixel position of the pixel point that is farther from the integer pixel.

[0311] Exemplarily, in the present application, when the encoder determines the second weight corresponding to the point-based prediction of a pixel in the current block, it can first calculate the absolute value of the motion vector of the pixel; then the second weight can be further determined according to the absolute value; wherein the absolute value is inversely proportional to the second weight, that is, the smaller the absolute value, the larger the corresponding second weight.

[0312] That is to say, in the present application, based on point prediction, the encoder gives a larger weight to the smaller absolute value of the motion vector of the pixel point, and a smaller weight to the larger absolute value of the motion vector of the pixel point. For example, the second weight corresponding to the motion vector of the pixel point is (1 / 4, 1 / 4), which is greater than the second weight corresponding to the motion vector of (1 / 2, 1 / 2).

[0313] It can be understood that, in the present application, the encoder may use different methods or the same method when setting the first weight and the second weight.

[0314] Furthermore, in the present application, the setting of the first weight based on the prediction of the sub-block can use a method similar to the interlaced prediction, that is, based on the prediction of the sub-block, a larger first weight is set for the pixel position that is closer to the pixel position used by the first motion vector of the sub-block, and a smaller first weight is set for the pixel position that is farther from the pixel position used by the first motion vector of the sub-block.

[0315] Furthermore, in the present application, the second weight for point-based prediction may be set using a different setting method from the second weight for sub-block-based prediction. For example, the encoder may set a larger second weight for a pixel point whose pixel position corresponds to a sub-pixel position in a reference image that is closer to an integer pixel, and set a smaller second weight for a pixel point whose pixel position corresponds to a sub-pixel position in a reference image that is farther from an integer pixel. Or, a larger second weight is set for a pixel point whose motion vector is closer to an integer, and a smaller second weight is set for a pixel point whose motion vector is farther from an integer.

[0316] Illustratively, in the present application, the second weight of the pixel position where the motion vector of the pixel is (1 / 4, 1 / 4) is greater than the second weight of the pixel position where the motion vector of the pixel is (1 / 2, 1 / 2).

[0317] Furthermore, in the present application, if the encoder first determines the position of the second filter, that is, the encoder first determines the motion vector of the integer pixel and the motion vector of the sub-pixel corresponding to the current pixel, then when setting the second weight based on point prediction, the pixel position with a larger absolute value of the sub-pixel motion vector is set to a smaller second weight, and the pixel position with a smaller absolute value of the sub-pixel motion vector is set to a larger second weight.

[0318] Exemplarily, in the present application, the method of deriving the second weight w from the sub-pixel motion vector (x, y) is:

[0319] w=clip3(0,1,1-abs(x)-abs(y))

[0320] Where abs represents the absolute value, and clip3(a, b, c) means that if c is less than a, the result is a, if c is greater than b, the result is b, otherwise the result is c. Specifically, this is a normalized expression, where 1 in x and y represents 1 pixel. For the convenience of calculation, this formula can also be deformed, such as by shifting to remove fractions / decimals.

[0321] It can be understood that, in the present application, in order to match the first weight based on the prediction value of the sub-block, the second weight w can be multiplied by a multiple, etc.

[0322] It can be understood that the inter-frame prediction method proposed in the embodiment of the present application does not limit the order in which the encoder executes step 402 and step 403. That is to say, in the present application, the encoder can first execute step 402 and then step 403, or first execute step 403 and then step 402, or execute step 402 and step 403 at the same time.

[0323] Step 404: determine the predicted value of the pixel based on the first weight, the second weight, the first predicted value, and the second predicted value.

[0324] In an embodiment of the present application, after the encoder obtains a first prediction value and a corresponding first weight based on the prediction of the sub-block, and obtains a second prediction value and a corresponding second weight based on the prediction of the point, the encoder can use the first weight, the second weight, the first prediction value and the second prediction value to obtain a prediction value of a pixel point in the current block.

[0325] It can be understood that in an embodiment of the present application, after performing sub-block-based prediction and point-based prediction, for a pixel point in the current block, the encoder can use the first weight and the second weight to perform weighted averaging on the first prediction value and the second prediction value after determining the first weight and the second weight respectively corresponding to the pixel point, thereby obtaining the prediction value of the pixel point.

[0326] Further, in the present application, before determining the third prediction value of the current block according to the prediction value of the pixel point, that is, before step 405, the method for the encoder to perform inter-frame prediction may further include the following steps:

[0327] Step 406: For a pixel in the current block, determine a prediction value of the pixel based on the first prediction value and the second prediction value.

[0328] In an embodiment of the present application, after the encoder obtains a first prediction value based on sub-block prediction and a second prediction value based on point prediction, it can use the first prediction value and the second prediction value to obtain a prediction value of a pixel point in the current block.

[0329] It can be understood that in the embodiment of the present application, compared with step 404, the encoder may also directly average the first prediction value and the second prediction value of the same pixel point after respectively obtaining the first prediction value and the second prediction value to obtain the prediction value of the pixel point.

[0330] That is to say, in the present application, the encoder may not perform a weighted average operation on the first prediction value and the second prediction value, but directly use the two prediction values ​​to determine the final prediction result.

[0331] It should be noted that, in the present application, the method proposed in step 406 can also be used as one of the cases of step 404, that is, the case where the first weight and the second weight are equal.

[0332] Step 405: Determine a third prediction value of the current block according to the prediction value of the pixel point; wherein the third prediction value is used to determine the residual of the current block.

[0333] In an embodiment of the present application, after the encoder determines the prediction value of the pixel based on the first weight, the second weight, the first prediction value and the second prediction value, it can further determine the third prediction value of the current block according to the prediction value of the pixel.

[0334] It should be noted that in the present application, the encoder can filter all the pixels in the current block according to the above method, and use the prediction values ​​of all the pixels in the current block to determine the third prediction value of the current block, or filter some of the pixels in the current block, and use the prediction values ​​of some of the pixels in the current block to determine the third prediction value of the sub-block.

[0335] It can be understood that, in the embodiment of the present application, the third prediction value is used to determine the residual of the current block.

[0336] Furthermore, in an embodiment of the present application, the encoder may, after traversing all or part of the pixels in the current block, perform an addition operation on the predicted values ​​of all or part of the pixels in the current block to obtain a sum result, and finally normalize the addition result to finally obtain a third predicted value of the current block.

[0337] In the embodiment of the present application, further, the inter-frame prediction method proposed in the above steps 401 to 406 can be applied to both unidirectional prediction and bidirectional prediction.

[0338] Furthermore, in an embodiment of the present application, the method for the encoder to perform inter-frame prediction may further include the following steps:

[0339] Step 501: Determine the inter-frame prediction direction.

[0340] Step 502: If the inter-frame prediction direction is bidirectional prediction, based on the first prediction direction, the first motion vector of each sub-block is used to determine the fourth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the fifth prediction value of each pixel point; based on the second prediction direction, the first motion vector of each sub-block is used to determine the sixth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the seventh prediction value of each pixel point.

[0341] Step 503 : For a pixel of the current block, determine a third weight and a fourth weight of the pixel based on the first prediction direction; and determine a fifth weight and a sixth weight of the pixel based on the second prediction direction.

[0342] Step 504: Based on the third weight, the fourth weight, the fourth prediction value and the fifth prediction value, obtain the prediction value of the pixel point corresponding to the first prediction direction; based on the fifth weight, the sixth weight, the sixth prediction value and the seventh prediction value, obtain the prediction value of the pixel point corresponding to the second prediction direction.

[0343] Step 505: Determine the prediction value of the pixel point according to the prediction value corresponding to the first prediction direction and the prediction value corresponding to the second prediction direction.

[0344] That is to say, in the present application, when the encoder performs bidirectional prediction, for each unidirectional prediction, a set of prediction values ​​is obtained by using sub-block-based prediction for the current block, and another set of prediction values ​​is obtained by using point-based prediction for the current block. The two prediction values ​​for the same pixel are weighted averaged to obtain the prediction value of the pixel. Then, the prediction values ​​of the two unidirectional predictions are averaged (or weighted averaged) to finally obtain the prediction value of the bidirectional prediction.

[0345] It should be noted that in the present application, the averaging operation or weighted averaging operation in the inter-frame prediction processing process can be performed after the entire prediction block that needs to be averaged or weighted averaged is obtained, or it can be performed after the prediction blocks of the sub-blocks that need to be averaged or weighted averaged are obtained, or it can be performed after the prediction blocks of the points that need to be averaged or weighted averaged are obtained, and the present application does not make specific limitations.

[0346] Furthermore, in an embodiment of the present application, the method for the encoder to perform inter-frame prediction may further include the following steps:

[0347] Step 501: Determine the inter-frame prediction direction.

[0348] Step 502: If the inter-frame prediction direction is bidirectional prediction, based on the first prediction direction, the first motion vector of each sub-block is used to determine the fourth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the fifth prediction value of each pixel point; based on the second prediction direction, the first motion vector of each sub-block is used to determine the sixth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the seventh prediction value of each pixel point.

[0349] Step 503 : For a pixel of the current block, determine a third weight and a fourth weight of the pixel based on the first prediction direction; and determine a fifth weight and a sixth weight of the pixel based on the second prediction direction.

[0350] Step 506: Determine the predicted value of the pixel based on the third weight, the fourth weight, the fifth weight, the sixth weight, the fourth predicted value, the fifth predicted value, the sixth predicted value and the seventh predicted value.

[0351] That is to say, in the present application, when the encoder performs bidirectional prediction, for each unidirectional prediction, a set of prediction values ​​is obtained for the current block using sub-block-based prediction, and another set of prediction values ​​is obtained for the current block using point-based prediction. The four prediction values ​​of the two unidirectional predictions are then averaged (or weighted averaged), and finally the prediction value of the bidirectional prediction can be obtained.

[0352] It can be seen that for bidirectional prediction, the encoder can determine the respective prediction values ​​of the two prediction directions and then perform averaging or weighted averaging operations, or it may not calculate the prediction value of each prediction direction separately, but directly obtain it by weighted averaging the prediction block based on sub-blocks in the first prediction direction, the prediction block based on points in the first prediction direction, the prediction block based on sub-blocks in the second prediction direction, and the prediction block based on points in the second prediction direction.

[0353] It is understandable that, in the present application, the averaging operation can be considered as a special weighted averaging operation, and the pixel point pixel position can be understood as a different expression of the same concept in some scenes.

[0354] Furthermore, the inter-frame prediction method proposed in the present application can be used only for the luminance component, or for the luminance component and the chrominance component, or for any other format, such as one, several, or all components of RGB, etc. The embodiment of the present application is described by taking the luminance component as an example, but is not limited to the luminance component.

[0355] That is to say, the inter-frame prediction method proposed in the present application can be applied to any image component. In the present embodiment, the prediction scheme is exemplarily used for the luminance component, but it can also be used for the chrominance component or any component in other formats. The inter-frame prediction method proposed in the present application can also be applied to any video format, including but not limited to the YUV format, including but not limited to the luminance component in the YUV format.

[0356] This embodiment provides an inter-frame prediction method, which can use sub-block-based prediction for the current block to obtain a set of prediction values, namely, first prediction values, and use point-based prediction for the current block to obtain another set of prediction values, namely, second prediction values, and after determining the first weight of sub-block-based prediction and the second weight of point-based prediction for the same pixel point in the current block, use the first weight and the second weight to perform weighted averaging on the first prediction value and the second prediction value, and finally obtain a new prediction value for the current block, thereby improving the accuracy of inter-frame prediction while reducing the complexity of calculation and reducing bandwidth, thereby greatly improving encoding performance and improving encoding and decoding efficiency.

[0357] Based on the above embodiment, in yet another embodiment of the present application, Fig.25 The structure of the decoder is shown in Figure 2. Figure 1 ,like Fig.25 As shown, the decoder 300 proposed in the embodiment of the present application may include a parsing part 301 and a first determining part 302;

[0358] The parsing part 301 is configured to parse the bitstream and obtain the prediction mode parameters of the current block;

[0359] The first determination part 302 is configured to, when the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, use the first motion vector of each sub-block of the current block to determine the first prediction value of the pixel points in each sub-block, and use the second motion vector of each pixel point of the current block to determine the second prediction value of each pixel point; wherein the current block includes one or more sub-blocks; the current block includes one or more pixels; for a pixel point in the current block, determine the first weight and the second weight of the pixel point; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; based on the first weight, the second weight, the first prediction value and the second prediction value, determine the prediction value of the pixel point; according to the prediction value of the pixel point, determine the third prediction value of the current block; wherein the third prediction value is used to determine the reconstruction value of the current block.

[0360] Fig.26 The structure of the decoder is shown in Figure 2. Figure 2 ,like Fig.26 As shown, the decoder 300 proposed in the embodiment of the present application may also include a first processor 303, a first memory 304 storing executable instructions of the first processor 303, a first communication interface 305, and a first bus 306 for connecting the first processor 303, the first memory 304 and the first communication interface 305.

[0361] Further, in an embodiment of the present application, the above-mentioned first processor 303 is used to parse the code stream and obtain the prediction mode parameters of the current block; when the prediction mode parameters indicate that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, the first prediction value of the pixel in each sub-block is determined using the first motion vector of each sub-block of the current block, and the second prediction value of each pixel in the current block is determined using the second motion vector; wherein the current block includes one or more sub-blocks; the current block includes one or more pixels; for a pixel in the current block, a first weight and a second weight of the pixel are determined; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; based on the first weight, the second weight, the first prediction value and the second prediction value, the prediction value of the pixel is determined; according to the prediction value of the pixel, a third prediction value of the current block is determined; wherein the third prediction value is used to determine the reconstruction value of the current block.

[0362] Fig. 27 The structure of the encoder is shown in Figure 2. Figure 1 ,like Fig. 27As shown, the encoder 400 proposed in the embodiment of the present application may include a second determining part 401;

[0363] The second determination part 401 is configured to determine the prediction mode parameters of the current block; when the prediction mode parameters indicate that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, the first prediction value of the pixel points in each sub-block is determined using the first motion vector of each sub-block of the current block, and the second prediction value of each pixel point of the current block is determined using the second motion vector of each pixel point of the current block; wherein the current block includes one or more sub-blocks; the current block includes one or more pixels; for a pixel point in the current block, a first weight and a second weight of the pixel point are determined; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; based on the first weight, the second weight, the first prediction value and the second prediction value, the prediction value of the pixel point is determined; according to the prediction value of the pixel point, a third prediction value of the current block is determined; wherein the third prediction value is used to determine the residual of the current block.

[0364] Fig.28 The structure of the encoder is shown in Figure 2. Figure 2 ,like Fig.28 As shown, the encoder 400 proposed in the embodiment of the present application may also include a second processor 402, a second memory 403 storing executable instructions of the second processor 402, a second communication interface 404, and a second bus 405 for connecting the second processor 402, the second memory 403 and the second communication interface 404.

[0365] Further, in an embodiment of the present application, the above-mentioned second processor 402 is configured to determine a prediction mode parameter of the current block; when the prediction mode parameter indicates that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, the first prediction value of the pixel in each sub-block is determined using the first motion vector of each sub-block of the current block, and the second prediction value of each pixel in the current block is determined using the second motion vector of each pixel; wherein the current block includes one or more sub-blocks; the current block includes one or more pixels; for a pixel in the current block, a first weight and a second weight of the pixel are determined; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; based on the first weight, the second weight, the first prediction value and the second prediction value, the prediction value of the pixel is determined; according to the prediction value of the pixel, a third prediction value of the current block is determined; wherein the third prediction value is used to determine the residual of the current block.

[0366] The embodiments of the present application provide an encoder and an encoder, which can use sub-block-based prediction for the current block to obtain a set of prediction values, namely, first prediction values, and use point-based prediction for the current block to obtain another set of prediction values, namely, second prediction values, and after determining the first weight of sub-block-based prediction and the second weight of point-based prediction for the same pixel point in the current block, use the first weight and the second weight to perform weighted averaging on the first prediction value and the second prediction value, and finally obtain a new prediction value for the current block, thereby improving the accuracy of inter-frame prediction while reducing the complexity of calculation and reducing bandwidth, thereby greatly improving the encoding performance and improving the encoding and decoding efficiency.

[0367] The embodiments of the present application provide a computer-readable storage medium and a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the method described in the above embodiments is implemented.

[0368] Specifically, the program instructions corresponding to an inter-frame prediction method in this embodiment may be stored in a storage medium such as an optical disk, a hard disk, or a USB flash drive. When the program instructions corresponding to an inter-frame prediction method in the storage medium are read or executed by an electronic device, the following steps are included:

[0369] Parse the bitstream and obtain the prediction mode parameters of the current block;

[0370] When the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, a first prediction value of a pixel in each sub-block is determined by using a first motion vector of each sub-block of the current block, and a second prediction value of each pixel in the current block is determined by using a second motion vector of each pixel; wherein the current block includes one or more sub-blocks; and the current block includes one or more pixels;

[0371] For a pixel point in the current block, determine a first weight and a second weight of the pixel point; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value;

[0372] Determine a predicted value of the pixel based on the first weight, the second weight, the first predicted value, and the second predicted value;

[0373] A third prediction value of the current block is determined according to the prediction value of the pixel point; wherein the third prediction value is used to determine a reconstruction value of the current block.

[0374] Specifically, the program instructions corresponding to an inter-frame prediction method in this embodiment may be stored in a storage medium such as an optical disk, a hard disk, or a USB flash drive. When the program instructions corresponding to an inter-frame prediction method in the storage medium are read or executed by an electronic device, the following steps are included:

[0375] Determining prediction mode parameters for the current block;

[0376] When the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, a first prediction value of a pixel in each sub-block is determined by using a first motion vector of each sub-block of the current block, and a second prediction value of each pixel in the current block is determined by using a second motion vector of each pixel; wherein the current block includes one or more sub-blocks; and the current block includes one or more pixels;

[0377] For a pixel point in the current block, determine a first weight and a second weight of the pixel point; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value;

[0378] Determine a predicted value of the pixel based on the first weight, the second weight, the first predicted value, and the second predicted value;

[0379] Determine a third prediction value of the current block according to the prediction value of the pixel point; wherein the third prediction value is used to determine the residual of the current block.

[0380] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of hardware embodiments, software embodiments, or embodiments in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0381] The present application is described with reference to implementation flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the process in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0382] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which is implemented in the implementation flow diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0383] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing the steps in the flowchart. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0384] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain a new method embodiment. The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain a new product embodiment. The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain a new method embodiment or device embodiment.

[0385] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0386] Industrial Applicability

[0387] The embodiments of the present application provide an inter-frame prediction method, an encoder, a decoder and a computer storage medium. The decoder parses a bit stream to obtain prediction mode parameters of a current block. When the prediction mode parameters indicate that an inter-frame prediction value of the current block is determined using an inter-frame prediction mode, a first prediction value of a pixel in each sub-block is determined using a first motion vector of each sub-block of the current block, and a second prediction value of each pixel in the current block is determined using a second motion vector of each pixel. The current block includes one or more sub-blocks. The current block includes one or more pixels. For a pixel in the current block, a first weight and a second weight of the pixel are determined. The first weight corresponds to a first prediction value, and the second weight corresponds to a second prediction value. The prediction value of the pixel is determined based on the first weight, the second weight, the first prediction value and the second prediction value. The third prediction value of the current block is determined based on the prediction value of the pixel. The third prediction value is used to determine a reconstruction value of the current block. That is to say, the inter-frame prediction method proposed in the present application can use sub-block-based prediction for the current block to obtain a set of prediction values, namely, the first prediction value, and use point-based prediction for the current block to obtain another set of prediction values, namely, the second prediction value, and after determining the first weight of the sub-block-based prediction and the second weight of the point-based prediction for the same pixel point in the current block, use the first weight and the second weight to perform weighted averaging on the first prediction value and the second prediction value, and finally obtain a new prediction value for the current block, thereby improving the accuracy of the inter-frame prediction while reducing the complexity of the calculation and reducing the bandwidth, thereby greatly improving the encoding performance and improving the encoding and decoding efficiency.

Claims

1. An inter-frame prediction method, applied to a decoder, the method comprising: Parse the bitstream and obtain the prediction mode parameters of the current block; When the prediction mode parameter indicates that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, a first prediction value of a pixel in each sub-block is determined using a first motion vector of each sub-block of the current block, and a second prediction value of each pixel in the current block is determined using a second motion vector of each pixel; wherein the current block includes one or more sub-blocks; the current block includes one or more pixels; the first prediction value is obtained based on a first filter, and the second prediction value is obtained based on a second filter; the first filter is different from the second filter; For a pixel point in the current block, determine a first weight and a second weight of the pixel point; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; Determine a predicted value of the pixel based on the first weight, the second weight, the first predicted value, and the second predicted value; A third prediction value of the current block is determined according to the prediction value of the pixel point; wherein the third prediction value is used to determine a reconstruction value of the current block.

2. The method according to claim 1, wherein: Before determining the third prediction value of the current block according to the prediction value of the pixel point, the method further includes: Parsing the bitstream to determine the inter-frame prediction direction; If the inter-frame prediction direction is bidirectional prediction, based on the first prediction direction, the first motion vector of each sub-block is used to determine the fourth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the fifth prediction value of each pixel point; based on the second prediction direction, the first motion vector of each sub-block is used to determine the sixth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the seventh prediction value of each pixel point; For a pixel of the current block, determining a third weight and a fourth weight of the pixel based on the first prediction direction; and determining a fifth weight and a sixth weight of the pixel based on the second prediction direction; Based on the third weight, the fourth weight, the fourth prediction value, and the fifth prediction value, obtaining a prediction value of the pixel corresponding to the first prediction direction; based on the fifth weight, the sixth weight, the sixth prediction value, and the seventh prediction value, obtaining a prediction value of the pixel corresponding to the second prediction direction; Determine the prediction value of the pixel point according to the prediction value corresponding to the first prediction direction and the prediction value corresponding to the second prediction direction.

3. The method according to claim 1, wherein: Before determining the third prediction value of the current block according to the prediction value of the pixel point, the method further includes: Parsing the bitstream to determine the inter-frame prediction direction; If the inter-frame prediction direction is bidirectional prediction, based on the first prediction direction, the first motion vector of each sub-block is used to determine the fourth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the fifth prediction value of each pixel point; based on the second prediction direction, the first motion vector of each sub-block is used to determine the sixth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the seventh prediction value of each pixel point; For a pixel of the current block, determining a third weight and a fourth weight of the pixel based on the first prediction direction; and determining a fifth weight and a sixth weight of the pixel based on the second prediction direction; Based on the third weight, the fourth weight, the fifth weight, the sixth weight, the fourth prediction value, the fifth prediction value, the sixth prediction value and the seventh prediction value, the prediction value of the pixel point is determined.

4. The method according to claim 1, wherein: The determining a first prediction value of a pixel point in each sub-block by using a first motion vector of each sub-block of the current block comprises: For a subblock of the current block, determining a first filtering parameter according to the first motion vector; Based on the first filtering parameter, the first filter is used to perform filtering processing to obtain the first prediction value.

5. The method according to claim 4, wherein: The first filter may be any one of the following filters: an n-tap interpolation filter, a separable two-dimensional filter, or a non-separable two-dimensional filter; wherein n is any one of the following values: 8, 6, 5, 4, 3, or 2.

6. The method according to claim 4, wherein: The determining a first filtering parameter according to the first motion vector comprises: determining a first scale parameter; The first filtering parameter is determined according to the first scale parameter and the first motion vector; wherein the first filtering parameter is a linear function, a quadratic function, or a higher-order function of the first motion vector.

7. The method according to claim 4, wherein: The determining a first filtering parameter according to the first motion vector comprises: Determine a mapping table of a first phase and a motion vector; The first filtering parameter is determined according to the first phase-motion vector mapping table and the first motion vector.

8. The method according to claim 1, wherein: The determining the second prediction value of each pixel point by using the second motion vector of each pixel point of the current block comprises: For a pixel point of the current block, determining a second filtering parameter according to the second motion vector; Based on the second filtering parameter, the second filter is used to perform filtering processing to obtain the second prediction value.

9. The method according to claim 8, wherein: The second filter may be any one of the following filters: an m-tap interpolation filter, a separable two-dimensional filter, or a non-separable two-dimensional filter, wherein m is any one of the following values: 8, 6, 5, 4, 3, or 2.

10. The method according to claim 8, wherein: The determining a first filtering parameter according to the first motion vector comprises: determining a second scale parameter; The second filtering parameter is determined according to the second scale parameter and the second motion vector; wherein the second filtering parameter is a linear function, a quadratic function, or a higher-order function of the second motion vector.

11. The method according to claim 8, wherein: The determining a second filtering parameter according to the second motion vector comprises: Determine a mapping table of a second phase and a motion vector; The second filtering parameter is determined according to the second phase-motion vector mapping table and the second motion vector.

12. The method according to claim 1, wherein: The step of determining, for a pixel point in the current block, a first weight and a second weight of the pixel point comprises: Determine a target sub-block corresponding to the pixel point; Determine a first distance between a pixel position of the pixel point and a reference position of the target sub-block; the reference position is a position used by the first motion vector of the target sub-block; The first weight is determined according to the first distance; wherein the first distance is inversely proportional to the first weight.

13. The method according to claim 1, wherein: The step of determining, for a pixel point in the current block, a first weight and a second weight of the pixel point comprises: Determine a target sub-block corresponding to the pixel point; Determine a deviation value between the motion vector of the pixel and the first motion vector of the target sub-block; The first weight is determined according to the deviation value; wherein the deviation value is inversely proportional to the first weight.

14. The method according to claim 1, wherein: The step of determining, for a pixel point in the current block, a first weight and a second weight of the pixel point comprises: Determine the sub-pixel and the whole pixel corresponding to the pixel point in the reference image; determining a second distance between the sub-pixel and the whole pixel; The second weight is determined according to the second distance; wherein the second distance is inversely proportional to the second weight.

15. The method according to claim 1, wherein: The step of determining, for a pixel point in the current block, a first weight and a second weight of the pixel point comprises: Determining the absolute value of the motion vector of the pixel; The second weight is determined according to the absolute value; wherein the absolute value is inversely proportional to the second weight.

16. The method according to any one of claims 4 to 11, wherein: Before determining the third prediction value of the current block according to the prediction value of the pixel point, the method further includes: For a pixel in the current block, a prediction value of the pixel is determined based on the first prediction value and the second prediction value.

17. The method according to claim 1, wherein: The method further comprises: Using a sub-block of the current block as a reference block; The first motion vector of the reference block is used to determine the first prediction value of the pixel points in the reference block, and the second motion vector of each pixel point of the reference block is used to determine the second prediction value of each pixel point; wherein the reference block includes one or more pixel points.

18. The method according to claim 9, wherein: If the second filter is a horizontally and vertically separable 3-tap filter, the method further includes: The sub-pixel motion vector corresponding to each pixel point belongs to the interval from -1 / 2 pixel to 1 / 2 pixel; or, The sub-pixel motion vector corresponding to each pixel point belongs to the interval from -1 pixel to 1 pixel.

19. The method according to claim 9, wherein: If the second filter is a horizontally and vertically separable 3-tap filter, the method further includes: Determine in the reference image a reference pixel point having the same position as each of the pixel points; The reference pixel point is determined as the center of the second filter.

20. An inter-frame prediction method, applied to an encoder, the method comprising: Determining prediction mode parameters for the current block; When the prediction mode parameter indicates that the inter-frame prediction value of the current block is determined using the inter-frame prediction mode, a first prediction value of a pixel in each sub-block is determined using a first motion vector of each sub-block of the current block, and a second prediction value of each pixel in the current block is determined using a second motion vector of each pixel; wherein the current block includes one or more sub-blocks; the current block includes one or more pixels; the first prediction value is obtained based on a first filter, and the second prediction value is obtained based on a second filter; the first filter is different from the second filter; For a pixel point in the current block, determine a first weight and a second weight of the pixel point; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; Determine a predicted value of the pixel based on the first weight, the second weight, the first predicted value, and the second predicted value; Determine a third prediction value of the current block according to the prediction value of the pixel point; wherein the third prediction value is used to determine the residual of the current block.

21. The method according to claim 20, wherein: Before determining the third prediction value of the current block according to the prediction value of the pixel point, the method further includes: Determine the inter-frame prediction direction; If the inter-frame prediction direction is bidirectional prediction, based on the first prediction direction, the first motion vector of each sub-block is used to determine the fourth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the fifth prediction value of each pixel point; based on the second prediction direction, the first motion vector of each sub-block is used to determine the sixth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the seventh prediction value of each pixel point; For a pixel of the current block, determining a third weight and a fourth weight of the pixel based on the first prediction direction; and determining a fifth weight and a sixth weight of the pixel based on the second prediction direction; Based on the third weight, the fourth weight, the fourth prediction value, and the fifth prediction value, obtaining a prediction value of the pixel corresponding to the first prediction direction; based on the fifth weight, the sixth weight, the sixth prediction value, and the seventh prediction value, obtaining a prediction value of the pixel corresponding to the second prediction direction; Determine the prediction value of the pixel point according to the prediction value corresponding to the first prediction direction and the prediction value corresponding to the second prediction direction.

22. The method according to claim 20, wherein: Before determining the third prediction value of the current block according to the prediction value of the pixel point, the method further includes: Determine the inter-frame prediction direction; If the inter-frame prediction direction is bidirectional prediction, based on the first prediction direction, the first motion vector of each sub-block is used to determine the fourth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the fifth prediction value of each pixel point; based on the second prediction direction, the first motion vector of each sub-block is used to determine the sixth prediction value of the pixel point in each sub-block, and the second motion vector of each pixel point is used to determine the seventh prediction value of each pixel point; For a pixel of the current block, determining a third weight and a fourth weight of the pixel based on the first prediction direction; and determining a fifth weight and a sixth weight of the pixel based on the second prediction direction; Based on the third weight, the fourth weight, the fifth weight, the sixth weight, the fourth prediction value, the fifth prediction value, the sixth prediction value and the seventh prediction value, the prediction value of the pixel point is determined.

23. The method according to claim 20, wherein: The determining a first prediction value of a pixel point in each sub-block by using a first motion vector of each sub-block of the current block comprises: For a subblock of the current block, determining a first filtering parameter according to the first motion vector; Based on the first filtering parameter, the first filter is used to perform filtering processing to obtain the first prediction value.

24. The method according to claim 23, wherein: The first filter may be any one of the following filters: an n-tap interpolation filter, a separable two-dimensional filter, or a non-separable two-dimensional filter; wherein n is any one of the following values: 8, 6, 5, 4, 3, or 2.

25. The method according to claim 23, wherein: The determining a first filtering parameter according to the first motion vector comprises: determining a first scale parameter; The first filtering parameter is determined according to the first scale parameter and the first motion vector; wherein the first filtering parameter is a linear function, a quadratic function, or a higher-order function of the first motion vector.

26. The method of claim 23, wherein: The determining a first filtering parameter according to the first motion vector comprises: Determine a mapping table of a first phase and a motion vector; The first filtering parameter is determined according to the first phase-motion vector mapping table and the first motion vector.

27. The method according to claim 20, wherein: The determining the second prediction value of each pixel point by using the second motion vector of each pixel point of the current block comprises: For a pixel point of the current block, determining a second filtering parameter according to the second motion vector; Based on the second filtering parameter, the second filter is used to perform filtering processing to obtain the second prediction value.

28. The method according to claim 27, wherein: The second filter may be any one of the following filters: an m-tap interpolation filter, a separable two-dimensional filter, or a non-separable two-dimensional filter, wherein m is any one of the following values: 8, 6, 5, 4, 3, or 2.

29. The method according to claim 27, wherein: The determining a first filtering parameter according to the first motion vector comprises: determining a second scale parameter; The second filtering parameter is determined according to the second scale parameter and the second motion vector; wherein the second filtering parameter is a linear function, a quadratic function, or a higher-order function of the second motion vector.

30. The method of claim 27, wherein: The determining a second filtering parameter according to the second motion vector comprises: Determine a mapping table of a second phase and a motion vector; The second filtering parameter is determined according to the second phase-motion vector mapping table and the second motion vector.

31. The method of claim 20, wherein: The step of determining, for a pixel point in the current block, a first weight and a second weight of the pixel point comprises: Determine a target sub-block corresponding to the pixel point; Determine a first distance between a pixel position of the pixel point and a reference position of the target sub-block; the reference position is a position used by the first motion vector of the target sub-block; The first weight is determined according to the first distance; wherein the first distance is inversely proportional to the first weight.

32. The method of claim 20, wherein: The step of determining, for a pixel point in the current block, a first weight and a second weight of the pixel point comprises: Determine a target sub-block corresponding to the pixel point; Determine a deviation value between the motion vector of the pixel and the first motion vector of the target sub-block; The first weight is determined according to the deviation value; wherein the deviation value is inversely proportional to the first weight.

33. The method of claim 20, wherein: The step of determining, for a pixel point in the current block, a first weight and a second weight of the pixel point comprises: Determine the sub-pixel and the whole pixel corresponding to the pixel point in the reference image; determining a second distance between the sub-pixel and the whole pixel; The second weight is determined according to the second distance; wherein the second distance is inversely proportional to the second weight.

34. The method of claim 20, wherein: The step of determining, for a pixel point in the current block, a first weight and a second weight of the pixel point comprises: Determining the absolute value of the motion vector of the pixel; The second weight is determined according to the absolute value; wherein the absolute value is inversely proportional to the second weight.

35. The method according to any one of claims 23 to 30, wherein: Before determining the third prediction value of the current block according to the prediction value of the pixel point, the method further includes: For a pixel in the current block, a prediction value of the pixel is determined based on the first prediction value and the second prediction value.

36. The method of claim 20, wherein: The method further comprises: Using a sub-block of the current block as a reference block; The first motion vector of the reference block is used to determine the first prediction value of the pixel points in the reference block, and the second motion vector of each pixel point of the reference block is used to determine the second prediction value of each pixel point; wherein the reference block includes one or more pixel points.

37. The method of claim 28, wherein: If the second filter is a horizontally and vertically separable 3-tap filter, the method further includes: The sub-pixel motion vector corresponding to each pixel point belongs to the interval from -1 / 2 pixel to 1 / 2 pixel; or, The sub-pixel motion vector corresponding to each pixel point belongs to the interval from -1 pixel to 1 pixel.

38. The method of claim 28, wherein: If the second filter is a horizontally and vertically separable 3-tap filter, the method further includes: Determine in the reference image a reference pixel point having the same position as each of the pixel points; The reference pixel point is determined as the center of the second filter.

39. A decoder, the decoder comprising a parsing part, a first determining part; The parsing part is configured to parse the bitstream and obtain the prediction mode parameters of the current block; The first determining part is configured to determine the first prediction value of the pixel points in each sub-block by using the first motion vector of each sub-block of the current block, and determine the second prediction value of each pixel point by using the second motion vector of each pixel point of the current block when the prediction mode parameter indicates that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block; wherein, The current block includes one or more sub-blocks; the current block includes one or more pixels; the first prediction value is obtained based on a first filter, and the second prediction value is obtained based on a second filter; the first filter is different from the second filter; for a pixel in the current block, determine a first weight and a second weight of the pixel; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; based on the first weight, the second weight, the first prediction value and the second prediction value, determine the prediction value of the pixel; according to the prediction value of the pixel, determine a third prediction value of the current block; wherein the third prediction value is used to determine a reconstruction value of the current block.

40. A decoder, comprising a first processor and a first memory storing instructions executable by the first processor, wherein when the instructions are executed, the first processor implements the method according to any one of claims 1 to 19.

41. An encoder comprising a second determining portion; The second determination part is configured to determine the prediction mode parameters of the current block; when the prediction mode parameters indicate that the inter-frame prediction mode is used to determine the inter-frame prediction value of the current block, the first prediction value of the pixel points in each sub-block of the current block is determined by using the first motion vector of each sub-block of the current block, and the second prediction value of each pixel point of the current block is determined by using the second motion vector of each pixel point of the current block; wherein, The current block includes one or more sub-blocks; the current block includes one or more pixels; the first prediction value is obtained based on a first filter, and the second prediction value is obtained based on a second filter; the first filter is different from the second filter; for a pixel in the current block, determine a first weight and a second weight of the pixel; wherein the first weight corresponds to the first prediction value, and the second weight corresponds to the second prediction value; based on the first weight, the second weight, the first prediction value and the second prediction value, determine the prediction value of the pixel; according to the prediction value of the pixel, determine a third prediction value of the current block; wherein the third prediction value is used to determine the residual of the current block.

42. An encoder, comprising a second processor and a second memory storing instructions executable by the second processor, wherein when the instructions are executed, the second processor implements the method according to any one of claims 20 to 38.

43. A computer storage medium storing a computer program, wherein the computer program implements the method according to any one of claims 1 to 19 when executed by a first processor, or implements the method according to any one of claims 20 to 38 when executed by a second processor.

Citation Information

Patent Citations

  • Method and apparatus for encoding / decoding image, and recording medium in which bit stream is stored

    CN110024394A