End-Motion Refinement in Video Coding / Decoding Systems
By simplifying the standard for motion vector refinement, the problem of high computational complexity in the prior art is solved, resource saving and encoding and decoding time are achieved, and the efficiency of video encoding and decoding is improved.
Patent Information
- Application Number
- CN201980087686.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-02
- Filing Date
- 2019-12-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-12-09
AI Technical Summary
In the prior art, in the process of video encoding and decoding, the calculation complexity of motion vector refinement is high, resulting in increased resource consumption and encoding and decoding time, affecting compression efficiency.
By adjusting the criteria for motion vector refinement, simplifying the computational complexity, and using alternative standards to decide whether to perform bidirectional optical flow (BIO) or other motion vector refinement techniques, reducing computing requirements.
Reduces the computational complexity of motion vector refinement, saves processing resources, reduces the time required for video sequence encoding or decoding, and ignores the negative impact on compression efficiency.
Smart Images

Figure CN113302935B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to methods and apparatus for video encoding and decoding. Background Art
[0002] A video sequence contains a sequence of pictures. Each picture is assigned a picture order count (POC) value, which indicates the order in which the picture is displayed in the sequence.
[0003] Each picture in the sequence includes one or more components. Each component can be described as a two-dimensional rectangular array of sample values. Typically, an image in a video sequence includes three components: a luminance component Y, where the sample values are luminance values; and two chrominance components Cb and Cr, where the sample values are chrominance values. Other examples include Y'CbCr, YUV, and ICTCP. In ICTCP, I is the "intensity brightness" component. In the following description, any luminance component Y', Y, or I will be referred to as Y or simply brightness. Typically, the size of the chrominance component is half of the luminance component in each dimension. For example, the size of the luminance component of a high-density (HD) image is 1920x1080, while the size of the chrominance components is 960x540. Chroma components are sometimes referred to as color components.
[0004] Video coding is used to compress a video sequence into a sequence of coded pictures. Typically, pictures are divided into blocks of size ranging from 4x4 to 128x128. These blocks are used as the basis for encoding. The video decoder then decodes the coded pictures into pictures containing sample values.
[0005] A block is a two-dimensional array of samples. In video coding, each component is divided into blocks, and the encoded video bitstream includes a series of blocks. Typically, in video coding, an image is divided into units covering a specific area of the image. Each unit includes all blocks from all components that make up the specific area, and each block belongs to one unit. Macroblocks in H.264 and coding units (CUs) in High Efficiency Video Coding (HEVC) are examples of units.
[0006] The draft VVC video coding standard uses a block structure called a quadtree plus binary tree plus ternary tree block structure (QTBT+TT), in which each picture is first partitioned into square blocks called coding tree units (CTUs). All CTUs are the same size, and the partitioning is done without any syntax to control it. Each CTU is further partitioned into coding units (CUs) that can have a square or rectangular shape. The CTU is first partitioned by a quadtree structure, and then it can be further partitioned vertically or horizontally using partitions of the same size in a binary structure to form coding units (CUs). Therefore, the block can have a square or rectangular shape. The depth of the quadtree and binary tree can be set by the encoder in the bitstream. Figure 1 An example of using QTBT to partition a CTU is shown. In particular, Figure 1 A picture 12 is shown divided into four CTUs 14. The CTUs are further partitioned into square or rectangular CUs 16. Figure 1 Also shown is the QTBT+TT 18 used to partition the picture. The ternary tree (TT) part increases the likelihood of dividing a CU into three partitions instead of two equally sized partitions; this increases the likelihood of using a block structure that is more suitable for the content structure in the picture.
[0007] Inter prediction
[0008] In order to achieve efficient compression in the temporal domain, inter-frame prediction techniques aim to exploit the similarity between pictures. Inter-frame prediction uses previously decoded pictures to predict blocks in the current picture. The previously decoded pictures are called reference pictures for the current picture.
[0009] In video encoders, a method called motion estimation is usually used to find the most similar block in a reference picture. The displacement between the current block and its reference block is a motion vector (MV). MV has two components, MV.x and MV.y, i.e., the x and y directions. Figure 2 An example of an MV between a current block 22 in a current picture 24 and a reference block 26 in a reference picture 28 is depicted. The MV is signaled to the decoder in the video bitstream. The current picture 24 has a picture order count POC0, while the reference picture 28 has a picture order count POC1. Note that the POC of the reference picture may be before or after the POC of the current picture (i.e., the current picture may be predicted from a picture that is after the current picture in the picture order count), because pictures in a video bitstream may not be encoded and / or decoded in the order in which they are displayed.
[0010] The video decoder decodes the MV from the video bitstream. The decoder then applies a method called motion compensation, which uses the MV to find the corresponding reference block in the reference picture.
[0011] If a block is predicted from at least one reference block in a reference picture, the block is called an inter block.
[0012] Bidirectional inter prediction
[0013] The number of reference blocks is not limited to one. In bidirectional motion compensation, two reference blocks can be used to further exploit temporal redundancy, i.e., the current block is predicted from two previously decoded blocks. Pictures using bidirectional motion compensation are called bidirectional prediction pictures (B-pictures). Figure 3 An example of a block 22 in a current picture 24 with bidirectional motion compensation based on reference blocks 26, 32 in reference pictures 28, 34, respectively, is depicted. In this example, a first motion vector (MV0) points from point C of the current block to point A of the first reference block 26, while a second motion vector (MV1) points from point C of the current block to point B of the second reference block 32.
[0014] The motion information set contains MV (MV.x and MV.y) and a reference picture with a POC number. If bidirectional motion compensation is used, there are two motion information sets, namely, set 0 with MV0, POC1 and associated block 26, and set 1 with MV1, POC2 and associated block 32, as shown in FIG. Figure 3 shown.
[0015] The temporal distance between the current B picture and the reference picture can be represented by the absolute POC difference between the pictures. Figure 3 If POC1=0 of reference picture 0 and POC0=8 of the current B-picture, the absolute temporal distance between the two pictures is |POC1-POC0|=8. The signed temporal distance between the two pictures is (POC1-POC0)=-8. The negative sign of the signed temporal distance between them indicates that the current B-picture is after the reference picture in the display sequence.
[0016] MV Difference and MV Scaling
[0017] A common method for comparing the similarity between two MVs is to calculate the absolute MV difference, i.e., |MV0-MV1|. Since both MVs originate from point C, one of the motion vectors must be rotated 180 degrees before calculating the MV difference (ΔMV), as Figure 4A and 4BAs shown (MV0 is rotated 180 degrees to become MV0'). Rotating 180 degrees can be accomplished by simply negating the values of the vector components (e.g., MV0'.x = -MV0.x and MV0'.y = -MV0.y). Then the value of ΔMV = (MV1.x - MV0'.x, MV1.y - MV0'.y) can be obtained.
[0018] When the reference pictures associated with two MVs have different absolute temporal distances from the current picture, motion vector scaling is required before the motion vector difference can be calculated.
[0019] For example, refer to Figure 5A , MV0 is associated with reference picture 28 having POC1=10, and MV1 is associated with reference picture 34 having POC2=16. Current picture 24 has POC0=12.
[0020] Therefore, assuming that ΔPOCN=POCN-POC0, ΔPOC1=-2 and ΔPOC2=4.
[0021] Therefore, when comparing the similarity between MV0 and MV1, one of the MVs needs to be scaled first by the ratio of ΔPOC2 / ΔPOC1=-2.
[0022] from Figure 5A As can be seen in , on the one hand, a motion vector MV can be viewed as a three-dimensional vector having x and y components (MV.x, MV.y) in the plane of the picture and having a z component in the time dimension corresponding to the ΔPOC1 associated with MV.
[0023] Figure 5B shows the x and y components of MV0 and MV1 plotted on the same xy plane, while Figure 5C The scaling of MV0 is shown to account for the different temporal distances of the current picture associated with the two motion vectors.
[0024] In the general case when ΔPOC is not the same, MV scaling depends on the POC difference (ΔPOC2 = POC2 - POC0) and (ΔPOC1 = POC1 - POC0). Assuming MV0 is the vector to be scaled, the components of the scaled vector MV0' can be calculated as:
[0025]
[0026] Note that in this example, scaling MV0 also has the effect of rotating MV0 180 degrees, since the ratio of ΔPOC2 / ΔPOC1 has a negative sign.
[0027] Bidirectional Optical Flow (BIO)
[0028] The BIO method described in [1] is a decoder-side technique for further refining the motion vectors used for bidirectional motion compensation. It uses the concept of optical flow combined with bidirectional prediction to predict the luminance value in the current block. BIO is applied after traditional bidirectional motion compensation as a pixel-level motion refinement.
[0029] Optical flow assumes that the brightness of an object does not change during a specific period of motion. It gives the optical flow equation:
[0030] I x v x +I y v y +I t ≈0, (2) where I is the brightness of the pixel, v_x is the velocity in the x direction, v_y is the velocity in the y direction, and and are the derivatives in the x direction, y direction, and with respect to time, respectively. In BIO, the motion is assumed to be steady and the velocity (v x ,v y ) and the speed in reference 1 (-v x ,-v y ) On the contrary, Figure 6 shown.
[0031] Let the current pixel be at position [i, j] in the B picture, I (0) [i,j] is the brightness at [i,j] in reference 0, I (1) [i,j] is the brightness at [i,j] in reference 1. Let the derivative be Where k = 0, 1 is the reference index. Based on (2), the authors of [1] defined the error of BIO as:
[0032]
[0033] Then, BIO can be formulated as a least squares problem:
[0034]
[0035] where [i′, j′]∈Ω is a sub-block containing neighboring pixels including [i, j] which is used to solve equation (4). The solution of equation (4) gives The optimal expression of the brightness difference An alternative solution is to use a moving window around [i, j] to compute the sum in Eq. (4). However, by using sub-blocks, two pixels belonging to the same sub-block will be summed over the same Ω, which means The calculation of can be reused between pixels belonging to the same sub-block.
[0036] Once the velocity is obtained, the next step is to predict the current pixel from the two references. Figure 6 As shown, the current pixel is predicted from two directions. The authors of [2] introduced a third-order polynomial function to interpolate the value between two reference pixels, that is,
[0037] P(t)=a0+a1t+a2t 2 +a3t 3 , (5) where t is time, and a0 to a3 are parameters. Let reference 0 be at time 0, B picture at time τ0, and reference 1 at time τ0+τ1. That is, Figure 6 In , the time difference between reference 0 and B picture is τ0, and the time difference between B picture and reference 1 is τ1. To find these four parameters, consider the following four equations:
[0038]
[0039] Equation (6) shows that the interpolation function P(t) needs to match the brightness value I (0) and I (1) and the derivatives at t = 0 and t = τ0 + τ1. Using the four equations, a0 to a3 can be solved.
[0040] In the general case, when τ0 = τ1 = τ, the interpolation function gives:
[0041]
[0042] Note that when This is different from just I (0) and I (1) Equation (7) can be viewed as a simple linear interpolation of the average value of I (0) and I (1) The refinement of , and this refinement helps to improve the interpolation accuracy. Using the optical flow equation I t =-I x v x -I y v y Replace I t , equation (7) can be written as:
[0043]
[0044] From equation (8), the brightness value I (0) and I (1) is known. The derivative can be estimated from the gradient using neighboring pixels. The velocity is solved in equation (4). Therefore, the predicted brightness value P(τ0) can be obtained.
[0045] In the implementation of BIO described in [3], criteria indicating whether BIO should be considered are checked in the decoder. The criteria is set to TRUE if all of the following conditions hold: a) the prediction is bi-directional and from opposite directions (e.g., POC1 < POC0 and POC2 > POC0), b) an affine motion model is not used, and c) advanced temporal motion prediction (ATMVP) is not used. If any of the conditions from a) to c) does not hold, the criteria is set to FALSE.
[0046] As shown in Equation (4), sub-block Ω is used to calculate each (v x , v y ) pair in BIO. For two pixels belonging to the same sub-block, the same (v x , v y ) vector is used in BIO refinement. In this implementation, the size of the sub-block is 4x4. For example, if the size of the block is 128x128, it contains 32x32 = 1024 sub-blocks and has 1024 (v x , v y ) pairs.
[0047] When the criteria for BIO is TRUE, the sum of absolute differences (SAD) between two reference blocks is calculated. SAD is obtained by calculating the absolute difference of each pixel between two reference blocks. In addition, SAD_sub between two reference sub-blocks (both of size 4x4) is also calculated. Then, BIO is applied when SAD and SAD_sub are greater than specific thresholds respectively. These thresholds depend on the block size and bit depth. On the other hand, if BIO is not applied, linear averaging is used to predict the signal.
[0048] Other motion vector refinement techniques are well-known, such as decoder-side motion vector refinement (DMVR). In DMVR, a bilateral template is generated as a weighted combination of two reference blocks associated with the initial motion vector. Then, bilateral template matching is performed to find the best matching block in the reference picture and identify the updated motion vector. SUMMARY OF THE INVENTION
[0049] A first aspect of the embodiments defines a method performed by a decoder for decoding a current block in a current picture of a video bitstream. The current picture has a current picture sequence count. The method includes decoding from the video bitstream a first motion vector for the current block relative to a first reference block of a first reference picture having a first picture sequence count. The method includes decoding from the video bitstream a second motion vector for the current block relative to a second reference block of a second reference picture having a second picture sequence count. The method also includes generating a similarity metric based on a comparison of the first motion vector and the second motion vector. The method also includes determining whether to refine the first motion vector based on the similarity metric. In response to determining whether to refine the first motion vector, the method also includes generating a first refined motion vector from the first motion vector. The method also includes performing motion compensation to derive a first reference block from the first reference picture using the first refined motion vector.
[0050] A second aspect of the embodiments defines a decoder for decoding a current block in a current picture of a video bitstream for use in a communication network according to some embodiments. The decoder comprises a processor circuit and a memory coupled to the processor circuit. The memory comprises instructions that, when executed by the processor circuit, cause the processor circuit to perform operations according to the first aspect.
[0051] A third aspect of the embodiments defines a computer program for a decoder. The computer program comprises computer executable instructions configured to, when executed on a processor circuit included in the decoder, cause the decoder to perform operations according to the second aspect.
[0052] A fourth aspect of the embodiments defines a computer program product comprising a non-transitory storage medium, the non-transitory storage medium comprising program code, the program code to be executed by at least one processor of a decoder, whereby execution of the program code causes the decoder to perform a method according to the first aspect.
[0053] A fifth aspect of the embodiments defines a method performed by a video encoder for encoding a current block in a current picture of a video bitstream. The current picture has a current picture sequence count. The method includes generating a first motion vector for the current block relative to a first reference block of a first reference picture having a first picture sequence count. The method includes generating a second motion vector for the current block relative to a second reference block of a second reference picture having a second picture sequence count. The method also includes generating a similarity metric based on a comparison of the first motion vector and the second motion vector. The method also includes determining whether to refine the first motion vector based on the similarity metric. In response to determining whether to refine the first motion vector, the method also includes generating a first refined motion vector from the first motion vector. The method also includes performing motion compensation using the first refined motion vector to derive a first reference block from the first reference picture.
[0054] A sixth aspect of the embodiments defines an encoder for encoding a current block in a current picture of a video bitstream. The encoder comprises a processor circuit and a memory coupled to the processor circuit, wherein the memory comprises instructions that, when executed by the processor circuit, cause the processor circuit to perform operations according to the fourth aspect.
[0055] A seventh aspect of the embodiments defines a computer program for an encoder. The computer program comprises computer executable instructions configured to, when executed on a processor circuit included in the encoder, cause the encoder to perform operations according to the fifth aspect.
[0056] An eighth aspect of the embodiment defines a computer program product comprising a non-transitory storage medium, wherein the non-transitory storage medium comprises program code, wherein the program code is to be executed by at least one processor of an encoder, whereby the execution of the program code causes the encoder to perform a method according to the fifth aspect.
[0057] One potential advantage that the inventive concept can provide includes reducing the computational complexity of the criteria for determining whether motion vector refinement (e.g., using a BIO processing algorithm) should be performed during encoding and / or decoding of a video sequence. This can save processing resources and / or reduce the time required to encode or decode a video sequence. In addition, some embodiments described herein have a negligible impact on compression efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings, which are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this application, illustrate certain non-limiting embodiments of the inventive concept. In the drawings:
[0059] Figure 1 is a block diagram illustrating partitioning of a picture into coding units;
[0060] Figure 2 is a block diagram illustrating an example of using a motion vector for encoding / decoding a block based on a reference block;
[0061] Figure 3 is a block diagram illustrating an example of bidirectional motion compensation;
[0062] Figure 4A and 4B A comparison of two motion vectors is shown;
[0063] Figure 5A , 5B and 5C show a comparison of the scaled and unscaled motion vectors;
[0064] Figure 6 An example of bidirectional optical flow processing is shown;
[0065] Figure 7 is a block diagram illustrating an example of an environment in which a system of an encoder and a decoder may be implemented according to some embodiments of the inventive concept;
[0066] Figure 8 is a block diagram illustrating an encoder according to some embodiments;
[0067] Fig. 9 is a block diagram illustrating a decoder according to some embodiments;
[0068] Figures 10 to 15 is a flowchart illustrating the operation of a decoder or an encoder according to some embodiments of the inventive concept. DETAILED DESCRIPTION
[0069] The inventive concept will now be described more fully below with reference to the accompanying drawings, in which examples of embodiments of the inventive concept are shown. However, the inventive concept may be embodied in many different forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the disclosure will be thorough and complete and fully convey the scope of the inventive concept to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Components from one embodiment may be assumed by default to be present in / used in another embodiment.
[0070] The following description sets forth various embodiments of the disclosed subject matter. These embodiments are presented as teaching examples and should not be construed as limiting the scope of the disclosed subject matter. For example, specific details of the described embodiments may be modified, omitted, or expanded without departing from the scope of the described subject matter.
[0071] As mentioned above, BIO processing can be applied after conventional bidirectional motion compensation as pixel-level motion refinement. As described in [3], BIO is enabled when SAD and SAD_sub are greater than a certain threshold. However, one disadvantage of calculating SAD is that it requires many arithmetic operations. To calculate the SAD between two blocks of size n×n, n 2 Subtraction, n 2 The absolute value and n 2 -1 addition. Taking the worst case as an example, if the block size is n = 128, the number of operations is 128 2 = 16384 subtractions, 128 2 = 16384 times the absolute value, and 128 2 -1 = 16383 additions. This is computationally expensive.
[0072] Another disadvantage is that the solution given in equation (8) is only valid for equal temporal distance τ_0 = τ_1. As mentioned above, in video coding, a B picture can have two reference pictures with different temporal distances, i.e., τ_0 ≠ τ_1 or |POC0-POC1| ≠ |POC0-POC2|. However, in the implementation [3], the solution in equation (8) is applied to all BIO cases regardless of the POC difference. Therefore, there is an inconsistency between the theoretical solution and the actual implementation, which may lead to inefficiency of the BIO method.
[0073] Some embodiments provide alternative criteria for determining whether to perform BIO or other motion vector refinement techniques as part of the encoding or decoding process.
[0074] Figure 7 An example of an operating environment for an encoder 800 that can be used to encode a bitstream as described herein is shown. The encoder 800 receives video from a network 702 and / or a storage device 704 and encodes the video into a bitstream as described below and sends the encoded video to a decoder 900 via a network 708. The storage device 704 can be part of a repository of multi-channel audio signals, such as a repository of a store or a streaming video service, a separate storage component, a component of a mobile device, etc. The decoder 900 can be part of a device 910 having a media player 912. The device 910 can be a mobile device, a set-top device, a desktop computer, etc.
[0075] Figure 8800 is a block diagram illustrating units of an encoder 800 configured to encode video frames according to some embodiments of the present inventive concept. As shown, the encoder 800 may include a network interface circuit 805 (also referred to as a network interface) configured to provide communication with other devices / entities / functions / etc. The encoder 800 may also include a processor circuit 801 (also referred to as a processor) coupled to the network interface circuit 805 and a memory circuit 803 (also referred to as a memory) coupled to the processor circuit. The memory circuit 803 may include a computer-readable program code that, when executed by the processor circuit 801, causes the processor circuit to perform operations according to the embodiments disclosed herein.
[0076] According to other embodiments, the processor circuit 801 may be defined to include a memory, thereby eliminating the need for a separate memory circuit. As discussed herein, the operations of the encoder 800 may be performed by the processor 801 and / or the network interface 805. For example, the processor 801 may control the network interface 805 to send communications to the decoder 900 and / or receive communications from one or more other network nodes / entities / servers (e.g., other encoder nodes, repository servers, etc.) through the network interface 802. In addition, modules may be stored in the memory 803, and these modules may provide instructions so that when the instructions of the modules are executed by the processor 801, the processor 801 performs the corresponding operations.
[0077] Fig. 9 is a block diagram illustrating units of a decoder 900 configured to decode a video frame according to some embodiments of the present inventive concept. As shown, the decoder 900 may include a network interface circuit 905 (also referred to as a network interface) configured to provide communication with other devices / entities / functions / etc. The decoder 900 may also include a processor circuit 901 (also referred to as a processor) coupled to the network interface circuit 905 and a memory circuit 903 (also referred to as a memory) coupled to the processor circuit. The memory circuit 903 may include a computer-readable program code that, when executed by the processor circuit 901, causes the processor circuit to perform operations according to the embodiments disclosed herein.
[0078] According to other embodiments, the processor circuit 901 may be defined to include a memory, thereby eliminating the need for a separate memory circuit. As discussed herein, the operations of the decoder 900 may be performed by the processor 901 and / or the network interface 905. For example, the processor 901 may control the network interface 905 to receive communications from the encoder 800. In addition, modules may be stored in the memory 903, and these modules may provide instructions such that when the instructions of the modules are executed by the processor 901, the processor 901 performs the corresponding operations.
[0079] Some embodiments of the inventive concept adjust the conditions of a criterion for enabling BIO during encoding and / or decoding of a video sequence and simplify the calculation of the criterion.
[0080] Decoder Operation
[0081] Fig.10 The operation of a decoder 900 for decoding a video bitstream according to some embodiments is shown. Fig.10 , the decoder 900 first selects a current block K in a current picture having a POC equal to POCn for processing (block 1002). The current picture is a B picture decoded using bidirectional prediction.
[0082] The decoder 900 then decodes a motion information set mA from the video bitstream, where mA contains a motion vector mvA and a first reference picture refA having a POC equal to POCa (block 1004).
[0083] The decoder 900 then decodes another motion information set mB from the video bitstream, where mB contains a motion vector mvB and a reference picture refB having a POC equal to POCb, (block 1006).
[0084] Next, the decoder 900 determines whether further refinement of mvA and / or mvB is needed based on some criterion C (block 1008). In some embodiments, the criterion C may be based on a comparison of one or more components of the motion vectors mvA, mvB, including a comparison of their x-components, y-components, and / or z-components (e.g., their ΔPOCs). In other embodiments, the criterion may be based on a comparison of reference blocks associated with the motion vectors.
[0085] When criterion C is met, further refinement is performed on mvA and / or mvB (block 1010). The refinement produces two new motion vectors mvA*, mvB*. These motion vectors mvA* and mvB* are then used in the motion compensation process to derive the corresponding reference blocks for predicting the sample values of block K (block 1012).
[0086] When criterion C is not met, no further refinement is performed on mvA and mvB. Both mvA and mvB are used directly in the motion compensation process to find the corresponding reference block for predicting the sample values of block K.
[0087] Finally, the decoder 900 decodes the current block using the reference block (block 1014).
[0088] In some embodiments, criterion C may include comparing the similarity between mvA and mvB. The similarity may be measured as described above by calculating the motion vector difference ΔMV between mvA and mvB, where one of the MVs is rotated 180 degrees. ΔMV may then be calculated by taking the absolute difference in the x and y directions, i.e., ΔMV.x = |mvA.x - mvB.x|, ΔMV.y = |mvA.y - mvB.y|.
[0089] In some embodiments, when calculating the motion vector difference, motion vector scaling according to equation (1) may be involved. When only one MV is scaled, |ΔMV| will differ depending on which MV is scaled. In some embodiments, the selection of which motion vector to scale may be based on the relative magnitude of the motion vectors. For example, if |mvA|<|mvB|, then mvA may be selected for scaling, and vice versa.
[0090] In other embodiments, the two motion vectors may be scaled as follows:
[0091] mvA' = mvA / ΔPOCa (9)
[0092] mvB'=mvB / ΔPOCb
[0093] Among them, ΔPOCa=POCa-POCn and ΔPOCb=POCb-POCn.
[0094] When scaled in this manner, the same threshold can be used to evaluate ΔMV regardless of the order in which the MVs are considered. Furthermore, since only one of ΔPOCa and ΔPOCb will be negative, only one MV will be rotated 180 degrees when scaling the MVs using equation (9).
[0095] Criterion C determines whether one or both of the ΔMV.x and ΔMV.y components of ΔMV are less than a first threshold, ie, ΔMV.x < threshold 1 and / or ΔMV.y < threshold 1. This refinement is performed when criterion C is met. Fig.14 As shown therein, the method may generate a similarity metric based on motion vectors mvA and mvB (block 1402). The similarity metric may include ΔMV.x, ΔMV.y, or a combination of ΔMV.x and ΔMV.y.
[0096] In some embodiments, a second threshold may be provided, wherein the second threshold is less than the first threshold. Criterion C may determine whether one or both of the components of ΔMV are less than the first threshold and greater than the second threshold. The refinement is performed when criterion C is met.
[0097] In some embodiments, the criterion C may include comparing the temporal distance ΔPOC between the reference picture refA, the current picture and the reference picture refB. The temporal distance may be calculated as described above using the corresponding POC values POC0, POCa and POCb.
[0098] Criterion C may include determining whether the absolute POC differences are equal, ie, |POCn-POCa|=|POCn-POCb|, and performing the refinement when criterion C is met.
[0099] In some embodiments, criterion C may include determining whether at least one of the following two conditions is met: (a) |POCn-POCa|≤Threshold_1; (b) |POCn-POCb|≤Threshold_2. Criterion C is met if one or both of the two conditions are met.
[0100] The temporal distance between two pictures is a function of both the POC difference and the frame rate (usually expressed in frames per second or fps). The thresholds for evaluating the POC difference and / or the MV difference can be based on a specific frame rate. In some embodiments, the encoder can explicitly signal the thresholds to the decoder in the video bitstream. In other embodiments, the decoder can scale the thresholds based on the frame rate of the video bitstream.
[0101] exist Fig.11 The operation of the decoder 900 according to other embodiments is shown in FIG. Fig.11 , the decoder 900 first selects a current block K in the current picture having a POC equal to POCn for processing (block 1102). The current picture is a B picture decoded using bidirectional prediction.
[0102] The decoder 900 then decodes a first motion information set mA from the video bitstream, wherein mA includes a motion vector mvA and a first reference picture refA having a POC equal to POCa, and decodes a second motion information set mB from the video bitstream, wherein mB includes a motion vector mvB and a reference picture refB having a POC equal to POCb (box 1104).
[0103] Next, the decoder 900 performs motion compensation using mvA and mvB to find reference blocks R0 and R1 in reference pictures refA and refB, respectively (block 1106).
[0104] Next, the decoder 900 determines whether further refinement of the reference blocks R0 and R1 is needed based on some criterion C (block 1108). In some embodiments, the criterion C may be based on a comparison of the reference blocks R0 and R1.
[0105] When criterion C is met, mvA and / or mvB are further refined (block 1110). The refinement produces two new motion vectors mvA*, mvB*. These motion vectors mvA* and mvB* are then used in the motion compensation process to derive corresponding refined reference blocks R0* and R1* for predicting sample values of the current block K (block 1112).
[0106] When criterion C is not met, no further refinement is performed. The resulting reference blocks R0 and R1 or R0* and R1* are then used to decode the current block K (block 1114).
[0107] Brief reference Fig.15 As shown therein, the method may include generating a similarity metric based on R0 and R1 (block 1502). The similarity metric may be generated by comparing only a limited set of sample values in R0 and R1, rather than performing a full SAD. In some embodiments, the criterion compares every mth and nth sample value of the reference blocks R1 and R0 in the x and y directions, respectively, and calculates a similarity metric for the reference blocks based on the selected samples. When the similarity is less than a first threshold, refinement may be performed.
[0108] In some embodiments, a similarity metric may be generated by comparing the mean or variance of the finite sample sets in R0 and R1. In other embodiments, the SAD of the finite sample sets may be calculated. That is, the similarity value may be calculated by the SAD method as the sum of the absolute differences of all co-located sample value pairs having coordinates included in the coordinate set.
[0109] In other embodiments, the mean square error (MSE) may be calculated for a finite set of sample values of R0 and R1.
[0110] Encoder Operation
[0111] Fig.12 FIG. 8 is a diagram illustrating the operation of an encoder 800 for decoding a video bitstream according to some embodiments. Fig.12 , the encoder 800 first selects a current block K in a current picture having a POC equal to POCn for processing (block 1202). The current picture is a B picture decoded using bidirectional prediction.
[0112] The encoder 800 then generates a motion information set mA from the video bitstream, where mA includes a motion vector mvA and a first reference picture refA having a POC equal to POCa (block 1204).
[0113] The encoder 800 then generates another motion information set mB from the video bitstream, where mB contains a motion vector mvB and a reference picture refB having a POC equal to POCb (block 1206).
[0114] Next, the encoder 800 determines whether further refinement of mvA and / or mvB is needed based on some criterion C (block 1208). In some embodiments, the criterion C may be based on a comparison of one or more components of the motion vectors mvA, mvB, including a comparison of their x-components, y-components, and / or z-components (e.g., their ΔPOCs). In other embodiments, the criterion may be based on a comparison of reference blocks associated with the motion vectors.
[0115] When criterion C is met, further refinement is performed on mvA and / or mvB (block 1210). The refinement produces two new motion vectors mvA*, mvB*. These motion vectors mvA* and mvB* are then used in the motion compensation process to derive the corresponding reference blocks for predicting the sample values of block K (block 1212).
[0116] When criterion C is not met, no further refinement is performed on mvA and mvB. Both mvA and mvB are directly used in the motion compensation process to find the corresponding reference block for predicting the sample values of block K.
[0117] Finally, the encoder 800 encodes the current block using the reference block (block 1214).
[0118] In some embodiments, criterion C may include comparing the similarity between mvA and mvB. The similarity may be measured as described above by calculating the motion vector difference ΔMV between mvA and mvB, where one of the MVs is rotated 180 degrees.
[0119] In some embodiments, when calculating motion vector differences, motion vector scaling according to equation (1) or (9) may be applied.
[0120] Criterion C determines whether one or both of the ΔMV.x and ΔMV.y components of ΔMV are less than a first threshold, ie, ΔMV.x<threshold 1 and / or ΔMV.y<threshold 1. When criterion C is met, refinement is performed.
[0121] In some embodiments, a second threshold may be provided, wherein the second threshold is less than the first threshold. Criterion C may determine whether both or one of the components of ΔMV is less than the first threshold and greater than the second threshold. The refinement is performed when criterion C is met.
[0122] In some embodiments, the criterion C may include comparing the temporal distance ΔPOC between the reference picture refA, the current picture and the reference picture refB. The temporal distance may be calculated as described above using the corresponding POC values POC0, POCa and POCb.
[0123] Criterion C may include determining whether the absolute POC differences are equal, ie, |POCn-POCa|=|POCn-POCb|, and performing the refinement when criterion C is met.
[0124] In some embodiments, criterion C may include determining whether at least one of the following two conditions is met: (a) |POCn-POCa|≤Threshold_1; (b) |POCn-POCb|≤Threshold_2. Criterion C is met if one or both of the two conditions are met.
[0125] In some embodiments, the encoder may explicitly signal the threshold value to the decoder in the video bitstream. In other embodiments, the decoder may signal the frame rate or a scaling factor that the decoder may use to scale the threshold value based on the frame rate of the video bitstream.
[0126] exist Fig.13 The operation of the encoder 800 according to other embodiments is shown in FIG. Fig.13 , the encoder 800 first selects a current block K in a current picture having a POC equal to POCn for processing (block 1302). The current picture is a B picture decoded using bidirectional prediction.
[0127] The encoder 800 then generates a first motion information set mA from the video bitstream, wherein mA includes a motion vector mvA and a first reference picture refA having a POC equal to POCa, and generates a second motion information set mB from the video bitstream, wherein mB includes a motion vector mvB and a reference picture refB having a POC equal to POCb (box 1304).
[0128] Next, the encoder 800 performs motion compensation using mvA and mvB to find reference blocks R0 and R1 in reference pictures refA and refB, respectively (block 1306).
[0129] Next, the encoder 800 determines whether further refinement of the reference blocks R0 and R1 is needed based on some criterion C (block 1308). In some embodiments, the criterion C may be based on a comparison of the reference blocks R0 and R1.
[0130] When criterion C is met, further refinement is performed on mvA and / or mvB (block 1310). The refinement produces two new motion vectors mvA*, mvB*. These motion vectors mvA* and mvB* are then used in the motion compensation process to derive corresponding refined reference blocks R0* and R1* for predicting sample values of the current block K (block 1312).
[0131] When criterion C is not met, no further refinement is performed. The resulting reference blocks R0 and R1 or R0* and R1* are then used to encode the current block K (block 1314).
[0132] In this embodiment, criterion C may include comparing only a limited set of sample values in R0 and R1, rather than performing a full SAD. In some embodiments, the criterion compares every mth and nth sample value of reference blocks R1 and R0 in the x and y directions, respectively, and calculates a similarity measure for the reference blocks. When the similarity is less than a first threshold, the refinement may be performed.
[0133] In some embodiments, a similarity metric may be generated by comparing the mean or variance of the finite sample sets in R0 and R1. In other embodiments, the SAD of the finite sample sets may be calculated. That is, the similarity value may be calculated by the SAD method as the sum of the absolute differences of all co-located sample value pairs having coordinates included in the coordinate set.
[0134] In other embodiments, the mean square error (MSE) may be calculated for a finite set of sample values of R0 and R1.
[0135] References:
[0136] [1] A. Alshin and E. Alshina, “Bi-directional optical flow”, Joint Collaboration Team on Video Coding (JCT-VC) of ITU-TSG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, JCTVC-C204, Guangzhou, China, October 10-15, 2010.
[0137] [2] A. Alexander and E. Alshina, “Bi-directional optical flow for future video codecs,” in 2016 Data Compression Conference (DCC), pages 83–90. IEEE, 2016.
[0138] [3] X. Xiu, Y. He, Y. Ye, “CE9-related: Complexity reduction and bit-width control for bi-directional optical flow (BIO)”, JVET input document, document number JVET-L0256.
Claims
1. A method for decoding a current block in a current picture of a video bitstream, the current picture having a current picture sequence count, the method comprising: decoding, from the video bitstream, a first motion vector for a first reference block of a first reference picture having a first picture order count for the current block; decoding, from the video bitstream, a second motion vector for a second reference block of a second reference picture having a second picture order count for the current block, wherein the first motion vector and the second motion vector each comprise a three-dimensional motion vector comprising an x-component in a plane of the current picture, a y-component in the plane of the current picture, and a z-component, wherein the z-component represents a temporal component; determining whether to refine the first motion vector based on a comparison of a difference between the z components of the first motion vector and the second motion vector and a third threshold, and determining to refine the first motion vector in response to the difference between the z components of the first motion vector and the second motion vector being less than the third threshold; In response to determining to refine the first motion vector, generating a first refined motion vector from the first motion vector; and Using the first refined motion vector, motion compensation is performed to derive an updated first reference block from the first reference picture.
2. The method according to claim 1, wherein: Determining whether to refine the first motion vector further includes comparing a difference between the x-components of the first motion vector and the second motion vector to a first threshold, and determining to refine the first motion vector in response to the difference between the x-components of the first motion vector and the second motion vector being less than the first threshold.
3. The method according to claim 2, further comprising: A difference between the x-components of the first motion vector and the second motion vector is generated based on a difference between an absolute value of the x-component of the first motion vector and an absolute value of the x-component of the second motion vector.
4. The method according to claim 1, wherein: Determining whether to refine the first motion vector further includes comparing a difference between the y components of the first motion vector and the second motion vector with a second threshold, and determining to refine the first motion vector in response to the difference between the y components of the first motion vector and the second motion vector being less than the second threshold.
5. The method according to claim 4, further comprising: A difference between the y components of the first motion vector and the second motion vector is generated based on a difference between an absolute value of the y component of the first motion vector and an absolute value of the y component of the second motion vector.
6. The method according to claim 1, wherein: the z component of the first motion vector comprises a difference between the current picture order count and the first picture order count, and the z component of the second motion vector comprises a difference between the current picture order count and the second picture order count; as well as The method further includes generating a difference between the z components of the first motion vector and the second motion vector based on a difference between an absolute value of the z component of the first motion vector and an absolute value of the z component of the second motion vector.
7. A decoder for decoding a current block in a current picture of a video bitstream, the current picture having a current picture sequence count, the decoder comprising: processor circuit; as well as a memory coupled to the processor circuit, wherein the memory includes instructions that, when executed by the processor circuit, cause the processor circuit to perform operations comprising: decoding, from the video bitstream, a first motion vector for a first reference block of a first reference picture having a first picture order count for the current block; decoding, from the video bitstream, a second motion vector for a second reference block of a second reference picture having a second picture order count for the current block, wherein the first motion vector and the second motion vector each comprise a three-dimensional motion vector comprising an x-component in a plane of the current picture, a y-component in the plane of the current picture, and a z-component, wherein the z-component represents a temporal component; determining whether to refine the first motion vector based on a comparison of a difference between the z components of the first motion vector and the second motion vector and a third threshold, and determining to refine the first motion vector in response to the difference between the z components of the first motion vector and the second motion vector being less than the third threshold; In response to determining to refine the first motion vector, generating a first refined motion vector from the first motion vector; and Using the first refined motion vector, motion compensation is performed to derive an updated first reference block from the first reference picture.
8. The decoder according to claim 7, wherein: The operation of determining whether to refine the first motion vector further includes comparing a difference between the x-components of the first motion vector and the second motion vector with a first threshold, and determining to refine the first motion vector in response to the difference between the x-components of the first motion vector and the second motion vector being less than the first threshold.
9. The decoder according to claim 8, wherein: The operations also include: A difference between the x-components of the first motion vector and the second motion vector is generated based on a difference between an absolute value of the x-component of the first motion vector and an absolute value of the x-component of the second motion vector.
10. The decoder according to claim 7, wherein: The operation of determining whether to refine the first motion vector also includes: comparing the difference between the y components of the first motion vector and the second motion vector with a second threshold, and determining to refine the first motion vector in response to the difference between the y components of the first motion vector and the second motion vector being less than the second threshold.
11. The decoder according to claim 10, wherein: The operations also include: A difference between the y components of the first motion vector and the second motion vector is generated based on a difference between an absolute value of the y component of the first motion vector and an absolute value of the y component of the second motion vector.
12. The decoder according to claim 7, wherein: the z component of the first motion vector comprises a difference between the current picture order count and the first picture order count, and the z component of the second motion vector comprises a difference between the current picture order count and the second picture order count; as well as The operations also include generating a difference between the z components of the first motion vector and the second motion vector based on a difference between an absolute value of the z component of the first motion vector and an absolute value of the z component of the second motion vector.
13. A method performed by an encoder for encoding a current block in a current picture of a video bitstream, the current picture having a current picture sequence count, the method comprising: generating a first motion vector for the current block relative to a first reference block of a first reference picture having a first picture order count; generating a second motion vector for the current block relative to a second reference block of a second reference picture having a second picture order count, wherein the first motion vector and the second motion vector each comprise a three-dimensional motion vector comprising an x-component in a plane of the current picture, a y-component in the plane of the current picture, and a z-component, wherein the z-component represents a temporal component; determining whether to refine the first motion vector based on a comparison of a difference between the z components of the first motion vector and the second motion vector and a third threshold, and determining to refine the first motion vector in response to the difference between the z components of the first motion vector and the second motion vector being less than the third threshold; In response to determining to refine the first motion vector, generating a first refined motion vector from the first motion vector; and Using the first refined motion vector, motion compensation is performed to derive an updated first reference block from the first reference picture.
14. The method according to claim 13, wherein: Determining whether to refine the first motion vector further includes comparing a difference between the x-components of the first motion vector and the second motion vector to a first threshold, and determining to refine the first motion vector in response to the difference between the x-components of the first motion vector and the second motion vector being less than the first threshold.
15. The method according to claim 14, further comprising: A difference between the x-components of the first motion vector and the second motion vector is generated based on a difference between an absolute value of the x-component of the first motion vector and an absolute value of the x-component of the second motion vector.
16. The method according to claim 13, wherein: Determining whether to refine the first motion vector further includes comparing a difference between the y components of the first motion vector and the second motion vector with a second threshold, and determining to refine the first motion vector in response to the difference between the y components of the first motion vector and the second motion vector being less than the second threshold.
17. The method according to claim 16, further comprising: A difference between the y components of the first motion vector and the second motion vector is generated based on a difference between an absolute value of the y component of the first motion vector and an absolute value of the y component of the second motion vector.
18. The method according to claim 13, wherein: the z component of the first motion vector comprises a difference between the current picture order count and the first picture order count, and the z component of the second motion vector comprises a difference between the current picture order count and the second picture order count; as well as The method also includes generating a difference between the z-components of the first motion vector and the second motion vector includes generating a difference between an absolute value of the z-component of the first motion vector and an absolute value of the z-component of the second motion vector.
19. The method according to claim 13, wherein: Generating the first refined motion vector includes: performing bidirectional optical flow (BIO) processing on the first motion vector.
20. A computer program product comprising computer executable instructions executable by a processor of a computing device and configured to perform the following operations, the computer executable instructions being for decoding a current block in a current picture of a video bitstream, the current picture having a current picture sequence count: decoding, from the video bitstream, a first motion vector for a first reference block of a first reference picture having a first picture order count for the current block; A second motion vector for a second reference block of a second reference picture having a second picture order count is decoded from the video bitstream, wherein: The first motion vector and the second motion vector each comprise a three-dimensional motion vector, the three-dimensional motion vector comprising an x component in a plane of the current picture, a y component in the plane of the current picture, and a z component, wherein the z component represents a temporal component; determining whether to refine the first motion vector based on a comparison of a difference between the z components of the first motion vector and the second motion vector and a third threshold, and determining to refine the first motion vector in response to the difference between the z components of the first motion vector and the second motion vector being less than the third threshold; In response to determining to refine the first motion vector, generating a first refined motion vector from the first motion vector; and Using the first refined motion vector, motion compensation is performed to derive an updated first reference block from the first reference picture.
21. A non-transitory computer-readable medium storing computer-executable instructions executable by a processor of a computing device and configured to perform the following operations, wherein the computer-executable instructions are for decoding a current block in a current picture of a video bitstream, the current picture having a current picture sequence count: decoding, from the video bitstream, a first motion vector for a first reference block of a first reference picture having a first picture order count for the current block; A second motion vector for a second reference block of a second reference picture having a second picture order count is decoded from the video bitstream, wherein: The first motion vector and the second motion vector each comprise a three-dimensional motion vector, the three-dimensional motion vector comprising an x component in a plane of the current picture, a y component in the plane of the current picture, and a z component, wherein the z component represents a temporal component; determining whether to refine the first motion vector based on a comparison of a difference between the z components of the first motion vector and the second motion vector and a third threshold, and determining to refine the first motion vector in response to the difference between the z components of the first motion vector and the second motion vector being less than the third threshold; In response to determining to refine the first motion vector, generating a first refined motion vector from the first motion vector; and Using the first refined motion vector, motion compensation is performed to derive an updated first reference block from the first reference picture.