Motion compensation in different directions
By employing optical flow derivation and interpolation methods different from those used in horizontal and vertical directions in video encoding and decoding, the inefficiency problem in existing technologies is solved, enabling efficient processing of texture regions in diagonal and anti-diagonal directions and improving encoding and decoding performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2021-01-26
- Publication Date
- 2026-05-26
Smart Images

Figure CN115104310B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] In accordance with the applicable Patent Law and / or the provisions of the Paris Convention, this application aims to promptly claim priority and benefit to International Patent Application No. PCT / CN2020 / 074052, filed on January 26, 2020. The entire disclosure of that International Patent Application No. PCT / CN2020 / 074052 is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to video encoding and decoding technologies, devices, and systems. Background Technology
[0004] Currently, efforts are underway to improve the performance of existing video codec technologies to provide better compression ratios or to offer video codec and decoding schemes that allow for lower complexity or parallel implementation. Industry experts have recently proposed several new video codec tools, which are currently being tested to determine their effectiveness. Summary of the Invention
[0005] This paper describes devices, systems, and methods related to digital video coding and decoding, specifically the management of motion vectors. The described methods can be applied to existing video coding and decoding standards (e.g., High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC)) and future video coding and decoding standards or codecs.
[0006] In one representative aspect, the disclosed technology can be used to provide a visual media processing method. The method includes: for a conversion between a current video block of visual media data and a bitstream representation of the current video block, determining one or more directional optical flows of a list of reference images associated with the current video block, wherein the one or more directional optical flows do not include horizontal and / or vertical directions.
[0007] In another representative aspect, the disclosed technology can be used to provide another method for visual media processing. This method includes: for a conversion between a current video block of visual media data and a bitstream representation of the current video block; determining one or more directional optical flows of a list of reference images associated with the current video block, wherein the one or more directional optical flows do not include horizontal and / or vertical directions; and using the one or more directional optical flows in multiple prediction refinements to generate a resulting prediction refinement.
[0008] In another representative aspect, the disclosed technology can be used to provide another method for visual media processing. This method includes: for a conversion between a current video block of visual media data and a bitstream representation of the current video block, selectively determining one or more directions or direction pairs included in a list of reference images associated with the current video block, wherein the one or more directional optical flows are used to generate predictive refinement, wherein the one or more directions or direction pairs vary from one region of the current video block to another.
[0009] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes: a conversion between a current video block and a bitstream representation of the current video block; determining an optical flow associated with the current video block during an optical flow-based motion refinement or prediction process, wherein the optical flow is derived along a direction different from the horizontal and / or vertical directions; and performing the conversion based on the optical flow.
[0010] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes: a conversion between a current video block and a bitstream representation of the current video block; determining, during an optical flow-based motion refinement or prediction process, a spatial gradient of a pair of orientations associated with the current video block, wherein the spatial gradient of the pair of orientations depends on the spatial gradients of the two directions of the pair; and performing the conversion based on the spatial gradient.
[0011] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes: a conversion between a current video block and a bitstream representation of the current video block; generating one or more prediction refinements associated with the current video block during an optical flow-based motion refinement or prediction process; generating a final prediction refinement associated with the current video block by combining multiple prediction refinements; and performing the conversion based on the final prediction refinement.
[0012] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes: for a conversion between a current video block and a bitstream representation of the current video block, determining a direction or direction pair associated with the current video block during an optical flow-based prediction refinement or prediction process, wherein the direction or direction pair changes from one video region of the current video block to another video region; and performing the conversion based on the direction or direction pair. In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes: for a conversion between a current video block and a bitstream representation of the current video block, performing interpolation of motion vectors associated with the current video block to generate an interpolation result during an optical flow-based motion refinement or prediction process, wherein the interpolation is performed along a direction different from the horizontal and / or vertical direction; and performing the conversion based on the interpolation result. In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes: performing interpolation of motion vectors associated with the current video block for a conversion between a current video block and a bitstream representation of the current video block, to generate one or more interpolation results in an optical flow-based motion refinement or prediction process; generating a final interpolation result associated with the current video block by combining multiple interpolation results; and performing the conversion based on the final interpolation result.
[0013] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes: for a conversion between a current video block and a bitstream representation of the current video block, performing interpolation of motion vectors associated with the current video block to generate an interpolation result in an optical flow-based motion refinement or prediction process, wherein the interpolation is performed along one or more directions or direction pairs, the one or more directions or direction pairs transforming from one video region of the current video block to another video region; and performing the conversion based on the interpolation result.
[0014] In another representative aspect, the disclosed technology can be used to provide a video processing method. The method includes: a conversion between a current video block and a bitstream representation of the current video block; determining an optical flow associated with the current video block during an optical flow-based motion refinement or prediction process, wherein the optical flow is derived along a direction different from the horizontal and / or vertical directions; generating the bitstream from the current video block based on the optical flow; and storing the bitstream in a non-transitory computer-readable storage medium.
[0015] Furthermore, in one representative aspect, an apparatus for a video system is disclosed, comprising a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform one or more of the disclosed methods.
[0016] In addition, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product including program code for implementing any one or more of the disclosed methods.
[0017] The above and other aspects and features of the present disclosure are described in more detail in the accompanying drawings, description and claims. Attached Figure Description
[0018] Figure 1 An example of interpolation of sample points at fractional locations is shown.
[0019] Figure 2 An example of an optical flow trajectory is shown.
[0020] Figure 3A and Figure 3B An example of BIO without block extension is shown.
[0021] Figure 4 An example of a sub-block MV VSB and a pixel is shown.
[0022] Figure 5 Examples of interpolation along the diagonal and anti-diagonal directions are shown.
[0023] Figure 6 A block diagram of an example hardware platform used to implement the visual media decoding or visual media encoding technologies described in this document.
[0024] Figure 7 A flowchart of an example method for video encoding and decoding is shown.
[0025] Figure 8 A flowchart of an example method for video encoding and decoding is shown.
[0026] Figure 9 A flowchart of an example method for video encoding and decoding is shown.
[0027] Figure 10 A flowchart of an example method for video encoding and decoding is shown.
[0028] Figure 11 A flowchart of an example method for video encoding and decoding is shown.
[0029] Figure 12 A flowchart of an example method for video encoding and decoding is shown.
[0030] Figure 13 A flowchart of an example method for video encoding and decoding is shown.
[0031] Figure 14 A flowchart of an example method for video encoding and decoding is shown.
[0032] Figure 15 A flowchart of an example method for video encoding and decoding is shown. Detailed Implementation
[0033] 1. Video encoding and decoding in HEVC / H.265
[0034] Video codec standards have primarily evolved from well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, which uses temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, a Joint Video Expert Team (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1SC29 / WG11 (MPEG) to study the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.
[0035] 2.1 Motion Compensation
[0036] In inter-frame encoding and decoding, if the block's motion vector points to a fractional position, the reference sample at the integer position is used to interpolate the reference sample at the fractional position. When the motion vector has fractional components in both the horizontal and vertical directions, the sample at the fractional horizontal position and the integer vertical position is interpolated first, and then used to interpolate the sample at the fractional horizontal position and the fractional vertical position. Figure 1 An example is shown below. Interpolation can be performed along the horizontal or vertical direction.
[0037] 2.2 Bidirectional optical flow
[0038] In BIO, motion compensation is first performed to generate the first prediction for the current block (in each prediction direction). The first prediction is used to derive the spatial gradient, temporal gradient, and optical flow for each sub-block / pixel within the block, and then it is used to generate the second prediction, which is the final prediction for the sub-block / pixel. A detailed description follows.
[0039] Bidirectional optical flow (BIO) is sample-level motion refinement performed on top of block-level motion compensation for bidirectional prediction. Sample-level motion refinement does not use signaling notification.
[0040] Will I (k) Let the brightness value after block motion compensation be the value from the reference k (k = 0, 1), and and I (k) The horizontal and vertical components of the gradient. Assuming optical flow is effective, the motion vector field (v) x ,v y The equation gives:
[0041]
[0042] Combining this optical flow equation with the Hermite interpolation of the motion trajectory at each sample point yields a unique third-order polynomial, which, along with the function values I at both ends... (k) and derivative and Simultaneous matching. The value of this polynomial at t=0 is the BIO prediction:
[0043]
[0044] Here, τ0 and τ1 represent the distances to the reference frame, such as... Figure 2 As shown. Distances τ0 and τ1 are calculated based on the POC of Ref0 and Ref1: τ0 = POC(current) - POC(Ref0) and τ1 = POC(Ref1) - POC(current). If both predictions originate from the same time direction (either both from the past or both from the future), they have different signs (i.e., τ0·τ1 < 0). In this case, BIO is applied only if the predictions do not originate from the same time (i.e., τ0 ≠ τ1), both reference regions have non-zero motion (MVx0, MVy0, MVx1, MVy1 ≠ 0), and the block motion vector is proportional to the time distance (MVx0 / MVx1 = MVy0 / MVy1 = -τ0 / τ1). To refine the forecast.
[0045] By minimizing points A and B ( Figure 9 The difference Δ between the values at the intersection of the motion trajectory and the reference frame plane is used to determine the motion vector field (v).x ,v y The model uses only the first linear term of the local Taylor expansion of Δ:
[0046]
[0047] All values in Equation 3 depend on the sample point location (i′,j′), which has been omitted from the notation so far. Assuming the motion is consistent in the local surrounding region, we minimize Δ, which lies within a (2M+1)×(2M+1) squared window Ω centered at the current predicted point (i,j), where M equals 2:
[0048]
[0049] For this optimization problem, JEM uses a simplification method, first minimizing in the vertical direction and then minimizing in the horizontal direction. This results in:
[0050]
[0051]
[0052] in,
[0053]
[0054]
[0055]
[0056] To avoid division by zero or very small values, regularization parameters r and m are introduced in Equations 5 and 6.
[0057] r = 500·4 d-8 (8)
[0058] m = 700·4 d-8 (9)
[0059] Here, d is the bit depth of the video sample.
[0060] To make BIO's memory access the same as regular bidirectional predictive motion compensation, all prediction and gradient values, I (k) , and Calculations are performed only for positions within the current block. In Equation 7, the (2M+1)×(2M+1) squared window Ω centered on the current prediction point on the boundary of the prediction block needs to access positions outside the block (e.g., ...). Figure 3A (As shown). In JEM, I outside the block (k) , and The value is set to be equal to the nearest available value within the block. For example, this can be implemented as padding, such as... Figure 3B As shown.
[0061] Using BIO, the motion field can be refined for each sample point. To reduce computational complexity, JEM uses a block-based BIO design. Motion refinement is calculated based on 4×4 blocks. In block-based BIO, the s in Formula 7 for all samples in the 4×4 block is... n Aggregate the values and then use s n The aggregated values are used to derive the BIO motion vector offset for a 4×4 block. More specifically, the following formula is used for block-based BIO derivation:
[0062]
[0063]
[0064]
[0065] Among them, b k Let represent the sample set belonging to the k-th 4×4 block of the prediction block. Combine s in Formulas 5 and 6. n Replace with (s) n,bk )>>4, to derive the associated motion vector offset.
[0066] In some cases, the MV Regiment of BIO may be unreliable due to noise or irregular motion. Therefore, in BIO, the size of the MV Regiment is cropped to a threshold thBIO. The threshold is determined based on whether all reference images of the current image come from the same direction. If all reference images of the current image come from the same direction, the threshold is set to 12×2. 14-d Otherwise, set it to 12×2 13-d .
[0067] The gradient of the BIO is computed simultaneously with motion compensation interpolation using the same operation as the HEVC motion compensation process (2D separable FIR). The input to this 2D separable FIR is the same reference frame sample as the input to the motion compensation process and the fractional positions (fracX, fracY) based on the fractional part of the block motion vector. (Horizontal gradient...) If the signal is first interpolated vertically using BIOfilterS corresponding to the fractional position fracY with a descaling displacement of d-8, then a gradient filter BIOfilterG is applied in the horizontal direction corresponding to the fractional position fracX with a descaling displacement of 18-d. In the vertical gradient... In this case, the gradient filter is first applied vertically using BIOfilterG corresponding to the fractional position fracY with a descaling displacement of d-8, and then the signal shift is performed using BIOfilterS in the horizontal direction corresponding to the fractional position fracX with a descaling displacement of 18-d. To maintain reasonable complexity, the interpolation filters for gradient computation BIOfilterG and signal shifting BIOfilterF are relatively short (6-tap). Table 1 shows the filters for gradient computation at different fractional positions of the block motion vector in BIO. Table 2 shows the interpolation filters for predictive signal generation in BIO.
[0068] Table 1. Filters for gradient calculation in BIO
[0069]
[0070]
[0071] Table 2 Interpolation filters for predictive signal generation in BIO
[0072] Fractional pixel position Interpolation filters for predictive signals (BIOfilters) 0 {0,0,64,0,0,0} 1 / 16 {1,-3,64,4,-2,0} 1 / 8 {1,-6,62,9,-3,1} 3 / 16 {2,-8,60,14,-5,1} 1 / 4 {2,-9,57,19,-7,2} 5 / 16 {3,-10,53,24,-8,2} 3 / 8 {3,-11,50,29,-9,2} 7 / 16 {3,-11,44,35,-10,3} 1 / 2 {3,-10,35,44,-11,3}
[0073] In JEM, BIO is applied to all bidirectional prediction blocks when two predictions come from different reference images. BIO is disabled when LIC is enabled for CU.
[0074] In JEM, OBMC is applied to blocks following the normal MC process. To reduce computational complexity, BIO is not applied during OBMC. This means that when using its own MV, BIO is only applied during the block's MC process, but when using the MV of an adjacent block during OBMC, BIO is not applied during the MC process.
[0075] Based on the similarity between two predicted signals, a two-stage early termination method is used to conditionally disable BIO operations. Early termination is first applied at the CU level and then at the sub-CU level. Specifically, the method first calculates the SAD between the L0 and L1 predicted signals at the CU level. Since BIO only applies to luminance, the SAD calculation only considers luminance samples. If the CU-level SAD is not greater than a predefined threshold, the BIO process is completely disabled for the entire CU. The CU-level threshold is set to 2 per sample. (BDepth-9) If the BIO process is not disabled at the CU level, and if the current CU contains multiple sub-CUs, the SAD (Severity Aspect Ratio) for each sub-CU within the CU is calculated. Then, based on a predefined sub-CU level SAD threshold set to 3*2 per sample point, a decision is made at the sub-CU level to enable or disable the BIO process. (BDepth-10) .
[0076] BIO is also known as BDOF (Bidirectional Optical Flow).
[0077] The specifications for BDOF are as follows:
[0078] 8.5.6.5 Bidirectional Optical Flow Prediction Process
[0079] The input to this process is:
[0080] --Two variables, nCbW and nCb, specify the width and height of the current encoding / decoding block;
[0081] --Two (nCbW+2)x(nCbH+2) brightness prediction sample arrays predSamplesL0 and predSamplesL1;
[0082] --The prediction list uses the flags predFlagL0 and predFlagL1;
[0083] --Refer to indices refIdxL0 and refIdxL1;
[0084] -- Bidirectional optical flow utilization flag sbBdofFlag.
[0085] The output of this process is an array of (nCbW)x(nCbH) sample values, pbSamples.
[0086] The derivation of variables shift1, shift2, shift3, shift4, offset4, and mvRefineThres is as follows:
[0087] -- Set the variable shift1 to equal 6.
[0088] -- Set the variable shift2 to equal 4.
[0089] -- Set the variable shift3 to equal 1.
[0090] -- The variable shift4 is set to equal Max(3, 15-BitDepth), and the variable offset4 is set to 1 << (shift4-1).
[0091] -- Set the variable mvRefineThres to equal 1 << 4.
[0092] For xIdx=0..(nCbW>>2)–1 and yIdx=0..(nCbH>>2)–1, the application is as follows:
[0093] -- The variable xSb is set to equal (xIdx<<2)+1, and the variable ySb is set to equal (yIdx<<2)+1.
[0094] --If sbBdofFlag equals FALSE, for x = xSb-1..xSb+2 and y = ySb-1..ySb+2, the derivation of the predicted sample values for the current sub-block is as follows:
[0095] pbSamples[x][y] = Clip3(0,(2) BitDepth )-1,(predSamplesL0[x+1][y+1]+offset4+predSamplesL1[x+1][y+1])>>shift4) (987)
[0096] --Otherwise (sbBdofFlag equals TRUE), the derivation of the predicted sample values for the current sub-block is as follows:
[0097] For x = xSb-1..xSb+4 and y = ySb-1..ySb+4, apply the following sequence of steps:
[0098] 1. Predict the position (h) of each corresponding sample point (x, y) within the sample point array. x ,v y The derivation is as follows:
[0099] h x =Clip3(1,nCbW,x) (988)
[0100] v y =Clip3(1,nCbH,y) (989)
[0101] 2. The derivation of variables gradientHL0[x][y], gradientVL0[x][y], gradientHL1[x][y], and gradientVL1[x][y] is as follows:
[0102] gradientHL0[x][y]=(predSamplesL0[h x +1][v y ]>>shift1)–(predSampleL0[h x -1][v y ])>>shift1)(990)
[0103] gradientVL0[x][y]=(predSampleL0[h x ][v y +1]>>shift1)–(predSampleL0[h x ][v y-1])>>shift1)(991)
[0104] gradientHL1[x][y] = (predSamplesL1[h x +1][v y >>shift1) – (predSampleL1[h x -1][v y )>>shift1)(992)
[0105] gradientVL1[x][y] = (predSampleL1[h x [v y +1]>>shift1) – (predSampleL1[h x [v y -1])>>shift1)(993)
[0106] 3. The variables diff[x][y], tempH[x][y], and tempV[x][y] are derived as follows:
[0107] diff[x][y] = (predSamplesL0[h x [v y >>shift2) - (predSamplesL1[h x [v y >>shift2) (994)
[0108] tempH[x][y] = (gradientHL0[x][y] + gradientHL1[x][y])>>shift3 (995)
[0110] tempV[x][y] = (gradientVL0[x][y] + gradientVL1[x][y])>>shift3 (996)
[0112] The variables sGx2, sGy2, sGxGy, sGxdI, and sGydI are derived as follows:
[0113] sGx2 = Σ i Σ j Abs(tempH[xSb + i][ySb + j]) with i, j = -1..4 (997)
[0114] sGy2 = Σ i Σ jAbs(tempV[xSb+i][ySb+j])with i,j=-1..4 (998)
[0115] sGxGy=Σ i Σ j (Sign(tempV[xSb+i][ySb+j])*tempH[xSb+i][ySb+j])withi,j=-1..4 (999)
[0116] sGxdI=Σ i Σ j (-Sign(tempH[xSb+i][ySb+j])*diff[xSb+i][ySb+j])withi,j=-1..4 (1000)
[0117] sGydI=Σ i Σ j (-Sign(tempV[xSb+i][ySb+j])*diff[xSb+i][ySb+j])withi,j=-1..4 (1001)
[0118] The horizontal and vertical motion offsets of the current sub-block are derived as follows:
[0119] v x =sGx2>0? Clip3(-mvRefineThres+1,mvRefineThres-1,(sGxdI<<2)>>Floor(Log2(sGx2))):0 (1002)
[0120] v y =sGy2>0? Clip3(-mvRefineThres+1,mvRefineThres-1,
[0121] ((sGydI<<2)-((v x *sGxGy)>>1))>>Floor(Log2(sGy2))):0 (1003)
[0123] – For x = xSb-1..xSb+2 and y = ySb-1..ySb+2, the predicted sample values for the current sub-block are derived as follows:
[0124] bdofOffset = v x *(gradientHL0[x+1][y+1]-gradientHL1[x+1][y+1])+v y*(gradientVL0[x+1][y+1]-gradientVL1[x+1][y+1]) (1004)
[0126] pbSamples[x][y] = Clip3(0,(2) BitDepth )-1,(predSamplesL0[x+1][y+1]+offset4+predSamplesL1[x+1][y+1]+bdofOffset)>>shift4) (1005)
[0128] 2.3 Refinement of Optical Flow Prediction
[0129] This paper proposes a method for refining sub-block-based affine motion compensation prediction using optical flow. After performing sub-block-based affine motion compensation, the predicted samples are refined by adding differences derived from the optical flow equation; this is called optical flow prediction refinement (PROF). This method can achieve pixel-level granularity inter-frame prediction without increasing memory access bandwidth.
[0130] To achieve finer motion compensation granularity, this paper proposes a method for refining sub-block-based affine motion compensation predictions using optical flow. After performing sub-block-based affine motion compensation, the brightness prediction samples are refined by adding differences derived from the optical flow equation. The proposed optical flow prediction refinement (PROF) is described in the following four steps.
[0131] Step 1): Perform sub-block-based affine motion compensation to generate sub-block prediction I(i,j).
[0132] Step 2): Calculate the spatial gradient g of the sub-block prediction at each sample location using a 3-tap filter [-1, 0, 1]. x (i,j) and g y (i,j).
[0133] g x (i,j)=I(i+1,j)-I(i-1,j)
[0134] g y (i,j)=I(i,j+1)-I(i,j-1)
[0135] For gradient calculation, sub-block predictions are extended by one pixel on each side. To reduce memory bandwidth and complexity, pixels on the extended boundary are copied from the nearest integer pixel position in the reference image. This avoids additional interpolation in the filled regions.
[0136] Step 3): Calculate the brightness prediction refinement (denoted as ΔI) using the optical flow equation.
[0137] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j)
[0138] Where delta MV (denoted as Δv(i,j)) is the pixel MV calculated at sample position (i,j) (where v(i,j) represents the difference between the pixel MV and the sub-block MV of the sub-block to which pixel (i,j) belongs, such as...). Figure 4 As shown. Figure 4 In the diagram, delta MV is represented by a small arrow.
[0139] Since the affine model parameters and the pixel position relative to the sub-block center do not change between sub-blocks, Δv(i,j) for the first sub-block can be calculated and reused for other sub-blocks in the same CU. Assuming x and y are the horizontal and vertical offsets from the pixel position to the sub-block center, Δv(x,y) can be derived using the following formula:
[0140] Δv x (x,y)=c*x+d*y
[0141] Δv y (x,y)=e*x+f*y
[0142] For a 4-parameter affine model:
[0143]
[0144]
[0145] For a 6-parameter affine model:
[0146]
[0147] Among them, (v 0x ,v 0y ),(v 1x ,v 1y ),(v 2x ,v 2y ) represents the motion vectors of the control points at the top left, top right, and bottom left, and w and h are the width and height of the CU.
[0148] Step 4): Finally, refine the brightness prediction and add it to the sub-block prediction I(i,j). The final prediction I' is generated by the following formula:
[0149] I′(i,j)=I(i,j)+ΔI(i,j)
[0150] 3. Disadvantages of existing implementation methods
[0151] The current designs of BDOF, PROF, and motion compensation have the following problems:
[0152] 1) In BDOF and PROF, the optical flow is derived along the horizontal and / or vertical direction. If the actual optical flow is along other directions, such as diagonal or anti-diagonal directions, it may be less efficient.
[0153] 2) In motion compensation, when interpolation of fractional samples is required, the interpolation is always performed along the horizontal and / or vertical direction. This can be inefficient for regions containing oriented textures along other directions (e.g., diagonal or anti-diagonal directions).
[0154] 4. Example technologies and implementation examples
[0155] The detailed embodiments described below should be considered as examples to explain general concepts. These embodiments should not be interpreted narrowly. Furthermore, these embodiments can be combined in any way.
[0156] In the following discussion, for a list of reference images X (X = 0, 1), the horizontal and vertical optical flows derived in motion-based or predictive thinning processes (e.g., BDOF, PROF) are represented as ofX. h (x,y) and ofX v (x,y). For example, of0 h (x,y) and of0 v (x, y) can be referenced from the reference image list 0, which can be referenced from v. x / 2 and v y / 2, for reference image list 1, please refer to -v x / 2 and -v y / 2, where for BDOF, "v x and v y "Defined in formulas 1002 and 1003 of section 8.5.6.5. And, ofX..." h (x,y) and ofX v (x,y) can be referenced in PROF under "Δv". x (x,y) and Δv y (x,y)”, where Δv is derived for each valid list of reference images. x (x,y) and Δv y (x,y).
[0157] In the following, "diagonal direction" refers to the horizontal direction rotated M degrees counterclockwise, and "anti-diagonal direction" refers to the vertical direction rotated N degrees counterclockwise. In one example, M and / or N equals 45. In one example, a direction pair can include two directions, such as horizontal and vertical directions or diagonal and anti-diagonal directions. The diagonal and anti-diagonal optical flows in the reference image list X (X = 0, 1) are represented as ofX. d (x,y) and ofX ad (x,y).
[0158] Let PX(x,y) be the predicted sample point of the sample point (x,y) in the reference image list X (X=0,1), and let gradX be the horizontal and vertical gradients of PX(x,y). h (x,y) and gradX v The diagonal and anti-diagonal gradients of PX(x,y) are denoted as gradX. d (x,y) and gradX ad (x,y).
[0159] The proposed method for PROF / BDOF can be applied to other types of encoding and decoding methods that use optical flow.
[0160] 1. It is proposed that optical flow can be derived along directions different from the horizontal and / or vertical directions.
[0161] a. Alternatively, the spatial gradient can be derived along the same direction used to derive the optical flow.
[0162] b. Alternatively, optical flow and spatial gradients derived in these directions can be used to generate prediction refinements.
[0163] c. In one example, optical flow and / or spatial gradient can be derived along the diagonal and anti-diagonal directions.
[0164] i. Alternatively, optical flow and / or spatial gradient can be derived for a pair of directions.
[0165] 2. It is proposed that the spatial gradient of a direction pair can depend on the spatial gradients of the two directions of the direction pair. For example, the spatial gradient can be calculated as a function of the gradients in the two directions.
[0166] a. In one example, the spatial gradient of a direction pair can be computed as the sum or weighted sum of the absolute gradients in both directions.
[0167] i. For example, the spatial gradient of the horizontal-vertical pair can be calculated as the sum of the absolute horizontal gradient and the absolute vertical gradient.
[0168] ii. For example, the spatial gradient of the diagonal-anti-diagonal direction pair can be calculated as the sum of the absolute diagonal gradient and the absolute anti-diagonal gradient.
[0169] b. In one example, the spatial gradient of a direction pair can be calculated as the larger or smaller or average of the absolute gradients in the two directions.
[0170] c. Alternatively, the spatial gradient of the direction pair can be used to determine which direction pair to choose for prediction refinement.
[0171] 3. A method is proposed to combine multiple prediction refinements to generate the final prediction refinement.
[0172] a. In one example, multiple prediction refinements can be derived in multiple directions or direction pairs.
[0173] b. In one example, the first prediction refinement can be derived in the horizontal-vertical direction, and the second prediction refinement can be derived in the diagonal-anti-diagonal direction.
[0174] i. In one example, the first prediction refinement of the reference image list X (X = 0, 1) can be defined as: ofX h (x,y)×gradX h (x,y)+ofX v (x,y)×gradX v (x,y).
[0175] ii. In one example, the second prediction refinement of the reference image list X (X = 0, 1) can be defined as: ofX d (x,y)×gradX d (x,y)+ofX ad (x,y)×gradX ad (x,y).
[0176] c. In one example, multiple prediction refinements can be weighted and averaged to generate the final prediction refinement.
[0177] i. In one example, the weights may depend on the gradient information of the prediction block.
[0178] (i) For example, spatial gradients can be computed for multiple orientation pairs, and smaller weights can be assigned to orientation pairs with smaller spatial gradients.
[0179] (ii) For example, spatial gradients can be computed for multiple orientation pairs, and smaller weights can be assigned to orientation pairs with larger spatial gradients.
[0180] ii. In one example, the weight of the first sample point in the first prediction refinement block may be different from the weight of the second sample point in the first prediction refinement block.
[0181] iii. In one example, default weights can be assigned to multiple prediction refinements. For example, 3 / 4 can be used for the first prediction refinement, and 1 / 4 can be used for the second prediction refinement.
[0182] iv. In one example, a final prediction refinement can be generated for each list of reference images X.
[0183] d. In one example, the reliability of optical flow can be defined, and the weights used for multiple prediction refinements can depend on the reliability of multiple optical flows.
[0184] i. In one example, in the case of bidirectional prediction, the predicted samples, optical flow, and spatial gradient of the predicted samples can be used to generate refined predicted samples in the reference image list X (X = 0, 1).
[0185] (i) For example, the refined prediction sample can be generated as the sum of the prediction sample and the prediction refinement.
[0186] (ii) For example, for a horizontal-vertical pair, the refined prediction samples in the reference image list X can be generated as: PX(x,y)+ofX h (x,y)×gradX h (x,y)+ofX v (x,y)×gradX v (x,y).
[0187] (iii) For example, for a diagonal-anti-diagonal direction pair, the refined prediction samples in the reference image list X can be generated as: PX(x,y)+ofX d (x,y)×gradX d (x,y)+ofX ad (x,y)×gradX ad (x,y).
[0188] ii. Reliability may depend on the difference between the refined predictions in two lists of reference images in bidirectional predictive coding.
[0189] (i) In one example, the reliability of each pixel can be derived.
[0190] (ii) In one example, the reliability of each block can be derived.
[0191] (iii) In one example, the reliability of each sub-block can be derived.
[0192] (iv) In one example, when deriving the reliability of a block or sub-block, the difference of some representative samples can be calculated.
[0193] (v) In one example, the difference can be the sum of absolute differences (SAD), the sum of squared errors (SSE), or the sum of absolute transformation differences (SATD).
[0194] (vi) In one example, higher reliability can be assigned to optical flow with a smaller difference between the refined predictions in two lists of reference images.
[0195] iii. In one example, a larger weight can be assigned to the refinement of predictions generated from optical flow with higher reliability.
[0196] (i) Alternatively, the weights can also depend on whether the prediction refinement comes from a horizontal-vertical direction pair or a diagonal-anti-diagonal direction pair. For example, assuming that the optical flow has the same reliability for both direction pairs, a larger weight can be assigned to the prediction refinement generated from the horizontal-vertical direction pair.
[0197] e. In one example, the weight of each sample point can be derived.
[0198] f. In one example, the weight of each block or sub-block (e.g., a 4×4 block) can be derived.
[0199] 4. Instead of performing optical flow-based prediction refinement along a fixed direction or direction pair, the direction or direction pair can be transformed from one video region to another.
[0200] a. In one example, a direction pair can be determined first, and an optical flow-based prediction refinement process can be performed along the determined direction pair.
[0201] b. In one example, the gradient of the predicted block can be used to determine the orientation pair.
[0202] c. For example, spatial gradients can be computed for multiple direction pairs, and prediction refinement can be performed on the direction pairs with the minimum spatial gradient.
[0203] d. For example, spatial gradients can be computed for multiple direction pairs, and prediction refinement can be performed on the direction pair with the largest spatial gradient.
[0204] 5. Propose that interpolation can be performed in directions different from the horizontal and / or vertical directions.
[0205] a. In one example, interpolation can be performed along two directions that are orthogonal to each other but different from the horizontal and vertical directions.
[0206] b. In one example, interpolation can be performed along the diagonal direction and / or the anti-diagonal direction.
[0207] c. In one example, an interpolation filter different from the one used in horizontal / vertical interpolation can be used for this type of direction.
[0208] d. In one example, when the motion vector contains fractional components in both the diagonal and anti-diagonal directions, intermediate samples can be interpolated first along the diagonal direction and then used to interpolate predicted samples along the anti-diagonal direction. An example is shown below. Figure 5 As shown.
[0209] i. Alternatively, when the motion vector contains fractional components in both the diagonal and anti-diagonal directions, intermediate samples can be interpolated first along the diagonal direction and then used to interpolate predicted samples along the anti-diagonal direction.
[0210] e. Motion vectors can be represented and / or signaled using coordinates along the interpolation direction and / or orthogonal to the interpolation direction.
[0211] f. The accuracy of motion vectors may differ when interpolation is performed in different directions.
[0212] 6. Propose combining multiple interpolation results to generate the final interpolation result.
[0213] a. In one example, multiple interpolation results can be derived in multiple directions or direction pairs.
[0214] b. In one example, a first interpolation result can be generated in the horizontal-vertical direction, and a second interpolation result can be derived in the diagonal-anti-diagonal direction.
[0215] c. In one example, multiple interpolation results can be weighted and averaged to generate the final interpolation result.
[0216] i. In one example, the weights may depend on the gradient information of the reference block.
[0217] (i) For example, spatial gradients can be computed for multiple orientation pairs, and smaller weights can be assigned to orientation pairs with smaller spatial gradients.
[0218] (ii) For example, spatial gradients can be computed for multiple orientation pairs, and smaller weights can be assigned to orientation pairs with larger spatial gradients.
[0219] ii. In one example, the weight of the first sample point in the first interpolation block may be different from the weight of the second sample point in the first interpolation block.
[0220] iii. In one example, the weight of each sample point can be derived.
[0221] iv. In one example, the weight of each block or sub-block (e.g., a 4×4 block) can be derived.
[0222] d. In one example, default weights can be assigned to multiple interpolation results. For example, 3 / 4 can be used for the first interpolation result, and 1 / 4 can be used for the second interpolation result.
[0223] 7. Instead of performing interpolation along a fixed direction or direction pair, it can change the direction or direction pair from one video region to another.
[0224] a. In one example, a direction pair can be determined first, and the interpolation process can be performed along the determined direction pair.
[0225] b. In one example, the gradient of the reference block can be used to determine the orientation pair.
[0226] i. For example, spatial gradients can be computed for multiple direction pairs, and interpolation can be performed on the direction pairs with the minimum spatial gradient.
[0227] ii. For example, spatial gradients can be computed for multiple direction pairs, and interpolation can be performed on the direction pair with the largest spatial gradient.
[0228] c. In one example, when the motion vector has only fractional components in one of the diagonal and anti-diagonal directions, interpolation can be performed on the diagonal-anti-diagonal direction pair.
[0229] 8. Whether and / or how the methods described above can be applied for signaling notification, whether explicitly or implicitly, or whether it depends on the encoding / decoding information.
[0230] a. In one example, it can be applied to certain block sizes / shapes and / or certain sub-block sizes, color components.
[0231] b. The proposed method can be applied to certain block sizes.
[0232] i. In one example, it can be applied only to blocks where max(W,H) / min(W,H)<=T, where W and H are the width and height of the current block.
[0233] ii. In one example, it can be applied only to blocks where max(W,H) / min(W,H)>=T, where W and H are the width and height of the current block.
[0234] iii. In one example, it can be applied only to blocks where W×H<=T, where W and H are the width and height of the current block.
[0235] iv. In one example, it can be applied only to blocks where H <= T or H == T, where W and H are the width and height of the current block.
[0236] v. In one example, it can be applied only to blocks where W <= T or W == T, where W and H are the width and height of the current block.
[0237] vi. In one example, it can be applied only to blocks where W <= T1 and H <= T2, where W and H are the width and height of the current block.
[0238] vii. In one example, it can be applied only to blocks where W>=T1 and H>=T2, where W and H are the width and height of the current block.
[0239] viii. In one example, it can be applied only to blocks where W×H>=T, where W and H are the width and height of the current block.
[0240] ix. In one example, it can be applied only to blocks where H>=T, where W and H are the width and height of the current block.
[0241] x. In one example, it can only be applied to blocks where W>=T, where W and H are the width and height of the current block.
[0242] c. The proposed method can be applied to certain color components, such as only the luminance component.
[0243] 5. Example implementations of the disclosed technology
[0244] Figure 6 This is a block diagram of a video processing apparatus 600. Apparatus 600 can be used to implement one or more methods described in this disclosure. Apparatus 600 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 600 may include one or more processors 602, one or more memories 604, and video processing hardware 606. Processor 602 can be configured to implement one or more methods described in this document. Memory 604 can be used to store data and code for implementing the methods and techniques described herein. In hardware circuitry, video processing hardware 606 can be used to implement some of the techniques described in this document and may be partly or entirely part of processor 602 (e.g., a graphics processing unit (GPU) or other signal processing circuitry).
[0245] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of the video to the corresponding bitstream, and vice versa. The bitstream representation of the current video block may, for example, correspond to bits co-located or distributed at different positions within the bitstream, as defined in the syntax. For example, macroblocks may be encoded based on the error residuals from the transform and encoding / decoding, and may also utilize bits in the header and other fields in the bitstream.
[0246] It should be understood that by allowing the use of the techniques disclosed in this document, the disclosed methods and techniques will be beneficial for incorporation into video encoder and / or decoder embodiments in video processing devices (such as smartphones, laptops, desktops and similar devices).
[0247] Figure 7 A flowchart of an example method 700 for video processing. Method 700 includes: at 702, determining one or more directional optical flows of a list of reference images associated with the current video block for a conversion between a current video block of visual media data and a bitstream representation of the current video block, wherein the one or more directional optical flows do not include the horizontal and / or vertical directions.
[0248] Some embodiments can be described using the following clause-based format.
[0249] 1. A visual data processing method, comprising:
[0250] For the conversion between the current video block and the bitstream representation of the current video block in the visual media data, determine one or more directional optical flows of a list of reference images associated with the current video block, wherein the one or more directional optical flows do not include the horizontal and / or vertical directions.
[0251] 2. The method according to claim 1, further comprising:
[0252] One or more directional optical flows are used to determine the spatial gradient associated with the motion vector of the current video block.
[0253] 3. The method according to claim 2, wherein the spatial gradient is oriented along the direction of optical flow in one or more directions.
[0254] 4. The method according to any one or more of claims 1-2, wherein one or more directional optical flows and / or spatial gradients are oriented toward a diagonal direction generated by rotation in the horizontal direction.
[0255] 5. The method according to any one or more of claims 1-2, wherein one or more directional optical flows and / or spatial gradients are oriented toward the anti-diagonal direction generated by rotation in the vertical direction.
[0256] 6. The method according to any one or more of claims 4-5, wherein the horizontal rotation and / or vertical rotation is clockwise.
[0257] 7. The method according to any one or more of claims 4-5, wherein the horizontal rotation and / or vertical rotation is counterclockwise.
[0258] 8. The method of claim 1, wherein one or more directional optical flows are oriented along a pair of directions.
[0259] 9. The method of claim 3, wherein the spatial gradient is oriented along a pair of directions.
[0260] 10. The method of claim 9, wherein the spatial gradient of the pair depends on the spatial gradients of the two directions included in the pair.
[0261] 11. The method of claim 10, wherein the spatial gradient of the pair is the sum or weighted sum of the absolute gradients in the two directions.
[0262] 12. The method of claim 11, wherein the spatial gradient of the pair is the sum or weighted sum of the absolute horizontal gradient and the absolute vertical gradient.
[0263] 13. The method of claim 11, wherein the spatial gradient of the pair is the sum or weighted sum of the absolute diagonal gradient and the absolute antidiagonal gradient.
[0264] 14. The method of claim 10, wherein the spatial gradient of the pair is the larger, smaller, or average value of the absolute gradients in the two directions.
[0265] 15. The method according to any one or more of claims 10-14, wherein the spatial gradient of the pair is used in the prediction refinement step.
[0266] 16. A visual data processing method, comprising:
[0267] For the conversion between the current video block and its bitstream representation in visual media data, determine one or more directional optical flows of a list of reference images associated with the current video block, wherein the one or more directional optical flows do not include the horizontal and / or vertical directions; and
[0268] One or more directional optical flows are used to generate the resulting prediction refinement in multiple prediction refinements.
[0269] 17. The method of claim 16, wherein the plurality of prediction refinements are associated with a plurality of directions or a plurality of direction pairs.
[0270] 18. The method according to any one or more of claims 16-17, wherein the plurality of prediction refinements includes a first prediction refinement in a horizontal-vertical direction pair and a second prediction refinement in a diagonal-anti-diagonal direction pair.
[0271] 19. The method according to claim 18, wherein the first prediction refinement of the reference image list X (X = 0,1) is defined as 〖ofX〗_h(x,y)×〖gradX〗_h(x,y)+〖ofX〗_v(x,y)×〖gradX〗_v(x,y), and the second prediction refinement of the reference image list X (X = 0,1) is defined as 〖ofX〗_d(x,y)×〖gradX〗_d(x,y)+〖ofX〗_ad(x,y)×〖gradX〗_ad(x,y).
[0272] 20. The method of claim 16, wherein the result prediction refinement is a weighted average of multiple prediction refinements.
[0273] 21. The method of claim 20, wherein the assigned weights are based on gradient information of the current video block.
[0274] 22. The method of claim 21, wherein smaller weights are assigned to direction pairs having smaller spatial gradients.
[0275] 23. The method of claim 21, wherein larger weights are assigned to direction pairs having larger spatial gradients.
[0276] 24. The method of claim 21, wherein the weight assigned to the first sample point in the current video block is different from the weight assigned to the second sample point in the current video block.
[0277] 25. The method of claim 21, wherein one or more default weights are assigned to a plurality of prediction refinements.
[0278] 26. The method of claim 25, wherein a default weight of 3 / 4 is assigned to the first prediction refinement, and a default weight of 1 / 4 is assigned to the second prediction refinement.
[0279] 27. The method of claim 16, wherein the result prediction refinement is associated with each list of reference images.
[0280] 28. The method of claim 16, wherein one or more directional optical flow reliability measures are defined, and the resulting prediction refinement is a weighted average of multiple prediction refinements, such that the weights assigned to the multiple prediction refinements are based on the reliability measures.
[0281] 29. The method of claim 28, wherein the result prediction refinement is based on prediction samples, directional optical flow in one or more directional optical flows, and at least one prediction refinement.
[0282] 30. The method of claim 28, wherein the result prediction refinement is the sum of the prediction samples and at least one prediction refinement.
[0283] 31. The method according to claim 28, wherein, for the horizontal-vertical direction pair, the result prediction refinement in the reference image list X is generated as 〖PX(x,y)+ofX〗_h(x,y)×〖gradX〗_h(x,y)+〖ofX〗_v(x,y)×〖gradX〗_v(x,y).
[0284] 32. The method according to claim 28, wherein, for the diagonal-anti-diagonal direction pair, the result prediction refinement in the reference image list X is generated as 〖PX(x,y)+ofX〗_d(x,y)×〖gradX〗_d(x,y)+〖ofX〗_ad(x,y)×〖gradX〗_ad(x,y).
[0285] 33. The method of claim 28, wherein, for the bidirectional prediction mode, the reliability metric is based on the difference between a first result prediction refinement and a second result prediction refinement associated with two lists of reference images.
[0286] 34. The method of claim 28, wherein a reliability metric is calculated for each pixel.
[0287] 35. The method of claim 28, wherein a reliability metric is calculated for each block.
[0288] 36. The method of claim 28, wherein a reliability metric is calculated for each sub-block.
[0289] 37. The method of claim 28, wherein a reliability metric is calculated for a subset of samples in a block or sub-block.
[0290] 38. The method of claim 33, wherein the difference is one of the following: sum of absolute differences (SAD), sum of squared errors (SSE), or sum of absolute transformation differences (SATD).
[0291] 39. The method of claim 33, wherein a higher reliability metric is calculated when the difference between the first result prediction refinement and the second result prediction refinement is smaller.
[0292] 40. The method of claim 28, wherein a larger weight is assigned to the directional optical flow having a higher reliability metric.
[0293] 41. The method of claim 28, wherein different weights are assigned to the multiple prediction refinements for the same reliability metric.
[0294] 42. The method of claim 28, wherein weights are assigned to each sample point.
[0295] 43. The method of claim 28, wherein weights are assigned to each block.
[0296] 44. The method of claim 28, wherein weights are assigned to each sub-block.
[0297] 45. A visual data processing method, comprising:
[0298] For the conversion between the current video block and the bitstream representation of the current video block of visual media data, one or more directions or direction pairs are selectively determined from the directional optical flow included in the list of reference images associated with the current video block, wherein the one or more directional optical flows are used to generate predictive refinement, wherein the one or more directions or direction pairs vary from one region of the current video block to another region.
[0299] 46. The method of claim 45, wherein prediction refinement is generated after the orientation pair is determined.
[0300] 47. The method of claim 46, wherein the spatial gradient information of the current video block is used to determine the orientation pair.
[0301] 48. The method of claim 47, wherein the generated prediction is refined along the direction having the minimum spatial gradient.
[0302] 49. The method of claim 47, wherein the generated prediction is refined along the direction having the maximum spatial gradient.
[0303] 50. The method of claim 45, further comprising:
[0304] In response to determining that one or more directions or direction pairs are oriented neither horizontally nor vertically, interpolation is performed along one or more directions or direction pairs.
[0305] 51. The method of claim 50, wherein the interpolation is performed along directions that are different from the horizontal and vertical directions and orthogonal to each other.
[0306] 52. The method of claim 50, wherein the interpolation is performed along the diagonal direction and / or the anti-diagonal direction.
[0307] 53. The method according to any one or more of claims 50-52, wherein the interpolation filter is used to perform interpolation.
[0308] 54. The method of claim 52, wherein when the motion vector of the current video block overlaps in the diagonal and anti-diagonal directions, interpolation is first performed along the diagonal direction, and then interpolation is performed along the anti-diagonal direction.
[0309] 55. The method of claim 54, wherein a first interpolation is performed on the intermediate sample points along the diagonal direction, and a second interpolation is performed on the predicted sample points along the anti-diagonal direction, wherein the second interpolation uses the intermediate sample points.
[0310] 56. The method of claim 50, wherein multiple interpolations are performed along one or more directions or direction pairs.
[0311] 57. The method of claim 56, wherein the precision of the motion vectors along one or more directions or directions is different.
[0312] 57. The method of claim 56, wherein multiple interpolations are combined to generate the resulting interpolation.
[0313] 58. The method of claim 56, wherein the interpolation results for different directions or pairs of directions are different, thereby generating a first interpolation result for horizontal-vertical direction pairs and a second interpolation result for diagonal-anti-diagonal direction pairs.
[0314] 59. The method of claim 56, wherein the resulting interpolation is a weighted average of multiple interpolations.
[0315] 60. The method of claim 59, wherein the weights used in the weighted average depend on the gradient information of the current video block.
[0316] 61. The method of claim 60, wherein the weights are allocated differently.
[0317] 62. The method of claim 60, wherein weights are assigned to each sample point.
[0318] 63. The method of claim 60, wherein weights are assigned to each block.
[0319] 64. The method of claim 60, wherein weights are assigned to each sub-block.
[0320] 65. The method of claim 58, wherein a default weight of 3 / 4 is assigned to the first interpolation result, and a default weight of 1 / 4 is assigned to the second interpolation result.
[0321] 66. The method of claim 56, wherein the plurality of interpolations are selectively varied from one region of the current video block to another region.
[0322] 67. The method according to any one or more of claims 1-66, wherein the shape, size, and / or color components of the current video block or its sub-blocks are used to determine the applicability of the method.
[0323] 68. The method of claim 67, wherein an indication of the applicability of the method is included as a field in the bitstream representation.
[0324] 69. The method according to any one or more of claims 1-68, wherein the conversion includes generating a bitstream representation from the current video block.
[0325] 70. The method according to any one or more of claims 1-68, wherein the conversion includes generating pixel values of the current video block from the bitstream representation.
[0326] 71. A video encoder apparatus, comprising a processor configured to implement the method of any one or more of claims 1-68.
[0327] 72. A video decoder apparatus, comprising a processor configured to implement the method of any one or more of claims 1-68.
[0328] 73. A computer-readable medium having code stored thereon, the code embodying processor-executable instructions for implementing the method of any one or more of claims 1-68.
[0329] Figure 8 A flowchart of an example method 800 for video processing. Method 800 includes: at 802, a conversion between a current video block and a bitstream representation of the current video block, determining the optical flow associated with the current video block during an optical flow-based motion refinement or prediction process, wherein the optical flow is derived along a direction different from the horizontal and / or vertical directions; and at 804, performing the conversion based on the optical flow.
[0330] In some examples, the spatial gradient associated with the current video block is derived along the same direction used to derive the optical flow.
[0331] In some examples, optical flow and spatial gradients derived in the direction are used to generate a prediction refinement associated with the current video patch.
[0332] In some examples, the optical flow and / or spatial gradient are derived along the diagonal and anti-diagonal directions, where the diagonal direction refers to the horizontal direction rotated M degrees counterclockwise, and the anti-diagonal direction refers to the vertical direction rotated N degrees counterclockwise, where M and N are integers.
[0333] In some examples, M and / or N equals 45.
[0334] In some examples, an optical flow and / or spatial gradient of a direction pair is derived, wherein a direction pair includes two directions, which include horizontal and vertical directions or diagonal and anti-diagonal directions.
[0335] Figure 9 A flowchart of an example method 900 for video processing. Method 900 includes: at 902, a conversion between a current video block and a bitstream representation of the current video block, determining the spatial gradient of a pair of orientations associated with the current video block during an optical flow-based motion refinement or prediction process, wherein the spatial gradient of the pair of orientations depends on the spatial gradients of the two orientations of the pair; and at 904, performing the conversion based on the spatial gradients.
[0336] In some examples, the spatial gradient of a direction pair is computed as a function of the spatial gradients in both directions of the direction pair.
[0337] In some examples, the spatial gradient of a direction pair is computed as the sum or weighted sum of the absolute gradients in both directions of the direction pair.
[0338] In some examples, the direction pair includes a horizontal and a vertical direction, and the spatial gradient of the direction pair is calculated as the sum of the absolute horizontal gradient and the absolute vertical gradient.
[0339] In some examples, the direction pair includes a diagonal direction and an anti-diagonal direction, and the spatial gradient of the direction pair is calculated as the sum of the absolute diagonal gradient and the absolute anti-diagonal gradient.
[0340] In some examples, the spatial gradient of a direction pair is calculated as the larger or smaller or average of the absolute gradients in the two directions of the direction pair.
[0341] In some examples, the spatial gradient of the orientation pair is used to determine which orientation pair to select to perform prediction refinement associated with the current video patch.
[0342] Figure 10A flowchart of an example method 1000 for video processing is provided. Method 1000 includes: at 1002, for a conversion between a current video block and a bitstream representation of the current video block, generating one or more prediction refinements associated with the current video block during an optical flow-based motion refinement or prediction process; at 1004, generating a final prediction refinement associated with the current video block by combining multiple prediction refinements; and at 1006, performing a conversion based on the final prediction refinement.
[0343] In some examples, multiple prediction refinements are derived in multiple directions or multiple direction pairs.
[0344] In some examples, a first prediction refinement with multiple prediction refinements is derived on a horizontal-vertical pair including both horizontal and vertical directions, and a second prediction refinement with multiple prediction refinements is derived on a diagonal-anti-diagonal pair including both diagonal and anti-diagonal directions.
[0345] In some examples, for a list of reference images X, the first prediction refinement is defined as:
[0346] ofX h (x,y)×gradX h (x,y)+ofX v (x,y)×gradX v (x,y),
[0347] Where X equals 0 or 1, ofX h (x,y) and ofX v (x, y) represent the horizontal and vertical optical flow of the reference image list X, respectively, and gradX h (x,y) and gradX v (x,y) represents the horizontal and vertical gradients of PX(x,y), and PX(x,y) represents the predicted sample points of sample points (x,y) in the reference image list X.
[0348] In some examples, for a list of reference images X (X equals 0 or 1), the second prediction refinement is defined as:
[0349] ofX d (x,y)×gradX d (x,y)+ofX ad (x,y)×gradX ad (x,y),
[0350] Where X equals 0 or 1, ofX d (x,y) and ofX ad (x, y) represent the diagonal and anti-diagonal optical flows in the reference image list X, respectively, and gradXd (x,y) and gradX ad (x,y) represents the diagonal and anti-diagonal gradients of PX(x,y), where PX(x,y) represents the predicted sample point of sample point (x,y) in the reference image list X.
[0351] In some examples, a weighted average of multiple prediction refinements is taken to generate the final prediction refinement.
[0352] In some examples, the weights of multiple prediction refinements depend on the gradient information of the prediction blocks associated with the current video block.
[0353] In some examples, the spatial gradients of multiple orientation pairs are computed, and smaller weights are assigned to orientation pairs with smaller spatial gradients.
[0354] In some examples, the spatial gradients of multiple orientation pairs are computed, and smaller weights are assigned to orientation pairs with larger spatial gradients.
[0355] In some examples, the weight of the first sample in the first prediction refinement block associated with the current video block is different from the weight of the second sample in the first prediction refinement block.
[0356] In some examples, default weights are assigned to multiple prediction refinements.
[0357] In some examples, 3 / 4 is used for the first prediction refinement and 1 / 4 is used for the second prediction refinement.
[0358] In some examples, a final prediction refinement is generated for each list of reference images X.
[0359] In some examples, the weights used for multiple prediction refinements depend on the reliability of multiple optical flows associated with the current video block.
[0360] In some examples, in the case of bidirectional prediction, refined prediction samples in a list of reference images X associated with the current video block are generated using the prediction samples, optical flow, and spatial gradients of the prediction samples, where X is 0 or 1.
[0361] In some examples, refined prediction samples are generated as a sum of prediction samples and prediction refinement.
[0362] In some examples, for horizontal-vertical pairs, the refined prediction samples in the reference image list X are generated as follows:
[0363] PX(x,y)+ofX h (x,y)×gradX h (x,y)+ofX v (x,y)×gradX v (x,y).
[0364] In some examples, for diagonal-anti-diagonal direction pairs, the refined prediction samples in the reference image list X are generated as follows:
[0365] PX(x,y)+ofX d (x,y)×gradX d (x,y)+ofX ad (x,y)×gradX ad (x,y).
[0366] In some examples, reliability depends on the difference between the refined predictions in two lists of reference images in bidirectional predictive encoding and decoding.
[0367] In some examples, the reliability of each pixel is derived.
[0368] In some examples, the reliability of each block or each sub-block is derived.
[0369] In some examples, when deriving the reliability of a block or sub-block, the difference between some representative samples is calculated.
[0370] In some examples, the difference is the absolute difference and SAD, the squared error and SSE, or the absolute transform difference and SATD.
[0371] In some examples, higher reliability is assigned to optical flow with smaller differences between refined predictions in two lists of reference images.
[0372] In some examples, larger weights are assigned to the refinement of predictions generated from optical flow with higher reliability.
[0373] In some examples, the weights also depend on whether the prediction refinement comes from horizontal-vertical pairs or from diagonal-anti-diagonal pairs.
[0374] Figure 11 A flowchart of an example method 1100 for video processing. Method 1100 includes: at 1102, for a conversion between a current video block and a bitstream representation of the current video block, determining a direction or direction pair associated with the current video block during an optical flow-based prediction refinement or prediction process, wherein the direction or direction pair changes from one video region of the current video block to another video region; and at 1104, performing a conversion based on the direction or direction pair.
[0375] In some examples, a direction pair is first determined, and an optical flow-based prediction refinement process is performed along the determined direction pair.
[0376] In some examples, the gradient of the predicted block associated with the current video block is used to determine the orientation pair.
[0377] In some examples, the spatial gradients of multiple orientation pairs are computed, and an optical flow-based prediction refinement process is performed on the orientation pair with the minimum spatial gradient.
[0378] In some examples, the spatial gradients of multiple orientation pairs are computed, and an optical flow-based prediction refinement process is performed on the orientation pair with the largest spatial gradient.
[0379] Figure 12 A flowchart of an example method 1200 for video processing. Method 1200 includes: at 1202, for a conversion between a current video block and a bitstream representation of the current video block, performing interpolation of motion vectors associated with the current video block to generate an interpolation result in an optical flow-based motion refinement or prediction process, wherein interpolation is performed along a direction different from the horizontal and / or vertical directions; and at 1204, performing a conversion based on the interpolation result.
[0380] In some examples, interpolation is performed along two directions that are orthogonal to each other, and these two directions are different from the horizontal and vertical directions.
[0381] In some examples, interpolation is performed along the diagonal direction and / or the anti-diagonal direction, where the diagonal direction refers to the horizontal direction rotated M degrees counterclockwise, and the anti-diagonal direction refers to the vertical direction rotated N degrees counterclockwise, where M and N are integers.
[0382] In some examples, the interpolation filter used is different from those used in horizontal and / or vertical interpolation for direction.
[0383] In some examples, when the motion vector contains fractional components in both the diagonal and anti-diagonal directions, intermediate samples are first interpolated along the diagonal direction, and then the intermediate samples are used to interpolate and predict samples along the anti-diagonal direction.
[0384] In some examples, when the motion vector contains fractional components in both the diagonal and anti-diagonal directions, intermediate samples are first interpolated along the anti-diagonal direction, and then the intermediate samples are used to interpolate and predict samples along the diagonal direction.
[0385] Figure 13 A flowchart of an example method 1300 for video processing is provided. Method 1300 includes: at 1302, for a conversion between a current video block and a bitstream representation of the current video block, performing interpolation of motion vectors associated with the current video block to generate one or more interpolation results in an optical flow-based motion refinement or prediction process; at 1304, generating a final interpolation result associated with the current video block by combining multiple interpolation results; and at 1306, performing a conversion based on the final interpolation result.
[0386] In some examples, multiple interpolation results are derived in multiple directions or direction pairs.
[0387] In some examples, a first interpolation result is generated on a horizontal-vertical pair that includes both horizontal and vertical directions, and a second interpolation result is derived on a diagonal-anti-diagonal pair that includes both diagonal and anti-diagonal directions.
[0388] In some examples, a weighted average of multiple interpolation results is taken to generate the final interpolation result.
[0389] In some examples, the weighting depends on the gradient information of the reference block associated with the current video block.
[0390] In some examples, the spatial gradients of multiple orientation pairs are computed, and smaller weights are assigned to orientation pairs with smaller spatial gradients.
[0391] In some examples, the spatial gradients of multiple orientation pairs are computed, and smaller weights are assigned to orientation pairs with larger spatial gradients.
[0392] In some examples, the weight of the first sample point in the first interpolation block is different from the weight of the second sample point in the first interpolation block.
[0393] In some examples, the weight of each sample point is derived.
[0394] In some examples, the weight of each block or sub-block is derived.
[0395] In some examples, default weights are assigned to multiple interpolation results.
[0396] In some examples, 3 / 4 is used for the first interpolation result and 1 / 4 is used for the second interpolation result.
[0397] Figure 14 A flowchart of an example method 1400 for video processing. Method 1400 includes: at 1402, for a conversion between a current video block and a bitstream representation of the current video block, performing interpolation of motion vectors associated with the current video block to generate an interpolation result in an optical flow-based motion refinement or prediction process, wherein interpolation is performed along one or more directions or direction pairs from one video region of the current video block to another video region; and at 1404, performing a conversion based on the interpolation result.
[0398] In some examples, a direction pair is first determined, and interpolation is performed along the determined direction pair.
[0399] In some examples, the gradient of a reference block associated with the current video block is used to determine the orientation pair.
[0400] In some examples, the spatial gradients of multiple direction pairs are computed, and interpolation is performed on the direction pairs with the minimum spatial gradients.
[0401] In some examples, the spatial gradients of multiple direction pairs are computed, and interpolation is performed on the direction pair with the largest spatial gradient.
[0402] In some examples, interpolation is performed on the diagonal-anti-diagonal direction pair when the motion vector has a fractional component in only one of the diagonal and anti-diagonal directions.
[0403] In some examples, whether and / or how the determination or execution process is applied is explicitly or implicitly signaled, or depends on the encoding / decoding information in the bitstream representation.
[0404] In some examples, the determination or execution process is applied to certain block sizes or shapes, and / or certain sub-block sizes and / or color components.
[0405] In some examples, certain block sizes include at least one of the following:
[0406] Blocks where max(W,H) / min(W,H)<=T;
[0407] Blocks where max(W,H) / min(W,H)>=T;
[0408] Blocks where W×H<=T;
[0409] Blocks where H<=T or H==T;
[0410] Blocks where W <= T or W == T;
[0411] A block where W <= T1 and H <= T2;
[0412] A block where W>=T1 and H>=T2;
[0413] Blocks where W×H>=T;
[0414] Blocks where H>=T;
[0415] Blocks where W>=T;
[0416] Where W and H are the width and height of the current video block, and T, T1, and T2 are predetermined thresholds.
[0417] In some examples, the color component includes only the luminance component.
[0418] In some examples, the motion refinement or prediction refinement process based on optical flow is PROF or BDOF.
[0419] In some examples, the conversion involves encoding the current video block into a bitstream.
[0420] In some examples, the conversion involves decoding the current video chunk from the bitstream.
[0421] In some examples, the conversion includes generating a bitstream from the current video block; the method also includes storing the bitstream in a non-transitory computer-readable storage medium.
[0422] Figure 15 A flowchart of an example method 1500 for video processing. Method 1500 includes: at 1502, a conversion between a current video block and a bitstream representation of the current video block, determining an optical flow associated with the current video block during an optical flow-based motion refinement or prediction process, wherein the optical flow is derived along a direction different from the horizontal and / or vertical directions; at 1504, generating a bitstream from the current video block based on the optical flow; and at 1506, storing the bitstream in a non-transitory computer-readable storage medium.
[0423] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium, executed or controlled by a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of substances that influences machine-readable propagated signals, or a combination thereof. The term "data processing apparatus" encompasses all means, devices, and machines that process data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0424] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple harmonizing files (e.g., files storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on a single computer or on multiple computers located in one location or distributed across multiple locations and interconnected via a communication network.
[0425] The processes and logic described in this document can be executed by one or more programmable processors to execute one or more computer programs, thereby performing functions by manipulating input data and producing outputs. The processes and logic can also be executed by dedicated logic circuits, and can be implemented as dedicated logic circuits, such as FPGAs (field-programmable gate arrays) or ASICs (application-specific integrated circuits).
[0426] Processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0427] Although this patent document contains numerous details, these details should not be construed as limiting any invention or the scope of the claims, but rather as a description of features that may be specific to particular embodiments of a particular invention. Certain features described in this patent document in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although the foregoing may describe features as functioning in certain combinations and even initially claimed in this way, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof.
[0428] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential order, or to perform all shown operations to achieve the desired effect. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0429] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.
Claims
1. A video processing method, comprising: For the conversion between the current video block and its bitstream representation, the optical flow associated with the current video block is determined during an optical flow-based motion refinement or prediction process, wherein the optical flow is derived along a direction different from the horizontal and / or vertical directions; and The conversion is performed based on the optical flow; The optical flow is derived along the diagonal and anti-diagonal directions; The diagonal direction includes a horizontal direction rotated counterclockwise by M degrees, and the anti-diagonal direction includes a vertical direction rotated counterclockwise by N degrees, where M and N equal 45.
2. The method according to claim 1, wherein, The spatial gradient associated with the current video block is derived along the same direction used to derive the optical flow.
3. The method according to claim 2, wherein, Using the optical flow and spatial gradient derived in the said direction, a predictive refinement associated with the current video block is generated.
4. The method according to claim 1, further comprising: In an optical flow-based motion refinement or prediction process, the spatial gradient of a first direction pair associated with the current video block is determined, wherein the spatial gradient of the first direction pair depends on the spatial gradients of the two directions of the first direction pair; and The transformation is performed based on the spatial gradient.
5. The method according to claim 4, wherein, The spatial gradient of the first direction pair is calculated as a function of the spatial gradients in the two directions of the first direction pair.
6. The method according to claim 5, wherein, The spatial gradient of the first direction pair is calculated as the sum or weighted sum of the absolute gradients in the two directions of the first direction pair.
7. The method according to claim 6, wherein, The first direction pair includes a horizontal direction and a vertical direction, and the spatial gradient of the first direction pair is calculated as the sum of the absolute horizontal gradient and the absolute vertical gradient.
8. The method according to claim 6, wherein, The first direction pair includes a diagonal direction and an anti-diagonal direction, and the spatial gradient of the first direction pair is calculated as the sum of the absolute diagonal gradient and the absolute anti-diagonal gradient.
9. The method according to claim 4, wherein, The spatial gradient of the first direction pair is calculated as the larger or smaller or average value of the absolute gradients in the two directions of the first direction pair.
10. The method according to claim 4, wherein, The spatial gradient of the first direction pair is used to determine which first direction pair to select to perform prediction refinement associated with the current video block.
11. The method according to claim 1, further comprising: One or more predictive refinements are generated in the optical flow-based motion refinement or prediction process, which are associated with the current video block. A final prediction refinement associated with the current video block is generated by combining multiple prediction refinements. as well as The transformation is performed based on the final prediction.
12. The method according to claim 11, wherein, The multiple prediction refinements are derived in multiple directions or multiple direction pairs.
13. The method according to claim 12, wherein, A first prediction refinement of the plurality of prediction refinements is derived on a horizontal-vertical pair including the horizontal and vertical directions, and a second prediction refinement of the plurality of prediction refinements is derived on a diagonal-anti-diagonal pair including the diagonal and anti-diagonal directions.
14. The method according to claim 13, wherein, For the reference image list X, the first prediction refinement is defined as: , Where X equals 0 or 1, and Let X represent the horizontal and vertical optical flow of the reference image list X, respectively, and gradX. h (x, y) and gradX v (x, y) represents the horizontal and vertical gradients of PX(x, y), and PX(x, y) represents the predicted sample point of sample point (x, y) in the reference image list X.
15. The method according to claim 14, wherein, For the reference image list X, the second prediction refinement is defined as: , Where X equals 0 or 1, and Let X represent the diagonal optical flow and anti-diagonal optical flow in the reference image list X, respectively, and gradX d (x, y) and gradX ad (x, y) represents the diagonal gradient and anti-diagonal gradient of PX(x, y), and PX(x, y) represents the predicted sample point of sample point (x, y) in the reference image list X.
16. The method according to any one of claims 13-15, wherein, The multiple prediction refinements are weighted and averaged to generate the final prediction refinement.
17. The method according to claim 16, wherein, The weights of the multiple prediction refinements depend on the gradient information of the prediction blocks associated with the current video block.
18. The method according to claim 17, wherein, The spatial gradient is calculated for the plurality of direction pairs, and smaller weights are assigned to direction pairs with smaller spatial gradients.
19. The method according to claim 17, wherein, The spatial gradient is calculated for the plurality of direction pairs, and smaller weights are assigned to direction pairs with larger spatial gradients.
20. The method of claim 16, wherein, The weight of the first sample in the first prediction refinement block associated with the current video block is different from the weight of the second sample in the first prediction refinement block.
21. The method according to claim 16, wherein, The default weights are assigned to the multiple prediction refinements.
22. The method according to claim 21, wherein, 3 / 4 is used for the first prediction refinement, and 1 / 4 is used for the second prediction refinement.
23. The method according to claim 16, wherein, For each reference image list X, the final prediction refinement is generated.
24. The method of claim 16, wherein, The weights used for the multiple prediction refinements depend on the reliability of the multiple optical flows associated with the current video block.
25. The method according to claim 24, wherein, In the case of bidirectional prediction, refined prediction samples in a list of reference images X associated with the current video block are generated using the prediction samples, the optical flow, and the spatial gradient of the prediction samples, where X is 0 or 1.
26. The method according to claim 25, wherein, The refined prediction sample points are generated as the sum of the prediction sample points and the prediction refinement.
27. The method according to claim 26, wherein, For the horizontal-vertical pair, the refined prediction samples in the reference image list X are generated as follows: 。 28. The method according to claim 27, wherein, For the diagonal-anti-diagonal direction pair, the refined prediction samples in the reference image list X are generated as follows: 。 29. The method according to claim 24, wherein, The reliability depends on the difference between the refined predictions in the two reference image lists in the bidirectional predictive encoding and decoding.
30. The method according to claim 29, wherein, The reliability is derived for each pixel.
31. The method according to claim 29, wherein, The reliability is derived for each block or each sub-block.
32. The method according to claim 29, wherein, When deriving the reliability of a block or sub-block, the difference between some representative samples is calculated.
33. The method according to claim 29, wherein, The difference is the absolute difference and SAD, the squared error and SSE, or the absolute transformation difference and SATD.
34. The method according to claim 29, wherein, Higher reliability is assigned to the optical flow that has a smaller difference between the refined predictions in the two lists of reference images.
35. The method according to claim 24, wherein, Larger weights are assigned to the prediction refinement generated from the optical flow, which has higher reliability.
36. The method according to claim 35, wherein, The weights also depend on whether the prediction refinement comes from horizontal-vertical pairs or from diagonal-anti-diagonal pairs.
37. The method according to claim 1, further comprising: In an optical flow-based prediction refinement process or prediction process, a direction or direction pair associated with the current video block is determined, wherein the direction or direction pair changes from one video region of the current video block to another video region; and The transformation is performed based on the direction or direction pair.
38. The method according to claim 37, wherein, First, a direction pair is determined, and the optical flow-based prediction refinement process is performed along the determined direction pair.
39. The method according to claim 37 or 38, wherein, The gradient of the prediction block associated with the current video block is used to determine the direction pair.
40. The method of claim 37, wherein, The spatial gradients of multiple direction pairs are calculated, and the optical flow-based prediction refinement process is performed on the direction pairs with the minimum spatial gradients.
41. The method according to claim 37, wherein, The spatial gradients of multiple direction pairs are calculated, and the optical flow-based prediction refinement process is performed on the direction pair with the largest spatial gradient.
42. The method according to claim 1, further comprising: Interpolation of the motion vectors associated with the current video block is performed to generate interpolation results in an optical flow-based motion refinement or prediction process, wherein the interpolation is performed along a direction different from the horizontal and / or vertical directions; and The transformation is performed based on the interpolation result.
43. The method according to claim 42, wherein, Interpolation is performed along two mutually orthogonal directions, which are different from the horizontal and vertical directions.
44. The method according to claim 42, wherein, Interpolation is performed along the diagonal direction and / or the anti-diagonal direction, wherein the diagonal direction refers to the horizontal direction rotated counterclockwise by M degrees, and the anti-diagonal direction refers to the vertical direction rotated counterclockwise by N degrees, where M and N are integers.
45. The method according to any one of claims 42-44, wherein, An interpolation filter, different from the interpolation filter used in horizontal and / or vertical interpolation, is used for the direction.
46. The method of claim 44, wherein, When the motion vector contains fractional components in the diagonal direction and the anti-diagonal direction, the intermediate sample points are first interpolated along the diagonal direction, and then the intermediate sample points are used to interpolate and predict sample points along the anti-diagonal direction.
47. The method of claim 44, wherein, When the motion vector contains fractional components in the diagonal direction and the anti-diagonal direction, the intermediate sample points are first interpolated along the anti-diagonal direction, and then the intermediate sample points are used to interpolate and predict sample points along the diagonal direction.
48. The method according to claim 1, further comprising: Perform interpolation of the motion vectors associated with the current video block to generate one or more interpolation results in an optical flow-based motion refinement or prediction process; By combining multiple interpolation results, a final interpolation result associated with the current video block is generated; as well as The transformation is performed based on the final interpolation result.
49. The method according to claim 48, wherein, The multiple interpolation results are derived in multiple directions or direction pairs.
50. The method according to claim 49, wherein, A first interpolation result is generated on a horizontal-vertical direction pair including the horizontal and vertical directions, and a second interpolation result is derived on a diagonal-anti-diagonal direction pair including the diagonal and anti-diagonal directions.
51. The method according to claim 50, wherein, The multiple interpolation results are weighted and averaged to generate the final interpolation result.
52. The method according to claim 51, wherein, The weighting depends on the gradient information of the reference block associated with the current video block.
53. The method according to claim 52, wherein, Calculate the spatial gradient of the plurality of direction pairs and assign smaller weights to direction pairs with smaller spatial gradients.
54. The method according to claim 52, wherein, Calculate the spatial gradient of the plurality of direction pairs and assign smaller weights to direction pairs with larger spatial gradients.
55. The method according to claim 51, wherein, The weight of the first sample point in the first interpolation block is different from the weight of the second sample point in the first interpolation block.
56. The method according to claim 51, wherein, The weights are derived for each sample point.
57. The method according to claim 51, wherein, Derive weights for each block or sub-block.
58. The method according to claim 57, wherein, Default weights are assigned to the multiple interpolation results.
59. The method according to claim 58, wherein, 3 / 4 is used for the first interpolation result, and 1 / 4 is used for the second interpolation result.
60. The method of claim 1, further comprising: Interpolation of motion vectors associated with the current video block is performed to generate interpolation results during an optical flow-based motion refinement or prediction process, wherein the interpolation is performed along one or more directions or direction pairs that transform one video region of the current video block into another video region; and The transformation is performed based on the interpolation result.
61. The method according to claim 60, wherein, First, a second direction pair is determined, and the interpolation is performed along the determined second direction pair.
62. The method according to claim 61, wherein, The gradient of the reference block associated with the current video block is used to determine the second direction pair.
63. The method according to claim 62, wherein, Calculate the spatial gradient of the plurality of direction pairs, and perform the interpolation on the direction pair with the minimum spatial gradient.
64. The method according to claim 62, wherein, Calculate the spatial gradient of the plurality of direction pairs, and perform the interpolation on the direction pair with the largest spatial gradient.
65. The method according to claim 61, wherein, When the motion vector has a fractional component in only one of the diagonal and anti-diagonal directions, the interpolation is performed on the diagonal-anti-diagonal direction pair.
66. The method according to claim 1, wherein, Whether and / or how the determination or execution process is applied is explicitly or implicitly signaled, or depends on the encoding / decoding information in the bitstream representation.
67. The method according to claim 66, wherein, The determination or execution process is applied to certain block sizes or shapes, and / or certain sub-block sizes and / or color components.
68. The method according to claim 67, wherein, The block sizes include at least one of the following: Blocks where max(W, H) / min(W, H) <= T; Blocks where max(W, H) / min(W, H) >= T; Blocks where W×H <= T; Blocks where H <= T or H == T; Blocks where W <= T or W == T; A block where W <= T1 and H <= T2; A block where W >= T1 and H >= T2; Blocks where W×H >= T; Blocks where H >= T; Blocks where W >= T; Where W and H are the width and height of the current video block, and T, T1, and T2 are predetermined thresholds.
69. The method according to claim 67, wherein, The color components include only the luminance component.
70. The method according to claim 1, wherein, The motion refinement or prediction refinement process based on optical flow is PROF or BDOF.
71. The method according to claim 1, wherein, The conversion includes encoding the current video block into the bitstream.
72. The method according to claim 1, wherein, The conversion includes decoding the current video block from the bitstream.
73. The method according to claim 1, wherein, The conversion includes generating the bitstream from the current video block; The method further includes: The bit stream is stored in a non-transitory computer-readable storage medium.
74. A video data processing apparatus, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: For the conversion between the current video block and its bitstream representation, the optical flow associated with the current video block is determined during an optical flow-based motion refinement or prediction process, wherein the optical flow is derived along a direction different from the horizontal and / or vertical directions; and The conversion is performed based on the optical flow; The optical flow is derived along the diagonal and anti-diagonal directions; The diagonal direction includes a horizontal direction rotated counterclockwise by M degrees, and the anti-diagonal direction includes a vertical direction rotated counterclockwise by N degrees, where M and N equal 45.
75. A non-transitory computer-readable storage medium storing instructions that cause a processor to: For the conversion between the current video block and its bitstream representation, the optical flow associated with the current video block is determined during the optical flow-based motion refinement or prediction process, wherein... The optical flow is derived along a direction different from the horizontal and / or vertical directions; and The conversion is performed based on the optical flow; The optical flow is derived along the diagonal and anti-diagonal directions; The diagonal direction includes a horizontal direction rotated counterclockwise by M degrees, and the anti-diagonal direction includes a vertical direction rotated counterclockwise by N degrees, where M and N equal 45.
76. A non-transitory computer-readable storage medium having stored thereon a computer program and a bitstream, wherein the computer program, when executed by a processor, implements the following video processing method to generate the bitstream, wherein... The video processing method includes: For the conversion between the current video block and its bitstream representation, the optical flow associated with the current video block is determined during an optical flow-based motion refinement or prediction process, wherein the optical flow is derived along a direction different from the horizontal and / or vertical directions; and The bitstream is generated from the current video block based on the optical flow; The optical flow is derived along the diagonal and anti-diagonal directions; The diagonal direction includes a horizontal direction rotated counterclockwise by M degrees, and the anti-diagonal direction includes a vertical direction rotated counterclockwise by N degrees, where M and N equal 45.
77. A method for storing a bitstream of video, wherein the bitstream is generated by performing the following video processing method: For the conversion between the current video block and its bitstream representation, the optical flow associated with the current video block is determined during the optical flow-based motion refinement or prediction process, wherein... The optical flow is derived along a direction different from the horizontal and / or vertical directions; The bitstream is generated from the current video block based on the optical flow. The optical flow is derived along the diagonal and anti-diagonal directions. Wherein, the diagonal direction includes a horizontal direction rotated counterclockwise by M degrees, and the anti-diagonal direction includes a vertical direction rotated counterclockwise by N degrees, where M and N equal 45°; and The bit stream is stored in a non-transitory computer-readable storage medium.