Systems, apparatuses and methods for inter prediction refinement with optical flow

By refining sub-block based motion prediction signals using spatial gradients and optical flow analysis, the method addresses challenges in video encoding related to complex motion and illumination changes, resulting in improved coding efficiency and compression performance.

JP2025087840AActive Publication Date: 2025-06-10INTERDIGITAL VC HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025035590
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-04-15
Filing Date
2025-03-06
Publication Date
2025-06-10
Estimated Expiration
2040-02-04

AI Technical Summary

Technical Problem

Current video encoding systems face challenges in efficiently compressing digital video signals, particularly in handling rapid changes in illumination and complex motion patterns, which can lead to suboptimal prediction techniques and increased bandwidth requirements.

Method used

The method involves obtaining a sub-block based motion prediction signal, calculating spatial gradients or motion vector difference values, and using these to refine the motion prediction signal through optical flow analysis, thereby enhancing the precision of motion compensation.

Benefits of technology

This approach improves the accuracy of motion prediction, reduces residual errors, and enhances coding efficiency by better handling complex motion and illumination changes, ultimately leading to more effective video compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087840000001_ABST
    Figure 2025087840000001_ABST
Patent Text Reader

Abstract

To provide methods, apparatuses and systems.SOLUTION: In one aspect, a decoding method includes obtaining a sub-block based motion prediction signal for a current block of video; obtaining one or more spatial gradients of the sub-block based motion prediction signal or one or more motion vector difference values; obtaining a refinement signal for the current block based on the one or more obtained spatial gradients or the one or more obtained motion vector difference values; obtaining a refined motion prediction signal for the current block based on the sub-block based motion prediction signal and the refinement signal; and decoding the current block based on the refined motion prediction signal.SELECTED DRAWING: Figure 18A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to video encoding, and more particularly to systems, apparatuses, and methods that use inter-prediction refinement using optical flow.

Background Art

[0002] Cross References This application claims the benefit of U.S. Provisional Patent Application No. 62 / 802,428, filed Feb. 7, 2019, U.S. Provisional Patent Application No. 62 / 814,611, filed Mar. 6, 2019, and U.S. Provisional Patent Application No. 62 / 883,999, filed Apr. 15, 2019, each of which is incorporated herein by reference in its entirety.

[0003] Prior Art Video encoding systems are widely used to compress digital video signals to reduce the storage amount and / or transmission bandwidth of such signals. Among various types of video encoding systems, such as block-based, wavelet-based, and object-based systems, currently block-based hybrid video encoding systems are the most widely used and deployed. Examples of block-based video encoding systems include international video encoding standards such as MPEG1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, and the latest video encoding standard called High Efficiency Video Coding (HEVC) developed by ITU-T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG's Joint Collaborative Team on Video Coding (JCT-VC).

Summary of the Invention

[0004] In one representative embodiment, a method of decoding includes obtaining a sub-block based motion prediction signal for a current block of video, obtaining one or more spatial gradients of the sub-block based motion prediction signal, or one or more motion vector difference values, obtaining a refinement signal for the current block based on the one or more obtained spatial gradients, or the one or more obtained motion vector difference values, obtaining a refined motion prediction signal for the current block based on the sub-block based motion prediction signal and the refinement signal, and decoding the current block based on the refined motion prediction signal. Various other embodiments are also disclosed herein.

[0005] A more detailed understanding can be obtained from the following detailed description, given by way of example and with reference to the drawings attached hereto. The figures in the description are examples. Accordingly, the figures and the detailed description are not to be regarded as limiting, and other equally effective examples are possible and contemplated. Further, like reference numerals in the figures indicate like elements.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8A

Figure 8B

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16

Figure 17A

Figure 17B

Figure 17C

Figure 18A

Figure 18B

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34A

Figure 34B

Figure 34C

Figure 34D

DETAILED DESCRIPTION OF THE INVENTION

[0007] Block-based Hybrid Video Encoding Procedure Similar to HEVC, VVC is built on a block-based hybrid video encoding framework.

[0008] FIG. 1 is a block diagram showing a general block-based hybrid video encoding system.

[0009] Referring to FIG. 1, the encoder 100 may be provided with an input video signal 102 that is processed block by block (referred to as a coding unit (CU)), and may be used to efficiently compress a high-resolution (1080p or higher) video signal. In HEVC, a CU can be up to 64×64 pixels. A CU can be further divided into prediction units, i.e., PUs, to which different prediction procedures can be applied. For each input video block (MB and / or CU), spatial prediction 160 and / or temporal prediction 162 can be performed. Spatial prediction (or "intra prediction") can predict the current video block using pixels from already encoded adjacent blocks within the same video picture / slice.

[0010] Spatial prediction can reduce the spatial redundancy inherent in the video signal. Temporal prediction (also called "inter prediction" or "motion compensation prediction") predicts the current video block using pixels from already encoded video pictures. Temporal prediction can reduce the temporal redundancy inherent in the video signal. The temporal prediction signal for a given video block can be signaled (e.g., typically) by one or more motion vectors (MVs) that can indicate the amount and / or direction of motion between the current block (CU) and its reference block.

[0011] If multiple reference pictures are supported (similar to recent video coding standards such as H.264 / AVC or HEVC), for each video block, the reference picture index of the video block can be transmitted (e.g., additionally transmitted), and / or a reference index can be used to identify from which reference picture in the reference picture store 164 the temporal prediction signal arrives. After spatial prediction and / or temporal prediction, the mode decision block 180 in the encoder 100 can select the best prediction mode, for example, based on a rate-distortion optimization method / procedure. The prediction block of either spatial prediction 160 or temporal prediction 162 can be subtracted from the current video block 116, and / or the prediction residue can be decorrelated using transformation 104 and quantization 106 to achieve the target bit rate. The quantized residue coefficients can be inverse quantized 110 and inverse transformed 112 to form the reconstructed residue, and the reconstructed residue can be added back to the prediction block at 126 to form the reconstructed video block. In-loop filtering 166 such as a deblocking filter and / or an adaptive loop filter can be applied to the reconstructed video block, and then it can be put into the reference picture store 164 and used to encode future video blocks. To form the output video bitstream 120, the coding mode (inter or intra), prediction mode information, motion information, and quantized residue coefficients are transmitted (e.g., all transmitted) to the entropy coding unit 108, and further compressed and / or packed to form the bitstream.

[0012] Encoder 100 can be implemented using a processor, a memory, and a transmitter that provide the various elements / modules / units disclosed above. For example, the transmitter can send bitstream 120 to a decoder, and one skilled in the art understands that (2) the processor can be configured to execute software to enable reception of input video 102 and execution of functions associated with the various blocks of encoder 100.

[0013] Figure 2 is a block diagram showing a block-based video decoder.

[0014] Referring to Figure 2, video decoder 200 can be provided with video bitstream 202 that can be unpacked and entropy decoded in entropy decoding unit 208. Encoding mode and prediction information can be sent to the appropriate one of spatial prediction unit 260 (for intra encoding mode) and / or temporal prediction unit 262 (for inter encoding mode) to form prediction blocks. Residual transform coefficients can be sent to inverse quantization unit 210 and inverse transform unit 212 to reconstruct residual blocks. The reconstructed blocks may further pass through in-loop filtering 266 before being stored in reference picture store 264. Reconstructed video 220 can be sent to be stored in reference picture store 264, for example, to drive a display device and for use in predicting future video blocks.

[0015] Decoder 200 can be implemented using a processor, a memory, and a receiver that can provide the various elements / modules / units disclosed above. For example, (1) the receiver can be configured to receive bitstream 202, and (2) the processor can be configured to execute software to enable the reception of bitstream 202, the output of the reconstructed video 220, and the execution of functions associated with the various blocks of decoder 200, which would be understood by those skilled in the art.

[0016] Those skilled in the art will understand that many of the functions / operations / processes of a block-based encoder and a block-based decoder are the same.

[0017] In modern video codecs, bidirectional motion compensation prediction (MCP) can be used for high efficiency in removing temporal redundancy by exploiting the temporal correlation between pictures. The bidirectional prediction signal can be formed by combining two single prediction signals using a weight value equal to 0.5, which may not be optimal for combining the single prediction signals, especially in conditions where the illuminance changes rapidly from one reference picture to another. Certain prediction techniques / operations and / or procedures can be implemented to compensate for the illuminance variation over time by applying some global / local weights and / or offset values to the sample values in the reference pictures (e.g., a part or each of the sample values in the reference pictures).

[0018] The use of bidirectional motion compensation prediction (MCP) in a video codec enables the removal of temporal redundancy by exploiting the temporal correlation between pictures. The bidirectional prediction signal can be formed by combining two uni-prediction signals using a weight value (e.g., 0.5). In certain videos, the illumination characteristics may change rapidly from one reference picture to another. Therefore, the prediction technique may compensate for the variation in illumination over time (e.g., a fading transition) by applying global or local weights and / or offset values to one or more sample values in the reference picture.

[0019] Generalized bidirectional prediction (GBi) may improve MCP for the bidirectional prediction mode. In the bidirectional prediction mode, the prediction signal at a given sample x can be calculated by Equation 1 as follows.

[0020] P[x]=w 0 *P 0 [x+v 0 +w 1 *P 1 [x+v 1 (1) In the above formula, P[x] can represent the predicted signal provided by sample x placed at picture position x. Pi[x + vi] can be the motion compensated prediction signal of x using the motion vector (MV) vi for the i-th list (for example, list 0, list 1, etc.). w0 and w1 can be two weight values shared across (for example, all) samples within a block. Based on this formula, various prediction signals can be obtained by adjusting the weight values w0 and w1. Some configurations of w0 and w1 can mean the same prediction as single prediction and bidirectional prediction. For example, (w0, w1) = (1, 0) can be used for single prediction using reference list L0. (w0, w1) = (0, 1) can be used for single prediction using reference list L1. (w0, w1) = (0.5, 0.5) can be used for bidirectional prediction using two reference lists. The weights can be signaled for each CU. To reduce the signaling overhead, a constraint such as w0 + w1 = 1 can be applied so that only one weight is signaled. Therefore, Equation 1 can be further simplified as shown in Equation 2 below.

[0021] P[x]=(1 - w 1 )*P 0 [x + v 0 +w 1 *P 1 [x + v 1 (2) To further reduce the signaling overhead, w1 can be discretized (for example, -2 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 10 / 8, etc.). In this way, each weight value can be represented by an index value within a (for example, small) limited range.

[0022] Figure 3 is a block diagram showing a representative block-based video encoder having GBi support.

[0023] Encoder 300 may include a mode decision module 304, a spatial prediction module 306, a motion prediction module 308, a transformation module 310, a quantization module 312, an inverse quantization module 316, an inverse transformation module 318, a loop filter 320, a reference picture store 322, and an entropy encoding module 314. Some or all of the modules or components of the encoder (e.g., the spatial prediction module 306) may be the same as or similar to those described in connection with FIG. 1. Further, the spatial prediction module 306 and the motion prediction module 308 may be pixel region prediction modules. Thus, the input video bitstream 302 may be processed in a manner similar to the input video bitstream 102, but the motion prediction module 308 may further include GBi support. In this way, the motion prediction module 308 can combine two separate prediction signals in a weighted average manner. Further, the selected weight index may be signaled in the output video bitstream 324.

[0024] Encoder 300 may be implemented using a processor, a memory, and a transmitter that provide the various elements / modules / units disclosed above. For example, the transmitter can send the bitstream 324 to a decoder, and (2) one of ordinary skill in the art will understand that the processor can be configured to execute software to enable reception of the input video 302 and execution of the functions associated with the various blocks of the encoder 300.

[0025] FIG. 4 is a diagram showing a representative GBi estimation module 400 that can be used in a motion prediction module of an encoder such as motion prediction module 308. The GBi estimation module 400 can include a weight value estimation module 402 and a motion estimation module 404. Thus, the GBi estimation module 400 can utilize a process (e.g., a two-step operation / process) for generating an inter-prediction signal such as a final inter-prediction signal. The motion estimation module 404 can perform motion estimation by searching for two optimal motion vectors (MVs) that point to (e.g., two) reference blocks using the input video block 401 and one or more reference pictures received from the reference picture store 406. The weight value estimation module 402 can receive (1) the output of the motion estimation module 404 (e.g., motion vectors v 0 and v 1 ), one or more reference pictures from the reference picture store 406, and weight information W, and search for an optimal weight index to minimize the weighted bi-directional prediction error between the current video block and the bi-directional prediction. The weight information W can describe a list of available weight values or weight sets so that the determined weight index and the weight information W can be used together to specify the weights w 0 and w 1 used in GBi. The prediction signal for generalized bi-directional prediction can be calculated as a weighted average of two prediction blocks. The output of the GBi estimation module 400 can include an inter-prediction signal, motion vectors v 0 and v 1 , and / or a weight index weight_idx, etc.).

[0026] FIG. 5 is a diagram showing a representative block-based video decoder having GBi support, which can decode a bitstream 502 that supports GBi (e.g., from an encoder), such as the bitstream 324 created by the encoder 300 described in connection with FIG. 3. As shown in FIG. 5, the video decoder 500 can include an entropy decoder 504, a spatial prediction module 506, a motion prediction module 508, a reference picture store 510, an inverse quantization module 512, an inverse transform module 514, and / or a loop filter module 518. Some or all of the modules of the decoder can be the same as or similar to those described in connection with FIG. 2, but the motion prediction module 508 can further include GBi support. Thus, the coding mode and prediction information can be used to derive a prediction signal using spatial prediction or MCP with GBi support. For GBi, block motion information and weight values (e.g., in the form of an index indicating a weight value) can be received and decoded to generate a prediction block.

[0027] The decoder 500 can be implemented using a processor, memory, and receiver that can provide the various elements / modules / units disclosed above. For example, those skilled in the art will understand that (1) the receiver can be configured to receive the bitstream 502, and (2) the processor can be configured to execute software to enable the reception of the bitstream 502, the output of the reconstructed video 520, and the execution of the functions associated with the various blocks of the decoder 500.

[0028] FIG. 6 is a diagram showing a representative GBi prediction module that can be utilized in a motion prediction module of a decoder, such as the motion prediction module 508.

[0029] Referring to FIG. 6, the GBi prediction module can include a weighted average module 602 and a motion compensation module 604. The motion compensation module 604 can receive one or more reference pictures from a reference picture store 606. The weighted average module 602 can receive the output of the motion compensation module 604, weight information W, and a weight index (e.g., weight_idx). The output of the motion compensation module 604 can include motion information corresponding to a block of the picture. The GBi prediction module 600 can calculate a prediction signal of GBi (e.g., an inter prediction signal 608) as a weighted average of (e.g., two) motion compensation prediction blocks using the block motion information and the weight values.

[0030] Typical bidirectional predictive prediction based on an optical flow model FIG. 7 is a diagram showing a typical bidirectional optical flow.

[0031] Referring to FIG. 7, the bidirectional predictive prediction can be based on an optical flow model. For example, the prediction associated with a current block (e.g., curblk700) can be the first prediction block I (0) 702 (e.g., a temporally previous prediction block shifted by only time τ 0 ), and the second prediction block I (1) 704 (e.g., time τ 1It can be based on an optical flow associated with a time-shifted, temporally future prediction block. Bidirectional prediction in video coding may be a combination of two temporal prediction blocks 702 and 704 obtained from already reconstructed reference pictures. Due to the limitations of block-based motion compensation (MC), there may be remaining small motions that can be observed between the samples of the two prediction blocks, thereby potentially reducing the efficiency of motion compensation prediction. To reduce the impact of such motion for all samples within a block, bidirectional optical flow (referred to as BIO or BDOF) can be applied. BIO can provide sample-by-sample motion refinement that can be performed in addition to block-based motion compensation prediction when bidirectional prediction is used. Regarding BIO, the derivation of refined motion vectors for each sample in a block can be based on classical optical flow models. For example, if I (k) (x,y) is the sample value at the coordinates (x,y) of a prediction block derived from reference picture list k (k = 0,1), and ∂I (k) (x,y) / ∂x and ∂I (k) (x,y) / ∂y are the horizontal and vertical gradients of the sample, assuming an optical flow model, the motion refinement (v x ,v y ) at (x,y) can be derived by Equation 3 below.

[0032]

Equation

[0033] In FIG. 7, (MV x0 ,MV y0 ) associated with the first prediction block 702 and (MV x1 ,MV y1 ) associated with the second prediction block 704 are the two prediction blocks I (0) and I (1)Shows block-level motion vectors that can be used to generate. The motion refinement (v x , v y ) at the sample position (x, y) can be calculated by minimizing the difference Δ between the values of the samples after motion refinement compensation (e.g., A and B in FIG. 7), as shown in Equation 4 below.

[0034]

Equation

[0035] For example, to ensure the regularity of the derived motion refinement, it is intended that the motion refinement be consistent for samples within one small unit (e.g., a 4×4 block or other small unit). In the Benchmark Set (BMS)-2.0, the values of (v x , v y ) are derived by minimizing Δ within a 6×6 window Ω around each 4×4 block, as shown in Equation 5 below.

[0036]

Equation

[0037] To solve the optimization specified in Equation 5, BIO can use a progressive method / operation / procedure that can optimize the motion refinement horizontally and vertically (e.g., then vertically). This can result in the following equations / inequalities 6 and 7.

[0038]

Equation

[0039]

Equation

[0040] Here,

[0041]

Number

[0042] can be a floor function that outputs the maximum value of the following input, and th BIO can be a motion refinement threshold, for example, for preventing error propagation due to coding noise and / or irregular local motion, and is equal to 2 18-BD . S 1 , S 2 , S 3 , S 5 and S 6 The values of can be further calculated as shown in the following equations 8 to 12. S 1 = Σ (i,j)∈Ω ψ x (i, j)·ψ x (i, j) (8) S 3 = Σ (i,j)∈Ω θ(i, j)·ψ x (i, j)·2 L (9) S 2 = Σ (i,j)∈Ω ψ x (i, j)·ψ y (i, j) (10) S 5 = Σ (i,j)∈Ω ψ y (i, j)·ψ y (i, j)·2 (11) S 6 = Σ (i,j)∈Ω θ(i, j)·ψ y (i, j)·2 L+1 (12) Here, various gradients can be shown by the following equations 13 to 15.

[0043]

Number

[0044]

Number

[0045]

Number

[0046] In BMS-2.0, the BIO gradients in both the horizontal and vertical directions in Equations 13 to 15 can be directly obtained by calculating the difference between two adjacent samples at one sample position of each L0 / L1 prediction block (e.g., horizontally or vertically depending on the direction of the derived gradient), as shown in the following Equations 16 and 17.

[0047]

Number

[0048]

Number

[0049] k = 0, 1 In Equations 8 to 12, L can be an increase in bit depth for internal BIO processing / procedures to maintain data accuracy, which can be set to 5, for example, in BMS-2.0. To avoid segmentation by smaller values, the adjustment parameters r and m in Equations 6 and 7 can be defined as shown in the following Equations 18 and 19. r = 500·4 BD-8 (18) m = 700·4 BD-8 (19) Here, BD can be the bit depth of the input video. Based on the motion refinement derived by Equations 4 and 5, the final bidirectional prediction signal of the current CU can be calculated by interpolating L0 / L1 prediction samples along the motion trajectory based on the optical flow Equation 3, as specified in the following Equations 20 and 21.

[0050] [Number] TIFF2025087840000015.tif6150

[0051] [Number]

[0052] Here, shift and ο offset can be right shift and offset applied to combine the L0 prediction signal and the L1 prediction signal for bidirectional prediction. For example, they can be set equal to 15 - BD and 1 << (14 - BD)+2·(1 << 13), respectively. rnd(·) is a rounding function that may round the input value to the nearest integer value.

[0053] Typical affine mode In HEVC, a translational motion (only translational motion) model is applied to motion - compensated prediction. In the real world, many types of motion (e.g., zoom - in / out, rotation, perspective motion, and other irregular motions) exist. In the VVC Test Model (VTM) - 2.0, affine motion - compensated prediction is applied. The affine motion model is either a 4 - parameter or a 6 - parameter model. The first flag for an inter - coded CU is signaled to indicate whether a translational motion model or an affine motion model is applied to inter - prediction. When an affine motion model is applied, a second flag is sent to indicate whether the model is a 4 - parameter model or a 6 - parameter model.

[0054] The 4-parameter affine motion model has two parameters for translational movement in the horizontal and vertical directions, one parameter for zoom movement in both directions, and one parameter for rotational movement in both directions. The horizontal zoom parameter is equal to the vertical zoom parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. The 4-parameter affine motion model is encoded in the VTM using two motion vectors at two control point positions defined at the upper left corner 810 and upper right corner 820 of the current CU. Other control point positions at other corners and / or edges of the current CU are also possible.

[0055] One affine motion model has been described above, but other affine models are equally possible and may be used in various embodiments of this specification.

[0056] Figures 8A and 8B are diagrams showing a representative 4-parameter affine model and sub-block level motion derivation of an affine block. Referring to Figures 8A and 8B, the affine motion field of a block is described at a first control point 810 (upper left corner of the current block) and a second control point 820 (upper right corner of the current block) by two control point motion vectors respectively. Based on the control point motion, the motion field (v x ,v y ) of one affine-encoded block is described as shown in the following equations 22 and 23.

[0057]

Equation

[0058]

Equation

[0059] Here, (v 0x ,v 0y ) can be the motion vector of the upper left corner control point 810, and (v 1x ,v1y ) can be the motion vector of the upper-right control point 820 as shown in FIG. 8A, and w can be the width of the CU. For example, the motion field of an affine-coded CU is derived at the 4×4 block level, that is, (v x , v y ) is derived for each 4×4 block within the current CU and applied to the corresponding 4×4 block.

[0060] The four parameters of the four-parameter affine model can be iteratively estimated. The MV pair at step k is

[0061]

Equation

[0062] shown as, the original signal (e.g., the luminance signal) is shown as I(i, j), and the predicted signal (e.g., the luminance signal) can be shown as I’ k (i, j). The spatial gradients g x (i, j) and g y (i, j) can be derived, for example, using Sobel filters applied to the predicted signal I’ k (i, j) in the horizontal and / or vertical directions respectively. The derivation of Equation 3 can be expressed as shown in the following Equations 24 and 25.

[0063]

Equation

[0064] Here, at step k, (a, b) can be the delta translation parameters, and (c, d) can be the delta zoom and rotation parameters. The delta MV at the control point can be derived using its coordinates as shown in the following Equations 26 - 29. For example, (0, 0), (w, 0) can be the coordinates of the upper-left control point 810 and the upper-right control point 820 respectively.

[0065] [Number]

[0066] [Number]

[0067] Based on the optical flow equation, the relationship between the change in intensity (e.g., luminance), the spatial gradient, and the temporal movement is formulated in Equation 30 as follows.

[0068] [Number]

[0069] [Number]

[0070] and

[0071] [Number]

[0072] By replacing with Equation 24, Equation 31 for the parameters (a, b, c, d) is obtained as follows. I’ k (i,j) - I(i,j) = (g x (i,j) * i + g y (i,j) * j) * c + (-g x (i,j) * j + g y (i,j) * i) * d + g x (i,j) * a + g y (i,j) * b (31)

[0073] Since the samples in the CU (e.g., all samples) satisfy Equation 31, the parameter set (e.g., a, b, c, d) can be solved using, for example, the least squares error method. The two control points at step (k + 1)

[0074] [Number]

[0075] The MVs in [it] can be solved using Equations 26 to 29, and they can be rounded to a specific precision (e.g., 1 / 4 pixel precision (pel) or other sub-pixel precision, etc.). Using iteration, the MVs at two control points can be refined, for example, until convergence (e.g., until all of the parameters (a, b, c, d) are zero, or until the iteration time reaches a predefined limit).

[0076] FIG. 9 is a diagram showing a typical six-parameter affine mode. In the figure, for example, V 0 , V 1 , and V 2 are motion vectors at control points 910, 920, and 930, respectively, and (MV x , MV y ) are the motion vectors of the sub-block centered at the position (x, y).

[0077] Referring to FIG. 9, an affine motion model (e.g., having six parameters) can have any of (1) a parameter for translational movement in the horizontal direction, (2) a parameter for translational movement in the vertical direction, (3) a parameter for zoom movement in the horizontal direction, (4) a parameter for rotational movement in the horizontal direction, (5) a parameter for zoom movement in the vertical direction, and / or (6) a parameter for rotational movement in the vertical direction. The six-parameter affine motion model can be encoded using three MVs at three control points 910, 920, and 930. As shown in FIG. 9, the three control points 910, 920, and 930 for the six-parameter affine-encoded CU are defined at the upper left corner, upper right corner, and lower left corner of the CU, respectively. The motion at the upper left control point 910 may be associated with translational motion, the motion at the upper right control point 920 may be associated with rotational motion in the horizontal direction and / or zoom motion in the horizontal direction, and the motion at the lower left control point 930 may be associated with rotational motion in the vertical direction and / or zoom motion in the vertical direction. In the six-parameter affine motion model, rotational motion and / or zoom motion in the horizontal direction may not be the same as the same motion in the vertical direction. The motion vector of each sub-block (v x ,v y ) can be derived using the three MVs at control points 910, 920, and 930, as shown in the following equations 32 and 33.

[0078]

Number

[0079]

Number

[0080] Here, (v 2x ,v 2y ) is the motion vector V of the lower left control point 930 2can be set, (x, y) can be the center position of the sub-block, w can be the width of the CU, and h can be the height of the CU.

[0081] The six parameters of the six-parameter affine model can be estimated in a similar way. Equations 24 and 25 can be changed as shown in the following equations 34 and 35.

[0082]

Equation

[0083] Here, at step k, (a, b) can be the delta translation parameter, (c, d) can be the delta zoom and rotation parameters for the horizontal direction, and (e, f) can be the delta zoom and rotation parameters for the vertical direction. Equation 31 can be changed as shown in the following equation 36. I’ k (i,j)-I(i,j)=(g x (i,j)*i)*c+(g x (i,j)*j)*d+(g y (i,j)*i)*e+(g y (i,j)*j)*f+g x (i,j)*a+g y (i,j)*b (36)

[0084] The parameter set (a, b, c, d, e, f) can be solved using the least squares method / procedure / operation by considering, for example, the samples (e.g., all samples) within the CU. The MV of the upper left control point

[0085]

Equation

[0086] can be calculated using equations 26 to 29. The MV of the upper right control point

[0087]

Number

[0088] can be calculated using equations 37 and 38 as shown below. The MV of the lower left control point

[0089]

Number

[0090] can be calculated using equations 39 and 40 as shown below.

[0091]

Number

[0092]

Number

[0093] In FIGS. 8A, 8B and 9, 4- and 6-parameter affine models are shown, but those skilled in the art will understand that affine models with different numbers of parameters and / or different control points are equally possible.

[0094] Although an affine model is described herein in connection with optical flow refinement, those skilled in the art will understand that other motion models associated with optical flow refinement are equally possible.

[0095] Typical interweave prediction for affine motion compensation In affine motion compensation (AMC), for example in VTM, an encoded block is partitioned into small sub-blocks of about 4×4, and each of them can be assigned an individual motion vector (MV) derived by an affine model as shown in, for example, FIGS. 8A and 8B or FIG. 9. In a 4-parameter or 6-parameter affine model, the MV can be derived from the MVs of two or three control points.

[0096] AMC may face a dilemma associated with the size of sub-blocks. With smaller sub-blocks, AMC can achieve better encoding performance but may be troubled by a greater burden of complexity.

[0097] FIG. 10 shows a typical interweave prediction procedure that can achieve finer granularity of MVs in exchange for a gentle increase in complexity, for example.

[0098] In FIG. 10, the encoded block 1010 can be partitioned into sub-blocks having two different partitioning patterns (for example, the first pattern 0 and the second pattern 1). As shown in FIG. 10, the first partitioning pattern 0 (for example, the first sub-block pattern, e.g., 4×4 sub-block pattern) may be the same as that in VTM, and the second partitioning pattern 1 (for example, the second sub-block pattern that overlaps and / or is interwoven) may partition the encoded block 1010 into 4×4 sub-blocks having a 2×2 offset from the first partitioning pattern 0. With AMC using two partitioning patterns (for example, the first partitioning pattern 0 and the second partitioning pattern 1), several auxiliary predictions (for example, two auxiliary predictions P 0 and P 1 ) can be generated. The MV of each sub-block in each of the partitioning patterns 0 and 1 can be derived from the control point motion vector (CPMV) by an affine model.

[0099] The final prediction P is the auxiliary prediction (for example, two auxiliary predictions P0 and P 1 can be calculated as the weighted sum of (

[0100] [Number]

[0101] FIG. 11 is a diagram showing representative weight values (e.g., associated with pixels) in a sub-block. Referring to FIG. 11, an auxiliary prediction sample arranged at the center (e.g., the center pixel) of sub-block 1100 may be associated with a weight value of 3, and an auxiliary prediction sample arranged at the boundary of sub-block 1100 may be associated with a weight value of 1.

[0102] FIG. 12 is a diagram showing a region where interweave prediction is applied and other regions where interweave prediction is not applied. Referring to FIG. 12, region 1200 can include a first region 1210 (not shaded as shown in FIG. 12) having, for example, 4×4 sub-blocks to which interweave prediction is applied, and a second region 1220 (shaded as shown in FIG. 12) where, for example, interweave prediction is not applied. To avoid small block motion compensation, interweave prediction can be applied only to regions where the size of the sub-block satisfies a threshold size (e.g., 4×4) for both the first segmentation pattern and the second segmentation pattern.

[0103] In VTM-3.0, the size of the sub-block may be 4×4 in the chrominance components, and the interweave prediction may be applied to the chrominance components and / or the luminance component. Since the regions used to perform motion compensation (MC) for the sub-blocks (e.g., all sub-blocks) can be taken out together as a whole in AMC, the bandwidth may not be increased by the interweave prediction. For flexibility, a flag may be signaled in the slice header to indicate whether the interweave prediction is used. For the interweave prediction, the flag may be signaled as a 1-bit flag (e.g., the first logical level that can always be signaled as 0 or 1).

[0104] Typical procedure for Sub-block-Based Temporal Motion Vector Prediction (SbTMVP) SbTMVP is supported by VTM. Similar to the Temporal Motion Vector Prediction (TMVP) in HEVC, SbTMVP can use the motion field in the picture at the same position, for example, to improve the motion vector prediction and merge mode of the CU in the current picture. The same picture at the same position used by TMVP may be used by SbTMVP. What SbTMVP differs from TMVP is that (1) TMVP can predict the motion at the CU level, and SbTMVP can predict the motion at the sub-CU level, and / or (2) TMVP can extract the temporal motion vector from the block at the same position within the picture at the same position (for example, the block at the same position may be the block at the lower right or the center with respect to the current CU), and SbTMVP can apply a motion shift before extracting the temporal motion information from the picture at the same position (for example, the motion shift can be obtained from the motion vector from one of the spatially adjacent blocks of the current CU).

[0105] FIG. 13A and FIG. 13B are diagrams showing SbTMVP processing. FIG. 13A shows the spatial adjacent blocks used by ATMVP, and FIG. 13B shows the derivation of the sub-CU motion field by applying a motion shift from spatial adjacency and scaling the motion information from the corresponding sub-CUs at the same position.

[0106] Referring to FIGS. 13A and 13B, SbTMVP can predict the motion vectors of sub-CUs within the current CU operation (e.g., in two operations). In the first operation, the spatial adjacent blocks A1, B1, B0, and A0 can be inspected in the order of A1, B1, B0, and A0. As soon as and / or after a first spatial adjacent block having a motion vector using the same-position picture as its reference picture is identified, this motion vector can be selected as the applied motion shift. If such a motion is not identified from the spatial adjacent blocks, the motion shift can be set to (0, 0). In the second operation, as shown in FIG. 13B, the motion shift identified in the first operation is applied (e.g., added to the coordinates of the current block), and motion information (e.g., motion vectors and reference indices) at the sub-CU level can be obtained from the same-position picture. The example of FIG. 13B shows the motion shift set for the motion of block A1. For each sub-CU, the motion information of its corresponding block (e.g., the smallest motion grid covering the central sample) in the same-position picture can be used to derive the motion information of that sub-CU. After the motion information of the sub-CUs at the same position is identified, in a manner similar to the TMVP processing of HEVC, the motion information can be converted into the motion vectors and reference indices of the current sub-CU. For example, temporal motion scaling can be applied to align the reference pictures of the temporal motion vectors with those of the current CU.

[0107] The combined sub-block based merge list may be used in VTM-3 and can contain or include both SbTMVP and affine merge candidates, for example, for use in signaling the sub-block based merge mode. The SbTMVP mode can be enabled / disabled by a sequence parameter set (SPS) flag. When the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry of the list of sub-block based merge candidates, and the affine merge candidates may follow. The size of the sub-block based merge list may be signaled in the SPS, and the maximum allowable size of the sub-block based merge list is an integer and may be set to 5 in VTM3 for example.

[0108] The sub-CU size used for SbTMVP may be fixed at, for example, 8×8 or another sub-CU size, and the SbTMVP mode may be applicable to CUs having both width and height that can be 8 or more (for example, applicable only to them), as done in the affine merge mode. The encoding logic for additional SbTMVP merge candidates may be the same as that for other merge candidates. For example, for each CU in a P or B slice, an additional rate distortion (RD) check may be performed to determine whether to use the SbTMVP candidate.

[0109] Representative Regression based Motion Vector Field To provide fine-grained granularity of motion vectors within a block, the regression based motion vector field (RMVF) tool may be implemented (e.g., in JVET-M0302), which attempts to model the motion vector of each block at the sub-block level based on spatially adjacent motion vectors.

[0110] FIG. 14 is a diagram showing adjacent motion blocks (e.g., 4×4 motion blocks) that can be used for deriving motion parameters. One row 1410 and one column 1420 of directly adjacent motion vectors based on 4×4 sub-blocks (and at their central positions) from each side of the block can be used in the regression process. For example, those adjacent motion vectors can be used in RMVF motion parameter derivation.

[0111] FIG. 15 is a diagram showing adjacent motion blocks that can be used for deriving motion parameters, where the adjacent motion information can be reduced (e.g., the number of adjacent motion blocks used in the regression process related to FIG. 14 can be reduced). The reduced amount of adjacent motion information for RMVF parameter derivation of adjacent 4×4 motion blocks can be used for deriving motion parameters (e.g., about half, for example, every other adjacent motion block can be used for deriving motion parameters). Specific adjacent motion blocks of row 1410 and column 1420 can be selected, determined, or pre-determined to reduce the adjacent motion information.

[0112] Although it is shown that about half of the adjacent motion blocks of row 1410 and column 1420 are selected, other ratios (including other motion block positions) may be selected to reduce the number of adjacent motion blocks used in the regression process, for example.

[0113] When collecting motion information for deriving motion parameters, the five regions shown in the figure (e.g., lower left, left, upper left, upper, upper right) can be used. The reference motion regions of the upper right and lower left can be limited to half (e.g., only half) of the corresponding width or height of the current block.

[0114] In the RMVF mode, the motion of a block can be defined by a six-parameter motion model. These parameters a xx 、a xy 、a yx 、a yy 、b x 、and b ycan be calculated by solving a linear regression model in the sense of the mean squared error (MSE). The input to the regression model can consist of, or include, the center positions (x, y) of the available adjacent 4×4 sub-blocks and / or motion vectors (mv x and mv y ), as defined above.

[0115] (X subPU ,Y subPU ) having the center position, the motion vectors (MV X_subPU ,MV Y_subPU ) of the 8×8 sub-block can be calculated as shown in Equation 43 below.

[0116]

Equation

[0117] The motion vector can be calculated for an 8×8 sub-block with respect to the center position of the sub-block (e.g., each sub-block). For example, motion compensation can be applied with 8×8 sub-block accuracy in the RMVF mode. To obtain an efficient modeling of the motion vector field, the RMVF tool is applied only when at least one motion vector from at least three candidate regions is available.

[0118] The affine motion model parameters can be used to derive the motion vectors of specific pixels (e.g., each pixel) in a CU. Although the complexity of generating pixel-based affine motion compensation prediction can be high (e.g., very high), since the memory access bandwidth requirement for this type of sample-based MC can be high, a sub-block-based affine motion compensation procedure / method may be implemented (e.g., by VVC). For example, a CU may be partitioned into sub-blocks (e.g., 4×4 sub-blocks, square sub-blocks, and / or non-square sub-blocks). Each of the sub-blocks can be assigned an MV that can be derived from the affine model parameters. The MV can be the MV at the center of the sub-block (or another position within the sub-block). The pixels within the sub-block (e.g., all the pixels within the sub-block) can share the sub-block MV. Sub-block-based affine motion compensation can be a trade-off between coding efficiency and complexity. To achieve finer-grained motion compensation, an interweave prediction for affine motion compensation may be implemented and can be generated by weighted averaging two sub-block motion compensation predictions. The interweave prediction requires and / or can use two or more motion compensation predictions per sub-block and thus may increase the memory bandwidth and complexity.

[0119] In certain representative embodiments, methods, apparatuses, procedures, and / or operations may be implemented to refine sub-block-based affine motion compensation prediction using optical flow (e.g., using optical flow and / or based on optical flow). For example, after sub-block-based affine motion compensation is performed, pixel intensities may be refined by adding difference values derived by an optical flow equation, which is referred to as prediction refinement with optical flow (PROF). PROF is capable of achieving pixel-level granularity without significantly increasing complexity and may maintain a worst-case memory access bandwidth similar to sub-block-based affine motion compensation. PROF may be applied in any scenario where a pixel-level motion vector field is available (e.g., can be calculated) in addition to a prediction signal (e.g., an unrefined motion prediction signal and / or a sub-block-based motion prediction signal). In addition to or instead of the affine mode, the prediction PROF procedure may be used in other sub-block prediction modes. The application of PROF in sub-block modes such as SbTMVP and / or RMVF may be implemented. The application of PROF in bidirectional prediction is described herein.

[0120] Representative PROF procedure for the affine mode In certain representative embodiments, methods, apparatuses, and / or procedures may be implemented to improve the granularity of sub-block-based affine motion compensation prediction, for example, by applying changes in pixel intensities derived from optical flow (e.g., an optical flow equation), and may use and / or require, for example, the same one motion compensation operation per sub-block (e.g., only one motion compensation operation per sub-block) as existing affine motion compensation in VVC.

[0121] FIG. 16 is a diagram showing sub-block MVs and pixel-level motion vector differences Δv(i,j) (which may also be referred to as, for example, refinement MVs for pixels) after sub-block-based affine motion compensation prediction.

[0122] Referring to FIG. 16, CU1600 can include sub-blocks 1610, 1620, 1630, and 1640. Each of the sub-blocks 1610, 1620, 1630, and 1640 can include a plurality of pixels (for example, 16 pixels within sub-block 1610). A sub-block MV1650 (for example, as a rough or average sub-block MV) associated with each pixel 1660(i,j) of sub-block 1610 is shown. For each respective pixel (i,j) within sub-block 1610, a refinement MV1670(i,j) can be determined, which can indicate the difference between the actual MV of pixel 1660(i,j) and sub-block MV1650 (where (i,j) defines the pixel position within sub-block 1610). For clarity of FIG. 16, only refinement MV1670(1,1) is labeled, but other individual pixel-level motions are shown. In a particular representative embodiment, refinement MV1670(i,j) can be determined as a pixel-level motion vector difference Δv(i,j) (which may also be referred to as a motion vector difference).

[0123] In a particular representative embodiment, a method, apparatus, procedure, and / or operation including any of the following operations can be implemented. (1) In a first operation, sub-block-based AMC can be performed as disclosed herein to generate a sub-block-based motion prediction I(i,j). (2) In a second operation, the spatial gradients g x (i,j) and g y(i,j) can be calculated (in one example, the spatial gradient can be generated using the same process as the gradient generation used in BDOF. For example, the horizontal gradient at a sample position can be calculated as the difference between its right adjacent sample and its left adjacent sample, and / or the vertical gradient at a sample position can be calculated as the difference between its lower adjacent sample and its upper adjacent sample. In another example, the spatial gradient can be generated using a Sobel filter). (3) In a third operation, for example, as shown in Equation 44 below, the optical flow equation can be used and / or thereby the luminance intensity change for each pixel in the CU can be calculated. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (44) Here, the value of the motion vector difference Δv(i,j) is the difference 1670 between the pixel-level MV v(i,j) calculated for the sample position (i,j) denoted by v(i,j) as shown in FIG. 16 and the sub-block level MV1650 of the sub-block covering pixel 1660(i,j). The pixel-level MV v(i,j) can be derived from the control point MV by Equations 22 and 23 for the four-parameter affine model or Equations 32 and 33 for the six-parameter affine model.

[0124] In certain representative embodiments, the motion vector difference value Δv(i,j) may be derived by or using equations 24 and 25 by affine model parameters, where x and y can be the offsets from the pixel position to the center of the sub-block. Since the affine model parameters and pixel offsets are not changed for each sub-block, the motion vector difference value Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks within the same CU. For example, since the translational affine parameters (a,b) may be the same for the pixel-level MV and the sub-block MV, the difference between the pixel-level MV and the sub-block level MV can be calculated using equations 45 and 46 as follows. (c,d,e,f) can be four additional affine parameters (e.g., four affine parameters other than the translational affine parameters).

[0125] [Number]

[0126] Here, (i,j) can be the pixel position relative to the upper left position of the sub-block, and (x sb ,y sb ) can be the center position of the sub-block relative to the upper left position of the sub-block.

[0127] FIG. 17A is a diagram showing a representative procedure for determining the MV corresponding to the actual center of the sub-block.

[0128] Referring to FIG. 17A, two sub-blocks SB 0 and SB 1is shown as a 4×4 sub-block. When the sub-block width is SW and the sub-block height is SH, the sub-block center position can be shown as ((SW−1) / 2,(SH−1) / 2). In other examples, the sub-block center position can be estimated based on the position shown as (SW / 2,SH / 2). The actual center point uses ((SW−1) / 2,(SH−1) / 2) and for the first sub-block SB 0 is P 0 ’ and for the second sub-block SB 1 is P 1 ’. The estimated center point uses (for example, in VVC) for example (SW / 2,SH / 2) and for the first sub-block SB 0 is P 0 and for the second sub-block SB 1 is P 1 . In certain representative embodiments, the MV of the sub-block can be based on the actual center position rather than the estimated center position (used in VVC).

[0129] FIG. 17B is a diagram showing the positions of chrominance samples in the 4:2:0 chrominance format. Referring to FIG. 17B, the chrominance sub-block MV can be derived from the MV of the luminance sub-block. For example, in the 4:2:0 chrominance format, one 4×4 chrominance sub-block can correspond to an 8×8 luminance region. Although representative embodiments are shown in connection with the 4:2:0 chrominance format, those skilled in the art will understand that other chrominance formats such as the 4:2:2 chrominance format can be used as well.

[0130] The chrominance sub-block MV can be derived by averaging the top-left 4×4 luminance sub-block MV and the bottom-right luminance sub-block MV. The derived chrominance sub-block MV may or may not be placed at the center of the chrominance sub-block for chrominance sample position types 0, 2, and / or 3. For chrominance sample position types 0, 2, and 3, the chrominance sub-block center position (x sb ,ysb ) may be adjusted by an offset or may need to be adjusted. For example, for 4:2:0 color difference sample position types 0, 2, and 3, adjustments may be applied as shown in the following equations 47 to 49.

[0131]

Number

[0132]

Number

[0133]

Number

[0134] The sub-block based motion prediction I(i,j) can be refined by adding intensity changes (e.g., luminance intensity changes as provided in Equation 44 as an example). The final (i.e., refined) prediction I’(i,j) can be generated by Equation 50 or using Equation 50 as follows. I’(i,j)=I(i,j)+ΔI(i,j) (50)

[0135] When refinement is applied, sub-block based affine motion compensation can achieve pixel-level granularity without increasing the worst-case bandwidth and / or memory bandwidth.

[0136] To maintain the accuracy of prediction and / or gradient calculation, the bit depth in the operational relationship performance of sub-block based AMC can be set to an intermediate bit depth that can be higher than the coding bit depth.

[0137] The above-described process may be used to refine the chroma intensity (e.g., in addition to or instead of refining the luminance intensity). In one example, the intensity difference used in Equation 50 may be multiplied by a weight factor w before being added to the prediction, as shown in Equation 51 below. I’(i,j)=I(i,j)+w·ΔI(i,j) (51) Here, w may be set to a value from 0 to 1 and w may be signaled at the CU level or picture level. For example, w may be signaled by a weight index. For example, Index Table 1 may be used to signal w.

[0138] [Table 1]

[0139] The encoder algorithm can select the value of w that results in the lowest rate-distortion cost.

[0140] The gradient of the predicted sample, e.g., g x and / or g y can be calculated in different ways. In certain representative embodiments, the predicted samples g x and g y can be calculated by applying a 2D Sobel filter. Examples of 3×3 Sobel filters for the horizontal and vertical gradients are shown below.

[0141] [Number]

[0142] [Number]

[0143] In other representative embodiments, the gradient can be calculated using a one-dimensional 3-tap filter. The example may include [-1 0 1], which can be simpler (e.g., much simpler) than the Sobel filter.

[0144] FIG. 17C is a diagram showing extended sub-block prediction. The shaded circle 1710 is a padding sample around a 4×4 sub-block (e.g., the unshaded circle 1720). As an example, the Sobel filter can be used to calculate the gradient of the samples within the box 1730 for the central sample 1740. The gradient can be calculated using the Sobel filter, but other filters such as a 3-tap filter are also possible.

[0145] In the case of the exemplary gradient filters described above, e.g., the 3×3 Sobel filter and the one-dimensional filter, extended sub-block prediction can be used and / or required for sub-block gradient calculation. One row at the upper and lower boundaries of the sub-block and one column at the left and right boundaries can be padded, for example, to calculate the gradient of those samples at the sub-block boundaries.

[0146] There may be different methods / procedures and / or operations for obtaining an extended sub-block prediction. In one representative embodiment, an N×M is given as the sub-block size, and an (N+2)×(M+2) extended sub-block prediction can be obtained by performing (N+2)×(M+2) block motion compensation using the sub-block MV. In this embodiment, the memory bandwidth may be increased. To avoid increasing the memory bandwidth, in certain representative embodiments, a K-tap interpolation filter in both the horizontal and vertical directions is provided, and for interpolation of the N×M sub-block, integer reference samples of (N+K-1)×(M+K-1) before interpolation may be taken, and boundary samples of the (N+K-1)×(M+K-1) block may be copied from adjacent samples of the (N+K-1)×(M+K-1) sub-block so that the extended region can be (N+K-1+2)×(M+K-1+2). The extended region may be used for interpolation of the (N+2)×(M+2) sub-block. These representative embodiments may further use and / or require additional interpolation operations to generate an (N+2)×(M+2) prediction when the sub-block MV points refer to fractional positions.

[0147] For example, to reduce the computational complexity, in other representative embodiments, sub-block prediction may be obtained by N×M block motion compensation using the sub-block MV. The boundary of the (N+2)×(M+2) prediction may be obtained without interpolation by any one of (1) integer motion compensation where the MV is the integer part of the sub-block MV, (2) integer motion compensation where the MV is the nearest integer MV of the sub-block MV, and / or (3) copying from the nearest adjacent samples in the N×M sub-block prediction.

[0148] For example, the accuracy and / or range of pixel-level refinement MV can affect the accuracy of PROF. In certain representative embodiments, a combination of a multi-bit fractional component and another multi-bit integer component can be implemented. For example, a 5-bit fractional component and an 11-bit integer component may be used. The combination of a 5-bit fractional component and an 11-bit integer component can represent an MV range of -1024 to 1023 with a total of 16 bits at 1 / 32 pel accuracy.

[0149] Gradient, e.g., g x and g y , as well as the accuracy of the intensity change ΔI, can affect the performance of PROF. In certain representative embodiments, the prediction sample accuracy can be maintained or held at a predetermined number or the number of signaled bits (e.g., the internal sample accuracy defined in the current VVC draft at 14 bits). In certain representative embodiments, the gradient and / or the intensity change ΔI can be maintained at the same accuracy as the prediction sample.

[0150] The range of the intensity change ΔI can affect the performance of PROF. The intensity change ΔI can be clipped to a smaller range to avoid incorrect values generated by an inaccurate affine model. In one example, the intensity change ΔI can be clipped to predition_bitdepth - 2.

[0151] Δv x and Δv y The combination of the number of bits of the fractional component of Δv, the number of bits of the fractional component of the gradient, and the number of bits of the intensity change ΔI can affect the complexity of a particular hardware or software implementation. In one representative embodiment, 5 bits can be used to represent the fractional component of Δv x and Δv y , 2 bits can be used to represent the fractional component of the gradient, and 12 bits can be used to represent ΔI, but they can be any number of bits.

[0152] In order to reduce the computational complexity, PROF may be omitted in certain situations. For example, if the magnitude of all pixel-based deltas (e.g., refinements) of the motion vectors (Δv(i,j)) within a 4×4 sub-block is smaller than a threshold, PROF may be omitted for the entire affine CU. If the gradient of all samples within a 4×4 sub-block is smaller than a threshold, PROF may be omitted. PROF may be applied to chrominance components such as the Cb and / or Cr components. The delta MV of the Cb and / or Cr components of a sub-block may reuse the delta MV of the sub-block (e.g., the delta MV calculated for different sub-blocks within the same CU may be reused).

[0153] The gradient procedures disclosed herein (e.g., expanding a sub-block for gradient calculation using copied reference samples) are shown to be used with the PROF operation, but the gradient procedures may also be used with other operations, particularly operations such as the BDOF operation and / or the affine motion estimation operation.

[0154] Typical PROF procedures for other sub-block modes PROF may be applied in any scenario where a pixel-level motion vector field is available (e.g., can be calculated) in addition to the prediction signal (e.g., the non-refined prediction signal). For example, in addition to the affine mode, prediction refinement using optical flow may be used in other sub-block prediction modes, such as the SbTMVP mode (e.g., the ATMVP mode in VVC), or the regression-based motion vector field (RMVF).

[0155] In certain representative embodiments, a method for applying PROF to SbTMVP may be implemented. For example, such a method may particularly include any of the following. (1) In a first operation, the sub-block level MV and sub-block prediction can be generated based on the existing SbTMVP processing described herein. (2) In the second operation, the affine model parameters can be estimated by the sub-block MV field using the linear regression method / procedure. (3) In the third operation, the pixel-level MV can be derived by the affine model parameters obtained in the second operation, and the associated pixel-level motion refinement vector (Δv(i,j)) for the sub-block MV can be calculated, and / or (4) In the fourth operation, prediction refinement using optical flow processing can be applied to generate, among other things, the final prediction.

[0156] In certain representative embodiments, a method for applying PROF to RMVF can be implemented. For example, such a method can include any of the following. (1) In the first operation, the sub-block level MV field, sub-block prediction, and / or the affine model parameters a xx , a xy , a yx , a yy , b x and b x can be generated based on the RMVF processing described herein. (2) In the second operation, the pixel-level MV offset from the sub-block level MV can be derived by the affine model parameters a xx , a xy , a yx , a yy , b x and b x as shown in Equation 52.

[0157]

Equation

[0158] Here, (i,j) is the pixel offset from the sub-block center. Since the affine model parameters and / or the pixel offset from the sub-block center do not change for each sub-block, the pixel MV offset may be calculated for the first sub-block (e.g., it needs to be calculated only for that, or should be calculated), and can be reused in other sub-blocks within the CU, and / or (3) In the third operation, the PROF process can be applied to generate a final prediction, for example, by applying Equations 44 and 50.

[0159] Typical PROF procedure for bidirectional prediction In addition to or instead of using PROF for single prediction as described herein, the PROF technique may be used for bidirectional prediction. When used in bidirectional prediction, PROF can be used to generate L0 prediction and / or L1 prediction, for example, to generate them before they are combined with weights. To reduce the computational complexity, PROF may be applied to one prediction such as L0 or L1 (e.g., it may be applied only to that). In certain representative embodiments, PROF may be applied to a list (e.g., a list associated with or together with the current picture being near and / or the nearest reference picture (e.g., within a threshold)).

[0160] Typical procedure for PROF activation PROF activation can be signaled in or within the sequence parameter set (SPS) header, the picture parameter set (PPS) header, and / or the tile group header. In certain embodiments, a flag can be signaled to indicate whether PROF is activated for affine mode. If the flag is set to a first logical level (e.g., "True"), PROF can be used for both single prediction and bidirectional prediction. In certain embodiments, if the first flag is set to "True", a second flag can be used to indicate whether PROF is activated or not activated for bidirectional prediction affine mode. If the first flag is set to a second logical level (e.g., "False"), the second flag can be presumed to be set to "False". Whether to apply PROF to the chrominance components can be signaled using a flag in or within the SPS header, the PPS header, and / or the tile group header if the first flag is set to "True", whereby the control of PROF for the luminance and chrominance components can be separated.

[0161] Typical procedures for conditionally activated PROF For example, to reduce complexity, PROF can be applied when certain conditions are met (e.g., only then). For example, for small CU sizes (e.g., below a threshold level), the advantage of applying PROF may be limited because the affine motion is relatively small. In certain representative embodiments, when the CU size is small (e.g., in the case of CU sizes of 16×16 or less such as 8×8, 8×16, 16×8), or under that condition, PROF can be disabled in affine motion compensation to reduce complexity for both the encoder and / or the decoder. In certain representative embodiments, when the CU size is small (less than the same or different threshold levels), PROF can be omitted, for example, in affine motion estimation (e.g., only affine motion estimation), to reduce the complexity of the encoder, and PROF can be executed in the decoder regardless of the CU size. For example, on the encoder side, after motion estimation to search for affine model parameters (e.g., control point MVs), the motion compensation (MC) procedure is activated and PROF can be executed. For each iteration in motion estimation, the MC procedure may be activated. In MC during motion estimation, PROF can be omitted to suppress complexity, but since the final MC in the encoder will execute PROF, there will be no prediction mismatch between the encoder and the decoder. That is, when the encoder searches for affine model parameters (e.g., affine MVs) to be used for prediction of the CU, PROF refinement may not be applied, and after the encoder completes the search or thereafter, the encoder can apply PROF to refine the prediction for the CU using the affine model parameters determined from the search.

[0162] In some representative embodiments, the difference between CPMVs can be used as a criterion to determine whether to enable PROF. When the difference between CPMVs is small (e.g., less than a threshold level), and thus the affine motion is small, the advantages of applying PROF may be limited, and PROF can be disabled for affine motion compensation and / or affine motion estimation. For example, in the four-parameter affine mode, PROF may be disabled if the following conditions are met (e.g., all of the following conditions are met).

[0163] [Number]

[0164] [Number]

[0165] In the six-parameter affine mode, in addition to or instead of the above conditions, PROF may be disabled if the following conditions are met (e.g., all of the following conditions are also met).

[0166] [Number]

[0167] [Number]

[0168] Here, T is a predefined threshold, e.g., 4. This CPMV- or affine parameter-based PROF omission procedure may be applied (e.g., only applied) in the encoder, and the decoder may or may not omit PROF.

[0169] Representative procedures for PROF in combination with or instead of the deblocking filter PROF can be a per-pixel refinement that can compensate for block-based MC, so that the motion difference between block boundaries can be reduced (e.g., significantly reduced). The encoder and / or decoder can omit the application of the deblocking filter and / or can apply a weaker filter at the sub-block boundary when PROF is applied. In a CU divided into multiple transform units (TUs), blocking artifacts may appear on the transform block boundary.

[0170] In certain representative embodiments, the encoder and / or decoder can omit the application of the deblocking filter and / or can apply one or more weaker filters at the sub-block boundary unless the sub-block boundary coincides with the TU boundary.

[0171] When PROF is applied to luminance (e.g., applied to luminance only) or under that condition, the encoder and / or decoder can omit the application of the deblocking filter and / or can apply one or more weaker filters on the sub-block boundary for luminance (e.g., luminance only). For example, the boundary strength parameter B can be used to apply a weaker deblocking filter.

[0172] For example, the encoder and / or decoder can omit the application of the deblocking filter on the sub-block boundary when PROF is applied unless the sub-block boundary coincides with the TU boundary. In that case, the deblocking filter may be applied to reduce or remove blocking artifacts that may occur along the TU boundary.

[0173] As another example, the encoder and / or decoder can apply a weaker deblocking filter on the sub-block boundary when PROF is applied, as long as the sub-block boundary does not coincide with the TU boundary. The "weaker" deblocking filter is intended to be a deblocking filter that is weaker than what can normally be applied to the sub-block boundary when PROF is not applied. When the sub-block boundary coincides with the TU boundary, a stronger deblocking filter is applied to reduce or remove blocking artifacts that are expected to be more visible along the sub-block boundary that coincides with the TU boundary.

[0174] In certain representative embodiments, when PROF is applied to (e.g., only applied to) luminance, or under that condition, the encoder and / or decoder can match the application of the deblocking filter for chrominance to luminance for design uniformity purposes, e.g., even though there is no application of PROF in chrominance. For example, when PROF is applied only to luminance, the normal application of the deblocking filter for luminance can be changed based on whether PROF has been applied (and optionally, based on whether there is a TU boundary at the sub-block boundary). In certain representative embodiments, rather than having separate / different logic for applying the deblocking filter to the corresponding chrominance pixels, the deblocking filter can be applied to the sub-block boundary for chrominance such that it aligns with (and / or mirrors) the procedure for luminance deblocking.

[0175] FIG. 18A is a flowchart showing a first representative encoding / decoding method.

[0176] Referring to FIG. 18A, a representative method 1800 for encoding and / or decoding can include, at block 1805, the encoder 100 or 300 and / or the decoder 200 or 500 obtaining a sub-block based motion prediction signal for a current block of, for example, video. At block 1810, the encoder 100 or 300 and / or the decoder 200 or 500 can obtain one or more spatial gradients of the sub-block based motion prediction signal for the current block, or one or more motion vector difference values associated with sub-blocks of the current block. At block 1815, the encoder 100 or 300 and / or the decoder 200 or 500 can obtain a refinement signal for the current block based on one or more of the obtained spatial gradients, or one or more motion vector difference values associated with sub-blocks of the current block. At block 1820, the encoder 100 or 300 and / or the decoder 200 or 500 can obtain a refined motion prediction signal for the current block based on the sub-block based motion prediction signal and the refinement signal. In certain embodiments, the encoder 100 or 300 can encode the current block based on the refined motion prediction signal, or the decoder 200 or 500 can decode the current block based on the refined motion prediction signal. The refined motion prediction signal can be a refined motion inter prediction signal generated (e.g., by the GBi encoder 300 and / or the GBi decoder 500), and one or more PROF operations can be used.

[0177] For example, in certain representative embodiments associated with other methods described herein including methods 1850 and 1900, obtaining a sub-block based motion prediction signal for a current block of video can include generating the sub-block based motion prediction signal.

[0178] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, obtaining one or more spatial gradients of a sub-block based motion prediction signal for a current block, or one or more motion vector difference values associated with sub-blocks of the current block, can include determining one or more spatial gradients (e.g., associated with a gradient filter) of the sub-block based motion prediction signal.

[0179] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, obtaining one or more spatial gradients of a sub-block based motion prediction signal for a current block, or one or more motion vector difference values associated with sub-blocks of the current block, can include determining one or more motion vector difference values associated with sub-blocks of the current block.

[0180] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, obtaining a refinement signal for a current block based on one or more determined spatial gradients or one or more determined motion vector difference values can include determining, based on the determined spatial gradients, a motion prediction refinement signal for the current block as the refinement signal.

[0181] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, obtaining a refinement signal for a current block based on one or more determined spatial gradients or one or more determined motion vector difference values can include determining, based on the determined motion vector difference values, a motion prediction refinement signal for the current block as the refinement signal.

[0182] The term "determine" or "determining" related to something such as information can generally include one or more of inferring, calculating, predicting, obtaining, and / or retrieving information. For example, determining can particularly refer to retrieving something from a memory or a bitstream.

[0183] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, obtaining a refined motion prediction signal for a current block based on a sub-block based motion prediction signal and a refinement signal can include combining (e.g., particularly adding or subtracting) the sub-block based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block.

[0184] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, encoding and / or decoding a current block based on the refined motion prediction signal can include using the refined motion prediction signal as a prediction for the current block to encode a video and / or using the refined motion prediction signal as a prediction for the current block to decode a video.

[0185] FIG. 18B is a flowchart showing a second representative encoding and / or decoding method.

[0186] Referring to FIG. 18B, a representative method 1850 for encoding and / or decoding video can include, at block 1855, the encoder 100 or 300 and / or the decoder 200 or 500 generating a sub-block based motion prediction signal. At block 1860, the encoder 100 or 300 and / or the decoder 200 or 500 can determine one or more spatial gradients (e.g., associated with a gradient filter) of the sub-block based motion prediction signal. At block 1865, the encoder 100 or 300 and / or the decoder 200 or 500 can determine a motion prediction refinement signal for the current block based on the determined spatial gradients. At block 1870, the encoder 100 or 300 and / or the decoder 200 or 500 can combine (e.g., specifically add or subtract) the sub-block based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. At block 1875, the encoder 100 or 300 can encode the video using the refined motion prediction signal as a prediction for the current block, and / or the decoder 200 or 500 can decode the video using the refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations at blocks 1810, 1820, 1830, and 1840 can be performed with respect to a current block, which is a block generally referring to the currently encoded or decoded block. The refined motion prediction signal can be a refined motion inter prediction signal generated (e.g., by the GBi encoder 300 and / or the GBi decoder 500) and may use one or more PROF operations.

[0187] For example, the determination of one or more spatial gradients of the sub-block based motion prediction signal by the encoder 100 or 300 and / or the decoder 200 or 500 can include the determination of a first set of spatial gradients associated with a first reference picture and a second set of spatial gradients associated with a second reference picture. The determination of the motion prediction refinement signal for the current block by the encoder 100 or 300 and / or the decoder 200 or 500 can be based on the determined spatial gradients, and can include the determination of the motion inter-prediction refinement signal (e.g., bidirectional prediction signal) for the current block based on the first set and the second set of spatial gradients, and can also be based on the weight information W (e.g., indicating or including one or more weight values associated with one or more reference pictures).

[0188] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the encoder 100 or 300 can generate, use, and / or transmit the weight information W to the decoder 200 or 500, and / or the decoder 200 or 500 can receive or obtain the weight information W. For example, the motion inter-prediction refinement signal for the current block can be based on (1) a first gradient value derived from the first set of spatial gradients and weighted according to a first weight coefficient indicated by the weight information W, and / or (2) a second gradient value derived from the second set of spatial gradients and weighted according to a second weight coefficient indicated by the weight information W.

[0189] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can further include that the encoder 100 or 300 and / or the decoder 200 or 500 determines affine motion model parameters for the current block of the video, such that a sub-block based motion prediction signal can be generated using the determined affine motion model parameters.

[0190] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include the determination of one or more spatial gradients of the sub-block based motion prediction signal by the encoder 100 or 300 and / or the decoder 200 or 500, which determination can include the calculation of at least one gradient value for each respective sample position, a part of each respective sample position, or each respective sample position in at least one sub-block of the sub-block based motion prediction signal. For example, the calculation of at least one gradient value for each respective sample position, a part of each respective sample position, or each respective sample position in at least one sub-block of the sub-block based motion prediction signal can include applying a gradient filter to each respective sample position, a part of each respective sample position, or each respective sample position with respect to the respective sample positions in at least one sub-block of the sub-block based motion prediction signal.

[0191] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can further include the encoder 100 or 300 and / or the decoder 200 or 500 determining a set of motion vector difference values associated with the sample positions of the first sub-block of the current block of the sub-block based motion prediction signal. In some examples, the difference values may be determined for a sub-block (e.g., the first sub-block) and reused for some or all of the other sub-blocks within the current block. In a particular example, an affine motion model or a different motion model (e.g., another sub-block based motion model such as the SbTMVP model) may be used to generate the sub-block based motion prediction signal and determine the set of motion vector difference values. By way of example, the set of motion vector difference values may be determined for the first sub-block of the current block and used to determine a motion prediction refinement signal for one or more additional sub-blocks of the current block.

[0192] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, one or more spatial gradients of the sub-block based motion prediction signal, and the set of motion vector difference values, can be used to determine a motion prediction refinement signal for the current block.

[0193] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, an affine motion model for the current block is used to generate the sub-block based motion prediction signal and determine the set of motion vector difference values.

[0194] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the determination of one or more spatial gradients of the sub-block based motion prediction signal can include determining an extended sub-block for each of one or more respective sub-blocks of the current block using the sub-block based motion prediction signal and neighboring reference samples that border and surround the sub-block, and determining the spatial gradient of each sub-block using the determined extended sub-block to determine a motion prediction refinement signal.

[0195] FIG. 19 is a flowchart illustrating a third representative encoding and / or decoding method.

[0196] Referring to FIG. 19, a representative method 1900 for encoding and / or decoding video may include, in block 1910, an encoder 100 or 300 and / or a decoder 200 or 500 generating a sub-block based motion prediction signal. In block 1920, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a set of motion vector difference values associated with sub-blocks of a current block (e.g., the set of motion vector difference values may be associated with all of the sub-blocks of the current block). In block 1930, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined set of motion vector difference values. In block 1940, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., specifically add or subtract) the sub-block based motion prediction signal and the motion prediction refinement signal to create or generate a refined motion prediction signal for the current block. In block 1950, the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block, and / or the decoder 200 or 500 may decode the video using the refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 1910, 1920, 1930, and 1940 may be performed with respect to a current block that generally refers to the block currently being encoded or decoded. In certain representative embodiments, the refined motion prediction signal may be a refined motion inter prediction signal generated (e.g., by the GBi encoder 300 and / or the GBi decoder 500) and may use one or more PROF operations.

[0197] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include that the encoder 100 or 300 and / or the decoder 200 or 500 determines motion model parameters (e.g., one or more affine motion model parameters) for the current block of the video, so that a sub-block based motion prediction signal can be generated using the determined motion model parameters (e.g., affine motion model parameters).

[0198] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include the determination of one or more spatial gradients of the sub-block based motion prediction signal by the encoder 100 or 300 and / or the decoder 200 or 500. For example, the determination of one or more spatial gradients of the sub-block based motion prediction signal can include the calculation of at least one gradient value for each respective sample position, a part of each respective sample position, or each respective sample position in at least one sub-block of the sub-block based motion prediction signal. For example, the calculation of at least one gradient value for each respective sample position, a part of each respective sample position, or each respective sample position in at least one sub-block of the sub-block based motion prediction signal can include applying a gradient filter to each respective sample position, a part of each respective sample position, or each respective sample position with respect to the respective sample position in at least one sub-block of the sub-block based motion prediction signal.

[0199] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include determining a motion prediction refinement signal for a current block using, by an encoder 100 or 300 and / or a decoder 200 or 500, a gradient value associated with each sample position of one of the current block, a part of each sample position, or a spatial gradient for each respective sample position, and a determined set of motion vector difference values associated with the sample positions of the sub-blocks (e.g., any sub-block) of the current block of the sub-block motion prediction signal.

[0200] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the determination of the motion prediction refinement signal for the current block can use a gradient value associated with one or more sample positions or the spatial gradient for each sample position of one or more sub-blocks of the current block, and a determined set of motion vector difference values.

[0201] FIG. 20 is a flowchart showing a fourth representative encoding and / or decoding method.

[0202] Referring to FIG. 20, a representative method 2000 for encoding and / or decoding video can include, in block 2010, the encoder 100 or 300 and / or the decoder 200 or 500 generating a sub-block based motion prediction signal using at least a first motion vector for a first sub-block of the current block and a further motion vector for a second sub-block of the current block. In block 2020, the encoder 100 or 300 and / or the decoder 200 or 500 can calculate a first set of gradient values for a first sample position in the first sub-block of the sub-block based motion prediction signal and a second different set of gradient values for a second sample position in the first sub-block of the sub-block based motion prediction signal. In block 2030, the encoder 100 or 300 and / or the decoder 200 or 500 can determine a first set of motion vector difference values for the first sample position and a second different set of motion vector difference values for the second sample position. For example, the first set of motion vector difference values for the first sample position can indicate the difference between the motion vector at the first sample position and the motion vector of the first sub-block, and the second set of motion vector difference values for the second sample position can indicate the difference between the motion vector at the second sample position and the motion vector of the first sub-block. In block 2040, the encoder 100 or 300 and / or the decoder 200 or 500 can use the first and second sets of gradient values and the first and second sets of motion vector difference values to determine a prediction refinement signal. In block 2050, the encoder 100 or 300 and / or the decoder 200 or 500 can combine (e.g., add or subtract, in particular) the sub-block based motion prediction signal with the prediction refinement signal to create a refined motion prediction signal.In block 2060, encoder 100 or 300 can encode video using a refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 can decode video using a refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 2010, 2020, 2030, 2040, and 2050 can be performed on a current block that includes a plurality of sub-blocks.

[0203] FIG. 21 is a flowchart illustrating a fifth exemplary encoding and / or decoding method.

[0204] Referring to FIG. 21, a representative method 2100 for encoding and / or decoding video may include, in block 2110, an encoder 100 or 300 and / or a decoder 200 or 500 generating a sub-block based motion prediction signal for a current block. In block 2120, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a prediction refinement signal using optical flow information indicating refined motion of a plurality of sample positions in the current block of the sub-block based motion prediction signal. In block 2130, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract in particular) the sub-block based motion prediction signal with the prediction refinement signal to create a refined motion prediction signal. In block 2140, the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block, and / or the decoder 200 or 500 may decode the video using the refined motion prediction signal as a prediction for the current block. For example, the current block may include a plurality of sub-blocks, and the sub-block based motion prediction signal may be generated using at least a first motion vector for a first sub-block of the current block and a further motion vector for a second sub-block of the current block.

[0205] In certain exemplary embodiments including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include a determination by encoder 100 or 300 and / or decoder 200 or 500 of a prediction refinement signal that can use optical flow information. This determination can include the calculation by encoder 100 or 300 and / or decoder 200 or 500 of a first set of gradient values for a first sample position in a first sub-block of a sub-block based motion prediction signal, and a second different set of gradient values for a second sample position in the first sub-block of the sub-block based motion prediction signal. A first set of motion vector difference values for the first sample position, and a second different set of motion vector difference values for the second sample position can be determined. For example, the first set of motion vector difference values for the first sample position can indicate the difference between the motion vector at the first sample position and the motion vector of the first sub-block, and the second set of motion vector difference values for the second sample position can indicate the difference between the motion vector at the second sample position and the motion vector of the first sub-block. Encoder 100 or 300 and / or decoder 200 or 500 can use the first and second sets of gradient values and the first and second sets of motion vector difference values to determine the prediction refinement signal.

[0206] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include the determination by encoder 100 or 300 and / or decoder 200 or 500 of a prediction refinement signal that can use optical flow information. This determination can include calculating a third set of gradient values for a first sample position in a second sub-block of the sub-block based motion prediction signal, and a fourth set of gradient values for a second sample position in the second sub-block of the sub-block based motion prediction signal. Encoder 100 or 300 and / or decoder 200 or 500 can use the third and fourth sets of gradient values and the first and second sets of motion vector difference values to determine a prediction refinement signal for the second sub-block.

[0207] FIG. 22 is a flowchart illustrating a sixth representative encoding and / or decoding method.

[0208] Referring to FIG. 22, a representative method 2200 for encoding and / or decoding video can include, in block 2210, an encoder 100 or 300 and / or a decoder 200 or 500 determining a motion model for a current block of video. The current block can include a plurality of sub-blocks. For example, the motion model can generate individual (e.g., per-sample) motion vectors for a plurality of sample positions in the current block. In block 2220, the encoder 100 or 300 and / or the decoder 200 or 500 can use the determined motion model to generate a sub-block-based motion prediction signal for the current block. The generated sub-block-based motion prediction signal can use one motion vector for each sub-block of the current block. In block 2230, the encoder 100 or 300 and / or the decoder 200 or 500 can calculate gradient values by applying a gradient filter to a portion of the plurality of sample positions of the sub-block-based motion prediction signal. In block 2240, the encoder 100 or 300 and / or the decoder 200 or 500 can determine motion vector difference values for a portion of the sample positions, each of the motion vector difference values indicating the difference between the motion vector (e.g., the individual motion vector) generated for each sample position according to the motion model and the motion vector used to create the sub-block-based motion prediction signal for the sub-block including each sample position. In block 2250, the encoder 100 or 300 and / or the decoder 200 or 500 can use the gradient values and the motion vector difference values to determine a prediction refinement signal. In block 2260, the encoder 100 or 300 and / or the decoder 200 or 500 can combine (e.g., add or subtract, in particular) the sub-block-based motion prediction signal with the prediction refinement signal to create a refined motion prediction signal for the current block.In block 2270, encoder 100 or 300 can encode video using a refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 can decode video using a refined motion prediction signal as a prediction for the current block.

[0209] FIG. 23 is a flowchart showing a seventh exemplary encoding and / or decoding method.

[0210] Referring to FIG. 23, a representative method 2300 for encoding and / or decoding video can include, in block 2310, the encoder 100 or 300 and / or the decoder 200 or 500 performing sub-block based motion compensation to generate a sub-block based motion prediction signal as a rough motion prediction signal. In block 2320, the encoder 100 or 300 and / or the decoder 200 or 500 can calculate one or more spatial gradients of the sub-block based motion prediction signal at the sample positions. In block 2330, the encoder 100 or 300 and / or the decoder 200 or 500 can calculate the per-pixel intensity change in the current block based on the calculated spatial gradients. In block 2340, the encoder 100 or 300 and / or the decoder 200 or 500 can determine a per-pixel based motion prediction signal as a refined motion prediction signal based on the calculated per-pixel intensity change. In block 2350, the encoder 100 or 300 and / or the decoder 200 or 500 can predict the current block using the rough motion prediction signal for each sub-block of the current block and using the refined motion prediction signal for each pixel of the current block. In certain embodiments, the operations in blocks 2310, 2320, 2330, 2340, and 2350 can be performed for at least one block (e.g., the current block) in the video. For example, calculating the per-pixel intensity change in the current block can include determining the luminance intensity change for each pixel in the current block according to an optical flow formula. Predicting the current block can include predicting the motion vector for each respective pixel in the current block by combining the rough motion prediction vector for the sub-block containing each pixel with the refined motion prediction vector for the rough motion prediction vector, where the refined motion prediction vector is associated with each pixel.

[0211] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, one or more spatial gradients of the sub-block based motion prediction signal can include either a horizontal gradient and / or a vertical gradient. For example, the horizontal gradient may be calculated as the difference in luminance or color difference between the right adjacent sample of the samples of the sub-block and the left adjacent sample of the samples of the sub-block, and / or the vertical gradient may be calculated as the difference in luminance or color difference between the lower adjacent sample of the samples of the sub-block and the upper adjacent sample of the samples of the sub-block.

[0212] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, one or more spatial gradients of the sub-block prediction can be generated using a Sobel filter.

[0213] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the coarse motion prediction signal can use either a 4-parameter affine model or a 6-parameter affine model. For example, sub-block based motion compensation can be either (1) affine sub-block based motion compensation, or (2) another compensation (e.g., sub-block based temporal motion vector prediction (SbTMVP) mode motion compensation, and / or regression based motion vector field (RMVF) mode based compensation). Given that SbTMVP mode based motion compensation is performed, this method can include estimating affine model parameters using a sub-block motion vector field by linear regression operations, and deriving pixel level motion vectors using the estimated affine model parameters. Given that RMVF mode based motion compensation is performed, this method can include estimating affine model parameters, and deriving pixel level motion vector offsets from sub-block level motion vectors using the estimated affine model parameters. For example, the pixel motion vector offset can be with respect to the center of the sub-block (e.g., the actual center, or the sample position closest to the actual center). For example, the coarse motion prediction vector for a sub-block can be based on the actual center position of the sub-block.

[0214] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods can include selecting, as the central position associated with the rough motion prediction vector for each sub-block (e.g., sub-block based motion prediction vector), either (1) the actual center of each sub-block, or (2) one of the pixel (e.g., sample) positions closest to the center of the sub-block. For example, predicting the current block using the rough motion prediction signal (e.g., sub-block based motion prediction signal) of the current block and using the refined motion prediction signal of each pixel (e.g., sample) of the current block can be based on the selected central position of each sub-block. For example, the encoder 100 or 300 and / or the decoder 200 or 500 can determine the central position associated with the color difference pixels of the sub-block and, based on the color difference position sample type associated with the color difference pixels, can determine an offset with respect to the central position of the color difference pixels of the sub-block. The rough motion prediction signal for the sub-block (e.g., sub-block based motion prediction signal) can be based on the actual position of the sub-block corresponding to the determined central position of the color difference pixels adjusted by the offset.

[0215] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods can include the encoder 100 or 300 generating or the decoder 200 or 500 receiving information indicating whether prediction refinement by optical flow (PROF) is enabled in one of (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header. For example, under the condition that PROF is enabled, a refined motion prediction operation may be performed, and thus, a coarse motion prediction signal (e.g., a sub-block based motion prediction signal) and a refined motion prediction signal may be used to predict the current block. As another example, under the condition that PROF is not enabled, the refined motion prediction operation is not performed, and thus, only a coarse motion prediction signal (e.g., a sub-block based motion prediction signal) may be used to predict the current block.

[0216] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 determining whether to perform a refined motion prediction operation for the current block or in affine motion estimation based on the attributes of the current block and / or the attributes of affine motion estimation.

[0217] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods can include determining whether to perform a refined motion prediction operation on a current block or in affine motion estimation based on attributes of the current block and / or attributes of affine motion estimation. For example, determining whether to perform a refined motion prediction operation on a current block based on attributes of the current block can include determining whether to perform a refined motion prediction operation on the current block based on either (1) whether the size of the current block exceeds a specific size and / or (2) whether the control point motion vector (CPMV) difference exceeds a threshold.

[0218] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 applying a first deblocking filter to one or more boundaries of sub-blocks of the current block that coincide with the transform unit boundary and applying a second different deblocking filter to other boundaries of sub-blocks of the current block that do not coincide with any transform unit boundary. For example, the first deblocking filter can be a stronger deblocking filter than the second deblocking filter.

[0219] FIG. 24 is a flowchart showing an eighth representative encoding and / or decoding method.

[0220] Referring to FIG. 24, a representative method 2400 for encoding and / or decoding video may include, in block 2410, the encoder 100 or 300 and / or the decoder 200 or 500 performing sub-block based motion compensation to generate a sub-block based motion prediction signal as a rough motion prediction signal. In block 2420, the encoder 100 or 300 and / or the decoder 200 or 500 can determine, for each respective boundary sample of a sub-block of the current block, one or more reference samples surrounding the sub-block corresponding to samples adjacent to each respective boundary sample as surrounding reference samples, and use the surrounding reference samples and the samples of the sub-block adjacent to each respective boundary sample to determine one or more spatial gradients associated with each respective boundary sample. In block 2430, the encoder 100 or 300 and / or the decoder 200 or 500 can determine, for each respective non-boundary sample in the sub-block, one or more spatial gradients associated with each respective non-boundary sample using the samples of the sub-block adjacent to each respective non-boundary sample. In block 2440, the encoder 100 or 300 and / or the decoder 200 or 500 can calculate the intensity change for each pixel in the current block using the determined spatial gradients of the sub-block. In block 2450, the encoder 100 or 300 and / or the decoder 200 or 500 can determine a pixel-by-pixel based motion prediction signal as a refined motion prediction signal based on the calculated intensity change for each pixel. In block 2460, the encoder 100 or 300 and / or the decoder 200 or 500 can predict the current block using the rough motion prediction signal associated with each sub-block of the current block and the refined motion prediction signal associated with each pixel of the current block.In certain embodiments, the operations at 2410, 2420, 2430, 2440, 2450, and 2460 may be performed on at least one block (e.g., the current block) in the video.

[0221] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, the determination of one or more spatial gradients of boundary samples and non-boundary samples can include calculating one or more spatial gradients using any of (1) a vertical Sobel filter, (2) a horizontal Sobel filter, or (3) a 3-tap filter.

[0222] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 copying the surrounding reference samples from the reference store without any further operation, and the determination of one or more spatial gradients associated with each boundary sample can use the copied, surrounding reference samples to determine one or more spatial gradients associated with each boundary sample.

[0223] FIG. 25 is a flowchart showing a representative gradient calculation method.

[0224] Referring to FIG. 25, a representative method 2500 of calculating the gradient of a sub-block using reference samples corresponding to samples near the boundary of the sub-block (e.g., used for encoding and / or decoding video) may include, in block 2510, the encoder 100 or 300 and / or the decoder 200 or 500 determining, for each respective boundary sample of the sub-blocks of the current block, one or more reference samples surrounding the sub-block corresponding to the samples near each respective boundary sample as surrounding reference samples, and determining one or more spatial gradients associated with each respective boundary sample using the surrounding reference samples and the samples of the sub-block near each respective boundary sample. In block 2520, the encoder 100 or 300 and / or the decoder 200 or 500 may determine one or more spatial gradients associated with each respective non-boundary sample of the sub-block using the samples of the sub-block near each respective non-boundary sample. In certain embodiments, the operations in blocks 2510 and 2520 may be performed on at least one block (e.g., the current block) in the video.

[0225] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, 2500, and 2600, the one or more determined spatial gradients may be used to predict the current block by any one of (1) a prediction refinement by optical flow (PROF) operation, (2) a bidirectional optical flow operation, or (3) an affine motion estimation operation.

[0226] FIG. 26 is a flowchart showing a ninth representative encoding and / or decoding method.

[0227] Referring to FIG. 26, a representative method 2600 for encoding and / or decoding video can include an encoder 100 or 300 and / or a decoder 200 or 500 generating a sub-block based motion prediction signal for a current block of the video. For example, the current block can include a plurality of sub-blocks. In block 2620, the encoder 100 or 300 and / or the decoder 200 or 500 can use, for one or more respective sub-blocks or each respective sub-block of the current block, a sub-block based motion prediction signal and proximity reference samples surrounding the sub-block adjacent to the sub-block to determine an extended sub-block, and use the determined extended sub-block to determine the spatial gradient of each sub-block. In block 2630, the encoder 100 or 300 and / or the decoder 200 or 500 can determine a motion prediction refinement signal for the current block based on the determined spatial gradient. In block 2640, the encoder 100 or 300 and / or the decoder 200 or 500 can combine (e.g., specifically add or subtract) the sub-block based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. In block 2650, the encoder 100 or 300 can use the refined motion prediction signal to encode the video as a prediction for the current block, and / or the decoder 200 or 500 can use the refined motion prediction signal to decode the video as a prediction for the current block. In certain embodiments, the operations in blocks 2610, 2620, 2630, 2640, and 2650 can be performed for at least one block (e.g., the current block) in the video.

[0228] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 copying proximity reference samples from a reference store without any further operations. For example, the determination of the spatial gradient of each sub-block can use the copied proximity reference samples to determine the gradient values associated with the sample positions on the boundaries of each sub-block. The proximity reference samples of the extended block can be copied from the nearest integer position in the reference picture including the current block. In a particular example, the proximity reference samples of the extended block have the nearest integer motion vector rounded from the original accuracy.

[0229] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 determining affine motion model parameters for the current block of the video and using the determined affine motion model parameters such that a sub-block based motion prediction signal can be generated.

[0230] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2500, and 2600, the determination of the spatial gradient of each sub-block can include calculating at least one gradient value for each respective sample position in each sub-block. For example, the calculation of at least one gradient value for each respective sample position in each sub-block can include applying a gradient filter to each respective sample position with respect to each respective sample position in each sub-block. As another example, the calculation of at least one gradient value for each respective sample position in each sub-block can include determining an intensity change for each respective sample position in each sub-block according to an optical flow formula.

[0231] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 determining a set of motion vector difference values associated with the sample positions of each sub-block. For example, an affine motion model can be used for the current block to generate a sub-block-based motion prediction signal, and a set of motion vector difference values can be determined.

[0232] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the set of motion vector difference values can be determined for each sub-block of the current block and can be used to determine a motion prediction refinement signal for the other remaining sub-blocks of the current block.

[0233] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, determination of the spatial gradient of each sub-block can include calculating the spatial gradient using any of (1) a vertical Sobel filter, (2) a horizontal Sobel filter, and / or (3) a 3-tap filter.

[0234] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, proximity reference samples adjacent to and surrounding each sub-block can use integer motion compensation.

[0235] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the spatial gradient of each sub-block can include either a horizontal gradient or a vertical gradient. For example, the horizontal gradient may be calculated as the difference in luminance or color difference between the right adjacent sample of each sample and the left adjacent sample of each sample, and / or the vertical gradient may be calculated as the difference in luminance or color difference between the bottom adjacent sample of each sample and the top adjacent sample of each sample.

[0236] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the sub-block based motion prediction signal can be generated using any of (1) a 4-parameter affine model, (2) a 6-parameter affine model, (3) sub-block based temporal motion vector prediction (SbTMVP) mode motion compensation, or (3) regression based motion compensation. For example, given the condition that SbTMVP mode motion compensation is performed, this method can include estimating affine model parameters using a linear regression operation with a sub-block motion vector field and / or deriving pixel level motion vectors using the estimated affine model parameters. As another example, given the condition that RMVF mode based motion compensation is performed, this method can include estimating affine model parameters and / or deriving pixel level motion vector offsets from sub-block level motion vectors using the estimated affine model parameters. The pixel motion vector offsets can be with respect to the center of each sub-block.

[0237] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the refined motion prediction signal for each sub-block can be based on the actual center position of each sub-block or on the sample position closest to the actual center of each sub-block.

[0238] For example, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 selecting, as the center position associated with the motion prediction vector for each respective sub-block, one of (1) the actual center of each respective sub-block or (2) the sample position closest to the actual center of each sub-block. The refined motion prediction signal can be based on the selected center position of each sub-block.

[0239] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 determining a central position associated with the color difference pixels of each sub-block and determining an offset of the color difference pixels of each sub-block relative to the central position based on the color difference position sample type associated with the color difference pixels. The refined prediction signal for each sub-block can be based on the actual position of the sub-block corresponding to the determined central position of the color difference pixels adjusted by the offset.

[0240] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the encoder 100 or 300 can generate and transmit information indicating whether prediction refinement by optical flow (PROF) is enabled in one of (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header, and / or the decoder 200 or 500 can receive information indicating whether PROF is enabled in one of (1) the SPS header, (2) the PPS header, or (3) the tile group header.

[0241] FIG. 27 is a flowchart showing a tenth representative encoding and / or decoding method.

[0242] Referring to FIG. 27, a representative method 2700 for encoding and / or decoding video may include, in block 2710, the encoder 100 or 300 and / or the decoder 200 or 500 determining the actual center position of each respective sub-block of the current block. In block 2720, the encoder 100 or 300 and / or the decoder 200 or 500 may use the actual center position of each respective sub-block of the current block to generate a sub-block based motion prediction signal or a refined motion prediction signal. In block 2730, (1) the encoder 100 or 300 may encode the video using the sub-block based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block, or (2) the decoder 200 or 500 may decode the video using the sub-block based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 2710, 2720, and 2730 may be performed for at least one block (e.g., the current block) in the video. For example, determining the actual center position of each respective sub-block of the current block may include determining the center position of the color difference of the color difference pixels associated with each respective sub-block, and the offset of the center position of the color difference with respect to the center position of each respective sub-block, based on the color difference position sample type of the color difference pixels. The sub-block based motion prediction signal or the refined motion prediction signal for each respective sub-block may be based on the actual center position of each respective sub-block corresponding to the determined center position of the color difference adjusted by the offset. Although the actual center of each respective sub-block of the current block is described as being determined / used for various operations, it is contemplated that one, some, or all of such center positions of the sub-blocks may be determined / used.

[0243] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the generation of a refined motion prediction signal can use a sub-block based motion prediction signal by determining one or more spatial gradients of the sub-block based motion prediction signal for each respective sub-block of the current block, determining a motion prediction refinement signal for the current block based on the determined spatial gradients, and / or combining the sub-block based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. For example, determining one or more spatial gradients of the sub-block based motion prediction signal can include determining an extended sub-block using the sub-block based motion prediction signal and neighboring reference samples that contact and surround the sub-block for each respective sub-block, and / or determining one or more spatial gradients of each respective sub-block using the determined extended sub-block.

[0244] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, determining the spatial gradient of each respective sub-block can include calculating at least one gradient value for each respective sample position in each respective sub-block. For example, calculating at least one gradient value for each respective sample position in each respective sub-block can include applying a gradient filter to each respective sample position with respect to each respective sample position in each respective sub-block.

[0245] As another example, calculating at least one gradient value for each respective sample position in each respective sub-block can include determining an intensity change for one or more respective sample positions in each respective sub-block according to an optical flow formula.

[0246] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, these methods can include the encoder 100 or 300 and / or the decoder 200 or 500 determining a set of motion vector difference values associated with the sample positions of the respective sub-blocks. An affine motion model can be used for the current block to generate a sub-block based motion prediction signal, and a set of motion vector difference values can be determined. In a particular example, the set of motion vector difference values can be determined for each sub-block of the current block and used (e.g., reused) to determine a motion prediction refinement signal for that sub-block and the other remaining sub-blocks of the current block. For example, the determination of the spatial gradient of each sub-block can include calculating the spatial gradient using any of (1) a vertical Sobel filter, (2) a horizontal Sobel filter, and / or (3) a 3-tap filter. Proximity reference samples surrounding the sub-block adjacent to each sub-block can use integer motion compensation.

[0247] In some embodiments, the spatial gradient of each sub-block can include either a horizontal gradient or a vertical gradient. For example, the horizontal gradient can be calculated as the difference in luminance or color difference between the right adjacent sample of each sample and the left adjacent sample of each sample. As another example, the vertical gradient can be calculated as the difference in luminance or color difference between the lower adjacent sample of each sub-block and the upper adjacent sample of each sub-block.

[0248] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the sub-block based motion prediction signal can be generated using any of (1) a 4-parameter affine model, (2) a 6-parameter affine model, (3) sub-block based temporal motion vector prediction (SbTMVP) mode motion compensation, and / or (3) regression based motion compensation. For example, given the condition that SbTMVP mode motion compensation is performed, this method can include estimating affine model parameters using a sub-block motion vector field by linear regression operations and / or deriving pixel-level motion vectors using the estimated affine model parameters. As another example, given the condition that regression motion vector field (RMVF) mode based motion compensation is performed, this method can include estimating affine model parameters and / or deriving pixel-level motion vector offsets from sub-block level motion vectors using the estimated affine model parameters, where the pixel motion vector offsets are with respect to the center of each sub-block.

[0249] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the refined motion prediction signal can be generated using a plurality of motion vectors associated with the control points of the current block.

[0250] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the encoder 100 or 300 can generate, encode, and transmit information indicating whether prediction refinement by optical flow (PROF) is enabled in one of (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header, and the decoder 200 or 500 can receive and decode the information.

[0251] Figure 28 is a flowchart showing an eleventh representative encoding and / or decoding method.

[0252] Referring to FIG. 28, a representative method 2800 for encoding and / or decoding video can include, in block 2810, the encoder 100 or 300 and / or the decoder 200 or 500 selecting, as a central position associated with a motion prediction vector for each respective sub-block, either (1) the actual center of each respective sub-block, or (2) one of the sample positions closest to the actual center of each respective sub-block. In block 2820, the encoder 100 or 300 and / or the decoder 200 or 500 can determine the selected central position for each respective sub-block of the current block. In block 2830, the encoder 100 or 300 and / or the decoder 200 or 500 can generate a sub-block based motion prediction signal or a refined motion prediction signal using the selected central position for each respective sub-block of the current block. In block 2840, (1) the encoder 100 or 300 can encode the video using the sub-block based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block, or (2) the decoder 200 or 500 can decode the video using the sub-block based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 2810, 2820, 2830, and 2840 can be performed for at least one block (e.g., the current block) in the video. Although the selection of the central position for each respective sub-block of the current block is described, it is contemplated that one, some, or all of the central positions of such sub-blocks can be selected / used in various operations.

[0253] FIG. 29 is a flowchart showing a representative encoding method.

[0254] Referring to FIG. 29, a representative method 2900 for encoding a video can include, at block 2910, the encoder 100 or 300 performing motion estimation for the current block of the video, which can include using an iterative motion compensation operation to determine the affine motion model parameters for the current block and using the determined affine motion model parameters to generate a sub-block based motion prediction signal for the current block. At block 2920, after the encoder 100 or 300 performs motion estimation for the current block, it can perform a prediction refinement by optical flow (PROF) operation to generate a refined motion prediction signal. At block 2930, the encoder 100 or 300 can encode the video using the refined motion prediction signal as a prediction for the current block. For example, the PROF operation can include determining one or more spatial gradients of the sub-block based motion prediction signal, determining a motion prediction refinement signal for the current block based on the determined spatial gradients, and / or combining the sub-block based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block.

[0255] In certain representative embodiments, including at least the representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, and 2900, the PROF operation can be performed after (e.g., only after) the iterative motion compensation operation is complete. For example, the PROF operation is not performed during motion estimation for the current block.

[0256] FIG. 30 is a flowchart showing another representative encoding method.

[0257] Referring to FIG. 30, a representative method 3000 for encoding a video can include, in block 3010, the encoder 100 or 300 determining affine motion model parameters using an iterative motion compensation operation during motion estimation for the current block, and generating a sub-block based motion prediction signal using the determined affine motion model parameters. In block 3020, the encoder 100 or 300 can perform a prediction refinement by optical flow (PROF) operation to generate a refined motion prediction signal, after motion estimation for the current block, on the condition that the size of the current block meets or exceeds a threshold size. In block 3030, the encoder 100 or 300 can encode the video using, as a prediction for the current block, (1) the refined motion prediction signal if the current block meets or exceeds the threshold size, or (2) the sub-block based motion prediction signal if the current block does not meet the threshold size.

[0258] FIG. 31 is a flowchart showing a twelfth representative encoding / decoding method.

[0259] Referring to FIG. 31, a representative method 3100 for encoding and / or decoding video may include, in block 3110, the encoder 100 or 300 determining or obtaining information indicating the size of the current block, or the decoder 200 or 500 receiving information indicating the size of the current block. In block 3120, the encoder 100 or 300 or the decoder 200 or 500 can generate a sub-block based motion prediction signal. In block 3130, the encoder 100 or 300 or the decoder 200 or 500 can perform a prediction refinement by optical flow (PROF) operation to generate a refined motion prediction signal on the condition that the size of the current block meets or exceeds a threshold size. In block 3140, the encoder 100 or 300 can encode the video using, as a prediction for the current block, (1) the refined motion prediction signal if the current block meets or exceeds the threshold size, or (2) the sub-block based motion prediction signal if the current block does not meet the threshold size, or the decoder 200 or 500 can decode the video using, as a prediction for the current block, (1) the refined motion prediction signal if the current block meets or exceeds the threshold size, or (2) the sub-block based motion prediction signal if the current block does not meet the threshold size.

[0260] FIG. 32 is a flowchart showing a 13th representative encoding / decoding method.

[0261] Referring to FIG. 32, a representative method 3200 for encoding and / or decoding video can include, at block 3210, the encoder 100 or 300 determining whether pixel-level motion compensation should be performed, or the decoder 200 or 500 receiving a flag indicating whether pixel-level motion compensation should be performed. At block 3220, the encoder 100 or 300 or the decoder 200 or 500 can generate a sub-block-based motion prediction signal. At block 3230, on the condition that pixel-level motion compensation should be performed, the encoder 100 or 300 or the decoder 200 or 500 determines one or more spatial gradients of the sub-block-based motion prediction signal, determines a motion prediction refinement signal for the current block based on the determined spatial gradients, and can combine the sub-block-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. At block 3240, in accordance with the determination of whether pixel-level motion compensation should be performed, the encoder 100 or 300 can encode the video using the sub-block-based motion prediction signal or the refined motion prediction signal as a prediction for the current block, or the decoder 200 or 500 can decode the video using the sub-block-based motion prediction signal or the refined motion prediction signal as a prediction for the current block in accordance with the indication of the flag. In certain embodiments, the operations at blocks 3220 and 3230 can be performed on blocks (e.g., the current block) in the video.

[0262] FIG. 33 is a flowchart showing a fourteenth representative encoding / decoding method.

[0263] Referring to FIG. 33, a representative method 3300 for encoding and / or decoding video can include, at block 3310, an encoder 100 or 300 determining or obtaining, or a decoder 200 or 500 receiving, inter-prediction weight information indicating one or more weights associated with a first reference picture and a second reference picture. At block 3320, an encoder 100 or 300 or a decoder 200 or 500 can generate a sub-block based motion inter-prediction signal for a current block of the video, determine a first set of spatial gradients associated with the first reference picture and a second set of spatial gradients associated with the second reference picture, and based on the first set and the second set of spatial gradients and the inter-prediction weight information, determine a motion inter-prediction refinement signal for the current block, and combine the sub-block based motion inter-prediction signal and the motion inter-prediction refinement signal to create a refined motion inter-prediction signal for the current block. At block 3330, an encoder 100 or 300 can encode the video using the refined motion inter-prediction signal as a prediction for the current block, or a decoder 200 or 500 can decode the video using the refined motion inter-prediction signal as a prediction for the current block. For example, the inter-prediction weight information can be either (1) an indicator indicating a first weight coefficient applied to the first reference picture and / or a second weight coefficient applied to the second reference picture, or (2) a weight index. In certain embodiments, the motion inter-prediction refinement signal for the current block can be based on (1) a first gradient value derived from the first set of spatial gradients and weighted according to the first weight coefficient indicated by the inter-prediction weight information, and (2) a second gradient value derived from the second set of spatial gradients and weighted according to the second weight coefficient indicated by the inter-prediction weight information.

[0264] Exemplary Network for Implementing an Aspect FIG. 34A is a diagram illustrating an exemplary communication system 3400 in which one or more of the disclosed aspects may be implemented. The communication system 3400 can be a multi-connection system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 3400 can enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, the communication system 3400 can utilize one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT-spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multi-carrier (FBMC), etc.

[0265] As shown in FIG. 34A, the communication system 3400 can include wireless transmit / receive units (WTRUs) 3402a, 3402b, 3402c, 3402d, a RAN 3404 / 3413, a CN 3406 / 3415, a public switched telephone network (PSTN) 3408, the Internet 3410, and other networks 3412, although it is understood that the disclosed aspects contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 3402a, 3402b, 3402c, 3402d can be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 3402a, 3402b, 3402c, 3402d may each be referred to as a “station” and / or “STA,” may be configured to transmit and / or receive wireless signals, and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscriber-based units, pagers, cellular telephones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in an industrial and / or automated processing chain context), home electronics, devices operating on commercial and / or industrial wireless networks, etc. The WTRUs 3402a, 3402b, 3402c, and 3402d may each be interchangeably referred to as a UE.

[0266] The communication system 3400 can also include base station 3414a and / or base station 3414b. Each of base stations 3414a, 3414b can be any type of device configured to wirelessly interface with at least one of WTRUs 3402a, 3402b, 3402c, 3402d to facilitate access to one or more communication networks such as CN 3406 / 3415, Internet 3410, and / or network 3412. By way of example, base stations 3414a, 3414b can be a transceiver base station (BTS), Node B, eNode B (eNB), home Node B (HNB), home eNode B (HeNB), gNB, NR Node B, site controller, access point (AP), and wireless router, among others. Although base stations 3414a, 3414b are each shown as a single element, it will be understood that base stations 3414a, 3414b can include any number of interconnected base stations and / or network elements.

[0267] The base station 3414a can be part of the RAN 3404 / 3413, which can also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. The base station 3414a and / or the base station 3414b can be configured to transmit and / or receive wireless signals at one or more carrier frequencies, sometimes referred to as a cell (not shown). These frequencies can be in the licensed spectrum, the unlicensed spectrum, or a combination of the licensed and unlicensed spectrums. A cell can provide coverage for wireless services for a particular geographic area that can be relatively fixed or can change over time. A cell can further be divided into cell sectors. For example, the cell associated with the base station 3414a can be divided into three sectors. Thus, in one embodiment, the base station 3414a can include three transceivers, i.e., one transceiver for each sector of the cell. In an embodiment, the base station 3414a can utilize multiple-input multiple-output (MIMO) technology and can utilize multiple transceivers for each sector of the cell. For example, beamforming can be used to transmit and / or receive signals in a desired spatial direction.

[0268] The base stations 3414a, 3414b can communicate with one or more of the WTRUs 3402a, 3402b, 3402c, 3402d via the air interface 3416, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 3416 can be established using any suitable radio access technology (RAT).

[0269] More specifically, as described above, the communication system 3400 can be a multi-connection system and can utilize one or more channel access methods such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base stations 3414a within RAN 3404 / 3413, and the WTRUs 3402a, 3402b, 3402c can implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) that can establish the air interfaces 3415 / 3416 / 3417 using Wideband CDMA (WCDMA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).

[0270] In an embodiment, the base station 3414a, and the WTRUs 3402a, 3402b, 3402c can implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA) that can establish the air interface 3416 using Long-Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-A Pro.

[0271] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c can implement radio technologies such as NR Radio Access that can establish the air interface 3416 using New Radio (NR).

[0272] In an embodiment, base station 3414a and WTRUs 3402a, 3402b, 3402c can implement multiple radio access technologies. For example, base station 3414a and WTRUs 3402a, 3402b, 3402c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Accordingly, the air interface utilized by WTRUs 3402a, 3402b, 3402c can be characterized by transmissions to / from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNB and gNB).

[0273] In other embodiments, base station 3414a and WTRUs 3402a, 3402b, 3402c can implement wireless technologies such as IEEE802.11 (i.e., Wireless Fidelity (WiFi)), IEEE802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0274] The base station 3414b in FIG. 34A can be, for example, a wireless router, a home node B, a home eNode B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in local areas such as workplaces, homes, vehicles, campuses, industrial facilities, aerial corridors (such as those used by drones), and roadways. In one embodiment, the base station 3414b and the WTRUs 3402c, 3402d can implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 3414b and the WTRUs 3402c, 3402d can implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 3414b and the WTRUs 3402c, 3402d can utilize a cellular-based RAT (such as, for example, WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As shown in FIG. 34A, the base station 3414b can have a direct connection to the Internet 3410. Thus, the base station 3414b may not be required to access the Internet 3410 via the CN 3406 / 3415.

[0275] RAN 3404 / 3413 can communicate with CN 3406 / 3415, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 3402a, 3402b, 3402c, 3402d. The data may have various Quality of Service (QoS) requirements such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN 3406 / 3415 can provide call control, billing services, mobile location-based services, prepaid calls, Internet connectivity, video delivery, etc., and / or perform high-level security functions such as user authentication. Although not shown in FIG. 34A, it will be understood that RAN 1084 / 3413 and / or CN 3406 / 3415 can communicate directly or indirectly with other RANs that utilize the same or a different Radio Access Technology (RAT) as RAN 3404 / 3413. For example, in addition to being connected to RAN 3404 / 3413 which may utilize New Radio (NR) radio technology, CN 3406 / 3415 can also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0276] CN3406 / 3415 can also serve as a gateway for WTRUs 3402a, 3402b, 3402c, 3402d to access the PSTN 3408, the Internet 3410, and / or other networks 3412. The PSTN 3408 can include a circuit-switched telephone network that provides basic telephone service (POTS). The Internet 3410 can include a global system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 3412 can include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 3412 can include another CN connected to one or more RANs that can utilize the same RAT or a different RAT as the RAN 3404 / 3413.

[0277] Some or all of the WTRUs 3402a, 3402b, 3402c, 3402d in the communication system 3400 can include multimode capabilities (e.g., the WTRUs 3402a, 3402b, 3402c, 3402d can include multiple transceivers for communicating with different wireless networks via different wireless links). For example, the WTRU 3402c shown in FIG. 34A can be configured to communicate with a base station 3414a that can utilize cellular-based wireless technology and a base station 3414b that can utilize IEEE 802 wireless technology.

[0278] Figure 34B is a system diagram showing an exemplary WTRU3402. As shown in Figure 34B, the WTRU3402 can include a processor 3418, a transceiver 3420, a transmit / receive element 3422, a speaker / microphone 3424, a keypad 3426, a display / touchpad 3428, a removable memory 3430, a detachable memory 3432, a power supply 3434, a global positioning system (GPS) chipset 3436, and / or other peripheral devices 3438, etc. It will be understood that the WTRU3402 can include any partial combination of the above elements while maintaining consistency with the embodiments.

[0279] The processor 3418 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. The processor 3418 can perform signal encoding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU3402 to operate in a wireless environment. The processor 3418 may be coupled to the transceiver 3420, and the transceiver 3420 may be coupled to the transmit / receive element 3422. Although Figure 34B shows the processor 3418 and the transceiver 3420 as separate components, it will be understood that the processor 3418 and the transceiver 3420 may be integrated together in an electronic package or chip. The processor 3418 may be configured to encode or decode video (e.g., video frames).

[0280] The transmit / receive element 3422 may be configured to transmit signals to, or receive signals from, a base station (e.g., base station 3414a) via the air interface 3416. For example, in one embodiment, the transmit / receive element 3422 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 3422 can be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 3422 can be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmit / receive element 3422 can be configured to transmit and / or receive any combination of wireless signals.

[0281] In FIG. 34B, the transmit / receive element 3422 is shown as a single element, but the WTRU 3402 can include any number of transmit / receive elements 3422. More specifically, the WTRU 3402 can utilize MIMO technology. Thus, in one embodiment, the WTRU 3402 can include two or more transmit / receive elements 3422 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 3416.

[0282] The transceiver 3420 can be configured to modulate signals transmitted by the transmit / receive element 3422 and demodulate signals received by the transmit / receive element 3422. As described above, the WTRU 3402 can have multi-mode capabilities. Thus, the transceiver 3420 can include multiple transceivers to enable the WTRU 3402 to communicate via multiple RATs such as, for example, NR and IEEE 802.11.

[0283] The processor 3418 of the WTRU 3402 can be coupled to and receive user input data from a speaker / microphone 3424, a keypad 3426, and / or a display / touchpad 3428 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 3418 can also output user data to the speaker / microphone 3424, the keypad 3426, and / or the display / touchpad 3428. In addition, the processor 3418 can access information from and store data in any type of suitable memory, such as a non-removable memory 3430 and / or a removable memory 3432. The non-removable memory 3430 can include a random access memory (RAM), a read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 3432 can include a subscriber identity module (SIM) card, a memory stick, and a secure digital (SD) memory card, among others. In other embodiments, the processor 3418 can access information from and store data in a memory that is not physically located on the WTRU 3402, such as on a server or a home computer (not shown).

[0284] The processor 3418 can receive power from a power source 3434 and can be configured to distribute and / or control the power to other components within the WTRU 3402. The power source 3434 can be any suitable device for powering the WTRU 3402. For example, the power source 3434 can include one or more dry cells (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), a solar cell, and a fuel cell, among others.

[0285] Processor 3418 may be coupled to a GPS chipset 3436, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 3402. In addition to or instead of information from the GPS chipset 3436, the WTRU 3402 may receive location information from a base station (e.g., base stations 3414a, 3414b) via an air interface 3416 and / or may determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 3402 may acquire location information by any suitable positioning method while maintaining consistency with the embodiments.

[0286] Processor 3418 may be further coupled to other peripheral devices 3438, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 3438 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a virtual reality and / or augmented reality (VR / AR) device, and an activity tracker, among others. The peripheral devices 3438 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geographical location sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0287] The processor 3418 of the WTRU 3402 can operably communicate with various peripheral devices 3438, including, for example, one or more accelerometers, one or more gyroscopes, a USB port, other communication interfaces / ports, a display, and / or other visual / audio indicators, to implement the representative embodiments disclosed herein.

[0288] The WTRU 3402 may include full-duplex radio, in which (for example, for both UL (e.g., for transmission) and downlink (e.g., for reception)) the transmission and reception of some or all of the signals associated with a particular subframe can be parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce or substantially eliminate self-interference via signal processing by hardware (e.g., a choke) or a processor (e.g., a separate processor (not shown) or the processor 3418). In an embodiment, the WTRU 3402 may include half-duplex radio for the transmission and reception of some or all of the signals (associated with a particular subframe for either UL (e.g., for transmission) or downlink (e.g., for reception)).

[0289] FIG. 34C is a system diagram showing the RAN 104 and the CN 3406 according to an aspect. As described above, the RAN 3404 can communicate with the WTRUs 3402a, 3402b, 3402c via the air interface 3416 using the E-UTRA radio technology. The RAN 3404 can also communicate with the CN 3406.

[0290] RAN 3404 can include eNodeBs 3460a, 3460b, 3460c, although it will be understood that RAN 3404 can include any number of eNodeBs while maintaining consistency with the embodiments. Each of eNodeBs 3460a, 3460b, 3460c can include one or more transceivers for communicating with WTRUs 3402a, 3402b, 3402c via air interface 3416. In one embodiment, eNodeBs 3460a, 3460b, 3460c can implement MIMO technology. Thus, eNodeB 3460a can transmit wireless signals to and / or receive wireless signals from WTRU 3402a, for example, using multiple antennas.

[0291] Each of eNodeBs 3460a, 3460b, 3460c can be associated with a particular cell (not shown) and can be configured to handle wireless resource management decisions, handover decisions, and user scheduling in the UL and / or DL, etc. As shown in FIG. 34C, eNodeBs 3460a, 3460b, 3460c can communicate with each other via the X2 interface.

[0292] CN 3406 shown in FIG. 34C can include a Mobility Management Entity (MME) 3462, a Serving Gateway (SGW) 3464, and a Packet Data Network (PDN) Gateway (or PGW) 3466. Although each of the above elements is shown as part of CN 3406, it will be understood that any of these elements can be owned and / or operated by an entity different from the CN operator.

[0293] The MME 3462 may be connected to each of the eNodeBs 3460a, 3460b, and 3460c within the RAN 3404 via the S1 interface and can serve as a control node. For example, the MME 3462 can be responsible for authenticating users of the WTRUs 3402a, 3402b, 3402c, bearer activation / deactivation, and selecting a specific serving gateway during the initial attach of the WTRUs 3402a, 3402b, 3402c. The MME 3462 can provide control plane functions for handover between the RAN 3404 and other RANs (not shown) that utilize other radio technologies such as GSM and / or WCDMA.

[0294] The SGW 3464 can be connected to each of the eNodeBs 3460a, 3460b, and 3460c within the RAN 104 via the S1 interface. The SGW 3464 can generally route and forward user data packets to / from the WTRUs 3402a, 3402b, 3402c. The SGW 3464 can also perform other functions such as anchoring the user plane during handover between eNodeBs, triggering paging when DL data is available to the WTRUs 3402a, 3402b, 3402c, and managing and storing the context of the WTRUs 3402a, 3402b, 3402c.

[0295] The SGW 3464 may be connected to the PGW 3466, which can provide access to a packet switched network such as the Internet 3410 to the WTRUs 3402a, 3402b, 3402c and facilitate communication between the WTRUs 3402a, 3402b, 3402c and IP-enabled devices.

[0296] CN3406 can facilitate communication with other networks. For example, CN106 can provide access to a circuit-switched network such as PSTN3408 to WTRUs 3402a, 3402b, 3402c, and can facilitate communication between WTRUs 3402a, 3402b, 3402c and conventional landline communication devices. For example, CN3406 can include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN3406 and PSTN3408. In addition, CN3406 can provide access to other network 3412 to WTRUs 3402a, 3402b, 3402c, and other network 3412 can include other wired and / or wireless networks owned and / or operated by other service providers.

[0297] In FIGS. 34A - 34D, the WTRU is described as a wireless terminal, but in certain representative embodiments, it is contemplated that such a terminal may use (e.g., temporarily or permanently) a wired communication interface with the communication network.

[0298] In a representative embodiment, other network 3412 may be a WLAN.

[0299] A WLAN in infrastructure basic service set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have a distribution system (DS) that carries traffic within and / or outside the BSS or an access or interface to another type of wired / wireless network. Traffic to an STA originating from outside the BSS can arrive via the AP and be delivered to the STA. Traffic originating from an STA destined for a destination outside the BSS can be sent to the AP and delivered to their respective destinations. Traffic between STAs within the BSS can be sent through the AP. For example, a source STA can send traffic to the AP, and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic can be sent between a source STA and a destination STA (e.g., directly between them) using a direct link setup (DLS). In certain representative embodiments, the DLS can use an 802.11e DLS or an 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode need not have an AP, and STAs within the IBSS or using the IBSS (e.g., all of the STAs) can communicate directly with each other. The IBSS communication mode may be referred to herein as the "ad hoc" communication mode.

[0300] When using the 802.11ac infrastructure operation mode or a similar operation mode, the AP can transmit beacons on a fixed channel such as the primary channel. The primary channel can be of a fixed width (e.g., a 20 MHz bandwidth), or a width dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In certain representative embodiments, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) can be implemented, for example, in an 802.11 system. In CSMA / CA, STAs including the AP (e.g., all STAs) can sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., a single station) may transmit at any given time in a given BSS.

[0301] High Throughput (HT) STAs can use a 40 MHz wide channel for communication by using, for example, a combination of adjacent or non - adjacent 20 MHz channels and the primary 20 MHz channel to form a 40 MHz wide channel.

[0302] Very High Throughput (VHT) STAs can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels can be formed by combining contiguous 20 MHz channels. 160 MHz channels can be formed by combining eight contiguous 20 MHz channels or by combining two non - contiguous 80 MHz channels sometimes called an 80 + 80 configuration. In an 80 + 80 configuration, data may be passed through a segment parser that can segment the data into two streams after channel encoding. Inverse Fast Fourier Transform (IFFT) processing and time - domain processing may be performed separately on each stream. The streams may be mapped onto two 80 MHz channels and the data may be transmitted by a transmitting STA. At the receiver of a receiving STA, the operations described above for the 80 + 80 configuration may be reversed and the combined data may be sent to the Media Access Control (MAC).

[0303] The sub - 1 GHz operation mode is supported by 802.11af and 802.11ah. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using the non - TVWS spectrum. According to an exemplary embodiment, 802.11ah can support meter - type control / machine - type communication such as MTC devices within a macro - coverage area. The MTC device may have limited capabilities including specific capabilities, e.g., support for a specific and / or limited bandwidth (e.g., support for only that). The MTC device may include a battery having a battery life above a threshold (e.g., to maintain a very long battery life).

[0304] A WLAN system that can support multiple channels and channel bandwidths such as 802.11n, 802.11ac, 802.11af, and 802.11ah includes a channel that can be designated as a primary channel. The primary channel can have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in a BSS. The bandwidth of the primary channel can be set and / or limited by the STA that supports the minimum bandwidth operating mode among all STAs operating in the BSS. In the example of 802.11ah, even if an AP and other STAs in a BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, the primary channel can be 1MHz wide for an STA (e.g., an MTC type device) that supports the 1MHz mode (e.g., supports only that). Carrier sensing and / or network allocation vector (NAV) setting can depend on the status of the primary channel. For example, if the primary channel is busy by an STA transmitting to an AP (supporting only the 1MHz operating mode), the entire available frequency band can be considered busy even if most of the frequency band remains available and idle.

[0305] In the United States, the available frequency band that can be used by 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11ah is from 6 MHz to 26 MHz depending on the country code.

[0306] FIG. 34D is a system diagram showing RAN3413 and CN3415 according to an aspect. As described above, RAN3413 can communicate with WTRUs 3402a, 3402b, 3402c via air interface 3416 using NR radio technology. RAN3413 can also communicate with CN3415.

[0307] RAN 3413 can include gNBs 3480a, 3480b, and 3480c, although it will be understood that RAN 3413 can include any number of gNBs while maintaining consistency with the embodiments. Each of gNBs 3480a, 3480b, and 3480c can include one or more transceivers for communicating with WTRUs 3402a, 3402b, and 3402c via air interface 3416. In one embodiment, gNBs 3480a, 3480b, and 3480c can implement MIMO technology. For example, gNBs 3480a and 3480b can use beamforming to transmit signals to and / or receive signals from gNBs 3480a, 3480b, and 3480c. Thus, gNB 3480a can transmit wireless signals to and / or receive wireless signals from WTRU 3402a, for example, using multiple antennas. In an embodiment, gNBs 3480a, 3480b, and 3480c can implement carrier aggregation technology. For example, gNB 3480a can transmit multiple component carriers to WTRU 3402a (not shown). A subset of these component carriers can be on unlicensed spectrum while the remaining component carriers can be on licensed spectrum. In an embodiment, gNBs 3480a, 3480b, and 3480c can implement coordinated multipoint (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNB 3480a and gNB 3480b (and / or gNB 3480c).

[0308] WTRU3402a, 3402b, and 3402c can communicate with gNB480a, 3480b, and 3480c using transmissions associated with scalable numerology. For example, the OFDM symbol interval and / or the OFDM sub-carrier interval can vary depending on different transmissions, different cells, and / or different portions of the wireless transmission spectrum. WTRU3402a, 3402b, and 3402c can communicate with gNB3480a, 3480b, and 3480c using sub-frames or transmission time intervals (TTIs) of various or scalable lengths (e.g., including various numbers of OFDM symbols and / or continuing through absolute times of various lengths).

[0309] gNBs 3480a, 3480b, and 3480c may be configured to communicate with WTRUs 3402a, 3402b, and 3402c in a stand-alone configuration and / or a non-stand-alone configuration. In a stand-alone configuration, WTRUs 3402a, 3402b, and 3402c may communicate with gNBs 3480a, 3480b, and 3480c without accessing other RANs (e.g., eNodeBs 3460a, 3460b, 3460c, etc.). In a stand-alone configuration, WTRUs 3402a, 3402b, and 3402c may utilize one or more of gNBs 3480a, 3480b, and 3480c as a mobility anchor point. In a stand-alone configuration, WTRUs 3402a, 3402b, and 3402c may communicate with gNBs 3480a, 3480b, and 3480c using signals in an unlicensed band. In a non-stand-alone configuration, WTRUs 3402a, 3402b, and 3402c may communicate / connect with gNBs 3480a, 3480b, and 3480c while also communicating / connecting with another RAN such as eNodeBs 3460a, 3460b, 3460c. For example, WTRUs 3402a, 3402b, and 3402c may implement the DC principle to communicate with one or more gNBs 3480a, 3480b, and 3480c and one or more eNodeBs 3460a, 3460b, 3460c substantially simultaneously. In a non-stand-alone configuration, eNodeBs 3460a, 3460b, 3460c may serve as a mobility anchor for WTRUs 3402a, 3402b, and 3402c, and gNBs 3480a, 3480b, and 3480c may provide additional coverage and / or throughput for serving WTRUs 3402a, 3402b, and 3402c.

[0310] Each of gNBs 3480a, 3480b, and 3480c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards user plane functions (UPFs) 3484a, 3484b, and routing of control plane information towards access and mobility management functions (AMFs) 3482a, 3482b, etc. As shown in FIG. 34D, gNBs 3480a, 3480b, and 3480c can communicate with each other via the Xn interface.

[0311] CN 3415 shown in FIG. 34D may include at least one AMF 3482a, 3482b, at least one UPF 3484a, 3484b, at least one session management function (SMF) 3483a, 3483b, and optionally data networks (DNs) 3485a, 3485b. Although each of the above elements is shown as part of CN 3415, it will be understood that any of these elements may be owned and / or operated by entities different from the CN operator.

[0312] AMF3482a and 3482b may be connected to one or more of gNB3480a, 3480b, and 3480c within RAN3413 via the N2 interface and can serve as control nodes. For example, AMF3482a and 3482b can authenticate users of WTRU3402a, 3402b, and 3402c, support network slicing (e.g., handling different protocol data unit (PDU) sessions with different requirements), select specific SMF3483a and 3483b, manage the registration area, terminate non-access stratum (NAS) signaling, and be responsible for mobility management, etc. Network slicing can be used by AMF3482a and 3482b to customize the CN support for WTRU3402a, 3402b, and 3402c based on the type of services utilized by WTRU3402a, 3402b, and 3402c. For example, different network slices can be established for different use cases such as services that rely on ultra-reliable low-latency communication (URLLC) access, services that rely on extended mobile (e.g., massive mobile) broadband (eMBB) access, and / or services for machine type communication (MTC) access. AMF3462 can provide control plane functions for switching between RAN3413 and other RANs (not shown) that utilize other radio technologies such as non-3GPP access technologies like LTE, LTE-A, LTE-A Pro, and / or WiFi.

[0313] SMF3483a and 3483b can be connected to AMF3482a and 3482b in CN3415 via the N11 interface. SMF3483a and 3483b can also be connected to UPF3484a and 3484b in CN3415 via the N4 interface. SMF3483a and 3483b can select and control UPF3484a and 3484b, and configure traffic routing via UPF3484a and 3484b. SMF3483a and 3483b can perform other functions such as managing and allocating UE IP addresses, managing PDU sessions, implementing policies and controlling QoS, and providing downlink data notifications. The PDU session type can be IP-based, non-IP-based, Ethernet-based, etc.

[0314] UPF3484a and 3484b may be connected to one or more of gNB3480a, 3480b, and 3480c in RAN3413 via the N3 interface, which can provide access to a packet-switched network such as the Internet 3410 to WTRU3402a, 3402b, and 3402c, and facilitate communication between WTRU3402a, 3402b, and 3402c and IP-compatible devices. UPF3484 and 3484b can perform other functions such as routing and forwarding packets, implementing user plane policies, supporting multi-home PDU sessions, processing user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0315] CN3415 can facilitate communication with other networks. For example, CN3415 can include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN3415 and the PSTN 408. In addition, CN3415 can provide access to another network 3412 to the WTRUs 3402a, 3402b, 3402c, where the other network 3412 can include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 3402a, 3402b, 3402c can be connected through the UPFs 3484a, 3484b to a local data network (DN) 3485a, 3485b via an N3 interface to the UPFs 3484a, 3484b, and an N6 interface between the UPFs 3484a, 3484b and the DNs 3485a, 3485b.

[0316] With reference to FIGS. 34A - 34D, and the corresponding descriptions of FIGS. 34A - 34D, one or more of the functions described herein with respect to one or more of the WTRUs 3402a - d, base stations 3414a - b, eNodeBs 3460a - c, MME 3462, SGW 3464, PGW 3466, gNBs 3480a - c, AMFs 3482a - b, UPFs 3484a - b, SMFs 3483a - b, DNs 3485a - b, and / or any other device described herein may be performed by one or more emulation devices (not shown). An emulation device can be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device can be used to test other devices and / or simulate network and / or WTRU functionality.

[0317] An emulation device can be designed to perform one or more tests of other devices in a laboratory environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more emulation devices may perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. An emulation device may be directly coupled to another device for testing and / or may perform tests using wireless wireless communication.

[0318] One or more emulation devices may perform one or more or all functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be utilized in a test scenario in a laboratory and / or a non-deployed (e.g., test) wired and / or wireless communication network to perform tests on one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via an RF circuit (which may include one or more antennas, for example) may be used by an emulation device to transmit and / or receive data.

[0319] The HEVC standard saves approximately 50% of the bitrate with equivalent perceptual quality compared to the conventional video coding standard H.264 / MPEG AVC. The HEVC standard provides significant coding improvements over its predecessor, but further coding efficiency improvements can be achieved using additional coding tools. The Joint Video Exploration Team (JVET), for example, initiated a project to develop a new generation of video coding standard called Versatile Video Coding (VVC) to provide such coding efficiency improvements, and a reference software codebase called the VVC Test Model (VTM) was established to demonstrate the reference implementation of the VVC standard. Another reference software base called the benchmark set (BMS) was also created to facilitate the evaluation of new coding tools. The BMS codebase includes a list of additional coding tools that provide higher coding efficiency and moderate implementation complexity in addition to the VTM, and is used as a benchmark when evaluating similar coding techniques in the VVC standardization process. The JEM coding tools integrated into BMS-2.0 include, for example, 4×4 non-separable secondary transform (NSST), generalized bi-prediction (GBi), bi-directional optical flow (BIO), decoder-side motion vector refinement (DMVR), and current picture referencing (CPR), in addition to the quantization tools of trellis coding.

[0320] The systems and methods for processing data according to representative embodiments may be implemented by one or more processors that execute a sequence of instructions included in a memory device. Such instructions may be read from another computer-readable medium, such as a secondary data storage device, into the memory device. Execution of the sequence of instructions included in the memory device causes, for example, the processor to operate as described above. In an alternative embodiment, a hardwired circuit may be used instead of or in combination with software instructions to implement one or more embodiments. Such software may operate on a processor housed within a robotic assist / device (RAA) and / or remotely within another mobile device. In the latter case, data may be transferred, wired or wirelessly, between the RAA or other mobile device that includes sensors and a remote device that includes a processor that executes software that performs scale estimation and compensation as described above. According to other representative embodiments, some of the processing described above with respect to localization may be performed by a device that includes sensors / cameras, while the remainder of the processing may be performed by a second device after receiving the partially processed data from the device that includes sensors / cameras.

[0321] Although features and elements are described above in certain combinations, one of ordinary skill in the art will appreciate that each feature or element can be used alone or in any combination with other features and elements. Also, the methods described herein may be implemented by a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, magnetic media such as read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). The processor associated with the software can be used to implement a radio frequency transceiver for use in a WTRU3402, UE, terminal, base station, RNC, or any host computer.

[0322] Furthermore, in the embodiments described above, reference is made to other devices including a processing platform, computing system, controller, and processor. These devices may include at least one central processing unit ("CPU") and memory. In accordance with the practice of those of ordinary skill in the computer programming art, references to acts and operations or symbolic representations of instructions may be performed by various CPUs and memories. Such acts and operations or instructions may be referred to as being "executed," "computer-executed," or "CPU-executed."

[0323] It will be understood by those skilled in the art that operations or instructions represented by actions and symbols include the manipulation of electrical signals by a CPU. An electrical system represents data bits that can result in the conversion or reduction of electrical signals, and the maintenance of data bits at memory locations within a memory system, thereby restructuring or otherwise changing the operations of the CPU and the processing of other signals. The memory location where a data bit is maintained is a physical location having specific electrical, magnetic, optical, or organic characteristics corresponding to or representing the data bit. It should be understood that the representative embodiments are not limited to the platforms or CPUs described above, and that the provided methods can be supported by other platforms and CPUs.

[0324] Also, data bits are maintained on a computer-readable medium, including magnetic disks, optical disks, and any other volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)) mass storage system readable by a CPU. The computer-readable medium may include cooperating or interconnected computer-readable media, which may exist solely on a processing system or be distributed among multiple interconnected processing systems that may be local or remote to the processing system. It will be understood that the representative embodiments are not limited to the memories described above, and that the described methods can be supported by other platforms and memories. It will be understood that the representative embodiments are not limited to the platforms and CPUs described above, and that the described methods can be supported by other platforms and CPUs.

[0325] In an exemplary embodiment, any of the operations, processes, etc. described herein may be implemented as computer-readable instructions stored on a computer-readable medium. The computer-readable instructions may be executed by a processor of a mobile unit, a network element, and / or any other computing device.

[0326] Between the hardware implementation and the software implementation of aspects of the system, little difference remains. The use of hardware or software is generally (but not always, as the choice between hardware and software can be important in certain situations) a design choice representing a cost - effectiveness trade - off. There can be various means (e.g., hardware, software, and / or firmware) by which the processes and / or systems and / or other technologies described herein can be affected, but the preferred means can vary depending on the situation in which the processes and / or systems and / or other technologies are deployed. For example, if the implementer determines that speed and accuracy are most important, the implementer can primarily select hardware and / or firmware means. If flexibility is most important, the implementer can primarily select a software implementation. Alternatively, the implementer can select some combination of hardware, software, and / or firmware.

[0327] In the above detailed description, various embodiments of devices and / or processes are shown by using block diagrams, flowcharts, and / or examples. As long as such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, it will be understood by those skilled in the art that each function and / or operation in such block diagrams, flowcharts, or examples can be individually and / or collectively implemented by various hardware, software, firmware, or substantially any combination thereof. Suitable processors include, by way of example, general - purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with DSP cores, controllers, microcontrollers, application - specific integrated circuits (ASICs), application - specific standard products (ASSPs), field - programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.

[0328] Although features and elements are provided above in a particular combination, it will be understood by those skilled in the art that each feature or element can be used alone or in any combination with other features and elements. The present disclosure is not limited to the perspective of the particular embodiments described in this application, which are intended as illustrations of various aspects. As will be apparent to those skilled in the art, many changes and modifications can be made without departing from the spirit and scope thereof. No element, act, or instruction used in the description of this application should be construed as critical or essential to the embodiments unless explicitly shown as such. In addition to those listed herein, functionally equivalent methods and apparatuses within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such changes and modifications are intended to be included within the scope of the appended claims. The present disclosure should be limited only by the terms of the appended claims and the full scope of equivalents to which such claims are entitled. It should be understood that the present disclosure is not limited to a particular method or system.

[0329] Also, the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used when referred to herein, the term "station" and its abbreviation "STA", the term "user equipment" and its abbreviation "UE" can mean (i) a wireless transmit and / or receive unit (WTRU) as described below, (ii) any of a plurality of embodiments of a WTRU as described below, (iii) a wireless compatible and / or wired compatible (e.g., tetherable) device configured to specifically have some or all of the structure and functionality of a WTRU as described below, (iii) a wireless compatible and / or wired compatible device configured to have structure and functionality that is less than all of a WTRU as described below, or (iv) something similar. Details of an exemplary WTRU that can represent any UE described herein are provided below with reference to FIGS. 34A-34D.

[0330] In certain representative embodiments, some portions of the subject matter described herein may be implemented using application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated forms. However, some aspects of the aspects disclosed herein may be implemented equivalently in integrated circuits as one or more computer programs executed in whole or in part on one or more computers (e.g., as one or more programs executed on one or more computer systems), as one or more programs executed on one or more processors (e.g., as one or more programs executed on one or more microprocessors), as firmware, or as substantially any combination thereof, and it will be understood by those skilled in the art that the design of circuits and / or the writing of code for software and / or firmware are well within the skill of those skilled in the art in light of this disclosure. Also, it will be understood by those skilled in the art that the mechanisms of the subject matter described herein may be distributed as various forms of program products, and that the exemplary embodiments of the subject matter described herein apply regardless of the particular type of signal carrying medium used to actually effect the distribution. Examples of signal carrying media include, but are not limited to, recordable media such as floppy disks, hard disk drives, CDs, DVDs, digital tapes, computer memories, and transmission media such as digital and / or analog communication media (e.g., optical fiber cables, waveguides, wired communication links, wireless communication links, etc.).

[0331] The subject matter described herein may exemplify different components that are included within or connected to other different components. It will be understood that such illustrated configurations are merely examples, and in fact, many other architectures that achieve the same functionality may be implemented. In a conceptual sense, components in any arrangement for achieving the same functionality are substantially "associated" so that the desired functionality can be achieved. Thus, any two components combined to achieve a particular functionality herein may be considered "associated" with each other so that the desired functionality is realized, regardless of the architecture or intervening components. Similarly, any two components so associated may be considered "operably connected" or "operably coupled" to each other to achieve the desired functionality, and any two components that can be so associated may be considered "operably couplable" to each other to achieve the desired functionality. Specific examples of being operably couplable include, but are not limited to, components that are physically engagable and / or physically interacting, and / or components that are wirelessly interactable and / or wirelessly interacting, and / or components that are logically interacting and / or logically interactable.

[0332] Regarding the use of substantially any plural and / or singular terms herein, those skilled in the art can appropriately convert from plural to singular and / or from singular to plural depending on the context and / or application. For clarity, various singular / plural replacements may be explicitly described herein.

[0333] Generally, the terms used in this specification and especially in the appended claims (e.g., the body of the appended claims) are generally intended to be open terms (e.g., the term "comprising" should be construed as "including but not limited to", the term "having" should be construed as "having at least", the term "including" should be construed as "including but not limited to", etc.), which will be understood by those skilled in the art. Further, when a specific number is intended in the introduced claim recitation, such intent is explicitly recited in the claim, and in the absence of such recitation, it will be understood by those skilled in the art that such intent does not exist. For example, if only one element is intended, the term "single" or a similar word may be used. For the sake of understanding, the following appended claims and / or the description in this specification may include the use of introductory phrases such as "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite article "a" or "an" limits any particular claim containing such introduced claim recitation to an embodiment containing only one such recitation, and this should not be so construed even when the same claim contains both the introductory phrase "one or more" or "at least one" and the indefinite article "a" or "an" (e.g., "a" and / or "an" should be construed to mean "at least one" or "one or more"). The same applies to the use of the definite article used to introduce claim recitations. Also, even when a specific number of the introduced claim recitation is explicitly recited, it will be recognized by those skilled in the art that such recitation should be construed to mean at least the recited number (e.g., a mere recitation of "two recitations" without other modifiers means at least two recitations or more than two recitations).Furthermore, in instances where a convention similar to “at least one of A, B, and C” is used, such syntax is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., “a system having at least one of A, B, and C” includes, but is not limited to, a system having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together). In instances where a convention similar to “at least one of A, B, or C” is used, such syntax is generally intended in the sense that one of ordinary skill in the art would understand the convention (e.g., “a system having at least one of A, B, or C” includes, but is not limited to, a system having only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, and C together). Further, it should be understood by one of ordinary skill in the art that substantially any disjunctive word and / or phrase presenting two or more alternative terms is contemplated to include one of those terms, any of those terms, or both terms in any of the specification, claims, or drawings. For example, the phrase “A or B” would be understood to include the possibilities of “A” or “B,” or “A and B.” Further, as used herein, the term “any of” following a listing of a plurality of elements and / or a plurality of categories of elements is intended to include “any,” “any combination,” “any plurality,” and / or “any combination of a plurality” of the elements and / or categories of elements, separately or in combination with other elements and / or other categories of elements. Further, as used herein, the term “set” or “group” is intended to include any number of elements, including zero. Further, as used herein, the term “number” is intended to include any number, including zero.

[0334] Furthermore, it will be understood by one of ordinary skill in the art that where a feature or aspect of the present disclosure is described in terms of a Markush group, the present disclosure is thereby also described in terms of any individual element of the Markush group or a subgroup of elements.

[0335] As will be understood by those skilled in the art, for all purposes, for example, from the perspective of providing a written description, all ranges disclosed in this specification include all possible sub-ranges and combinations of sub-ranges thereof. Any recited range can be readily recognized as being capable of fully describing the same range decomposed into at least equal halves, thirds, fourths, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily decomposed into, for example, a lower third, a middle third, and an upper third. Also, as will be understood by those skilled in the art, all words such as "up to", "at least", "greater than", "less than", etc. include the recited numbers and refer to ranges that can then be decomposed into sub-ranges as described above. Finally, as will be understood by those skilled in the art, a range includes each individual element. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, and so on.

[0336] Furthermore, the claims should not be construed as being limited to the order or elements provided unless otherwise stated therein. Also, the use of the term "means" in any claim is intended to assert the means-plus-function claim format of 35 U.S.C. § 112, paragraph 6, and no claim without the term "means" is so intended.

[0337] A processor related to software can be used to implement a radio frequency transceiver for use in a wireless transmit / receive unit (WTRU), a user equipment (UE), a terminal, a base station, a mobility management entity (MME) or evolved packet core (EPC), or any host computer. The WTRU can be used in association with modules implemented in software including hardware and / or software radios (SDR), as well as other components such as cameras, video camera modules, videophones, speakerphones, vibrating devices, speakers, microphones, television transceivers, hands-free headsets, keyboards, Bluetooth® modules, frequency modulation (FM) radio units, near field communication (NFC) modules, liquid crystal display (LCD) display units, organic light emitting diode (OLED) display units, digital music players, media players, video game player modules, Internet browsers, and / or wireless local area network (WLAN) or ultra wideband (UWB) modules.

[0338] Throughout this disclosure, one of ordinary skill in the art will understand that certain representative embodiments may be used selectively or in combination with other representative embodiments.

[0339] Also, the methods described herein may be implemented in a computer program, software, or firmware incorporated in a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, magnetic media such as read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor related to software can be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer.

Description of Symbols

[0340] 3400 Communication System 3402a WTRU 3402b WTRU 3402c WTRU 3402d WTRU 3408 PSTN 3410 Internet 3412 Network 3414a Base Station 3414b Base Station 3416 Air Interface 3418 Processor 3420 Transceiver 3422 Receiver Element 3424 Microphone 3426 Keypad 3428 Touch Pad 3430 Non-removable Memory 3432 Removable Memory 3434 Power Supply 3436 Chipset 3438 Peripheral Devices 3460a eNodeB 3460b eNodeB 3460c eNodeB 3462 MME 3464 SGW 3466 PGW 3482a AMF 3482b AMF 3483a SMF 3483b SMF 3484a UPF 3484b UPF 3485a DN 3485b DN

Claims

1. 1. A method of decoding a video, comprising: generating a sub-block based motion prediction signal for a sub-block of the block of the picture based on an affine motion model associated with the block; determining a set of pixel-level motion vector differential values ​​for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block-based motion prediction signal for each sample position of the sub-block, where an extended sub-block is formed to include the sub-block-based motion prediction signal and a plurality of samples surrounding the sub-block, each of the plurality of samples surrounding the sub-block being obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; decoding the video using the refined motion prediction signal; and 23. A method comprising:

2. 2. The method of claim 1, wherein the integer motion compensation is based on integer portions of motion vectors of the subblocks.

3. 2. The method of claim 1, wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the subblock.

4. 2. The method of claim 1, wherein the sub-block is expanded by one sample in each direction forming the expanded sub-block.

5. 13. A computer readable medium having instructions stored thereon for decoding video data in accordance with the method of claim 1.

6. 1. A method of encoding video, comprising the steps of: generating a sub-block based motion prediction signal for a sub-block of the block of the picture based on an affine motion model associated with the block; determining a set of pixel-level motion vector differential values ​​for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block-based motion prediction signal for each sample position of the sub-block, where an extended sub-block is formed to include the sub-block-based motion prediction signal and a plurality of samples surrounding the sub-block, each of the plurality of samples surrounding the sub-block being obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; encoding the video using the refined motion prediction signal; and 23. A method comprising:

7. 7. The method of claim 6, wherein the integer motion compensation is based on integer portions of motion vectors of the subblocks.

8. 7. The method of claim 6, wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the subblock.

9. 7. The method of claim 6, wherein the sub-block is expanded by one sample in each direction forming the expanded sub-block.

10. 7. A computer readable medium having instructions stored thereon for encoding video data according to the method of claim 6.

11. 1. An apparatus for decoding video, comprising: generating a sub-block based motion prediction signal for a sub-block of the block of the picture based on an affine motion model associated with the block; determining a set of pixel-level motion vector differential values ​​for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block-based motion prediction signal for each sample position of the sub-block, an extended sub-block being formed to include the sub-block-based motion prediction signal and a plurality of samples surrounding the sub-block, each of the plurality of samples surrounding the sub-block being obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; Decoding the video using the refined motion prediction signal. An apparatus comprising a processor configured to:

12. The apparatus of claim 11 , wherein the integer motion compensation is based on integer portions of motion vectors of the subblocks.

13. The apparatus of claim 11 , wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the subblock.

14. 12. The apparatus of claim 11, wherein the sub-block is expanded by one sample in each direction forming the expanded sub-block.

15. 1. An apparatus for encoding video, comprising: generating a sub-block based motion prediction signal for a sub-block of the block of the picture based on an affine motion model associated with the block; determining a set of pixel-level motion vector differential values ​​for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block-based motion prediction signal for each sample position of the sub-block, an extended sub-block being formed to include the sub-block-based motion prediction signal and a plurality of samples surrounding the sub-block, each of the plurality of samples surrounding the sub-block being obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; Encoding the video using the refined motion prediction signal. An apparatus comprising a processor configured to:

16. The apparatus of claim 15 , wherein the integer motion compensation is based on integer portions of motion vectors of the subblocks.

17. The apparatus of claim 15 , wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the subblock.

18. 16. The apparatus of claim 15, wherein the sub-block is expanded by one sample in each direction forming the expanded sub-block.

Citation Information

Patent Citations

  • Motion-compensation prediction based on BI-directional optical flow

    WO2019010156A1