Systems, apparatuses, and methods for inter-frame prediction refinement using optical flow

By calculating the spatial gradient or motion vector difference value using the optical flow model in video encoding, and generating a refined signal, the problem of insufficient refinement of inter-prediction signals in the prior art is solved, and the efficiency and accuracy of video encoding are improved.

CN113383551BActive Publication Date: 2025-06-17INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080012137.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-15
Filing Date
2020-02-04
Publication Date
2025-06-17
Estimated Expiration
2040-02-04

AI Technical Summary

Technical Problem

The existing video encoding technology is difficult to effectively utilize optical flow information in inter-frame prediction, resulting in insufficient refinement of motion prediction signals and affecting video encoding efficiency.

Method used

By obtaining the motion prediction signal based on the sub-block of the video block, and calculating the spatial gradient or motion vector difference using the optical flow model, a refinement signal is generated, and the inter prediction signal is then refined.

Benefits of technology

Improves the accuracy and efficiency of inter-frame prediction and enhances the performance of video encoding, especially when dealing with rapidly changing illumination conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113383551B_ABST
    Figure CN113383551B_ABST
Patent Text Reader

Abstract

Methods, apparatuses, and systems are disclosed. In one embodiment, a decoding method includes: obtaining a sub-block-based motion prediction signal for a current block of a video; obtaining one or more motion vector differences or one or more spatial gradients of the sub-block-based motion prediction signal; obtaining a refinement signal for the current block based on the one or more obtained spatial gradients or the one or more obtained motion vector differences; obtaining a refined motion prediction signal for the current block based on the sub-block-based motion prediction signal and the refinement signal; and decoding the current block based on the refined motion prediction signal.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 802,428, filed on February 7, 2019, U.S. Provisional Patent Application No. 62 / 814,611, filed on March 6, 2019, and U.S. Provisional Patent Application No. 62 / 883,999, filed on April 15, 2019, the respective contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to video coding, and in particular, to systems, apparatuses, and methods for using inter-frame prediction refinement using optical flow.

[0004] Related Art

[0005] Video coding systems are widely used to compress digital video signals to reduce the storage and / or transmission bandwidth of such signals. Among various types of video coding systems, such as block-based systems, wavelet-based systems, and object-based systems, currently block-based hybrid video coding systems are the most widely used and deployed. Examples of block-based video coding systems include various international video coding standards, such as MPEG1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, and the latest video coding standard known as High Efficiency Video Coding (HEVC), which was developed by the Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG. Summary of the Invention

[0006] In one representative embodiment, a decoding method includes: obtaining a sub-block based motion prediction signal of a current block of a video; obtaining one or more spatial gradients or one or more motion vector differences of the sub-block based motion prediction signal; obtaining a refinement signal of the current block based on the one or more obtained spatial gradients or the one or more obtained motion vector differences; obtaining a refined motion prediction signal of the current block based on the sub-block based motion prediction signal and the refinement signal; and decoding the current block based on the refined motion prediction signal. Various other embodiments are also disclosed herein. Brief Description of the Drawings

[0007] A more detailed understanding can be obtained from the following detailed description given by way of example in conjunction with the accompanying drawings. The drawings in the specification are examples. Therefore, the drawings and the detailed description should not be considered restrictive, and other equivalent examples are feasible and possible. In addition, the same reference numerals in the figures indicate the same elements, and wherein:

[0008] FIG. 1 is a block diagram showing a representative block-based video coding system;

[0009] FIG. 2 is a block diagram showing a representative block-based video decoder;

[0010] FIG. 3 is a block diagram showing a representative block-based video encoder with generalized bi-prediction (GBi) support;

[0011] FIG. 4 is a schematic diagram showing a representative GBi module for an encoder;

[0012] FIG. 5 is a schematic diagram showing a representative block-based video decoder with GBi support;

[0013] FIG. 6 is a schematic diagram showing a representative GBi module for a decoder;

[0014] FIG. 7 is a schematic diagram showing a representative bidirectional optical flow;

[0015] FIGS. 8A and 8B are schematic diagrams showing a representative four-parameter affine mode;

[0016] FIG. 9 is a schematic diagram showing a representative six-parameter affine mode;

[0017] FIG. 10 is a schematic diagram showing a representative interleaved prediction process;

[0018] FIG. 11 is a schematic diagram showing representative weight values (e.g., associated with pixels) in a sub-block;

[0019] FIG. 12 is a schematic diagram showing some regions where interleaved prediction is applied and other regions where the interleaved prediction is not applied;

[0020] FIGS. 13A and 13B are schematic diagrams showing the SbTMVP process;

[0021] FIG. 14 is a schematic diagram showing neighboring motion blocks (e.g., 4×4 motion blocks) that can be used for motion parameter derivation;

[0022] FIG. 15 is a schematic diagram showing neighboring motion blocks that can be used for motion parameter derivation;

[0023] FIG. 16 is a schematic diagram showing the difference Δv(i,j) between the sub-block MV and the pixel-level MV after sub-block-based affine motion compensation prediction;

[0024] FIG. 17A is a schematic diagram showing a representative process for determining the MV corresponding to the actual center of a sub-block;

[0025] FIG. 17B is a schematic diagram showing the positions of chrominance samples in a 4:2:0 chrominance format;

[0026] Figure 17C is a schematic diagram showing an extended prediction sub-block;

[0027] Figure 18A is a flowchart showing a first representative encoding / decoding method;

[0028] Figure 18B is a flowchart showing a second representative encoding / decoding method;

[0029] Figure 19 is a flowchart showing a third representative encoding / decoding method;

[0030] Figure 20 is a flowchart showing a fourth representative encoding / decoding method;

[0031] Figure 21 is a flowchart showing a fifth representative encoding / decoding method;

[0032] Figure 22 is a flowchart showing a sixth representative encoding / decoding method;

[0033] Figure 23 is a flowchart showing a seventh representative encoding / decoding method;

[0034] Figure 24 is a flowchart showing an eighth representative encoding / decoding method;

[0035] Figure 25 is a flowchart showing a representative gradient calculation method;

[0036] Figure 26 is a flowchart showing a ninth representative encoding / decoding method;

[0037] Figure 27 is a flowchart showing a tenth representative encoding / decoding method;

[0038] Figure 28 is a flowchart showing an eleventh representative encoding / decoding method;

[0039] Figure 29 is a flowchart showing a representative encoding method;

[0040] Figure 30 is a flowchart showing another representative encoding method;

[0041] Figure 31 is a flowchart showing a twelfth representative encoding / decoding method;

[0042] Figure 32 is a flowchart showing a thirteenth representative encoding / decoding method;

[0043] Figure 33 is a flowchart showing a fourteenth representative encoding / decoding method;

[0044] Figure 34A is a system diagram showing an exemplary communication system in which one or more disclosed embodiments can be implemented;

[0045] Figure 34B is shown according to an embodiment and can be in Figure 34A a system diagram showing an exemplary wireless transmit / receive unit (WTRU) that can be used inside the shown communication system;

[0046] Figure 34C is shown according to an embodiment and can be in Figure 34A a system diagram showing an exemplary radio access network (RAN) and an exemplary core network (CN) that can be used inside the shown communication system; and

[0047] Figure 34D is shown according to an embodiment and can be in Figure 34A a system diagram showing another exemplary RAN and another exemplary CN that can be used inside the shown communication system. Detailed Description

[0048] Block - based Hybrid Video Coding Process

[0049] Similar to HEVC, VVC is built on a block - based hybrid video coding framework.

[0050] FIG. 1 is a block diagram showing a general block - based hybrid video coding system.

[0051] Referring to FIG. 1, encoder 100 can be provided with an input video signal 102 that is processed block - by - block (referred to as coding units (CUs)), and can be used to efficiently compress high - resolution (1080p and above) video signals. In HEVC, a CU can be up to 64×64 pixels. A CU can be further divided into prediction units or PUs, to which separate prediction processes can be applied. For each input video block (MB and / or CU), spatial prediction 160 and / or temporal prediction 162 can be performed. Spatial prediction (or “intra - frame prediction”) can use pixels from already - encoded adjacent blocks in the same video picture / slice to predict the current video block.

[0052] Spatial prediction can reduce the spatial redundancy inherent in a video signal. Temporal prediction (also referred to as “inter-frame prediction” or “motion-compensated prediction”) uses pixels from previously encoded video pictures to predict the current video block. Temporal prediction can reduce the temporal redundancy inherent in a video signal. The temporal prediction signal for a given video block can (e.g., typically can) be signaled by one or more motion vectors (MVs), which can indicate the amount and / or direction of motion between the current block (CU) and its reference block.

[0053] If multiple reference pictures are supported (as is the case for recent video coding standards such as H.264 / AVC or HEVC), then for each video block, its reference picture index can be sent (e.g., additionally sent); and / or this reference index can be used to identify which reference picture in the reference picture buffer 164 the temporal prediction signal is from. After spatial and / or temporal prediction, the mode decision block 180 in the encoder 100 can select the best prediction mode, e.g., based on a rate-distortion optimization method / process. The prediction block from the spatial prediction 160 or the temporal prediction 162 can be subtracted from the current video block 116; and / or the prediction residuals can be decorrelated using the transform 104 and quantized 106 to achieve a target bit rate. The quantized residual coefficients can be inverse quantized 110 and inverse transformed 112 to form a reconstructed residual, which can be added back to the prediction block at 126 to form a reconstructed video block. Additionally, in-loop filtering 166, such as a deblocking filter and an adaptive loop filter, can be applied to the reconstructed video block before it is placed in the reference picture buffer 164 and can be used to encode future video blocks. To form the output video bitstream 120, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and the quantized residual coefficients can be sent (e.g., all sent) to the entropy coding unit 108 to be further compressed and / or packetized to form the bitstream.

[0054] The encoder 100 can be implemented using a processor, a memory, and a transmitter that provide the various elements / modules / units described above. For example, those skilled in the art will understand that: (1) the transmitter can send the bitstream 120 to the decoder; and (2) the processor can be configured to execute software to enable receiving the input video 102 and performing functions associated with the various blocks of the encoder 100.

[0055] FIG. 2 is a block diagram illustrating a block-based video decoder.

[0056] Referring to FIG. 2, the video decoder 200 may be provided with a video bitstream 202, which may be unpacked and entropy decoded at the entropy decoding unit 208. The coding mode and prediction information may be sent to the appropriate one of the spatial prediction unit 260 (for intra-coding mode) and / or the temporal prediction unit 262 (for inter-coding mode) to form a prediction block. The residual transform coefficients may be sent to the inverse quantization unit 210 and the inverse transform unit 212 to reconstruct the residual block. The reconstructed block may be further subjected to in-loop filtering 266 before being stored in the reference picture repository 264. In addition to being saved in the reference picture repository 264 for predicting future video blocks, the reconstructed video 220 may be sent out, for example, to drive a display device.

[0057] The decoder 200 may be implemented using a processor, a memory, and a receiver, which may provide the various elements / modules / units disclosed above. For example, those skilled in the art understand that: (1) the receiver may be configured to receive the bitstream 202; and (2) the processor may be configured to execute software to enable receiving the bitstream 202 and outputting the reconstructed video 220 and to perform functions associated with the respective blocks of the decoder 200.

[0058] Those skilled in the art understand that many functions / operations / procedures of block-based encoders and block-based decoders are the same.

[0059] In modern video codecs, bidirectional motion compensation prediction (MCP) may be used to efficiently remove temporal redundancy by exploiting the temporal correlation between pictures. The bi-prediction signal may be formed by combining two uni-prediction signals using a weight value equal to 0.5, which may not be optimal for combining uni-prediction signals, especially under some conditions where the luminance changes rapidly from one reference picture to another. Certain prediction techniques / operations and / or procedures may be implemented to compensate for the luminance change over time by applying some global / local weights and / or offset values to the sample values in the reference pictures (e.g., some or each of the sample values in the reference pictures).

[0060] Using bidirectional motion compensation prediction (MCP) in a video codec enables removing temporal redundancy by exploiting the temporal correlation between pictures. The bi-prediction signal may be formed by combining two uni-prediction signals using a weight value (e.g., 0.5). In some videos, the luminance characteristics may change rapidly from one reference picture to another. Therefore, prediction techniques may compensate for the change in luminance over time (e.g., fade transition) by applying global or local weights and / or offset values to one or more sample values in the reference pictures.

[0061] Generalized Bi-Prediction (GBi) can improve the MCP for the bi-prediction mode. In the bi-prediction mode, the predicted signal at a given sample x can be calculated by Equation 1 as follows:

[0062] P[x] = w0 * P0[x + v0] + w1 * P1[x + v1] (1)

[0063] In the above equation, P[x] can represent the resulting predicted signal for sample x at picture position x. Pi[x + vi] can be the motion-compensated predicted signal for x using the motion vector (MV) vi of the i-th list (e.g., list 0, list 1, etc.). w0 and w1 can be two weight values shared among samples in a block (e.g., among all samples). Based on this equation, various predicted signals can be obtained by adjusting the weight values w0 and w1. Some configurations of w0 and w1 can mean the same prediction as single prediction and bi-prediction. For example, (w0, w1) = (0, 1) can be used for single prediction using reference list L0. (w0, w1) = (0, 1) can be used for single prediction using reference list L1. (w0, w1) = (0.5, 0.5) can be used for bi-prediction using two reference lists. The weights can be signaled for each CU. To reduce the signaling overhead, a constraint such as w0 + w1 = 1 can be applied so that only one weight needs to be signaled. Thus, Equation 1 can be further simplified as set forth in Equation 2 below:

[0064] P[x] = (1 - w1) * P0[x + v0] + w1 * P1[x + v1] (2)

[0065] To further reduce the weight signaling overhead, w1 can be discretized (e.g., -2 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 10 / 8, etc.). Then, each weight value can be indicated by an index value within a (e.g., small) finite range.

[0066] Figure 3 is a block diagram showing a representative block-based video encoder with GBi support.

[0067] The encoder 300 may include a mode decision module 304, a spatial prediction module 306, a motion prediction module 308, a transform module 310, a quantization module 312, an inverse quantization module 316, an inverse transform module 318, a loop filter 320, a reference picture repository 322, and an entropy encoding module 314. Some or all of the modules or components of the encoder (e.g., the spatial prediction module 306) may be the same as or similar to those described in connection with FIG. 1. Additionally, the spatial prediction module 306 and the motion prediction module 308 may be pixel domain prediction modules. Thus, the input video bitstream 302 may be processed in a manner similar to the input video bitstream 102, although the motion prediction module 308 may also include GBi support. Thus, the motion prediction module 308 may combine two separate prediction signals in a weighted average manner. Additionally, a selected weight index may be signaled in the output video bitstream 324.

[0068] The encoder 300 may be implemented using a processor, a memory, and a transmitter that provide the various elements / modules / units described above. For example, those skilled in the art will understand that: (1) the transmitter may send the bitstream 324 to the decoder; and (2) the processor may be configured to execute software to enable receiving the input video 302 and performing functions associated with the respective blocks of the encoder 300.

[0069] FIG. 4 is a schematic diagram illustrating a representative GBi estimation module 400 that may be employed in a motion prediction module of an encoder such as the motion prediction module 308. The GBi estimation module 400 may include a weight value estimation module 402 and a motion estimation module 404. Thus, the GBi estimation module 400 may utilize processing (e.g., a two-step operation / processing) to generate an inter-frame prediction signal, such as a final inter-frame prediction signal. The motion estimation module 404 may perform motion estimation using an input video block 401 and one or more reference pictures received from a reference picture repository 406 and by searching for two best motion vectors (MVs) that point to (e.g., two) reference blocks. The weight value estimation module 402 may receive: (1) the output of the motion estimation module 404 (e.g., motion vectors v0 and v1), one or more reference pictures from the reference picture repository 406, and weight information W, and may search for an optimal weight index to minimize the weighted bi-prediction error between the current video block and the bi-prediction. It is contemplated that the weight information W may describe a list or set of available weight values such that the determined weight index and the weight information W may be used together to specify the weights w0 and w1 used in the GBi. The prediction signal of the generalized bi-prediction may be calculated as a weighted average of two prediction blocks. The output of the GBi estimation module 400 may include an inter-frame prediction signal, motion vectors v0 and v1, and / or a weight index weight_idx, etc.

[0070] FIG. 5 is a schematic diagram showing a representative block-based video decoder with GBi support, which can decode a bitstream 502 supporting GBi (e.g., a bitstream from an encoder), such as the bitstream 324 generated by the encoder 300 described in connection with FIG. 3. As shown in FIG. 5, the video decoder 500 may include an entropy decoder 504, a spatial prediction module 506, a motion prediction module 508, a reference picture repository 510, an inverse quantization module 512, an inverse transform module 514, and / or a loop filter module 518. Some or all of the modules of the decoder may be the same or similar to those described in connection with FIG. 2, although the motion prediction module 508 may also include GBi support. Thus, the coding mode and prediction information are used to derive a prediction signal by using spatial prediction or GBi-supported MCP. For GBi, the block motion information and weight values (e.g., in the form of an index indicating the weight value) may be received and decoded to generate the predicted block.

[0071] The decoder 500 may be implemented using a processor, a memory, and a receiver, which may provide the various elements / modules / units disclosed above. For example, those skilled in the art understand that: (1) the receiver may be configured to receive the bitstream 502; and (2) the processor may be configured to execute software to enable receiving the bitstream 502 and outputting the reconstructed video 520 and performing functions associated with the respective blocks of the decoder 500.

[0072] FIG. 6 is a schematic diagram showing a representative GBi prediction module that may be employed in the motion prediction module (e.g., motion prediction module 508) of a decoder.

[0073] Referring to FIG. 6, the GBi prediction module may include a weighted average module 602 and a motion compensation module 604, which may receive one or more reference pictures from a reference picture repository 606. The weighted average module 602 may receive the output of the motion compensation module 604, weight information W, and a weight index (e.g., weight_idx). The output of the motion compensation module 604 may contain motion information corresponding to blocks of the picture. The GBi prediction module 600 may use the block motion information and weight values to calculate a prediction signal for GBi (e.g., an inter-frame prediction signal 608) as a weighted average of (e.g., two) motion-compensated prediction blocks.

[0074] Representative dual prediction based on an optical flow model

[0075] FIG. 7 is a schematic diagram showing a representative bidirectional optical flow.

[0076] Referring to FIG. 7, dual prediction may be based on an optical flow model. For example, the prediction associated with a current block (e.g., current block 700) may be based on the optical flow associated with a first predicted block I (0) 702 (e.g., a temporally previous predicted block, e.g., shifted in time by τ0) and a second predicted block I (1) 704 (e.g., a temporally future block, e.g., shifted in time by τ1). Dual prediction in video coding may be a combination of two temporal predicted blocks 702 and 704 obtained from reconstructed reference pictures. Due to the limitations of block-based motion compensation (MC), there may be small remaining motions that can be observed between the samples of the two predicted blocks, thus reducing the efficiency of motion compensation prediction. Bidirectional optical flow (BIO, or BDOF) may be applied to reduce the impact of such motions for each sample within a block. BIO may provide per-sample motion refinement, which may be performed on top of block-based motion compensation prediction when dual prediction is used. For BIO, a refined motion vector for each sample in a block may be derived based on a classical optical flow model. For example, when I (k) (x,y) is the sample value at the coordinates (x,y) of a predicted block derived from reference picture list k (k = 0,1) and and are the horizontal and vertical gradients of the sample, given an optical flow model, the motion refinement (v x ,v y ) at (x,y) can be derived by Equation 3 below:

[0077]

[0078] In FIG. 7, the (MV x0 ,MV y0 ) associated with the first predicted block 702 and the (MV x1 ,MV y1 ) associated with the second predicted block 704 indicate the block-level motion vectors available for generating two predicted blocks I (0) and I (1) . The motion refinement (v x ,v y ) at the sample position (x,y) can be calculated by minimizing the difference Δ between the sample values (e.g., A and B in FIG. 7) after motion refinement compensation, as shown in Equation 4 below:

[0079]

[0080] For example, to ensure the regularity of the derived motion refinement, it can be envisioned that the motion refinement is consistent for samples within a small unit (e.g., a 4×4 block or other small unit). In the Benchmark Set (BMS)-2.0, the value (v x ,v y ) is derived by minimizing Δ within a 6×6 window Ω around each 4×4 block, as elaborated in Equation 5 below:

[0081]

[0082] To solve the optimization specified in Equation 5, the BIO can use a progressive method / operation / process, which can optimize the motion refinement in the horizontal and vertical directions (e.g., then the vertical direction). This may result in the following Equations / Inequalities 6 and 7:

[0083]

[0084]

[0085] where can be the floor function, which can output the largest value less than or equal to the input, and th BIO can be the motion refinement threshold, e.g., to prevent error propagation due to coding noise and / or irregular local motion, which is equal to 2 18-BD . The values of S1, S2, S3, S5, and S6 can be further calculated as elaborated in Equations 8 - 12 below:

[0086] S1 = ∑ (i,j)∈Ω ψ x (i,j)·ψ x (i,j), (8)

[0087] S3 = ∑ (i,j)∈Ω θ(i,j)·ψ x (i,j)·2 L (9)

[0088] S2 = ∑ (i,j)∈Ω ψ x (i,j)·ψ y (i,j) (10)

[0089] S5 = ∑ (i,j)∈Ω ψ y (i,j)·ψ y (i,j)·2 (11)

[0090] S6 = ∑ (i,j)∈Ω θ(i,j)·ψ y (i,j)·2 L+1 (12)

[0091] Among them, various gradients can be described in the following equations 13 - 15:

[0092]

[0093]

[0094] θ(i,j) = I (1) (i,j) - I (0) (i,j) (15)

[0095] For example, in BMS - 2.0, the BIO gradients in the horizontal and vertical directions in equations 13 - 15 can be directly obtained by calculating the difference between two adjacent samples at a sample position in each L0 / L1 prediction block (e.g., horizontal or vertical, depending on the direction of the gradient to be derived), as described in the following equations 16 and 17:

[0096]

[0097]

[0098] In equations 8 - 12, L can be the bit - depth increase of the internal BIO process / program to maintain data precision. For example, in BMS - 2.0, it can be set to 5. To avoid division by a smaller value, the adjustment parameters r and m in equations 6 and 7 can be defined as in the following equations 18 and 19:

[0099] r = 500·4 BD-8 (18)

[0100] m = 700·4 BD-8 (19)

[0101] Where BD can be the bit - depth of the input video. Based on the motion refinement derived from equations 4 and 5, the final dual - prediction signal of the current CU can be calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on the optical - flow equation 3, as specified in the following equations 20 and 21:

[0102]

[0103]

[0104] Where shift and o offset Can be the right - shift and offset that can be applied to combine the L0 and L1 prediction signals for dual - prediction. For example, they can be set to be equal to 15 - BD and 1 << (14 - BD)+2·(1 << 13) respectively. rnd(.) is a rounding function that can round the input value to the nearest integer value.

[0105] Representative affine mode

[0106] In HEVC, a translational motion (only translational motion) model is applied to motion compensation prediction. In the real world, there are many kinds of motions (e.g., zoom in / out, rotation, perspective motion, and other irregular motions). In the Versatile Video Coding (VVC) Test Model (VTM)-2.0, affine motion compensation prediction can be applied. This affine motion model is either 4-parameter or 6-parameter. A first flag is signaled for each inter-coded CU to indicate whether the translational motion model or the affine motion model is applied to inter-frame prediction. If the affine motion model is applied, then a second flag is sent to indicate whether the model is a 4-parameter model or a 6-parameter model.

[0107] The 4-parameter affine motion model has the following parameters: two parameters for translational motion in the horizontal and vertical directions, one parameter for scaling motion in two directions, and one parameter for rotational motion in two directions. The horizontal scaling parameter is equal to the vertical scaling parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. In the VTM, the four-parameter affine motion model is coded using two motion vectors at two control point positions defined at the upper left corner 810 and the upper right corner 820 of the current CU. Other control point positions are also possible, such as at other corners and / or edges of the current CU.

[0108] Although one affine motion model is described above, other affine models are equally possible and can be used in various embodiments herein.

[0109] Figures 8A and 8B are schematic diagrams showing a representative four-parameter affine model and sub-block level motion derivation for an affine block. Referring to Figures 8A and 8B, the affine motion field of the block is described by two control point motion vectors at a first control point 810 (at the upper left corner of the current block) and a second control point 820 (at the upper right corner of the current block), respectively. Based on the control point motion, the motion field (v x ,v y ) of an affine-coded block is described as set forth in Equations 22 and 23 below:

[0110]

[0111]

[0112] where (v 0x ,v 0y ) can be the motion vector of the upper left corner control point 810, (v 1x ,v 1y) can be the motion vector of the upper-right control point 820, as shown in FIG. 8A, and w can be the width of the CU. For example, in VTM-2.0, the motion field of an affine-coded CU can be derived at the 4×4 block level; that is, (v x , v y ) can be derived for each 4×4 block within the current CU and will be applied to the corresponding 4×4 block.

[0113] The four parameters of the 4-parameter affine model can be estimated iteratively. The MV pair at step k can be expressed as The original signal (e.g., the illuminance signal) can be expressed as I(i, j), and the predicted signal (e.g., the illuminance signal) can be expressed as I′ k (i, j). The spatial gradients g x (i, j) and g y (i, j) can be derived, for example, by applying Sobel filters to the predicted signal I′ k (i, j) in the horizontal and / or vertical directions, respectively. The derivatives of Equation 3 can be expressed as set forth in the following Equations 24 and 25:

[0114]

[0115] where (a, b) can be the delta translation parameters, and (c, d) can be the delta scaling and rotation parameters at step k. The delta MVs at the control points can be derived using their coordinates, as described in the following Equations 26-29. For example, (0,0), (w,0) can be the coordinates of the upper-left and upper-right control points 810 and 820, respectively.

[0116]

[0117]

[0118] Based on the optical flow equation, the relationship between the change in intensity (e.g., luminance) and the spatial gradient and the temporal shift is formulated as follows in Equation 30:

[0119]

[0120] By substituting Equations 24 and 25 for and Equation 31 for the parameters (a, b, c, d) is obtained as follows:

[0121] I′ k (i, j) - I(i, j) = (g x (i, j) * i + g y (i, j) * j) * c + (-g x(i,j)*j + g y ((i,j)*i)*d + g x (i,j)*a + g y (i,j)*b(31)

[0122] Since the samples in the CU (e.g., all samples) satisfy Equation 31, the parameter set (e.g., a, b, c, d) can be solved using, for example, the least - squares error method. At step (k + 1), the values at the two control points can be solved using Equations 26 - 29 and can be rounded to a specific precision (e.g., 1 / 4 pixel precision (pel) or other sub - pixel precision, etc.). By using iteration, the MVs at the two control points can be refined, for example, until convergence (e.g., when all of the parameters (a, b, c, d) are zero or the iteration time meets a predetermined limit).

[0123] FIG. 9 is a schematic diagram showing a representative six - parameter affine pattern, where, for example: V0, V1, and V2 are the motion vectors at control points 910, 920, and 930 respectively, and (MV x , MV y ) is the motion vector of the sub - block centered at position (x, y).

[0124] Referring to FIG. 9, the affine motion model (e.g., with 6 parameters) can have any of the following parameters: (1) a parameter for translational movement in the horizontal direction; (2) a parameter for translational movement in the vertical direction; (3) a parameter for scaling motion in the horizontal direction; (4) a parameter for rotational motion in the horizontal direction; (5) a parameter for scaling motion in the vertical direction; and / or (6) a parameter for rotational motion in the vertical direction. The 6 - parameter affine motion model can be encoded using three MVs at three control points 910, 920, and 930. As shown in FIG. 9, the three control points 910, 920, and 930 of the 6 - parameter affine - coded CU are defined at the upper - left corner, upper - right corner, and lower - left corner of the CU respectively. The motion at the upper - left control point 910 can be related to translational motion, and the motion at the upper - right control point 920 can be related to rotational motion and / or scaling motion in the horizontal direction, and the motion at the lower - left control point 930 can be related to rotational and / or scaling motion in the vertical direction. For the 6 - parameter affine motion model, the rotational motion and / or scaling motion in the horizontal direction can be different from the same motion in the vertical direction. The motion vector (v x , v y ) of each sub - block can be derived using the three MVs at control points 910, 920, and 930, as set forth in Equations 32 and 33 below:

[0125]

[0126]

[0127] where (v 2x , v 2y ) is the motion vector V2 of the lower left control point 930, (x, y) can be the center position of the sub-block, w can be the width of the CU, and h can be the height of the CU.

[0128] The six parameters of the 6-parameter affine model can be estimated in a similar manner. Equations 24 and 25 can be changed as described in Equations 34 and 35 as follows.

[0129]

[0130] where at step k, (a, b) can be the differential translation parameters, (c, d) can be the differential scaling and rotation parameters in the horizontal direction, and (e, f) can be the differential scaling and rotation parameters in the vertical direction. Equation 31 can be changed as described in Equation 36, as follows:

[0131] I′ k (i, j) - I(i, j) = (g x (i, j) * i) * c + (g x (i, j) * j) * d + (g y (i, j) * i) * e + (g y (i, j) * j) * f + g x (i, j) * a + g y (i, j) * b (36)

[0132] The parameter set (a, b, c, d, e, f) can be solved by considering the samples (e.g., all samples) within the CU, for example, using the least squares method / process / operation. The of the upper left control point can be calculated using Equations 26 - 29. The of the upper right control point can be calculated using Equations 37 and 38 described below. The of the lower left control point can be calculated using Equations 39 and 40 described below.

[0133]

[0134]

[0135] Although the 4-parameter affine model and the 6-parameter affine model are shown in FIGS. 8A, 8B, and 9, those skilled in the art can understand that affine models with different numbers of parameters and / or different control points are equally possible.

[0136] Although the affine model is described herein in connection with optical flow refinement, one of ordinary skill in the art will appreciate that other motion models in combination with optical flow refinement are also possible.

[0137] Representative Interleaved Prediction for Affine Motion Compensation

[0138] With affine motion compensation (AMC), e.g., in VTM, the coding block is divided into sub-blocks as small as 4×4, and each sub-block can be assigned an individual motion vector (MV) derived from an affine model, e.g., as shown in FIGS. 8A and 8B or FIG. 9. Using a 4-parameter or 6-parameter affine model, the MV can be derived from the MVs of two or three control points.

[0139] AMC may face a dilemma associated with the size of the sub-blocks. Using smaller sub-blocks, AMC can achieve better coding performance, but may suffer from a higher complexity burden.

[0140] FIG. 10 is a schematic diagram showing a representative interleaved prediction process that can achieve a finer-grained MV, e.g., in exchange for a moderate increase in complexity.

[0141] In FIG. 10, the coding block 1010 can be partitioned into sub-blocks with two different partitioning patterns (e.g., the first and second patterns 0 and 1). The first partitioning pattern 0 (e.g., the first sub-block pattern, e.g., the 4×4 sub-block pattern) can be the same as in VTM, and the second partitioning pattern 1 (e.g., the second sub-block pattern of overlapping and / or interleaving) can partition the coding block 1010 into 4×4 sub-blocks with a 2×2 offset from the first partitioning pattern 0, as shown in FIG. 10. AMC can utilize the two partitioning patterns (e.g., the first and second partitioning patterns 0 and 1) to generate several auxiliary predictions (e.g., two auxiliary predictions P0 and P1). The MV of each sub-block in each of the partitioning patterns 0 and 1 can be derived from the control point motion vectors (CPMV) through the affine model.

[0142] The final prediction P can be calculated as a weighted sum of the auxiliary predictions (e.g., two auxiliary predictions P0 and P1), which is formulated as shown in Equations 41 and 42 below:

[0143]

[0144] FIG. 11 is a schematic diagram showing representative weight values (e.g., associated with pixels) in a sub-block. Referring to FIG. 11, the auxiliary prediction sample located at the center (e.g., the center pixel) of the sub-block 1100 can be associated with a weight value of 3, and the auxiliary prediction sample located at the boundary of the sub-block 1100 can be associated with a weight value of 1.

[0145] FIG. 12 is a schematic diagram showing a region where interlaced prediction is applied and other regions where interlaced prediction is not applied. Referring to FIG. 12, region 1200 may include a first region 1210 (shown as non-cross-hatched in FIG. 12) having, for example, 4×4 sub-blocks where interlaced prediction is applied and a second region 1220 (shown as cross-hatched in FIG. 12) where, for example, interlaced prediction is not applied. To avoid small-block motion compensation, for example, for both the first and second partitioning modes, the interlaced prediction may be applied only to regions where the size of the sub-blocks meets a threshold size (e.g., which is 4×4).

[0146] In VTM-3.0, the size of the sub-blocks may be 4×4 for the chrominance component, and the interlaced prediction may be applied to the chrominance component and / or the luminance component. Since in AMC, the regions for performing motion compensation (MC) on sub-blocks (e.g., all sub-blocks) can be extracted as a whole together, the interlaced prediction does not increase the bandwidth. For flexibility, a flag may be signaled in the slice header to indicate whether the interlaced prediction is used or not. For the interlaced prediction, the flag may be signaled as a 1-bit flag (e.g., a first logic level that can always be signaled as 0 or 1).

[0147] Representative process for sub-block based temporal motion vector prediction (SbTMVP)

[0148] SbTMVP is supported by VTM. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP can use the motion field in the collocated picture, for example, to improve the merge mode and motion vector prediction of the CUs in the current picture. The same collocated picture used by TMVP can be used for SbTMVP. SbTMVP differs from TMVP as follows: (1) TMVP can predict the motion at the CU level, while SbTMVP can predict the motion at the sub-CU level; and / or (2) TMVP can extract the temporal motion vector from the collocated block in the collocated picture (e.g., the collocated block can be the lower-right or center block relative to the current CU), while SbTMVP can apply a motion shift (e.g., the motion shift can be obtained from the motion vector of one of the spatial neighboring blocks of the current CU) before extracting the temporal motion information from the collocated picture, etc.

[0149] FIGS. 13A and 13B are schematic diagrams showing the SbTMVP process. FIG. 13A shows the spatial neighboring blocks used by ATMVP, and FIG. 13B shows the derivation of the sub-CU motion field by applying a motion shift from the spatial neighbor and scaling the motion information from the corresponding collocated sub-CU.

[0150] Referring to FIGS. 13A and 13B, SbTMVP can predict the motion vector of a sub-CU within the current CU operation (e.g., in two operations). In the first operation, the spatially adjacent blocks A1, B1, B0, and A0 can be examined in the order of A1, B1, B0, and A0. Once the first spatially adjacent block having a motion vector using the co-located picture as its reference picture and / or after identifying the first spatially adjacent block having a motion vector using the co-located picture as its reference picture, that motion vector can be selected as the motion shift to be applied. If no such motion is identified from the spatially adjacent blocks, the motion shift can be set to (0, 0). In the second operation, the motion shift identified in the first operation can be applied (e.g., added to the coordinates of the current block) to obtain sub-CU level motion information (e.g., motion vector and reference index) from the co-located picture, as shown in FIG. 13B. The example in FIG. 13B shows the motion shift set for the motion of block A1. For each sub-CU, the motion information of its corresponding block (e.g., the smallest motion grid covering the central sample) in the co-located picture can be used to derive the motion information for the sub-CU. After identifying the motion information of the co-located sub-CU, the motion information can be converted to the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC. For example, temporal motion scaling can be applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.

[0151] The combined sub-block based merge list can be used in VTM-3 and can contain or include both SbTMVP merge candidates and affine merge candidates, e.g., for signaling the sub-block based merge mode. The SbTMVP mode can be enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor can be added as the first entry in the list of sub-block based merge candidates, followed by the affine merge candidates. The size of the sub-block based merge list can be signaled in the SPS, and the maximum allowed size of the sub-block based merge list can be an integer, e.g., 5 in VTM3.

[0152] The sub-CU size used in SbTMVP can be fixed, e.g., 8×8 or other sub-CU sizes, and for the affine merge mode, the SbTMVP mode can be applicable to (e.g., can be applicable only to) CUs whose width and height are both greater than or equal to 8.

[0153] The encoding logic for additional SbTMVP merge candidates can be the same as that for other merge candidates. For example, for each CU in a P or B slice, an additional rate distortion (RD) check can be performed to decide whether to use the SbTMVP candidate.

[0154] Representative regression-based motion vector field

[0155] To provide fine-grained motion vectors within a block, a regression-based motion vector field (RMVF) tool (e.g., in JVET-M0302) can be implemented, which can attempt to model the motion vectors of each block at the sub-block level based on spatially adjacent motion vectors.

[0156] FIG. 14 is a schematic diagram showing adjacent motion blocks (e.g., 4×4 motion blocks) that can be used for motion parameter derivation. During the regression process, a row of adjacent motion vectors 1410 and a column of adjacent motion vectors 1420 based on 4×4 sub-blocks (and their center positions) from each side of the block can be used. For example, the adjacent motion vectors can be used for RMVF motion parameter derivation.

[0157] FIG. 15 is a schematic diagram showing adjacent motion blocks that can be used for motion parameter derivation to reduce adjacent motion information (e.g., the number of adjacent motion blocks used during the regression process relative to FIG. 14 can be reduced). The reduction amount of adjacent motion information for RMVF parameter derivation of adjacent 4×4 motion blocks can be used for motion parameter derivation (e.g., about half, e.g., about every other adjacent motion block can be used for motion parameter derivation). Certain adjacent motion blocks of the row 1410 and the column 1420 can be selected, determined, or predetermined to reduce adjacent motion information.

[0158] Although about half of the adjacent motion blocks of the row 1410 and the column 1420 are shown as being selected, other percentages (with other motion block positions) can be selected, e.g., to reduce the number of adjacent motion blocks to be used during the regression process.

[0159] When collecting motion information for motion parameter derivation, five regions as shown in the figure (e.g., lower left, upper left, upper right) can be used. The upper right reference motion region and the lower left reference motion region can be limited to half of the corresponding width or height of the current block (e.g., only half).

[0160] In the RMVF mode, the motion of the block can be defined by a 6-parameter motion model. These parameters a xx , a xy , a yx , a yy , b x and b y can be calculated by solving a linear regression model in the sense of mean square error (MSE). The input of the regression model can be the center positions (x, y) of the available adjacent 4×4 sub-blocks and / or the motion vectors (mv x and mv y) consists of, or may include, the central positions (x, y) of available adjacent 4×4 sub-blocks as defined above and / or motion vectors (mv x and mv y ).

[0161] The motion vectors (MV subPU , MV subPU ) of an 8×8 sub-block at the central position (X X_subPU , Y Y_subPU ) can be calculated as described in Equation 43 below:

[0162]

[0163] The motion vectors can be calculated for the 8×8 sub-block relative to the central position of the sub-block (e.g., each sub-block). For example, in the RMVF mode, motion compensation can be applied with 8×8 sub-block precision. To effectively model the motion vector field, the RMVF tool is applied only when at least one motion vector from at least three candidate regions is available.

[0164] Affine motion model parameters can be used to derive the motion vectors of certain pixels (e.g., each pixel) in a CU. Although the complexity of generating pixel-based affine motion compensation predictions may be high (e.g., very high), and also because the memory access bandwidth requirements for such sample-based MC may be high, a sub-block-based affine motion compensation process / method can be implemented (e.g., via VVC). For example, a CU can be divided into several sub-blocks (e.g., 4×4 sub-blocks, square sub-blocks, and / or non-square sub-blocks). Each sub-block can be assigned an MV that can be derived from the affine model parameters. The MV can be the MV at the center of the sub-block (or another position in the sub-block). The pixels in the sub-block (e.g., all pixels in the sub-block) can share the sub-block MV. Sub-block-based affine motion compensation can be a trade-off between coding efficiency and complexity. To achieve a finer granularity of motion compensation, interleaved prediction for affine motion compensation can be implemented, and the interleaved prediction for affine motion compensation can be generated by weighted averaging of two sub-block motion compensation predictions. Interleaved prediction may require and / or use two or more motion compensation predictions per sub-block, and thus may increase the memory bandwidth and complexity.

[0165] In some representative embodiments, methods, apparatuses, processes, and / or operations may be implemented to refine sub-block based affine motion compensation prediction using optical flow (e.g., using and / or based on optical flow). For example, after performing sub-block based affine motion compensation, pixel intensities may be refined by adding a difference derived from an optical flow equation, which is referred to as prediction refinement using optical flow (PROF). PROF may achieve pixel-level granularity without significantly increasing complexity and may maintain a worst-case memory access bandwidth comparable to sub-block based affine motion compensation. PROF may be applied to any scenario where a pixel-level motion vector field is available (e.g., computable) in addition to a prediction signal (e.g., an unrefined motion prediction signal and / or a sub-block based motion prediction signal). In addition to the affine mode, the PROF process may also be used in other sub-block prediction modes. The application of PROF to sub-block modes, such as SbTMVP and / or RMVF, may be implemented. The application of PROF in dual prediction is described herein.

[0166] Representative PROF process for affine mode

[0167] In some representative embodiments, a method, apparatus, and / or process may be implemented to improve the granularity of sub-block based affine motion compensation prediction, for example, by applying a change in pixel intensity derived from optical flow (e.g., an optical flow equation), and may use and / or require one motion compensation operation per sub-block (e.g., only one motion compensation operation per sub-block), which is the same as existing affine motion compensation in, for example, VVC.

[0168] FIG. 16 is a schematic diagram showing a sub-block MV and a pixel-level motion vector difference Δv(i,j) (e.g., which is sometimes also referred to as a refined MV of a pixel) after sub-block based affine motion compensation prediction.

[0169] Referring to FIG. 16, CU 1600 may include sub - blocks 1610, 1620, 1630, and 1640. Each sub - block 1610, 1620, 1630, and 1640 may include a plurality of pixels (e.g., 16 pixels in sub - block 1610). Shown in the figure is a sub - block MV 1650 (e.g., as a rough or average sub - block MV) associated with each pixel 1660(i,j) of sub - block 1610. For each corresponding pixel (i,j) in sub - block 1610, a refined MV 1670(i,j) can be determined (which may indicate the difference between the actual MV of pixel 1660(i,j) and sub - block MV 1650, where (i,j) defines the pixel position in sub - block 1610). For clarity of FIG. 16, only the refined MV 1670(1,1) is labeled, although other individual pixel - level motions are shown. In some representative embodiments, the refined MV 1670(i,j) can be determined as the pixel - level motion vector difference Δv(i,j) (sometimes referred to as the motion vector difference).

[0170] In some representative embodiments, a method, apparatus, process, and / or operation including any of the following operations can be implemented:

[0171] (1) In a first operation: sub - block - based AMC can be performed as disclosed herein to generate a sub - block - based motion prediction I(i,j);

[0172] (2) In a second operation: the spatial gradients g x (i,j) and g y (i,j) of the sub - block - based motion prediction I(i,j) at each sample position can be calculated (in one example, the same process used to generate gradients in BDOF can be used to generate the spatial gradients). For example, the horizontal gradient at a sample position can be calculated as the difference between its right - hand adjacent sample and its left - hand adjacent sample, and / or the vertical gradient at a sample position can be calculated as the difference between its bottom - hand adjacent sample and its top - hand adjacent sample. In another example, a Sobel filter can be used to generate the spatial gradients;

[0173] (3) In a third operation: the change in illuminance intensity of each pixel in the CU can be calculated using and / or via the optical flow equation, for example, as set forth in Equation 44 below:

[0174] ΔI(i,j) = g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (44)

[0175] The value of the motion vector difference Δv(i,j) is the difference 1670 between the pixel-level MV (represented by v(i,j)) calculated for the sample position (i,j) and the sub-block level MV 1650 of the sub-block covering the pixel 1660(i,j), as shown in FIG. 16. The pixel-level MV v(i,j) can be derived from the control point MVs through equations 22 and 23 for the 4-parameter affine model or through equations 32 and 33 for the 6-parameter affine model.

[0176] In some representative embodiments, the motion vector difference value Δv(i,j) can be derived from the affine model parameters through or using equations 24 and 25, where x and y can be the offsets from the pixel position to the center of the sub-block. Since the affine model parameters and pixel offsets do not change between sub-blocks, the motion vector difference Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. For example, equations 45 and 46 below can be used to calculate the difference between the pixel-level MV and the sub-block level MV, because the translational affine parameters (a, b) can be the same for the pixel-level MV and the sub-block MV. (c, d, e, f) can be four additional affine parameters (e.g., four affine parameters other than the translational affine parameters)

[0177]

[0178] where (i,j) can be the pixel position relative to the upper left position of the sub-block, and (x sb ,y sb ) can be the center position of the sub-block relative to the upper left position of the sub-block.

[0179] FIG. 17A is a schematic diagram of a representative procedure for determining the motion vector corresponding to the actual center of a sub-block.

[0180] Referring to FIG. 17A, the two sub-blocks SB0 and SB1 shown in the figure are 4×4 sub-blocks. If the sub-block width is SW and the sub-block height is SH, the sub-block center position can be set to ((SW-1) / 2, (SH-1) / 2). In other examples, the sub-block center position can be estimated based on a position such as (SW / 2, SH / 2). The actual center point of the first sub-block SB0 is P0', and the actual center point of the second sub-block SB1 is P1', which uses ((SW-1) / 2, (SH-1) / 2). By using, for example, (SW / 2, SH / 2) (e.g., in VVC), the estimated center point of the first sub-block SB0 is P0, and the estimated center point of the second sub-block SB1 is P1. In some representative embodiments, the MV of the sub-block can be more accurately based on the actual center position rather than the estimated center position (which is used in VVC).

[0181] Figure 17B is a schematic diagram showing the positions of chrominance samples in a 4:2:0 chrominance format. Referring to Figure 17B, the chrominance sub-block MV can be derived from the MV of the luminance sub-block. For example, in the 4:2:0 chrominance format, a 4×4 chrominance sub-block can correspond to an 8×8 luminance region. Although representative embodiments are shown in conjunction with the 4:2:0 chrominance format, those skilled in the art will understand that other chrominance formats, such as the 4:2:2 chrominance format, can be equivalently used.

[0182] The chrominance sub-block MV can be derived by averaging the upper-left 4×4 luminance sub-block MV and the lower-right luminance sub-block MV. For chrominance sample position types 0, 2, and / or 3, the derived chrominance sub-block MV may or may not be located at the center of the chrominance sub-block. For chrominance sample position types 0, 2, and 3, the chrominance sub-block center position (x sb , y sb ) may or may need to be adjusted by an offset. For example, for 4:2:0 chrominance sample position types 0, 2, and 3, the adjustment can be applied as set forth in the following equations 47-49:

[0183]

[0184]

[0185]

[0186] The sub-block based motion prediction I(i,j) can be refined by adding an intensity change (e.g., an illumination intensity change, e.g., as provided in equation 44). The final (i.e., refined) prediction I′(i,j) can be generated by or by using equation 50 below.

[0187] I′(i,j) = I(i,j) + ΔI(i,j) (50)

[0188] When applying the refinement, the sub-block based affine motion compensation can achieve pixel-level granularity without increasing the worst-case bandwidth and / or memory bandwidth.

[0189] To maintain the accuracy of prediction and / or gradient calculation, the bit depth in the operation-related performance of the sub-block based AMC can be an intermediate bit depth, which can be higher than the coding bit depth.

[0190] The above process can be used to refine the chrominance intensity (e.g., in addition to or instead of refining the luminance intensity). In one example, the intensity difference used in equation 50 can be multiplied by a weight factor w before being added to the prediction, as shown in equation 51 below:

[0191] I′(i,j) = I(i,j) + w·ΔI(i,j) (51)

[0192] Where w can be set to a value between 0 and 1, inclusive. w can be signaled at the CU level or at the picture level. For example, w can be signaled via a weight index. For instance, index table 1 can be used to signal w.

[0193] Index 0 1 2 3 4 Weight 1 / 2 3 / 4 1 / 4 1 0

[0194] Index table 1

[0195] The encoder algorithm can select the value of w that results in the lowest rate-distortion cost.

[0196] For example, the gradient of the predicted sample can be calculated in different ways, such as g x and / or g y . In some representative embodiments, the predicted samples g x and g y can be calculated by applying a 2D Sobel filter. Examples of 3×3 Sobel filters for horizontal and vertical gradients are as follows:

[0197] Horizontal Sobel filter:

[0198] Vertical Sobel filter:

[0199] In other representative embodiments, the gradient can be calculated using a one-dimensional 3-tap filter. An example can include [-1 0 1], which is a simpler (e.g., much simpler) filter than the Sobel filter.

[0200] FIG. 17C is a schematic diagram showing extended sub-block prediction. The shaded circles 1710 are padding samples around a 4×4 sub-block (e.g., the non-shaded circles 1720). Using the Sobel filter as an example, the samples in box 1730 can be used to calculate the gradient of the sample 1740 at the center. Although the Sobel filter can be used to calculate the gradient, other filters such as 3-tap filters are also possible.

[0201] For the above example gradient filters, such as the 3x3 Sobel filter and the one-dimensional filter, extended sub-block prediction can be used and / or required for sub-block gradient calculation. For example, one row at the top and bottom boundaries of the sub-block and one column at the left and right boundaries can be padded to calculate the gradients of those samples at the sub-block boundaries.

[0202] There can be different methods / processes and / or operations to obtain the extended sub-block prediction. In a representative embodiment, given a sub-block size of N×M, the (N+2)×(M+2) extended sub-block prediction can be obtained by performing (N+2)×(M+2) block motion compensation using the sub-block MV. Using this embodiment, the memory bandwidth may increase. To avoid the increase in memory bandwidth, in some representative embodiments, given a K-tap interpolation filter in both the horizontal and vertical directions, the (N+K-1)×(M+K-1) integer reference samples before interpolation can be extracted for interpolation of the N×M sub-block, and the boundary samples of the (N+K-1)×(M+K-1) block can be copied from the neighboring samples of the (N+K-1)×(M+K-1) sub-block, such that the extended region can be (N+K-1+2)×(M+K-1+2). The extended region can be used for interpolation of the (N+2)×(M+2) sub-block. If the sub-block MV points to a fractional position, these representative embodiments can still use and / or require additional interpolation operations to generate the (N+2)×(M+2) prediction.

[0203] For example, to reduce the computational complexity, in other representative embodiments, the sub-block prediction can be obtained by performing N×M block motion compensation using the sub-block MV. The boundary of the (N+2)×(M+2) prediction can be obtained without interpolation by any of the following: (1) integer motion compensation, where the MV is the integer part of the sub-block MV; (2) integer motion compensation, where the MV is the nearest integer MV of the sub-block MV; and / or (3) copying from the nearest neighboring samples in the N×M sub-block prediction.

[0204] The accuracy and / or range of the pixel-level refined MV (e.g., Δv x and Δv y ) may affect the accuracy of the PROF. In some representative embodiments, a combination of a multi-bit fractional component and another multi-bit integer component can be implemented. For example, a 5-bit fractional component and an 11-bit integer component can be used. The combination of the 5-bit fractional component and the 11-bit integer component can utilize a total of 16 bits to represent the MV range from -1024 to 1023 with 1 / 32 pixel accuracy.

[0205] The accuracy of the gradient (e.g., g x and g y ) and the accuracy of the intensity change ΔI may affect the performance of the PROF. In some representative embodiments, the prediction sample accuracy can be maintained at or kept at a predetermined number or a signal-informed number of bits (e.g., the internal sample accuracy defined in the current VVC draft, which is 14 bits). In some representative embodiments, the gradient and / or the intensity change ΔI can have the same accuracy as the prediction samples.

[0206] The range of the intensity change ΔI may affect the performance of PROF. The intensity change ΔI can be clipped to a smaller range to avoid false values caused by inaccurate affine models. In one example, the intensity change ΔI can be clipped to predition_bitdepth - 2.

[0207] Δv x and Δv y The combination of the number of bits of the fractional components of Δv x and Δv y the number of bits of the fractional components of the gradient, and the number of bits of the intensity change ΔI can together affect the complexity of certain hardware or software implementations. In a representative embodiment, 5 bits can be used to represent the fractional components of Δv

[0208] To reduce the computational complexity, in some cases, PROF can be skipped. For example, if the magnitude of all pixel - based differences (e.g., refinements) MV(Δv(i,j)) within a 4×4 sub - block is less than a threshold, then for the entire affine CU, PROF can be skipped. If the gradient of all samples within a 4×4 sub - block is less than a threshold, then PROF can be skipped.

[0209] PROF can be applied to chrominance components, such as the Cb and / or Cr components. The difference MV of the Cb and / or Cr components of a sub - block can reuse the difference MV of the sub - block (e.g., the difference MV calculated for different sub - blocks within the same CU can be reused).

[0210] Although the gradient process disclosed herein (e.g., using replicated reference samples to extend the sub - block for gradient calculation) is shown to be used with PROF operations, the gradient process can be used with other operations such as BDOF operations and / or affine motion estimation operations.

[0211] Representative PROF processes for other sub - block modes

[0212] PROF can be applied to any scenario where a pixel - level motion vector field can be obtained (e.g., calculated) in addition to the prediction signal (e.g., unrefined prediction signal). For example, in addition to the affine mode, prediction refinement using optical flow can be used in other sub - block prediction modes, such as the SbTMVP mode (e.g., the ATMVP mode in VVC), or the regression - based motion vector field (RMVF)

[0213] In some representative embodiments, a method can be implemented to apply PROF to SbTMVP. For example, such a method can include any of the following:

[0214] (1) In a first operation, sub-block level MVs and sub-block predictions can be generated based on the existing SbTMVP process described herein;

[0215] (2) In a second operation, affine model parameters can be estimated through the block MV field by using a linear regression method / process;

[0216] (3) In a third operation, pixel-level MVs can be derived by the affine model parameters obtained in the second operation, and an associated pixel-level motion refinement vector (Δv(i,j)) relative to the sub-block MV can be calculated; and / or

[0217] (4) In a fourth operation, a prediction refinement process utilizing optical flow can be applied to generate a final prediction, etc.

[0218] In some representative embodiments, a method can be implemented to apply PROF to RMVF. For example, such a method can include any of the following:

[0219] (1) In a first operation, the sub-block level MV field, sub-block predictions, and / or affine model parameters a xx ,a xy ,a yx ,a yy ,b x and b x can be generated based on the RMVF process described herein;

[0220] (2) In a second operation, the pixel-level MV offset (Δv(i,j)) from the sub-block level MV can be derived by the affine model parameters a xx ,a xy ,a yx ,a yy ,b x and b x through Equation 52 below:

[0221]

[0222] where (i,j) is the pixel offset from the sub-block center. Since the affine parameters and / or the pixel offset from the sub-block center do not change between sub-blocks, the pixel MV offset can be calculated for the first sub-block (e.g., only the pixel MV offset for the first sub-block needs to be or will be calculated), and this pixel MV offset can be reused for other sub-blocks in the CU; and / or

[0223] (3) In a third operation, the PROF process can be applied to generate a final prediction, for example, by applying Equations 44 and 50.

[0224] Representative PROF methods for dual prediction

[0225] In addition to or instead of using PROF in the single prediction described herein, the PROF technique can be used for dual prediction. When used in dual prediction, the PROF can be used to generate L0 prediction and / or L1 prediction, e.g., before they are combined by weights. To reduce computational complexity, PROF can be applied (e.g., can be applied only) to one prediction, such as L0 or L1. In some representative embodiments, PROF can be applied (e.g., can be applied only) to a list (e.g., a list associated with a reference picture that is close (e.g., within a threshold) and / or closest to the current picture).

[0226] Representative processes for PROF enabling

[0227] PROF enabling can be signaled at or in the sequence parameter set (SPS) header, picture parameter set (PPS) header, and / or tile group header. In some embodiments, a flag can be signaled to indicate whether the PROF is enabled for the affine mode. If the flag is set to a first logical level (e.g., "true"), then the PROF can be used for single prediction and dual prediction. In some embodiments, if the first flag is set to "true", then a second flag can be used to indicate whether the PROF is enabled or not for the dual prediction affine mode. If the first flag is set to a second logical level (e.g., "false"), then it can be inferred that the second flag is set to "false". If the first flag is set to "true", then a flag can be signaled at or in the SPS header, PPS header, and / or tile group header to indicate whether the PROF is applied to the chrominance component, so that the control of the PROF for the luminance component and the chrominance component can be separated.

[0228] Representative methods for conditionally enabling PROF

[0229] For example, to reduce complexity, PROF can be applied when certain conditions are met (e.g., only when certain conditions are met). For example, for small CU sizes (e.g., below a threshold level), the affine motion may be relatively small, such that the benefits of applying PROF may be limited. In some representative embodiments, when the CU size is small (e.g., for CU sizes not greater than 16x16, such as 8x8, 8x16, 16x8) or in such cases, PROF can be deactivated in affine motion compensation to reduce the complexity of both the encoder and / or the decoder. In some representative embodiments, when the CU size is small (below the same or a different threshold level), PROF can be skipped in affine motion estimation (e.g., only skipped in affine motion estimation), e.g., to reduce encoder complexity, and PROF can be performed at the decoder regardless of the CU size. For example, on the encoder side, after motion estimation that searches for affine model parameters (e.g., control point MVs), the motion compensation (MC) process can be invoked and PROF can be performed. For each iteration during motion estimation, the MC process can also be invoked. In the MC during motion estimation, PROF can be skipped to save complexity, and there will be no prediction mismatch between the encoder and the decoder because the final MC in the encoder will run PROF. That is, when the encoder searches for affine model parameters (e.g., affine MVs) for prediction of a CU, PROF refinement may not be applied, and once the encoder has completed the search or after the encoder has completed the search, the encoder can apply PROF to refine the prediction of the CU using the affine model parameters determined from the search.

[0230] In some representative embodiments, the difference between CPMVs can be used as a criterion to determine whether to enable PROF. When the difference between CPMVs is small (e.g., below a threshold level) such that the affine motion is small, the benefits of applying PROF may be limited, and PROF can be deactivated for affine motion compensation and / or affine motion estimation. For example, for a 4-parameter affine mode, if the following conditions are met (e.g., all of the following conditions are met), then PROF can be deactivated:

[0231]

[0232]

[0233] For a 6-parameter affine mode, in addition to or instead of the above conditions, if the following conditions are met (e.g., all of the following conditions are also met), then PROF can be deactivated:

[0234]

[0235]

[0236] Where T is a predefined threshold, e.g., 4. This PROF skipping process based on CPMV or affine parameters can be applied at the encoder (e.g., also only applied at the encoder), and the decoder may or may not skip the PROF.

[0237] Representative processes of PROF in combination with or instead of the deblocking filter

[0238] Since PROF can be a pixel-by-pixel refinement compensating for block-based MC, the motion difference between block boundaries can be reduced (e.g., can be greatly reduced). When PROF is applied, the encoder and / or decoder may skip the application of the deblocking filter, and / or may apply a weaker filter to sub-block boundaries. For a CU divided into multiple transform units (TUs), blocking effects may appear on the transform block boundaries.

[0239] In some representative embodiments, unless the sub-block boundary coincides with the TU boundary, the encoder and / or decoder may skip the application of the deblocking filter or may apply one or more weaker filters on the sub-block boundary.

[0240] When PROF is applied to luminance (e.g., only applied to luminance) or under the condition that PROF is applied to luminance (e.g., only applied to luminance), the encoder and / or decoder may skip the application of the deblocking filter and / or may apply one or more weaker filters for luminance (e.g., only luminance) on the sub-block boundary. For example, the boundary strength parameter Bs can be used to apply a weaker deblocking filter.

[0241] For example, when PROF is applied, the encoder and / or decoder may skip applying the deblocking filter to the sub-block boundary unless the sub-block boundary coincides with the TU boundary. In this case, the deblocking filter can be applied to reduce or remove the blocking effects that may occur along the TU boundary.

[0242] As another example, unless the sub-block boundary coincides with the TU boundary, when PROF is applied, the encoder and / or decoder may apply a weaker deblocking filter to the sub-block boundary. It is contemplated that the "weaker" deblocking filter can be a deblocking filter weaker than the one typically applied to the sub-block boundary when PROF is not applied. When the sub-block boundary coincides with the TU boundary, a stronger deblocking filter can be applied to reduce or remove the blocking effects that are expected to be more visible along the sub-block boundary coinciding with the TU boundary.

[0243] In some representative embodiments, when applying the PROF (e.g., only applying) to luminance or under the condition of applying the PROF (e.g., only applying) to luminance, the encoder and / or decoder may align the application of the deblocking filter for chrominance to luminance for design uniformity purposes, for example, despite the lack of applying the PROF to chrominance. For example, if the PROF is only applied to luminance, the normal application of the deblocking filter for luminance may be changed based on whether the PROF is applied (and possibly based on whether there is a TU boundary at the sub-block boundary). In some representative embodiments, the deblocking filter may be applied to the sub-block boundary of chrominance to match (and / or mirror) the process of luminance deblocking, rather than having separate / different logic for applying the deblocking filter to the corresponding chrominance pixels.

[0244] Figure 18A is a flowchart showing a first representative encoding and / or decoding method.

[0245] Reference Figure 18A , a representative method 1800 for encoding and / or decoding may include: at block 1805, the encoder 100 or 300 and / or the decoder 200 or 500 obtain a sub-block based motion prediction signal for a current block of, for example, video. At block 1810, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain one or more spatial gradients of the sub-block based motion prediction signal for the current block or one or more motion vector differences associated with the sub-blocks of the current block. At block 1815, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain a refinement signal for the current block based on the one or more obtained spatial gradients or the one or more obtained motion vector differences associated with the sub-blocks of the current block. At block 1820, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain a refined motion prediction signal for the current block based on the sub-block based motion prediction signal and the refinement signal. In some embodiments, the encoder 100 or 300 may encode the current block based on the refined motion prediction signal, or the decoder 200 or 500 may decode the current block based on the refined motion prediction signal. The refined motion prediction signal may be, for example, a refined motion inter prediction signal generated by the GBi encoder 300 and / or the GBi decoder 500, and may use one or more PROF operations.

[0246] In some representative embodiments, such as representative embodiments related to other methods including methods 1850 and 1900 described herein, obtaining the sub-block based motion prediction signal for the current block of the video may include generating the sub-block based motion prediction signal.

[0247] In certain representative embodiments, for example, with respect to other methods described herein that particularly include methods 1850 and 1900, obtaining one or more spatial gradients of the sub-block based motion prediction signal for the current block or one or more motion vector differences associated with sub-blocks of the current block may include: determining the one or more spatial gradients of the sub-block based motion prediction signal (e.g., associated with a gradient filter).

[0248] In certain representative embodiments, for example, in representative embodiments with respect to other methods described herein that particularly include methods 1850 and 1900, obtaining one or more spatial gradients of the sub-block based motion prediction signal for the current block or one or more motion vector differences associated with sub-blocks of the current block may include: determining the one or more motion vector differences associated with sub-blocks of the current block.

[0249] In certain representative embodiments, for example, in representative embodiments related to other methods described herein that particularly include methods 1850 and 1900, obtaining the refinement signal for the current block based on the one or more determined spatial gradients or the one or more determined motion vector differences may include: determining a motion prediction refinement signal for the current block as the refinement signal based on the determined spatial gradients.

[0250] In certain representative embodiments, for example, in representative embodiments related to other methods described herein including methods 1850 and 1900, obtaining the refinement signal for the current block based on the one or more determined spatial gradients or the one or more determined motion vector differences may include: determining a motion prediction refinement signal for the current block as the refinement signal based on the determined motion vector differences.

[0251] The term "determine" or "determining" when referring to something such as information may generally include one or more of the following: estimating, calculating, predicting, obtaining, and / or retrieving the information. For example, determine may refer to retrieving something from a memory or a bitstream, etc.

[0252] In certain representative embodiments, for example, in representative embodiments related to other methods described herein including methods 1850 and 1900, obtaining the refined motion prediction signal for the current block based on the sub-block based motion prediction signal and the refinement signal may include: combining (e.g., adding or subtracting, etc.) the sub-block based motion prediction signal and the motion prediction refinement signal to produce the refined motion prediction signal for the current block.

[0253] In certain representative embodiments, for example, in representative embodiments related to other methods including methods 1850 and 1900 described herein, encoding and / or decoding the current block based on the refined motion prediction signal may include: encoding the video using the refined motion prediction signal as a prediction for the current block, and / or decoding the video using the refined motion prediction signal as a prediction for the current block.

[0254] Figure 18B is a flowchart showing a second representative encoding and / or decoding method.

[0255] Referring Figure 18B , a representative method 1850 for encoding and / or decoding a video may include: at block 1855, an encoder 100 or 300 and / or a decoder 200 or 500 generate a sub-block-based motion prediction signal. At block 1860, the encoder 100 or 300 and / or the decoder 200 or 500 may determine one or more spatial gradients of the sub-block-based motion prediction signal (e.g., associated with a gradient filter). At block 1865, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined spatial gradients. At block 1870, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block-based motion prediction signal and the motion prediction refinement signal to produce a refined motion prediction signal for the current block. At block 1875, the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block, and / or the decoder 200 or 500 may decode the video using the refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations at blocks 1810, 1820, 1830, and 1840 may be performed for the current block, which generally refers to the block currently being encoded or decoded. The refined motion prediction signal may be a refined motion inter-prediction signal generated (e.g., by a GBi encoder 300 and / or a GBi decoder 500) and may use one or more PROF operations.

[0256] For example, the determination of the one or more spatial gradients of the sub-block based motion prediction signal by the encoder 100 or 300 and / or the decoder 200 or 500 may include: determining a first set of spatial gradients associated with a first reference picture and a second set of spatial gradients associated with a second reference picture. The determination of the motion prediction refinement signal of the current block by the encoder 100 or 300 and / or the decoder 200 or 500 may be based on the determined spatial gradients, and may include: determining an inter-frame motion prediction refinement signal (e.g., a bi-prediction signal) of the current block based on the first and second sets of spatial gradients, and may also be based on the weighting information W (e.g., indicating or including one or more weighting values associated with one or more reference pictures).

[0257] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the encoder 100 or 300 may generate, use, and / or transmit the weighting information W to the decoder 200 or 500, and / or the decoder 200 or 500 may receive or obtain the weighting information W. For example, the inter-frame motion prediction refinement signal of the current block may be based on: (1) a first gradient value derived from the first set of spatial gradients and weighted according to a first weighting factor indicated by the weighting information W, and / or (2) a second gradient value derived from the second set of spatial gradients and weighted according to a second weighting factor indicated by the weighting information W.

[0258] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may further include the encoder 100 or 300 and / or the decoder 200 or 500 determining affine motion model parameters of the current block of the video such that the determined affine motion model parameters can be used to generate the sub-block based motion prediction signal.

[0259] In certain representative embodiments that include representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, one or more spatial gradients of the sub-block-based motion prediction signal, which may include calculating at least one gradient value for a corresponding sample position, a partial corresponding sample position, or each corresponding sample position in at least one sub-block of the sub-block-based motion prediction signal. For example, the calculation of at least one gradient value for a corresponding sample position, a partial corresponding sample position, or each corresponding sample position in at least one sub-block of the sub-block-based motion prediction signal may include applying a gradient filter to the corresponding sample position in the at least one sub-block of the sub-block-based motion prediction signal for a corresponding sample position, a partial corresponding sample position, or each corresponding sample position.

[0260] In certain representative embodiments that include representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may further include encoder 100 or 300 and / or decoder 200 or 500 determining a set of motion vector differences associated with the sample positions of a first sub-block of the current block of the sub-block-based motion prediction signal. In some examples, the differences may be determined for a sub-block (e.g., the first sub-block) and the differences may be reused for some or all of the other sub-blocks in the current block. In certain examples, the set of motion vector differences may be determined by using an affine motion model or a different motion model (e.g., another sub-block-based motion model, such as the SbTMVP model), and the sub-block-based motion prediction signal may be generated. As an example, the set of motion vector differences may be determined for the first sub-block of the current block and may be used to determine the motion prediction refinement signal for one or more further sub-blocks of the current block.

[0261] In certain representative embodiments that include representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the one or more spatial gradients of the sub-block-based motion prediction signal and the set of motion vector differences may be used to determine the motion prediction refinement signal of the current block.

[0262] In certain representative embodiments that include representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, by using the affine motion model of the current block, the set of motion vector differences may be determined and the sub-block-based motion prediction signal may be generated.

[0263] In certain representative embodiments including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, determining the one or more spatial gradients of the sub-block based motion prediction signal may include: for one or more respective sub-blocks of the current block: using the sub-block based motion prediction signal and neighboring reference samples adjacent to and surrounding the respective sub-block to determine an extended sub-block; and using the determined extended sub-block to determine the spatial gradient of the respective sub-block to determine the motion prediction refinement signal.

[0264] Figure 19 is a flowchart showing a third representative encoding and / or decoding method.

[0265] Referring to Figure 19 , representative method 1900 for encoding and / or decoding video may include: at block 1910, encoder 100 or 300 and / or decoder 200 or 500 generate a sub-block based motion prediction signal. At block 1920, encoder 100 or 300 and / or decoder 200 or 500 may determine a set of motion vector differences associated with the sub-blocks of the current block (e.g., the set of motion vector differences may be associated with, for example, all sub-blocks of the current block). At block 1930, encoder 100 or 300 and / or decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined set of motion vector differences. At block 1940, encoder 100 or 300 and / or decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block based motion prediction signal and the motion prediction refinement signal to produce or generate a refined motion prediction signal for the current block. At block 1950, encoder 100 or 300 may use the refined motion prediction signal as a prediction for the current block to encode video, and / or decoder 200 or 500 may use the refined motion prediction signal as a prediction for the current block to decode video. In certain embodiments, the operations at blocks 1910, 1920, 1930, and 1940 may be performed for a current block that generally refers to the block currently being encoded or decoded. In certain representative embodiments, the refined motion prediction signal may be a refined motion inter prediction signal (e.g., generated by GBi encoder 300 or GBi decoder 500) and may use one or more PROF operations.

[0266] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, motion model parameters (e.g., one or more affine motion model parameters) of the current block of the video such that a sub-block-based motion prediction signal can be generated using the determined motion model parameters (e.g., affine motion model parameters).

[0267] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, one or more spatial gradients of the sub-block-based motion prediction signal. For example, the determining of the one or more spatial gradients of the sub-block-based motion prediction signal may include: calculating at least one gradient value of a corresponding sample position, a portion of corresponding sample positions, or each corresponding sample position in at least one sub-block of the sub-block-based motion prediction signal. For example, the calculating of at least one gradient value of a corresponding sample position, a portion of corresponding sample positions, or each corresponding sample position in at least one sub-block of the sub-block-based motion prediction signal may include: applying a gradient filter to the corresponding sample position in the at least one sub-block of the sub-block-based motion prediction signal for a corresponding sample position, a portion of corresponding sample positions, or each corresponding sample position.

[0268] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, the motion prediction refinement signal of the current block using gradient values associated with spatial gradients of a corresponding sample position, a portion of corresponding sample positions, or each corresponding sample position of the current block and a set of motion vector differences determined to be associated with sample positions of sub-blocks (e.g., any sub-block) of the current block of the sub-block motion prediction signal.

[0269] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the determining of the motion prediction refinement signal of the current block may use gradient values associated with spatial gradients of one or more corresponding sample positions or each sample position of one or more sub-blocks of the current block and a determined set of motion vector differences.

[0270] Figure 20is a flowchart showing a fourth representative encoding and / or decoding method.

[0271] Reference Figure 20 , a representative method 2000 for encoding and / or decoding video may include: at block 2010, an encoder 100 or 300 and / or a decoder 200 or 500 generate a sub-block based motion prediction signal using at least a first motion vector for a first sub-block of the current block and a further motion vector for a second sub-block of the current block. At block 2020, the encoder 100 or 300 and / or the decoder 200 or 500 may compute a first set of gradient values at a first sample position in the first sub-block of the sub-block based motion prediction signal and a different second set of gradient values at a second sample position in the first sub-block of the sub-block based motion prediction signal. At block 2030, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a first set of motion vector differences at the first sample position and a different second set of motion vector differences at the second sample position. For example, the first set of motion vector differences at the first sample position may indicate a difference between a motion vector at the first sample position and a motion vector of the first sub-block, and the second set of motion vector differences at the second sample position may indicate a difference between a motion vector at the second sample position and a motion vector of the first sub-block. At block 2040, the encoder 100 or 300 and / or the decoder 200 or 500 may use the first and second sets of gradient values and the first and second sets of motion vector differences to determine a prediction refinement signal. At block 2050, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block based motion prediction signal and the prediction refinement signal to produce a refined motion prediction signal. At block 2060, the encoder 100 or 300 may use the refined motion prediction signal as a prediction for the current block to encode video, and / or the decoder 200 or 500 may use the refined motion prediction signal as a prediction for the current block to decode video. In some embodiments, the operations at blocks 2010, 2020, 2030, 2040, and 2050 may be performed for a current block including multiple sub-blocks.

[0272] Figure 21 is a flowchart showing a fifth representative encoding and / or decoding method.

[0273] Reference Figure 21, a representative method 2100 for encoding and / or decoding video may include: at block 2110, encoder 100 or 300 and / or decoder 200 or 500 generates a sub-block based motion prediction signal for a current block. At block 2120, encoder 100 or 300 and / or decoder 200 or 500 may use optical flow information indicating refined motion of a plurality of sample positions in the current block of the sub-block based motion prediction signal to determine a prediction refinement signal. At block 2130, encoder 100 or 300 and / or decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block based motion prediction signal with the prediction refinement signal to produce a refined motion prediction signal. At block 2140, encoder 100 or 300 may use the refined motion prediction signal as a prediction for the current block to encode video, and / or decoder 200 or 500 may use the refined motion prediction signal as a prediction for the current block to decode video. For example, the current block may include a plurality of sub-blocks, and the sub-block based motion prediction signal may be generated using at least a first motion vector of a first sub-block of the current block and a further motion vector of a second sub-block of the current block.

[0274] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, a prediction refinement signal that may use optical flow information. This determination may include calculating, by encoder 100 or 300 and / or decoder 200 or 500, a first set of gradient values for a first sample position in the first sub-block of the sub-block based motion prediction signal and a different second set of gradient values for a second sample position in the first sub-block of the sub-block based motion prediction signal. A first set of motion vector differences for the first sample position and a different second set of motion vector differences for the second sample position may be determined. For example, the first set of motion vector differences for the first sample position may indicate a difference between a motion vector at the first sample position and the motion vector of the first sub-block, and the second set of motion vector differences for the second sample position may indicate a difference between a motion vector at the second sample position and the motion vector of the first sub-block. Encoder 100 or 300 and / or decoder 200 or 500 may use the first and second sets of gradient values and the first and second sets of motion vector differences to determine the prediction refinement signal.

[0275] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, the prediction refinement signal that may use the optical flow information. This determination may include calculating a third set of gradient values at a first sample position in a second sub-block of the sub-block-based motion prediction signal and a fourth set of gradient values at a second sample position in the second sub-block of the sub-block-based motion prediction signal. Encoder 100 or 300 and / or decoder 200 or 500 may use the third and fourth sets of gradient values and the first and second sets of motion vector differences to determine the prediction refinement signal for the second sub-block.

[0276] Figure 22 is a flowchart showing a sixth representative encoding and / or decoding method.

[0277] Reference Figure 22, A representative method 2200 for encoding and / or decoding video may include: At block 2210, encoder 100 or 300 and / or decoder 200 or 500 determines a motion model for a current block of the video. The current block may include a plurality of sub-blocks. For example, the motion model may generate separate (e.g., per-sample) motion vectors for a plurality of sample positions in the current block. At block 2220, encoder 100 or 300 and / or decoder 200 or 500 may use the determined motion model to generate a sub-block-based motion prediction signal for the current block. The generated sub-block-based motion prediction signal may use one motion vector for each sub-block of the current block. At block 2230, encoder 100 or 300 and / or decoder 200 or 500 may calculate gradient values by applying a gradient filter to a portion of the plurality of sample positions of the sub-block-based motion prediction signal. At block 2240, encoder 100 or 300 and / or decoder 200 or 500 may determine a motion vector difference for the portion of the sample positions, each of which may indicate the difference between the motion vector (e.g., individual motion vector) generated according to the motion model for the corresponding sample position and the motion vector used to generate the sub-block-based motion prediction signal for the sub-block containing the corresponding sample position. At block 2250, encoder 100 or 300 and / or decoder 200 or 500 may use the gradient values and the motion vector differences to determine a prediction refinement signal. At block 2260, encoder 100 or 300 and / or decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block-based motion prediction signal and the prediction refinement signal to produce a refined motion prediction signal for the current block. At block 2270, encoder 100 or 300 may use the refined motion prediction signal as a prediction for the current block to encode the video, and / or decoder 200 or 500 may use the refined motion prediction signal as a prediction for the current block to decode the video.

[0278] Figure 23 is a flowchart showing a seventh representative encoding and / or decoding method.

[0279] Reference Figure 23, a representative method 2300 for encoding and / or decoding video may include: at block 2310, encoder 100 or 300 and / or decoder 200 or 500 perform sub-block based motion compensation to generate a sub-block based motion prediction signal as a rough motion prediction signal. At block 2320, encoder 100 or 300 and / or decoder 200 or 500 may compute one or more spatial gradients of the sub-block based motion prediction signal at a plurality of sample positions. At block 2330, encoder 100 or 300 and / or decoder 200 or 500 may compute the intensity change per pixel in the current block based on the computed spatial gradients. At block 2340, encoder 100 or 300 and / or decoder 200 or 500 may determine a per-pixel based motion prediction signal as a refined motion prediction signal based on the computed intensity change per pixel. At block 2350, encoder 100 or 300 and / or decoder 200 or 500 may predict the current block using the rough motion prediction signal for each sub-block of the current block and using the refined motion prediction signal for each pixel of the current block. In certain embodiments, the operations at blocks 2310, 2320, 2330, 2340, and 2350 may be performed for at least one block (e.g., the current block) in the video. For example, computing the intensity change per pixel in the current block may include determining the illuminance intensity change for each pixel in the current block according to an optical flow equation. Prediction of the current block may include predicting the motion vector for each corresponding pixel in the current block by combining the rough motion prediction vector of the sub-block including the corresponding pixel with a refined motion prediction vector relative to the rough motion prediction vector and associated with the corresponding pixel.

[0280] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the one or more spatial gradients of the sub-block based motion prediction signal may include any of the following: a horizontal gradient and / or a vertical gradient, and for example, the horizontal gradient may be computed as the luminance difference or chrominance difference between the right adjacent sample of the samples of the sub-block and the left adjacent sample of the samples of the sub-block, and / or the vertical gradient may be computed as the luminance difference or chrominance difference between the bottom adjacent sample of the samples of the sub-block and the top adjacent sample of the samples of the sub-block.

[0281] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, one or more spatial gradients of the sub-block prediction may be generated using a Sobel filter.

[0282] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the coarse motion prediction signal may use one of the following: a 4-parameter affine model or a 6-parameter affine model. For example, the block-based motion compensation may be one of the following: (1) motion compensation based on an affine block; or (2) another compensation (e.g., block-based temporal motion vector prediction (SbTMVP) mode motion compensation; and / or regression-based motion vector field (RMVF) mode compensation). Under the condition of performing motion compensation in the SbTMVP mode, the method may include: estimating affine model parameters using a block motion vector field through a linear regression operation; and deriving pixel-level motion vectors using the estimated affine model parameters. In the case of performing motion compensation in the RMVF mode, the method may include: estimating affine model parameters; and deriving a pixel-level motion vector offset from block-level motion vectors using the estimated affine model parameters. For example, the pixel motion vector offset may be relative to the center of the block (e.g., the actual center or the sample position closest to the actual center). For example, the coarse motion prediction vector of the block may be based on the actual center position of the block.

[0283] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include: encoder 100 or 300 or decoder 200 or 500 selects one of the following as the center position associated with the coarse motion prediction vector (e.g., block-based motion prediction vector) of each block: (1) the actual center of each block; or (2) one of the pixel (e.g., sample) positions closest to the center of the block. For example, predicting the current block using the coarse motion prediction signal of the current block (e.g., block-based motion prediction signal) and using the refined motion prediction signal of each pixel (e.g., sample) of the current block may be based on the selected center position of each block. For example, encoder 100 or 300 and / or decoder 200 or 500 may determine the center position associated with the chrominance pixels of the block; and may determine the offset to the center position of the chrominance pixels of the block based on the chrominance position sample type associated with the chrominance pixels. The coarse motion prediction signal for the block (e.g., the block-based motion prediction signal) may be based on the actual position of the block corresponding to the center position of the determined chrominance pixels adjusted by the offset.

[0284] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 generating or the decoder 200 or 500 receiving information indicating whether Prediction Refinement using Optical Flow (PROF) is enabled in one of the following: (1) the Sequence Parameter Set (SPS) header, (2) the Picture Parameter Set (PPS) header, or (3) the Tile Group header. For example, under the condition that PROF is enabled, a refinement motion prediction operation may be performed such that the current block can be predicted using the rough motion prediction signal (e.g., a sub-block based motion prediction signal) and the refinement motion prediction signal. As another example, under the condition that the PROF is not enabled, the refinement motion prediction operation is not performed such that the current block can be predicted only using the rough motion prediction signal (e.g., a sub-block based motion prediction signal).

[0285] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining whether to perform a refinement motion prediction operation on the current block or in the affine motion estimation based on the attributes of the current block and / or the attributes of the affine motion estimation.

[0286] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining whether to perform a refinement motion prediction operation on the current block or in the affine motion estimation based on the attributes of the current block and / or the attributes of the affine motion estimation. For example, determining whether to perform a refinement motion prediction operation on the current block based on the attributes of the current block may include determining whether to perform a refinement motion prediction operation on the current block based on any of the following: (1) the size of the current block exceeds a specific size; and / or (2) the Control Point Motion Vector (CPMV) difference exceeds a threshold.

[0287] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 applying a first deblocking filter to one or more boundaries of sub-blocks of the current block that coincide with transform unit boundaries, and applying a second different deblocking filter to other boundaries of the sub-blocks of the current block that do not coincide with any transform unit boundaries. For example, the first deblocking filter may be a stronger deblocking filter than the second deblocking filter.

[0288] Figure 24 is a flowchart showing an eighth representative encoding and / or decoding method.

[0289] Reference Figure 24 , a representative method 2400 for encoding and / or decoding video may include: at block 2410, the encoder 100 or 300 and / or the decoder 200 or 500 performing sub-block based motion compensation to generate a sub-block based motion prediction signal as a rough motion prediction signal. At block 2420, the encoder 100 or 300 and / or the decoder 200 or 500 may determine, for each respective boundary sample of the sub-blocks of the current block, samples corresponding to samples adjacent to the respective boundary sample and one or more reference samples around the sub-block as surrounding reference samples, and may use the surrounding reference samples and the samples of the sub-block adjacent to the respective boundary sample to determine one or more spatial gradients associated with the respective boundary sample. At block 2430, the encoder 100 or 300 and / or the decoder 200 or 500 may use the samples of the sub-block adjacent to each respective non-boundary sample of the sub-block to determine one or more spatial gradients associated with the respective non-boundary sample. At block 2440, the encoder 100 or 300 and / or the decoder 200 or 500 may use the determined spatial gradients of the sub-block to calculate the intensity change per pixel in the current block. At block 2450, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a per-pixel based motion prediction signal as a refined motion prediction signal based on the calculated intensity change per pixel. At block 2460, the encoder 100 or 300 and / or the decoder 200 or 500 may use the rough motion prediction signal associated with each sub-block of the current block and the refined motion prediction signal associated with each pixel of the current block to predict the current block. In certain embodiments, the operations at blocks 2410, 2420, 2430, 2440, 2450, and 2460 may be performed for at least one block (e.g., the current block) in the video.

[0290] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, the determination of the one or more spatial gradients of the boundary samples and the non-boundary samples may include: calculating the one or more spatial gradients using any of the following: (1) a vertical Sobel filter; (2) a horizontal Sobel filter; or (3) a 3-tap filter.

[0291] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 copying around a reference sample from a reference repository without any further operation, and the determination of the one or more spatial gradients associated with the corresponding boundary sample may use the copied reference sample around to determine the one or more spatial gradients associated with the corresponding boundary sample.

[0292] Figure 25 is a flowchart showing representative gradient calculation methods.

[0293] Reference Figure 25 , a representative method 2500 for calculating the gradient of a sub-block using a reference sample corresponding to a sample adjacent to the boundary of the sub-block (e.g., used in encoding and / or decoding video) may include: at block 2510, the encoder 100 or 300 and / or the decoder 200 or 500 for each corresponding boundary sample of the sub-block of the current block, determines one or more reference samples corresponding to the sample adjacent to the corresponding boundary sample and surrounding the sub-block as the reference sample around, and uses the reference sample around and the sample of the sub-block adjacent to the corresponding boundary sample to determine one or more spatial gradients associated with the corresponding boundary sample. At block 2520, the encoder 100 or 300 and / or the decoder 200 or 500 for each corresponding non-boundary sample in the sub-block, may use the sample of the sub-block adjacent to the corresponding non-boundary sample to determine one or more spatial gradients associated with the corresponding non-boundary sample. In certain embodiments, the operations at blocks 2510 and 2520 may be performed for at least one block in the video (e.g., the current block).

[0294] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, 2500, and 2600, one or more determined spatial gradients can be used to predict a current block by any of the following: (1) a prediction refinement using optical flow (PROF) operation; (2) a bidirectional optical flow operation; or (3) an affine motion estimation operation.

[0295] Figure 26 is a flowchart showing a ninth representative encoding and / or decoding method.

[0296] Refer to Figure 26 , a representative method 2600 for encoding and / or decoding video can include: At block 2610, encoder 100 or 300 and / or decoder 200 or 500 generates a sub-block based motion prediction signal for a current block of the video. For example, the current block can include a plurality of sub-blocks. At block 2620, encoder 100 or 300 and / or decoder 200 or 500 can, for one or more or each respective sub-block of the current block, use the sub-block based motion prediction signal and neighboring reference samples adjacent to and surrounding the respective sub-block to determine an extended sub-block, and use the determined extended sub-block to determine the spatial gradient of the respective sub-block. At block 2630, encoder 100 or 300 and / or decoder 200 or 500 can determine a motion prediction refinement signal for the current block based on the determined spatial gradient. At block 2640, encoder 100 or 300 and / or decoder 200 or 500 can combine (e.g., add or subtract, etc.) the sub-block based motion prediction signal and the motion prediction refinement signal to produce a refined motion prediction signal for the current block. At block 2650, encoder 100 or 300 can use the refined motion prediction signal as a prediction for the current block to encode the video, and / or decoder 200 or 500 can use the refined motion prediction signal as a prediction for the current block to decode the video. In certain embodiments, the operations at blocks 2610, 2620, 2630, 2640, and 2650 can be performed for at least one block (e.g., the current block) in the video.

[0297] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 copying the neighboring reference samples from a reference repository without any further manipulation. For example, the determination of the spatial gradient of the corresponding sub-block may use the copied neighboring reference samples to determine the gradient values associated with the sample positions on the boundary of the corresponding sub-block. The neighboring reference samples of the extended block may be copied from the nearest integer positions in the reference picture containing the current block. In some examples, the neighboring reference samples of the extended block have the nearest integer motion vectors rounded from the original precision.

[0298] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining the affine motion model parameters of the current block of the video such that the determined affine motion model parameters can be used to generate the sub-block based motion prediction signal.

[0299] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2500, and 2600, the determination of the spatial gradient of the corresponding sub-block may include: calculating at least one gradient value for each corresponding sample position in the corresponding sub-block. For example, the calculation of the at least one gradient value for each corresponding sample position in the corresponding sub-block may include: applying a gradient filter to the corresponding sample position in the corresponding sub-block for each corresponding sample position. As another example, the calculation of the at least one gradient value for each corresponding sample position in the corresponding sub-block may include: determining the intensity change for each corresponding sample position in the corresponding sub-block according to the optical flow equation.

[0300] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining a set of motion vector differences associated with the sample positions of the corresponding sub-block. For example, by using the affine motion model of the current block, the sub-block based motion prediction signal may be generated and the set of motion vector differences may be determined.

[0301] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the set of motion vector differences may be determined for the corresponding sub-blocks of the current block, and the motion prediction refinement signals for the other remaining sub-blocks of the current block may be determined using the set of motion vector differences.

[0302] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the determination of the spatial gradient of the corresponding sub-block may include: calculating the spatial gradient using any of the following: (1) a vertical Sobel filter; (2) a horizontal Sobel filter; and / or (3) a 3-tap filter.

[0303] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, integer motion compensation may be used for neighboring and surrounding neighboring reference samples of the corresponding sub-block.

[0304] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the spatial gradient of the corresponding sub-block may include any of the following: a horizontal gradient or a vertical gradient. For example, the horizontal gradient may be calculated as the luminance difference or chrominance difference between the right adjacent sample of the corresponding sample and the left adjacent sample of the corresponding sample; and / or the vertical gradient may be calculated as the luminance difference or chrominance difference between the bottom adjacent sample of the corresponding sample and the top adjacent sample of the corresponding sample.

[0305] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, any of the following may be used to generate the sub-block-based motion prediction signal: (1) a 4-parameter affine model; (2) a 6-parameter affine model; (3) sub-block-based temporal motion vector prediction (SbTMVP) mode motion compensation; or (4) regression-based motion compensation. For example, under the condition of performing SbTMVP mode motion compensation, the method may include: estimating affine model parameters using a linear regression operation with the sub-block motion vector field; and / or using the estimated affine model parameters to derive pixel-level motion vectors. As another example, under the condition of performing motion compensation based on the RMVF mode, the method may include: estimating affine model parameters; and / or using the estimated affine model parameters to derive pixel-level motion vector offsets from the sub-block-level motion vectors. The pixel motion vector offsets may be relative to the center of the corresponding sub-block.

[0306] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the refined motion prediction signal for the respective sub-block may be based on the actual center position of the respective sub-block or may be based on a sample position closest to the actual center of the respective sub-block.

[0307] For example, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 selecting one of the following as the center position associated with the motion prediction vector for each respective sub-block: (1) the actual center of each respective sub-block, or (2) a sample position closest to the actual center of the respective sub-block. The refined motion prediction signal may be based on the selected center position of each sub-block.

[0308] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining a center position associated with the chrominance pixels of the respective sub-block; and determining an offset to the center position of the chrominance pixels of the respective sub-block based on the chrominance position sample type associated with the chrominance pixels. The refined motion prediction signal for the respective sub-block may be based on the actual position of the sub-block, which corresponds to the determined center position of the chrominance pixels adjusted by the offset.

[0309] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the encoder 100 or 300 may generate and transmit information indicating whether prediction refinement using optical flow (PROF) is enabled in one of the following: (1) the sequence parameter set (SPS) header, (2) the picture parameter set (PPS) header, or (3) the tile group header, and / or the decoder 200 or 500 may receive information indicating whether PROF is enabled in one of the following: (1) the SPS header, (2) the PPS header, or (3) the tile group header.

[0310] Figure 27 is a flowchart showing a tenth representative encoding and / or decoding method.

[0311] Reference Figure 27, a representative method 2700 for encoding and / or decoding video may include: at block 2710, encoder 100 or 300 and / or decoder 200 or 500 determine the actual center position of each respective sub-block of the current block. At block 2720, encoder 100 or 300 and / or decoder 200 or 500 may use the actual center position of each respective sub-block of the current block to generate a sub-block-based motion prediction signal or a refined motion prediction signal. At block 2730, (1) encoder 100 or 300 may use the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block to encode the video, or (2) decoder 200 or 500 may use the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block to decode the video. In certain embodiments, the operations at blocks 2710, 2720, and 2730 may be performed for at least one block in the video (e.g., the current block). For example, the determination of the actual center position of each respective sub-block of the current block may include: based on the chroma position sample type of the chroma pixels, determining the chroma center position associated with the chroma pixels of the respective sub-block and the offset of the chroma center position relative to the center position of the respective sub-block. The sub-block-based motion prediction signal or the refined motion prediction signal for the respective sub-block may be based on the actual center position of the respective sub-block, which corresponds to the determined chroma center position adjusted by the offset. Although the actual center of each respective sub-block of the current block is described as being determined / used for various operations, it is contemplated that one, some, or all of the center positions of such sub-blocks may be determined / used.

[0312] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the generation of the refined motion prediction signal may use the sub-block-based motion prediction signal by: for each respective sub-block of the current block, determining one or more spatial gradients of the sub-block-based motion prediction signal, determining a motion prediction refinement signal for the current block based on the determined spatial gradients, and / or combining the sub-block-based motion prediction signal with the motion prediction refinement signal to produce the refined motion prediction signal for the current block. For example, the determination of the one or more spatial gradients of the sub-block-based motion prediction signal may include: using the sub-block-based motion prediction signal and neighboring reference samples adjacent to and surrounding the respective sub-block to determine an extended sub-block, and / or using the determined extended sub-block to determine the one or more spatial gradients of the respective sub-block.

[0313] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the determination of the spatial gradient of the corresponding sub-block may include: calculating at least one gradient value for each corresponding sample position in the corresponding sub-block. For example, the calculating of the at least one gradient value for each corresponding sample position in the corresponding sub-block may include: for each corresponding sample position, applying a gradient filter to the corresponding sample position in the corresponding sub-block.

[0314] As another example, the calculating of the at least one gradient value for each corresponding sample position in the corresponding sub-block may include: determining an intensity change for one or more corresponding sample positions in the corresponding sub-block according to an optical flow equation.

[0315] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, these methods may include an encoder 100 or 300 and / or a decoder 200 or 500 determining a set of motion vector differences associated with the sample positions of the corresponding sub-block. By using the affine motion model of the current block, a sub-block-based motion prediction signal may be generated, and the set of motion vector differences may be determined. In certain examples, the set of motion vector differences may be determined for the corresponding sub-block of the current block, and the set of motion vector differences may be used (e.g., reused) to determine a motion prediction refinement signal for this sub-block and other remaining sub-blocks of the current block. For example, the determination of the spatial gradient of the corresponding sub-block may include calculating the spatial gradient using any of the following: (1) a vertical Sobel filter; (2) a horizontal Sobel filter; and / or (3) a 3-tap filter. Adjacent and neighboring reference samples around the corresponding sub-block may use integer motion compensation.

[0316] In some embodiments, the spatial gradient of the corresponding sub-block may include any of the following: a horizontal gradient or a vertical gradient. For example, the horizontal gradient may be calculated as the luminance difference or chrominance difference between the right adjacent sample and the left adjacent sample of the corresponding sample. As another example, the vertical gradient may be calculated as the luminance difference or chrominance difference between the bottom adjacent sample and the top adjacent sample of the corresponding sample.

[0317] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, any of the following may be used to generate the sub-block based motion prediction signal: (1) a 4-parameter affine model; (2) a 6-parameter affine model; (3) sub-block based temporal motion vector prediction (SbTMVP) mode motion compensation; and / or (4) regression based motion compensation. For example, under the condition of performing SbTMVP mode motion compensation, the method may include: using a sub-block motion vector field to estimate affine model parameters through a linear regression operation; and / or using the estimated affine model parameters to derive pixel-level motion vectors. As another example, under the condition of performing motion compensation in a regression based motion vector field (RMVF) mode, the method may include: estimating affine model parameters; and / or using the estimated affine model parameters to derive a pixel-level motion vector offset from sub-block level motion vectors, where the pixel motion vector offset is relative to the center of the corresponding sub-block.

[0318] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, multiple motion vectors associated with the control points of the current block may be used to generate the refined motion prediction signal.

[0319] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the encoder 100 or 300 may generate, encode, and send an information in one of the following, and the decoder 200 or 500 may receive and decode the information in one of the following, where the information indicates whether prediction refinement using optical flow (PROF) is enabled: (1) sequence parameter set (SPS) header, (2) picture parameter set (PPS) header, or (3) tile group header.

[0320] Figure 28 is a flowchart showing an eleventh representative encoding and / or decoding method.

[0321] Reference Figure 28, A representative method 2800 for encoding and / or decoding video may include: At block 2810, encoder 100 or 300 and / or decoder 200 or 500 selects one of the following as the center position associated with the motion prediction vector of each corresponding sub-block: (1) the actual center of each corresponding sub-block, or (2) the sample position closest to the actual center of the corresponding sub-block. At block 2820, encoder 100 or 300 and / or decoder 200 or 500 may determine the selected center position of each corresponding sub-block of the current block. At block 2830, encoder 100 or 300 and / or decoder 200 or 500 may use the selected center position of each corresponding sub-block of the current block to generate a sub-block-based motion prediction signal or a refined motion prediction signal. At block 2840, (1) encoder 100 or 300 may use the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction of the current block to encode the video, or (2) decoder 200 or 500 may use the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction of the current block to decode the video. In some embodiments, the operations at blocks 2810, 2820, 2830, and 2840 may be performed for at least one block (e.g., the current block) in the video. Although the selection of the center position is described with respect to each corresponding sub-block of the current block, it is contemplated that one, a portion, or all of the center positions of such sub-blocks may be selected / used in various operations.

[0322] Figure 29 is a flowchart showing a representative encoding method.

[0323] See Figure 29 , A representative method 2900 for encoding video may include: At block 2910, encoder 100 or 300 performs motion estimation on the current block of the video, which includes determining the affine motion model parameters of the current block using iterative motion compensation operations and generating a sub-block-based motion prediction signal for the current block using the determined affine motion model parameters. At block 2920, after performing motion estimation on the current block, encoder 100 or 300 may perform a prediction refinement using optical flow (PROF) operation to generate a refined motion prediction signal. At block 2930, encoder 100 or 300 may use the refined motion prediction signal as a prediction of the current block to encode the video. For example, the PROF operation may include: determining one or more spatial gradients of the sub-block-based motion prediction signal; determining a motion prediction refinement signal for the current block based on the determined spatial gradients; and / or combining the sub-block-based motion prediction signal and the motion prediction refinement signal to produce a refined motion prediction signal for the current block.

[0324] In certain representative embodiments that include at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, and 2900, the PROF operation may be performed after (e.g., only after) the iterative motion compensation operation is completed. For example, during the motion estimation of the current block, the PROF operation is not performed.

[0325] Figure 30 is a flowchart showing another representative coding method.

[0326] Refer to Figure 30 , a representative method 3000 for encoding video may include: at block 3010, encoder 100 or 300, during the motion estimation of the current block, uses an iterative motion compensation operation to determine affine motion model parameters and generates a sub-block-based motion prediction signal using the determined affine motion model parameters. At block 3020, after the motion estimation of the current block, under the condition that the size of the current block meets or exceeds a threshold size, encoder 100 or 300 may perform a prediction refinement (PROF) operation using optical flow to generate a refined motion prediction signal. At block 3030, encoder 100 or 300 may encode the video by: (1) using the refined motion prediction signal as a prediction for the current block under the condition that the current block meets or exceeds the threshold size, or (2) using the sub-block-based motion prediction signal as a prediction for the current block under the condition that the current block does not meet the threshold size.

[0327] Figure 31 is a flowchart showing the twelfth representative coding / decoding method.

[0328] Refer to Figure 31, A representative method 3100 for encoding and / or decoding video may include: At block 3110, encoder 100 or 300 determines or obtains information indicating the size of the current block, or decoder 200 or 500 receives information indicating the size of the current block. At block 3120, encoder 100 or 300 or decoder 200 or 500 may generate a sub-block based motion prediction signal. At block 3130, under the condition that the size of the current block meets or exceeds a threshold size, encoder 100 or 300 or decoder 200 or 500 may perform a prediction refinement using optical flow (PROF) operation to generate a refined motion prediction signal. At block 3140, encoder 100 or 300 may encode the video by: (1) using the refined motion prediction signal as a prediction for the current block under the condition that the current block meets or exceeds the threshold size, or (2) using the sub-block based motion prediction signal as a prediction for the current block under the condition that the current block does not meet the threshold size, or decoder 200 or 500 may decode the video by: (1) using the refined motion prediction signal as a prediction for the current block under the condition that the current block meets or exceeds the threshold size, or (2) using the sub-block based motion prediction signal as a prediction for the current block under the condition that the current block does not meet the threshold size.

[0329] Figure 32 is a flowchart showing a thirteenth representative encoding / decoding method.

[0330] Reference Figure 32, a representative method 3200 for encoding and / or decoding video may include: at block 3210, the encoder 100 or 300 determines whether to perform pixel-level motion compensation, or the decoder 200 or 500 receives a flag indicating whether to perform pixel-level motion compensation. At block 3220, the encoder 100 or 300 or the decoder 200 or 500 may generate a sub-block-based motion prediction signal. At block 3230, under the condition that the pixel-level motion compensation is to be performed, the encoder 100 or 300 or the decoder 200 or 500 may: determine one or more spatial gradients of the sub-block-based motion prediction signal, determine a motion prediction refinement signal for the current block based on the determined spatial gradients, and combine the sub-block-based motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the current block. At block 3240, depending on whether the pixel-level motion compensation is to be performed, the encoder 100 or 300 may use the sub-block-based motion prediction signal or the refined motion prediction signal as a prediction for the current block to encode the video, or the decoder 200 or 500 may use the sub-block-based motion prediction signal or the refined motion prediction signal as a prediction for the current block to decode the video according to the indication of the flag. In some embodiments, the operations at blocks 3220 and 3230 may be performed for a block (e.g., the current block) in the video.

[0331] Figure 33 is a flowchart showing a fourteenth representative encoding / decoding method.

[0332] Reference Figure 33, a representative method 3300 for encoding and / or decoding video may include: At block 3310, encoder 100 or 300 determines or obtains or decoder 200 or 500 receives indication of inter-frame prediction weight information that indicates one or more weights associated with first and second reference pictures. At block 3320, encoder 100 or 300 or decoder 200 or 500 may generate, for a current block of the video, a sub-block-based motion inter-frame prediction signal, may determine a first set of spatial gradients associated with the first reference picture and a second set of spatial gradients associated with the second reference picture, may determine a motion inter-frame prediction refinement signal for the current block based on the first set of spatial gradients, the second set of spatial gradients, and the inter-frame prediction weight information, and may combine the sub-block-based motion inter-frame prediction signal and the motion inter-frame prediction refinement signal to produce a refined motion inter-frame prediction signal for the current block. At block 3330, encoder 100 or 300 may encode the video using the refined motion inter-frame prediction signal as a prediction for the current block, or decoder 200 or 500 may decode the video using the refined motion inter-frame prediction signal as a prediction for the current block. For example, the inter-frame prediction weight information is any of the following: (1) an indicator that indicates a first weighting factor to be applied to the first reference picture and / or a second weighting factor to be applied to the second reference picture; or (2) a weight index. In some embodiments, the motion inter-frame prediction refinement signal for the current block may be based on: (1) a first gradient value derived from the first set of spatial gradients and weighted according to a first weight factor indicated by the inter-frame prediction weight information, and (2) a second gradient value derived from the second set of spatial gradients and weighted according to a second weight factor indicated by the inter-frame prediction weight information.

[0333] Example network for implementation of embodiments

[0334] Figure 34AFIG. is a schematic diagram of an exemplary communication system 3400 that can implement one or more of the disclosed embodiments. The communication system 3400 can be a multi-access system that provides content such as voice, data, video, messaging, broadcasting, etc. to a plurality of wireless users. The communication system 3400 can enable a plurality of wireless users to access such content by sharing system resources including wireless bandwidth. For example, the communication system 3400 can use one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT-spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, and filter bank multicarrier (FBMC), etc.

[0335] As Figure 34A shown, the communication system 3400 can include wireless transmit / receive units (WTRUs) 3402a, 3402b, 3402c, 3402d, RAN 3404 / 3413, CN 3406 / 3415, public switched telephone network (PSTN) 3408, Internet 3410, and other networks 3412. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network components. Each of the WTRUs 3402a, 3402b, 3402c, 3402d can be any type of device configured to operate and / or communicate in a wireless environment. For example, any one of the WTRUs 3402a, 3402b, 3402c, 3402d can be referred to as a "station" and / or "STA", which can be configured to transmit and / or receive wireless signals, and can include a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smart phone, a laptop computer, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable device, a head-mounted display (HMD), a vehicle, a drone, a medical device and application (such as remote surgery), an industrial device and application (such as a robot and / or other wireless devices operating in an industrial and / or automated processing chain environment), a consumer electronic device, and a device operating on a commercial and / or industrial wireless network, etc. Any one of the WTRUs 3402a, 3402b, 3402c, 3402d can be interchangeably referred to as a UE.

[0336] The communication system 3400 may also include base station 3414a and / or base station 3414b. Each of base stations 3414a, 3414b may be any type of device configured to facilitate access to one or more communication networks (such as CN 3406 / 3415, Internet 3410, and / or other networks 3412) by wirelessly docking with at least one of WTRUs 3402a, 3402b, 3402c, 3402d in a wireless manner. For example, base stations 3414a, 3414b may be a base transceiver station (BTS), Node B, eNode B (terminal), home Node B (HNB), home eNode B (HeNB), gNB, NR Node B, site controller, access point (AP), and wireless router, etc. Although each of base stations 3414a, 3414b is described as a single component, it should be understood that base stations 3414a, 3414b may include any number of interconnected base stations and / or network components.

[0337] Base station 3414a may be part of RAN 3404 / 3413, and the RAN may also include other base stations and / or network components (not shown), such as a base station controller (BSC), radio network controller (RNC), relay node, etc. Base station 3414a and / or base station 3414b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies in a cell (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide wireless service coverage for a relatively fixed or possibly time-varying specific geographical area. A cell may be further divided into cell sectors. For example, the cell associated with base station 3414a may be divided into three sectors. Thus, in one embodiment, base station 3414a may include three transceivers, that is, each transceiver corresponds to a sector of the cell. In an embodiment, base station 3414a may use multiple-input multiple-output (MIMO) technology and may use multiple transceivers for each sector of the cell. For example, by using beamforming, signals may be transmitted and / or received in a desired spatial direction.

[0338] Base stations 3414a, 3414b may communicate with one or more of WTRUs 3402a, 3402b, 3402c, 3402d via air interface 3416, where the air interface may be any suitable wireless communication link (such as radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 3416 may be established using any suitable radio access technology (RAT).

[0339] More specifically, as described above, the communication system 3400 can be a multi-access system and can use one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA, etc. For example, the base station 3414a in the RAN 3404 / 3413 and the WTRUs 3402a, 3402b, 3402c can implement a certain radio technology, such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), where the technology can use Wideband CDMA (WCDMA) to establish the air interfaces 3415 / 3416 / 3417. WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).

[0340] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c can implement a certain radio technology, such as Evolved UMTS Terrestrial Radio Access (E-UTRA), where the technology can use Long-Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-A Pro) to establish the air interface 3416.

[0341] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c can implement a certain radio technology that can use New Radio (NR) to establish the air interface 3416, such as NR radio access.

[0342] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c can implement multiple radio access technologies. For example, the base station 3414a and the WTRUs 3402a, 3402b, 3402c can jointly implement LTE radio access and NR radio access (e.g., using the Dual Connectivity (DC) principle). Thus, the air interfaces used by the WTRUs 3402a, 3402b, 3402c can be characterized by multiple types of radio access technologies and / or transmissions to / from multiple types of base stations (e.g., terminals and gNBs).

[0343] In other embodiments, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement the following radio technologies, such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN), and so on.

[0344] Figure 34A The base station 3414b in may be, for example, a wireless router, a home Node B, a home eNode B, or an access point, and may use any suitable RAT to facilitate wireless connections in a local area, such as business premises, residences, vehicles, campuses, industrial facilities, air corridors (e.g., for drones), and roads, and so on. In one embodiment, the base station 3414b and the WTRUs 3402c, 3402d may establish a Wireless Local Area Network (WLAN) by implementing a radio technology such as IEEE 802.11. In an embodiment, the base station 3414b and the WTRUs 3402c, 3402d may establish a Wireless Personal Area Network (WPAN) by implementing a radio technology such as IEEE 802.15. In yet another embodiment, the base station 3414b and the WTRUs 3402c, 3402d may establish a pico cell or a femto cell by using a cellular-based RAT (such as WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, and so on). As Figure 34A shown, the base station 3414b may be directly connected to the Internet 3410. Thus, the base station 3414b does not need to access the Internet 3410 via the CN 3406 / 3415.

[0345] RAN 3404 / 3413 can communicate with CN 3406 / 3415, and the CN can be any type of network configured to provide voice, data, applications, and / or voice over Internet Protocol (VoIP) services to one or more of WTRU 3402a, 3402b, 3402c, 3402d. The data can have different quality of service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements, etc. CN 3406 / 3415 can provide call control, accounting services, mobile location-based services, prepaid calls, Internet connection, video distribution, etc., and / or can perform advanced security functions such as user authentication. Although not shown in Figure 34A , it should be understood that RAN 1084 / 3413 and / or CN 3406 / 3415 can communicate directly or indirectly with other RANs that use the same or different radio access technologies (RATs) as RAN 3404 / 3413. For example, in addition to being connected to RAN 3404 / 3413 that uses NR radio technology, CN 3406 / 3415 can also communicate with other RANs (not shown) that use GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.

[0346] CN 3406 / 3415 can also act as a gateway for WTRU 3402a, 3402b, 3402c, 3402d to access the public switched telephone network (PSTN) 3408, the Internet 3410, and / or other networks 3412. The PSTN 3408 can include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 3410 can include a global interconnected computer network device system that uses common communication protocols (such as TCP, UDP, and / or IP in the Transmission Control Protocol / Internet Protocol (TCP / IP) Internet protocol family). The network 3412 can include a wired or wireless communication network owned and / or operated by other service providers. For example, the network 3412 can include another CN connected to one or more RANs, where the one or more RANs can use the same or different RATs as RAN 3404 / 3413.

[0347] Some or all of the WTRU 3402a, 3402b, 3402c, 3402d in the communication system 3400 can include multi-mode capabilities (e.g., WTRU 3402a, 3402b, 3402c, 3402d can include multiple transceivers for communicating with different wireless networks on different wireless links). For example, Figure 34AThe illustrated WTRU 3402c can be configured to communicate with a base station 3414a using a cellular-based radio technology and with a base station 3414b that can use IEEE 802 radio technology.

[0348] Figure 34B is a system diagram showing an exemplary WTRU 3402. As Figure 34B shown, the WTRU 3402 can include a processor 3418, a transceiver 3420, a transmit / receive component 3422, a speaker / microphone 3424, a keyboard 3426, a display / touchpad 3428, a non-removable memory 3430, a removable memory 3432, a power supply 3434, a global positioning system (GPS) chipset 3436, and / or peripheral devices 3438. It should be understood that the WTRU 3402 can also include any sub-combination of the foregoing components while remaining compliant with the embodiments.

[0349] The processor 3418 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. The processor 3418 can perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable the WTRU 3402 to operate in a wireless environment. The processor 3418 can be coupled to the transceiver 3420, and the transceiver 3420 can be coupled to the transmit / receive component 3422. Although Figure 34B the processor 3418 and the transceiver 3420 are described as separate components, it should be understood that the processor 3418 and the transceiver 3420 can also be integrated together in an electronic component or chip. The processor 3418 can be configured to encode or decode video (e.g., video frames).

[0350] The transmit / receive component 3422 can be configured to transmit or receive signals to or from a base station (e.g., base station 3414a) via an air interface 3416. For example, in one embodiment, the transmit / receive component 3422 can be an antenna configured to transmit and / or receive RF signals. As an example, in another embodiment, the transmit / receive component 3422 can be a radiator / detector configured to transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, the transmit / receive component 3422 can be configured to transmit and / or receive RF and optical signals. It should be understood that the transmit / receive component 3422 can be configured to transmit and / or receive any combination of wireless signals.

[0351] Although inFigure 34B The transmit / receive component 3422 is described as a single component, but the WTRU 3402 can include any number of transmit / receive components 3422. More specifically, the WTRU 3402 can utilize MIMO technology. Thus, in one embodiment, the WTRU 3402 can include two or more transmit / receive components 3422 (e.g., multiple antennas) that transmit and receive wireless signals via the air interface 3416.

[0352] The transceiver 3420 can be configured to modulate the signals to be transmitted by the transmit / receive component 3422 and to demodulate the signals received by the transmit / receive component 3422. As described above, the WTRU 3402 can have multi-mode capabilities. Accordingly, the transceiver 3420 can include multiple transceivers that allow the WTRU 3402 to communicate via multiple RATs (e.g., NR and IEEE 802.11).

[0353] The processor 3418 of the WTRU 3402 can be coupled to the speaker / microphone 3424, the numeric keypad 3426, and / or the display / touchpad 3428 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit) and can receive user input data from these components. The processor 3418 can also output user data to the speaker / microphone 3424, the keypad 3426, and / or the display / touchpad 3428. In addition, the processor 3418 can access information from and store information in any suitable memory such as non-removable memory 3430 and / or removable memory 3432. The non-removable memory 3430 can include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 3432 can include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and so on. In other embodiments, the processor 3418 can access information from and store data in memories that are not actually located in the WTRU 3402. By way of example, such memories can be located on a server or a home computer (not shown).

[0354] The processor 3418 can receive power from the power supply 3434 and can be configured to distribute and / or control the power for other components in the WTRU 3402. The power supply 3434 can be any suitable device for powering the WTRU 3402. For example, the power supply 3434 can include one or more dry battery packs (such as nickel cadmium (Ni-Cd), nickel zinc (Ni-Zn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), solar cells, and fuel cells, among others.

[0355] The processor 3418 may also be coupled to a GPS chipset 3436, which may be configured to provide location information (e.g., longitude and latitude) related to the current location of the WTRU 3402. As a supplement or replacement to the information from the GPS chipset 3436, the WTRU 3402 may receive location information from a base station (e.g., base stations 3414a, 3414b) via the air interface 3416, and / or determine its location based on the signal timing received from two or more nearby base stations. It should be understood that the WTRU 3402 may obtain location information by means of any suitable positioning method while remaining compliant with the embodiments.

[0356] The processor 3418 may also be coupled to other peripheral devices 3438, where the peripheral devices may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connections. For example, the peripheral devices 3438 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, modules, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game console module, an Internet browser, a virtual reality and / or augmented reality (VR / AR) device, and an activity tracker, among others. The peripheral devices 3438 may include one or more sensors, which may be one or more of the following: a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geographical location sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor, etc.

[0357] The processor 3418 of the WTRU 3402 may operably communicate with various peripheral devices 3438, which include, for example, any one of the following: the one or more accelerometers, the one or more gyroscopes, the USB port, other communication interfaces / ports, the display, and / or other video / audio indicators, to implement the representative embodiments disclosed herein.

[0358] The WTRU 3402 may include a full-duplex radio device, for which the reception or transmission of some or all signals (e.g., associated with a particular subframe for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio device may include an interference management unit that reduces and / or substantially eliminates self-interference either by means of hardware (e.g., a choke coil) or by signal processing of a processor (e.g., a separate processor (not shown) or by means of processor 3418). In an embodiment, the WTRU 3402 may include a half-duplex radio device that transmits and receives some or all signals (e.g., associated with a particular subframe for UL (e.g., for transmission) or downlink (e.g., for reception)).

[0359] Figure 34C FIG. is a system diagram showing the RAN 3404 and the CN 3406 according to an embodiment. As described above, the RAN 3404 may communicate with the WTRU 3402a, 3402b, 3402c using E-UTRA radio technology over the air interface 3416. The RAN 3404 may also communicate with the CN 3406.

[0360] The RAN 3404 may include eNodeBs 3460a, 3460b, 3460c. However, it should be understood that the RAN 3404 may include any number of eNodeBs while remaining compliant with the embodiment. Each of the eNodeBs 3460a, 3460b, 3460c may include one or more transceivers that communicate with the WTRU 3402a, 3402b, 3402c over the air interface 3416. In one embodiment, the eNodeBs 3460a, 3460b, 3460c may implement MIMO technology. Thus, for example, the eNodeB 3460a may use multiple antennas to transmit wireless signals to the WTRU 3402a and / or receive wireless signals from the WTRU 3402a.

[0361] Each of the eNodeBs 3460a, 3460b, 3460c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, etc. As Figure 34C shown, the eNodeBs 3460a, 3460b, 3460c may communicate with each other via the X2 interface.

[0362] Figure 34CThe illustrated CN 3406 may include a Mobility Management Entity (MME) 3462, a Serving Gateway (SGW) 3464, and a Packet Data Network (PDN) Gateway (or PGW) 3466. Although each of the foregoing components is described as being part of the CN 3406, it should be understood that any of these components may be owned and / or operated by an entity other than the CN operator.

[0363] The MME 3462 may be connected to each of the eNodeBs 3460a, 3460b, 3460c in the RAN 3404 via the S1 interface and may act as a control node. For example, the MME 3462 may be responsible for authenticating users of the WTRUs 3402a, 3402b, 3402c, performing bearer activation / deactivation procedures, and selecting a specific serving gateway during the initial attachment process of the WTRUs 3402a, 3402b, 3402c, etc. The MME 3462 may provide control plane functions for handovers between the RAN 3404 and other RANs (not shown) using other radio technologies, such as GSM and / or WCDMA.

[0364] The SGW 3464 may be connected to each of the eNodeBs 3460a, 3460b, 3460c in the RAN 3404 via the S1 interface. The SGW 3464 may typically route and forward user data packets to / from the WTRUs 3402a, 3402b, 3402c. Also, the SGW 3464 may perform other functions, such as anchoring the user plane during handover between eNBs, triggering paging procedures when DL data is available for the WTRUs 3402a, 3402b, 3402c, and managing and storing the context of the WTRUs 3402a, 3402b, 3402c, etc.

[0365] The SGW 3464 may be connected to the PGW 146, which may provide access to a packet switched network (such as the Internet 3410) for the WTRUs 3402a, 3402b, 3402c to facilitate communication between the WTRUs 3402a, 3402b, 3402c and IP-enabled devices.

[0366] CN 3406 can facilitate communication with other networks. For example, CN 3406 can provide the WTRUs 3402a, 3402b, 3402c with access to a circuit-switched network (such as the PSTN 3408) in order to facilitate communication between the WTRUs 3402a, 3402b, 3402c and traditional landline communication devices. For example, CN 3406 can include or communicate with an IP gateway (such as an IP Multimedia Subsystem (IMS) server), and this IP gateway can act as an interface between CN 3406 and the PSTN 3408. In addition, CN 3406 can provide the WTRUs 3402a, 3402b, 3402c with access to the other network 3412, where this network can include other wired and / or wireless networks owned and / or operated by other service providers.

[0367] Although the WTRU is described as a wireless terminal in Figures 34A - 34D it should be appreciated that in some representative embodiments, such terminals and the communication network can use (e.g., temporarily or permanently) a wired communication interface.

[0368] In a representative embodiment, the other network 3412 can be a WLAN.

[0369] A WLAN using the infrastructure basic service set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can access or interface with a distributed system (DS) or another type of wired / wireless network that sends traffic into and / or out of the BSS. Traffic originating from outside the BSS and going to an STA can reach and be delivered to the STA through the AP. Traffic originating from an STA and going to a destination outside the BSS can be sent to the AP for delivery to the corresponding destination. Traffic between STAs within the BSS can be sent through the AP, for example, in a case where the source STA can send traffic to the AP and the AP can deliver the traffic to the destination STA. Traffic between STAs within the BSS can be considered and / or referred to as point-to-point traffic. The point-to-point traffic can be sent using a direct link setup (DLS) between the source and destination STAs (e.g., directly therebetween). In some representative embodiments, the DLS can use 802.11e DLS or 802.11z channelized DLS (TDLS). For example, a WLAN using the independent BSS (IBSS) mode does not have an AP, and STAs within or using the IBSS (such as all STAs) can communicate directly with each other. Here, the IBSS communication mode can sometimes be referred to as an "Ad-hoc" communication mode.

[0370] When operating in an 802.11ac infrastructure mode or a similar mode, the AP can transmit beacons on a fixed channel (e.g., the primary channel). The primary channel can have a fixed width (e.g., a bandwidth of 20 MHz) or a width dynamically set via signaling. The primary channel can be the operating channel of the BSS and can be used by the STA to establish a connection with the AP. In some representative embodiments, Carrier Sense Multiple Access with Collision Avoidance (CSMA / CA) (e.g., in an 802.11 system) can be implemented. For CSMA / CA, STAs including the AP (e.g., each STA) can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, then the particular STA can back off. In a specified BSS, at any given time, there is one STA (e.g., only one station) transmitting.

[0371] High Throughput (HT) STAs can use a channel with a width of 40 MHz for communication (e.g., by combining a 20-MHz primary channel with an adjacent or non-adjacent 20-MHz channel to form a 40-MHz channel).

[0372] Very High Throughput (VHT) STAs can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40-MHz and / or 80-MHz channels can be formed by combining consecutive 20-MHz channels. A 160-MHz channel can be formed by combining eight consecutive 20-MHz channels or by combining two non-consecutive 80-MHz channels (this combination can be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, the data can be passed and go through a segmentation parser, which can split the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing can be performed separately on each stream. The streams can be mapped on two 80-MHz channels, and the data can be transmitted by the STA performing the transmission. On the receiver of the STA performing the reception, the above operations for the 80+80 configuration can be reversed, and the combined data can be sent to the Media Access Control (MAC).

[0373] 802.11af and 802.11ah support operating modes below 1 GHz. Compared with 802.11n and 802.11ac, the channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV white space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah can support meter type control / machine type communication (MTC) (e.g., MTC devices in a macro coverage area). The MTC device may have certain capabilities, such as limited capabilities including supporting (e.g., only supporting) certain and / or limited bandwidths. The MTC device may include a battery, and the battery life of the battery is higher than a threshold (e.g., for maintaining a long battery life).

[0374] For WLAN systems (such as 802.11n, 802.11ac, 802.11af, and 802.11ah) that can support multiple channels and channel bandwidths, these systems include channels that can be designated as the primary channel. The bandwidth of the primary channel can be equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or restricted by a certain STA, where the STA is from all STAs operating in the BSS that support the minimum bandwidth operating mode. In an example regarding 802.11ah, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes, for an STA that supports (e.g., only supports) the 1 MHz mode (such as an MTC type device), the width of the primary channel can be 1 MHz. Carrier sensing and / or network allocation vector (NAV) settings can depend on the state of the primary channel. If the primary channel is busy (e.g., because an STA (which only supports the 1 MHz operating mode) transmits to the AP), then the entire available frequency band can be considered busy even if most of the available frequency band remains idle and available for use.

[0375] In the United States, the available frequency band for 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. According to the country code, the total bandwidth available for 802.11ah is from 6 MHz to 26 MHz.

[0376] Figure 34DFIG. 0 is a system diagram showing RAN 3413 and CN 3415 according to an embodiment. As described above, RAN 3413 can communicate with WTRU 3402a, 3402b, 3402c using NR radio technology via air interface 3416. RAN 3413 can also communicate with CN 3415.

[0377] RAN 3413 may include gNBs 3480a, 3480b, 3480c, but it should be understood that RAN 3413 may include any number of gNBs while remaining compliant with the embodiment. Each of gNBs 3480a, 3480b, 3480c may include one or more transceivers to communicate with WTRU 3402a, 3402b, 3402c via air interface 3416. In one embodiment, gNBs 3480a, 3480b, 3480c may implement MIMO technology. For example, gNBs 3480a, 3480b may use beamforming processing to transmit and / or receive signals to and / or from gNBs 3480a, 3480b, 3480c. Thus, for example, gNB 3480a may use multiple antennas to transmit wireless signals to WTRU 3402a and receive wireless signals from WTRU 3402a. In an embodiment, gNBs 3480a, 3480b, 3480c may implement carrier aggregation technology. For example, gNB 3480a may transmit multiple component carriers to WTRU 3402a (not shown). A subset of these component carriers may be on unlicensed spectrum while the remaining component carriers may be on licensed spectrum. In an embodiment, gNBs 3480a, 3480b, 3480c may implement coordinated multipoint (CoMP) technology. For example, WTRU 3402a may receive a coordinated transmission from gNB 3480a and gNB 3480b (and / or gNB 3480c).

[0378] WTRU 3402a, 3402b, 3402c may communicate with gNBs 3480a, 3480b, 3480c using transmissions associated with a scalable digital configuration. For example, the OFDM symbol interval and / or the OFDM subcarrier interval may be different for different transmissions, different cells, and / or different wireless transmission spectrum portions. WTRU 3402a, 3402b, 3402c may communicate with gNBs 3480a, 3480b, 3480c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., containing different numbers of OFDM symbols and / or lasting different absolute time lengths).

[0379] gNBs 3480a, 3480b, 3480c can be configured to communicate with WTRUs 3402a, 3402b, 3402c operating in a Standalone configuration and / or a Non-Standalone configuration. In the Standalone configuration, WTRUs 3402a, 3402b, 3402c can communicate with gNBs 3480a, 3480b, 3480c without accessing other RANs (e.g., eNodeBs 3460a, 3460b, 3460c). In the Standalone configuration, WTRUs 3402a, 3402b, 3402c can use one or more of gNBs 3480a, 3480b, 3480c as a mobility anchor. In the Standalone configuration, WTRUs 3402a, 3402b, 3402c can use signals in the unlicensed band to communicate with gNBs 3480a, 3480b, 3480c. In the Non-Standalone configuration, WTRUs 3402a, 3402b, 3402c communicate / connect with gNBs 3480a, 3480b, 3480c while communicating / connecting with another RAN (e.g., eNodeBs 3460a, 3460b, 3460c). For example, WTRUs 3402a, 3402b, 3402c can communicate with one or more gNBs 3480a, 3480b, 3480c and one or more eNodeBs 3460a, 3460b, 3460c in a substantially simultaneous manner by implementing the DC principle. In the Non-Standalone configuration, eNodeBs 3460a, 3460b, 3460c can act as the mobility anchor for WTRUs 3402a, 3402b, 3402c, and gNBs 3480a, 3480b, 3480c can provide additional coverage and / or throughput to serve WTRUs 3402a, 3402b, 3402c.

[0380] Each of gNBs 3480a, 3480b, 3480c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support network slicing, dual connectivity, implement interworking between NR and E-UTRA, route user plane data to user plane functions (UPFs) 3484a, 3484b, and route control plane information to access and mobility management functions (AMFs) 3482a, 3482b, etc. As Figure 34D shown, gNBs 3480a, 3480b, 3480c can communicate with each other via the Xn interface.

[0381] Figure 34DThe illustrated CN 3415 may include at least one AMF 3482a, 3482b, at least one UPF 3484a, 3484b, at least one session management function (SMF) 3483a, 3483b, and may possibly include data networks (DN) 3485a, 3485b. Although each of the foregoing components has been described as part of CN 3415, it should be understood that any of these components may be owned and / or operated by entities other than the CN operator.

[0382] AMF 3482a, 3482b may be connected to one or more of gNB 3480a, 3480b, 3480c in RAN 3413 via the N2 interface and may act as a control node. For example, AMF 3482a, 3482b may be responsible for authenticating users of WTRU 3402a, 3402b, 3402c, supporting network slicing (e.g., handling different protocol data unit (PDU) sessions with different requirements), selecting a particular SMF 3483a, 3483b, managing the registration area, terminating non-access stratum (NAS) signaling, and mobility management, etc. AMF 3482a, 3482b may use network slicing processing in order to customize the CN support provided to WTRU 3402a, 3402b, 3402c based on the type of service used by WTRU 3402a, 3402b, 3402c. As an example, for different use cases, different network slices may be established, such as services relying on ultra-reliable low-latency communication (URLLC) access, services relying on enhanced mobile (e.g., massive mobile) broadband (eMBB) access, and / or services for machine-type communication (MTC) access, etc. AMF 3462 may provide control plane functions for handover between RAN 3413 and other RANs (not shown) using other radio technologies (e.g., LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi).

[0383] SMF 3483a, 3483b can be connected to AMF 3482a, 3482b in CN 3415 via N11 interface. SMF3483a, 3483b can also be connected to UPF 3484a, 3484b in CN 3415 via N4 interface. SMF 3483a, 3483b can select and control UPF 3484a, 3484b, and can configure traffic routing through UPF 3484a, 3484b. SMF3483a, 3483b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy implementation and QoS, and providing downlink data notification, etc. The PDU session type can be IP-based, non-IP-based, Ethernet-based, etc.

[0384] UPF 3484a, 3484b can be connected to one or more gNBs 3480a, 3480b, 3480c in RAN 3413 via the N3 interface, which can provide WTRU 3402a, 3402b, 3402c with access to packet-switched networks (such as the Internet 3410) to facilitate communication between WTRU 3402a, 3402b, 3402c and IP-enabled devices. UPF 3484, 3484b can perform other functions such as routing and forwarding packets, implementing user plane policies, supporting multi-host PDU sessions, processing user plane QoS, buffering downlink packets, and providing mobility anchor processing, etc.

[0385] The CN 3415 may facilitate communications with other networks. For example, the CN 3415 may include or may communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 3415 and the PSTN 3408. In addition, the CN 3415 may provide the WTRUs 3402a, 3402b, 3402c with access to other networks 3412, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 3402a, 3402b, 3402c may be connected to the local data networks (DNs) 3485a, 3485b via the N3 interface connected to the UPFs 3484a, 3484b and the N6 interface between the UPFs 3484a, 3484b and the DNs 3485a, 3485b and through the UPFs 3484a, 3484b.

[0386] In view of Figures 34A - 34D and about Figures 34A - 34DThe corresponding description, where one or more or all of the functions corresponding to one or more of the following descriptions can be performed by one or more emulation devices (not shown): WTRU 3402a-d, base station 3414a-b, eNode B 3460a-c, MME 3462, SGW 3464, PGW 3466, gNB 3480a-c, AMF 3482a-b, UPF 3484a-b, SMF 3483a-b, DN 3485a-b, and / or any one or more of the other devices described herein. These emulation devices can be one or more devices configured to simulate one or more or all of the functions described herein. For example, these emulation devices can be used to test other devices and / or simulate network and / or WTRU functions.

[0387] The emulation devices can be designed to perform one or more tests on other devices in a laboratory environment and / or an operator network environment. For example, the one or more emulation devices can perform one or more or all of the functions while being implemented and / or deployed, in whole or in part, as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more emulation devices can perform one or more or all of the functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation devices can be directly coupled to other devices to perform tests, and / or can use over-the-air wireless communication to perform tests.

[0388] One or more emulation devices can perform one or more functions, including all functions, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation device can be used in a test laboratory and / or a test scenario of a wired and / or wireless communication network that is not deployed (e.g., for testing) to perform tests on one or more components. The one or more emulation devices can be test devices. The emulation devices can transmit and / or receive data using direct RF coupling and / or wireless communication via an RF circuit (e.g., the circuit can include one or more antennas).

[0389] Compared with the previous-generation video coding standard H.264 / MPEG AVC, the HEVC standard provides approximately 50% bitrate savings for equivalent perceptual quality. Although the HEVC standard provides significant coding improvements compared to its predecessor, additional coding efficiency improvements can be achieved with additional coding tools. The Joint Video Exploration Team (JVET) initiated a project to develop a new generation of video coding standard (referred to as Versatile Video Coding (VVC)) to provide such coding efficiency improvements, and a reference software codebase called the VVC Test Model (VTM) was established for demonstrating a reference implementation of the VVC standard. To facilitate the evaluation of new coding tools, another reference software library called the Benchmark Set (BMS) was also generated. In the BMS codebase, a list of additional coding tools that provide higher coding efficiency and moderate implementation complexity is included on top of the VTM and is used as a benchmark when evaluating similar coding techniques during the VVC standardization process. In addition to the JEM coding tools integrated in BMS-2.0 (e.g., 4×4 non-separable second-order transform (NSST), generalized bi-prediction (GBi), bi-directional optical flow (BIO), decoder-side motion vector refinement (DMVR), and current picture reference (CPR)), it also includes trellis-coded quantization tools.

[0390] A system and method for processing data according to a representative embodiment can be executed by one or more processors that execute a sequence of instructions contained in a storage device. These instructions can be read into the storage device from other computer-readable media such as auxiliary data storage device(s). Execution of the sequence of instructions contained in the storage device causes the processor to operate as described above, for example. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions to implement one or more embodiments. Such software can run on a processor that is remotely housed within a robotic assist / device (RAA) and / or another mobile device. In the latter case, data can be transmitted between the RAA or other mobile device containing sensors and the remote device containing the processor that runs software performing ratio estimation and compensation as described above, either via wired or wireless means. According to other representative embodiments, some of the processing described above with respect to localization can be performed in a device containing sensors / cameras, while the remaining processing can be performed in a second device after receiving the partially processed data from the device containing the sensors / cameras.

[0391] Although the features and elements are described above in terms of specific combinations, one of ordinary skill in the art will appreciate that each feature or element can be used alone or in any combination with other features and elements. Additionally, the methods described herein can be implemented in a computer program, software, or firmware executed by a computer or processor and embedded in a computer-readable medium. Examples of non-transitory computer-readable media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memories, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital versatile discs (DVDs). A processor associated with software can be used to implement a radio frequency transceiver used in a WTRU 3402, UE, terminal, base station, RNC, or any host computer.

[0392] In addition, in the above-described embodiments, reference was made to a processing platform, a computing system, a controller, and other devices that include a processor. These devices can include at least one central processing unit (“CPU”) and a memory. In accordance with the practice of those skilled in the art of computer programming, references to symbolic descriptions of acts and operations or instructions can be performed by various CPUs and memories. These acts and operations or instructions can be referred to as “being executed,” “computer-executed,” or “CPU-executed.”

[0393] One of ordinary skill in the art will appreciate that acts and symbolic descriptions of operations or instructions include manipulation of electrical signals by a CPU. An electrical system representation can identify data bits that cause a transformation or reduction of an electrical signal and the maintenance of the storage location of the data bits in a storage system thereby to reconfigure or otherwise alter the operation of the CPU and other processing of the signal. The maintenance of the storage location of the data bits is by having a specific electrical, magnetic, optical, or organic property corresponding to or representing the data bits. It should be understood that the representative embodiments are not limited to the above-described platforms or CPUs and that other platforms and CPUs can support the provided methods.

[0394] The data bits can also be maintained on a computer-readable medium that includes magnetic disks, optical disks, and any other large storage system that is CPU-readable, whether volatile (e.g., random access memory (“RAM”)) or non-volatile (e.g., read-only memory (“ROM”)). The computer-readable medium can include cooperative or interconnected computer-readable media that are specifically present on a processor system or distributed among multiple interconnected processing systems that can be local or remote to the processing system. It can be understood that the representative embodiments are not limited to the above-described memories and that other platforms and memories can support the described methods. It should be understood that the representative embodiments are not limited to the above-described platforms or CPUs, and other platforms and CPUs can also support the provided methods.

[0395] In the illustrated embodiments, any of the operations, processes, etc. described herein may be implemented as computer-readable instructions stored on a computer-readable medium. The computer-readable instructions may be executed by a processor of a mobile unit, a network element, and / or any other computing device.

[0396] There is a difference between hardware and software implementations in terms of the system. The use of hardware or software is generally (but not always, as the choice between hardware and software can be important in some environments) a design choice that takes into account the cost-efficiency trade-off. There can be various tools (e.g., hardware, software, and / or firmware) that can affect the processes and / or systems and / or other technologies described herein, and the preferred tool can vary depending on the context of the process and / or system and / or other technology being deployed. For example, if the implementer determines that speed and accuracy are the most important, the implementer can choose mainly hardware and / or firmware tools. If flexibility is the most important, the implementer can choose mainly software implementations. Alternatively, the implementer can choose some combination of hardware, software, and / or firmware.

[0397] The foregoing detailed description has presented various embodiments of the apparatus and / or process by using block diagrams, flowcharts, and / or examples. To the extent that these block diagrams, flowcharts, and / or examples contain one or more functions and / or operations, those skilled in the art will appreciate that each function and / or operation within these block diagrams, flowcharts, or examples can be implemented individually and / or together by a wide range of hardware, software, or firmware, or substantially any combination thereof. Suitable processors include, for example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs); field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.

[0398] While the features and elements are provided above in particular combinations, those skilled in the art will appreciate that each feature or element can be used alone or in combination with other features and elements. The present disclosure is not limited to the specific embodiments described in this application, which are intended to be examples of various aspects. Many modifications and variations can be made without departing from its essence and scope, which are known to those skilled in the art. The elements, acts, or instructions used in the description of this application should not be construed as critical or essential to the embodiments unless explicitly stated. In addition to the methods and apparatuses enumerated herein, those skilled in the art will also know, based on the above description, methods and apparatuses that are functionally equivalent within the scope of the present disclosure. These modifications and variations should also fall within the scope of the appended claims. The present disclosure is limited only by the appended claims, including their full scope of equivalents. It should be understood that the present disclosure is not limited to a particular method or system.

[0399] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, when reference is made to the terms "station" and its abbreviation "STA", "user equipment" and its abbreviation "UE", it can mean: (i) a wireless transmit and / or receive unit (WTRU), such as those described below; (ii) any of a plurality of embodiments of a WTRU, such as those described below; (iii) a wireless and / or wired (e.g., wirelessly communicable) device configured with some or all of the structures and functions of a WTRU, such as those described below; (iii) a device having wireless capabilities and / or wired capabilities configured with structures and functions having less than all of the structures and functions of a WTRU, such as those described below; or (iv) the like. Details of an exemplary WTRU are provided below, which exemplary WTRU can represent any UE described herein. Figures 34A - 34D Details of an exemplary WTRU are provided below, which exemplary WTRU can represent any UE described herein.

[0400] In some representative embodiments, some portions of the subject matter described herein may be implemented via application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, those skilled in the art will appreciate that some aspects of the embodiments disclosed herein, in whole or in part, may equivalently be implemented by integrated circuits, as one or more computer programs running on one or more computers (e.g., one or more programs running on one or more computer systems), one or more programs running on one or more processors (e.g., one or more programs running on one or more microprocessors), firmware, or substantially any combination thereof, and designing circuits and / or writing code for the software and / or firmware in accordance with the present disclosure is known to those skilled in the art. Further, those skilled in the art will appreciate that the mechanisms of the subject matter described herein may be distributed in a variety of forms of program products, and that the exemplary embodiments of the subject matter described herein apply, regardless of the particular type of signal bearing medium used to actually effect such distribution. Examples of signal bearing media include, but are not limited to, the following: recordable type media such as floppy disks, hard disk drives, CDs, DVDs, digital tapes, computer memories, etc., and transmission type media such as digital and / or analog communication media (e.g., optical fibers, waveguides, wired communication links, wireless communication links, etc.).

[0401] The subject matter described herein is sometimes shown with different components that are included in or connected to different other components. It will be appreciated that these depicted architectures are merely examples, and that many other architectures may be implemented in practice that perform the same functionality. Conceptually, any arrangement of components that perform the same functionality effectively “associates” therewith such that the desired functionality may be implemented. Thus, any two components that are combined to perform a particular function may be viewed as being “associated” with each other such that the desired functionality is implemented, regardless of the architecture or intervening components. Similarly, any two components that are associated may also be viewed as being “operably connected” or “operably coupled” to each other to perform the desired functionality, and any two components that are capable of being so associated may also be viewed as being “operably couplable” to each other to perform the desired functionality. Specific examples of operably couplable include, but are not limited to, components that are physically mateable and / or physically interactive and / or wirelessly interactive and / or wirelessly interactable and / or logically interactive and / or logically interactable.

[0402] Regarding the use of substantially any plural and / or singular terms herein, those skilled in the art may translate from the plural to the singular and / or from the singular to the plural as appropriate to the context and / or application. For clarity, various singular / plural permutations may be explicitly set forth herein.

[0403] Those skilled in the art can understand that generally the terms used herein and especially the terms used in the claims (e.g., the main part of the claims) are generally "open" terms (e.g., the term "comprising" should be understood as "comprising but not limited to", the term "having" should be understood as "having at least", the term "including" should be understood as "including but not limited to", etc.). Those skilled in the art can also understand that if a claim is to describe a specific quantity, it will be explicitly described in the claim, and in the absence of such a description, there is no such meaning. For example, if only one item is to be indicated, the term "single" or similar language can be used. To assist understanding, the following claims and / or the description herein may include the use of the introductory phrase "at least one" or "one or more" to introduce the claim description. However, the use of these phrases should not be understood as implying that a claim description introduced by the indefinite article "a" will limit any particular claim including such an introduced claim description to an embodiment including only one such description, even when the same claim includes the introductory phrase "one or more" or "at least one" and the indefinite article (e.g., "a" should be understood as meaning "at least one" or "one or more"). The same is true for the use of the definite article to introduce a claim description. In addition, even if the specific quantity of the introduced claim description is explicitly described, those skilled in the art can understand that such a description should be understood as indicating at least the described quantity (e.g., simply describing "two descriptions" without other modifiers means at least two descriptions, or two or more descriptions). In addition, in these examples using conventions such as "at least one of A, B, and C, etc.", generally such conventions are understood by those skilled in the art (e.g., "The system has at least one of A, B, and C" can include but is not limited to the system having only A, only B, only C, A and B, A and C, B and C, and / or A, B, and C, etc.). In these examples using conventions such as "at least one of A, B, or C, etc.", generally such conventions are understood by those skilled in the art (e.g., "The system has at least one of A, B, or C" can include but is not limited to the system having only A, only B, only C, A and B, A and C, B and C, and / or A, B, and C, etc.). Those skilled in the art can also understand that substantially any separated words and / or phrases representing two or more alternative items, whether in the specification, the claims, or the drawings, should be understood as including the possibility of including one of the two items, either one, or both items. For example, the phrase "A or B" is understood to include the possibility of "A" or "B" or "A" and "B". In addition, the term "any" used herein followed by a list of multiple items and / or multiple types of items is intended to include "any", "any combination", "any number", and / or "any combination of numbers" of the multiple items and / or multiple types of items, alone or in combination with other items and / or other types of items.In addition, the terms "set" or "group" as used herein are intended to include any number of items, including zero. Further, the term "number" as used herein is intended to include any number, including zero.

[0404] In addition, if the features or aspects of the present disclosure are described in terms of a Markush group, those skilled in the art will appreciate that the present disclosure can also be described in terms of any individual member or subgroup of members of the Markush group.

[0405] Those skilled in the art will appreciate that, for any and all purposes, e.g., to provide a written description, all ranges disclosed herein also include any and all possible subranges and combinations thereof. Any listed range can be readily understood to describe and enable the same range being broken into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range described herein can be readily broken into a lower third, middle third, and upper third, etc. Those skilled in the art will also appreciate that all language such as "up to," "at least," "greater than," "less than," etc. including the recited numbers and ranges can then be broken into the above-described subranges. Finally, those skilled in the art will appreciate that ranges include each individual member. Thus, for example, a group and / or set having 1 - 3 cells refers to a group / set having 1, 2, or 3 cells. Similarly, a group / set having 1 - 5 cells refers to a group / set having 1, 2, 3, 4, or 5 cells, and so on.

[0406] In addition, the claims should not be construed as limited to the recited order or elements unless so described. Further, the use of the term "means for" in any claim is intended to invoke 35 U.S.C.§112, or the means - plus - function claim format, and any claim that does not have the term "means for" is not so intended.

[0407] A processor associated with software can be used to implement a radio frequency transceiver used in a wireless transmit / receive unit (WTRU), user equipment (UE), terminal, base station, mobility management entity (MME), or evolved packet core (EPC), or any host computer. The WTRU can be combined with modules implemented in hardware and / or software (including software - defined radio (SDR)) and other components such as, for example, a camera, video camera module, video phone, walkie - talkie, vibrating device, speaker, microphone, television transceiver, hands - free headset, keyboard, Module, frequency modulation (FM) radio unit, near field communication (NFC) module, liquid crystal display (LCD) display unit, organic light emitting diode (OLED) display unit, digital music player, media player, video game console module, Internet browser, and / or any wireless local area network (WLAN) or ultra-wideband (UWB) module.

[0408] Throughout the disclosure, those skilled in the art will understand that certain representative embodiments may alternatively or in combination with other representative embodiments be used.

[0409] Additionally, the methods described herein may be embodied in a computer program, software, or firmware incorporated in a computer-readable medium and executed by a computer or processor. Examples of non-transitory computer-readable media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital versatile discs (DVDs). A processor associated with the software may be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, and any host computer.

Claims

1. A method for decoding a video, the method comprising: For the current block of the video: Generate a sub-block based motion prediction signal for sub-blocks of the current block based on an affine motion model associated with the current block, Determine a set of pixel-level motion vector differences for the sub-blocks using the affine motion model associated with the current block, wherein the pixel-level motion vector difference for a sample is based on the difference between the position of the sample and the center position of the sub-block, Determine the spatial gradient of the sub-block based motion prediction signal for each sample position of the sub-block, Determine a motion prediction refinement signal for the sub-blocks of the current block based on the determined set of pixel-level motion vector differences and the determined spatial gradient, and Combine the sub-block based motion prediction signal and the motion prediction refinement signal to produce a refined motion prediction signal for the sub-blocks of the current block; And Decode the video using the refined motion prediction signal as a prediction for the current block.

2. A method for encoding a video, the method comprising: For the current block of the video: Generate a sub-block based motion prediction signal for sub-blocks of the current block based on an affine motion model associated with the current block, Determine a set of pixel-level motion vector differences for the sub-blocks using the affine motion model associated with the current block, wherein the pixel-level motion vector difference for a sample is based on the difference between the position of the sample and the center position of the sub-block, Determine the spatial gradient of the sub-block based motion prediction signal for each sample position of the sub-block, Determine a motion prediction refinement signal for the sub-blocks of the current block based on the determined set of motion vector differences and the determined spatial gradient, and Combine the sub-block based motion prediction signal and the motion prediction refinement signal to produce a refined motion prediction signal for the sub-blocks of the current block; And Encode the video using the refined motion prediction signal as a prediction for the current block.

3. The method according to claim 1 or 2, wherein determining the spatial gradient of the sub-block-based motion prediction signal comprises: For one or more corresponding sub-blocks of the current block: Use the sub-block based motion prediction signal and neighboring reference samples adjacent to and surrounding the corresponding sub-block to determine an extended sub-block; and Use the determined extended sub-block to determine the spatial gradient of the corresponding sub-block to determine the motion prediction refinement signal.

4. The method according to claim 1 or 2, wherein for sub-blocks of the current block, the set of pixel-level motion vector differences is determined and used to determine the motion prediction refinement signal for one or more further sub-blocks of the current block.

5. The method according to claim 1 or 2, further comprising determining affine motion model parameters for the current block of the video such that the sub-block-based motion prediction signal is generated by using the determined affine motion model parameters.

6. A decoder configured to decode a video, comprising: A processor, configured to: For the current block of the video: Generate a sub-block based motion prediction signal for sub-blocks of the current block based on an affine motion model associated with the current block, Determine a set of pixel-level motion vector differences for the sub-blocks using the affine motion model associated with the current block, wherein the pixel-level motion vector difference for a sample is based on the difference between the position of the sample and the center position of the sub-block, Determine the spatial gradient of the sub-block based motion prediction signal for each sample position of the sub-block, Determine a motion prediction refinement signal for the current block based on the determined set of pixel-level motion vector differences and the determined spatial gradient, and Combine the sub-block based motion prediction signal and the motion prediction refinement signal to produce a refined motion prediction signal for the sub-blocks of the current block;And Decode the video using the refined motion prediction signal as a prediction for the current block.

7. An encoder configured to encode video, comprising: A processor, configured to: For the current block of the video: Generate a sub-block based motion prediction signal for sub-blocks of a current block based on an affine motion model associated with the current block. Determine a set of pixel-level motion vector differences for the sub-blocks using the affine motion model associated with the current block, wherein the pixel-level motion vector difference for a sample is based on the difference between the position of the sample and the center position of the sub-block. Determine a spatial gradient of the sub-block based motion prediction signal for each sample position of the sub-block. Determine a motion prediction refinement signal for the sub-blocks of the current block based on the determined set of pixel-level motion vector differences, and Combine the sub-block based motion prediction signal and the motion prediction refinement signal to produce a refined motion prediction signal for the sub-blocks of the current block; And Encode the video using the refined motion prediction signal as a prediction for the current block.

8. The decoder according to claim 6 or the encoder according to claim 7, wherein the processor is configured to: For one or more corresponding sub - blocks of the current block: Determine an extended sub - block using the sub - block - based motion prediction signal and neighboring reference samples adjacent to and surrounding the corresponding sub - block; and Use the determined extended sub - block to determine the spatial gradient of the corresponding sub - block to determine the motion prediction refinement signal.

9. The decoder according to claim 6 or the encoder according to claim 7, wherein the processor is configured to determine the set of pixel - level motion vector differences of the sub - blocks of the current block, which are used to determine the motion prediction refinement signal for one or more further sub - blocks of the current block.

10. The decoder according to claim 6 or the encoder according to claim 7, wherein the processor is configured to determine the affine motion model parameters of the current block of the video such that the sub - block - based motion prediction signal is generated by using the determined affine motion model parameters.

11. A non - transitory computer - readable medium having instructions that, when executed by a computer, implement the method according to any one of claims 1 - 3.