Systems, apparatus, and methods for inter-prediction refinement using optical flow

Inter-prediction refinement using optical flow models addresses inefficiencies in video coding by deriving refined motion vectors, improving compression performance and video quality through enhanced motion compensation.

JP7801515B2Active Publication Date: 2026-01-16INTERDIGITAL VC HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025035590
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-04-15
Filing Date
2025-03-06
Publication Date
2026-01-16
Estimated Expiration
2040-02-04

AI Technical Summary

Technical Problem

Existing video coding systems face inefficiencies in removing temporal redundancy due to rapid illumination changes and limitations in block-based motion compensation, leading to suboptimal prediction techniques that do not adequately compensate for illumination variations.

Method used

Implementing inter-prediction refinement using optical flow models, such as bidirectional optical flow (BIO) and affine motion compensation, to derive refined motion vectors for each sample within a block, enhancing the efficiency of motion-compensated prediction by applying per-sample motion refinement.

Benefits of technology

Improves the accuracy and efficiency of video coding by reducing residual motion within blocks, thereby enhancing the compression performance and quality of video signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007801515000049
    Figure 0007801515000049
  • Figure 0007801515000050
    Figure 0007801515000050
  • Figure 0007801515000051
    Figure 0007801515000051
Patent Text Reader

Abstract

To provide methods, apparatuses and systems.SOLUTION: In one aspect, a decoding method includes obtaining a sub-block based motion prediction signal for a current block of video; obtaining one or more spatial gradients of the sub-block based motion prediction signal or one or more motion vector difference values; obtaining a refinement signal for the current block based on the one or more obtained spatial gradients or the one or more obtained motion vector difference values; obtaining a refined motion prediction signal for the current block based on the sub-block based motion prediction signal and the refinement signal; and decoding the current block based on the refined motion prediction signal.SELECTED DRAWING: Figure 18A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to video coding, and more particularly to systems, apparatus, and methods that use inter-prediction refinement using optical flow. [Background technology]

[0002] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 802,428, filed February 7, 2019, U.S. Provisional Patent Application No. 62 / 814,611, filed March 6, 2019, and U.S. Provisional Patent Application No. 62 / 883,999, filed April 15, 2019, the contents of each of which are incorporated herein by reference.

[0003] Prior art Video coding systems are widely used to compress digital video signals to reduce the storage and / or transmission bandwidth of such signals. Among various types of video coding systems, such as block-based, wavelet-based, and object-based systems, block-based hybrid video coding systems are currently the most widely used and deployed. Examples of block-based video coding systems include international video coding standards such as MPEG1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, and the latest video coding standard called High Efficiency Video Coding (HEVC), developed by the Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG. Summary of the Invention

[0004] In one exemplary embodiment, a decoding method includes obtaining a sub-block-based motion prediction signal for a current block of a video, obtaining one or more spatial gradients or one or more motion vector differential values ​​of the sub-block-based motion prediction signal, obtaining a refinement signal for the current block based on the one or more obtained spatial gradients or the one or more obtained motion vector differential values, obtaining a refined motion prediction signal for the current block based on the sub-block-based motion prediction signal and the refinement signal, and decoding the current block based on the refined motion prediction signal. Various other embodiments are also disclosed herein.

[0005] A more detailed understanding can be had from the following detailed description, given in conjunction with the drawings attached hereto, by way of example. The figures in the description are examples. Therefore, the figures and detailed description should not be considered limiting, as other equally effective examples are possible and anticipated. Furthermore, like reference numerals in the figures indicate like elements. [Brief explanation of the drawings]

[0006] [Figure 1] FIG. 1 is a block diagram illustrating a typical block-based video encoding system. [Figure 2] FIG. 1 is a block diagram illustrating a typical block-based video decoder. [Figure 3] 1 is a block diagram illustrating an exemplary block-based video encoder with generalized bi-prediction (GBi) support. [Figure 4] FIG. 1 illustrates a typical GBi module for an encoder. [Figure 5] FIG. 1 illustrates an exemplary block-based video decoder with GBi support. [Figure 6]FIG. 1 illustrates a typical GBi module for a decoder. [Figure 7] FIG. 1 illustrates a typical bidirectional optical flow. [Figure 8A] FIG. 1 illustrates a typical four-parameter affine mode. [Figure 8B] FIG. 1 illustrates a typical four-parameter affine mode. [Figure 9] FIG. 1 illustrates a typical six-parameter affine mode. [Figure 10] FIG. 1 illustrates a typical interweaved prediction procedure. [Figure 11] FIG. 10 illustrates exemplary weight values ​​(eg, associated with pixels) in a sub-block. [Figure 12] 10 is a diagram illustrating an example of an area where interweave prediction is applied and another area where interweave prediction is not applied. [Figure 13A] FIG. 1 illustrates an example of SbTMVP processing. [Figure 13B] FIG. 1 illustrates an example of SbTMVP processing. [Figure 14] FIG. 10 illustrates adjacent motion blocks (e.g., 4x4 motion blocks) that may be used for motion parameter derivation. [Figure 15] FIG. 10 illustrates adjacent motion blocks that can be used for motion parameter derivation. [Figure 16] FIG. 10 illustrates sub-block MVs and pixel-level MV differences Δv(i,j) after sub-block-based affine motion compensation prediction. [Figure 17A] FIG. 1 illustrates an exemplary procedure for determining the MV corresponding to the actual center of a sub-block. [Figure 17B] FIG. 10 is a diagram illustrating the locations of chrominance samples in a 4:2:0 chrominance format. [Figure 17C] FIG. 10 illustrates an example of an extended prediction sub-block. [Figure 18A] 1 is a flowchart illustrating a first exemplary encoding / decoding method. [Figure 18B] 10 is a flowchart illustrating a second exemplary encoding / decoding method. [Figure 19] 10 is a flowchart illustrating a third exemplary encoding / decoding method. [Figure 20] 10 is a flowchart illustrating a fourth exemplary encoding / decoding method. [Figure 21] 10 is a flowchart illustrating a fifth exemplary encoding / decoding method. [Figure 22] 10 is a flowchart illustrating a sixth exemplary encoding / decoding method. [Figure 23] 10 is a flowchart illustrating a seventh exemplary encoding / decoding method. [Figure 24] 13 is a flowchart illustrating an eighth exemplary encoding / decoding method. [Figure 25] 1 is a flowchart illustrating an exemplary gradient calculation method. [Figure 26] 13 is a flowchart illustrating a ninth exemplary encoding / decoding method. [Figure 27] 13 is a flowchart illustrating a tenth exemplary encoding / decoding method. [Figure 28] 13 is a flowchart illustrating an eleventh exemplary encoding / decoding method. [Figure 29] 1 is a flowchart illustrating an exemplary encoding method. [Figure 30] 10 is a flowchart illustrating another exemplary encoding method. [Figure 31] 16 is a flowchart illustrating a twelfth exemplary encoding / decoding method. [Figure 32] 13 is a flowchart illustrating a thirteenth exemplary encoding / decoding method. [Figure 33]14 is a flowchart illustrating a fourteenth exemplary encoding / decoding method. [Figure 34A] FIG. 1 is a system diagram illustrating an example communication system in which one or more disclosed aspects may be implemented. [Figure 34B] FIG. 34B is a system diagram illustrating an example wireless transmit / receive unit (WTRU) that may be used within the communication system illustrated in FIG. 34A according to an embodiment. [Figure 34C] A system diagram illustrating an example radio access network (RAN) and an example core network (CN) that may be used within the communication system shown in FIG. 34A according to an embodiment. [Figure 34D] FIG. 34B is a system diagram illustrating a further exemplary RAN and a further exemplary CN that may be used within the communications system shown in FIG. 34A according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0007] Block-based hybrid video coding procedure Like HEVC, VVC is built on a block-based hybrid video coding framework.

[0008] FIG. 1 is a block diagram illustrating a typical block-based hybrid video encoding system.

[0009] Referring to FIG. 1, an encoder 100 may be provided with an input video signal 102 that is processed block by block (called a coding unit (CU)) and may be used to efficiently compress high-resolution (1080p and above) video signals. In HEVC, a CU may be up to 64×64 pixels. A CU may be further divided into prediction units or PUs, to which separate prediction procedures may be applied. For each input video block (MB and / or CU), spatial prediction 160 and / or temporal prediction 162 may be performed. Spatial prediction (or “intra prediction”) may predict a current video block using pixels from already-encoded neighboring blocks within the same video picture / slice.

[0010] Spatial prediction can reduce spatial redundancy inherent in a video signal. Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") predicts a current video block using pixels from an already-encoded video picture. Temporal prediction can reduce temporal redundancy inherent in a video signal. A temporal prediction signal for a given video block may (e.g., typically) be signaled by one or more motion vectors (MVs), which may indicate the amount and / or direction of motion between a current block (CU) and its reference block.

[0011] If multiple reference pictures are supported (as in recent video coding standards such as H.264 / AVC or HEVC), for each video block, a reference picture index for the video block may be transmitted (e.g., may be transmitted additionally), and / or the reference index may be used to identify which reference picture in reference picture store 164 the temporal prediction signal comes from. After spatial prediction and / or temporal prediction, mode decision block 180 in encoder 100 may choose the best prediction mode based, for example, on a rate-distortion optimization method / procedure. A prediction block of either spatial prediction 160 or temporal prediction 162 may be subtracted from the current video block 116, and / or the prediction residual may be decorrelated using transform 104 and quantization 106 to achieve a target bitrate. The quantized residual coefficients may be inverse quantized 110 and inverse transformed 112 to form a reconstructed residual, which may be added back to the prediction block at 126 to form a reconstructed video block. In-loop filtering 166, such as a deblocking filter and / or an adaptive loop filter, may be applied to the reconstructed video block, which may then be placed in a reference picture store 164 and used to encode future video blocks. To form the output video bitstream 120, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients may be transmitted (e.g., transmitted all together) to the entropy coding unit 108 and further compressed and / or packed to form the bitstream.

[0012] Those skilled in the art will appreciate that the encoder 100 may be implemented using a processor, memory, and transmitter providing the various elements / modules / units disclosed above. For example, the transmitter may send the bitstream 120 to a decoder, and the processor may be configured to execute software to enable receiving the input video 102 and performing the functions associated with the various blocks of the encoder 100.

[0013] FIG. 2 is a block diagram illustrating a block-based video decoder.

[0014] Referring to FIG. 2, video decoder 200 may be provided with a video bitstream 202, which may be unpacked and entropy decoded in entropy decoding unit 208. Coding mode and prediction information may be sent to appropriate one of spatial prediction unit 260 (for intra-coding modes) and / or temporal prediction unit 262 (for inter-coding modes) to form a prediction block. Residual transform coefficients may be sent to inverse quantization unit 210 and inverse transform unit 212 to reconstruct the residual block. The reconstructed block may further pass through in-loop filtering 266 before being stored in reference picture store 264. Reconstructed video 220 may be sent to be saved in reference picture store 264, for example, to drive a display device and for use in predicting future video blocks.

[0015] Decoder 200 may be implemented using a processor, memory, and receiver capable of providing the various elements / modules / units disclosed above. For example, those skilled in the art will understand that (1) the receiver may be configured to receive bitstream 202, and (2) the processor may be configured to execute software to enable reception of bitstream 202 and output of reconstructed video 220, and performance of the functions associated with the various blocks of decoder 200.

[0016] Those skilled in the art will appreciate that many of the functions / operations / processing of a block-based encoder and a block-based decoder are the same.

[0017] In modern video codecs, bidirectional motion compensated prediction (MCP) can be used for its high efficiency in removing temporal redundancy by exploiting the temporal correlation between pictures. A bidirectionally predicted signal can be formed by combining two uni-predictive signals using a weight value equal to 0.5, but this may not be optimal for combining uni-predictive signals, especially in conditions where illumination changes rapidly from one reference picture to another. Specific prediction techniques / operations and / or procedures can be implemented to compensate for illumination variations over time by applying some global / local weights and / or offset values ​​to sample values ​​in reference pictures (e.g., some or each of the sample values ​​in the reference picture).

[0018] The use of bidirectional motion compensated prediction (MCP) in video codecs enables the removal of temporal redundancy by exploiting the temporal correlation between pictures. A bidirectionally predicted signal may be formed by combining two unipredictive signals using a weighting value (e.g., 0.5). In certain videos, illumination characteristics may change rapidly from one reference picture to another. Therefore, prediction techniques may compensate for illumination variations over time (e.g., fading transitions) by applying global or local weights and / or offset values ​​to one or more sample values ​​in the reference pictures.

[0019] Generalized bi-prediction (GBi) may improve MCP for bi-prediction mode. In bi-prediction mode, the predicted signal at a given sample x may be calculated by Equation 1 as follows:

[0020] P[x]=w0*P0[x+v0]+w1*P1[x+v1] (1) In the above equation, P[x] may indicate the resulting prediction signal of sample x located at picture position x. Pi[x+vi] may be a motion-compensated prediction signal of x using motion vector (MV) vi for the i-th list (e.g., list 0, list 1, etc.). w0 and w1 may be two weight values ​​shared across (e.g., all) samples in a block. Based on this equation, various prediction signals may be obtained by adjusting the weight values ​​w0 and w1. Some configurations of w0 and w1 may mean the same prediction as uni-prediction and bi-prediction. For example, (w0, w1) = (1, 0) may be used for uni-prediction using reference list L0. (w0, w1) = (0, 1) may be used for uni-prediction using reference list L1. (w0, w1) = (0.5, 0.5) may be used for bi-prediction using two reference lists. Weights may be signaled per CU. To reduce signaling overhead, a constraint such as w0+w1=1 may be applied so that one weight may be signaled. Thus, Equation 1 may be further simplified as shown in Equation 2 below.

[0021] P[x]=(1-w1)*P0[x+v0]+w1*P1[x+v1] (2) To further reduce signaling overhead, w1 may be discretized (e.g., −2 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 10 / 8, etc.), such that each weight value may be represented by an index value within a (e.g., small) restricted range.

[0022] FIG. 3 is a block diagram illustrating a representative block-based video encoder with GBi support.

[0023] The encoder 300 may include a mode decision module 304, a spatial prediction module 306, a motion prediction module 308, a transform module 310, a quantization module 312, an inverse quantization module 316, an inverse transform module 318, a loop filter 320, a reference picture store 322, and an entropy coding module 314. Some or all of the encoder's modules or components (e.g., the spatial prediction module 306) may be the same as or similar to those described in connection with FIG. 1. Furthermore, the spatial prediction module 306 and the motion prediction module 308 may be pixel-domain prediction modules. Thus, the input video bitstream 302 may be processed in a manner similar to the input video bitstream 102, but the motion prediction module 308 may further include GBi support. In this manner, the motion prediction module 308 can combine two separate prediction signals in a weighted average manner. Furthermore, the selected weight index may be signaled in the output video bitstream 324.

[0024] Those skilled in the art will appreciate that the encoder 300 may be implemented using a processor, memory, and transmitter providing the various elements / modules / units disclosed above. For example, the transmitter may send the bitstream 324 to a decoder, and the processor may be configured to execute software to enable receiving the input video 302 and performing the functions associated with the various blocks of the encoder 300.

[0025] 4 illustrates a representative GBi estimation module 400 that may be utilized in a motion prediction module of an encoder, such as motion prediction module 308. GBi estimation module 400 may include a weight value estimation module 402 and a motion estimation module 404. Thus, GBi estimation module 400 may utilize a process (e.g., a two-step operation / process) to generate an inter-predicted signal, such as a final inter-predicted signal. Motion estimation module 404 may perform motion estimation by searching for two optimal motion vectors (MVs) that point to (e.g., two) reference blocks using an input video block 401 and one or more reference pictures received from a reference picture store 406. Weight value estimation module 402 may (1) receive the output of motion estimation module 404 (e.g., motion vectors v0 and v1), one or more reference pictures from reference picture store 406, and weight information W, and may (2) search for an optimal weight index to minimize the weighted bidirectional prediction error between the current video block and the bidirectional prediction. It is contemplated that the weight information W may describe a list of available weight values ​​or weight sets, such that the determined weight index and the weight information W may be used together to specify the weights w0 and w1 to be used in GBi. The prediction signal for generalized bidirectional prediction may be calculated as a weighted average of two prediction blocks. The output of the GBi estimation module 400 may include an inter prediction signal, motion vectors v0 and v1, and / or a weight index weight_idx, etc.

[0026] FIG. 5 illustrates a representative block-based video decoder with GBi support, which can decode a bitstream 502 (e.g., from an encoder) that supports GBi, e.g., the bitstream 324 created by the encoder 300 described in connection with FIG. 3. As shown in FIG. 5, the video decoder 500 may include an entropy decoder 504, a spatial prediction module 506, a motion prediction module 508, a reference picture store 510, a dequantization module 512, an inverse transform module 514, and / or a loop filter module 518. Some or all of the decoder's modules may be the same or similar to those described in connection with FIG. 2, except that the motion prediction module 508 may further include GBi support. Thus, the coding mode and prediction information may be used to derive a prediction signal using spatial prediction or an MCP with GBi support. For GBi, block motion information and weight values ​​(e.g., in the form of indices indicating the weight values) may be received and decoded to generate a prediction block.

[0027] Decoder 500 may be implemented using a processor, memory, and receiver capable of providing the various elements / modules / units disclosed above. For example, those skilled in the art will understand that (1) the receiver may be configured to receive bitstream 502, and (2) the processor may be configured to execute software to enable reception of bitstream 502 and output of reconstructed video 520, and performance of the functions associated with the various blocks of decoder 500.

[0028] FIG. 6 illustrates a representative GBi prediction module that may be used in a motion estimation module of a decoder, such as motion estimation module 508.

[0029] Referring to Figure 6, the GBi prediction module may include a weighted average module 602 and a motion compensation module 604, where the motion compensation module 604 may receive one or more reference pictures from a reference picture store 606. The weighted average module 602 may receive the output of the motion compensation module 604, weight information W, and a weight index (e.g., weight_idx). The output of the motion compensation module 604 may include motion information that may correspond to blocks of the picture. The GBi prediction module 600 may use the block motion information and weight values ​​to calculate a prediction signal (e.g., inter prediction signal 608) for GBi as a weighted average of (e.g., two) motion-compensated prediction blocks.

[0030] Representative bidirectional predictive prediction based on optical flow model FIG. 7 is a diagram illustrating a typical bidirectional optical flow.

[0031] 7, bi-predictive prediction can be based on an optical flow model. For example, the prediction associated with the current block (e.g., curblk700) is based on the first predicted block I (0) 702 (e.g., a temporally earlier predicted block shifted by time τ0) and a second predicted block I (1)704 (e.g., a temporally future prediction block shifted by time τ1). Bidirectional prediction in video coding may be a combination of two temporal prediction blocks 702 and 704 obtained from an already reconstructed reference picture. Due to limitations of block-based motion compensation (MC), there may be residual small motion that can be observed between samples of the two prediction blocks, thereby reducing the efficiency of motion-compensated prediction. To reduce the effect of such motion for all samples in a block, bidirectional optical flow (referred to as BIO or BDOF) may be applied. BIO can provide per-sample motion refinement that can be performed in addition to block-based motion-compensated prediction when bidirectional prediction is used. With regard to BIO, the derivation of refined motion vectors for each sample in a block can be based on a classical optical flow model. For example, (k) (x, y) is the sample value at coordinates (x, y) of the prediction block derived from reference picture list k (k=0, 1), and ∂I (k) (x,y) / ∂x and ∂I (k) If (x,y) / ∂y is the horizontal and vertical gradient of the sample, then given the optical flow model, the motion refinement (v x ,v y ) can be derived by the following Equation 3:

[0032]

number

[0033] In FIG. 7, the (MV x0 ,MV y0 ) and (MV x1 ,MV y1 ) are the two predicted blocks I (0) and I (1)The motion refinement (v x ,v y ) may be calculated by minimizing the difference Δ between the values ​​of the samples after motion refinement compensation (eg, A and B in FIG. 7), as shown in Equation 4 below:

[0034]

number

[0035] For example, to ensure the regularity of the derived motion refinement, it is intended that the motion refinement be consistent for samples within one small unit (e.g., a 4x4 block or other small unit). In the benchmark set (BMS)-2.0, (v x ,v y The value of ) is derived by minimizing Δ within a 6×6 window Ω around each 4×4 block, as shown in Equation 5 below.

[0036]

number

[0037] To solve the optimization specified in Equation 5, BIO may use an incremental method / operation / procedure that can optimize the motion refinement horizontally and vertically (e.g., then vertically). This may result in Equations / Inequalities 6 and 7 as follows:

[0038]

number

[0039]

number

[0040] where:

[0041]

number

[0042] can be a floor function that can output the maximum value less than or equal to the input, and th BIO can be a motion refinement threshold, e.g., to prevent error propagation due to coding noise and / or irregular local motion, and 2 18-BD The values ​​of S1, S2, S3, S5 and S6 can be further calculated as shown in Equations 8-12 below. S1=Σ (i,j)∈Ω ψ x (i,j)·ψ x (i,j) (8) S3=Σ (i,j)∈Ω θ(i,j) ψ x (i,j) 2 L (9) S2=Σ (i,j)∈Ω ψ x (i,j)·ψ y (i,j) (10) S5=Σ (i,j)∈Ω ψ y (i,j)·ψ y (i,j)·2 (11) S6=Σ (i,j)∈Ω θ(i,j) ψ y (i,j) 2 L+1 (12) Here, the various gradients can be expressed by the following equations 13 to 15.

[0043]

number

[0044]

number

[0045]

number

[0046] In BMS-2.0, the BIO gradients in Equations 13-15 in both the horizontal and vertical directions can be obtained directly by calculating the difference between two adjacent samples at one sample position of each L0 / L1 prediction block (e.g., horizontally or vertically depending on the direction of the derived gradient), as shown in Equations 16 and 17 below.

[0047]

number

[0048]

number

[0049] k=0,1 In Equations 8-12, L may be the bit depth increase for internal BIO processing / procedures to maintain data precision, which may be set to 5 in BMS-2.0, for example. To avoid partitioning by smaller values, the adjustment parameters r and m in Equations 6 and 7 may be defined as shown in Equations 18 and 19 below. r=500·4 BD-8 (18) m=700·4 BD-8 (19) Here, BD may be the bit depth of the input video. Based on the motion refinement derived by Equations 4 and 5, the final bidirectional predicted signal of the current CU can be calculated by interpolating L0 / L1 prediction samples along the motion trajectory based on the optical flow Equation 3, as specified in the following Equations 20 and 21.

[0050]

number

[0051]

number

[0052] where shift and ο offset may be a right shift and offset that may be applied to combine the L0 and L1 predicted signals for bidirectional prediction, and may be set equal to, for example, 15−BD and 1<<(14−BD)+2·(1<<13), respectively. rnd(·) is a rounding function that may round the input value to the nearest integer value.

[0053] Typical affine modes In HEVC, a translational motion (translational motion only) model is applied for motion compensated prediction. In the real world, many types of motion exist (e.g., zoom in / out, rotation, perspective motion, and other irregular motion). In VVC Test Model (VTM)-2.0, affine motion compensated prediction is applied. Affine motion models are either four-parameter or six-parameter. A first flag for an inter-coded CU is signaled to indicate whether a translational motion model or an affine motion model is applied for inter prediction. If an affine motion model is applied, a second flag is sent to indicate whether the model is a four-parameter model or a six-parameter model.

[0054] The four-parameter affine motion model has two parameters for horizontal and vertical translation, one parameter for zoom motion in both directions, and one parameter for rotation motion in both directions. The horizontal zoom parameter is equal to the vertical zoom parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. The four-parameter affine motion model is coded in the VTM using two motion vectors at two control point positions defined at the top-left corner 810 and top-right corner 820 of the current CU. Other control point positions, such as at other corners and / or edges of the current CU, are also possible.

[0055] Although one affine motion model is described above, other affine models are possible as well and may be used in various embodiments herein.

[0056] 8A and 8B are diagrams illustrating a representative four-parameter affine model and sub-block level motion derivation of an affine block. Referring to FIGS. 8A and 8B, the affine motion field of a block is described by two control point motion vectors at a first control point 810 (the upper left corner of the current block) and a second control point 820 (the upper right corner of the current block), respectively. Based on the control point motion, the motion field (v x ,v y ) can be written as shown in Equations 22 and 23 below.

[0057]

number

[0058]

number

[0059] where (v 0x ,v 0y ) can be the motion vector of the top left corner control point 810, and (v 1x ,v1y ) may be the motion vector of the top right corner control point 820 as shown in FIG. 8A, and w may be the width of the CU. For example, the motion field of an affine coded CU is derived at the 4x4 block level, i.e., (v x ,v y ) is derived for each 4x4 block in the current CU and applied to the corresponding 4x4 block.

[0060] The four parameters of the four-parameter affine model can be estimated iteratively. The MV pair at step k is

[0061]

number

[0062] , the original signal (e.g., luminance signal) is denoted as I(i,j), and the predicted signal (e.g., luminance signal) is denoted as I' k (i,j). The spatial gradient g x (i,j) and g y (i,j) are, for example, the horizontal and / or vertical prediction signals I' k (i,j) can be derived using a Sobel filter applied to (i,j). The derivation of Equation 3 can be expressed as shown in Equations 24 and 25 below.

[0063]

number

[0064] where, in step k, (a, b) may be delta translation parameters, and (c, d) may be delta zoom and rotation parameters. The delta MV at a control point may be derived using its coordinates as shown in Equations 26-29 below. For example, (0, 0) and (w, 0) may be the coordinates of the top-left control point 810 and the top-right control point 820, respectively.

[0065]

number

[0066]

number

[0067] Based on the optical flow equation, the relationship between intensity (e.g., luminance) changes and spatial gradients and temporal movements is formulated in Equation 30 as follows:

[0068]

number

[0069]

number

[0070] and

[0071]

number

[0072] By substituting into Equation 24, Equation 31 for parameters (a, b, c, d) is obtained as follows: I' k (i,j)-I(i,j)=(g x (i,j)*i+g y (i,j)*j)*c+(-g x (i,j)*j+g y (i,j)*i)*d+g x (i,j)*a+g y (i,j)*b (31)

[0073] Since the samples in the CU (e.g., all samples) satisfy Equation 31, the parameter set (e.g., a, b, c, d) can be solved using, for example, the least squares error method. Two control points at step (k+1)

[0074]

number

[0075] The MVs at can be solved using Equations 26-29, which can be rounded to a particular precision (e.g., quarter-pixel precision (pel) or other sub-pixel precision, etc.). Using iterations, the MVs at the two control points can be refined, for example, until convergence (e.g., when the parameters (a, b, c, d) are all zero or when the iteration time reaches a predefined limit).

[0076] FIG. 9 shows a diagram illustrating a representative six-parameter affine mode, where, for example, V0, V1, and V2 are motion vectors at control points 910, 920, and 930, respectively, and (MV x , MV y ) is the motion vector of the sub-block centered at position (x,y).

[0077] 9, an affine motion model (e.g., having six parameters) can have any of (1) a parameter for horizontal translation, (2) a parameter for vertical translation, (3) a parameter for horizontal zoom, (4) a parameter for horizontal rotation, (5) a parameter for vertical zoom, and / or (6) a parameter for vertical rotation. The six-parameter affine motion model can be coded using three MVs at three control points 910, 920, and 930. As shown in FIG. 9, the three control points 910, 920, and 930 for a six-parameter affine-coded CU are defined at the top-left corner, top-right corner, and bottom-left corner of the CU, respectively. The motion at the top-left control point 910 may be associated with a translational motion, the motion at the top-right control point 920 may be associated with a horizontal rotational motion and / or a horizontal zoom motion, and the motion at the bottom-left control point 930 may be associated with a vertical rotational motion and / or a vertical zoom motion. In a six-parameter affine motion model, the horizontal rotational motion and / or zoom motion may not be the same as the same motion in the vertical direction. x ,v y ) may be derived using the three MVs at control points 910, 920, and 930, as shown in Equations 32 and 33 below.

[0078]

number

[0079]

number

[0080] where (v 2x ,v 2y) may be the motion vector V2 of the bottom-left control point 930, (x, y) may be the center position of the sub-block, w may be the width of the CU, and h may be the height of the CU.

[0081] The six parameters of the six-parameter affine model can be estimated in a similar manner. Equations 24 and 25 can be modified as shown in Equations 34 and 35 below.

[0082]

number

[0083] where, in step k, (a, b) may be the delta translation parameters, (c, d) may be the delta zoom and rotation parameters for the horizontal direction, and (e, f) may be the delta zoom and rotation parameters for the vertical direction. Equation 31 may be modified as shown in Equation 36 below. I' k (i,j)-I(i,j)=(g x (i,j)*i)*c+(g x (i,j)*j)*d+(g y (i,j)*i)*e+(g y (i,j)*j)*f+g x (i,j)*a+g y (i,j)*b (36)

[0084] The parameter set (a, b, c, d, e, f) can be solved using, for example, a least squares method / procedure / operation by considering the samples within the CU (e.g., all samples). MV of the top-left control point

[0085]

number

[0086] can be calculated using Equations 26 to 29. MV of the upper right control point

[0087]

number

[0088] can be calculated using equations 37 and 38 as shown below: MV of the bottom left control point

[0089]

number

[0090] can be calculated using equations 39 and 40 as shown below:

[0091]

number

[0092]

number

[0093] Although four and six parameter affine models are shown in Figures 8A, 8B and 9, one skilled in the art will appreciate that affine models with different numbers of parameters and / or different control points are similarly possible.

[0094] Although an affine model is described herein in connection with optical flow refinement, those skilled in the art will appreciate that other motion models in connection with optical flow refinement are possible as well.

[0095] Representative interweave prediction for affine motion compensation. In affine motion compensation (AMC), for example in VTM, a coding block is partitioned into small sub-blocks of about 4x4, each of which may be assigned an individual motion vector (MV) derived by an affine model, for example as shown in Figures 8A and 8B or 9. In a four-parameter or six-parameter affine model, the MV may be derived from the MVs of two or three control points.

[0096] AMC can face a dilemma associated with the size of the sub-blocks: with smaller sub-blocks, AMC may achieve better coding performance, but may suffer from a greater complexity burden.

[0097] FIG. 10 illustrates a representative interweave prediction procedure that can achieve finer granularity of MVs, eg, at the expense of a moderate increase in complexity.

[0098] In FIG. 10 , a coding block 1010 may be partitioned into sub-blocks having two different partitioning patterns (e.g., a first pattern 0 and a second pattern 1). As shown in FIG. 10 , the first partitioning pattern 0 (e.g., a first sub-block pattern, e.g., a 4×4 sub-block pattern) may be the same as that in VTM, and the second partitioning pattern 1 (e.g., an overlapping and / or interwoven second sub-block pattern) may partition the coding block 1010 into 4×4 sub-blocks having a 2×2 offset from the first partitioning pattern 0. AMC using two partitioning patterns (e.g., the first partitioning pattern 0 and the second partitioning pattern 1) may generate several auxiliary predictions (e.g., two auxiliary predictions P0 and P1). The motion vectors of each sub-block in each of partitioning patterns 0 and 1 may be derived from control point motion vectors (CPMVs) by an affine model.

[0099] The final prediction P may be calculated as a weighted sum of auxiliary predictions (eg, two auxiliary predictions P0 and P1) formulated as shown in Equations 41 and 42 below.

[0100]

number

[0101] 11 is a diagram illustrating representative weight values ​​(e.g., associated with pixels) in a sub-block. Referring to FIG. 11, an auxiliary prediction sample located at the center (e.g., center pixel) of a sub-block 1100 may be associated with a weight value of 3, and an auxiliary prediction sample located at a boundary of the sub-block 1100 may be associated with a weight value of 1.

[0102] 12 is a diagram illustrating regions where interweave prediction is applied and other regions where interweave prediction is not applied. Referring to FIG. 12, region 1200 may include a first region 1210 (not shaded as shown in FIG. 12) having, for example, 4×4 sub-blocks where interweave prediction is applied, and a second region 1220 (shaded as shown in FIG. 12) where, for example, interweave prediction is not applied. To avoid small block motion compensation, interweave prediction may be applied only to regions where the size of the sub-blocks meets a threshold size (e.g., 4×4) for, for example, both the first partitioning pattern and the second partitioning pattern.

[0103] In VTM-3.0, the size of a sub-block may be 4x4 in the chrominance component, and interweave prediction may be applied to the chrominance component and / or luma component. Because the regions used to perform motion compensation (MC) for a sub-block (e.g., all sub-blocks) can be retrieved together as a whole in AMC, bandwidth may not be increased by interweave prediction. For flexibility, a flag may be signaled in the slice header to indicate whether interweave prediction is used. For interweave prediction, the flag may be signaled as a 1-bit flag (e.g., a first logical level that can always be signaled as 0 or 1).

[0104] Representative Procedure for Sub-block-Based Temporal Motion Vector Prediction (SbTMVP) SbTMVP is supported by VTM. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP can use motion fields in co-located pictures to improve, for example, the motion vector prediction and merge mode of a CU in the current picture. The same co-located pictures used by TMVP may be used for SbTMVP. SbTMVP may differ from TMVP in that (1) TMVP can predict motion at the CU level, while SbTMVP can predict motion at the sub-CU level, and / or (2) TMVP can retrieve temporal motion vectors from co-located blocks in the co-located picture (e.g., the co-located block may be the bottom-right or center block with respect to the current CU), while SbTMVP can apply a motion shift before retrieving temporal motion information from the co-located picture (e.g., the motion shift may be obtained from a motion vector from one of the spatially neighboring blocks of the current CU), etc.

[0105] Figures 13A and 13B illustrate SbTMVP processing: Figure 13A shows the spatial neighboring blocks used by ATMVP, and Figure 13B shows the derivation of sub-CU motion fields by applying motion shifts from spatial neighbors and scaling motion information from corresponding co-located sub-CUs.

[0106] 13A and 13B, SbTMVP can predict the motion vector of a sub-CU within a current CU operation (e.g., in two operations). In the first operation, spatial neighboring blocks A1, B1, B0, and A0 may be examined in the order of A1, B1, B0, and A0. As soon as and / or after the first spatial neighboring block having a motion vector using the co-located picture as its reference picture is identified, this motion vector may be selected as the motion shift to be applied. If no such motion is identified from the spatial neighboring blocks, the motion shift may be set to (0,0). In the second operation, as shown in FIG. 13B, the motion shift identified in the first operation may be applied (e.g., added to the coordinates of the current block) to obtain sub-CU level motion information (e.g., motion vector and reference index) from the co-located picture. The example in FIG. 13B shows the motion shift set to the motion of block A1. For each sub-CU, the motion information of its corresponding block in the co-located picture (e.g., the smallest motion grid covering the center sample) may be used to derive the motion information of that sub-CU. After the motion information of the co-located sub-CU is identified, the motion information may be converted into a motion vector and reference index for the current sub-CU in a manner similar to the TMVP processing of HEVC. For example, temporal motion scaling may be applied to align the reference pictures of the temporal motion vectors with those of the current CU.

[0107] A combined sub-block-based merge list may be used in VTM-3, e.g., to signal sub-block-based merge mode, and may contain or include both SbTMVP and affine merge candidates. SbTMVP mode may be enabled / disabled by a sequence parameter set (SPS) flag. When SbTMVP mode is enabled, an SbTMVP predictor may be added as the first entry in the list of sub-block-based merge candidates, followed by affine merge candidates. The size of the sub-block-based merge list may be signaled in the SPS, and the maximum allowed size of the sub-block-based merge list may be an integer, e.g., 5 in VTM3.

[0108] The sub-CU size used in SbTMVP may be fixed, for example, at 8x8 or another sub-CU size, and as is done in affine merge mode, SbTMVP mode may be applicable to (e.g., only to) CUs having both a width and a height that may be equal to or greater than 8. The encoding logic for the additional SbTMVP merge candidates may be the same as for other merge candidates. For example, for each CU in a P or B slice, an additional rate-distortion (RD) check may be performed to determine whether to use an SbTMVP candidate.

[0109] Typical Regression-based Motion Vector Field To provide fine granularity of motion vectors within a block, a regression-based motion vector field (RMVF) tool may be implemented (e.g., in JVET-M0302), which attempts to model the motion vector of each block at the sub-block level based on the motion vectors of its spatial neighbors.

[0110] 14 is a diagram illustrating adjacent motion blocks (e.g., a 4x4 motion block) that may be used for motion parameter derivation. One row 1410 and one column 1420 of 4x4 sub-block-based (and at their center positions) directly adjacent motion vectors from each side of the block may be used in the regression process. For example, these adjacent motion vectors may be used in RMVF motion parameter derivation.

[0111] 15 is a diagram illustrating adjacent motion blocks that may be used for motion parameter derivation, reducing adjacent motion information (e.g., the number of adjacent motion blocks used in the regression process for FIG. 14 may be reduced). A reduced amount of adjacent motion information for RMVF parameter derivation of adjacent 4×4 motion blocks may be used for motion parameter derivation (e.g., approximately half, e.g., every other adjacent motion block may be used for motion parameter derivation). Particular adjacent motion blocks in rows 1410 and columns 1420 may be selected, determined, or predetermined to reduce adjacent motion information.

[0112] Although approximately half of the adjacent motion blocks in row 1410 and column 1420 are shown selected, other percentages (including other motion block positions) may be selected, for example, to reduce the number of adjacent motion blocks used in the regression process.

[0113] When collecting motion information for motion parameter derivation, the five regions shown in the figure (e.g., bottom-left, left, top-left, top, top-right) may be used. The top-right and bottom-left reference motion regions may be limited to half (e.g., only half) of the corresponding width or height of the current block.

[0114] In RMVF mode, the motion of a block can be defined by a six-parameter motion model. These parameters a xx , a xy , a yx , a yy , b x , and b ycan be calculated by solving a linear regression model in the mean squared error (MSE) sense. The inputs to the regression model are the center positions (x,y) and / or motion vectors (mv) of the available adjacent 4x4 sub-blocks, as defined above. x and mv y ) may be composed of or may include them.

[0115] (X subPU ,Y subPU ) is a motion vector (MV X_subPU ,MV Y_subPU ) can be calculated as shown in Equation 43 below:

[0116]

number

[0117] The motion vectors can be calculated for 8x8 sub-blocks relative to the center position of the sub-block (e.g., each sub-block). For example, motion compensation can be applied with 8x8 sub-block precision in RMVF mode. To obtain efficient modeling of the motion vector field, the RMVF tool is applied only if at least one motion vector from at least three of the candidate regions is available.

[0118] The affine motion model parameters can be used to derive a motion vector for a particular pixel (e.g., each pixel) in a CU. Because the complexity of generating a pixel-based affine motion compensation prediction can be high (e.g., very high), but the memory access bandwidth requirement for this type of sample-based MC can be high, a sub-block-based affine motion compensation procedure / method may be implemented (e.g., by VVC). For example, a CU may be partitioned into sub-blocks (e.g., 4x4 sub-blocks, square sub-blocks, and / or non-square sub-blocks). Each of the sub-blocks may be assigned an MV that can be derived from the affine model parameters. The MV may be the MV at the center of the sub-block (or another position within the sub-block). Pixels within a sub-block (e.g., all pixels within the sub-block) may share the sub-block MV. Sub-block-based affine motion compensation may be a trade-off between coding efficiency and complexity. To achieve finer-granularity motion compensation, inter-weave prediction for affine motion compensation may be implemented, which may be generated by weighted averaging two sub-block motion compensation predictions. Interweave prediction may require and / or use more than one motion compensated prediction per sub-block, thus increasing memory bandwidth and complexity.

[0119] In certain representative embodiments, methods, apparatuses, procedures, and / or operations may be implemented to refine sub-block-based affine motion compensation prediction using optical flow (e.g., using and / or based on optical flow). For example, after sub-block-based affine motion compensation is performed, pixel intensities may be refined by adding difference values ​​derived by an optical flow equation, which is referred to as prediction refinement with optical flow (PROF). PROF can achieve pixel-level granularity without significantly increasing complexity and may maintain similar worst-case memory access bandwidth as sub-block-based affine motion compensation. PROF may be applied in any scenario where a pixel-level motion vector field is available (e.g., can be calculated) in addition to a prediction signal (e.g., an unrefined motion prediction signal and / or a sub-block-based motion prediction signal). In addition to or other than the affine mode, the predictive PROF procedure may be used in other sub-block prediction modes. Application of PROF in sub-block modes such as SbTMVP and / or RMVF may be implemented. The application of PROF in bi-prediction is described herein.

[0120] A typical PROF procedure for affine modes In certain representative embodiments, methods, apparatus, and / or procedures may be implemented to improve the granularity of sub-block-based affine motion compensation prediction, for example, by applying pixel intensity variations derived from optical flow (e.g., optical flow equations), and may use and / or require one motion compensation operation per sub-block (e.g., only one motion compensation operation per sub-block), the same as existing affine motion compensation in, for example, VVC.

[0121] FIG. 16 is a diagram illustrating sub-block MVs and pixel-level motion vector differentials Δv(i,j) (eg, sometimes referred to as refinement MVs for pixels) after sub-block-based affine motion compensation prediction.

[0122] 16, CU 1600 may include sub-blocks 1610, 1620, 1630, and 1640. Each sub-block 1610, 1620, 1630, and 1640 may include multiple pixels (e.g., 16 pixels in sub-block 1610). A sub-block MV 1650 (e.g., as a coarse or average sub-block MV) associated with each pixel 1660(i,j) of sub-block 1610 is shown. For each respective pixel (i,j) in the sub-block 1610, a refinement MV 1670(i,j) may be determined, which may indicate the difference between the actual MV of pixel 1660(i,j) and the sub-block MV 1650 (where (i,j) defines the pixel position within the sub-block 1610). For clarity in FIG. 16, only refinement MV 1670(1,1) is labeled, although other individual pixel-level motions are shown. In certain representative embodiments, refinement MV 1670(i,j) may be determined as the pixel-level motion vector difference Δv(i,j) (sometimes referred to as the motion vector difference).

[0123] In certain exemplary embodiments, methods, apparatus, procedures and / or actions may be implemented that include any of the following actions: (1) In the first operation, sub-block-based AMC can be performed as disclosed herein to generate sub-block-based motion estimation I(i,j). (2) In the second operation, the spatial gradient g of the subblock-based motion estimation I(i,j) at each sample position is calculated. x (i,j) and g y(i,j) can be calculated (in one example, spatial gradients can be generated using the same process as the gradient generation used in BDOF. For example, the horizontal gradient at a sample location can be calculated as the difference between its right neighbor and its left neighbor, and / or the vertical gradient at a sample location can be calculated as the difference between its below neighbor and its above neighbor. In another example, spatial gradients can be generated using a Sobel filter). (3) In a third operation, the luminance intensity change for each pixel in the CU can be calculated using and / or by an optical flow equation, for example, as shown in Equation 44 below. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (44) Here, the value of the motion vector difference Δv(i,j) is the difference 1670 between the pixel-level MV calculated for sample position (i,j) denoted by v(i,j) and the sub-block-level MV 1650 of the sub-block covering pixel 1660(i,j), as shown in Figure 16. The pixel-level MV v(i,j) may be derived from the control point MV by Equations 22 and 23 for the four-parameter affine model, or Equations 32 and 33 for the six-parameter affine model.

[0124] In certain representative embodiments, the motion vector difference value Δv(i,j) may be derived by affine model parameters according to or using Equations 24 and 25, where x and y may be offsets from the pixel location to the center of the sub-block. Because the affine model parameters and pixel offsets do not change for each sub-block, the motion vector difference value Δv(i,j) may be calculated for the first sub-block and reused for other sub-blocks within the same CU. For example, because the translational affine parameters (a, b) may be the same for the pixel-level MV and the sub-block MV, the difference between the pixel-level MV and the sub-block-level MV may be calculated using Equations 45 and 46 as follows: (c, d, e, f) may be four additional affine parameters (e.g., four affine parameters other than the translational affine parameters).

[0125]

number

[0126] where (i, j) can be the pixel position relative to the top left position of the sub-block, and (x sb ,y sb ) can be the center position of the sub-block relative to the top left position of the sub-block.

[0127] FIG. 17A illustrates an exemplary procedure for determining the MV that corresponds to the actual center of a sub-block.

[0128] Referring to FIG. 17A, two sub-blocks SB0 and SB1 are shown as 4×4 sub-blocks. If the sub-block width is SW and the sub-block height is SH, the sub-block center positions may be denoted as ((SW−1) / 2, (SH−1) / 2). In another example, the sub-block center positions may be estimated based on the positions denoted as (SW / 2, SH / 2). The actual center points are P0′ for the first sub-block SB0 and P1′ for the second sub-block SB1 using ((SW−1) / 2, (SH−1) / 2). The estimated center points are (e.g., in VVC) P0 for the first sub-block SB0 and P1 for the second sub-block SB1, for example, using (SW / 2, SH / 2). In certain exemplary embodiments, the MVs of the sub-blocks may be more accurately based on the actual center positions rather than the estimated center positions (used in VVC).

[0129] 17B is a diagram illustrating the locations of chrominance samples in a 4:2:0 chrominance format. Referring to FIG. 17B, the chrominance sub-block MVs may be derived from the luminance sub-block MVs. For example, in a 4:2:0 chrominance format, one 4x4 chrominance sub-block may correspond to an 8x8 luminance region. Although exemplary embodiments are illustrated in relation to the 4:2:0 chrominance format, those skilled in the art will appreciate that other chrominance formats, such as a 4:2:2 chrominance format, may be used as well.

[0130] The chrominance sub-block MV may be derived by averaging the top-left 4x4 luma sub-block MV and the bottom-right luma sub-block MV. The derived chrominance sub-block MV may or may not be located at the center of the chrominance sub-block for chrominance sample position types 0, 2, and / or 3. For chrominance sample position types 0, 2, and 3, the chrominance sub-block center position (x sb ,y sb) may or may need to be adjusted by an offset. For example, for 4:2:0 chrominance sample position types 0, 2, and 3, adjustments may be applied as shown in Equations 47-49 below.

[0131]

number

[0132]

number

[0133]

number

[0134] The sub-block-based motion prediction I(i,j) may be refined by adding intensity variations (e.g., luminance intensity variations as provided in Equation 44 as an example). The final (i.e., refined) prediction I′(i,j) may be generated by or using Equation 50 as follows: I'(i,j)=I(i,j)+ΔI(i,j) (50)

[0135] When refinement is applied, sub-block-based affine motion compensation may achieve pixel-level granularity without increasing worst-case bandwidth and / or memory bandwidth.

[0136] To maintain the accuracy of prediction and / or gradient calculation, the bit depth in the motion-related performance of sub-block-based AMC may be an intermediate bit depth, which may be higher than the coding bit depth.

[0137] The above-described processing may be used to refine the chrominance intensity (e.g., in addition to or instead of refining the luma intensity). In one example, the intensity difference used in Equation 50 may be multiplied by a weighting factor w before being added to the prediction, as shown in Equation 51 below: I'(i,j)=I(i,j)+w·ΔI(i,j) (51) Here, w may be set to a value from 0 to 1, and w may be signaled at the CU level or the picture level. For example, w may be signaled by a weight index. For example, index table 1 may be used to signal w.

[0138] [Table 1]

[0139] The encoder algorithm can choose the value of w that results in the lowest rate-distortion cost.

[0140] The gradient of the predicted sample, e.g., g x and / or g y can be calculated in different ways. In certain exemplary embodiments, the predicted sample g x and g y can be calculated by applying a 2-dimensional Sobel filter. An example of a 3x3 Sobel filter for horizontal and vertical gradients is shown below:

[0141]

number

[0142]

number

[0143] In other exemplary embodiments, the gradient may be calculated using a one-dimensional 3-tap filter. Examples may include [-1 0 1], which may be simpler (e.g., much simpler) than a Sobel filter.

[0144] 17C is a diagram illustrating extended sub-block prediction. The shaded circle 1710 is a padding sample around a 4x4 sub-block (e.g., the unshaded circle 1720). As an example, using a Sobel filter, the samples in box 1730 can be used to calculate the gradient of the central sample 1740. The gradient can be calculated using a Sobel filter, although other filters, such as a 3-tap filter, are also possible.

[0145] For the above example gradient filters, e.g., 3x3 Sobel filters and 1-dimensional filters, extended sub-block prediction may be used and / or required for sub-block gradient calculation: one row at the top and bottom boundaries and one column at the left and right boundaries of the sub-block may be padded, for example, to calculate the gradients of those samples at the sub-block boundary.

[0146] There may be different methods / procedures and / or operations for obtaining extended sub-block predictions. In one exemplary embodiment, given a sub-block size of N×M, an (N+2)×(M+2) extended sub-block prediction may be obtained by performing (N+2)×(M+2) block motion compensation using the sub-block MV. In this embodiment, memory bandwidth may be increased. To avoid memory bandwidth increase, in a specific exemplary embodiment, a K-tap interpolation filter in both the horizontal and vertical directions may be provided, and for the interpolation of an N×M sub-block, (N+K−1)×(M+K−1) integer reference samples before interpolation may be retrieved, and boundary samples of the (N+K−1)×(M+K−1) block may be copied from neighboring samples of the (N+K−1)×(M+K−1) sub-block so that the extended region may be (N+K−1+2)×(M+K−1+2). The extended region may be used for the interpolation of an (N+2)×(M+2) sub-block. These exemplary embodiments may further use and / or require additional interpolation operations to generate the (N+2)×(M+2) prediction if the sub-block MV points point to fractional positions.

[0147] For example, to reduce computational complexity, in other representative embodiments, sub-block prediction may be obtained by N×M block motion compensation using sub-block MVs. The boundaries of the (N+2)×(M+2) prediction may be obtained without interpolation by either (1) integer motion compensation where MVs are integer parts of the sub-block MVs, (2) integer motion compensation where MVs are the nearest integer MVs of the sub-block MVs, and / or (3) copying from the nearest neighboring samples in the N×M sub-block prediction.

[0148] For example, the precision and / or range of the pixel-level refinement MV may affect the accuracy of the PROF. In certain exemplary embodiments, a combination of a multi-bit fractional component and another multi-bit integer component may be implemented. For example, a 5-bit fractional component and an 11-bit integer component may be used. The combination of the 5-bit fractional component and the 11-bit integer component can represent a MV range of -1024 to 1023 with a total of 16 bits of 1 / 32 pel precision.

[0149] Gradient, e.g. g x and g y , as well as the precision of the intensity change ΔΙ can affect the performance of PROF. In certain representative embodiments, the prediction sample precision may be maintained or retained to a predetermined or signaled number of bits (e.g., the internal sample precision defined in the current VVC draft, which is 14 bits). In certain representative embodiments, the slope and / or intensity change ΔΙ may be maintained to the same precision as the prediction sample.

[0150] The range of the intensity change ΔΙ can affect the performance of PROF. The intensity change ΔΙ can be clipped to a smaller range to avoid erroneous values ​​generated by an inaccurate affine model. In one example, the intensity change ΔΙ can be clipped to predition_bitdepth-2.

[0151] Δv x and Δv y The combination of the number of bits for the fractional component of , the number of bits for the fractional component of the gradient, and the number of bits for the intensity change ΔΙ may affect the complexity of a particular hardware or software implementation. In one exemplary embodiment, 5 bits are used to represent Δv x and Δv y , 2 bits can be used to represent the fractional component of the slope, and 12 bits can be used to represent ΔI, although they can be any number of bits.

[0152] To reduce computational complexity, PROF may be omitted in certain situations. For example, if the magnitude of all pixel-based delta (e.g., refinement) MV(Δv(i,j)) in a 4x4 sub-block is less than a threshold, PROF may be omitted for the entire affine CU. If the gradient of all samples in a 4x4 sub-block is less than a threshold, PROF may be omitted. PROF may be applied to chrominance components, such as the Cb and / or Cr components. The delta MVs of the Cb and / or Cr components of a sub-block may reuse the delta MVs of the sub-block (e.g., may reuse delta MVs calculated for different sub-blocks within the same CU).

[0153] Although the gradient procedures disclosed herein (e.g., using copied reference samples to expand sub-blocks for gradient calculation) are shown as being used with PROF operations, the gradient procedures may also be used with other operations, such as BDOF operations and / or affine motion estimation operations, among others.

[0154] Representative PROF Procedures for Other Sub-Block Modes PROF can be applied in any scenario where a pixel-level motion vector field is available (e.g., can be calculated) in addition to a prediction signal (e.g., an unrefined prediction signal). For example, besides affine mode, prediction refinement using optical flow may be used in other sub-block prediction modes, such as SbTMVP mode (e.g., ATMVP mode in VVC), or regression-based motion vector field (RMVF).

[0155] In certain exemplary embodiments, a method for applying PROF to SbTMVP may be implemented. For example, such a method may include, among other things: (1) In the first operation, sub-block level MVs and sub-block predictions can be generated based on the existing SbTMVP processing described herein; (2) In the second operation, the affine model parameters can be estimated by the sub-block MV fields using a linear regression method / procedure; (3) In a third operation, pixel-level MVs can be derived according to the affine model parameters obtained in the second operation, and associated pixel-level motion refinement vectors (Δv(i,j)) for the sub-block MVs can be calculated; and / or (4) In a fourth operation, prediction refinement using optical flow processing can be applied to generate, among other things, the final prediction.

[0156] In certain exemplary embodiments, a method for applying the PROF to the RMVF may be implemented. For example, such a method may include any of the following: (1) In the first operation, the sub-block level MV field, the sub-block prediction, and / or the affine model parameters a xx , a xy , a yx , a yy , b x and b x can be generated based on the RMVF processing described herein, (2) In the second operation, the pixel-level MV offset from the sub-block-level MV is calculated as the affine model parameter a xx , a xy , a yx , a yy , b x and b x It can be derived by:

[0157]

number

[0158] where (i,j) is a pixel offset from the sub-block center. Because the affine model parameters and / or pixel offset from the sub-block center do not change from sub-block to sub-block, the pixel MV offset may be calculated for (e.g., only need to be or should be calculated for) the first sub-block and may be reused for other sub-blocks within the CU; and / or (3) In the third operation, PROF processing can be applied to generate a final prediction, for example, by applying Equations 44 and 50.

[0159] A representative PROF procedure for bidirectional prediction. In addition to or instead of using PROF for uni-prediction as described herein, PROF techniques may be used for bi-prediction. When used in bi-prediction, PROF may be used to generate an L0 prediction and / or an L1 prediction, for example, before they are combined with weights. To reduce computational complexity, PROF may be applied to (e.g., only) one prediction, such as L0 or L1. In certain representative embodiments, PROF may be applied to a list (e.g., a list with or associated with a reference picture to which the current picture is close and / or nearest (e.g., within a threshold)).

[0160] Typical steps for enabling PROF PROF enablement may be signaled in or within a sequence parameter set (SPS) header, a picture parameter set (PPS) header, and / or a tile group header. In particular embodiments, a flag may be signaled to indicate whether PROF is enabled for affine mode. If the flag is set to a first logical level (e.g., “True”), PROF may be used for both uni-prediction and bi-prediction. In particular embodiments, if the first flag is set to “True,” a second flag may be used to indicate whether PROF is enabled or not for bi-predictive affine mode. If the first flag is set to a second logical level (e.g., “False”), the second flag may be inferred to be set to “False.” Whether PROF applies to the chroma component may be signaled using a flag in or within an SPS header, a PPS header, and / or a tile group header when the first flag is set to “True,” thereby separating the control of PROF for the luma and chroma components.

[0161] Typical steps for conditionally enabled PROF For example, to reduce complexity, PROF may be applied when (e.g., only when) certain conditions are met. For example, for small CU sizes (e.g., below a threshold level), the benefit of applying PROF may be limited because affine motion is relatively small. In certain representative embodiments, when the CU size is small (e.g., for CU sizes of 16×16 or less, such as 8×8, 8×16, 16×8), or under such conditions, PROF may be disabled in affine motion compensation to reduce complexity for both the encoder and / or decoder. In certain representative embodiments, when the CU size is small (below the same or a different threshold level), PROF may be omitted in affine motion estimation (e.g., affine motion estimation only) to reduce the complexity of the encoder, and PROF may be performed in the decoder regardless of the CU size. For example, on the encoder side, after motion estimation to search for affine model parameters (e.g., control points MV), a motion compensation (MC) procedure may be invoked and PROF may be performed. The MC procedure may be invoked for each iteration in motion estimation. In MC during motion estimation, PROF may be omitted to reduce complexity, but the final MC in the encoder will perform PROF, so that prediction mismatch does not occur between the encoder and the decoder. That is, PROF refinement may not be applied when the encoder searches for affine model parameters (e.g., affine MVs) to use for predicting a CU, and once or after the encoder completes the search, the encoder can apply PROF to refine the prediction for the CU using the affine model parameters determined from the search.

[0162] In some representative embodiments, the difference between CPMVs may be used as a criterion to determine whether to enable PROF. When the difference between CPMVs is small (e.g., below a threshold level) and therefore the affine motion is small, the benefit of applying PROF may be limited, and PROF may be disabled for affine motion compensation and / or affine motion estimation. For example, in a four-parameter affine mode, PROF may be disabled if the following conditions are met (e.g., all of the following conditions are met):

[0163]

number

[0164]

number

[0165] In the six-parameter affine mode, PROF may be disabled if the following conditions are met (e.g., all of the following conditions are also met), in addition to or instead of the above conditions:

[0166]

number

[0167]

number

[0168] where T is a predefined threshold, e.g., 4. This CPMV or affine parameter-based PROF omission procedure may be applied (e.g., may only be applied) at the encoder, and the decoder may or may not omit PROF.

[0169] Representative Procedure for PROF in Combination with or Instead of a Deblocking Filter PROF may be a pixel-by-pixel refinement that can compensate for block-based MC, so that motion differences between block boundaries may be reduced (e.g., significantly reduced). The encoder and / or decoder may omit applying a deblocking filter and / or apply a weaker filter at sub-block boundaries when PROF is applied. In a CU that is divided into multiple transform units (TUs), blocking artifacts may appear on transform block boundaries.

[0170] In certain representative embodiments, the encoder and / or decoder may omit applying a deblocking filter unless the sub-block boundary coincides with a TU boundary, or may apply one or more weaker filters to the sub-block boundary.

[0171] When or under what conditions PROF is applied to luma (e.g., applied only to luma), the encoder and / or decoder may omit application of a deblocking filter and / or apply one or more weaker filters on sub-block boundaries for luma (e.g., luma only). For example, a boundary strength parameter B may be used to apply a weaker deblocking filter.

[0172] For example, the encoder and / or decoder may omit applying a deblocking filter on sub-block boundaries when PROF is applied unless the sub-block boundaries coincide with TU boundaries, in which case the deblocking filter may be applied to reduce or remove blocking artifacts that may occur along TU boundaries.

[0173] As another example, the encoder and / or decoder may apply a weaker deblocking filter on sub-block boundaries when PROF is applied, unless the sub-block boundaries coincide with TU boundaries. A "weak" deblocking filter is intended to be a weaker deblocking filter than might normally be applied to sub-block boundaries when PROF is not applied. When a sub-block boundary coincides with a TU boundary, a stronger deblocking filter may be applied to reduce or remove blocking artifacts that are likely to be more visible along the sub-block boundaries that coincide with the TU boundaries.

[0174] In certain representative embodiments, when or under conditions when PROF is applied to luma (e.g., only applied), the encoder and / or decoder may, for design uniformity purposes, align the application of the deblocking filter for chroma to luma, e.g., despite the absence of application of PROF in chroma. For example, if PROF is applied only to luma, the normal application of the deblocking filter for luma may be modified based on whether PROF has been applied (and possibly based on whether the sub-block boundary is at a TU boundary). In certain representative embodiments, rather than having separate / different logic for applying the deblocking filter to corresponding chroma pixels, the deblocking filter may be applied to the sub-block boundary for chroma in a manner that is consistent with (and / or mirrors) the procedure for luma deblocking.

[0175] FIG. 18A is a flowchart illustrating a first exemplary encoding / decoding method.

[0176] 18A , a representative method 1800 of encoding and / or decoding may include, at block 1805, the encoder 100 or 300 and / or the decoder 200 or 500 obtaining a sub-block-based motion prediction signal for a current block, e.g., of a video. At block 1810, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain one or more spatial gradients of the sub-block-based motion prediction signal for the current block or one or more motion vector differential values ​​associated with sub-blocks of the current block. At block 1815, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain a refinement signal for the current block based on the one or more obtained spatial gradients or one or more motion vector differential values ​​associated with sub-blocks of the current block. At block 1820, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain a refined motion prediction signal for the current block based on the sub-block-based motion prediction signal and the refinement signal. In particular embodiments, the encoder 100 or 300 may encode the current block based on the refined motion prediction signal, or the decoder 200 or 500 may decode the current block based on the refined motion prediction signal. The refined motion prediction signal may be a refined motion inter prediction signal generated (e.g., by the GBi encoder 300 and / or the GBi decoder 500) and may use one or more PROF operations.

[0177] For example, in certain representative embodiments related to other methods described herein, including methods 1850 and 1900, obtaining a sub-block-based motion prediction signal for a current block of video may include generating a sub-block-based motion prediction signal.

[0178] For example, in certain representative embodiments related to other methods described herein, including particularly methods 1850 and 1900, obtaining one or more spatial gradients of a sub-block-based motion prediction signal for the current block or one or more motion vector difference values ​​associated with a sub-block of the current block may include determining one or more spatial gradients (e.g., associated with a gradient filter) of the sub-block-based motion prediction signal.

[0179] For example, in certain representative embodiments related to other methods described herein, including particularly methods 1850 and 1900, obtaining one or more spatial gradients of a sub-block-based motion prediction signal for the current block, or one or more motion vector differential values ​​associated with sub-blocks of the current block, may include determining one or more motion vector differential values ​​associated with sub-blocks of the current block.

[0180] For example, in certain representative embodiments related to other methods described herein, including particularly methods 1850 and 1900, obtaining a refinement signal for the current block based on one or more determined spatial gradients or one or more determined motion vector differential values ​​may include determining a motion prediction refinement signal for the current block as the refinement signal based on the determined spatial gradients.

[0181] For example, in certain representative embodiments related to other methods described herein, including particularly methods 1850 and 1900, obtaining a refinement signal for the current block based on one or more determined spatial gradients or one or more determined motion vector difference values ​​may include determining a motion prediction refinement signal for the current block as the refinement signal based on the determined motion vector difference values.

[0182] The terms "determine" or "determining" in relation to something such as information may generally include one or more of estimating, calculating, predicting, obtaining, and / or retrieving information. For example, determining may refer to retrieving something from a memory or a bitstream, among other things.

[0183] For example, in certain representative embodiments related to other methods described herein, including particularly methods 1850 and 1900, obtaining a refined motion prediction signal for the current block based on a sub-block-based motion prediction signal and a refinement signal may include combining (e.g., particularly by adding or subtracting) the sub-block-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block.

[0184] For example, in certain representative embodiments related to other methods described herein, including particularly methods 1850 and 1900, encoding and / or decoding the current block based on the refined motion prediction signal may include encoding video using the refined motion prediction signal as a prediction for the current block and / or decoding video using the refined motion prediction signal as a prediction for the current block.

[0185] FIG. 18B is a flowchart illustrating a second exemplary encoding and / or decoding method.

[0186] Referring to FIG. 18B , a representative method 1850 for encoding and / or decoding video may include, at block 1855, the encoder 100 or 300 and / or decoder 200 or 500 generating a sub-block-based motion prediction signal. At block 1860, the encoder 100 or 300 and / or decoder 200 or 500 may determine one or more spatial gradients (e.g., associated with a gradient filter) of the sub-block-based motion prediction signal. At block 1865, the encoder 100 or 300 and / or decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined spatial gradients. At block 1870, the encoder 100 or 300 and / or decoder 200 or 500 may combine (e.g., particularly by adding or subtracting) the sub-block-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. In block 1875, encoder 100 or 300 may encode video using the refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 may decode video using the refined motion prediction signal as a prediction for the current block. In particular embodiments, the operations in blocks 1810, 1820, 1830, and 1840 may be performed on a current block, which is generally a block that refers to the block currently being coded or decoded. The refined motion prediction signal may be a refined motion inter prediction signal generated (e.g., by GBi encoder 300 and / or GBi decoder 500) and may use one or more PROF operations.

[0187] For example, the determination of one or more spatial gradients of the sub-block-based motion prediction signal by the encoder 100 or 300 and / or the decoder 200 or 500 may include determining a first set of spatial gradients associated with a first reference picture and a second set of spatial gradients associated with a second reference picture. The determination of the motion prediction refinement signal for the current block by the encoder 100 or 300 and / or the decoder 200 or 500 may be based on the determined spatial gradients, may include determining a motion inter-prediction refinement signal (e.g., a bi-prediction signal) for the current block based on the first and second sets of spatial gradients, and may also be based on weight information W (e.g., indicating or including one or more weight values ​​associated with one or more reference pictures).

[0188] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, encoder 100 or 300 may generate, use, and / or transmit weighting information W to decoder 200 or 500, and / or decoder 200 or 500 may receive or obtain weighting information W. For example, the motion inter-prediction refinement signal for the current block may be based on (1) first gradient values ​​derived from a first set of spatial gradients and weighted according to a first weighting factor indicated by weighting information W, and / or (2) second gradient values ​​derived from a second set of spatial gradients and weighted according to a second weighting factor indicated by weighting information W.

[0189] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may further include the encoder 100 or 300 and / or the decoder 200 or 500 determining affine motion model parameters for the current block of video such that a sub-block-based motion prediction signal may be generated using the determined affine motion model parameters.

[0190] In certain exemplary embodiments, including exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, one or more spatial gradients of the sub-block-based motion prediction signal, which may include calculating at least one gradient value for one respective sample position, a portion of each sample position, or each respective sample position in at least one sub-block of the sub-block-based motion prediction signal. For example, calculating at least one gradient value for one respective sample position, a portion of each sample position, or each respective sample position in at least one sub-block of the sub-block-based motion prediction signal may include applying a gradient filter to the one respective sample position, a portion of each sample position, or each respective sample position in the at least one sub-block of the sub-block-based motion prediction signal.

[0191] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may further include encoder 100 or 300 and / or decoder 200 or 500 determining a set of motion vector difference values ​​associated with sample positions of a first sub-block of a current block of a sub-block-based motion prediction signal. In some examples, difference values ​​may be determined for a sub-block (e.g., the first sub-block) and may be reused for some or all other sub-blocks within the current block. In particular examples, a sub-block-based motion prediction signal may be generated and a set of motion vector difference values ​​may be determined using an affine motion model or a different motion model (e.g., another sub-block-based motion model such as the SbTMVP model). By way of example, a set of motion vector difference values ​​may be determined for a first sub-block of a current block and used to determine a motion prediction refinement signal for one or more further sub-blocks of the current block.

[0192] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, one or more spatial gradients of a sub-block-based motion prediction signal and a set of motion vector difference values ​​may be used to determine a motion prediction refinement signal for the current block.

[0193] In certain exemplary embodiments, including exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, an affine motion model for the current block is used to generate a sub-block-based motion prediction signal and determine a set of motion vector differential values.

[0194] In certain exemplary embodiments, including exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, determining one or more spatial gradients of the sub-block-based motion prediction signal may include, for one or more respective sub-blocks of the current block, determining an extended sub-block using the sub-block-based motion prediction signal and neighboring reference samples that border and surround the respective sub-block, and determining a spatial gradient of the respective sub-block using the determined extended sub-block to determine a motion prediction refinement signal.

[0195] FIG. 19 is a flowchart illustrating a third exemplary encoding and / or decoding method.

[0196] 19, a representative method 1900 for encoding and / or decoding video may include, at block 1910, the encoder 100 or 300 and / or the decoder 200 or 500 generating a sub-block-based motion prediction signal. At block 1920, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a set of motion vector difference values ​​associated with the sub-blocks of the current block (e.g., the set of motion vector difference values ​​may be associated with all of the sub-blocks of the current block, for example). At block 1930, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined set of motion vector difference values. In block 1940, encoder 100 or 300 and / or decoder 200 or 500 may combine (e.g., particularly add or subtract) the sub-block-based motion prediction signal and the motion prediction refinement signal to create or generate a refined motion prediction signal for the current block. In block 1950, encoder 100 or 300 may encode video using the refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 may decode video using the refined motion prediction signal as a prediction for the current block. In particular embodiments, the operations in blocks 1910, 1920, 1930, and 1940 may be performed on a current block, which generally refers to the block currently being coded or decoded. In certain exemplary embodiments, the refined motion prediction signal may be a refined motion inter prediction signal generated (e.g., by GBi encoder 300 and / or GBi decoder 500) and may use one or more PROF operations.

[0197] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 determining motion model parameters (e.g., one or more affine motion model parameters) for a current block of video such that a sub-block-based motion prediction signal may be generated using the determined motion model parameters (e.g., affine motion model parameters).

[0198] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, one or more spatial gradients of the sub-block-based motion prediction signal. For example, determining the one or more spatial gradients of the sub-block-based motion prediction signal may include calculating at least one gradient value for one respective sample position, a portion of each sample position, or each respective sample position in at least one sub-block of the sub-block-based motion prediction signal. For example, calculating the at least one gradient value for one respective sample position, a portion of each sample position, or each respective sample position in at least one sub-block of the sub-block-based motion prediction signal may include applying a gradient filter to the one respective sample position, a portion of each sample position, or each respective sample position in the at least one sub-block of the sub-block-based motion prediction signal.

[0199] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, a motion prediction refinement signal for the current block using gradient values ​​associated with spatial gradients for one respective sample position, a portion of each respective sample position, or each respective sample position of the current block, and a determined set of motion vector difference values ​​associated with sample positions of a sub-block (e.g., any sub-block) of the current block of the sub-block motion prediction signal.

[0200] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, determining a motion prediction refinement signal for the current block may use gradient values ​​associated with spatial gradients for one or more respective sample positions or each sample position of one or more sub-blocks of the current block and a determined set of motion vector difference values.

[0201] FIG. 20 is a flowchart illustrating a fourth exemplary encoding and / or decoding method.

[0202] 20 , a representative method 2000 for encoding and / or decoding video may include, at block 2010, the encoder 100 or 300 and / or the decoder 200 or 500 generating a sub-block-based motion prediction signal using at least a first motion vector for a first sub-block of a current block and an additional motion vector for a second sub-block of the current block. At block 2020, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate a first set of gradient values ​​for a first sample position in the first sub-block of the sub-block-based motion prediction signal and a second, different set of gradient values ​​for a second sample position in the first sub-block of the sub-block-based motion prediction signal. At block 2030, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a first set of motion vector differential values ​​for the first sample position and a second, different set of motion vector differential values ​​for the second sample position. For example, a first set of motion vector difference values ​​for a first sample location may indicate a difference between the motion vector at the first sample location and the motion vector of the first sub-block, and a second set of motion vector difference values ​​for a second sample location may indicate a difference between the motion vector at the second sample location and the motion vector of the first sub-block. At block 2040, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a prediction refinement signal using the first and second sets of gradient values ​​and the first and second sets of motion vector difference values. At block 2050, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., particularly by adding or subtracting) the sub-block-based motion prediction signal with the prediction refinement signal to create a refined motion prediction signal.In block 2060, the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block, and / or the decoder 200 or 500 may decode the video using the refined motion prediction signal as a prediction for the current block. In particular embodiments, the operations in blocks 2010, 2020, 2030, 2040, and 2050 may be performed on a current block that includes multiple sub-blocks.

[0203] FIG. 21 is a flowchart illustrating a fifth exemplary encoding and / or decoding method.

[0204] 21, a representative method 2100 for encoding and / or decoding video may include, at block 2110, the encoder 100 or 300 and / or the decoder 200 or 500 generating a sub-block-based motion prediction signal for a current block. At block 2120, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a prediction refinement signal using optical flow information indicating refined motion of multiple sample positions in the current block of the sub-block-based motion prediction signal. At block 2130, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., particularly by adding or subtracting) the sub-block-based motion prediction signal with the prediction refinement signal to create a refined motion prediction signal. In block 2140, encoder 100 or 300 may encode video using the refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 may decode video using the refined motion prediction signal as a prediction for the current block. For example, the current block may include multiple sub-blocks, and a sub-block-based motion prediction signal may be generated using at least a first motion vector for a first sub-block of the current block and an additional motion vector for a second sub-block of the current block.

[0205] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, a prediction refinement signal that can use optical flow information. This determining may include calculating, by encoder 100 or 300 and / or decoder 200 or 500, a first set of gradient values ​​for a first sample location in a first sub-block of the sub-block-based motion prediction signal and a second, different set of gradient values ​​for a second sample location in the first sub-block of the sub-block-based motion prediction signal. The first set of motion vector differential values ​​for the first sample location and the second, different set of motion vector differential values ​​for the second sample location may be determined. For example, a first set of motion vector difference values ​​for a first sample location may indicate a difference between the motion vector at the first sample location and the motion vector of the first sub-block, and a second set of motion vector difference values ​​for a second sample location may indicate a difference between the motion vector at the second sample location and the motion vector of the first sub-block. The encoder 100 or 300 and / or the decoder 200 or 500 may determine a prediction refinement signal using the first and second sets of gradient values ​​and the first and second sets of motion vector difference values.

[0206] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, a prediction refinement signal that can use optical flow information. This determination may include calculating a third set of gradient values ​​for a first sample location in a second sub-block of the sub-block-based motion prediction signal and a fourth set of gradient values ​​for a second sample location in the second sub-block of the sub-block-based motion prediction signal. Encoder 100 or 300 and / or decoder 200 or 500 may determine a prediction refinement signal for the second sub-block using the third and fourth sets of gradient values ​​and the first and second sets of motion vector difference values.

[0207] FIG. 22 is a flowchart illustrating a sixth exemplary encoding and / or decoding method.

[0208] Referring to FIG. 22, a representative method 2200 for encoding and / or decoding video may include, at block 2210, the encoder 100 or 300 and / or the decoder 200 or 500 determining a motion model for a current block of the video. The current block may include multiple sub-blocks. For example, the motion model may generate individual (e.g., per-sample) motion vectors for multiple sample locations in the current block. At block 2220, the encoder 100 or 300 and / or the decoder 200 or 500 may generate a sub-block-based motion prediction signal for the current block using the determined motion model. The generated sub-block-based motion prediction signal may use one motion vector for each sub-block of the current block. At block 2230, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate gradient values ​​by applying a gradient filter to some of the multiple sample locations of the sub-block-based motion prediction signal. In block 2240, the encoder 100 or 300 and / or the decoder 200 or 500 may determine motion vector difference values ​​for some of the sample locations, each of which may indicate a difference between a motion vector generated for the respective sample location (e.g., an individual motion vector) according to a motion model and a motion vector used to create a sub-block-based motion prediction signal for the sub-block that includes the respective sample location. In block 2250, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a prediction refinement signal using the gradient values ​​and the motion vector difference values. In block 2260, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., particularly by adding or subtracting) the sub-block-based motion prediction signal with the prediction refinement signal to create a refined motion prediction signal for the current block.In block 2270, the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block, and / or the decoder 200 or 500 may decode the video using the refined motion prediction signal as a prediction for the current block.

[0209] FIG. 23 is a flowchart illustrating a seventh exemplary encoding and / or decoding method.

[0210] Referring to FIG. 23, a representative method 2300 for encoding and / or decoding video may include, at block 2310, the encoder 100 or 300 and / or decoder 200 or 500 performing sub-block-based motion compensation to generate a sub-block-based motion prediction signal as a coarse motion prediction signal. At block 2320, the encoder 100 or 300 and / or decoder 200 or 500 may calculate one or more spatial gradients of the sub-block-based motion prediction signal at the sample position. At block 2330, the encoder 100 or 300 and / or decoder 200 or 500 may calculate a pixel-by-pixel intensity change in the current block based on the calculated spatial gradient. At block 2340, the encoder 100 or 300 and / or decoder 200 or 500 may determine a pixel-by-pixel-based motion prediction signal as a refined motion prediction signal based on the calculated pixel-by-pixel intensity change. In block 2350, the encoder 100 or 300 and / or the decoder 200 or 500 may predict the current block using the coarse motion prediction signal for each sub-block of the current block and using the refined motion prediction signal for each pixel of the current block. In particular embodiments, the operations in blocks 2310, 2320, 2330, 2340, and 2350 may be performed for at least one block in the video (e.g., the current block). For example, calculating the intensity change per pixel in the current block may include determining the luminance intensity change for each pixel in the current block according to an optical flow equation. Predicting the current block may include predicting a motion vector for each respective pixel in the current block by combining a coarse motion prediction vector for a sub-block that includes the respective pixel with a refined motion prediction vector for the coarse motion prediction vector associated with the respective pixel.

[0211] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, one or more spatial gradients of the sub-block-based motion prediction signal may include either a horizontal gradient and / or a vertical gradient, e.g., a horizontal gradient may be calculated as a luminance or chrominance difference between a right-neighboring sample of a sample of the sub-block and a left-neighboring sample of the sub-block, and / or a vertical gradient may be calculated as a luminance or chrominance difference between a bottom-neighboring sample of a sample of the sub-block and an top-neighboring sample of the sub-block.

[0212] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, one or more spatial gradients of the sub-block predictions may be generated using a Sobel filter.

[0213] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the coarse motion prediction signal can use either a four-parameter affine model or a six-parameter affine model. For example, sub-block-based motion compensation can be one of (1) affine sub-block-based motion compensation or (2) another compensation (e.g., sub-block-based temporal motion vector prediction (SbTMVP) mode motion compensation and / or regression-based motion vector field (RMVF) mode-based compensation). Provided that SbTMVP mode-based motion compensation is performed, the method can include estimating affine model parameters using the sub-block motion vector field by a linear regression operation and deriving pixel-level motion vectors using the estimated affine model parameters. Provided that RMVF mode-based motion compensation is performed, the method can include estimating affine model parameters and deriving pixel-level motion vector offsets from sub-block-level motion vectors using the estimated affine model parameters. For example, the pixel motion vector offset may be relative to the center of the sub-block (e.g., the actual center or the sample position closest to the actual center). For example, the coarse motion prediction vector for the sub-block may be based on the actual center position of the sub-block.

[0214] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the methods may include encoder 100 or 300 or decoder 200 or 500 selecting, as a center position associated with a coarse motion prediction vector (e.g., a subblock-based motion prediction vector) for each subblock, either (1) the actual center of each subblock or (2) one of the pixel (e.g., sample) positions closest to the center of the subblock. For example, predicting the current block using the coarse motion prediction signal (e.g., subblock-based motion prediction signal) of the current block and using the refined motion prediction signal for each pixel (e.g., sample) of the current block may be based on the selected center position of each subblock. For example, the encoder 100 or 300 and / or the decoder 200 or 500 may determine center positions associated with the chrominance pixels of the sub-block and may determine offsets for the center positions of the chrominance pixels of the sub-block based on the chrominance position sample types associated with the chrominance pixels. A coarse motion prediction signal (e.g., a sub-block-based motion prediction signal) for the sub-block may be based on the actual positions of the sub-block corresponding to the determined center positions of the chrominance pixels adjusted by the offset.

[0215] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the methods may include encoder 100 or 300 generating, or decoder 200 or 500 receiving, information indicating whether optical flow prediction refinement (PROF) is enabled in one of (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header. For example, under the condition that PROF is enabled, a refined motion estimation operation may be performed, and thus both a coarse motion estimation signal (e.g., a sub-block-based motion estimation signal) and the refined motion estimation signal may be used to predict the current block. As another example, under the condition that PROF is not enabled, a refined motion estimation operation may not be performed, and thus only a coarse motion estimation signal (e.g., a sub-block-based motion estimation signal) may be used to predict the current block.

[0216] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining whether to perform a refined motion estimation operation on the current block or on the affine motion estimation based on attributes of the current block and / or attributes of the affine motion estimation.

[0217] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include determining whether to perform a refined motion estimation operation on the current block or perform affine motion estimation based on attributes of the current block and / or attributes of the affine motion estimation. For example, determining whether to perform a refined motion estimation operation on the current block based on attributes of the current block may include determining whether to perform a refined motion estimation operation on the current block based on either (1) whether the size of the current block exceeds a certain size and / or (2) whether a control point motion vector (CPMV) difference exceeds a threshold.

[0218] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the methods may include encoder 100 or 300 and / or decoder 200 or 500 applying a first deblocking filter to one or more boundaries of sub-blocks of the current block that coincide with a transform unit boundary and applying a second, different deblocking filter to other boundaries of the sub-blocks of the current block that do not coincide with any transform unit boundary. For example, the first deblocking filter may be a stronger deblocking filter than the second deblocking filter.

[0219] FIG. 24 is a flowchart illustrating an eighth exemplary encoding and / or decoding method.

[0220] Referring to Figure 24, a representative method 2400 for encoding and / or decoding video may include, at block 2410, the encoder 100 or 300 and / or decoder 200 or 500 performing sub-block-based motion compensation to generate a sub-block-based motion prediction signal as a coarse motion prediction signal. At block 2420, the encoder 100 or 300 and / or decoder 200 or 500 may determine, for each respective boundary sample of a sub-block of a current block, one or more reference samples surrounding the sub-block that correspond to samples neighboring the respective boundary sample as surrounding reference samples, and may determine one or more spatial gradients associated with the respective boundary sample using the surrounding reference samples and the samples of the sub-block neighboring the respective boundary sample. At block 2430, the encoder 100 or 300 and / or decoder 200 or 500 may determine, for each respective non-boundary sample in the sub-block, one or more spatial gradients associated with the respective non-boundary sample using samples of the sub-block neighboring the respective non-boundary sample. In block 2440, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate per-pixel intensity changes in the current block using the determined spatial gradients of the sub-blocks. In block 2450, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a per-pixel based motion prediction signal as a refined motion prediction signal based on the calculated per-pixel intensity changes. In block 2460, the encoder 100 or 300 and / or the decoder 200 or 500 may predict the current block using a coarse motion prediction signal associated with each sub-block of the current block and using a refined motion prediction signal associated with each pixel of the current block.In particular embodiments, the operations at 2410, 2420, 2430, 2440, 2450, and 2460 may be performed on at least one block (e.g., the current block) in the video.

[0221] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, determining one or more spatial gradients of boundary samples and non-boundary samples may include calculating one or more spatial gradients using either (1) a vertical Sobel filter, (2) a horizontal Sobel filter, or (3) a 3-tap filter.

[0222] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 copying surrounding reference samples from a reference store without any further manipulation, and determining one or more spatial gradients associated with each boundary sample may use the copied surrounding reference samples to determine one or more spatial gradients associated with each boundary sample.

[0223] FIG. 25 is a flowchart illustrating a representative gradient calculation method.

[0224] 25, a representative method 2500 for calculating gradients of sub-blocks using reference samples corresponding to samples proximate a boundary of the sub-block (e.g., used in encoding and / or decoding video) may include, at block 2510, the encoder 100 or 300 and / or decoder 200 or 500 determining, for each respective boundary sample of a sub-block of a current block, one or more reference samples that correspond to samples proximate the respective boundary sample and surround the sub-block as surrounding reference samples, and determining one or more spatial gradients associated with each boundary sample using the surrounding reference samples and the samples of the sub-block proximate the respective boundary samples. At block 2520, the encoder 100 or 300 and / or decoder 200 or 500 may determine, for each respective non-boundary sample in the sub-block, one or more spatial gradients associated with each non-boundary sample using the samples of the sub-block proximate the respective non-boundary sample. In particular embodiments, the operations at blocks 2510 and 2520 may be performed on at least one block (eg, the current block) in the video.

[0225] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, 2500, and 2600, the determined spatial gradient or gradients may be used to predict the current block by either (1) a prediction refinement by optical flow (PROF) operation, (2) a bidirectional optical flow operation, or (3) an affine motion estimation operation.

[0226] FIG. 26 is a flowchart illustrating a ninth exemplary encoding and / or decoding method.

[0227] Referring to FIG. 26, a representative method 2600 for encoding and / or decoding video may include an encoder 100 or 300 and / or a decoder 200 or 500 generating a sub-block-based motion prediction signal for a current block of video. For example, the current block may include multiple sub-blocks. In block 2620, the encoder 100 or 300 and / or the decoder 200 or 500 may determine, for one or more respective or each sub-block of the current block, an extended sub-block using the sub-block-based motion prediction signal and neighboring reference samples that border and surround the respective sub-block, and may determine a spatial gradient for the respective sub-block using the determined extended sub-block. In block 2630, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined spatial gradient. At block 2640, encoder 100 or 300 and / or decoder 200 or 500 may combine (e.g., particularly add or subtract) the sub-block-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. At block 2650, encoder 100 or 300 may encode video using the refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 may decode video using the refined motion prediction signal as a prediction for the current block. In particular embodiments, the operations at blocks 2610, 2620, 2630, 2640, and 2650 may be performed for at least one block in the video (e.g., the current block).

[0228] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 copying neighboring reference samples from the reference store without any further operations. For example, determining the spatial gradient of each sub-block may use the copied neighboring reference samples to determine gradient values ​​associated with sample positions on the boundary of each sub-block. The neighboring reference samples of the extended block may be copied from the nearest integer position in the reference picture that contains the current block. In certain examples, the neighboring reference samples of the extended block have motion vectors that are rounded down from the original precision to the nearest integer.

[0229] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include encoder 100 or 300 and / or decoder 200 or 500 determining affine motion model parameters for a current block of video such that a sub-block-based motion prediction signal may be generated using the determined affine motion model parameters.

[0230] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2500, and 2600, determining the spatial gradient for each sub-block may include calculating at least one gradient value for each respective sample position in the respective sub-block. For example, calculating the at least one gradient value for each respective sample position in the respective sub-block may include applying a gradient filter to the respective sample position in the respective sub-block. As another example, calculating the at least one gradient value for each respective sample position in the respective sub-block may include determining an intensity change for each respective sample position in the respective sub-block according to an optical flow equation.

[0231] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include encoder 100 or 300 and / or decoder 200 or 500 determining a set of motion vector differential values ​​associated with sample positions of each sub-block. For example, a sub-block-based motion prediction signal may be generated and a set of motion vector differential values ​​may be determined using an affine motion model for the current block.

[0232] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, a set of motion vector difference values ​​may be determined for each sub-block of the current block and may be used to determine motion prediction refinement signals for the other remaining sub-blocks of the current block.

[0233] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, determining the spatial gradient for each sub-block may include calculating the spatial gradient using either (1) a vertical Sobel filter, (2) a horizontal Sobel filter, and / or (3) a 3-tap filter.

[0234] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, neighboring reference samples adjacent to and surrounding each sub-block may use integer motion compensation.

[0235] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the spatial gradient of each sub-block may include either a horizontal gradient or a vertical gradient. For example, a horizontal gradient may be calculated as the luminance or chrominance difference between the right-neighboring sample of each sample and the left-neighboring sample of each sample, and / or a vertical gradient may be calculated as the luminance or chrominance difference between the bottom-neighboring sample of each sample and the top-neighboring sample of each sample.

[0236] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the sub-block-based motion prediction signal may be generated using any of: (1) a four-parameter affine model; (2) a six-parameter affine model; (3) sub-block-based temporal motion vector prediction (SbTMVP) mode motion compensation; or (4) regression-based motion compensation. For example, provided that SbTMVP mode motion compensation is performed, the method may include estimating affine model parameters using the sub-block motion vector field via a linear regression operation and / or deriving pixel-level motion vectors using the estimated affine model parameters. As another example, provided that RMVF mode-based motion compensation is performed, the method may include estimating affine model parameters and / or deriving pixel-level motion vector offsets from the sub-block-level motion vectors using the estimated affine model parameters. The pixel motion vector offsets may be relative to the center of the respective sub-block.

[0237] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the refined motion prediction signal for each sub-block may be based on the actual center position of each sub-block or may be based on the sample position closest to the actual center of each sub-block.

[0238] For example, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 selecting one of (1) the actual center of each respective sub-block or (2) a sample location closest to the actual center of each respective sub-block as the center position associated with the motion prediction vector for each respective sub-block. The refined motion prediction signal may be based on the selected center position of each sub-block.

[0239] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include encoder 100 or 300 and / or decoder 200 or 500 determining center positions associated with chrominance pixels of each sub-block and determining offsets for the center positions of the chrominance pixels of each sub-block based on the chrominance position sample types associated with the chrominance pixels. The refined prediction signal for each sub-block may be based on the actual positions of the sub-block corresponding to the determined center positions of the chrominance pixels adjusted by the offset.

[0240] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, encoder 100 or 300 may generate and transmit information indicating whether optical flow prediction refinement (PROF) is enabled in one of (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header, and / or decoder 200 or 500 may receive information indicating whether PROF is enabled in one of (1) an SPS header, (2) a PPS header, or (3) a tile group header.

[0241] FIG. 27 is a flowchart illustrating a tenth exemplary encoding and / or decoding method.

[0242] 27, a representative method 2700 for encoding and / or decoding video may include, at block 2710, the encoder 100 or 300 and / or the decoder 200 or 500 determining an actual center position of each respective sub-block of the current block. At block 2720, the encoder 100 or 300 and / or the decoder 200 or 500 may generate a sub-block-based motion prediction signal or a refined motion prediction signal using the actual center position of each respective sub-block of the current block. At block 2730, (1) the encoder 100 or 300 may encode the video using the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block, or (2) the decoder 200 or 500 may decode the video using the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block. In particular embodiments, the operations in blocks 2710, 2720, and 2730 may be performed for at least one block (e.g., the current block) in the video. For example, determining the actual center position of each respective sub-block of the current block may include determining a chrominance center position associated with a chrominance pixel of the respective sub-block based on a chrominance position sample type of the chrominance pixel, and an offset of the chrominance center position relative to the center position of the respective sub-block. A sub-block-based motion prediction signal or a refined motion prediction signal for each sub-block may be based on the actual center position of the respective sub-block, which corresponds to the determined chrominance center position adjusted by the offset. While the actual center of each respective sub-block of the current block is described as being determined / used for various operations, it is contemplated that one, some, or all of such sub-block center positions may be determined / used.

[0243] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, generating the refined motion prediction signal may use the sub-block-based motion prediction signal by determining, for each respective sub-block of the current block, one or more spatial gradients of the sub-block-based motion prediction signal, determining a motion prediction refinement signal for the current block based on the determined spatial gradients, and / or combining the sub-block-based motion prediction signal and the motion prediction refinement signal to create the refined motion prediction signal for the current block. For example, determining one or more spatial gradients of the sub-block-based motion prediction signal may include determining extended sub-blocks using the sub-block-based motion prediction signal and neighboring reference samples that border and surround each sub-block, and / or determining one or more spatial gradients of each sub-block using the determined extended sub-blocks.

[0244] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, determining the spatial gradient for each sub-block may include calculating at least one gradient value for each respective sample position in the respective sub-block. For example, calculating the at least one gradient value for each respective sample position in the respective sub-block may include, for each respective sample position, applying a gradient filter to the respective sample position in the respective sub-block.

[0245] As another example, calculating at least one gradient value for each respective sample position in the respective sub-block may include determining an intensity change for one or more respective sample positions in the respective sub-block according to an optical flow equation.

[0246] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, these methods may include encoder 100 or 300 and / or decoder 200 or 500 determining a set of motion vector difference values ​​associated with sample positions of each sub-block. A sub-block-based motion prediction signal may be generated using an affine motion model for the current block, and a set of motion vector difference values ​​may be determined. In certain examples, a set of motion vector difference values ​​may be determined for each sub-block of the current block and used (e.g., reused) to determine a motion prediction refinement signal for that sub-block and other remaining sub-blocks of the current block. For example, determining the spatial gradient for each sub-block may include calculating the spatial gradient using either (1) a vertical Sobel filter, (2) a horizontal Sobel filter, and / or (3) a 3-tap filter. The neighboring reference samples that border and surround each sub-block can use integer motion compensation.

[0247] In some embodiments, the spatial gradient of each sub-block may include either a horizontal gradient or a vertical gradient. For example, the horizontal gradient may be calculated as the luminance or chrominance difference between the right-neighboring sample of each sample and the left-neighboring sample of each sample. As another example, the vertical gradient may be calculated as the luminance or chrominance difference between the bottom-neighboring sample of each sub-block and the top-neighboring sample of each sub-block.

[0248] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the sub-block-based motion prediction signal may be generated using any of: (1) a four-parameter affine model, (2) a six-parameter affine model, (3) sub-block-based temporal motion vector prediction (SbTMVP) mode motion compensation, and / or (4) regression-based motion compensation. For example, provided that SbTMVP mode motion compensation is performed, the method may include estimating affine model parameters using the sub-block motion vector field by a linear regression operation and / or deriving pixel-level motion vectors using the estimated affine model parameters. As another example, provided that regression motion vector field (RMVF) mode-based motion compensation is performed, the method may include estimating affine model parameters and / or using the estimated affine model parameters to derive pixel-level motion vector offsets from sub-block-level motion vectors, where the pixel motion vector offsets are relative to the centers of the respective sub-blocks.

[0249] In certain exemplary embodiments, including at least exemplary methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, a refined motion prediction signal may be generated using multiple motion vectors associated with control points of the current block.

[0250] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, encoder 100 or 300 can generate, encode, and transmit, and decoder 200 or 500 can receive and decode, information indicating whether optical flow prediction refinement (PROF) is enabled in one of: (1) a sequence parameter set (SPS) header; (2) a picture parameter set (PPS) header; or (3) a tile group header.

[0251] FIG. 28 is a flowchart illustrating an eleventh exemplary encoding and / or decoding method.

[0252] 28, a representative method 2800 for encoding and / or decoding video may include, at block 2810, the encoder 100 or 300 and / or the decoder 200 or 500 selecting one of (1) the actual center of each respective sub-block or (2) a sample position closest to the actual center of each respective sub-block as a center position associated with a motion prediction vector for each respective sub-block. At block 2820, the encoder 100 or 300 and / or the decoder 200 or 500 may determine the selected center position of each respective sub-block of the current block. At block 2830, the encoder 100 or 300 and / or the decoder 200 or 500 may generate a sub-block-based motion prediction signal or a refined motion prediction signal using the selected center position of each respective sub-block of the current block. In block 2840, (1) the encoder 100 or 300 may encode the video using the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block, or (2) the decoder 200 or 500 may decode the video using the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block. In particular embodiments, the operations in blocks 2810, 2820, 2830, and 2840 may be performed for at least one block in the video (e.g., the current block). Although selection of a center position is described with respect to each respective sub-block of the current block, it is contemplated that one, some, or all of the center positions of such sub-blocks may be selected / used in various operations.

[0253] FIG. 29 is a flowchart showing a representative encoding method.

[0254] 29, a representative method 2900 for encoding video may include, at block 2910, the encoder 100 or 300 performing motion estimation for a current block of video, which may include determining affine motion model parameters for the current block using an iterative motion compensation operation and generating a sub-block-based motion prediction signal for the current block using the determined affine motion model parameters. In block 2920, after performing motion estimation for the current block, the encoder 100 or 300 performs an optical flow prediction refinement (PROF) operation to generate a refined motion prediction signal. In block 2930, the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block. For example, the PROF operation may include determining one or more spatial gradients of the sub-block-based motion prediction signal, determining a motion prediction refinement signal for the current block based on the determined spatial gradients, and / or combining the sub-block-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block.

[0255] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, and 2900, a PROF operation may be performed after (e.g., only after) an iterative motion compensation operation is completed. For example, a PROF operation is not performed during motion estimation for the current block.

[0256] FIG. 30 is a flowchart illustrating another exemplary encoding method.

[0257] Referring to Figure 30, a representative method 3000 for encoding video may include, at block 3010, the encoder 100 or 300 determining affine motion model parameters using an iterative motion compensation operation during motion estimation for a current block and generating a sub-block-based motion prediction signal using the determined affine motion model parameters. At block 3020, the encoder 100 or 300 may perform an optical flow prediction refinement (PROF) operation after motion estimation for the current block to generate a refined motion prediction signal on the condition that the size of the current block meets or exceeds a threshold size. At block 3030, the encoder 100 or 300 may encode the video using (1) the refined motion prediction signal as a prediction for the current block on the condition that the current block meets or exceeds the threshold size, or (2) the sub-block-based motion prediction signal as a prediction for the current block on the condition that the current block does not meet the threshold size.

[0258] FIG. 31 is a flowchart illustrating a twelfth exemplary encoding / decoding method.

[0259] 31 , a representative method 3100 for encoding and / or decoding video may include, at block 3110, the encoder 100 or 300 determining or obtaining information indicating a size of a current block, or the decoder 200 or 500 receiving information indicating the size of the current block. At block 3120, the encoder 100 or 300 or the decoder 200 or 500 may generate a sub-block-based motion prediction signal. At block 3130, the encoder 100 or 300 or the decoder 200 or 500 may perform an optical flow prediction refinement (PROF) operation to generate a refined motion prediction signal, on the condition that the size of the current block meets or exceeds a threshold size. In block 3140, the encoder 100 or 300 can encode the video using the refined motion prediction signal as a prediction for the current block, provided that the current block meets or exceeds the threshold size, or (2) the sub-block-based motion prediction signal as a prediction for the current block, provided that the current block does not meet the threshold size, or the decoder 200 or 500 can decode the video using the refined motion prediction signal as a prediction for the current block, provided that the current block meets or exceeds the threshold size, or (2) the sub-block-based motion prediction signal as a prediction for the current block, provided that the current block does not meet the threshold size.

[0260] FIG. 32 is a flowchart illustrating a thirteenth exemplary encoding / decoding method.

[0261] 32 , a representative method 3200 for encoding and / or decoding video may include, at block 3210, the encoder 100 or 300 determining whether pixel-level motion compensation should be performed, or the decoder 200 or 500 receiving a flag indicating whether pixel-level motion compensation should be performed. At block 3220, the encoder 100 or 300 or the decoder 200 or 500 may generate a sub-block-based motion prediction signal. At block 3230, on the condition that pixel-level motion compensation should be performed, the encoder 100 or 300 or the decoder 200 or 500 may determine one or more spatial gradients of the sub-block-based motion prediction signal, determine a motion prediction refinement signal for the current block based on the determined spatial gradients, and combine the sub-block-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. In accordance with the determination of whether pixel-level motion compensation should be performed in block 3240, encoder 100 or 300 may encode the video using the sub-block-based motion prediction signal or the refined motion prediction signal as a prediction for the current block, or decoder 200 or 500 may decode the video using the sub-block-based motion prediction signal or the refined motion prediction signal as a prediction for the current block according to the indication of the flag. In particular embodiments, the operations in blocks 3220 and 3230 may be performed on a block in the video (e.g., the current block).

[0262] FIG. 33 is a flowchart illustrating a fourteenth exemplary encoding / decoding method.

[0263] 33, a representative method 3300 for encoding and / or decoding video may include, at block 3310, the encoder 100 or 300 determining or obtaining, or the decoder 200 or 500 receiving, inter-prediction weight information indicating one or more weights associated with a first reference picture and a second reference picture. At block 3320, the encoder 100 or 300 or the decoder 200 or 500 may generate a sub-block-based motion inter prediction signal for a current block of the video, may determine a first set of spatial gradients associated with the first reference picture and a second set of spatial gradients associated with the second reference picture, may determine a motion inter prediction refinement signal for the current block based on the first and second sets of spatial gradients and the inter prediction weight information, and may combine the sub-block-based motion inter prediction signal and the motion inter prediction refinement signal to create a refined motion inter prediction signal for the current block. At block 3330, the encoder 100 or 300 may encode video using the refined motion inter prediction signal as a prediction for the current block, or the decoder 200 or 500 may decode video using the refined motion inter prediction signal as a prediction for the current block. For example, the inter prediction weight information is either (1) an indicator indicating a first weighting factor to be applied to a first reference picture and / or a second weighting factor to be applied to a second reference picture, or (2) a weight index. In particular embodiments, the motion inter prediction refinement signal for the current block may be based on (1) a first gradient value derived from a first set of spatial gradients and weighted according to the first weighting factor indicated by the inter prediction weight information, and (2) a second gradient value derived from a second set of spatial gradients and weighted according to the second weighting factor indicated by the inter prediction weight information.

[0264] Exemplary Network for Implementation of Aspects 34A illustrates an example communication system 3400 in which one or more disclosed aspects may be implemented. The communication system 3400 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcasts, etc., to multiple wireless users. The communication system 3400 may enable the multiple wireless users to access such content through sharing of system resources, including wireless bandwidth. For example, the communication system 3400 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tailed unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.

[0265] 34A, the communications system 3400 may include wireless transmit / receive units (WTRUs) 3402a, 3402b, 3402c, 3402d, RANs 3404 / 3413, CNs 3406 / 3415, a public switched telephone network (PSTN) 3408, the Internet 3410, and other networks 3412, although it will be understood that the disclosed aspects contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 3402a, 3402b, 3402c, 3402d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 3402a, 3402b, 3402c, 3402d, any of which may be referred to as a “station” and / or “STA,” may be configured to transmit and / or receive wireless signals and may include user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable, a head-mounted display (HMD), a vehicle, a drone, a medical device and application (e.g., remote surgery), an industrial device and application (e.g., robots and / or other wireless devices operating in an industrial and / or automated processing chain context), a consumer electronic device, a device operating on a commercial and / or industrial wireless network, etc. The WTRUs 3402a, 3402b, 3402c, and 3402d may all be referred to interchangeably as UEs.

[0266] The communications system 3400 may also include a base station 3414a and / or a base station 3414b. Each of the base stations 3414a, 3414b may be any type of device configured to wirelessly interface with at least one of the WTRUs 3402a, 3402b, 3402c, 3402d to facilitate access to one or more communications networks, such as the CN 3406 / 3415, the Internet 3410, and / or the network 3412. By way of example, the base stations 3414a, 3414b may be a Base Transceiver Station (BTS), a Node B, an eNodeB(end), a Home Node B (HNB), a Home eNodeB (HeNB), a gNB, an NR Node B, a site controller, an access point (AP), a wireless router, etc. Although the base stations 3414a, 3414b are each shown as a single element, it will be appreciated that the base stations 3414a, 3414b may include any number of interconnected base stations and / or network elements.

[0267] The base station 3414a may be part of the RAN 3404 / 3413, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 3414a and / or base station 3414b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for wireless service for a particular geographic area, which may be relatively fixed or may change over time. A cell may be further divided into cell sectors. For example, the cell associated with the base station 3414a may be divided into three sectors. Thus, in one embodiment, the base station 3414a may include three transceivers, i.e., one transceiver for each sector of the cell. In an embodiment, the base station 3414a may utilize multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers per sector of the cell. For example, beamforming may be used to transmit and / or receive signals in desired spatial directions.

[0268] The base stations 3414a, 3414b may communicate with one or more of the WTRUs 3402a, 3402b, 3402c, 3402d over an air interface 3416, which may be any suitable wireless communications link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 3416 may be established using any suitable radio access technology (RAT).

[0269] More specifically, as noted above, the communication system 3400 may be a multiple-access system and may utilize one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base stations 3414a and the WTRUs 3402a, 3402b, and 3402c in the RAN 3404 / 3413 may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 3415 / 3416 / 3417 using Wideband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High Speed ​​UL Packet Access (HSUPA).

[0270] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 3416 using Long Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).

[0271] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement a radio technology such as NewRadio (NR) radio access that may establish the air interface 3416 using NR.

[0272] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement multiple radio access technologies. For example, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may jointly implement LTE and NR radio access, e.g., using a dual connectivity (DC) principle. Thus, the air interface utilized by the WTRUs 3402a, 3402b, 3402c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., end and gNB).

[0273] In other embodiments, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement a radio technology such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0274] The base station 3414b in FIG. 34A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a workplace, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by drones), and a roadway. In one embodiment, the base station 3414b and the WTRUs 3402c, 3402d may implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 3414b and the WTRUs 3402c, 3402d may implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 3414b and the WTRUs 3402c, 3402d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. 34A, base station 3414b may have a direct connection to the Internet 3410. Thus, base station 3414b may not be required to access the Internet 3410 via CN 3406 / 3415.

[0275] The RAN 3404 / 3413 can communicate with the CN 3406 / 3415, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 3402a, 3402b, 3402c, 3402d. Data may have various Quality of Service (QoS) requirements, such as different throughput, latency, error tolerance, reliability, data throughput, and mobility requirements. The CN 3406 / 3415 can provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, such as user authentication. 34A, it will be understood that the RAN 1084 / 3413 and / or the CN 3406 / 3415 may communicate directly or indirectly with other RANs that utilize the same RAT as the RAN 3404 / 3413 or a different RAT. For example, in addition to being connected to the RAN 3404 / 3413, which may utilize NR radio technology, the CN 3406 / 3415 may also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0276] The CNs 3406 / 3415 may also act as gateways for the WTRUs 3402a, 3402b, 3402c, 3402d to access the PSTN 3408, the Internet 3410, and / or other networks 3412. The PSTN 3408 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 3410 may include a global system of interconnected computer networks and devices that use common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet protocol suite. The networks 3412 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the networks 3412 may include another CN connected to one or more RANs that may utilize the same RAT as the RANs 3404 / 3413 or a different RAT.

[0277] Some or all of the WTRUs 3402a, 3402b, 3402c, 3402d in the communications system 3400 may include multi-mode capabilities (e.g., the WTRUs 3402a, 3402b, 3402c, 3402d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU 3402c shown in FIG. 34A may be configured to communicate with a base station 3414a that can employ cellular-based wireless technology and a base station 3414b that can employ IEEE 802 wireless technology.

[0278] 34B is a system diagram illustrating an example WTRU 3402. As shown in FIG. 34B, the WTRU 3402 may include a processor 3418, a transceiver 3420, a transmit / receive element 3422, a speaker / microphone 3424, a keypad 3426, a display / touchpad 3428, non-removable memory 3430, removable memory 3432, a power source 3434, a global positioning system (GPS) chipset 3436, and / or other peripherals 3438, etc. It will be understood that the WTRU 3402 may include any sub-combination of the above elements while remaining consistent with an embodiment.

[0279] The processor 3418 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors in conjunction with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 3418 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 3402 to operate in a wireless environment. The processor 3418 may be coupled to a transceiver 3420, which may be coupled to the transmit / receive element 3422. While FIG. 34B depicts the processor 3418 and the transceiver 3420 as separate components, it will be understood that the processor 3418 and the transceiver 3420 may be integrated together in an electronic package or chip. The processor 3418 may be configured to encode or decode video (e.g., video frames).

[0280] The transmit / receive element 3422 may be configured to transmit signals to or receive signals from a base station (e.g., base station 3414a) over the air interface 3416. For example, in one embodiment, the transmit / receive element 3422 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 3422 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In yet another embodiment, the transmit / receive element 3422 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 3422 may be configured to transmit and / or receive any combination of wireless signals.

[0281] 34B as a single element, the WTRU 3402 may include any number of transmit / receive elements 3422. More specifically, the WTRU 3402 may utilize MIMO technology. Thus, in one embodiment, the WTRU 3402 may include two or more transmit / receive elements 3422 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 3416.

[0282] The transceiver 3420 may be configured to modulate signals transmitted by the transmit / receive element 3422 and demodulate signals received by the transmit / receive element 3422. As mentioned above, the WTRU 3402 may have multi-mode capabilities. Thus, the transceiver 3420 may include multiple transceivers to enable the WTRU 3402 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.

[0283] The processor 3418 of the WTRU 3402 may be coupled to and may receive user input data from a speaker / microphone 3424, a keypad 3426, and / or a display / touchpad 3428 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 3418 may also output user data to the speaker / microphone 3424, the keypad 3426, and / or the display / touchpad 3428. Additionally, the processor 3418 may access information from and store data in any type of suitable memory, such as non-removable memory 3430 and / or removable memory 3432. The non-removable memory 3430 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 3432 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 3418 may access information from and store data in memory that is not physically located on the WTRU 3402, such as on a server or home computer (not shown).

[0284] The processor 3418 may receive power from a power source 3434 and may be configured to distribute and / or control the power to other components within the WTRU 3402. The power source 3434 may be any suitable device for powering the WTRU 3402. For example, the power source 3434 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0285] The processor 3418 may be coupled to a GPS chipset 3436, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 3402. In addition to or in place of information from the GPS chipset 3436, the WTRU 3402 may receive location information from a base station (e.g., base stations 3414a, 3414b) over the air interface 3416 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 3402 may acquire location information by way of any suitable location-determination method while remaining consistent with an embodiment.

[0286] The processor 3418 may further be coupled to other peripherals 3438, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 3438 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, and the like. The peripheral device 3438 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0287] The processor 3418 of the WTRU 3402 may be in operative communication with various peripherals 3438 including, for example, one or more accelerometers, one or more gyroscopes, a USB port, other communication interfaces / ports, a display, and / or any other visual / audio indicators to implement the exemplary embodiments disclosed herein.

[0288] The WTRU 3402 may include a full-duplex radio, in which transmission and reception of some or all of the signals associated with a particular subframe (e.g., for both the UL (e.g., for transmission) and downlink (e.g., for reception)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit for reducing or substantially eliminating self-interference through signal processing by hardware (e.g., a choke) or a processor (e.g., a separate processor (not shown) or the processor 3418). In an embodiment, the WTRU 3402 may include a half-duplex radio for transmission and reception of some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or downlink (e.g., for reception)).

[0289] 34C is a system diagram illustrating the RAN 104 and the CN 3406 according to an embodiment. As noted above, the RAN 3404 can communicate with the WTRUs 3402a, 3402b, and 3402c over the air interface 3416 using E-UTRA radio technology. The RAN 3404 can also communicate with the CN 3406.

[0290] While the RAN 3404 may include eNodeBs 3460a, 3460b, and 3460c, it will be appreciated that the RAN 3404 may include any number of eNodeBs while remaining consistent with an embodiment. The eNodeBs 3460a, 3460b, and 3460c may each include one or more transceivers for communicating with the WTRUs 3402a, 3402b, and 3402c over the air interface 3416. In one embodiment, the eNodeBs 3460a, 3460b, and 3460c may implement MIMO technology. Thus, the eNodeB 3460a, for example, may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 3402a.

[0291] Each of the eNodeBs 3460a, 3460b, 3460c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, and scheduling of users in the UL and / or DL, etc. As shown in Figure 34C, the eNodeBs 3460a, 3460b, 3460c may communicate with each other over an X2 interface.

[0292] The CN 3406 shown in Figure 34C may include a mobility management entity (MME) 3462, a serving gateway (SGW) 3464, and a packet data network (PDN) gateway (or PGW) 3466. While each of the above elements is shown as part of the CN 3406, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.

[0293] The MME 3462 may be connected to each of the eNodeBs 3460a, 3460b, 3460c in the RAN 3404 via an S1 interface and may act as a control node. For example, the MME 3462 may be responsible for authenticating users of the WTRUs 3402a, 3402b, 3402c, bearer activation / deactivation, selecting a specific serving gateway during initial attach of the WTRUs 3402a, 3402b, 3402c, etc. The MME 3462 may provide a control plane function for switching between the RAN 3404 and other RANs (not shown) that employ other radio technologies, such as GSM and / or WCDMA.

[0294] The SGW 3464 may be connected to each of the eNodeBs 3460a, 3460b, 3460c in the RAN 104 via an S1 interface. The SGW 3464 may generally route and forward user data packets to / from the WTRUs 3402a, 3402b, 3402c. The SGW 3464 may also perform other functions, such as anchoring the user plane during inter-eNodeB handover, triggering paging when DL data is available to the WTRUs 3402a, 3402b, 3402c, and managing and storing the context of the WTRUs 3402a, 3402b, 3402c.

[0295] The SGW 3464 may be connected to a PGW 3466, which may provide the WTRUs 3402a, 3402b, 3402c with access to packet-switched networks, such as the Internet 3410, to facilitate communications between the WTRUs 3402a, 3402b, 3402c and IP-enabled devices.

[0296] The CN 3406 may facilitate communications with other networks. For example, the CN 106 may provide the WTRUs 3402a, 3402b, 3402c with access to circuit-switched networks, such as the PSTN 3408, to facilitate communications between the WTRUs 3402a, 3402b, 3402c and traditional land-line communications devices. For example, the CN 3406 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the CN 3406 and the PSTN 3408. In addition, the CN 3406 may provide the WTRUs 3402a, 3402b, 3402c with access to other networks 3412, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0297] Although the WTRUs are described in Figures 34A-34D as wireless terminals, it is contemplated that in certain representative embodiments, such terminals may use a wired communication interface (e.g., temporarily or permanently) with the communication network.

[0298] In an exemplary embodiment, the other network 3412 may be a WLAN.

[0299] A WLAN in infrastructure basic service set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access to or interface with a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic to a STA originating from outside the BSS may arrive through the AP and be delivered to the STA. Traffic originating from a STA destined for a destination outside the BSS may be sent to the AP for delivery to the respective destination. Traffic between STAs within a BSS may be sent through the AP; e.g., a source STA may send traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within a BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between (e.g., directly between) a source STA and a destination STA using direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and the STAs within or using the IBSS (e.g., all of the STAs) may communicate directly with each other. The IBSS communication mode may sometimes be referred to herein as an "ad hoc" communication mode.

[0300] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a 20 MHz wide bandwidth) or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In certain representative embodiments, carrier sense multiple access with collision avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. In CSMA / CA, STAs (e.g., all STAs), including the AP, may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., a single station) may transmit at any given time in a given BSS.

[0301] High-throughput (HT) STAs may use 40 MHz wide channels for communication, for example, using a combination of adjacent or non-adjacent 20 MHz channels and the primary 20 MHz channel to form a 40 MHz wide channel.

[0302] A Very High Throughput (VHT) STA can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz channels can be formed by combining contiguous 20 MHz channels. A 160 MHz channel can be formed by combining eight contiguous 20 MHz channels or by combining two non-contiguous 80 MHz channels, sometimes referred to as an 80+80 configuration. In the 80+80 configuration, after channel encoding, the data may be passed through a segment parser that can partition the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing may be performed separately on each stream. The streams may be mapped onto two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80+80 configuration may be reversed, and the combined data may be sent to the Medium Access Control (MAC).

[0303] Sub-1 GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidths and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah can support meter-type control / machine-type communications, such as MTC devices, within a macro coverage area. MTC devices may have limited capabilities, including specific capabilities, such as support for (e.g., only) specific and / or limited bandwidths. MTC devices may include batteries with above-threshold battery life (e.g., to maintain very long battery life).

[0304] WLAN systems, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, that may support multiple channels and channel bandwidths include a channel that may be designated as a primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In an 802.11ah example, the primary channel may be 1 MHz wide for STAs (e.g., MTC-type devices) that support (e.g., only) the 1 MHz mode, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or network allocation vector (NAV) setting may depend on the status of the primary channel. For example, if a STA transmitting to an AP (that only supports 1 MHz operating mode) has a busy primary channel, the entire available frequency band may be considered busy even though most of the frequency band may remain idle and available.

[0305] In the United States, the available frequency band that can be used by 802.11ah is 902 MHz to 928 MHz. In South Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. The total available bandwidth for 802.11ah is 6 MHz to 26 MHz depending on the country code.

[0306] 34D is a system diagram illustrating a RAN 3413 and a CN 3415 according to an embodiment. As described above, the RAN 3413 can communicate with the WTRUs 3402a, 3402b, and 3402c over the air interface 3416 using NR radio technology. The RAN 3413 can also communicate with the CN 3415.

[0307] While the RAN 3413 may include gNBs 3480a, 3480b, and 3480c, it will be understood that the RAN 3413 may include any number of gNBs while remaining consistent with an embodiment. The gNBs 3480a, 3480b, and 3480c may each include one or more transceivers for communicating with the WTRUs 3402a, 3402b, and 3402c over the air interface 3416. In one embodiment, the gNBs 3480a, 3480b, and 3480c may implement MIMO technology. For example, the gNBs 3480a, 3480b may utilize beamforming to transmit signals to and / or receive signals from the gNBs 3480a, 3480b, and 3480c. Thus, the gNB 3480a may, for example, use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 3402a. In an embodiment, the gNBs 3480a, 3480b, 3480c may implement carrier aggregation technology. For example, the gNB 3480a may transmit multiple component carriers to the WTRU 3402a (not shown). A subset of these component carriers may be on an unlicensed spectrum, while the remaining component carriers may be on a licensed spectrum. In an embodiment, the gNBs 3480a, 3480b, 3480c may implement coordinated multipoint (CoMP) technology. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 3480a and 3480b (and / or gNB 3480c).

[0308] The WTRUs 3402a, 3402b, 3402c can communicate with the gNBs 480a, 3480b, 3480c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. The WTRUs 3402a, 3402b, 3402c can communicate with the gNBs 3480a, 3480b, 3480c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., including different numbers of OFDM symbols and / or lasting through different lengths of absolute time).

[0309] The gNBs 3480a, 3480b, 3480c may be configured to communicate with the WTRUs 3402a, 3402b, 3402c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 3402a, 3402b, 3402c may communicate with the gNBs 3480a, 3480b, 3480c without accessing another RAN (e.g., eNodeBs 3460a, 3460b, 3460c, etc.). In a standalone configuration, the WTRUs 3402a, 3402b, 3402c may utilize one or more of the gNBs 3480a, 3480b, 3480c as mobility anchor points. In a standalone configuration, the WTRUs 3402a, 3402b, 3402c may communicate with the gNBs 3480a, 3480b, 3480c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 3402a, 3402b, 3402c may communicate / connect with the gNBs 3480a, 3480b, 3480c while also communicating / connecting with another RAN, such as eNodeBs 3460a, 3460b, 3460c. For example, the WTRUs 3402a, 3402b, 3402c may implement DC principles to communicate with one or more gNBs 3480a, 3480b, 3480c and one or more eNodeBs 3460a, 3460b, 3460c substantially simultaneously. In a non-standalone configuration, the eNodeBs 3460a, 3460b, 3460c may act as mobility anchors for the WTRUs 3402a, 3402b, 3402c, and the gNBs 3480a, 3480b, 3480c may provide additional coverage and / or throughput for serving the WTRUs 3402a, 3402b, 3402c.

[0310] Each of the gNBs 3480a, 3480b, 3480c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data towards user plane functions (UPFs) 3484a, 3484b, and routing of control plane information towards access and mobility management functions (AMFs) 3482a, 3482b, etc. As shown in FIG. 34D, the gNBs 3480a, 3480b, 3480c may communicate with one another via an Xn interface.

[0311] The CN 3415 shown in Figure 34D may include at least one AMF 3482a, 3482b, at least one UPF 3484a, 3484b, at least one Session Management Function (SMF) 3483a, 3483b, and possibly a Data Network (DN) 3485a, 3485b. While each of the above elements is shown as part of the CN 3415, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.

[0312] The AMF 3482a, 3482b may be connected to one or more of the gNBs 3480a, 3480b, 3480c in the RAN 3413 via an N2 interface and may act as a control node. For example, the AMF 3482a, 3482b may be responsible for authenticating users of the WTRUs 3402a, 3402b, 3402c, supporting network slicing (e.g., handling different protocol data unit (PDU) sessions with different requirements), selecting a particular SMF 3483a, 3483b, managing registration areas, terminating non-access stratum (NAS) signaling, mobility management, etc. Network slicing may be used by the AMF 3482a, 3482b to customize the CN support of the WTRUs 3402a, 3402b, 3402c based on the type of service being utilized for the WTRUs 3402a, 3402b, 3402c. For example, different network slices may be established for different use cases, such as services relying on Ultra-Reliable Low-Latency Communications (URLLC) access, services relying on enhanced Mobile (e.g., High-Capacity Mobile) Broadband (eMBB) access, and / or services with Machine-Type Communications (MTC) access. The AMF 3462 may provide a control plane function for switching between the RAN 3413 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.

[0313] The SMFs 3483a and 3483b may be connected to the AMFs 3482a and 3482b in the CN 3415 via an N11 interface. The SMFs 3483a and 3483b may also be connected to the UPFs 3484a and 3484b in the CN 3415 via an N4 interface. The SMFs 3483a and 3483b may select and control the UPFs 3484a and 3484b and configure traffic routing through the UPFs 3484a and 3484b. The SMFs 3483a and 3483b may perform other functions, such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notification. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.

[0314] The UPFs 3484a, 3484b may be connected to one or more of the gNBs 3480a, 3480b, 3480c in the RAN 3413 via an N3 interface, which may provide the WTRUs 3402a, 3402b, 3402c with access to packet-switched networks, such as the Internet 3410, to facilitate communications between the WTRUs 3402a, 3402b, 3402c and IP-enabled devices. The UPFs 3484, 3484b may perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, etc.

[0315] The CN 3415 may facilitate communication with other networks. For example, the CN 3415 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 3415 and the PSTN 408. In addition, the CN 3415 may provide the WTRUs 3402a, 3402b, 3402cc with access to other networks 3412, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 3402a, 3402b, 3402c may be connected to local data networks (DNs) 3485a, 3485b through the UPFs 3484a, 3484b via an N3 interface to the UPFs 3484a, 3484b and an N6 interface between the UPFs 3484a, 3484b and the DNs 3485a, 3485b.

[0316] 34A-34D and the corresponding description of Figures 34A-34D, one or more or all of the functions described herein with respect to one or more of the WTRUs 3402a-d, base stations 3414a-b, eNodeBs 3460a-c, MME 3462, SGW 3464, PGW 3466, gNBs 3480a-c, AMFs 3482a-b, UPFs 3484a-b, SMFs 3483a-b, DNs 3485a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or simulate network and / or WTRU functionality.

[0317] The emulation device may be designed to perform one or more tests of other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one or more or all functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communications network to test other devices in the communications network. One or more emulation devices may perform one or more or all functions while temporarily implemented / deployed as part of a wired and / or wireless communications network. The emulation device may be directly coupled to another device for testing and / or may perform testing using over-the-air wireless communications.

[0318] The one or more emulation devices may perform one or more functions, inclusive, without being implemented / deployed as part of a wired and / or wireless communications network. For example, the emulation devices may be utilized in test scenarios in a testing laboratory and / or in an undeployed (e.g., test) wired and / or wireless communications network to perform tests of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may, for example, include one or more antennas) may be used by the emulation devices to transmit and / or receive data.

[0319] The HEVC standard offers approximately 50% bitrate savings at comparable perceptual quality compared to the traditional video coding standards H.264 / MPEG AVC. While the HEVC standard offers significant coding improvements over its predecessor, further coding efficiency improvements can be achieved using additional coding tools. For example, the Joint Video Exploration Team (JVET) initiated a project to develop a new generation video coding standard called Versatile Video Coding (VVC) to provide such coding efficiency improvements. A reference software code base called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard. To facilitate the evaluation of new coding tools, another reference software base called the benchmark set (BMS) was also created. In addition to the VTM, the BMS code base includes a list of additional coding tools that offer higher coding efficiency and moderate implementation complexity, and will be used as a benchmark for evaluating similar coding technologies in the VVC standardization process. JEM coding tools integrated into BMS-2.0 (e.g., 4x4 non-separable secondary transform (NSST), generalized bi-prediction (GBi), bidirectional optical flow (BIO), decoder-side motion vector refinement (DMVR), and current picture referencing (CPR) as well as quantization tools for trellis coding.

[0320] Systems and methods for processing data according to representative embodiments may be performed by one or more processors executing sequences of instructions contained in a memory device. Such instructions may be loaded into the memory device from another computer-readable medium, such as a secondary data storage device. Execution of the sequences of instructions contained in the memory device causes the processor to operate, for example, as described above. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions to implement one or more embodiments. Such software may run on a processor housed within a robotic assistance / apparatus (RAA) and / or remotely within another mobile device. In the latter case, data may be transferred wired or wirelessly between the RAA or other mobile device including the sensor and a remote device including a processor executing software that performs scale estimation and compensation as described above. According to other representative embodiments, some of the processing described above with respect to localization may be performed in the device including the sensor / camera, while the remainder of the processing may be performed in the second device after receiving partially processed data from the device including the sensor / camera.

[0321] While features and elements are described above in particular combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. Additionally, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). Software and an associated processor may be used to implement a radio frequency transceiver for use in the WTRU 3402, a UE, a terminal, a base station, an RNC, or any host computer.

[0322] Furthermore, in the above-described embodiments, reference is made to processing platforms, computing systems, controllers, and other devices that include processors. These devices may include at least one central processing unit ("CPU") and memory. References to symbolic representations of operations and operations or instructions may be performed by various CPUs and memories, in accordance with the practices of those skilled in the art of computer programming. Such operations and operations or instructions may be referred to as "executing," "executing on a computer," or "executing on a CPU."

[0323] Those skilled in the art will understand that the operations and symbolically expressed operations or instructions include the manipulation of electrical signals by a CPU. The electrical system represents data bits, which can result in a resulting transformation or reduction of the electrical signals, and the maintenance of the data bits in memory locations within a memory system, thereby reconfiguring or otherwise altering the CPU's operations and processing of other signals. The memory locations in which the data bits are maintained are physical locations that have particular electrical, magnetic, optical, or organic properties that correspond to or represent the data bits. It should be understood that exemplary embodiments are not limited to the platforms or CPUs described above, and that other platforms and CPUs may support the provided methods.

[0324] Additionally, data bits are maintained on computer-readable media, including magnetic disks, optical disks, and any other volatile (e.g., random access memory ("RAM")) or non-volatile (e.g., read-only memory ("ROM")) mass storage system readable by a CPU. The computer-readable media may include cooperating or interconnected computer-readable media, which reside solely on a processing system or are distributed among multiple interconnected processing systems, which may be local or remote to a processing system. It will be understood that representative embodiments are not limited to the memories described above, and that other platforms and memories may support the described methods. It will be understood that representative embodiments are not limited to the platforms and CPUs described above, and that other platforms and CPUs may support the described methods.

[0325] In an example embodiment, any of the operations, processes, etc. described herein may be implemented as computer-readable instructions stored on a computer-readable medium, which may be executed by a processor of a mobile unit, a network element, and / or any other computing device.

[0326] There is little distinction left between hardware and software implementations of aspects of the system. The use of hardware or software is generally (though not always, as the choice between hardware and software can be important in certain situations) a design choice representing a cost-effectiveness trade-off. There may be various means (e.g., hardware, software, and / or firmware) by which the processes and / or systems and / or other techniques described herein may be effected, and the preferred means may vary depending on the context in which the processes and / or systems and / or other techniques are deployed. For example, if an implementer determines that speed and accuracy are most important, the implementer may choose a primarily hardware and / or firmware means. If flexibility is most important, the implementer may choose a primarily software implementation. Alternatively, the implementer may choose some combination of hardware, software, and / or firmware.

[0327] The above detailed description illustrates various embodiments of devices and / or processes through the use of block diagrams, flowcharts, and / or examples. To the extent that such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, those skilled in the art will understand that each function and / or operation in such block diagrams, flowcharts, or examples can be individually and / or collectively implemented by various hardware, software, firmware, or substantially any combination thereof. Suitable processors include, by way of example, general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors in association with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.

[0328] While features and elements are provided above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with the other features and elements. The present disclosure is not limited in terms of the specific embodiments described herein, which are intended as illustrations of various aspects. It will be apparent to those skilled in the art that many modifications and variations can be made without departing from its spirit and scope. No element, act, or instruction used in the description of the present application should be construed as critical or essential to an embodiment unless expressly indicated as such. Functionally equivalent methods and apparatuses within the scope of the present disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to be included within the scope of the appended claims. The present disclosure should be limited only by the terms of the appended claims, and the full range of equivalents to which such claims are entitled. It should be understood that the present disclosure is not limited to any particular method or system.

[0329] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the terms “station” and its abbreviation “STA,” “user equipment” and its abbreviation “UE” may mean (i) a wireless transmit and / or receive unit (WTRU) as described below, (ii) any of multiple embodiments of a WTRU as described below, (iii) a wireless-enabled and / or wired-enabled (e.g., tetherable) device configured specifically to have some or all of the structure and functionality of a WTRU as described below, (iii) a wireless-enabled and / or wired-enabled device configured to have less than all of the structure and functionality of a WTRU as described below, or (iv) the like. Details of an example WTRU that may represent any UE described herein are provided below with reference to Figures 34A-34D.

[0330] In certain exemplary embodiments, portions of the subject matter described herein may be implemented using application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated forms. However, some aspects of the aspects disclosed herein may equivalently be implemented in an integrated circuit, in whole or in part, as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or as substantially any combination thereof, and those skilled in the art will understand that designing circuits and / or writing code for software and / or firmware is well within the skill of those skilled in the art in light of this disclosure. Those skilled in the art will also understand that the mechanisms of the subject matter described herein can be distributed as program products in various forms, and that exemplary embodiments of the subject matter described herein apply regardless of the particular type of signal-bearing medium used to actually effect the distribution. Examples of signal-bearing media include, but are not limited to, recordable-type media such as floppy disks, hard disk drives, CDs, DVDs, digital tape, computer memory, and transmission-type media such as digital and / or analog communications media (e.g., fiber optic cables, wave guides, wired communications links, wireless communications links, etc.).

[0331] The subject matter described herein may illustrate different components contained within or connected to different other components. It will be understood that such depicted configurations are merely examples, and that in fact many other architectures that achieve the same functionality may be implemented. In a conceptual sense, components in any arrangement to achieve the same functionality are substantially “associated” such that the desired functionality can be achieved. Thus, any two components combined herein to achieve specific functionality may be considered to be “associated” with each other such that the desired functionality is achieved, regardless of the architecture or intervening components. Similarly, any two components so associated may be considered to be “operably connected” or “operably coupled” to each other to achieve the desired functionality, and any two components capable of being so associated may be considered to be “operably coupleable” to each other to achieve the desired functionality. Specific examples of operably coupleable include, but are not limited to, components that are physically engageable and / or physically interacting, and / or components that are wirelessly interacting and / or wirelessly interacting, and / or components that are logically interacting and / or logically interacting.

[0332] With respect to the use of virtually any plural and / or singular term herein, those skilled in the art will be able to convert from plural to singular and / or from singular to plural as appropriate depending on the context and / or application. For clarity, various singular / plural permutations may be expressly set forth herein.

[0333] Those skilled in the art will understand that, in general, the terms used in this specification, and particularly in the appended claims (e.g., the body of the appended claims), are generally intended as “open” terms (e.g., the term “comprising” should be interpreted as “including, but not limited to,” the term “having” should be interpreted as “having at least,” the term “including” should be interpreted as “including, but not limited to,” etc.). Furthermore, where a specific number is intended by the introduced claim language, such intention will be explicitly set forth in the claim; in the absence of such a statement, those skilled in the art will understand that no such intention exists. For example, where only one element is intended, the term “single” or similar language may be used. As an aid to understanding, the following appended claims and / or description of this specification may include the use of the introductory phrases “at least one” and “one or more” to introduce claim language. However, the use of such phrases should not be construed as implying that the introduction of a claim recitation by the indefinite article "a" or "an" limits any particular claim containing such an introduced claim recitation to embodiments containing only one such recitation, even when the same claim includes the introductory phrase "one or more" or "at least one" and the indefinite article "a" or "an" (e.g., "a" and / or "an" should be construed to mean "at least one" or "one or more"). The same applies to the use of definite articles used to introduce claim recitations. Those skilled in the art will also recognize that even when a specific number of introduced claim recitations is explicitly recited, such recitation should be construed to mean at least the recited number (e.g., the mere recitation of "two recitations" without any other modifier means at least two recitations or more than two recitations).Furthermore, in instances where a convention similar to "at least one of A, B, and C, etc." is used, such syntax is generally intended in the sense that one of ordinary skill in the art would understand that convention (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). In instances where a convention similar to "at least one of A, B, or C, etc." is used, such syntax is generally intended in the sense that one of ordinary skill in the art would understand that convention (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems having A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Furthermore, those skilled in the art will understand that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of those terms, either of those terms, or both terms. For example, the phrase "A or B" will be understood to include the possibility of "A" or "B," or "A and B." Furthermore, as used herein, the term "any of," followed by a listing of multiple elements and / or multiple categories of elements, is intended to include "any," "any combination," "any plurality," and / or "any combination of multiple" of the elements and / or categories of elements, separately or in conjunction with other elements and / or other categories of elements. Furthermore, as used herein, the term "set" or "group" is intended to include any number of elements, including zero. Furthermore, as used herein, the term "number" is intended to include any number, including zero.

[0334] Furthermore, when features or aspects of the disclosure are described in terms of a Markush group, those skilled in the art will understand that the disclosure is also described in terms of any individual element or subgroup of elements of the Markush group.

[0335] As will be understood by those skilled in the art, for all purposes, e.g., in terms of providing a written description, all ranges disclosed herein encompass all possible subranges and combinations of subranges. Any recited range can be readily recognized as fully describing and allowing for the same range to be broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third, and upper third, etc. As will also be understood by those skilled in the art, all terms such as "up to," "at least," "greater than," "less than," etc., refer to ranges that are inclusive of the recited numbers and that can then be broken down into subranges as described above. Finally, as will be understood by those skilled in the art, ranges include each individual element. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, and so on.

[0336] Moreover, the claims should not be construed as limited to the order or elements provided unless expressly stated to that effect. Also, the use of the term "means for" in any claim is intended to invoke 35 U.S.C. § 112(6), i.e., means-plus-function claim format, and any claim without the term "means for" is not so intended.

[0337] A processor in association with software may be used to implement a wireless transmit receive unit (WTRU), user equipment (UE), terminal, base station, mobility management entity (MME) or evolved packet core (EPC), or radio frequency transceiver for use in any host computer. The WTRU may be used in association with modules implemented in hardware and / or software, including software defined radios (SDRs), as well as other components, such as cameras, video camera modules, videophones, speakerphones, vibration devices, speakers, microphones, television transceivers, hands-free headsets, keyboards, Bluetooth modules, frequency modulation (FM) radio units, near field communication (NFC) modules, liquid crystal display (LCD) display units, organic light emitting diode (OLED) display units, digital music players, media players, video game player modules, internet browsers, and / or wireless local area network (WLAN) or ultra-wideband (UWB) modules.

[0338] Throughout this disclosure, those skilled in the art will understand that certain exemplary embodiments may be used in the alternative or in combination with other exemplary embodiments.

[0339] Additionally, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). Software and associated processors may be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer. [Explanation of symbols]

[0340] 3400 Communication Systems 3402a WTRU 3402b WTRU 3402c WTRU 3402d WTRU 3408 PSTN 3410 Internet 3412 Network 3414a base station 3414b base station 3416 Air Interface 3418 processor 3420 Walkie-Talkie 3422 receiving element 3424 Microphone 3426 keypad 3428 Touchpad 3430 Non-removable Memory 3432 Removable Memory 3434 Power supply 3436 chipset 3438 Peripheral Equipment 3460a eNodeB 3460b eNodeB 3460c eNodeB 3462 MME 3464 SGW 3466 PGW 3482a AMF 3482b AMF 3483a SMF 3483b SMF 3484a UPF 3484b UPF 3485a DN 3485b DN

Claims

1. 1. A method of decoding video, comprising: generating a sub-block-based motion prediction signal for a sub-block of the block of the picture based on an affine motion model associated with the block; determining a set of pixel-level motion vector differential values ​​for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block-based motion prediction signal for each sample position of the sub-block, wherein an extended sub-block is formed to include the sub-block-based motion prediction signal and a plurality of samples surrounding the sub-block, each of the plurality of samples surrounding the sub-block being obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; decoding the video using the refined motion prediction signal; and A method comprising:

2. 2. The method of claim 1, wherein the integer motion compensation is based on integer portions of motion vectors of the sub-blocks.

3. 2. The method of claim 1, wherein the integer motion compensation is based on the nearest integer motion vector of the motion vector of the sub-block.

4. 2. The method of claim 1, wherein the sub-block is expanded by one sample in each direction forming the expanded sub-block.

5. 10. A computer-readable medium having instructions stored thereon for decoding video data according to the method of claim 1.

6. 1. A method of encoding video, comprising: generating a sub-block-based motion prediction signal for a sub-block of the block of the picture based on an affine motion model associated with the block; determining a set of pixel-level motion vector differential values ​​for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block-based motion prediction signal for each sample position of the sub-block, wherein an extended sub-block is formed to include the sub-block-based motion prediction signal and a plurality of samples surrounding the sub-block, each of the plurality of samples surrounding the sub-block being obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; encoding the video using the refined motion prediction signal; and A method comprising:

7. 7. The method of claim 6, wherein the integer motion compensation is based on integer portions of motion vectors of the sub-blocks.

8. 7. The method of claim 6, wherein the integer motion compensation is based on the nearest integer motion vector of the motion vector of the sub-block.

9. 7. The method of claim 6, wherein the sub-block is expanded by one sample in each direction forming the expanded sub-block.

10. 7. A computer-readable medium having instructions stored thereon for encoding video data according to the method of claim 6.

11. 1. An apparatus for decoding video, comprising: generating a sub-block-based motion prediction signal for a sub-block of the block of the picture based on an affine motion model associated with the block; determining a set of pixel-level motion vector differential values ​​for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block-based motion prediction signal for each sample position of the sub-block, wherein an extended sub-block is formed to include the sub-block-based motion prediction signal and a plurality of samples surrounding the sub-block, each of the plurality of samples surrounding the sub-block being obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; Decoding the video using the refined motion estimation signal. An apparatus comprising a processor configured to:

12. The apparatus of claim 11 , wherein the integer motion compensation is based on integer portions of motion vectors of the sub-blocks.

13. The apparatus of claim 11 , wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the sub-block.

14. 12. The apparatus of claim 11, wherein the sub-block is expanded by one sample in each direction forming the expanded sub-block.

15. 1. An apparatus for encoding video, comprising: generating a sub-block-based motion prediction signal for a sub-block of the block of the picture based on an affine motion model associated with the block; determining a set of pixel-level motion vector differential values ​​for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block-based motion prediction signal for each sample position of the sub-block, wherein an extended sub-block is formed to include the sub-block-based motion prediction signal and a plurality of samples surrounding the sub-block, each of the plurality of samples surrounding the sub-block being obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; Encoding the video using the refined motion estimation signal. An apparatus comprising a processor configured to:

16. The apparatus of claim 15, wherein the integer motion compensation is based on integer portions of motion vectors of the sub-blocks.

17. 16. The apparatus of claim 15, wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the sub-block.

18. 16. The apparatus of claim 15, wherein the sub-block is expanded by one sample in each direction forming the expanded sub-block.

Citation Information

Patent Citations

  • Motion-compensation prediction based on BI-directional optical flow

    WO2019010156A1