System, apparatus, and method for interpredictive refinement using optical flow

Interpredictive refinement using optical flow models addresses inefficiencies in video coding by refining motion vectors and applying weighted averaging, enhancing compression performance and reducing overhead.

JP2026062911APending Publication Date: 2026-04-10INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERDIGITAL VC HOLDINGS INC
Filing Date
2026-01-05
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing video coding systems face inefficiencies in eliminating temporal redundancy due to rapid illuminance changes and complex motion patterns, particularly in bidirectional prediction modes, leading to suboptimal compression and increased signaling overhead.

Method used

Implementing interpredictive refinement techniques using optical flow models, such as bidirectional optical flow (BIO) and affine motion compensation, to refine motion vectors per sample and apply weighted averaging, reducing residual errors and signaling overhead.

Benefits of technology

Enhances video coding efficiency by improving temporal redundancy reduction and reducing signaling overhead, leading to better compression performance and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062911000001_ABST
    Figure 2026062911000001_ABST
Patent Text Reader

Abstract

Methods, apparatus, and systems are disclosed. [Solution] In one embodiment, the decoding method includes: acquiring a subblock-based motion prediction signal for the current block of a video; acquiring one or more spatial gradients or one or more motion vector difference values ​​of the subblock-based motion prediction signal; acquiring a refinement signal for the current block based on one or more acquired spatial gradients or one or more acquired motion vector difference values; acquiring a refined motion prediction signal for the current block based on the subblock-based motion prediction signal and the refinement signal; and decoding the current block based on the refined motion prediction signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates to a system, apparatus, and method for video coding, particularly for interpredictive refinement using optical flow. [Background technology]

[0002] cross reference This application claims the benefits of U.S. Provisional Patent Application No. 62 / 802,428, filed 7 February 2019, U.S. Provisional Patent Application No. 62 / 814,611, filed 6 March 2019, and U.S. Provisional Patent Application No. 62 / 883,999, filed 15 April 2019, the contents of which are incorporated herein by reference.

[0003] Conventional technology Video coding systems are widely used to compress digital video signals to reduce the storage and / or transmission bandwidth of such signals. Among the various types of video coding systems, including block-based, wavelet-based, and object-based systems, block-based hybrid video coding systems are currently the most widely used and deployed. Examples of block-based video coding systems include international video coding standards such as MPEG1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, and the latest video coding standard called High Efficiency Video Coding (HEVC), developed by the ITU-T / SG16 / Q.6 / VCEG and the JCT-VC (Joint Collaborative Team on Video Coding) of ISO / IEC / MPEG. [Overview of the project]

[0004] In one representative embodiment, the decoding method includes the steps of: obtaining a subblock-based motion prediction signal for the current block of video; obtaining one or more spatial gradients or one or more motion vector difference values ​​for the subblock-based motion prediction signal; obtaining a refinement signal for the current block based on the obtained spatial gradients or one or more obtained motion vector difference values; obtaining a refined motion prediction signal for the current block based on the subblock-based motion prediction signal and the refinement signal; and decoding the current block based on the refined motion prediction signal. Various other embodiments are also disclosed herein.

[0005] A more detailed understanding can be obtained from the following detailed description, provided as an example with the drawings attached to this specification. The figures in the description are illustrative; therefore, the figures and detailed description should not be considered limiting, and other similarly effective examples are possible and expected. Furthermore, similar reference numbers in the figures indicate similar elements. [Brief explanation of the drawing]

[0006] [Figure 1] This is a block diagram illustrating a typical block-based video encoding system. [Figure 2] This is a block diagram illustrating a typical block-based video decoder. [Figure 3] This block diagram illustrates a typical block-based video encoder with generalized bi-prediction (GBi) support. [Figure 4] This diagram illustrates a typical GBi module for an encoder. [Figure 5] This figure illustrates a typical block-based video decoder with GBi support. [Figure 6]This diagram illustrates a typical GBi module for a decoder. [Figure 7] This diagram illustrates a typical bidirectional optical flow. [Figure 8A] This diagram illustrates a typical four-parameter affine mode. [Figure 8B] This diagram illustrates a typical four-parameter affine mode. [Figure 9] This diagram illustrates a typical 6-parameter affine mode. [Figure 10] This diagram illustrates a typical interweaved prediction procedure. [Figure 11] This diagram illustrates typical weight values ​​(for example, those associated with pixels) in a subblock. [Figure 12] This diagram illustrates areas where interweave prediction is applied and other areas where it is not applied. [Figure 13A] This diagram illustrates the SbTMVP process. [Figure 13B] This diagram illustrates the SbTMVP process. [Figure 14] This diagram illustrates adjacent motion blocks (for example, a 4x4 motion block) that can be used to derive motion parameters. [Figure 15] This diagram illustrates adjacent motion blocks that can be used to derive motion parameters. [Figure 16] This figure illustrates the subblock MV and pixel-level MV difference Δv(i,j) after subblock-based affine motion compensation prediction. [Figure 17A] This diagram illustrates a typical procedure for determining the MV that corresponds to the actual center of a subblock. [Figure 17B] This diagram illustrates the position of color difference samples in a 4:2:0 color difference format. [Figure 17C] This diagram illustrates an extended prediction subblock. [Figure 18A] This flowchart illustrates the first typical encoding / decoding method. [Figure 18B] This flowchart illustrates a second typical encoding / decoding method. [Figure 19] This flowchart illustrates a third typical encoding / decoding method. [Figure 20] This flowchart illustrates a fourth typical encoding / decoding method. [Figure 21] This flowchart illustrates a fifth typical encoding / decoding method. [Figure 22] This flowchart illustrates a sixth typical encoding / decoding method. [Figure 23] This flowchart illustrates a seventh typical encoding / decoding method. [Figure 24] This flowchart illustrates the eighth typical encoding / decoding method. [Figure 25] This flowchart illustrates a typical gradient calculation method. [Figure 26] This flowchart illustrates the ninth typical encoding / decoding method. [Figure 27] This flowchart illustrates the tenth typical encoding / decoding method. [Figure 28] This flowchart illustrates the eleventh typical encoding / decoding method. [Figure 29] This is a flowchart illustrating a typical encoding method. [Figure 30] This flowchart illustrates another typical encoding method. [Figure 31] This flowchart illustrates the 12th typical encoding / decoding method. [Figure 32] This flowchart illustrates the 13th typical encoding / decoding method. [Figure 33]This flowchart illustrates the 14th typical encoding / decoding method. [Figure 34A] This is a system diagram illustrating an exemplary communication system in which one or more of the disclosed embodiments may be implemented. [Figure 34B] This is a system diagram illustrating an exemplary wireless transmit / receive unit (WTRU) that may be used in the communication system illustrated in Figure 34A relating to the embodiment. [Figure 34C] This is a system diagram illustrating exemplary radio access networks (RANs) and exemplary core networks (CNs) that may be used within the communication system shown in Figure 34A relating to the embodiment. [Figure 34D] This is a system diagram illustrating further exemplary RANs and further exemplary CNs that may be used in the communication system shown in Figure 34A according to an embodiment. [Modes for carrying out the invention]

[0007] Block-based hybrid video encoding procedure Like HEVC, VVC is built on a block-based hybrid video encoding framework.

[0008] Figure 1 is a block diagram showing a typical block-based hybrid video encoding system.

[0009] Referring to Figure 1, the encoder 100 may be provided with an input video signal 102 that is processed block by block (called coding units (CUs)) and can be used to efficiently compress high-resolution (1080p or higher) video signals. In HEVC, a CU can be up to 64x64 pixels. A CU can be further divided into prediction units, or PUs, to which separate prediction procedures can be applied. For each input video block (MB and / or CU), spatial prediction 160 and / or temporal prediction 162 may be performed. Spatial prediction (or "intra prediction") can predict the current video block using pixels from already encoded adjacent blocks within the same video picture / slice.

[0010] Spatial prediction can reduce the spatial redundancy inherent in video signals. Temporal prediction (also called "interpretation" or "motion-compensated prediction") predicts the current video block using pixels from an already encoded video picture. Temporal prediction can reduce the temporal redundancy inherent in video signals. A temporal prediction signal for a given video block can (for example, typically) be signaled by one or more motion vectors (MVs) that may indicate the amount and / or direction of motion between the current block (CU) and its reference block.

[0011] If multiple reference pictures are supported (as in the case of recent video encoding standards such as H.264 / AVC or HEVC), for each video block, the reference picture index of the video block may be transmitted (for example, additionally), and / or the reference index may be used to identify from which reference picture in the reference picture store 164 the time prediction signal is coming. After spatial and / or time prediction, the mode determination block 180 in encoder 100 can select the best prediction mode, for example, based on a rate-distortion optimization method / procedure. Either the spatial prediction block 160 or the time prediction block 162 may be subtracted from the current video block 116, and / or the prediction residuals may be decorrelated using transformation 104 and quantization 106 to achieve the target bitrate. The quantized residual coefficients can be inversely quantized 110 and inversely transformed 112 to form reconstructed residuals, which can be added back to the prediction blocks in 126 to form reconstructed video blocks. In-loop filtering 166, such as a deblocking filter and / or adaptive loop filter, can be applied to the reconstructed video block, which is then placed in the reference picture store 164 and can be used to encode future video blocks. To form the output video bitstream 120, the encoding mode (inter or intra), predictive mode information, motion information, and quantized residual coefficients are sent to the entropy encoding unit 108 (e.g., all are sent) and can be further compressed and / or packed to form the bitstream.

[0012] The encoder 100 may be implemented using a processor, memory, and transmitter that provide the various elements / modules / units disclosed above. For example, a transmitter may transmit a bitstream 120 to a decoder, and (2) a processor may be configured to run software that enables the reception of input video 102 and the execution of functions associated with the various blocks of the encoder 100.

[0013] Figure 2 is a block diagram of a block-based video decoder.

[0014] Referring to Figure 2, the video decoder 200 may be provided with a video bitstream 202 that can be unpacked and entropy decoded in the entropy decode unit 208. The encoding mode and prediction information may be sent to the appropriate unit among the spatial prediction unit 260 (in the case of intra-encoding mode) and / or the temporal prediction unit 262 (in the case of inter-encoding mode) to form prediction blocks. The residual transformation coefficients may be sent to the inverse quantization unit 210 and the inverse transformation unit 212 to reconstruct the residual blocks. The reconstructed blocks may further pass through the in-loop filtering 266 before being stored in the reference picture store 264. The reconstructed video 220 may be sent to be stored in the reference picture store 264, for example, to drive a display device and to be used when predicting future video blocks.

[0015] The decoder 200 may be implemented using a processor, memory, and receiver that can provide the various elements / modules / units disclosed above. For example, a person skilled in the art will understand that (1) the receiver may be configured to receive the bitstream 202, and (2) the processor may be configured to run software that enables the reception of the bitstream 202 and the output of the reconstructed video 220, and the execution of functions associated with the various blocks of the decoder 200.

[0016] Those skilled in the art will understand that many of the functions, operations, and processes of block-based encoders and block-based decoders are the same.

[0017] In modern video codecs, bidirectional motion-compensated prediction (MCP) can be used for high efficiency in eliminating temporal redundancy by leveraging the temporal correlation between pictures. A bidirectional prediction signal can be formed by combining two single prediction signals with weight values ​​equal to 0.5, but this may not be optimal for combining single prediction signals, especially under conditions where illuminance changes rapidly from one reference picture to another. Certain prediction techniques / operations and / or procedures can be implemented to compensate for illuminance fluctuations over time by applying several global / local weights and / or offset values ​​to the sample values ​​in the reference pictures (e.g., some or each of the sample values ​​in the reference pictures).

[0018] The use of bidirectional motion-compensated prediction (MCP) in video codecs enables the elimination of temporal redundancy by leveraging temporal correlation between pictures. A bidirectional prediction signal can be formed by combining two single prediction signals using weight values ​​(e.g., 0.5). In certain videos, illuminance characteristics may change rapidly from one reference picture to another. Therefore, prediction techniques may compensate for temporal variations in illuminance (e.g., fading transitions) by applying global or local weights and / or offset values ​​to one or more sample values ​​in the reference pictures.

[0019] Generalized bidirectional prediction (GBi) can improve MCP for bidirectional prediction mode. In bidirectional prediction mode, the predicted signal for a given sample x can be calculated by Equation 1, as follows:

[0020] P[x]=w0*P0[x+v0]+w1*P1[x+v1] (1) In the above equation, P[x] can represent the predicted signal obtained from sample x placed at picture position x. Pi[x+vi] may be a motion-compensated predicted signal of x using motion vector (MV) vi for the i-th list (e.g., list 0, list 1, etc.). w0 and w1 may be two weight values ​​shared across all samples in a block (e.g., all). Based on this equation, various predicted signals can be obtained by adjusting the weight values ​​w0 and w1. Some configurations of w0 and w1 may mean the same prediction for single and bidirectional predictions. For example, (w0,w1)=(1,0) may be used for single prediction using reference list L0. (w0,w1)=(0,1) may be used for single prediction using reference list L1. (w0,w1)=(0.5,0.5) may be used for bidirectional prediction using two reference lists. Weights may be signaled per CU. To reduce signaling overhead, a constraint such as w0 + w1 = 1 may be applied so that only one weight can be signaled. Thus, equation 1 may be further simplified as shown in equation 2 below.

[0021] P[x]=(1-w1)*P0[x+v0]+w1*P1[x+v1] (2) To further reduce signaling overhead, w1 can be discretized (e.g., -2 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 10 / 8, etc.). In this case, each weight value can be represented by an index value within a (small) limited range.

[0022] Figure 3 is a block diagram showing a typical block-based video encoder with GBi support.

[0023] The encoder 300 may include a mode determination module 304, a spatial prediction module 306, a motion prediction module 308, a transformation module 310, a quantization module 312, an inverse quantization module 316, an inverse transformation module 318, a loop filter 320, a reference picture store 322, and an entropy coding module 314. Some or all of the encoder's modules or components (for example, the spatial prediction module 306) may be the same as or similar to those described in relation to Figure 1. Furthermore, the spatial prediction module 306 and the motion prediction module 308 may be pixel region prediction modules. Thus, the input video bitstream 302 may be processed in a similar manner to the input video bitstream 102, but the motion prediction module 308 may further include GBi support. In this way, the motion prediction module 308 can combine two separate prediction signals in a weighted average manner. Furthermore, the selected weight indices may be signaled in the output video bitstream 324.

[0024] Those skilled in the art will understand that the encoder 300 may be implemented using a processor, memory, and transmitter that provide the various elements / modules / units disclosed above. For example, the transmitter may transmit the bitstream 324 to the decoder, and (2) the processor may be configured to run software that enables the reception of the input video 302 and the execution of functions associated with the various blocks of the encoder 300.

[0025] Figure 4 shows a typical GBi estimation module 400 that may be used in an encoder motion prediction module such as motion prediction module 308. The GBi estimation module 400 may include a weight estimation module 402 and a motion estimation module 404. Thus, the GBi estimation module 400 can utilize processing (e.g., two-step operation / processing) to generate an inter-prediction signal such as the final inter-prediction signal. The motion estimation module 404 can perform motion estimation by using the input video block 401 and one or more reference pictures received from the reference picture store 406 to find two optimal motion vectors (MV) pointing to reference blocks (e.g., two). The weight estimation module 402 can (1) receive the output of the motion estimation module 404 (e.g., motion vectors v0 and v1), one or more reference pictures from the reference picture store 406, and weight information W, and can find the optimal weight index to minimize the weighted bidirectional prediction error between the current video block and the bidirectional prediction. The weight information W can be thought to describe a list of available weight values ​​or weight sets, so that the determined weight index and the weight information W can be used together to specify the weights w0 and w1 used in GBi. The prediction signal for generalized bidirectional prediction can be calculated as a weighted average of two prediction blocks. The output of the GBi estimation module 400 may include the inter-prediction signal, motion vectors v0 and v1, and / or weight index weight_idx, etc.

[0026] Figure 5 shows a typical block-based video decoder with GBi support, which can decode a bitstream 502 (e.g., from an encoder) that supports GBi, for example, a bitstream 324 created by an encoder 300 described in relation to Figure 3. As shown in Figure 5, the video decoder 500 may include an entropy decoder 504, a spatial prediction module 506, a motion prediction module 508, a reference picture store 510, an inverse quantization module 512, an inverse transform module 514, and / or a loop filter module 518. Some or all of the decoder's modules may be the same as or similar to those described in relation to Figure 2, except that the motion prediction module 508 may further include GBi support. Thus, the coding mode and prediction information can be used to derive a prediction signal using an MCP with spatial prediction or GBi support. For GBi, block motion information and weight values ​​(e.g., in the form of an index indicating the weight values) are received and can be decoded to generate prediction blocks.

[0027] The decoder 500 may be implemented using a processor, memory, and receiver that can provide the various elements / modules / units disclosed above. For example, a person skilled in the art will understand that (1) the receiver may be configured to receive the bitstream 502, and (2) the processor may be configured to run software that enables the reception of the bitstream 502 and the output of the reconstructed video 520, and the execution of functions associated with the various blocks of the decoder 500.

[0028] Figure 6 shows a typical GBi prediction module that can be used in a decoder's motion prediction module, such as the motion prediction module 508.

[0029] Referring to Figure 6, the GBi prediction module may include a weighted averaging module 602 and a motion compensation module 604, the motion compensation module 604 which can receive one or more reference pictures from the reference picture store 606. The weighted averaging module 602 can receive the output of the motion compensation module 604, weight information W, and weight index (e.g., weight_idx). The output of the motion compensation module 604 may include motion information that may correspond to blocks of pictures. The GBi prediction module 600 can use the block motion information and weight values ​​to calculate a GBi prediction signal (e.g., interpretation signal 608) as a weighted average of (e.g., two) motion-compensated prediction blocks.

[0030] Typical bidirectional predictive prediction based on optical flow models Figure 7 shows a typical bidirectional optical flow.

[0031] Referring to Figure 7, bidirectional predictive predictions can be based on an optical flow model. For example, the prediction associated with the current block (e.g., curblk700) is the first prediction block I (0) 702 (for example, a time-previous prediction block shifted by time τ0) and the second prediction block I (1)It can be based on the optical flow associated with 704 (e.g., a temporally future prediction block shifted by time τ1). Bidirectional prediction in video coding may be a combination of two temporal prediction blocks 702 and 704 obtained from already reconstructed reference pictures. Due to the limitations of block-based motion compensation (MC), there may be remaining small motions that can be observed between the samples of the two prediction blocks, thereby potentially reducing the efficiency of motion compensation prediction. To reduce the impact of such motion for all samples within one block, bidirectional optical flow (referred to as BIO or BDOF) can be applied. BIO can provide per-sample motion refinement that can be executed in addition to block-based motion compensation prediction when bidirectional prediction is used. Regarding BIO, the derivation of refined motion vectors per sample in one block can be based on a classical optical flow model. For example, if I (k) (x,y) is the sample value at the coordinates (x,y) of the prediction block derived from the reference picture list k (k = 0,1), and ∂I (k) (x,y) / ∂x and ∂I (k) (x,y) / ∂y are the horizontal and vertical gradients of the samples, assuming an optical flow model, the motion refinement (v x ,v y ) at (x,y) can be derived by Equation 3 below.

[0032]

Equation

[0033] In FIG. 7, (MV x0 ,MV y0 ) associated with the first prediction block 702 and (MV x1 ,MV y1 ) associated with the second prediction block 704 are the two prediction blocks I (0) and I (1)This shows block-level motion vectors that can be used to generate the motion refinement (v) at the sample position (x,y). x ,v y ) can be calculated by minimizing the difference Δ between the motion-refined-compensated sample values ​​(e.g., A and B in Figure 7), as shown in Equation 4 below.

[0034]

number

[0035] For example, to ensure the regularity of the derived motion refinement, the motion refinement is intended to be consistent with respect to samples within a single small unit (e.g., a 4x4 block or other small unit). In Benchmark Set (BMS)-2.0, (v x ,v y The value of ) is derived by minimizing Δ in the 6x6 window Ω around each 4x4 block, as shown in Equation 5 below.

[0036]

number

[0037] To solve the optimization specified in Equation 5, BIO can use a incremental method / operation / procedure that allows it to optimize the refinement by moving horizontally and vertically (for example, then vertically). This may result in equations / inequalities 6 and 7 as follows:

[0038]

number

[0039]

number

[0040] Here,

[0041]

number

[0042] This can be a floor function that can output the maximum value less than or equal to the input, th BIO This can be used as a motion refinement threshold, for example, to prevent error propagation due to coding noise and / or irregular local motion, and 2 18-BD This is equal to the following. The values ​​of S1, S2, S3, S5, and S6 can be further calculated as shown in equations 8-12 below. S1 = Σ (i,j)∈Ω ψ x (i,j)·ψ x (i,j) (8) S3 = Σ (i,j)∈Ω θ(i,j)·ψ x (i,j)·2 L (9) S2 = Σ (i,j)∈Ω ψ x (i,j)·ψ y (i,j) (10) S5=Σ (i,j)∈Ω ψ y (i,j)·ψ y (i,j)·2 (11) S6=Σ (i,j)∈Ω θ(i,j)·ψ y (i,j)·2 L+1 (12) Here, various gradients can be expressed by the following equations 13-15.

[0043]

number

[0044]

number

[0045]

number

[0046] In BMS-2.0, the BIO gradients in equations 13-15, both horizontally and vertically, can be directly obtained by calculating the difference between two adjacent samples (for example, horizontally or vertically, depending on the direction of the derived gradient) at one sample position in each L0 / L1 prediction block, as shown in equations 16 and 17 below.

[0047]

number

[0048]

number

[0049] k=0,1 In equations 8-12, L may be an increase in bit depth related to internal BIO processing / procedures to maintain data accuracy, which may be set to 5 in BMS-2.0, for example. To avoid distinctions by smaller values, the adjustment parameters r and m in equations 6 and 7 may be defined as shown in equations 18 and 19 below. r = 500·4 BD-8 (18) m = 700·4 BD-8 (19) Here, BD may be the bit depth of the input video. Based on the motion refinement derived by equations 4 and 5, the final bidirectional predicted signal of the current CU can be calculated by interpolating the L0 / L1 predicted samples along the motion trajectory based on optical flow equation 3, as specified in equations 20 and 21 below.

[0050]

number

[0051]

number

[0052] Here, shift and ο offset `L0` and `L1` can be right shifts and offsets that can be applied to combine the L0 and L1 prediction signals for bidirectional prediction, and can be set to, for example, equal to 15-BD and 1<<(14-BD)+2·(1<<13), respectively. `rnd(·)` is a rounding function that may round the input value to the nearest integer.

[0053] Typical affine modes In HEVC, a translational motion (translational motion only) model is applied to motion-compensated prediction. In the real world, many types of motion exist (e.g., zoom in / out, rotation, perspective motion, and other irregular motions). In the VVC test model (VTM)-2.0, affine motion-compensated prediction is applied. The affine motion model is either 4-parameter or 6-parameter. A first flag for the inter-encoded CU is signaled to indicate whether a translational motion model or an affine motion model is applied to the inter-prediction. If an affine motion model is applied, a second flag is sent to indicate whether the model is a 4-parameter or 6-parameter model.

[0054] The 4-parameter affine motion model has two parameters for horizontal and vertical translational movement, one parameter for bidirectional zoom movement, and one parameter for bidirectional rotational movement. The horizontal zoom parameter is equal to the vertical zoom parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. The 4-parameter affine motion model is encoded in VTM using two motion vectors at two control point positions defined at the upper left corner 810 and the upper right corner 820 of the current CU. Other control point positions at other corners and / or edges of the current CU are also possible.

[0055] Although one affine motion model is described above, other affine models are equally possible and may be used in various embodiments herein.

[0056] Figures 8A and 8B show the derivation of a typical four-parameter affine model and the subblock-level motion of an affine block. Referring to Figures 8A and 8B, the affine motion field of the block is described by two control point motion vectors at the first control point 810 (upper left corner of the current block) and the second control point 820 (upper right corner of the current block), respectively. Based on the control point motion, one affine-encoded motion field of the block (v x ,v y ) is described as shown in the following equations 22 and 23.

[0057]

number

[0058]

number

[0059] Here, (v 0x ,v 0y ) can be the motion vector of the upper left corner control point 810, and (v 1x ,v 1y ) can be the motion vector of the upper right corner control point 820 as shown in Figure 8A, and w can be the width of the CU. For example, the motion field of the affine-encoded CU is derived at the 4x4 block level, i.e., (v x ,v y This is derived for each 4x4 block in the current CU and applied to the corresponding 4x4 block.

[0060] The four parameters of a four-parameter affine model can be estimated iteratively. The MV pair at step k is:

[0061]

number

[0062] The original signal (e.g., luminance signal) is shown as I(i,j), and the predicted signal (e.g., luminance signal) is shown as I'. k It can be shown as (i,j). Spatial gradient g x (i,j) and g y (i,j) are, for example, the predicted signal I' in the horizontal and / or vertical directions, respectively. k It can be derived using the Sobel filter applied to (i,j). The derivation of Equation 3 can be expressed as shown in Equations 24 and 25 below.

[0063]

number

[0064] Here, in step k, (a,b) can be the delta translation parameter, and (c,d) can be the delta zoom and rotation parameters. The delta MV at the control point can be derived using its coordinates, as shown in equations 26-29 below. For example, (0,0) and (w,0) can be the coordinates of the upper left control point 810 and the upper right control point 820, respectively.

[0065]

number

[0066]

number

[0067] Based on the optical flow equation, the relationship between changes in intensity (e.g., luminance) and spatial gradients and temporal shifts is formulated in Equation 30 as follows:

[0068]

number

[0069]

number

[0070] and

[0071]

number

[0072] By replacing with equation 24, equation 31 for the parameters (a, b, c, d) is obtained as follows. I' k (i,j)-I(i,j)=(g x (i,j)*i+g y (i,j)*j)*c+(-g x (i,j)*j+g y (i,j)*i)*d+g x (i,j)*a+g y (i,j)*b (31)

[0073] Since the samples in CU (e.g., all samples) satisfy equation 31, the parameter set (e.g., a, b, c, d) can be solved, for example, using the least squares error method. Two control points at step (k+1)

[0074]

number

[0075] The MVs in can be solved using equations 26-29, and they can be rounded to a specific precision (e.g., 1 / 4 pixel precision (pel) or other sub-pixel precision). Using iteration, the MVs at two control points can be refined, for example, until convergence (e.g., when all parameters (a, b, c, d) are zero, or when the iteration time reaches a predefined limit).

[0076] Figure 9 shows a typical 6-parameter affine mode, where, for example, V0, V1, and V2 are motion vectors at control points 910, 920, and 930, respectively (MV x , MV y ) represents the motion vector of a subblock centered on position (x,y).

[0077] Referring to Figure 9, an affine motion model (for example, having six parameters) can have any of the following: (1) a parameter for horizontal translation, (2) a parameter for vertical translation, (3) a parameter for horizontal zoom motion, (4) a parameter for horizontal rotation motion, (5) a parameter for vertical zoom motion, and / or (6) a parameter for vertical rotation motion. A six-parameter affine motion model can be encoded using three MVs at three control points 910, 920, and 930. As shown in Figure 9, the three control points 910, 920, and 930 for a six-parameter affine-encoded CU are defined as the upper left corner, upper right corner, and lower left corner of the CU, respectively. The motion at the upper left control point 910 may be related to translational motion, the motion at the upper right control point 920 may be related to horizontal rotational motion and / or horizontal zoom motion, and the motion at the lower left control point 930 may be related to vertical rotation and / or vertical zoom motion. In a 6-parameter affine motion model, horizontal rotational motion and / or zoom motion may not be the same as the same motion in the vertical direction. Each subblock (v x ,v yThe motion vector of ) can be derived using three MVs at control points 910, 920, and 930, as shown in equations 32 and 33 below.

[0078]

number

[0079]

number

[0080] Here, (v 2x ,v 2y ) can be the motion vector V2 of the lower left control point 930, (x, y) can be the center position of the subblock, w can be the width of CU, and h can be the height of CU.

[0081] The six parameters of a six-parameter affine model can be estimated in a similar manner. Equations 24 and 25 can be modified as shown in equations 34 and 35 below.

[0082]

number

[0083] Here, in step k, (a,b) can be delta translation parameters, (c,d) can be delta zoom and rotation parameters for the horizontal direction, and (e,f) can be delta zoom and rotation parameters for the vertical direction. Equation 31 can be modified as shown in Equation 36 below. I' k (i,j)-I(i,j)=(g x (i,j)*i)*c+(g x (i,j)*j)*d+(g y (i,j)*i)*e+(g y (i,j)*j)*f+g x (i,j)*a+g y(i,j)*b (36)

[0084] The parameter set (a, b, c, d, e, f) can be solved using least squares / procedure / operation by considering, for example, the samples within the CU (e.g., all samples). MV of the upper left control point

[0085]

number

[0086] This can be calculated using equations 26-29. MV of the upper right control point

[0087]

number

[0088] This can be calculated using equations 37 and 38 as shown below. MV of the lower left control point

[0089]

number

[0090] This can be calculated using equations 39 and 40, as shown below.

[0091]

number

[0092]

number

[0093] While Figures 8A, 8B, and 9 show 4-parameter and 6-parameter affine models, those skilled in the art will understand that affine models with different numbers of parameters and / or different control points are equally possible.

[0094] While affine models are described herein in relation to optical flow refinement, those skilled in the art will understand that other motion models related to optical flow refinement are equally possible.

[0095] Typical interweave prediction for affine motion compensation In affine motion compensation (AMC), for example in a VTM, the encoded block is divided into small subblocks of about 4x4, each of which can be assigned an individual motion vector (MV) derived by the affine model, as shown in Figures 8A and 8B or Figure 9. In a 4-parameter or 6-parameter affine model, the MV can be derived from the MVs of two or three control points.

[0096] AMC may face a dilemma associated with subblock size. While AMC can achieve better encoding performance with smaller subblocks, it may suffer from the burden of greater complexity.

[0097] Figure 10 shows a typical interweave prediction procedure that can achieve finer granularity of MV, for example, at the cost of a moderate increase in complexity.

[0098] In Figure 10, the coded block 1010 may be divided into subblocks having two different partitioning patterns (e.g., a first pattern 0 and a second pattern 1). As shown in Figure 10, the first partitioning pattern 0 (e.g., a first subblock pattern, for example, a 4x4 subblock pattern) may be the same as that in the VTM, and the second partitioning pattern 1 (e.g., an overlapping and / or interwoven second subblock pattern) may divide the coded block 1010 into a 4x4 subblock with a 2x2 offset from the first partitioning pattern 0. Several auxiliary predictions (e.g., two auxiliary predictions P0 and P1) may be generated by the AMC using the two partitioning patterns (e.g., the first partitioning pattern 0 and the second partitioning pattern 1). The MV of each subblock in partitioning patterns 0 and 1 can be derived from the control point motion vector (CPMV) by the affine model.

[0099] The final prediction P can be calculated as a weighted sum of auxiliary predictions (e.g., two auxiliary predictions P0 and P1) formulated as shown in equations 41 and 42 below.

[0100]

number

[0101] Figure 11 shows typical weight values ​​(for example, associated with pixels) in a subblock. Referring to Figure 11, an auxiliary prediction sample located at the center of subblock 1100 (for example, the central pixel) may be associated with weight value 3, and an auxiliary prediction sample located at the boundary of subblock 1100 may be associated with weight value 1.

[0102] Figure 12 shows regions to which interweave prediction is applied and other regions to which it is not applied. Referring to Figure 12, region 1200 may include a first region 1210 (not shaded as shown in Figure 12) which has, for example, a 4x4 subblock to which interweave prediction is applied, and a second region 1220 (shaded as shown in Figure 12) to which interweave prediction is not applied. To avoid small block motion compensation, interweave prediction may be applied only to regions where the subblock size satisfies a threshold size (e.g., 4x4) for both the first and second partitioning patterns, for example.

[0103] In VTM-3.0, the subblock size may be 4x4 in the chrominance component, and interweave prediction may be applied to the chrominance component and / or luminance component. The area used for motion compensation (MC) for subblocks (e.g., all subblocks) may be taken together as a whole in AMC, so bandwidth may not be increased by interweave prediction. For flexibility, a flag may be signaled in the slice header to indicate whether interweave prediction is used. For interweave prediction, the flag may be signaled as a 1-bit flag (e.g., a first logic level that can always be signaled as 0 or 1).

[0104] Typical procedure for sub-block-based temporal motion vector prediction (SbTMVP) SbTMVP is supported by VTM. Similar to Temporal Motion Vector Prediction (TMVP) in HEVC, SbTMVP can use motion fields in co-located pictures to improve, for example, the motion vector prediction and merging modes of the CU in the current picture. The same co-located pictures used by TMVP may be used for SbTMVP. SbTMVP differs from TMVP in that (1) TMVP can predict motion at the CU level, while SbTMVP can predict motion at the sub-CU level, and / or (2) TMVP can extract temporal motion vectors from co-located blocks within co-located pictures (for example, the co-located block may be the bottom-right or center block relative to the current CU), and SbTMVP can apply motion shifts before extracting temporal motion information from co-located pictures (for example, motion shifts may be obtained from motion vectors from one of the spatially adjacent blocks of the current CU).

[0105] Figures 13A and 13B illustrate the SbTMVP process. Figure 13A shows the spatially adjacent blocks used by ATMVP, and Figure 13B shows the derivation of the subCU motion field by applying motion shifts from spatial adjacents and scaling motion information from corresponding co-located subCUs.

[0106] Referring to Figures 13A and 13B, SbTMVP can predict the motion vector of a subCU within the current CU motion (for example, in two motions). In the first motion, spatially adjacent blocks A1, B1, B0, and A0 may be examined in the order A1, B1, B0, and A0. As soon as and / or after a first spatially adjacent block is identified that has a motion vector using a picture at the same location as its reference picture, this motion vector may be selected as the motion shift to be applied. If no such motion is identified from a spatially adjacent block, the motion shift may be set to (0,0). In the second motion, as shown in Figure 13B, the motion shift identified in the first motion is applied (for example, added to the coordinates of the current block) so that subCU-level motion information (for example, motion vector and reference index) can be obtained from the picture at the same location. An example in Figure 13B shows the motion shift set for the motion of block A1. For each subCU, motion information from its corresponding block in the same location picture (e.g., the smallest motion grid covering the central sample) may be used to derive motion information for that subCU. After the motion information for the same location subCU is identified, the motion information may be converted into motion vectors and reference indices for the current subCU in a manner similar to HEVC's TMVP processing. For example, temporal motion scaling may be applied to align the reference pictures of the temporal motion vectors with those of the current CU.

[0107] A combined subblock-based merge list may be used in VTM-3 and may contain or include both SbTMVP and affine merge candidates, for example, for signaling the subblock-based merge mode. The SbTMVP mode can be enabled / disabled by the Sequence Parameter Set (SPS) flag. When the SbTMVP mode is enabled, an SbTMVP predictor is added as the first entry in the list of subblock-based merge candidates, followed by affine merge candidates. The size of the subblock-based merge list may be signaled by the SPS, and the maximum allowable size of the subblock-based merge list is an integer, which may be 5 in VTM3, for example.

[0108] The subCU size used in SbTMVP may be fixed, for example, 8x8 or another subCU size, and SbTMVP mode may be applicable to CUs that have both a width and height of 8 or more (for example, it may be applicable only to them), as is done in affine merge mode. The encoding logic for additional SbTMVP merge candidates may be the same as for other merge candidates. For example, for each CU in a P or B slice, an additional rate distortion (RD) check may be performed to determine whether to use the SbTMVP candidate.

[0109] Typical Regression-Based Motion Vector Fields To provide finer granularity for motion vectors within a block, a regression-based motion vector field (RMVF) tool may be implemented (for example, in JVET-M0302), which can attempt to model the motion vectors of each block at the sub-block level based on spatially adjacent motion vectors.

[0110] Figure 14 shows adjacent motion blocks (e.g., 4x4 motion blocks) that can be used for motion parameter derivation. One row 1410 and one column 1420 of the directly adjacent motion vectors (and their center positions) based on 4x4 subblocks from each side of the block can be used in the regression process. For example, those adjacent motion vectors can be used for RMVF motion parameter derivation.

[0111] Figure 15 shows adjacent motion blocks that can be used for motion parameter derivation, reducing adjacent motion information (for example, the number of adjacent motion blocks used in the regression processing for Figure 14 may be reduced). The amount of adjacent motion information reduced for RMVF parameter derivation of adjacent 4x4 motion blocks can be used for motion parameter derivation (for example, about half, e.g., every other adjacent motion block may be used for motion parameter derivation). Certain adjacent motion blocks in row 1410 and column 1420 may be selected, determined, or pre-determined to reduce adjacent motion information.

[0112] While it is shown that approximately half of the adjacent motion blocks in row 1410 and column 1420 have been selected, other proportions (including other motion block locations) may be selected, for example, to reduce the number of adjacent motion blocks used in the regression process.

[0113] When collecting motion information for deriving motion parameters, the five regions shown in the figure (for example, bottom left, left, top left, top, and top right) may be used. The top right and bottom left reference motion regions may be limited to half (for example, only half) of the corresponding width or height of the current block.

[0114] In RMVF mode, the block's motion can be defined by a 6-parameter motion model. These parameters a xx a xy a yx a yy , b x , and b yThis can be calculated by solving a linear regression model in the sense of mean squared error (MSE). The input to the regression model is the center positions (x,y) and / or motion vectors (mv) of the available adjacent 4x4 subblocks, as defined above. x and mv y ) may consist of, or include, these elements.

[0115] (X subPU ,Y subPU ) Motion vector (MV) of an 8x8 subblock with its center position at ). X_subPU MV Y_subPU ) can be calculated as shown in Equation 43 below.

[0116]

number

[0117] Motion vectors can be calculated for 8x8 subblocks relative to the center position of each subblock. For example, motion compensation can be applied with 8x8 subblock accuracy in RMVF mode. To obtain efficient modeling of the motion vector field, the RMVF tool is applied only when at least one motion vector from at least three candidate regions is available.

[0118] Affine motion model parameters can be used to derive the motion vector of a particular pixel (e.g., each pixel) in the CU. While the complexity of generating pixel-based affine motion compensation predictions can be high (e.g., very high), and the memory access bandwidth requirements for this type of sample-based MC can be high, subblock-based affine motion compensation procedures / methods may be implemented (e.g., by VVC). For example, the CU may be divided into subblocks (e.g., 4x4 subblocks, square subblocks, and / or non-square subblocks). Each subblock may be assigned an MV, which can be derived from the affine model parameters. The MV may be at the center of the subblock (or at another location within the subblock). Pixels within a subblock (e.g., all pixels within the subblock) may share a subblock MV. Subblock-based affine motion compensation may involve a trade-off between coding efficiency and complexity. To achieve finer-grained motion compensation, interweave predictions for affine motion compensation may be implemented and can be generated by weighting two subblock motion compensation predictions. Interweave prediction requires and / or can use two or more motion-compensated predictions per subblock, which can therefore increase memory bandwidth and complexity.

[0119] In certain representative embodiments, methods, apparatus, procedures, and / or operations may be implemented to refine subblock-based affine motion compensation predictions using optical flow (e.g., using and / or based on optical flow). For example, after subblock-based affine motion compensation has been performed, the pixel intensity may be refined by adding a difference value derived by an optical flow equation, which is called prediction refinement with optical flow (PROF). PROF can achieve pixel-level granularity without significantly increasing complexity and may maintain the same worst-case memory access bandwidth as subblock-based affine motion compensation. PROF may be applied in any scenario where a pixel-level motion vector field is available (e.g., can be computed) in addition to the prediction signal (e.g., the unrefined motion prediction signal and / or the subblock-based motion prediction signal). In addition to or otherwise than the affine mode, the prediction PROF procedure may be used in other subblock prediction modes. The application of PROF in subblock modes such as SbTMVP and / or RMVF may be implemented. The application of PROF in bidirectional prediction is described herein.

[0120] Typical PROF procedures for affine mode In certain representative embodiments, methods, apparatus, and / or procedures may be implemented to improve the granularity of subblock-based affine motion compensation prediction by, for example, applying a change in pixel intensity derived from optical flow (e.g., an optical flow formula), and may use and / or require, for example, one motion compensation operation per subblock (e.g., only one motion compensation operation per subblock), the same as existing affine motion compensation in VVC.

[0121] Figure 16 shows the subblock MV and pixel-level motion vector difference Δv(i,j) (sometimes called refinement MV for pixels, for example) after subblock-based affine motion compensation prediction.

[0122] Referring to Figure 16, CU1600 can contain subblocks 1610, 1620, 1630, and 1640. Each subblock 1610, 1620, 1630, and 1640 can contain multiple pixels (for example, 16 pixels within subblock 1610). Subblock MV1650 (for example, as a coarse or average subblock MV) associated with each pixel 1660(i,j) of subblock 1610 is shown. For each individual pixel (i,j) within subblock 1610, a refinement MV1670(i,j) can be determined (this can represent the difference between the actual MV of pixel 1660(i,j) and subblock MV1650, where (i,j) defines the pixel position within subblock 1610). For clarity in Figure 16, only the refinement MV1670(1,1) is labeled, while other individual pixel-level motions are shown. In certain representative embodiments, the refinement MV1670(i,j) can be determined as a pixel-level motion vector difference Δv(i,j) (sometimes called the motion vector difference).

[0123] In certain representative embodiments, methods, apparatus, procedures, and / or operations including any of the following operations may be implemented: (1) In the first operation, a subblock-based AMC can be performed as disclosed herein to generate a subblock-based motion prediction I(i,j). (2) In the second operation, the spatial gradient g of the subblock-based motion prediction I(i,j) at each sample position. x (i,j) and g y(i,j) can be calculated (for example, the spatial gradient can be generated using the same process as gradient generation used in BDOF. For instance, the horizontal gradient at a sample location can be calculated as the difference between the adjacent sample to its right and the adjacent sample to its left, and / or the vertical gradient at a sample location can be calculated as the difference between the adjacent sample below and the adjacent sample above. In another example, the spatial gradient can be generated using a Sobel filter). (3) In the third operation, the change in brightness intensity per pixel in the CU can be calculated using and / or by the optical flow formula, for example, as shown in Equation 44 below. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (44) Here, the value of the motion vector difference Δv(i,j) is 1670, which is the difference between the pixel-level MV calculated for the sample position (i,j) indicated by v(i,j), and the subblock-level MV 1650 of the subblock covering pixel 1660(i,j), as shown in Figure 16. The pixel-level MV v(i,j) can be derived from the control point MV by equations 22 and 23 for the four-parameter affine model, or by equations 32 and 33 for the six-parameter affine model.

[0124] In certain representative embodiments, the motion vector difference value Δv(i,j) may be derived by the affine model parameters using or by equations 24 and 25, where x and y may be offsets from the pixel position to the center of the subblock. Since the affine model parameters and pixel offsets do not change per subblock, the motion vector difference value Δv(i,j) can be calculated for a first subblock and reused for other subblocks in the same CU. For example, the translational affine parameters (a,b) may be the same for the pixel-level MV and the subblock-level MV, so the difference between the pixel-level MV and the subblock-level MV can be calculated using equations 45 and 46 as follows: (c,d,e,f) may be four additional affine parameters (e.g., four affine parameters other than the translational affine parameters).

[0125]

number

[0126] Here, (i, j) can be the pixel position relative to the top-left position of the subblock, and (x sb ,y sb ) can be the center position of the subblock relative to the upper left position of the subblock.

[0127] Figure 17A shows a typical procedure for determining the MV corresponding to the actual center of a subblock.

[0128] Referring to FIG. 17A, two sub-blocks SB0 and SB1 are shown as 4×4 sub-blocks. When the sub-block width is SW and the sub-block height is SH, the sub-block center position can be shown as ((SW−1) / 2,(SH−1) / 2). In other examples, the sub-block center position can be estimated based on the position shown as (SW / 2,SH / 2). The actual center point is P0’ for the first sub-block SB0 and P1’ for the second sub-block SB1 using ((SW−1) / 2,(SH−1) / 2). The estimated center point is P0 for the first sub-block SB0 and P1 for the second sub-block SB1 using, for example, (SW / 2,SH / 2) (e.g., in VVC). In certain representative embodiments, the MV of the sub-block can be based on the more accurate actual center position rather than the estimated center position (used in VVC).

[0129] FIG. 17B is a diagram showing the positions of chroma samples in the 4:2:0 chroma format. Referring to FIG. 17B, the chroma sub-block MV can be derived from the MV of the luma sub-block. For example, in the 4:2:0 chroma format, one 4×4 chroma sub-block can correspond to an 8×8 luma area. Representative embodiments are shown in relation to the 4:2:0 chroma format, but those skilled in the art will understand that other chroma formats such as the 4:2:2 chroma format can be used similarly.

[0130] The chroma sub-block MV can be derived by averaging the top-left 4×4 luma sub-block MV and the bottom-right luma sub-block MV. The derived chroma sub-block MV may or may not be centered within the chroma sub-block for chroma sample position types 0, 2, and / or 3. For chroma sample position types 0, 2, and 3, the chroma sub-block center position (x sb ,y sb) may or may need to be adjusted by an offset. For example, for 4:2:0 chromatic difference sample position types 0, 2, and 3, adjustments may be applied as shown in equations 47-49 below.

[0131]

number

[0132]

number

[0133]

number

[0134] The subblock-based motion prediction I(i,j) can be refined by adding intensity changes (for example, luminance intensity changes as provided in Equation 44 as an example). The final (i.e., refined) prediction I'(i,j) can be generated by or using Equation 50 as follows: I'(i,j)=I(i,j)+ΔI(i,j) (50)

[0135] When refinement is applied, subblock-based affine motion compensation can achieve pixel-level granularity without increasing worst-case bandwidth and / or memory bandwidth.

[0136] To maintain accuracy in prediction and / or gradient calculation, the bit depth in the operational relational performance of subblock-based AMC may be set to an intermediate bit depth that can be higher than the encoded bit depth.

[0137] The processes described above may be used to refine chromatic difference intensity (for example, in addition to or instead of refining luminance intensity). For example, the intensity difference used in Equation 50 may be multiplied by a weighting coefficient w before being added to the prediction, as shown in Equation 51 below. I'(i,j)=I(i,j)+w·ΔI(i,j) (51) Here, w may be set to a value between 0 and 1, and w may be signaled at the CU level or the picture level. For example, w may be signaled by a weight index. For example, index table 1 may be used to signal w.

[0138] [Table 1]

[0139] The encoder algorithm can choose the value of w that yields the lowest rate distortion cost.

[0140] The gradient of the predicted sample, for example, g x and / or g y This can be calculated in different ways. In a particular representative embodiment, the predicted sample g x and g y This can be calculated by applying a two-dimensional Sobel filter. Examples of 3x3 Sobel filters for horizontal and vertical gradients are shown below.

[0141]

number

[0142]

number

[0143] In other representative embodiments, the gradient can be calculated using a one-dimensional three-tap filter. An example may include [-1 0 1], which may be simpler (for example, considerably simpler) than a Sobel filter.

[0144] Figure 17C shows an extended subblock prediction. The shaded circle 1710 is a padding sample around a 4x4 subblock (e.g., the unshaded circle 1720). As an example, a Sobel filter may be used to calculate the gradient of the sample in box 1730 to the central sample 1740. While the gradient can be calculated using a Sobel filter, other filters such as a 3-tap filter are also possible.

[0145] For the exemplary gradient filters described above, such as the 3x3 Sobel filter and the one-dimensional filter, extended subblock prediction may be used and / or required for subblock gradient calculation. One row at the top and bottom boundaries and one column at the left and right boundaries of the subblock may be padded, for example, to calculate the gradient of those samples at the subblock boundaries.

[0146] Different methods / procedures and / or operations may exist for obtaining extended subblock predictions. In one representative embodiment, given N×M as the subblock size, an (N+2)×(M+2) extended subblock prediction can be obtained by performing (N+2)×(M+2) block motion compensation using the subblock MV. In this embodiment, memory bandwidth may be increased. To avoid increased memory bandwidth, in certain representative embodiments, given K-tap interpolation filters in both the horizontal and vertical directions, integer reference samples of (N+K-1)×(M+K-1) before interpolation may be taken for interpolation of the N×M subblock, and boundary samples of the (N+K-1)×(M+K-1) block may be copied from adjacent samples of the (N+K-1)×(M+K-1) subblock so that the extended region may be (N+K-1+2)×(M+K-1+2). The extended region may be used for interpolation of the (N+2)×(M+2) subblock. These representative embodiments may further use and / or require additional interpolation operations to generate (N+2)×(M+2) predictions when the subblock MV points to fractional positions.

[0147] For example, to reduce computational complexity, in other representative embodiments, subblock predictions may be obtained by N×M block motion compensation using subblock MVs. The boundaries of the (N+2)×(M+2) predictions may be obtained without interpolation by one of the following: (1) integer motion compensation where MV is an integer part of the subblock MV, (2) integer motion compensation where MV is the nearest integer MV of the subblock MV, and / or (3) a copy from the nearest neighbor sample in the N×M subblock prediction.

[0148] For example, the accuracy and / or range of pixel-level refinement MV can affect the accuracy of PROF. In certain representative embodiments, a combination of a multi-bit fractional component and another multi-bit integer component can be implemented. For example, a 5-bit fractional component and an 11-bit integer component may be used. The combination of a 5-bit fractional component and an 11-bit integer component can represent an MV range of -1024 to 1023 with a total of 16 bits of 1 / 32 pel accuracy.

[0149] Gradient, e.g., g x and g y and the accuracy of the intensity change ΔI can affect the performance of PROF. In certain representative embodiments, the prediction sample accuracy can be maintained or held at a predetermined number or the number of bits signaled (e.g., the internal sample accuracy defined in the current VVC draft which is 14 bits). In certain representative embodiments, the gradient and / or the intensity change ΔI can be maintained at the same accuracy as the prediction sample.

[0150] The range of the intensity change ΔI can affect the performance of PROF. The intensity change ΔI can be clipped to a smaller range to avoid incorrect values generated by an inaccurate affine model. In one example, the intensity change ΔI can be clipped to predition_bitdepth - 2.

[0151] Δv x and Δv y The combination of the number of bits of the fractional component of Δv, the number of bits of the fractional component of the gradient, and the number of bits of the intensity change ΔI can affect the complexity of a particular hardware or software implementation. In one representative embodiment, 5 bits can be used to represent the fractional components of Δv x and Δv y and 2 bits can be used to represent the fractional component of the gradient and 12 bits can be used to represent ΔI, but they can be any number of bits.

[0152] To reduce computational complexity, PROF may be omitted in certain situations. For example, if the magnitude of all pixel-based deltas (e.g., refinement) MV(Δv(i,j)) within a 4x4 subblock is less than a threshold, PROF may be omitted for the entire affine CU. Similarly, if the gradient of all samples within a 4x4 subblock is less than a threshold, PROF may be omitted. PROF can be applied to chrominance components such as Cb and / or Cr components. The delta MV of the Cb and / or Cr components of a subblock may reuse the delta MV of the subblock (for example, the delta MV calculated for different subblocks within the same CU may be reused).

[0153] The gradient procedures disclosed herein (for example, extending subblocks for gradient calculation using copied reference samples) are shown to be used with PROF operations, but gradient procedures may also be used with other operations, particularly BDOF operations and / or affine motion estimation operations.

[0154] Typical PROF procedures for other subblock modes PROF can be applied in any scenario where a pixel-level motion vector field is available (e.g., can be computed) in addition to the prediction signal (e.g., the unrefined prediction signal). For example, in addition to affine modes, prediction refinement using optical flow may be used in other subblock prediction modes, such as SbTMVP mode (e.g., ATMVP mode in VVC) or regression-based motion vector field (RMVF).

[0155] In certain representative embodiments, methods for applying PROF to SbTMVP may be implemented. For example, such methods may include, in particular, any of the following: (1) In the first operation, subblock level MV and subblock prediction can be generated based on existing SbTMVP processing described herein. (2) In the second operation, the affine model parameters can be estimated by the subblock MV field using a linear regression method / procedure. (3) In the third operation, the pixel-level MV can be derived from the affine model parameters obtained in the second operation, and the associated pixel-level motion refinement vector (Δv(i,j)) for the subblock MV can be calculated, and / or (4) In the fourth operation, prediction refinement using optical flow processing can be applied to generate the final prediction.

[0156] In certain representative embodiments, methods for applying PROF to RMVF may be implemented. For example, such methods may include any of the following: (1) In the first operation, the subblock level MV field, the subblock prediction, and / or the affine model parameter a xx a xy a yx a yy , b x and b x However, it can be generated based on the RMVF processing described herein. (2) In the second operation, the pixel-level MV offset from the subblock-level MV is given by the affine model parameter a by equation 52 as follows: xx a xy a yx a yy , b x and b x This can be derived by [the following method].

[0157]

number

[0158] Here, (i,j) is the pixel offset from the subblock center. Since the affine model parameters and / or the pixel offset from the subblock center do not change per subblock, the pixel MV offset may be calculated for the first subblock (for example, it may need to be calculated only for that subblock, or should be calculated only for that subblock), and may be reused in other subblocks within the CU, and / or (3) In the third operation, the PROF process is applied, and the final prediction can be generated by applying, for example, equations 44 and 50.

[0159] Typical PROF procedures for bidirectional prediction In addition to or instead of using PROF for single prediction as described herein, the PROF technique may be used for bidirectional prediction. When used for bidirectional prediction, PROF may be used to generate L0 and / or L1 predictions, for example, before they are combined with weights. To reduce computational complexity, PROF may be applied to a single prediction, such as L0 or L1 (for example, to only that). In certain representative embodiments, PROF may be applied to a list (for example, a list of the current picture along with or associated with the closest and / or nearest reference pictures (for example, within a threshold)).

[0160] Typical steps for enabling PROF PROF activation may be signaled in or within the Sequence Parameter Set (SPS) header, Picture Parameter Set (PPS) header, and / or Tile Group header. In certain embodiments, a flag may be signaled to indicate whether PROF is enabled for affine mode. If the flag is set to a first logical level (e.g., "True"), PROF may be used for both single and bidirectional prediction. In certain embodiments, if the first flag is set to "True", a second flag may be used to indicate whether PROF is enabled or not for bidirectional prediction affine mode. If the first flag is set to a second logical level (e.g., "False"), the second flag may be inferred to be set to "False". Whether PROF is applied to the chrominance component may be signaled using a flag in or within the SPS header, PPS header, and / or Tile Group header if the first flag is set to "True", thereby separating PROF control for luminance and chrominance components.

[0161] Typical steps for conditionally activating PROF For example, to reduce complexity, PROF may be applied only when certain conditions are met (e.g., only when). For example, with respect to small CU sizes (e.g., below a threshold level), the benefits of applying PROF may be limited because the affine motion is relatively small. In certain representative embodiments, when the CU size is small (e.g., for CU sizes of 16x16 or less, such as 8x8, 8x16, 16x8) or under those conditions, PROF may be disabled in affine motion compensation to reduce complexity with respect to both the encoder and / or decoder. In certain representative embodiments, when the CU size is small (below the same or a different threshold level), PROF may be omitted in affine motion estimation (e.g., affine motion estimation only) to reduce the complexity of the encoder, for example, while PROF may be performed in the decoder regardless of the CU size. For example, on the encoder side, after motion estimation to find affine model parameters (e.g., control point MV), a motion compensation (MC) procedure may be invoked and PROF may be executed. The MC procedure may be invoked for each iteration in motion estimation. In motion estimation, PROF can be omitted in the MC to reduce complexity, but since the final MC within the encoder will perform PROF, there will be no prediction mismatch between the encoder and decoder. In other words, PROF refinement does not need to be applied when the encoder searches for the affine model parameters (e.g., affine MV) to use for predicting the CU, and once the encoder has completed the search, or afterward, the encoder can apply PROF to refine the prediction about the CU using the affine model parameters determined from the search.

[0162] In some representative embodiments, the difference between CPMVs can be used as a criterion to determine whether to enable PROF. When the difference between CPMVs is small (e.g., less than a threshold level), and thus the affine motion is small, the advantages of applying PROF may be limited, and PROF may be disabled for affine motion compensation and / or affine motion estimation. For example, in the four-parameter affine mode, PROF may be disabled if the following conditions are met (e.g., all of the following conditions are met).

[0163]

Number

[0164]

Number

[0165] In the six-parameter affine mode, in addition to or instead of the above conditions, PROF may be disabled if the following conditions are met (e.g., all of the following conditions are also met).

[0166]

Number

[0167]

Number

[0168] Here, T is a predefined threshold, e.g., 4. This CPMV- or affine parameter-based PROF omission procedure may be applied in the encoder (e.g., may only be applied), and the decoder may or may not omit PROF.

[0169] Representative procedures for PROF in combination with or instead of the deblocking filter PROF can be a pixel-by-pixel refinement that can compensate for block-based MC, thus reducing (for example, significantly) the motion difference between block boundaries. Encoders and / or decoders can omit the application of deblocking filters and / or apply weaker filters at subblock boundaries when PROF is applied. In CUs divided into multiple transform units (TUs), blocking artifacts may appear on transform block boundaries.

[0170] In certain representative embodiments, the encoder and / or decoder may omit the application of a deblocking filter, or may apply one or more weaker filters to the subblock boundaries, as long as the subblock boundaries do not coincide with the TU boundaries.

[0171] When PROF is applied to luminance (for example, only to luminance), or under such conditions, the encoder and / or decoder may omit the application of a deblocking filter and / or apply one or more weaker filters on the subblock boundary for luminance (for example, only to luminance). For example, a boundary strength parameter B may be used to apply a weaker deblocking filter.

[0172] For example, an encoder and / or decoder may omit the application of a deblocking filter on a subblock boundary when PROF is applied, as long as the subblock boundary does not coincide with a TU boundary. In that case, a deblocking filter may be applied to reduce or eliminate blocking artifacts that may occur along the TU boundary.

[0173] As another example, an encoder and / or decoder may apply a weaker deblocking filter on subblock boundaries when PROF is applied, as long as the subblock boundary does not coincide with a TU boundary. The "weaker" deblocking filter is intended to be weaker than the one that would normally be applied to a subblock boundary when PROF is not applied. When a subblock boundary coincides with a TU boundary, a stronger deblocking filter is applied, which can reduce or eliminate blocking artifacts that would be expected to be more visible along the subblock boundary that coincides with the TU boundary.

[0174] In certain representative embodiments, when PROF is applied to luminance (e.g., only), or under such conditions, the encoder and / or decoder may, for design uniformity purposes, align the application of the deblocking filter to chrominance with that of luminance, for example, even though PROF is not applied to chrominance. For example, when PROF is applied only to luminance, the normal application of the deblocking filter to luminance may be modified based on whether PROF is applied (and possibly based on whether there is a TU boundary at the subblock boundary). In certain representative embodiments, rather than having separate / different logic for applying the deblocking filter to the corresponding chrominance pixels, the deblocking filter may be applied to the subblock boundary for chrominance in a way that is consistent with (and / mirrors) the procedure for luminance deblocking.

[0175] Figure 18A is a flowchart showing the first representative encoding / decoding method.

[0176] Referring to Figure 18A, a typical encoding and / or decoding method 1800 may include, in block 1805, the encoder 100 or 300 and / or the decoder 200 or 500 acquiring a subblock-based motion prediction signal for the current block, for example, of video. In block 1810, the encoder 100 or 300 and / or the decoder 200 or 500 may acquire one or more spatial gradients of the subblock-based motion prediction signal for the current block, or one or more motion vector difference values ​​associated with the subblocks of the current block. In block 1815, the encoder 100 or 300 and / or the decoder 200 or 500 may acquire a refinement signal for the current block based on one or more acquired spatial gradients or one or more motion vector difference values ​​associated with the subblocks of the current block. In block 1820, encoder 100 or 300 and / or decoder 200 or 500 can obtain a refined motion prediction signal for the current block based on a subblock-based motion prediction signal and a refinement signal. In certain embodiments, encoder 100 or 300 can encode the current block based on the refined motion prediction signal, or decoder 200 or 500 can decode the current block based on the refined motion prediction signal. The refined motion prediction signal can be a refined motion interpretation signal generated (for example by GBi encoder 300 and / or GBi decoder 500), and one or more PROF operations may be used.

[0177] For example, in certain representative embodiments relating to other methods described herein, including methods 1850 and 1900, obtaining a subblock-based motion prediction signal for the current block of video may include generating a subblock-based motion prediction signal.

[0178] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, obtaining one or more spatial gradients of a sub-block based motion prediction signal for a current block, or one or more motion vector difference values associated with sub-blocks of the current block, can include determining one or more spatial gradients (e.g., associated with a gradient filter) of the sub-block based motion prediction signal.

[0179] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, obtaining one or more spatial gradients of a sub-block based motion prediction signal for a current block, or one or more motion vector difference values associated with sub-blocks of the current block, can include determining one or more motion vector difference values associated with sub-blocks of the current block.

[0180] For example, in certain representative embodiments associated with other methods described herein, particularly methods 1850 and 1900, obtaining a refinement signal for a current block based on one or more determined spatial gradients or one or more determined motion vector difference values can include determining, based on the determined spatial gradients, a motion prediction refinement signal for the current block as the refinement signal.

[0181] For example, in certain representative embodiments relating to other methods described herein, particularly methods 1850 and 1900, obtaining a refinement signal for a current block based on one or more determined spatial gradients or one or more determined motion vector difference values ​​may include determining a motion prediction refinement signal for the current block as a refinement signal based on the determined motion vector difference values.

[0182] The term "determine" or "decide" in relation to something like information can generally include one or more of the following actions regarding information: estimation, calculation, prediction, acquisition, and / or retrieval. For example, determining can refer, among other things, to retrieving something from memory or a bitstream.

[0183] For example, in certain representative embodiments relating to other methods described herein, particularly methods 1850 and 1900, obtaining a refined motion prediction signal for the current block based on a subblock-based motion prediction signal and a refinement signal may include combining (e.g., by adding or subtracting) the subblock-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block.

[0184] For example, in certain representative embodiments relating to other methods described herein, particularly methods 1850 and 1900, encoding and / or decoding a current block based on a refined motion prediction signal may include encoding video using the refined motion prediction signal as a prediction for the current block, and / or decoding video using the refined motion prediction signal as a prediction for the current block.

[0185] Figure 18B is a flowchart showing a second typical encoding and / or decoding method.

[0186] Referring to Figure 18B, a typical method 1850 for encoding and / or decoding video may include, in block 1855, having encoders 100 or 300 and / or decoders 200 or 500 generate a subblock-based motion prediction signal. In block 1860, encoders 100 or 300 and / or decoders 200 or 500 may determine one or more spatial gradients (e.g., associated with gradient filters) of the subblock-based motion prediction signal. In block 1865, encoders 100 or 300 and / or decoders 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined spatial gradients. In block 1870, encoders 100 or 300 and / or decoders 200 or 500 may combine (e.g., by adding or subtracting) the subblock-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. In block 1875, encoder 100 or 300 may encode video using a refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 may decode video using a refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 1810, 1820, 1830, and 1840 may be performed on the current block, which is generally the block that points to the block currently being encoded or decoded. The refined motion prediction signal may be a refined motion interpretation signal generated (for example, by GBi encoder 300 and / or GBi decoder 500), and may use one or more PROF operations.

[0187] For example, the determination of one or more spatial gradients for a subblock-based motion prediction signal by encoder 100 or 300 and / or decoder 200 or 500 may include the determination of a first set of spatial gradients associated with a first reference picture and a second set of spatial gradients associated with a second reference picture. The determination of a motion prediction refinement signal for the current block by encoder 100 or 300 and / or decoder 200 or 500 may be based on the determined spatial gradients and may include the determination of a motion interpretation refinement signal (e.g., a bidirectional prediction signal) for the current block based on the first and second sets of spatial gradients, and may also be based on weight information W (e.g., indicating or including one or more weight values ​​associated with one or more reference pictures).

[0188] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, an encoder 100 or 300 may generate, use, and / or transmit weight information W to a decoder 200 or 500, and / or the decoder 200 or 500 may receive or acquire the weight information W. For example, a motion interprediction refinement signal for a current block may be based on (1) a first gradient value derived from a first set of spatial gradients and weighted according to a first weight coefficient indicated by the weight information W, and / or (2) a second gradient value derived from a second set of spatial gradients and weighted according to a second weight coefficient indicated by the weight information W.

[0189] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may further include encoder 100 or 300 and / or decoder 200 or 500 determining affine motion model parameters for the current block of video so that a subblock-based motion prediction signal can be generated using the determined affine motion model parameters.

[0190] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining one or more spatial gradients of a subblock-based motion prediction signal by an encoder 100 or 300 and / or a decoder 200 or 500, the determination of which may include calculating at least one gradient value for each sample position, a portion of each sample position, or each respective sample position in at least one subblock of the subblock-based motion prediction signal. For example, calculating at least one gradient value for each sample position, a portion of each sample position, or each respective sample position in at least one subblock of the subblock-based motion prediction signal may include applying a gradient filter to each sample position in at least one subblock of the subblock-based motion prediction signal for each sample position, a portion of each sample position, or each respective sample position.

[0191] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may further include encoder 100 or 300 and / or decoder 200 or 500 determining a set of motion vector difference values ​​associated with the sample position of a first subblock of the current block in the subblock-based motion prediction signal. In some examples, the difference values ​​may be determined for a subblock (e.g., a first subblock) and may be reused for some or all other subblocks in the current block. In certain examples, an affine motion model or a different motion model (e.g., another subblock-based motion model such as the SbTMVP model) may be used to generate the subblock-based motion prediction signal and determine a set of motion vector difference values. For example, the set of motion vector difference values ​​may be determined for a first subblock of the current block and may be used to determine a motion prediction refinement signal for one or more further subblocks of the current block.

[0192] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, one or more spatial gradients of subblock-based motion prediction signals and a set of motion vector difference values ​​may be used to determine a motion prediction refinement signal for the current block.

[0193] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, an affine motion model for the current block is used to generate a subblock-based motion prediction signal and determine a set of motion vector difference values.

[0194] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, determining the spatial gradient of one or more subblock-based motion prediction signals may include determining an expanded subblock for each of one or more subblocks of the current block using the subblock-based motion prediction signal and a proximity reference sample adjacent to and surrounding the subblock, and determining the spatial gradient of each subblock using the determined expanded subblock in order to determine the motion prediction refinement signal.

[0195] Figure 19 is a flowchart showing a third representative encoding and / or decoding method.

[0196] Referring to Figure 19, a typical method 1900 for encoding and / or decoding video may include, in block 1910, encoders 100 or 300 and / or decoders 200 or 500 generating subblock-based motion prediction signals. In block 1920, encoders 100 or 300 and / or decoders 200 or 500 may determine a set of motion vector difference values ​​associated with the subblocks of the current block (for example, the set of motion vector difference values ​​may be associated with all of the subblocks of the current block). In block 1930, encoders 100 or 300 and / or decoders 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined set of motion vector difference values. In block 1940, encoder 100 or 300 and / or decoder 200 or 500 can combine (for example, by adding or subtracting) a subblock-based motion prediction signal and a motion prediction refinement signal to create or generate a refined motion prediction signal for the current block. In block 1950, encoder 100 or 300 can encode video using the refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 can decode video using the refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 1910, 1920, 1930, and 1940 may be performed on the current block, which generally refers to the block currently being encoded or decoded. In certain representative embodiments, the refined motion prediction signal may be a refined motion interpretation signal generated (for example, by a GBi encoder 300 and / or a GBi decoder 500), and may use one or more PROF operations.

[0197] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 determining motion model parameters (e.g., one or more affine motion model parameters) for the current block of video, so that a subblock-based motion prediction signal can be generated using the determined motion model parameters (e.g., affine motion model parameters).

[0198] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining one or more spatial gradients of a subblock-based motion prediction signal by an encoder 100 or 300 and / or a decoder 200 or 500. For example, determining one or more spatial gradients of a subblock-based motion prediction signal may include calculating at least one gradient value for each sample position, a portion of each sample position, or each respective sample position in at least one subblock of the subblock-based motion prediction signal. For example, calculating at least one gradient value for each sample position, a portion of each sample position, or each respective sample position in at least one subblock of the subblock-based motion prediction signal may include applying a gradient filter to each sample position in at least one subblock of the subblock-based motion prediction signal for each sample position, a portion of each sample position, or each respective sample position.

[0199] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining a motion prediction refinement signal for the current block using an encoder 100 or 300 and / or a decoder 200 or 500, a gradient value associated with the spatial gradient for each sample position of one of the current blocks, a portion of each sample position, or each respective sample position, and a determined set of motion vector difference values ​​associated with the sample positions of subblocks of the current block (e.g., any subblock) of the subblock motion prediction signal.

[0200] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the determination of the motion prediction refinement signal for the current block can be performed using a determined set of motion vector difference values ​​and gradient values ​​associated with the spatial gradient for one or more sample positions or each sample position of one or more sub-blocks of the current block.

[0201] Figure 20 is a flowchart showing a fourth representative encoding and / or decoding method.

[0202] Referring to Figure 20, a typical method 2000 for encoding and / or decoding video may include, in block 2010, encoder 100 or 300 and / or decoder 200 or 500 generating a subblock-based motion prediction signal using at least a first motion vector for a first subblock of the current block and further motion vectors for a second subblock of the current block. In block 2020, encoder 100 or 300 and / or decoder 200 or 500 may calculate a first set of gradient values ​​for a first sample position in the first subblock of the subblock-based motion prediction signal, and a second distinct set of gradient values ​​for a second sample position in the first subblock of the subblock-based motion prediction signal. In block 2030, encoder 100 or 300 and / or decoder 200 or 500 may determine a first set of motion vector difference values ​​for a first sample position, and a second distinct set of motion vector difference values ​​for a second sample position. For example, a first set of motion vector difference values ​​for a first sample position can represent the difference between the motion vector at the first sample position and the motion vector of the first subblock, and a second set of motion vector difference values ​​for a second sample position can represent the difference between the motion vector at the second sample position and the motion vector of the first subblock. In block 2040, encoders 100 or 300 and / or decoders 200 or 500 can use the first and second sets of gradient values ​​and the first and second sets of motion vector difference values ​​to determine the prediction refinement signal. In block 2050, encoders 100 or 300 and / or decoders 200 or 500 can combine the subblock-based motion prediction signal with the prediction refinement signal (for example, by adding or subtracting) to create a refined motion prediction signal.In block 2060, encoder 100 or 300 can encode video using a refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 can decode video using a refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 2010, 2020, 2030, 2040, and 2050 may be performed for a current block that includes multiple subblocks.

[0203] Figure 21 is a flowchart showing a fifth representative encoding and / or decoding method.

[0204] Referring to Figure 21, a typical method 2100 for encoding and / or decoding video may include, in block 2110, having encoders 100 or 300 and / or decoders 200 or 500 generate a subblock-based motion prediction signal for the current block. In block 2120, encoders 100 or 300 and / or decoders 200 or 500 may determine a prediction refinement signal using optical flow information that indicates the refined motion of multiple sample positions in the current block of the subblock-based motion prediction signal. In block 2130, encoders 100 or 300 and / or decoders 200 or 500 may combine the subblock-based motion prediction signal with the prediction refinement signal (e.g., by adding or subtracting) to create a refined motion prediction signal. In block 2140, encoder 100 or 300 can encode video using a refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 can decode video using a refined motion prediction signal as a prediction for the current block. For example, the current block may contain multiple subblocks, and a subblock-based motion prediction signal may be generated using at least a first motion vector for a first subblock of the current block and further motion vectors for a second subblock of the current block.

[0205] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include a determination by encoder 100 or 300 and / or decoder 200 or 500 of a predictive refinement signal that can use optical flow information. This determination may include the encoder 100 or 300 and / or decoder 200 or 500 calculating a first set of gradient values ​​for a first sample position in a first subblock of the subblock-based motion prediction signal, and a second distinct set of gradient values ​​for a second sample position in a first subblock of the subblock-based motion prediction signal. A first set of motion vector difference values ​​for the first sample position and a second distinct set of motion vector difference values ​​for the second sample position may be determined. For example, a first set of motion vector difference values ​​for a first sample position can represent the difference between the motion vector at the first sample position and the motion vector of a first subblock, and a second set of motion vector difference values ​​for a second sample position can represent the difference between the motion vector at the second sample position and the motion vector of the first subblock. Encoders 100 or 300 and / or decoders 200 or 500 can use the first and second sets of gradient values ​​and the first and second sets of motion vector difference values ​​to determine the predicted refinement signal.

[0206] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include a determination by encoder 100 or 300 and / or decoder 200 or 500 of a predictive refinement signal that can utilize optical flow information. This determination may include calculating a third set of gradient values ​​for a first sample position in a second subblock of the subblock-based motion prediction signal, and a fourth set of gradient values ​​for a second sample position in a second subblock of the subblock-based motion prediction signal. Encoder 100 or 300 and / or decoder 200 or 500 may use the third and fourth sets of gradient values ​​and the first and second sets of motion vector difference values ​​to determine a predictive refinement signal for a second subblock.

[0207] Figure 22 is a flowchart showing a sixth representative encoding and / or decoding method.

[0208] Referring to Figure 22, a typical method 2200 for encoding and / or decoding video may include, in block 2210, the encoder 100 or 300 and / or the decoder 200 or 500 determining a motion model for the current block of video. The current block may contain multiple subblocks. For example, the motion model may generate individual (e.g., sample-by-sample) motion vectors for multiple sample positions in the current block. In block 2220, the encoder 100 or 300 and / or the decoder 200 or 500 may use the determined motion model to generate a subblock-based motion prediction signal for the current block. The generated subblock-based motion prediction signal may use one motion vector for each subblock of the current block. In block 2230, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate gradient values ​​by applying a gradient filter to some of the multiple sample positions in the subblock-based motion prediction signal. In block 2240, encoder 100 or 300 and / or decoder 200 or 500 can determine motion vector difference values ​​for a subset of sample locations, each of which can represent the difference between the motion vector generated for each sample location according to the motion model (e.g., individual motion vectors) and the motion vector used to create a subblock-based motion prediction signal for the subblock containing each sample location. In block 2250, encoder 100 or 300 and / or decoder 200 or 500 can use the gradient values ​​and motion vector difference values ​​to determine a prediction refinement signal. In block 2260, encoder 100 or 300 and / or decoder 200 or 500 can combine the subblock-based motion prediction signal with the prediction refinement signal (e.g., by adding or subtracting, in particular) to create a refined motion prediction signal for the current block.In block 2270, encoder 100 or 300 can encode video using a refined motion prediction signal as a prediction for the current block, and / or decoder 200 or 500 can decode video using a refined motion prediction signal as a prediction for the current block.

[0209] Figure 23 is a flowchart showing the seventh representative encoding and / or decoding method.

[0210] Referring to Figure 23, a typical method 2300 for encoding and / or decoding video may include, in block 2310, having encoders 100 or 300 and / or decoders 200 or 500 perform subblock-based motion compensation to generate a subblock-based motion prediction signal as a coarse motion prediction signal. In block 2320, encoders 100 or 300 and / or decoders 200 or 500 may calculate one or more spatial gradients of the subblock-based motion prediction signal at the sample location. In block 2330, encoders 100 or 300 and / or decoders 200 or 500 may calculate the per-pixel intensity change in the current block based on the calculated spatial gradient. In block 2340, encoders 100 or 300 and / or decoders 200 or 500 may determine a pixel-based motion prediction signal as a refined motion prediction signal based on the calculated per-pixel intensity change. In block 2350, encoder 100 or 300 and / or decoder 200 or 500 can predict the current block using coarse motion prediction signals for each subblock of the current block and using refined motion prediction signals for each pixel of the current block. In certain embodiments, the operations in blocks 2310, 2320, 2330, 2340, and 2350 may be performed for at least one block in the video (e.g., the current block). For example, calculating the intensity change per pixel in the current block may include determining the luminance intensity change for each pixel in the current block according to an optical flow formula. Predicting the current block may include predicting the motion vector for each individual pixel in the current block by combining a coarse motion prediction vector for the subblock containing each pixel with a refined motion prediction vector for the coarse motion prediction vector, which is associated with each pixel.

[0211] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, one or more spatial gradients of the subblock-based motion prediction signal may include either a horizontal gradient and / or a vertical gradient, for example, the horizontal gradient may be calculated as the difference in luminance or chrominance between the right-adjacent sample of the subblock sample and the left-adjacent sample of the subblock sample, and / or the vertical gradient may be calculated as the difference in luminance or chrominance between the lower-adjacent sample of the subblock sample and the upper-adjacent sample of the subblock sample.

[0212] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, one or more spatial gradients of the subblock predictions may be generated using a Sobel filter.

[0213] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the coarse motion prediction signal can use either a four-parameter affine model or a six-parameter affine model. For example, subblock-based motion compensation may be either (1) affine subblock-based motion compensation, or (2) another type of compensation (e.g., subblock-based time-motion vector prediction (SbTMVP) mode-motion compensation, and / or regression-based motion vector field (RMVF) mode-based compensation). Provided that SbTMVP mode-based motion compensation is performed, this method may include estimating affine model parameters using the subblock motion vector field by linear regression and deriving pixel-level motion vectors using the estimated affine model parameters. Provided that RMVF mode-based motion compensation is performed, this method may include estimating affine model parameters and deriving pixel-level motion vector offsets from the subblock-level motion vectors using the estimated affine model parameters. For example, the pixel motion vector offset may be relative to the center of the subblock (e.g., the actual center, or the sample position closest to the actual center). For example, the coarse motion prediction vector for a subblock may be based on the actual center position of the subblock.

[0214] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 or decoder 200 or 500 selecting one of the following as the center position associated with a coarse motion prediction vector (e.g., a subblock-based motion prediction vector) for each subblock: (1) the actual center of each subblock, or (2) one of the pixel (e.g., sample) positions closest to the center of the subblock. For example, predicting the current block using a coarse motion prediction signal (e.g., a subblock-based motion prediction signal) of the current block, and using a refined motion prediction signal for each pixel (e.g., sample) of the current block, may be based on the selected center position of each subblock. For example, encoder 100 or 300 and / or decoder 200 or 500 can determine the center position associated with the color difference pixels of a subblock, and can determine an offset relative to the center position of the color difference pixels of the subblock based on the color difference position sample type associated with the color difference pixels. A coarse motion prediction signal for the subblock (e.g., a subblock-based motion prediction signal) can be based on the actual position of the subblock corresponding to the determined center position of the color difference pixels adjusted by the offset.

[0215] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 generating information in one of the following: (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header, or the decoder 200 or 500 receiving information indicating whether optical flow prediction refinement (PROF) is enabled. For example, under the condition that PROF is enabled, refined motion prediction operation may be performed, and therefore a coarse motion prediction signal (e.g., a subblock-based motion prediction signal) and a refined motion prediction signal may be used to predict the current block. As another example, under the condition that PROF is not enabled, refined motion prediction operation is not performed, and therefore only a coarse motion prediction signal (e.g., a subblock-based motion prediction signal) may be used to predict the current block.

[0216] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 and / or decoder 200 or 500 determining, based on the attributes of the current block and / or the attributes of the affine motion estimation, whether to perform a refined motion prediction action on the current block or in the affine motion estimation.

[0217] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include determining whether to perform a refined motion prediction action on the current block or in the affine motion estimation based on the attributes of the current block and / or the attributes of the affine motion estimation. For example, the decision on whether to perform a refined motion prediction action on the current block based on the attributes of the current block may include determining whether to perform a refined motion prediction action on the current block based on either (1) whether the size of the current block exceeds a certain size, and / or (2) whether the control point motion vector (CPMV) difference exceeds a threshold.

[0218] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 applying a first deblocking filter to one or more boundaries of subblocks of the current block that coincide with a conversion unit boundary, and applying a second different deblocking filter to other boundaries of subblocks of the current block that do not coincide with any conversion unit boundary. For example, the first deblocking filter may be a stronger deblocking filter than the second deblocking filter.

[0219] Figure 24 is a flowchart showing the eighth representative encoding and / or decoding method.

[0220] Referring to Figure 24, a typical method 2400 for encoding and / or decoding video may include, in block 2410, having encoder 100 or 300 and / or decoder 200 or 500 perform subblock-based motion compensation to generate a subblock-based motion prediction signal as a coarse motion prediction signal. In block 2420, encoder 100 or 300 and / or decoder 200 or 500 may determine, for each boundary sample of a subblock in the current block, one or more reference samples surrounding the subblock, corresponding to samples adjacent to each boundary sample, as surrounding reference samples, and may determine one or more spatial gradients associated with each boundary sample using the surrounding reference samples and samples of the subblock adjacent to each boundary sample. In block 2430, encoder 100 or 300 and / or decoder 200 or 500 may determine, for each non-boundary sample in a subblock, one or more spatial gradients associated with each non-boundary sample using samples of the subblock adjacent to each non-boundary sample. In block 2440, encoder 100 or 300 and / or decoder 200 or 500 can calculate the per-pixel intensity change in the current block using the determined spatial gradient of the subblock. In block 2450, encoder 100 or 300 and / or decoder 200 or 500 can determine a pixel-based motion prediction signal as a refined motion prediction signal based on the calculated per-pixel intensity change. In block 2460, encoder 100 or 300 and / or decoder 200 or 500 can predict the current block using the coarse motion prediction signal associated with each subblock of the current block and the refined motion prediction signal associated with each pixel of the current block.In certain embodiments, the operations in 2410, 2420, 2430, 2440, 2450, and 2460 may be performed on at least one block of video (e.g., the current block).

[0221] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, determining one or more spatial gradients of boundary and non-boundary samples may include calculating one or more spatial gradients using (1) a vertical Sobel filter, (2) a horizontal Sobel filter, or (3) a 3-tap filter.

[0222] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 copying the surrounding reference samples from a reference store without any further operation, and determining one or more spatial gradients associated with each boundary sample can be done using the copied, surrounding reference samples to determine one or more spatial gradients associated with each boundary sample.

[0223] Figure 25 is a flowchart showing typical gradient calculation methods.

[0224] Referring to Figure 25, a typical method 2500 for calculating the gradient of a subblock using reference samples corresponding to samples adjacent to the boundaries of the subblock (for example, used to encode and / or decode video) may include, in block 2510, encoder 100 or 300 and / or decoder 200 or 500 determining, for each boundary sample of a subblock in the current block, one or more reference samples that surround the subblock, corresponding to samples adjacent to each boundary sample, as surrounding reference samples, and determining one or more spatial gradients associated with each boundary sample using the surrounding reference samples and samples of the subblock adjacent to each boundary sample. In block 2520, encoder 100 or 300 and / or decoder 200 or 500 may determine, for each non-boundary sample in the subblock, one or more spatial gradients associated with each non-boundary sample using samples of the subblock adjacent to each non-boundary sample. In certain embodiments, the operations in blocks 2510 and 2520 may be performed on at least one block in the video (e.g., the current block).

[0225] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, 2500, and 2600, one or more determined spatial gradients may be used to predict the current block by any of the following: (1) predictive refinement (PROF) operation by optical flow, (2) bidirectional optical flow operation, or (3) affine motion estimation operation.

[0226] Figure 26 is a flowchart showing the ninth representative encoding and / or decoding method.

[0227] Referring to Figure 26, a typical method 2600 for encoding and / or decoding video may include the encoder 100 or 300 and / or the decoder 200 or 500 generating a subblock-based motion prediction signal for the current block of video. For example, the current block may include multiple subblocks. In block 2620, the encoder 100 or 300 and / or the decoder 200 or 500 may determine an expanded subblock for one or more subblocks of the current block, or for each individual subblock, using the subblock-based motion prediction signal and the proximity reference samples adjacent to and surrounding each subblock, and then determine the spatial gradient of each subblock using the determined expanded subblock. In block 2630, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined spatial gradient. In block 2640, encoder 100 or 300 and / or decoder 200 or 500 can combine (for example, by adding or subtracting) the subblock-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. In block 2650, encoder 100 or 300 can encode the video using the refined motion prediction signal as the prediction for the current block, and / or decoder 200 or 500 can decode the video using the refined motion prediction signal as the prediction for the current block. In certain embodiments, the operations in blocks 2610, 2620, 2630, 2640 and 2650 may be performed for at least one block in the video (for example, the current block).

[0228] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 copying nearby reference samples from a reference store without any further operation. For example, determining the spatial gradient of each subblock can be done by using the copied nearby reference samples to determine the gradient value associated with the sample position on the boundary of each subblock. Nearby reference samples of an extended block may be copied from the nearest integer position in the reference picture containing the current block. In certain examples, nearby reference samples of an extended block have the nearest integer motion vector rounded from the original precision.

[0229] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 determining affine motion model parameters for the current block of video so that a subblock-based motion prediction signal can be generated using the determined affine motion model parameters.

[0230] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2500, and 2600, determining the spatial gradient of each subblock may include calculating at least one gradient value for each respective sample position in each subblock. For example, calculating at least one gradient value for each respective sample position in each subblock may include applying a gradient filter to each respective sample position in each subblock for each respective sample position. As another example, calculating at least one gradient value for each respective sample position in each subblock may include determining the intensity change for each respective sample position in each subblock according to an optical flow formula.

[0231] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 determining a set of motion vector difference values ​​associated with the sample position of each subblock. For example, an affine motion model may be used for the current block to generate a subblock-based motion prediction signal and determine a set of motion vector difference values.

[0232] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, a set of motion vector difference values ​​may be determined for each subblock of the current block and used to determine motion prediction refinement signals for the other remaining subblocks of the current block.

[0233] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, determining the spatial gradient of each subblock may include calculating the spatial gradient using one of the following: (1) a vertical Sobel filter, (2) a horizontal Sobel filter, and / or (3) a three-tap filter.

[0234] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, proximity reference samples adjacent to and surrounding each subblock can use integer motion compensation.

[0235] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the spatial gradient of each subblock may include either a horizontal or vertical gradient. For example, the horizontal gradient may be calculated as the difference in luminance or chrominance between the right-adjacent sample of each sample and the left-adjacent sample of each sample, and / or the vertical gradient may be calculated as the difference in luminance or chrominance between the bottom-adjacent sample of each sample and the top-adjacent sample of each sample.

[0236] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the subblock-based motion prediction signal may be generated using any of the following: (1) a four-parameter affine model, (2) a six-parameter affine model, (3) subblock-based time-motion vector prediction (SbTMVP) mode motion compensation, or (4) regression-based motion compensation. For example, provided that SbTMVP mode motion compensation is performed, this method may include estimating affine model parameters using the subblock motion vector field by a linear regression operation, and / or deriving pixel-level motion vectors using the estimated affine model parameters. As another example, provided that RMVF mode-based motion compensation is performed, this method may include estimating affine model parameters, and / or deriving pixel-level motion vector offsets from the subblock-level motion vectors using the estimated affine model parameters. The pixel motion vector offsets may be relative to the center of each subblock.

[0237] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the refined motion prediction signal for each subblock can be based on the actual center position of each subblock, or on the sample position closest to the actual center of each subblock.

[0238] For example, these methods may include the encoder 100 or 300 and / or decoder 200 or 500 selecting either (1) the actual center of each subblock, or (2) the sample position closest to the actual center of each subblock, as the center position associated with the motion prediction vector for each subblock. The refined motion prediction signal may be based on the selected center position for each subblock.

[0239] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 determining the center position associated with the color difference pixels of each subblock and determining an offset relative to the center position of the color difference pixels of each subblock based on the color difference position sample type associated with the color difference pixels. The refined prediction signal for each subblock may be based on the actual position of the subblock corresponding to the determined center position of the color difference pixels adjusted by the offset.

[0240] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, encoder 100 or 300 may generate and transmit information in one of (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header indicating whether optical flow predictive refinement (PROF) is enabled, and / or decoder 200 or 500 may receive information in one of (1) an SPS header, (2) a PPS header, or (3) a tile group header indicating whether PROF is enabled.

[0241] Figure 27 is a flowchart showing the 10th representative encoding and / or decoding method.

[0242] Referring to Figure 27, a typical method 2700 for encoding and / or decoding video may include, in block 2710, encoder 100 or 300 and / or decoder 200 or 500 determining the actual center position of each subblock of the current block. In block 2720, encoder 100 or 300 and / or decoder 200 or 500 may use the actual center position of each subblock of the current block to generate a subblock-based motion prediction signal or a refined motion prediction signal. In block 2730, (1) encoder 100 or 300 may encode video using the subblock-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block, or (2) decoder 200 or 500 may decode video using the subblock-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 2710, 2720, and 2730 may be performed for at least one block in the video (e.g., the current block). For example, determining the actual center position of each subblock in the current block may include determining the center position of the color difference associated with the color difference pixels in each subblock, and the offset of the center position of the color difference relative to the center position of each subblock, based on the color difference position sample type of the color difference pixels. A subblock-based motion prediction signal or a refined motion prediction signal for each subblock may be based on the actual center position of each subblock, corresponding to the determined center position of the color difference adjusted by the offset. Although the actual center of each subblock in the current block is described as being determined / used for various operations, it is intended that one, some, or all of the center positions of such subblocks may be determined / used.

[0243] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the generation of a refined motion prediction signal can be achieved by using the subblock-based motion prediction signal by determining one or more spatial gradients of the subblock-based motion prediction signal for each subblock of the current block, determining a motion prediction refinement signal for the current block based on the determined spatial gradients, and / or combining the subblock-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. For example, determining one or more spatial gradients of the subblock-based motion prediction signal may include determining an extended subblock using the subblock-based motion prediction signal and proximity reference samples adjacent to and surrounding each subblock, and / or determining one or more spatial gradients for each subblock using the determined extended subblock.

[0244] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, determining the spatial gradient of each subblock may include calculating at least one gradient value for each respective sample position in each subblock. For example, calculating at least one gradient value for each respective sample position in each subblock may include applying a gradient filter to each respective sample position in each subblock for each respective sample position.

[0245] As another example, the calculation of at least one gradient value for each sample location in each subblock may include determining the intensity change for one or more sample locations in each subblock according to an optical flow formula.

[0246] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, these methods may include encoder 100 or 300 and / or decoder 200 or 500 determining a set of motion vector difference values ​​associated with the sample position of each subblock. An affine motion model may be used for the current block to generate a subblock-based motion prediction signal and determine a set of motion vector difference values. In certain examples, the set of motion vector difference values ​​may be determined for each subblock of the current block and may be used (e.g., reused) to determine a motion prediction refinement signal for that subblock and other remaining subblocks of the current block. For example, determining the spatial gradient of each subblock may include calculating the spatial gradient using one of (1) a vertical Sobel filter, (2) a horizontal Sobel filter, and / or (3) a three-tap filter. Each adjacent subblock and its surrounding neighbor reference samples can use integer motion compensation.

[0247] In some embodiments, the spatial gradient of each subblock may include either a horizontal or vertical gradient. For example, the horizontal gradient may be calculated as the difference in luminance or chrominance between the right-adjacent sample of each sample and the left-adjacent sample of each sample. As another example, the vertical gradient may be calculated as the difference in luminance or chrominance between the bottom-adjacent sample of each subblock and the top-adjacent sample of each subblock.

[0248] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the subblock-based motion prediction signal may be generated using any of the following: (1) a four-parameter affine model, (2) a six-parameter affine model, (3) subblock-based time-motion vector prediction (SbTMVP) mode motion compensation, and / or (4) regression-based motion compensation. For example, provided that SbTMVP mode motion compensation is performed, this method may include estimating affine model parameters using the subblock motion vector field by a linear regression operation, and / or deriving pixel-level motion vectors using the estimated affine model parameters. As another example, given that regression motion vector field (RMVF) mode-based motion compensation is performed, this method may involve estimating affine model parameters and / or using the estimated affine model parameters to derive pixel-level motion vector offsets from subblock-level motion vectors, where the pixel motion vector offsets are relative to the center of each subblock.

[0249] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the refined motion prediction signal may be generated using multiple motion vectors associated with the control points of the current block.

[0250] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, encoder 100 or 300 can generate, encode, and transmit information in one of (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header, indicating whether optical flow predictive refinement (PROF) is enabled, which decoder 200 or 500 can receive and decode.

[0251] Figure 28 is a flowchart showing the 11th representative encoding and / or decoding method.

[0252] Referring to Figure 28, a typical method 2800 for encoding and / or decoding video may include, in block 2810, having encoders 100 or 300 and / or decoders 200 or 500 select either (1) the actual center of each subblock, or (2) the sample position closest to the actual center of each subblock, as the center position associated with the motion prediction vector for each subblock. In block 2820, encoders 100 or 300 and / or decoders 200 or 500 can determine the selected center position for each subblock of the current block. In block 2830, encoders 100 or 300 and / or decoders 200 or 500 can use the selected center position for each subblock of the current block to generate a subblock-based motion prediction signal or a refined motion prediction signal. In block 2840, (1) encoder 100 or 300 may encode the video using a subblock-based motion prediction signal or a generated refined motion prediction signal as a prediction for the current block, or (2) decoder 200 or 500 may decode the video using a subblock-based motion prediction signal or a generated refined motion prediction signal as a prediction for the current block. In certain embodiments, the operations in blocks 2810, 2820, 2830, and 2840 may be performed for at least one block in the video (e.g., the current block). While the selection of a central position is described for each subblock of the current block, it is intended that one, some, or all of such subblock central positions may be selected / used in various operations.

[0253] Figure 29 shows a flowchart of typical encoding methods.

[0254] Referring to Figure 29, a typical method 2900 for encoding video may include, in block 2910, encoder 100 or 300 performing motion estimation for the current block of video, which includes determining affine motion model parameters for the current block using iterative motion compensation operations and generating a subblock-based motion prediction signal for the current block using the determined affine motion model parameters. In block 2920, after performing motion estimation for the current block, encoder 100 or 300 performs optical flow prediction refinement (PROF) operations to generate a refined motion prediction signal. In block 2930, encoder 100 or 300 can encode video using the refined motion prediction signal as the prediction for the current block. For example, PROF operation may include determining one or more spatial gradients of subblock-based motion prediction signals, determining a motion prediction refinement signal for the current block based on the determined spatial gradients, and / or combining the subblock-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block.

[0255] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, and 2900, the PROF operation may be performed only after the iterative motion compensation operation is completed. For example, the PROF operation is not performed during motion estimation for the current block.

[0256] Figure 30 is a flowchart showing another typical encoding method.

[0257] Referring to Figure 30, a typical method 3000 for encoding video may include, in block 3010, encoder 100 or 300 determining affine motion model parameters using iterative motion compensation during motion estimation for the current block, and generating a subblock-based motion prediction signal using the determined affine motion model parameters. In block 3020, encoder 100 or 300 may, after motion estimation for the current block, perform optical flow prediction refinement (PROF) operation to generate a refined motion prediction signal, provided that the size of the current block meets or exceeds a threshold size. In block 3030, encoder 100 or 300 may encode video using (1) the refined motion prediction signal as the prediction for the current block, provided that the current block meets or exceeds a threshold size, or (2) the subblock-based motion prediction signal as the prediction for the current block, provided that the current block does not meet a threshold size.

[0258] Figure 31 is a flowchart showing the 12th representative encoding / decoding method.

[0259] Referring to Figure 31, a typical method 3100 for encoding and / or decoding video may include, in block 3110, an encoder 100 or 300 determining or acquiring information indicating the size of the current block, or a decoder 200 or 500 receiving information indicating the size of the current block. In block 3120, the encoder 100 or 300 or the decoder 200 or 500 may generate a subblock-based motion prediction signal. In block 3130, the encoder 100 or 300 or the decoder 200 or 500 may perform an optical flow prediction refinement (PROF) operation to generate a refined motion prediction signal, provided that the size of the current block meets or exceeds a threshold size. In block 3140, encoder 100 or 300 can encode video using a refined motion prediction signal as a prediction for the current block if (1) the current block meets or exceeds a threshold size, or (2) the current block does not meet a threshold size, or decoder 200 or 500 can decode video using a refined motion prediction signal as a prediction for the current block if (1) the current block meets or exceeds a threshold size, or (2) the current block does not meet a threshold size, or the current block does not meet a threshold size, or decoder 200 or 500 can decode video using a refined motion prediction signal as a prediction for the current block if (1) the current block meets or exceeds a threshold size, or (2) the current block does not meet a threshold size, or

[0260] Figure 32 is a flowchart showing the 13th representative encoding / decoding method.

[0261] Referring to Figure 32, a typical method 3200 for encoding and / or decoding video may include, in block 3210, an encoder 100 or 300 determining whether pixel-level motion compensation should be performed, or a decoder 200 or 500 receiving a flag indicating whether pixel-level motion compensation should be performed. In block 3220, the encoder 100 or 300 or the decoder 200 or 500 may generate a subblock-based motion prediction signal. In block 3230, given that pixel-level motion compensation should be performed, the encoder 100 or 300 or the decoder 200 or 500 may determine one or more spatial gradients of the subblock-based motion prediction signal, determine a motion prediction refinement signal for the current block based on the determined spatial gradients, and combine the subblock-based motion prediction signal and the motion prediction refinement signal to create a refined motion prediction signal for the current block. In block 3240, depending on the decision of whether pixel-level motion compensation should be performed, encoder 100 or 300 may encode the video using a subblock-based motion prediction signal or a refined motion prediction signal as a prediction for the current block, or decoder 200 or 500 may decode the video using a subblock-based motion prediction signal or a refined motion prediction signal as a prediction for the current block, depending on the display of a flag. In certain embodiments, the operations in blocks 3220 and 3230 may be performed on a block in the video (e.g., the current block).

[0262] Figure 33 is a flowchart showing the 14th representative encoding / decoding method.

[0263] Referring to Figure 33, a typical method 3300 for encoding and / or decoding video may include, in block 3310, encoder 100 or 300 determining or acquiring interprediction weight information indicating one or more weights associated with a first reference picture and a second reference picture, or decoder 200 or 500 receiving it. In block 3320, encoder 100 or 300 or decoder 200 or 500 may generate a subblock-based motion interprediction signal for the current block of video, determine a first set of spatial gradients associated with the first reference picture and a second set of spatial gradients associated with the second reference picture, determine a motion interprediction refinement signal for the current block based on the first and second sets of spatial gradients and the interprediction weight information, and combine the subblock-based motion interprediction signal and the motion interprediction refinement signal to create a refined motion interprediction signal for the current block. In block 3330, encoder 100 or 300 can encode video using a refined motion interprediction signal as a prediction for the current block, or decoder 200 or 500 can decode video using a refined motion interprediction signal as a prediction for the current block. For example, the interprediction weight information is either (1) an indicator showing a first weight coefficient applied to a first reference picture and / or a second weight coefficient applied to a second reference picture, or (2) a weight index. In certain embodiments, the motion interprediction refinement signal for the current block may be based on (1) a first gradient value derived from a first set of spatial gradients and weighted according to a first weight coefficient indicated by the interprediction weight information, and (2) a second gradient value derived from a second set of spatial gradients and weighted according to a second weight coefficient indicated by the interprediction weight information.

[0264] Exemplary Network for Implementation of the Embodiment Figure 34A shows an exemplary communication system 3400 in which one or more disclosed embodiments may be implemented. The communication system 3400 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcast to multiple wireless users. The communication system 3400 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 3400 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), quadrature FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT spread OFDM (ZT UW DTS-s OFDM), unique-word OFDM (UW-OFDM), resource-block filtered OFDM, and filtered-bank multi-carrier (FBMC).

[0265] As shown in Figure 34A, the communication system 3400 may include wireless transmit / receive units (WTRUs) 3402a, 3402b, 3402c, 3402d, RAN 3404 / 3413, CN 3406 / 3415, public switched telephone network (PSTN) 3408, the Internet 3410, and other networks 3412, but it will be understood that the disclosed embodiments intend any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 3402a, 3402b, 3402c, and 3402d can be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRU3402a, 3402b, 3402c, and 3402d may all be referred to as “stations” and / or “STAs,” which may be configured to transmit and / or receive wireless signals, and may include user equipment (UEs), mobile stations, fixed or mobile subscriber units, subscriber-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in an industrial and / or automated processing chain context), consumer electronics, and devices operating on commercial and / or industrial wireless networks. WTRU3402a, 3402b, 3402c, and 3402d may all be interchangeably referred to as UEs.

[0266] The communication system 3400 may also include base stations 3414a and / or base station 3414b. Each of the base stations 3414a and 3414b can be any type of device configured to wirelessly interface with at least one of the WTRUs 3402a, 3402b, 3402c, and 3402d to facilitate access to one or more communication networks such as CN3406 / 3415, the Internet 3410, and / or network 3412. For example, base stations 3414a and 3414b may be transceiver base stations (BTS), node B, enode B (end), home node B (HNB), home enode B (HeNB), gNB, NR node B, site controller, access point (AP), and wireless router, etc. Although base stations 3414a and 3414b are shown as single elements, it will be understood that base stations 3414a and 3414b can contain any number of interconnected base stations and / or network elements.

[0267] Base station 3414a may be part of RAN 3404 / 3413, which may also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), and relay nodes. Base stations 3414a and / or base stations 3414b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be called cells (not shown). These frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell can provide coverage for wireless service to a particular geographic area that may be relatively fixed or change over time. A cell may be further divided into cell sectors. For example, a cell associated with base station 3414a may be divided into three sectors. Thus, in one embodiment, base station 3414a may include three transceivers, i.e., one transceiver per sector of the cell. In an embodiment, the base station 3414a can utilize multiple-input multiple-output (MIMO) technology and utilize multiple transceivers per sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0268] Base stations 3414a and 3414b can communicate with one or more WTRUs 3402a, 3402b, 3402c, and 3402d via an air interface 3416, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 3416 can be established using any suitable radio access technology (RAT).

[0269] More specifically, as described above, the communication system 3400 can be a multiple access system and can utilize one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, base stations 3414a and WTRU 3402a, 3402b, and 3402c in RAN 3404 / 3413 can implement radio technologies such as Universal Mobile Communications System (UMTS) Terrestrial Radio Access (UTRA) which can establish air interfaces 3415 / 3416 / 3417 using broadband CDMA (WCDMA). WCDMA may include communication protocols such as High Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed ​​Downlink (DL) Packet Access (HSDPA) and / or High Speed ​​UL Packet Access (HSUPA).

[0270] In embodiments, base stations 3414a and WTRUs 3402a, 3402b, and 3402c can implement radio technologies such as Advanced UMTS Terrestrial Radio Access (E-UTRA) that can establish an air interface 3416 using Long-Term Evolution (LTE) and / or LTE Advanced (LTE-A) and / or LTE Advanced Pro (LTE-A Pro).

[0271] In this embodiment, base stations 3414a and WTRUs 3402a, 3402b, and 3402c can implement radio technologies such as NR radio access, which can establish an air interface 3416 using NewRadio (NR).

[0272] In embodiments, base stations 3414a and WTRUs 3402a, 3402b, and 3402c can implement multiple radio access technologies. For example, base stations 3414a and WTRUs 3402a, 3402b, and 3402c can implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Thus, the air interface utilized by WTRUs 3402a, 3402b, and 3402c may be characterized by multiple types of radio access technologies and / or transmissions sent to and from multiple types of base stations (e.g., end and gNB).

[0273] In other embodiments, base stations 3414a and WTRUs 3402a, 3402b, and 3402c can implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Global Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), Extended Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0274] In Figure 34A, base station 3414b can be, for example, a wireless router, home node B, home e-node B, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in localized areas such as workplaces, homes, vehicles, campuses, industrial facilities, aerial walkways (used by drones, for example), and roadways. In one embodiment, base stations 3414b and WTRU3402c, 3402d can establish a wireless local area network (WLAN) by implementing radio technologies such as IEEE 802.11. In another embodiment, base stations 3414b and WTRU3402c, 3402d can establish a wireless personal area network (WPAN) by implementing radio technologies such as IEEE 802.15. In yet another embodiment, base stations 3414b and WTRU3402c, 3402d can establish a picocell or femtocell using cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.). As shown in Figure 34A, base station 3414b can have a direct connection to the internet 3410. Therefore, base station 3414b does not need to be required to access the internet 3410 via CN3406 / 3415.

[0275] RAN3404 / 3413 can communicate with CN3406 / 3415, which can be any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRU3402a, 3402b, 3402c, and 3402d. The data may have various Quality of Service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN3406 / 3415 can provide call control, billing services, mobile location-based services, prepaid calls, internet connectivity, video distribution, and / or perform high-level security functions such as user authentication. Although not shown in Figure 34A, it will be understood that RAN1084 / 3413 and / or CN3406 / 3415 are capable of direct or indirect communication with other RANs that utilize the same or different RATs as RAN3404 / 3413. For example, in addition to connecting to RAN3404 / 3413 which may utilize NR radio technology, CN3406 / 3415 can also communicate with other RANs (not shown) that utilize GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0276] CN3406 / 3415 can also act as a gateway for WTRU3402a, 3402b, 3402c, 3402d to access PSTN3408, the Internet3410, and / or other networks3412. PSTN3408 may include a circuit-switched telephone network providing basic telephone services (POTS). The Internet3410 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet Protocol suite. Network3412 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network3412 may include another CN connected to one or more RANs that have the same or different RATs as RAN3404 / 3413.

[0277] Some or all of the WTRUs 3402a, 3402b, 3402c, and 3402d in the communication system 3400 can include multimode capability (for example, WTRUs 3402a, 3402b, 3402c, and 3402d can include multiple transceivers for communicating with different wireless networks via different wireless links). For example, WTRU 3402c, shown in Figure 34A, may be configured to communicate with base station 3414a, which can utilize cellular-based radio technology, and base station 3414b, which can utilize IEEE 802 radio technology.

[0278] Figure 34B is a system diagram showing an exemplary WTRU3402. As shown in Figure 34B, the WTRU3402 may include a processor 3418, a transceiver 3420, a transmit / receive element 3422, a speaker / microphone 3424, a keypad 3426, a display / touchpad 3428, a non-removable memory 3430, a removable memory 3432, a power supply 3434, a Global Positioning System (GPS) chipset 3436, and / or other peripherals 3438, etc. It will be understood that the WTRU3402 may include any partial combination of the above elements while maintaining consistency with the embodiment.

[0279] The processor 3418 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. The processor 3418 can perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU3402 to operate in a wireless environment. The processor 3418 may be coupled to a transceiver 3420, and the transceiver 3420 may be coupled to a transmit / receive element 3422. Although Figure 34B shows the processor 3418 and transceiver 3420 as separate components, it will be understood that the processor 3418 and transceiver 3420 may be integrated together in an electronic package or chip. The processor 3418 may be configured to encode or decode video (e.g., video frames).

[0280] The transmit / receive element 3422 may be configured to transmit signals to or receive signals from a base station (e.g., base station 3414a) via the air interface 3416. For example, in one embodiment, the transmit / receive element 3422 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 3422 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 3422 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmit / receive element 3422 may be configured to transmit and / or receive any combination of wireless signals.

[0281] In Figure 34B, the transmit / receive element 3422 is shown as a single element, but the WTRU 3402 can include any number of transmit / receive elements 3422. More specifically, the WTRU 3402 can utilize MIMO technology. Thus, in one embodiment, the WTRU 3402 can include two or more transmit / receive elements 3422 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 3416.

[0282] Transceiver 3420 may be configured to modulate the signal transmitted by the transmit / receive element 3422 and demodulate the signal received by the transmit / receive element 3422. As described above, WTRU3402 may have multimode capability. Therefore, transceiver 3420 may include multiple transceivers to enable WTRU3402 to communicate via multiple RATs, such as NR and IEEE802.11.

[0283] The processor 3418 of the WTRU3402 can be coupled to a speaker / microphone 3424, a keypad 3426, and / or a display / touchpad 3428 (for example, a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and can receive user input data from them. The processor 3418 can also output user data to the speaker / microphone 3424, the keypad 3426, and / or the display / touchpad 3428. In addition, the processor 3418 can access information from any type of suitable memory, such as non-removable memory 3430 and / or removable memory 3432, and can store data in them. Non-removable memory 3430 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. Removable memory 3432 may include subscriber identification module (SIM) cards, memory sticks, and secure digital (SD) memory cards, etc. In other embodiments, the processor 3418 can access information from memory located on a server or home computer (not shown) that is not physically located on the WTRU 3402, and can store data therein.

[0284] The processor 3418 can receive power from the power supply 3434 and may be configured to distribute and / or control power to other components within the WTRU 3402. The power supply 3434 can be any suitable device for supplying power to the WTRU 3402. For example, the power supply 3434 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, and a fuel cell.

[0285] The processor 3418 may be coupled to a GPS chipset 3436, which may be configured to provide location information (e.g., longitude and latitude) about the current location of the WTRU 3402. In addition to or instead of the information from the GPS chipset 3436, the WTRU 3402 can receive location information from base stations (e.g., base stations 3414a, 3414b) via the air interface 3416 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 3402 can acquire location information by any suitable location determination method while maintaining consistency with the embodiments.

[0286] The processor 3418 may be further coupled to other peripherals 3438, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripherals 3438 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, and an activity tracker. The peripheral device 3438 may include one or more sensors, one or more of which are gyroscopes, accelerometers, Hall effect sensors, magnetometers, compass sensors, proximity sensors, temperature sensors, time sensors, geographic position sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors.

[0287] The processor 3418 of the WTRU3402 can operately communicate with a variety of peripheral devices 3438, including, for example, one or more accelerometers, one or more gyroscopes, a USB port, other communication interfaces / ports, a display, and / or other visual / audio indicators, in order to implement a typical embodiment disclosed herein.

[0288] WTRU3402 may include a full-duplex radio, in which the transmission and reception of some or all of the signals associated with a particular subframe for both UL (e.g., for transmission) and downlink (e.g., for reception) may be in parallel and / or simultaneous. The full-duplex radio may include an interference management unit for reducing or substantially eliminating self-interference through signal processing by hardware (e.g., chokes) or a processor (e.g., a separate processor (not shown) or processor 3418). In embodiments, WTRU3402 may include a half-duplex radio for the transmission and reception of some or all of the signals (e.g., associated with a particular subframe for either UL (e.g., for transmission) or downlink (e.g., for reception)).

[0289] Figure 34C is a system diagram showing RAN104 and CN3406 according to the embodiment. As described above, RAN3404 can communicate with WTRU3402a, 3402b, and 3402c via air interface 3416 using E-UTRA wireless technology. RAN3404 can also communicate with CN3406.

[0290] RAN3404 may include enodes B3460a, 3460b, and 3460c, but it will be understood that RAN3404 may include any number of enodes B while maintaining consistency with the embodiment. Each of enodes B3460a, 3460b, and 3460c may include one or more transceivers for communicating with WTRU3402a, 3402b, and 3402c via the air interface 3416. In one embodiment, enodes B3460a, 3460b, and 3460c can implement MIMO technology. Thus, enode B3460a may, for example, use multiple antennas to transmit a wireless signal to and / or receive a wireless signal from WTRU3402a.

[0291] Each of the e-nodes B3460a, 3460b, and 3460c may be associated with a specific cell (not shown) and may be configured to handle wireless resource management decisions, handover decisions, and user scheduling in UL and / or DL. As shown in Figure 34C, the e-nodes B3460a, 3460b, and 3460c can communicate with each other via the X2 interface.

[0292] The CN3406 shown in Figure 34C may include a Mobility Management Entity (MME) 3462, a Serving Gateway (SGW) 3464, and a Packet Data Network (PDN) Gateway (or PGW) 3466. Although each of the above elements is shown as part of CN3406, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.

[0293] The MME3462 may be connected to each of the e-nodes B3460a, 3460b, and 3460c within RAN3404 via the S1 interface and can act as a control node. For example, the MME3462 can be responsible for authenticating users of WTRU3402a, 3402b, and 3402c, activating / deactivating bearers, and selecting a specific serving gateway during the initial attachment of WTRU3402a, 3402b, and 3402c. The MME3462 can provide control plane functionality for switching between RAN3404 and other RANs (not shown) utilizing other radio technologies such as GSM and / or WCDMA.

[0294] The SGW3464 can be connected to each of the e-nodes B3460a, 3460b, and 3460c in RAN104 via the S1 interface. The SGW3464 can generally route and forward user data packets to and from WTRU3402a, 3402b, and 3402c. The SGW3464 can also perform other functions, such as anchoring the user plane during e-node B handovers, triggering paging when DL data is available to WTRU3402a, 3402b, and 3402c, and managing and remembering the context of WTRU3402a, 3402b, and 3402c.

[0295] SGW3464 may be connected to PGW3466, which provides WTRU3402a, 3402b, and 3402c with access to packet-switched networks such as the Internet 3410, thereby facilitating communication between WTRU3402a, 3402b, and 3402c and IP-enabled devices.

[0296] CN3406 can facilitate communication with other networks. For example, CN106 can provide WTRU3402a, 3402b, and 3402c with access to circuit-switched networks such as PSTN3408, thereby facilitating communication between WTRU3402a, 3402b, and 3402c and conventional land-line communication devices. For example, CN3406 may include, or communicate with, an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) acting as an interface between CN3406 and PSTN3408. In addition, CN3406 can provide WTRU3402a, 3402b, and 3402c with access to other networks 3412, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0297] Although the WTRU is described as a wireless terminal in Figures 34A to 34D, in certain representative embodiments, such a terminal is intended to be able to use a wired communication interface with a communication network (for example, temporarily or permanently).

[0298] In a typical embodiment, the other network 3412 may be a WLAN.

[0299] In Infrastructure Basic Service Set (BSS) mode, a WLAN may have access points (APs) for the BSS and one or more stations (STAs) associated with the APs. APs may have access to or interfaces with a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic originating from outside the BSS to an STA may arrive via the AP and be delivered to the STA. Traffic originating from an STA destined for a destination outside the BSS may be sent to the AP and delivered to its respective destination. Traffic between STAs within the BSS may be sent through the AP; for example, a source STA may send traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between a source STA and a destination STA (e.g., directly between them) using a Direct Link Setup (DLS). In certain representative embodiments, the DLS may be an 802.11e DLS or an 802.11z Tunnel DLS (TDLS). A WLAN using Independent BSS (IBSS) mode does not need to have an AP, and STAs within or using IBSS (e.g., all STAs) may communicate directly with each other. The IBSS communication mode may be referred to as the “ad hoc” communication mode in this specification.

[0300] When using the 802.11ac infrastructure operating mode or a similar operating mode, an AP can transmit beacons on a fixed channel, such as the primary channel. The primary channel may have a fixed width (e.g., a 20 MHz bandwidth) or a dynamically set width via signaling. The primary channel may be the operating channel of the BSS and may be used by the STA to establish a connection with the AP. In certain representative embodiments, Carrier Sensitivity Multiple Access / Collision Avoidance (CSMA / CA) may be implemented, for example, in an 802.11 system. In CSMA / CA, the STA, including the AP (e.g., all STAs), can sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, that particular STA may back off. One STA (e.g., a single station) may transmit at any given time on a given BSS.

[0301] High-throughput (HT) STAs can use a 40MHz wide channel for communication by, for example, using a combination of adjacent or non-adjacent 20MHz channels and a primary 20MHz channel to form a 40MHz wide channel.

[0302] Ultra-high throughput (VHT) STAs can support 20MHz, 40MHz, 80MHz, and / or 160MHz wide channels. 40MHz and / or 80MHz channels may be formed by combining consecutive 20MHz channels. 160MHz channels may be formed by combining eight consecutive 20MHz channels, or by combining two non-contiguous 80MHz channels, sometimes referred to as an 80+80 configuration. In an 80+80 configuration, data may be passed through a segment parser that can separate the data into two streams after channel encoding. Inverse fast Fourier transform (IFFT) processing and time-domain processing may be performed separately for each stream. The streams may be mapped onto two 80MHz channels, and the data may be transmitted by a transmitting STA. At the receiver of a receiving STA, the operation described above for the 80+80 configuration may be reversed, and the combined data may be sent to a media access control (MAC).

[0303] Sub-1GHz operating modes are supported by 802.11af and 802.11ah. Channel operating bandwidth and carrier are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5MHz, 10MHz, and 20MHz bandwidths in the TV white space (TVWS) spectrum, while 802.11ah supports 1MHz, 2MHz, 4MHz, 8MHz, and 16MHz bandwidths using the non-TVWS spectrum. According to a typical embodiment, 802.11ah can support meter-type control / machine-type communications, such as MTC devices in a macro coverage area. MTC devices may have limited capabilities, including support for specific and / or limited bandwidths (e.g., support for only that). MTC devices may include batteries with a battery life above a threshold (e.g., to maintain a very long battery life).

[0304] A WLAN system that can support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, includes a channel that can be designated as the primary channel. The primary channel may have a bandwidth equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In the 802.11ah example, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, the primary channel may be 1MHz wide for an STA (e.g., an MTC type device) that supports (e.g., only) 1MHz mode. Carrier discovery and / or network allocation vector (NAV) settings may depend on the status of the primary channel. For example, if an STA (which only supports 1MHz operation mode) is transmitting to an AP and the primary channel is busy, the entire available frequency band may be considered busy, even if a large portion of the frequency band could remain idle and available.

[0305] In the United States, the available frequency band that can be used by 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11ah is from 6 MHz to 26 MHz, depending on the country code.

[0306] Figure 34D is a system diagram showing RAN3413 and CN3415 in an embodiment. As described above, RAN3413 can communicate with WTRU3402a, 3402b, and 3402c via air interface 3416 using NR radio technology. RAN3413 can also communicate with CN3415.

[0307] While RAN3413 may include gNB3480a, 3480b, and 3480c, it will be understood that RAN3413 may include any number of gNBs while maintaining consistency with the embodiment. Each of gNB3480a, 3480b, and 3480c may include one or more transceivers for communicating with WTRU3402a, 3402b, and 3402c via the air interface 3416. In one embodiment, gNB3480a, 3480b, and 3480c can implement MIMO technology. For example, gNB3480a and 3480b can utilize beamforming to transmit signals to and / or receive signals from gNB3480a, 3480b, and 3480c. Therefore, gNB3480a can, for example, use multiple antennas to transmit a wireless signal to WTRU3402a and / or receive a wireless signal from WTRU3402a. In embodiments, gNB3480a, 3480b, and 3480c can implement carrier aggregation technology. For example, gNB3480a can transmit multiple component carriers to WTRU3402a (not shown). A subset of these component carriers may be on the unlicensed spectrum, while the remaining component carriers may be on the licensed spectrum. In embodiments, gNB3480a, 3480b, and 3480c can implement coordinated multipoint (CoMP) technology. For example, WTRU102a can receive coordinated transmissions from gNB3480a and gNB3480b (and / or gNB3480c).

[0308] WTRU3402a, 3402b, and 3402c can communicate with gNB480a, 3480b, and 3480c using transmissions associated with scalable numerology. For example, OFDM symbol intervals and / or OFDM subcarrier intervals may vary by different transmissions, different cells, and / or different parts of the wireless transmission spectrum. WTRU3402a, 3402b, and 3402c can communicate with gNB3480a, 3480b, and 3480c using subframes or transmit time intervals (TTIs) of varying or scalable lengths (e.g., containing varying numbers of OFDM symbols and / or continuing through varying lengths of absolute time).

[0309] The gNB3480a, 3480b, and 3480c can be configured to communicate with WTRU3402a, 3402b, and 3402c in standalone and / or non-standalone configurations. In a standalone configuration, the WTRU3402a, 3402b, and 3402c can communicate with the gNB3480a, 3480b, and 3480c without accessing other RANs (e.g., e-nodes B3460a, 3460b, and 3460c). In a standalone configuration, the WTRU3402a, 3402b, and 3402c can use one or more of the gNB3480a, 3480b, and 3480c as mobility anchor points. In a standalone configuration, WTRU3402a, 3402b, and 3402c can communicate with gNB3480a, 3480b, and 3480c using signals in the unlicensed band. In a non-standalone configuration, WTRU3402a, 3402b, and 3402c can communicate with gNB3480a, 3480b, and 3480c while also communicating with other RANs such as enodes B3460a, 3460b, and 3460c. For example, WTRU3402a, 3402b, and 3402c can implement the DC principle to communicate substantially simultaneously with one or more gNB3480a, 3480b, and 3480c and one or more enodes B3460a, 3460b, and 3460c. In a non-standalone configuration, e-nodes B3460a, 3460b, and 3460c can act as mobility anchors for WTRU3402a, 3402b, and 3402c, while gNB3480a, 3480b, and 3480c can provide additional coverage and / or throughput to service WTRU3402a, 3402b, and 3402c.

[0310] Each of the gNB3480a, 3480b, and 3480c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPF) 3484a and 3484b, and routing of control plane information to access and mobility management functions (AMF) 3482a and 3482b. As shown in Figure 34D, the gNB3480a, 3480b, and 3480c can communicate with each other via the Xn interface.

[0311] The CN3415 shown in Figure 34D may include at least one AMF3482a, 3482b, at least one UPF3484a, 3484b, at least one Session Management Function (SMF)3483a, 3483b, and possibly a Data Network (DN)3485a, 3485b. Although each of the above elements is shown as part of the CN3415, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.

[0312] AMF3482a and 3482b may be connected to one or more of gNB3480a, 3480b, and 3480c within RAN3413 via the N2 interface and can act as control nodes. For example, AMF3482a and 3482b may be responsible for authenticating users of WTRU3402a, 3402b, and 3402c, supporting network slicing (e.g., handling different protocol data unit (PDU) sessions with different requirements), selecting specific SMF3483a and 3483b, managing registration areas, terminating non-accessible tier (NAS) signaling, and mobility management. Network slicing may be used by AMF3482a and 3482b to customize CN support for WTRU3402a, 3402b, and 3402c based on the type of service being utilized by WTRU3402a, 3402b, and 3402c. For example, different network slices may be established for different use cases, such as services that rely on ultra-high reliability low latency (URLLC) access, services that rely on enhanced mobile (e.g., high-capacity mobile) broadband (eMBB) access, and / or services with machine-type communication (MTC) access. The AMF3462 can provide control plane functionality for switching between the RAN3413 and other RANs (not shown) that utilize other radio technologies such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi.

[0313] SMF3483a and 3483b can be connected to AMF3482a and 3482b in CN3415 via the N11 interface. SMF3483a and 3483b can also be connected to UPF3484a and 3484b in CN3415 via the N4 interface. SMF3483a and 3483b can select and control UPF3484a and 3484b and configure traffic routing through them. SMF3483a and 3483b can perform other functions, such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types can be IP-based, non-IP-based, Ethernet-based, etc.

[0314] UPF3484a, 3484b may be connected via the N3 interface to one or more of gNB3480a, 3480b, 3480c in RAN3413, which can provide WTRU3402a, 3402b, 3402c with access to packet-switched networks such as the Internet 3410, facilitating communication between WTRU3402a, 3402b, 3402c and IP-enabled devices. UPF3484, 3484b can perform other functions such as routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0315] CN3415 can facilitate communication with other networks. For example, CN3415 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between CN3415 and PSTN408, or may communicate with such an IP gateway. In addition, CN3415 can provide access to other networks 3412 to WTRU3402a, 3402b, 3402cc, and other networks 3412 may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU3402a, 3402b, 3402c may be connected to the Local Data Network (DN) 3485a, 3485b through UPF3484a, 3484b via an N3 interface to UPF3484a, 3484b, and an N6 interface between UPF3484a, 3484b and DN3485a, 3485b.

[0316] In view of Figures 34A–34D and the corresponding descriptions thereof, one or more of the functions described herein with respect to one or more of the WTRU3402a–d, base stations 3414a–b, e-nodes B3460a–c, MME3462, SGW3464, PGW3466, gNB3480a–c, AMF3482a–b, UPF3484a–b, SMF3483a–b, DN3485a–b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.

[0317] Emulation devices may be designed to perform one or more tests on other devices in a lab environment and / or an operator network environment. For example, one or more emulation devices may perform one, more or all of their functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices in a communication network. One or more emulation devices may perform one, more or all of their functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. Emulation devices may be directly coupled to another device for testing and / or perform testing using wireless communication.

[0318] One or more emulation devices may perform one or more functions, including all of the above, without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a test scenario in a test laboratory and / or an undeployed (e.g., test) wired and / or wireless communication network to perform testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, for example, one or more antennas) may be used by the emulation device to transmit and / or receive data.

[0319] The HEVC standard saves approximately 50% bitrate compared to the conventional video encoding standard H.264 / MPEG AVC while maintaining comparable perceived quality. While the HEVC standard offers a significant improvement over its predecessor, further improvements in encoding efficiency can be achieved using additional encoding tools. The Joint Video Exploration Team (JVET), for example, initiated a project to develop a new generation video encoding standard called Versatile Video Coding (VVC) to provide such improvements in encoding efficiency, and established a reference software codebase called the VVC Test Model (VTM) to demonstrate a reference implementation of the VVC standard. Another reference software base called the benchmark set (BMS) was also created to facilitate the evaluation of new encoding tools. The BMS codebase includes a list of additional encoding tools that offer higher encoding efficiency and moderate implementation complexity, in addition to the VTM, and is used as a benchmark when evaluating similar encoding techniques in the VVC standardization process. BMS-2.0 integrates JEM coding tools, including trellis coding quantization tools such as 4x4 non-separable secondary transform (NSST), generalized bi-prediction (GBi), bi-directional optical flow (BIO), decoder-side motion vector refinement (DMVR), and current picture referencing (CPR).

[0320] A system and method for processing data according to a representative embodiment may be executed by one or more processors that execute a sequence of instructions contained in a memory device. Such instructions may be read into the memory device from another computer-readable medium, such as a secondary data storage device. The execution of the sequence of instructions contained in the memory device operates the processor, for example, as described above. In alternative embodiments, hardwired circuits may be used instead of or in combination with software instructions to implement one or more embodiments. Such software may run on a processor housed in a robotic assistance / device (RAA) and / or remotely in another mobile device. In the latter case, data may be transferred wired or wirelessly between the RAA or other mobile device including sensors and a remote device including a processor that runs software performing scale estimation and compensation as described above. According to another representative embodiment, some of the processing described above with respect to localization may be performed in a device including sensors / cameras, while the remainder of the processing may be performed in a second device after receiving partially processed data from the device including sensors / cameras.

[0321] While features and elements are described above in specific combinations, it will be understood by those skilled in the art that each feature or element can be used alone or in any combination with other features and elements. Furthermore, the methods described herein may be implemented in computer programs, software, or firmware embedded in computer-readable media for execution by a computer or processor. Examples of non-temporary computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multi-purpose disks (DVDs). Software and associated processors may be used to implement radio frequency transceivers for use in WTRU3402, UEs, terminals, base stations, RNCs, or any host computer.

[0322] Furthermore, in the embodiments described above, other devices including processing platforms, computing systems, controllers, and processors may be referred to. These devices may include at least one central processing unit ("CPU") and memory. In accordance with the practice of those skilled in the field of computer programming, references to symbolic representations of operations and arithmetic or instructions may be performed by various CPUs and memories. Such operations and arithmetic or instructions may be referred to as "executed," "executed by the computer," or "executed by the CPU."

[0323] Those skilled in the art will understand that operations or instructions expressed in terms of actions and symbols involve the manipulation of electrical signals by the CPU. The electrical system represents data bits that can result in the resulting transformation or reduction of electrical signals, and the preservation of data bits in memory locations within the memory system, thereby reconfiguring or otherwise modifying the CPU's operations and processing of other signals. The memory locations where data bits are preserved are physical locations having specific electrical, magnetic, optical, or organic properties that correspond to or represent the data bits. It should be understood that representative embodiments are not limited to the platforms or CPUs described above, and that other platforms and CPUs may support the methods provided.

[0324] Furthermore, data bits are maintained on computer-readable media, including magnetic disks, optical disks, and any other volatile (e.g., Random Access Memory ("RAM")) or non-volatile (e.g., Read-Only Memory ("ROM")) mass storage systems readable by a CPU. Computer-readable media may include cooperating or interconnected computer-readable media, which may reside exclusively on a processing system or be distributed among multiple interconnected processing systems, which may be local or remote to a processing system. It will be understood that representative embodiments are not limited to the memories described above, and that other platforms and memories may support the methods described.

[0325] In exemplary embodiments, any of the operations, processes, etc., described herein may be implemented as computer-readable instructions stored on a computer-readable medium. These computer-readable instructions may be executed by the processor, network elements, and / or any other computing devices of the mobile unit.

[0326] There is little difference between hardware and software implementations of the system's aspects. The use of hardware or software is generally (but not always) a design choice representing a cost-benefit trade-off, as the choice between hardware and software can be important in certain situations. There may be various means (e.g., hardware, software, and / or firmware) that can affect the processes and / or systems and / or other technologies described herein, but the preferred means may differ depending on the context in which the processes and / or systems and / or other technologies are deployed. For example, if the implementer determines that speed and accuracy are paramount, the implementer may choose primarily hardware and / or firmware means. If flexibility is paramount, the implementer may choose primarily software implementation. Alternatively, the implementer may choose any combination of hardware, software, and / or firmware.

[0327] The detailed description above illustrates various embodiments of devices and / or processes by using block diagrams, flowcharts, and / or examples. It will be understood by those skilled in the art that, insofar as such block diagrams, flowcharts, and / or examples include one or more functions and / or operations, each function and / or operation in such block diagrams, flowcharts, or examples can be implemented individually and / or collectively by various hardware, software, firmware, or substantially any combination thereof. Suitable processors include, by example, general-purpose processors, dedicated processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.

[0328] While features and elements are provided above in specific combinations, it will be understood by those skilled in the art that each feature or element can be used alone or in any combination with other features and elements. This disclosure is not limited to the specific embodiments described in this application, which are intended as examples of various aspects. As will be obvious to those skilled in the art, many modifications and variations can be made without departing from its spirit and scope. No element, operation, or command used in the description of this application should be construed as important or essential to the embodiments unless expressly indicated so. In addition to those enumerated herein, functionally equivalent methods and apparatus within the scope of this disclosure will be obvious to those skilled in the art from the above description. Such modifications and variations are intended to be included within the scope of the appended claims. This disclosure should be limited only by the terms of the appended claims and the entire scope of the equivalents to which such claims are entitled. This disclosure should be understood as not being limited to any particular method or system.

[0329] Furthermore, the terminology used herein should be understood to be solely for the purpose of describing specific embodiments and not intended to limit them. The terms “station” and its abbreviation “STA,” and the terms “user equipment” and its abbreviation “UE,” as used herein, may mean (i) a wireless transmit and / or receive unit (WTRU) as described below, (ii) any of the multiple embodiments of a WTRU as described below, (iii) a wireless-enabled and / or wired (e.g., tetherable) device configured to have some or all of the structure and functionality of a WTRU as described below, (iii) a wireless-enabled and / or wired device configured to have less than all of the structure and functionality of a WTRU as described below, or (iv) something similar. Details of exemplary WTRUs that may represent any UE described herein are provided below with reference to Figures 34A to 34D.

[0330] In certain representative embodiments, some parts of the subject matter described herein may be implemented using application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated forms. However, some aspects of the embodiments disclosed herein may be equivalently implemented in an integrated circuit, in whole or in part, as one or more computer programs running on one or more computers (e.g., as one or more programs running on one or more computer systems), as one or more programs running on one or more processors (e.g., as one or more programs running on one or more microprocessors), as firmware, or substantially any combination thereof, and it will be understood by those skilled in the art that the design of circuits and / or the writing of code for software and / or firmware will be well within the scope of the art of those skilled in the art in light of this disclosure. It will also be understood by those skilled in the art that mechanisms of the subject matter described herein may be distributed as various forms of program products, and that exemplary embodiments of the subject matter described herein apply regardless of the specific type of signal-carrying medium used to actually carry out the distribution. Examples of signal-carrying media include, but are not limited to, recordable media such as floppy disks, hard disk drives, CDs, DVDs, digital tapes, and computer memory, as well as transmission media such as digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.).

[0331] The subject matter described herein may illustrate different components that are contained within or connected to other different components. Such shown configurations are merely examples, and it will be understood that in practice many other architectures can be implemented to achieve the same functionality. In a conceptual sense, components in any arrangement to achieve the same functionality are substantially “associated” in such a way that the desired functionality can be achieved. Thus, any two components combined herein to achieve a particular functionality may be considered “associated” with each other, regardless of the architecture or intervening components, in such a way that the desired functionality can be achieved. Similarly, any two components thus associated may be considered “operably connected” or “operably coupled” with each other to achieve the desired functionality, and any two components that can be associated in such a way may be considered “operably coupled” with each other to achieve the desired functionality. Specific examples of operably coupled components include, but are not limited to, components that are physically engageable and / or physically interacting, as well as components that are / or wirelessly interactable and / or wirelessly interacting, as well as components that are / or logically interacting and / or logically interacting.

[0332] With regard to the use of substantially any plural and / or singular terms herein, those skilled in the art can appropriately convert from plural to singular and / or singular to plural depending on the context and / or use. For clarity, various singular / plural substitutions may be explicitly stated herein.

[0333] It will be understood by those skilled in the art that, in general, the terms used herein and especially in the appended claims (e.g., the body of the appended claims) are intended to be generally “open” terms (for example, the term “contains” should be interpreted as “contains but not limited,” the term “has” should be interpreted as “has at least,” and the term “includes” should be interpreted as “includes but not limited,” etc.). Furthermore, it will be understood by those skilled in the art that if a particular number is intended in the introduced claim description, such intention is explicitly stated in the claim, and if such statement is not made, such intention does not exist. For example, if only one element is intended, the term “single” or similar word may be used. For the sake of understanding, the following description of the appended claims and / or herein may include the use of the introductory phrases “at least one” and “one or more” to introduce the claim description. However, the use of such phrases should not be interpreted as implying that the introduction of a claim description by the indefinite article "a" or "an" limits any particular claim containing such introduced description to embodiments containing only one such description, even when the same claim contains the introductory phrase "one or more" or "at least one" and the indefinite article "a" or "an" (for example, "a" and / or "an" should be interpreted as meaning "at least one" or "one or more"). The same applies to the use of the definite article used to introduce a claim description. Furthermore, even if a specific number of introduced claim descriptions is explicitly stated, it will be recognized by those skilled in the art that such a statement should be interpreted as meaning at least that number (for example, the mere statement "two descriptions" without other modifiers means at least two descriptions or two or more descriptions).Furthermore, in cases where a convention similar to "at least one of A, B, and C" is used, such constructions are generally intended to mean that a person skilled in the art would understand the convention (for example, "a system having at least one of A, B, and C" includes, but is not limited to, systems having only A, only B, only C, both A and B, both A and C, both B and C, and / or systems having A, B, and C together). Furthermore, it will be understood by those skilled in the art that any substantially separate word and / or phrase presenting two or more alternative terms should be understood in the specification, claims, or drawings as construing the possibility of including one of those terms, either of those terms, or both of those terms. For example, the phrase “A or B” will be understood as including the possibility of “A” or “B,” or “A and B.” Furthermore, as used herein, the term “any of,” followed by an enumeration of multiple elements and / or multiple categories of elements, is intended to include “any,” “any combination,” “any number,” and / or “any combination of multiple” of the elements and / or categories of elements, separately or in conjunction with other elements and / or other categories of elements. Furthermore, as used herein, the term “set” or “group” is intended to include any number of elements, including zero. Furthermore, as used herein, the term “number” is intended to include any number, including zero.

[0334] Furthermore, if any feature or aspect of this disclosure is described in terms of the Markush group, it will be understood by those skilled in the art that this disclosure can also be described in terms of any individual element or subgroup of any element of the Markush group.

[0335] As will be understood by those skilled in the art, for all purposes, for example, in terms of providing written explanations, all scopes disclosed herein also encompass all possible sub-scopes and combinations of sub-scopes. Any enumerated scope can be readily recognized as being able to adequately describe and enable the same scope when broken down into at least equal 1 / 2, 1 / 3, 1 / 4, 1 / 5, 1 / 10, etc. As a non-restrictive example, each scope discussed herein can be readily broken down into a lower third, a middle third, an upper third, etc. Again, as will be understood by those skilled in the art, all words such as “at most,” “at least,” “greater than,” and “less than” refer to scopes that include the stated numbers and can then be broken down into sub-scopes as described above. Finally, as will be understood by those skilled in the art, a scope includes each individual element. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, and so on.

[0336] Furthermore, the scope of the claims should not be construed as being limited to the order or elements provided unless otherwise stated. Also, the use of the term “means” in any claim is intended to exercise the means-plus-function claim form under 35 United States Code § 112(6), and no claim without the term “means” is intended to do so.

[0337] Software-related processors may be used to implement radio frequency transceivers for use in wireless transmit / receive units (WTRUs), user equipment (UEs), terminals, base stations, mobility management entities (MMEs) or evolved packet cores (EPCs), or any host computer. WTRUs may be used in conjunction with modules implemented in hardware and / or software-defined radio (SDR), as well as other components such as cameras, video camera modules, videophones, speakerphones, vibration devices, speakers, microphones, television transceivers, hands-free headsets, keyboards, Bluetooth® modules, frequency modulation (FM) radio units, near-field communication (NFC) modules, liquid crystal display (LCD) display units, organic light-emitting diode (OLED) display units, digital music players, media players, video game player modules, internet browsers, and / or wireless local area network (WLAN) or ultra-wideband (UWB) modules.

[0338] Through this disclosure, those skilled in the art will understand that certain representative embodiments may be used selectively or in combination with other representative embodiments.

[0339] Furthermore, the methods described herein may be implemented in computer programs, software, or firmware embedded in computer-readable media for execution by a computer or processor. Examples of non-temporary computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multipurpose disks (DVDs). Software and associated processors may be used to implement radio frequency transceivers for use in WTRUs, UEs, terminals, base stations, RNCs, or any host computer. [Explanation of symbols]

[0340] 3400 Communication Systems 3402a WTRU 3402b WTRU 3402c WTRU 3402d WTRU 3408 PSTN 3410 Internet 3412 Network 3414a base station 3414b base station 3416 Air Interface 3418 Processor 3420 Transceiver 3422 Receiving element 3424 Microphone 3426 Keypad 3428 Touchpad 3430 Non-removable memory 3432 Removable Memory 3434 Power supply 3436 Chipset 3438 Peripherals 3460a eNode B 3460b eNode B 3460c eNode B 3462 MME 3464 SGW 3466 PGW 3482a AMF 3482b AMF 3483a SMF 3483b SMF 3484a UPF 3484b UPF 3485a DN 3485b DN

Claims

1. A method for decoding video, Determine a syntax element that indicates that optical flow-based predictive refinement is used for the video, With respect to the current block of the video, the block includes multiple subblocks, To generate the motion vector of the subblock of the current block, Using the motion vector of the subblock, a subblock-based motion prediction signal is generated. Determining a set of pixel-level motion vector difference values ​​associated with the subblock of the current block, Determining the spatial gradient of the subblock-based motion prediction signal at each sample position of the subblock, Based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient, a motion prediction refinement signal is determined for the current block. The subblock-based motion prediction signal and the motion prediction refinement signal are combined to generate a refined motion prediction signal for the current block, The prediction for the current block involves decoding the video using the refined motion prediction signal, wherein the motion vector of the subblock is generated, and the set of pixel-level motion vector difference values ​​is determined using an affine motion model relating to the current block. A method characterized by comprising:

2. A method for encoding video, Encoding a syntax element indicating that optical flow prediction refinement is used for the video, With respect to the current block of the video, the block includes multiple subblocks, Using the motion vector of the aforementioned subblock, a subblock-based motion prediction signal is generated, Determining a set of pixel-level motion vector difference values ​​associated with the subblock of the current block, Determining the spatial gradient of the subblock-based motion prediction signal at each sample position of the subblock, Based on the determined set of pixel-level motion vector difference values ​​and the determined spatial gradient, a motion prediction refinement signal is determined for the current block. The subblock-based motion prediction signal and the motion prediction refinement signal are combined to generate a refined motion prediction signal for the current block, The prediction for the current block involves encoding the video using the refined motion prediction signal, wherein the motion vector of the subblock is generated, and the set of pixel-level motion vector difference values ​​is determined using an affine motion model relating to the current block. A method characterized by comprising:

3. The method according to 1 or 2, characterized in that the syntax element is signaled in at least one of the sequence parameter set header, picture parameter set header, and tile group header.

4. The method according to 1 or 2, characterized in that a second syntax element is used to signal whether the prediction refinement by optical flow is used for bidirectional prediction or single prediction.

5. The method according to 1 or 2, characterized in that a third syntax element is signaled to indicate whether to apply the predictive refinement by optical flow to the color difference components of the video.

6. A decoder configured to decode video, Determine the syntax element that indicates that optical flow-based predictive refinement is used for the video. With respect to the current block of the video, the block includes multiple subblocks, Using the motion vector of the aforementioned subblock, generate the motion vector of the subblock. Using the motion vector of the subblock, a subblock-based motion prediction signal is generated. Determine the set of pixel-level motion vector difference values ​​associated with the subblock of the current block. The spatial gradient of the subblock base motion prediction signal at each sample position of the subblock is determined. Based on the set of pixel-level motion vector difference values ​​and the determined spatial gradient, a motion prediction refinement signal is determined for the current block. The subblock-based motion prediction signal and the motion prediction refinement signal are combined to generate a refined motion prediction signal for the current block. The video is decoded using the refined motion prediction signal as the prediction for the current block. Processors configured in this way Equipped with, The processor is configured to generate the motion vector of the subblock using an affine motion model for the current block and to determine the set of pixel-level motion vector difference values. A decoder characterized by the following features.

7. An encoder configured to encode video, Encode a syntax element indicating that optical flow prediction refinement is used for the video, With respect to the current block of the video, the block includes multiple subblocks, Generate the motion vector of the subblock, Using the motion vector of the subblock, a subblock-based motion prediction signal is generated. Determine the set of pixel-level motion vector difference values ​​associated with the subblock of the current block. The spatial gradient of the subblock base motion prediction signal at each sample position of the subblock is determined. Based on the set of pixel-level motion vector difference values ​​and the determined spatial gradient, a motion prediction refinement signal is determined for the current block. The subblock-based motion prediction signal and the motion prediction refinement signal are combined to generate a refined motion prediction signal for the current block. The video is encoded using the refined motion prediction signal as the prediction for the current block. Processors configured in this way Equipped with, The processor is configured to generate the motion vector of the subblock using an affine motion model for the current block and to determine the set of pixel-level motion vector difference values. An encoder characterized by the following features.

8. The decoder according to claim 6 or the encoder according to claim 7, characterized in that the syntax element is signaled by at least one of the sequence parameter set header, the picture parameter set header, and the tile group header.

9. The decoder according to claim 6 or the encoder according to claim 7, characterized in that a second syntax element is used to signal whether the prediction refinement by optical flow is used for bidirectional prediction or single prediction.

10. The decoder according to claim 6 or the encoder according to claim 7, characterized in that a third syntax element is signaled to indicate whether the predictive refinement by optical flow is applied to the color difference components of the video.