System, apparatus, and method for inter prediction refinement with optical flow
Through the optical flow model and affine motion compensation technology, the motion prediction in video coding is optimized, the problem of low efficiency of bidirectional motion compensation prediction is solved, and more efficient temporal redundancy removal and coding quality improvement are achieved.
Patent Information
- Application Number
- CN202510697945.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-15
- Filing Date
- 2020-02-04
- Publication Date
- 2025-09-19
AI Technical Summary
When existing video coding technologies process videos with rapid illumination changes, the efficiency of bidirectional motion compensation prediction is low, and it is difficult to effectively remove temporal redundancy.
The motion prediction process is optimized by adopting bidirectional prediction technology based on optical flow model, through sample-by-sample motion refinement and affine motion compensation, combined with interleaved prediction and sub-block-based temporal motion vector prediction.
It improves the video coding efficiency, especially under conditions of drastic illumination changes, reduces time redundancy, and improves coding quality and compression performance.
Smart Images

Figure CN120676166A_ABST
Abstract
Description
Cross-references
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 802,428, filed February 7, 2019, U.S. Provisional Patent Application No. 62 / 814,611, filed March 6, 2019, and U.S. Provisional Patent Application No. 62 / 883,999, filed April 15, 2019, the contents of each of which are incorporated herein by reference. Technical Field
[0002] The present application relates to video coding, and in particular, to systems, apparatus, and methods using inter-frame prediction refinement with optical flow. Related fields
[0003] Video coding systems are widely used to compress digital video signals to reduce the storage and / or transmission bandwidth of such signals. Among various types of video coding systems, such as block-based systems, wavelet-based systems, and object-based systems, block-based hybrid video coding systems are the most widely used and deployed today. Examples of block-based video coding systems include various international video coding standards, such as MPEG1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, and the latest video coding standard called High Efficiency Video Coding (HEVC), which was developed by the JCT-VC (Joint Collaboration on Video Coding) of ITU-T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG. Summary of the Invention
[0004] In a representative embodiment, a decoding method includes: obtaining a sub-block-based motion prediction signal for a current block of a video; obtaining one or more spatial gradients or one or more motion vector differences of the sub-block-based motion prediction signal; obtaining a refinement signal for the current block based on the one or more obtained spatial gradients or the one or more obtained motion vector differences; obtaining a refined motion prediction signal for the current block based on the sub-block-based motion prediction signal and the refinement signal; and decoding the current block based on the refined motion prediction signal. Various other embodiments are also disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0005] A more detailed understanding can be obtained from the following detailed description given by way of example in conjunction with the accompanying drawings. The drawings in the specification are examples. Therefore, the drawings and detailed description should not be considered limiting, and other equivalent examples are feasible and possible. In addition, the same reference numerals in the drawings indicate the same elements, and among them: Figure 1 is a block diagram illustrating a representative block-based video encoding system; Figure 2 is a block diagram illustrating a representative block-based video decoder; Figure 3 is a block diagram illustrating a representative block-based video encoder with generalized bi-prediction (GBi) support; Figure 4 is a schematic diagram showing a representative GBi module for an encoder; Figure 5 is a schematic diagram illustrating a representative block-based video decoder with GBi support; Figure 6 is a schematic diagram showing a representative GBi module for a decoder; Figure 7 is a schematic diagram showing a representative bidirectional optical flow; Figure 8A and 8B is a schematic diagram showing a representative four-parameter affine pattern; Figure 9 is a schematic diagram showing a representative six-parameter affine pattern; Figure 10 is a schematic diagram showing a representative interleaved prediction process; Figure 11 is a schematic diagram illustrating representative weight values (eg, associated with pixels) in a sub-block; Figure 12 is a schematic diagram showing some areas in which interleaved prediction is applied and other areas in which the interleaved prediction is not applied; Figure 13A and 13B is a schematic diagram showing the SbTMVP process; Figure 14 is a schematic diagram showing adjacent motion blocks (eg, 4×4 motion blocks) that can be used for motion parameter derivation; Figure 15 is a schematic diagram showing adjacent motion blocks that can be used for motion parameter derivation; Figure 16 is a schematic diagram showing the difference Δv(i, j) between the sub-block MV (arrow) and the pixel-level MV (arrow with circle) after sub-block-based affine motion compensation prediction; Figure 17A is a schematic diagram showing a representative process of determining the MV corresponding to the actual center of the sub-block; Figure 17B is a schematic diagram showing the positions of chroma samples in a 4:2:0 chroma format; Figure 17C is a schematic diagram showing an extended prediction sub-block; Figure 18A is a flowchart illustrating a first representative encoding / decoding method; Figure 18B is a flowchart illustrating a second representative encoding / decoding method; Figure 19 is a flowchart illustrating a third representative encoding / decoding method; Figure 20 is a flowchart illustrating a fourth representative encoding / decoding method; Figure 21 is a flowchart illustrating a fifth representative encoding / decoding method; Figure 22 is a flowchart illustrating a sixth representative encoding / decoding method; Figure 23 is a flowchart illustrating a seventh representative encoding / decoding method; Figure 24 is a flowchart illustrating an eighth representative encoding / decoding method; Figure 25 is a flowchart illustrating a representative gradient calculation method; Figure 26 is a flowchart illustrating a ninth representative encoding / decoding method; Figure 27 is a flowchart illustrating a tenth representative encoding / decoding method; Figure 28 is a flowchart illustrating an eleventh representative encoding / decoding method; Figure 29 is a flowchart illustrating a representative encoding method; Figure 30 is a flowchart illustrating another representative encoding method; Figure 31 is a flowchart illustrating a twelfth representative encoding / decoding method; Figure 32 is a flowchart illustrating a thirteenth representative encoding / decoding method; Figure 33 is a flowchart illustrating a fourteenth representative encoding / decoding method; Figure 34A is a system diagram illustrating an exemplary communication system in which one or more disclosed embodiments may be implemented; Figure 34B is a diagram showing that according to an embodiment, Figure 34A A system diagram of an exemplary wireless transmit / receive unit (WTRU) for use within the illustrated communication system; Figure 34C is a diagram showing that according to an embodiment, Figure 34Aa system diagram of an exemplary radio access network (RAN) and an exemplary core network (CN) used within the illustrated communication system; and Figure 34D is a diagram showing that according to an embodiment, Figure 34A A system diagram of another exemplary RAN and another exemplary CN used within the communication system is shown. DETAILED DESCRIPTION Block-based hybrid video coding process
[0006] Similar to HEVC, VVC is built on a block-based hybrid video coding framework.
[0007] Figure 1 is a block diagram illustrating a general block-based hybrid video coding system.
[0008] refer to Figure 1 , the encoder 100 may be provided with an input video signal 102 that is processed block by block (referred to as a coding unit (CU)) and may be used to efficiently compress high-resolution (1080p and above) video signals. In HEVC, a CU may be up to 64×64 pixels. A CU may be further partitioned into prediction units or PUs, to which a separate prediction process may be applied. For each input video block (MB and / or CU), spatial prediction 160 and / or temporal prediction 162 may be performed. Spatial prediction (or "intra-frame prediction") may use pixels from already coded neighboring blocks in the same video picture / slice to predict the current video block.
[0009] Spatial prediction can reduce spatial redundancy inherent in video signals. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses pixels from an already encoded video picture to predict the current video block. Temporal prediction can reduce temporal redundancy inherent in video signals. The temporal prediction signal for a given video block can (e.g., typically can) be signaled by one or more motion vectors (MVs), which can indicate the amount and / or direction of motion between the current block (CU) and its reference blocks.
[0010] If multiple reference pictures are supported (as is the case for recent video coding standards such as H.264 / AVC or HEVC), for each video block, its reference picture index can be sent (e.g., can be sent in addition); and / or this reference index can be used to identify which reference picture in the reference picture store 164 the temporal prediction signal comes from. After spatial and / or temporal prediction, a mode decision block 180 in the encoder 100 can select the best prediction mode, for example, based on a rate-distortion optimization method / process. The prediction block from the spatial prediction 160 or the temporal prediction 162 can be subtracted from the current video block 116; and / or the prediction residual can be decorrelated using a transform 104 and quantized 106 to achieve the target bit rate. The quantized residual coefficients can be inverse quantized 110 and inverse transformed 112 to form a reconstructed residual, which can be added back to the prediction block at 126 to form a reconstructed video block. Furthermore, in-loop filtering 166, such as a deblocking filter and an adaptive loop filter, may be applied to the reconstructed video block before it is placed in a reference picture store 164 and can be used to encode future video blocks. To form the output video bitstream 120, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients may be sent (e.g., all) to the entropy coding unit 108 for further compression and / or packing to form the bitstream.
[0011] The encoder 100 can be implemented using a processor, memory, and transmitter that provide the various elements / modules / units described above. For example, one skilled in the art will appreciate that: (1) the transmitter can send a bitstream 120 to a decoder; and (2) the processor can be configured to execute software to receive the input video 102 and perform the functions associated with the various blocks of the encoder 100.
[0012] Figure 2 is a block diagram illustrating a block-based video decoder.
[0013] refer to Figure 2 , the video decoder 200 may be provided with a video bitstream 202, which may be unpacked and entropy decoded at an entropy decoding unit 208. The coding mode and prediction information may be sent to the appropriate one of a spatial prediction unit 260 (for intra coding mode) and / or a temporal prediction unit 262 (for inter coding mode) to form a prediction block. The residual transform coefficients may be sent to an inverse quantization unit 210 and an inverse transform unit 212 to reconstruct the residual block. The reconstructed block may be further subjected to an in-loop filter 266 before being stored in a reference picture store 264. In addition to being stored in the reference picture store 264 for use in predicting future video blocks, the reconstructed video 220 may be sent, for example, to drive a display device.
[0014] The decoder 200 can be implemented using a processor, a memory, and a receiver, which can provide various elements / modules / units disclosed above. For example, those skilled in the art will appreciate that: (1) the receiver can be configured to receive a bitstream 202; and (2) the processor can be configured to execute software to enable receiving the bitstream 202 and outputting the reconstructed video 220 and performing functions associated with various blocks of the decoder 200.
[0015] Those skilled in the art will appreciate that many functions / operations / processes of a block-based encoder and a block-based decoder are the same.
[0016] In modern video codecs, bidirectional motion compensated prediction (MCP) can be used to efficiently remove temporal redundancy by exploiting temporal correlation between pictures. A bi-prediction signal can be formed by combining two uni-prediction signals using a weight value equal to 0.5, which may not be optimal for combining uni-prediction signals, especially under some conditions where the luminance changes rapidly from one reference picture to another. Certain prediction techniques / operations and / or processes can be implemented to compensate for luminance changes over time by applying some global / local weights and / or offset values to sample values in a reference picture (e.g., some or each of the sample values in a reference picture).
[0017] The use of bidirectional motion compensated prediction (MCP) in video codecs enables the removal of temporal redundancy by exploiting temporal correlations between pictures. A bi-prediction signal can be formed by combining two uni-prediction signals using a weight value (e.g., 0.5). In some videos, illumination characteristics can change rapidly from one reference picture to another. Therefore, prediction techniques can compensate for changes in illumination over time (e.g., fading transitions) by applying global or local weights and / or offset values to one or more sample values in a reference picture.
[0018] Generalized bi-prediction (GBi) can improve the MCP of bi-prediction mode. In bi-prediction mode, the prediction signal at a given sample x can be calculated by the following equation 1: P[x] = w0*P0[x+v0]+w1*P1[x+v1] (1)
[0019] In the above equation, P[x] can represent the resulting prediction signal of sample x at position x in the picture. Pi[x+vi] can be the motion compensated prediction signal of x using the motion vector (MV) vi of the i-th list (e.g., list 0, list 1, etc.). w0 and w1 can be two weight values shared between samples in a block (e.g., between all samples). Based on this equation, various prediction signals can be obtained by adjusting the weight values w0 and w1. Some configurations of w0 and w1 can mean the same prediction as uni-prediction and bi-prediction. For example, (w0, w1) = (0, 1) can be used for uni-prediction using reference list L0. (w0, w1) = (0, 1) can be used for uni-prediction using reference list L1. (w0, w1) = (0.5, 0.5) can be used for bi-prediction using two reference lists. The weights can be signaled for each CU. In order to reduce signaling overhead, a constraint such as w0 + w1 = 1 can be applied so that one weight can be signaled. Therefore, Equation 1 can be further simplified as set forth in Equation 2 below: P[x] = (1-w1)*P0[x+v0]+w1*P1[x+v1] (2)
[0020] To further reduce weight signaling overhead, w1 can be discretized (e.g., -2 / 8, 2 / 8, 3 / 8, 4 / 8, 5 / 8, 6 / 8, 10 / 8, etc.) Each weight value can then be indicated by an index value within a (e.g., small) limited range.
[0021] Figure 3 is a block diagram illustrating a representative block-based video encoder with GBi support.
[0022] The encoder 300 may include a mode decision module 304, a spatial prediction module 306, a motion prediction module 308, a transform module 310, a quantization module 312, an inverse quantization module 316, an inverse transform module 318, a loop filter 320, a reference picture store 322, and an entropy coding module 314. Some or all of the modules or components of the encoder (e.g., the spatial prediction module 306) may be combined with Figure 1 The spatial prediction module 306 and the motion prediction module 308 may be the same or similar to those described above. Alternatively, the spatial prediction module 306 and the motion prediction module 308 may be pixel-domain prediction modules. Thus, the input video bitstream 302 may be processed in a manner similar to the input video bitstream 102, although the motion prediction module 308 may also include GBi support. Thus, the motion prediction module 308 may combine two separate prediction signals in a weighted average manner. Furthermore, a selected weight index may be signaled in the output video bitstream 324.
[0023] The encoder 300 can be implemented using a processor, memory, and transmitter that provide the various elements / modules / units described above. For example, one skilled in the art will appreciate that: (1) the transmitter can send a bitstream 324 to a decoder; and (2) the processor can be configured to execute software to receive input video 302 and perform functions associated with the various blocks of the encoder 300.
[0024] Figure 4 is a schematic diagram illustrating a representative GBi estimation module 400 that may be employed in a motion prediction module of an encoder, such as the motion prediction module 308. The GBi estimation module 400 may include a weight value estimation module 402 and a motion estimation module 404. In this manner, the GBi estimation module 400 may utilize a process (e.g., a two-step operation / process) to generate an inter-frame prediction signal, such as a final inter-frame prediction signal. The motion estimation module 404 may use an input video block 401 and one or more reference pictures received from a reference picture repository 406 and perform motion estimation by searching for two best motion vectors (MVs) pointing to (e.g., two) reference blocks. The weight value estimation module 402 may receive: (1) the output of the motion estimation module 404 (e.g., motion vectors v0 and v1), one or more reference pictures from the reference picture repository 406, and weight information W, and may search for an optimal weight index to minimize a weighted bi-prediction error between the current video block and a bi-prediction. It is contemplated that the weight information W may describe a list of available weight values or a set of weights, such that the determined weight index and the weight information W may be used together to specify the weights w0 and w1 used in GBi. The prediction signal for the generalized bi-prediction may be calculated as a weighted average of the two prediction blocks. The output of the GBi estimation module 400 may include an inter-frame prediction signal, motion vectors v0 and v1, and / or a weight index weight_idx, etc.).
[0025] Figure 5 is a diagram illustrating a representative block-based video decoder with GBi support that can decode a GBi-enabled bitstream 502 (e.g., a bitstream from an encoder) such as that provided by a combination of Figure 3 The encoder 300 described above generates a bit stream 324. Figure 5 As shown, the video decoder 500 may include an entropy decoder 504, a spatial prediction module 506, a motion prediction module 508, a reference picture store 510, an inverse quantization module 512, an inverse transform module 514 and / or a loop filter module 518. Some or all of the modules of the decoder may be combined with Figure 2The motion prediction module 508 may be the same or similar to those described above, although it may also include GBi support. In this way, the coding mode and prediction information are used to derive a prediction signal by using spatial prediction or MCP supported by GBi. For GBi, the block motion information and weight values (e.g., in the form of an index indicating the weight values) may be received and decoded to generate the predicted block.
[0026] The decoder 500 can be implemented using a processor, a memory, and a receiver, which can provide various elements / modules / units disclosed above. For example, those skilled in the art will appreciate that: (1) the receiver can be configured to receive a bitstream 502; and (2) the processor can be configured to execute software to enable receiving the bitstream 502 and outputting a reconstructed video 520, as well as performing functions associated with the various blocks of the decoder 500.
[0027] Figure 6 is a diagram illustrating a representative GBi prediction module that may be employed in a motion prediction module (eg, motion prediction module 508 ) of a decoder.
[0028] Reference Figure 6 The GBi prediction module may include a weighted average module 602 and a motion compensation module 604, which may receive one or more reference pictures from a reference picture repository 606. The weighted average module 602 may receive the output of the motion compensation module 604, weight information W, and a weight index (e.g., weight_idx). The output of the motion compensation module 604 may include motion information of a block that may correspond to the picture. The GBi prediction module 600 may use the block motion information and weight values to calculate a prediction signal (e.g., an inter-frame prediction signal 608) for GBi as a weighted average of (e.g., two) motion-compensated prediction blocks. Representative dual prediction based on optical flow model
[0029] Figure 7 is a schematic diagram showing representative bidirectional optical flow.
[0030] refer to Figure 7 , the dual prediction can be based on an optical flow model. For example, the prediction associated with the current block (eg, current block 700) can be based on the first prediction block I (0) 702 (eg, a temporally preceding prediction block, eg, temporally shifted by τ0) and a second prediction block I (1)704 (e.g., a temporally future block, e.g., shifted in time by τ1). Bi-prediction in video coding may be a combination of two temporal prediction blocks 702 and 704 obtained from a reconstructed reference picture. Due to the limitations of block-based motion compensation (MC), there may be residual small motion that can be observed between the samples of the two prediction blocks, thus reducing the efficiency of the motion compensated prediction. Bidirectional optical flow (BIO, or BDOF) may be applied to reduce the impact of such motion for each sample within a block. BIO may provide sample-by-sample motion refinement, which may be performed on top of block-based motion compensated prediction when bi-prediction is used. For BIO, a refined motion vector for each sample in a block may be derived based on a classical optical flow model. For example, in I (k) (x, y) is the sample value at coordinate (x, y) of the prediction block derived from reference picture list k (k=0, 1) and and Where the horizontal gradient and vertical gradient of the sample are given, given the optical flow model, the motion refinement (v) at (x, y) can be derived by the following equation 3 x ,v y ):
[0031] exist Figure 7 , the (MV x0 ,MV y0 ) and (MV associated with the second prediction block 704 x1 ,MV y1 ) indicates that two prediction blocks can be generated. (0) and I (1) The block-level motion vector can be obtained by minimizing the sample value after motion refinement compensation (e.g., Figure 7 The difference Δ between A and B in (x, y) is used to calculate the motion refinement (v) at the sample position (x, y) x ,v y ), as shown in Equation 4 below:
[0032] For example, to ensure the regularity of the derived motion refinement, it can be assumed that the motion refinement is consistent for samples within a small unit (e.g., a 4×4 block or other small unit). In the benchmark set (BMS)-2.0, the value (v x ,v y ), as described in Equation 5 below:
[0033] To solve the optimization specified in Equation 5, the BIO may use an incremental approach / operation / process that optimizes motion refinement in both the horizontal and vertical directions (e.g., then the vertical direction). This may result in the following equations / inequalities 6 and 7: in can be a floor function that can output a maximum value less than or equal to the input, and th BIO can be a motion refinement threshold, e.g. to prevent error propagation due to coding noise and / or irregular local motion, which is equal to 2 18-BD The values of S1, S2, S3, S5 and S6 can be further calculated as set forth in equations 8-12 below: S1=∑ (i,j)∈Ω ψ x (i,j)·ψ x (i,j), (8) S3=∑ (i,j)∈Ω θ(i,j)·ψ x (i,j)·2 L (9) S2=∑ (i,j)∈Ω ψ x (i,j)·ψ y (i,j) (10) S5=∑ (i,j)∈Ω ψ y (i,j)·ψ y (i,j)·2 (11) S6=∑ (i,j)∈Ω θ(i,j)·ψ y (i,j)·2 L+1 (12) The various gradients can be described in the following equations 13-15: θ(i,j)=I (1) (i,j)-I (0) (i,j) (15)
[0034] For example, in BMS-2.0, the BIO gradients in the horizontal and vertical directions in Equations 13-15 can be directly obtained by calculating the difference (e.g., horizontal or vertical, depending on the direction of the gradient being derived) between two adjacent samples at a sample position of each L0 / L1 prediction block, as described in Equations 16 and 17 below:
[0035] In equations 8-12, L can be the bit depth increase of the internal BIO process / program to maintain data accuracy, for example, it can be set to 5 in BMS-2.0. To avoid division by small values, the adjustment parameters r and m in equations 6 and 7 can be defined as follows in equations 18 and 19: r=500·4 BD-8 (18) m=700·4 BD-8 (19) Where BD can be the bit depth of the input video. Based on the motion refinement derived by Equations 4 and 5, the final dual prediction signal of the current CU can be calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on the optical flow Equation 3, as specified in the following Equations 20 and 21: Among them, shift and o offset rnd(.) is a rounding function that rounds the input value to the nearest integer value. Representative affine patterns
[0036] In HEVC, a translational motion (and only translational motion) model is applied to motion-compensated prediction. In the real world, there are many kinds of motion (e.g., zooming in / out, rotation, perspective motion, and other irregular motions). In VVC Test Model (VTM)-2.0, affine motion-compensated prediction can be applied. The affine motion model is either 4-parameter or 6-parameter. A first flag is signaled for each inter-coded CU to indicate whether a translational motion model or an affine motion model is applied to inter prediction. If the affine motion model is applied, a second flag is sent to indicate whether the model is a 4-parameter model or a 6-parameter model.
[0037] The four-parameter affine motion model has the following parameters: two parameters for translation in the horizontal and vertical directions, one parameter for scaling in both directions, and one parameter for rotation in both directions. The horizontal scaling parameter is equal to the vertical scaling parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. The four-parameter affine motion model is encoded in the VTM using two motion vectors at two control point locations defined at the top left corner 810 and the top right corner 820 of the current CU. Other control point locations are also possible, such as at other corners and / or edges of the current CU.
[0038] Although one affine motion model is described above, other affine models are also possible and may be used in various embodiments herein.
[0039] Figure 8A and 8B is a schematic diagram showing a representative four-parameter affine model and sub-block-level motion derivation for affine blocks. Figure 8A and 8B The affine motion field of the block is described by two control point motion vectors at the first control point 810 (at the top left corner of the current block) and the second control point 820 (at the top right corner of the current block). Based on the control point motion, the motion field (v x ,v y ) is described as set forth in Equations 22 and 23 below: Where (v 0x ,v 0y ) can be the motion vector of the upper left corner control point 810, (v 1x ,v 1y ) can be the motion vector of the upper right corner control point 820, such as Figure 8A As shown, and w can be the width of the CU. For example, in VTM-2.0, the motion field of an affine-coded CU can be derived at the 4×4 block level; that is, (v x ,v y ) can be derived for each 4x4 block within the current CU and will be applied to the corresponding 4x4 block.
[0040] The four parameters of the 4-parameter affine model can be estimated iteratively. The MV pair in step k can be expressed as The original signal (eg, luminance signal) can be represented as I(i, j), and the predicted signal (eg, luminance signal) can be represented as I′ k (i,j). Spatial gradient g x (i,j) and g y (i, j) may be used, for example, to apply the prediction signal I' in the horizontal and / or vertical directions, respectively. k The derivative of Equation 3 can be expressed as set forth in Equations 24 and 25 below: Where (a, b) can be the delta translation parameters, and (c, d) can be the delta scale and rotation parameters at step k. The delta MV at a control point can be derived using its coordinates, as described in equations 26-29 below. For example, (0, 0) and (w, 0) can be the coordinates of the top left and top right control points 810 and 820, respectively.
[0041] Based on the optical flow equation, the relationship between the change in intensity (e.g., brightness) and the spatial gradient and temporal movement is equated in Equation 30 as follows:
[0042] By replacing Equations 24 and 25 and Equation 31 for the parameters (a, b, c, d) is obtained as follows: I′ k (i,j)-I(i,j)=(g x (i,j)*i+g y (i,j)*j)*c+(-g x (i,j)*j+g y (i,j)*i)*d +g x (i,j)*a+g y (i,j)*b(31)
[0043] Since the samples in the CU (e.g., all samples) satisfy Equation 31, the parameter set (e.g., a, b, c, d) can be solved using, for example, the least square error method. In step (k+1), the two control points Equations 26-29 can be used to solve the problem and can be rounded to a specific accuracy (e.g., 1 / 4 pixel accuracy (pel) or other sub-pixel accuracy, etc.). By using iterations, the MV at the two control points can be refined, for example, until convergence (e.g., when the parameters (a, b, c, d) are all zero or the iteration time meets a predetermined limit).
[0044] Figure 9 is a diagram showing a representative six-parameter affine pattern, where for example: V0, V1, and V2 are motion vectors at control points 910, 920, and 930, respectively, and (MV x ,MV y ) is the motion vector of the sub-block centered at position (x, y).
[0045] See Figure 9, an affine motion model (e.g., with 6 parameters) may have any of the following parameters: (1) parameters for translational movement in the horizontal direction; (2) parameters for translational movement in the vertical direction; (3) parameters for scaling movement in the horizontal direction; (4) parameters for rotational movement in the horizontal direction; (5) parameters for scaling movement in the vertical direction; and / or (6) parameters for rotational movement in the vertical direction. The 6-parameter affine motion model may be encoded using three MVs at three control points 910, 920, and 930. As Figure 9 As shown, the three control points 910, 920 and 930 of the 6-parameter affine-coded CU are defined at the upper left corner, upper right corner and lower left corner of the CU, respectively. The motion at the upper left control point 910 can be related to translation motion, and the motion at the upper right control point 920 can be related to rotation motion in the horizontal direction and / or scaling motion in the horizontal direction, and the motion at the lower left control point 930 can be related to rotation motion in the vertical direction and / or scaling motion in the vertical direction. For the 6-parameter affine motion model, the rotation motion and / or scaling motion in the horizontal direction may be different from the same motion in the vertical direction. The motion vector (v) of each sub-block x ,v y ) can be derived using the three MVs at control points 910, 920, and 930, as set forth in equations 32 and 33 below: Where (v 2x ,v 2y ) is the motion vector V2 of the lower left control point 930, (x, y) may be the center position of the sub-block, w may be the width of the CU, and h may be the height of the CU.
[0046] The six parameters of the 6-parameter affine model can be estimated in a similar manner. Equations 24 and 25 can be changed as described in equations 34 and 35 as follows. Wherein, in step k, (a, b) may be differential translation parameters, (c, d) may be differential scaling and rotation parameters in the horizontal direction, and (e, f) may be differential scaling and rotation parameters in the vertical direction. Equation 31 may be modified as described in Equation 36, as follows: I′ k (i,j)-I(i,j)=(g x (i,j)*i)*c+(g x (i,j)*j)*d+(g y (i,j)*i)*e+(g y (i,j)*j)*f +gx (i,j)*a+g y (i,j)*b(36)
[0047] The parameter set (a, b, c, d, e, f) can be solved by considering the samples (e.g., all samples) within the CU, for example, using a least squares method / process / operation. It can be calculated using equations 26-29. It can be calculated using equations 37 and 38 as described below. You can use the following Calculated using Equations 39 and 40.
[0048] Despite Figure 8A 、 8B 9 show a 4-parameter affine model and a 6-parameter affine model, but the skilled person will appreciate that affine models with different numbers of parameters and / or different control points are also possible.
[0049] Although the affine model is described herein in conjunction with optical flow refinement, a skilled person will appreciate that other motion models combined with optical flow refinement are also possible. Representative interlaced prediction for affine motion compensation
[0050] With affine motion compensation (AMC), such as in VTM, the coding block is divided into sub-blocks as small as 4×4, each of which can be assigned an individual motion vector (MV) derived from an affine model, such as Figure 8A and 8B or Figure 9 Using a 4-parameter or 6-parameter affine model, the MV can be derived from the MVs of two or three control points.
[0051] AMC may face a dilemma associated with the size of the sub-blocks. Using smaller sub-blocks, AMC may achieve better coding performance, but may suffer from a higher complexity burden.
[0052] Figure 10 is a schematic diagram illustrating a representative interleaved prediction process that can achieve finer-grained MVs, eg, in exchange for a moderate increase in complexity.
[0053] exist Figure 10In the VTM, the coding block 1010 can be divided into sub-blocks with two different division modes (e.g., first and second modes 0 and 1). The first division mode 0 (e.g., a first sub-block mode, such as a 4×4 sub-block mode) can be the same as in the VTM, and the second division mode 1 (e.g., an overlapping and / or interleaved second sub-block mode) can divide the coding block 1010 into 4×4 sub-blocks with a 2×2 offset from the first division mode 0, such as Figure 10 As shown. AMC can generate several auxiliary predictions (e.g., two auxiliary predictions P0 and P1) using the two partition modes (e.g., first and second partition modes 0 and 1). The MV of each subblock in each of partition modes 0 and 1 can be derived from the control point motion vector (CPMV) through the affine model.
[0054] The final prediction P can be calculated as the weighted sum of the auxiliary predictions (e.g., two auxiliary predictions P0 and P1), which is equated as shown in Equations 41 and 42 as follows:
[0055] Figure 11 is a schematic diagram showing representative weight values (eg, associated with pixels) in a sub-block. Figure 11 , auxiliary prediction samples located at the center (eg, center pixel) of the sub-block 1100 may be associated with a weight value of 3, and auxiliary prediction samples located at the boundary of the sub-block 1100 may be associated with a weight value of 1.
[0056] Figure 12 is a schematic diagram showing a region in which interleaved prediction is applied and other regions in which interleaved prediction is not applied. Figure 12 , the region 1200 may include a first region 1210 (e.g., having 4×4 sub-blocks) in which interleaved prediction is applied. Figure 12 ) and a second region 1220 (shown as non-cross-hatched in FIG) where, for example, interlaced prediction is not applied Figure 12 To avoid small block motion compensation, for example, for both the first and second partitioning modes, the interlaced prediction may be applied only to regions where the size of the sub-blocks meets a threshold size (eg, 4×4).
[0057] In VTM-3.0, the size of the sub-block can be 4×4 for the chroma component, and interleaved prediction can be applied to the chroma component and / or the luma component. Since the area used for motion compensation (MC) of the sub-block (e.g., all sub-blocks) can be extracted as a whole in AMC, interleaved prediction does not increase bandwidth. For flexibility, a flag can be signaled in the slice header to indicate whether interleaved prediction is used or not. For interleaved prediction, the flag can be signaled as a 1-bit flag (e.g., a first logic level that can always be signaled as 0 or 1). Representative process for sub-block based temporal motion vector prediction (SbTMVP)
[0058] SbTMVP is supported by VTM. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP can use the motion field in the collocated picture, for example, to improve the merge mode and motion vector prediction of the CU in the current picture. The same collocated picture used by TMVP can be used for SbTMVP. The differences between SbTMVP and TMVP are as follows: (1) TMVP can predict the motion at the CU level, while SbTMVP can predict the motion at the sub-CU level; and / or (2) TMVP can extract the temporal motion vector from the collocated block in the collocated picture (for example, the collocated block can be the bottom right or center block relative to the current CU), while SbTMVP can apply motion shifting (for example, the motion shifting can be obtained from the motion vector of one of the spatially neighboring blocks of the current CU) before extracting the temporal motion information from the collocated picture, etc.
[0059] Figure 13A and 13B is a schematic diagram illustrating the SbTMVP process. Figure 13A shows the spatial neighboring blocks used by ATMVP, Figure 13B It is shown that the sub-CU motion field is derived by applying the motion shift from the spatial neighbors and scaling the motion information from the corresponding co-located sub-CU.
[0060] See Figure 13A and 13B, SbTMVP may predict the motion vector of the sub-CU within the current CU operation (e.g., in two operations). In the first operation, the spatial neighboring blocks A1, B1, B0, and A0 may be checked in the order of A1, B1, B0, and A0. Once the first spatial neighboring block having a motion vector that uses the co-located picture as its reference picture is identified and / or after the first spatial neighboring block having a motion vector that uses the co-located picture as its reference picture is identified, the motion vector may be selected as the motion shift to be applied. If no such motion is identified from the spatial neighboring blocks, the motion shift may be set to (0, 0). In the second operation, the motion shift identified in the first operation may be applied (e.g., added to the coordinates of the current block) to obtain sub-CU level motion information (e.g., motion vector and reference index) from the co-located picture, such as Figure 13B As shown in . Figure 13B The example in shows motion shifting set to the motion of block A1. For each sub-CU, the motion information of its corresponding block in the co-located picture (e.g., the minimum motion grid covering the center sample) can be used to derive the motion information for the sub-CU. After the motion information of the co-located sub-CU is identified, the motion information can be converted into a motion vector and reference index for the current sub-CU in a manner similar to the TMVP process of HEVC. For example, temporal motion scaling can be applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.
[0061] A combined sub-block based merge list may be used in VTM-3 and may contain or include both SbTMVP merge candidates and affine merge candidates, for example, for signaling a sub-block based merge mode. The SbTMVP mode may be enabled / disabled by a sequence parameter set (SPS) flag. If SbTMVP mode is enabled, the SbTMVP predictor may be added as the first entry in the list of sub-block based merge candidates, followed by the affine merge candidates. The size of the sub-block based merge list may be signaled in the SPS, and the maximum allowed size of the sub-block based merge list may be an integer, for example 5 in VTM3.
[0062] The sub-CU size used in SbTMVP may be fixed, eg, 8×8 or other sub-CU size, and performed for affine merge mode, which may be applicable to (eg, applicable only to) CUs whose width and height may both be greater than or equal to 8. The encoding logic for the additional SbTMVP merge candidates may be the same as that for other merge candidates. For example, for each CU in a P or B slice, an additional rate-distortion (RD) check may be performed to decide whether to use the SbTMVP candidate. Representative regression-based motion vector fields
[0063] To provide fine granularity of intra-block motion vectors, a regression-based motion vector field (RMVF) tool (e.g., in JVET-M0302) can be implemented, which can attempt to model the motion vector of each block at the sub-block level based on spatially neighboring motion vectors.
[0064] Figure 14 is a schematic diagram showing adjacent motion blocks (eg, 4×4 motion blocks) that can be used for motion parameter derivation. Figure 14 Neighboring motion vectors for RMVF motion parameter derivation are shown. In the regression process, a row of adjacent motion vectors 1410 and a column of adjacent motion vectors 1420 based on 4×4 sub-blocks (and their center positions) from each side of the block can be used. For example, the adjacent motion vectors can be used for RMVF motion parameter derivation.
[0065] Figure 15 is a diagram showing a method for deriving motion parameters to reduce neighboring motion information (e.g., to reduce the relative Figure 14 Schematic diagram of adjacent motion blocks (the number of adjacent motion blocks used in the regression process). Figure 15 A reduced number of neighboring motion vector candidates for RMVF motion parameter derivation is shown. A reduced amount of neighboring motion information for RMVF parameter derivation of neighboring 4×4 motion blocks is available for motion parameter derivation (e.g., approximately half, such as approximately every other neighboring motion block, is available for motion parameter derivation). Certain neighboring motion blocks of the rows 1410 and columns 1420 may be selected, determined, or predetermined for reduction of neighboring motion information.
[0066] While approximately half of the neighboring motion blocks of the rows 1410 and columns 1420 are shown as being selected, other percentages (with other motion block positions) may be selected, for example, to reduce the number of neighboring motion blocks to use in the regression process.
[0067] When collecting motion information for motion parameter derivation, five regions (e.g., bottom left, top left, top right) as shown in the figure may be used. The top right reference motion region and the bottom left reference motion region may be limited to half (e.g., only half) of the corresponding width or height of the current block.
[0068] In RMVF mode, the motion of the block can be defined by a 6-parameter motion model. These parameters are xx ,a xy ,a yx ,a yy ,b x and b yIt can be calculated by solving a linear regression model in the sense of mean square error (MSE). The input of the regression model can be the center position (x, y) and / or motion vector (mv) of the available adjacent 4×4 sub-blocks defined above. x and mv y ), or may comprise the center position (x, y) and / or motion vector (mv) of the available adjacent 4×4 sub-blocks as defined above. x and mv y )).
[0069] The center is (X subPU ,Y subPU ) of the 8×8 sub-block (MV X_subPU ,MV Y_subPU ) can be calculated as described in Equation 43 below:
[0070] The motion vectors may be calculated for 8x8 sub-blocks relative to the center position of the sub-block (e.g., each sub-block). For example, in RMVF mode, motion compensation may be applied with 8x8 sub-block accuracy. To efficiently model the motion vector field, the RMVF tool is applied only when at least one motion vector from at least three candidate regions is available.
[0071] Affine motion model parameters can be used to derive motion vectors for certain pixels (e.g., every pixel) in a CU. Although the complexity of generating pixel-based affine motion compensation predictions may be high (e.g., very high), and also because the memory access bandwidth requirements of such sample-based MCs may be high, a sub-block-based affine motion compensation process / method (e.g., via VVC) may be implemented. For example, a CU may be divided into several sub-blocks (e.g., 4×4 sub-blocks, square sub-blocks, and / or non-square sub-blocks). Each sub-block may be assigned an MV that can be derived from the affine model parameters. This MV may be the MV at the center of the sub-block (or another location within the sub-block). The pixels in the sub-block (e.g., all pixels in the sub-block) may share the sub-block MV. Sub-block-based affine motion compensation may be a trade-off between coding efficiency and complexity. To achieve finer granularity in motion compensation, interleaved prediction for affine motion compensation may be implemented, and this interleaved prediction for affine motion compensation may be generated by taking a weighted average of two sub-block motion compensation predictions. Interleaved prediction may require and / or use two or more motion compensated predictions per sub-block and thus may increase memory bandwidth and complexity.
[0072] In certain representative embodiments, methods, apparatus, processes, and / or operations may be implemented to refine sub-block based affine motion compensation predictions using optical flow (e.g., using and / or based on optical flow). For example, after performing sub-block based affine motion compensation, pixel intensities may be refined by adding differences derived from the optical flow equation, which is referred to as prediction refinement using optical flow (PROF). PROF may achieve pixel-level granularity without significantly increasing complexity and may maintain a worst-case memory access bandwidth comparable to sub-block based affine motion compensation. PROF may be applied to any scenario where a pixel-level motion vector field is available (e.g., computable) in addition to a prediction signal (e.g., an unrefined motion prediction signal and / or a sub-block based motion prediction signal). In addition to the described affine mode, the prediction PROF process may also be applied to other sub-block prediction modes. Application of PROF to sub-block modes, such as SbTMVP and / or RMVF, may be implemented. Application of PROF to dual prediction is described herein. Representative PROF procedure for affine mode
[0073] In certain representative embodiments, methods, apparatus, and / or processes may be implemented to improve the granularity of sub-block based affine motion compensated prediction, e.g., by applying changes in pixel intensities derived from optical flow (e.g., an optical flow equation), and may use and / or require one motion compensation operation per sub-block (e.g., only one motion compensation operation per sub-block), which is the same as existing affine motion compensation, e.g., in VVC.
[0074] Figure 16 is a schematic diagram illustrating a sub-block MV and a pixel-level motion vector difference Δv(i,j) (eg, which is sometimes also referred to as a refined MV of a pixel) after sub-block-based affine motion compensated prediction.
[0075] See Figure 16 , the CU 1600 may include sub-blocks 1610, 1620, 1630, and 1640. Each sub-block 1610, 1620, 1630, and 1640 may include a plurality of pixels (e.g., 16 pixels in the sub-block 1610). A sub-block MV 1650 (e.g., as a coarse or average sub-block MV) associated with each pixel 1660(i, j) of the sub-block 1610 is shown. For each corresponding pixel (i, j) in the sub-block 1610, a refined MV 1670(i, j) may be determined (which may indicate the difference between the actual MV of the pixel 1660(i, j) and the sub-block MV 1650 (where (i, j) defines the pixel position in the sub-block 1610). In order Figure 16For clarity, only the refined MV 1670(1, 1) is labeled, although pixel-level motion of other individuals is shown. In some representative embodiments, the refined MV 1670(i, j) can be determined as the pixel-level motion vector difference Δv(i, j) (sometimes referred to as the motion vector difference).
[0076] In certain representative embodiments, methods, apparatus, processes, and / or operations may be implemented that include any of the following operations: (1) In a first operation: sub-block based AMC may be performed as disclosed herein to generate a sub-block based motion prediction I(i, j); (2) In the second operation: the spatial gradient g of the sub-block-based motion prediction I(i,j) at each sample position can be calculated x (i,j) and g y (i, j) (In one example, the spatial gradient can be generated using the same process as the gradient generation used in BDOF. For example, the horizontal gradient at a sample position can be calculated as the difference between its right neighbor and its left neighbor, and / or the vertical gradient at a sample position can be calculated as the difference between its bottom neighbor and its top neighbor. In another example, a Sobel filter can be used to generate the spatial gradient; (3) In a third operation: the illumination intensity change of each pixel in the CU may be calculated using and / or by an optical flow equation, for example, as set forth in Equation 44 below: ΔI(i,j)= g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (44) where the value of the motion vector difference Δv(i,j) is the difference 1670 between the pixel-level MV calculated for the sample position (i,j) (denoted by v(i,j)) and the sub-block-level MV 1650 of the sub-block covering the pixel 1660 (i,j), as Figure 16 The pixel-level MV v(i, j) can be derived from the control point MV by equations 22 and 23 for a 4-parameter affine model or by equations 32 and 33 for a 6-parameter affine model.
[0077] In certain representative embodiments, the motion vector difference Δv(i, j) may be derived from the affine model parameters by or using equations 24 and 25, where x and y may be offsets from the pixel position to the center of the sub-block. Since the affine model parameters and the pixel offsets do not change between sub-blocks, the motion vector difference Δv(i, j) may be calculated for the first sub-block and reused in other sub-blocks in the same CU. For example, the difference between the pixel-level MV and the sub-block-level MV may be calculated using equations 45 and 46 below, because the translation affine parameters (a, b) may be the same for the pixel-level MV and the sub-block MV. (c, d, e, f) may be four additional affine parameters (e.g., four affine parameters in addition to the translation affine parameters) where (i, j) can be the pixel position relative to the top left position of the sub-block, and (x sb ,y sb ) can be the center position of the sub-block relative to the upper left position of the sub-block.
[0078] Figure 17A FIG. 1 is a diagram illustrating a representative procedure for determining a motion vector corresponding to the actual center of a sub-block. Figure 17A The MV derivation for sub-blocks of one 8x4 CU is shown.
[0079] Please refer to Figure 17A , the two sub-blocks SB0 and SB1 shown in the figure are 4×4 sub-blocks. If the sub-block width is SW and the sub-block height is SH, the sub-block center position can be set to ((SW-1) / 2, (SH-1) / 2). In other examples, the sub-block center position can be estimated based on the position as described by (SW / 2, SH / 2). The actual center point of the first sub-block SB0 is P0', and the actual center point of the second sub-block SB1 is P1', which uses ((SW-1) / 2, (SH-1) / 2). By using, for example, (SW / 2, SH / 2) (for example, in VVC), the estimated center point of the first sub-block SB0 is P0, and the estimated center point of the second sub-block SB1 is P1. In some representative embodiments, the MV of the sub-block can be more accurately based on the actual center position rather than the estimated center position (which is used in VVC).
[0080] Figure 17B is a diagram showing the positions of chroma samples in a 4:2:0 chroma format. Figure 17B, the chroma sub-block MV can be derived from the MV of the luma sub-block. For example, in a 4:2:0 chroma format, a 4×4 chroma sub-block can correspond to an 8×8 luma region. Although representative embodiments are shown in conjunction with the 4:2:0 chroma format, skilled artisans understand that other chroma formats, such as the 4:2:2 chroma format, can be used equally.
[0081] The chroma sub-block MV can be derived by averaging the upper left 4×4 luma sub-block MV and the lower right luma sub-block MV. For chroma sample position types 0, 2, and / or 3, the derived chroma sub-block MV may or may not be located at the center of the chroma sub-block. For chroma sample position types 0, 2, and 3, the chroma sub-block center position (x sb ,y sb ) may or may need to be adjusted by an offset. For example, for 4:2:0 chroma sample position types 0, 2, and 3, the adjustments may be applied as set forth in equations 47-49 below:
[0082] The sub-block based motion prediction I(i,j) may be refined by adding intensity variations (e.g., luminance intensity variations, e.g., as provided in Equation 44). The final (ie, refined) prediction I'(i,j) may be generated by or using Equation 50 as follows. I′(i,j)=I(i,j)+ΔI(i,j) (50)
[0083] When the refinement is applied, the sub-block based affine motion compensation can achieve pixel-level granularity without increasing the worst-case bandwidth and / or memory bandwidth.
[0084] To maintain the accuracy of prediction and / or gradient calculation, the bit depth in the operation-related performance of sub-block based AMC may be an intermediate bit depth, which may be higher than the encoding bit depth.
[0085] The above process can be used to refine chroma intensities (e.g., in addition to or instead of refining luma intensities). In one example, the intensity difference used in Equation 50 can be multiplied by a weighting factor w before being added to the prediction, as shown in Equation 51 below: I′(i,j)=I(i,j)+w·ΔI(i,j) (51) Where w can be set to a value between 0 and 1, inclusive. w can be signaled at the CU level or the picture level. For example, w can be signaled via a weight index. For example, Index Table 1 can be used to signal w. index 0 1 2 3 4 Weight 1 / 2 3 / 4 1 / 4 1 0 Index Table 1
[0086] The encoder algorithm can choose the value of w that results in the lowest rate-distortion cost.
[0087] For example, the gradient of the predicted sample can be calculated in different ways, such as g x and / or g y In certain representative embodiments, the prediction sample g x and g y It can be calculated by applying a 2D Sobel filter. An example of a 3×3 Sobel filter for horizontal and vertical gradients is as follows: Horizontal Sobel filter: Vertical Sobel filter:
[0088] In other representative embodiments, the gradient may be calculated using a one-dimensional 3-tap filter. One example may include [-1 0 1], which is a simpler (eg, much simpler) filter than a Sobel filter.
[0089] Figure 17C is a schematic diagram illustrating extended sub-block prediction. Shaded circle 1710 is a padded sample around a 4×4 sub-block (e.g., non-shaded circle 1720). Using a Sobel filter as an example, the samples in box 1730 can be used to calculate the gradient of sample 1740 at the center. Although a Sobel filter can be used to calculate the gradient, other filters such as a 3-tap filter are also possible.
[0090] For the above example gradient filters, such as 3x3 Sobel filter and one-dimensional filter, extended sub-block prediction may be used and / or required for sub-block gradient calculation. For example, a row at the top and bottom boundaries of the sub-block and a column at the left and right boundaries may be filled to calculate the gradient of those samples at the sub-block boundaries.
[0091] Different methods / processes and / or operations may be used to obtain the extended sub-block prediction. In one representative embodiment, given a sub-block size of N×M, an (N+2)×(M+2) extended sub-block prediction may be obtained by performing (N+2)×(M+2) block motion compensation using the sub-block MV. Utilizing this embodiment, memory bandwidth may be increased. To avoid an increase in memory bandwidth, in certain representative embodiments, given a K-tap interpolation filter in both the horizontal and vertical directions, (N+K-1)×(M+K-1) integer reference samples prior to interpolation may be extracted for interpolation of the N×M sub-block, and boundary samples of the (N+K-1)×(M+K-1) block may be copied from neighboring samples of the (N+K-1)×(M+K-1) sub-block, such that an extended region may be (N+K-1+2)×(M+K-1+2). The extended region may be used for interpolation of the (N+2)×(M+2) sub-block. If the sub-block MV points to a fractional location, these representative embodiments may still use and / or require additional interpolation operations to generate the (N+2)×(M+2) prediction.
[0092] For example, to reduce computational complexity, in other representative embodiments, the sub-block prediction may be obtained by using N×M block motion compensation of the sub-block MV. The boundaries of the (N+2)×(M+2) prediction may be obtained without interpolation by any of the following: (1) integer motion compensation, where MV is the integer part of the sub-block MV; (2) integer motion compensation, where MV is the nearest integer MV of the sub-block MV; and / or (3) copying from the nearest neighboring sample in the N×M sub-block prediction.
[0093] Pixel-level refinement MV (e.g., Δv x and Δv y ) may affect the accuracy of PROF. In certain representative embodiments, a combination of a multi-bit fractional component and another multi-bit integer component may be implemented. For example, a 5-bit fractional component and an 11-bit integer component may be used. This combination of a 5-bit fractional component and an 11-bit integer component can use a total of 16 bits to represent an MV range from -1024 to 1023 with 1 / 32 pixel precision.
[0094] The gradient (e.g., g x and g y ) and the accuracy of the intensity variation ΔI may affect the performance of PROF. In certain representative embodiments, the prediction sample accuracy may be maintained or maintained at a predetermined number or a signaled number of bits (e.g., the internal sample accuracy defined in the current VVC draft, which is 14 bits). In certain representative embodiments, the gradient and / or intensity variation ΔI may be maintained at the same accuracy as the prediction sample.
[0095] The range of the intensity variation ΔI may affect the performance of PROF. The intensity variation ΔI may be clipped to a smaller range to avoid false values generated by an inaccurate affine model. In one example, the intensity variation ΔI may be clipped to predition_bitdepth-2.
[0096] Δv x and Δv y The combination of the number of bits for the fractional component of Δv, the number of bits for the fractional component of the gradient, and the number of bits for the intensity change ΔI may together affect the complexity of certain hardware or software implementations. In a representative embodiment, 5 bits may be used to represent Δv x and Δv y , 2 bits can be used to represent the fractional component of the gradient, and 12 bits can be used to represent ΔI, however, they can be any number of bits.
[0097] To reduce computational complexity, PROF can be skipped in some cases. For example, if the magnitude of all pixel-based deltas (e.g., refinements) MV(Δv(i,j)) within a 4×4 sub-block is less than a threshold, PROF can be skipped for the entire affine CU. If the gradients of all samples within a 4×4 sub-block are less than a threshold, PROF can be skipped. The PROF can be applied to chrominance components, such as Cb and / or Cr components. The difference MV of the Cb and / or Cr components of a sub-block can reuse the difference MV of the sub-block (for example, the difference MV calculated for different sub-blocks in the same CU can be reused).
[0098] Although the gradient process disclosed herein (e.g., using replicated reference samples to extend sub-blocks for gradient calculation) is shown as being used with a PROF operation, the gradient process may be used with other operations such as a BDOF operation and / or an affine motion estimation operation. Representative PROF procedures for other sub-block modes
[0099] PROF can be applied to any scenario where a pixel-level motion vector field is available (e.g., can be calculated) in addition to a prediction signal (e.g., an unrefined prediction signal). For example, in addition to the affine mode, prediction refinement using optical flow can be used in other sub-block prediction modes, such as SbTMVP mode (e.g., ATMVP mode in VVC), or regression-based motion vector field (RMVF).
[0100] In certain representative embodiments, a method may be implemented to apply PROF to SbTMVP. For example, such a method may include any of the following: (1) In the first operation, sub-block level MV and sub-block prediction can be generated based on the existing SbTMVP process described herein; (2) in a second operation, the affine model parameters may be estimated by using the block MV field using a linear regression method / process; (3) In a third operation, a pixel-level MV may be derived from the affine model parameters obtained in the second operation, and an associated pixel-level motion refinement vector (Δv(i,j)) relative to the sub-block MV may be calculated; and / or (4) In a fourth operation, a prediction refinement process using optical flow may be applied to generate a final prediction, etc.
[0101] In certain representative embodiments, a method may be implemented to apply PROF to RMVF. For example, such a method may include any of the following: (1) In the first operation, the sub-block level MV field, sub-block prediction and / or affine model parameters a xx ,a xy ,a yx ,a yy ,b x and b x can be generated based on the RMVF process described herein; (2) In the second operation, the pixel-level MV offset (Δv(i,j)) from the sub-block-level MV can be obtained by the affine model parameter a xx ,a xy ,a yx ,a yy ,b x and b x It is derived as follows by equation 52: Where (i, j) is the pixel offset from the sub-block center. Since the affine parameters and / or the pixel offset from the sub-block center do not change between sub-blocks, the pixel MV offset may be calculated for the first sub-block (e.g., only the pixel MV offset for the first sub-block needs to be calculated) and may be reused for other sub-blocks in the CU; and / or (3) In a third operation, the PROF process may be applied to generate a final prediction, for example by applying Equations 44 and 50. Representative PROF methods for dual prediction
[0102] In addition to or in lieu of using PROF in the uni-prediction described herein, the PROF technique can be used for bi-prediction. When used in bi-prediction, the PROF can be used to generate an L0 prediction and / or an L1 prediction, for example, before they are combined by weights. To reduce computational complexity, PROF can be applied (e.g., can be applied only) to one prediction, such as L0 or L1. In certain representative embodiments, PROF can be applied (e.g., can be applied only) to a list (e.g., a list associated with reference pictures that are close (e.g., within a threshold) and / or closest to the current picture). Representative process for PROF enablement
[0103] PROF enablement may be signaled at or in a sequence parameter set (SPS) header, a picture parameter set (PPS) header, and / or a tile group header. In some embodiments, a flag may be signaled to indicate whether PROF is enabled for affine mode. If the flag is set to a first logic level (e.g., "true"), PROF may be used for both uni-prediction and bi-prediction. In some embodiments, if the first flag is set to "true," a second flag may be used to indicate whether PROF is enabled or not for bi-prediction affine mode. If the first flag is set to a second logic level (e.g., "false"), it may be inferred that the second flag is set to "false." If the first flag is set to "true," a flag may be used at or in the SPS header, PPS header, and / or tile group header to signal whether PROF applies to chroma components, allowing PROF control of luma and chroma components to be separated. Representative methods for conditionally enabled PROF
[0104] For example, to reduce complexity, PROF may be applied when (e.g., only when) certain conditions are met. For example, for small CU sizes (e.g., below a threshold level), affine motion may be relatively small, thus limiting the benefits of applying PROF. In certain representative embodiments, PROF may be disabled in affine motion compensation when or in situations where the CU size is small (e.g., for CU sizes no larger than 16x16, such as 8x8, 8x16, or 16x8) to reduce complexity at both the encoder and / or decoder. In certain representative embodiments, PROF may be skipped in affine motion estimation (e.g., only during affine motion estimation) when the CU size is small (below the same or a different threshold level), for example, to reduce encoder complexity, and PROF may be performed at the decoder regardless of CU size. For example, on the encoder side, after motion estimation to search for affine model parameters (e.g., control point MVs), a motion compensation (MC) process may be invoked and PROF may be performed. The MC process may also be invoked for each iteration during motion estimation. In MC in motion estimation, PROF can be skipped to save complexity, and there will be no prediction mismatch between the encoder and decoder because the final MC in the encoder will run PROF. That is, when the encoder searches for affine model parameters (e.g., affine MV) for prediction of a CU, PROF refinement may not be applied, and once or after the encoder completes the search, the encoder may apply PROF to refine the prediction of the CU using the affine model parameters determined from the search.
[0105] In some representative embodiments, the difference between the CPMVs can be used as a criterion for determining whether to enable PROF. When the difference between the CPMVs is small (e.g., below a threshold level) such that the affine motion is small, the benefit of applying PROF may be limited, and PROF may be disabled for affine motion compensation and / or affine motion estimation. For example, for a 4-parameter affine mode, PROF may be disabled if the following conditions (e.g., all of the following conditions) are met:
[0106] For the 6-parameter affine mode, in addition to or instead of the above conditions, PROF may be disabled if the following conditions are met (e.g., all of the following conditions are also met): Wherein T is a predefined threshold, for example, 4. The PROF skipping process based on CPMV or affine parameters can be applied at the encoder (for example, also only applied at the encoder), and the decoder may or may not skip PROF. Representative process of PROF combined with or replacing the deblocking filter
[0107] Because PROF can be a pixel-by-pixel refinement that can compensate for block-based MC, motion differences across block boundaries can be reduced (e.g., significantly reduced). When applying PROF, the encoder and / or decoder can skip the application of deblocking filters and / or apply weaker filters to sub-block boundaries. For a CU that is partitioned into multiple transform units (TUs), blocking effects may appear at transform block boundaries.
[0108] In certain representative embodiments, the encoder and / or decoder may skip application of a deblocking filter or may apply one or more weaker filters across a subblock boundary unless the subblock boundary coincides with a TU boundary.
[0109] When PROF is applied to luma (e.g., only to luma) or under the condition that PROF is applied to luma (e.g., only to luma), the encoder and / or decoder may skip the application of the deblocking filter and / or may apply one or more weaker filters for luma (e.g., only luma) at sub-block boundaries. For example, the boundary strength parameter Bs may be used to apply a weaker deblocking filter.
[0110] For example, when PROF is applied, the encoder and / or decoder may skip applying a deblocking filter to subblock boundaries unless the subblock boundaries coincide with TU boundaries. In this case, a deblocking filter may be applied to reduce or remove blocking artifacts that may occur along TU boundaries.
[0111] As another example, unless the sub-block boundary coincides with a TU boundary, when PROF is applied, the encoder and / or decoder may apply a weaker deblocking filter to the sub-block boundary. It is contemplated that the "weaker" deblocking filter may be a weaker deblocking filter than the deblocking filter that would normally be applied to the sub-block boundary when PROF is not applied. When the sub-block boundary coincides with a TU boundary, a stronger deblocking filter may be applied to reduce or remove blocking artifacts that are expected to be more visible along the sub-block boundary that coincides with the TU boundary.
[0112] In certain representative embodiments, when PROF is applied (e.g., only) to luma or under the condition that PROF is applied (e.g., only) to luma, the encoder and / or decoder may, for example, align the application of a deblocking filter for chroma to luma for design consistency purposes despite the lack of application of PROF to chroma. For example, if PROF is applied only to luma, the normal application of a deblocking filter for luma may be altered based on whether PROF is applied (and possibly based on whether a TU boundary exists at a sub-block boundary). In certain representative embodiments, a deblocking filter may be applied to chroma sub-block boundaries to match (and / or mirror) the process of luma deblocking, rather than having separate / different logic for applying a deblocking filter to the corresponding chroma pixels.
[0113] Figure 18A is a flowchart illustrating a first representative encoding and / or decoding method.
[0114] refer to Figure 18A The representative method 1800 for encoding and / or decoding may include: at block 1805, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain a sub-block-based motion prediction signal for a current block of a video, for example. At block 1810, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain one or more spatial gradients of the sub-block-based motion prediction signal for the current block or one or more motion vector differences associated with sub-blocks of the current block. At block 1815, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain a refinement signal for the current block based on the one or more obtained spatial gradients or the one or more obtained motion vector differences associated with the sub-blocks of the current block. At block 1820, the encoder 100 or 300 and / or the decoder 200 or 500 may obtain a refined motion prediction signal for the current block based on the sub-block-based motion prediction signal and the refinement signal. In some embodiments, the encoder 100 or 300 may encode the current block based on the refined motion prediction signal, or the decoder 200 or 500 may decode the current block based on the refined motion prediction signal. The refined motion prediction signal may be a refined motion inter-frame prediction signal generated (e.g., by the GBi encoder 300 and / or the GBi decoder 500), and may use one or more PROF operations.
[0115] In certain representative embodiments, such as representative embodiments related to other methods described herein including methods 1850 and 1900, obtaining the sub-block based motion prediction signal for the current block of the video may include generating the sub-block based motion prediction signal.
[0116] In certain representative embodiments, e.g., with respect to other methods described herein, including in particular methods 1850 and 1900, obtaining the one or more spatial gradients of the sub-block based motion prediction signal of the current block or the one or more motion vector difference values associated with the sub-blocks of the current block may include determining the one or more spatial gradients of the sub-block based motion prediction signal (e.g., associated with a gradient filter).
[0117] In certain representative embodiments, for example, representative embodiments with respect to other methods described herein, particularly including methods 1850 and 1900, obtaining the one or more spatial gradients of the sub-block-based motion prediction signal of the current block or the one or more motion vector difference values associated with the sub-blocks of the current block may include: determining the one or more motion vector difference values associated with the sub-blocks of the current block.
[0118] In certain representative embodiments, such as representative embodiments related to other methods described herein, including in particular methods 1850 and 1900, obtaining the refinement signal for the current block based on the one or more determined spatial gradients or the one or more determined motion vector differences may include determining a motion prediction refinement signal for the current block as the refinement signal based on the determined spatial gradients.
[0119] In certain representative embodiments, such as representative embodiments related to other methods described herein, including methods 1850 and 1900, obtaining the refinement signal for the current block based on the one or more determined spatial gradients or the one or more determined motion vector difference values may include determining, based on the determined motion vector difference values, a motion prediction refinement signal for the current block as the refinement signal.
[0120] The terms "determine" or "determining" when referring to something such as information may generally include one or more of: estimating, calculating, predicting, obtaining, and / or retrieving the information. For example, determining may refer to retrieving something from a memory or a bitstream, etc.
[0121] In certain representative embodiments, for example, representative embodiments related to other methods described herein, including methods 1850 and 1900, obtaining the refined motion prediction signal for the current block based on the sub-block based motion prediction signal and the refinement signal may include combining (e.g., adding or subtracting, etc.) the sub-block based motion prediction signal and the motion prediction refinement signal to generate the refined motion prediction signal for the current block.
[0122] In certain representative embodiments, for example, representative embodiments related to other methods described herein, including methods 1850 and 1900, encoding and / or decoding the current block based on the refined motion prediction signal may include: encoding the video using the refined motion prediction signal as a prediction for the current block, and / or decoding the video using the refined motion prediction signal as a prediction for the current block.
[0123] Figure 18B is a flowchart illustrating a second representative encoding and / or decoding method.
[0124] refer to Figure 18B , a representative method 1850 for encoding and / or decoding a video may include: at block 1855, the encoder 100 or 300 and / or the decoder 200 or 500 generates a sub-block-based motion prediction signal. At block 1860, the encoder 100 or 300 and / or the decoder 200 or 500 may determine one or more spatial gradients (e.g., associated with a gradient filter) of the sub-block-based motion prediction signal. At block 1865, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined spatial gradients. At block 1870, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block-based motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the current block. At block 1875, the encoder 100 or 300 may use the refined motion prediction signal as a prediction for the current block to encode the video, and / or the decoder 200 or 500 may use the refined motion prediction signal as a prediction for the current block to decode the video. In some embodiments, the operations at blocks 1810, 1820, 1830, and 1840 may be performed for a current block, which generally refers to a block currently being encoded or decoded. The refined motion prediction signal may be a refined motion inter prediction signal generated (e.g., by the GBi encoder 300 and / or GBi decoder 500), and may use one or more PROF operations.
[0125] For example, the determination of the one or more spatial gradients of the sub-block-based motion prediction signal by the encoder 100 or 300 and / or the decoder 200 or 500 may include determining a first set of spatial gradients associated with a first reference picture and a second set of spatial gradients associated with a second reference picture. The determination of the motion prediction refinement signal of the current block by the encoder 100 or 300 and / or the decoder 200 or 500 may be based on the determined spatial gradients and may include determining an inter-motion prediction refinement signal (e.g., a bi-prediction signal) for the current block based on the first and second sets of spatial gradients, and may also be based on weighting information W (e.g., indicating or including one or more weighting values associated with one or more reference pictures).
[0126] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the encoder 100 or 300 may generate, use, and / or send weight information W to the decoder 200 or 500, and / or the decoder 200 or 500 may receive or obtain the weight information W. For example, the inter-motion prediction refinement signal for the current block may be based on: (1) a first gradient value derived from the first set of spatial gradients and weighted according to a first weight factor indicated by the weight information W, and / or (2) a second gradient value derived from the second set of spatial gradients and weighted according to a second weight factor indicated by the weight information W.
[0127] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200 and 2600, these methods may further include the encoder 100 or 300 and / or the decoder 200 or 500 determining affine motion model parameters for the current block of the video, so that the determined affine motion model parameters can be used to generate the sub-block based motion prediction signal.
[0128] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by the encoder 100 or 300 and / or the decoder 200 or 500, one or more spatial gradients of the sub-block based motion prediction signal, which may include calculating at least one gradient value for one corresponding sample position, a portion of the corresponding sample positions, or each corresponding sample position in at least one sub-block of the sub-block based motion prediction signal. For example, calculating the at least one gradient value for one corresponding sample position, a portion of the corresponding sample positions, or each corresponding sample position in at least one sub-block of the sub-block based motion prediction signal may include applying a gradient filter to the corresponding sample position in the at least one sub-block of the sub-block based motion prediction signal.
[0129] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may further include encoder 100 or 300 and / or decoder 200 or 500 determining a set of motion vector difference values associated with sample positions of a first subblock of the current block of the subblock-based motion prediction signal. In some examples, the difference values may be determined for a subblock (e.g., the first subblock) and may be reused for some or all of the other subblocks in the current block. In some examples, the set of motion vector difference values may be determined using an affine motion model or a different motion model (e.g., another subblock-based motion model, such as the SbTMVP model), and the subblock-based motion prediction signal may be generated. As an example, the set of motion vector difference values may be determined for the first subblock of the current block and used to determine the motion prediction refinement signal for one or more further subblocks of the current block.
[0130] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the one or more spatial gradients of the sub-block based motion prediction signal and the set of motion vector difference values may be used to determine the motion prediction refinement signal for the current block.
[0131] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the set of motion vector differences may be determined and the sub-block based motion prediction signal may be generated using an affine motion model of the current block.
[0132] In certain representative embodiments, including representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, determining the one or more spatial gradients of the sub-block based motion prediction signal may include: for one or more corresponding sub-blocks of the current block: determining an extended sub-block using the sub-block based motion prediction signal and neighboring reference samples adjacent to and surrounding the corresponding sub-block; and determining the spatial gradient of the corresponding sub-block using the determined extended sub-block to determine the motion prediction refinement signal.
[0133] Figure 19 is a flowchart illustrating a third representative encoding and / or decoding method.
[0134] Reference Figure 19 The representative method 1900 for encoding and / or decoding video may include: at block 1910, the encoder 100 or 300 and / or the decoder 200 or 500 may generate a sub-block-based motion prediction signal. At block 1920, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a set of motion vector difference values associated with the sub-blocks of the current block (e.g., the set of motion vector difference values may be associated with, for example, all sub-blocks of the current block). At block 1930, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined set of motion vector difference values. At block 1940, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block-based motion prediction signal and the motion prediction refinement signal to produce or generate a refined motion prediction signal for the current block. At block 1950, the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block, and / or the decoder 200 or 500 may decode the video using the refined motion prediction signal as a prediction for the current block. In some embodiments, the operations at blocks 1910, 1920, 1930, and 1940 may be performed for a current block, which generally refers to a block currently being encoded or decoded. In some representative embodiments, the refined motion prediction signal may be a refined motion inter prediction signal generated (e.g., by the GBi encoder 300 or GBi decoder 500), and one or more PROF operations may be used.
[0135] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by the encoder 100 or 300 and / or the decoder 200 or 500, motion model parameters (e.g., one or more affine motion model parameters) of the current block of the video, such that the sub-block-based motion prediction signal can be generated using the determined motion model parameters (e.g., affine motion model parameters).
[0136] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by the encoder 100 or 300 and / or the decoder 200 or 500, one or more spatial gradients of the sub-block based motion prediction signal. For example, the determining of the one or more spatial gradients of the sub-block based motion prediction signal may include calculating at least one gradient value for one corresponding sample position, a portion of corresponding sample positions, or each corresponding sample position in at least one sub-block of the sub-block based motion prediction signal. For example, the calculating of at least one gradient value for one corresponding sample position, a portion of corresponding sample positions, or each corresponding sample position in at least one sub-block of the sub-block based motion prediction signal may include applying a gradient filter to the corresponding sample position in the at least one sub-block of the sub-block based motion prediction signal for the one corresponding sample position, a portion of corresponding sample positions, or each corresponding sample position.
[0137] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include determining, by the encoder 100 or 300 and / or the decoder 200 or 500, the motion prediction refinement signal for the current block using gradient values associated with spatial gradients of one, a portion of, or each corresponding sample position of the current block and a determined set of motion vector difference values associated with the sample positions of a subblock (e.g., any subblock) of the current block of the subblock motion prediction signal.
[0138] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the determination of the motion prediction refinement signal for the current block may use gradient values associated with spatial gradients of one or more corresponding sample positions or each sample position of one or more sub-blocks of the current block and the determined set of motion vector difference values.
[0139] Figure 20is a flowchart illustrating a fourth representative encoding and / or decoding method.
[0140] refer to Figure 20 The representative method 2000 for encoding and / or decoding video may include: at block 2010, the encoder 100 or 300 and / or the decoder 200 or 500 may generate a sub-block-based motion prediction signal using at least a first motion vector for a first sub-block of the current block and a further motion vector for a second sub-block of the current block. At block 2020, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate a first set of gradient values for a first sample position in the first sub-block of the sub-block-based motion prediction signal and a second, different set of gradient values for a second sample position in the first sub-block of the sub-block-based motion prediction signal. At block 2030, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a first set of motion vector difference values for the first sample position and a second, different set of motion vector difference values for the second sample position. For example, the first set of motion vector difference values for the first sample position may indicate the difference between the motion vector at the first sample position and the motion vector of the first sub-block, and the second set of motion vector difference values for the second sample position may indicate the difference between the motion vector at the second sample position and the motion vector of the first sub-block. At block 2040, the encoder 100 or 300 and / or the decoder 200 or 500 may use the first and second sets of gradient values and the first and second sets of motion vector difference values to determine a prediction refinement signal. At block 2050, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block-based motion prediction signal with the prediction refinement signal to generate a refined motion prediction signal. At block 2060, the encoder 100 or 300 may use the refined motion prediction signal as a prediction for the current block to encode the video, and / or the decoder 200 or 500 may use the refined motion prediction signal as a prediction for the current block to decode the video. In some embodiments, the operations at blocks 2010 , 2020 , 2030 , 2040 , and 2050 may be performed for a current block including a plurality of sub-blocks.
[0141] Figure 21 is a flowchart illustrating a fifth representative encoding and / or decoding method.
[0142] refer to Figure 21The representative method 2100 for encoding and / or decoding video may include: at block 2110, the encoder 100 or 300 and / or the decoder 200 or 500 may generate a sub-block-based motion prediction signal for a current block. At block 2120, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a prediction refinement signal using optical flow information indicating refined motion of a plurality of sample positions in the current block of the sub-block-based motion prediction signal. At block 2130, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block-based motion prediction signal with the prediction refinement signal to generate a refined motion prediction signal. At block 2140, the encoder 100 or 300 may use the refined motion prediction signal as a prediction for the current block to encode the video, and / or the decoder 200 or 500 may use the refined motion prediction signal as a prediction for the current block to decode the video. For example, the current block may include a plurality of subblocks, and the subblock-based motion prediction signal may be generated using at least a first motion vector of a first subblock of the current block and a further motion vector of a second subblock of the current block.
[0143] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, a prediction refinement signal that can use optical flow information. This determination may include calculating, by encoder 100 or 300 and / or decoder 200 or 500, a first set of gradient values for a first sample position in the first sub-block of the sub-block-based motion prediction signal and a second, different set of gradient values for a second sample position in the first sub-block of the sub-block-based motion prediction signal. A first set of motion vector difference values for the first sample position and a second, different set of motion vector difference values for the second sample position may be determined. For example, the first set of motion vector difference values for the first sample position may indicate a difference between a motion vector at the first sample position and a motion vector of the first sub-block, and the second set of motion vector difference values for the second sample position may indicate a difference between a motion vector at the second sample position and a motion vector of the first sub-block. The encoder 100 or 300 and / or the decoder 200 or 500 may use the first and second sets of gradient values and the first and second sets of motion vector difference values to determine the prediction refinement signal.
[0144] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include determining, by encoder 100 or 300 and / or decoder 200 or 500, the prediction refinement signal that may use optical flow information. This determination may include calculating a third set of gradient values for a first sample position in a second sub-block of the sub-block-based motion prediction signal and a fourth set of gradient values for a second sample position in the second sub-block of the sub-block-based motion prediction signal. The encoder 100 or 300 and / or decoder 200 or 500 may use the third and fourth sets of gradient values and the first and second sets of motion vector difference values to determine the prediction refinement signal for the second sub-block.
[0145] Figure 22 is a flowchart illustrating a sixth representative encoding and / or decoding method.
[0146] refer to Figure 22, a representative method 2200 for encoding and / or decoding a video may include: at block 2210, the encoder 100 or 300 and / or the decoder 200 or 500 determines a motion model for a current block of the video. The current block may include multiple sub-blocks. For example, the motion model may generate separate (e.g., per-sample) motion vectors for multiple sample positions in the current block. At block 2220, the encoder 100 or 300 and / or the decoder 200 or 500 may use the determined motion model to generate a sub-block-based motion prediction signal for the current block. The generated sub-block-based motion prediction signal may use one motion vector for each sub-block of the current block. At block 2230, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate a gradient value by applying a gradient filter to a portion of the multiple sample positions of the sub-block-based motion prediction signal. At block 2240, the encoder 100 or 300 and / or the decoder 200 or 500 may determine motion vector difference values for the portion of the sample positions, each of the motion vector difference values indicating a difference between a motion vector (e.g., an individual motion vector) generated for the corresponding sample position according to the motion model and the motion vector used to generate the sub-block-based motion prediction signal for the sub-block including the corresponding sample position. At block 2250, the encoder 100 or 300 and / or the decoder 200 or 500 may use the gradient value and the motion vector difference values to determine a prediction refinement signal. At block 2260, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block-based motion prediction signal with the prediction refinement signal to generate a refined motion prediction signal for the current block. In block 2270 , the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block, and / or the decoder 200 or 500 may decode the video using the refined motion prediction signal as a prediction for the current block.
[0147] Figure 23 is a flowchart illustrating a seventh representative encoding and / or decoding method.
[0148] refer to Figure 23, a representative method 2300 for encoding and / or decoding video may include: at block 2310, the encoder 100 or 300 and / or the decoder 200 or 500 performs sub-block-based motion compensation to generate a sub-block-based motion prediction signal as a coarse motion prediction signal. At block 2320, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate one or more spatial gradients of the sub-block-based motion prediction signal at a plurality of sample positions. At block 2330, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate a per-pixel intensity change in the current block based on the calculated spatial gradients. At block 2340, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a per-pixel motion prediction signal as a refined motion prediction signal based on the calculated per-pixel intensity change. At block 2350, the encoder 100 or 300 and / or the decoder 200 or 500 may predict the current block using the coarse motion prediction signal for each sub-block of the current block and the refined motion prediction signal for each pixel of the current block. In some embodiments, the operations at blocks 2310, 2320, 2330, 2340, and 2350 may be performed for at least one block in a video (e.g., the current block). For example, calculating the intensity change for each pixel in the current block may include determining the luminance intensity change for each pixel in the current block according to an optical flow equation. Predicting the current block may include predicting a motion vector for each corresponding pixel in the current block by combining a coarse motion prediction vector for a sub-block including the corresponding pixel with a refined motion prediction vector relative to the coarse motion prediction vector and associated with the corresponding pixel.
[0149] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300 and 2600, the one or more spatial gradients of the sub-block-based motion prediction signal may include any of the following: a horizontal gradient and / or a vertical gradient, and for example, the horizontal gradient may be calculated as a luma difference or a chroma difference between a right neighboring sample of a sample of the sub-block and a left neighboring sample of the sample of the sub-block, and / or the vertical gradient may be calculated as a luma difference or a chroma difference between a bottom neighboring sample of the sample of the sub-block and a top neighboring sample of the sample of the sub-block.
[0150] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, a Sobel filter may be used to generate the one or more spatial gradients of the sub-block prediction.
[0151] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the coarse motion prediction signal may use one of the following: a 4-parameter affine model or a 6-parameter affine model. For example, the sub-block based motion compensation may be one of the following: (1) affine sub-block based motion compensation; or (2) another compensation (e.g., sub-block based temporal motion vector prediction (SbTMVP) mode motion compensation; and / or regression based motion vector field (RMVF) mode compensation). Under the condition that motion compensation based on the SbTMVP mode is performed, the method may include: estimating affine model parameters using the sub-block motion vector field through a linear regression operation; and deriving pixel-level motion vectors using the estimated affine model parameters. Under the condition that motion compensation based on the RMVF mode is performed, the method may include: estimating affine model parameters; and deriving pixel-level motion vector offsets from the sub-block level motion vectors using the estimated affine model parameters. For example, the pixel motion vector offset may be relative to the center of the sub-block (eg, the actual center or the sample position closest to the actual center). For example, the coarse motion prediction vector of the sub-block may be based on the actual center position of the sub-block.
[0152] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the methods may include: the encoder 100 or 300 or the decoder 200 or 500 selecting one of the following as the center position associated with the coarse motion prediction vector (e.g., sub-block-based motion prediction vector) for each sub-block: (1) the actual center of each sub-block; or (2) one of the pixel (e.g., sample) position closest to the center of the sub-block. For example, predicting the current block using the coarse motion prediction signal (e.g., sub-block-based motion prediction signal) of the current block and using the refined motion prediction signal for each pixel (e.g., sample) of the current block may be based on the selected center position of each sub-block. For example, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a center position associated with a chroma pixel of the sub-block, and may determine an offset to the center position of the chroma pixel of the sub-block based on a chroma position sample type associated with the chroma pixel. The coarse motion prediction signal for the sub-block (e.g., the sub-block-based motion prediction signal) may be based on an actual position of the sub-block corresponding to the determined center position of the chroma pixel adjusted by the offset.
[0153] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, the methods may include the encoder 100 or 300 generating or the decoder 200 or 500 receiving information indicating whether prediction refinement using optical flow (PROF) is enabled in one of the following: (1) a sequence parameter set (SPS) header, (2) a picture parameter set (PPS) header, or (3) a tile group header. For example, if PROF is enabled, a refined motion prediction operation may be performed such that the coarse motion prediction signal (e.g., a sub-block based motion prediction signal) and the refined motion prediction signal may be used to predict the current block. As another example, if PROF is not enabled, the refined motion prediction operation is not performed such that only the coarse motion prediction signal (e.g., a sub-block based motion prediction signal) may be used to predict the current block.
[0154] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining, based on properties of the current block and / or properties of the affine motion estimate, whether to perform a refined motion prediction operation on the current block or to perform a refined motion prediction operation in the affine motion estimate.
[0155] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining, based on properties of the current block and / or properties of an affine motion estimate, whether to perform a refined motion prediction operation on the current block or to perform a refined motion prediction operation in the affine motion estimate. For example, determining whether to perform a refined motion prediction operation on the current block based on the properties of the current block may include determining whether to perform a refined motion prediction operation on the current block based on any of the following: (1) a size of the current block exceeds a specific size; and / or (2) a control point motion vector (CPMV) difference exceeds a threshold.
[0156] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 applying a first deblocking filter to one or more boundaries of subblocks of the current block that coincide with transform unit boundaries, and applying a second, different deblocking filter to other boundaries of the subblocks of the current block that do not coincide with any transform unit boundaries. For example, the first deblocking filter may be a stronger deblocking filter than the second deblocking filter.
[0157] Figure 24 is a flowchart illustrating an eighth representative encoding and / or decoding method.
[0158] refer to Figure 24 The representative method 2400 for encoding and / or decoding video may include: at block 2410, the encoder 100 or 300 and / or the decoder 200 or 500 may perform subblock-based motion compensation to generate a subblock-based motion prediction signal as a coarse motion prediction signal. At block 2420, the encoder 100 or 300 and / or the decoder 200 or 500 may determine, for each corresponding boundary sample of a subblock of the current block, one or more reference samples that are adjacent to the corresponding boundary sample and surround the subblock as surrounding reference samples, and may use the surrounding reference samples and samples of the subblock adjacent to the corresponding boundary sample to determine one or more spatial gradients associated with the corresponding boundary sample. At block 2430, the encoder 100 or 300 and / or the decoder 200 or 500 may use samples of the subblock adjacent to the corresponding non-boundary sample to determine one or more spatial gradients associated with the corresponding non-boundary sample. At block 2440, the encoder 100 or 300 and / or the decoder 200 or 500 may calculate an intensity change per pixel in the current block using the determined spatial gradients of the sub-blocks. At block 2450, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a per-pixel motion prediction signal as a refined motion prediction signal based on the calculated per-pixel intensity change. At block 2460, the encoder 100 or 300 and / or the decoder 200 or 500 may predict the current block using the coarse motion prediction signal associated with each sub-block of the current block and using the refined motion prediction signal associated with each pixel of the current block. In some embodiments, the operations at blocks 2410, 2420, 2430, 2440, 2450, and 2460 may be performed for at least one block in a video (e.g., the current block).
[0159] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400 and 2600, the determining of the one or more spatial gradients of the boundary samples and the non-boundary samples may include: calculating the one or more spatial gradients using any of the following: (1) a vertical Sobel filter; (2) a horizontal Sobel filter; or (3) a 3-tap filter.
[0160] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 copying surrounding reference samples from a reference repository without any further operation, and the determination of the one or more spatial gradients associated with the corresponding boundary samples may use the copied surrounding reference samples to determine the one or more spatial gradients associated with the corresponding boundary samples.
[0161] Figure 25 is a flowchart illustrating a representative gradient calculation method.
[0162] refer to Figure 25 A representative method 2500 for calculating gradients for a sub-block using reference samples corresponding to samples adjacent to a boundary of the sub-block (e.g., used in encoding and / or decoding a video) may include: at block 2510, the encoder 100 or 300 and / or decoder 200 or 500 may determine, for each corresponding boundary sample of the sub-block of the current block, one or more reference samples corresponding to samples adjacent to the corresponding boundary sample and surrounding the sub-block as surrounding reference samples, and use the surrounding reference samples and samples of the sub-block adjacent to the corresponding boundary sample to determine one or more spatial gradients associated with the corresponding boundary sample. At block 2520, the encoder 100 or 300 and / or decoder 200 or 500 may determine, for each corresponding non-boundary sample in the sub-block, one or more spatial gradients associated with the corresponding non-boundary sample using samples of the sub-block adjacent to the corresponding non-boundary sample. In some embodiments, the operations at blocks 2510 and 2520 may be performed for at least one block in a video (e.g., the current block).
[0163] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2400, 2500, and 2600, the determined one or more spatial gradients may be used to predict the current block by any of: (1) a prediction refinement (PROF) operation using optical flow; (2) a bidirectional optical flow operation; or (3) an affine motion estimation operation.
[0164] Figure 26 is a flowchart illustrating a ninth representative encoding and / or decoding method.
[0165] refer to Figure 26 , a representative method 2600 for encoding and / or decoding a video may include: at block 2610, the encoder 100 or 300 and / or the decoder 200 or 500 generates a sub-block-based motion prediction signal for a current block of the video. For example, the current block may include multiple sub-blocks. At block 2620, the encoder 100 or 300 and / or the decoder 200 or 500 may determine, for one or more or each corresponding sub-block of the current block, an extended sub-block using the sub-block-based motion prediction signal and neighboring reference samples adjacent to and surrounding the corresponding sub-block, and determine a spatial gradient of the corresponding sub-block using the determined extended sub-block. At block 2630, the encoder 100 or 300 and / or the decoder 200 or 500 may determine a motion prediction refinement signal for the current block based on the determined spatial gradient. At block 2640, the encoder 100 or 300 and / or the decoder 200 or 500 may combine (e.g., add or subtract, etc.) the sub-block-based motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the current block. At block 2650, the encoder 100 or 300 may use the refined motion prediction signal as a prediction for the current block to encode the video, and / or the decoder 200 or 500 may use the refined motion prediction signal as a prediction for the current block to decode the video. In some embodiments, the operations at blocks 2610, 2620, 2630, 2640, and 2650 may be performed for at least one block in the video (e.g., the current block).
[0166] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include encoder 100 or 300 and / or decoder 200 or 500 copying the neighboring reference samples from a reference store without any further manipulation. For example, the determination of the spatial gradient of the corresponding sub-block may use the copied neighboring reference samples to determine gradient values associated with sample positions on a boundary of the corresponding sub-block. The neighboring reference samples of the extended block may be copied from the nearest integer position in a reference picture containing the current block. In some examples, the neighboring reference samples of the extended block have the nearest integer motion vector rounded from the original precision.
[0167] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining affine motion model parameters for the current block of the video, such that the determined affine motion model parameters can be used to generate the sub-block based motion prediction signal.
[0168] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2300, 2500, and 2600, determining the spatial gradient of the corresponding sub-block may include calculating at least one gradient value for each corresponding sample position in the corresponding sub-block. For example, calculating the at least one gradient value for each corresponding sample position in the corresponding sub-block may include applying a gradient filter to the corresponding sample position in the corresponding sub-block for each corresponding sample position. As another example, calculating the at least one gradient value for each corresponding sample position in the corresponding sub-block may include determining an intensity change for each corresponding sample position in the corresponding sub-block according to an optical flow equation.
[0169] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, these methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining a set of motion vector difference values associated with the sample positions of the corresponding sub-blocks. For example, by using an affine motion model of the current block, the sub-block-based motion prediction signal may be generated and the set of motion vector difference values may be determined.
[0170] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the set of motion vector difference values may be determined for the corresponding sub-block of the current block, and the set of motion vector difference values may be used to determine motion prediction refinement signals for other remaining sub-blocks of the current block.
[0171] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the determining of the spatial gradient of the corresponding sub-block may include: calculating the spatial gradient using any of the following: (1) a vertical Sobel filter; (2) a horizontal Sobel filter; and / or (3) a 3-tap filter.
[0172] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, neighboring reference samples adjacent to and surrounding the respective sub-block may use integer motion compensation.
[0173] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the spatial gradient of the corresponding sub-block may include any of the following: a horizontal gradient or a vertical gradient. For example, the horizontal gradient may be calculated as a luma difference or a chroma difference between a right neighboring sample of the corresponding sample and a left neighboring sample of the corresponding sample; and / or the vertical gradient may be calculated as a luma difference or a chroma difference between a bottom neighboring sample of the corresponding sample and a top neighboring sample of the corresponding sample.
[0174] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the sub-block-based motion prediction signal may be generated using any of the following: (1) a 4-parameter affine model; (2) a 6-parameter affine model; (3) sub-block-based temporal motion vector prediction (SbTMVP) mode motion compensation; or (4) regression-based motion compensation. For example, when performing SbTMVP mode motion compensation, the method may include: estimating affine model parameters using the sub-block motion vector field through a linear regression operation; and / or deriving pixel-level motion vectors using the estimated affine model parameters. As another example, when performing RMVF mode-based motion compensation, the method may include: estimating affine model parameters; and / or deriving pixel-level motion vector offsets from the sub-block-level motion vectors using the estimated affine model parameters. The pixel motion vector offsets may be relative to the center of the corresponding sub-block.
[0175] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the refined motion prediction signal for the respective sub-block may be based on the actual center position of the respective sub-block or may be based on the sample position closest to the actual center of the respective sub-block.
[0176] For example, the methods may include the encoder 100 or 300 and / or the decoder 200 or 500 selecting one of the following as the center position associated with the motion prediction vector for each respective sub-block: (1) the actual center of each respective sub-block, or (2) the sample position closest to the actual center of the respective sub-block. The refined motion prediction signal may be based on the selected center position of each sub-block.
[0177] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the methods may include encoder 100 or 300 and / or decoder 200 or 500 determining a center position associated with a chroma pixel of the corresponding sub-block; and determining an offset to the center position of the chroma pixel of the corresponding sub-block based on a chroma position sample type associated with the chroma pixel. A refined motion prediction signal for the corresponding sub-block may be based on an actual position of the sub-block corresponding to the determined center position of the chroma pixel adjusted by the offset.
[0178] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, and 2600, the encoder 100 or 300 may generate and send information indicating whether prediction refinement with optical flow (PROF) is enabled in one of the following: (1) a sequence parameter set SPS header, (2) a picture parameter set PPS header, or (3) a tile group header, and / or the decoder 200 or 500 may receive information indicating whether PROF is enabled in one of the following: (1) the SPS header, (2) the PPS header, or (3) the tile group header.
[0179] Figure 27 is a flowchart illustrating a tenth representative encoding and / or decoding method.
[0180] refer to Figure 27, a representative method 2700 for encoding and / or decoding a video may include: at block 2710, the encoder 100 or 300 and / or the decoder 200 or 500 determines an actual center position of each corresponding sub-block of the current block. At block 2720, the encoder 100 or 300 and / or the decoder 200 or 500 may use the actual center position of each corresponding sub-block of the current block to generate a sub-block-based motion prediction signal or a refined motion prediction signal. At block 2730, (1) the encoder 100 or 300 may use the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block to encode the video, or (2) the decoder 200 or 500 may use the sub-block-based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block to decode the video. In some embodiments, the operations at blocks 2710, 2720, and 2730 may be performed for at least one block in the video (e.g., the current block). For example, the determination of the actual center position of each corresponding sub-block of the current block may include: determining a chroma center position associated with a chroma pixel of the corresponding sub-block and an offset of the chroma center position relative to the center position of the corresponding sub-block based on the chroma position sample type of the chroma pixel. The sub-block-based motion prediction signal or the refined motion prediction signal for the corresponding sub-block may be based on the actual center position of the corresponding sub-block, which corresponds to the determined chroma center position adjusted by the offset. Although the actual center of each corresponding sub-block of the current block is described as being determined / used for various operations, it is contemplated that one, some, or all of the center positions of such sub-blocks may be determined / used.
[0181] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the generation of the refined motion prediction signal may be performed using the sub-block-based motion prediction signal by: determining one or more spatial gradients of the sub-block-based motion prediction signal for each corresponding sub-block of the current block, determining a motion prediction refinement signal for the current block based on the determined spatial gradients, and / or combining the sub-block-based motion prediction signal with the motion prediction refinement signal to generate the refined motion prediction signal for the current block. For example, the determining of the one or more spatial gradients of the sub-block-based motion prediction signal may include: determining an extended sub-block using the sub-block-based motion prediction signal and neighboring reference samples adjacent to and surrounding the corresponding sub-block, and / or determining the one or more spatial gradients of the corresponding sub-block using the determined extended sub-block.
[0182] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, determining the spatial gradient of the corresponding sub-block may include calculating at least one gradient value for each corresponding sample position in the corresponding sub-block. For example, calculating the at least one gradient value for each corresponding sample position in the corresponding sub-block may include applying a gradient filter to each corresponding sample position in the corresponding sub-block.
[0183] As another example, calculating the at least one gradient value for each corresponding sample position in the corresponding sub-block may include: determining intensity changes of one or more corresponding sample positions in the corresponding sub-block according to an optical flow equation.
[0184] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the methods may include the encoder 100 or 300 and / or the decoder 200 or 500 determining a set of motion vector difference values associated with the sample positions of the corresponding sub-block. The sub-block-based motion prediction signal may be generated using an affine motion model of the current block, and the set of motion vector difference values may be determined. In some examples, the set of motion vector difference values may be determined for the corresponding sub-block of the current block, and the set of motion vector difference values may be used (e.g., reused) to determine a motion prediction refinement signal for the sub-block and other remaining sub-blocks of the current block. For example, determining the spatial gradient of the corresponding sub-block may include calculating the spatial gradient using any of the following: (1) a vertical Sobel filter; (2) a horizontal Sobel filter; and / or (3) a 3-tap filter. Neighboring reference samples adjacent to and surrounding the corresponding sub-block may use integer motion compensation.
[0185] In some embodiments, the spatial gradient of the corresponding sub-block may include any of the following: a horizontal gradient or a vertical gradient. For example, the horizontal gradient may be calculated as the luma difference or chroma difference between the right adjacent sample of the corresponding sample and the left adjacent sample of the corresponding sample. As another example, the vertical gradient may be calculated as the luma difference or chroma difference between the bottom adjacent sample of the corresponding sample and the top adjacent sample of the corresponding sample.
[0186] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the sub-block-based motion prediction signal may be generated using any of the following: (1) a 4-parameter affine model; (2) a 6-parameter affine model; (3) sub-block-based temporal motion vector prediction (SbTMVP) mode motion compensation; and / or (4) regression-based motion compensation. For example, when performing SbTMVP mode motion compensation, the method may include: estimating affine model parameters using a sub-block motion vector field through a linear regression operation; and / or deriving pixel-level motion vectors using the estimated affine model parameters. As another example, when performing regression motion vector field (RMVF) mode motion compensation, the method may include: estimating affine model parameters; and / or deriving pixel-level motion vector offsets from sub-block-level motion vectors using the estimated affine model parameters, wherein the pixel motion vector offsets are relative to the center of the corresponding sub-block.
[0187] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the refined motion prediction signal may be generated using multiple motion vectors associated with control points of the current block.
[0188] In certain representative embodiments including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, 2700, and 2800, the encoder 100 or 300 may generate, encode, and send information in one of the following, and the decoder 200 or 500 may receive and decode the information in one of the following, indicating whether prediction refinement (PROF) with optical flow is enabled: (1) a sequence parameter set SPS header, (2) a picture parameter set PPS header, or (3) a tile group header.
[0189] Figure 28 is a flowchart illustrating an eleventh representative encoding and / or decoding method.
[0190] refer to Figure 28, a representative method 2800 for encoding and / or decoding video may include: at block 2810, the encoder 100 or 300 and / or the decoder 200 or 500 selects one of the following as a center position associated with a motion prediction vector for each corresponding sub-block: (1) the actual center of each corresponding sub-block, or (2) a sample position closest to the actual center of the corresponding sub-block. At block 2820, the encoder 100 or 300 and / or the decoder 200 or 500 may determine the selected center position of each corresponding sub-block of the current block. At block 2830, the encoder 100 or 300 and / or the decoder 200 or 500 may use the selected center position of each corresponding sub-block of the current block to generate a sub-block-based motion prediction signal or a refined motion prediction signal. At block 2840, (1) the encoder 100 or 300 may encode the video using the sub-block based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block, or (2) the decoder 200 or 500 may decode the video using the sub-block based motion prediction signal or the generated refined motion prediction signal as a prediction for the current block. In some embodiments, the operations at blocks 2810, 2820, 2830, and 2840 may be performed for at least one block in the video (e.g., the current block). Although the selection of the center position is described with respect to each corresponding sub-block of the current block, it is contemplated that one, a portion, or all of the center positions of such sub-blocks may be selected / used in various operations.
[0191] Figure 29 is a flowchart illustrating a representative encoding method.
[0192] See Figure 29 A representative method 2900 for encoding a video may include, at block 2910, the encoder 100 or 300 performing motion estimation on a current block of the video, including determining affine motion model parameters for the current block using an iterative motion compensation operation, and generating a sub-block-based motion prediction signal for the current block using the determined affine motion model parameters. At block 2920, after performing motion estimation on the current block, the encoder 100 or 300 may perform a prediction refinement with optical flow (PROF) operation to generate a refined motion prediction signal. At block 2930, the encoder 100 or 300 may encode the video using the refined motion prediction signal as a prediction for the current block. For example, the PROF operation may include determining one or more spatial gradients of the sub-block-based motion prediction signal; determining a motion prediction refinement signal for the current block based on the determined spatial gradients; and / or combining the sub-block-based motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the current block.
[0193] In certain representative embodiments, including at least representative methods 1800, 1850, 1900, 2000, 2100, 2200, 2600, and 2900, the PROF operation may be performed after (e.g., only after) an iterative motion compensation operation is completed. For example, the PROF operation is not performed during motion estimation of the current block.
[0194] Figure 30 is a flowchart illustrating another representative encoding method.
[0195] See Figure 30 , a representative method 3000 for encoding a video may include: at block 3010, the encoder 100 or 300, during motion estimation of the current block, using an iterative motion compensation operation to determine affine motion model parameters, and using the determined affine motion model parameters to generate a sub-block based motion prediction signal. At block 3020, after the motion estimation of the current block, the encoder 100 or 300 may perform a prediction refinement (PROF) operation using optical flow to generate a refined motion prediction signal if the size of the current block meets or exceeds a threshold size. At block 3030, the encoder 100 or 300 may encode the video by: (1) using the refined motion prediction signal as a prediction for the current block if the current block meets or exceeds the threshold size, or (2) using the sub-block based motion prediction signal as a prediction for the current block if the current block does not meet the threshold size.
[0196] Figure 31 is a flowchart illustrating a twelfth representative encoding / decoding method.
[0197] refer to Figure 31, a representative method 3100 for encoding and / or decoding video may include: at block 3110, the encoder 100 or 300 determines or obtains information indicating the size of a current block, or the decoder 200 or 500 receives information indicating the size of a current block. At block 3120, the encoder 100 or 300 or the decoder 200 or 500 may generate a sub-block-based motion prediction signal. At block 3130, under the condition that the size of the current block meets or exceeds a threshold size, the encoder 100 or 300 or the decoder 200 or 500 may perform a prediction refinement (PROF) operation using optical flow to generate a refined motion prediction signal. At box 3140, the encoder 100 or 300 can encode the video by: (1) using the refined motion prediction signal as a prediction for the current block if the current block meets or exceeds a threshold size, or (2) using the sub-block based motion prediction signal as a prediction for the current block if the current block does not meet the threshold size, or the decoder 200 or 500 can decode the video by: (1) using the refined motion prediction signal as a prediction for the current block if the current block meets or exceeds the threshold size, or (2) using the sub-block based motion prediction signal as a prediction for the current block if the current block does not meet the threshold size.
[0198] Figure 32 is a flowchart illustrating a thirteenth representative encoding / decoding method.
[0199] refer to Figure 32, a representative method 3200 for encoding and / or decoding video may include: at block 3210, the encoder 100 or 300 determines whether to perform pixel-level motion compensation, or the decoder 200 or 500 receives a flag indicating whether to perform pixel-level motion compensation. At block 3220, the encoder 100 or 300 or the decoder 200 or 500 may generate a sub-block-based motion prediction signal. At block 3230, under the condition that the pixel-level motion compensation is to be performed, the encoder 100 or 300 or the decoder 200 or 500 may: determine one or more spatial gradients of the sub-block-based motion prediction signal, determine a motion prediction refinement signal for the current block based on the determined spatial gradients, and combine the sub-block-based motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the current block. At block 3240, based on whether the pixel-level motion compensation is to be performed, the encoder 100 or 300 may use the sub-block-based motion prediction signal or the refined motion prediction signal as a prediction for the current block to encode the video, or the decoder 200 or 500 may use the sub-block-based motion prediction signal or the refined motion prediction signal as a prediction for the current block to decode the video based on the indication of the flag. In some embodiments, the operations at blocks 3220 and 3230 may be performed for a block in the video (e.g., the current block).
[0200] Figure 33 is a flowchart illustrating a fourteenth representative encoding / decoding method.
[0201] refer to Figure 33The representative method 3300 for encoding and / or decoding video may include: at block 3310, the encoder 100 or 300 determines or obtains, or the decoder 200 or 500 receives, information indicating inter-frame prediction weights, which indicates one or more weights associated with first and second reference pictures. At block 3320, the encoder 100 or 300 or the decoder 200 or 500 may generate a sub-block-based inter-frame prediction signal for a current block of the video, may determine a first set of spatial gradients associated with the first reference picture and a second set of spatial gradients associated with the second reference picture, may determine an inter-frame prediction refinement signal for the current block based on the first set of spatial gradients and the second set of spatial gradients and the inter-frame prediction weight information, and may combine the sub-block-based inter-frame prediction signal and the inter-frame prediction refinement signal to generate a refined inter-frame prediction signal for the current block. In block 3330, the encoder 100 or 300 may encode the video using the refined inter-motion prediction signal as a prediction for the current block, or the decoder 200 or 500 may decode the video using the refined inter-motion prediction signal as a prediction for the current block. For example, the inter-frame prediction weight information is any of the following: (1) an indicator indicating a first weighting factor to be applied to the first reference picture and / or a second weighting factor to be applied to the second reference picture; or (2) a weight index. In some embodiments, the inter-motion prediction refinement signal for the current block may be based on: (1) a first gradient value derived from the first set of spatial gradients and weighted according to a first weighting factor indicated by the inter-frame prediction weight information, and (2) a second gradient value derived from the second set of spatial gradients and weighted according to a second weighting factor indicated by the inter-frame prediction weight information. Example network for implementation of embodiments
[0202] Figure 34A 3400 is a schematic diagram illustrating an exemplary communication system 3400 in which one or more disclosed embodiments may be implemented. The communication system 3400 may be a multiple access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communication system 3400 may enable multiple wireless users to access such content by sharing system resources, including wireless bandwidth. For example, the communication system 3400 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tailing unique word DFT-spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, and filter bank multi-carrier (FBMC).
[0203] like Figure 34A As shown, the communication system 3400 may include wireless transmit / receive units (WTRUs) 3402a, 3402b, 3402c, 3402d, the RAN 3404 / 3413, the CN 3406 / 3415, the public switched telephone network (PSTN) 3408, the Internet 3410, and other networks 3412. However, it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network components. Each of the WTRUs 3402a, 3402b, 3402c, and 3402d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, any of the WTRUs 3402a, 3402b, 3402c, and 3402d may be referred to as a “station” and / or “STA,” which may be configured to transmit and / or receive wireless signals and may include user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smartphone, a laptop, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable device, a head-mounted display (HMD), a vehicle, a drone, medical equipment and applications (e.g., remote surgery), industrial equipment and applications (e.g., robots and / or other wireless devices operating in industrial and / or automated process chain environments), consumer electronic devices, and devices operating on commercial and / or industrial wireless networks, etc. Any of the WTRUs 3402a, 3402b, 3402c, and 3402d may be interchangeably referred to as a UE.
[0204] The communication system 3400 may also include a base station 3414a and / or a base station 3414b. Each of the base stations 3414a and 3414b may be any type of device configured to facilitate access to one or more communication networks (e.g., CN 3406 / 3415, Internet 3410, and / or other networks 3412) by wirelessly interfacing with at least one of the WTRUs 3402a, 3402b, 3402c, and 3402d. For example, the base stations 3414a and 3414b may be base transceiver stations (BTSs), Node Bs, eNode Bs (terminals), Home Node Bs (HNBs), Home eNode Bs (HeNBs), gNBs, NR Node Bs, site controllers, access points (APs), wireless routers, and the like. Although each of the base stations 3414a and 3414b is depicted as a single component, it should be understood that the base stations 3414a and 3414b may include any number of interconnected base stations and / or network components.
[0205] Base station 3414a may be part of RAN 3404 / 3413, which may also include other base stations and / or network components (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, and the like. Base station 3414a and / or base station 3414b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of both. A cell may provide wireless service coverage for a specific geographic area that may be relatively fixed or potentially change over time. A cell may be further divided into cell sectors. For example, the cell associated with base station 3414a may be divided into three sectors. Thus, in one embodiment, base station 3414a may include three transceivers, one for each sector of the cell. In an embodiment, base station 3414a may utilize multiple-input, multiple-output (MIMO) technology and may use multiple transceivers for each sector of the cell. For example, by utilizing beamforming, signals may be transmitted and / or received in a desired spatial direction.
[0206] The base stations 3414a and 3414b may communicate with one or more of the WTRUs 3402a, 3402b, 3402c, and 3402d over an air interface 3416, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 3416 may be established using any suitable radio access technology (RAT).
[0207] More specifically, as described above, the communication system 3400 may be a multiple access system and may utilize one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 3414a in the RAN 3404 / 3413 and the WTRUs 3402a, 3402b, and 3402c may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may utilize Wideband CDMA (WCDMA) to establish the air interface 3415 / 3416 / 3417. WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink (DL) Packet Access (HSDPA) and / or High Speed UL Packet Access (HSUPA).
[0208] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which may establish the air interface 3416 using Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) and / or Advanced LTE Pro (LTE-APro).
[0209] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement a radio technology that may establish the air interface 3416 using New Radio (NR), such as NR radio access.
[0210] In an embodiment, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement multiple radio access technologies. For example, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may jointly implement LTE radio access and NR radio access (e.g., using dual connectivity (DC) principles). Thus, the air interface used by the WTRUs 3402a, 3402b, 3402c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., terminals and gNBs).
[0211] In other embodiments, the base station 3414a and the WTRUs 3402a, 3402b, 3402c may implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN), among others.
[0212] Figure 34AThe base station 3414b in the may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point, and may use any appropriate RAT to facilitate wireless connectivity in a local area, such as a business location, a residence, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a road, and the like. In one embodiment, the base station 3414b and the WTRUs 3402c, 3402d may establish a wireless local area network (WLAN) by implementing a radio technology such as IEEE 802.11. In an embodiment, the base station 3414b and the WTRUs 3402c, 3402d may establish a wireless personal area network (WPAN) by implementing a radio technology such as IEEE 802.15. In another embodiment, the base station 3414b and the WTRUs 3402c, 3402d may establish a microcell or a femtocell by using a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, and the like). As Figure 34A As shown, base station 3414b can be directly connected to the Internet 3410. Therefore, base station 3414b does not need to access the Internet 3410 via CN 3406 / 3415.
[0213] The RAN 3404 / 3413 may communicate with the CN 3406 / 3415, which may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 3402a, 3402b, 3402c, 3402d. The data may have different quality of service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and the like. The CN 3406 / 3415 may provide call control, billing services, mobile location-based services, prepaid calls, Internet connectivity, video distribution, and the like, and / or may perform advanced security functions such as user authentication. Although in Figure 34A Although not shown, it will be appreciated that the RAN 1084 / 3413 and / or the CN 3406 / 3415 may be in direct or indirect communication with other RANs that employ the same RAT or a different RAT as the RAN 3404 / 3413. For example, in addition to being connected to the RAN 3404 / 3413 employing NR radio technology, the CN 3406 / 3415 may also be in communication with another RAN (not shown) employing GSM, UMTS, CDMA 2000, WiMAX, E-UTRA, or WiFi radio technology.
[0214] The CN 3406 / 3415 may also act as a gateway for the WTRUs 3402a, 3402b, 3402c, 3402d to access the PSTN 3408, the Internet 3410, and / or other networks 3412. The PSTN 3408 may include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 3410 may include a global system of interconnected computer network devices that uses common communication protocols, such as TCP, User Datagram Protocol (UDP), and / or IP from the Transmission Control Protocol / Internet Protocol (TCP / IP) suite of Internet protocols. The networks 3412 may include wired or wireless communication networks owned and / or operated by other service providers. For example, the networks 3412 may include another CN connected to one or more RANs, which may use the same RAT as the RAN 3404 / 3413 or a different RAT.
[0215] Some or all of the WTRUs 3402a, 3402b, 3402c, 3402d in the communication system 3400 may include multi-mode capabilities (e.g., the WTRUs 3402a, 3402b, 3402c, 3402d may include multiple transceivers for communicating with different wireless networks over different wireless links). Figure 34A The WTRU 3402c shown may be configured to communicate with the base station 3414a, which may employ a cellular-based radio technology, and with the base station 3414b, which may employ an IEEE 802 radio technology.
[0216] Figure 34B is a system diagram illustrating an exemplary WTRU 3402. Figure 34B As shown, the WTRU 3402 may include a processor 3418, a transceiver 3420, a transmit / receive element 3422, a speaker / microphone 3424, a keypad 3426, a display / touchpad 3428, non-removable memory 3430, removable memory 3432, a power supply 3434, a global positioning system (GPS) chipset 3436, and / or peripherals 3438. It will be appreciated that the WTRU 3402 may include any sub-combination of the foregoing components while remaining consistent with an embodiment.
[0217] The processor 3418 may be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, and the like. The processor 3418 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 3402 to operate in a wireless environment. The processor 3418 may be coupled to the transceiver 3420, which may be coupled to the transmit / receive element 3422. Although Figure 34B The processor 3418 and the transceiver 3420 are depicted as separate components, however, it should be understood that the processor 3418 and the transceiver 3420 may also be integrated together in an electronic component or chip. The processor 3418 may be configured to encode or decode video (eg, video frames).
[0218] The transmit / receive element 3422 may be configured to transmit or receive signals to or from a base station (e.g., base station 3414a) via the air interface 3416. For example, in one embodiment, the transmit / receive element 3422 may be an antenna configured to transmit and / or receive RF signals. As another example, in another embodiment, the transmit / receive element 3422 may be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 3422 may be configured to transmit and / or receive RF and light signals. It should be understood that the transmit / receive element 3422 may be configured to transmit and / or receive any combination of wireless signals.
[0219] Although Figure 34B 34. Although the transmit / receive element 3422 is depicted as a single element in the WTRU 3402, the WTRU 3402 may include any number of transmit / receive elements 3422. More specifically, the WTRU 3402 may employ MIMO technology. Thus, in one embodiment, the WTRU 3402 may include two or more transmit / receive elements 3422 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 3416.
[0220] The transceiver 3420 may be configured to modulate signals to be transmitted by the transmit / receive element 3422 and to demodulate signals received by the transmit / receive element 3422. As described above, the WTRU 3402 may have multi-mode capabilities. Thus, the transceiver 3420 may include multiple transceivers that allow the WTRU 3402 to communicate via multiple RATs, such as NR and IEEE 802.11.
[0221] The processor 3418 of the WTRU 3402 may be coupled to a speaker / microphone 3424, a numeric keypad 3426, and / or a display / touchpad 3428 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit) and may receive user input data from these components. The processor 3418 may also output user data to the speaker / microphone 3424, the keypad 3426, and / or the display / touchpad 3428. Furthermore, the processor 3418 may access information from, and store information in, any suitable memory, such as non-removable memory 3430 and / or removable memory 3432. The non-removable memory 3430 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 3432 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 3418 may access information from, and store data in, memory that is not physically located on the WTRU 3402, such as on a server or a home computer (not shown).
[0222] The processor 3418 may receive power from the power source 3434 and may be configured to distribute and / or control power to the other components in the WTRU 3402. The power source 3434 may be any suitable device for powering the WTRU 3402. For example, the power source 3434 may include one or more dry cell batteries (e.g., nickel-cadmium (Ni-Cd), nickel-zinc (Ni-Zn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.
[0223] The processor 3418 may also be coupled to the GPS chipset 3436, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 3402. In addition to or in lieu of the information from the GPS chipset 3436, the WTRU 3402 may receive location information from a base station (e.g., base stations 3414a, 3414b) via the air interface 3416 and / or determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 3402 may acquire location information using any suitable positioning method while remaining consistent with the embodiments.
[0224] The processor 3418 may also be coupled to other peripheral devices 3438, wherein the peripheral devices may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 3438 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, Modules, frequency modulation (FM) radio units, digital music players, media players, video game console modules, Internet browsers, virtual reality and / or augmented reality (VR / AR) devices, and activity trackers, etc. The peripheral devices 3438 may include one or more sensors, which may be one or more of the following: gyroscopes, accelerometers, Hall effect sensors, magnetometers, orientation sensors, proximity sensors, temperature sensors, time sensors, geolocation sensors, altimeters, light sensors, touch sensors, magnetometers, barometers, gesture sensors, biometric sensors, and / or humidity sensors, etc.
[0225] The processor 3418 of the WTRU 3402 can be operatively in communication with various peripheral devices 3438, including, for example, any of: the one or more accelerometers, the one or more gyroscopes, the USB port, other communication interfaces / ports, the display and / or other video / audio indicators, to implement the representative embodiments disclosed herein.
[0226] The WTRU 3402 may include a full-duplex radio in which the reception or transmission of some or all signals (e.g., associated with specific subframes used for UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. A full-duplex radio may include an interference management unit that reduces and / or substantially eliminates self-interference by means of hardware (e.g., chokes) or by signal processing by a processor (e.g., a separate processor (not shown) or by the processor 3418). In an embodiment, the WTRU 3402 may include a half-duplex radio that transmits and receives some or all signals (e.g., associated with specific subframes used for UL (e.g., for transmission) or downlink (e.g., for reception)).
[0227] Figure 34C 3406. As described above, the RAN 3404 may communicate with the WTRUs 3402a, 3402b, and 3402c using an E-UTRA radio technology over the air interface 3416. The RAN 3404 may also communicate with the CN 3406.
[0228] The RAN 3404 may include eNode-Bs 3460a, 3460b, and 3460c, though it will be appreciated that the RAN 3404 may include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 3460a, 3460b, and 3460c may each include one or more transceivers for communicating with the WTRUs 3402a, 3402b, and 3402c over the air interface 3416. In one embodiment, the eNode-Bs 3460a, 3460b, and 3460c may implement MIMO technology. Thus, for example, the eNode-B 3460a may use multiple antennas to transmit wireless signals to and / or receive wireless signals from the WTRU 3402a.
[0229] Each of the eNodeBs 3460a, 3460b, 3460c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in the UL and / or DL, and the like. Figure 34C As shown, the eNode-Bs 3460a, 3460b, and 3460c may communicate with each other via an X2 interface.
[0230] Figure 34C The illustrated CN 3406 may include a mobility management entity (MME) 3462, a serving gateway (SGW) 3464, and a packet data network (PDN) gateway (or PGW) 3466. While each of the aforementioned components is depicted as being part of the CN 3406, it should be understood that any of these components may be owned and / or operated by an entity other than the CN operator.
[0231] The MME 3462 may be connected to each of the eNode-Bs 3460a, 3460b, 3460c in the RAN 3404 via an S1 interface and may serve as a control node. For example, the MME 3462 may be responsible for authenticating users of the WTRUs 3402a, 3402b, 3402c, performing bearer activation / deactivation processing, selecting a specific serving gateway during an initial attach of the WTRUs 3402a, 3402b, 3402c, and the like. The MME 3462 may also provide control plane functions for switching between the RAN 3404 and other RANs (not shown) employing other radio technologies, such as GSM and / or WCDMA.
[0232] The SGW 3464 may be connected to each of the eNode-Bs 3460a, 3460b, 3460c in the RAN 3404 via an S1 interface. The SGW 3464 may generally route and forward user data packets to and from the WTRUs 3402a, 3402b, 3402c. The SGW 3464 may also perform other functions, such as anchoring the user plane during inter-eNB handovers, triggering paging when downlink data is available for the WTRUs 3402a, 3402b, 3402c, and managing and storing the context of the WTRUs 3402a, 3402b, 3402c.
[0233] The SGW 3464 may be connected to the PGW 146, which may provide the WTRUs 3402a, 3402b, 3402c with access to packet-switched networks, such as the Internet 3410, to facilitate communications between the WTRUs 3402a, 3402b, 3402c and IP-enabled devices.
[0234] The CN 3406 may facilitate communications with other networks. For example, the CN 3406 may provide the WTRUs 3402a, 3402b, 3402c with access to circuit-switched networks, such as the PSTN 3408, to facilitate communications between the WTRUs 3402a, 3402b, 3402c and traditional landline communications devices. For example, the CN 3406 may include or communicate with an IP gateway, such as an IP Multimedia Subsystem (IMS) server, which may serve as an interface between the CN 3406 and the PSTN 3408. The CN 3406 may also provide the WTRUs 3402a, 3402b, 3402c with access to other networks 3412, which may include other wired and / or wireless networks owned and / or operated by other service providers.
[0235] Although Figures 34A-34D While the WTRU is described as a wireless terminal, it is appreciated that in certain representative embodiments, such a terminal may utilize a (eg, temporary or permanent) wired communication interface with a communication network.
[0236] In a representative embodiment, the other network 3412 may be a WLAN.
[0237] A WLAN employing an infrastructure basic service set (BSS) model may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may access or interface with a distributed system (DS) or other type of wired / wireless network that routes traffic into and / or out of the BSS. Traffic originating from outside the BSS and destined for a STA may be routed through the AP and delivered to the STA. Traffic originating from a STA and destined for a destination outside the BSS may be sent to the AP for delivery to the destination. Traffic between STAs within the BSS may be routed through the AP, for example, where a source STA may send traffic to the AP and the AP may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or referred to as point-to-point traffic. This point-to-point traffic may be routed between (e.g., directly between) the source and destination STAs using a direct link setup (DLS). In certain representative embodiments, the DLS may utilize 802.11e DLS or 802.11z channelized DLS (TDLS). For example, a WLAN using an independent BSS (IBSS) mode does not have an AP, and STAs (e.g., all STAs) within or using the IBSS can communicate directly with each other. The IBSS communication mode may sometimes be referred to as an "ad-hoc" communication mode.
[0238] When using 802.11ac infrastructure mode or a similar mode of operation, the AP may transmit beacons on a fixed channel (e.g., a primary channel). The primary channel may have a fixed width (e.g., a 20 MHz bandwidth) or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In certain representative embodiments, carrier sense multiple access with collision avoidance (CSMA / CA) (e.g., in an 802.11 system) may be implemented. For CSMA / CA, STAs (e.g., each STA), including the AP, may sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, the particular STA may back off. In a given BSS, at any given time, one STA (e.g., only one station) may transmit.
[0239] High throughput (HT) STAs may communicate using a 40 MHz wide channel (eg, by combining a 20 MHz wide primary channel with adjacent or non-adjacent 20 MHz wide channels to form a 40 MHz wide channel).
[0240] Very High Throughput (VHT) STAs can support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels can be formed by combining contiguous 20 MHz channels. A 160 MHz channel can be formed by combining eight contiguous 20 MHz channels or by combining two discontinuous 80 MHz channels (this combination may be referred to as an 80+80 configuration). For the 80+80 configuration, after channel coding, the data can be passed and passed through a segment parser that can split the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time domain processing can be performed separately on each stream. The streams can be mapped onto two 80 MHz channels, and the data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above operations for the 80+80 configuration can be reversed, and the combined data can be sent to the medium access control (MAC).
[0241] 802.11af and 802.11ah support operating modes below 1 GHz. Compared to 802.11n and 802.11ac, the channel operating bandwidth and carrier used in 802.11af and 802.11ah are reduced. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TV White Space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative embodiment, 802.11ah can support meter type control / machine type communication (MTC) (e.g., MTC devices in macro coverage areas). MTC devices can have certain capabilities, such as limited capabilities including support for (e.g., only support for) certain and / or limited bandwidths. MTC devices can include a battery, and the battery life of the battery is above a threshold (e.g., for maintaining a very long battery life).
[0242] For WLAN systems that support multiple channels and channel bandwidths (e.g., 802.11n, 802.11ac, 802.11af, and 802.11ah), these systems include a channel that can be designated as a primary channel. The bandwidth of the primary channel can be equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by a STA, where the STA is derived from all STAs operating in the BSS that supports the minimum bandwidth operating mode. In the example of 802.11ah, even if the AP and other STAs in the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operating modes, for STAs that support (e.g., only) 1MHz mode (e.g., MTC-type devices), the primary channel width can be 1MHz. Carrier sensing and / or network allocation vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy (e.g., because a STA (which only supports 1MHz operation) is transmitting to the AP), the entire available frequency band can be considered busy, even if most of the available frequency band remains idle and available for use.
[0243] In the United States, the available frequency band for 802.11ah is 902MHz to 928MHz. In South Korea, the available frequency band is 917.5MHz to 923.5MHz. In Japan, the available frequency band is 916.5MHz to 927.5MHz. Depending on the country code, the total bandwidth available for 802.11ah ranges from 6MHz to 26MHz.
[0244] Figure 34D 3415 according to an embodiment. As described above, the RAN 3413 may communicate with the WTRUs 3402a, 3402b, 3402c using NR radio technology over the air interface 3416. The RAN 3413 may also communicate with the CN 3415.
[0245] The RAN 3413 may include gNBs 3480a, 3480b, and 3480c, though it will be appreciated that the RAN 3413 may include any number of gNBs while remaining consistent with an embodiment. Each of the gNBs 3480a, 3480b, and 3480c may include one or more transceivers for communicating with the WTRUs 3402a, 3402b, and 3402c over the air interface 3416. In one embodiment, the gNBs 3480a, 3480b, and 3480c may implement MIMO technology. For example, the gNBs 3480a and 3480b may utilize beamforming to transmit and / or receive signals to and / or from the gNBs 3480a, 3480b, and 3480c. Thus, for example, the gNB 3480a may utilize multiple antennas to transmit wireless signals to and receive wireless signals from the WTRU 3402a. In an embodiment, gNBs 3480a, 3480b, and 3480c may implement carrier aggregation techniques. For example, gNB 3480a may transmit multiple component carriers to WTRU 3402a (not shown). A subset of these component carriers may be in unlicensed spectrum, while the remaining component carriers may be in licensed spectrum. In an embodiment, gNBs 3480a, 3480b, and 3480c may implement coordinated multi-point (CoMP) techniques. For example, WTRU 3402a may receive coordinated transmissions from gNB 3480a and gNB 3480b (and / or gNB 3480c).
[0246] The WTRUs 3402a, 3402b, 3402c may communicate with the gNBs 3480a, 3480b, 3480c using transmissions associated with a scalable digital configuration. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may be different for different transmissions, different cells, and / or different portions of the radio transmission spectrum. The WTRUs 3402a, 3402b, 3402c may communicate with the gNBs 3480a, 3480b, 3480c using subframes or transmission time intervals (TTIs) of different or scalable lengths (e.g., containing different numbers of OFDM symbols and / or lasting different absolute time lengths).
[0247] The gNBs 3480a, 3480b, 3480c may be configured to communicate with the WTRUs 3402a, 3402b, 3402c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 3402a, 3402b, 3402c may communicate with the gNBs 3480a, 3480b, 3480c without accessing other RANs (e.g., eNode-Bs 3460a, 3460b, 3460c). In a standalone configuration, the WTRUs 3402a, 3402b, 3402c may use one or more of the gNBs 3480a, 3480b, 3480c as mobility anchors. In a standalone configuration, the WTRUs 3402a, 3402b, 3402c may communicate with the gNBs 3480a, 3480b, 3480c using signals in an unlicensed band. In a non-standalone configuration, the WTRUs 3402a, 3402b, 3402c may communicate with the gNBs 3480a, 3480b, 3480c while communicating with another RAN (e.g., an eNode-B 3460a, 3460b, 3460c). For example, the WTRUs 3402a, 3402b, 3402c may communicate with one or more gNBs 3480a, 3480b, 3480c and one or more eNode-Bs 3460a, 3460b, 3460c substantially simultaneously by implementing the DC principle. In a non-standalone configuration, the eNode-B 3460a, 3460b, 3460c may act as a mobility anchor for the WTRUs 3402a, 3402b, 3402c and the gNBs 3480a, 3480b, 3480c may provide additional coverage and / or throughput to serve the WTRUs 3402a, 3402b, 3402c.
[0248] Each gNB 3480a, 3480b, 3480c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support network slicing, dual connectivity, implement interworking between NR and E-UTRA, route user plane data to user plane functions (UPFs) 3484a, 3484b, and route control plane information to access and mobility management functions (AMFs) 3482a, 3482b, etc. Figure 34D As shown, gNBs 3480a, 3480b, and 3480c can communicate with each other through the Xn interface.
[0249] Figure 34DThe illustrated CN 3415 may include at least one AMF 3482a, 3482b, at least one UPF 3484a, 3484b, at least one Session Management Function (SMF) 3483a, 3483b, and may include Data Networks (DNs) 3485a, 3485b. While each of the aforementioned components is described as part of the CN 3415, it should be understood that any of these components may be owned and / or operated by an entity other than the CN operator.
[0250] The AMF 3482a, 3482b may be connected to one or more of the gNBs 3480a, 3480b, 3480c in the RAN 3413 via the N2 interface and may act as a control node. For example, the AMF 3482a, 3482b may be responsible for authenticating users of the WTRUs 3402a, 3402b, 3402c, supporting network slicing (e.g., handling different protocol data unit (PDU) sessions with different requirements), selecting a specific SMF 3483a, 3483b, managing registration areas, terminating non-access stratum (NAS) signaling, and mobility management, among others. The AMF 3482a, 3482b may use network slicing processing to customize the CN support provided to the WTRUs 3402a, 3402b, 3402c based on the service type used by the WTRUs 3402a, 3402b, 3402c. As an example, different network slices may be established for different use cases, such as services relying on Ultra-Reliable Low Latency Communication (URLLC) access, services relying on Enhanced Mobile (e.g., Massive Mobile) Broadband (eMBB) access, and / or services for Machine Type Communication (MTC) access, etc. The AMF 3462 may provide a control plane function for switching between the RAN 3413 and other RANs (not shown) using other radio technologies (e.g., LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies such as WiFi).
[0251] The SMF 3483a, 3483b can be connected to the AMF 3482a, 3482b in the CN 3415 via the N11 interface. The SMF 3483a, 3483b can also be connected to the UPF 3484a, 3484b in the CN 3415 via the N4 interface. The SMF 3483a, 3483b can select and control the UPF 3484a, 3484b, and can configure traffic routing through the UPF 3484a, 3484b. The SMF 3483a, 3483b can perform other functions, such as managing and allocating UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications, etc. The PDU session type can be IP-based, non-IP-based, Ethernet-based, etc.
[0252] The UPF 3484a, 3484b can be connected to one or more of the gNBs 3480a, 3480b, 3480c in the RAN 3413 via the N3 interface, which can provide the WTRU 3402a, 3402b, 3402c with access to packet-switched networks (such as the Internet 3410) to facilitate communication between the WTRU 3402a, 3402b, 3402c and IP-enabled devices. The UPF 3484, 3484b can perform other functions such as routing and forwarding packets, implementing user plane policies, supporting multi-host PDU sessions, processing user plane QoS, buffering downlink packets, and providing mobility anchoring processing, etc.
[0253] The CN 3415 may facilitate communications with other networks. For example, the CN 3415 may include or may communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 3415 and the PSTN 3408. Furthermore, the CN 3415 may provide the WTRUs 3402a, 3402b, 3402c with access to other networks 3412, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 3402a, 3402b, 3402c may connect to the local data networks (DNs) 3485a, 3485b through the UPFs 3484a, 3484b via the N3 interface to the UPFs 3484a, 3484b and the N6 interface between the UPFs 3484a, 3484b and the DNs 3485a, 3485b.
[0254] In view of Figures 34A-34D and about Figures 34A-34D
[00105] As described herein, one or more or all of the functions described herein with respect to one or more of the following may be performed by one or more emulated devices (not shown): WTRU 3402a-d, base station 3414a-b, eNodeB 3460a-c, MME 3462, SGW 3464, PGW 3466, gNB 3480a-c, AMF 3482a-b, UPF 3484a-b, SMF 3483a-b, DN 3485a-b, and / or one or more other devices described herein. These emulated devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, these emulated devices may be used to test other devices and / or emulate network and / or WTRU functions.
[0255] The simulation device can be designed to perform one or more tests on other devices in a laboratory environment and / or an operator network environment. For example, the one or more simulation devices can perform one or more or all functions while being implemented and / or deployed in whole or in part as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more simulation devices can perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The simulation device can be directly coupled to other devices to perform the test, and / or can use over-the-air wireless communication to perform the test.
[0256] One or more emulation devices can perform one or more functions, including all functions, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation device can be used in a test lab and / or a test scenario of a wired and / or wireless communication network that has not been deployed (e.g., tested) to perform tests on one or more components. The one or more emulation devices can be test devices. The emulation device can transmit and / or receive data using direct RF coupling and / or wireless communication with the aid of RF circuitry (e.g., which can include one or more antennas).
[0257] Compared to the previous generation video coding standard H.264 / MPEG AVC, the HEVC standard provides approximately 50% bitrate savings for equivalent perceptual quality. Although the HEVC standard provides significant coding improvements compared to its predecessor, additional coding efficiency improvements can be achieved with additional coding tools. The Joint Video Exploration Team (JVET) initiated a project to develop a new generation video coding standard, called Versatile Video Coding (VVC), for example to provide such coding efficiency improvements, and a reference software code base called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard. To facilitate the evaluation of new coding tools, another reference software library called the Benchmark Set (BMS) was also generated. In the BMS code base, a list of additional coding tools that provide higher coding efficiency and moderate implementation complexity is included on top of the VTM and is used as a benchmark when evaluating similar coding techniques during the VVC standardization process. In addition to the JEM coding tools integrated in BMS-2.0 (e.g., 4×4 Non-Separable Secondary Transform (NSST), Generalized Bi-Prediction (GBi), Bidirectional Optical Flow (BIO), Decoder-Side Motion Vector Refinement (DMVR), and Current Picture Reference (CPR)), it also includes a trellis coded quantization tool.
[0258] The system and method for processing data according to a representative embodiment can be executed by one or more processors that execute a sequence of instructions contained in a storage device. These instructions can be read into the storage device from other computer-readable media such as an auxiliary data storage device (one or more). The execution of the sequence of instructions contained in the storage device causes the processor to operate, for example, as described above. In an alternative embodiment, hard-wired circuits can be used instead of software instructions or in combination with software instructions to implement one or more embodiments. Such software can be run on a processor that is remotely housed in a robotic assistance / apparatus (RAA) and / or another mobile device. In the latter case, data can be transmitted between the RAA or other mobile device containing the sensor and a remote device containing the processor via a wired or wireless manner, and the processor runs the software that performs the scale estimation and compensation as described above. According to other representative embodiments, some of the processing described above with respect to positioning can be performed in the device containing the sensor / camera, while the remaining processing can be performed in a second device after receiving partially processed data from the device containing the sensor / camera.
[0259] Although features and elements are described above in specific combinations, it will be understood by those skilled in the art that each feature or element may be used alone or in any combination with the other features and elements. In addition, the methods described herein may be implemented in a computer program, software, or firmware embedded in a computer-readable medium and executed by a computer or processor. Examples of non-transitory computer-readable media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROMs and digital versatile discs (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for use in a WTRU 3402, UE, terminal, base station, RNC, or any host computer.
[0260] In addition, in the above-mentioned embodiments, processing platforms, computing systems, controllers and other devices containing processors are mentioned. These devices may include at least one central processing unit ("CPU") and memory. According to the practice of those skilled in the art of computer programming, references to symbolic descriptions of actions and operations or instructions can be performed by various CPUs and memories. These actions and operations or instructions can be referred to as being "executed," "computer-executed," or "CPU-executed."
[0261] Those skilled in the art will appreciate that the actions and symbols depicting operations or instructions include manipulation of electrical signals by a CPU. The electrical system representation can identify data bits that cause the electrical signals to be transformed or restored, and the maintenance of the storage locations of the data bits in the memory system, thereby reconfiguring or otherwise changing the operation of the CPU and other processing of the signals. Maintaining the storage locations of the data bits has specific electrical, magnetic, optical, or organic properties corresponding to or representing the data bits. It should be understood that representative embodiments are not limited to the aforementioned platforms or CPUs and that other platforms and CPUs can support the provided methods.
[0262] The data bits may also be maintained on a computer-readable medium, including magnetic disks, optical disks, and any other volatile (e.g., random access memory ("RAM")) or non-volatile (e.g., read-only memory ("ROM")) CPU-readable mass storage system. The computer-readable medium may include cooperating or interconnected computer-readable media that reside exclusively on a processor system or distributed across multiple interconnected processing systems that may be local or remote to the processing system. It will be understood that the representative embodiments are not limited to the aforementioned memories and that other platforms and memories may support the described methods. It will be understood that the representative embodiments are not limited to the aforementioned platforms or CPUs and that other platforms and CPUs may also support the provided methods.
[0263] In the illustrated embodiment, any of the operations, processes, etc. described herein may be implemented as computer-readable instructions stored on a computer-readable medium. The computer-readable instructions may be executed by a processor of a mobile unit, a network element, and / or any other computing device.
[0264] There is a slight distinction between hardware and software implementations of systems. The use of hardware or software is generally (but not always, as the choice between hardware and software can be important in certain environments) a design choice that considers a trade-off between cost and efficiency. There can be various tools (e.g., hardware, software, and / or firmware) that affect the processes and / or systems and / or other technologies described herein, and the preferred tools can change with the context of the deployed processes and / or systems and / or other technologies. For example, if the implementer determines that speed and accuracy are most important, the implementer may choose to use primarily hardware and / or firmware tools. If flexibility is most important, the implementer may choose to use primarily software implementation. Alternatively, the implementer may choose some combination of hardware, software, and / or firmware.
[0265] The above detailed description has presented various embodiments of the apparatus and / or process through the use of block diagrams, flow charts, and / or examples. To the extent that these block diagrams, flow charts, and / or examples include one or more functions and / or operations, those skilled in the art will appreciate that each function and / or operation within these block diagrams, flow charts, or examples can be implemented individually and / or together using a wide range of hardware, software, or firmware, or substantially any combination thereof. Suitable processors include, for example, general-purpose processors, special-purpose processors, conventional processors, digital signal processors (DSPs), multiple microprocessors, one or more microprocessors associated with a DSP core, controllers, microcontrollers, application-specific integrated circuits (ASICs), application-specific standard products (ASSPs); field-programmable gate array (FPGA) circuits, any other type of integrated circuit (IC), and / or state machines.
[0266] Although features and elements are provided above in specific combinations, it will be understood by those skilled in the art that each feature or element can be used alone or in combination with other features and elements. The present disclosure is not limited to the specific embodiments described in this application, which are intended to be examples of various aspects. Many modifications and variations can be made without departing from its essence and scope, which are known to those skilled in the art. The elements, actions or instructions used in the description of this application should not be understood as being critical or necessary to the embodiments unless explicitly stated. In addition to the methods and devices listed herein, those skilled in the art can also know functionally equivalent methods and devices within the scope of this disclosure based on the above description. These modifications and variations should also fall within the scope of the appended claims. The present disclosure is limited only by the appended claims, including the full scope of their equivalents. It should be understood that the present disclosure is not limited to a specific method or system.
[0267] It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting. As used herein, when referring to the term "station" and its abbreviation "STA", "user equipment" and its abbreviation "UE" herein, it can mean: (i) a wireless transmit and / or receive unit (WTRU), such as described below; (ii) any of a plurality of embodiments of a WTRU, such as described below; (iii) a wireless and / or wired (e.g., wirelessly communicable) device configured with some or all of the structure and functionality of a WTRU, such as described below; (iii) a device having wireless capabilities and / or wired capabilities that is configured with less than all of the structure and functionality of a WTRU, such as described below; or (iv) the like. Reference is made below to Figures 34A-34D Details are provided for an example WTRU that may represent any of the UEs described herein.
[0268] In certain representative embodiments, some parts of the subject matter described herein can be implemented via application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), and / or other integrated formats. However, those skilled in the art will appreciate that some aspects of the embodiments disclosed herein, in whole or in part, can equally be implemented by integrated circuits as one or more computer programs running on one or more computers (e.g., one or more programs running on one or more computer systems), one or more programs running on one or more processors (e.g., one or more programs running on one or more microprocessors), firmware, or substantially any combination of these, and that designing circuits and / or writing code for the software and / or firmware according to the present disclosure is known to those skilled in the art. In addition, those skilled in the art will appreciate that the mechanisms of the subject matter described herein can be distributed as program products in various forms, and that the exemplary embodiments of the subject matter described herein are applicable, regardless of the specific type of signal-bearing medium used to actually perform the distribution. Examples of signal-bearing media include, but are not limited to, the following: recordable media, such as floppy disks, hard disks, CDs, DVDs, digital tapes, computer memories, and the like, and transmission media, such as digital and / or analog communication media (e.g., optical cables, waveguides, wired communication links, wireless communication links, and the like).
[0269] The subject matter described herein sometimes shows different components, which are contained in or connected to different other components. It will be understood that the architectures described here are only examples, and many other architectures that implement the same function in practice can be implemented. Conceptually, any arrangement of components that implement the same function is effectively "associated" so that the desired function can be implemented. Therefore, any two components combined here to implement a specific function can be considered to be "associated" to each other so as to implement the desired function, regardless of the architecture or intermediate components. Similarly, any two components that are associated can also be considered to be "operationally connected" or "operationally coupled" to implement the desired function, and any two components that can be associated in this way can also be considered to be "operationally couplable" to each other to implement the desired function. Specific examples of operational coupling include but are not limited to physically pairable and / or physically interactive components and / or wirelessly interactive and / or wirelessly interactive components and / or logically interactive and / or logically interactive components.
[0270] With respect to the use of substantially any plural and / or singular terms herein, those skilled in the art can shift from the plural to the singular and / or from the singular to the plural as appropriate to the context and / or application. For the sake of clarity, various singular / plural permutations may be explicitly set forth herein.
[0271] Those skilled in the art will appreciate that the terms used herein generally, and particularly in the claims (e.g., the body of the claims), are generally "open" terms (e.g., the term "including" should be understood as "including but not limited to," the term "having" should be understood as "having at least," the term "comprising" should be understood as "including but not limited to," etc.). Those skilled in the art will also understand that if a claim is to recite a specific quantity, it will be explicitly stated in the claim, and in the absence of such a statement, no such meaning is intended. For example, if only one item is intended, the term "single" or similar language may be used. To aid understanding, the following claims and / or the description herein may include the use of the prepositional phrases "at least one" or "one or more" to introduce a claim recitation. However, the use of these phrases should not be construed to imply that a claim recitation introduced by the indefinite article "a" or "an" limits any particular claim containing such introduced claim recitation to embodiments containing only one such recitation, even if the same claim includes the prepositional phrases "one or more" or "at least one" and an indefinite article (e.g., "a") (e.g., "a" should be understood to mean "at least one" or "one or more"). The same is true for the use of definite articles to introduce claim recitations. Furthermore, even if a specific quantity of a claim description is explicitly described, a person skilled in the art will understand that such description should be understood to mean at least the quantity described (e.g., simply describing "two descriptions" without other modifiers means at least two descriptions, or two or more descriptions). Furthermore, in those examples where a convention similar to "at least one of A, B, and C, etc." is used, generally speaking, such a convention is a convention understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" may include, but is not limited to, a system having only A, only B, only C, A and B, A and C, B and C, and / or A, B and C, etc.). In those examples where a convention similar to "at least one of A, B, or C, etc." is used, generally speaking, such a convention is a convention understood by a person skilled in the art (e.g., "a system having at least one of A, B, or C" may include, but is not limited to, a system having only A, only B, only C, A and B, A and C, B and C, and / or A, B and C, etc.). Those skilled in the art will also understand that substantially any separated word and / or phrase representing two or more alternatives, whether in the specification, claims or drawings, should be understood to include the possibility of including one of the two items, any one or both items. For example, the phrase "A or B" is understood to include the possibility of "A" or "B" or "A" and "B". In addition, the term "any" as used herein followed by a plurality of items and / or a variety of items is intended to include "any," "any combination," "any plurality" and / or "any combination of a plurality" of the plurality of items and / or a variety of items, alone or in combination with other items and / or other items.Furthermore, the terms "set" or "group" as used herein are intended to include any number of items, including zero. Furthermore, the term "number" as used herein is intended to include any number, including zero.
[0272] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also described in terms of any individual member or subgroup of members of the Markush group.
[0273] Those skilled in the art will appreciate that, for any and all purposes, such as for providing a written description, all ranges disclosed herein also include any and all possible subranges and combinations of subranges thereof. Any listed range can be readily understood to be sufficient to describe and implement the same range divided into at least equal halves, thirds, fourths, fifths, tenths, etc. As non-limiting examples, each range described herein can be readily divided into a lower third, a middle third, and an upper third, etc. Those skilled in the art will also appreciate that all language such as "up to," "at least," "greater than," "less than," etc. includes the number being described and up to which the range can then be divided into the subranges described above. Finally, those skilled in the art will appreciate that a range includes each individual member. Thus, for example, a group and / or set having 1-3 cells refers to a group / set having 1, 2, or 3 cells. Similarly, a group / set having 1-5 cells refers to a group / set having 1, 2, 3, 4, or 5 cells, etc.
[0274] Furthermore, the claims should not be read as limited to the order or elements provided unless described to that effect. Furthermore, the use of the term "means for..." in any claim is intended to invoke 35 U.S.C. § 112, 6 or means-plus-function claim format, any claim without the term “means for…” does not have such an intention.
[0275] A processor in association with software may be used to implement a radio frequency transceiver for use in a wireless transmit / receive unit (WTRU), user equipment (UE), terminal, base station, mobility management entity (MME) or evolved packet core (EPC), or any host computer. The WTRU may incorporate modules implemented in hardware and / or software, including software defined radio (SDR), and other components such as a camera, a video camera module, a video phone, an intercom phone, a vibration device, a speaker, a microphone, a television transceiver, a hands-free headset, a keyboard, module, a frequency modulation (FM) radio unit, a near field communication (NFC) module, a liquid crystal display (LCD) display unit, an organic light emitting diode (OLED) display unit, a digital music player, a media player, a video game console module, an Internet browser, and / or any wireless local area network (WLAN) or ultra-wideband (UWB) module.
[0276] Throughout this disclosure, skilled artisans understand that certain representative embodiments may be used instead of or in combination with other representative embodiments.
[0277] Additionally, the methods described herein may be implemented in a computer program, software, or firmware incorporated into a computer-readable medium for execution by a computer or processor. Examples of non-transitory computer-readable media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROMs and digital versatile discs (DVDs). A processor associated with the software may be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method for decoding a video, comprising: generating, for a sub-block of a block of the picture, a sub-block based motion prediction signal based on an affine motion model associated with the block; determining a set of pixel-level motion vector differences for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block based motion prediction signal for each sample position of the sub-block, wherein an extended sub-block is formed to include the sub-block based motion prediction signal and a plurality of samples surrounding the sub-block, and wherein each of the plurality of samples surrounding the sub-block is obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector differences and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; and Decoding video using a refined motion prediction signal. 2 . The method of claim 1 , wherein the integer motion compensation is based on an integer portion of a motion vector for the sub-block. 3 . The method of claim 1 , wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the sub-block. The method according to claim 1 , wherein the sub-block is extended by one sample in each direction to form the extended sub-block.
5. A computer-readable storage medium having stored thereon instructions for decoding video data according to the method of claim 1.
6. A method for encoding a video, comprising: generating, for a sub-block of a block of the picture, a sub-block based motion prediction signal based on an affine motion model associated with the block; determining a set of pixel-level motion vector differences for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block based motion prediction signal for each sample position of the sub-block, wherein an extended sub-block is formed to include the sub-block based motion prediction signal and a plurality of samples surrounding the sub-block, and wherein each of the plurality of samples surrounding the sub-block is obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector differences and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; as well as Video is encoded using a refined motion prediction signal.
7. The method of claim 6, wherein the integer motion compensation is based on an integer portion of a motion vector for the sub-block.
8. The method of claim 6, wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the sub-block.
9. The method of claim 6, wherein the sub-block is extended by one sample in each direction to form the extended sub-block.
10. A computer-readable storage medium having stored thereon instructions for decoding video data according to the method of claim 6.
11. An apparatus for decoding a video, comprising a processor configured to: generating, for a sub-block of a block of the picture, a sub-block based motion prediction signal based on an affine motion model associated with the block; determining a set of pixel-level motion vector differences for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block based motion prediction signal for each sample position of the sub-block, wherein an extended sub-block is formed to include the sub-block based motion prediction signal and a plurality of samples surrounding the sub-block, and wherein each of the plurality of samples surrounding the sub-block is obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector differences and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; and Decoding video using a refined motion prediction signal.
12. The apparatus of claim 11, wherein the integer motion compensation is based on an integer portion of a motion vector for the sub-block.
13. The device of claim 11, wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the sub-block. The apparatus according to claim 11 , wherein the sub-block is extended by one sample in each direction to form the extended sub-block.
15. An apparatus for encoding a video, comprising a processor configured to: generating, for a sub-block of a block of the picture, a sub-block based motion prediction signal based on an affine motion model associated with the block; determining a set of pixel-level motion vector differences for the sub-block using the affine motion model associated with the block; determining a spatial gradient of the sub-block based motion prediction signal for each sample position of the sub-block, wherein an extended sub-block is formed to include the sub-block based motion prediction signal and a plurality of samples surrounding the sub-block, and wherein each of the plurality of samples surrounding the sub-block is obtained based on integer motion compensation; determining a motion prediction refinement signal for the sub-block based on the determined set of pixel-level motion vector differences and the determined spatial gradient; combining the motion prediction signal and the motion prediction refinement signal to generate a refined motion prediction signal for the sub-block; as well as Video is encoded using a refined motion prediction signal.
16. The apparatus of claim 15, wherein the integer motion compensation is based on an integer portion of a motion vector for the sub-block.
17. The device of claim 15, wherein the integer motion compensation is based on a nearest integer motion vector of the motion vector of the sub-block. The apparatus according to claim 15 , wherein the sub-block is extended by one sample in each direction to form the extended sub-block.