Motion-compensation prediction based on bi-directional optical flow
By determining the similarity of prediction signals within a video coding system, the device adaptively enables or disables bi-directional optical flow, enhancing coding efficiency and prediction accuracy.
Patent Information
- Application Number
- JP2025037401
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-12-15
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2038-07-03
AI Technical Summary
Current video coding systems face challenges in efficiently determining when to enable or disable bi-directional optical flow (BIO) for coding units, which affects prediction accuracy and coding efficiency.
A device configured to determine whether to enable or disable BIO for a current coding unit by calculating the prediction difference between two prediction signals and assessing their similarity, allowing for adaptive reconstruction with BIO enabled or disabled based on the similarity threshold.
This approach enhances coding efficiency by selectively applying BIO only when necessary, improving prediction accuracy and reducing computational complexity.
Smart Images

Figure 2025090700000001_ABST
Abstract
Description
Background Art
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 528,296, filed Jul. 3, 2017; U.S. Provisional Patent Application No. 62 / 560,823, filed Sep. 20, 2017; U.S. Provisional Patent Application No. 62 / 564,598, filed Sep. 28, 2017; U.S. Provisional Patent Application No. 62 / 579,559, filed Oct. 31, 2017; and U.S. Provisional Patent Application No. 62 / 599,241, filed Dec. 15, 2017, the contents of which are incorporated herein by reference.
[0002] Video coding systems are widely used to compress digital video signals to reduce the storage requirements and / or transmission bandwidth of such signals. There are various types of video coding systems, such as block-based, wavelet-based, and object-based systems. Currently, block-based hybrid video coding systems are widely used and / or deployed. Examples of block-based video coding systems include international video coding standards such as MPEG1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1, and the latest video coding standard called High Efficiency Video Coding (HEVC) developed by the Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG.
Summary of the Invention
Means for Solving the Problems
[0003] A device for performing video data coding may be configured to determine whether to enable or disable bi-directional optical flow (BIO) for a current coding unit (e.g., a block and / or a sub-block). Prediction information for the current coding unit may be identified. The prediction information may include a prediction signal associated with a first reference block (or, e.g., a sub-block) and a prediction signal associated with a second reference block (or, e.g., a sub-block). A prediction difference between the two prediction signals may be calculated. A similarity between the two prediction signals may be determined based on the prediction difference. The current coding unit may be reconstructed based on the similarity of the two prediction signals. For example, whether to reconstruct the current coding unit with BIO enabled or with BIO disabled may be determined based on whether the two prediction signals are sufficiently similar. When it is determined that the two prediction signals are not similar (e.g., dissimilar), it may be determined to enable BIO for the current coding unit. For example, when it is determined that the two prediction signals are similar, the current coding unit may be reconstructed with BIO disabled.
[0004] The prediction difference that may be used to determine the similarity between the two prediction signals may be determined in a plurality of ways. For example, calculating the prediction difference may include calculating an average difference between the sample values of each of the two reference blocks associated with the two prediction signals. The sample values from each reference block may be interpolated. For example, calculating the prediction difference may include calculating an average motion vector difference between the motion vectors of each of the two reference blocks associated with the two prediction signals. The motion vectors may be scaled based on the temporal distance between the reference picture and the current coding unit.
[0005] The similarity between two prediction signals may be determined by comparing the prediction difference between the two prediction signals with a threshold. When the prediction difference is less than or equal to the threshold, the two prediction signals may be determined to be similar. When the prediction difference is greater than the threshold, the two prediction signals may not be determined to be sufficiently similar (e.g., dissimilar). The threshold may be determined by and / or received at the video coding device. The threshold may be determined based on a desired complexity level and / or a desired coding efficiency.
[0006] A device for performing video data coding may be configured to group one or more sub-blocks into sub-block groups. For example, consecutive sub-blocks having similar motion information may be grouped together into a sub-block group. The sub-block group may vary in shape and size and may be formed based on the shape and / or size of the current coding unit. The sub-blocks may be grouped horizontally and / or vertically. A motion compensation operation (e.g., a single motion compensation operation) may be performed on the sub-block group. BIO refinement may be performed on the sub-block group. For example, the BIO refinement may be based on the gradient values of the sub-blocks of the sub-block group.
[0007] The BIO gradient may be derived such that single instruction multiple data (SIMD)-based acceleration can be utilized. In one or more techniques, the BIO gradient may be derived by applying an interpolation filter and a gradient filter, in which case horizontal filtering may be performed followed by vertical filtering. In the BIO gradient derivation, a rounding operation that may be performed by addition and right shift may be performed on the input value.
[0008] Devices, processes, and means for skipping BIO operations in a (e.g., normal) motion compensation (MC) stage (e.g., block level) of a video encoder and / or decoder are disclosed. In one or more techniques, for one or more blocks / sub-blocks for which one or more factors / conditions may be satisfied, the BIO operation may be (e.g., partially or fully) disabled. For blocks / sub-blocks coded in a frame rate up-conversion (FRUC) bilateral mode / by the FRUC bilateral mode, the BIO may be disabled. For blocks / sub-blocks predicted by at least two motion vectors that are approximately proportional in the temporal domain, the BIO may be disabled. When the average difference between at least two prediction blocks is below a predefined / pre-determined threshold, the BIO may be disabled. The BIO may be disabled based on gradient information.
[0009] A decoding device for video data coding may include a memory. The decoding device may include a processor. The processor may be configured to identify a plurality of sub-blocks of at least one coding unit (CU). The processor may be configured to select one or more of the plurality of sub-blocks for MC. The processor may be configured to determine whether the status of the MC condition is satisfied or not. When the status of the MC condition is satisfied, the processor may be configured to initiate motion compensation without BIO motion refinement processing for one or more sub-blocks. When the status of the MC condition is not satisfied, the processor may be configured to initiate motion compensation with BIO motion refinement processing for one or more sub-blocks.
[0010] Like reference numerals in the figures indicate like elements.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 9A
Figure 9B
Figure 10A
Figure 10B
Figure 11
Figure 12A
Figure 12B
Figure 13
Figure 14
Figure 15A
Figure 15B
Figure 16A
Figure 16B
Figure 16C
Figure 17A
Figure 17B
Figure 18A
Figure 18B
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28A
Figure 28B
Figure 28C
Figure 28D
DETAILED DESCRIPTION OF THE INVENTION
[0012] A detailed description of the illustrative embodiments will now be made with reference to the various figures. It should be noted that this description provides detailed examples of possible implementations, but the details are intended to be exemplary and in no way limit the scope of the present application.
[0013] FIG. 1 shows an exemplary block diagram of a block-based hybrid video encoding system. The input video signal 1102 may be processed block by block. To efficiently compress high-resolution (1080p and above) video signals, an extended block size (referred to as a “coding unit” or CU) may be used. The CU can be up to 64×64 pixels. The CU can be further partitioned into prediction units (PUs) to which different prediction methods are applied. Spatial prediction (1160) and / or temporal prediction (1162) may be performed on one or more or each input video block (MB or CU). Spatial prediction (or “intra prediction”) uses pixels from samples of already-coded neighboring blocks (e.g., called reference samples) within the same video picture / slice to predict the current video block. Spatial prediction can reduce the spatial redundancy inherent in the video signal. Temporal prediction (also called, e.g., “inter prediction” and / or “motion-compensated prediction”) may use reconstructed pixels from already-coded video pictures to predict the current video block. Temporal prediction can reduce the temporal redundancy inherent in the video signal. The temporal prediction signal for a given video block may be conveyed by one or more motion vectors indicating the amount and / or direction of motion between the current block and its reference block. Also, if multiple reference pictures are supported, one or more or each video block, its reference picture index may be transmitted. A reference index may be used to identify from which reference picture within the reference picture store (1164) the temporal prediction signal comes. After spatial and / or temporal prediction, a mode decision block (1180) in the encoder may select the best prediction mode, e.g., based on a rate-distortion optimization method.
[0014] The prediction block may be subtracted from the current video block (1116). The prediction residual may be decorrelated using transform (1104) and / or quantization (1106). The quantized residual coefficients may be inverse quantized (1110) and / or inverse transformed (1112) to form a reconstructed residual for addition to the prediction block to form a reconstructed video block (1126). Further, in-loop filtering such as deblocking filtering and / or adaptive loop filtering may be applied (1166) to the reconstructed video block, perhaps before it is stored in the reference picture store (1164) and / or before it is used to code future video blocks. Coding mode (e.g., inter and / or intra), prediction mode information, motion information, and / or the quantized residual coefficients may be further compressed and / or packed and sent to an entropy coding unit (1108) to form the output video bit stream 1120.
[0015] FIG. 2 shows an overall block diagram of a block-based video decoder. The video bit stream 202 may be unpacked and / or entropy decoded in an entropy decoding unit 208. Coding mode and / or prediction information may be sent to a spatial prediction unit 260 (e.g., if intra-coded) and / or a temporal prediction unit 262 (e.g., if inter-coded) to form a prediction block. Residual transform coefficients may be sent to an inverse quantization unit 210 and / or an inverse transform unit 212 to reconstruct a residual block. At 226, the prediction block and the residual block may be added together. The reconstructed block may pass through in-loop filtering, perhaps before it can be stored within the reference picture store 264. The reconstructed video within the reference picture store may be sent to drive a display device and / or may be used to predict future video blocks.
[0016] As shown in FIG. 1 and / or FIG. 2, spatial prediction (e.g., intra prediction), temporal prediction (e.g., inter prediction), transformation, quantization, entropy coding, and / or loop filtering may be performed or used. Bi-prediction in video coding may include a combination of two temporal prediction blocks obtained from a reference picture that may already be reconstructed using averaging. For example, due to the limitations of block-based motion compensation (MC), there may still be a small observable motion remaining between two prediction blocks. To compensate for such motion, bidirectional optical flow (BIO) may be applied to one or more or all samples within at least one block. BIO may be a motion refinement for samples that may be performed on block-based motion-compensated prediction when bi-prediction is used. The derivation of refined motion vectors for one or more or each sample within at least one block may be based on a classical optical flow model. I (k) Let (x, y) be the sample value at the coordinates (x, y) of a prediction block derived from the reference picture list k (k = 0, 1), and ∂I (k) (x, y) / ∂x and ∂I (k) (x, y) / ∂y may indicate the horizontal and vertical gradients of the sample. Assuming that the optical flow model is valid, the motion refinement (v x , v y ) at (x, y) can be
[0017] [Number]
[0018] derived by. Using a combination of the optical flow equation (1) and interpolation of the prediction block along a motion trajectory (as shown, for example, in FIG. 3), the BIO prediction is
[0019] [Number]
[0020] may be obtained as, where τ0 and τ1 are I (0) and I (1) may indicate the temporal distance to the current picture CurPic of the reference pictures Ref0 and Ref1 associated with, for example,
[0021] [Number]
[0022] is.
[0023] In FIG. 3, (MV x0 , MV y0 ) and (MV x1 , MV y1 ) may indicate block-level motion vectors that may be used to generate two prediction blocks I (0) and I (1) . Further, the motion refinement (v x , v y ) at the sample location (x, y) may be
[0024] [Number]
[0025] calculated by minimizing the difference Δ between the values of the samples after motion refinement compensation (e.g., A and B in FIG. 3) as shown as.
[0026] Perhaps, for example, to ensure the regularity of the derived motion refinement, it may be assumed that the motion refinement is consistent within a local surrounding area centered on (x, y). In an exemplary BIO design, the value of (v x , v y ) is
[0027]
Number
[0028] As shown in (x,y), it may be derived by minimizing Δ within the 5×5 window Ω around the current sample. BIO may be applied to the dual prediction block, which may be predicted by two reference blocks from temporally adjacent pictures. BIO may be enabled without transmitting additional information from the encoder to the decoder. BIO may be applied to some bi-directionally predicted blocks that have both forward and backward prediction signals (e.g., τ0·τ1>0). For example, if the two prediction blocks of the current block are from the same direction (either forward or backward, e.g., τ0·τ1<0), BIO is associated with non-zero motion, e.g., abs(MV x0 )+abs(MV y0 )≠0, and abs(MV x1 )+abs(MV y1 )≠0, and the two motion vectors are proportional to the temporal distance between the current picture and the reference picture, e.g.,
[0029]
Number
[0030] It may be applied when. For example, if the two prediction blocks of the current block are from the same reference picture (e.g., τ0=τ1), BIO may be disabled. When local illumination compensation (LIC) is used for the current block, BIO may be disabled.
[0031] (2) and (4) As shown, perhaps in addition to block-level MC, in BIO (e.g., to derive local motion refinement and / or generate the final prediction at its sample location), the motion-compensated block (e.g., I (0) and I(1) ) The gradient for the sample of may be derived. In BIO, the horizontal and vertical gradients of the samples within the prediction block (e.g.,
[0032]
Number
[0033] 、
[0034]
Number
[0035] 、and
[0036]
Number
[0037] 、
[0038]
Number
[0039] ) may be calculated simultaneously when the prediction signal is generated based on a filtering process (e.g., a 2D separable finite impulse response (FIR) filter) that may be consistent with motion compensation interpolation. The input to the gradient derivation process is the same reference samples used for motion compensation and the fractional components (fracX, fracY) of the input motion (MV x0 / x1 、MV y0 / y1 ).
[0040] To derive the gradient values at the sample positions (e.g., each sample position), different filters (e.g., one is the interpolation filter h L 、one is the gradient filter h G ) may be applied separately, perhaps in a different order for each direction of the gradient that may be calculated. The horizontal gradient (e.g.,
[0041] [Number]
[0042] and
[0043] [Number]
[0044] When deriving (), to derive the sample value at the vertical fractional position in fracY, the interpolation filter h L may be applied vertically to the samples inside the prediction block. Perhaps, based on the value of fracX, in order to calculate the horizontal gradient value, then the gradient filter h G may be applied horizontally to the generated vertical fractional samples. Vertical gradient (e.g.,
[0045] [Number]
[0046] and
[0047] [Number]
[0048] When deriving ()), to calculate the intermediate vertical gradient corresponding to fracY, the gradient filter h G may be applied vertically on the prediction samples, and the horizontal interpolation of the intermediate vertical gradient follows the value of fracX and uses the interpolation filter h L . The length of both the gradient filter and the interpolation filter may be 6 - tap or 8 - tap. Tables 1 and 2 show, respectively, according to the accuracy of the block - level motion vector, h G and h LIllustrative filter coefficients that may be used for
[0049] [Table 1]
[0050] [Table 2]
[0051] Figures 4A and 4B show an illustrative gradient derivation process applied in BIO. Sample values at integer sample positions are shown as patterned squares, and sample values at fractional sample positions are shown as blank squares. The motion vector accuracy may be increased to 1 / 16 pixel. In Figures 4A and 4B, 255 fractional sample positions defined within the region of integer samples may exist. The subscript coordinates (x, y) represent the corresponding horizontal and vertical fractional positions of the sample (e.g., the coordinate (0, 0) corresponds to the sample at the integer position). At the fractional position (1, 1) (e.g., a 1,1 ), horizontal and vertical gradient values may be calculated. According to Figures 4A and 4B, for horizontal gradient derivation, by applying the interpolation filter h L vertically, the fractional samples f 0,1 , e 0,1 , a 0,1 , b 0,1 , c 0,1 , and d 0,1 may be derived. For example,
[0052] [Equation]
[0053] where B is the bit depth of the input signal and OffSet0 is
[0054] [Equation]
[0055] which may be equal to, is a rounding offset. f 0,1 , e 0,1 , a 0,1 , b 0,1 , c 0,1 , and d 0,1 The accuracy of may be 14 bits. a 1,1 The horizontal gradient of is calculated by horizontally applying the corresponding gradient filter h G to the derived fractional samples. This is
[0056]
Number
[0057] As shown by, it may be done by calculating the unrounded gradient value in the middle 20 bits. The final horizontal gradient is
[0058]
Number
[0059] It may be calculated by shifting the intermediate gradient value to the output accuracy as, where sign(·) and abs(·) are functions that return the sign and absolute value of the input signal, and OffSet1 is 2 17-B which may be calculated as, is a rounding offset.
[0060] When deriving the vertical gradient value at (1,1), the intermediate vertical gradient value at the fractional position (0,1) may be derived, for example,
[0061]
Number
[0062] is. The intermediate vertical gradient value is then
[0063]
Number
[0064] It may be adjusted by shifting to a 14-bit value. The vertical gradient value at the fractional position (1, 1) may be obtained by the interpolation filter h on top of the intermediate gradient value at the fractional position (0, 1). This may be done by calculating the gradient value that is not rounded at 20 bits, which is then L performed. It may then be adjusted to the output bit depth through a shift operation, as shown as
[0065]
Number
[0066] as shown.
[0067] (5) As shown, perhaps, for deriving local motion refinement at one position, sample values and gradient values may be calculated for several samples within a surrounding window Ω around the sample. The window size may be (2M + 1)×(2M + 1), where M = 2. As described herein, gradient derivation may access additional reference samples within an extended area of the current block. Assuming that the length T of the interpolation filter and the gradient filter may be 6, the corresponding extended block size may be equal to T - 1 = 5. For a given W×H block, the memory access required by BIO may be (W + T - 1 + 2M)×(H + T - 1 + 2M) = (W + 9)×(H + 9), which may be larger than the memory access (W + 7)×(H + 7) used by motion compensation. Block expansion constraints may be used to control the memory access of BIO. For example, if block constraints are applied, perhaps, local motion refinement (v x , v yTo calculate [[ID=]], neighboring samples that may be within the current block may be used. FIGS. 5A and 5B compare the size of the memory access area for BIO before and after the block expansion constraint is applied.
[0068] In Advanced Temporal Motion Vector Prediction (ATMVP), temporal motion vector prediction may be improved by enabling the block to derive a plurality of motion information (e.g., motion vectors and reference indices) for sub-blocks within the block from a plurality of smaller blocks of temporally neighboring pictures of the current picture. ATMVP may derive motion information for sub-blocks within the block by identifying the corresponding block of the current block (which may be referred to as a collocated block) within the temporal reference picture. The selected temporal reference picture may be referred to as a collocated picture. The current block may be divided into sub-blocks, and the motion information for each sub-block may be derived from the corresponding smaller block within the collocated picture, as shown in FIG. 6.
[0069] Collocated blocks and collocated pictures may be identified by the motion information of spatially neighboring blocks of the current block. FIG. 6 shows a process in which available candidates within the merge candidate list are considered. It may be assumed that block A is identified as an available merge candidate for the current block based on the scanning order of the merge candidate list. The corresponding motion vector of block A (e.g., MV A ) and its reference index may be used to identify the collocated picture and the collocated block. The location of the collocated block within the collocated picture may be determined by adding the motion vector of block A (e.g., MV A ) to the coordinates of the current block.
[0070] For sub - blocks within a current block, in order to derive motion information of the sub - blocks within the current block, motion information of its corresponding small block within the collocated block (as indicated by the arrows in FIG. 6) may be used. When the motion information of each small block within the collocated block is identified, it may be converted into the motion vector and reference index of the corresponding sub - block within the current block (in the same way as, for example, temporal motion vector prediction (TMVP) where temporal motion vector scaling may be applied).
[0071] In spatio - temporal motion vector prediction (STVMP), motion information of sub - blocks within a coding block may be derived in a recursive manner, an example of which is shown in FIG. 7. FIG. 7 shows one example for illustrative purposes. As shown in FIG. 7, a current block may include four sub - blocks A, B, C, D. The neighboring small blocks that are spatial neighbors of the current block are labeled a, b, c, d respectively. For motion derivation for sub - block A, its two spatial neighbors may be identified. The first neighbor of sub - block A may be neighbor c. For example, if small block c is not available or not intra - coded, the next neighboring small block above (from left to right) of the current block may be checked. The second neighbor of sub - block A may be left neighbor b. For example, if small block b is not available or not intra - coded, the next neighboring small block to the left (from top to bottom) of the current block may be checked. Motion information of the temporal neighbors of sub - block A may be obtained by following a similar procedure to the TMVP process in HEVC. Motion information of available spatial and temporal neighbors (for example, up to three) may be averaged and used as the motion information of sub - block A. Based on raster scan order, the above - described STMVP process may be repeated to derive motion information of other sub - blocks within the current video block.
[0072] Frame rate up-conversion mode (FRUC) may be supported for inter-coded blocks. When this mode is enabled, for example, motion information of coded blocks (e.g., including motion vectors and / or reference indices) may not be transmitted. The information may be derived at the decoder side by template matching and / or bilateral matching techniques. Perhaps, for example, during the motion derivation process in the decoder, a set of preliminary motion vectors generated from the merge candidate list of the block and / or the motion vectors of the temporally collocated blocks of the current block may be checked. A candidate that results in the minimum sum of absolute differences (SAD) may be selected as the starting point. Local search based on template matching and / or bilateral matching around the starting point may be performed. The MV that results in the minimum SAD may be taken as the MV for all blocks. For better motion compensation efficiency, the motion information may be further refined at the sub-block level.
[0073] Figures 8A and 8B show examples of the FRUC process. As shown in Figure 8A, template matching may be used to derive the motion information of the current block by finding the (e.g., best) match between a template in the current picture (e.g., the upper and / or left neighboring blocks of the current block) and a block in the reference picture (e.g., of the same size as the template). In Figure 8B, bilateral matching may be used to derive the motion information of the current block by finding the (e.g., best) match between two blocks along the motion trajectory of the current block in two different reference pictures. The bilateral matching motion search process may be based on the motion trajectory. For example, the motion vectors MV0 and MV1 pointing to the two reference blocks may be proportional to the temporal distances (e.g., T0 and T1) between the current picture and one or each of the two reference pictures.
[0074] For motion compensation prediction, a translational motion model may be applied. Many types of motion exist, such as zoom in / out, rotation, perspective motion, and other irregular motions. Affine transform motion compensation prediction may be applied. As shown in FIG. 9A, the affine motion field of a block may be described by several (e.g., two) control point motion vectors. Based on the motion of the control points, the motion field of the affine block may be
[0075]
Number
[0076] described as follows. Here, as shown in FIG. 9A, (v 0x , v 0y ) may be the motion vector of the control point at the upper left corner, and (v 1x , v 1y ) may be the motion vector of the control point at the upper right corner. Perhaps, for example, when a video block is coded in affine mode, its motion field may be derived based on the granularity of 4×4 blocks. To derive the motion vectors of 4×4 blocks, the motion vectors of the central samples of each sub-block, as shown in FIG. 9B, may be calculated according to (15) and rounded to an accuracy of 1 / 16 pixel. The derived motion vectors may be used in the motion compensation stage to generate the prediction signal of the sub-blocks inside the current block.
[0077] To accelerate the processing speed of both encoding and decoding, in the software / hardware design of the latest video codecs, single instruction multiple data (SIMD) instructions may be used, and SIMD may execute the same operation on multiple data elements simultaneously, perhaps by using a single instruction. The SIMD width defines the number of data elements that may be processed in parallel by the registers. In a general-purpose central processing unit (CPU), 128-bit SIMD may be used. Graphics processing units (GPUs) can support wider SIMD implementations, for example, using 512-bit registers to support arithmetic, logical, load, and store instructions.
[0078] As described herein, to reduce the number of filtering operations, the BIO implementation may use in the gradient derivation process a 2D separable FIR filter, for example, a combination of a 1D low-pass interpolation filter and a 1D high-pass gradient filter. The selection of the corresponding filter coefficients may be based on the fractional position of the target sample. Because of such characteristics (e.g., 2D separable filters), some computational operations may be performed in parallel on multiple samples. The gradient derivation process may be suitable for SIMD acceleration.
[0079] For horizontal and vertical gradient derivation, vertical filtering may be applied, followed by horizontal filtering. For example, to calculate the horizontal gradient, vertical interpolation using the interpolation filter h L may be performed to generate intermediate samples, followed by the gradient filter h G being applied horizontally on the intermediate samples. For calculating the vertical gradient, the gradient filter h G may be applied vertically to calculate an intermediate vertical gradient, which may then be input to the horizontal interpolation filter h L . h L and h GAssuming that both filter lengths are T = 6, the horizontal filtering process may generate additional intermediate data (e.g., intermediate samples for horizontal gradient calculation and intermediate gradients for vertical gradient calculation) in the horizontal expansion region of the current block, providing sufficient reference data for the subsequent horizontal filtering process. FIGS. 10A and 10B show a 2D gradient filtering process that may be applied to BIO, and the dashed lines indicate the respective directions in which each filtering process may be applied. As shown in FIGS. 10A and 10B, for a W×H block, the size of the intermediate data is (W+T−1)×H=(W + 5)×H. In HEVC and JEM, the width of the coding block may be a power of 2, e.g., 4, 8, 16, 32, etc. Samples may be stored in memory by 1 byte (for 8-bit video) or 2 bytes (for video signals with more than 8 bits). As described herein, the width of the intermediate data input to the horizontal filtering process may be W + 5. Given an existing SIMD width, the SIMD register may not be fully utilized during the horizontal filtering process, which can reduce the parallel efficiency of the SIMD implementation. For example, for a coding block with a width of 8, the width of the intermediate data may be 8 + 5 = 13. Assuming a 128-bit SIMD implementation and a 10-bit input video, it may require two SIMD operation loops to process each line of the intermediate data during the horizontal filtering process. For example, the first SIMD loop may use the payload of the 128-bit register by filtering 8 samples in parallel, while in the second loop, the remaining 5 samples (e.g., 5×16 bits = 80 bits) may exist.
[0080] As shown in (10), (12), and (14), the rounding operation during the gradient derivation process may be performed by calculating the absolute value of the input data, adding one offset, and then right-shifting to round the absolute value, and multiplying the sign of the input data by the rounded absolute value. As described herein, one or more sub-block coding modes (e.g., ATMVP, STMVP, FRUC, and affine mode) may be used. When the sub-block level coding mode is enabled, the current coding block may be further divided into a plurality of small sub-blocks, and the motion information for each sub-block may be derived separately. Since the motion vectors of the sub-blocks within one coding block may be different, motion compensation may be performed separately for each sub-block. Assuming that the current block is coded by the sub-block mode, FIG. 11 shows an exemplary process used to generate the prediction signal of the block using BIO-related operations. As shown in FIG. 11, motion vectors may be derived for some (e.g., all) of the sub-blocks of the current block. Thereafter, normal MC may be applied to generate the motion-compensated prediction signal (e.g., Predi) for the sub-blocks within the block. For example, if BIO is used, BIO-based motion refinement may be further performed to obtain the modified prediction signal PredBIOi for the sub-blocks. This may result in multiple activations of BIO to generate the prediction signal for each sub-block. To perform BIO at the sub-block level, interpolation filtering and gradient filtering may access additional reference samples (depending on the filter length). The number of sub-blocks included in the block may be relatively large. Frequent switching between the motion compensation operation and the use of different motion vectors may occur.
[0081] For efficient SIMD implementation, BIO prediction derivation may be performed.
[0082] In an exemplary BIO implementation, the BIO prediction may be derived using Equation (2). In an exemplary BIO implementation, the BIO prediction may include one or more steps (e.g., two steps). A step (e.g., the first step) may be to derive an adjustment from high accuracy (e.g., using Equation (16)). A step (e.g., the second step) may be to derive the BIO prediction by combining predictions (e.g., two predictions) from lists (e.g., two lists) with the adjustment, as seen in Equation (17).
[0083]
Number
[0084] The parameter round1 may be equal to (1<<(shift1 - 1)) for 0.5 rounding.
[0085]
Number
[0086] The parameter round2 may be equal to (1<<(shift2 - 1)) for 0.5 rounding.
[0087] The rounding in Equation (16) may calculate the absolute value and sign from a variable and combine the sign with the intermediate result after right shift. The rounding in Equation (16) may use multiple operations.
[0088] The BIO gradient may be derived so that SIMD-based acceleration can be utilized. For example, the BIO gradient may be derived by applying horizontal filtering and subsequently applying vertical filtering. The length of the intermediate data input to the second filtering process (e.g., vertical filtering) may be a multiple of the length of the SIMD register, perhaps to fully utilize the parallelism capabilities of the SIMD. In an example, the rounding operation on the input values may be performed directly. This may be done by addition and right shift.
[0089] In the BIO gradient derivation process, horizontal filtering may follow vertical filtering. The width of the intermediate block may not often be well-aligned with the length of the common SIMD register. The rounding operation during the gradient derivation process may be performed based on the absolute value, which may introduce costly calculations (e.g., absolute value calculation and multiplication) for SIMD implementation.
[0090] As shown in FIGS. 10A and 10B, in the BIO gradient calculation process, horizontal filtering may follow vertical filtering. Perhaps due to the lengths of the interpolation filter and gradient filter that may be used, the width of the intermediate data after vertical filtering may be W + 5. Such a width may or may not be aligned with the width of the SIMD register actually used.
[0091] For both horizontal and vertical gradient derivation, horizontal filtering may be performed, followed by vertical filtering. To calculate the horizontal gradient, the horizontal gradient filter h G may be executed to generate an intermediate horizontal gradient based on the fractional horizontal movement fracX, and subsequently, the interpolation filter h L may be applied vertically on the intermediate horizontal gradient according to the fractional vertical movement fracY. For the calculation of the vertical gradient, the interpolation filter h Lmay be applied horizontally based on the value of fracX, perhaps, to calculate an intermediate sample. Gradient filter h G may be applied vertically to the intermediate sample, perhaps, according to the value of fracY. FIGS. 12A and 12B show the corresponding 2D gradient derivation process after filtering. As shown in FIGS. 12A and 12B, the size of the intermediate data (e.g., the intermediate gradient for horizontal gradient calculation and the intermediate sample for vertical gradient calculation) may be W×(H+T−1)=W×(H+5). The filtering processes of FIGS. 12A and 12B may ensure that the width of the intermediate data (e.g., W) can align with the SIMD register lengths (e.g., 128 bits and 512 bits) that may be used in implementation. For illustration, take the same example as in FIGS. 10A and 10B, assuming 128-bit SIMD implementation, 10-bit input video, and block width W = 8. As seen in FIGS. 10A and 10B, two sets of SIMD operations may be used to process the data within each line of the intermediate block (e.g., W+5). The first set may use the payload of a 128-bit SIMD register (e.g., 100% utilization), while the second set may use 80 bits out of the 128-bit payload (e.g., 62.5% utilization). The set of SIMD operations may utilize the 128-bit register capacity, e.g., 100% utilization.
[0092] As described herein, BIO gradient derivation may round the input value based on the absolute value, which may minimize the rounding error. For example, the absolute value of the input may be calculated by rounding the absolute value and multiplying the rounded absolute value by the sign of the input. This rounding may be
[0093]
Number
[0094] described as being like this, where σ i and σ rmay be the corresponding values of the input signal and the quantized signal, and o and shift may be the number of offset and right shift that may be applied during quantization. When deriving the gradient for the BIO block, a quantization operation on the input value may be performed, for example,
[0095] [Number]
[0096] is.
[0097] Figures 13 and 14 compare the mapping functions of different quantization methods. As can be seen in Figures 13 and 14, the difference between the quantized values calculated by the two methods may be small. The difference is σ i when the input values of, probably, are rounded to integers -1, -2, -3,... by the quantization method of Figure 13 and can be rounded to integers 0, -1, -2,... by the quantization method of Figure 14, and are equal to -0.5, -1.5, -2.5,... There may be. The coding performance impact introduced by the quantization method of Figure 14 may be negligible. As seen in (17), the quantization method of Figure 14 may be terminated by a single step and may be implemented by addition and right shift, neither of which can be more costly than the absolute value calculation and multiplication that may be used in (16).
[0098] As described herein, the order of the 2D separable filter and / or the use of any quantization method may affect BIO gradient derivation. Figures 15A and 15B show an exemplary gradient derivation process in which the order of the 2D separable filter and the use of any quantization method may affect BIO gradient derivation. For example, when deriving the horizontal gradient, the fractional sample s 1,0 、g 1,0 、a 1,0 、m 1,0 、y 1,0 The horizontal gradient value at is the gradient filter h in the horizontal direction GBy applying, it may be derived, for example,
[0099]
Number
[0100] is. a 1,1 The horizontal gradient of, for example, gH_a’ 1,1 is
[0101]
Number
[0102] As shown, the interpolation filter h L By applying it vertically, it may be interpolated from those intermediate horizontal gradient values. The vertical gradient may be calculated by applying the interpolation filter h L horizontally and interpolating the sample value at the fractional position (1,0), for example,
[0103]
Number
[0104] is. The vertical gradient value at a1,1 is
[0105]
Number
[0106] As shown, it may be obtained by vertically executing the gradient filter h G at the intermediate fractional position at (1,0).
[0107] The bit-depth increase caused by the interpolation filter and the gradient filter may be the same (e.g., 6 bits as shown in Tables 1 and 2). Changing the filtering order may not affect the internal bit depth.
[0108] As described herein, one or more coding tools based on sub-block level motion compensation (e.g., ATMVP, STMVP, FRUC, and affine mode) may be used. When these coding tools are enabled, the coding block may be divided into a plurality of small sub-blocks (e.g., 4×4 blocks) and may derive its own motion information (e.g., reference picture and motion vector) for use in the motion compensation stage. Motion compensation may be performed separately for each sub-block. Additional reference samples may be fetched to perform motion compensation for each sub-block. Region-based motion compensation based on variable block size may be applied to merge consecutive sub-blocks presenting the same motion information within the coding block for the motion compensation process. This may reduce the number of motion compensation processes and BIO processes applied within the current block. Different schemes may be used to merge neighboring sub-blocks. Line-based sub-block merging and 2D sub-block merging may be performed.
[0109] For blocks coded in sub-block mode, motion-compensated prediction may be performed. Variable block size motion compensation may be applied by merging consecutive sub-blocks having the same motion information into sub-block groups. A single motion compensation may be performed for each sub-block group.
[0110] In line-based sub-block merging, adjacent sub-blocks may be merged by placing the same sub-block lines within the current coding block having the same motion into one group and performing a single motion compensation for the sub-blocks within the group. FIG. 16A shows an example where the current coding block consists of 16 sub-blocks, and each block may be associated with a specific motion vector. Based on the existing sub-block-based motion compensation method (as shown in FIG. 11), perhaps both normal motion compensation and BIO motion refinement may be performed separately for each sub-block to generate the prediction signal for the current block. Correspondingly, 16 activations of the motion compensation operation may exist (each operation includes both normal motion compensation and BIO). FIG. 16B shows the sub-block motion compensation process after the line-based sub-block merging scheme is applied. As shown in FIG. 16B, after horizontally merging the sub-blocks having the same motion, the number of motion compensation operations may be reduced to 6.
[0111] Sub-block merging may depend on the shape of the block.
[0112] As shown in FIG. 16B, for merging in the motion compensation stage, the motion of neighboring sub-blocks in the horizontal direction may be considered. For example, for merging in the motion compensation stage, sub-blocks within a (e.g., the same) sub-block row within a CU may be considered. For partitioning blocks within a (e.g., one) picture, a (e.g., one) quadtree plus binary tree (QTBT) structure may be applied. In the QTBT structure, (e.g., each) coding unit tree (CTU) may be partitioned using a quadtree implementation. (e.g., each) quadtree leaf node may be partitioned by a binary tree. This partitioning may occur in the horizontal and / or vertical directions. For intra-coding and / or inter-coding, rectangular and / or square-shaped coding blocks may be used. This may be due to binary tree partitioning. For example, when such a block partitioning scheme is implemented and a line-based sub-block merging method is applied, the sub-blocks may have similar (e.g., the same) motion in the horizontal direction. For example, when a rectangular block is oriented vertically (e.g., the block height is greater than the block width), adjacent sub-blocks arranged within the same sub-block column may have a greater correlation than adjacent sub-blocks arranged within the same sub-block row. In such a case, the sub-blocks may be merged in the vertical direction.
[0113] A sub-block merging scheme dependent on block shape may be used. For example, when the width of a CU is greater than or equal to its height, a sub-block merging scheme for rows may be used to jointly predict sub-blocks having the same motion in the horizontal direction (e.g., sub-blocks arranged in the same sub-block row). This may be performed using (e.g., one) motion compensation operation. For example, when the height of a CU is greater than its width, a sub-block merging scheme for columns may be used to merge adjacent sub-blocks having the same motion and arranged in the same sub-block column within the current CU. This may be performed using (e.g., one) motion compensation operation. FIGS. 18A and 18B show exemplary implementations of adaptive sub-block merging based on block shape.
[0114] In the line / column-based sub-block merging scheme described herein, in the motion compensation stage, for merging sub-blocks, the motion consistency of neighboring sub-blocks in the horizontal and / or vertical directions may be considered. In practice, the motion information of adjacent sub-blocks may have a high correlation in the vertical direction. For example, as shown in FIG. 16A, the motion vectors of the first three sub-blocks in the first sub-block row and the second sub-block row may be the same. In such a case, for more efficient motion compensation, both horizontal and vertical motion consistency may be considered when merging sub-blocks. A 2D sub-block merging scheme may be used, and adjacent sub-blocks in both the vertical and horizontal directions may be merged into sub-block groups. To calculate the block size for each motion compensation, an incremental search method may be used to merge sub-blocks both horizontally and vertically. Given a sub-block position, it calculates the maximum number of consecutive sub-blocks in a sub-block row (e.g., each sub-block row) that can be merged into a single motion compensation stage, compares the motion compensation block size that may be achieved in the current sub-block row with that calculated in the last sub-block row, and / or may function by merging sub-blocks. This may perhaps proceed by repeating the above steps until the motion compensation block size can no longer be increased after merging additional sub-blocks in a given sub-block row.
[0115] The search method described herein may be summarized by the following exemplary procedure. Given a sub-block position b in the i-th sub-block row and the j-th sub-block column i,j , the number of consecutive sub-blocks in the i-th sub-block row having the same motion as the current sub-block (e.g., N i ) may be calculated. The corresponding motion compensation block size is S i = N iSet k = i. Proceed to the (k + 1)-th sub-block row and calculate the number of consecutive sub-blocks that are allowed to be merged (e.g., N k+1 ). N k+1 = min(N k , N k+1 ) and update it. Calculate the corresponding motion compensation block size as S k+1 = N k+1 ·(k - i + 1). If S k+1 ≥ S k , set N k+1 = N k , k = k + 1, proceed to the (k + 1)-th sub-block row, and calculate the number of consecutive sub-blocks that are allowed to be merged (e.g., N k+1 ). N k+1 = min(N k , N k+1 ) and update it. Calculate the corresponding motion compensation block size as S k+1 = N k+1 ·(k - i + 1). Otherwise, end.
[0116] Figure 16C shows the corresponding sub-block motion compensation process after the 2D sub-block merging scheme is applied. As can be seen in Figure 16C, the number of motion compensation operations may be 3 (e.g., motion compensation operations for three sub-block groups).
[0117] As described herein, sub-block merging may be performed. Block extension constraints may be applied to the BIO gradient derivation process. Neighboring samples within the current block may be used to calculate local motion refinement for sample positions within the block. The local motion refinement of the samples may be calculated by considering neighboring samples within the 5×5 surrounding area of the samples (as shown in (5)), and the length of the current gradient / interpolation filter may be 6. The local motion refinement values derived for samples in the first / last two rows and the first / last two columns of the current block may perhaps not be as accurate when compared to the motion refinement values derived at other sample positions, for the copying of the horizontal / vertical gradients of the four corner sample positions to the extended area of the current block. Using a larger block size in the motion compensation stage may reduce the number of samples affected by the BIO block extension constraints. Assuming that the size of the sub-block is 8, FIGS. 17A and 17B compare the number of samples affected due to the BIO block extension constraints when the use of the sub-block motion compensation method (FIG. 17A) and the use of the 2D sub-block merging scheme (FIG. 17B) are respectively applied. As seen in FIGS. 17A and 17B, 2D sub-block merging can reduce the number of samples affected and can also minimize the impact caused by inaccurate motion refinement calculations.
[0118] In the derivation of the BIO prediction, rounding may be performed. Rounding may be applied to the absolute value (e.g., it may be applied first). The sign may be applied (e.g., in that case, it may be applied after a right shift). Rounding in the adjustment derivation for the BIO prediction may be applied as shown in Equation (24). The right shift in Equation (24) may be an arithmetic right shift (e.g., the sign of the variable may remain unchanged after the right shift). In SIMD, the adjustment adj(x,y) may be calculated. The adjustment may be calculated using one addition and one right shift operation. An exemplary difference between the rounding in Equation (16) and the rounding in Equation (24) may be shown in FIG. 13 and / or FIG. 14.
[0119] [Number]
[0120] (24) The rounding method may perform two rounding operations (e.g., round1 in (24) and round2 in (17)) on the original input value. The rounding method seen in (24) may merge two right shift operations (e.g., shift1 in (24) and shift2 in (17)) into one right shift. The final prediction generated by BIO may be seen in Equation (25),
[0121] [Number]
[0122] where round3 is equal to (1<<(shift1+shift2-1)).
[0123] (It may be set to 21 bits and 1 bit may be used for the sign) The current bit depth for deriving the adjustment value may be higher than the intermediate bit depth of the double prediction. As shown in (16), a rounding operation (e.g., round1) is performed on the adjustment value (e.g., adj hpIt may be applied to (x, y). The intermediate bit depth may be applied to generate a dual-prediction signal (e.g., may be set to 14 bits). An absolute-value-based right shift (e.g., shift1 = 20 - 14 = 6) may be performed in Equation (16). The bit depth of the derived BIO adjustment value may be reduced, for example, from 21 bits to 15 bits, so that a first rounding operation (e.g., (e.g., round1)) may be skipped during the BIO process. The BIO generation processes seen in (16) and (17) may be changed as seen in (26) and (27),
[0124] [Number]
[0125] Here,
[0126] [Number]
[0127] and
[0128] [Number]
[0129] may be horizontal and vertical prediction gradients derived with reduced precision, and v LP x and v LP y may be the corresponding local motions at a lower bit depth. As seen in (27), one rounding operation may be used in the BIO generation process.
[0130] In addition to being applied separately, it should be noted that the methods described herein may be applied in combination. For example, the BIO gradient derivation and sub-block motion compensation described herein may be combined. The methods described herein may be jointly enabled at the motion compensation stage.
[0131] To remove blocking artifacts in the MC stage, overlapped block motion compensation (OBMC) may be performed. OBMC may be performed, for example, for one or more or all block boundaries, except perhaps for the right and / or bottom boundaries of one block. When one video block is coded in one sub-block mode, the sub-block mode may refer to a coding mode (e.g., FRUC mode) that allows sub-blocks within the current block to have their own motion. OBMC may be performed for one or more or all four of the sub-block boundaries. FIG. 19 shows an example of the concept of OBMC. When OBMC is applied in one sub-block (e.g., sub-block A in FIG. 19), perhaps in addition to the motion vector of the current sub-block, the motion vectors of up to four neighboring sub-blocks may also be used to derive the prediction signal of the current sub-block. One or more or a number of prediction blocks that use the motion vectors of the neighboring sub-blocks may be averaged to generate the final prediction signal of the current sub-block.
[0132] In order to generate the prediction signal of one or more blocks, in OBMC, weighted averaging may be used. The prediction signal using the motion vector of at least one neighboring sub-block may be referred to as PN, and / or the prediction signal using the motion vector of the current sub-block may be referred to as PC. When OBMC is applied, the samples in the first / last 4 rows / columns of PN may be weighted-averaged using the samples at the same positions in PC. The samples to which weighted averaging is applied may be determined according to the location of the corresponding neighboring sub-block. When the neighboring sub-block is, for example, the upper neighbor (e.g., sub-block b in FIG. 19), the samples in the first X rows of the current sub-block may be adjusted. When the neighboring sub-block is, for example, the lower neighbor (e.g., sub-block d in FIG. 19), the samples in the last X rows of the current sub-block may be adjusted. When the neighboring sub-block is, for example, the left neighbor (e.g., sub-block a in FIG. 19), the samples in the first X columns of the current block may be adjusted. Perhaps, when the neighboring sub-block is, for example, the right neighbor (e.g., sub-block c in FIG. 19), the samples in the last X columns of the current sub-block may be adjusted.
[0133] The value of X and / or the weights may be determined based on the coding mode used to code the current block. For example, when the current block is not coded in sub-block mode, for at least the first 4 rows / columns of PN, the weight coefficients {1 / 4, 1 / 8, 1 / 16, 1 / 32} may be used, and / or for the first 4 rows / columns of PC, the weight coefficients {3 / 4, 7 / 8, 15 / 16, 31 / 32} may be used. For example, when the current block is coded in sub-block mode, the first 2 rows / columns of PN and PC may be averaged. In such a scenario, in particular, for PN, the weight coefficients {1 / 4, 1 / 8} may be used, and / or for PC, the weight coefficients {3 / 4, 7 / 8} may be used.
[0134] As described herein, BIO can be regarded as one enhancement of normal MC by improving the granularity and / or accuracy of the motion vectors used in the MC stage. Assuming that the CU includes a plurality of sub-blocks, FIG. 20 shows an example of a process for generating a prediction signal for the CU using BIO-related operations. As shown in FIG. 20, the motion vectors may be derived for one or more or all of the sub-blocks of the current CU. Normal MC may be applied to generate a motion-compensated prediction signal (e.g., Pred i ) for one or more or each sub-block within the CU. Perhaps, for example, when BIO is used, BIO-based motion refinement may be performed to obtain a modified prediction signal Pred BIO i for the sub-block. For example, when OBMC is used, it may be performed on one or more or each sub-block of the CU, and then the same procedures described herein continue to generate the corresponding OBMC prediction signal. In some scenarios, motion vectors of spatial neighboring sub-blocks may be used (e.g., perhaps instead of the motion vector of the current sub-block) to derive the prediction signal.
[0135] In FIG. 20, for example, when at least one sub-block is bi-predicted, BIO may be used in the normal MC stage and / or the OBMC stage. BIO may be activated to generate a prediction signal for the sub-block. FIG. 21 shows an exemplary flowchart of a prediction generation process that may be performed without using BIO after OBMC. In FIG. 21, perhaps, for example, in the normal MC stage, BIO may still follow after motion-compensated prediction. As described herein, the derivation of BIO-based motion refinement may be a sample-based operation.
[0136] BIO may be applied to a current CU coded using a sub-block mode (e.g., FRUC, affine mode, ATMVP, and / or STMVP) in a normal MC process. For a CU coded by one or more or any of those sub-block modes, the CU may be further divided into one or more or a number of sub-blocks, and one or more or each sub-block may be assigned one or more unique motion vectors (e.g., single prediction and / or dual prediction). Perhaps, for example, when BIO is enabled, the determination of whether to apply BIO or not, and / or the BIO operation itself may be performed separately for one or more or each of the sub-blocks.
[0137] As described herein, one or more techniques are contemplated for skipping the BIO operation in an MC stage (e.g., a normal MC stage). For example, the core design of BIO (e.g., gradient calculation and / or refined motion vectors) may be kept the same and / or substantially similar. In one or more techniques, for a block / sub-block for which one or more factors or conditions may be satisfied, the BIO operation may be (e.g., partially or fully) disabled. In some examples, MC may be performed without BIO.
[0138] In sub-block mode, a CU may be permitted to be divided into two or more sub-blocks, and / or one or more different motion vectors may be associated with one or more or each sub-block. For a CU and / or (e.g., if the CU is enabled for sub-block coding) for two or more sub-blocks, motion-compensated prediction and / or BIO operations may be performed. One or more of the techniques described herein may be applicable to video blocks that may not be coded in sub-block mode (e.g., that are not divided and / or (e.g., have a single) motion vector).
[0139] As described herein, BIO may compensate for (e.g., small) motion that may remain between (e.g., at least) two predicted blocks generated by conventional block-based MC. As shown in FIGS. 8A and 8B, FRUC bilateral matching may be used to estimate motion vectors based on temporal symmetry along the motion trajectory between predicted blocks in the forward and / or backward reference pictures. For example, the values of the motion vectors associated with two predicted blocks may be proportional to the temporal distance between the current picture and their respective reference pictures. Bilateral matching-based motion estimation may perhaps provide one or more (e.g., reliable) motion vectors when there may be (e.g., only) small translational motion between two reference blocks (e.g., coding blocks in the highest temporal layer in a random access configuration).
[0140] For example, among other scenarios, when at least one sub-block is coded in the FRUC bilateral mode, one or more true motion vectors of the samples within the sub-block may be coherent (e.g., should be coherent). In one or more techniques, during the normal MC process for one or more sub-blocks coded in the FRUC bilateral mode, BIO may be disabled. FIG. 22 shows an exemplary diagram of the prediction generation process after disabling the BIO process for FRUC bilateral blocks in the MC stage. For example, on the decoder side, perhaps once the decoder determines that the FRUC bilateral mode has been used / is being used to code a block, the BIO process may be bypassed.
[0141] As described herein, BIO may be skipped for one or more FRUC bilateral sub-blocks where two motion vectors may be symmetric in the time domain (e.g., may always be symmetric). Perhaps among other scenarios, in order to achieve further complexity reduction, based on the (e.g., absolute) difference between at least two motion vectors of at least one doubly predicted sub-block, the BIO process may be skipped in the MC stage. For example, for one or more sub-blocks that may be predicted by two motion vectors that may be approximately proportional in the time domain, it may be reasonably assumed that the two prediction blocks have a high correlation and / or the motion vectors used in sub-block level MC may be sufficient to accurately reflect the true motion between the prediction blocks. In one or more scenarios, the BIO process may be skipped for those sub-blocks. In the case of a scenario of a doubly predicted sub-block where its motion vectors may not be temporally proportional (e.g., may never be proportional), BIO may be executed on sub-block level motion compensated prediction.
[0142] Using the same notation as in FIG. 3, for example, (MV x0 , MV y0 ) and / or (MV x1 , MV y1 ) may represent sub-block level motion vectors (e.g., prediction signals) used to generate two prediction blocks. Also, τ0 and τ1 represent the temporal distances from the current picture to the forward and / or backward temporal reference pictures. Also, (MV s x1 , MV s y1 ) may be calculated as a scaled version of (MV x1 , MV y1 ),
[0143]
Number
[0144] and may be generated based on τ0 and τ1, as in
[0145] . Based on (28), perhaps when one or more contemplated techniques may be applied, the BIO process may be skipped for at least one block and / or sub-block. For example, based on (28), two prediction blocks (e.g., reference blocks) may be determined to be similar or dissimilar (e.g., determined based on a prediction difference). If two prediction blocks (e.g., reference blocks) are similar, the BIO process may be skipped for at least one block or sub-block. If two prediction blocks (e.g., reference blocks) are dissimilar, the BIO process may not need to be skipped for at least one block or sub-block. For example, two prediction blocks may be determined to be similar when the following conditions are satisfied.
[0146]
Number
[0147] The variable thres may indicate a predefined / pre-determined threshold of the motion vector difference. Otherwise, the motion vectors that may be used for sub-block level MC of the current sub-block may be considered inaccurate. In particular, in such scenarios, BIO may still be applied to the sub-blocks. For example, the variable thres may be transmitted and / or determined based on desired coding performance (e.g., determined by the decoder).
[0148] FIG. 23 shows an exemplary prediction generation process in which BIO may be disabled based on the motion vector difference criterion of (29). As can be seen from (29), in the MC stage, a threshold of the motion vector difference (e.g., thres) may be used to determine whether at least one block or sub-block can skip the BIO process. In other words, a threshold may be used to identify where skipping the BIO process may or may not have a negligible impact on the overall encoding / decoding complexity. A threshold may be used to determine whether the BIO process should be skipped or not. In one or more techniques, the same threshold of the motion vector difference may be used for one or more or all pictures in the encoder and / or decoder. In one or more techniques, one or more different thresholds may be used for different pictures. For example, in a random access configuration, a relatively small threshold may be used for pictures in a high temporal layer (e.g., perhaps due to less motion). For example, a relatively large threshold may be used for pictures in a low temporal layer (e.g., due to more motion).
[0149] As shown in (28) and (29), motion vector scaling is used to calculate the motion vector difference (MVx1 , MV y1 ) may be applicable. For example, when the motion vector scaling in (28) is applied and / or τ0 > τ1 is assumed, the error caused by the motion estimation of (MV x1 , MV y1 ) may be amplified. Motion vector scaling may be applied to the motion vectors associated with the reference picture that has a relatively large temporal distance from the current picture (e.g., may always be applied). For example, when τ0 ≤ τ1, in order to calculate the motion vector difference (as shown, for example, by (28) and (29)), the motion vectors (MV x1 , MV y1 ) may be scaled. Otherwise, for example, when τ0 > τ1,
[0150]
Number
[0151] as shown, in order to calculate the motion vector difference, the motion vectors (MV x0 , MV y0 ) may be scaled.
[0152] As described herein, motion vector scaling may be applied when the temporal distance from two reference pictures to the current picture is different (e.g., τ0≠τ1). Since motion vector scaling can introduce additional errors (e.g., due to division and / or rounding operations), it can affect (e.g., reduce) the accuracy of the scaled motion vectors. Among other scenarios, perhaps in order to avoid or reduce such errors, for example, when the temporal distance from at least two or at most two reference pictures (e.g., reference blocks) to the current sub-block can be the same or substantially similar (e.g., τ0 = τ1) (e.g., only then), a motion vector difference (as shown in (30) and (31)) may be used (e.g., only used) to disable BIO in the MC stage.
[0153] As described herein, a motion vector difference may be used as a measurement to determine whether the BIO process can be skipped for at least one sub-block in the MC stage (e.g., based on whether two reference blocks are similar). When the motion vector difference between two reference blocks is relatively small (e.g., below a threshold), it may be reasonable to assume that the two predicted blocks may be similar (e.g., have a high correlation) so that BIO can be disabled without incurring a (substantial) coding loss. The motion vector difference may be one of many ways to measure the similarity (e.g., correlation) between two predicted blocks (e.g., reference blocks). In one or more techniques, the correlation between two predicted blocks may be determined, for example, by calculating the average difference between the two predicted blocks, by
[0154]
Number
[0155] and may be determined.
[0156] Variable I (0) (x, y) and I (1) (x, y) are sample values at the coordinates (x, y) of a motion-compensated block derived from a forward and / or backward reference picture (e.g., a reference block). The sample values may be associated with the luma value of their respective reference blocks. The sample values may be interpolated from their respective reference blocks. Variables B, and N are, respectively, a set of sample coordinates, and the number of samples defined within the current block or sub-block. Variable D is a distortion measure to which one or more different measurements / metrics, such as sum of squared errors (SSE), sum of absolute differences (SAD), and / or sum of absolute transformed differences (SATD), may be applied. Given (32), perhaps, for example, the difference measurement may be less than one or more pre-defined / pre-determined thresholds, e.g., Diff ≦ D thres When that is the case, BIO can be skipped in the MC stage. Otherwise, the two prediction signals of the current sub-block may be considered dissimilar (e.g., less correlated), to which BIO may be applied (e.g., may still be applied). As explained herein, D thres may be transmitted or determined by the decoder (e.g., based on desired coding performance). FIG. 24 shows an example of the prediction generation process after BIO has been skipped, based on measuring the difference between two prediction signals.
[0157] As described herein, the BIO process may be conditionally skipped at the CU or sub-CU level (e.g., in a sub-block having a CU). For example, the encoder and / or decoder may determine that the BIO may be skipped for the current CU. As described herein, this determination may be based on the similarity of two reference blocks associated with the current CU (e.g., using (28), (29), (30), (31), and / or (32)). If the encoder and / or decoder determine that the BIO should be skipped for the current CU, the encoder and / or decoder may determine whether the current CU is coded using enabled sub-block coding. If the current CU is coded using enabled sub-block coding, the encoder and / or decoder may determine that the BIO may be skipped for the sub-blocks within the CU.
[0158] As shown in FIG. 24, the BIO process may be conditionally skipped for the current CU or a sub-block within the current CU where the distortion between its two prediction signals can be below a threshold. The calculation of the distortion measurement and the BIO process may be performed based on sub-blocks and may be frequently activated for the sub-blocks within the current CU. CU-level distortion measurements may be used to determine whether the BIO process should be skipped for the current CU. Based on the distortion values calculated from different block levels, multi-stage early termination may be performed where the BIO process may be skipped.
[0159] The skew may be calculated considering some (e.g., all) samples within the current CU. For example, if the skew at the CU level is small enough (e.g., below a predefined CU-level threshold), the BIO process may be skipped for the CU, and otherwise, at the sub-block level, the skew for each sub-block within the current CU may be calculated and used to determine whether to skip the BIO process. FIG. 26 shows an exemplary motion-compensated prediction process to which multi-stage early termination is applied to BIO. The variables in FIG. 26
[0160]
Number
[0161] and
[0162]
Number
[0163] represent the prediction signals generated for the current CU from the reference picture lists L0 and L1,
[0164]
Number
[0165] and
[0166]
Number
[0167] represent the corresponding prediction signals generated for the i-th sub-block within the CU.
[0168] As described herein, for example, for current CU, CU-level distortion may be calculated to determine whether the BIO operation may be disabled. The motion information of sub-blocks within the CU may have a high correlation or may not have a high correlation. The distortion of sub-blocks within the CU may vary. Multi-stage early termination may be performed. For example, the CU may be divided into a plurality of groups of sub-blocks, and the groups may include consecutive sub-blocks having similar (e.g., the same) motion information (e.g., the same reference picture index and the same motion vector). For each sub-block group, a distortion measurement may be calculated. For example, if the distortion of one sub-block group is sufficiently small (e.g., below a predefined threshold), the BIO process may be skipped for samples within the sub-block group, and otherwise, for each sub-block within the sub-block group, the distortion may be calculated and used to determine whether to skip the BIO process for the sub-blocks.
[0169] (32) In, I (0) (x, y) and I (1) (x, y) may refer to the value of the motion-compensated sample (e.g., luma value) at the coordinates (x, y) obtained from the reference picture lists L0 and L1. The value of the motion-compensated sample may be defined with the precision of the input signal bit depth (e.g., 8 bits or 10 bits if the input signal is 8-bit or 10-bit video). The prediction signal of one bi-prediction block may be generated by averaging two prediction signals from L0 and L1 with the precision of the input bit depth.
[0170] For example, if the MV points to a fractional sample position, I (0) (x, y) and I (1)(x, y) is obtained using interpolation at an intermediate precision (which may be higher than the input bit depth), and then the intermediate prediction signal is rounded to the input bit depth prior to the averaging operation. For example, for one block, if fractional MVs are used, the two prediction signals at the precision of the input bit depth may be averaged at an intermediate precision (e.g., the intermediate precision specified in H.265 / HEVC and JEM). For example, I (0) (x, y) and I (1) (x, y), if the MVs used to obtain them correspond to fractional sample positions, the interpolation filter process may maintain intermediate values at a high precision (e.g., intermediate bit depth). For example, if one of the two MVs is an integer motion (e.g., the corresponding prediction is generated without applying interpolation), the precision of the corresponding prediction may be increased to the intermediate bit depth before the averaging is applied. FIG. 27 shows an exemplary dual-prediction process when averaging two intermediate prediction signals at a high precision,
[0171]
Number
[0172] and
[0173]
Number
[0174] refer to two prediction signals obtained from lists L0 and L1 at an intermediate bit depth (e.g., 14 bits as specified in HEVC and JEM), and BitDepth indicates the bit depth of the input video.
[0175] Given the dual-prediction signals generated at a high bit depth, the corresponding distortion between the two prediction blocks in Equation (32) is
[0176]
Number
[0177] May be calculated at an intermediate precision as specified herein, where
[0178] [Number]
[0179] and
[0180] [Number]
[0181] are, respectively, the high-precision difference sample values at the coordinates (x, y) of the prediction blocks generated from L0 and L1, Diff h represents the corresponding distortion measurement calculated at the intermediate bit depth. Additionally, perhaps due to the high-bit-depth distortion in (33), the CU-level threshold and sub-block-level threshold used by BIO early termination (e.g., as described herein with respect to FIGS. 24 and 26) may be adjusted as they would be defined at the same bit depth of the predicted signal. As an example, using the L1 norm distortion (e.g., SAD), the following equation may describe how to adjust the distortion threshold from the input bit depth to the intermediate bit depth.
[0182] [Number]
[0183] Variable
[0184] [Number]
[0185] and
[0186] [Number]
[0187] are the CU level and sub-block distortion thresholds at the precision of the input bit depth,
[0188]
Number
[0189] and
[0190]
Number
[0191] are the corresponding distortion thresholds at the precision of the intermediate bit depth.
[0192] As described herein, the BIO may provide motion refinement for samples, which may be calculated based on local gradient information at one or more or each sample location within at least one motion-compensated block. For sub-blocks within one region (e.g., a flat area) that may not contain much high-frequency detail, the gradients that may be derived using the gradient filter in Table 1 may tend to be small. As shown in Equation (4), when the local gradients (e.g., ∂I (k) (x,y) / ∂x and ∂I (k) (x,y) / ∂y) are close to zero, the final prediction signal obtained from the BIO may be approximately equal to the prediction signal generated by conventional dual prediction, e.g.,
[0193]
Number
[0194] is.
[0195] In one or more techniques, BIO may be applied (e.g., only applied) to one or more predicted samples of sub-blocks that may contain sufficient and / or abundant high-frequency information. The determination of whether the predicted signal of at least one sub-block contains high-frequency information can be made, for example, based on the magnitude of the average of the gradients for the samples within at least one sub-block. Perhaps, for example, if the average is less than at least one threshold, the sub-block may be classified as a flat area and / or BIO may not need to be applied to the sub-block. Otherwise, at least one sub-block may be considered to contain sufficient high-frequency information / details for which BIO may still be applicable. The value(s) of one or more thresholds for one or more or each reference picture may perhaps be determined a priori and / or adaptively, for example, based on the average gradient in that reference picture. In one or more techniques, gradient information may be calculated (e.g., only calculated) to determine whether the current sub-block may belong to one flat area. In some techniques, this gradient information may not need to be as accurate as the gradient values that may be used to generate the BIO prediction signal (as shown in (4)). A relatively simple gradient filter may be used to determine whether to apply BIO to at least one sub-block. For example, a 2-tap filter [-1,1] may be applied to derive horizontal and / or vertical gradients. FIG. 25 shows an exemplary flowchart of the prediction generation process after applying gradient-based BIO skip.
[0196] In the MC stage, BIO may be skipped, and one or more of the techniques individually described herein may be applied jointly (e.g., two or more techniques may be combined, etc.). One or more of the techniques, thresholds, equations, and / or factors / conditions, etc. described herein may be freely combined. One or more different combinations of the techniques, thresholds, equations, and / or factors / conditions, etc. described herein can provide different trade-offs with respect to coding performance and / or reduction of encoding / decoding complexity. For example, one or more of the techniques, thresholds, equations, and / or factors / conditions, etc. described herein may be implemented jointly or separately such that the BIO process may be disabled (e.g., completely disabled) for at least one block / sub-block.
[0197] FIG. 28A is a diagram showing an exemplary communication system 100 in which one or more of the disclosed embodiments may be implemented. The communication system 100 may be a multi-connection system that provides content such as voice, data, video, messaging, and broadcasting to a plurality of wireless users. The communication system 100 may enable a plurality of wireless users to access such content through sharing of system resources including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, and filter bank multicarrier (FBMC).
[0198] As shown in FIG. 28A, communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, public switched telephone network (PSTN) 108, Internet 110, and other networks 112, although the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, each of them, which may be referred to as a “station” and / or “STA,” WTRUs 102a, 102b, 102c, 102d may be configured to transmit and / or receive wireless signals and may be a user equipment (UE), mobile station, fixed or mobile subscriber unit, subscription-based unit, pager, cellular phone, personal digital assistant (PDA), smartphone, laptop, netbook, personal computer, wireless sensor, hotspot or Mi-Fi device, machine-to-internet (IoT) device, watch or other wearable, head-mounted display (HMD), vehicle, drone, medical device and application (e.g., telesurgery), industrial device and application (e.g., robots and / or other wireless devices operating in industrial and / or automated processing chain scenarios), home appliance device, and devices operating on commercial and / or industrial wireless networks, etc. Any of WTRUs 102a, 102b, 102c, 102d may alternatively be referred to as a UE.
[0199] The communication system 100 may also include base station 114a and / or base station 114b. Each of base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks such as CN106 / 115, Internet 110, and / or other network 112. By way of example, base stations 114a, 114b may be a base transceiver station (BTS), Node B, eNodeB, home Node B, home eNodeB, gNB, NR Node B, site controller, access point (AP), and wireless router, among others. Although base stations 114a, 114b are each depicted as a single element, it will be understood that base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0200] Base station 114a may be part of RAN104 / 113, and RAN104 / 113 may also include other base stations and / or network elements (not shown) such as a base station controller (BSC), a radio network controller (RNC), a relay node, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive radio signals on one or more carrier frequencies, which may be referred to as a cell (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. The cell may provide coverage for wireless services in a particular geographic area that may be relatively fixed or may change over time. The cell may further be divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one for each sector of the cell. In an embodiment, base station 114a may utilize multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0201] Base stations 114a, 114b may communicate with one or more of WTRUs 102a, 102b, 102c, 102d over air interface 116, and air interface 116 may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interface 116 may be established using any suitable radio access technology (RAT).
[0202] More specifically, as mentioned above, the communication system 100 may be a multi-connection system and may utilize one or more channel access methods such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 114a within RAN104 / 113 and the WTRUs 102a, 102b, 102c may establish the air interfaces 115 / 116 / 117 using wideband CDMA (WCDMA) and implement radio technologies such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA). WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed Uplink (UL) Packet Access (HSUPA).
[0203] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may establish the air interface 116 using Long-Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro) and implement radio technologies such as Evolved UMTS Terrestrial Radio Access (E-UTRA).
[0204] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may establish the air interface 116 using New Radio (NR) and implement radio technologies such as NR radio access.
[0205] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may implement LTE radio access and NR radio access together, for example, using the dual connectivity (DC) principle. Accordingly, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by transmissions to / from multiple types of radio access technologies and / or multiple types of base stations (e.g., eNBs and gNBs).
[0206] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement wireless technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM) for mobile communications, High-Speed Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN).
[0207] The base station 114b in Fig. 28A may be, for example, a wireless router, a home Node B, a home eNodeB, or an access point, and may utilize any suitable RAT to facilitate wireless connectivity in a localized area such as an office, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., used by a drone), and a roadway. In one embodiment, the base station 114b and the WTRUs 102c, 102d may implement a wireless technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d may implement a wireless technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As shown in Fig. 28A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106 / 115.
[0208] RAN 104 / 113 may communicate with CN 106 / 115, and CN 106 / 115 may be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more of WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid originating calls, Internet connectivity, video distribution, etc., and / or may perform high-level security functions, such as user authentication. Although not shown in FIG. 28A, it will be understood that RAN 104 / 113 and / or CN 106 / 115 may communicate directly or indirectly with other RANs that utilize the same or a different radio access technology (RAT) as RAN 104 / 113. For example, in addition to being connected to RAN 104 / 113, which may utilize New Radio (NR) radio technology, CN 106 / 115 may also communicate with another RAN (not shown) that utilizes GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0209] CN106 / 115 may also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network that provides basic telephone service. The Internet 110 may include a worldwide system of interconnected computer networks and devices that use common communication protocols such as the Transmission Control Protocol (TCP), the User Datagram Protocol (UDP), and / or the Internet Protocol (IP) within the TCP / IP Internet protocol suite. The network 112 may include wired and / or wireless communication networks that are owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs that may utilize the same or a different RAT as the RAN 104 / 113.
[0210] Some or all of the WTRUs 102a, 102b, 102c, 102d within the communication system 100 may include a multimode function (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with different wireless networks over different wireless links). For example, the WTRU 102c shown in FIG. 28A may be configured to communicate with a base station 114a that may utilize cellular-based wireless technology and with a base station 114b that may utilize IEEE 802 wireless technology.
[0211] Figure 28B is a system diagram showing an exemplary WTRU 102. As shown in Figure 28B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a global positioning system (GPS) chipset 136, and / or other peripheral devices 138. It will be understood that the WTRU 102 may include any sub-combination of the above elements while maintaining consistency with the embodiments.
[0212] The processor 118 may be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors cooperating with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to the transceiver 120, and the transceiver 120 may be coupled to the transmit / receive element 122. Although Figure 28B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.
[0213] The transmitting / receiving element 122 may be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) over the air interface 116. For example, in one embodiment, the transmitting / receiving element 122 may be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmitting / receiving element 122 may be a radiator / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmitting / receiving element 122 may be configured to transmit and / or receive both RF signals and optical signals. It will be understood that the transmitting / receiving element 122 may be configured to transmit and / or receive any combination of wireless signals.
[0214] In FIG. 28B, the transmitting / receiving element 122 is depicted as a single element, but the WTRU 102 may include any number of transmitting / receiving elements 122. More specifically, the WTRU 102 may utilize MIMO technology. Thus, in one embodiment, the WTRU 102 may include two or more transmitting / receiving elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0215] The transceiver 120 may be configured to modulate signals that are to be transmitted by the transmitting / receiving element 122 and demodulate signals received by the transmitting / receiving element 122. As mentioned above, the WTRU 102 may have a multimode function. Thus, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as, for example, NR and IEEE 802.11.
[0216] The processor 118 of the WTRU 102 may be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit) and may receive user input data therefrom. The processor 118 may output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may obtain information from and store data in any type of suitable memory, such as a non-removable memory 130 and / or a removable memory 132. The non-removable memory 130 may include a random access memory (RAM), a read only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include, for example, a subscriber identity module (SIM) card, a memory stick, and a secure digital (SD) memory card. In other embodiments, the processor 118 may obtain information from and store data in a memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0217] The processor 118 may receive power from a power source 134 and may be configured to distribute power to and / or control the power to other components within the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry cells (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), a solar cell, and a fuel cell.
[0218] Processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of WTRU 102. In addition to, or instead of, information from the GPS chipset 136, WTRU 102 may receive location information on air interface 116 from a base station (e.g., base stations 114a, 114b), and / or may determine its location based on the timing of signals received from two or more nearby base stations. It will be understood that WTRU 102 may obtain location information using any suitable location determination method while maintaining consistency with the embodiments.
[0219] Processor 118 may also be coupled to other peripheral devices 138, which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripheral devices 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or video), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a Frequency Modulation (FM) radio unit, a digital music player, a media player, a video game player module, an Internet browser, a Virtual Reality and / or Augmented Reality (VR / AR) device, and an activity tracker, among others. Peripheral devices 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0220] The WTRU 102 may include full-duplex radio in which some or all of the transmission and reception of some signals associated with a particular subframe for both the uplink (e.g., for transmission) and the downlink (e.g., for reception) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit 139 to reduce and / or substantially eliminate self-interference, either via hardware (e.g., a choke) or via signal processing through a processor (e.g., a separate processor (not shown) or the processor 118). In embodiments, the WTRU 102 may include half-duplex radio for some or all of the transmission and reception of some signals associated with a particular subframe for either the uplink (e.g., for transmission) or the downlink (e.g., for reception).
[0221] FIG. 28C is a system diagram showing the RAN 104 and the CN 106, in accordance with an embodiment. As mentioned above, the RAN 104 may communicate with the WTRU 102a, 102b, 102c over the air interface 116 using E-UTRA radio technology. The RAN 104 may also communicate with the CN 106.
[0222] The RAN 104 may include eNodeBs 160a, 160b, 160c, although it will be understood that the RAN 104 may include any number of eNodeBs while maintaining consistency with the embodiments. The eNodeBs 160a, 160b, 160c may each include one or more transceivers for communicating with the WTRU 102a, 102b, 102c over the air interface 116. In one embodiment, the eNodeBs 160a, 160b, 160c may implement MIMO technology. Thus, the eNodeB 160a may, for example, transmit wireless signals to and / or receive wireless signals from the WTRU 102a using multiple antennas.
[0223] Each of eNodeBs 160a, 160b, and 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, and scheduling of users in the UL and / or DL. As shown in FIG. 28C, eNodeBs 160a, 160b, and 160c may communicate with each other over the X2 interface.
[0224] CN 106 shown in FIG. 28C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. Although each of the above elements is depicted as part of CN 106, it will be understood that any of these elements may be owned and / or operated by an entity different from the CN operator.
[0225] MME 162 may be connected to each of eNodeBs 160a, 160b, and 160c within RAN 104 via the S1 interface and may act as a control node. For example, MME 162 may be responsible for authenticating users of WTRUs 102a, 102b, and 102c, bearer activation / deactivation, and selecting a particular serving gateway during the initial attach of WTRUs 102a, 102b, and 102c. MME 162 may provide control plane functions for exchanges between RAN 104 and other RANs (not shown) that utilize other radio technologies such as GSM and / or WCDMA.
[0226] SGW164 may be connected to each of eNodeBs 160a, 160b, and 160c within RAN104 via the S1 interface. SGW164 may generally route and transfer user data packets to / from WTRUs 102a, 102b, and 102c. SGW164 may perform other functions such as anchoring the user plane during eNodeB handover, triggering paging when DL data is available to WTRUs 102a, 102b, and 102c, and managing and storing the contexts of WTRUs 102a, 102b, and 102c.
[0227] SGW164 may be connected to PGW166, which may provide access to a packet switched network, such as the Internet 110, to WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and IP-enabled devices.
[0228] CN106 may facilitate communication with other networks. For example, CN106 may provide access to a circuit switched network, such as PSTN 108, to WTRUs 102a, 102b, and 102c to facilitate communication between WTRUs 102a, 102b, and 102c and conventional fixed line telephone communication devices. For example, CN106 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN106 and PSTN 108. Additionally, CN106 may provide access to other network 112 to WTRUs 102a, 102b, and 102c, where other network 112 may include other wired and / or wireless networks owned and / or operated by other service providers.
[0229] In FIGS. 28A - 28D, the WTRU is described as a wireless terminal, but in certain representative embodiments, it is contemplated that such a terminal may (e.g., temporarily or permanently) use a wired communication interface to a communication network.
[0230] In a representative embodiment, another network 112 may be a WLAN.
[0231] A WLAN in infrastructure basic service set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have an access or interface to a distribution system (DS) or another type of wired / wireless network that carries traffic within and / or outside the BSS. Traffic to an STA originating from outside the BSS may arrive through the AP and be delivered to the STA. Traffic transmitted from an STA to a destination outside the BSS may be transmitted to the AP for delivery to each destination. Traffic between STAs within the BSS may be transmitted through the AP; for example, the source STA may transmit the traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within the BSS may be considered peer - to - peer traffic and / or may be called peer - to - peer traffic. Peer - to - peer traffic may be transmitted (e.g., directly) between the source STA and the destination STA using direct link setup (DLS). In certain representative embodiments, the DLS may use 802.11e DLS or 802.11z tunnel DLS (TDLS). A WLAN using independent BSS (IBSS) mode may not have an AP, and STAs within the IBSS or using the IBSS (e.g., all of the STAs) may communicate directly with each other. Communication in IBSS mode may sometimes be referred to herein as "ad - hoc" mode communication.
[0232] When using the operation of 802.11ac infrastructure mode or the operation of a similar mode, the AP may transmit beacons on a fixed channel such as the primary channel. The primary channel may be of a fixed width (e.g., 20 MHz bandwidth), or a width dynamically set via signaling. The primary channel may be the operating channel of the BSS and may be used by the STA to establish a connection with the AP. In a representative embodiment, for example, in an 802.11 system, Carrier Sense Multiple Access / Collision Avoidance (CSMA / CA) may be implemented. In the case of CSMA / CA, STAs including the AP (e.g., any STA) may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. Within a given BSS, at any given time, one STA (e.g., only one station) may transmit.
[0233] A high throughput (HT) STA may use a 40 MHz wide channel for communication, for example, by combining the primary 20 MHz channel with adjacent or non - adjacent 20 MHz channels to form a 40 MHz wide channel.
[0234] Very High Throughput (VHT) STAs may support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels may be formed by combining consecutive 20 MHz channels. A 160 MHz channel may be formed by combining eight consecutive 20 MHz channels, or may be formed by combining two non - consecutive 80 MHz channels, which may be referred to as an 80 + 80 configuration. In the case of the 80 + 80 configuration, after channel encoding, the data may be passed through a segment parser that can split the data into two streams. For each stream separately, an Inverse Fast Fourier Transform (IFFT) process and a time - domain process may be performed. The streams may be mapped onto two 80 MHz channels and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations described above for the 80 + 80 configuration may be reversed and the combined data may be transmitted to the Media Access Control (MAC).
[0235] Operations in the sub - 1 GHz mode are supported by 802.11af and 802.11ah. The channel operating bandwidth and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports bandwidths of 5 MHz, 10 MHz, and 20 MHz in the TV White Space (TVWS) spectrum, and 802.11ah supports bandwidths of 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz using non - TVWS spectrum. According to an exemplary embodiment, 802.11ah may support Meter Type Control / Machine Type Communication, such as MTC devices in a macro - coverage area. MTC devices may have limited capabilities, including for example, support for a certain bandwidth and / or limited bandwidth support (e.g., only that support). MTC devices may include a battery with a battery life exceeding a threshold (e.g., to maintain a very long battery life).
[0236] A WLAN system that may support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, includes a channel that may be designated as a primary channel. The primary channel may have a bandwidth equal to the maximum common operating bandwidth supported by all STAs within a BSS. The bandwidth of the primary channel may be set and / or limited by an STA that supports the minimum bandwidth operation mode among all STAs operating within the BSS. In the example of 802.11ah, for an STA (e.g., an MTC type device) that supports (e.g., only supports) the 1MHz mode, even if the AP and other STAs within the BSS support 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidth operation modes, the primary channel may be 1MHz wide. Carrier sensing and / or network allocation vector (NAV) setting may depend on the status of the primary channel. For example, if the primary channel is busy because an STA (that only supports the 1MHz operation mode) is transmitting to an AP, the entire available frequency band may be considered busy, even if most of the frequency band remains idle and available.
[0237] In the United States, the available frequency band that may be used by 802.11ah is from 902MHz to 928MHz. In Korea, the available frequency band is from 917.5MHz to 923.5MHz. In Japan, the available frequency band is from 916.5MHz to 927.5MHz. The total bandwidth available for 802.11ah is from 6MHz to 26MHz, depending on national regulations.
[0238] FIG. 28D is a system diagram showing RAN 113 and CN 115 according to an embodiment. As mentioned above, RAN 113 may communicate with WTRUs 102a, 102b, 102c over air interface 116 using NR radio technology. RAN 113 may also communicate with CN 115.
[0239] RAN 113 may include gNBs 180a, 180b, 180c, although it will be understood that RAN 113 may include any number of gNBs while maintaining consistency with the embodiment. Each of gNBs 180a, 180b, 180c may include one or more transceivers for communicating with WTRUs 102a, 102b, 102c over air interface 116. In one embodiment, gNBs 180a, 180b, 180c may implement MIMO technology. For example, gNBs 180a, 108b may utilize beamforming to transmit signals to and / or receive signals from gNBs 180a, 180b, 180c. Thus, gNB 180a, for example, may transmit wireless signals to and / or receive wireless signals from WTRU 102a using multiple antennas. In an embodiment, gNBs 180a, 180b, 180c may implement carrier aggregation technology. For example, gNB 180a may transmit multiple component carriers to WTRU 102a (not shown). A subset of these component carriers may be in unlicensed spectrum, while the remaining component carriers may be in licensed spectrum. In an embodiment, gNBs 180a, 180b, 180c may implement multi-site coordinated (CoMP) technology. For example, WTRU 102a may receive coordinated transmissions from gNB 180a and gNB 180b (and / or gNB 180c).
[0240] WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using transmissions associated with scalable numerology. For example, the OFDM symbol interval and / or the OFDM subcarrier interval may vary for different transmissions, different cells, and / or different portions of the radio transmission spectrum. WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., including various numbers of OFDM symbols and / or lasting for various lengths of absolute time).
[0241] gNBs 180a, 180b, and 180c may be configured to communicate with WTRUs 102a, 102b, and 102c in a stand-alone configuration and / or a non-stand-alone configuration. In a stand-alone configuration, WTRUs 102a, 102b, and 102c may communicate with gNBs 180a, 180b, and 180c without accessing another RAN (such as eNodeBs 160a, 160b, and 160c for example). In a stand-alone configuration, WTRUs 102a, 102b, and 102c may utilize one or more of gNBs 180a, 180b, and 180c as a mobility anchor point. In a stand-alone configuration, WTRUs 102a, 102b, and 102c may communicate with gNBs 180a, 180b, and 180c using signals within an unlicensed band. In a non-stand-alone configuration, WTRUs 102a, 102b, and 102c may communicate with / connect to gNBs 180a, 180b, and 180c while also communicating with / connecting to another RAN such as eNodeBs 160a, 160b, and 160c. For example, WTRUs 102a, 102b, and 102c may implement the DC principle to communicate with one or more gNBs 180a, 180b, and 180c and one or more eNodeBs 160a, 160b, and 160c substantially simultaneously. In a non-stand-alone configuration, eNodeBs 160a, 160b, and 160c may serve as a mobility anchor for WTRUs 102a, 102b, and 102c, and gNBs 180a, 180b, and 180c may provide additional coverage and / or throughput for serving WTRUs 102a, 102b, and 102c.
[0242] Each of gNBs 180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to user plane functions (UPFs) 184a, 184b, and routing of control plane information to access and mobility management functions (AMFs) 182a, 182b. As shown in FIG. 28D, gNBs 180a, 180b, and 180c may communicate with each other over the Xn interface.
[0243] CN 115 shown in FIG. 28D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one session management function (SMF) 183a, 183b, and possibly data networks (DNs) 185a, 185b. Although each of the above elements is depicted as part of CN 115, it will be understood that any of these elements may be owned and / or operated by entities different from the CN operator.
[0244] AMF 182a and 182b may be connected to one or more of gNBs 180a, 180b, and 180c within RAN 113 via the N2 interface and may serve as control nodes. For example, AMF 182a and 182b may authenticate users of WTRUs 102a, 102b, and 102c, support network slicing (e.g., handling different PDU sessions with different requirements), select specific SMFs 183a and 183b, manage the registration area, terminate NAS signaling, and perform mobility management, etc. Network slicing may be used by AMF 182a and 182b to customize the CN support for WTRUs 102a, 102b, and 102c based on the type of service utilized by WTRUs 102a, 102b, and 102c. For example, different network slices may be established for different use cases such as services that rely on ultra-reliable low-latency (URLLC) access, services that rely on high-speed large-capacity mobile broadband (eMBB) access, and / or services for machine type communication (MTC) access. AMF 162 may provide control plane functions for exchanges between RAN 113 and other RANs (not shown) that utilize other radio technologies such as non-3GPP access technologies like LTE, LTE-A, LTE-A Pro, and / or WiFi.
[0245] SMF183a and 183b may be connected to AMF182a and 182b within CN115 via the N11 interface. SMF183a and 183b may also be connected to UPF184a and 184b within CN115 via the N4 interface. SMF183a and 183b may select and control UPF184a and 184b and configure the routing of traffic through UPF184a and 184b. SMF183a and 183b may perform other functions such as managing and allocating UE IP addresses, managing PDU sessions, implementing policies and controlling QoS, and providing downlink data notifications. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.
[0246] UPF184a and 184b may be connected to one or more of gNB180a, 180b, and 180c within RAN113 via the N3 interface, and they may provide access to a packet-switched network such as the Internet 110 to WTRU102a, 102b, and 102c to facilitate communication between WTRU102a, 102b, and 102c and IP-corresponding devices. UPF184a and 184b may perform other functions such as routing and forwarding packets, implementing user plane policies, supporting multi-homing PDU sessions, processing user plane QoS, buffering downlink packets, and providing mobility anchoring.
[0247] CN115 may facilitate communication with other networks. For example, CN115 may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between CN115 and PSTN108. Additionally, CN115 may provide access to other network 112 to WTRU102a, 102b, 102c, and other network 112 may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a, 102b, 102c may be connected to local data network (DN) 185a, 185b through UPF184a, 184b via an N3 interface to UPF184a, 184b and an N6 interface between UPF184a, 184b and DN185a, 185b.
[0248] With reference to FIGS. 28A-28D and the corresponding descriptions thereof, one or more of the functions described herein with respect to one or more of WTRU102a-d, base stations 114a-b, eNodeB160a-c, MME162, SGW164, PGW166, gNB180a-c, AMF182a-b, UPF184a-b, SMF183a-b, DN185a-b, and / or any other devices described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more or all of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate network and / or WTRU functionality.
[0249] An emulation device may be designed to perform one or more tests of other devices in a laboratory environment and / or in an operator network environment. For example, one or more emulation devices may perform one or more or all functions while being fully or partially implemented and / or deployed as part of a wired and / or wireless communication network to test other devices within the communication network. One or more emulation devices may perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. An emulation device may be directly coupled to another device for testing purposes and / or may use over-the-air wireless communication to perform the test.
[0250] One or more emulation devices may perform one or more functions including all functions without being implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be utilized in a test scenario in a test laboratory and / or in a non-deployed (e.g., test) wired and / or wireless communication network to perform tests of one or more components. One or more emulation devices may be test equipment. For transmitting and / or receiving data, direct RF coupling and / or wireless communication via an RF circuit (which may include one or more antennas, for example) may be used by an emulation device.
[0251] Above, features and elements have been described in specific combinations, but one of ordinary skill in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Additionally, the methods described herein may be implemented by a computer program, software, or firmware included in a computer-readable medium and executed by a computer and / or processor. Examples of computer-readable media include electronic signals (transmitted over a wired or wireless connection) and computer-readable storage media. Examples of computer-readable storage media include, without limitation, magnetic media such as read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, internal hard disks, and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any host computer.
Description of Reference Numerals
[0252] 100 Communication system 102a Receiving unit (WTRU) 102b Receiving unit (WTRU) 102c Receiving unit (WTRU) 102d Receiving unit (WTRU) 108 Public switched telephone network (PSTN) 110 Internet 112 Network 114a Base station 114b Base station 116 Air interface 118 Processor 120 Transceiver 122 Receiving element 124 Microphone 126 Keypad 128 Touch pad 130 Non-removable memory 132 Removable memory 134 Power supply 136 Chipset 138 Peripheral devices 139 Interference management unit 162 Mobility Management Entity (MME) 164 Serving Gateway (SGW) 166 Gateway (or PGW) 182a Access and Mobility Management Function (AMF) 182b Access and Mobility Management Function (AMF) 183a Session Management Function (SMF) 183b Session Management Function (SMF) 184a User Plane Function (UPF) 184b User Plane Function (UPF) 185a Data Network (DN) 185b Data Network (DN)
Claims
[Claim 1] Obtaining a first plurality of sample values in a first reference block of a current sub-block and a second plurality of sample values in a second reference block of the current sub-block; obtaining a sum of absolute differences (SAD) based on the first plurality of sample values and the second plurality of sample values; determining to disable bidirectional optical flow (BIO) for the current sub-block based on the SAD being less than a value; Decoding the current sub-block based on the determination of disabling BIO for the current sub-block. Processor configured to 1. A device for video decoding, comprising:
Citation Information
Patent Citations
Bi-directional optical flow for video coding
US20170094305A1