System and method for general multi-hypothesis prediction for video coding

By introducing a general-purpose multi-assumption prediction system and method in video encoding technology, using motion compensation prediction and block-level weight values ​​for linear combination, the problem of poor prediction performance in the prior art under illumination changes is solved, and more efficient video encoding and decoding is achieved.

CN115118971BActive Publication Date: 2025-06-17INTERDIGITAL VC HOLDINGS INC

Patent Information

Application Number
CN202210588340.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-10-31
Filing Date
2017-05-11
Publication Date
2025-06-17
Estimated Expiration
2037-05-11

AI Technical Summary

Technical Problem

The existing video encoding technology has poor prediction performance when processing illuminance changes over time, especially when the illuminance changes rapidly. The existing weighted double prediction and local illuminance compensation technology are difficult to completely solve this problem.

Method used

A general-format multi-assumption prediction system and method are proposed. Multiple prediction signals are linearly combined by using motion compensation prediction and block-level weight values. Specifically, the general-format double prediction framework is adopted. The limited weight set is used at the sequence level, picture level and slice level, and the weight value is optimized through the encoding method of signal transmission weight values.

Benefits of technology

By improving the prediction efficiency of weighted motion compensation prediction, reducing the number of bits required to encode and decode video, improving the operational performance of video encoder and decoder, and enhancing prediction capabilities in the case of illumination variation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115118971B_ABST
    Figure CN115118971B_ABST
Patent Text Reader

Abstract

Systems and methods for video encoding using general dual prediction are described. In an exemplary embodiment, to encode a current block of video in a bitstream, a first reference block is selected from a first reference picture and a second reference block is selected from a second reference picture. Each reference block is associated with a weight, where the weight is any weight in a range, for example, between 0 and 1. The current block is predicted by using a weighted sum of the reference blocks. The weight can be selected from a plurality of candidate weights. The candidate weights can be signaled within the bitstream or can be implicitly derived based on a template. The candidate weights can be pruned to avoid out-of-range or substantially duplicate candidate weights. General dual prediction can be additionally used within frame rate upconversion.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 201780042257.5, titled "System and Method for Generalized Multi-Hypothesis Prediction for Video Coding", filed on May 11, 2017, the content of which is incorporated herein by reference in its entirety.

[0002] Cross-reference to Related Applications

[0003] This application is a non-provisional application of the following U.S. Provisional Patent Applications and claims the benefit of these provisional patent applications under 35 U.S.C. § 119(c): Application No. 62 / 336,227, titled "System and Method for Generalized Multi-Hypothesis Prediction for Video Coding", filed on May 13, 2016; Application No. 62 / 342,772, titled "System and Method for Generalized Multi-Hypothesis Prediction for Video Coding", filed on May 27, 2016; Application No. 62 / 399,234, titled "System and Method for Generalized Multi-Hypothesis Prediction for Video Coding", filed on September 23, 2016; and Application No. 62 / 415,187, titled "System and Method for Generalized Multi-Hypothesis Prediction for Video Coding", filed on October 31, 2016. The entire contents of all of these applications are incorporated herein by reference. Background Art

[0004] Video coding systems are widely used to compress digital video signals to reduce the storage requirements and / or transmission bandwidth of such signals. Among various types of video coding systems, such as block-based, wavelet-based, and object-based systems, currently block-based hybrid video coding systems are the most widely used and deployed. Examples of block-based video coding systems include international video coding standards (such as MPEG-1 / 2 / 4 Part 2, H.264 / MPEG-4 Part 10 AVC, VC-1) and the latest video coding standard known as High Efficiency Video Coding (HEVC), which was developed by the Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T / SG16 / Q.6 / VCEG and ISO / IEC / MPEG.

[0005] Video encoded using block-based coding occupies a major portion of the data transmitted electronically, such as over the Internet. There is a desire to increase video compression efficiency so that high-quality video content can be stored and transmitted using fewer bits. Summary of the Invention

[0006] In an exemplary embodiment, a system and method for performing Generalized Bi-Prediction (GBi) are described. The exemplary method includes encoding and decoding (collectively referred to as “coding”) a video including a plurality of pictures, the plurality of pictures including a current picture, a first reference picture, and a second reference picture, each picture including a plurality of blocks. In the exemplary method, for at least a current block within the current picture, a block-level index is encoded to identify a first weight and a second weight within a set of weights, where at least one weight within the set of weights has a value other than 0, 0.5, or 1. The current block is predicted as a weighted sum of a first reference block within the first reference picture and a second reference block within the second reference picture, where the first reference block is weighted by the first weight and the second block is weighted by the second weight.

[0007] In some embodiments (or for some blocks), the information identifying the first weight and the second weight block-level information may be encoded in other ways than encoding an index for the block for the current block. For example, the block may be encoded in merge mode. In this case, the block-level information may be information identifying a candidate block from a plurality of merge candidate blocks. Thus, the first weight and the second weight may be identified based on the weights used to encode the identified candidate block.

[0008] In some embodiments, the first reference block and the second reference block may be further scaled by at least one scaling factor signaled within the bitstream for the current picture.

[0009] In some embodiments, the set of weights is encoded within the bitstream, allowing different sets of weights to be applicable to different slices, pictures, or sequences. In other embodiments, the set of weights is pre-determined. In some embodiments, only one of the two weights is signaled within the bitstream, and the other weight is derived by subtracting the signaled weight from 1.

[0010] In some embodiments, codewords are assigned to the respective weights, and the weights are identified by using the corresponding codewords. The assignment of codewords to weights may be a pre-determined assignment, or the assignment may be adaptively adjusted based on the weights used within previously encoded blocks.

[0011] An exemplary encoder and decoder for performing Generalized Bi-Prediction are also described herein.

[0012] The systems and methods described herein provide novel techniques for predicting blocks having sample values. This technique can be used by encoders and decoders. In the encoding method, the prediction of a block can result in the block containing sample values ​​being subtracted from the original input block to determine the residual encoded in the bitstream. In the decoding method, the residual can be decoded from the bitstream and added to the predicted block to obtain a reconstructed block that is the same as or similar to the original input block. Thus, the prediction method described herein can improve the operation of video encoders and decoders by reducing the number of bits required to encode and decode video in at least some implementations. Further advantages of the exemplary prediction method for the operation of video encoders and decoders will be given in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] A more detailed understanding may be obtained from the following description given by way of example in conjunction with the accompanying drawings, which are first briefly described below.

[0014] Figure 1 is a functional block diagram illustrating an example of a block-based video encoder.

[0015] Figure 2 is a functional block diagram illustrating an example of a block-based video decoder.

[0016] Figure 3 Is to use template T c And a schematic diagram of the prediction of the associated prediction blocks T0 and T1.

[0017] Figure 4 is a graph that provides a schematic representation of the variation of illumination over time.

[0018] Figure 5 is a functional block diagram illustrating a video encoder configured to use generalized bi-prediction in accordance with some embodiments.

[0019] Figure 6 is a functional block diagram of an exemplary generalized bi-prediction module for use within a video encoder.

[0020] Figure 7 is a schematic diagram of an exemplary decoder-side derivation of implicit weight values ​​for generalized bi-prediction.

[0021] Figure 8 is a schematic diagram of a tree structure for binarizing weight indices, where each circle represents a bit to be signaled.

[0022] Figure 9 is a functional block diagram illustrating a video decoder configured to use generalized bi-prediction in accordance with some embodiments.

[0023] Figure 10It is a functional block diagram of an exemplary general dual prediction module used within a video decoder.

[0024] Figure 11A and 11B provides a schematic diagram of the following codeword assignment methods: constant assignment ( Figure 11A ) and replaceable assignment ( Figure 11B ).

[0025] Figure 12A and 12B provides a schematic diagram of an example of block adaptive codeword assignment: weight value field ( Figure 12A ) and the final codeword assignment updated from constant assignment ( Figure 12B ).

[0026] Figure 13 is a schematic diagram of merge candidate positions.

[0027] Figure 14 is a schematic diagram of an example of overlapped block motion compensation (OBMC), where m is the basic processing unit for performing OBMC, N1 to N8 are causally adjacent sub-blocks, and B1 to B7 are sub-blocks within the current block.

[0028] Figure 15 shows an example of frame rate up-conversion (FRUC), where v0 is a given motion vector corresponding to reference list L0, and v1 is an MV scaled based on v0 and the temporal distance.

[0029] Figure 16 is a schematic diagram showing an example of the encoded bitstream structure.

[0030] Figure 17 is a diagram showing an illustration of a communication system.

[0031] Figure 18 is a diagram showing an illustration of an exemplary wireless transmit / receive unit (WTRU) that can be used as an encoder or a decoder in some embodiments. Detailed Description

[0032] Block-Based Coding

[0033] Figure 1Block diagram of a general block-based hybrid video coding system 100. An input video signal 102 is processed block by block. In HEVC, extended block sizes (referred to as "coding units" or CUs) are used to efficiently compress high-resolution (1080p and above) video signals. In HEVC, a CU can be up to 64x64 pixels. A CU can be further partitioned into prediction units or PUs, for which separate prediction methods can be applied. For each input video block (MB or CU), spatial prediction (160) and / or temporal prediction (162) can be performed. Spatial prediction (or "intra prediction") can use pixels from already-encoded neighboring blocks within the same video picture / slice to predict the current video block. Spatial prediction can reduce the spatial redundancy inherent in the video signal. Temporal prediction (also referred to as "inter prediction" or "motion-compensated prediction") uses pixels from already-encoded video pictures to predict the current video block. Temporal prediction can reduce the temporal redundancy inherent in the video signal. The temporal prediction signal for a given video block can be signaled by one or more motion vectors indicating the direction and amount of motion between the current block and its reference blocks. Additionally, if multiple reference pictures are supported (e.g., for recent video coding standards such as H.264 / AVC or HEVC), a reference index for each video block can also be sent. The reference index is used to identify which reference picture within the reference picture buffer (164) the temporal prediction signal is from. After spatial and / or temporal prediction, a mode decision block (180) within the encoder can select the best prediction mode, for example, based on a rate-distortion optimization method. Then, the prediction block can be subtracted from the current video block (116), and the prediction residual is decorrelated by using transform (104) and quantization (106) to achieve the target bitrate. The quantized residual coefficients can be inverse quantized (110) and inverse transformed (112) to form a reconstructed residual, which can then be added back to the prediction block (126) to form a reconstructed video block. Further in-loop filtering (such as a deblocking filter and an adaptive loop filter) can be applied (166) to the reconstructed video block before the reconstructed video block is placed in the reference picture buffer (164) and used to encode future video blocks. To form the output video bitstream 120, the coding mode (intra or inter), prediction mode information, motion information, and quantized residual coefficients can all be fed into an entropy coding unit (108) to be further compressed and packetized to form the bitstream.

[0034] Block-based decoding

[0035] Figure 2A main block diagram of a block-based video decoder 200 is given. At the entropy decoding unit 208, the video bitstream 202 is unpacked and entropy decoded. The coding mode and prediction information are sent to the spatial prediction unit 260 (if intra-coded) or the spatial prediction unit 262 (if inter-coded) to form a prediction block. The residual transform coefficients are sent to the inverse quantization unit 210 and the inverse transform unit 212 to reconstruct the residual block. Then, at 226, the prediction block and the residual block are added together. The reconstructed block can be further loop-filtered before being stored in the reference picture buffer 264. Then, the reconstructed video in the reference picture buffer can be sent out to drive a display device and be used to predict future video blocks.

[0036] In modern video codecs, bidirectional motion compensated prediction (MCP) is well-known for its efficiency in removing temporal redundancy by exploiting the temporal correlation between pictures and has been widely used in most state-of-the-art video codecs. However, the dual prediction signal is simply formed by merging the single prediction signals using a weight value equal to 0.5. From the perspective of merging the single prediction signals, this may not necessarily be optimal, especially in some cases where the illuminance changes rapidly from one reference picture to another. Therefore, several prediction techniques have been developed to compensate for the change of illuminance over time by applying some global or local weight and offset values to each sample value in the reference pictures.

[0037] Weighted dual prediction

[0038] Weighted dual prediction is an encoding tool mainly used to compensate for the illuminance change over time (such as fade transitions) when performing motion compensation. For each slice, two sets of multiplicative weight values and additive offset values are explicitly indicated and applied separately to the motion compensated prediction, one set at a time for each reference list. This technique is most effective when the illuminance changes linearly and evenly from one picture to another.

[0039] Local illuminance compensation

[0040] In local illuminance compensation, the parameters (two pairs of multiplicative weight values and additive offset values) can be adaptively modulated block by block. Different from weighted dual prediction that indicates these parameters at the slice level, this technique relies on adapting the optimized parameters to the illuminance change between the reconstructed signal of the template (T C ) and the prediction signals of the template (T0 and T1). The final parameters can be optimized by separately minimizing the illuminance difference between T C and T0 (for the first pair of weights and offsets) and between T C and T1 (for the second pair of weights and offsets). Then, a motion compensation process similar to weighted dual prediction can be performed using the obtained parameters.

[0041] Effect of Illumination Change

[0042] Illumination changes in space and time can seriously affect the performance of motion compensation prediction. As Figure 4 shown, when the illumination degenerates along the time direction, motion compensation prediction cannot provide good performance. For example, when an object sample travels along the time period from t-3 to t, and the intensity value of the sample changes from v t-3 to v t along its motion trajectory. Assuming that the sample is to be predicted at the t-th picture, its predicted value will be limited within v t-3 and v t-1 , resulting in poor motion compensation prediction. The above weighted double prediction and local illumination compensation techniques cannot completely solve this problem. Since the illumination fluctuates violently within a picture, weighted double prediction fails. Due to the low illumination correlation between a block and its associated template block, local illumination compensation sometimes produces poor estimates of weight and offset values. These examples show that both global descriptions and template-based local descriptions are insufficient to represent illumination changes in space and time.

[0043] Exemplary Embodiments

[0044] The exemplary embodiments described herein can improve the prediction efficiency of weighted motion compensation prediction. In some embodiments, systems and methods for general multi-hypothesis prediction are proposed, which linearly combine multi-hypothesis prediction signals by using motion compensation prediction and block-level weight values. In some embodiments, a general double prediction framework is described by using weight values. In some embodiments, a finite set of weights is used at the sequence level, picture level, and slice level, and the construction process for this set of weights is described. In some embodiments, the weight values are determined based on a given set of weights and optimized while considering the signals of the current block and its reference blocks. An exemplary encoding method for signaling the weight values is described. An exemplary encoder search criterion is described for the motion estimation process for the proposed prediction, and the proposed prediction process used in combination with the disclosed temporal prediction techniques is described.

[0045] In the present disclosure, systems and methods for temporal prediction by using general multi-hypothesis prediction are described. Referring to Figure 5 and Figure 9An exemplary encoder and decoder using general dual prediction are described. The systems and methods disclosed herein are organized as follows. The section "General multi-hypothesis prediction" describes exemplary embodiments using general multi-hypothesis prediction. The section "General dual prediction" describes an exemplary framework and prediction process regarding general dual prediction. The sections "Construction of weight sets" and "Weight index coding" respectively describe an exemplary construction process of weight sets and describe an exemplary technique for signaling the weight selection within the set. In the section "Extension to advanced temporal prediction techniques", systems and methods for combining the exemplary proposed prediction methods with advanced inter-frame prediction techniques (including local illumination compensation and weighted dual prediction, merge mode, overlapping block motion compensation, affine prediction, bidirectional optical flow, and a decoder-side motion vector derivation technique called frame rate up-conversion dual prediction) are described. In the section "GBi prediction search strategy", an exemplary encoder-only method for enhancing the efficiency of the exemplary prediction method is described.

[0046] General multi-hypothesis prediction

[0047] The exemplary systems and methods described herein employ general multi-hypothesis prediction. General multi-hypothesis prediction is described as a general form of multi-hypothesis prediction to provide an estimate of the intensity value of a pixel based on a linear combination of multiple motion-compensated prediction signals. General multi-hypothesis prediction can take advantage of them by combining multiple predictions with different amounts. For the purpose of accurate estimation, the motion-compensated prediction signals can be processed by a predefined function f(·) (e.g., gamma correction, local illumination correction, dynamic range conversion), and then can be linearly combined. General multi-hypothesis prediction can be described with reference to Equation (1):

[0048]

[0049] where P[x] represents the final prediction signal of sample x at picture position x, w i represents the weight value applied to the i-th motion hypothesis from the i-th reference picture, P i [x + v i is the motion-compensated prediction signal of x using the motion vector (MV) v i , and n is the total number of motion hypotheses.

[0050] For motion compensated prediction, one factor to consider is how to balance the accuracy of the motion field and the required motion overhead to achieve maximum rate-distortion performance. An accurate motion field means better prediction, while the required motion overhead can sometimes be more important than the benefit of prediction accuracy. Thus, in an exemplary embodiment, the proposed video encoder can adaptively switch between different numbers of motion hypotheses n and can find the value of n that provides the best rate-distortion performance for each corresponding PU. To help explain how general multi-hypothesis prediction works, the value of n = 2 is chosen as an example in the following sections because two motion hypotheses are the most commonly used in current video coding standards. Nevertheless, other values of n can be alternatively used. To simplify the understanding of the exemplary embodiment, the function f(·) can be considered an identity function and thus is not explicitly given. It will be apparent to those skilled in the art that after understanding the present disclosure, the systems and methods disclosed herein can be applied to cases where f(·) is not an identity function.

[0051] Generalized Bi-prediction

[0052] As used herein, the term Generalized Bi-prediction (GBi) refers to a special case of general multi-hypothesis prediction where the number of motion predictions is limited to 2, i.e., n = 2. In this case, the predicted signal at sample x given by Equation (1) can be simplified to:

[0053] P[x] = w0*P0[x + v0] + w1*P1[x + v1] (2)

[0054] where w0 and w1 are weight values shared among all samples in the block. Based on this equation, a very large number of predicted signals can be generated by adjusting the weight values w0 and w1. Some configurations of w0 and w1 can result in the same predictions as traditional single-prediction and bi-prediction. For example, (w0, w1) = (1, 0) can be used to achieve single-prediction using reference list L0, (w0, w1) = (0, 1) can be used to achieve single-prediction using reference list L1, and (w0, w1) = (0.5, 0.5) can be used to achieve bi-prediction using two reference lists. In the cases of (w0, w1) = (1, 0) and (w0, w1) = (0, 1), only one set of motion information is signaled because the other set associated with a weight value equal to 0 does not contribute to the predicted signal P[x].

[0055] Flexibility regarding the values of w0 and w1 may incur a high signaling overhead cost, especially at high precision. To save signaling overhead, in some embodiments, a unit gain constraint may be applied, i.e., w0 + w1 = 1, such that only one weight value for a block is explicitly signaled for a GBi-coded PU. To further reduce the weight signaling overhead, the weight value may be signaled at the CU level rather than the PU level. For simplicity of explanation, in the discussion of the present disclosure, w1 is signaled, such that Equation (2) can be further simplified to:

[0056] P[x] = (1 - w1)*P0[x + v0] + w1*P1[x + v1]. (3)

[0057] In an example embodiment, to further limit the signaling overhead, frequently used weight values may be arranged within a set (hereinafter referred to as W L1 ), such that within a limited range, each weight value can be indicated by an index value, i.e., weight_idx points to the entry it occupies within W L1 ).

[0058] In an exemplary embodiment, for supporting the weighted averaging of two reference blocks, general dual prediction does not introduce an additional decoding burden. Since most current video standards (e.g., AVC, HEVC) support weighted dual prediction, the same prediction model can be adaptively adjusted for GBi prediction. In an exemplary embodiment, general dual prediction can be applied not only to traditional single prediction and dual prediction, but also to other advanced temporal prediction techniques, such as affine prediction, advanced temporal motion vector derivation, and bidirectional optical flow. These techniques all aim to obtain a more refined representation of the motion field at a finer unit (e.g., 4x4) at the cost of very low motion overhead. Affine prediction is a model-based motion field coding method, where the motion of each unit within a PU can be derived based on model parameters. Advanced temporal motion vector derivation involves deriving the motion of each unit from the motion field of a temporal reference picture. Bidirectional optical flow involves obtaining motion refinement for each pixel by using an optical flow model. Regardless of the size of the unit, once the weight value is specified at the block level, the given video codec can perform general dual prediction per unit by using those derived motions and the given weight values.

[0059] Exemplary encoders and decoders using general dual prediction will be described in more detail below.

[0060] Exemplary Encoder for General Dual Prediction

[0061] Figure 5 Block diagram of an exemplary video encoder adapted to perform general dual prediction. AndFigure 1 Similar to the video encoder shown, spatial prediction and temporal prediction are two basic pixel-domain prediction models within an exemplary video encoder. The spatial prediction model can be the same as the Figure 1 spatial prediction model shown. Figure 1 The temporal prediction model labeled "motion prediction" in can be replaced by a general bi-prediction (GBi) model 502. The general bi-prediction (GBi) model can be operable to combine two separate motion-compensated prediction (MCP) signals in a weighted average manner. As Figure 6 shown, the GBi model can implement a process to generate the following inter-frame reference signal. The GBi model can perform motion estimation within the reference picture(s) to search for two optimal motion vectors (MVs) pointing to two reference blocks, which can minimize the weighted bi-prediction error between the current video block and the bi-prediction prediction. The GBi model can obtain these two prediction blocks by performing motion compensation using those two optimal MVs. Subsequently, the GBi model can calculate the prediction signal of the general bi-prediction based on the weighted average of the two prediction blocks.

[0062] In some embodiments, all available weighting values can be specified within a single set. Since the weighting values would consume a large number of bits if signaled for two reference lists at the PU level (which means that two separate weighting values for each bi-prediction PU would be signaled), a unit gain constraint (the sum of the weight values equals 1) can be applied. Under this constraint, only a single weight value for each PU is signaled, and the other weight value can be derived by subtracting the signaled weight value from 1. For ease of explanation, in the present disclosure, the weight value associated with reference list L1 is signaled, and the set of weight values can be represented by W L1 For further saving signaling overhead, the weight values can be encoded by an index value weight_idx pointing to the entry position within W L1 By making appropriate assignments to W L1 conventional single prediction (weight equals 0 for one reference list and weight equals 1 for the other list) and conventional bi-prediction (weight values equal 0.5 for both reference lists

[0063] can both be represented within the GBi framework. In the special case of W L1 ={0, 0.5, 1}

[0064] the GBi model can achieve the same function as the Figure 1 motion prediction model depicted.

[0065] In addition to {0, 0.5, 1}, for W L1The additional weight values can be specified at the slice level, picture level, or sequence level by a non - negative integer extra_number_of_weights indicating their number, so that there are extra_number_of_weights + 3 individual weights within the GBi framework. More specifically, in an exemplary embodiment, when extra_number_of_weights is greater than 0, one of these additional weight values can be obtained block - by - block, depending on the control of an implicit_weight_flag given at the slice level, picture level, or sequence level. When this flag is set to equal 1, this particular weight value may not be signaled, but can be obtained as shown in Figure 7 by finding a weight value that minimizes the difference between the general - form dual - prediction signal of a directly adjacent inverse L - shape (referred to as a template) and the reconstructed signal of the template. The above - mentioned process of constructing W L1 can be performed by the weight set construction model 504.

[0066] In order to adapt the additional weight values within W L1 to pictures with highly dynamic illumination changes, two scaling factors (GBi_scaling_factors) can be applied and signaled at the picture level. With these two scaling factors, the weight set construction model can scale the values of the additional weights for GBi prediction. After inter - frame prediction (i.e., GBi prediction within the proposed video encoder) and intra - frame prediction, the original signal can be subtracted from the final prediction signal, and thus a final prediction residual signal for encoding can be generated.

[0067] In an exemplary proposed video encoder, block motion (motion vector and reference picture index) and weight value index are the only block - level information to be indicated for each intra - coded PU.

[0068] In an exemplary embodiment, the block motion information of GBi prediction can be encoded in the same way as a basic video codec. Except when weight_idx is associated with a weight equal to 0 or 1 (i.e., equivalent to the case of single prediction), two sets of motion information for the PU are signaled.

[0069] In an exemplary video encoder, the weight index encoding module 506 is used to binarize the weight_idx of each PU. The output of the weight index encoding module can be the unique binary representation binary_weight_idx of the weight_idx. Figure 8Shows the tree structure of an exemplary binarization scheme. Similar to traditional inter prediction, the first bit of binary_weight_idx can distinguish between single prediction (weight index associated with a weight value equal to 0 or 1) and dual prediction (weight index associated with a weight value other than 0 and 1 within W L1 for each inter PU. In the single prediction branch, another bit can be signaled to indicate whether the L0 reference list (weight index associated with a weight value equal to 0) or the L1 reference list (weight index associated with a weight value equal to 1) is being referenced. In the dual prediction branch, each leaf node can be assigned a unique weight index value associated with one of the remaining weight values (i.e., the non-0 and 1 weights within W L1 ). At the slice level or picture level, the exemplary video encoder can adaptively switch between several predefined assignment methods, or dynamically assign each weight to a unique leaf node on a per-PU basis based on the usage of weight values from previously encoded blocks. Generally, frequently used weight indices are assigned to leaf nodes closer to the root within the dual prediction branch, while other weight indices are instead assigned to deeper leaf nodes farther from the root. By traversing Figure 8 the tree within, each weight_idx can be converted to a unique binary_weight_idx for entropy coding.

[0070] Decoding Framework for Generalized Dual Prediction

[0071] Figure 9 Is a block diagram of a video decoder in some embodiments. Figure 9 The decoder can be operable to decode Figure 5 the bitstream produced by the video encoder shown within. The coding mode and prediction information can be used to obtain a prediction signal using spatial prediction or generalized dual prediction. For generalized dual prediction, block motion information and weight values can be received and decoded.

[0072] The weight index decoding module 902 can decode the weight index encoded by the weight index encoding module 506 within the proposed video encoder. This weight index decoding module 902 can reconstruct a similar Figure 8 tree structure as specified, and in the same manner as the proposed video encoder, each leaf node on the tree is assigned a unique weight_idx. In this way, the tree can be synchronized between the proposed video encoder and decoder. By traversing this tree, each received binary_weight_idx can find its associated weight_idx at a certain leaf node on the tree. Similar to Figure 5 the video encoder, the exemplary video decoder includes means for constructing a weight set W L1The weight set construction module 904. When implicit_weight_flag is equal to 1, one of the extra weight values within W L1 can be derived rather than being signaled explicitly, and all of the extra weight values within W L1 can be further scaled by using the scaling factors indicated by gbi_scaling_factors. After that, the reconstruction of the weight values can be done by retrieving the weight value pointed to by weight_idx from W L1 .

[0073] Depending on the weight value selection at each block, the decoder can receive one or two sets of motion information. When the reconstructed weight value is non-0 or 1, two sets of motion information can be received; otherwise (when it is 0 or 1), only one set of motion information associated with non-zero weights can be received. For example, if the weight value is equal to 0, then only the motion information for reference list L0 will be signaled; otherwise if the weight value is equal to 1, then only the motion information for reference list L1 will be signaled.

[0074] By utilizing the block motion information and the weight values, Figure 10 the general dual-prediction module 1050 shown can be operative to calculate a prediction signal for general dual-prediction based on a weighted average of two motion-compensated prediction blocks.

[0075] Depending on the coding mode, the spatial prediction signal or the general dual-prediction signal can be added to the reconstructed residual signal to obtain the reconstructed video block signal.

[0076] Construction of the weight set

[0077] The following will describe an exemplary construction process for the weight set W using explicitly signaled weights, weights derived at the decoder side, and scaled weights, as well as an exemplary pruning process for compressing the size of the weight set W L1 . L1 .

[0078] Explicit weight values

[0079] Explicit weight values can be signaled and managed hierarchically at each sequence level, picture level, and slice level. The weights specified at a lower level can replace those at a higher level. Assuming the number of explicit weights at a higher level is p and the number of explicit weights at a relatively lower level is q, when constructing the list of weight values at the lower level, the following replacement rules can be applied:

[0080] · When p > q, the last q weights at the higher level can be replaced by the q weights at the lower level.

[0081] · When p ≤ q, all the weights at the higher level are replaced with those specified at the lower level.

[0082] The number of explicit weight values can be indicated by extra_number_of_weights at each of the sequence level, picture level, and slice level. In some embodiments, at the slice level, the base weight set always contains three default values {0, 0.5, 1} for GBi to support traditional single prediction and dual prediction, so that a total of (extra_number_of_weights + 3) weights can be used for each block. For example, when the values of extra_number_of_weights given at the sequence level, picture level, and slice level are 2 (e.g., w A , w B ), 1 (e.g., w C ), and 3 (e.g., w D , w E , w F ) respectively, the available weight values at the sequence level, picture level, and slice level are {w A , w B}, {w A , w C}, and {0, 0.5, 1} ∪ {w D , w E , w F} respectively. In this example, the W L1 mentioned in the "Generalized Dual Prediction" section is the slice-level weight set.

[0083] Derivation process of implicit weight values

[0084] In some embodiments, the weight values in the slice-level weight set W L1 can be derived through template matching at the encoder and decoder without signaling. As Figure 7 shown, the implicit weight value can be derived by minimizing the difference between the prediction signals (T0 and T1) of the template using the motion information of the current block and the reconstructed signal of the template (i.e., T c ). This problem can be formulated as:

[0085] w * = argmin w ∑ x (T c [x] - (1 - w) * T0[x + v0] - w * T1[x + v1]) 2 (4)

[0086] where v0 and v1 are the motion vectors of the current block. Since formula (4) is a quadratic function, if T0 and T1 are not exactly the same, a closed-form expression for the derived weights can be obtained, that is:

[0087]

[0088] When the weight value of the current block signal is related to the weight value of the associated template prediction signal, the effectiveness of the method can be seen. However, this is not always guaranteed, especially when the pixels within the current block and their associated templates are in different moving objects. To maximize the prediction performance of GBi, when extra_number_of_weights ≥ 1, the signaling flag implicit_weight_flag can be used at the slice level, picture level, or sequence level to determine whether the implicit weights are used. Once this is set to 1, the last slice-level weight value within W L1 can be derived and thus does not need to be signaled. For example, the w F mentioned in the above "explicit weight value" section does not need to be signaled, and when implicit_weight_flag is equal to 1, the weights of the block can be implicitly derived.

[0089] Regarding the scaling process of the weight values

[0090] In some embodiments, the explicit weight values can be further scaled by using two scaling factors gbi_scaling_factors indicated at the picture level. Due to the possible highly dynamic illumination changes over time within the picture, the dynamic range of these weight values may not be sufficient to cover all these cases. Although weighted bi-prediction can compensate for the illumination differences between pictures, it is not guaranteed that it can always be enabled within the base video codec. Therefore, when weighted bi-prediction is not used, those scaling factors can be used to adjust the illumination differences across multiple reference pictures.

[0091] The first scaling factor can scale each explicit weight value within W L1 . By this, the prediction function of GBi in formula (3) can be expressed as:

[0092]

[0093] where α is the first scaling factor of the current picture, and w1' represents the scaled weight value (that is, α * (w1 - 0.5) + 0.5). The first formula in formula (6) can be expressed in the same form as formula (3). The only difference lies in the weight values applied to formula (6) and (3).

[0094] The second scaling factor can be used to reduce the illumination difference between the associated reference pictures P0 and P1. With this scaling factor, formula (6) can be further reformulated as:

[0095]

[0096] where s, s0, and s1 represent the second scaling factors signaled at the current picture and its two reference pictures, respectively. According to formula (7), an optimal assignment for the variable s is the average value of the samples within the current picture. Thus, after applying the second scaling factor, the average values of the reference pictures can be expected to be similar. Due to the commutative property, applying the scaling factors to P0 and P1 is equivalent to applying them to the weight values, and thus formula (7) can be reinterpreted as:

[0097]

[0098] Therefore, the construction process regarding the weight set can be expressed as a function of the explicit weights, implicit weights, scaling factors, and reference pictures. For example, the above slice-level weight set W L1 becomes {0, 0.5, 1} ∪ {(s / s1) * w D ’, (s / s1) * w E ’, (s / s1) * w F ’}, and the weight set for L0 becomes {1, 0.5, 1} ∪ {(s / s0) * (1 - w D ’), (s / s0) * (1 - w E ’), (s / s0) * (1 - w F ’)}, where s1 is the average sample value of the reference picture within the list L1 of the current block, and s0 is the average sample value of the reference picture within the list L0 of the current block.

[0099] Pruning of Weight Values

[0100] Exemplary embodiments are operable to further reduce the number of weight values within W L1 . Two exemplary methods for pruning weight values are described below. The first method operates in response to the motion compensation prediction result, and the second method operates based on weight values outside the range between 0 and 1.

[0101] Prediction-Based Method. When the motion vector of a given PU is available, not every weight can produce a substantially different dual prediction from other weights. Exemplary embodiments utilize this property to prune redundant weight values (which produce similar dual prediction signals) and retain only one weight among multiple redundant values, such that W L1More compact. To do so, a function can be used to calculate the similarity between two dual-prediction signals with different weight values. The function can be, but is not limited to, the cosine similarity function, which can operate as follows:

[0102]

[0103] where w (i) and w (j) are two independent weight values within W L1 , v0 and v1 are given dual-prediction motion information, and P[x; w, v0, v1] represents the same prediction function specified in formulas (3), (6), and (8) given w, v0, and v1. When the value of formula (9) is below a given threshold (indicated by the slice-level weight_pruning_threshold), one of the weights can be pruned depending on the slice-level syntax pruning_smaller_weight_flag. If the flag is set to equal 1, the pruning process can remove the smaller weight from W L1 between w (i) and w (j) . Otherwise (when the flag is set to equal 0), the larger one can be removed. In an exemplary embodiment, the pruning process can be applied to each pair of weight values within W L1 , and finally, there will be no two weight values within the final W L1 that produce similar dual-prediction signals. The similarity between two weight values can also be evaluated by using the sum of absolute transform differences (SATD). To reduce the computational complexity, the similarity can be evaluated by using two sub-sampled prediction blocks. For example, it can be calculated by taking advantage of sub-sampled rows or sub-sampled columns of samples in the horizontal and vertical directions.

[0104] Weight value-based method. Depending on the coding performance under different coding structures (e.g., hierarchical structure or low-latency structure), weight values outside the range between 0 and 1 (or simply out-of-range weights) can behave differently. To take advantage of this fact, an exemplary embodiment employs a set of sequence-level indices (weight_control_idx) to limit the use of out-of-range weights for each time layer individually. In these embodiments, each weight_control_idx is associated with all the pictures of a specific time layer. Depending on how the index is configured, the out-of-range weights can be conditionally used or pruned as follows:

[0105] · For weight_control_idx = 0, W L1 remains unchanged for the associated pictures.

[0106] · For weight_control_idx = 1, weights outside the range within W L1 are not available for associated pictures.

[0107] · For weight_control_idx = 2, weights outside the range within W L1 are available only for some associated pictures whose reference frames are purely from the past (e.g., low-delay configurations in HEVC and JEM).

[0108] · For weight_control_idx = 3, weights outside the range within W L1 are available for associated pictures only when the slice-level flag mvd_l1_zero_flag in HEVC and JEM is enabled.

[0109] Weight index coding

[0110] Exemplary systems and methods for binarizing and codeword assignment for weight index coding are described in more detail below.

[0111] Binarization process for weight index coding

[0112] In an exemplary embodiment, each weight index (weight_idx) can be converted to a unique binary representation (binary_weight_idx) by a systematic code before entropy coding. For illustrative purposes, Figure 8 a tree structure of the proposed binarization method is shown in. The first bit of binary_weight_idx is used to distinguish between single prediction (associated with weights equal to 0 or 1) and double prediction. Another bit signaled within the single prediction branch indicates which of the two reference lists is referenced, reference list L0 (associated with weight indices pointing to weight values equal to 0) or reference list L1 (associated with weight indices pointing to weight values equal to 1). In the double prediction branch, each leaf node can be assigned a unique weight index associated with one of the remaining weight values in W L1 that are not 0 or 1. Exemplary video codecs can support various systematic codes to binarize the double prediction branch, such as truncated unary codes (e.g., Figure 8 ) and exponential Golomb codes. Exemplary techniques for assigning a unique weight_idx to each leaf node within the double prediction branch are described in more detail below. By looking up this tree structure, each weight index can be mapped to a unique codeword or recovered from a unique codeword (e.g., binary_weight_idx).

[0113] Adaptive codeword assignment for weight index coding

[0114] In an exemplary binary tree structure, each leaf node corresponds to a codeword. To reduce the signaling overhead of the weight index, various adaptive codeword assignment methods can be used to map each leaf node within a dual-prediction branch to a unique weight index. Exemplary methods include predetermined codeword assignment, block-adaptive codeword assignment, time-layer-based codeword assignment, and time-delay CTU-adaptive codeword assignment. These exemplary methods can update the codeword assignment within the dual-prediction branch based on the presence of weight values used within previously encoded blocks. Frequently used weights can be assigned to codewords with shorter lengths (e.g., shallower leaf nodes within the dual-prediction branch), while other weights can be assigned to codewords with relatively longer lengths.

[0115] 1) Predetermined codeword assignment. By using predetermined codeword assignment, a constant codeword assignment can be provided for the leaf nodes within a dual-prediction branch. In this method, the weight index associated with the weight 0.5 can be assigned the shortest codeword, i.e., node i in, for example, Figure 8 . Weight values other than 0.5 can be divided into two sets: Set 1 contains all values greater than 0.5, which are sorted in ascending order; Set 2 contains all values less than 0.5, which are sorted in descending order. These two sets are then interleaved to form Set 3, which can start with Set 1 or Set 2. All the remaining codewords with lengths from short to long are sequentially assigned to the weight values within Set 3. For example, when the set of all possible weight values within a dual-prediction branch is {0.1, 0.3, 0.5, 0.7, 0.9}. Set 1 is {0.7, 0.9}, Set 2 is {0.3, 0.1}, and if the interleaving starts with Set 1, Set 3 is {0.7, 0.3, 0.9, 0.1}. Codewords with lengths from short to long are sequentially assigned to 0.5, 0.7, 0.3, 0.9, and 0.1.

[0116] Some codecs may discard a motion vector difference when two sets of motion information are sent, in which case the assignment can be changed. For example, this behavior can be found in HEVC through the slice-level flag mvd_l1_zero_flag. In this case, an alternative codeword assignment can assign the weight index associated with weight values greater than and close to 0.5 (e.g., w + ). Subsequently, the weight index associated with the nth smallest (or largest) weight value among those greater than (or less than) w + can be assigned the (2n + 1)th shortest (or 2nth shortest) codeword. Based on the previous example, codewords with lengths from short to long can be sequentially assigned to 0.7, 0.5, 0.9, 0.3, and 0.1. Figures 11A - 11B shows the final assignment for these two examples.

[0117] 2) Block adaptive codeword assignment using causal - neighboring weights. The weight values used within causal neighboring blocks can be related to the weight values used for the current block. Based on this knowledge and a given codeword assignment method (e.g., constant assignment or alternative assignment), the weight indices found from causal neighboring blocks can be promoted to the leaf nodes with shorter codeword lengths in the dual - prediction branches. Similar to the construction process of the motion vector prediction list, causal neighboring blocks can be accessed in the sorting order shown in Figure 12A and at most two weight indices can be promoted. As can be seen from the figure, from the bottom - left block to the left block, the first available weight index (if any) can be promoted to have the shortest codeword length; from the top - right block to the top - left block, the first available weight index (if any) can be promoted to have the second - shortest codeword length. For other weight indices, according to their codeword lengths in the original given assignment, they can be assigned to the remaining leaf nodes from the shallowest to the deepest. Figure 12B Fig. Figure 12B gives an example of how a given codeword assignment adjusts itself to adapt to causal neighboring weights. In this example, constant assignment can be used, and weight values equal to 0.3 and 0.9 can be promoted.

[0118] 3) Temporal - layer - based codeword assignment. In an exemplary method of using temporal - layer - based codeword assignment, the proposed video encoder can adaptively switch between constant codeword assignment and alternative codeword assignment. By using the weight indices of previously encoded pictures at the same temporal layer or by using the same QP value, the optimal codeword assignment method with the minimum expected codeword length of the weight indices can be found by:

[0119]

[0120] where \(L\) m (w) represents the codeword length of \(w\) using a certain codeword assignment method \(m\), \(\Omega\) is the set of weights only set for dual - prediction, and \(Prob\) k (w) represents the cumulative probability of \(w\) over \(k\) pictures at the temporal layer. Once the best codeword assignment method is determined, it can be applied to encode the weight indices or parse the binary codeword indices for the current picture.

[0121] Several different methods can be considered to accumulate the use of weight indices over temporal pictures. An exemplary method can be formulated as a common formula:

[0122]

[0123] where \(w\) i is a certain weight within \(W\) L1 , \(Count\) j(w) represents the presence of a certain weight value at the j-th picture of the temporal layer, n determines the number of pictures to be stored most recently, and λ is the forgetting term. Since n and λ are parameters only for the encoder, they can be adaptively adjusted at each picture to adapt to various encoding conditions, such as n = 0 for scene changes and smaller λ for motion videos.

[0124] In some embodiments, the selection of the codeword assignment method can be explicitly indicated by using slice-level syntax elements. Thus, the decoder does not need to maintain the use of the weight index over time, and thus the parsing of the dependency on the weight index on the temporal pictures can be completely avoided. This method can also improve decoding robustness.

[0125] 4) CTU adaptive codeword assignment. Switching between different codeword assignment methods based solely on the weight usage of the previously encoded pictures may not always match well with the codeword assignment of the current picture. This can be attributed to the lack of consideration of the weight usage of the current picture. In an exemplary embodiment using CTU adaptive codeword assignment, Prob k (w i ) can be updated based on the weight usage of the encoded blocks within the current CTU row and directly above the CTU row. Assuming the current picture is the (k + 1)-th picture within the temporal layer, then Prob k (w i ) can be updated per CTU as follows:

[0126]

[0127] where B represents the set of encoded CTUs within the current CTU row and directly above the CTU row, and Count' j (w) represents the presence of a certain weight value at the j-th CTU collected within the set B. Once Prob k (w i ) is updated, it can be applied to formula (10), and the optimal codeword assignment method can thus be determined.

[0128] Extension of advanced temporal prediction techniques

[0129] The embodiments discussed below are used to extend the application of general dual prediction with other encoding techniques (including local illumination compensation, weighted dual prediction, merge mode, bidirectional optical flow, affine motion prediction, overlapping block motion compensation, and frame rate up-conversion dual prediction).

[0130] Local illumination compensation and weighted dual prediction

[0131] Exemplary general dual prediction techniques may be performed based on local illumination compensation (IC) and / or weighted dual prediction or other techniques. Both IC and weighted dual prediction are operable to compensate for illumination changes on a reference block. One difference between them is that when using IC, the weights (c0 and c1) and offset values (o0 and o1) are derived by performing template matching block by block; when using weighted dual prediction, these parameters are signaled explicitly slice by slice. By using these parameters (c0, c1, o0, o1), the prediction signal of GBi can be calculated as:

[0132]

[0133] where the scaling process of the weight values described in the "Scaling Process for Weight Values" section above may be applied. When this scaling process is not applied, the prediction signal of GBi can be calculated as:

[0134] P[x] = (1 - w1)*(c0*P0[x + v0] + o0) + w1*(c1*P1[x + v1] + o1). (14)

[0135] For example, the use of the combined prediction processes such as those given in formulas (13) or (14) may be signaled at the sequence level, picture level, or slice level. Signaling may be performed separately for the combination of GBi and IC and for the combination of GBi and weighted dual prediction. In some embodiments, the combined prediction process of formula (13) or (14) may be applied only when the weight value (w1) is not equal to 0, 0.5, or 1. Specifically, when the use of the combined prediction process is active, the value of the block-level IC flag (which can be used to indicate the use of IC) determines whether GBi (w1≠0, 0.5, 1) is combined with IC. Otherwise, when the combined prediction process is not used, GBi (w1≠0, 0.5, 1) and IC are performed as two independent prediction modes, and for each block, the block-level IC flag does not need to be signaled and can thus be inferred as 0.

[0136] In some embodiments, whether GBi can be combined with IC or weighted dual prediction can be signaled by using a set of sequence parameters (SPS), a set of picture parameters (PPS), or high-level syntax at the slice header by using flags (such as, GBi_IC_combination_flag (gbi_ic_comb_flag) and GBi_weighted_dual_prediction_combination_flag (gbi_wb_comb_flag)). In some embodiments, if gbi_ic_comb_flag is equal to 0, GBi and IC are not combined, and for any dual-predicted coding unit, the GBi weight value (w1≠0, 0.5, 1) and the IC flag will not coexist. For example, in some embodiments, if a GBi weight (w1≠0, 0.5, 1) is signaled for a coding unit, then no IC flag will be signaled, and the flag value can be inferred to be 0; otherwise the IC flag can be signaled explicitly. In some embodiments, if gbi_ic_comb_flag is equal to 1, GBi and IC can be combined, and the GBi weight and the IC flag can be signaled independently for a coding unit. The same syntax can be applied to gbi_wb_comb_flag.

[0137] Merge mode

[0138] In some embodiments, the merge mode can be used to infer not only motion information from causally adjacent blocks, but also to infer the weight index of the block simultaneously. The access order to causally adjacent blocks (as Figure 13 illustrated) can be the same as that specified in HEVC, where spatial blocks are accessed in the order of the left block, the upper block, the upper-right block, the lower-left block, and the upper-right block, while temporal blocks are accessed in the order of the lower-right block and the center block. In some embodiments, up to five merge candidates can be constructed by using up to four blocks from spatial blocks and up to one block from temporal blocks. Given a merge candidate, the GBi prediction process specified in formulas (3), (8), (13), or (14) can be applied. It should be noted that the weight index does not need to be signaled because it can be inferred from the weight information of the selected merge candidate.

[0139] In the JEM platform, an additional merge mode called Advanced Temporal Motion Vector Prediction (ATMVP) can be provided. In some embodiments of the present disclosure, ATMVP can be combined with GBi prediction. In ATMVP, the motion information of each 4x4 unit within a CU can be derived from the motion field of a temporal reference picture. In an exemplary embodiment using ATMVP, when the GBi prediction mode is enabled (for example, when extra_number_of_weights is greater than 0), the weight index of each 4x4 unit can also be inferred from the weight index of the corresponding temporal block within the temporal reference picture.

[0140] Bidirectional optical flow

[0141] In some embodiments, the weight value of GBi can be applied to the bidirectional optical flow (BIO) model. Based on the motion-compensated prediction signals (P0[x+v0] and P1[x+v1]), BIO can estimate the offset value o BIO [x] to reduce the difference between two corresponding samples within L0 and L1 (in accordance with their spatial vertical and horizontal gradient values). To combine this offset value with the GBi prediction, Equation (3) can be reformulated as:

[0142] P[x] = (1 - w1)*P0[x+v0] + w1*P1[x+v1] + o BIO [x], (15)

[0143] where w1 is the weight value used to perform the GBi prediction. This offset value can also be applied as an additional offset to other GBi variants after the prediction signals within P0 and P1 are scaled, similar to Equation (8), (13), or (14).

[0144] Affine prediction

[0145] In an exemplary embodiment, the GBi prediction can be combined with the affine prediction in a manner similar to the extension of the traditional dual prediction. However, there are differences in the basic processing units used to perform motion compensation. Affine prediction is a model-based motion field derivation technique for forming a fine-grained motion field representation of a PU, where the motion field representation of each 4x4 unit can be derived based on unidirectional or bidirectional transform motion vectors and given model parameters. Since all motion vectors point to the same reference picture, there is no need to adjust the weight value to adapt to each 4x4 unit. Therefore, the weight value can be shared between each unit and only one weight index for the PU needs to be signaled. By utilizing the motion vectors and weight values at the 4x4 unit, GBi is performed on a per-unit basis, and thus the same Equations (3), (8), (13), and (14) can be directly applied without any modification.

[0146] Overlapped block motion compensation

[0147] Overlapped block motion compensation (OBMC) is a method for providing a prediction of the intensity value of a sample based on a motion-compensated signal derived from the motion vectors of the sample itself and those within its causal neighbors. In an exemplary embodiment of GBi, the weight value can also be considered in the motion compensation for OBMC. Figure 14An example is shown where sub-block B1 within the current block has three motion-compensated prediction blocks, each motion-compensated prediction block being formed by using motion information from N1, N5, or B1 itself and a weight value, and the final prediction signal for B1 can be a weighted average of the three motion-compensated prediction blocks.

[0148] Frame Rate Up-Conversion

[0149] In some embodiments, GBi can operate together with Frame Rate Up-Conversion (FRUC). Two different modes can be used for FRUC. If the current picture falls between the first reference picture in L0 and the first reference picture in L1, the dual-prediction mode can be used. If both the first reference picture in L0 and the first reference picture in L1 are forward reference pictures or backward reference pictures, the single-prediction mode can be used. The dual-prediction case within FRUC will be discussed in detail below. In JEM, equal weights (i.e., 0.5) can be used for FRUC dual-prediction. Although the quality of the two predictors within FRUC dual-prediction may vary, combining two predictors with unequal prediction quality by using equal weights may be sub-optimal. Due to the use of unequal weights, the use of GBi can improve the final dual-prediction quality. In an exemplary embodiment, for blocks encoded using FRUC dual-prediction, the weight value of GBi can be derived and thus does not need to be signaled. For each 4x4 sub-block within a PU, each weight value within W L1 can be individually evaluated through the MV derivation process of FRUC dual-prediction. The weight value that results in the minimum bilateral matching error for the 4x4 block (i.e., the absolute difference between two uni-directional motion-compensated predictors associated with two reference lists) can be selected.

[0150] In an exemplary embodiment, FRUC dual-prediction is a decoder-side MV derivation technique that derives the MV by using bilateral matching. For each PU, a list of candidate MVs collected from causally adjacent blocks can be formed. Under the assumption of constant motion compensation, each candidate MV can be linearly projected onto the first reference picture within the other reference list, where the scaling factor of the projection is set to be proportional to the temporal distance between the reference picture (e.g., at time t0 or t1) and the current picture (t c )). As an example, where v0 is the candidate MV associated with reference list L0, v1 is calculated as v0*(t1 - t Figure 15 ) / (t0 - t c ). Therefore, the bilateral matching error can still be calculated for each candidate MV, and the initial MV that achieves the minimum bilateral matching error can be selected from the candidate list. This initial MV can be denoted as v0 c . From this initial MV v0 INIT . From this initial MV v0 INITStarting from the pointed position, decoder-side motion estimation can be performed to find the MV within a predefined search range, and the MV that can achieve the minimum bilateral matching error can be selected as the PU-level MV. Assuming v1 is the projected MV, the optimization process can be formulated as:

[0151]

[0152] where FRUC dual prediction can be combined with GBi, and the search process in formula (16) can be reformulated using the weight value w within W L1 That is:

[0153]

[0154] The PU-level v0 can be further individually refined by using the same bilateral matching in formula (17) for each 4x4 sub-block within the PU, as shown in formula (18):

[0155]

[0156] For each available weight value w within W L1 Formula (18) can be evaluated, and the weight value that minimizes the bilateral matching error can be selected as the optimal weight. At the end of the evaluation process, each 4x4 sub-block within the PU has its own dual prediction MV and weight value for performing general dual prediction. The complexity of this exhaustive search method may be high because the weights and motion vectors are searched jointly. In another embodiment, the search for the optimal motion vector and the optimal weight can be performed in two steps. In the first step, the motion vector of each 4x4 block can be obtained by using formula (18) and setting w to an initial value (e.g., w = 0.5). In the second step, the optimal weight can be searched given the optimal motion vector.

[0157] In yet another embodiment, to improve the motion search accuracy, three steps can be applied. In the first step, the initial weight can be searched by using the initial motion vector v0 INIT This initial optimal weight can be denoted as w INIT . In the second step, the motion vector of each 4x4 block can be obtained by using formula (18) and setting w to w INIT . In the third step, the final optimal weight can be searched given the optimal motion vector.

[0158] Through formulas (17) and (18), the goal is to minimize the difference between two weighted predictors respectively associated with two reference lists. For this purpose, negative weights may be inappropriate. In one embodiment, the GBi mode based on FRUC will only evaluate weight values greater than zero. To reduce complexity, the calculation of the sum of absolute differences can be performed by using partial samples within each sub-block. For example, the sum of absolute differences can be calculated by using only the samples located at even-numbered rows and columns (or, alternatively, odd-numbered rows and columns).

[0159] GBi prediction search strategy

[0160] Initial reference lists for dual prediction search

[0161] The following describes a method for improving the prediction performance of GBi by determining which of the two reference lists should be searched first in the motion estimation (ME) stage of dual prediction. As in traditional dual prediction, two motion vectors respectively associated with reference list L0 and reference list L1 need to be determined to minimize the ME stage cost, that is:

[0162] Cost(t i ,u j )=∑ x |I[x]-P[x]|+λ*Bits(t i ,u j ,weight index) (19)

[0163] where I[x] is the original signal of sample x located at x within the current picture, P[x] is the prediction signal of GBi, and t i and u j are motion vectors respectively pointing to the i-th reference picture within L0 and the j-th reference picture within L1, λ is the Lagrangian parameter used in the ME stage, and the Bits(·) function estimates the number of bits used to encode the input variable. Each of formulas (3), (8), (13), and (14) can be applied to replace P[x] within formula (19). For the purpose of simplifying the description, in the following process, formula (3) can be considered as an example. Therefore, the cost function within formula (19) can be rewritten as:

[0164]

[0165] Since there are two parameters (t i and u j ) to be determined, an iterative process can be adopted. The first such process can follow the following rules:

[0166] 1. Use Optimal motion within, optimize t i ,

[0167] 2. Use Optimal motion within, optimize u j ,

[0168] 3. Repeat steps 1 and 2 until t i and u j no longer change or reach the maximum number of iterations. The second exemplary iterative process can be carried out as follows:

[0169] 1. Use Optimal motion within, optimize u j ,

[0170] 2. Use Optimal motion within, optimize t i ,

[0171] 3. Repeat steps 1 and 2 until u j and t i no longer change or reach the maximum number of iterations.

[0172] The choice of which iterative process can depend solely on the ME stage costs of t i and u j , that is:

[0173]

[0174] Where the ME stage cost function can be as follows:

[0175] Cost(t i ) = ∑ x |I[x] - P0[x + t i | + λ * Bits(t i ). (22)

[0176] Cost(u j ) = ∑ x |I[x] - P1[x + u j | + λ * Bits(u j ). (23)

[0177] However, this initialization process may not be optimal when 1 - w1 is not equal to w1. A typical example is when one of the weight values is extremely close to 0, such as w1 = lim w→0 w, and the ME stage cost of its associated motion happens to be lower than the other. In this case, formula (20) degenerates to:

[0178] Cost(t i ,u j )=∑ x |I[x]-P0[x+t i |+λ*Bits(t i ,u j ,weight index). (24)

[0180] For u j the cost incurred will not help the predicted signal, and ultimately lead to poor search results for GBi. In the present disclosure, the magnitude of the weight value can be used to replace formula (21), specifically:

[0181]

[0182] Binary search for weight index

[0183] Since the number of weight values to be evaluated may introduce additional complexity to the encoder, the exemplary embodiment adopts a binary search method to prune less likely weight values early in the encoding. In one such search method, traditional single prediction (associated with 0 and 1 weights) and double prediction (associated with 0.5 weight) can be performed at the beginning, and the weight values within W L1 can be divided into 4 groups, that is, A = [w min ,0], B = [0,0.5], C = [0.5,1] and D = [1,w max . w min and w max represent the minimum weight value and the maximum weight value within W L1 respectively, and without loss of generality, it can be assumed that w min <0 and w max > 1. The following rules can be applied to determine the range of possible weight values.

[0184] · If w = 0 gives a better ME stage cost than w = 1, the following rules can be applied:

[0185] o If w = 0.5 gives a better ME stage cost than w = 0 and w = 1, the weight set W (0) can be formed based on the weight values within B.

[0186] o Otherwise, W (0) can be formed based on the weight values within A.

[0187] · Otherwise (if w = 1 gives a better ME stage cost than w = 0), the following rules can be applied:

[0188] o If the ME stage cost is better when w = 0.5 than when w = 0 and w = 1, the weight set W can be formed based on the weight values within C (0) .

[0189] o Otherwise, W can be formed based on the weight values within D (0) .

[0190] After forming W (0) , w min and w max values can be reset according to the minimum and maximum values within W (0) respectively. If W (0) is associated with A and D, the ME stage cost of w min within A and the ME stage cost of w max within D can be calculated respectively.

[0191] The iterative process can be operated to keep updating W (k) until there are more than 2 weight values remaining in the set during the k-th iteration. Assuming the process is for k iterations, the iterative process can be specified as follows:

[0192] 1. Execute GBi using the weight value w min closest to (w max + w middle ).

[0193] 2. If w middle gives a better ME stage cost than w min and w max , a recursive process can be called for W (k+1) to separately test [w min , w middle and [w middle , w max , and the iterative process jumps to step 6.

[0194] 3. Otherwise, if w middle gives a worse ME stage cost than w min and w max , the iterative process terminates.

[0195] 4. Otherwise, if w min gives a better ME stage cost than w max , form W min based on the weight values within [w middle , w (k+1) , and the iterative process jumps to step 6.

[0196] 5. Otherwise (if w min is better than w maxgives a worse ME stage cost), then based on [w middle , w max within the weight values to form W (k+1) , and the iteration process jumps to step 6.

[0197] 6. If the number of remaining weight values within W (k+1) is greater than 2, then w min and w max can be reset according to the maximum and minimum values within W (k+1) , and the iteration process jumps to step 1; otherwise, the iteration process terminates.

[0198] After the iteration process stops, the weight value that can achieve the lowest ME stage cost among all test values can be selected to perform general dual prediction.

[0199] Weight value estimation for non - 2Nx2N partitions

[0200] In some embodiments, after testing each weight value for 2Nx2N partitions, the best - performing weight values other than 0, 0.5, and 1 can be used as an estimate of the optimal weight value for non - 2Nx2N partitions. In some embodiments, assuming there are n unique estimates, these n unique estimates as well as the weight values equal to only 0, 0.5, and 1 can be evaluated for non - 2Nx2N partitions.

[0201] Partition size estimation for non - 2Nx2N partitions

[0202] In some embodiments, not all non - 2Nx2N partitions are tested by the exemplary video encoder. Non - 2Nx2N partitions can be divided into two sub - categories: symmetric motion partitions (SMP) with 2NxN and Nx2N partition types and asymmetric motion partitions (AMP) with 2NxnU, 2NxnD, nLx2N, and nRx2N partition types. If the rate - distortion (R - D) cost of the partitions within SMP is less than the distortion cost of 2Nx2N, then some partition types within AMP can be evaluated at the encoder. The decision on which partition types within AMP to test can depend on which of 2NxN and Nx2N can exhibit better performance in terms of R - D cost. If the rate - distortion cost of 2NxN is smaller, then the partition types 2NxnU and 2NxnD can be further examined, otherwise (if the cost of Nx2N is smaller), then the partition types nLx2N and nRx2N can be further examined.

[0203] Fast parameter estimation for multi - channel coding

[0204] In an exemplary embodiment of using a multi-channel encoder, prediction parameters (e.g., block motion and weight values) optimized from an earlier encoding channel can be adopted as initial parameter estimates at a subsequent encoding channel. In this encoder, encoding blocks partitioned from a picture can be predicted and encoded two or more times, which ultimately leads to a significant increase in encoding complexity. A technique to reduce this complexity is to cache the optimized prediction parameters from the initial encoding channel and use them as initial parameter estimates for further refinement within subsequent encoding channels. For example, if the inter-frame prediction mode happens to be the best mode at the initial channel, the encoder can evaluate only the inter-frame prediction mode at the remaining encoding channels. In some embodiments, caching can be performed for prediction parameters related to GBi, such as the selection of weight values within W L1 the selection of weight values within, the dual-prediction MVs associated with the selected weight values, the IC flag, the OBMC flag, the integer motion vector (IMV) flag, and the coded block flag (CBF). In such embodiments, the values of these cached parameters can be reused or refined at subsequent encoding channels. More specifically, when the above dual-prediction MVs are adopted, these MVs can be used as the initial search positions for dual-prediction searches. Subsequently, they can be refined during the motion estimation stage and then used as the initial search positions for the next encoding channel.

[0205] Exemplary bitstream communication architecture

[0206] Figure 16 is a schematic diagram showing an example of the structure of an encoded bitstream. The encoded bitstream 1000 includes a plurality of NAL (Network Abstraction Layer) units 1001. The NAL units can contain encoded sample data (e.g., encoded slices 1006) or high-level syntax metadata (e.g., parameter set data, slice header data 1005, or supplementary enhancement information data 1007 (which can be referred to as SEI messages)). A parameter set is a high-level syntax structure containing basic syntax elements, where the basic syntax elements can be applied to multiple bitstream layers (e.g., video parameter set 1002 (VPS)), or to an encoded video sequence within a layer (e.g., sequence parameter set 1003 (SPS)), or to multiple encoded pictures within an encoded video sequence (e.g., picture parameter set 1004 (PPS)). The parameter set can be sent together with the encoded pictures of the video bitstream or in other ways (including out-of-band transmission using a reliable channel, hard coding, etc.). The slice header 1005 is also a high-level syntax structure, which can contain some relatively small or picture-related information that is only related to certain slices or picture types. The SEI message 1007 carries information that is not necessarily required for decoding processing but can be used for various other purposes (e.g., picture output timing or display, and loss detection and concealment).

[0207] Figure 17is a schematic diagram showing an example of a communication system. The communication system 1300 may include an encoder 1302, a communication network 1304, and a decoder 1306. The encoder 1302 may communicate with the network 1304 via a connection 1308, which may be a wired connection or a wireless connection. The encoder 1302 may be similar to a block-based video encoder Figure 1 . The encoder 1302 may include a single-layer codec (e.g., Figure 1 ) or a multi-layer codec. For example, the encoder 1302 may be a multi-layer (e.g., two-layer) scalable coding system that supports picture-level ILP. The decoder 1306 may communicate with the network 1304 via a connection 1310, which may be a wired connection or a wireless connection. The decoder 1306 may be similar to a block-based video decoder Figure 2 . The decoder 1306 may include a single-layer codec (e.g., Figure 2 ) or a multi-layer codec. As an example, the decoder 1306 may be a multi-layer (e.g., two-layer) scalable decoding system that supports picture-level ILP.

[0208] The encoder 1302 and / or the decoder 1306 may be incorporated into various wired communication devices and / or wireless transmit / receive units (WTRUs), such as digital televisions, wireless broadcast systems, network components / terminals, servers (e.g., content or network servers (e.g., Hypertext Transfer Protocol (HTTP) servers)), personal digital assistants (PDAs), laptop or desktop computers, tablet computers, digital cameras, digital recording devices, video game devices, video game consoles, cellular or satellite wireless telephones, and / or digital media players, etc., but are not limited thereto.

[0209] The communication network 1304 may be a suitable type of communication network. For example, the communication network 1304 may be a multiple access system that provides content (e.g., voice, data, video, messages, broadcasts, etc.) to multiple wireless users. The communication network 1304 enables multiple wireless users to access such content by sharing system resources including wireless bandwidth. As an example, the communication network 1304 may use one or more channel access methods, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal FDMA (OFDMA), and / or Single-Carrier FDMA (SC-FDMA), etc. The communication network 1304 may include multiple interconnected communication networks. The communication network 1304 may include the Internet and / or one or more private commercial networks, such as cellular networks, WiFi hotspots, and / or Internet Service Provider (ISP) networks, etc.

[0210] Figure 18System diagram of an exemplary WTRU that can implement the encoders or decoders described herein. As shown, the exemplary WTRU 1202 can include a processor 1218, a transceiver 1220, a transmit / receive component 1222, a speaker / microphone 1224, a keyboard or numeric keypad 1226, a display / touchpad 1228, a non-removable memory 1230, a removable memory 1232, a power supply 1234, a global positioning system (GPS) chipset 1236, and / or other peripheral devices 1238. It should be understood that the WTRU 1202 can also include any sub-combination of the foregoing components while remaining consistent with the embodiments. Further, a terminal integrated with an encoder (e.g., encoder 100) and / or a decoder (e.g., decoder 200) can include some or all of the components depicted in the Figure 18 WTRU 1202 depicted herein and described herein with reference to Figure 18 the WTRU 1202.

[0211] The processor 1218 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), a graphics processing unit (GPU), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 1218 can perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable the WTRU 1500 to operate in a wired and / or wireless environment. The processor 1218 can be coupled to the transceiver 1220, which can in turn be coupled to the transmit / receive component 1222. Although Figure 18 the processor 1218 and the transceiver 1220 are described as separate components, it should be understood that the processor 118 and the transceiver 1220 can also be integrated together in an electronic component and / or chip.

[0212] The transmit / receive component 1222 can be configured to transmit and / or receive signals to / from another terminal via an air interface 1215. For example, in one or more embodiments, the transmit / receive component 1222 can be an antenna configured to transmit and / or receive RF signals. As an example, in one or more embodiments, the transmit / receive component 1222 can be a radiator / detector configured to transmit and / or receive IR, UV, or visible light signals. In one or more embodiments, the transmit / receive component 1222 can be configured to transmit and / or receive both RF and optical signals. It should be understood that the transmit / receive component 1222 can be configured to transmit and / or receive any combination of wireless signals.

[0213] In addition, although inFigure 18 The transmit / receive component 1222 is described as a single component, but the WTRU 1202 can include any number of transmit / receive components 1222. More specifically, the WTRU 1202 can utilize MIMO technology. Thus, in one embodiment, the WTRU 1202 can include two or more transmit / receive components 1222 (e.g., multiple antennas) that transmit and receive wireless signals via the air interface 1215.

[0214] The transceiver 1220 can be configured to modulate the signals to be transmitted by the transmit / receive component 1222, and / or demodulate the signals received by the transmit / receive component 1222. As described above, the WTRU 1202 can have multi-mode capabilities. Thus, the transceiver 1220 can include multiple transceivers that allow the WTRU 1202 to communicate via multiple RATs (e.g., UTRA and IEEE 802.11).

[0215] The processor 1218 of the WTRU 1202 can be coupled to the speaker / microphone 1224, the digital keypad 1226, and / or the display / touchpad 1228 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit), and can receive user input data from these components. The processor 1218 can also output user data to the speaker / microphone 1224, the digital keypad 1226, and / or the display / touchpad 1228. In addition, the processor 1218 can access information from any suitable memory (e.g., non-removable memory 1230 and / or removable memory 1232), and store information in these memories. The non-removable memory 1230 can include random access memory (RAM), read only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 1232 can include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In one or more embodiments, the processor 1218 can access information from, and store data in, memories that are not actually located within the WTRU 1202. By way of example, such memories can be located on a server or a home computer (not shown).

[0216] The processor 1218 can receive power from the power supply 1234 and can be configured to distribute and / or control the power for other components within the WTRU 1202. The power supply 1234 can be any suitable device for powering the WTRU 1202. For example, the power supply 1234 can include one or more dry battery packs (e.g., nickel cadmium (Ni-Cd), nickel zinc (Ni-Zn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), solar cells, and fuel cells, etc.

[0217] The processor 1218 may also be coupled to a GPS chipset 1236 that may be configured to provide location information (e.g., longitude and latitude) related to the current location of the WTRU 1202. As a supplement or replacement to the information from the GPS chipset 1236, the WTRU 1202 may receive location information from a terminal (e.g., a base station) via the air interface 1215, and / or determine its location based on the signal timing received from two or more nearby base stations. It should be understood that the WTRU 1202 may obtain location information by means of any suitable positioning method while remaining consistent with the embodiments.

[0218] The processor 1218 may be further coupled to other peripheral devices 1238 that may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connections. For example, the peripheral devices 1238 may include an accelerometer, an orientation sensor, a motion sensor, a proximity sensor, an electronic compass, a satellite transceiver, a digital camera and / or a video recorder (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth module, a frequency modulation (FM) radio unit, and software modules (such as a digital music player, a media player, a video game console module, and an Internet browser, etc.).

[0219] As an example, the WTRU 1202 may be configured to transmit and / or receive wireless signals and may include a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a pager, a cellular phone, a personal digital assistant (PDA), a smart phone, a laptop computer, a netbook, a tablet computer, a personal computer, a wireless sensor, a consumer electronic product, or any other terminal capable of receiving and processing compressed video communications.

[0220] The WTRU 1202 and / or a communication network (e.g., the communication network 804) may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) that may use Wideband Code Division Multiple Access (WCDMA) to establish the air interface 1215. WCDMA may include communication protocols such as High Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High Speed Downlink Packet Access (HSDPA) and / or High Speed Uplink Packet Access (HSUPA). The WTRU 1202 and / or a communication network (e.g., the communication network 804) may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA) that may use Long Term Evolution (LTE) and / or Advanced LTE (LTE-A) to establish the air interface 1215.

[0221] The WTRU 1202 and / or the communication network (such as communication network 804) may implement radio technologies such as IEEE 802.16 (such as Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN), etc. The WTRU 1202 and / or the communication network (such as communication network 804) may implement radio technologies such as IEEE 802.11 or IEEE 802.15, etc.

[0222] It should be noted that different hardware components of one or more of the described embodiments are referred to as "modules", and the module refers to a "module" for performing (i.e., implementing, running, etc.) different functions described herein in connection with the corresponding module. The modules used herein include hardware (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices) that are considered suitable by those skilled in the relevant art for the specified implementation. Each of the described modules may also include instructions that can be executed to implement one or more functions described as being performed by the corresponding module, and it should be noted that these instructions may be in the form of hardware (i.e., hardwired) instructions, firmware instructions, and / or software instructions or include these instructions, and may be stored in any suitable non-transitory computer-readable medium or media, such as media or media commonly referred to as RAM, ROM, etc.

[0223] Although specific combinations of features and elements have been described above, those of ordinary skill in the art will recognize that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented in a computer program, software, or firmware that is introduced into a computer-readable medium for operation by a computer or a processor. Examples of computer-readable media include electrical signals (transmitted via wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, buffer memories, semiconductor storage devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM discs and digital versatile discs (DVDs). A processor associated with the software can be used to implement a radio frequency transceiver used in a WTRU, UE, terminal, base station, RNC, or any computer host.

Claims

1. A video encoding method, comprising: Encode a single block-level information for at least the current block in the current picture, the single block-level information identifying a first weight and a second weight, wherein at least one of the first weight and the second weight has a value other than 0, 0.5, or 1; And For each sub-block in the current block: Obtain a first sub-block motion vector and a second sub-block motion vector based on the affine motion model of the current block; And Predict the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, wherein the first reference block is weighted by the first weight and the second reference block is weighted by the second weight; Wherein the first weight and the second weight are shared among each of the sub-blocks in the current block.

2. The method according to claim 1, wherein encoding the single block-level information includes mapping a weight index to a codeword using a truncated unary code and entropy encoding the codeword in a bitstream.

3. The method according to claim 1, wherein the prediction of the current block includes the prediction of each sub-block in the current block, and the method further comprises: Subtract the prediction of the current block from the input block to generate a residual; And Encode the residual in the bitstream.

4. A video decoding method, comprising: Decode a single block-level information for at least the current block in the current picture, the single block-level information identifying a first weight and a second weight, wherein at least one of the first weight and the second weight has a value other than 0, 0.5, or 1; And For each sub-block in the current block: Obtain a first sub-block motion vector and a second sub-block motion vector based on the affine motion model of the current block; And Predict the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, wherein the first reference block is weighted by the first weight and the second reference block is weighted by the second weight; Wherein the first weight and the second weight are shared among each of the sub-blocks in the current block.

5. The method according to claim 4, wherein the decoding of the single block-level information includes entropy decoding a codeword from the bitstream and recovering a weight index from the codeword using a truncated unary code.

6. The method according to claim 4, wherein the prediction of the current block includes the prediction of each sub-block in the current block, and the method further comprises: Decode the residual of the current block from the bitstream; And Add the residual to the predicted current block to generate a reconstructed block.

7. The method according to claim 1 or 4, wherein, The second weight is identified by subtracting the first weight from one.

8. The method according to claim 1 or 4, wherein, The first weight and the second weight are identified from a predetermined set of weights.

9. A video encoding device, comprising a processor configured to at least execute: Encode a single block-level information for at least the current block in the current picture, the single block-level information identifying a first weight and a second weight, wherein at least one of the first weight and the second weight has a value other than 0, 0.5, or 1; And For each sub-block in the current block: Obtain a first sub-block motion vector and a second sub-block motion vector based on the affine motion model of the current block; And Predict the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, wherein the first reference block is weighted by the first weight and the second reference block is weighted by the second weight; Wherein the first weight and the second weight are shared among each of the sub-blocks in the current block.

10. The apparatus according to claim 9, wherein the encoding of the single block-level information includes mapping a weight index to a codeword using a truncated unary code and entropy encoding the codeword in a bitstream.

11. The apparatus according to claim 9, wherein the prediction of the current block includes the prediction of each sub-block in the current block, and the processor is further configured to perform: Subtract the prediction of the current block from an input block to generate a residual; and Encode the residual in the bitstream.

12. A video decoding apparatus, comprising a processor configured to at least perform the following operations: Decode a single block-level information for at least the current block in the current picture, the single block-level information identifying a first weight and a second weight, wherein at least one of the first weight and the second weight has a value other than 0, 0.5, or 1; And For each sub-block in the current block: Obtain a first sub-block motion vector and a second sub-block motion vector based on the affine motion model of the current block; And Predict the sub-block as a weighted sum of a first reference block in a first reference picture pointed to by the first sub-block motion vector and a second reference block in a second reference picture pointed to by the second sub-block motion vector, where the first reference block is weighted by the first weight, and the second reference block is weighted by the second weight; where the first weight and the second weight are shared among each of the sub-blocks in the current block.

13. The apparatus according to claim 12, wherein the decoding of the single block-level information includes entropy decoding a codeword from the bitstream and recovering a weight index from the codeword using a truncated unary code.

14. The apparatus according to claim 12, wherein the prediction of the current block includes the prediction of each sub-block in the current block, and the processor is further configured to perform: Decode the residual of the current block from the bitstream; and Add the residual to the predicted current block to generate a reconstructed block.

15. The apparatus according to claim 9 or 12, wherein, The second weight is identified by subtracting the first weight from one.

16. The apparatus according to claim 9 or 12, wherein,The first weight and the second weight are identified from a predetermined set of weights.

17. A computer-readable medium comprising instructions for causing one or more processors to perform the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method and apparatus for coding motion and prediction weighting parameters

    CN101176350A

  • Method for encoding to produce video signal data for a picture with many image blocks

    CN101902645A

Cited By

  • Systems and methods for universal multi-hypothesis prediction for video coding

    CN120676141A