Encoding / decoding video picture data

By introducing affine motion mode based on sub-blocks into the video compression system, and defining the affine motion field using control point motion vectors, the problem of difficulty in capturing complex motion in the prior art is solved, and the video encoding and decoding efficiency and quality are improved.

CN120513623APending Publication Date: 2025-08-19BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380089445.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-09
Filing Date
2023-06-08
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Existing video compression systems are difficult to effectively capture complex motions such as enlargement, reduction, rotation and perspective motion in inter-frame prediction, resulting in inefficient encoding and decoding.

Method used

Affine motion mode based on sub-blocks is adopted to define the affine motion field through two or three control point motion vectors, for inter-frame prediction of video blocks, and motion compensation is performed by combining the affine motion model and control point motion vectors.

Benefits of technology

It improves the efficiency of video encoding and decoding, can capture complex motion more accurately, and improves encoding quality and compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513623A_ABST
    Figure CN120513623A_ABST
Patent Text Reader

Abstract

The present application relates to a method of decoding a video picture, decoding temporal bi-directional prediction by means of motion compensation of inter-coded blocks, motion compensated temporal bidirectional prediction uses two reference video pictures in two separate reference video picture lists and two affine motion fields defined by at least two control point motion vectors, where the control point motion vectors are represented as CPMV, where the at least two control point motion vectors are associated with each reference picture, the refined CPMV is obtained as an output of a first step (171), or optionally as an output of a second step (172) following the first step (171). -in said first step (171), for each CPMV, performing bilateral matching on a block centered on said CPMV to derive at least two CPMVs refined with integer precision, and selecting a set of CPMVs comprising non-refined CPMVs and / or CPMVs refined with integer precision to produce an overall predicted block with minimum bilateral matching cost,-in said second step (172), selecting a set of CPMVs comprising non-refined CPMVs and / or CPMVs refined with integer precision to produce an overall predicted block with minimum bilateral matching cost. For each successive CPMV in the selected set of CPMVs associated with the block, each CPMV is refined at subsample accuracy to minimize bilateral matching costs for the block, where the second step (172) is bypassed according to a comparison between the bilateral matching costs associated with the CPMVs performed in the first step and a threshold.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure claims priority to and the benefit of European Patent Application No. 23305027.7, filed on January 9, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present application generally relates to video picture encoding and decoding. In particular, the technical field of the present application relates to inter-frame prediction of video picture blocks, but is not limited thereto. Background Art

[0004] This section is intended to introduce the reader to various aspects of the art that may be related to various aspects of at least one embodiment of the present application described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present application. Therefore, it should be understood that these statements are to be read in this light, and not as admissions of the related art.

[0005] In state-of-the-art video compression systems, such as HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en) or VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), low-level and high-level picture partitioning are provided to divide the video picture into picture blocks, so-called Codec Tree Units (CTUs), whose size is typically between 16x16 and 64x64 pixels for HEVC and between 32x32, 64x64 or 128x128 pixels for VVC.

[0006] The CTU partitioning of a video picture forms a grid of fixed-size CTUs, i.e., a CTU grid, whose upper and left boundaries spatially coincide with the upper and left boundaries of the video picture. The CTU grid represents the spatial partitioning of the video picture.

[0007] In VVC and HEVC, the CTU sizes (CTU width and CTU height) of all CTUs of the CTU grid are equal to the same default CTU size (default CTU width CTU DW and default CTU height CTU DH). For example, the default CTU size (default CTU height, default CTU width) can be equal to 128 (CTU DW = CTU DH = 128). The default CTU size (height, width) is encoded into the bitstream, for example, at the sequence level in the sequence parameter set (SPS).

[0008] The spatial position of a CTU in the CTU grid is determined by the CTU address ctuAddr, which defines the spatial position of the upper left corner of the CTU relative to the origin. Figure 1 As shown in , a CTU address may define a spatial position starting from the upper left corner of a higher-level spatial structure S that contains the CTU.

[0009] Each CTU is associated with a codec tree to determine the tree partitioning of the CTU.

[0010] like Figure 1 As shown above, in HEVC, the codec tree is a quadtree partition of CTU, where each leaf is called a codec unit (CU). The spatial position of the CU in the video picture is defined by the CU index cuIdx, which indicates the spatial position starting from the upper left corner of the CTU. The CU is spatially partitioned into one or more prediction units (PUs). The spatial position of the PU in the video picture VP is defined by the PU index puIdx, which defines the spatial position starting from the upper left corner of the CTU, and the spatial position of the elements of the partitioned PU is defined by the PU partition index puPartIdx, which defines the spatial position starting from the upper left corner of the PU. Each PU is assigned some intra-frame or inter-frame prediction data.

[0011] The codec mode intra or inter is assigned at the CU level. This means that although the prediction parameters vary from PU to PU, the same intra / inter codec mode is assigned to each PU of a CU.

[0012] A CU can also be spatially partitioned into one or more transform units (TUs) according to a quadtree called a transform tree. A transform unit is a leaf of the transform tree. The spatial position of a TU in a video picture is defined by a TU index tuIdx, which defines the spatial position starting from the upper left corner of the CU. Each TU is assigned some transform parameters. The transform type is assigned at the TU level, and 2D individual transforms are performed at the TU level during encoding or decoding of a picture block.

[0013] Figure 2The above diagram shows the existing PU partition types in HEVC. They include square partitions (2Nx2N and NxN), which are the only partitions used in both intra-frame and inter-frame predicted CUs, symmetric non-square partitions (2NxN, Nx2N, used only in inter-frame predicted CUs) and asymmetric partitions (used only in inter-frame predicted CUs). For example, PU type 2NxnU represents an asymmetric horizontal partition of the PU, where the smaller partition is located at the top of the PU. According to another example, PU type 2NxnL represents an asymmetric horizontal partition of the PU, where the smaller partition is located at the top of the PU.

[0014] like Figure 3 As shown above, in VVC, the codec tree starts from the root node (i.e., CTU). Next, the quadtree (or quadtree) splits the root node into 4 nodes, corresponding to 4 sub-blocks of equal size (solid lines). Next, the quadtree (or quadtree) leaves can be further partitioned by the so-called multi-type tree, which involves Figure 4 One of the four splitting modes shown above performs a binary or trifurcated split. These split types are the vertical and horizontal binary splitting modes (denoted as SBTV and SBTH) and the vertical and horizontal trifurcated splitting modes SPTTV and STTH.

[0015] In the case of a joint codec tree shared by luma and chroma components, the leaves of the codec tree of a CTU are CUs.

[0016] In contrast to HEVC, in VVC, CU, PU and TU have equal sizes in most cases, which means that the codec unit is generally not partitioned into PUs or TUs except in some specific codec modes.

[0017] Figure 5 and Figure 6 An overview of video encoding / decoding methods, such as used in current video standard compression systems like HEVC or VVC, is provided.

[0018] Figure 5 A schematic block diagram showing steps of a method 100 for encoding a video picture VP according to the related art is shown.

[0019] In step 110, the video picture VP is partitioned into blocks of samples and partition information data is signaled into the bitstream. Each block comprises samples of one component of the video picture VP. Thus, the blocks comprise samples of each component defining the video picture VP.

[0020] For example, in HEVC, a picture is divided into codec tree units (CTUs). Each CTU can be further subdivided using quadtree partitioning, where each leaf of the quadtree is represented as a codec unit (CU). The partition information data may then include data describing the CTU and the quadtree subdivision of each CTU.

[0021] Thus, each block of samples (block for short) may be either a CU (if the CU comprises a single PU) or a PU of the CU.

[0022] Each block is encoded along the coding loop (also called "in-loop") using either intra or inter prediction mode.

[0023] Intra-frame prediction (step 120) uses intra-frame prediction data. Intra-frame prediction refers to predicting the current block with the help of intra-frame prediction blocks based on already encoded, decoded, and reconstructed samples. These samples are located around the current block, usually at the top and left of the current block. Intra-frame prediction is performed in the spatial domain.

[0024] In inter-frame prediction mode, motion estimation (step 130) and motion compensation (135) are performed. Motion estimation searches for candidate reference blocks that are good predictors of the current block in one or more reference video pictures used to predictively encode the current video picture. For example, a good predictor of the current block is a predictor that is similar to the current block. The output of the motion estimation step 130 is inter-frame prediction data, which includes motion information associated with the current block (usually one or more motion vectors and one or more reference video picture indices) and other information used to obtain the same predicted block on the encoding / decoding side. Next, motion compensation (step 135) obtains the predicted block with the help of the (one or more) motion vectors and the reference video picture index (multiple indices) determined by the motion estimation step 130. Basically, the block belonging to the selected reference video picture and pointed to by the motion vector can be used as the predicted block of the current block. In addition, since the motion vector is expressed as a fraction of an integer pixel position (this is called sub-pixel accurate motion vector representation), motion compensation generally involves spatial interpolation of some reconstructed samples of the reference video picture to calculate the predicted block.

[0025] The prediction information data is signaled into the bitstream.The prediction information may include prediction mode (intra, inter or skip), intra / inter prediction data and any other information used to obtain the same predicted CU at the decoding side.

[0026] Method 100 selects a prediction mode (intra-frame or inter-frame prediction mode) by optimizing the following rate-distortion trade-off: the encoding of the calculated prediction residual block (for example, by subtracting the candidate predicted block from the current block) and the signaling of the prediction information data required to determine the candidate predicted block at the decoding side.

[0027] Typically, the best prediction mode is given as the prediction mode of the best codec mode p* of the current block given by:

[0028]

[0029] Where P is the set of all candidate codec modes for the current block, p represents the candidate codec mode in the set, and RD cost (p) is the rate-distortion cost of candidate codec mode p, usually expressed as:

[0030] RD cost(p) =D(p)+λ.R(p)

[0031] D(p) is the distortion between the current block and the reconstructed block obtained after encoding / decoding the current block with the candidate codec mode p, R(p) is the rate cost associated with encoding and decoding the current block with the codec mode p, and λ is a Lagrangian parameter representing the rate constraint for encoding and decoding the current block and is typically derived from the quantization parameter used to encode the current block.

[0032] The current block is usually coded from a prediction residual block PR. More precisely, the prediction residual block PR is calculated, for example, by subtracting the best predicted block from the current block. The prediction residual block PR is then transformed (step 140) using, for example, a DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform) type transform or any other suitable transform, and the resulting transformed coefficient block is quantized (step 150).

[0033] In a variant, the method 100 may also skip the transform step 140 and directly apply quantization to the prediction residual block PR (step 150 ), according to the so-called transform-skip coding mode.

[0034] The block of quantized transform coefficients (or the block of quantized prediction residuals) is entropy encoded into a bitstream (step 160).

[0035] Next, as part of the encoding loop, the quantized transform coefficient block (or quantized residual block) is dequantized (step 170) and inverse transformed (step 180) (or not) to obtain a decoded prediction residual block. The decoded prediction residual block and the predicted block are then combined (typically summed) to provide a reconstructed block.

[0036] Other information data may also be entropy encoded in step 160 to encode the current block of the video picture VP.

[0037] A loop filter (step 190) may be applied to the reconstructed picture (including the reconstructed blocks) to reduce compression artifacts. The loop filter may be applied after all picture blocks have been reconstructed. For example, these include a deblocking filter, sample adaptive offset (SAO), or an adaptive loop filter.

[0038] The reconstructed block or the filtered reconstructed block forms a reference picture, which can be stored in a decoded picture buffer (DPB) so that it can be used as a reference picture for encoding the next current block of a video picture VP or the next video picture to be encoded.

[0039] Figure 6 A schematic block diagram showing steps of a method 200 for decoding a video picture VP according to the related art is shown.

[0040] In step 210 , partition information data, prediction information data, and quantized transform coefficient blocks (or quantized residual blocks) are obtained by entropy decoding a bitstream of the encoded video picture data. For example, this bitstream has been generated according to method 100 .

[0041] Other information data may also be entropy decoded to decode the current block of the video picture VP from the bitstream.

[0042] In step 220, the reconstructed picture is divided into current blocks based on the partition information. Each current block is entropy decoded from the bitstream along a decoding loop (also referred to as "in-loop"). Each decoded current block is either a block of quantized transform coefficients or a block of quantized prediction residuals.

[0043] In step 230, the current block is dequantized and possibly inversely transformed (step 240) to obtain a decoded prediction residual block.

[0044] On the other hand, the prediction information data is used to predict the current block. The predicted block is obtained by intra-frame prediction (step 250) or temporal prediction (step 260) of the current block. The prediction process performed on the decoding side is exactly the same as the prediction process on the encoding side.

[0045] Next, the decoded prediction residual block and the predicted block are combined (typically summed), which provides the reconstructed block.

[0046] In step 270, the in-loop filter may be applied to the reconstructed picture (including the reconstructed block), and the reconstructed block or the filtered reconstructed block forms a reference picture, which may be stored in a decoded picture buffer (DPB), as discussed above ( Figure 5 ).

[0047] exist Figure 5 Steps 130 / 135 or Figure 6 In step 260, an inter-predicted block is defined based on inter-prediction data associated with the current block of the video picture (CU or PU of a CU). This inter-prediction data includes motion information, which can be represented (encoded) according to a so-called AMVP mode (Adaptive Motion Vector Prediction) based on the entire block or a so-called Merge mode based on the entire block.

[0048] In HEVC, in the full-block AMVP mode, the motion information of a block used to define inter-frame prediction is represented by up to two reference video picture indices, which are associated with up to two reference video picture lists (usually denoted as L0 and L1). Reference video picture reference list L0 includes at least one reference video picture and reference video picture reference list L1 includes at least one reference video picture. Each reference video picture index is temporally predicted relative to the current block. The motion information also includes up to two motion vectors, each associated with a reference video picture index in one of the two reference video picture lists. Each motion vector is predictively coded and signaled in the bitstream, i.e., a motion vector difference MVd is derived from the motion vector, and an AMVP (Adaptive Motion Vector Predictor) candidate is selected from the full-block AMVP candidate list (constructed on the encoding and decoding side) and the MVd is signaled in the bitstream. The index of the AMVP candidate selected in the full-block AMVP candidate list is also signaled in the bitstream.

[0049] Figure 7 An illustrative example is shown for constructing a whole-block based AMVP candidate list, which is used to define a block for inter-frame prediction of a current block of a current video picture.

[0050] The whole-block based AMVP candidate list may include two spatial MVP (motion vector predictor) candidates derived from the current video picture. The first spatial MVP candidate is derived from the motion information associated with the inter-frame predicted block (if any) located at the neighboring positions A0 and A1 to the left of the current block, and the second MVP candidate is derived from the motion information associated with the inter-frame predicted block (if any) located at the neighboring positions B0, B1, and B2 at the top of the current block. A redundancy check is then performed between the derived spatial MVP candidates, i.e., duplicate derived MVP candidates are discarded. The whole-block based AMVP candidate list may also include a temporal MVP candidate, which is derived from the motion information associated with the co-located block (if any) located at spatial position H or, in other cases, at spatial position C in the reference video picture. The temporal MVP candidate is scaled according to the temporal distance between the current video picture and the reference video picture. Finally, if the whole-block based AMVP candidate list includes fewer than two MVP candidates, the whole-block based AMVP candidate list is padded with zero motion vectors.

[0051] In HEVC, in whole-block merge mode, the motion information of a block used to define inter prediction is represented by a merge index of the whole-block merge MVP candidate list. Each merge index points to the motion predictor information indicating which MVP is used to derive the motion information. The motion information is represented by a unidirectional or bidirectional temporal prediction type, up to 2 reference video picture indices, and up to 2 motion vectors, each motion vector being associated with a reference video picture index of one of the two reference video picture lists (L0 or L1).

[0052] No other information is signaled except the merge index. This means that the motion vector of the current block is set equal to the motion vector of the whole-block MVP candidate indicated by the merge index. Therefore, in contrast to the whole-block merge mode, the MVd and reference picture index are not signaled in the bitstream. Only the index of the selected merge candidate in the whole-block merge candidate list is also signaled in the bitstream.

[0053] Therefore, in contrast to AMVP mode, MVd and reference picture signals are not signaled in merge mode. Only the merge index is signaled into the bitstream.

[0054] The whole-block-based merged MVP candidate list may include five spatial MVP candidates derived from the current video picture, such as Figure 7As shown above. The first spatial MVP candidate is derived from the motion information associated with the inter-frame predicted block located at the left neighboring position A1 (if any); the second spatial MVP candidate is derived from the motion information associated with the inter-frame predicted block located at the upper neighboring position B1 (if any); the third spatial MVP candidate is derived from the motion information associated with the inter-frame predicted block located at the upper right neighboring position B0 (if any); the fourth spatial MVP candidate is derived from the motion information associated with the inter-frame predicted block located at the lower left neighboring position A0 (if any); and the fifth spatial MVP candidate is derived from the motion information associated with the inter-frame predicted block located at the left neighboring position B2 (if any). Redundancy is then checked between the derived spatial MVPs, i.e., duplicate derived MVP candidates are discarded. The whole-block based merge candidate list may also include a temporal MVP candidate, referred to as a TMVP candidate, which is derived from the motion information associated with the co-located block (if any) located at position H or, in other cases, at the center spatial position "C" in the reference video picture. Redundancy is then checked between the derived spatial MVPs, i.e., duplicate derived MVP candidates are discarded. Finally, when using the bidirectional temporal prediction type, if the whole-block based merge candidate list includes less than 5 MVP candidates, a combined candidate is added to the whole-block based merge candidate list. The combined candidate is derived from the motion information associated with one reference video picture list and relative to an MVP candidate already in the whole-block based merge candidate list, and the motion information is associated with another reference video picture list and relative to another MVP candidate already in the whole-block based merge candidate list. Finally, if the whole-block based merge candidate list is still not full (5 merge candidates), the whole-block based merge candidate list is filled with zero motion vectors.

[0055] Encoding and decoding motion information according to VVC provides a richer representation of motion information than in HEVC.

[0056] In VVC, motion information can be represented (coded) according to an entire block-based inter prediction mode that provides an entire block-based motion representation or a sub-block-based inter prediction mode that provides a sub-block-based motion representation.

[0057] In VVC, whole-block based motion representation can be provided according to either whole-block based AMVP mode or whole-block based merge mode.

[0058] In VVC, in the whole-block based AMVP mode, the motion information of the block used to define the inter-frame prediction is represented as in the whole-block based AMVP mode of HEVC. The whole-block based AMVP candidate list can include spatial and temporal MVP candidates (if present) as in HEVC. The whole-block based AMVP candidate list can also include four additional HMVP (history-based motion vector prediction) candidates (if present). Finally, if the whole-block based AMVP candidate list includes less than 2 MVP candidates, the whole-block based AMVP candidate list is filled with zero motion vectors.

[0059] The HMVP candidate is derived from the previously coded MVP associated with adjacent blocks or non-adjacent blocks relative to the current block. To this end, a table of HMVP candidates is maintained on both the encoder and decoder sides and is updated at any time as a first-in-first-out (FIFO) buffer for MVP. There are up to five HMVP candidates in the table. After encoding and decoding a block, the table is updated by appending the associated motion information as a new HMVP candidate to the end of the table. The FIFO rule is applied to manage the table, in which, in addition to the basic FIFO mechanism, redundant candidates in the HMVP table are removed first instead of the first candidate. The table is reset at each CTU row to enable parallel processing.

[0060] In addition to the block-based AMVP candidate list, the block-based AMVP mode also uses the following new tools compared to HEVC:

[0061] -SMVD (Symmetrical Motion Vector Difference): For bidirectional blocks, SMVD consists of setting the MVd associated with one reference video picture list for the current block to be equal to the inverse of the MVd associated with the other reference video picture list for the current block. Furthermore, the reference video pictures used in SMVD mode are derived by the decoder according to some predefined rules. SMVD enables reducing the rate cost for encoding and decoding MVd and reference picture index information and making selections at the block level.

[0062] -AMVR (Adaptive Motion Vector Resolution). The AMVR tool allows signaling MVd at quarter-pixel, half-pixel, integer-pixel, or quad-pixel luma sample resolution. This also allows saving bits in encoding and decoding MVd information. In AMVR, the motion vector resolution is selected at the block level.

[0063] - BCW (Bidirectional Prediction with Codec Weights) enables bidirectional prediction for blocks with unequal weights and signals this at the block level.

[0064] Finally, the internal motion vector representation is implemented with 1 / 16 luma sample accuracy, rather than 1 / 4 luma sample accuracy in HEVC.

[0065] In VVC, in the full-block based merge mode, the motion information of the block used to define the inter prediction is represented as in the HEVC full-block based merge mode.

[0066] The whole-block based merge candidate list is different from the whole-block based merge candidate list used in HEVC.

[0067] In VVC, a whole-block based merge candidate list can be constructed for MMVD (Merge Mode with MV Difference) mode, GPM (Geometric Partitioning Mode) or CIIP (Combined Intra / Inter Prediction) mode.

[0068] The MMVD mode allows encoding and decoding of limited motion vector differences (MVd) on top of the selected merge candidate to represent the motion information associated with the inter-predicted block. The MMVD codec is limited to 4 vector directions and 8 magnitudes, from 1 / 4 luma samples to 32 luma samples, e.g. Figure 8 The MMVD mode provides a medium level of accuracy, thus making a medium trade-off between rate cost and MV (motion vector) accuracy to signal motion information.

[0069] GPM refers to partitioning the current block PG into two motion partitions MV0 and MV1 along a straight line, such as Figure 9 Each motion vector MV0 and MV1 points to a block P0, P1 of a reference video picture of one of the two reference video picture lists.

[0070] The geometric partition can be non-rectangular or rectangular. If rectangular, an asymmetric split is performed to avoid redundancy by binary splitting at the CU level. A CU-level flag is used to signal the use of GPM as a specific whole-block based merge mode.

[0071] The orientation and position of the split line relative to the center of the current block are signaled at the CU level through a dedicated GPM index. Geometric partitioning mode is used for each possible CU size w×h=2 m ×2 n , where m,n∈{3…6} (excluding 8×64 and 64×8) supports a total of 64 partitions. For example, Figure 10 The diagram illustrates various split lines that can be achieved at different position offsets from the CU center and for various split line angles.

[0072] Each part of the geometric partition in the current block PG is inter-predicted using its own motion. Only unidirectional prediction is allowed for each partition, i.e., each part has only one motion vector and one reference video picture index. A unidirectional inter prediction constraint is applied to ensure that only two motion-compensated predictions are required for each current block PG, similar to conventional bidirectional inter prediction.

[0073] The motion vectors MV0 and MV1 of each partition are derived from the two merge indices of each partition, respectively.

[0074] After predicting each part of the geometric partition, a hybrid process with adaptive weights is used to adjust the sample values along the edges of the geometric partition. This is the prediction signal for the entire current block PG, and like other prediction modes based on the entire block, the transform and quantization process will be applied to the entire current block PG (rather than for each partition).

[0075] CIIP refers to combining the inter prediction signal with the intra prediction signal to predict the current block. The inter prediction signal in CIIP mode is derived using the same inter prediction process as applied to the whole block based merge mode; and the intra prediction signal is derived following the conventional intra prediction process with planar mode. The intra prediction signal and the inter prediction signal are then combined using a weighted average, where the weight values are calculated based on the codec mode of the top neighboring block and the left neighboring block:

[0076] PCIIP=(Wmerge×Pmerge+Wintra×Pintra+2)>>2

[0077] Weight W merge and weight W intra The sum of is equal to 4, and these weights are constant throughout the current block.

[0078] In summary, in VVC, the whole-block based merge candidate list can include spatial MVP candidates similar to HEVC, except that the first two candidates are swapped: during the construction of the whole-block based merge candidate list, candidate B1 is considered before candidate A1. The whole-block based merge candidate list can also include TMVP candidates similar to HEVC, as well as HMVP candidates as in the VVC AMVP mode. Several HMVP candidates are inserted into the whole-block based merge candidate list so that the merge candidate list reaches the maximum number of MVP candidates allowed (minus 1). The whole-block based merge candidate list can also include up to 1 pairwise average candidate calculated as follows: consider the first two merge candidates present in the whole-block based merge candidate list and average their motion vectors. This averaging is calculated separately for each reference video picture list L0 and L1. Therefore, if both MVPs are bidirectional, the motion vectors associated with list L0 and list L1 are averaged. If there is only one motion vector in the reference video picture list, it is used to form a pairwise candidate. Finally, if the merge candidate list does not reach the maximum number of MVP candidates allowed, the whole-block based merge candidate list is filled with a zero motion vector.

[0079] In VVC, subblock-based inter prediction basically means that the current block is divided into subblocks, usually into 4×4 or 8×8 luma sample subblocks, and the subblock-based motion representation of the block used to define the inter prediction is represented by the motion information associated with all subblocks.

[0080] In HEVC, a translation-only motion model is applied for temporal prediction of motion compensation. This translational motion model cannot capture some types of motion, such as zooming in, zooming out, rotation, perspective motion, and irregular motion.

[0081] To solve this problem, in VVC, a sub-block based affine motion model can be used at the CU level. Figure 11 As shown above, the affine motion field of the current block is defined by the translation motion vector of two control points (4-parameter affine motion model) or three control points (6-parameter affine motion model). Figure 11 The translation motion vectors v0, v1, v2 of the control points located at the corners of the current block are used to derive the affine motion field of the current block.

[0082] For the 4-parameter affine motion model, the motion vector (mv x ,mv y ) is derived as follows:

[0083]

[0084] Among them (mv 0x ,mv 0y) is the translation motion vector of the upper left control point, and (mv 1x ,mv 1y ) is the motion vector of the upper right control point.

[0085] For the 6-parameter affine motion model, the motion vector at the sample location (x, y) in the 4×4 sub-block of the current block is derived as follows:

[0086]

[0087] Among them (mv 0x ,mv 0y ) is the translation motion vector of the upper left control point, (mv 1x ,mv 1y ) is the translation motion vector of the upper right control point, and (mv 2x ,mv 2y ) is the translation motion vector of the lower left control point. Parameters and Represents non-translation parameters (rotation, scaling).

[0088] In order to derive the motion vector of each sub-block (usually the luminance sub-block) of the current block, the motion vector of the center sample of each sub-block is calculated according to Equation 1 or 2, as Figure 12 , and rounded to 1 / 16 decimal accuracy. A motion-compensated interpolation filter is then applied to produce a prediction for each subblock using the derived motion vector, thereby performing inter-frame prediction on the current block. The subblock size in the chroma component is also 4×4. The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the top-left and bottom-right luminance subblocks in the co-located 8×8 luminance region.

[0089] Hereinafter, the translation motion vector of the control point is named CPMV (Control Point Motion Vector).

[0090] The sub-block based motion representation for defining inter-predicted blocks may be provided according to the sub-block based affine AMVP mode or the sub-block based affine merge mode.

[0091] The sub-block based affine merge mode can be applied to current blocks with width and height both greater than or equal to 8. An affine merge candidate list of up to 5 sub-block based motion candidates is used. The sub-block based motion candidates are CPMVP candidates based on motion information associated with spatially neighboring blocks. The index of the selected sub-block based motion candidate in the affine merge candidate list is signaled into the bitstream. Finally, if the affine merge candidate list is not full (5 CPMVP candidates), the affine merge candidate list is filled with zero motion vectors.

[0092] The CPMVP candidate can be either a so-called inherited CPMVP candidate derived by extrapolating the affine motion model (CPMV) of the neighboring blocks, or a so-called constructed CPMVP which is a CPMVP candidate derived by combining the translation MVs of the neighboring sub-blocks.

[0093] In VVC, there are at most two inherited CPMVP candidates, which are derived by extrapolating the affine motion model of the left neighbor block and the affine motion model of the upper neighbor block. For the left CPMVP candidate, if the affine motion model of the left neighbor A0 exists, the left CPMVP candidate is extrapolated from the CPMV of the affine motion model. Otherwise, the left CPMVP candidate is extrapolated from the CPMV of the affine motion model of the left neighbor A1 (if it exists). For the upper CPMVP candidate, if the affine motion model of the upper neighbor B0 exists, the upper CPMVP candidate is extrapolated from the CPMV of the affine motion model. Otherwise, if the affine motion model of the upper neighbor B1 exists, the upper CPMVP candidate is extrapolated from the CPMV of the affine motion model. Otherwise, the CPMVP candidate is extrapolated from the CPMV of the affine motion model of the left neighbor B2 (if it exists).

[0094] Extrapolating the affine motion models of neighboring blocks means deriving inherited CPMVP candidates by combining the CPMVs of the affine motion models.

[0095] like Figure 13 As shown in , if an affine motion model of the lower left neighbor sub-block A exists, then the translation motion vectors v2, v3, and v4 of the affine motion model relative to the upper left corner, upper right corner, and lower left corner of the neighboring block containing sub-block A are obtained. When sub-block A is encoded and decoded using the above 4-parameter affine model, the left CPMVP candidate of the current block is inherited (calculated) by combining the translation motion vectors v2 and v3 according to the following two equations.

[0096]

[0097] Where (X curr ,Y curr ) is the current block position, (X neighb ,Y neighb ) is the location of the neighboring block, and (W neighb ,H neighb ) is the size of the neighboring block.

[0098] In the case where sub-block A is encoded and decoded using the above 6-parameter affine model, the three CPMVs of the current block are calculated by combining the translation motion vectors v2, v3 and v4 according to the following equations.

[0099]

[0100] In VVC, the constructed affine candidates are derived by combining the non-affine MVs of neighboring blocks. The MVs of control points are derived from spatial and temporal neighboring sub-blocks.

[0101] Figure 14 An example of deriving a constructed CPMVP candidate by combining the translated MVs of neighboring blocks is shown.

[0102] CPMV k (k=1, 2, 3, 4) represents the translation MV associated with the k-th control point of the current block. For CPMV1, if the translation MV of the neighbor sub-block B2 exists, then CPMV1 is equal to the translation MV. Otherwise, if the translation MV of the neighbor sub-block B3 exists, then CPMV1 is equal to the translation MV. Otherwise, if the translation MV of the neighbor sub-block A2 exists, then CPMV1 is equal to the translation MV. For CPMV2, if the translation MV of the neighbor sub-block B1 exists, then CPMV2 is equal to the translation MV. Otherwise, if the translation MV of the neighbor sub-block B0 exists, then CPMV2 is equal to the translation MV. For CPMV3, if the translation MV of the neighbor sub-block A1 exists, then CPMV3 is equal to the translation MV. Otherwise, if the translation MV of the neighbor sub-block A0 exists, then CPMV3 is equal to the translation MV. For CPMV4, if the translation MV of the co-located neighbor sub-block exists in the reference video picture at position T, then CPMV4 is equal to the translation MV.

[0103] The constructed CPMVP candidate is constructed by combining CPMV k Derived: If CPMV1, CPMV2, and CPMV3 exist, then the constructed CPMVP candidate is derived by combining {CPMV1, CPMV2, CPMV3}. Otherwise, if CPMV1, CPMV2, and CPMV4 exist, then the constructed CPMVP candidate is derived by combining {CPMV1, CPMV2, CPMV4}. Otherwise, if CPMV1, CPMV3, and CPMV4 exist, then the constructed CPMVP candidate is derived by combining {CPMV1, CPMV3, CPMV4}. Otherwise, if CPMV2, CPMV3, and CPMV4 exist, then the constructed CPMVP candidate is derived by combining {CPMV2, CPMV3, CPMV4}. Otherwise, if CPMV1 and CPMV2 exist, then the constructed CPMVP candidate is derived by combining {CPMV1, CPMV2}. Otherwise, if CPMV1 and CPMV3 exist, then the constructed CPMVP candidate is derived by combining {CPMV1, CPMV3}.

[0104] To avoid CPMV kMotion scaling occurs when pointing to different reference video frames, as with the CPMV k The combination of associated control point MVs is discarded.

[0105] Combine 3 CPMVs k This results in a 6-parameter affine motion model (Eq. 2), and combining the two CPMVs results in a 4-parameter affine motion model (Eq. 1).

[0106] The sub-block based affine merge model uses an affine merge candidate list of CPMVP candidates. The merge index mergeIdx indicates the CPMVP candidate used to derive the motion information of the current block. The best CPMVP candidate in the affine merge candidate list is determined for the current block as the candidate that minimizes the rate-distortion cost for encoding and decoding the given block.

[0107] The affine merge candidate list may include SbTMVP (sub-block based temporal motion vector prediction) candidates and / or inherited CPMVP candidates and / or constructed CPMVP candidates, as discussed above. Finally, if the affine merge candidate list does not reach the maximum number of candidates allowed, the affine merge candidate list is filled with zero motion vectors.

[0108] SbTMVP candidates can be derived as follows: Similar to TMVP candidates, SbTMVP candidates are derived from the motion information of the co-located sub-blocks (if any) in the reference video picture. SbTMVP candidates differ from TMVP candidates in the following two main aspects:

[0109] -TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub-block level (sub-CU level);

[0110] -TMVP obtains the temporal motion vector from the co-located block in the reference video picture, and SbTMVP applies a motion shift before obtaining the temporal motion information from the reference video picture, where the motion shift is obtained from the motion vector of one of the spatially neighboring sub-blocks of the current block.

[0111] The sub-block based affine AMVP mode can be applied to the current block with width and height greater than or equal to 16 in VVC. In the sub-block based affine AMVP mode, the CPMV difference MVd between the MV of the control point at the corner of the current block and its CPMVP (control point motion vector predictor) is signaled to the bitstream. Finally, the CPMVP index pointing to the CPMVP candidate in the affine AMVP candidate list is also signaled to the bitstream.

[0112] The CPMVP for predicting the MV of a control point located at a corner of the current block is derived from an affine AMVP candidate list including two CPMVP candidates.

[0113] The affine AMVP candidate list may include inherited CPMVP candidates and / or constructed CPMVP candidates.

[0114] In ECM ("Algorithm description of Enhanced Compression Model 7 (ECM 7)", M.Coban, F.Le Léannec, R.-L.Liao, K.Naser, J. L. Zhang, document JVET-AB2025, ITU-TSG 16WP 3 and ISO / IEC JTC 1 / SC 29 Joint Video Experts Group (JVET), 28th Meeting, Germany, October 21-28, 2022, https: / / jvet-experts.org / doc_end_user / documents / 28_Mainz / wg11 / JVET-AB2025-v1.zip), the motion information of blocks used to define inter prediction can also be represented according to the so-called bilateral matching AMVP merge mode (BM-AMVP merge mode).

[0115] The BM-AMVP merge mode is allowed when at least one reference video picture is available in one reference picture list (along one so-called inter direction) of the video picture including the current block, and at least one reference video picture is available in the other reference picture list (along another inter direction) of the video picture. In the BM-AMVP merge mode, the motion information of the block used to define the inter prediction includes first motion information associated with a reference video picture of the reference video picture list and second motion information associated with a reference video picture of the other reference video picture list.

[0116] In the BM-AMVP merge mode as defined in the ECM, the first motion information is a first whole-block based motion vector (MV) associated with a reference video picture of a reference video picture list associated with an inter-frame direction, and the second motion information is a second whole-block based MV associated with a reference video picture of another reference video picture list associated with the other inter-frame direction.

[0117] The first full-block based MV is represented as discussed above and is related to the full-block based AMVP mode in VVC, so it is derived from the full-block based motion vector predictor (MVP) obtained by the full-block based AMVP mode, the signaled reference picture index and the signaled motion vector difference.

[0118] The second entire block based MV is represented as discussed above, and is associated with the entire block based merge mode, so it is derived from the entire block based motion vector predictor (MVP) obtained by the entire block based merge mode.

[0119] The block-based MVP used to derive the second block-based MV (i.e., the mergeIdx index in the block-based merge MVP candidate list indicated by the refListMerge index) is derived from the following bilateral matching method given by the following formula, which minimizes the bilateral matching cost Bmcost:

[0120]

[0121] Wherein mIdx[refListMerge] is an index pointing to an MVP candidate in the block-based merged MVP candidate list indicated by the refListMerge index, and MvpIdx[refListAmvp] is the block-based MVP for predicting the first block-based MV of the BM-AMVP merge mode.

[0122] For each MVP candidate in the block-based merged MVP candidate list, a bilateral matching cost is calculated using the MVP candidate and MvpIdx[refListAmvp]. The MVP candidate with the minimum cost is selected as the block-based MVP for deriving the second block-based MV.

[0123] In VVC, in order to improve the accuracy of the MV in the merge mode, the decoder-side motion vector refinement method based on bilateral matching (BM) (called DMVR method) is applied to the current block to obtain the motion vector derived by the AMVP mode based on the entire block ( Figure 15 MV0 on the MV) and the second whole-block-based MV derived by the whole-block-based merge mode ( Figure 15MV1 in ) as a starting point. The principle is as follows. In the bidirectional prediction operation, the refined motion vectors MV0' and MV1' are derived around the two initial motion vectors MV0 and MV1 of the reference blocks pointing to the reference video pictures of the reference video picture lists L0 and L1 by minimizing the bilateral matching cost between the two reference blocks (minimizing the motion vector difference Mdiff). Basically, the local search is to derive the MV refinement offset MV_offset. The local search usually applies an n×n (n is an integer value) square search pattern to loop through the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction. The bilateral matching cost (BM cost) is calculated as the sum of mvDistanceCost and sadCost, where mvDistanceCost represents the motion vector encoding and decoding cost and increases according to the motion vector magnitude, and sadCost is a measure of the difference between the two reference blocks. For example, sadCost is the sum of the absolute difference (SAD) of the samples of the two reference blocks. When the BM cost of the center point of the n×n search pattern has the minimum cost, the local search terminates. Otherwise, the current minimum cost search point becomes the new center point of the n×n search pattern and continues searching for the minimum cost until the end of the search range is reached.

[0124] In DVMR, the search point surrounds the initial MV and the MV refinement offset follows the MV difference mirror rule. In other words, any point examined by DMVR (represented by a candidate MV pair (MV0, MV1)) obeys the following two equations:

[0125] MV0′=MV0+MV_offset (4)

[0126] MV1 ′ =MV1-MV_offset (5)

[0127] Wherein MV_offset represents the MV refinement offset between the initial MV and the refined MV in one of the reference pictures.

[0128] The DMVR search consists of an integer sample offset search step (the first step of DMVR) followed by a fractional refinement step (the second step of DMVR). The refinement search range is two integer luma samples from the initial MV.

[0129] The refined MV derived by the DMVR method is used to generate inter-frame prediction samples and is also used in temporal motion vector prediction for future picture coding, while the original MV is used in the deblocking process and is also used in spatial motion vector prediction for future CU coding.

[0130] In VVC, a 25-point full search is applied for the first step of the DMVR method. The BM cost of the initial MV pair is first calculated. If the BM cost of the initial MV pair is less than a threshold, the first step of the DMVR method is terminated. Otherwise, the BM costs of the remaining 24 points are calculated and checked in raster scan order. The search point with the minimum BM cost is selected as the output of the first step of the DMVR method.

[0131] In order to reduce the penalty caused by the uncertainty of DMVR refinement, we propose to bias the original MV during the DMVR process, reducing the SAD between the reference blocks pointed to by the initial MV candidates by 1 / 4 of the SAD value.

[0132] To reduce computational complexity, the second step of the DMVR method is derived using the parametric error surface equation rather than an additional search using SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the first step of the DMVR method. Fractional sample refinement (the second step of the DMVR method) is further applied when the first step of the DMVR method terminates at a center with the minimum SAD in either the first iteration or the second iteration of the first step of the DMVR method.

[0133] In the sub-pixel offset estimation based on the parametric error surface, the BM cost at the center position and the BM costs at the four neighboring positions from the center are used to fit the 2D parabolic error surface equation of the following form:

[0134] E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C (6)

[0135] Where (x min ,y min ) corresponds to the fractional position with the least BM cost, and C corresponds to the minimum BM cost value. By solving the above equation using the BM cost values of the five search points, (x min ,y min ) is calculated as follows:

[0136] x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (7)

[0137] y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (8)

[0138] x min and y minThe value of is automatically constrained to be between -8 and 8, since all BM cost values are positive and the minimum value is E(0,0). This corresponds to a half-pixel offset of 1 / 16 pixel MV accuracy in VVC. min ,y min ) is added to the integer distance refinement MV to obtain the sub-pixel accurate refinement increment MV.

[0139] In VVC, the application of DMVR is restricted and only applies to CUs that meet a specific DMVR condition.

[0140] The DMVR condition is met when the CU is encoded or decoded using one of the following codec modes and features:

[0141] -CU level merge mode with bi-predictive MV;

[0142] -Relative to the current picture, one reference picture is in the past and the other reference picture is in the future;

[0143] - The distances from the two reference pictures to the current picture (i.e., POC difference) are the same;

[0144] - Both reference pictures are short-term reference pictures;

[0145] -CU has more than 64 luminance samples;

[0146] -CU height and CU width are both greater than or equal to 8 luma samples;

[0147] - BCW weight index indicates equal weight;

[0148] - Weighted prediction (WP) is not enabled for the current block; weighted prediction allows combining a multiplicative weighting factor and an additive offset to apply to motion compensated prediction.

[0149] ("Weighted prediction in the H.264 / MPEG AVC video coding standard", JMBoyce, May 2004, 10.1109 / iscas.2004.1328865).

[0150] - Do not use CIIP mode for the current block.

[0151] In VVC, the resolution of MV is 1 / 16 luma sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points surround the initial fractional pixel MV, which has an integer sample offset, so the samples at those fractional positions need to be interpolated for the DMVR search process. In order to reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important impact is that by using a bilinear filter with a 2-sample search range, DVMR does not access more reference samples than the ordinary motion compensation process. After the refined MV is obtained using the DMVR search process, an ordinary 8-tap interpolation filter is applied to generate the final prediction. In order not to access more reference samples in the ordinary MC (motion compensation) process, samples that are not required by the interpolation process based on the original MV but are required by the interpolation process based on the refined MV are filled in from those available samples.

[0152] When the width and / or height of a CU is greater than 16 luma samples, it is further split into sub-blocks with width and / or height equal to 16 luma samples. The maximum unit size for the DMVR search process is limited to 16×16.

[0153] In ECM, DMVR is applied to affine merged codecs when the DMVR conditions are met. This is so-called affine DMVR. The first step of affine DMVR (the integer sample offset search step) is applied to the translational part of the affine motion (Equation 3), so that if a candidate meets the DMVR conditions, the translational MV offset is added to all CPMVs of that candidate in the affine merge list. The MV offset is derived by minimizing the BM cost, which is the same as conventional DMVR.

[0154] Affine DMVR consists of a 3×3 square search pattern (8 search points) that is used to loop over the search range set to [-3, 3] to find the best integer MV refinement offset (the first step of affine DMVR).

[0155] Figure 16 An example of the scheme and search point order used by the affine DMVR method is shown. The cross × represents the initial position of the search, and positions 0, ..., 7 represent all eight positions of the search point.

[0156] Then, fractional refinement is performed around the best integer position (the second step of affine DMVR), and finally an error surface estimation is performed to find the optimal MV refinement offset with an accuracy of 1 / 16.

[0157] In affine DMVR, there is no bias towards the original MV and all search points are evaluated.

[0158] In JVET contribution JVET-AB0178 (https: / / jvet-experts.org / doc_end_user / documents / 28_Mainz / wg11 / JVET-AB2028-v1.zip), CPMV refinement of affine DMVR is proposed to further improve affine DMVR, where different MV refinement offsets can be applied to different CPMVs.

[0159] The method as described in JVET-AB0178 can be applied to merge candidates of affine merge mode and affine MMVD mode.

[0160] Figure 17 A block diagram schematically illustrates the steps of the CPMV thinning method as described in JVET-AB0178.

[0161] The CPMV refinement method follows the affine DMVR method as described in the related art.

[0162] In a first step 171, for each control point motion vector initCpMvLX[cpIdx], where cpIdx = 0..numCpMv-1, and where numCpMv is the number of CPMVs of the current affine codec block, the method performs bilateral matching on the block centered at the control point to derive a refined CPMV bmRefinedCpMvLx[cpIdx].

[0163] Next, loop through the combination of initCpMvLX[cpIdx] and bmRefinedCpMvLx[cpIdx] to derive the optimal CPMV set that minimizes the BM cost of the current block.

[0164] In the second step 172 (optional), the CPMVs in the best CPMV set are iteratively refined to further minimize the BM cost of the current block, which follows the principle of the second step of affine DMVR. In each iteration, a single CPMV is refined, while the other CPMVs are fixed.

[0165] Figure 17 The approach is advantageous because it improves the compression efficiency of state-of-the-art video codecs.

[0166] However, a disadvantage is that the trade-off between compression efficiency improvement and increased encoder and decoder complexity is not very attractive, since this approach implies a significant increase in codec complexity.

[0167] The goal of this application is to improve the trade-off between compression efficiency gain and encoding time overhead of the CPMV refinement method proposed in JVET-AB0178.

[0168] At least one embodiment of the present application is designed based on the above content. Summary of the Invention

[0169] The following section provides a brief overview of at least one embodiment to provide a basic understanding of some aspects of the present application. This overview is not an exhaustive overview of the embodiments. It is not intended to identify key or important elements of the embodiments. The following overview merely provides some aspects of at least one embodiment in a simplified form as a prelude to the more detailed descriptions provided elsewhere in this document.

[0170] According to a first aspect of the present application, a method for decoding video pictures is provided, wherein the decoding is performed by means of temporal bidirectional prediction of motion-compensated inter-coded blocks, wherein the temporal bidirectional prediction of motion compensation uses two reference video pictures in two separate reference video picture lists and two affine motion fields defined by at least two control point motion vectors, wherein the control point motion vectors are represented as CPMVs, wherein at least two control point motion vectors are associated with each reference picture, and the refined CPMVs are obtained as the output of the following first step, or optionally, as the output of a second step following the first step.

[0171] - in said first step, for each CPMV, performing bilateral matching on a block centered around the CPMV to derive at least two CPMVs refined with integer precision, and selecting a set of CPMVs including non-refined CPMVs and / or CPMVs refined with integer precision to produce an overall predicted block with a minimum bilateral matching cost,

[0172] - in said second step, for each successive CPMV of the selected set of CPMVs associated with a block, refining each CPMV with subsample accuracy to minimize the bilateral matching cost of said block,

[0173] The second step is bypassed based on a comparison between the two-sided matching cost associated with the CPMV performed in the first step and a threshold.

[0174] In one embodiment, when the bilateral matching cost associated with a CPMV performed in the first step is below a first threshold, the second step is bypassed for the CPMV.

[0175] In one embodiment, when the bilateral matching cost associated with a CPMV performed in the first step is higher than a second threshold, the second step is bypassed for the CPMV.

[0176] In one embodiment, the first threshold and / or the second threshold are fixed or adapted to the prediction unit size.

[0177] In one embodiment, two affine motion fields are defined by first, second, and third CPMVs, and wherein the second step is bypassed for the third CPMV when the bilateral matching cost associated with the first or second CPMV satisfies a bypass condition.

[0178] In one embodiment, the bypass condition is met when the bilateral cost associated with the first or second CPMV performed in the second step is higher than the bilateral matching cost associated with the first or second CPMV performed in the first step.

[0179] In one embodiment, the bypass condition is met when the bilateral matching cost associated with the first or second CPMV performed in the second step is lower than the bilateral matching cost associated with the first or second CPMV performed in the first step, and if the reduction in the bilateral cost is lower than a third threshold.

[0180] In one embodiment, the first and second steps are disabled for certain block shapes.

[0181] In one embodiment, when the width or height of the block to be inter-coded is higher than a first value when the width and height are equal, the first step and the second step are disabled.

[0182] In one embodiment, when the width or height of the block to be inter-coded is higher than a second value when the width and height are different, the first step and the second step are disabled.

[0183] In one embodiment, the block shape is defined according to the width or height of the inter-coded block, and wherein at least one block shape disabling the first and second steps is signaled in the bitstream.

[0184] In one embodiment, the at least one block shape disabling the first and second steps is signaled by signaling a maximum and a minimum block width and / or height.

[0185] In one embodiment, the at least one block shape disabling the first and second steps is signaled by signaling a minimum and / or maximum block ratio between a maximum value of block width and block height and a minimum value of block width and block height.

[0186] According to a second aspect of the present application, a device is provided, comprising components for performing one of the methods according to the first aspect of the present application.

[0187] According to a third aspect of the present application, a computer program product comprising instructions is provided. When the program is executed by one or more processors, the one or more processors are caused to perform the method according to the first aspect of the present application.

[0188] According to a fourth aspect of the present application, there is provided a non-transitory storage medium carrying instructions of a program code for executing the method according to the first aspect of the present application.

[0189] According to a fifth aspect of the present application, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor. The processor is configured to execute the method according to the first aspect of the present application.

[0190] The specific nature of at least one of the embodiments and other objects, advantages, features and uses of the at least one of the embodiments will become apparent from the following description of the examples taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0191] Reference will now be made, by way of example, to the accompanying drawings which show embodiments of the present application, in which:

[0192] Figure 1 An example of a codec tree unit according to HEVC is shown;

[0193] Figure 2 An example of partitioning a codec unit into prediction units according to HEVC is shown;

[0194] Figure 3 An example of CTU partitioning according to VVC is shown;

[0195] Figure 4 shows an example of splitting modes supported in multi-type tree partitioning according to VVC;

[0196] Figure 5 A schematic block diagram showing steps of a method 100 for encoding a video picture VP according to the related art;

[0197] Figure 6 A schematic block diagram showing steps of a method 200 for decoding a video picture VP according to the related art;

[0198] Figure 7 An illustrative example is shown for constructing a whole-block based AMVP candidate list for defining a block for inter-frame prediction of a current block of a current video picture;

[0199] Figure 8 An example of limited motion vector difference according to related art is shown;

[0200] Figure 9 An example of geometrically partitioning a video picture block according to related art is shown;

[0201] Figure 10 An example of geometric block partitioning according to the related art is shown;

[0202] Figure 11 An example of representing an affine motion field of a video picture block according to the related art is shown;

[0203] Figure 12 shows an example of motion vectors associated with sub-blocks of a video picture block according to the related art;

[0204] Figure 13 An example of deriving an affine motion model from a translation motion vector relative to a control point located at a corner of a video picture block according to the related art is shown;

[0205] Figure 14 An example of deriving a constructed CPMVP candidate by combining the translation MVs of neighboring blocks according to the related art is shown;

[0206] Figure 15 A diagram showing a bilateral matching refinement method according to the related art is shown;

[0207] Figure 16 An example of the scheme and search point order used by affine DMVR is shown;

[0208] Figure 17 A block diagram schematically illustrates the steps of the CPMV thinning method as described in JVET-AB0178;

[0209] Figure 18 shows the location of a search point according to the related art; and

[0210] Figure 19 A schematic block diagram illustrating an example of a system in which various aspects and embodiments are implemented is illustrated.

[0211] Similar or identical elements are referenced with the same reference numerals. DETAILED DESCRIPTION

[0212] At least one of the embodiments will be described more fully below with reference to the accompanying drawings, in which examples of at least one of the embodiments are depicted. However, the embodiments may be implemented in a variety of alternative forms and should not be construed as being limited to the examples set forth herein. Thus, it should be understood that the present invention is not intended to limit the embodiments to the specific forms disclosed. On the contrary, this application is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the application.

[0213] At least one of the aspects generally relates to video picture encoding and decoding, another aspect generally relates to transmitting provided or encoded bitstreams, and another aspect relates to receiving / accessing decoded bitstreams.

[0214] At least one of the embodiments is described with respect to encoding / decoding one video picture, but extends to encoding / decoding multiple video pictures (sequences of pictures) in that each video picture is encoded / decoded sequentially as described below.

[0215] Moreover, for example, at least one embodiment is not limited to MPEG standards such as AVC (ISO / IEC 14496-10 Advanced Video Coding for generic audio-visual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Essential video coding), HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en), VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), but can be applied to other standards and recommendations, such as AV1 (AOMedia Video 1, http: / / aomedia.org / av1 / specification / ). At least one embodiment can be applicable to existing or future developments and extensions of any such standards and recommendations. Unless otherwise specified or technically excluded, the various aspects described in this application can be used alone or in combination.

[0216] A pixel corresponds to the smallest display unit on the screen, which can be composed of one or more light sources (1 for a monochrome screen and 3 or more for a color screen).

[0217] A video picture, also called a frame or picture frame, comprises at least one component (also called picture component or channel) determined by a specific picture / video format, which specifies all information related to pixel values and all information that can be used by a display unit and / or any other device to display and / or decode video picture data related to the video picture.

[0218] A video picture comprises at least one component which is usually represented in the form of an array of samples.

[0219] A monochrome video picture includes a single component, while a color video picture may include three components.

[0220] For example, when the picture / video format is the well-known (Y, Cb, Cr) format, a color video picture may include one luminance (or luminance) component and two chrominance components, and when the picture / video format is the well-known (R, G, B) format, a color video picture may include three color components (one for red, one for green, and one for blue).

[0221] Each component of a video picture may comprise a number of samples relative to the number of pixels of the screen on which the video picture is to be displayed. In a variant, the number of samples comprised in a component may be a multiple (or fraction) of the number of samples comprised in another component of the same video picture.

[0222] For example, in case the video format includes one luma component and two chroma components (such as the (Y, Cb, Cr) format), the chroma components may contain half the number of samples in width and / or height relative to the luma component, depending on the color format considered.

[0223] A sample is the smallest visual information unit that makes up a component of a video picture. A sample value can be, for example, a luminance or chrominance value, or a color value in (R, G, B) format.

[0224] A pixel value is the value of a pixel on the screen. For a monochrome video screen, a pixel value can be represented by a single sample, while for a color video screen, a pixel value can be represented by multiple co-located samples. The co-located samples associated with a pixel refer to samples corresponding to the position of the pixel on the screen.

[0225] A video frame is usually viewed as a set of pixel values, with each pixel represented by at least one sample.

[0226] A block of a video picture is a set of samples of a component of a video picture. When the picture / video format is the well-known (Y, Cb, Cr) format, it can be considered a block of at least one luma sample or a block of at least one chroma sample, or when the picture / video format is the well-known (R, G, B) format, it can be considered a block of at least one color sample.

[0227] At least one embodiment is not limited to a particular picture / video format.

[0228] In general, the present application relates to various embodiments to reduce the complexity of the CPMV refinement method of JVET-AB0178 while preserving as much as possible the compression efficiency achieved with JVET-AB0178.

[0229] In a first embodiment, the MV / CPMV refinement search is optimized by systematically storing the results of the affine DMVR, ie the best MV and lowest BM cost obtained (output of the first or second step of the affine DMVR method).

[0230] In ECM, in the second step of affine DMVR and steps 171 and 172, the affine DMVR search scheme is processed and sometimes repeated several times to improve compression efficiency by selecting MVs that provide a good BM cost. When repeated several times, if the previous iteration does not provide any better solution, the second step of the affine DMVR method can be skipped.

[0231] When several iterations of the affine DMVR search scheme are used, the initial MV and the best BM cost are updated with the best results (best CPMV and best BM cost) obtained in the first few iterations.

[0232] Assume that position 0 is the best position. Figure 18 Shows the positions that will be evaluated in the next iteration. The positions corresponding to the gray rectangles are the positions evaluated in the previous iteration. The positions corresponding to the diagonal rectangles are the positions evaluated in the previous iteration and will be reused in the current iteration. Positions 0, 1, 2, 7, and 6 will be evaluated in the new iteration.

[0233] In ECM, when multiple iterations of the refinement process are triggered, the results of the iteration (best MV / CPMV and lowest distortion) are stored (saved) only if it is the first iteration. By doing so, when this refinement is combined with other algorithms, the subsequent algorithms will not be initialized with the best identified MV / CPMV or will not use the lowest achieved BM cost.

[0234] According to a first embodiment, after each iteration of the first or second step of the affine DMVR method, the results (best MV and lowest BM cost) are systematically stored.

[0235] In a second embodiment, the search scheme can be divided into two positions:

[0236] • Primary positions, which may be, for example, left, top, right and bottom, i.e. Figure 16 The positions in {1,3,5,7}.

[0237] Corner positions, which may be, for example, the upper left corner, the upper right corner, the lower right corner, and the lower left corner, i.e., Figure 16 The positions in {0,2,4,6}.

[0238] According to a second embodiment, the main position is evaluated first and the next corner position is evaluated based on the results of the neighbouring positions of said corner position.

[0239] In a first variation of the second embodiment, if none of the neighboring positions of a corner position is the current best position, then the evaluation of the neighboring positions of the corner position is skipped.

[0240] For example, according to Figure 16 The evaluation of position 0 (the corner position) is processed based on the results of positions 1, 7, and the initial position x (the neighboring position of corner position 0). If position 1, position 7, and the initial position x are not the current best position of the iteration, then the probability that position 0 is the best position is low. In this case, it is beneficial to skip the evaluation of position 0 to save runtime.

[0241] In a second variation of the second embodiment, a corner position is evaluated if the BM costs of its two neighboring positions are both within a given range of the current best BM cost.

[0242] The range may be defined as from 0 to the best distortion weighted by a factor strictly higher than one.

[0243] For example, let's consider position 7 (the neighboring position of corner position 0) as the best position. Using the first variant of the second embodiment, positions 2 and 4 are skipped and positions 1, 7, and x are evaluated. However, if the BM cost of position 1 is much higher than the current best BM cost (i.e., the BM cost of position 7), then the probability that position 0 is the best position is low. If the BM costs of positions 1, 7, and x are all within a given range (not low), then the probability that position 0 is the best position is high.

[0244] In a third embodiment, between the two reference video picture lists L0 and L1, the best MV that provides the lowest BM cost is selected. Motion vector candidate refinement (the second step of affine DMVR) is performed for a given list (corresponding to one of the reference video picture lists (either L0 or L1)), i.e., only one MV is refined. The MV refinement offset follows the conventional scheme and search point order used by affine DMVR, as described above in conjunction with Figure 16 discussed.

[0245] At the encoder side, the two reference video picture lists MV are sequentially and independently refined to identify and select the best prediction mode.

[0246] In order to measure the BM cost, the reference video picture list prediction needs to be available.

[0247] In the prior art, for a given PU, the distortion is evaluated (e.g., by SAD) per sub-block (from PU size to 4x4 block) and two reference video picture list predictions are calculated. A sub-block is an aggregation of consecutive 4x4 blocks (at least one) that share the same MV.

[0248] In order to optimize this estimation of distortion (e.g., by SAD evaluation), the prediction for the unrefined reference video picture list remains unchanged / fixed and can be calculated only once if it is stored for a given PU in an intelligent way. For example, for the first defined MV refinement offset, both predictions are calculated, while in the following MV refinement offsets, only the prediction for the refined video picture list is calculated, thus saving half the prediction calculations.

[0249] Furthermore, instead of calculating distortion at the sub-block level (e.g., by SAD evaluation), it can be calculated at the PU level. In the prior art, the distortion between the predictions of the dual reference video picture lists (e.g., by SAD evaluation) is calculated per sub-block and summed over all sub-blocks to obtain the distortion at the PU level (e.g., by SAD evaluation). It is then possible to generate predictions per sub-block and then measure the distortion at the PU level. This helps reduce the number of function calls (once per PU, rather than once per sub-block) and uses more efficient intrinsics to reduce runtime processing when processing larger pixel surfaces.

[0250] In a fourth embodiment, step 172 is bypassed for a CPMV based on a comparison between a two-sided matching cost associated with the CPMV and a threshold value performed in the first step 171 .

[0251] This fourth embodiment avoids spending time evaluating the merge candidates, thereby reducing the runtime consumption of the affine DMVR method, because step 172 is the most time-consuming step in the affine DMVR method, as described in relation to Figure 17 described.

[0252] In a first variant of the fourth embodiment, if the cost of the BM performed for a CPMV in step 171 is below a threshold value TH1 , step 172 is bypassed for said CPMV.

[0253] Then, the CPMV is partially optimized (refined).

[0254] This early termination assumes that the BM cost (distortion) is already low and the refinement effect is significantly better than the rate-distortion trade-off of the current merge candidate.

[0255] In a second variant of the fourth embodiment, if the cost of the BM performed for a CPMV in step 171 is higher than a threshold value TH2, step 172 is bypassed for said CPMV.

[0256] Then, the CPMV is partially optimized (refined).

[0257] This early termination assumes that the BM cost (distortion) is quite high and the merge candidates are not competitive enough.

[0258] In a third variant of the fourth embodiment, the thresholds TH1 and / or TH2 are either adapted to the PU size or are fixed.

[0259] During iterative refinement (steps 171 and 172), the CPMVs are sequentially refined to improve compression efficiency. The affine DMVR method then attempts to refine the first CPMV, then the second CPMV, and finally, in the case of a 6-parameter affine model, the third CPMV. Refinement is performed exhaustively, i.e., all CPMVs are refined regardless of the result of the previous CPMV refinement.

[0260] In a third variation of the fourth embodiment, when the BM cost (distortion) associated with either the first or second CPMV satisfies the bypass condition, the second step 172 is bypassed for the third CPMV.

[0261] In an example embodiment of the third variation of the fourth embodiment, the bypass condition is met when the BM cost associated with the first or second CPMV performed in the second step 172 is higher than the BM cost associated with the first or second CPMV performed in the first step 171 .

[0262] In another example embodiment of the third variant of the fourth embodiment, the bypass condition is met when the BM cost associated with the first or second CPMV performed in the second step 172 is lower than the BM cost associated with the first or second CPMV performed in the first step 171 and if the reduction in the bilateral cost is lower than a third threshold TH3.

[0263] In a fifth embodiment, the affine DMVR method is disabled for specific block (CU) shapes.

[0264] In the prior art, affine DMVR is performed on each CU whose CU height and CU width are both greater than or equal to 8 luma samples. However, the affine DMVR approach can offer different runtime / compression efficiency tradeoffs for all CU shapes. Therefore, disabling affine DMVR for specific CU shapes is worth considering, for example, when the runtime / compression efficiency tradeoff is unsatisfactory. By identifying CU shapes for which affine DMVR performs poorly, affine DMVR can be disabled to save runtime.

[0265] This fifth embodiment is further advantageous because when affine DMVR is bypassed for a CU with a specific shape, the CU syntax signaled for this CU may also be bypassed to further reduce the bitrate.

[0266] In a first variation of the fifth embodiment, when the width or height of a block (CU) to be inter-coded is higher than a first value (e.g., 128) when the width and height are equal, the affine DMVR method (steps 171 and 172) may be disabled.

[0267] In a second variation of the fifth embodiment, the affine DMVR method may be disabled (steps 171 and 172 ) when the width or height of the block (CU) to be inter-coded is higher than a second value (e.g., 64) when the width is different from the height.

[0268] In variations of the first and second variations of the fifth embodiment, the CU shape (width and height of the inter-coded CU) with affine DMVR disabled may be signaled into the bitstream to be copied at the decoder side.

[0269] In a first embodiment of the variant, at least one specific CU shape for which affine DMVR is disabled (steps 171 and 172) may be signaled by signaling a maximum and / or minimum CU size (i.e., a maximum and minimum CU width and / or height expressed as, for example, a number of pixels).

[0270] In a second embodiment of the variant, the at least one block shape for which affine DMVR is disabled (steps 171 and 172) may be signaled by signaling a minimum and / or maximum CU ratio between a maximum value of CU width and block height and a minimum value of CU width and CU height.

[0271] In the prior art, for each inter mode, a list of candidates is generated using neighbor blocks, temporal information, and combined spatial and temporal information, as discussed in the introductory part of this application.

[0272] Regarding the affine DMVR method, available spatial and temporal information is used to generate a merge candidate list. The merge candidate list is not optimal and the same merge candidate may be repeated in the merge candidate list.

[0273] Duplicate merge candidates are suboptimal for several reasons.

[0274] First, duplicate merge candidates are evaluated by the DMVR method like merge candidates, but are not selected because they provide lower compression efficiency due to their higher merge index compared to merge candidates with lower merge indexes. Then, evaluating duplicate merge candidates increases the runtime of the DMVR method without increasing compression.

[0275] Secondly, the presence of duplicate merge candidates in the merge candidate list degrades the rate-distortion performance of the merge candidates following the duplicate merge candidate, since it uselessly increases the merge index associated with the subsequent merge candidate.

[0276] Third, when the merge candidate list is bounded, several potential merge candidates are discarded when the maximum size of the list is reached while there are useless duplicate merge candidates in the merge candidate list.

[0277] In order to overcome such disadvantages, in a sixth embodiment, it is checked whether a new merge candidate intended to be added to the merge candidate list already exists in the merge candidate list.

[0278] The sixth embodiment helps limit the runtime overhead of the DMVE method while providing better compression performance compared to a merge candidate list including at least one duplicate merge candidate.

[0279] In a variation of the sixth embodiment, the reference video picture index of the new merge candidate is compared with the reference video picture index of the merge candidate in the merge candidate list and the MVs of the two reference video picture lists (L0 and L1). In the case of affine blocks, all CPMVs of the two lists are compared. More precisely, each affine merge index points to motion predictor information, which indicates which CPMVs are used to derive motion information. The motion information is represented by a unidirectional or bidirectional temporal prediction type, up to 2 reference video picture indices and up to 3 control point motion vectors (CPMVs), each CPMV being associated with a reference video picture index of one of the two reference video picture lists (L0 or L1). Then, when comparing two affine merge candidates, the motion information of the two lists (L0 or / and L1) is compared. For the lists, this includes comparing the reference video picture indices as well as the CPMVs.

[0280] If the reference index and MV / CPMV are exactly the same, then the potential candidate is identified as a duplicate.

[0281] A variation of the sixth embodiment discards only strictly duplicate merge candidates. It can be enhanced by discarding potential candidates whose MV / CPMV is too close to a candidate in a list.

[0282] For example, if all differences between the MV / CPMV of a potential candidate and the MV / CPMV of one of the merge candidates in the list are below a threshold, the potential candidate is discarded.

[0283] The comparison involves comparing reference indices, calculating the MV / CPMV difference, and comparing the MV / CPMV difference to a threshold. In this example, the goal is to compare the motion information of two candidates. Unlike the previous description, the CPMV criterion is amplified, resulting in the rejection of further candidates. This now involves rejecting candidates with similar CPMVs to those of the listed candidates, i.e., all CPMV differences between the two merged candidates are below the threshold.

[0284] If the reference index is exactly the same and the MV / CPMV difference is lower than the threshold TH4, the potential candidate will not be added to the merge candidate list. The threshold TH4 can be either fixed or adaptive, such as adaptively adjusted according to the PU size.

[0285] Regarding the DMVR method, at the encoder side, all merge candidates in the merge candidate list are refined, that is, the final merge candidate may be different from the merge candidates of the list.

[0286] Therefore, the merge candidates in the list can be matched with the merge candidates refined by affine DMVR from the list. More precisely, the encoder will perform the affine DMVR algorithm on all affine merge candidates in the list. As described below, affine DMVR refines the CPMV of the affine merge candidates, thereby changing the CPMV. As a result, the Nth refined affine merge candidate in the list is likely to obtain motion information (video picture reference index and CPMV) that is equal to or similar (i.e., has a close CPMV) to the affine merge candidates with a merge index higher than N.

[0287] In this case, applying the DMVR method to matching merge candidates is inefficient. Therefore, on the encoder side (which needs to evaluate each merge candidate), it is possible to bypass the merge candidate DMVR refinement when the merge candidate is similar to the DMVR-refined merge candidate. A similar strategy as above can be used to discard the merge candidate DMVR refinement.

[0288] Figure 19 Shown is a schematic block diagram illustrating an example of a system 600 in which various aspects and embodiments are implemented.

[0289] System 600 may be embedded as one or more devices, including various components described below. In various embodiments, system 600 may be configured to implement one or more aspects described herein.

[0290] Examples of equipment that may constitute all or part of system 600 include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMDs, perspective glasses), projectors (projectors), "caves" (systems comprising multiple displays), servers, video encoders, video decoders, post-processors for processing outputs from video decoders, pre-processors for providing inputs to video encoders, web servers, video servers (e.g., broadcast servers, video-on-demand servers, or web servers), still or video cameras, encoding or decoding chips, or any other communication devices. The elements of system 600 may be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 600 may be distributed across multiple ICs and / or discrete components. In various embodiments, system 600 may be communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports.

[0291] The system 600 may include at least one processor 610 configured to execute instructions loaded therein to implement, for example, various aspects described herein. The processor 610 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 600 may include at least one memory 620 (e.g., a volatile memory device and / or a non-volatile memory device). The system 600 may include a storage device 640, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, magnetic disk drive, and / or optical disk drive. As non-limiting examples, the storage device 640 may include an internal storage device, an attached storage device, and / or a network accessible storage device.

[0292] The system 600 may include an encoder / decoder module 630, which is configured to, for example, process data to provide encoded / decoded video picture data, and the encoder / decoder module 630 may include its own processor and memory. The encoder / decoder module 630 may represent a module(s) that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. In addition, the encoder / decoder module 630 may be implemented as a separate element of the system 600, or may be incorporated into the processor 610 as a combination of hardware and software known to those skilled in the art.

[0293] Program code to be loaded onto the processor 610 or the encoder / decoder 630 to perform various aspects described herein may be stored in the storage device 640 and subsequently loaded onto the memory 620 for execution by the processor 610. According to various embodiments, during execution of the processes described herein, one or more of the processor 610, the memory 620, the storage device 640, and the encoder / decoder module 630 may store one or more of various items. Such stored items may include, but are not limited to, video picture data, information data for encoding / decoding video picture data, bitstreams, matrices, variables, and intermediate or final results of processing of equations, formulas, operations, and operational logic.

[0294] In several embodiments, memory internal to the processor 610 and / or encoder / decoder module 630 may be used to store instructions and provide working memory for processes that may be performed during encoding or decoding.

[0295] However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 610 or the encoder / decoder module 630) may be used for one or more of these functions. The external memory may be a memory 620 and / or a storage device 640, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, the external non-volatile flash memory may be used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM may be used as working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 Video), AVC, HEVC, EVC, VVC, AV1, etc.

[0296] As indicated in block 690, input to the elements of system 600 may be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that may receive an RF signal transmitted over the air, for example, by a broadcast device, (ii) a composite input terminal, (iii) a USB input terminal, (iv) an HDMI input terminal, (v) a bus such as a CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data-rate), FlexRay (ISO 17458), or Ethernet (ISO / IEC 802-3) bus when the present application is implemented in the automotive field.

[0297] In various embodiments, the input device of block 690 may have associated corresponding input processing elements, as known in the art. For example, the RF section may be associated with elements necessary for: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting the signal to a narrower frequency band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired packet stream. The RF section of various embodiments may include one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband.

[0298] In a set-top box embodiment, the RF section and its associated input processing elements can receive RF signals transmitted on a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.

[0299] Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0300] Adding an element may include inserting an element between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter.In various embodiments, the RF portion may include an antenna.

[0301] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 600 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 610, if necessary. Similarly, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 610, if necessary. The demodulated, error-corrected, and demultiplexed streams may be provided to various processing elements, including, for example, the processor 610 and an encoder / decoder 630, which operate in conjunction with memory and storage elements to process the data streams for presentation on an output device, if necessary.

[0302] The various elements of system 600 may be provided within an integrated housing. Within the integrated housing, a suitable connection arrangement 690, such as an internal bus (including an I2C bus), wiring, and printed circuit boards known in the art, may be used to interconnect the various elements and transfer data between them.

[0303] System 600 may include a communication interface 650 that enables communication with other devices via a communication channel 651. Communication interface 650 may include, but is not limited to, a transceiver configured to send and receive data over communication channel 651. Communication interface 650 may include, but is not limited to, a modem or a network card, and communication channel 651 may be implemented, for example, within a wired and / or wireless medium.

[0304] In various embodiments, a Wi-Fi network such as IEEE 802.11 may be used to stream data to system 600. Wi-Fi signals of these embodiments may be received via a communication channel 651 adapted for Wi-Fi communications and a communication interface 650. Communication channel 651 of these embodiments may typically be connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications.

[0305] Other embodiments may provide streamed data to the system 600 using a set-top box that delivers the data through the HDMI connection of the input box 690 .

[0306] Still other embodiments may use the RF connection of input box 690 to provide streaming data to system 600 .

[0307] The streamed data may be used in the form of signaling information used by the system 600. The signaling information may include the bitstream B and / or information such as the number of pixels of a video picture and / or any encoding / decoding setting parameters.

[0308] It should be appreciated that signaling may be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. may be used to signal information to a corresponding decoder.

[0309] The system 600 can provide output signals to various output devices, including a display 661, speakers 671, and other peripheral devices 681. In various examples of embodiments, the other peripheral devices 681 can include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of the system 600.

[0310] In various embodiments, control signals may be communicated between the system 600 and the display 661, speakers 671, or other peripheral devices 681 using signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.

[0311] Output devices may be communicatively coupled to system 600 via dedicated connections through respective interfaces 660 , 670 , and 680 .

[0312] Alternatively, the output device may be connected to the system 600 via the communication interface 650 using the communication channel 651. The display 661 and the speaker 671 may be integrated with the other components of the system 600 in a single unit in an electronic device such as, for example, a television.

[0313] In various embodiments, the display interface 660 may include a display driver, such as, for example, a timing controller (TCon) chip.

[0314] For example, if the RF portion of input 690 is part of a separate set-top box, then display 661 and speaker 671 may optionally be separate from one or more of the other components. In various embodiments where display 661 and speaker 671 may be external components, the output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.

[0315] exist Figure 1-19 In the present invention, various methods are described herein, and each method includes one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.

[0316] Some examples are described with respect to block diagrams and / or operational flow charts. Each block represents a portion of a circuit element, module, or code that includes one or more executable instructions for implementing (one or more) specified logical functions. It should also be noted that, in other embodiments, the (one or more) functions marked in the block may not occur in the order indicated. For example, depending on the functions involved, two blocks shown in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order.

[0317] The various embodiments and aspects described herein may be implemented in, for example, a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if only discussed in the context of a single form of embodiment (e.g., discussed only as a method), the embodiments of the features discussed may also be implemented in other forms (e.g., an apparatus or a computer program).

[0318] The method may be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes a communication device.

[0319] In addition, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored on a computer-readable storage medium. A computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having a computer-readable program code implemented thereon that can be executed by a computer. Considering the inherent ability to store information therein and the inherent ability to provide retrieval of information therefrom, a computer-readable storage medium as used herein can be considered to be a non-transitory storage medium. A computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or apparatus, or any suitable combination of the foregoing. It should be appreciated that although more specific examples of computer-readable storage media to which the various embodiments of the present application can be applied are provided below, as those of ordinary skill in the art will readily appreciate, they are merely illustrative and not an exhaustive list: a portable computer floppy disk; a hard disk; a read-only memory (ROM); an erasable programmable read-only memory (EPROM or flash memory); a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing.

[0320] The instructions may form an application program tangibly embodied on a processor-readable medium.

[0321] For example, instructions may be in hardware, firmware, software, or a combination thereof. For example, instructions may be found in an operating system, a separate application, or a combination of the two. Thus, a processor may be characterized as, for example, a device configured to perform a process and a device that includes a processor-readable medium (such as a storage device) having instructions for performing the process. Additionally, in addition to or in lieu of instructions, a processor-readable medium may store data values generated by an embodiment.

[0322] Device can be realized in suitable hardware, software and firmware for example.The example of this device comprises personal computer, laptop computer, smart phone, tablet computer, digital multimedia set-top box, digital television receiver, personal video recording system, the household appliance of connection, head-mounted display device (HMD, perspective glasses), projector (projector), " cave " (comprising the system of multiple displays), server, video encoder, video decoder, process the post-processor of the output from video decoder, provide the pre-processor of input to video encoder, web server, set-top box, and any other equipment for processing video picture, or other communication equipment.It should be clear that equipment can be mobile and even be installed in mobile vehicle.

[0323] Computer software may be implemented by processor 610 or by hardware, or by a combination of hardware and software. As a non-limiting example, various embodiments may also be implemented by one or more integrated circuits. Memory 620 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). Processor 610 may be of any type suitable for the technical environment and may encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.

[0324] An embodiment of the present application further provides an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor. The processor is configured to execute the method described in any of the above embodiments.

[0325] As will be apparent to one of ordinary skill in the art, embodiments may generate various signals formatted to carry information that can be stored or transmitted, for example. The information may include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.

[0326] The terms used herein are only used to describe the purpose of specific embodiments and are not intended to be limiting. As used herein, the singular forms "one", "a kind of" and "the / said" may also be intended to include plural forms, unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms "include / comprise" and / or "including / comprising" may specify the existence of stated, for example, features, integers, steps, operations, elements and / or components, but do not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. Moreover, when an element is referred to as "in response to" or "connected to" another element or "associated with another element", it may directly respond to or be connected to another element, or there may be an intermediate element. In contrast, when an element is referred to as "directly responding to" or "directly connected to" another element or "directly associated with another element", there is no intermediate element.

[0327] It should be appreciated that use of any of the symbols / terms " / ," "and / or," and "at least one of" in the context of, for example, "A / B," "A and / or B," and "at least one of A and B" may be intended to encompass selection of only the first listed option (A), or only the second listed option (B), or both options (A and B). As a further example, in the context of "A, B, and / or C" and "at least one of A, B, and C," such wording is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). This may be extended to as many items as listed, as will be apparent to one of ordinary skill in this and related arts.

[0328] Various numerical values may be used in this application. Specific values may be used for illustrative purposes and the described aspects are not limited to these specific values.

[0329] It will be understood that although the terms first, second, etc. can be used to describe various elements in this article, these elements are not limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the teachings of this application, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. There is no suggestion of sorting between the first element and the second element.

[0330] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in conjunction with the embodiment / implementation) is included in at least one embodiment / implementation. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any other variations thereof appearing in various places in this application are not necessarily all referring to the same embodiment.

[0331] Similarly, references herein to "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in conjunction with an embodiment / example / implementation) may be included in at least one embodiment / example / implementation. Thus, the phrases "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" appearing in various places in this application are not necessarily all referring to the same embodiment / example / implementation, nor are separate or alternative embodiments / examples / implementations necessarily mutually exclusive of other embodiments / examples / implementations.

[0332] Reference numerals appearing in the claims are for illustration purposes only and have no limiting effect on the scope of the claims. Although not explicitly described, the various embodiments / examples and variations of the present application may be employed in any combination or subcombination.

[0333] When a figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.

[0334] While some diagrams include arrows on communication paths to illustrate a primary direction of communication, it should be understood that communication can occur in the opposite direction to the depicted arrows.

[0335] Various embodiments relate to decoding. As used in this application, "decoding" can encompass, for example, all or part of a process performed on a received video picture (which may include a received bitstream encoding one or more video pictures) to produce a final output suitable for display or further processing in a reconstructed video domain. In various embodiments, such a process includes one or more of the processes typically performed by a decoder. In various embodiments, for example, such a process also or alternatively includes a process performed by a decoder of the various embodiments described herein.

[0336] As a further example, in one embodiment, "decoding" may refer only to dequantization, in one embodiment, "decoding" may refer to entropy decoding, in another embodiment, "decoding" may refer only to differential decoding, and in another embodiment, "decoding" may refer to a combination of dequantization, entropy decoding, and differential decoding. Depending on the context of the particular description, whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will be clear and is believed to be well understood by those skilled in the art.

[0337] Various embodiments relate to encoding. In a manner similar to the discussion above regarding "decoding," "encoding" as used herein may encompass, for example, all or part of a process performed on an input video frame to produce an output bitstream. In various embodiments, such a process includes one or more of the processes typically performed by an encoder. In various embodiments, such a process also includes or alternatively includes a process performed by an encoder of the various embodiments described herein.

[0338] As a further example, in one embodiment, "encoding" may refer only to quantization, in one embodiment, "encoding" may refer only to entropy coding, in another embodiment, "encoding" may refer only to differential coding, and in another embodiment, "encoding" may refer to a combination of quantization, differential coding, and entropy coding. Based on the context of the particular description, whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process will be clear and is believed to be well understood by those skilled in the art.

[0339] Furthermore, this application may refer to "obtaining" various information. Obtaining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory, processing information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0340] Furthermore, this application may refer to "receiving" various pieces of information. Receiving information may include, for example, one or more of: accessing the information or receiving the information from a communication network.

[0341] Furthermore, as used herein, the term "signal" specifically refers to indicating something to a corresponding decoder. For example, in some embodiments, an encoder signals specific information, such as codec parameters or encoded video frame data. In this way, in embodiments, the same parameters can be used on both the encoder and decoder sides. Thus, for example, an encoder can transmit (explicit signaling) specific parameters to a decoder so that the decoder can use the same specific parameters. Conversely, if the decoder already has specific parameters along with other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the specific parameters. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It should be appreciated that signaling can be accomplished in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the verb form of the term "signal" is mentioned above, the term "signal" can also be used as a noun in this document.

[0342] A number of embodiments have been described. However, it should be understood that various modifications may be made. For example, elements of different embodiments may be combined, supplemented, modified, or removed to produce other embodiments. Furthermore, it will be understood by those skilled in the art that other structures and processes may replace the disclosed structures and processes, and that the resulting embodiments will perform at least substantially the same function(s) in at least substantially the same manner(s) to achieve at least substantially the same result(s) as the disclosed embodiments. Thus, these and other embodiments are contemplated herein.

Claims

1. A method for decoding a video image, wherein The decoding is done by means of temporal bidirectional prediction with motion compensation for inter-coded blocks. The motion compensated temporal bidirectional prediction uses: two reference video pictures in two separate reference video picture lists; and two affine motion fields defined by at least two control point motion vectors, wherein the control point motion vectors are denoted as CPMVs, wherein the at least two control point motion vectors are associated with each reference picture, The refined CPMV is obtained as an output of the following first step (171), or alternatively, as an output of a second step (172) following said first step (171), - in said first step (171), for each CPMV, performing bilateral matching on a block centered around said CPMV to derive at least two CPMVs refined with integer precision, and selecting a set of CPMVs comprising non-refined CPMVs and / or CPMVs refined with integer precision to produce an overall predicted block with a minimum bilateral matching cost, - in said second step (172), for each successive CPMV in the selected set of CPMVs associated with a block, refining each CPMV with subsample accuracy to minimize the bilateral matching cost of said block, wherein the second step (172) is bypassed based on a comparison between a two-sided matching cost associated with the CPMV performed in the first step and a threshold value.

2. The method of claim 1, wherein the second step (172) is bypassed for the CPMV when the bilateral matching cost associated with the CPMV performed in the first step (171) is below a first threshold (TH1).

3. The method of claim 1, wherein when a bilateral matching cost associated with a CPMV performed in the first step (171) is higher than a second threshold (TH2), the second step (172) is bypassed for the CPMV.

4. The method according to claim 2 or 3, wherein the first threshold (TH1) and / or the second threshold (TH2) are fixed or adapted to the prediction unit size.

5. A method as described in any one of claims 1 to 4, wherein the two affine motion fields are defined by a first CPMV, a second CPMV and a third CPMV, and wherein the second step is bypassed for the third CPMV when the bilateral matching cost associated with the first CPMV or the second CPMV meets a bypass condition.

6. The method of claim 5, wherein the bypass condition is satisfied when the bilateral cost associated with the first CPMV or the second CPMV performed in the second step (172) is higher than the bilateral matching cost associated with the first CPMV or the second CPMV performed in the first step (171).

7. The method of claim 5, wherein the bypass condition is satisfied when the bilateral matching cost associated with the first CPMV or the second CPMV performed in the second step (172) is lower than the bilateral matching cost associated with the first CPMV or the second CPMV performed in the first step (171), and if the reduction in the bilateral cost is lower than a third threshold (TH3).

8. The method of any one of the preceding claims 1 to 7, wherein the first and second steps are disabled for certain block shapes.

9. The method of claim 8, wherein the first step and the second step are disabled when a width or a height of a block to be inter-coded is higher than a first value when the width is equal to the height.

10. The method of claim 8 or 9, wherein the first step and the second step are disabled when a width or a height of a block to be inter-coded is higher than a second value when the width is different from the height.

11. The method of any one of claims 8 to 10, wherein a block shape is defined according to a width or a height of the inter-coded block, and wherein at least one block shape disabling the first step and the second step is signaled in the bitstream.

12. The method of claim 11, wherein the at least one block shape disabling the first and second steps is signaled by signaling a maximum and a minimum block width and / or height.

13. The method of claim 11 , wherein the at least one block shape disabling the first and second steps is signaled by signaling a minimum and / or maximum block ratio between a maximum value of block width and block height and a minimum value of block width and block height.

14. An apparatus comprising means for performing one of the methods according to any one of claims 1 to 13.

15. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 13.

16. A non-transitory storage medium carrying program code instructions for executing the method according to any one of claims 1 to 13.

17. An electronic device comprising: processor; and a memory for storing instructions executable by said processor; The processor is configured to perform the method according to any one of claims 1 to 13.