Video decoding method and related electronic device

By implementing mutual exclusion rules of codec mode or tool in the high-efficiency video encoding and decoding standard HEVC, the problems of low encoding efficiency and high hardware complexity in the inter prediction mode are solved, and more efficient encoding and simplified hardware design are achieved.

CN120302056APending Publication Date: 2025-07-11HFI INNOVATION INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510574082.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-02-26
Filing Date
2020-02-27
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing high-efficiency video encoding and decoding standard HEVC has problems with low encoding efficiency in inter prediction mode, especially in motion vector prediction and inter prediction mode selection, resulting in high hardware implementation complexity and increased pipeline delay.

Method used

The mutual exclusion rules are used to limit the cascade of different codec tools or modes, ensuring that only one codec mode is enabled in the same current codec unit CU. By implementing the mutual exclusion group of codec mode or tool, hardware design is simplified, pipeline stage is reduced, and hardware utilization is improved.

Benefits of technology

By limiting the simultaneous activation of multiple codec tools or modes, simplifying hardware design, reducing pipeline latency, improving hardware utilization, and improving coding efficiency and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302056A_ABST
    Figure CN120302056A_ABST
Patent Text Reader

Abstract

A video decoder implementing a codec mode mutually exclusive group is provided. The video decoder receives data of a block of pixels to be decoded as a current block of a current picture of a video. When a first codec mode of the current block is enabled, a second codec mode of the current block is disabled, where the first codec mode and the second codec mode specify different methods for calculating an inter prediction of the current block. The current block is decoded by using inter prediction calculated according to the enabled codec mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to video processing. Specifically, the present invention relates to a method for signaling encoding and decoding modes. Background Art

[0002] Unless otherwise indicated herein, the schemes described in this section are not prior art to the claims listed below and are not admitted to be prior art by inclusion in this section.

[0003] High Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). HEVC is based on a hybrid block-based motion compensated DCT-like transform coding architecture. The basic unit of compression is a 2N×2N square block, termed a coding unit (CU), and each CU can be recursively split into four smaller CUs until a predetermined minimum size is reached. Each CU contains one or more prediction units (PUs).

[0004] To achieve the best coding efficiency of the hybrid coding architecture in HEVC, each PU has two types of prediction modes, which are intra prediction and inter prediction. For the intra prediction mode, spatially adjacent reconstructed pixels can be used to generate a directional prediction. There are up to 35 directions in HEVC. For the inter prediction mode, temporally reconstructed reference frames can be used to generate motion compensated predictions. There are three different modes, including Skip, Merge, and Advanced Motion Vector Prediction (AMVP) mode.

[0005] When a PU is coded and decoded in the inter AMVP mode, motion compensated prediction is performed with the transmitted Motion Vector Difference (MVD), and the MVD can be used together with a Motion Vector Predictor (MVP) to generate a Motion Vector (MV). To determine the MVP in the inter AMVP mode, the Advanced Motion Vector Prediction (AMVP) scheme is used to select a motion vector predictor from an AMVP candidate set including two spatial MVPs and one temporal MVP. Thus, in the AMVP mode, the MVP index of the MVP and the corresponding MVD need to be coded and transmitted. In addition, the inter prediction direction specifying the prediction direction in bidirectional prediction and unidirectional prediction (which are List 0 (L0) and List 1 (L1)) and the reference frame index of each list should also be coded and transmitted.

[0006] When a PU is decoded or encoded in skip or merge mode, no motion information is transmitted except for the merge index of the selected candidate. This is because the skip and merge modes utilize motion inference methods (MV = MVP + MVD, where MVD = 0) to obtain motion information from spatially adjacent blocks (spatial candidates) or temporal blocks (temporal candidates) located in the collocated picture, where the collocated picture is the first reference picture in list 0 or list 1, which is signaled in the slice header. In the case of a skipped PU, the residual signal is also omitted. To determine the merge index for the skip and merge modes, a merge scheme is used to select a motion vector predictor from a set of merge candidates that includes four spatial MVPs and one temporal MVP. SUMMARY OF THE INVENTION

[0007] The following summary is illustrative only and is not intended to be limiting in any way. That is, the following summary is provided to introduce the concepts, highlights, benefits, and advantages of the novel and non-obvious techniques described herein. Selected, rather than all, embodiments are further described in the detailed description. Accordingly, the following summary is not intended to identify the essential features of the claimed subject matter, nor is it intended to be used to determine the scope of the claimed subject matter.

[0008] Embodiments of the present invention provide a video decoder that implements mutually exclusive groups of coding modes or tools. The decoder receives data for a block of a current picture that is to be decoded as a video. When a first coding mode for the current block is enabled, the decoder disables a second coding mode for the current block, where the first coding mode and the second coding mode specify different methods for calculating the inter-picture prediction for the current block. In other words, the second coding mode for the current block can be applied only when the first coding mode is disabled. The decoder decodes the current block by using the inter-picture prediction calculated according to the enabled coding mode.

[0009] The present invention proposes a mutually exclusive rule for the exclusion setting of some coding tools for the current CU, so that the pipeline stage can be shorter and the hardware utilization rate can be higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings are included to provide a further understanding of the present invention and are incorporated into and constitute a part of this specification. The drawings illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention. Since some components may be shown out of proportion to their actual dimensions for the purpose of clearly illustrating the concepts of the present invention, the drawings are not necessarily drawn to scale.

[0011] Figure 1 Motion candidates for the merge mode are marked.

[0012] Figure 2Conceptually shows encoding or decoding a current block using decoder-side motion vector refinement based on bilateral-matching.

[0013] Figure 3 Shows the search process of Decoder Motion Vector Refinement (DMVR).

[0014] Figure 4 Marks the DMVR integer luminance sample search pattern.

[0015] Figure 5 Conceptually shows deriving a lighting-based prediction offset.

[0016] Figure 6 Shows an example derivation of the prediction offset.

[0017] Figure 7 Shows the extended CU region used by BDOF for encoding / decoding a CU.

[0018] Figure 8 Shows an exemplary 8×8 transform unit block and a bi-directional filter aperture.

[0019] Figure 9 Shows the filtering process under a Hadamard transform domain filter.

[0020] Figure 10 Shows the adaptive weighting applied along the diagonal edge between two triangular prediction units.

[0021] Figure 11a Conceptually shows encoding or decoding a pixel block by using the MH mode for intra.

[0022] Figure 11b Conceptually shows encoding / decoding the current block by using the MH mode for inter.

[0023] Figure 12 Marks an exemplary video encoder that can implement mutually exclusive groups of encoding / decoding modes or tools.

[0024] Figure 13 Marks a part of the video encoder that can implement mutually exclusive groups of encoding / decoding modes or tools.

[0025] Figure 14Conceptually illustrates a process for implementing mutually exclusive groups of coding modes or tools in a video encoder.

[0026] Figure 15 An exemplary video decoder that can implement mutually exclusive groups of coding modes or tools is marked.

[0027] Figure 16 Shows a part of the video decoder that can implement mutually exclusive groupings of coding modes or tools.

[0028] Figure 17 Conceptually illustrates a process for implementing mutually exclusive groups of coding modes or tools in a video decoder.

[0029] Figure 18 Conceptually illustrates an electronic system with which some embodiments of the present invention are implemented. Detailed Description

[0030] In the following detailed description, many specific series are given by way of example to provide a thorough understanding of the relevant teachings. Various variations, derivations, and / or extensions based on the teachings described herein are within the scope of the present invention. In some cases, well-known methods, processes, components, and / or circuits related to one or more exemplary embodiments disclosed herein may be described at a relatively high level without details to avoid unnecessarily obscuring aspects of the teachings of the present invention.

[0031] I. Merge Mode

[0032] Figure 1 The motion candidates for the merge mode are marked. As shown, up to four spatial MV candidates are derived from A0, A1, B0, and B1, and one temporal MV candidate is derived from T BR or T CTR (first using T BR , and if T BR is not available, using T CTR ). If any of the four spatial MV candidates are not available, then the position B2 is used to derive an MV candidate as an alternative. After the derivation process of the four spatial MV candidates and one temporal MV candidate, in some embodiments, redundancy removal (pruning) is applied to remove redundant MV candidates. If, after redundancy removal (pruning), the number of available MV candidates is less than 5, three types of additional candidates are derived and added to the candidate set (candidate list). The video encoder decides, based on rate-distortion optimization (RDO), to select a final candidate within the candidate set of the skip or merge mode, and transmits an index to the video decoder. (The skip mode and the merge mode are collectively referred to as the "merge mode" herein).

[0033] II. Decoder-Side Motion Vector Refinement (DMVR)

[0034] To increase the accuracy of the MVs in the merge mode, in some embodiments, a decoder-side motion vector refinement or DMVR based on bidirectional matching is applied. In the bidirectional prediction operation, the video codec searches for refined MVs around the initial MVs in the reference picture list L0 and the reference picture list L1. The bidirectional matching method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1.

[0035] Figure 2 Conceptually shows encoding or decoding the current block 200 using decoder-side motion vector refinement based on bidirectional matching. As shown, based on the difference between the pixel blocks referred to by these MV candidates (e.g., R0' and R1'), the SAD (Sum of Absolute Differences) is calculated for the MV candidates (e.g., MV0' and MV1') around the initial MVs (e.g., MV0 and MV1). The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal.

[0036] In some embodiments, DMVR is applied as follows. For DMVR where the luma CB width or height > 16, the CU is split into multiple 16x16, 16x8, or 8x16 luma sub-blocks (and corresponding chroma sub-blocks). Next, when the SAD at the zero MVD position (indicated by the initial MVs, labeled as MV0 and MV1) between list 0 and list 1 is small, the DMVR for each sub-block or small CU is terminated early. Based on a 25-point SAD integer-step search (i.e., a ±2 integer-step refinement search range), the search range fractional samples are generated by bilinear interpolation.

[0037] In some embodiments, when the enabling conditions for DMVR are met, DMVR is applied to the encoded / decoded CU. In some embodiments, the enabling conditions for DMVR can be any subset of (i) to (v). (i) CU-level merge mode with bidirectional prediction MVs; (ii) one reference picture is in the past pictures with respect to the current picture and the other reference picture is in the future pictures with respect to the current picture; (iii) the distances from the two reference pictures to the current picture (e.g., picture order count or POC difference) are the same; (iv) the CU has more than 64 luma samples; (v) both the CU height and the CU width are greater than or equal to 8 luma samples.

[0038] The refined MVs derived by the DMVR process are used to generate inter-prediction samples and also for temporal motion vector prediction for future picture coding. While the original MVs are used for the deblocking process and also for spatial motion vector prediction for future CU coding.

[0039] a. Search scheme

[0040] As Figure 2 shown, the search points surrounding the initial MV and the MV offsets follow the MV difference mirroring principle. In other words, any point examined by the DMVR, labeled as a candidate MV pair (MV0, MV1), follows the following two equations:

[0041] MV0′ = MV0 + MV_offset

[0042] MV1′ = MV1 - MV_offset

[0043] where MV_offset represents the refined offset between the initial MV and the refined MV in one of the reference images. In some embodiments, the refined search range is two integer luminance samples from the initial MV.

[0044] Figure 3 Illustrates the search process of the DMVR. As shown, the search includes an integer sample offset search phase and a fractional sample refinement phase.

[0045] Figure 4 Marks the DMVR integer sample search pattern. As shown, a 25-point full search is applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample phase of the DMVR is ended. Otherwise, the SAD of the remaining 24 points is calculated and examined in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search phase. To reduce the penalty of DMVR refinement uncertainty, it is proposed to prefer the original MV in the DMVR process. The SAD calculated from the initial MV pair will be reduced to 1 / 4 of its SAD value.

[0046] Return to Figure 3 . The integer sample search is followed by the fractional sample refinement. To save computational complexity, the fractional sample refinement is derived by using a parametric error surface equation instead of an additional search using SAD comparison. Based on the output of the integer sample search phase, the fractional sample refinement is conditionally invoked. When the integer sample search phase ends with the minimum SAD at the center in the first iteration or the second iteration, the fractional sample refinement is further applied.

[0047] In the sub-pixel offset estimation based on the parametric error surface, the current position cost and the costs from the center to four adjacent positions are used to fit a 2-D parabolic error surface equation of the following form:

[0048] E(x,y) = A(x - xmin ) 2 +B(y - y min ) 2 +C

[0049] where (x min , y min ) corresponds to the fractional position with the minimum cost and C corresponds to the minimum cost value. By parsing the above equation using the cost values of five search points, (x min , y min ) is calculated as:

[0050] x min = (E(-1, 0) - E(1, 0)) / (2(E(-1, 0) + E(1, 0) - 2E(0, 0)))

[0051] y min = (E(0, -1) - E(0, 1)) / (2((E(0, -1) + E(0, 1) - 2E(0, 0)))

[0052] Since all cost values are integers and the minimum value is E(0, 0), the values of x min and y min are usually automatically constrained to be between -8 and 8. This corresponds to the half peal offset with 1 / 16 pixel MV accuracy in VTM4. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain the sub - pixel accurate refined δMV.

[0053] b. Bilinear interpolation and sample filling

[0054] In some embodiments, the resolution of the MV is 1 / 16 luminance samples. An 8 - tap interpolation filter is used to interpolate the samples at the fractional positions. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so the samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect of using the bilinear filter is that with a 2 - sample search range, DMVR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV with the DMVR search process, a normal 8 - tap interpolation filter is applied to generate the final prediction. To not access more reference samples to the normal MC process, samples that are not needed for the interpolation process based on the original MV but are needed for the interpolation process based on the refined MV are filled from these available samples.

[0055] c. Maximum DMVR processing unit

[0056] In some embodiments, when the width and / or height of a CU is greater than 16 luma samples, it is further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16.

[0057] III. Weighted Prediction (WP)

[0058] Weighted Prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards to efficiently codec video content with fill. Support for WP has also been added to the VVC standard. WP allows weighted parameters (weights and offsets) to be signaled for each reference image in each reference image list L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference images are applied.

[0059] IV. Illumination-based Prediction Offset

[0060] As mentioned before, inter-frame prediction explores the pixel correlation between frames and if the scene is stationary, the correlation will be valid, and motion estimation can easily find similar blocks with similar pixel values in temporally adjacent frames. However, in some practical cases, multiple frames will be captured under different lighting conditions. Even if the content is similar and the scene is stationary, the pixel values between multiple frames will be different.

[0061] In some embodiments, the Neighboring-derived Prediction Offset (NPO) is used to add a prediction offset to improve the motion compensation predictor. According to this offset, different lighting conditions between multiple frames can be considered. The offset is derived using neighboring reconstructed pixels (NRP) and an extended motion compensated predictor (EMCP).

[0062] Figure 5 Conceptually shows the derivation of the illumination-based prediction offset. The patterns selected for NRP and EMCP are N pixels to the left and M pixels above the current PU, where N and M are predetermined values. The pattern can be of any size and shape and can be determined according to any coding parameter, such as the PU or CU size, as long as they are the same for both NRP and EMCP. The offset is calculated as the average pixel value of NRP minus the average pixel value of EMCP. The derived offset is unique for the PU and is applied to the entire PU along with the motion compensation predictor.

[0063] Figure 6An example derivation of the prediction offset is shown. First, for each adjacent position (to the left and above the boundary, shaded in gray), the individual offset is calculated as the corresponding pixel in the NRP minus the pixel in the EMCP. In this example, offset values 6, 4, 2, -2 are generated for the above adjacent positions and 6, 6, 6, 6 for the left adjacent positions. Second, when all individual offsets are calculated and obtained, the derived offset for each position in the current PU will be the average of the offsets from the left and above positions. For example, in the first position in the upper left corner, an offset 6 is generated by averaging the offsets from the left and above. For the next position, the offset is equal to (6 + 4) / 2, i.e., 5. The offsets for each position can be processed and generated in raster scan order. Since adjacent pixels are more highly correlated with the boundary pixels, so are the offsets. This method can adapt the offset according to the pixel position. The derived offset will be adapted to the entire PU and will be applied to each PU position separately together with the motion compensation predictor.

[0064] In some embodiments, local illumination compensation (LIC) is used to correct the result of inter prediction. LIC is a method of inter prediction that uses the neighboring samples of the current block and the reference block to generate a linear model, which is characterized by a scaling factor a and an offset b. The scaling factor a and the offset b are derived by referring to the neighboring samples of the current block and the reference block. For each CU, the LIC mode can be adaptively enabled or disabled.

[0065] V. Generalized Bi-Prediction (GBI)

[0066] Generalized bi-prediction (GBI) is a method of inter prediction that uses different weights for the predictors from L0 and L1, instead of using equal weights as in traditional bi-prediction. GBI is also known as bi-prediction with weighted average (BMA) or bi-prediction with CU-level weights (BCW). In HEVC, a bi-prediction signal is generated by averaging two prediction signals obtained from two different reference images and / or using two different motion vectors. In some embodiments, the bi-prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals.

[0067] P bi-pred = ((8 - w)*P0 + w*P1 + 4) >> 3

[0068] In some embodiments, five different possible weights are allowed in weighted average bidirectional prediction, or w ∈ {-2, 3, 4, 5, 10}. For each bidirectional prediction CU, the weight w is determined in one of two ways: 1) for non-merged CUs, the weight index is signaled after the motion vector difference; 2) for merged CUs, the weighted index is inferred from adjacent blocks based on the merge candidate index. The weighted average of bidirectional prediction is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only three different possible weights are used (w ∈ {3, 4, 5}).

[0069] In some embodiments, in the video encoder, a fast search algorithm is applied to find the weighted index without significantly increasing the encoder complexity. When combined with AMVR, which allows the MVD of CUs to be encoded and decoded with different precisions, unequal weights (the weight used for L0 is not equal to the weight used for L1) are conditionally used for 1-pixel and 4-pixel motion vector precisions if the current picture is a low-delay picture. When combined with affine, the affine motion estimation (ME) will use unequal weights only if the affine mode is selected as the current best mode. Unequal weights are conditionally used when the two reference pictures for bidirectional prediction are the same. When certain conditions are not met, unequal weights are not searched according to the POC (picture order count) between the current picture and its reference pictures, the encoded and decoded QP (quantization parameter), and the temporal level.

[0070] VI. Bidirectional Displacement Optical Flow (BDOF)

[0071] In some embodiments, bidirectional displacement optical flow (BDOF), also known as BIO, is used to refine the bidirectional prediction signal of CUs at the 4x4 sub-block level. In particular, the video codec refines the bidirectional prediction signal by using sample gradients and a set of derived displacements.

[0072] BDOF is applied when the enabling conditions are met. In some embodiments, the enabling conditions for BDOF can be any subset of (1) to (4). (1) Both the CU height and CU width are greater than or equal to 8 luma samples; (2) the CU is not encoded and decoded using the affine mode or the ATMVP merge mode, which belongs to the sub-block merge mode; (3) the CU is encoded and decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in the display order and the other reference picture is after the current picture in the display order; (4) the CU has more than 64 luma samples. In some embodiments, BDOF is applied to the luma component.

[0073] The BDOF mode is based on the concept of optical flow, which assumes that the motion of the object is smooth. For each 4x4 sub-block, the motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 and L1 predicted samples. The motion refinement is then used to adjust the bidirectional predicted sample values in the 4x4 sub-block. The subsequent steps are applied to the BDOF process.

[0074] First, the horizontal and vertical gradients of the two prediction signals are calculated by directly computing the difference between two adjacent samples and k = 0, 1, that is:

[0075]

[0076] where I (k) (i,j) is the sample value at the coordinate (i,j) of the prediction signal in the list k (k = 0, 1), and shift1 is calculated based on the luminance bit depth (bitDepth), such as shift1 = max(6, bitDepth - 6). Then, the auto- and cross-correlations S1, S2, S3, S5, and S6 of the gradients are calculated as follows:

[0077] S1 = ∑ (i,j)∈Ω Abs(ψ x (i,j)), S3 = ∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j))

[0078]

[0079] S5 = ∑ (i , j)∈Ω Abs(ψ y (i,j)),

[0080]

[0081] where

[0082]

[0083] θ(i,j) = (I (1) (i,j) >> n b ) - (I (0) (i,j) >> n b )

[0084] where Ω is a 6x6 window around the 4x4 sub-block and the values of na and nb are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8) respectively. Then motion refinement (v x , v y ) is derived using cross and auto-correlation terms as follows:

[0085]

[0086] Finally, the BDOF samples of the CU are calculated by adjusting the bi-predicted samples as follows:

[0087] pred BDOF (x, y) = (I (0) (x, y) + I (1) (x, y) + b(x, y) + o pffset ) >> shift

[0088] In some embodiments, the values of na, nb, and ns2 are equal to 3, 6, and 12 respectively. In some embodiments, these values are selected such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process is kept within 32 bits. To derive the gradient values, some predicted samples I (k) (i, j) in the list k (k = 0, 1) outside the current CU boundary need to be generated.

[0089] In some embodiments, BDOF uses an extended column / row around the CU boundary. Figure 7 Shows the extended CU region used by BDOF for encoding / decoding a CU. To control the computational complexity of generating out-of-boundary predicted samples, a linear filter is required to generate predicted samples in the extended region (the white positions of the CU), and a normal 8-tap motion compensation interpolation filter is used to generate predicted samples within the CU (the shaded positions of the CU). These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any samples and gradient values outside the CU boundary are needed, they are filled (e.g., copied) from their nearest neighboring blocks.

[0090] VII. Combined Inter-Frame and Intra-Frame Prediction (CIIP)

[0091] In some embodiments, when the enabling conditions of CIIP are satisfied, the CU-level syntax of CIIP is signaled. For example, additional flags are signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. The enabling conditions may include that the CU is coded in the merge mode, and the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64). To form the CIIP prediction, an intra prediction mode is required. One or more possible intra prediction modes may be used: for example, DC, planar, horizontal, or vertical. Then, the inter prediction and the intra prediction signals are derived using the conventional intra and inter decoding processes. Finally, a weighted average of the inter and intra prediction signals is performed to obtain the CIIP prediction.

[0092] In some embodiments, if only one intra prediction mode (e.g., planar) is available for CIIP, the intra prediction mode for CIIP may be implicitly assigned to that mode (e.g., planar). In some embodiments, up to four intra prediction modes (including DC, PLANAR, horizontal, and vertical modes) may be used to predict the luma component in the CIIP mode. For example, if the CU shape is very wide (i.e., the width is greater than twice the height), then the horizontal mode is not allowed; if the CU shape is very narrow (i.e., the height is greater than twice the width), then the vertical mode is not allowed. In these cases, only 3 intra prediction modes are allowed. The CIIP mode may use three most probable modes (MPM) for intra prediction. If the CU shape is very wide or very narrow as defined above, the MPM flag is inferred as 1 in the case of no signaling. Otherwise, the MPM flag is signaled to indicate whether the CIIP intra prediction mode is one of the multiple CIIP MPM candidate modes. If the MPM flag is 1, the MPM index is further signaled to indicate which MPM candidate mode is used in the CIIP intra prediction. Otherwise, if the MPM flag is 0, the intra prediction mode is set to the "lost" mode in the MPM candidate list. For example, if the PLANAR mode is not in the MPM candidate list, then PLANAR is the lost mode, and the intra prediction mode is set to PLANAR. Since 4 possible intra prediction modes are allowed in CIIP, and the MPM candidate list contains only 3 intra prediction modes, one of the 4 possible modes must be the lost mode. The intra prediction mode of the coded CU of CIIP will be retained and used for the intra mode coding of future adjacent CUs.

[0093] The same inter - frame prediction process applied to the regular merge mode is used to derive the inter - frame prediction signal (or inter - frame prediction) in the CIIP mode Pinter, and the CIIP intra - frame prediction mode following the regular intra - frame prediction process is used to derive the intra - frame prediction or intra - frame prediction signal Pintra. The intra - frame and inter - frame prediction signals are then combined using weighted averaging, where the weighting values depend on adjacent blocks, on the intra - frame prediction mode, or on where the samples are located within the coded block. In some embodiments, if the intra - frame prediction mode is DC or planar mode, or if the block width or height is less than 4, then equal weights are applied to the intra - frame prediction and the inter - frame prediction signals. Otherwise, the weights are determined based on the intra - frame prediction mode (horizontal mode or vertical mode in this case) and the sample positions within the block. Starting from the nearest part of the intra - frame prediction reference samples and ending at the farthest part of the intra - frame prediction reference samples, the weights wt of each 4 - region are set to 6, 5, 3, and 2 respectively. In some embodiments, the CIIP prediction or the CIIP prediction signal PCIIP is derived as follows:

[0094] P CIIP = ((N1 - wt)*P inter + wt*P intra + N2) >> N3

[0095] where (N1, N2, N3) = (8, 4, 3) or (N1, N2, N3) = (4, 2, 2). When (N1, N2, N3) = (4, 2, 2), wt is selected from 1, 2, or 3.

[0096] VIII. Diffusion Filter (DIF)

[0097] The diffusion filter for video coding and decoding is to apply a diffusion filter to the prediction signal in video coding and decoding. Assume that pred is the prediction signal on a given block obtained by intra - frame or motion - compensated prediction. To handle the boundary points of the filter, the prediction signal is extended to the prediction signal predext. The extended prediction is formed by adding a line of reconstructed samples on the left and above of the block to the prediction signal and then the generated signal is mirrored in all directions.

[0098] The uniform diffusion filter is implemented by convolving the prediction signal with a fixed mask hI. In some embodiments, the prediction signal pred is replaced by h I *pred, using the boundary extension mentioned later. This time, the filter mask hI is defined as:

[0099]

[0100] Oriented diffusion filters such as the horizontal filter hhor and the vertical filter hver are used, which have fixed masks. The filtering is restricted to being applied only along the vertical or along the horizontal direction. The vertical filter is implemented by applying the fixed filter mask hver to the predicted signal and the horizontal filter is implemented by using the transposed mask as shown below.

[0101] The prediction signal is extended in the same manner as the uniform diffusion filter.

[0102] IX. Bilateral Filtering (BIF)

[0103] Quantizing in the transform domain is a well-known technique to better preserve information in images and videos compared to quantizing in the pixel domain. However, it is also well-known that quantized transform blocks can generate ringing artifacts around the edges of still images and moving objects in videos. Applying a bilateral filter can significantly reduce the ringing artifacts. In some embodiments, after the inverse transform has been performed and combined with the predicted sample values, a small, low-complexity bilateral filter is directly applied to the reconstructed samples.

[0104] In some embodiments, when applying bilateral filtering, each sample in the reconstructed image is replaced by a weighted average of itself and its neighboring samples. The weights are calculated based on the distance from the central sample and the difference in sample values. Figure 8 An exemplary 8x8 transform unit block and a bilateral filter aperture are shown. The filter aperture is used for the sample located at (1,1). As shown, since the filter is in the shape of a small plus sign as Figure 1 shown, all distances are 0 or 1. The sample located at (i,j) is filtered using its neighboring samples (k,l). The weight w(i,j,k,l) is the weight assigned to the sample (k,l) to filter the sample (i,j), and it is defined as follows:

[0105]

[0106] I(i,j) and I(k,l) are the original reconstructed intensity values of the samples (i,j) and (k,l) respectively. σ d is the spatial parameter, and σ r is the range parameter. The properties (or strength) of the bilateral filter are controlled by these two parameters. Samples that are closer to the sample being filtered and samples that have a smaller intensity difference with the sample being filtered will be filtered compared to samples that are farther away and samples that have a larger intensity difference. In some embodiments, σ d is set based on the transform unit size, and σ r is set based on the QP used for the current block.

[0107]

[0108] In some embodiments, after the inverse transform in both the encoder and the decoder, the bi-directional filter is directly applied to each TU block. As a result, subsequent intra-coded blocks are predicted from sample values that have been filtered by the bi-directional filter. This also makes it possible to include the bi-directional filtering operation in the rate-distortion decision in the encoder.

[0109] In some embodiments, each sample in the transform unit is filtered using only its directly adjacent samples. The filter has a plus-shaped filtering aperture centered on the sample to be filtered. The output filtered sample value ID(i,j) is calculated as follows:

[0110]

[0111] For TU sizes larger than 16x16, the block is treated as multiple 16x16 blocks with TU block width = TU block height = 16. Additionally, rectangular blocks are given as several examples of square blocks. In some embodiments, to reduce the amount of computation, the bi-directional filter is implemented using a lookup table (LUT) that stores all weights for a specific QP in a two-dimensional array. The LUT uses the intensity difference between the sample to be filtered and the reference samples as the index for the LUT in one dimension, and the TU size as the index in the other dimension. For efficient storage of the LUT, in some embodiments, the weights are rounded to 8-bit precision.

[0112] X. Hadamard transform domain filter (HAD)

[0113] In some embodiments, the Hadamard transform domain filter (HAD) is applied to the luminance reconstruction blocks with non-0 transform coefficients, and 4x4 blocks are excluded if the quantization parameter is greater than 17. The filter parameters are explicitly derived from the coding information. If the HAD filter is applied, the HAD is performed on the decoded samples after block reconstruction. The filtering result is used for both output and spatial and temporal prediction. The filter has the same implementation for both intra and inter CU filtering. According to the HAD filter, for each pixel from the pixels of the reconstructed block, the process includes the following steps: (1) scanning 4 adjacent pixels around the pixel to be processed including the current pixel according to a scanning pattern, (2) reading the 4-point Hadamard transform of the pixel, and (3) spectral filtering based on the following formula:

[0114]

[0115] where (i) is the index of the spectral component in the Hadamard spectrum, R(i) is the spectral component of the reconstructed pixel corresponding to the index, m = 4 is the normalization constant equal to the number of spectral components, and σ is a list parameter derived from the codec quantization parameter QP using the following equation:

[0116] σ = 2.64 * 2 (0.1296*(QP-11))

[0117] The first spectral component corresponding to the DC value is bypassed without filtering. The inverse 4-point Hadamard transform of the filtered spectrum is used. After the filtering step, the filtered pixels are placed back into their original positions in the accumulation buffer. After pixel filtering is complete, the accumulated values are normalized by the number of processing groups used for each pixel filtering. Since a one-sample padding is used around the block, the number of processing groups is equal to 4 for each pixel in the block, and normalization is performed by a right shift by 2 bits.

[0118] Figure 9 The list process under the Hadamard transform domain filter is shown. As shown, the equal filter shape is 3x3 pixels. In some embodiments, all pixels in the block can be processed independently for maximum parallelism. The results of 2x2 grouped filtering can be reused for spatially co-located samples. In some embodiments, a 2x2 filter is performed for each new pixel of the block, and the remaining three are used.

[0119] XI. Triangle Prediction Unit Mode (TPM)

[0120] In some embodiments, the Triangle Prediction Unit Mode (TPM) is used to perform inter-frame prediction of a CU. Under TPM, the CU is split into two triangle prediction units either diagonally or anti-diagonally. Each triangle prediction unit in the CU uses its own unidirectional motion vector and a reference frame for inter-frame prediction. In other words, the CU is split along the line dividing the current block. The transform and quantization processes are then applied to the entire CU. In some embodiments, this mode is only applied to the skip and merge modes. In some embodiments, TPM can be extended to split the CU into two prediction units with a line, which can be represented by an angle and a distance. The split line can be indicated by the signaled index and the signaled index is then mapped to an angle and a distance. Additionally, one or more indices are signaled to indicate the motion candidates for the two splits. After predicting each prediction unit, an adaptive weighting process is applied to the edges of the line dividing the current block between the two prediction units to derive the final prediction for the entire CU.

[0121] Figure 10Shows the adaptive weighting applied along the diagonal edge between two triangular prediction units of a CU. The first set of weighting factors {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8, 4 / 8, 1 / 8} are used for luminance and chrominance samples respectively. The second set of weighting factors {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} are used for luminance and chrominance samples respectively. A set of weighting factors is selected based on the comparison of the motion vectors of the two triangular prediction units. When the reference images of the two triangular prediction units are different from each other or the difference in their motion vectors is greater than 16 pixels, the second set of weighting factors is used. Otherwise, the first set of weighting factors is used.

[0122] XII. Mutually Exclusive Groups

[0123] In some embodiments, to simplify the hardware implementation complexity, mutually exclusive rules are implemented to limit the cascading of different tools or coding / decoding modes described in Parts I to XI. The cascaded hardware implementation of tools or coding / decoding modes makes the hardware design more complex and results in longer pipeline latency. By implementing the mutually exclusive rules, the pipeline stages can be made shorter and the hardware utilization can be made higher (i.e., less idle hardware). Generally, the mutually exclusive rules are used to ensure that tools or coding / decoding modes in certain sets of two or more tools or coding / decoding modes are not enabled simultaneously for coding / decoding a current CU.

[0124] In some embodiments, mutually exclusive groups of multiple (e.g., four) tools or coding / decoding modes are implemented. The mutually exclusive groups can include some or all of the following coding / decoding modes or tools: GBI (Generalized Bi-Directional Prediction), CIIP (Combined Inter and Intra Prediction), (BDOF) Bi-Directional Optical Flow, DMVR (Decoder-Side Motion Vector Refinement), and Weighted Prediction (WP).

[0125] In some embodiments, for any CU, the prediction stage of a video codec (video encoder or video decoder) implements mutually exclusive rules among some prediction tools or coding / decoding modes. Mutually exclusive means that only one of these coding / decoding tools is individually activated for coding / decoding, rather than two coding / decoding tools being activated for the same CU. In particular, it can define mutually exclusive groups of tools, where for any CU, only one tool in the group is activated for coding / decoding and no two tools belonging to the same mutually exclusive group are activated for the same CU. In some embodiments, different CUs can have different activated tools.

[0126] For some embodiments, the mutually exclusive group includes GBI, BDOF, DMVR, CIIP, WP. That is, for any CU, among GBI, BDOF, DMVR, CIIP, and WP, only one tool in the group is activated for encoding / decoding, rather than two of them being activated for the same CU. For example, when the CIIP flag is equal to 1, DMVR / BDOF / GBI is not applied. In other words, when the CIIP flag is equal to 0, DMVR / BDOF / GBI can be applied (if the enabling conditions for DMVR / BDOF / GBI are met). In some embodiments, the mutually exclusive group includes any two or three or some subsets of GBI, BDOF, DMVR, CIIP, and WP. In particular, the mutually exclusive group can include BDOF, DMVR, CIIP; the mutually exclusive group can include GBI, DMVR, CIIP; the mutually exclusive group can include GBI, BDOF, CIIP; the mutually exclusive group can include GBI, BDOF, DMVR; the mutually exclusive group can include GBI and BDOF; the mutually exclusive group can include GBI and DMVR; the mutually exclusive group can include GBI and CIIP; the mutually exclusive group can include BDOF, DMVR, CIIP; the mutually exclusive group can include BDOF and CIIP; the mutually exclusive group can include DMVR and CIIP. For example, if the mutually exclusive group includes GBI, BDOF, CIIP, when CIIP is enabled (ciip_flag equals 1), BDOF is off and GBI is off (which means equal weights are used for mixing inter-predicted frames from list 0 and list 1 regardless of the BCW weight index). For example, for the mutually exclusive group including GBI, DMVR, CIIP, when CIIP is enabled (ciip_flag equals 1), DMVR is off and GBI is off (which means equal weights are used for mixing inter-predicted frames from list 0 and list 1 regardless of the BCW weight index). For example, for the mutually exclusive group including GBI, DMVR, if GBI is enabled (the GBI weight index indicates unequal weights), DMVR is off. For example, for the mutually exclusive group including GBI, BDOF, if GBI is enabled (the GBI weight index indicates unequal weights), BDOF is off.

[0127] According to the mutually exclusive rules, relevant syntax elements can be saved (or omitted from the bitstream). For example, if the mutually exclusive group includes GBI, BDOF, DMVR, CIIP, then, if the CIIP mode is not enabled for the current CU (e.g., it is GBI or BDOF or DMVR), since CIIP is turned off by the exclusion rule, the CIIP flag or syntax elements can be saved or ignored (instead of being sent from the encoder to the decoder) for this CU. For some other embodiments of the mutually exclusive group, for a tool excluded or disabled for a certain CU, the relevant syntax elements can be saved or ignored.

[0128] In some embodiments, a priority order is applied to the mutually exclusive group. In some embodiments, each tool within the mutually exclusive group has a certain original or traditional enabling condition. The enabling condition is the original enabling rule for each tool before mutual exclusion. For example, the enabling conditions of DMVR include true bidirectional prediction and equal POC distance between the current picture and the L0 picture / L1 picture and others; the enabling conditions of GBI include bidirectional prediction and the GBI index from the syntax (when AMVP) or the inherited GBI index (when in merge mode).

[0129] A priority rule can be predefined for the mutually exclusive group. Tools or coding / decoding modes within the mutually exclusive group have a priority number for each tool. If both tool A and tool B can be activated (i.e., their enabling conditions are satisfied) for the same CU, but tool A has a better predefined priority than tool B (defined as tool A > tool B), then, if tool A is activated or enabled, tool B is turned off or disabled.

[0130] Different embodiments have different precedence rules for mutually exclusive groups such as those including GBI, DMVR, BDOF, CIIP, and WP or any subset of {GBI, DMVR, BDOF, CIIP, WP}. For example, in some embodiments, the precedence rule specifies GBI > DMVR > BDOF > CIIP. In some embodiments, the precedence rule specifies GBI > DMVR > BDOF. In some embodiments, the precedence rule specifies DMVR > GBI > BDOF. In some embodiments, the precedence rule specifies DMVR > GBI. In some embodiments, the precedence rule specifies GBI > BDOF. In some embodiments, the precedence rule specifies GBI > DMVR. In some embodiments, the precedence rule specifies DMVR > GBI > BDOF > CIIP. In some embodiments, the precedence rule specifies DMVR > BDOF > GBI > CIIP. In some embodiments, the precedence rule specifies CIIP > GBI > BDOF. In some embodiments, the precedence rule specifies CIIP > GBI > DMVR. The predefined rules also specify any other order in any subset of GBI, BDOF, DMVR, CIIP. For another example, the mutually exclusive group includes {GBI, CIIP} and the precedence rule specifies CIIP > GBI, so when CIIP is used (ciip_flag equals 1), GBI is turned off (or disabled), which means equal weights are applied to mix predictors from list 0 and list 1. For another example, the mutually exclusive group includes {DMVR, CIIP} and the precedence rule specifies CIIP > DMVR, so when CIIP is used (ciip_flag equals 1), DMVR is not used. For another example, the mutually exclusive group includes {BDOF, CIIP} and the precedence rule specifies CIIP > BDOF, so when CIIP is used (ciip_flag equals 1), BDOF is not used. For another example, the mutually exclusive group includes {BDOF, GBI} and the precedence rule specifies GBI > BDOF, so when GBI is used (the GBI index indicates unequal weights are used to mix predictors from list 0 and list 1), BDOF is not used. For another example, the mutually exclusive group includes {DMVR, GBI} and the precedence rule specifies GBI > DMVR, so when GBI is used (the GBI index indicates unequal weights are used to mix predictors from list 0 and list 1), DMVR is not used.

[0131] In some embodiments, the priority rules for mutually exclusive groups are not predefined but are also based on some parameters of the current CU (such as CU size or the current MV). For example, for a mutually exclusive group including DMVR and BDOF, there can be an exclusion rule based on the CU size or other parameters of the CU. When the enabling conditions for both DMVR and BDOF are satisfied, a given priority is assigned to DMVR or BDOF. For example, in some embodiments, if the current CU size is greater than a threshold, the priority of DMVR (for tool exclusion) is higher than that of BDOF. In some embodiments, if the current CU size is greater than a threshold, the priority of BDOF (for tool exclusion) is higher than that of DMVR.

[0132] In some embodiments, if the current CU aspect ratio is greater than a threshold, the priority of DMVR (for tool exclusion) is greater than that of BDOF. If CU_width > CU_height, the aspect ratio is defined as CU_width / CU_height, or if CU_height >= CU_width, the aspect ratio is defined as CU_height / CU_width. In some embodiments, if the current CU aspect ratio is greater than a threshold, the priority of BDOF (for tool exclusion) is greater than the priority of DMVR. In some embodiments, for some merge mode candidates (if selected for inter prediction), the priority of DMVR (for tool exclusion) is higher than the priority of BDOF, while for other merge candidates (if selected for inter prediction), the priority of BDOF (for tool exclusion) is higher than the priority of DMVR.

[0133] In some embodiments, for a true bi - directional prediction merge candidate, if the mirror (and subsequent scaling) of the L0 MV is very similar to the L1 MV, then the priority of DMVR (for tool exclusion) is higher than the priority of BDOF. In some embodiments, for a true bi - directional prediction merge candidate, if the mirror (and subsequent scaling) of the L0 MV is very similar to the L1 MV, then the priority of BDOF (for tool exclusion) is higher than the priority of DMVR.

[0134] In some embodiments, the mutually exclusive group may include some or all of the following coding and decoding modes or tools: LIC (Local Luminance Compensation), DIF (Uniform Luminance Inter - frame Prediction Filter or Diffusion Filter), BIF (Bidirectional Filter), HAD filter (Hadamard Transform Domain Filter). These tools or coding and decoding modes are applied to the residual signal or the prediction signal or the reconstructed signal, that is, they act on the "post - stage". The post - stage is defined as the pipeline stage after prediction (intra - frame / inter - frame prediction) or after reference decoding or after both. In some embodiments, the mutually exclusive group may also include pre - stage tools or coding and decoding modes different from LIC, DIF, BIF, and HAD.

[0135] In some embodiments, the mutually exclusive group may include all or a subset of the following 8 coding and decoding modes or tools: GBI, BDOF, DMVR, CIIP, LIC, DIF, BIF, HAD. That is, for any CU, only one of them is activated for encoding, rather than two coding and decoding modes or tools among GBI, BDOF, DMVR, CIIP, LIC, DIF, BIF, HAD being activated for the same CU. In some embodiments, the mutually exclusive modes include LIC, DIF, BIF, HAD. In some embodiments, the mutually exclusive group includes DIF, BIF, HAD. In some embodiments, the mutually exclusive group includes LIC, BIF, HAD. In some embodiments, the mutually exclusive group includes LIC, DIF, HAD. In some embodiments, the mutually exclusive group includes LIC, DIF, BIF. In some embodiments, the mutually exclusive group includes LIC and DIF. In some embodiments, the mutually exclusive group includes LIC, BIF. In some embodiments, the mutually exclusive group includes LIC and HAD. In some embodiments, the mutually exclusive group includes DIF and HAD. In some embodiments, the mutually exclusive group includes DIF and HAD. In some embodiments, the mutually exclusive group includes BIF and HAD.

[0136] XIII. Signaling of Multi - Hypothesis Prediction Modes

[0137] Both CIIP and TPM generate the final prediction of the current CU with two candidates. Either CIIP or TPM can be regarded as a type of multi - hypothesis prediction merge mode, where one hypothesis of the prediction is generated by one candidate and the other hypothesis of the prediction is generated by another candidate. For CIIP, one candidate comes from the intra - frame mode and the other candidate comes from the merge mode. For TPM, the two candidates come from the candidate list of the merge mode.

[0138] In some embodiments, the multi-hypothesis mode is used to improve inter-frame prediction, which is an improvement method for the skip and / or merge modes. In the original skip and merge modes, one merge index is used to select a motion candidate, which can be a unidirectional prediction or a bidirectional prediction derived from the candidate itself from the merge candidate list. The generated motion compensation predictor is referred to as the first hypothesis (or first prediction) in some embodiments. In the multi-hypothesis mode, in addition to the first hypothesis, a second hypothesis is also generated. The second hypothesis of the predictor can be generated by motion compensation from motion candidates based on the inter-frame prediction mode (merge or skip mode), or by intra-frame prediction based on the intra-frame prediction mode.

[0139] When the second hypothesis (or second prediction) is generated by the intra-frame prediction mode, the multi-hypothesis mode is referred to as the intra-frame MH mode or MH mode intra-frame or MH intra-frame or inter-intra mode. The CU encoded and decoded by CIIP is encoded and decoded using the intra-frame MH mode. When the second hypothesis is generated by motion compensation from motion candidates or the inter-frame prediction mode (e.g., merge or skip mode), the multi-hypothesis mode is referred to as the inter-frame MH mode or MH mode inter-frame or MH inter-frame (or also referred to as the merged MH mode or MH merge). The diagonal edge region of the CU encoded and decoded by TPM is encoded and decoded using the inter-frame MH mode.

[0140] For the multi-hypothesis mode, each multi-hypothesis candidate (or each candidate with multi-hypotheses) includes one or more candidates (i.e., the first hypothesis) and / or an intra-frame prediction mode (i.e., the second hypothesis), where the motion candidate is selected from candidate list I and / or the intra-frame prediction mode is selected from candidate list II. For the intra-frame MH mode, each multi-hypothesis candidate (or each candidate with multi-hypotheses) includes a motion candidate and an intra-frame prediction mode, where the motion candidate is selected from candidate list I and the intra-frame prediction mode is fixed to one mode (e.g., planar) or selected from candidate list II. The inter-frame MH mode uses two motion candidates, and at least one of the two motion candidates is derived from candidate list I. In some embodiments, candidate list I is equal to the merge candidate list of the current block and both of the two motion candidates of the multi-hypothesis candidate of the inter-frame MH mode are selected from candidate list I. In some embodiments, candidate list 1 is a subset of the merge candidate list. In some embodiments, for the inter-frame MH mode, each of the two motions used to generate the prediction of each prediction unit is indicated by the transmitted index. When the index refers to a bidirectional prediction motion candidate in candidate list 1, the motion in list 0 or list 1 is selected according to the index. When the index refers to a unidirectional prediction motion candidate in candidate list I, the unidirectional prediction motion is used.

[0141] Figure 11aConceptually shows encoding or decoding a pixel block by using an intra MH mode. The figure shows a video image 1100 currently being encoded or decoded by a video encoder. The video image 1100 includes a pixel block 1110 currently being encoded or decoded as a current block. The current block 1110 is encoded by an intra MH mode. In particular, a combined prediction 1120 is generated based on a first prediction 1122 (first hypothesis) of the current block 1110 and a second prediction 1124 (second hypothesis) of the current block 1110. The combined prediction 1120 is then used to reconstruct the current block 1110.

[0142] The current block 1110 is encoded and decoded by using an intra MH mode. In particular, a first prediction is obtained by inter prediction based on at least one reference frame 1102 and 1104. A second prediction 1124 is obtained by intra prediction based on neighboring pixels 1106 of the current block 1110. As shown, the first prediction 1122 is generated based on an inter prediction mode or a motion candidate 1142 selected from a first candidate list 1132 (candidate list I), and the first candidate list 1132 has one or more candidate inter prediction modes. The candidate list I may be a merge candidate list of the current block 1110. The second prediction 1124 is generated based on an intra prediction mode 1144, which is predefined as an intra prediction mode (e.g., planar) or selected from a second candidate list 1134 (candidate list II) having one or more candidate intra prediction modes. If only one intra prediction mode (e.g., planar) is used for intra MH, the intra prediction mode for intra MH is set to an intra prediction mode that does not need to be signaled.

[0143] Figure 11b Shows a current block 1110 being encoded and decoded by using an inter MH mode. In particular, a first prediction 1122 is obtained by inter prediction based on at least one reference frame 1102 and 1104. A second prediction 1124 is obtained by inter prediction based on at least one reference frame 1106 and 1108. As shown, the first prediction 1122 is generated based on an inter prediction mode or a motion candidate 1142 (first prediction mode), and the motion candidate 1142 is selected from a first candidate list 1132 (candidate list I). The second prediction 1124 is generated based on an inter prediction mode or a motion candidate 1146, and the motion candidate 1146 is also selected from the first candidate list 1132 (candidate list I). The candidate list 1 may be a merge candidate list of the current block.

[0144] In some embodiments, when the in - frame MH mode is currently supported, in addition to the original syntax of the merge mode, a flag is signaled (e.g., to indicate whether the in - frame MH mode is applied). This flag can be represented or indicated by a syntax element in the bitstream. In some embodiments, if the flag is present, an additional in - frame mode index is signaled to indicate the in - frame prediction mode from candidate list II. In some embodiments, if the flag is on, the in - frame prediction mode of the in - frame MH mode (e.g., CIIP, or any of the in - frame MH modes) is implicitly selected from candidate list II or an in - frame prediction mode is implicitly assigned without an additional in - frame mode index. In some embodiments, when the flag is off, an inter - frame MH mode (e.g., TPM, or any other inter - frame MH mode with a different prediction unit shape) can be used.

[0145] In some embodiments, the video encoder (video encoder or video decoder) removes all bi - directional prediction cases in CIIP. That is, the video encoder activates CIIP only when the current merge candidate is a uni - directional prediction. In some embodiments, the video encoder removes all bi - directional prediction candidates for CIIP merge candidates. In some embodiments, the video encoder retrieves the L0 information of a bi - directional prediction (merge candidate) and changes it into a uni - directional prediction candidate for CIIP. In some embodiments, the video encoder retrieves the L1 information of a bi - directional prediction (merge candidate) and changes it into a uni - directional prediction candidate for CIIP. By removing all bi - directional prediction behaviors of CIIP, related syntax elements can be saved or omitted from the transmission.

[0146] In some embodiments, when generating an inter prediction for the CIIP mode, according to a predetermined rule, a motion candidate with bi - directional prediction is changed to uni - directional prediction. In some embodiments, based on the POC distance, the predetermined rule designates or selects a list 0 or list 1 motion vector. When the distance (denoted as D1) between the current POC (or the POC of the current picture) and the POC of the reference picture referenced by the list x (where x is 0 or 1) motion vector is less than the distance (denoted as D2) between the current POC and the POC of the reference picture referenced by the list y (where y is 0 or 1 and y is not equal to x) motion vector, the list x motion vector is selected to generate the inter prediction for CIIP. If D1 is the same as D2 or the difference between D1 and D2 is less than a threshold, the list x (where x is predetermined to be 0 or 1) motion vector is selected to generate the inter prediction for CIIP. In some other embodiments, the predetermined rule generally selects the list x motion vector, where x is predetermined to be 0 or 1. In some other embodiments, this bi - directional to uni - directional prediction scheme can be applied to motion compensation to generate the prediction. When the motion information of the currently encoded CIIP CU is saved for reference by subsequent or next CUs, the motion information before applying this bi - directional to uni - directional prediction scheme is used. In some embodiments, after generating the merge candidate list for CIIP, this bi - directional to uni - directional prediction scheme is applied. Processes such as motion compensation and / or motion information saving and / or de - blocking can use the generated uni - directional prediction motion information.

[0147] In some embodiments, a new candidate list formed by uni - directional prediction motion candidates is constructed for CIIP. In some embodiments, according to a predetermined rule, this candidate list can be generated from the merge candidate list for the regular merge mode. For example, when generating the candidate list as done in the regular merge mode, the predetermined rule can specify that bi - directional prediction motion candidates can be ignored. The length of this new candidate list for CIIP can be equal to or less than that of the regular merge mode. For another example, the predetermined rule can specify that the candidate list for CIIP re - uses the candidate list of TPM or the candidate list for CIIP can be re - used for TPM. The methods proposed above can be combined with implicit rules or explicit rules. The implicit rule can depend on the block width or height or area and the explicit rule can signal a flag at the CU, CTU, slice, tile, tile group, SPS, PPS level, etc.

[0148] In some embodiments, CIIP and TPM are classified into a group of combined prediction modes and the syntax of CIIP and TPM is also unified instead of using two respective flags to determine whether to use CIIP and whether to use TPM. The unified scheme is as follows: When the enabling condition for the group of combined prediction modes is satisfied (e.g., the unified set of CIIP and TPM enabling conditions, including high-level syntax, size constraints, supported modes, or stripe type), CIIP or TPM can be enabled or disabled with the unified syntax. First, a first bin is signaled (or a first flag is signaled using the first bin) to indicate whether to apply the multi-hypothesis prediction mode. Second, if the first bin indicates to apply the multi-hypothesis prediction mode, a second bin is signaled (or a second flag is signaled using the second bin) to indicate that one of CIIP and TPM is applied. For example, when the first bin (or the first flag) is equal to 0, a non-multi-hypothesis prediction mode such as the regular merge mode is applied, otherwise, a multi-hypothesis prediction mode such as CIIP or TPM is applied. When the first bin (or the first flag) indicates that the multi-hypothesis prediction mode is applied (regular_merge_flag is equal to 0), the second flag is signaled. When the second bin (or the second flag) is equal to 0, TPM is applied and additional syntax for TPM is required (e.g., the additional syntax for TPM is to indicate two motion candidates for TPM or the split direction of TPM). When the second bin (or the second flag) is equal to 1, CIIP is applied and additional syntax for CIIP may be required (e.g., the additional syntax for CIIP to indicate two candidates for CIIP). Examples of the enabling conditions for the group of combined prediction modes include (1) high-level syntax CIIP and (2) TPM being enabled.

[0149] Signaling of XIV.LIC

[0150] In some embodiments, all bidirectional predictions are removed for the LIC mode. In some embodiments, LIC is allowed only when the current merge candidate is a unidirectional prediction. In some embodiments, the video encoder retrieves a bidirectional prediction L0 information (candidate), changes the current merge candidate to a unidirectional prediction candidate, and then applies LIC. In some embodiments, the video encoder retrieves a bidirectional prediction L1 information (candidate), changes it to a unidirectional prediction candidate, and then applies LIC.

[0151] In some embodiments, when generating inter prediction for the LIC mode, according to a predetermined rule, a motion candidate with bidirectional prediction is transformed into a unidirectional prediction. In some embodiments, the predetermined rule specifies or selects a list 0 or list 1 motion vector based on the POC distance. When the distance (denoted as D1) between the current POC and the POC of the (reference image) referenced by the list x (where x is 0 or 1) motion vector is less than the distance (denoted as D2) between the current POC and the POC referenced by the list y (where y is 0 or 1 and y is not equal to x) motion vector, then the list x motion vector is selected for refining the inter prediction by applying LIC. If D1 is the same as D2 or the difference between D1 and D2 is less than a threshold, then the list x (where x is predetermined to be 0 or 1) motion vector is selected for refining the inter prediction by using LIC. In some embodiments, the predetermined rule formulates or selects the list x motion vector, where x is predetermined to be 0 or 1. In some embodiments, this bidirectional-to-unidirectional prediction scheme can be applied only to motion compensation to generate the prediction. When the motion information of the current encoded LIC CU is saved for reference to subsequent or later CUs, the motion information between applying this bidirectional-to-unidirectional prediction scheme is used. In some embodiments, after generating the merge candidate list of LIC, this bidirectional-to-unidirectional prediction scheme is applied. Such as the process of motion compensation and / or the process of motion information saving uses the generated unidirectional prediction motion information.

[0152] In some embodiments, a new candidate list formed by unidirectional prediction motion candidates is constructed for LIC. In some embodiments, according to a predetermined rule, the candidate list can be generated from the merge candidate list for the regular merge mode. For example, the predetermined rule can ignore bidirectional prediction motion candidates during the generation of the candidate list like the regular merge mode. The length of this new candidate list for LIC can be equal to or less than the regular merge mode.

[0153] In some embodiments, in the merge mode, the criteria for enabling LIC depend not only on the LIC flag of the merge candidate but also on the number of neighboring merge candidates using LIC or historical statistics. For example, if the number of candidates in the merge list using LIC is greater than a predetermined threshold, then LIC is enabled for the current block regardless of whether the LIC flag of the merge candidate is on or off. As another example, the historical FIFO buffer records the LIC mode usage of the most recently encoded blocks. Assuming the recorded size in the FIFO buffer is M, if N out of M used the LIC mode, then LIC is enabled for the current block. In addition, this embodiment can also be combined with the aforementioned bi - directional to uni - directional prediction scheme for LIC, i.e., if the LIC flag of the current block is enabled due to the number of neighboring merge candidates using LIC being greater than the threshold or N out of M records in the historical FIFO buffer using the LIC mode, and the merge candidate uses bi - directional prediction, then the list x motion vector is selected, where x is predetermined to be 0 or 1.

[0154] All of the above combinations can be determined by implicit rules or explicit rules. The implicit rules can depend on block width, height, area, block size aspect ratio, color component, or image type. The explicit rule can be signaled by a flag at the CU, CTU, slice, tile, tile group, image, SPS, PPS level, etc.

[0155] XV. Exemplary Video Encoder

[0156] Figure 12 An exemplary video encoder 1200 that indicates mutually exclusive groups of codec modes or tools that can be implemented is shown. As shown, the video encoder 1200 receives an input video signal from a video source 1205 and encodes the signal into a bitstream 1295. The video encoder 1200 has various components or modules for encoding the signal from the video source 1205, including at least some components selected from a transform module 1210, a quantization module 1211, an inverse quantization module 1214, an inverse transform module 1215, an intra - picture estimation module 1220, an intra - prediction module 1225, a motion compensation module 1230, a motion estimation module 1235, a loop filter 1245, a reconstructed picture buffer 1250, an MV buffer 1265, an MV prediction module 1275, and an entropy encoder 1290. The motion compensation module 1230 and the motion estimation module 1235 are part of the inter - prediction module 1240.

[0157] In some embodiments, modules 1210 - 1290 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device or an electronic device. In some embodiments, modules 1210 - 1290 are modules of hardware circuits implemented by one or more integrated circuits of an electronic device. Although modules 1210 - 1290 are shown as separate modules, some modules may be combined into a single module.

[0158] Video source 1205 provides an original video signal representing pixel data of each uncompressed video frame. Subtractor 1208 calculates the difference between the original video pixel data of video source 1205 and the predicted pixel data 1213 from motion compensation module 1230 or intra prediction module 1225. Transform module 1210 converts the difference (or residual pixel data or residual signal 1209) into transform coefficients 1216 (e.g., by performing a discrete cosine transform or DCT). Quantization module 1211 quantizes the transform coefficients into quantized data (or quantized coefficients) 1212, which are encoded into bitstream 1295 by entropy encoder 1290.

[0159] Inverse quantization module 1214 dequantizes the quantized data (or quantized coefficients) 1212 to obtain transform coefficients, and inverse transform module 1215 performs an inverse transform on the transform coefficients to generate a reconstructed residual 1219. The reconstructed residual 1219 is added to the predicted pixel data 1213 to generate reconstructed pixel data 1217. In some embodiments, the reconstructed pixel data 1217 is temporarily stored in a linear buffer (not shown) for intra-image prediction and spatial MV prediction. The reconstructed pixels are filtered by loop filter 1245 and stored in reconstructed image buffer 1250. In some embodiments, the reconstructed image buffer 1250 is an external storage area of video encoder 1200. In some embodiments, the reconstructed image buffer 1250 is an internal storage area of video encoder 1200.

[0160] Intra-image estimation module 1220 performs intra prediction based on the reconstructed pixel data 1217 to generate intra prediction data. The intra prediction data is provided to entropy encoder 1290 to be encoded into bitstream 1295. The intra prediction data is also used by intra prediction module 1225 to generate predicted pixel data 1213.

[0161] Motion estimation module 1235 performs inter prediction by generating MVs to reference the pixel data of previously decoded frames stored in reconstructed image buffer 1250. These MVs are provided to motion compensation module 1230 to generate predicted pixel data.

[0162] In addition to encoding the complete actual MVs in the bitstream, the video encoder 1200 uses MV prediction to generate predicted MVs, and the differences between the MVs for motion compensation and the predicted MVs are encoded as residual motion data and stored in the bitstream 1295.

[0163] The MV prediction module 1275 generates predicted MVs based on reference MVs, which are generated for encoding previous video frames, i.e., the motion compensation MVs used for performing motion compensation. The MV prediction module 1275 retrieves the reference MVs from the previous video frames from the MV buffer 1265. The video encoder 1200 stores the MVs generated for the current video frame in the MV buffer 1265 as reference MVs for generating predicted MVs.

[0164] The MV prediction module 1275 uses the reference MVs to create predicted MVs. The predicted MVs can be calculated by spatial MV or temporal MV prediction. The difference (residual motion data) between the predicted MV of the current frame and the motion compensation MV (MC MV) is encoded by the entropy encoder 1290 into the bitstream 1295.

[0165] The entropy encoder 1290 encodes various parameters and data into the bitstream 1295 by using entropy coding techniques, such as context-adaptive binary arithmetic coding (CABAC) or Huffman coding. The entropy encoder 1290 encodes various header elements, flags, and the quantized coefficients 1212, as well as the residual motion data, as syntax elements into the bitstream 1295. The bitstream 1295 is in turn stored in a storage device or transmitted to a decoder through a communication medium such as a network.

[0166] The loop filter 1245 performs filtering or smoothing operations on the reconstructed pixel data 1217 to reduce coding artifacts, particularly at the boundaries of pixel blocks. In some embodiments, the filtering operations performed include sample adaptive offset (SAO). In some embodiments, the filtering operations include adaptive loop filter (ALF).

[0167] Figure 13 Parts of the video encoder 1200 that mark mutually exclusive groups of encoding / decoding modes or tools are shown. As shown, the video encoder 1200 implements a combined prediction module 1310, which can receive the intra-prediction values generated by the intra-image prediction module 1225. The combined prediction module 1310 can also receive inter-prediction values from the motion compensation module 1230 and the second motion compensation module 1330. The combined prediction module 1310 in turn generates predicted pixel data 1213, which can be further filtered by a set of prediction filters 1350.

[0168] The MV buffer 1265 provides merge candidates to the motion compensation modules 1230 and 1330. The MV buffer 1265 also stores the motion information and motion directions for encoding the current block for use by subsequent blocks. The merge candidates may be changed, extended, and / or refined by the MV refinement module 1365.

[0169] The codec mode (or tool) control module 1300 controls the operations of the intra-image prediction module 1225, the motion compensation module 1230, the second motion compensation module 1330, the MV refinement module 1365, the combined prediction module 1310, and the prediction filter 1350.

[0170] The codec module control 1300 may enable the MV refinement mode 1365 to perform MV refinement operations by searching for refined MVs (e.g., for DMVR) or adjust the computed gradient based on the MVs (e.g., for BDOF). The codec mode control module 1300 may enable the intra-prediction module 1225 and the motion compensation module 1230 to implement the MH mode intra (or inter-intra) mode (e.g., CIIP). The codec mode control module 1300 may enable the motion compensation module 1230 and the second motion compensation module 1330 to implement the MH mode inter mode (e.g., for the diagonal edge region of TPM). When combining the prediction signals from the intra-image prediction module 1225, the motion compensation module 1230, and / or the second motion compensation module 1330 to implement codec modes such as CIIP, TPM, GBI, and / or WP, the codec mode control module 1300 may enable the combined prediction module 1310 to employ different weighting schemes. The codec mode control 1300 may also enable the prediction filter 1350 to apply LIC, DIF, BIF, and / or HAD filters on the predicted pixel data or the reconstructed pixel data 1217.

[0171] The codec mode control mode 1300 also determines which codec modes to enable and / or disable for encoding / decoding the current block. The codec mode control module 1300 then controls the operations of the intra-image prediction module 1225, the motion compensation module 1230, the second motion compensation module 1330, the MV refinement module 1365, the combined prediction module 1310, and the prediction filter 1350 to enable and / or disable the specific codec modes.

[0172] In some embodiments, the codec mode control 1300 enables only a subset (one or more) of multiple codec modes from a specific set of two or more codec modes for encoding the current block or CU. This specific set of codec modes includes all or any subset of the following codec modes: CIIP, TPM, BDOF, DMVR, GBI, WP, LIC, DIF, BIF, and HAD. In some embodiments, when a first condition for enabling a first codec mode of the current block is satisfied, the codec mode control 1300 disables a second codec mode of the current block.

[0173] In some embodiments, when the condition for enabling the first codec mode is satisfied and the first codec mode is enabled, in addition to the first codec mode, the codec mode control 1300 disables all modes in a specific set of codec modes. In some embodiments, when both a first condition for enabling the first codec mode of the current block and a second condition for enabling the second codec mode of the current block are satisfied and the first codec mode is enabled, the codec mode control 1300 disables the second codec mode. For example, in some embodiments, when the conditions for the codec mode control 1300 to decide to enable GBI and BDOF are both satisfied and the GBI index indicates unequal weights for mixing prediction lists 0 and 1, the codec mode control 1300 will disable BDOF. As another example, in some embodiments, when the conditions for the codec mode control 1300 to decide to enable GBI and DMVR are both satisfied and the GBI index indicates unequal weights for mixing prediction lists 0 and 1, the codec mode control 1300 will disable DMVR.

[0174] In some embodiments, the codec mode control 1300 identifies the highest-priority codec mode from one or more codec modes. If the highest-priority codec mode is enabled, the codec mode control 1300 then disables all other codec modes in the specific set of codec modes, regardless of whether the enabling conditions for each other codec mode are satisfied. In some embodiments, each codec mode in the specific set of codec modes is assigned a priority according to priority rules defined based on parameters of the current block (such as the size or aspect ratio of the current block).

[0175] The codec mode control 1300 generates or signals syntax elements 1390 to the entropy coder 1290 to indicate that one or more codec modes are enabled. The video encoder 1200 may also disable one or more other codec modes in a particular set of codec modes without signaling syntax elements for disabling the one or more other codec modes. In some embodiments, a first syntax element (e.g., a first flag) is used to indicate whether a multi-hypothesis prediction mode is applied and a second syntax element (e.g., a second flag) is used to indicate whether CIIP or TPM is applied. The first and second syntax elements are correspondingly encoded and decoded by the entropy coder 1290 into a first bin and a second bin. In some embodiments, the second bin for deciding between CIIP and TPM is signaled only if the first bin indicates that the multi-hypothesis mode is enabled.

[0176] Figure 14 Conceptually illustrates a process 1400 for implementing mutually exclusive groups of codec modes or tools. In some embodiments, one or more processing units (or processors) of a computing device implementing the encoder 1200 execute the process 1400 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing the encoder 1200 executes the process 1400.

[0177] The encoder receives (at block 1410) data of a pixel block of a current block of a current image to be encoded as video.

[0178] The encoder identifies (at block 1430) a highest-priority codec mode among one or more codec modes. In some embodiments, each codec mode in a particular set of codec modes is assigned a priority according to a priority rule defined based on the current block parameters.

[0179] If the highest-priority codec mode is enabled, the encoder disables (at block 1140) all other codec modes in the codec-mode specific set. The conditions for enabling the various codec modes are described in the above paragraphs related to these codec modes. The conditions for enabling a codec mode can include receiving an explicit syntax element from the bitstream for the codec mode. The conditions for enabling a codec mode can also include having specific characteristics or parameters of the current block being decoded (e.g., size, aspect ratio). For example, when the specific set of codec modes includes a first codec mode assigned a higher priority and a second codec mode assigned a lower priority, and when the first codec mode is enabled, the encoder disables (at block 1445) the second codec mode for the current block. In some embodiments, when the first codec mode is enabled, in addition to the first codec mode, the encoder disables all codec modes in the codec-mode specific set. In some embodiments, if the GBI weight index indicates unequal weights, the encoder enables GBI (which means using unequal weights to blend the inter-prediction from list 0 and list 1), and disables BDOF because GBI is assigned a higher priority than BDOF. As another example, in some embodiments, because GBI is assigned a higher priority than DMVR, if the GBI weight index indicates unequal weights, the encoder enables GBI (which means using unequal weights to blend the inter-prediction from list 0 and list 1), but disables DMVR. As another example, in some embodiments, because CIIP is assigned a higher priority than the disabled tools, if the CIIP flag is equal to 1, the encoder enables CIIP, but disables GBI, BDOF, and / or DMVR.

[0180] The encoder encodes (at block 1450) the current block in the bitstream by using an inter-prediction that is calculated according to the enabled codec mode.

[0181] XVI. Exemplary Video Decoder

[0182] Figure 15Exemplary video decoder 1500 that indicates mutually exclusive groups in which codec modes or tools can be implemented. As shown, video decoder 1500 is an image decoding or video decoding circuit that receives bitstream 1595 and decodes the content of the bitstream into pixel data of video frames for display. Video decoder 1500 has several components or modules for decoding bitstream 1595, including some components selected from inverse quantization module 1505, inverse transform module 1510, intra prediction module 1525, motion compensation module 1530, loop filter 1545, decoded image buffer 1550, MV buffer 1565, MV prediction module 1575, and parser 1590. Motion compensation module 1530 is part of inter prediction module 1540.

[0183] In some embodiments, modules 1510 - 1590 are modules of software instructions executed by one or more processing units (such as processors) of a computing device. In some embodiments, modules 1510 - 1590 are modules of hardware circuits implemented by one or more ICs of an electronic device. Although modules 1510 - 1590 are shown as separate modules, some modules can be combined into a single module.

[0184] According to the syntax defined by video coding or image coding standards, parser 1590 (or entropy decoder) receives bitstream 1595 and performs an initial parse. The parsed syntax elements include various header elements, flags, and quantized data (or quantized coefficients) 1512. Parser 1590 parses out various syntax elements by using entropy coding techniques such as context - adaptive arithmetic coding (CABAC) or Huffman coding.

[0185] Inverse quantization module 1505 de - quantizes the quantized data (or quantized coefficients) 1512 to obtain transform coefficients, and inverse transform module 1510 performs an inverse transform on transform coefficients 1516 to generate a reconstructed residual signal 1519. Reconstructed residual signal 1519 is added to predicted pixel data 1513 from intra prediction module 1525 or motion compensation module 1530 to generate decoded pixel data 1517. Decoded pixel data is filtered by loop filter 1545 and stored in decoded image buffer 1550. In some embodiments, decoded image buffer 1550 is external storage of video decoder 1550. In some embodiments, decoder image buffer 1550 is internal storage of video decoder 1550.

[0186] The intra prediction module 1525 receives intra prediction data from the bitstream 1595 and, based thereon, generates predicted pixel data 1513 from the decoded pixel data 1517 stored in the decoded picture buffer 1550. In some embodiments, the decoded pixel data 1517 is also stored in a linear buffer (not shown) for inter picture prediction and spatial MV prediction.

[0187] In some embodiments, the content of the decoded picture buffer 1550 is used for display. The display device 1555 retrieves the content of the decoded picture buffer 1550 for direct display or retrieves the content of the decoded picture buffer to a display buffer. In some embodiments, the display device receives pixel values from the decoded picture buffer 1550 by pixel transfer.

[0188] The motion compensation module 1530 generates predicted pixel data 1513 from the decoded pixel data 1517 stored in the decoded picture buffer 1550 according to motion compensation MVs (MC MVs). These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 1595 to the predicted MVs received from the MV prediction module 1575.

[0189] The MV prediction module 1575 generates predicted MVs based on reference MVs, which are generated for decoding previous video frames, e.g., motion compensation MVs for performing motion compensation. The MV prediction module 1575 retrieves the reference MVs of previous video frames from the MV buffer 1565. The video decoder 1500 stores the motion compensation MVs generated for decoding the current video frame in the MV buffer 1565 as reference MVs for generating predicted MVs.

[0190] The loop filter 1545 performs filtering or smoothing operations on the decoded pixel data 1517 to reduce coding artifacts, especially at the boundaries of pixel blocks. In some embodiments, the filtering operations performed include sample adaptive offset (SAO). In some embodiments, the filtering operations include adaptive loop filter (ALF).

[0191] Figure 16 Parts of the video decoder 1500 that mark mutually exclusive groups of coding modes or tools are shown. As shown, the video decoder 1500 implements a combined prediction module 1610, which receives the intra prediction values generated by the intra picture prediction module 1525. The combined prediction module 1610 may also receive inter prediction values from the motion compensation module 1530 and a second motion compensation module 1630. The combined prediction module 1610 in turn generates predicted pixel data 1513, which may be further filtered by a set of prediction filters 1650.

[0192] The MV buffer provides merge candidates to the motion compensation modules 1530 and 1630. The MV buffer 1565 also stores the motion information and mode directions for decoding the current block for use by subsequent blocks. The merge candidates can be changed, extended, and / or refined by the MV refinement module 1665.

[0193] The codec mode (or tool) control 1600 controls the operations of the intra-image prediction module 1525, the motion compensation module 1530, the second motion compensation module 1630, the MV refinement module 1665, the combined prediction module 1610, and the prediction filter 1650.

[0194] The codec mode control 1600 may enable the MV refinement module 1665 to perform MV refinement (e.g., for DMVR) operations by searching for refined MVs or calculate gradients based on MV adjustments (e.g., for BDOF). The codec mode control module 1600 may enable the intra-prediction module 1525 and the motion compensation module 1530 to implement the MH-mode intra (or inter-intra) mode (e.g., CIIP). The codec mode control module 1600 may enable the motion compensation module 1530 and the second motion compensation module 1630 to implement the MH-mode inter mode (e.g., for the diagonal edge region of TPM). When combining the prediction signals from the intra-image prediction module 1525, the motion compensation module 1530, and / or the second motion compensation module 1630 to implement codec modes such as CIIP, TPM, GBI, and / or WP, the codec mode control module 1600 may enable the combined prediction module 1610 to adopt different weight schemes. The codec mode control 1600 may also enable the prediction filter 1650 to apply LIC, DIF, BIF, and / or HAD filters to the predicted pixel data 1513 or the decoded pixel data 1517.

[0195] The codec mode control module 1600 also determines which codec modes to enable and / or disable for coding / decoding the current block. The codec mode control module 1600 then controls the operations of the intra-image prediction module 1525, the motion compensation module 1530, the second motion compensation module 1630, the MV refinement module 1665, the combined prediction module 1610, and the prediction filter 1650 to enable and / or disable the specific codec modes.

[0196] In some embodiments, the codec mode control 1600 enables only a subset (one or more) of the codec modes from a specific set of two or more codec modes for encoding / decoding the current block or CU. This specific set of codec modes may include all or any subset of the following codec modes: CIIP, TPM, BDOF, DMVR, GBI, WP, LIC, DIF, BIF, and HAD. In some embodiments, when a first condition for enabling a codec mode of the current block is satisfied, the codec mode control 1600 disables a second codec mode of the current block.

[0197] In some embodiments, when the condition for enabling the first codec mode is satisfied and the first codec mode is enabled, in addition to the first codec mode, the codec mode control 1600 disables all codec modes in the specific subset of codec modes. In some embodiments, when both a first condition for enabling the first codec mode of the current block and a second condition for enabling the second codec mode of the current block are satisfied and the first codec mode is enabled, the codec mode control 1600 disables the second codec mode. For example, in some embodiments, when the codec mode control 1600 determines that the conditions for enabling GBI and BDOF are both satisfied and the GBI index indicates unequal weights for mixing prediction in list 0 and list 1, the codec mode control 1600 will disable BDOF. As another example, in some embodiments, when the codec mode control 1600 determines that the conditions for enabling GBI and DMVR are both satisfied and the GBI index indicates unequal weights for mixing prediction in list 0 and list 1, the codec mode control 1600 will disable DMVR.

[0198] In some embodiments, the codec mode control 1600 identifies the highest priority codec mode from one or more codec modes. If the highest priority codec mode is enabled, the codec mode control 1600 then disables all other codec modes in the specific set of codec modes, regardless of whether the enabling conditions for each other codec mode are satisfied. In some embodiments, according to the priority rules defined based on the parameters of the current block, such as the size or aspect ratio of the current block, each codec mode in the specific set of codec modes is assigned a priority.

[0199] The codec mode control 1600 receives a syntax element 1690 from the entropy decoder 1590 to indicate that one or more codec modes are enabled. Without receiving a syntax element to disable one or more other codec modes, the video decoder 1500 may also disable one or more other codec modes. In some embodiments, a first syntax element (e.g., a first flag) is used to indicate whether to apply the multi-hypothesis prediction mode and a second syntax element (e.g., a second flag) is used to indicate whether to apply the CIIP or TPM mode. Correspondingly, the first and second elements are decoded from a first bin and a second bin in the bitstream 1595. In some embodiments, the second bin for deciding between CIIP and TPM is signaled only if the first bin indicates that the multi-hypothesis mode is enabled.

[0200] Figure 17 Conceptually illustrates a process 1700 for implementing mutually exclusive groups of codec modes or tools. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing the decoder 1500 execute the process 1700 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing the decoder 1500 executes the process 1700.

[0201] The decoder receives (at block 1710) data for a pixel block of a current block of a current image to be decoded as a video.

[0202] The decoder identifies (at block 1730) a highest-priority codec mode among one or more codec modes. In some embodiments, each codec mode in a particular set of codec modes is assigned a priority according to a priority rule defined based on parameters of the current block.

[0203] If the highest-priority codec mode is enabled, the decoder disables (at block 1740) all other codec modes of the codec mode specific set. The conditions for enabling the various codec modes are described in the above paragraphs associated with those codec modes. The conditions for enabling a codec mode may include receiving explicit syntax elements from the bitstream for the codec mode. The conditions for enabling a codec mode may also include having specific characteristics or parameters of the current block being decoded (e.g., size, aspect ratio). For example, when the codec mode specific set includes a first codec mode assigned a higher priority and a second codec mode assigned a lower priority, and when the first codec mode is enabled, the decoder disables (at block 1745) the second codec mode for the current block. In some embodiments, when the first codec mode is enabled, in addition to the first codec mode, the decoder disables all other codec modes in the codec mode specific set. In some embodiments, because GBI is assigned a higher priority than BDOF, if the GBI weight index indicates unequal weights, the decoder enables GBI (which means using unequal weights to blend the inter-predictors from list 0 and list 1), and disables BDOF. Right for example, in some embodiments, because GBI is assigned a higher priority than DMVR, if the GBI weight index indicates unequal weights, the decoder enables GBI (which means using unequal weights to blend the inter-predictions from list 0 and list 1), and disables DMVR. Again for example, in some embodiments, because CIIP is assigned a higher priority than other disabled tools, if the CIIP flag is equal to 1, the encoder enables CIIP, but disables GBI, BDOF, and / or DMVR.

[0204] The decoder decodes (at block 1750) the current block by using inter-prediction calculated according to the enabled codec mode.

[0205] XVII. Exemplary Electronic System

[0206] Many of the features and applications described above are implemented as software processes specified as a set of instructions recorded on a computer-readable storage medium (also referred to as a computer-readable medium). When these instructions are executed by one or more computing or processing units (e.g., one or more processors, processor cores, or other processing units), they cause the processing unit to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard disk drives, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. Computer-readable media include, but are not limited to, carrier waves transmitted wirelessly or via a wired connection and electronic signals.

[0207] In this specification, the term "software" is intended to include firmware resident in read-only memory or applications stored on magnetic storage that can be read into storage for processing by a processor. Additionally, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while remaining distinct software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Ultimately, any combination of separate programs together implement the software inventions described within the scope of this invention. In some embodiments, when installed to operate one or more electronic systems, a software program defines one or more specific machine implementations that run and execute the operations of the software program.

[0208] Figure 18 Conceptually illustrated is an electronic system 1800 in which some embodiments of the present invention are implemented. The electronic system 1800 can be a computer (e.g., a desktop computer, a personal computer, a tablet computer, etc.), a telephone, a PDA, or any other suitable electronic device. Such an electronic system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 1800 includes a bus 1805, a processing unit 1810, an image processing unit (GPU) 1815, a system memory 1820, a network 1825, a read-only memory 1830, a permanent storage device 1835, an input device 1840, and an output device 1845.

[0209] The bus 1805 collectively represents all system, peripheral, and chipset buses that communicatively connect multiple internal devices of the electronic system 1800. For example, the bus 1805 communicatively connects the processing unit 1810 with the GPU 1815, the read-only memory 1830, the system memory 1820, and the permanent storage device 1835.

[0210] From these various memory units, the processing unit 1810 retrieves the instructions to be executed and the data to be processed to execute the processes of the present invention. The processing unit can be a single processor or a multi-core processor in different embodiments. Some embodiments are transmitted for execution by the GPU 1815. The GPU 1815 can offload various computations provided by the processing unit 1810 or implement image processing.

[0211] The read-only memory (ROM) 830 stores data and instructions used by the processing unit 1810 and other modules of the electronic system. On the other hand, the permanent storage device 1835 is a read-write storage device. This device is a non-volatile memory that stores instructions and data even when the electronic system 1800 is turned off. Some embodiments of the present invention use a mass storage device (such as a magnetic or optical disk and its corresponding hard disk drive) as the permanent storage device 1835.

[0212] Other embodiments use removable storage devices (such as floppy disks, flash storage devices, etc. and their corresponding hard disk drives) as the permanent storage device. Like the permanent storage device 1835, the system memory 1820 is a read-write storage device. However, unlike the storage device 1835, the system memory 1820 is a volatile read-write memory, such as random access memory. The system memory 1820 stores some instructions and data that the processor uses during operation. In some embodiments, the processes according to the present invention are stored in the system memory 1820, the permanent storage device 1835, and / or the read-only memory 1830. For example, various storage units include instructions for processing multimedia video according to some embodiments. From these various storage units, the processing unit 1810 retrieves the instructions to be executed and the data to be processed to perform the processing of some embodiments.

[0213] The bus 1805 is also connected to the input and output devices 1840 and 1845. The input device 1840 enables the user to communicate information and select commands with the electronic system. The input device 1840 includes an alphanumeric keyboard and a pointing device (also called a "cursor control device"), a camera (e.g., a network camera), a microphone, or a similar device for receiving voice commands, etc. The output device 1845 displays the images or other output data generated by the electronic system. The output device 1845 includes a printer and a display device, such as a cathode ray tube (CRT) or a liquid crystal display (LCD), and a speaker or a similar sound output device. Some embodiments include a device such as a touch screen that serves as both an input and an output device.

[0214] Finally, as Figure 18 shown, the bus 1805 also couples the electronic system 1800 to the network 1825 through a network adapter (not shown). In this way, the computer can be part of a computer network (such as a local area network ("LAN"), a wide area network ("WAN"), or an internal network, or a network of networks, such as the Internet). Any or all components of the electronic system 1800 can be used in combination with the present invention.

[0215] Some embodiments include electronic components such as microprocessors, computer program instructions stored in the form of a machine-readable or computer-readable medium (or referred to as computer-readable storage medium, machine-readable storage medium, or machine-readable storage medium), and memory. Some examples of such computer-readable media include RAM, ROM, compact disc read-only memory (CD-ROM), recordable compact disc (CD-R), rewritable compact disc (CD-RW), digital versatile disc read-only (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini-SD card, micro-SD card, etc.), magnetic and / or solid state disk drives, read-only and recordable Blu-ray discs, ultra density optical discs, any other optical or magnetic media, and floppy disks. The computer-readable media can store computer programs executed by at least one processing unit and a set of instructions including those for performing various operations. Examples of computer programs or computer code include machine code (such as generated by a compiler) and files including high-level code executed by a computer, an electronic component, or a microprocessor using an interpreter.

[0216] Although the above description mainly refers to microprocessors or multi-core processors that execute software, many of the features and applications described above are performed by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions stored within the circuit itself. Additionally, some embodiments execute software in programmable logic devices (PLDs), ROM, or RAM devices.

[0217] As used in this specification and any claims of this application, the terms "computer", "server", "processor", and "memory" all refer to electronic or other technological devices. These terms exclude humans or groups of humans. For illustrative purposes, the term display or displaying means displaying on an electronic device. As used in this specification and any claims of this application, the terms "computer-readable medium", "computer-readable media", and "machine-readable medium" are restricted to tangible, physical objects that store information in a computer-readable form. These terms exclude any wireless signals, wired download signals, and any other transient signals.

[0218] Although the present invention has been described with reference to various specific details, those of ordinary skill in the art will recognize that the present invention may be embodied in other specific forms without departing from the spirit of the present invention. In addition, the various diagrams (including FIGS. 14 and 17) conceptually illustrate processes. The specific operations of these processes may be performed in the exact order shown and described. The specific operations may not be performed in a continuous series of operations, and different specific operations may be performed in different embodiments. In addition, the processes may be implemented using various subprocesses or as part of a larger macro process. Thus, those of ordinary skill in the art will understand that the present invention is not limited by the foregoing illustrative details, but is defined by the appended claims for patent.

[0219] Note

[0220] The subject matter described herein is sometimes shown including different other components or connected to different components. It can be understood that such depicted architectures are merely examples, and in fact many other architectures can be implemented to achieve the same function. Conceptually, any arrangement of components that achieves the same function is effectively "associated" so as to achieve the desired function. Thus, any two components combined herein to achieve a particular function can be regarded as "associated" with each other so as to achieve the desired function, regardless of the architecture or intermediate components. Similarly, any two components so associated can also be regarded as "operably connected" or "operably coupled" to each other to achieve the desired function, and any two components that can be so associated can also be regarded as "operably coupled" to each other to achieve the desired function. Specific examples of operably coupling include, but are not limited to, components that physically mate and / or physically interact and / or wirelessly communicate and / or wirelessly interact and / or logically interact and / or logically communicate.

[0221] In addition, with respect to the use of substantially any plural and / or singular terms herein, those of ordinary skill in the art can appropriately convert them from plural to singular and / or from singular to plural according to the context and application. For clarity, various singular / plural permutations may be explicitly set forth herein.

[0222] In addition, those skilled in the art will understand that, generally, the terms used herein, particularly the terms used in the appended claims (such as the body of the appended claims), are generally intended to be "open" terms. For example, the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "having at least", the term "includes" should be interpreted as "including but not limited to", etc. Those skilled in the art will further understand that if a specific number of recited claim limitations is intended, such intent will be expressly recited in the claims, and in the absence of such recitation, such intent is not present. For example, for purposes of illustration, the following appended claims may contain the use of introductory phrases "at least one" and "one or more" to introduce claim limitations. However, the use of such phrases should not be construed as implying that a claim limitation introduced by the indefinite article "a" or "an" limits any particular claim containing such introduced claim limitation to only embodiments containing one such element, even when the same claim includes an introductory phrase "one or more" or "at least one" as well as an indefinite article such as "a" or "an", "a" and / or "an" should be interpreted to mean "at least one" or "one or more", and the same applies to the definite article introducing a claim limitation. In addition, even if a specific number of introduced claim limitations are expressly recited, those of ordinary skill in the art will recognize that such recitations should be interpreted to mean at least one of the recited numbers, e.g., the mere recitation of "two elements" without further modification means at least two elements, or two or more elements. Further, in instances where a convention such as "at least one of A, B, and C, etc." is used, generally such construction is intended to be understood by those of ordinary skill in the art, e.g., "the system has at least one of A, B, and C" will include but not be limited to the system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. In instances where a convention such as "at least one of A, B, or C" is used, generally such construction is intended to be understood by those of ordinary skill in the art, e.g., "the system has at least one of A, B, or C" will include but not be limited to the system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. Those skilled in the art will further understand that in fact, any separator and / or phrase that represents two or more alternative terms in a description, claim, or drawing will be understood to contemplate the possibility of including one of the terms, any one of the terms, or both terms. For example, the phrase "A or B" will be understood to contemplate the possibility of "A or B" or "A and B".

[0223] As can be understood from the foregoing, for purposes of illustration, various embodiments of the present invention have been described herein, and various modifications may be made without departing from the scope and spirit of the present invention. Accordingly, the various embodiments described herein are not intended to be limiting, and the true scope and spirit are indicated by the appended claims.

Claims

1. A video decoding method, comprising: Receiving data of a pixel block of a current block of a current image to be decoded into a video; When a first codec mode of the current block is enabled, disabling a second codec mode of the current block, wherein the first codec mode and the second codec mode specify different methods for calculating an inter-frame prediction of the current block, wherein a specific set of two or more codec modes includes the first codec mode and the second codec mode, and the specific set of codec modes is a subset of bi-directional prediction with CU-level weights, decoder-side motion vector refinement, bi-directional optical flow, weighted prediction, and combined inter-frame and intra-frame prediction, and the number of codec modes included in the specific set is two or more; And Decoding the current block by using an inter-frame prediction calculated according to the enabled codec mode.

2. The video decoding method according to claim 1, wherein Wherein when the first codec mode is enabled, all other codec modes in the specific set of codec modes are disabled except the first codec mode.

3. The video decoding method according to claim 2, wherein The specific set of codec modes includes bi-directional prediction with CU-level weights, decoder-side motion vector refinement, and combined inter-frame and intra-frame prediction, and wherein: Bi-directional prediction with CU-level weights is a codec mode in which a video decoder performs a weighted average of two prediction signals in two different directions to generate the inter-frame prediction; Decoder-side motion vector refinement is a codec mode in which the video decoder searches for a refined motion vector around an initial motion vector and uses the refined motion vector to generate the inter-frame prediction; and Combined inter-frame and intra-frame prediction is a codec mode in which the video decoder combines an inter-frame prediction signal and an intra-frame prediction signal to generate the inter-frame prediction.

4. The video decoding method according to claim 2, wherein The specific set of codec modes includes bi-directional prediction with CU-level weights, bi-directional optical flow, and combined inter-frame prediction and intra-frame prediction, and wherein: Bi-directional prediction with CU-level weights is a codec mode in which a video decoder performs a weighted average of two prediction signals in two different directions to generate the inter-frame prediction; Bi-directional optical flow is a codec mode in which the video decoder calculates motion refinement to minimize distortion between prediction samples in different directions and adjusts the inter-frame prediction based on the calculated refinement; and Combined inter-frame prediction and intra-frame prediction is a codec mode in which the video decoder combines an inter-frame prediction signal and an intra-frame prediction signal to generate the inter-frame prediction.

5. The video decoding method according to claim 1, wherein The first codec mode is combined inter-frame and intra-frame prediction and the second codec mode is bi-directional prediction with CU-level weights.

6. The video decoding method according to claim 1, wherein The first codec mode is bi-directional prediction with CU-level weights and the second codec mode is bi-directional optical flow.

7. The video decoding method according to claim 1, wherein The first codec mode is bi-directional prediction with CU-level weights and the second codec mode is decoder-side motion vector refinement.

8. The video decoding method according to claim 1, wherein The first codec mode is combined inter-frame and intra-frame prediction and the second codec mode is bi-directional optical flow.

9. The video decoding method according to claim 1, wherein The first codec mode is combined inter-frame and intra-frame prediction and the second codec mode is decoder-side motion vector refinement.

10. An electronic device, comprising: A video decoder circuit for performing operations, including: Receive data of a pixel block of a current block of a current picture to be decoded into video; When a first codec mode of the current block is enabled, disable a second codec mode of the current block, where the first codec mode and the second codec mode specify different methods for calculating inter-frame prediction of the current block, where a specific set of two or more codec modes includes the first codec mode and the second codec mode, and the specific set of codec modes is a subset of bi-directional prediction with CU-level weights, decoder-side motion vector refinement, bi-directional optical flow, weighted prediction, and combined inter-frame and intra-frame prediction, and the number of codec modes included in the specific set is two or more; and Decode the current block by using the inter-frame prediction calculated according to the enabled codec mode.

11. The electronic device according to claim 10, wherein A specific set of two or more codec modes includes the first codec mode and the second codec mode, and where when the first codec mode is enabled, all other codec modes in the specific set of codec modes are disabled except the first codec mode.

12. The electronic device according to claim 10, wherein The specific set of codec modes includes bi-directional prediction with CU-level weights, decoder-side motion vector refinement, and combined inter-frame and intra-frame prediction, and where: Bi-directional prediction with CU-level weights is a codec mode in which the video decoder circuit performs a weighted average of two prediction signals in two different directions to generate the inter-frame prediction; Decoder-side motion vector refinement is a codec mode in which the video decoder circuit searches for a refined motion vector around an initial motion vector and uses the refined motion vector to generate the inter-frame prediction; and Combined inter-frame and intra-frame prediction is a codec mode in which the video decoder circuit combines an inter-frame prediction signal and an intra-frame prediction signal to generate the inter-frame prediction.

13. The electronic device according to claim 10, wherein The specific set of codec modes includes bi-directional prediction with CU-level weights, bi-directional optical flow, and combined inter-frame and intra-frame prediction, and where: Bi-directional prediction with CU-level weights is a codec mode in which the video decoder circuit performs a weighted average of two prediction signals in two different directions to generate the inter-frame prediction; Bi-directional optical flow is a codec mode in which the video decoder circuit calculates motion refinement to minimize the distortion between prediction samples in different directions and adjusts the inter-frame prediction based on the calculated motion refinement; and Combined inter-frame and intra-frame prediction is a codec mode in which the video decoder circuit combines an inter-frame prediction signal and an intra-frame prediction signal to generate the inter-frame prediction.