Video decoding method and related electronic device

By introducing mutual exclusion rules in the video decoder and disabling unnecessary encoding and decoding modes, the problem of inefficient inter prediction mode in the prior art is solved, and higher hardware utilization and video decoding efficiency are achieved.

CN113853794BActive Publication Date: 2025-05-23HFI INNOVATION INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080016256.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-26
Filing Date
2020-02-27
Publication Date
2025-05-23
Estimated Expiration
2040-02-27

AI Technical Summary

Technical Problem

Existing video encoding and decoding techniques have problems with inefficiency in inter prediction mode, especially in motion vector prediction and predict direction coding.

Method used

A mutual exclusion rule is proposed, disabling the second codec mode of the current block, and applying the rule only when the first codec mode is enabled, to improve the hardware utilization of the video decoder and simplification of the pipeline stage.

Benefits of technology

By disabling unnecessary codec modes, the hardware idle time and pipeline delay are reduced, and the video decoding efficiency and hardware resource utilization are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113853794B_ABST
    Figure CN113853794B_ABST
Patent Text Reader

Abstract

A video decoder implementing mutually exclusive groups of codec modes is provided. The video decoder receives data of a pixel block of a current block to be decoded as a current image of a video. When a first codec mode of the current block is enabled, a second codec mode of the current block is disabled, wherein the first codec mode and the second codec mode specify different methods for calculating an inter-frame prediction of the current block. The current block is decoded by using the inter-frame prediction calculated according to the enabled codec mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to video processing and more particularly to a method of signaling a codec mode. Background Art

[0002] Unless otherwise indicated herein, the approaches described in this section are not prior art to the claims listed below and are not admitted to be prior art by inclusion in this section.

[0003] High Efficiency Video Codec (HEVC) is an international video codec standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). HEVC is a hybrid block-based motion compensated DCT-like transform codec architecture. The basic unit of compression is a 2N×2N square block, termed a coding unit (CU), and each CU can be recursively split into four smaller CUs until a predetermined minimum size is reached. Each CU contains one or more prediction units (PUs).

[0004] In order to achieve the best codec efficiency of the hybrid codec architecture in HEVC, each PU has two types of prediction modes, which are intra prediction and inter prediction. For intra prediction mode, spatially adjacent reconstructed pixels can be used to generate directional predictions. There are up to 35 directions in HEVC. For inter prediction mode, temporally reconstructed reference frames can be used to generate motion compensated predictions. There are three different modes, including Skip, Merge, and Inter Advanced Motion Vector Prediction (AMVP) mode.

[0005] When the PU is encoded and decoded in the inter-frame AMVP mode, motion compensated prediction is performed with the transmitted motion vector difference (MVD), which can be used together with the motion vector predictor (MVP) to generate a motion vector (MV). In order to determine the MVP in the inter-frame AMVP mode, the advanced motion vector prediction (AMVP) scheme is used to select a motion vector predictor from an AMVP candidate set including two spatial MVPs and one temporal MVP. Therefore, in the AMVP mode, the MVP index of the MVP and the corresponding MVD need to be encoded and transmitted. In addition, the inter-frame prediction direction and the reference frame index of each list to specify the prediction direction in bidirectional prediction and unidirectional prediction (which are list 0 (L0) and list 1 (L1)) should also be encoded and transmitted.

[0006] When a PU is encoded or decoded in skip or merge mode, no motion information is transmitted except the merge index of the selected candidate. This is because the skip and merge modes utilize motion inference methods (MV=MVP+MVD, where MVD=0) to obtain motion information from spatial neighboring blocks (spatial candidates) or temporal blocks (temporal candidates) located in a collocated picture, where the collocated picture is the first reference picture in list 0 or list 1, which is signaled in the slice header. In case of skip PU, the residual signal is also omitted. To determine the merge index for skip and merge modes, a merge scheme is used to select a motion vector predictor from a merge candidate set consisting of four spatial MVPs and one temporal MVP. Summary of the invention

[0007] The subsequent overview is illustrative only and is not intended to be limiting in any way. That is, the subsequent overview is provided to introduce the concepts, highlights, benefits, and advantages of the novel and non-obvious technologies described herein. Selected but not all embodiments are further described in the detailed description. Therefore, the subsequent overview is not intended to identify essential features of the claimed subject matter, or is not intended to be used to determine the scope of the claimed subject matter.

[0008] Embodiments of the present invention provide a video decoder that implements mutually exclusive groups of codec modes or tools. The decoder receives data of a block of a current block of a current image to be decoded as a video. When a first codec mode of the current block is enabled, the decoder disables a second codec mode of the current block, wherein the first codec mode and the second codec mode specify different methods for calculating an inter-frame prediction of the current block. In other words, the second codec mode of the current block can be applied only when the first codec mode is disabled. The decoder decodes the current block by using the inter-frame prediction calculated according to the enabled codec mode.

[0009] The present invention proposes a mutual exclusion rule for the exclusion setting of some encoding and decoding tools of the current CU, so that the pipeline stage can be shortened and the hardware utilization rate can be higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings are included to provide a further understanding of the present invention and are incorporated into and constitute a part of the present invention. The drawings illustrate embodiments of the present invention and together with the description are used to explain the principles of the present invention. Because some components may be shown as not proportional to the size in the actual embodiment in order to clearly illustrate the concept of the present invention, the drawings are not necessarily drawn to scale.

[0011] Figure 1 Motion candidates for merge mode are marked.

[0012] Figure 2Conceptually illustrated is the use of decoder-side motion vector refinement based on bilateral-matching to encode or decode a current block.

[0013] Figure 3 The search process of Decoder Motion Vector Refinement (DMVR) is shown.

[0014] Figure 4 The DMVR integer brightness sample search pattern is marked.

[0015] Figure 5 Deriving a lighting-based prediction offset is conceptually illustrated.

[0016] Figure 6 An example derivation of prediction offsets is shown.

[0017] Figure 7 The extended CU region used by BDOF for encoding and decoding CUs is shown.

[0018] Figure 8 An exemplary 8x8 transform unit block and a bidirectional filter aperture are shown.

[0019] Fig. 9 The filtering process under the Hadamard transform domain filter is shown.

[0020] Fig.10 The adaptive weighting applied along the diagonal edge between two triangular prediction units is shown.

[0021] Fig.11a It is conceptually shown that a pixel block is encoded or decoded by using the MH mode for intra.

[0022] Fig.11b It is conceptually shown that a current block is encoded and decoded by using an MH mode for inter.

[0023] Fig.12 Exemplary video encoders that can implement mutually exclusive sets of codec modes or tools are identified.

[0024] Fig.13 A portion of the video encoder that can implement mutually exclusive sets of codec modes or tools is marked.

[0025] Fig.14A process for implementing mutually exclusive sets of codec modes or tools at a video encoder is conceptually illustrated.

[0026] Fig.15 Exemplary video decoders that can implement mutually exclusive sets of codec modes or tools are identified.

[0027] Fig.16 A portion of the video decoder is shown that can implement mutually exclusive groupings of codec modes or tools.

[0028] Fig.17 A process for implementing mutually exclusive sets of codec modes or tools at a video decoder is conceptually illustrated.

[0029] Fig.18 An electronic system is conceptually illustrated in which some embodiments of the present invention may be implemented. DETAILED DESCRIPTION

[0030] In the subsequent detailed description, many specific series are given by way of example to provide a thorough understanding of the relevant teachings. Various variations, derivatives and / or extensions based on the teachings described herein are within the scope of protection of the present invention. In some cases, well-known methods, processes, components and / or circuits related to one or more example embodiments disclosed herein may be described at a relatively high level without details to avoid unnecessary confusion of aspects of the teachings of the present invention.

[0031] I. Merge Mode

[0032] Figure 1 The motion candidates of the merge mode are marked. 0 , A 1 , B 0 and B 1 Derive up to four spatial MV candidates, and from T BR or T CTR Derive a temporal MV candidate (first use T BR , if T BR Not available, use T CTR ). If any of the four spatial MV candidates is unavailable, then position B 2Used to derive MV candidates as an alternative. After the derivation process of four spatial MV candidates and one temporal MV candidate, redundancy removal (pruning) is applied in some embodiments to remove redundant MV candidates. If the number of available MV candidates is less than 5 after removing redundancy (pruning), three types of additional candidates are derived and added to the candidate set (candidate list). The video encoder decides to select a final candidate within the candidate set in skip or merge mode based on rate-distortion optimization (RDO), and transmits the index to the video decoder. (Skip mode and merge mode are collectively referred to as "merge mode" in this article).

[0033] II. Decoder-side Motion Vector Refinement (DMVR)

[0034] To increase the accuracy of the MV in merge mode, in some embodiments, decoder-side motion vector refinement or DMVR based on bidirectional matching is applied. In bidirectional prediction operation, the video codec searches for a refined MV around the initial MV in reference image list L0 and reference image list L1. The bidirectional matching method calculates the distortion between two candidate blocks in reference image list L0 and list L1.

[0035] Figure 2 Conceptually illustrates encoding or decoding a current block 200 using decoder-side motion vector refinement based on bidirectional matching. As shown, SAD (Sum of Absolute Difference) is calculated for MV candidates (e.g., MV0' and MV1') around an initial MV (e.g., MV0 and MV1) based on the differences between the pixel blocks referenced by these MV candidates (e.g., R0' and R1'). The MV candidate with the lowest SAD becomes the refined MV and is used to generate a bidirectional prediction signal.

[0036] In some embodiments, DMVR is applied as follows. For DMVR with luma CB width or height > 16, the CU is split into multiple 16x16, 16x8, or 8x16 luma sub-blocks (and corresponding chroma sub-blocks). Next, when the SAD of the zero MVD position between list 0 and list 1 (indicated by the initial MV, labeled MV0 and MV1) is small, the DMVR of each sub-block or small CU is ended early. Based on an integer-step search of 25-point SAD (i.e., ±2 integer-step refinement of the search range), the search range fractional samples are generated by bilinear interpolation.

[0037] In some embodiments, when the enabling conditions for DMVR are met, DMVR is applied to the coded and decoded CUs. In some embodiments, the enabling conditions for DMVR can be any subset of (i) to (v). (i) CU-level merge mode with bi-predicted MVs; (ii) one reference image among the past images with respect to the current image and another reference image among the future images with respect to the current image; (iii) the distances from the two reference images to the current image (e.g., picture order count or POC difference) are the same; (iv) the CU has more than 64 luma samples; (v) both the CU height and the CU width are greater than or equal to 8 luma samples.

[0038] The refined MVs derived by the DMVR process are used to generate inter-prediction samples and also for temporal motion vector prediction for future image coding. While the original MVs are used for the deblocking process and also for spatial motion vector prediction for future CU coding.

[0039] a. Search scheme

[0040] As Figure 2 shown, the search points surrounding the initial MV and the MV offset follow the MV difference mirroring principle. In other words, any point examined by DMVR, labeled as a candidate MV pair (MV0, MV1) follows the following two equations:

[0041] MV0′ = MV0 + MV_offset

[0042] MV1′ = MV1 - MV_offset

[0043] where MV_offset represents the refined offset between the initial MV and the refined MV in one of the reference images. In some embodiments, the refined search range is two integer luma samples from the initial MV.

[0044] Figure 3 Shows the search process of DMVR. As shown, the search includes an integer sample offset search stage and a fractional sample refinement stage.

[0045] Figure 4 The DMVR integer sample search pattern is marked. As shown, a 25-point full search is applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR is ended. Otherwise, the SADs of the remaining 24 points are calculated and examined in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search stage. To reduce the penalty for DMVR refinement uncertainty, it is proposed to prefer the original MV in the DMVR process. The SAD calculated from the initial MV pair will be reduced to 1 / 4 of its SAD value.

[0046] Back to Figure 3 . The integer sample search is followed by the fractional sample refinement. To save computational complexity, the fractional sample refinement is derived by using the parametric error surface equation instead of an additional search using SAD comparison. The fractional sample refinement is conditionally called based on the output of the integer sample search stage. The fractional sample refinement is further applied when the integer sample search stage ends with the center having the minimum SAD in the first iteration or the second iteration.

[0047] In the sub-pixel offset estimation based on the parameterized error surface, the cost at the current position and the costs from the center to the four neighboring positions are used to fit a 2-D parabolic error surface equation of the following form:

[0048] E(x,y)=A(x min ) 2 +B(yy min ) 2 +C

[0049] Where (x min ,y min ) corresponds to the fractional position with the minimum cost and C corresponds to the minimum cost value. By parsing the above equation using the cost values ​​of the five search points, (x min ,y min ) is calculated as:

[0050] x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0)))

[0051] y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0)))

[0052] Since all cost values ​​are integers and the minimum is E(0,0), x min and min The value of is usually automatically constrained to be between -8 and 8. This corresponds to a half peak offset with 1 / 16 pixel MV accuracy in VTM4. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to obtain the sub-pixel accurate refinement δMV.

[0053] b. Bilinear interpolation and sample filling

[0054] In some embodiments, the resolution of the MV is 1 / 16 luma sample. An 8-tap interpolation filter is used to interpolate samples at fractional positions. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so the samples at these fractional positions need to be interpolated for the DMVR search process. In order to reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect of using a bilinear filter is that with a 2-sample search range, DMVR does not access more reference samples than a normal motion compensation process. After obtaining the refined MV with the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. In order not to access more reference samples to the normal MC process, samples that are not needed for the interpolation process based on the original MV but are needed for the interpolation process based on the refined MV are filled from these available samples.

[0055] c. Maximum DMVR processing unit

[0056] In some embodiments, when the width and / or height of the CU is greater than 16 luma samples, it further enters a sub-block with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16.

[0057] III. Weighted Prediction (WP)

[0058] Weighted prediction (WP) is a codec tool supported by H.264 / AVC and HEVC standards to efficiently encode and decode video content with padding. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in each reference picture list L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied.

[0059] IV. Lighting-Based Prediction Offset

[0060] As mentioned before, inter prediction explores pixel correlation between frames and if the scene is static, the correlation will be valid and motion estimation can easily find similar blocks with similar pixel values ​​in temporally adjacent frames. However, in some real cases, multiple frames will be shot with different lighting conditions. Even if the content is similar and the scene is static, the pixel values ​​between multiple frames will be different.

[0061] In some embodiments, a Neighboring-derived PredictionOffset (NPO) is used to add a prediction offset to improve the motion compensated predictor. Based on this offset, different lighting conditions between multiple frames can be taken into account. The offset is derived using neighboring reconstructed pixels (NRP) and an extended motion compensated predictor (EMCP).

[0062] Figure 5 Conceptually illustrated is the derivation of illumination based prediction offsets. The pattern selected for the NRP and EMCP is N pixels to the left and M pixels above the current PU, where N and M are predetermined values. The pattern can be of any size and shape and can be determined based on any coding parameters, such as PU or CU size, as long as they are the same for both the NRP and EMCP. The offset is calculated as the mean pixel value of the NRP minus the mean pixel value of the EMCP. The derived offset is unique across the PU and is applied to the entire PU with the motion compensated predictor.

[0063] Figure 6 An example derivation of prediction offsets is shown. First, for each neighboring position (left and above the boundary, shaded in gray), the individual offset is calculated as the corresponding pixel in the NRP minus the pixel in the EMCP. In this example, offset values ​​6, 4, 2, -2 are generated for the above neighboring position and 6, 6, 6, 6 for the left neighboring position. Second, when all individual offsets are calculated and obtained, the derived offset for each position in the current PU will be the average of the offsets from the left and above positions. For example, in the first position in the upper left corner, an offset of 6 is generated by averaging the offsets from the left and above. For the next position, the offset is equal to (6+4) / 2, which is 5. The offsets for each position can be processed and generated sequentially in raster scan order. Because neighboring pixels are more highly correlated to boundary pixels, so are the offsets. This method can adapt the offsets according to the pixel position. The derived offsets will be adapted to the entire PU and will be applied to each PU position individually together with the motion compensation predictor.

[0064] In some embodiments, local illumination compensation (LIC) is used to correct the result of inter prediction. LIC is a method of inter prediction that uses neighboring samples of the current block and the reference block to generate a linear model, which is characterized by a scaling factor a and an offset b. The scaling factor a and the offset b are derived by referring to the neighboring samples of the current block and the reference block. For each CU, the LIC mode can be adaptively enabled or disabled.

[0065] V. Generalized Bidirectional Inference (GBI)

[0066] Generalized bi-prediction (GBI) is a method of inter-frame prediction that uses different weights for the predictors from L0 and L1, instead of using equal weights as in traditional bi-prediction. GBI is also called bi-directional prediction with weighted average (BMA) or bi-directional prediction with CU-level weights (BCW). In HEVC, a bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference images and / or using two different motion vectors. In some embodiments, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of two prediction signals.

[0067] P bi-pred =((8-w)*P 0 +w*P 1 +4)>>3

[0068] In some embodiments, five different possible weights, or w∈{-2,3,4,5,10}, are allowed in the weighted average bidirectional prediction. For each bidirectionally predicted CU, the weight w is determined in one of two ways: 1) for non-merged CUs, the weight index is signaled after the motion vector difference; 2) for merged CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. The weighted average of bidirectional prediction is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency images, all 5 weights are used. For non-low-latency images, only three different possible weights are used (w∈(3,4,5)).

[0069] In some embodiments, in the video encoder, a fast search algorithm is applied to find the weighted index without significantly increasing the encoder complexity. When combined with AMVR, it allows encoding and decoding the MVD of the CU with different precision. If the current image is a low-delay image, unequal weights (making the weight used for L0 not equal to the weight used for L1) are conditionally used for 1 pixel and 4 pixel motion vector precision. When combined with affine, affine motion estimation (ME) will use unequal weights only when the affine mode is selected as the current best mode. When the two reference images used for bidirectional prediction are the same, unequal weights are used conditionally. When certain conditions are not met, unequal weights are not searched according to the POC (picture order count) between the current image and its reference image, the QP (quantization parameter) of the codec, and the temporal level.

[0070] VI. Bidirectional Optical Flow (BDOF)

[0071] In some embodiments, bidirectional optical flow (BDOF), also known as BIO, is used to refine the bidirectional prediction signal of a CU at a 4x4 sub-block level. In particular, the video codec refines the bidirectional prediction signal by using sample gradients and a set of derived displacements.

[0072] BDOF is applied when the enabling conditions are met. In some embodiments, the enabling conditions for BDOF can be any subset of (1) to (4). (1) Both the CU height and the CU width are greater than or equal to 8 luma samples; (2) The CU is not encoded or decoded using the affine mode or ATMVP merge mode, which belongs to the sub-block merge mode; (3) The CU is encoded or decoded using the "true" bidirectional prediction mode, that is, one of the two reference images is before the current image in the display order and the other reference image is after the current image in the display order; (4) The CU has more than 64 luma samples. In some embodiments, BDOF is applied to the luma component.

[0073] The BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4x4 sub-block, the motion refinement (v) is calculated by minimizing the difference between the L0 and L1 prediction samples. x ,v y ). Motion refinement is then used to adjust the bidirectional prediction sample values ​​in the 4x4 sub-blocks. The subsequent steps are applied to the BDOF process.

[0074] First, the horizontal and vertical gradients of the two prediction signals are calculated by directly calculating the difference between two adjacent samples. as well as Right now:

[0075]

[0076] Among them I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k (k = 0, 1), and shift1 is calculated based on the brightness bit depth (bitDepth), such as shift1 = max(6, bitDepth-6). Then, the automatic and cross-correlations of the gradients S1, S2, S3, S5 and S6 are calculated as follows:

[0077] S 1 =∑ (i,j)∈Ω Abs(ψ x (i,j)),S 3 =∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j))

[0078]

[0079] in

[0080]

[0081] θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>n b )

[0082] where Ω is a 6x6 window around the 4x4 sub-block and the values ​​of na and nb are set equal to min(1, bitDepth-11) and min(4, bitDepth-8) respectively. The motion refinement (v) is then derived using the cross and auto-correlation terms. x ,v y ),as follows:

[0083]

[0084] Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows:

[0085] pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift

[0086] In some embodiments, the values ​​of na, nb, and ns2 are equal to 3, 6, and 12, respectively. In some embodiments, these values ​​are selected so that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits. In order to derive the gradient value, some prediction samples I in the list k (k=0,1) outside the boundary of the current CU need to be generated. (k) (i,j).

[0087] In some embodiments, BDOF uses an extended row / column around the CU boundary. Figure 7 The extended CU region used by BDOF for encoding and decoding a CU is shown. In order to control the computational complexity of generating prediction samples beyond the boundary, a linear filter is required to generate prediction samples in the extended area (white positions of the CU), and a normal 8-tap motion compensated interpolation filter is used to generate prediction samples within the CU (shadow positions of the CU). These extended sample values ​​are only used for gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled (e.g., copied) from their nearest neighboring blocks.

[0088] VII. Combined Inter and Intra Prediction (CIIP)

[0089] In some embodiments, when the enabling conditions of CIIP are met, the CU-level syntax of CIIP is signaled. For example, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. The enabling conditions may include that the CU is encoded and decoded in merge mode, and that the CU contains at least 64 luminance samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64). In order to form a CIIP prediction, an intra prediction mode is required. One or more possible intra prediction modes may be used: for example, DC, plane, horizontal or vertical. Then, the inter prediction and intra prediction signals are derived using conventional intra and inter decoding processes. Finally, a weighted average of the inter and intra prediction signals is performed to obtain the CIIP prediction.

[0090] In some embodiments, if only one intra prediction mode (e.g., plane) is available for CIIP, the intra prediction mode for CIIP can be implicitly assigned to the mode (e.g., plane). In some embodiments, up to four intra prediction modes (including DC, PLANAR, horizontal, and vertical modes) can be used to predict the luminance component in CIIP mode. For example, if the CU shape is very wide (i.e., the width is greater than twice the height), the horizontal mode is not allowed; if the CU shape is very narrow (i.e., the height is greater than twice the width), the vertical mode is not allowed. In these cases, only 3 intra prediction modes are allowed. CIIP mode can use three most probable modes (MPM) for intra prediction. If the CU shape is very wide or very narrow as defined above, the MPM flag is inferred to be 1 without signaling. Otherwise, the MPM flag is signaled to indicate whether the CIIP intra prediction mode is one of multiple CIIP MPM candidate modes. If the MPM flag is 1, the MPM index is further signaled to indicate which MPM candidate mode is used in CIIP intra prediction. Otherwise, if the MPM flag is 0, the intra prediction mode is set to the "lost" mode in the MPM candidate list. For example, if the PLANAR mode is not in the MPM candidate list, then PLANAR is the lost mode, and the intra prediction mode is set to PLANAR. Because 4 possible intra prediction modes are allowed in CIIP, and the MPM candidate list contains only 3 intra prediction modes, one of the 4 possible modes must be the lost mode. The intra prediction mode of the CIIP-encoded CU will be retained and used for future intra mode encoding and decoding of neighboring CUs.

[0091] The inter prediction signal (or inter prediction) in CIIP mode Pinter is derived using the same inter prediction process applied to the conventional merge mode, and the intra prediction or intra prediction signal Pintra is derived using the CIIP intra prediction mode following the conventional intra prediction process. The intra and inter prediction signals are then combined using a weighted average, where the weight value depends on the neighboring blocks, on the intra prediction mode, or on where the sample is located in the coding block. In some embodiments, if the intra prediction mode is DC or planar mode, or if the block width or height is less than 4, equal weights are applied to the intra prediction and inter prediction signals. Otherwise, the weights are determined based on the intra prediction mode (horizontal mode or vertical mode in this case) and the sample position in the block. Starting from the nearest part of the intra prediction reference sample and ending at the farthest part of the intra prediction reference sample, the weight wt of each 4 region is set to 6, 5, 3 and 2 respectively. In some embodiments, the CIIP prediction or CIIP prediction signal PCIIP is derived according to the following:

[0092] P CIIP =((N1-wt)*P inter +wt*P intra +N2)>>N3

[0093] Where (N1, N2, N3) = (8, 4, 3) or (N1, N2, N3) = (4, 2, 2). When (N1, N2, N3) = (4, 2, 2), wt is selected from 1, 2 or 3.

[0094] VIII. Diffusion filter (DIF)

[0095] Diffusion filter for video codec is to use diffusion filter to apply to the prediction signal in video codec. Assume that pred is the prediction signal on a given block obtained by intra-frame or motion compensation prediction. In order to process the boundary points of the filter, the prediction signal is extended to the prediction signal predext. The extended prediction is formed by adding a line of reconstructed samples from the left and top of the block to the prediction signal and then the generated signal is mirrored in all directions.

[0096] The uniform diffusion filter is implemented by convolving the prediction signal with a fixed mask hI. In some embodiments, the prediction signal pred is composed of h I *pred, using the boundary extension mentioned later. This time, the filter mask hI is defined as:

[0097]

[0098] Directed diffusion filters such as the horizontal filter hhor and the vertical filter hver are used with fixed masks. The filtering is restricted to be applied only along the vertical or along the horizontal direction. The vertical filter is implemented by applying the fixed filter mask hver to the prediction signal and by using the transposed mask Implements a horizontal filter.

[0099] The expansion of the prediction signal is performed in the same way as for the uniform diffusion filter.

[0100] IX. Bilateral filtering (BIF)

[0101] It is a known technique to perform quantization in the transform domain to better preserve information in images and videos compared to quantization in the pixel domain. However, it is also known that quantized transform blocks can generate edge ringing artifacts around the edges of still images and moving objects in the video. Applying a bidirectional filter can significantly reduce the edge ringing artifact. In some embodiments, a small, low-complexity bidirectional filter is applied directly to the reconstructed samples after the inverse transform has been performed and combined with the predicted sample values.

[0102] In some embodiments, when applying bilateral filtering, each sample in the reconstructed image is replaced by a weighted average of itself and its neighboring samples. The weights are calculated based on the distance from the center sample and the difference in sample values. Figure 8 An exemplary 8x8 transform unit block and a bidirectional filter aperture are shown. The filter aperture is used for samples located at (1,1). As shown in the figure, since the filter is Figure 1 Within the small plus sign shape shown, all distances are either 0 or 1. The sample at (i,j) is filtered using its neighboring sample (k,l). The weight w(i,j,k,l) ​​is the weight assigned to sample (k,l) to filter sample (i,j), and is defined as follows:

[0103]

[0104] I(i,j) and I(k,l) are the original reconstructed intensity values ​​of samples (i,j) and (k,l) respectively. d is the spatial parameter, and σ r is a range parameter. The properties (or strength) of the bilateral filter are controlled by these two parameters. Samples located closer to a sample will be filtered, and samples with a smaller intensity difference to a sample will be filtered compared to samples further away and samples with a larger intensity difference. In some embodiments, σ ​​is set based on the transform unit size d , and based on the QP setting σ for the current block r .

[0105]

[0106] In some embodiments, a bidirectional filter is applied directly to each TU block after the inverse transform in both the encoder and the decoder. As a result, subsequent intra-coded blocks are predicted from sample values ​​that have been bidirectionally filtered. This also makes it possible to include the bidirectional filtering operation in the rate-distortion decision of the encoder.

[0107] In some embodiments, each sample in the transform unit is filtered using only its immediate neighbors. The filter has a plus-shaped filter aperture centered on the sample to be filtered. The output filtered sample value ID(i,j) is calculated as follows:

[0108]

[0109] For TU sizes greater than 16x16, blocks are treated as multiple 16x16 blocks using TU block width = TU block height = 16. In addition, rectangular blocks are treated as several examples of square blocks. In some embodiments, to reduce the number of calculations, a bidirectional filter is implemented using a lookup table (LUT) that stores all weights for a specific QP in a two-dimensional array. The LUT uses the intensity difference between the sample to be filtered and the reference sample as the index to the LUT in one dimension, and the TU size in another dimension. For efficient storage of the LUT, in some embodiments, the weights are rounded to 8 bits of precision.

[0110] X. Hadamard transform domain filter (HAD)

[0111] In some embodiments, a Hadard transform domain filter (HAD) is applied to luma reconstructed blocks with non-zero transform coefficients and to exclude 4x4 blocks if the quantization parameter is greater than 17. The filter parameters are explicitly derived from the codec information. If the HAD filter is applied, HAD is performed on the decoded samples after block reconstruction. The filter result is used for both output and spatial and temporal prediction. The filter has the same implementation for both intra and inter CU filtering. According to the HAD filter, for each pixel from the reconstructed block pixel, the process includes the following steps: (1) scanning the 4 neighboring pixels around the pixel including the current pixel for processing according to a scanning pattern, (2) reading the 4-point Hadard transform of the pixel, and (3) spectral filtering based on the following formula:

[0112]

[0113] where (i) is the index of the spectral component in the Hadron coded spectrum, R(i) is the spectral component of the reconstructed pixel corresponding to the index, m=4 is a normalization constant equal to the number of spectral components, and σ is a tabulated parameter derived from the codec quantization parameter QP using the following equation:

[0114] σ=2.64*2 (0.1296*(QP-11))

[0115] The first spectral component corresponding to the DC value is bypassed without filtering. An inverse 4-point Haad Code transform of the filtered spectrum is used. After the filtering step, the filtered pixels are placed into the original position in the accumulation buffer. After pixel filtering is completed, the accumulated values ​​are normalized by the number of processing groups used for filtering each pixel. Since a padding of one sample is used around the block, the number of processing groups is equal to 4 for each pixel in the block, and the normalization is performed by right shifting on 2 bits.

[0116] Fig. 9 The list process under the Hadcode transform domain filter is shown. As shown, the equivalent filter shape is 3x3 pixels. In some embodiments, all pixels in the block can be processed independently for maximum parallelism. The results of the 2x2 group filtering can be reused for spatial parallel samples. In some embodiments, one 2x2 filter is performed for each new pixel of the block, and the remaining three are used.

[0117] XI. Triangle Prediction Mode (TPM)

[0118] In some embodiments, the triangle prediction unit mode (TPM) is used to perform inter-frame prediction of the CU. Under TPM, the CU is split into two triangle prediction units in the diagonal or anti-diagonal direction. Each triangle prediction unit in the CU uses its own unidirectional motion vector and reference frame for inter-frame prediction. In other words, the CU is split along the straight line that divides the current block. The conversion and quantization process is then applied to the entire CU. In some embodiments, this mode is only applied to the skip and merge modes. In some embodiments, TPM can be extended to split the CU into two prediction units with a straight line, which can be represented by angles and distances. The split line can be indicated by the signaled index and the signaled index is then mapped to the angle and distance. In addition, one or more indexes are signaled to indicate the motion candidates of the two partitions. After predicting each prediction unit, an adaptive weighting process is applied to the edge of the straight line that divides the current block between the two prediction units to derive the final prediction of the entire CU.

[0119] Fig.10Adaptive weighting applied along the diagonal edge between two triangular prediction units of a CU is shown. The first weighting factor group {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8, 4 / 8, 1 / 8} are used for luma and chroma samples, respectively. The second weighting factor group {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} are used for luma and chroma samples, respectively. A weighting factor group is selected based on a comparison of the motion vectors of the two triangular prediction units. The second weighting factor group is used when the reference images of the two triangular prediction units are different from each other or their motion vectors differ by more than 16 pixels. Otherwise, the first weighting factor group is used.

[0120] XII. Mutually Exclusive Groups

[0121] In some embodiments, to simplify hardware implementation complexity, mutually exclusive rules are implemented to limit the cascading of different tools or codec modes described in Sections I to XI. The cascading hardware implementation of tools or codec modes makes the hardware design more complex and results in longer pipeline latency. By implementing mutually exclusive rules, pipeline stages can be made shorter and hardware utilization can be made higher (i.e., less idle hardware). Typically, mutually exclusive rules are used to ensure that tools or codec modes in certain sets of two or more tools or codec modes are not enabled at the same time for encoding and decoding a current CU.

[0122] In some embodiments, a plurality (e.g., four) mutually exclusive groups of tools or codec modes are implemented. A mutually exclusive group may include some or all of the following codec modes or tools: GBI (generic bidirectional prediction), CIIP (combined inter and intra prediction), (BDOF) bidirectional optical flow, DMVR (decoder-side motion vector refinement), and weighted prediction (WP).

[0123] In some embodiments, for any CU, the prediction stage of the video codec (video encoder or video decoder) implements mutually exclusive rules among some prediction tools or codec modes. Mutual exclusion means that only one of these codec tools is started separately for coding and decoding, rather than two codec tools being started for the same CU. In particular, it can define a mutually exclusive group of tools, in which, for any CU, only one tool in the group is started for coding and decoding, and no two tools belonging to the same mutually exclusive group are started for the same CU. In some embodiments, different CUs can have different startup tools.

[0124] For some embodiments, the mutually exclusive group includes GBI, BDOF, DMVR, CIIP, WP. That is, for any CU, among GBI, BDOF, DMVR, CIIP and WP, ​​only one tool in the group is enabled for encoding and decoding, rather than two of them being enabled for the same CU. For example, when the CIIP flag is equal to 1, DMVR / BDOF / GBI is not applied. In other words, when the CIIP flag is equal to 0, DMVR / BDOF / GBI can be applied (if the enabling conditions of DMVR / BDOF / GBI are met). In some embodiments, the mutually exclusive group includes any two or three or some subsets of GBI, BDOF, DMVR, CIIP and WP. In particular, mutually exclusive groups may include BDOF, DMVR, CIIP; mutually exclusive groups may include GBI, DMVR, CIIP; mutually exclusive groups may include GBI, BDOF, CIIP; mutually exclusive groups may include GBI, BDOF, DMVR; mutually exclusive groups may include GBI and BDOF; mutually exclusive groups may include GBI and DMVR; mutually exclusive groups may include GBI and CIIP; mutually exclusive groups may include BDOF, DMVR, CIIP; mutually exclusive groups may include BDOF and CIIP; mutually exclusive groups may include DMVR and CIIP. For example, mutually exclusive groups include GBI, BDOF, CIIP, if CIIP is enabled (ciip_flag is equal to 1), BDOF is off and GBI is off (which means that equal weights are used to mix inter-frame predictions from list 0 and list 1 regardless of the BCW weight index). For example, a mutually exclusive group including GBI, DMVR, CIIP, if CIIP is enabled (ciip_flag is equal to 1), DMVR is off and GBI is off (which means equal weights are used to mix inter predictions from list 0 and list 1 regardless of the BCW weight index). For example, a mutually exclusive group including GBI, DMVR, if GBI is enabled (GBI weight index indicates unequal weights), DMVR is off. For example, a mutually exclusive group including GBI, BDOF, if GBI is enabled (GBI weight index indicates unequal weights), BDOF is off.

[0125] According to mutually exclusive rules, related syntax elements can be saved (or omitted from the bitstream). For example, if a mutually exclusive group includes GBI, BDOF, DMVR, CIIP, then if CIIP mode is not enabled for the current CU (e.g., it is GBI or BDOF or DMVR) because CIIP is turned off by the exclusion rule, the CIIP flag or syntax element can be saved or ignored (rather than sent from the encoder to the decoder) for this CU. For some other embodiments of mutually exclusive groups, related syntax elements can be saved or ignored for tools that are excluded or disabled for a certain CU.

[0126] In some embodiments, the priority is applied to mutually exclusive groups. In some embodiments, each tool in a mutually exclusive group has a certain original or traditional enabling condition. The enabling condition is the original enabling rule used for each tool before mutual exclusion. For example, the enabling conditions of DMVR include true bidirectional prediction and equal POC distance between the current image and L0 image / L1 image, among others; the enabling conditions of GBI include bidirectional prediction and GBI index from syntax (when AMVP) or inherited GBI index (when merge mode).

[0127] A priority rule may be predefined for mutually exclusive groups. Tools or codec modes within a mutually exclusive group have a priority number for each tool. If both tool A and B can be enabled (i.e., their enabling conditions are met) for the same CU, but tool A has a better predefined priority than tool B (defined as tool A>tool B), then if tool A is enabled or enabled, tool B is turned off or disabled.

[0128] Different embodiments have different precedence rules such as mutually exclusive groups including GBI, DMVR, BDOF, CIIP, and WP or any subset of {GBI, DMVR, BDOF, CIIP, WP}. For example, in some embodiments, the precedence rule specifies GBI>DMVR>BDOF>CIIP. In some embodiments, the precedence rule specifies GBI>DMVR>BDOF. In some embodiments, the precedence rule specifies DMVR>GBI>BDOF. In some embodiments, the precedence rule specifies DMVR>GBI. In some embodiments, the precedence rule specifies GBI>BDOF. In some embodiments, the precedence rule specifies GBI>DMVR. In some embodiments, the precedence rule specifies DMVR>GBI>BDOF>CIIP. In some embodiments, the precedence rule specifies DMVR>BDOF>GBI>CIIP. In some embodiments, the precedence rule specifies CIIP>GBI>BDOF. In some embodiments, the precedence rule specifies CIIP>GBI>DMVR. The predetermined rule also specifies any other order in any subset of GBI, BDOF, DMVR, CIIP. For another example, the exclusion group includes {GBI, CIIP} and the priority rule specifies CIIP>GBI, so when CIIP is used (ciip_flag is equal to 1), GBI is turned off (or disabled), which means that equal weights are applied to mix predictors from list 0 and list 1. For another example, the exclusion group includes {DMVR, CIIP} and the priority rule specifies CIIP>DMVR, so when CIIP is used (ciip_flag is equal to 1), DMVR is not used. For another example, the exclusion group includes {BDOF, CIIP} and the priority rule specifies CIIP>BDOF, so when CIIP is used (ciip_flag is equal to 1), BDOF is not used. For another example, the exclusion group includes {BDOF, GBI} and the priority rule specifies GBI>BDOF, so when GBI is used (GBI index indicates unequal weights for mixing predictions from list 0 and list 1), BDOF is not used. As another example, the exclusion group includes {DMVR, GBI} and the priority rule specifies GBI>DMVR, so when GBI is used (the GBI index indicates unequal weights for mixing predictions from list 0 and list 1), DMVR is not used.

[0129] In some embodiments, the priority rules of mutually exclusive groups are not predefined, but are also based on some parameters of the current CU (such as CU size or current MV). For example, for a mutually exclusive group including DMVR and BDOF, there may be exclusion rules based on CU size or other parameters of the CU, and when the enabling conditions of DMVR and BDOF are met, priority is given to DMVR or BDOF. For example, in some embodiments, if the current CU size is greater than a threshold, the priority of DMVR (for tool exclusion) is higher than BDOF. In some embodiments, if the current CU size is greater than a threshold, the priority of BDOF (for tool exclusion) is higher than DMVR.

[0130] In some embodiments, if the current CU aspect ratio is greater than a threshold, DMVR has a higher priority (for tool exclusion) than BDOF. If CU_width>CU_height, the aspect ratio is defined as CU_width / CU_height or if CU_height>=CU_width, the aspect ratio is defined as CU_height / CU_width. In some embodiments, if the current CU aspect ratio is greater than a threshold, BDOF has a higher priority (for tool exclusion) than DMVR. In some embodiments, for some merge mode candidates (if selected for inter-frame prediction), DMVR has a higher priority (for tool exclusion) than BDOF, while for other merge candidates (if selected for inter-frame prediction), BDOF has a higher priority (for tool exclusion) than DMVR.

[0131] In some embodiments, for a true bi-predictive merge candidate, if the mirror image (and subsequently scaled) MV of L0 MV is very similar to L1 MV, then DMVR has a higher priority (for tool exclusion) than BDOF. In some implementations, for a true bi-predictive merge candidate, if the mirror image (and subsequently scaled) MV of L0 MV is very similar to L1 MV, then BDOF has a higher priority (for tool exclusion) than DMVR.

[0132] In some embodiments, mutually exclusive groups may include some or all of the subsequent codec modes or tools: LIC (local illumination compensation), DIF (uniform illumination interframe prediction filter or diffusion filter), BIF (bidirectional filter), HAD filter (Haad code transform domain filter). These tools or codec modes are applied to the residual signal or the prediction signal or the reconstructed signal, that is, they act on the "later stage". The latter stage is defined as the pipeline stage after prediction (intra / inter prediction) or after reference decoding or after both. In some embodiments, mutually exclusive groups may also include front-stage tools or codec modes other than LIC, DIF, BIF and HAD.

[0133] In some embodiments, mutually exclusive groups may include all or a subset of the following 8 codec modes or tools: GBI, BDOF, DMVR, CIIP, LIC, DIF, BIF, HAD. That is, for any CU, only one of them is activated for encoding, rather than two codec modes or tools in GBI, BDOF, DMVR, CIIP, LIC, DIF, BIF, HAD being activated for the same CU. In some embodiments, mutually exclusive modes include LIC, DIF, BIF, HAD. In some embodiments, mutually exclusive groups include DIF, BIF, HAD. In some embodiments, mutually exclusive groups include LIC, BIF, HAD. In some embodiments, mutually exclusive groups include LIC, DIF, HAD. In some embodiments, mutually exclusive groups include LIC, DIF, BIF. In some embodiments, mutually exclusive groups include LIC and DIF. In some embodiments, mutually exclusive groups include LIC, BIF. In some embodiments, mutually exclusive groups include LIC and HAD. In some embodiments, mutually exclusive groups include DIF and HAD. In some embodiments, mutually exclusive groups include BIF and HAD.

[0134] XIII. Published Multi-hypothesis Forecasting Model

[0135] Both CIIP and TPM generate the final prediction of the current CU with two candidates. Either CIIP or TPM can be viewed as a type of multi-hypothesis prediction merge mode, where one hypothesis of the prediction is generated by one candidate and the other hypothesis of the prediction is generated by another candidate. For CIIP, one candidate comes from the intra mode and the other candidate comes from the merge mode. For TPM, the two candidates come from the candidate list of the merge mode.

[0136] In some embodiments, the multi-hypothesis mode is used to improve inter-frame prediction, which is an improvement method for skip and / or merge mode. In the original skip and merge modes, a merge index is used to select a motion candidate, which can be a unidirectional prediction or a bidirectional prediction derived from the merge candidate list by the candidate itself. The generated motion compensation predictor is referred to as the first hypothesis (or first prediction) in some embodiments. In the multi-hypothesis mode, a second hypothesis is generated in addition to the first hypothesis. The second hypothesis of the predictor can be generated by motion compensation from a motion candidate based on an inter-frame prediction mode (merge or skip mode), or by intra-frame prediction based on an intra-frame prediction mode.

[0137] When the second hypothesis (or second prediction) is generated by an intra prediction mode, the multiple hypothesis mode is called intra MH mode or MH mode intra or MH intra or inter-intra mode. A CU encoded and decoded by CIIP is encoded and decoded using the intra MH mode. When the second hypothesis is generated by motion compensation by a motion candidate or an inter prediction mode (such as merge or skip mode), the multiple hypothesis mode is called inter MH mode or MH mode inter or MH inter (or also called merged MH mode or MH merge). The diagonal edge area of ​​a CU encoded and decoded by TPM is encoded and decoded using the inter MH mode.

[0138] For multiple hypothesis mode, each multiple hypothesis candidate (or each candidate with multiple hypothesis) includes one or more candidates (i.e., first hypothesis) and / or one intra prediction mode (i.e., second hypothesis), wherein the motion candidate is selected from candidate list I and / or the intra prediction mode is selected from candidate list II. For intra MH mode, each multiple hypothesis candidate (or each candidate with multiple hypothesis) includes one motion candidate and one intra prediction mode, wherein the motion candidate is selected from candidate list I and the intra prediction mode is fixed to one mode (e.g., plane) or selected from candidate list II. Inter MH mode uses two motion candidates, and at least one of the two motion candidates is derived from candidate list I. In some embodiments, candidate list I is equal to the merge candidate list of the current block and both motion candidates of the multiple hypothesis candidate of the inter MH mode are selected from candidate list I. In some embodiments, candidate list 1 is a subset of the merge candidate list. In some embodiments, for inter MH mode, each of the two motions used to generate the prediction of each prediction unit is indicated by a signaled index. When the index refers to a bidirectional prediction motion candidate in candidate list 1, the motion of list 0 or list 1 is selected according to the index. When the index refers to a unidirectional prediction motion candidate in candidate list 1, the unidirectional prediction motion is used.

[0139] Fig.11aConceptually, encoding or decoding a pixel block by using an intra MH mode is shown. The diagram shows a video image 1100 currently encoded or decoded by a video encoder. The video image 1100 includes a pixel block 1110 currently encoded or decoded as a current block. The current block 1110 is encoded by the intra MH mode, and in particular, a combined prediction 1120 is generated based on a first prediction 1122 (first hypothesis) of the current block 1110 and a second prediction 1124 (second hypothesis) of the current block 1110. The combined prediction 1120 is then used to reconstruct the current block 1110.

[0140] The current block 1110 is encoded and decoded by using an intra MH mode. In particular, a first prediction is obtained by inter prediction based on at least one reference frame 1102 and 1104. A second prediction 1124 is obtained by intra prediction based on neighboring pixels 1106 of the current block 1110. As shown in the figure, the first prediction 1122 is generated based on an inter prediction mode or a motion candidate 1142 selected from a first candidate list 1132 (candidate list I), and the first candidate list 1132 has one or more candidate inter prediction modes. The candidate list I can be a merge candidate list of the current block 1110. The second prediction 1124 is generated based on an intra prediction mode 1144, which is predefined as an intra prediction mode (e.g., plane) or selected from a second candidate list 1134 (candidate list II) having one or more candidate intra prediction modes. If only one intra prediction mode (e.g., plane) is used for intra MH, the intra prediction mode for intra MH is set to an intra prediction mode that does not need to be signaled.

[0141] Fig.11b A current block 1110 is shown that is encoded and decoded using an inter-frame MH mode. In particular, a first prediction 1122 is obtained by inter-frame prediction based on at least one reference frame 1102 and 1104. A second prediction 1124 is obtained by inter-frame prediction based on at least one reference frame 1106 and 1108. As shown, the first prediction 1122 is generated based on an inter-frame prediction mode or motion candidate 1142 (first prediction mode), which is selected from a first candidate list 1132 (candidate list I). The second prediction 1124 is generated based on an inter-frame prediction mode or motion candidate 1146, which is also selected from the first candidate list 1132 (candidate list I). The candidate list 1 can be a merge candidate list for the current block.

[0142] In some embodiments, when intra MH mode is currently supported, in addition to the original syntax of merge mode, a flag is signaled (e.g., to indicate whether intra MH mode is applied). This flag can be represented or indicated by a syntax element in the bitstream. In some embodiments, if the flag is present, an additional intra mode index is signaled to indicate the intra prediction mode from candidate list II. In some embodiments, if the flag is on, the intra prediction mode of intra MH mode (e.g., CIIP, or any intra MH mode) is implicitly selected from candidate list II or implicitly assigned an intra prediction mode without an additional intra mode index. In some embodiments, when the flag is off, inter MH mode (e.g., TPM, or any other inter MH mode with a different prediction unit shape) can be used.

[0143] In some embodiments, the video encoder (video encoder or video decoder) removes all bidirectional prediction use cases in CIIP. That is, only when the current merge candidate is a unidirectional prediction, the video encoder starts CIIP. In some embodiments, the video encoder removes all bidirectional prediction candidates for the CIIP merge candidate. In some embodiments, the video encoder retrieves the L0 information of a bidirectional prediction (merge candidate) and changes it into a unidirectional prediction candidate and is used for CIIP. In some embodiments, the video encoder retrieves the L1 information of a bidirectional prediction (merge candidate) and changes it into a unidirectional prediction candidate of CIIP. By removing all bidirectional prediction behaviors of CIIP, relevant syntax elements can be saved or omitted from transmission.

[0144] In some embodiments, when generating inter-frame prediction of CIIP mode, motion candidates with bidirectional prediction are changed to unidirectional prediction according to a predetermined rule. In some embodiments, based on POC distance, the predetermined rule specifies or selects list 0 or list 1 motion vectors. When the distance (marked as D1) between the current POC (or the POC of the current image) and the POC (of the reference image) referenced by the list x (where x is 0 or 1) motion vector is less than the distance (marked as D2) between the current POC and the POC referenced by the list y (where y is 0 or 1 and y is not equal to x), the list x motion vector is selected to generate inter-frame prediction of CIIP. If D1 is the same as D2 or the difference between D1 and D2 is less than a threshold, the list x (where x is predetermined to be 0 or 1) motion vector is selected to generate inter-frame prediction of CIIP. In some other embodiments, the predetermined rule generally selects the list x motion vector, where x is predetermined to be 0 or 1. In some other embodiments, this bidirectional to unidirectional prediction scheme can be applied to motion compensation to generate the prediction. When the motion information of the currently coded CIIP CU is saved for reference by a subsequent or next CU, the motion information before applying this bidirectional to unidirectional prediction scheme is used. In some embodiments, after generating the merge candidate list of CIIP, this bidirectional to unidirectional prediction scheme is applied. Processes such as motion compensation and / or motion information saving and / or deblocking can use the generated unidirectional prediction motion information.

[0145] In some embodiments, a new candidate list formed by unidirectional prediction motion candidates is constructed for CIIP. In some embodiments, this candidate list can be generated from the merge candidate list for conventional merge mode according to a predetermined rule. For example, when generating a candidate list as in the conventional merge mode, the predetermined rule may specify that bidirectional prediction motion candidates can be ignored. The length of this new candidate list of CIIP may be equal to or less than the conventional merge mode. For another example, a predetermined rule may specify that the candidate list of CIIP reuses the candidate list of TPM or that the candidate list of CIIP can be reused for TPM. The above-mentioned proposed method can be combined with implicit rules or explicit rules. Implicit rules can depend on block width or height or area and explicit rules can signal a flag at CU, CTU, stripe, tile, tile group, SPS, PPS level, etc.

[0146] In some embodiments, CIIP and TPM are classified into a group of combined prediction modes and the syntax of CIIP and TPM is also unified instead of using two separate flags to determine whether to use CIIP and whether to use TPM. The unified scheme is as follows: when the enabling conditions of the group for the combined prediction mode are met (e.g., a unified set of CIIP and TPM enabling conditions, including high-level syntax, size constraints, supported modes, or slice types), CIIP or TPM can be enabled or disabled using a unified syntax. First, a first bin is signaled (or a first flag is signaled using the first bin) to indicate whether the multi-hypothesis prediction mode is applied. Second, if the first bin indicates that the multi-hypothesis prediction mode is applied, a second bin is signaled (or a second flag is signaled using the second bin) to indicate that one of CIIP and TPM is applied. For example, when the first bin (or the first flag) is equal to 0, a non-multi-hypothesis prediction mode such as a regular merge mode is applied, otherwise, a multi-hypothesis prediction mode such as CIIP or TPM is applied. When the first box (or the first flag) indicates that the multi-hypothesis prediction mode is applied (regular_merge_flag is equal to 0), the second flag is signaled. When the second box (or the second flag) is equal to 0, TPM is applied and additional syntax of TPM is required (e.g., the additional syntax of TPM is to indicate two motion candidates of TPM or the segmentation direction of TPM). When the second box (or the second flag) is equal to 1, CIIP is applied and additional syntax of CIIP may be required (e.g., the additional syntax of CIIP is to indicate two candidates of CIIP). Examples of the enabling conditions of this group for the combined prediction mode include (1) high-level syntax CIIP and (2) TPM is enabled.

[0147] XIV.LIC's letter

[0148] In some embodiments, all bidirectional predictions are removed for LIC mode. In some embodiments, LIC is allowed only when the current merge candidate is unidirectional prediction. In some embodiments, the video encoder retrieves a bidirectionally predicted L0 information (candidate), changes the current merge candidate to a unidirectional prediction candidate, and then applies LIC. In some embodiments, the video encoder retrieves a bidirectionally predicted L1 information (candidate), changes it to a unidirectional prediction candidate, and then applies LIC.

[0149] In some embodiments, when generating inter-frame prediction for LIC mode, motion candidates with bidirectional prediction are converted to unidirectional prediction according to a predetermined rule. In some embodiments, the predetermined rule specifies or selects list 0 or list 1 motion vectors based on POC distance. When the distance (labeled as D1) between the current POC and the POC (of the reference image) referenced by the list x (where x is 0 or 1) motion vector is less than the distance (labeled as D2) between the current POC and the POC referenced by the list y (where y is 0 or 1 and y is not equal to x) motion vector, then the list x motion vector is selected for refining the inter-frame prediction by applying LIC. If D1 is the same as D2 or the difference between D1 and D2 is less than a threshold, then the list x (where x is predetermined to be 0 or 1) motion vector is selected for refining the inter-frame prediction by using LIC. In some embodiments, the predetermined rule formulates or selects the list x motion vector, where x is predetermined to be 0 or 1. In some embodiments, this bidirectional to unidirectional prediction scheme can be applied only to motion compensation to generate the prediction. When the motion information of the currently coded LIC CU is saved for reference to a subsequent or subsequent CU, the motion information before applying this bidirectional to unidirectional prediction scheme is used. In some embodiments, after generating a merge candidate list of the LIC, this bidirectional to unidirectional prediction scheme is applied. The generated unidirectional prediction motion information is used by processes such as motion compensation and / or motion information saving.

[0150] In some embodiments, a new candidate list formed of unidirectional prediction motion candidates is constructed for LIC. In some embodiments, a candidate list can be generated from a merge candidate list for regular merge mode according to a predetermined rule. For example, the predetermined rule can ignore bidirectional prediction motion candidates during candidate list generation as in regular merge mode. The length of this new candidate list for LIC can be equal to or less than regular merge mode.

[0151] In some embodiments, in merge mode, the criteria for enabling LIC depends not only on the LIC flag of the merge candidate, but also on the number of neighboring merge candidates using LIC or historical statistics. For example, if the number of candidates in the merge list using LIC is greater than a predetermined threshold, then LIC is enabled for the current block regardless of whether the LIC flag of the merge candidate is on or off. For another example, the history FIFO buffer records the use of LIC mode of the most recently coded block, assuming that the record size in the FIFO buffer is M, if N of the M use the LIC mode, then LIC is enabled for the current block. In addition, this embodiment can also be combined with the mentioned bidirectional to unidirectional prediction scheme for LIC, that is, if the LIC flag of the current block is enabled because the number of neighboring merge candidates using LIC is greater than a threshold or N of the M records in the history FIFO buffer use the LIC mode, and the merge candidate uses bidirectional prediction, then the list x motion vector is selected, where x is predetermined to be 0 or 1.

[0152] All the above combinations can be determined using implicit rules or explicit rules. The implicit rule can depend on block width, height, area, block size aspect ratio, color component or picture type. The explicit rule can be signaled in a flag at CU, CTU, slice, tile, tile group, picture, SPS, PPS level, etc.

[0153] XV. Exemplary Video Encoders

[0154] Fig.12 An exemplary video encoder 1200 that can implement mutually exclusive groups of coding modes or tools is marked. As shown, the video encoder 1200 receives an input video signal from a video source 1205 and encodes the signal into a bitstream 1295. The video encoder 1200 has various components or modules for encoding the signal from the video source 1205, including at least some components selected from a transform module 1210, a quantization module 1211, an inverse quantization module 1214, an inverse transform module 1215, an intra-frame image estimation module 1220, an intra-frame prediction module 1225, a motion compensation module 1230, a motion estimation module 1235, a loop filter 1245, a reconstructed image buffer 1250, an MV buffer 1265, and an MV prediction module 1275 and an entropy encoder 1290. The motion compensation module 1230 and the motion estimation module 1235 are part of the inter-frame prediction module 1240.

[0155] In some embodiments, modules 1210-1290 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device or electronic device. In some embodiments, modules 1210-1290 are modules of hardware circuits implemented by one or more integrated circuits of an electronic device. Although modules 1210-1290 are shown as separate modules, some modules may be combined into a single module.

[0156] The video source 1205 provides a raw video signal representing pixel data for each video frame that is not compressed. A subtractor 1208 calculates the difference between the raw video pixel data of the video source 1205 and the predicted pixel data 1213 from the motion compensation module 1230 or the intra prediction module 1225. The transform module 1210 converts the difference (or residual pixel data or residual signal 1209) into transform coefficients 1216 (e.g., by performing a discrete cosine transform or DCT). The quantization module 1211 quantizes the transform coefficients into quantized data (or quantized coefficients) 1212, which are encoded into a bitstream 1295 by an entropy encoder 1290.

[0157] The inverse quantization module 1214 dequantizes the quantized data (or quantized coefficients) 1212 to obtain transform coefficients, and the inverse transform module 1215 performs an inverse transform on the transform coefficients to generate a reconstructed residual 1219. The reconstructed residual 1219 is added to the predicted pixel data 1213 to generate reconstructed pixel data 1217. In some embodiments, the reconstructed pixel data 1217 is temporarily stored in a linear buffer (not shown) for intra-frame image prediction and spatial MV prediction. The reconstructed pixels are filtered by the loop filter 1245 and stored in the reconstructed image buffer 1250. In some embodiments, the reconstructed image buffer 1250 is an external storage area of ​​the video encoder 1200. In some embodiments, the reconstructed image buffer 1250 is an internal storage area of ​​the video encoder 1200.

[0158] The intra picture estimation module 1220 performs intra prediction based on the reconstructed pixel data 1217 to generate intra prediction data. The intra prediction data is provided to the entropy encoder 1290 to be encoded into a bitstream 1295. The intra prediction data is also used by the intra prediction module 1225 to generate predicted pixel data 1213.

[0159] The motion estimation module 1235 performs inter-frame prediction by generating MVs to refer to pixel data of a previously decoded frame stored in the reconstructed image buffer 1250. These MVs are provided to the motion compensation module 1230 to generate predicted pixel data.

[0160] In addition to encoding the complete actual MV in the bitstream, the video encoder 1200 uses MV prediction to generate a predicted MV, and the difference between the MV used for motion compensation and the predicted MV is encoded as residual motion data and stored in the bitstream 1295 .

[0161] The MV prediction module 1275 generates a predicted MV based on a reference MV that is generated for encoding a previous video frame, i.e., a motion compensated MV for performing motion compensation. The MV prediction module 1275 retrieves the reference MV from the previous video frame from the MV buffer 1265. The video encoder 1200 stores the MV generated for the current video frame in the MV buffer 1265 as a reference MV for generating the predicted MV.

[0162] The MV prediction module 1275 uses the reference MV to create a predicted MV. The predicted MV can be calculated by spatial MV or temporal MV prediction. The difference (residual motion data) between the predicted MV and the motion compensated MV (MC MV) of the current frame is encoded into the bitstream 1295 by the entropy encoder 1290.

[0163] The entropy encoder 1290 encodes various parameters and data into a bitstream 1295 using entropy coding techniques, such as context adaptive binary arithmetic coding and decoding (CABAC) or Huffman coding. The entropy encoder 1290 encodes various header elements, flags and quantized coefficients 1212 and residual motion data as syntax elements into the bitstream 1295. The bitstream 1295 is in turn stored in a storage device or transmitted to a decoder via a communication medium such as a network.

[0164] The loop filter 1245 performs a filtering or smoothing operation on the reconstructed pixel data 1217 to reduce coding artifacts, particularly at the boundaries of pixel blocks. In some embodiments, the filtering operation performed includes sample adaptive offset (SAO). In some embodiments, the filtering operation includes an adaptive loop filter (ALF).

[0165] Fig.13 Portions of the video encoder 1200 that implement mutually exclusive groups of codec modes or tools are labeled. As shown, the video encoder 1200 implements a combined prediction module 1310 that can receive intra-frame prediction values ​​generated by the intra-frame image prediction module 1225. The combined prediction module 1310 can also receive inter-frame prediction values ​​from the motion compensation module 1230 and the second motion compensation module 1330. The combined prediction module 1310 in turn generates predicted pixel data 1213, which can be further filtered by a set of prediction filters 1350.

[0166] MV buffer 1265 provides merge candidates to motion compensation modules 1230 and 1330. MV buffer 1265 also stores motion information and motion direction used to encode the current block for use by subsequent blocks. Merge candidates may be changed, expanded and / or refined by MV refinement module 1365.

[0167] The codec mode (or tool) control module 1300 controls the operations of the intra picture prediction module 1225 , the motion compensation module 1230 , the second motion compensation module 1330 , the MV refinement module 1365 , the combined prediction module 1310 , and the prediction filter 1350 .

[0168] The codec module control 1300 may enable the MV refinement mode 1365 to perform MV refinement operations by searching for refined MVs (e.g., for DMVR) or to adjust the calculated gradients based on the MVs (e.g., for BDOF). The codec mode control module 1300 may enable the intra prediction module 1225 and the motion compensation module 1230 to implement an MH mode intra (or inter-intra) mode (e.g., CIIP). The codec mode control module 1300 may enable the motion compensation module 1230 and the second motion compensation module 1330 to implement an MH mode inter mode (e.g., for diagonal edge regions of TPM). When combining the prediction signals from the intra picture prediction module 1225, the motion compensation module 1230, and / or the second motion compensation module 1330 to implement a codec mode such as CIIP, TPM, GBI, and / or WP, the codec mode control module 1300 may enable the combined prediction module 1310 to adopt different weighting schemes. The codec mode control 1300 may also enable the prediction filter 1350 to apply LIC, DIF, BIF and / or HAD filters on the predicted pixel data or the reconstructed pixel data 1217 .

[0169] The codec mode control module 1300 also determines which codec mode is enabled and / or disabled for encoding the current block. The codec mode control module 1300 then controls the operation of the intra picture prediction module 1225, the motion compensation module 1230, the second motion compensation module 1330, the MV refinement module 1365, the combined prediction module 1310, and the prediction filter 1350 to enable and / or disable a specific codec mode.

[0170] In some embodiments, the codec mode control 1300 enables only a subset (one or more) of a plurality of codec modes from a specific set of two or more codec modes for encoding the current block or CU. This specific set of codec modes includes all or any subset of the following codec modes: CIIP, TPM, BDOF, DMVR, GBI, WP, LIC, DIF, BIF, and HAD. In some embodiments, when a first condition for enabling a first codec mode of the current block is met, the codec mode control 1300 disables a second codec mode of the current block.

[0171] In some embodiments, when the condition for enabling the first codec mode is met and the first codec mode is enabled, the codec mode control 1300 disables all modes in the specific set of codec modes except the first codec mode. In some embodiments, when the first condition for enabling the first codec mode of the current block and the second condition for enabling the second codec mode of the current block are both met and the first codec mode is enabled, the codec mode control 1300 disables the second codec mode. For example, in some embodiments, when the codec mode control 1300 determines that the conditions for enabling GBI and BDOF are both met and the GBI index indicates unequal weights to mix the predictions of list 0 and list 1, the codec mode control 1300 will disable BDOF. For another example, in some embodiments, when the codec mode control 1300 determines that the conditions for enabling GBI and DMVR are both met and the GBI index indicates unequal weights to mix the predictions of list 0 and list 1, the codec mode control 1300 will disable DMVR.

[0172] In some embodiments, the codec mode control 1300 identifies a highest priority codec mode from the one or more codec modes. If the highest priority codec mode is enabled, the codec mode control 1300 then disables all other codec modes in the specific set of codec modes, regardless of whether the enabling conditions of each other codec mode are met. In some embodiments, each codec mode of the specific set of codec modes is assigned a priority according to a priority rule defined based on parameters of the current block (such as the size or aspect ratio of the current block).

[0173] The codec mode control 1300 generates or signals a syntax element 1390 to the entropy encoder 1290 to indicate that one or more codec modes are enabled. The video encoder 1200 may also disable one or more other codec modes in a particular set of codec modes without signaling a syntax element for disabling the one or more other codec modes. In some embodiments, a first syntax element (e.g., a first flag) is used to indicate whether a multi-hypothesis prediction mode is applied and a second syntax element (e.g., a second flag) is used to indicate whether CIIP or TPM is applied. The first and second syntax elements are encoded and decoded by the entropy encoder 1290 into a first bin and a second bin, respectively. In some embodiments, a second bin for deciding between CIIP and TPM is signaled only when the first bin indicates that the multi-hypothesis mode is enabled.

[0174] Fig.14 A process 1400 for implementing a mutually exclusive group of encoding and decoding modes or tools is conceptually illustrated. In some embodiments, one or more processing units (or processors) of a computing device implementing the encoder 1200 perform the process 1400 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing the encoder 1200 performs the process 1400.

[0175] The encoder receives (at block 1410) data for a block of pixels to be encoded as a current block of a current picture of a video.

[0176] The encoder identifies (at block 1430) a highest priority codec mode among the one or more codec modes. In some embodiments, each codec mode of the particular set of codec modes is assigned a priority according to a priority rule defined based on the current block parameters.

[0177] If the highest priority codec mode is enabled, the encoder disables (at block 1140) all other codec modes in the specific set of codec modes. The conditions for enabling various codec modes are described in the above paragraphs related to these codec modes. The conditions for enabling a codec mode may include receiving an explicit syntax element from the bitstream for the codec mode. The conditions for enabling a codec mode may also include having specific characteristics or parameters (e.g., size, aspect ratio) of the current block being encoded. For example, when the specific set of codec modes includes a first codec mode assigned a higher priority and a second codec mode assigned a lower priority, and when the first codec mode is enabled, the encoder disables (at block 1445) the second codec mode of the current block. In some embodiments, when the first codec mode is enabled, the encoder disables all codec modes in the specific set of codec modes except the first codec mode. In some embodiments, if the GBI weight index indicates unequal weights, the encoder enables GBI (which means that unequal weights are used to mix inter predictions from list 0 and list 1), and disables BDOF because GBI is assigned a higher priority than BDOF. For another example, in some embodiments, because GBI is assigned a higher priority than DMVR, if the GBI weight index indicates unequal weights, the encoder enables GBI (which means that unequal weights are used to mix inter predictions from list 0 and list 1), but disables DMVR. For another example, in some embodiments, because CIIP is assigned a higher priority than disabled tools, if the CIIP flag is equal to 1, the encoder enables CIIP, but disables GBI, BDOF, and / or DMVR.

[0178] The encoder encodes (at block 1450) the current block in the bitstream using an inter prediction that is calculated according to the enabled codec mode.

[0179] XVI. Exemplary Video Decoder

[0180] Fig.15An exemplary video decoder 1500 that can implement mutually exclusive groups of codec modes or tools is marked. As shown, the video decoder 1500 is an image decoding or video decoding circuit that receives a bitstream 1595 and decodes the contents of the bitstream into pixel data of a video frame for display. The video decoder 1500 has several components or modules for decoding the bitstream 1595, including some components selected from an inverse quantization module 1505, an inverse transform module 1510, an intra-frame prediction module 1525, a motion compensation module 1530, a loop filter 1545, a decoded image buffer 1550, an MV buffer 1565, an MV prediction module 1575, and a parser 1590. The motion compensation module 1530 is part of the inter-frame prediction module 1540.

[0181] In some embodiments, modules 1510-1590 are modules of software instructions executed by one or more processing units (such as processors) of a computing device. In some embodiments, modules 1510-1590 are modules of hardware circuits implemented by one or more ICs of an electronic device. Although modules 1510-1590 are shown as separate modules, some modules may be combined into a single module.

[0182] The parser 1590 (or entropy decoder) receives the bitstream 1595 and performs initial parsing according to the syntax defined by the video codec or image codec standard. The parsed syntax elements include various header elements, flags, and quantized data (or quantized coefficients) 1512. The parser 1590 parses out the various syntax elements by using entropy coding techniques such as context adaptive arithmetic coding (CABAC) or Huffman coding.

[0183] The inverse quantization module 1505 dequantizes the quantized data (or quantized coefficients) 1512 to obtain transform coefficients, and the inverse transform module 1510 performs an inverse transform on the transform coefficients 1516 to generate a reconstructed residual signal 1519. The reconstructed residual signal 1519 is added to the predicted pixel data 1513 from the intra prediction module 1525 or the motion compensation module 1530 to generate decoded pixel data 1517. The decoded pixel data is filtered by the loop filter 1545 and stored in the decoded picture buffer 1550. In some embodiments, the decoded picture buffer 1550 is an external storage of the video decoder 1550. In some embodiments, the decoder picture buffer 1550 is an internal storage of the video decoder 1550.

[0184] The intra prediction module 1525 receives intra prediction data from the bitstream 1595 and generates predicted pixel data 1513 based thereon from decoded pixel data 1517 stored in the decoded picture buffer 1550. In some embodiments, the decoded pixel data 1517 is also stored in a linear buffer (not shown) for inter picture prediction and spatial MV prediction.

[0185] In some embodiments, the contents of the decoded image buffer 1550 are used for display. The display device 1555 retrieves the contents of the decoded image buffer 1550 for direct display or retrieves the contents of the decoded image buffer to a display buffer. In some embodiments, the display device receives pixel values ​​from the decoded image buffer 1550 via pixel transfer.

[0186] The motion compensation module 1530 generates predicted pixel data 1513 from decoded pixel data 1517 stored in the decoded picture buffer 1550 according to motion compensated MVs (MC MVs). These motion compensated MVs are decoded by adding residual motion data received from the bitstream 1595 to predicted MVs received from the MV prediction module 1575.

[0187] The MV prediction module 1575 generates a predicted MV based on a reference MV, which is generated for decoding a previous video frame, such as a motion compensation MV for performing motion compensation. The MV prediction module 1575 retrieves the reference MV of the previous video frame from the MV buffer 1565. The video decoder 1500 stores the motion compensation MV generated for decoding the current video frame in the MV buffer 1565 as a reference MV for generating the predicted MV.

[0188] The loop filter 1545 performs a filtering or smoothing operation on the decoded pixel data 1517 to reduce encoding and decoding artifacts, especially at the boundaries of pixel blocks. In some embodiments, the filtering operation performed includes sample adaptive offset (SAO). In some embodiments, the filtering operation includes an adaptive loop filter (ALF).

[0189] Fig.16 Portions of the video decoder 1500 that implement mutually exclusive groups of codec modes or tools are labeled. As shown, the video decoder 1500 implements a combined prediction module 1610 that receives intra-frame prediction values ​​generated by the intra picture prediction module 1525. The combined prediction module 1610 may also receive inter-frame prediction values ​​from the motion compensation module 1530 and the second motion compensation module 1630. The combined prediction module 1610 in turn generates predicted pixel data 1513, which may be further filtered by a set of prediction filters 1650.

[0190] The MV buffer provides merge candidates to the motion compensation modules 1530 and 1630. The MV buffer 1565 also stores the motion information and mode direction used to decode the current block for use by subsequent blocks. The merge candidates may be modified, expanded and / or refined by the MV refinement module 1665.

[0191] The codec mode (or tool) control 1600 controls the operations of the intra picture prediction module 1525 , the motion compensation module 1530 , the second motion compensation module 1630 , the MV refinement module 1665 , the combined prediction module 1610 , and the prediction filter 1650 .

[0192] The codec mode control 1600 may enable the MV refinement module 1665 to perform MV refinement (e.g., for DMVR) operations by searching for refined MVs or to adjust calculated gradients based on MVs (e.g., for BDOF). The codec mode control module 1600 may enable the intra prediction module 1525 and the motion compensation module 1530 to implement an MH mode intra (or inter-intra) mode (e.g., CIIP). The codec mode control module 1600 may enable the motion compensation module 1530 and the second motion compensation module 1630 to implement an MH mode inter mode (e.g., for diagonal edge regions of TPM). When combining prediction signals from the intra picture prediction module 1525, the motion compensation module 1530, and / or the second motion compensation module 1630 to implement a codec mode such as CIIP, TPM, GBI, and / or WP, the codec mode control module 1600 may enable the combined prediction module 1610 to adopt different weighting schemes. The codec mode control 1600 may also enable the prediction filter 1650 to apply LIC, DIF, BIF and / or HAD filters to the predicted pixel data 1513 or the decoded pixel data 1517 .

[0193] The codec mode control module 1600 also determines which codec mode is enabled and / or disabled for encoding the current block. The codec mode control module 1600 then controls the operation of the intra picture prediction module 1525, the motion compensation module 1530, the second motion compensation module 1630, the MV refinement module 1665, the combined prediction module 1610, and the prediction filter 1650 to enable and / or disable a specific codec mode.

[0194] In some embodiments, the codec mode control 1600 enables only a subset (one or more) of codec modes from a specific set of two or more codec modes for encoding the current block or CU. This specific set of codec modes may include all or any subset of the following codec modes: CIIP, TPM, BDOF, DMVR, GBI, WP, LIC, DIF, BIF, and HAD. In some embodiments, when a first condition for enabling a codec mode of the current block is met, the codec mode control 1600 disables a second codec mode of the current block.

[0195] In some embodiments, when the condition for enabling the first codec mode is met and the first codec mode is enabled, the codec mode control 1600 disables all codec modes in the specific subset of codec modes except the first codec mode. In some embodiments, when the first condition for enabling the first codec mode of the current block and the second condition for enabling the second codec mode of the current block are both met and the first codec mode is enabled, the codec mode control 1600 disables the second codec mode. For example, in some embodiments, when the codec mode control 1600 determines that the conditions for enabling GBI and BDOF are both met and the GBI index indicates unequal weights to mix the predictions of list 0 and list 1, the codec mode control 1600 will disable BDOF. For another example, in some embodiments, when the codec mode control 1600 determines that the conditions for enabling GBI and DMVR are both met and the GBI index indicates unequal weights to mix the predictions of list 0 and list 1, the codec mode control 1600 will disable DMVR.

[0196] In some embodiments, the codec mode control 1600 identifies a highest priority codec mode from the one or more codec modes. If the highest priority codec mode is enabled, the codec mode control 1600 then disables all other codec modes in the codec mode specific set, regardless of whether the enabling conditions of each other codec mode are met. In some embodiments, each codec mode in the codec mode specific set is assigned a priority according to a priority rule defined based on a parameter of the current block, such as a size or aspect ratio of the current block.

[0197] The codec mode control 1600 receives a syntax element 1690 from the entropy decoder 1590 to indicate that one or more codec modes are enabled. The video decoder 1500 may also disable one or more other codec modes without receiving a syntax element for disabling one or more other codec modes. In some embodiments, a first syntax element (e.g., a first flag) is used to indicate whether a multi-hypothesis prediction mode is applied and a second syntax element (e.g., a second flag) is used to indicate whether a CIIP or TPM mode is applied. The first and second elements are decoded from a first box and a second box in the bitstream 1595, respectively. In some embodiments, the second box for deciding between CIIP and TPM is signaled only when the first box indicates that the multi-hypothesis mode is enabled.

[0198] Fig.17 Conceptually, a process 1700 for implementing a mutually exclusive group of codec modes or tools is shown. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing decoder 1500 perform process 1700 by executing instructions stored on a computer-readable medium. In some embodiments, an electronic device implementing decoder 1500 performs process 1700.

[0199] The decoder receives (at block 1710) data for a block of pixels to be decoded as a current block of a current picture of a video.

[0200] The decoder identifies (at block 1730) a highest priority codec mode among the one or more codec modes. In some embodiments, each codec mode of the specific set of codec modes is assigned a priority according to a priority rule defined based on parameters of the current block.

[0201] If the highest priority codec mode is enabled, the decoder disables (at block 1740) all other codec modes of the codec mode specific set. The conditions for enabling various codec modes are described in the above paragraphs related to these codec modes. The conditions for enabling a codec mode may include receiving explicit syntax elements from the bitstream for the codec mode. The conditions for enabling a codec mode may also include having specific characteristics or parameters (e.g., size, aspect ratio) of the current block being encoded. For example, when the codec mode specific set includes a first codec mode assigned a higher priority and a second codec mode assigned a lower priority, and when the first codec mode is enabled, the decoder disables (at block 1745) the second codec mode of the current block. In some embodiments, when the first codec mode is enabled, the decoder disables all other codec modes in the codec mode specific set except the first codec mode. In some embodiments, because GBI is assigned a higher priority than BDOF, if the GBI weight index indicates unequal weights, the decoder enables GBI (which means that unequal weights are used to mix inter-frame predictors from list 0 and list 1), but disables BDOF. For example, in some embodiments, because GBI is assigned a higher priority than DMVR, if the GBI weight index indicates unequal weights, the decoder enables GBI (which means that unequal weights are used to mix inter-frame predictors from list 0 and list 1), but disables DMVR. For another example, in some embodiments, because CIIP is assigned a higher priority than other disabled tools, if the CIIP flag is equal to 1, the encoder enables CIIP, but disables GBI, BDOF and / or DMVR.

[0202] The decoder decodes (at block 1750) the current block using the inter prediction calculated according to the enabled codec mode.

[0203] XVII. Exemplary Electronic Systems

[0204] Many of the features and applications described above are implemented as software processes designated as a set of instructions recorded on a computer-readable storage medium (also referred to as a computer-readable medium). When these instructions are executed by one or more computing or processing units (e.g., one or more processors, processor cores, or other processing units), they cause the processing units to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard drives, erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), and the like. Computer-readable media include, but are not limited to, carrier waves and electronic signals transmitted wirelessly or through wired connections.

[0205] In this specification, the term "software" is intended to include firmware residing in a read-only memory or an application stored in a magnetic storage, which can be read into storage and processed by a processor. In addition, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while maintaining unique software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Ultimately, any combination of separate programs implements the software inventions described within the scope of the present invention together. In some embodiments, when installed to operate one or more electronic systems, the software program defines one or more specific machine implementations that run and execute the operations of the software program.

[0206] Fig.18 An electronic system 1800 is conceptually shown with which some embodiments of the present invention are implemented. The electronic system 1800 may be a computer (e.g., a desktop computer, a personal computer, a tablet computer, etc.), a phone, a PDA, or any other suitable electronic device. Such an electronic system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 1800 includes a bus 1805, a processing unit 1810, a graphics processing unit (GPU) 1815, a system memory 1820, a network 1825, a read-only memory 1830, a permanent storage device 1835, an input device 1840, and an output device 1845.

[0207] Bus 1805 collectively represents all system, peripheral, and chipset buses that communicatively connect the various internal devices of electronic system 1800. For example, bus 1805 communicatively connects processing unit 1810 with GPU 1815, read-only memory 1830, system memory 1820, and permanent storage 1835.

[0208] From these various memory units, the processing unit 1810 retrieves instructions to be executed and data to be processed to perform the processes of the present invention. The processing unit can be a single processor or a multi-core processor in different embodiments. Some embodiments are transmitted to be executed by the GPU 1815. The GPU 1815 can offload various calculations provided by the processing unit 1810 or perform image processing.

[0209] Read-only memory (ROM) 830 stores data and instructions used by processing unit 1810 and other modules of the electronic system. On the other hand, permanent storage device 1835 is a read-write storage device. This device is a non-volatile memory that stores instructions and data even when electronic system 1800 is turned off. Some embodiments of the present invention use a mass storage device (such as a magnetic or optical disk and its corresponding hard disk drive) as permanent storage device 1835.

[0210] Other embodiments use removable storage devices (such as floppy disks, flash storage devices, etc. and their corresponding hard disk drives) as permanent storage devices. Like permanent storage device 1835, system memory 1820 is a read-write storage device. However, unlike storage device 1835, system memory 1820 is a volatile read-write memory, such as random access memory. System memory 1820 stores some instructions and data used by the processor at runtime. In some embodiments, processes according to the present invention are stored in system memory 1820, permanent storage device 1835 and / or read-only memory 1830. For example, various storage units include instructions for processing multimedia videos according to some embodiments. From these various storage units, processing unit 1810 retrieves instructions to be executed and data to be processed to perform the processing of some embodiments.

[0211] The bus 1805 is also connected to input and output devices 1840 and 1845. The input device 1840 enables a user to communicate information and select commands with the electronic system. The input device 1840 includes an alphabetic keyboard and a pointing device (also called a "cursor control device"), a camera (e.g., a webcam), a microphone or a similar device for receiving voice commands, etc. The output device 1845 displays images or other output data generated by the electronic system. The output device 1845 includes a printer and a display device, such as a cathode ray tube (CRT) or a liquid crystal display (LCD) and a speaker or similar sound output device. Some embodiments include devices such as a touch screen that are both input and output devices.

[0212] Finally, if Fig.18 As shown, bus 1805 also couples electronic system 1800 to a network 1825 via a network adapter (not shown). In this manner, the computer may be part of a computer network (such as a local area network ("LAN"), a wide area network ("WAN"), or an intranet, or a network of networks, such as the Internet). Any or all components of electronic system 1800 may be used in conjunction with the present invention.

[0213] Some embodiments include electronic components, such as microprocessors, storage of computer program instructions in the form of machine-readable or computer-readable media (or computer-readable storage media, machine-readable storage media, or machine-readable storage media), and memory. Some examples of such computer-readable media include RAM, ROM, compact disk read-only (CD-ROM), recordable compact disk (CD-R), rewritable compact disk (CD-RW), read-only digital versatile disk (e.g., DVD-ROM, double-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini SD card, micro mini SD card, etc.), magnetic and / or solid-state hard drives, read-only and recordable Blu-ray discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. Computer-readable media can store computer programs executed by at least one processing unit and a collection of instructions for performing various operations. Examples of computer programs or computer code include machine code (e.g., generated by a compiler) and files including high-level code executed by a computer, electronic component, or microprocessor using an annotator.

[0214] Although the above description refers primarily to microprocessors or multi-core processors executing software, many of the features and applications described above are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions stored in the circuits themselves. In addition, some embodiments execute software in programmable logic devices (PLDs), ROM, or RAM devices.

[0215] As used in this specification and any claims herein, the terms "computer," "server," "processor," and "memory" refer to electronic or other technological devices. These terms exclude persons or groups of persons. For purposes of this description, the terms display or displaying mean displaying on an electronic device. As used in this specification and any claims herein, the terms "computer-readable medium," "computer-readable media," and "machine-readable medium" are limited to tangible, physical objects that store information in a computer-readable form. These terms exclude any wireless signals, wired download signals, and any other transient signals.

[0216] Although the present invention has been described with reference to various specific details, those skilled in the art will recognize that the present invention may be implemented in other specific forms without departing from the spirit of the present invention. In addition, various diagrams (including Figures 14 and 17) conceptually illustrate processes. The specific operations of these processes may be performed in the exact order shown and described. Specific operations may not be performed as a continuous series of operations, and different specific operations may be performed in different embodiments. In addition, processes may be implemented using various sub-processes or as part of a larger macro process. Therefore, those skilled in the art will understand that the present invention is not limited by the foregoing illustrative details, but is defined by the scope of the attached patent application.

[0217] Notes

[0218] The subject matter described herein sometimes shows different components included in or connected to different other components. It can be understood that the architecture described is only an example, and in fact many other architectures that realize the same function can be implemented. Conceptually, any arrangement of components that realize the same function is effectively "associated" so as to realize the desired function. Therefore, any two components combined herein to realize a specific function can be regarded as "associated" to each other so as to realize the desired function, regardless of the architecture or intermediate components. Similarly, any two components so associated can also be regarded as "operably connected" or "operably coupled" to each other to realize the desired function, and any two components that can be so associated can also be regarded as "operably coupled" to each other to realize the desired function. Specific examples of operably coupled include but are not limited to physically matchable and / or physically interactive components and / or wirelessly understandable and / or wirelessly interactive components and / or logically interactive and / or logically interactive components.

[0219] In addition, with respect to the use of substantially any plural and / or singular terms herein, those having ordinary knowledge in the art may appropriately convert from the plural to the singular and / or from the singular to the plural, depending on the context and application. For the sake of clarity, various singular / plural permutations may be explicitly set forth herein.

[0220] In addition, those skilled in the art will understand that, in general, the terms used herein, especially the terms used in the appended claims (such as the body of the appended claims), are generally intended to be "open" terms, such as, the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "having at least", the term "includes" should be interpreted as "including but not limited to", etc. Those skilled in the art will further understand that if a specific number of cited claims is intended, such intention will be explicitly listed in the claims, and that such intention does not exist in the absence of such a statement. For example, to aid understanding, the subsequent appended claims may include the use of the introductory phrases "at least one" and "one or more" to introduce the claim statements. However, the use of such phrases should not be construed as implying that a claim statement introduced by the indefinite article "a" or "an" limits any particular claim statement containing such introduced claim statement to embodiments containing only one such statement, even when the same claim statement includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an", "a" and / or "an" should be construed to mean "at least one" or "one or more", and the same applies to the definite article introducing the claim statement. Furthermore, even if a specific number of introduced claim statements is explicitly recited, one of ordinary skill in the art will recognize that such a statement should be construed to mean at least one of the recited number, such as the mere statement "two statements" without other modifications means at least two statements, or two or more statements. Furthermore, where a convention similar to “at least one A, B, and C, etc.” is used, generally such construction is intended to be understood by one of ordinary skill in the art, such as “a system has at least one A, B, and C” would include but is not limited to a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. In such cases where a convention similar to “at least one A, B, or C” is used, generally such construction is intended to be understood by one of ordinary skill in the art, such as “a system has at least one A, B, or C” would include but is not limited to a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. Those skilled in the art will further appreciate that, in fact, any separator and / or phrase in the description, claims, or figures indicating two or more alternative terms will be understood to contemplate the possibility of including one, either, or both of the terms. For example, the phrase "A or B" will be understood to include the possibilities of "A or B" or "A and B."

[0221] It can be understood from the above that various embodiments of the present invention have been described herein for illustrative purposes, and various modifications may be made without departing from the scope and spirit of the present invention. Therefore, the various embodiments described herein are not intended to be limited, and the true scope and spirit are indicated by the scope of subsequent patent applications.

Claims

1. A video decoding method, include: receiving data of a block of pixels to be decoded as a current block of a current image of a video; When the first codec mode of the current block is enabled, disabling the second codec mode of the current block, wherein the first codec mode and the second codec mode specify different methods for calculating inter-frame prediction of the current block, wherein a specific set of two or more codec modes includes the first codec mode and the second codec mode, the specific set of codec modes is a subset of generalized bidirectional prediction, decoder-side motion vector refinement, bidirectional optical flow, weighted prediction, and combined inter-frame and intra-frame prediction, and the number of codec modes included in the specific set is two or more; as well as The current block is decoded by using the inter prediction calculated according to the enabled codec mode.

2. The video decoding method according to claim 1, It is characterized in that When the first codec mode is enabled, all other codec modes in the codec mode specific set are disabled except the first codec mode.

3. The video decoding method according to claim 2, It is characterized in that The specific set of codec modes includes generalized bi-directional prediction, decoder-side motion vector refinement, and combined inter and intra prediction, and wherein: Generalized bidirectional prediction is a coding mode in which the video decoder performs a weighted average of two prediction signals in two different directions to generate the inter-frame prediction. Decoder-side motion vector refinement is a codec mode in which the video decoder searches for refined motion vectors around the initial motion vector and uses the refined motion vectors to generate the inter-frame prediction, and Combined inter-frame and intra-frame prediction is a coding mode in which the video decoder combines an inter-frame prediction signal with an intra-frame prediction signal to generate the inter-frame prediction.

4. The video decoding method according to claim 2, It is characterized in that The specific set of codec modes includes generalized bidirectional prediction, bidirectional optical flow, and combined inter- and intra-prediction, and where: Generalized bidirectional prediction is a coding mode in which the video decoder performs a weighted average of two prediction signals in two different directions to generate the inter-frame prediction. Bidirectional optical flow is a codec mode in which the video decoder computes motion refinements to minimize distortion between prediction samples in different directions and adjusts the inter-prediction based on the computed refinements, and Combined inter prediction and intra prediction is a coding mode in which the video decoder combines an inter prediction signal with an intra prediction signal to generate the inter prediction.

5. The video decoding method according to claim 1, It is characterized in that The first codec mode is combined inter and intra prediction and the second codec mode is generalized bi-directional prediction.

6. The video decoding method according to claim 1, It is characterized in that The first coding mode is generalized bidirectional prediction and the second coding mode is bidirectional optical flow.

7. The video decoding method according to claim 1, It is characterized in that The first coding mode is generalized bi-directional prediction and the second coding mode is decoder-side motion vector refinement.

8. The video decoding method according to claim 1, It is characterized in that The first codec mode is combined inter and intra prediction and the second codec mode is bidirectional optical flow.

9. The video decoding method according to claim 1, It is characterized in that The first codec mode is combined inter and intra prediction and the second codec mode is decoder side motion vector refinement.

10. An electronic device, include: A video decoder circuit for performing operations including: receiving data of a block of pixels to be decoded as a current block of a current image of a video; When a first codec mode of the current block is enabled, disabling a second codec mode of the current block, wherein the first codec mode and the second codec mode specify different methods for calculating inter prediction of the current block, wherein a specific set of two or more codec modes includes the first codec mode and the second codec mode, the specific set of codec modes is a subset of generalized bidirectional prediction, decoder-side motion vector refinement, bidirectional optical flow, weighted prediction, and combined inter and intra prediction, and the number of codec modes included in the specific set is two or more; and The current block is decoded by using the inter prediction calculated according to the enabled codec mode.

11. The electronic device according to claim 10, It is characterized in that A specific set of two or more codec modes includes the first codec mode and the second codec mode, and wherein when the first codec mode is enabled, all other codec modes in the specific set of codec modes are disabled except the first codec mode.

12. The electronic device according to claim 10, It is characterized in that The specific set of codec modes includes generalized bi-directional prediction, decoder-side motion vector refinement, and combined inter and intra prediction, and wherein: Generalized bidirectional prediction is a coding mode in which the video decoder circuit performs a weighted average of two prediction signals in two different directions to generate the inter-frame prediction. Decoder-side motion vector refinement is the video decoder circuit searching for refined motion vectors around the initial motion vector and using the refined motion vectors to generate the inter-frame prediction codec mode, and Combined inter-frame and intra-frame prediction is a codec mode in which the video decoder circuit combines an inter-frame prediction signal with an intra-frame prediction signal to generate the inter-frame prediction.

13. The electronic device according to claim 10, It is characterized in that The specific set of codec modes includes generalized bidirectional prediction, bidirectional optical flow, and combined inter and intra prediction, and among them: Generalized bidirectional prediction is a coding mode in which the video decoder circuit performs a weighted average of two prediction signals in two different directions to generate the inter-frame prediction. Bidirectional optical flow is the video decoder circuit calculating motion refinement to minimize the distortion between prediction samples in different directions and adjusting the inter-frame prediction codec mode based on the calculated motion refinement, and Combined inter-frame and intra-frame prediction is a codec mode in which the video decoder circuit combines an inter-frame prediction signal with an intra-frame prediction signal to generate the inter-frame prediction.

Citation Information

Patent Citations

  • Systems and methods of determining illumination compensation status for video coding

    CN107710764A

  • Constraining motion vector information derived by decoder-side motion vector derivation

    US20180278949A1

  • Decoder-side motion vector derivation

    WO2018175756A1

  • A memory-bandwidth-efficient design for BI-directional optical flow (BIO)

    WO2018237303A1