Methods and apparatus for multiple transform set improvements in video coding systems

By deriving inter MTS sets using Virtual Intra Prediction Modes based on Histogram of Gradients, the VVC standard achieves improved coding efficiency and reduced complexity in video coding systems.

WO2026098425A1PCT designated stage Publication Date: 2026-05-15MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MEDIATEK INC
Filing Date
2025-11-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

The existing Versatile Video Coding (VVC) standard lacks efficient methods for determining Multiple Transform Sets (MTS) in inter prediction, leading to suboptimal coding efficiency and increased computational complexity.

Method used

Derive inter MTS sets based on Virtual Intra Prediction Modes (VIPMs) using Histogram of Gradients (HoG) from predicted samples, applying forward or inverse transforms according to selected kernels for improved transform selection.

Benefits of technology

Enhances coding efficiency and reduces computational complexity by optimizing transform selection in inter prediction, leading to better video quality and encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025132305_15052026_PF_FP_ABST
    Figure CN2025132305_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus of deriving inter MTS sets for video coding are disclosed. According to this method, one or more VIPMs (Virtual Intra Prediction Modes) are derived based on HoG (Histogram of Gradients) calculated from predicted samples of the current block or neighbouring reconstructed samples. One or more target inter MTS (Multiple Transform Selection) sets are derived according to said one or more VIPMs. Forward transform is applied to the residual data according to a selected transform kernel of said one or more target inter MTS sets at the encoder side to generate transformed data at the encoder side, or inverse transform is applied to the coded transformed residual data according to the selected transform kernel of said one or more target inter MTS sets to derive reconstructed residual data. The transformed data at the encoder side or the reconstructed residual data at the decoder side is provided.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND APPARATUS FOR MULTIPLE TRANSFORM SET IMPROVEMENTS IN VIDEO CODING SYSTEMSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 716,765, filed on November 6, 2024. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates to video coding system. In particular, the present invention relates to deriving inter MTS (Multiple Transform Selection) sets according to VIPMs (Virtual Intra Prediction Modes) derived based on HoG (Histogram of Gradients) calculated from predicted samples of the current block or neighbouring reconstructed samples. BACKGROUND AND RELATED ART

[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology-Coded representation of immersive media-Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and to handle various types of video sources including 3-dimensional (3D) video signals.

[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter-Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.

[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.

[0006] The decoder, as shown in Fig. 1B, can use some of the functional blocks as the encoder. For example, the decoder can reuse Inverse Quantization 124 and Inverse Transform 126; however, Transform 118 and Quantization 120 are not needed at the decoder. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.

[0007] According to VVC, an input picture is partitioned into non-overlapped square block regions referred as CTUs (Coding Tree Units) , similar to HEVC. Each CTU can be partitioned into one or multiple smaller size coding units (CUs) . The resulting CU partitions can be in square or rectangular shapes. In addition, VVC divides a CTU into prediction units (PUs) as a unit to apply prediction process, such as Inter prediction, Intra prediction, etc.

[0008] 1. Multiple Transform Set (MTS)

[0009] The MTS block in VVC involves three transform types including The Discrete Cosine Transform (DCT) -II, DCT-VIII and Discrete Sine Transform (DST) -VII. For Luma blocks of size lower than 64, the MTS scheme selects a set of transforms that minimizes the rate distortion cost among five transform sets and the skip configuration. However, only DCT-II is considered for chroma components and luma blocks of size 64. The sps_mts_enabled_flag flag defined in the Sequence Parameter Set (SPS) enables activation of the MTS at the encoder side. Two other flags are defined at the SPS level to indicate whether implicit or explicit MTS signalling is used for Intra and Inter coded blocks, respectively. For the explicit signalling, used by default in the reference software, a syntax element signals the selected horizontal and vertical transforms. To reduce the computational cost of large block–size transforms, the effective height M and width N of the coding block (CB) are reduced depending on the CB size and transform type. The sample values beyond the limits of the effective N and M are considered to be zero, thus reducing the computational cost of the 64-pt DCT-II and 32-pt DCT-VIII / DST-VII transforms. This concept is called zeroing in the VVC specification.

[0010] 2. Enhanced MTS for Intra Coding

[0011] The current VVC design for MTS utilizes only DST-7 and DCT-8 transform kernels for intra and inter coding.

[0012] Additional primary transforms including DCT5, DST4, DST1, and identity transform (IDT) are employed. In addition, MTS set is made dependent on the TU size and intra mode information. For blocks predicted via IntraTMP, DIMD process is used on the prediction block to derive an intra mode that is used for transform selection. Specifically, a horizontal gradient and a vertical gradient are calculated for each predicted sample to build a HoG (Histogram of Gradients) . Then the intra prediction mode with the highest histogram amplitude values is used to select the MTS transform set.

[0013] Overall, 16 different TU sizes are considered, and for each TU size, 5 different classes are considered depending on intra-mode information. For each class, 1, 4 or 6 different transform pairs are considered. Number of intra MTS candidates are adaptively selected (between 1, 4 and 6 MTS candidates) depending on the sum of absolute value of transform coefficients. The sum is compared against the two fixed thresholds to determine the total number of allowed MTS candidates: 1 candidate: sum <=th0 4 candidates: th0 < sum <=th1 6 candidates: sum > th1.

[0014] Note, although 80 different classes are considered, some of those different classes often share exactly same transform set. Therefore, there are 58 (less than 80) unique entries in the resultant LUT.

[0015] For angular modes, joint symmetry over TU shape and intra prediction is considered. Therefore, mode i (i > 34) with TU shape A×B will be mapped to the same class corresponding to mode j=(68–i) with TU shape B×A. However, for each transform pair, the order of the horizontal and vertical transform kernel is swapped. For example, for a 16x4 block with mode 18 (horizontal prediction) and a 4x16 block with mode 50 (vertical prediction) are mapped to the same class. However, the vertical and horizontal transform kernels are swapped. For the wide-angle modes the nearest conventional angular mode is used for the transform set determination. For example, mode 2 is used for all the modes between -2 and -14. Similarly, mode 66 is used for mode 67 to mode 80.

[0016] 3, Inter MTS Optimization

[0017] For the MTS of inter-coded CUs, four candidates: { (DST7, DST7) , (DST7, DCT8) , (DCT8, DST7) , (DCT8, DCT8) } are used for each CU. For higher resolution sequences (width > 1080) , maximum CU size for Inter-MTS usage is set to 32 (i.e., Inter-MTS is used for CU with width <=32 and height <=32) . For remaining sequences (smaller resolution) , it is set to 16. For 4-pt, 8-pt and 16-pt transforms, the current AMT transform cores, i.e., DST-7 and DCT-8, are replaced with separable KLTs, as proposed in JVET-J0021.

[0018] 4. Low-Frequency Non-Separable Transform (LFNST)

[0019] In VVC, LFNST is applied between forward primary transform and quantization (at encoder) and between de-quantization and inverse primary transform (at decoder side) as shown in Fig. 2. As shown in Fig. 2, after Forward Primary Transform 210, Forward Low-Frequency Non-Separable Transform LFNST 220 is applied to top-left region 222 of the Forward Primary Transform output, for example, 16 coefficients for 4x4 forward LFNST and / or 64 coefficients for 8x8 forward LFNST. In LFNST, 4x4 non-separable transform or 8x8 non-separable transform is applied according to block size. For example, 4x4 LFNST is applied for small blocks (i.e., min (width, height) < 8) and 8x8 LFNST is applied for larger blocks (i.e., min (width, height) > 4) . After LFNST, the transform coefficients are quantized by Quantization 230. To reconstruct the input signal, the quantized transform coefficients are de-quantized using De-Quantization 240 to obtain the de-quantized transform coefficients. Inverse LFNST 250 is applied to the top-left region 252 (8 coefficients for 4x4 inverse LFNST or 16 coefficients for 8x8 inverse LFNST) . After invers LFNST, inverse Primary Transform 260 is applied to recover the input signal.

[0020] Application of a non-separable transform, which is being used in LFNST, is described as follows using input as an example. To apply 4x4 LFNST, the 4x4 input block X, is first represented as a vector

[0021] The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16x16 transform matrix. The 16x1 coefficient vector is subsequently re-organized as 4x4 block using the scanning order for that block (horizontal, vertical or diagonal) . The coefficients with smaller index will be placed with the smaller scanning index in the 4x4 coefficient block.

[0022] 4.1 Reduced non-separable transform

[0023] LFNST is based on direct matrix multiplication approach to apply non-separable transform so that it is implemented in a single pass without multiple iterations. However, the non-separable transform matrix dimension needs to be reduced to minimize computational complexity and memory space to store the transform coefficients. Hence, reduced non-separable transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N (N is commonly equal to 64 for 8x8 NSST) dimensional vector to an R dimensional vector in a different space, where N / R (R < N) is the reduction factor. Hence, instead of NxN matrix, RST matrix becomes an R×N matrix as follows: where the R rows of the transform are R bases of the N dimensional space.

[0024] The inverse transform matrix for RT is the transpose of its forward transform. For 8x8 LFNST, a reduction factor of 4 is applied, and 64x64 direct matrix, which is conventional 8x8 non-separable transform matrix size, is reduced to16x48 direct matrix. Hence, the 48×16 inverse RST matrix is used at the decoder side to generate core (primary) transform coefficients in 8×8 top-left regions. When16x48 matrices are applied instead of 16x64 with the same transform set configuration, each of which takes 48 input data from three 4x4 blocks in a top-left 8x8 block excluding right-bottom 4x4 block. With the help of the reduced dimension, memory usage for storing all LFNST matrices is reduced from 10KB to 8KB with reasonable performance drop. In order to reduce the complexity, LFNST is restricted to be applicable only if all coefficients outside the first coefficient sub-group are non-significant. Hence, all primary-only transform coefficients have to be zero when LFNST is applied. This allows conditioning the LFNST index signalling on the last-significant position, and hence avoids the extra coefficient scanning in the current LFNST design, which is needed for checking for significant coefficients at specific positions only.

[0025] The worst-case handling of LFNST (in terms of multiplications per pixel) restricts the non-separable transforms for 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In those cases, the last-significant scan position has to be less than 8 when LFNST is applied, for other sizes less than 16. For blocks with a shape of 4xN and Nx4 and N > 8, the proposed restriction implies that the LFNST is now applied only once, and to the top-left 4x4 region only. As all primary-only coefficients are zero when LFNST is applied, the number of operations needed for the primary transforms is reduced in such cases. From encoder perspective, the quantization of coefficients is remarkably simplified when LFNST transforms are tested. A rate-distortion optimized quantization has to be done at maximum for the first 16 coefficients (in scan order) , the remaining coefficients are enforced to be zero.

[0026] 4.2 LFNST transform selection

[0027] There are 4 transform sets and 2 non-separable transform matrices (kernels) in total per transform set are used in LFNST. The mapping from the intra prediction mode to the transform set is pre-defined as shown in Table 1. If one of three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM or INTRA_L_CCLM) is used for the current block (i.e., 81 <=predModeIntra <=83) , transform set 0 is selected for the current chroma block. For each transform set, the selected non-separable secondary transform candidate is further specified by the explicitly signalled LFNST index. The index is signalled in a bit-stream once per Intra CU after transform coefficients. Table 1. Transform selection table

[0028] 4.3 LFNST index signalling and interaction with other tools

[0029] Since LFNST is restricted to be applicable only if all coefficients outside the first coefficient sub-group are non-significant, LFNST index coding depends on the position of the last significant coefficient. In addition, the LFNST index is context coded, but does not depend on intra prediction mode, and only the first bin is context coded. Furthermore, LFNST is applied for intra CU in both intra and inter slices, and for both luma and chroma. If a dual tree is enabled, LFNST indices for luma and chroma are signalled separately. For inter slice (the dual tree is disabled) , a single LFNST index is signalled and used for both luma and chroma.

[0030] Considering that a large CU greater than 64x64 is implicitly split (TU tiling) due to the existing maximum transform size restriction (64x64) , an LFNST index search could increase data buffering by four times for a certain number of decode pipeline stages. Therefore, the maximum size that LFNST is allowed is restricted to 64x64. Note that LFNST is enabled with DCT2 only. The LFNST index signalling is placed before MTS (Multiple Transform Selection) index signalling.

[0031] The use of scaling matrices for perceptual quantization is not evident that the scaling matrices that are specified for the primary matrices may be useful for LFNST coefficients. Hence, the use of the scaling matrices for LFNST coefficients is not allowed. For single-tree partition mode, chroma LFNST is not applied.

[0032] 5. Secondary Transformation: LFNST Extension with Large Kernel

[0033] The LFNST design in VVC is extended as follows: ·The number of LFNST sets (S) and candidates (C) are extended to S=35 and C=3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: οFor predModeIntra < 2, lfnstTrSetIdx is equal to 2 οlfnstTrSetIdx =predModeIntra, for predModeIntra in [0, 34]  οlfnstTrSetIdx =68 –predModeIntra, for predModeIntra in [35, 66]  ·Three different kernels, LFNST4, LFNST8, and LFNST16, are defined to indicate LFNST kernel sets, which are applied to 4xN / Nx4 (N≥4) , 8xN / Nx8 (N≥8) , and MxN (M, N≥16) , respectively.

[0034] The kernel dimensions are specified by: (LFSNT4, LFNST8*, LFNST16*) =(16x16, 32x64, 32x96) .

[0035] The forward LFNST is applied to top-left low frequency region, which is called Region-Of-Interest (ROI) . When LFNST is applied, primary-transformed coefficients that exist in the region other than ROI are zeroed out, which is not changed from the VVC standard.

[0036] The ROI for LFNST16 is depicted in Fig. 3A. It consists of six 4x4 sub-blocks, which are consecutive in scan order. Since the number of input samples is 96, transform matrix for forward LFNST16 can be Rx96. R is chosen to be 32 in this case, 32 coefficients (two 4x4 sub-blocks) are generated from forward LFNST16 accordingly, which are placed following the coefficient scan order.

[0037] The ROI for LFNST8 is shown in Fig. 3B. The forward LFNST8 matrix can be Rx64 and R is chosen to be 32. The generated coefficients are located in the same manner as with LFNST16.

[0038] The mapping from intra prediction modes to these sets is shown in Table 2. Table 2. Mapping of intra prediction modes to LFNST set index

[0039] 6. Non-Separable Primary Transform (NSPT) for Intra Coding

[0040] The separable DCT-II plus LFNST transform combinations are replaced with NSPT for the block shapes 4x4, 4x8, 8x4 and 8x8, 4x16, 16x4, 8x16 and 16x8.

[0041] The affected block sizes are summarized in Fig. 4.

[0042] All NSPTs consist of 35 sets and 3 candidates (similar to the current LFNST) . The kernels of NSPTs have the following shapes: ·NSPT4x4: 16x16 ·NSPT4x8 / NSPT8x4: 32x20 ·NSPT8x8: 64x32 ·NSPT4x16 / NSPT16x4: 64x24 ·NSPT8x16 / NSPT16x8: 128x40 ·NSPT4x32 / NSPT32x4: 128x20 ·NSPT8x32 / NSPT32x8: 256x24.

[0043] Therefore, 12, 32, 40 and 88 coefficients are zeroed-out using NSPT4x8 / NSPT8x4, NSPT8x8, NSPT4x16 / NSPT16x4 and NSPT8x16 / NSPT16x8 respectively. For NSPT4x32 / NSPT32x4 and NSPT8x32 / NSPT32x8, remaining 108 and 232 positions in each transform block are zeroed-out, respectively.

[0044] 7. LFNST / NSPT for Inter Coding

[0045] For inter coded block, an intra prediction mode is first derived according to the inter prediction block with a DIMD-like process applied to the prediction. Then the derived intra prediction mode is used to select an LFNST / NSPT transform set and the transform is processed with the selected LFNST / NSPT kernel, like in the intra coding process.

[0046] The signalling of inter LFNST / NSPT index is different from that of intra LFNST / NSPT index. The intra LFNST / NSPT index binarization employs two context coded bins for each symbol, while truncated unary code with different context models for inter LFNST / NSPT index coding is used. LFNST / NSPT index signalling is prohibited in the case of SBT, of which index value is inferred as 0.

[0047] Encoding fast algorithm, such as inter MTS, can be applied to reduce the encoding time.

[0048] 8. Multiple Transform Set Selection (MTSS) for LFNST / NSPT Proposed in JVET-AH0056

[0049] The first candidate transform set remains the same as the current ECM.

[0050] To derive the second candidate: -For TIMD, if fusion is applied, the second TIMD IPM is first considered; -For SGPM, the 2 IPMs that SGPM uses are first considered; -For MIP and EIP, PLANAR is first considered;

[0051] The first and second candidate should be different transform sets. If the second candidate is not yet decided by the above process, it is derived by the HoG of neighbouring reconstructed pixels.

[0052] To reduce the encoding time, CU size restriction is used for MTSS. The CU size is measured with the number of pixels within the CU. MTSS can be applied to CUs with size equal or larger than the threshold.

[0053] 9. Multiple Transform Set Selection for Intra MTS Proposed in JVET-AH0307

[0054] An alternative transform set is used in explicit intra MTS for DIMD, TIMD, intraTMP and EIP modes. For these modes, the initial transform set remains unchanged, and the alternative transform set corresponds to the planar transform set in the lookup-table.

[0055] The number of candidates in the alternative transform set depends on the number of candidates in the initial transform set. -If 1 candidate is used in the initial transform set, 1 candidate is used in the alternative transform set. -If 4 candidates are used in the initial transform set, 2 candidates are used in the alternative transform set. -If 6 candidates are used in the initial transform set, 3 candidates are used in the alternative transform set.

[0056] The usage of the alternative transform set is signalled by a CU-level flag to the decoder. If the alternative transform set is selected by the encoder, the flag is set to 1.

[0057] In the present invention, inter MTS (Multiple Transform Selection) sets are determined according to VIPMs (Virtual Intra Prediction Modes) derived based on HoG (Histogram of Gradients) calculated from predicted samples of the current block or neighbouring reconstructed samples. BRIEF SUMMARY OF THE INVENTION

[0058] A method and apparatus of deriving inter MTS sets for video coding are disclosed. According to this method, input data is received, wherein the input data comprises residual data for a current block at an encoder side or coded transformed residual data for the current block at a decoder side, wherein the current block is coded in inter prediction. One or more VIPMs (Virtual Intra Prediction Modes) are derived based on HoG (Histogram of Gradients) calculated from predicted samples of the current block or neighbouring reconstructed samples. One or more target inter MTS (Multiple Transform Selection) sets are derived according to said one or more VIPMs. Forward transform is applied to the residual data according to a selected transform kernel of said one or more target inter MTS sets at the encoder side to generate transformed data at the encoder side, or inverse transform is applied to the coded transformed residual data according to the selected transform kernel of said one or more target inter MTS sets to derive reconstructed residual data. The transformed data at the encoder side or the reconstructed residual data at the decoder side is provided.

[0059] In one embodiment, M target inter MTS sets are used in an inter TB (Transform Block) , and wherein the N target inter MTS sets are determined by N VIPMs and other target inter MTS sets are determined by one or more predefined methods, and wherein M and N are integers, M >=N and N >=0. In one embodiment, the N VIPMs for the inter TB correspond to N intra directional modes with largest histogram amplitudes. In one embodiment, a flag is signalled to indicate whether the N target inter MTS sets are determined by the N VIPMs or determined by said one or more predefined methods.

[0060] In one embodiment, said one or more target inter MTS sets are determined for an inter TB according to TB or CB (Coding Block) size, TB or CB area, TB or CB shape, inter prediction mode, VIPM of the inter TB, SBT (Subblock Transform) type, QP (Quantization Parameter) value, neighbouring prediction mode, motion information, absolute sum of quantized coefficients, number of non-zero quantized coefficients, difference of derived intra angels, gradient characteristic of current block samples or neighbouring samples, or any combination thereof.

[0061] In one embodiment, said one or more target inter MTS sets are derived also depending on difference between two dominant directions of current TB or CB from the current block, wherein the two dominant directions of the current TB or CB are derived based on two highest HoG values calculated using prediction samples of the current block or neighbouring samples of the current block. In one embodiment, when a first condition is satisfied, said one or more target inter MTS sets derived based on said one or more VIPMs or the difference between the two dominant directions of current TB or CB are used to derive the selected transform kernel of said one or more target inter MTS sets for encoding or decoding the current block, and wherein the first condition corresponds to a first result by comparing an absolute sum of quantized coefficients of the current TB with a first threshold.

[0062] In one embodiment, when the first condition is not satisfied, a first explicit MTS sets or second explicit MTS sets are selected depending on whether a second condition is satisfied or not, and wherein the second condition corresponds to a second result by comparing the absolute sum of quantized coefficients of the current TB with a second threshold. In one embodiment, said one or more target inter MTS sets are determined for an inter TB according to TB or CB (Coding Block) size, TB or CB area, TB or CB shape, inter prediction mode, VIPM of the inter TB, SBT (Subblock Transform) type, QP (Quantization Parameter) value, neighbouring prediction mode, motion information, absolute sum of quantized coefficients, number of non-zero quantized coefficients, difference of derived intra angels, gradient characteristic of current block samples or neighbouring samples, or any combination thereof.BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.

[0064] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.

[0065] Fig. 2 illustrates an example of Low-Frequency Non-Separable Transform (LFNST) process.

[0066] Fig. 3A illustrates an example of Region-Of-Interest (ROI) for LFNST16.

[0067] Fig. 3B illustrates an example of Region-Of-Interest (ROI) for LFNST8.

[0068] Fig. 4 illustrates an example of NSPT and LFNST used for various block sizes.

[0069] Fig. 5 illustrates a flowchart of an exemplary video coding system that derives inter MTS sets by using VIPMs (Virtual Intra Prediction Modes) according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION

[0070] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0071] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is provided by way of example only and is intended to illustrate certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.

[0072] Implicit MTS Improvements Proposed in JVET-AJ0259

[0073] An improved implicit MTS method is proposed by considering, in addition to block shape, intra mode type (whether it is regular intra-prediction, DIMD, TIMD, SGPM, Intra-TMP, MIP or EIP) , intra prediction direction as in EE-4.2, and the difference between the two dominant directions within the coding block. Two dominant directions are estimated in a similar fashion as the multiple transform set selection for LFNST / NSPT method described in Section 8. The existing DCT / DST kernels in ECM are being used for this method, thus no new transform kernels are added in ECM.

[0074] Determining Inter MTS Set by Virtual Intra Prediction Mode

[0075] In the present invention, the inter MTS sets used in an inter transform block (TB) can be determined by the virtual intra prediction modes (VIPM) derived from the prediction samples of the current block or neighbouring reconstructed samples. Specifically, the VIPMs can be derived by the HOG calculated from the prediction samples of current block or neighbouring reconstructed samples. The N intra directional modes with the largest histogram amplitude are selected as the VIPMs of current inter TB, where N is greater than or equal to 1. The MTS sets used in the current TB can then be determined by the intra MTS sets corresponding to the selected VIPMs.

[0076] In one embodiment, N (where N >=1) inter MTS sets used in an inter TB are determined by N VIPMs and other inter MTS sets are determined by the predefined methods (e.g., {(DST7, DST7) , (DST7, DCT8) , (DCT8, DST7) , (DCT8, DCT8) } are used or separable / non-separable KLTs are used) . In other words, if there are a total of M inter MTS sets, N inter MTS sets are determined based on VIPMs and the remaining inter MTS sets are determined by pre-defined methods, where M and N are integers and M >=N. In another embodiment, all the inter MTS sets used in an inter TB are determined by VIPMs.

[0077] In another embodiment, the methods of determining inter MTS sets used in an inter TB are conditioned by the TB or CB size, TB or CB area, TB or CB shape, inter prediction mode, VIPM of the inter TB, SBT type, QP value, neighbouring prediction mode, motion information, absolute sum of quantized coefficients, number of non-zero quantized coefficients, difference of derived intra angels, gradient characteristic of current block samples or neighbouring samples, or any combination thereof. In one example, when the TB area is smaller or larger than a threshold, the inter MTS sets are determined by the VIPMs. Otherwise, the inter MTS sets are determined by the predefined methods.

[0078] In another embodiment, N (where N >=1) inter MTS sets used in an inter TB are determined by N VIPMs and the other inter MTS sets are determined by the predefined methods. The number N (where N >=0) are conditioned by the TB or CB size, TB or CB area, TB or CB shape, inter prediction mode, VIPM of the inter TB, SBT type, QP value, neighbouring prediction mode, motion information, absolute sum of quantized coefficients, number of non-zero quantized coefficients, difference of derived intra angels, gradient characteristic of current block samples or neighbouring samples, or any combination thereof. In one example, when the absolute sum of quantized coefficients of current TB is smaller or larger than a first threshold, N is set to T1 (where T1 >=0) . Otherwise, when the absolute sum of quantized coefficients of current TB is smaller or larger than a second threshold, N is set to T2 (where T2 >=0) . Otherwise, N is set to T3 (where T3 >=0) . One or more syntax / index can be signalled to select the implicit MTS set derived from VIPM or predefined MTS sets.

[0079] In another embodiment, a flag is signalled to indicate whether the inter MTS sets are determined by the VIPMs or determined by the predefined methods (e.g., { (DST7, DST7) , (DST7, DCT8) , (DCT8, DST7) , (DCT8, DCT8) } are used or separable KLTs are used) . When the flag is true or false, the inter MTS sets are determined by the VIPMs. Otherwise, the inter MTS sets are determined by the predefined methods.

[0080] In another embodiment, N (where N >=1) inter MTS sets used in an inter TB are determined by N VIPMs and the other inter MTS sets are determined by the predefined methods. All of the inter MTS sets form an inter MTS candidate list. An index is signalled to indicate the best inter MTS set used in current TB.

[0081] The contexts of the flag indicating the usage of determining the inter MTS set by VIPMs in current TB and / or the index indicating the best inter MTS set among all the inter MTS sets can be conditioned by the current or neighbouring coding mode, neighbouring MTS set, neighbouring transform type, or any coding information in the current or neighbouring TBs. In one example, when the inter MTS sets of both top and left neighbouring TBs are determined by VIPMs. A certain context model is used for coding the flag that indicates the usage of determining the inter MTS set by VIPMs of current TB. When the inter MTS set of one of the top and left neighbouring TBs is determined by VIPMs, another certain context model is used for coding the flag that indicates the usage of determining the inter MTS set by VIPMs of current TB. When none of the inter MTS set of top and left neighbouring TBs is determined by VIPMs. Another certain context model is used for coding the flag that indicates the usage of determining the inter MTS set by VIPMs of current TB.

[0082] Implicit MTS Pair Determined by Intra Prediction Mode

[0083] In one embodiment, the inter MTS pair used in an inter TB can be determined implicitly by the VIPM (i.e., intra direction) derived by the HOG calculated from the prediction samples of current block or neighbouring reconstructed samples similar as the method mentioned in Section 10. The derived VIPM (i.e., intra direction) with the largest histogram amplitude, the TB / CB shape, and / or the difference between two dominant directions of current TB / CB can then be used to determine the MTS pair of current TB.

[0084] In one embodiment, the usage of the proposed implicit inter MTS method and the implicit intra MTS method can be conditioned by the TB or CB size, TB or CB area, TB or CB shape, inter prediction mode, VIPM of the inter TB, SBT type, QP value, neighbouring prediction mode, motion information, absolute sum of quantized coefficients, number of non-zero quantized coefficients, difference of derived intra angels, gradient characteristic of current block samples or neighbouring samples, or any combination thereof. In one embodiment, the transform sets used for inter and intra mode for a VIPM are the same. In another embodiment, the transform sets used for inter and intra mode for a VIPM are the different. Different transform set are used for inter and intra mode. In one example, the inter and / or intra implicit MTS methods can be applied to the TBs with area smaller or larger than a threshold. Otherwise, the explicit MTS method is applied. In another example, when the absolute sum of quantized coefficients of current TB is smaller or larger than a first threshold, the inter and / or intra implicit MTS methods can be applied. Otherwise, when the absolute sum of quantized coefficients of current TB is smaller or larger than a second threshold, the explicit MTS methods can be applied with M1 (M1 >=1) MTS sets to be selected. Otherwise, the explicit MTS methods can be applied with M2 (M2 >=1) MTS sets to be selected.

[0085] Any of the foregoing proposed methods can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in transform module of an encoder and / or a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to transform module of the encoder and / or the decoder.

[0086] With reference to the exemplary encoder or decoder in Fig. 1A and Fig. 1B, any of the proposed methods can be implemented in a transform module and / or an Intra / Inter coding module (e.g. IT 126, Intra Pred. 150 / MC 152 in Fig. 1B) in a decoder or an Intra / Inter coding module in an encoder (e.g. T 118, Intra Pred. 110 / Inter Pred. 112 in Fig. 1A) . Any of the proposed methods can also be implemented as circuits coupled to the intra coding module at the decoder or the encoder. However, the decoder or encoder may also use additional processing unit to implement the required processing. While the transform module and / or Intra / Inter Pred. units (e.g. T 118, Intra Pred. 110 / Inter Pred. 112 in Fig. 1A and IT 126, Intra Pred. 150 / MC 152 in Fig. 1B) are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .

[0087] Fig. 5 illustrates a flowchart of an exemplary video coding system that derives inter MTS sets by using VIPMs (Virtual Intra Prediction Modes) according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to the method, input data is received in step 510, wherein the input data comprises residual data for a current block at an encoder side or coded transformed residual data for the current block at a decoder side, wherein the current block is coded in inter prediction. One or more VIPMs (Virtual Intra Prediction Modes) are derived based on HoG (Histogram of Gradients) calculated from predicted samples of the current block or neighbouring reconstructed samples in step 520. One or more target inter MTS (Multiple Transform Selection) sets are derived according to said one or more VIPMs in step 530. Forward transform is applied to the residual data according to a selected transform kernel of said one or more target inter MTS sets at the encoder side to generate transformed data at the encoder side, or inverse transform is applied to the coded transformed residual data according to the selected transform kernel of said one or more target inter MTS sets to derive reconstructed residual data in step 540. The transformed data at the encoder side or the reconstructed residual data at the decoder side is provided in step 550.

[0088] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0089] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.

[0090] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.

[0091] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1.A method of video coding, the method comprising:receiving input data, wherein the input data comprises residual data for a current block at an encoder side or coded transformed residual data for the current block at a decoder side, wherein the current block is coded in inter prediction;deriving one or more VIPMs (Virtual Intra Prediction Modes) based on HoG (Histogram of Gradients) calculated from predicted samples of the current block or neighbouring reconstructed samples;deriving one or more target inter MTS (Multiple Transform Selection) sets according to said one or more VIPMs;applying forward transform to the residual data according to a selected transform kernel of said one or more target inter MTS sets at the encoder side to generate transformed data at the encoder side, or applying inverse transform to the coded transformed residual data according to the selected transform kernel of said one or more target inter MTS sets to derive reconstructed residual data; andproviding the transformed data at the encoder side or the reconstructed residual data at the decoder side.2.The method of Claim 1, wherein M target inter MTS sets are used in an inter TB (Transform Block) , and wherein the N target inter MTS sets are determined by N VIPMs and other target inter MTS sets are determined by one or more predefined methods, and wherein M and N are integers, M >= N and N >=1.3.The method of Claim 2, wherein the N VIPMs for the inter TB correspond to N intra directional modes with largest histogram amplitudes.4.The method of Claim 2, wherein a flag is signalled to indicate whether the N target inter MTS sets are determined by the N VIPMs or determined by said one or more predefined methods.5.The method of Claim 1, wherein said one or more target inter MTS sets are determined for an inter TB according to TB or CB (Coding Block) size, TB or CB area, TB or CB shape, inter prediction mode, VIPM of the inter TB, SBT (Subblock Transform) type, QP (Quantization Parameter) value, neighbouring prediction mode, motion information, absolute sum of quantized coefficients, number of non-zero quantized coefficients, difference of derived intra angels, gradient characteristic of current block samples or neighbouring samples, or any combination thereof.6.The method of Claim 1, wherein said one or more target inter MTS sets are derived also depending on difference between two dominant directions of current TB or CB from the current block, wherein the two dominant directions of the current TB or CB are derived based on two highest HoG values calculated using prediction samples of the current block or neighbouring samples of the current block.7.The method of Claim 6, wherein when a first condition is satisfied, said one or more target inter MTS sets derived based on said one or more VIPMs or the difference between the two dominant directions of current TB or CB are used to derive the selected transform kernel of said one or more target inter MTS sets for encoding or decoding the current block, and wherein the first condition corresponds to a first result by comparing an absolute sum of quantized coefficients of the current TB with a first threshold.8.The method of Claim 7, wherein when the first condition is not satisfied, a first explicit MTS sets or second explicit MTS sets are selected depending on whether a second condition is satisfied or not, and wherein the second condition corresponds to a second result by comparing the absolute sum of quantized coefficients of the current TB with a second threshold.9.The method of Claim 6, wherein said one or more target inter MTS sets are determined for an inter TB according to TB or CB (Coding Block) size, TB or CB area, TB or CB shape, inter prediction mode, VIPM of the inter TB, SBT (Subblock Transform) type, QP (Quantization Parameter) value, neighbouring prediction mode, motion information, absolute sum of quantized coefficients, number of non-zero quantized coefficients, difference of derived intra angels, gradient characteristic of current block samples or neighbouring samples, or any combination thereof.10.An apparatus of video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data, wherein the input data comprises residual data for a current block at an encoder side or coded transformed residual data for the current block at a decoder side, wherein the current block is coded in inter prediction;derive one or more VIPMs (Virtual Intra Prediction Modes) based on HoG (Histogram of Gradients) calculated from predicted samples of the current block or neighbouring reconstructed samples;derive one or more target inter MTS (Multiple Transform Selection) sets according to said one or more VIPMs;apply forward transform to the residual data according to a selected transform kernel of said one or more target inter MTS sets at the encoder side to generate transformed data at the encoder side, or applying inverse transform to the coded transformed residual data according to the selected transform kernel of said one or more target inter MTS sets to derive reconstructed residual data; andprovide the transformed data at the encoder side or the reconstructed residual data at the decoder side.