Method and apparatus of deriving transform type for subblock coding in video coding systems
Patent Information
- Application Number
- PCT/CN2026/084447
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-19
- Publication Date
- 2026-09-24
Smart Images

Figure CN2026084447_24092026_PF_FP_ABST
Abstract
Description
METHOD AND APPARATUS OF DERIVING TRANSFORM TYPE FOR SUBBLOCK CODING IN VIDEO CODING SYSTEMSCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 774,191, filed on March 19, 2025. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION
[0002] The present invention relates to video coding system. In particular, the present invention relates to transform pair for SBT (Subblock Transform) coded blocks in a video coding system. BACKGROUND AND RELATED ART
[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.
[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction 110, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and in-loop filter 130, is provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.
[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, in-loop filter 130 is often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, Loop filter 130 is applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.
[0006] The decoder, as shown in Fig. 1B, can use some of the functional blocks as the encoder. For example, the decoder can reuse Inverse Quantization 124 and Inverse Transform 126; however, Transform 118 and Quantization 120 are not needed at the decoder. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.
[0007] In VVC, the Sequence Parameter Set (SPS) and the Picture Parameter Set (PPS) contain high-level syntax elements that apply to entire coded video sequences and pictures, respectively. The Picture Header (PH) and Slice Header (SH) contain high-level syntax elements that apply to a current coded picture and a current coded slice, respectively.
[0008] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A bi-predictive (B) slice may be decoded using intra prediction or inter prediction with at most two motion vectors and reference indices to predict the sample values of each block. A predictive (P) slice is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using intra prediction only.
[0009] In VVC, each CTU can be partitioned into one or multiple non-overlapped coding units (CUs) using a quaternary tree (QT) with nested multi-type-tree (MTT) structure. The partitioning information is signalled by a coding tree syntax structure, where each CTU is treated as the root of a coding tree. The CTUs may be first partitioned by the quaternary tree (a.k. a. quadtree) structure, as shown 210 in Fig. 2. Then the quaternary tree leaf nodes can be further partitioned by a MTT structure, as shown in Fig. 2. There are four splitting types in multi-type tree structure: vertical binary splitting (SPLIT_BT_VER) 220 as shown in Fig. 2, horizontal binary splitting (SPLIT_BT_HOR) 230 as shown in Fig. 2, vertical ternary splitting (SPLIT_TT_VER) 240 as shown in Fig. 2, and horizontal ternary splitting (SPLIT_TT_HOR) 250 as shown in Fig. 2. Each quadtree child node may be further split into smaller coding tree nodes using any one of five splitting types in Fig. 2. However, each multi-type-tree child node is only allowed to be further split by one of four MTT splitting types. The coding tree leaf nodes correspond to the coding units (CUs) . Fig. 3 provides an example of a CTU recursively partitioned by QT with the nested MTT, where the bold block edges represent QT partitioning and the remaining edges represent multi-type tree partitioning.
[0010] Each CU contains one or more prediction units (PUs) . The prediction unit, together with the associated CU syntax, works as a basic unit for signalling the predictor information. The specified prediction process is employed to predict the values of the associated pixel samples inside the PU. Each CU may contain one or more transform units (TUs) for representing the prediction residual blocks. A transform unit (TU) is comprised of one transform block (TB) of luma samples and two corresponding transform blocks of chroma samples. Each TB corresponds to one residual block of samples from one colour component. An integer transform is applied to a transform block. The level values of quantized coefficients together with other side information are entropy coded in the bitstream. The terms coding tree block (CTB) , coding block (CB) , prediction block (PB) , and transform block (TB) are defined to specify the 2-D sample array of one colour component associated with CTU, CU, PU, and TU, respectively. Thus, a CTU consists of one luma CTB, two chroma CTBs, and associated syntax elements. A similar relationship is valid for CU, PU, and TU.
[0011] In VVC, an intra dual tree mode can be employed for coding intra slices. When the intra dual tree mode is applied, the luma CTB is partitioned into CUs by one coding tree structure, and the two chroma CTBs are partitioned into chroma CUs by another coding tree structure. There are two coding tree syntax structures thus signalled for luma and chroma, respectively, in a CTU. As a result, each CU consists of either one coding block from the luma component or two coding blocks, respectively from the two chroma components, when the intra dual tree mode is applied.
[0012] For achieving high compression efficiency, the context-based adaptive binary arithmetic coding (CABAC) mode, or known as regular mode, is employed for entropy coding of the values of the syntax elements in VVC. Fig. 4 provides the block diagram of the CABAC process. As the arithmetic coder in the CABAC engine can only encode the binary symbol values, the CABAC operation first needs to convert the value of a syntax element into a binary string using a binarizer (410) , the process commonly referred to as binarization. During the coding process, the probability models are gradually built up from the coded symbols for the different contexts. The selection of the modelling context, for coding the next binary symbol can be determined the coded information. Symbols can be coded without the context modelling stage and assume predetermined probability distribution, commonly referred to as the bypass mode, for improving bitstream parsing throughput rate.
[0013] In Fig. 4, the context modeller (420) serves the modelling purpose. During normal context based coding, the regular coding engine (430) is used, which corresponds to a binary arithmetic coder. For the bypassed symbols, a bypass coding engine (440) may be used. As shown in Fig. 4, switches (S1, S2 and S3) are used to direct the data flow between the regular CABAC mode and the bypass mode. When the regular CABAC mode is selected, the switches are flipped to the upper contacts. When the bypass mode is selected, the switches are flipped to the lower contacts as shown in Fig. 4.
[0014] Joint Video Expert Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 are currently in the process of exploring the next-generation video coding standard. Some promising new coding tools have been adopted into Enhanced Compression Model (ECM) (M. Coban, R. -L. Liao, K. Naser, J. L. Zhang “Algorithm description of Enhanced Compression Model 14 (ECM 14) ” , Joint Video Expert Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, Doc. 35th Meeting: Sapporo, JP, 12–19 July 2024) to further improve VVC.
[0015] In the present invention, methods and apparatus of alternative transform pair for SBT (Subblock Transform) coded blocks are disclosed. BRIEF SUMMARY OF THE INVENTION
[0016] A method and apparatus of alternative transform pair for SBT (Subblock Transform) coded blocks are disclosed. According to this method, input data associated with a current block is received, wherein the input data comprise residual data to be processed at an encoder side or coded transform data associated with the current block to be decoded at a decoder side, and wherein the current block is coded in SBT (Subblock Transform) mode. The current block is partitioned vertically or horizontally into two subblocks. A transform is applied process according to a target transform pair to a non-zero subblock of the two subblocks to generate a transform processed subblock, wherein the target transform pair consists of DCT-2 applied to a first spatial direction and a second transform type different from the DCT-2 applied to a second spatial direction. Transformed output comprising the transform processed subblock is provided, wherein the transform processed subblock comprises transform coefficients of the non-zero subblock at the encoder side or recovered residual data at the decoder side.
[0017] In one embodiment, the second transform type comprises DCT-8 or DCT-7.
[0018] In one embodiment, whether to apply the target transform pair to the current block is dependent on contextual information associated the current block. In one embodiment, the contextual information comprises block dimension, nonzero subblock dimension, SBT split direction, SBT split type, SBT position, or a combination thereof. In one embodiment, when block height or width of the current block coded in an SBT-V or SBT-H mode, respectively, exceeds a predefined threshold, DCT-2 is applied along a vertical or horizontal dimension of the target transform pair. In one embodiment, when the block height or width of the current block coded in the SBT-V or SBT-H mode, respectively, does not exceed the predefined threshold, DST-7 is applied along the vertical or horizontal dimension of the target transform pair.
[0019] In one embodiment, the method 1 further comprises signalling one or more syntax elements in one or more high-level syntax sets to indicate whether the target transform pair is enabled or disabled for the current block. In one embodiment, said one or more high-level syntax sets are signalled or parsed in SPS (Sequence Parameter Set) , PPS (Picture Parameter Set) , PH (Picture Header) , SH (Slice Header) , or a combination thereof.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.
[0021] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.
[0022] Fig. 2 illustrates that a CU can be split into smaller CUs using one of the five splitting types (quad-tree partitioning, vertical binary tree partitioning, horizontal binary tree partitioning, vertical centre-side triple-tree partitioning, and horizontal centre-side triple-tree partitioning) .
[0023] Fig. 3 illustrates an example of a CTU being recursively partitioned by QT with the nested MTT.
[0024] Fig. 4 illustrates an exemplary block diagram of the CABAC process.
[0025] Fig. 5 illustrates SBT splitting types and transform types in VVC.
[0026] Fig. 6 illustrates an example of corner subblocks (only top left is shown) .
[0027] Fig. 7 illustrates a flowchart of an exemplary video coding system that uses alternative transform pair for SBT (Subblock Transform) coded blocks according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION
[0028] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.
[0029] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.
[0030] PROPOSED METHODS
[0031] In addition to DCT-2, which has been employed in HEVC, a Multiple Transform Selection (MTS) scheme is used for residual coding for both inter and intra coded luma blocks in VVC. Additional separable primary transform pairs derived from DCT-8 and DST-7 can be selected for transformation of residual blocks. The selected horizontal and vertical primary transform pair for a current block is indicated by a syntax element mts_idx, specified as follows: Table 1: Selected primary transform type by mts_idx where trTypeHor and trTypeVer in a separable primary transform pair (trTypeHor, trTypeVer) indicate the selected horizontal primary transform type and vertical primary transform type, respectively, with 0, 1, and 2 corresponding to DCT-2, DST-7, and DCT-8, respectively. When mts_idx is greater than 0, one of four MTS candidates { (DST-7, DST-7) , (DCT-8, DST-7) , (DST-7, DCT-8) , (DCT-8, DCT-8) } can be selected for a residual luma block.
[0032] An intra block can be coded in an implicit MTS mode without explicitly signalling mts_idx, where the selected MTS primary transform pair is determined by the width and the height of a coded current block. In ECM, for 4-pt, 8-pt and 16-pt transforms, the MTS transform cores DST-7 and DCT-8 in VVC can be replaced with two separable Karhunen–Loève transforms (KLTs) , KLT-0 and KLT-1, for coding inter luma blocks.
[0033] In VVC, a subblock transform (SBT) mode is introduced for an inter-predicted CU.In this transform mode, only a sub-part of the residual block is coded for the CU. When an inter-predicted CU is coded with nonzero residual (indicated by the syntax flag cu_coded_flag equal to 1) , a syntax flag cu_sbt_flag may be signalled to indicate whether the whole residual block or a sub-part of the residual block is coded. In the former case, inter MTS information may be explicitly signalled to determine the selected MTS transform pair for the CU. In the latter case, one part of the residual block is nonzero and is coded in an implicit MTS mode, wherein the selected MTS transform pair can be derived by the encoder and the decoder without explicitly signalling any syntax information. The other part of the residual block is simply zeroed out.
[0034] When a SBT mode is used for an inter-coded CU, information indicating the SBT type and the SBT position is signalled in the bitstream. There are 8 SBT types derived from two SBT split directions, two SBT split types and two SBT positions, as indicated in Fig. 5. For SBT-V (510 and 520) or SBT-H (530 and 540) , the TU width (or height) may equal to half of the CU width (or height) or 1 / 4 of the CU width (or height) , resulting in 2: 2 split or 1:3 / 3: 1 split. Position-dependent transform type selection is applied on luma transform blocks in SBT-V and SBT-H (chroma TB always using DCT-2) . The two positions of SBT-H and SBT-V are associated with different core transforms. More specifically, the horizontal and vertical transform pair for each SBT position is specified in Fig. 5. For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7, respectively. In ECM, four corner types are introduced to SBT shown in Fig. 6. The size of the nonzero subblock can be width / 2 x height / 2 or width / 4 x height / 4 of the whole block, the latter can be applied for transform blocks 16x16 and larger. The MTS transform pairs are implicitly derived from combinations of DCT-8 and DST-7 for corner subblocks.
[0035] In VVC and ECM, when a separable primary transform pair is applied to inter encoding or decoding a current block in SBT mode, the current block is coded in implicit MTS mode and the adopted transform pair is implicitly selected from the set of fives MTS transform pairs listed in Table 1, comprising (DCT-2, DCT-2) , (DST-7, DST-7) , (DCT-8, DST-7) , (DST-7, DCT-8) , and (DCT-8, DCT-8) . According to an aspect of the present invention, a video coder may further adopt a separable primary transform pair with DCT-2 applied to only one of the two spatial dimensions for transformation of a nonzero subblock in a block coded in SBT mode.
[0036] In the proposed method, a video coder may encode or decode a current block in SBT mode by using a separable primary transform pair with DCT-2 applied to one spatial dimension and a different transform type applied to the other spatial dimension for transformation of a nonzero subblock in the current block. In some embodiments, a video coder may encode or decode a current block in an SBT-V or SBT-H type by adopting DCT-2 in the spatial dimension along the SBT block splitting direction while using the implicitly selected transform type in the other spatial dimension. For example, a video coder may adopt a transform pair (DCT-8, DCT-2) , instead of (DCT-8, DST-7) , for transformation of a nonzero subblock in a current block coded in an SBT-V type at SBT position 0. The video coder may adopt a transform pair (DST-7, DCT-2) , instead of (DST-7, DST-7) , for transformation of a nonzero subblock in a current block coded in an SBT-V type at SBT position 1. The video coder may adopt a transform pair (DCT-2, DCT-8) , instead of (DST-7, DCT-8) , for transformation of a nonzero subblock in a current block coded in an SBT-H type at SBT position 0. The video coder may adopt a transform pair (DCT-2, DST-7) , instead of (DST-7, DST-7) , for transformation of a nonzero subblock in a current block coded in an SBT-H type at SBT position 1.
[0037] According to a further aspect of the present invention, a video coder may adaptively determine whether to adopt a proposed separable primary transform pair with DCT-2 applied to only one of the two spatial dimensions for encoding or decoding a current block in a particular SBT type further considering contextual information such as block dimension (width, height, shape, and / or area size) , nonzero subblock dimension, SBT split direction, SBT split type, SBT position or some combination thereof.
[0038] In some embodiments, a video coder may determine whether to adopt a proposed transform pair for inter encoding or decoding a current block in a particular SBT-V or SBT-H type by comparing the block size (width or height) in the spatial dimension along the current block splitting direction with a specified threshold. In some specific embodiments, when the block height (or width) of a current block coded in an SBT-V (or SBT-H) type is greater than a specified threshold Tsbt, DCT-2 is applied to the vertical (or horizontal) dimension of the separable transform pair. Otherwise, DST-7 is applied to vertical (or horizontal) dimension of the separable transform pair. For example, for inter encoding or decoding a current block in an SBT-V type at position 0, when the block height is greater than a specified threshold Tsbt, the selected separable primary transform pair is (DCT-8, DCT-2) . Otherwise, the selected separable primary transform pair is (DCT-8, DST-7) . Some preferred values of Tsbt can be equal to 4, 8, 16, 32, or 64.
[0039] In some embodiments, a video coder may determine whether to adopt a proposed separable primary transform pair with DCT-2 applied to only one of the two spatial dimensions for inter encoding or decoding a current block in a particular SBT type by comparing the nonzero subblock width and / or height of the current block with one or more specified thresholds. In some embodiments, the selected transform type for a horizontal or vertical spatial dimension is determined by comparing the nonzero subblock width or height with a specified threshold.
[0040] In some preferred embodiments for inter encoding or decoding a current block in an SBT-V or SBT-H type, different thresholds may be adopted for the nonzero subblock width and height of the block, respectively, depending on whether the nonzero subblock width or height is less than the current block width or height. For example, when the current block is split vertically (or horizontally) and the resulting nonzero subblock height (or width) is equal to the height (or width) of the current block, a video coder may determine whether to adopt a proposed separable transform pair in which DCT-2 is applied to only one of the two spatial dimensions for inter encoding or decoding of the current block in a particular SBT-V or SBT-H type, based on a comparison of the nonzero subblock height (or width) with a first threshold and a comparison of the nonzero subblock width (or height) with a second threshold..
[0041] In some specific embodiments, when the nonzero subblock height (or width) of the coded block in an SBT-V (or SBT-H) type is greater than a specified threshold Tsbt1, DCT-2 is applied to the vertical (or horizontal) dimension of the separable transform pair. Otherwise, DST-7 is applied to vertical (or horizontal) dimension of the separable transform pair. When the nonzero subblock width (or height) of the coded block in an SBT-V (or SBT-H) type is greater than a specified threshold Tsbt2, DCT-2 is applied to the horizontal (or vertical) dimension of the separable transform pair. Otherwise, DCT-8 and DST-7 are applied to the horizontal (or vertical) dimension of the separable transform pair for SBT positions 0 and 1, respectively.
[0042] In some other embodiments, for a current block coded in an SBT-V (or SBT-H) type, when the nonzero subblock height (or width) of the current block is greater than a specified threshold Tsbt1 and the nonzero subblock width (or height) of the current block is greater than a specified threshold Tsbt2, DCT-2 is applied to the vertical (or horizontal) dimension of the separable transform pair. Otherwise, DST-7 is applied to vertical (or horizontal) dimension of the separable transform pair.
[0043] The proposed methods may further comprise signalling one or more syntax elements in one or more high-level syntax sets to indicate whether any of the proposed methods is enabled or disable in a current video data unit, wherein the one or more high-level syntax sets may comprise SPS, PPS, PH, SH, or a combination thereof.
[0044] Any of the foregoing proposed methods of alternative transform pair for SBT (Subblock Transform) coded blocks can be implemented in encoders and / or decoders. For example, any of the proposed methods or combination thereof can be implemented in a transform module of an encoder, and / or an inverse transform module of a decoder. Alternatively, any of the proposed methods or combination thereof can be implemented as a circuit integrated to the transform module of the encoder and / or the transform module of the decoder. The proposed aspects, methods and related embodiments can be implemented individually or jointly in a video coding system.
[0045] With reference to the exemplary encoder and decoder in Fig. 1A and Fig. 1B, any of the proposed methods of adaptive entropy coding of block partitioning information can be implemented in an Intra / Inter / Entropy coding module (e.g. Intra Pred. 150 / MC 150 / Entropy Decoder 140 in Fig. 1B) in a decoder or an Intra / Inter / Entropy coding module (e.g. Intra Pred. 110 / Inter Pred. 112 / Entropy Encoder 122 in Fig. 1A) in an encoder. Any of the proposed methods can also be implemented as a circuit coupled to the intra / inter coding module at the decoder or the encoder. However, the decoder or encoder may also use additional processing unit to implement the required processing. While the proposed methods can be implanted using various processing modules as shown in Fig. 1A and Fig. 1B, the proposed methods may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .
[0046] Fig. 7 illustrates a flowchart of an exemplary video coding system that uses alternative transform pair for SBT (Subblock Transform) coded blocks according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to this method, input data associated with a current block is received in step 710, wherein the input data comprise residual data to be processed at an encoder side or coded transform data associated with the current block to be decoded at a decoder side, and wherein the current block is coded in SBT (Subblock Transform) mode. The current block is partitioned vertically or horizontally into two subblocks in step 720. A transform is applied process according to a target transform pair to a non-zero subblock of the two subblocks to generate a transform processed subblock in step 730, wherein the target transform pair consists of DCT-2 applied to a first spatial direction and a second transform type different from the DCT-2 applied to a second spatial direction. Transformed output comprising the transform processed subblock is provided in step 740, wherein the transform processed subblock comprises transform coefficients of the non-zero subblock at the encoder side or recovered residual data at the decoder side.
[0047] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.
[0048] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.
[0049] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.
[0050] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
1.A method of video coding, the method comprising:receiving input data associated with a current block, wherein the input data comprise residual data to be processed at an encoder side or coded transform data associated with the current block to be decoded at a decoder side, and wherein the current block is coded in SBT (Subblock Transform) mode;partitioning the current block vertically or horizontally into two subblocks;applying a transform process according to a target transform pair to a non-zero subblock of the two subblocks to generate a transform processed subblock, wherein the target transform pair consists of DCT-2 applied to a first spatial direction and a second transform type different from the DCT-2 applied to a second spatial direction; andproviding transformed output comprising the transform processed subblock, wherein the transform processed subblock comprises transform coefficients of the non-zero subblock at the encoder side or recovered residual data at the decoder side.2.The method of Claim 1, wherein the second transform type comprises DCT-8 or DCT-7.3.The method of Claim 1, wherein whether to apply the target transform pair to the current block is dependent on contextual information associated the current block.4.The method of Claim 3, wherein the contextual information comprises block dimension, nonzero subblock dimension, SBT split direction, SBT split type, SBT position, or a combination thereof.5.The method of Claim 4, wherein when block height or width of the current block coded in an SBT-V or SBT-H mode, respectively, exceeds a predefined threshold, DCT-2 is applied along a vertical or horizontal dimension of the target transform pair.6.The method of Claim 5, wherein when the block height or width of the current block coded in the SBT-V or SBT-H mode, respectively, does not exceed the predefined threshold, DST-7 is applied along the vertical or horizontal dimension of the target transform pair.7.The method of Claim 1 further comprises signalling one or more syntax elements in one or more high-level syntax sets to indicate whether the target transform pair is enabled or disabled for the current block.8.The method of Claim 7, wherein said one or more high-level syntax sets are signalled or parsed in SPS (Sequence Parameter Set) , PPS (Picture Parameter Set) , PH (Picture Header) , SH (Slice Header) , or a combination thereof.9.An apparatus for video coding, the apparatus comprising one or more electronics or processors arranged to:receive input data associated with a current block, wherein the input data comprise residual data to be processed at an encoder side or coded transform data associated with the current block to be decoded at a decoder side, and wherein the current block is coded in SBT (Subblock Transform) mode;partition the current block vertically or horizontally into two subblocks;apply a transform process according to a target transform pair to a non-zero subblock of the two subblocks to generate a transform processed subblock, wherein the target transform pair consists of DCT-2 applied to a first spatial direction and a second transform type different from the DCT-2 applied to a second spatial direction; andprovide transformed output comprising the transform processed subblock, wherein the transform processed subblock comprises transform coefficients of the non-zero subblock at the encoder side or recovered residual data at the decoder side.