Method and apparatus for predicting signs of transform coefficients in image or video coding systems

The method optimizes transform coefficient sign prediction and encoding in video coding systems by constraining the number of predicted signs and using entropy coding with bin budget constraints, addressing inefficiencies in high-complexity scenarios like VVC, thus enhancing compression performance and reducing computational load.

WO2025209325A1PCT designated stage Publication Date: 2025-10-09MEDIATEK INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/085595
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-01
Filing Date
2025-03-28
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing video coding systems face inefficiencies in predicting and encoding the signs of transform coefficients, particularly in high-complexity scenarios like Versatile Video Coding (VVC), leading to increased computational load and suboptimal compression performance.

Method used

A method and apparatus for predicting and encoding transform coefficient signs by determining a maximum allowed number of predicted signs and applying sign prediction based on specified rules, using a residual coding process that includes entropy coding and constrained bin budgets to optimize the encoding process.

Benefits of technology

Improves encoding efficiency by reducing the need for revisiting transform blocks and minimizing computational overhead, thereby enhancing compression performance and reducing complexity in video coding systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025085595_09102025_PF_FP_ABST
    Figure CN2025085595_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for a residue process of multiple residue blocks including entropy coding a collection of coefficients for jointly sign coding are disclosed. According to this method, a constraint on a maximum allowed number of predicted signs for the current block is determined. A set of sign-predicted transform coefficients is determined for the current block based on one or more specified rules. For each of the multiple residue blocks: coefficient values in said each of the multiple residue blocks are encoded or decoded; sign prediction is derived for signs of the set of sign-predicted transform coefficients in said each of the multiple residue blocks; and entropy encoding or decoding is applied to the signs of the set of sign-predicted transform coefficients in said each of the multiple residue blocks using the sign prediction.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR PREDICTING SIGNS OF TRANSFORM COEFFICIENTS IN IMAGE OR VIDEO CODING SYSTEMSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present invention is a non-Provisional Application of and claims priority to U.S. Provisional Patent Application No. 63 / 572,386, filed on April 1, 2024. The U.S. Provisional Patent Application is hereby incorporated by reference in its entirety.FIELD OF THE INVENTION

[0002] The present invention relates to video coding system. In particular, the present invention relates to coding of signs of transform coefficients for residual blocks in a video coding system.BACKGROUND

[0003] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) . The standard has been published as an ISO standard: ISO / IEC 23090-3: 2021, Information technology -Coded representation of immersive media -Part 3: Versatile video coding, published Feb. 2021. VVC is developed based on its predecessor HEVC (High Efficiency Video Coding) by adding more coding tools to improve coding efficiency and also to handle various types of video sources including 3-dimensional (3D) video signals.

[0004] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing. For Intra Prediction, the prediction data is derived based on previously coded video data in the current picture. For Inter Prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based of the result of ME to provide prediction data derived from other picture (s) and motion data. Switch 114 selects Intra Prediction 110 or Inter-Prediction 112 and the selected prediction data is supplied to Adder 116 to form prediction errors, also called residues. The prediction error is then processed by Transform (T) 118 followed by Quantization (Q) 120. The transformed and quantized residues are then coded by Entropy Encoder 122 to be included in a video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packed with side information such as motion and coding modes associated with Intra prediction and Inter prediction, and other information such as parameters associated with loop filters applied to underlying image area. The side information associated with Intra Prediction 110, Inter prediction 112 and deblocking filter (DF) 130 / non-deblocking filters (NDFs) 132, are provided to Entropy Encoder 122 as shown in Fig. 1A. When an Inter-prediction mode is used, a reference picture or pictures have to be reconstructed at the encoder end as well. Consequently, the transformed and quantized residues are processed by Inverse Quantization (IQ) 124 and Inverse Transformation (IT) 126 to recover the residues. The residues are then added back to prediction data 136 at Reconstruction (REC) 128 to reconstruct video data. The reconstructed video data may be stored in Reference Picture Buffer 134 and used for prediction of other frames.

[0005] As shown in Fig. 1A, incoming video data undergoes a series of processing in the encoding system. The reconstructed video data from REC 128 may be subject to various impairments due to a series of processing. Accordingly, deblocking filter (DF) 130 / non-deblocking filters (NDFs) are often applied to the reconstructed video data before the reconstructed video data are stored in the Reference Picture Buffer 134 in order to improve video quality. For example, deblocking filter (DF) , Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF) may be used. The loop filter information may need to be incorporated in the bitstream so that a decoder can properly recover the required information. Therefore, loop filter information is also provided to Entropy Encoder 122 for incorporation into the bitstream. In Fig. 1A, deblocking filter (DF) 130 / non-deblocking filters (NDFs) 132 are applied to the reconstructed video before the reconstructed samples are stored in the reference picture buffer 134. The system in Fig. 1A is intended to illustrate an exemplary structure of a typical video encoder. It may correspond to the High Efficiency Video Coding (HEVC) system, VP8, VP9, H. 264 or VVC.

[0006] The decoder, as shown in Fig. 1B, can use some of the functional blocks as the encoder. For example, the decoder can reuse Inverse Quantization 124 and Inverse Transform 126; however, Transform 118 and Quantization 120 are not needed at the decoder. Instead of Entropy Encoder 122, the decoder uses an Entropy Decoder 140 to decode the video bitstream into quantized transform coefficients and needed coding information (e.g. ILPF information, Intra prediction information and Inter prediction information) . The Intra prediction 150 at the decoder side does not need to perform the mode search. Instead, the decoder only needs to generate Intra prediction according to Intra prediction information received from the Entropy Decoder 140. Furthermore, for Inter prediction, the decoder only needs to perform motion compensation (MC 152) according to Inter prediction information received from the Entropy Decoder 140 without the need for motion estimation.

[0007] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs. The individual CTUs in a slice are processed in raster-scan order. A bi-predictive (B) slice may be decoded using intra prediction or inter prediction with at most two motion vectors and reference indices to predict the sample values of each block. A predictive (P) slice is decoded using intra prediction or inter prediction with at most one motion vector and reference index to predict the sample values of each block. An intra (I) slice is decoded using intra prediction only.

[0008] A CTU can be partitioned into one or multiple non-overlapped coding units (CUs) using the quatree (QT) with nested multi-type-tree (MTT) structure to adapt to various local motion and texture characteristics. A CU can be further split into smaller CUs using one of the five split types (quad-tree partitioning 210, vertical binary tree partitioning 220, horizontal binary tree partitioning 230, vertical centre-side triple-tree partitioning 240, horizontal centre-side triple-tree partitioning 250) illustrated in Fig. 2. Fig. 3 provides an example of a CTU recursively partitioned by QT with the nested MTT. Each CU contains one or more prediction units (PUs) . The prediction unit, together with the associated CU syntax, works as a basic unit for signalling the predictor information. The specified prediction process is employed to predict the values of the associated pixel samples inside the PU. Each CU may contain one or more transform units (TUs) for representing the prediction residual blocks. A transform unit (TU) comprises of a transform block (TB) of luma samples and two corresponding transform blocks of chroma samples and each TB correspond to one residual block of samples from one colour component. An integer transform is applied to a transform block. The level values of quantized coefficients together with other side information are entropy coded in the bitstream. The terms coding tree block (CTB) , coding block (CB) , prediction block (PB) , and transform block (TB) are defined to specify the 2-D sample array of one colour component associated with CTU, CU, PU, and TU, respectively. Thus, a CTU consists of one luma CTB, two chroma CTBs, and associated syntax elements. A similar relationship is valid for CU, PU, and TU.

[0009] For achieving high compression efficiency, the context-based adaptive binary arithmetic coding (CABAC) mode, or known as the regular mode, is employed for entropy coding of the values of the syntax elements in HEVC and VVC. Fig. 4 illustrates an exemplary block diagram of the CABAC process. Since the arithmetic coder in the CABAC engine can only encode the binary symbol values, the CABAC process needs to convert the values of the syntax elements into a binary string using a binarizer (410) . The conversion process is commonly referred to as binarization. During the coding process, the probability models are gradually built up from the coded symbols for the different contexts. The context modeller (420) serves the modelling purpose. During normal context based coding, the regular coding engine (430) is used, which corresponds to a binary arithmetic coder. The selection of the modelling context for coding the next binary symbol can be determined by the coded information. Symbols can also be encoded without the context modelling stage and assume an equal probability distribution, commonly referred to as the bypass mode, for reduced complexity. For the bypassed symbols, a bypass coding engine (440) may be used. As shown in Fig. 4, switches (S1, S2 and S3) are used to direct the data flow between the regular CABAC mode and the bypass mode. When the regular CABAC mode is selected, the switches are flipped to the upper contacts. When the bypass mode is selected, the switches are flipped to the lower contacts as shown in Fig. 4.

[0010] In VVC, the transform coefficients may be quantized by dependent scalar quantization. The selection of one of the two quantizers is determined by a state machine with four states. The state for a current transform coefficient is determined by the state and the parity of the absolute level value associated with the preceding coded transform coefficient. The quantized transform coefficient levels in a transform block are entropy coded by a residual coding process. In the process, each transform block is further partitioned into non-overlapped subblocks and individual subblocks are coded one by one following a backward diagonal scan order (from highest frequencies toward lowest frequencies with decreasing indices) . In each subblock, transform coefficient levels are entropy coded using multiple subblock coding passes following a backward diagonal scan order. However, transform coefficients in a transform block are indexed in forward diagonal scan order with index 0 indicating the transform coefficient at the top-left transform block origin. Fig. 5 shows a diagonal scan order for a 4x4 block in forward scan direction.

[0011] In the first subblock coding pass, the syntax elements sig_coeff_flag, abs_level_gt1_flag, par_level_flag and abs_level_gt3_flag are all entropy coded in the regular mode. The elements abs_level_gt1_flag and abs_level_gt3_flag indicate whether the absolute value of the current coefficient level is greater than 1 and greater than 3, respectively. The syntax element par_level_flag indicates the party bit of the absolute value of the current level. The partially reconstructed absolute value of a transform coefficient level from the first pass is given by AbsLevelPass1 = sig_coeff_flag + par_level_flag + abs_level_gt1_flag + 2 * abs_level_gt3_flag

[0012] The context selection for entropy coding sig_coeff_flag is dependent on the state for the current coefficient. The syntax element par_level_flag is thus signaled in the first coding pass for deriving the state for the next coefficient. The syntax elements abs_remainder and coeff_sign_flag are further coded in the bypass mode in the second and third subblock coding passes, respectively, to indicate the remaining coefficient level values and signs, respectively. The fully reconstructed absolute value of a transform coefficient level is given by AbsLevel = AbsLevelPass1 + 2 *abs_remainder              (A)

[0013] The transform coefficient level is given by TransCoeffLevel = (2 *AbsLevel - (QState > 1 ? 1 : 0 ) ) * (1 -2 * coeff_sign_flag) ,                                    (B) where QState indicates the state for the current transform coefficient.

[0014] To control the worst-case bitstream parsing throughput rate, the number of regular bins for entropy coding the absolute values of transform coefficient levels in a transform block is constrained to be less than or equal to 1.75 regular bins per sample. At the beginning of the residual coding process, a counter keeping the remaining regular bin budget for remaining transform coefficients in a transform block is set equal to the transform block size multiplied by 1.75. The counter is decremented by one after consuming each regular bin for entropy coding any syntax element related to signalling the absolute level of a quantized transform coefficient in the first subblock coding pass. When the counter indicates the remaining regular bin budget is less than 4, the first sub-bock coding pass is skipped immediately, and the remaining absolute values of transform coefficient levels are entropy coded in the second subblock pass in the bypass mode.

[0015] The present invention is intended to further improve the performance of the sign coding for transform coefficients of residual data in a video coding system. BRIEF SUMMARY OF THE INVENTION

[0016] A method and apparatus of sign coding for transform coefficients of residual data in a video coding system are disclosed. According to this method, input data is received, wherein the input data comprises coded transform coefficients of residue data for a current block in a decoder side or the input data comprises transform coefficients of residue data for the current block at an encoder side, and wherein the current block comprises multiple residue blocks. For each of the multiple residue blocks: a maximum allowed number of predicted signs in said each of the multiple residue blocks is determined; a set of sign-predicted transform coefficients for said each of the multiple residue blocks is determined based on one or more specified rules, wherein a number of the set of sign-predicted transform coefficients is limited by the maximum allowed number of predicted signs; absolute values of transform coefficients in said each of the multiple residue blocks are encoded or decoded; signs of non-sign-predicted transform coefficients in said each of the multiple residue blocks are encoded or decoded; and entropy encoding or decoding is applied to the signs of the set of sign-predicted transform coefficients in said each of the multiple residue blocks using the sign prediction.

[0017] In one embodiment, said determining the constraint on the maximum allowed number of predicted signs in said each of the multiple residue blocks is dependent on regular bin budget for the current block or remaining regular bin budget before entropy coding any sign prediction residue corresponding to difference between a target sign of the set of sign-predicted transform coefficients and a corresponding sign prediction.

[0018] In one embodiment, the maximum allowed number of predicted signs in said each of the multiple residue blocks is set to a minimum value of a second maximum allowed number of predicted signs and the remaining regular bin budget, and wherein the second maximum allowed number of predicted signs is specified by a high-level syntax or other method.

[0019] In one embodiment, when the regular bin budget runs out before said entropy coding any sign prediction residue in said each of the multiple residue blocks, the sign prediction is disabled for the current block.BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Fig. 1A illustrates an exemplary adaptive Inter / Intra video encoding system incorporating loop processing.

[0021] Fig. 1B illustrates a corresponding decoder for the encoder in Fig. 1A.

[0022] Fig. 2 illustrates that a CU can be split into smaller CUs using one of the five split types (quad-tree partitioning, vertical binary tree partitioning, horizontal binary tree partitioning, vertical centre-side triple-tree partitioning, and horizontal centre-side triple-tree partitioning) .

[0023] Fig. 3 illustrates an example of a CTU being recursively partitioned by QT with the nested MTT.

[0024] Fig. 4 illustrates an exemplary block diagram of the CABAC process.

[0025] Fig. 5 illustrates an example of collecting the signs of first Nsp coefficients in a forward subblock-wise diagonal scan order starting from the DC coefficient according to one embodiment of the present invention.

[0026] Fig. 6 illustrates the cost function calculation to derive a best sign prediction hypothesis for a residual transform block according to Enhanced Compression Model 11 (ECM 11) .

[0027] Fig. 7 illustrates a flowchart of an exemplary video coding system that applies a residue process, including entropy coding a collection of coefficients for jointly sign coding, to multiple residue blocks in a transform block to avoid the need for revisiting the multiple residue blocks for coding signs of the collection of coefficients according to an embodiment of the present invention.DETAILED DESCRIPTION OF THE INVENTION

[0028] It will be readily understood that the components of the present invention, as generally described and illustrated in the figures herein, may be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the systems and methods of the present invention, as represented in the figures, is not intended to limit the scope of the invention, as claimed, but is merely representative of selected embodiments of the invention. References throughout this specification to “one embodiment, ” “an embodiment, ” or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. Thus, appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment.

[0029] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. One skilled in the relevant art will recognize, however, that the invention can be practiced without one or more of the specific details, or with other methods, components, etc. In other instances, well-known structures, or operations are not shown or described in detail to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by reference to the drawings, wherein like parts are designated by like numerals throughout. The following description is intended only by way of example, and simply illustrates certain selected embodiments of apparatus and methods that are consistent with the invention as claimed herein.

[0030] Joint Video Expert Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 are currently in the process of exploring the next-generation video coding standard. Some promising new coding tools have been adopted into Enhanced Compression Model 11 (ECM 11) (M. Coban, et al., “Algorithm description of Enhanced Compression Model 11 (ECM 11) , ” Joint Video Expert Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 32nd Meeting, Hannover, DE, 13–20 October 2023, Doc. JVET-AF2025) to further improve VVC. The adopted new tools have been implemented in the reference software ECM-11.0 [3] . Particularly, a new method for jointly predicting a collection of signs of transform coefficient levels in a residual transform block has been developed. In ECM 11, to derive a best sign prediction hypothesis for a residual transform block, a cost function is defined as discontinuity measure across block boundaries shown on Fig. 6. The cost function is calculated as a sum of absolute second derivatives in the residual domain for the above row and left column as follows:

[0031] In the above equation, R is reconstructed neighbours, P is prediction of the current block, and r is the residual hypothesis. The allowed maximum number of the predicted signs Nsp for each sign prediction hypothesis in a transform block is signalled in the sequence parameter set (SPS) and is constrained to be less than or equal to 8 in ECM-11.0.

[0032] When the number of jointly predicted signs is equal to N, a sign hypothesis can be indicated by a bin string of length N with each bin indicating a predicted sign for an associated transform coefficient. 2N different sign hypotheses can be formed by different combinations of predicted signs. The cost function is measured for all residual hypotheses (each derived from a sign prediction hypothesis) , and the sign hypothesis leading to the smallest cost is selected as a predictor for coefficient signs.

[0033] In ECM-11.0, a transform block area allowed for sign prediction is further specified by a SPS syntax element, corresponding to a top-left (lowest-frequency) MxM subblock region in a current transform block with M indicating region width and height. The number of predicted sign, Nsp, in a current transform block is determined by Min (Nmax, numNonzeroCoeffs ) , where the variable numNonzeroCoeffs indicates the number of nonzero transform coefficients excluding transform coefficients with data hiding being applied in the specified sign prediction region of the current transform block, and the function Min (x, y ) returns a smaller value between x and y. The signs of first Nsp nonzero coefficients are collected and predicted according to a raster-scan order from the specified sign prediction area. For those transform coefficients subject to sign prediction, a residual bin of sign prediction, indicating whether the sign of a current transform coefficient is equal to the predicted sign by the selected sign hypothesis, is entropy coded in the regular mode. Context selection for entropy coding the sign prediction residue of a transform coefficient is determined by whether the absolute value of the transform coefficient is greater than one or not. Transform blocks from CUs coded in inter and intra prediction modes are assigned separate sets of contexts. Luma and chroma transform blocks are further assigned separate sets of contexts. For those other transform coefficients not being assigned for sign prediction, the corresponding signs are just entropy coded in the bypass mode.

[0034] In ECM-11.0, when sign prediction is enabled for a transform block, the top-left 32x32 transform block region may be applicable for sign prediction. As a result, when a current transform block may be allowed for sign prediction, the subblock coding pass for entropy coding signs of the nonzero transform coefficients is skipped for each of subblocks within the top-left 32x32 transform block region during the residual coding process. After entropy coding all the transform blocks (by the residual coding process individually) , a separate sign coding process is added for signalling sign information in all the subblocks skipping the subblock sign coding pass in the prior residual coding process. In this way, a video coder needs to revisit individual transform blocks for sign coding after entropy coding each of transform blocks by the residual coding process.

[0035] In the present invention, new methods are disclosed for improving prediction of transform coefficient signs in implementation complexity. In one method, a video coding coder comprises entropy coding a plurality of residual blocks in a current CU, wherein each of the plurality of residual blocks is coded by a residual coding process. The video coder further comprises a scheme for predicting one or more signs of transform coefficients jointly in a residual block. For example, sign prediction can be achieved by comparing block boundary matching costs among different sign hypotheses. The said residual coding process may further comprise determining a constraint on the maximum allowed number of predicted signs and determining the number of predicted signs for a current transform block. The residual coding process may further comprise determining a collection of transform coefficients subject to sign prediction in the current transform block based on some specified rule. The residual coding process further comprises entropy coding sign information for the collection of transform coefficients in the current residual block based on sign prediction. For example, a residual coding process may further comprise entropy coding sign prediction residual bins in a transform block, wherein each of sign prediction residual bins indicates whether sign prediction for an associated transform coefficient in a sign hypothesis is correct or not. In this way, the separate sign coding process in ECM-11.0 for entropy coding sign prediction residual bins may no longer be needed. A bitstream parser does not need to revisit any transform block for parsing syntax elements related to deriving signs of transform coefficients after entropy coding individual transform blocks by a residual process in a CU.

[0036] In some embodiments, a residual coding process in a video coder may further comprise some constraint on the number of consumed CABAC regular bins for entropy coding a transform block or transform unit. In the video coder, some sign information based on sign prediction such as sign prediction residues are entropy coded in the regular mode. The consumed regular bins for entropy coding sign information are also counted by the regular bin counter under the constraint on the regular bin budget assigned for a current transform block or transform unit. The proposed method may further comprise determining the maximum allowed number of the predicted signs in the transform block further depending on the regular bin budget for the transform block or the remaining regular bin budget before entropy coding any sign prediction residues. For example, a video coder may signal sign information based on sign prediction by entropy coding sign prediction residual bins and the maximum allowed number of the predicted signs in a current transform block may set equal to Min (maxPredSigns, remainingRegBins ) , where the variable maxPredSigns indicates the maximum allowed number of predicted signs specified by a high-level syntax or other method, and the variable remainingRegBins indicates the remaining regular bin budget before entropy coding any sign prediction residues in the current transform block. When the regular bin budget is run out before entropy coding any sign prediction residues in a current transform block, sign prediction can be disabled for the current transform block.

[0037] In some embodiments, each residual block is further divided into one or more non-overlapped subblocks and the residual coding process comprises entropy coding individual subblocks in a current transform block according to a specified subblock scan order. The residual coding process may further comprise one or more subblock coding passes for entropy coding information about the absolute values of the transform coefficient levels in a current subblock. For a nonzero transform coefficient subject to sign prediction, a syntax flag is entropy coded in the regular mode to indicate a sign prediction residue for an associated transform coefficient in a sign hypothesis. For a nonzero transform coefficient neither subject to sign data hiding nor subject to sign prediction, a syntax flag is entropy coded in the bypass mode to indicate a sign of an associated transform coefficient.

[0038] In some embodiments, the residual coding process further comprises a mixed sign coding pass for entropy coding sign information in a transform block after entropy coding significances or absolute values of all transform coefficients in a transform block, wherein the mixed sign coding pass comprises entropy coding signs or sign prediction residues of individual nonzero transform coefficients following a specified scan order. In one embodiment, entropy coding signs or sign prediction residues of nonzero transform coefficients follows a forward raster scan order from the block origin. The signs of first Nsp nonzero transform coefficients are subject to sign prediction and the residues of sign prediction are entropy coded in the regular mode. The signs of remaining nonzero transform coefficients are entropy coded in the bypass mode. In another embodiment, the mixed sign coding pass just follows the same scan order used for entropy coding the absolute values of the individual transform coefficient levels, and proceeds in either forward or backward scan direction in a transform block. When the sign coding pass proceeds in a backward scan direction from the highest frequency toward the lowest frequency (with decreasing scan indices) , the last N nonzero transform coefficients, excluding transform coefficients with sign data hiding being applied, are subject to sign prediction.

[0039] In some embodiments, the residual coding process may further comprise a subblock sign coding pass for entropy coding sign information in a transform subblock after entropy coding significances or absolute values of all transform coefficients in a transform subblock, wherein the subblock sign coding pass comprises entropy coding signs of individual nonzero transform coefficients following a specified scan order. In some embodiments, the residual coding process comprises determining a sign prediction residue region. When a subblock is outside of the sign residue region, signs of the individual nonzero transform coefficients in the subblock are entropy coded in the bypass mode by a subblock sign coding pass. The subblock sign coding pass may follow the second subblock pass for entropy coding remaining absolute values of transform coefficient levels according to the same scan order. The residual coding process may further comprise the mixed sign coding pass for entropy coding sign information in the sign prediction residue region after entropy coding significances or absolute values of all transform coefficients in the sign prediction residue region. The mixed sign coding pass comprises entropy coding signs or sign prediction residues of individual nonzero transform coefficients in a block region following a specified scan order as described previously. In this way, all sign prediction residues are entropy coded in the specified sign prediction residue region. Signs of transform coefficients can be entropy coded in the bypass mode immediately after the subblock coding passes for entropy coding absolute values of transform coefficients in each of subblocks not subject to sign prediction, similar to the conventional method used in VVC.

[0040] In some embodiments, the sign prediction residue region is fixed to be the maximum supported sign prediction area by a video coding system, corresponding to the top-left 32x32 block region in ECM-11.0. In some embodiments, the sign prediction residue region is set to be the maximum allowed sign prediction area specified by one or more high-level syntax elements. In some embodiments, the sign prediction residue region may be smaller than the maximum allowed sign prediction area. For example, the sign prediction residue region can comprise the R lowest-frequency subblocks in a specified scan order or the subblocks in the top-left MxN subblock region in a transform block, where the constraint parameters R or M and N can either be pre-defined by a video coding system or additionally signalled by using one or more high-level syntax elements. In some preferred embodiments, the sign prediction residue region is the top-left lowest-frequency subblock. In this scheme, the sign of a nonzero transform coefficients outside the sign prediction residue region but inside the sign prediction region can be mapped to a nonzero transform coefficient within the sign prediction residue region and entropy coded with sign prediction. The maximum allowed number of predicted signs in a current transform block is further constrained to be less than or equal to the total number of nonzero transform coefficients, excluding transform coefficients with sign data hiding being applied, in the sign prediction residue region of the current transform block.

[0041] In ECM-11.0, all nonzero transform coefficients, excluding transform coefficients with sign data hiding being applied, from a specified sign prediction area in a transform block are sorted according to a decreasing order of the absolute values of transform coefficient levels. The signs of first Nsp transform coefficients in the sorted list are subject to sign prediction. The best sign hypothesis is selected from all sign hypotheses based on the block boundary matching costs calculated by eqn. (1) . The signs of transform coefficients in the sorted list are mapped to the nonzero transform coefficients, excluding transform coefficients with sign data hiding being applied, in the sign prediction area one by one in a raster scan order.

[0042] The mapped signs of first Nsp transform coefficients in a raster scan order are subject to sign prediction. Each of sign prediction residues of the first Nsp transform coefficients in the raster scan order are entropy coded in the regular mode with context selection depending on the absolute level of the associated transform coefficients after mapping. The mapped signs of the remaining transform coefficients in the sign prediction area are entropy coded in the bypass mode following the raster scan order. For reconstruction of transform coefficients in the sign prediction area, the mapped signs of the first Nsp nonzero transform coefficients in the scan order are recovered from decoded sign prediction residues and the best sign prediction hypothesis. The mapped signs of the remaining nonzero transform coefficients are set to corresponding decoded signs. The signs of nonzero transform coefficients in the sign prediction area are derived by inverse mapping the decoded signs to originally associated transform coefficients. In this scheme, a video coder may need to maintain a mapping table, wherein each of coded signs shall need a mapping table entry for storing the position of a transform coefficient originally associated with a coded sign (before sign mapping) . The video coder needs to perform inverse sign mapping for each of nonzero transform coefficients in the sign prediction area.

[0043] In the proposed method, a video coder may first find the Nsp leading transform coefficients in a sorted list of nonzero transform coefficients, excluding coefficients with sign data hiding being applied, from a specified sign prediction area in a transform block according to a decreasing order of the absolute values of transform coefficient levels. The signs of the first Nsp transform coefficients in the sorted list are mapped to the first Nsp nonzero transform coefficients (excluding transform coefficients with sign data hiding being applied) in the sign prediction area one by one according to the specified scan order.

[0044] The mapped signs of first Nsp transform coefficients in a specified scan order are subject to sign prediction. Each of sign prediction residues of the first Nsp transform coefficients in a raster scan order are entropy coded in the regular mode with context selection depending on the absolute level of the associated transform coefficients after mapping. When there are K transform coefficients among the first Nsp nonzero transform coefficients in the specified scan order are not included in the first Nsp transform coefficients in the sorted list, the original signs of these K transform coefficients are not coded with the first Nsp transform coefficients in the scan order. The original signs of these K transform coefficients are then mapped to the K transform coefficients in the sorted list corresponding to K highest scanning indices. The signs of other remaining nonzero transform coefficients in the sign prediction area are not mapped. Those transform coefficients not subject to sign prediction are entropy coded in the bypass mode. For reconstruction of transform coefficients in the sign prediction area, the signs of the first Nsp nonzero transform coefficients in the scan order are recovered from decoded sign prediction residues and predicted signs. For those transform coefficients coded with mapped signs, the original signs of those transform coefficients can be recovered by inverse sign mapping. In this way, a video coder only needs a small mapping table with up to Nsp entries for storing the positions of the first Nsp transform coefficients in the sorted list. The video coder needs to perform sign mapping for up to 2·Nsp transform coefficients associated with mapped signs and does not need to perform any sign mapping for all other nonzero transform coefficients in the sign prediction area.

[0045] Any of the foregoing proposed methods of initial residual hypothesis derivation can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / IBC / prediction / transform module of an encoder, and / or an inter / intra / IBC / prediction / transform module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / IBC / prediction / transform module of the encoder and / or the inter / intra / IBC / prediction / transform module of the decoder, so as to provide the information needed by the inter / intra / IBC / prediction / transform module.

[0046] With reference to the exemplary encoder and decoder in Fig. 1A and Fig. 1B, any of the proposed methods of applying residue coding process, including entropy sign coding for a collection of sign-predicted coefficients, to multiple residue blocks can be implemented in an Intra / Inter / Entropy coding module (e.g. Intra Pred. 150 / MC 150 / Entropy Decoder 140 in Fig. 1B) in a decoder or an Intra / Inter / Entropy coding module (e.g. Intra Pred. 110 / Inter Pred. 112 / Entropy Encoder 122 in Fig. 1A) in an encoder. Any of the proposed methods can also be implemented as a circuit coupled to the intra / inter coding module at the decoder or the encoder. However, the decoder or encoder may also use additional processing unit to implement the required processing. While the Intra Pred. units (e.g. unit 110 in Fig. 1A and unit 150 in Fig. 1B) are shown as individual processing units, they may correspond to executable software or firmware codes stored on a media, such as hard disk or flash memory, for a CPU (Central Processing Unit) or programmable devices (e.g. DSP (Digital Signal Processor) or FPGA (Field Programmable Gate Array) ) .

[0047] Fig. 7 illustrates a flowchart of an exemplary video coding system that applies a residue process, including entropy coding a collection of coefficients for jointly sign coding, to multiple residue blocks in a transform block to avoid the need for revisiting the multiple residue blocks for coding signs of the collection of coefficients according to an embodiment of the present invention. The steps shown in the flowchart may be implemented as program codes executable on one or more processors (e.g., one or more CPUs) at the encoder side. The steps shown in the flowchart may also be implemented based hardware such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to this method, input data is received in step 710, wherein the input data comprises coded transform coefficients of residue data for a current block in a decoder side or the input data comprises transform coefficients of residue data for the current block at an encoder side, and wherein the current block comprises multiple residue blocks. In step 720: for each of the multiple residue blocks: a maximum allowed number of predicted signs in said each of the multiple residue blocks is determined; a set of sign-predicted transform coefficients for said each of the multiple residue blocks is determined based on one or more specified rules, wherein a number of the set of sign-predicted transform coefficients is limited by the maximum allowed number of predicted signs; absolute values of transform coefficients in said each of the multiple residue blocks are encoded or decoded; signs of non-sign-predicted transform coefficients in said each of the multiple residue blocks are encoded or decoded; and entropy encoding or decoding is applied to the signs of the set of sign-predicted transform coefficients in said each of the multiple residue blocks using the sign prediction.

[0048] The flowchart shown is intended to illustrate an example of video coding according to the present invention. A person skilled in the art may modify each step, re-arranges the steps, split a step, or combine steps to practice the present invention without departing from the spirit of the present invention. In the disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present invention. A skilled person may practice the present invention by substituting the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present invention.

[0049] The above description is presented to enable a person of ordinary skill in the art to practice the present invention as provided in the context of a particular application and its requirement. Various modifications to the described embodiments will be apparent to those with skill in the art, and the general principles defined herein may be applied to other embodiments. Therefore, the present invention is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present invention. Nevertheless, it will be understood by those skilled in the art that the present invention may be practiced.

[0050] Embodiment of the present invention as described above may be implemented in various hardware, software codes, or a combination of both. For example, an embodiment of the present invention can be one or more circuit circuits integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. An embodiment of the present invention may also be program code to be executed on a Digital Signal Processor (DSP) to perform the processing described herein. The invention may also involve a number of functions to be performed by a computer processor, a digital signal processor, a microprocessor, or field programmable gate array (FPGA) . These processors can be configured to perform particular tasks according to the invention, by executing machine-readable software code or firmware code that defines the particular methods embodied by the invention. The software code or firmware code may be developed in different programming languages and different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles and languages of software codes and other means of configuring code to perform the tasks in accordance with the invention will not depart from the spirit and scope of the invention.

[0051] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is therefore, indicated by the appended claims rather than by the foregoing description. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

1.A method of video coding, the method comprising:receiving input data, wherein the input data comprises coded transform coefficients of residue data for a current block in a decoder side or the input data comprises transform coefficients of residue data for the current block at an encoder side, and wherein the current block comprises multiple residue blocks;for each of the multiple residue blocks:determining a maximum allowed number of predicted signs in said each of the multiple residue blocks;determining a set of sign-predicted transform coefficients for said each of the multiple residue blocks based on one or more specified rules, wherein a number of the set of sign-predicted transform coefficients is limited by the maximum allowed number of predicted signs;encoding or decoding absolute values of transform coefficients in said each of the multiple residue blocks;encoding or decoding signs of non-sign-predicted transform coefficients in said each of the multiple residue blocks; andapplying entropy encoding or decoding to the signs of the set of sign-predicted transform coefficients in said each of the multiple residue blocks using the sign prediction.2.The method of Claim 1, wherein said determining the constraint on the maximum allowed number of predicted signs in said each of the multiple residue blocks is dependent on regular bin budget for the current block or remaining regular bin budget before entropy coding any sign prediction residue corresponding to difference between a target sign of the set of sign-predicted transform coefficients and a corresponding sign prediction.3.The method of Claim 2, wherein the maximum allowed number of predicted signs in said each of the multiple residue blocks is set to a minimum value of a second maximum allowed number of predicted signs and the remaining regular bin budget, and wherein the second maximum allowed number of predicted signs is specified by a high-level syntax or other method.4.The method of Claim 2, wherein when the regular bin budget runs out before said entropy coding any sign prediction residue in said each of the multiple residue blocks, the sign prediction is disabled for the current block.5.An apparatus of video coding, the apparatus comprising one or more electronic circuits or processors arranged to:receive input data, wherein the input data comprises coded transform coefficients of residue data for a current block in a decoder side or the input data comprises transform coefficients of residue data for the current block at an encoder side, and wherein the current block comprises multiple residue blocks;for each of the multiple residue blocks:determine a maximum allowed number of predicted signs in said each of the multiple residue blocks;determine a set of sign-predicted transform coefficients for said each of the multiple residue blocks based on one or more specified rules, wherein a number of the set of sign-predicted transform coefficients is limited by the maximum allowed number of predicted signs;encode or decode absolute values of transform coefficients in said each of the multiple residue blocks;encode or decode signs of non-sign-predicted transform coefficients in said each of the multiple residue blocks; andapply entropy encoding or decoding to the signs of the set of sign-predicted transform coefficients in said each of the multiple residue blocks using the sign prediction.

Citation Information

Patent Citations

  • Signaling residual signs predicted in transform domain

    CN111819853A

  • Transforms and Sign Prediction

    US20240040122A1

  • Sign prediction for block-based video coding

    WO2023023039A1

  • Video signal processing method and device therefor

    WO2023080691A1

  • Method and apparatus for sign coding of transform coefficients in video coding system

    WO2023103521A1