Video coding method and device and video decoding method and device
By using joint prediction and contextual coding techniques, the symbol encoding and decoding of the transform coefficients of the residual block are optimized, which solves the problem of low symbol encoding and decoding efficiency in VVC and achieves more efficient video encoding and decoding performance.
Patent Information
- Application Number
- CN202511182187.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2022-12-08
- Publication Date
- 2025-12-09
AI Technical Summary
Existing video coding and decoding technologies suffer from inefficiency in transform coefficient coding and decoding of residual blocks, especially in the high-performance video coding and decoding standard VVC, where the performance of symbol coding and decoding needs to be improved.
A joint prediction method is adopted to encode and decode the symbols of the transform coefficients of the residual block. By determining the maximum allowable number and selecting the hypothesis that achieves the minimum cost from multiple hypotheses for symbol prediction, the symbol prediction process is optimized by combining contextual encoding and decoding techniques.
It improves the performance of transform coefficient symbol encoding and decoding, thereby enhancing the compression efficiency and quality of the video encoding and decoding system.
Smart Images

Figure CN121099050A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to video coding systems. In particular, the present invention relates to the coding of signs of transform coefficients of residual blocks in video coding systems.
BACKGROUND
[0002] Versatile Video Coding (VVC) is the latest international video coding standard developed by the Joint Video Expert Team (JVET) of ITU-T Video Coding Experts Group (VCEG) and ISO / IEC Moving Picture Experts Group (MPEG). The standard has been published as an ISO standard: ISO / IEC 23090-3:2021, Information technology - Coding of audio-visual objects - Part 3: Versatile Video Coding, published in February 2021. VVC is developed based on its predecessor, HEVC (High Efficiency Video coding), with the addition of more coding tools to improve coding efficiency and handle various video sources including 3-dimensional (3D) video signals.
[0003] Figure 1A An exemplary adaptive inter / intra video decoding system incorporating loop processing is illustrated. For intra prediction, prediction data is derived based on previously coded video data in the current picture. For inter prediction 112, Motion Estimation (ME) is performed at the encoder side and Motion Compensation (MC) is performed based on the results of ME to provide prediction data derived from other pictures and motion data. A switch 114 selects either intra prediction 110 or inter prediction 112, and the selected prediction data is provided to a summer 116 to form prediction error, also referred to as residual. The prediction error is then processed by a transform (T) 118 and subsequent quantization (Q) 120. The transformed and quantized residual is then encoded by an entropy coder 122 to be included in a video bitstream corresponding to compressed video data. The bitstream associated with the transform coefficients is then packed with auxiliary information such as motion and coding mode associated with intra and inter prediction, and other information such as parameters associated with in-loop filters applied to underlying image regions, and so on. As Figure 1AAs shown, side information associated with intra-frame prediction 110, inter-frame prediction 112, and loop filter 130 is provided to entropy encoder 122. When using inter-frame prediction mode, the reference picture must also be reconstructed at the encoder. Therefore, the residuals from transform and quantization are processed by inverse quantization (IQ) 124 and inverse transform (IT) 126 to recover the residuals. Then, in reconstruction (REC) 128, the residuals are added back to the prediction data 136 to reconstruct the video data. The reconstructed video data can be stored in the reference picture buffer 134 and used for prediction of other frames.
[0004] like Figure 1A As shown, the input video data undergoes a series of processes in the encoding system. Due to these processes, the reconstructed video data from REC 128 may suffer various forms of degradation. Therefore, a loop filter 130 is often applied to the reconstructed video data to improve video quality before storing the reconstructed video data in the reference picture buffer 134. For example, a deblocking filter (DF), sample adaptive offset (SAO), and adaptive loop filter (ALF) can be used. It may be necessary to incorporate the loop filter information into the bitstream so that the decoder can correctly recover the required information. Therefore, the loop filter information is also provided to the entropy encoder 122 to be incorporated into the bitstream. Figure 1A In the process, before storing the reconstructed sample in the reference image buffer 134, the loop filter 130 is applied to the reconstructed video. Figure 1A The system described herein is intended to illustrate an exemplary architecture of a typical video encoder. It may correspond to a High Efficiency Video Decoding (HEVC) system, VP8, VP9, H.264, or VVC.
[0005] Figure 1B Another example of a decoding system is shown. For example... Figure 1B As shown, the decoder can use function blocks similar to or partially identical to the encoder, except for transform 118 and quantization 120, since the decoder only needs inverse quantization 124 and inverse transform 126. Instead of entropy encoder 122, the decoder uses entropy decoder 140 to decode the video bitstream into quantized transform coefficients and the required decoding information (e.g., ILPF information, intra-frame prediction information, and inter-frame prediction information). Intra-frame prediction 150 on the decoder side does not require mode search. Instead, the decoder only needs to generate intra-frame predictions based on the intra-frame prediction information received from entropy decoder 140. Furthermore, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the intra-frame prediction information received from entropy decoder 140, without motion estimation.
[0006] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by associated coding tree units (CTUs). A coded picture can be represented by a set of slices, each including an integer number of CTUs. Each CTU in a slice is processed in a raster scan order. Bi-predictive (B) slices are decoded using intra prediction or inter prediction with up to two motion vectors and reference indices to predict sample values of each block. Predictive (P) slices are decoded using intra prediction or inter prediction with up to one motion vector and reference index to predict sample values of each block. Intra (I) slices are decoded using only intra prediction.
[0007] A CTU can be partitioned into one or more non-overlapped coding units (CUs) using quad-tree (QT) with nested multi-type-tree (MTT) structure to adapt to various local motion and texture characteristics. A CU can be further split into smaller CUs using one of the five split types shown in FIG. 2 (quad-tree partitioning 210, vertical binary tree partitioning 220, horizontal binary tree partitioning 230, vertical center-side triple-tree partitioning 240, horizontal center-side triple-tree partitioning 250). Figure 2 A CU can be further split into smaller CUs using one of the five split types shown in FIG. 2 (quad-tree partitioning 210, vertical binary tree partitioning 220, horizontal binary tree partitioning 230, vertical center-side triple-tree partitioning 240, horizontal center-side triple-tree partitioning 250). Figure 3An example of a QT recursive partitioning of a CTU with nested MTTs is provided. Each CU contains one or more prediction units (PUs). The prediction units together with the associated CU syntax serve as the basic unit for signaling prediction sub-information. A specified prediction process is used to predict the values of the related pixel samples within a PU. Each CU can contain one or more transform units (TUs) for representing a prediction residual block. A transform unit (TU) includes a transform block (TB) of luma samples and two corresponding transform blocks of chroma samples, and each TB corresponds to one residual sample block from one color component. An integer transform is applied to the transform block. The level values of the quantized coefficients are entropy coded in the bitstream together with other side information. The terms coding tree block (CTB), coding block (CB), prediction block (PB), and transform block (TB) are defined as a two-dimensional array of samples of one color component associated with a CTU, CU, PU, and TU, respectively. Thus, one CTU consists of one luma CTB, two chroma CTBs, and related syntax elements. Similar relationships apply to CUs, PUs, and TUs.
[0008] To achieve high compression efficiency, the values of the syntax elements in HEVC and VVC are entropy coded using a context-based adaptive binary arithmetic coding (CABAC) mode, or regular mode. Figure 4 An exemplary block diagram of the CABAC process is illustrated. Since the arithmetic coder in the CABAC engine can only code binary symbol values, the CABAC process needs to convert the values of the syntax elements into a binary string using a binarizer (410). The conversion process is commonly referred to as binarization. During the coding process, the probability models are gradually built up from the coded symbols of different contexts. A context modeler (420) is used for modeling purposes. During the coding process based on the normal context, a regular coding engine (430) corresponding to the binary arithmetic coder is used. The selection of the modeling context for the next binary symbol can be determined from the coded information. The symbols can also be coded without a context modeling stage and assume an equal probability distribution, commonly referred to as bypass mode, to reduce complexity. For the symbols that are bypassed, a bypass coding engine (440) can be used. As Figure 4 shown, switches (S1, S2, and S3) are used to direct the data flow between the regular CABAC mode and the bypass mode. When the regular CABAC mode is selected, the switches are toggled to the upper contact. When the bypass mode is selected, the switches are flipped to the lower contact, as Figure 4 shown.
[0009] In VVC, a transform coefficient can be quantized using a dependent scalar quantization. The selection of one of two quantizers is determined by a state machine with four states. The state of a current transform coefficient is determined by the state of the absolute level value and the parity of the previous transform coefficient in the scan order. A transform block is partitioned into non-overlapping sub-blocks. The transform coefficient levels in each sub-block are entropy coded using multiple sub-block coding passes. The syntax elements sig_coeff_flag, abs_level_gt1_flag, par_level_flag, and abs_level_gt3_flag are coded in the first sub-block coding pass in regular mode. The elements abs_level_gt1_flag and abs_level_gt3_flag indicate whether the absolute value of the current coefficient level is greater than 1 and 3, respectively. The syntax element par_level_flag represents the parity bit of the absolute value of the current level. The partially reconstructed absolute value of the transform coefficient levels in the first pass is given by
[0010] AbsLevelPass1 = sig_coeff_flag + par_level_flag + abs_level_gt1_flag + 2 * abs_level_gt3_flag. (1)
[0011] The context selection for entropy coding sig_coeff_flag depends on the state of the current coefficient. Therefore, par_level_flag is signaled in the first coding pass for deriving the state of the next coefficient. The syntax elements abs_remainder and coeff_sign_flag are further coded in the subsequent sub-block coding passes in bypass mode to indicate the coefficient level value and the sign of the residual, respectively. The fully reconstructed absolute value of the transform coefficient levels is given by
[0012] AbsLevel = AbsLevelPass1 + 2 * abs_remainder. (2)
[0013] The transform coefficient level is given by
[0014] TransCoeffLevel = (2 * AbsLevel - (QState > 1? 1 : 0)) * (1 - 2 * coeff_sign_flag), (3)
[0015] where QState represents the state of the current transform coefficient.
[0016] The present application aims to further improve the performance of coding transform coefficients of residual data in a video coding system. SUMMARY
[0017] In view of the above, the present application provides the following technical solutions:
[0018] A video decoding method is provided, comprising: receiving coded transform coefficients of a residual block corresponding to a current block, wherein the coded transform coefficients include coded sign residues associated with a set of jointly predicted coefficient signs; determining a maximum allowed number according to a coding context associated with the current block; decoding the coded sign residues into decoded sign residues, wherein a total number of the set of jointly predicted coefficient signs is equal to or less than the maximum allowed number; determining a joint sign prediction of the set of jointly predicted coefficient signs by selecting one of a set of hypotheses of the set of jointly predicted coefficient signs corresponding to a hypothesis that achieves a minimum cost, wherein the cost between a boundary pixel of the current block and a corresponding neighboring pixel of the current block for each hypothesis of the set of hypotheses is calculated separately, and the boundary pixel of the current block associated with each hypothesis is reconstructed using information including the each hypothesis; and reconstructing the set of jointly predicted coefficient signs based on the decoded sign residues and the joint sign prediction.
[0019] A video decoding apparatus is provided, comprising one or more electronic circuits or processors configured to: receive coded transform coefficients of a residual block corresponding to a current block, wherein the coded transform coefficients include coded sign residues associated with a set of jointly predicted coefficient signs; determine a maximum allowed number according to a coding context associated with the current block; decode the coded sign residues into decoded sign residues, wherein a total number of the set of jointly predicted coefficient signs is equal to or less than the maximum allowed number; determine a joint sign prediction of the set of jointly predicted coefficient signs by selecting one of a set of hypotheses of the set of jointly predicted coefficient signs corresponding to a hypothesis that achieves a minimum cost, wherein the cost between a boundary pixel of the current block and a corresponding neighboring pixel of the current block for each hypothesis of the set of hypotheses is calculated separately, and the boundary pixel of the current block associated with each hypothesis is reconstructed using information including the each hypothesis; and reconstruct the set of jointly predicted coefficient signs based on the decoded sign residues and the joint sign prediction.
[0020] The present application provides a video encoding method, the method comprising: receiving transform coefficients corresponding to a residual block of a current block; determining a maximum allowed number according to a coding context associated with the current block; determining a set of jointly predicted coefficient signs associated with a set of selected transform coefficients, wherein a total number of the set of jointly predicted coefficient signs is equal to or less than the maximum allowed number; determining a joint sign prediction of the set of jointly predicted coefficient signs by selecting one of a set of hypotheses of the set of jointly predicted coefficient signs corresponding to a hypothesis that achieves a minimum cost, wherein the cost between a boundary pixel of the current block and a corresponding neighboring pixel of the current block for each hypothesis of the set of hypotheses is calculated separately using information including the each hypothesis to reconstruct the boundary pixel of the current block associated with the each hypothesis; determining a sign residual between the set of jointly predicted coefficient signs and the joint sign prediction; and applying a context coding to the sign residual to generate a coded sign residual.
[0021] The present application also provides a video encoding apparatus, the apparatus comprising one or more electronic circuits or processors configured to: receive transform coefficients corresponding to a residual block of a current block; determine a maximum allowed number according to a coding context associated with the current block; determine a set of jointly predicted coefficient signs associated with a set of selected transform coefficients, wherein a total number of the set of jointly predicted coefficient signs is equal to or less than the maximum allowed number; determine a joint sign prediction of the set of jointly predicted coefficient signs by selecting one of a set of hypotheses of the set of jointly predicted coefficient signs corresponding to a hypothesis that achieves a minimum cost, wherein the cost between a boundary pixel of the current block and a corresponding neighboring pixel of the current block for each hypothesis of the set of hypotheses is calculated separately using information including the each hypothesis to reconstruct the boundary pixel of the current block associated with the each hypothesis; determine a sign residual between the set of jointly predicted coefficient signs and the joint sign prediction; and apply a context coding to the sign residual to generate a coded sign residual.
[0022] The video encoding method and apparatus and the video decoding method and apparatus of the present application can improve the performance of transform coefficient sign coding. BRIEF DESCRIPTION OF DRAWINGS
[0023] The accompanying drawings incorporated in and forming a part of the specification illustrate embodiments of the present application and, together with the description, serve to explain the principles of the application:
[0024] Figure 1A An exemplary adaptive inter / intra video decoding system incorporating loop processing is shown.
[0025] Figure 1B Another exemplary decoding system is shown.
[0026] Figure 2 An illustration of a CU that can be split into smaller CUs using one of five split types is shown.
[0027] Figure 3 An example of a CTU recursively partitioned by QT with nested MTT is provided.
[0028] Figure 4 An example block diagram of the CABAC process is illustrated.
[0029] Figure 5 An illustration of a cost function calculation to derive the best sign prediction hypothesis for a residual transform block according to enhanced compression model 2 (ECM 2) is shown.
[0030] Figure 6 An example of collecting the signs of the first Nsp coefficients in a diagonal scan order in a forward sub-block manner starting from the DC coefficient according to one embodiment of the present application is illustrated.
[0031] Figure 7 A flowchart of an example video decoding system utilizing joint sign prediction according to an embodiment of the present application is shown.
[0032] Figure 8 A flowchart of an example video encoding system utilizing joint sign prediction according to an embodiment of the present application is shown.
DETAILED DESCRIPTION
[0033] It will be readily understood that the components of the application, as generally described and illustrated in the Figures herein, can be arranged and designed in a wide variety of different configurations. Thus, the following more detailed description of the embodiments of the system and method of the application, as represented in the Figures, is not intended to limit the scope of the application, as claimed, but is merely representative of selected embodiments of the application. The
[0034] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. However, those skilled in the art will recognize that the invention can be practiced without one or more specific details, or using other methods, components, etc. In other instances, well-known structures or operations are not shown or illustrated. Detailed description is provided to avoid obscuring aspects of the invention. The illustrated embodiments of the invention will be best understood by referring to the accompanying drawings, wherein like parts are represented by like numbers throughout. The following description is by way of example only and simply illustrates certain selected embodiments of devices and methods consistent with the invention as claimed herein.
[0035] The Joint Video Experts Group (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 is currently exploring next-generation video codec standards. Enhanced Compression Model 2 (ECM 2) incorporates several promising new codec tools (M. Coban et al., “Algorithm description of Enhanced Compression Model 2 (ECM 2)”, Joint Video Experts Group of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, 23rd meeting, teleconference, July 7-16, 2021, document JVET-W2025) to further improve VVC. These new tools have been implemented in the reference software ECM-2.0 (ECM Reference Software ECM-2.0, available at https: / / vcgit.hhi.fraunhofer.de / ecm / ECM [online]). In particular, a novel method has been developed for jointly predicting the symbol set at the transform coefficient level in residual transform blocks (JVET-D0031, Felix Henry et al., “Residual Coefficient Sign Prediction”, ITU-TSG16 WP3 and ISO / IEC JTC1 / SC29 / WG11 Joint Video Expert Group, 4th Meeting: Chengdu, China, October 15-21, 2016, document JVET-D0031). In ECM 2, to derive the optimal symbol prediction assumption for residual transform blocks, the cost function is defined across... Figure 5The discontinuity measure of the block boundaries is shown, where block 510 corresponds to the transform block, circle 520 corresponds to the neighboring sample, and circle 530 corresponds to the reconstructed sample associated with the sign candidate of block 510. The cost function is defined as the sum of the absolute second derivatives in the residual domain of the above rows and left columns, as follows:
[0036]
[0037] In the above equation, R is the reconstructed neighbor, P is the prediction of the current block, and r is the residual hypothesis. The maximum allowed number of predicted symbols N for each symbol prediction hypothesis in the transmit transform block within the Sequence Parameter Set (SPS) is... sp Furthermore, in ECM-2.0, it is restricted to less than or equal to 8. The cost function of all hypotheses is measured, and the one with the lowest cost is selected as the predictor with the lowest coefficient sign. Only coefficients from the top-left 4x4 transform sub-block region (coefficients with the lowest frequency) are allowed to be included in the hypothesis. The first N values are collected and encoded / decoded according to the raster scan order on the top-left 4x4 sub-block. sp The first N non-zero coefficients sp The symbols for non-zero coefficients are sent (if available). For those predicted coefficients, a sign prediction bin is sent, indicating whether the predicted symbol equals the selected hypothesis, without sending the coefficient symbol. This sign prediction bin is context-coded, where the selected context is derived based on whether the coefficient is DC. For intra-frame and inter-frame blocks, and for luma and chroma components, the context is separate. For other coefficients without sign prediction, the corresponding symbols are encoded and decoded by CABAC in bypass mode.
[0038] According to an aspect of the present disclosure, the context modeling for the entropy coding of the sign prediction bits of the current coefficient can be further conditioned on information about the absolute value of the current transform coefficient level. This is because coefficients with larger absolute level values have a larger impact on the output value of the cost function and thus tend to have a higher correct prediction rate. In the proposed method, the process of the video codec includes coding syntax information related to the sign of the transform coefficient level associated with the sign prediction using a plurality of context variables, wherein the selection of the context variable to code the sign of the current coefficient level can further depend on the absolute value of the current transform coefficient level. In some embodiments, the context selection for the entropy coding of the sign prediction bits of certain coefficients further depends on whether the absolute value of the current transform coefficient level is greater than or less than one or more threshold values. In one example, the context selection for the entropy coding of the sign prediction bits of certain coefficients further depends on whether the absolute value of the current transform coefficient level is greater than a first threshold value T1. In some preferred embodiments, T1 can be equal to 1, 2, 3, or 4. In another example, the context selection for the entropy coding of the sign prediction bits of certain coefficients further depends on whether the absolute value of the current transform coefficient level is greater than a second threshold value T2, where T2 is greater than T1. In some preferred embodiments, (T1, T2) can be equal to (1, 2), (1, 3), or (2, 4).
[0039] According to yet another aspect of the present disclosure, the process for the video codec can further include adaptively setting the values of the one or more threshold values considering the coding context of the current transform block. In some embodiments, the derivation of the one or more threshold values can further depend on the transform block dimension, the transform type, the color component index, the number of predicted signs, the number of non-zero coefficients, or the position of the last significant coefficient associated with the current transform block. The derivation of the one or more threshold values can further depend on the prediction mode of the current CU. The derivation of the one or more threshold values can further depend on the position or index associated with the current coefficient in the transform block. The derivation of the one or more threshold values can further depend on the sum of the absolute values of the sign predicted coefficients in the current transform block.
[0040] According to another aspect of the present disclosure, the context modeling for the entropy coding of the sign prediction bits of the current coefficient can be further conditioned on information derived from the absolute values of the current coefficient level and other coefficient levels in the current transform block. In some embodiments, the context selection for the entropy coding of the sign of the coefficients in the current transform block can further depend on the sum of the absolute values of the sign predicted coefficients in the current transform block. In some embodiments, the context selection for the entropy coding of the sign of the coefficients in the current transform block can further depend on the absolute value of the next coefficient or the sum of the absolute values of the remaining coefficients that are sign predicted in the current transform block.
[0041] In some proposed embodiments, the context selection depending on the absolute coefficient level can be employed only for a specified number of transform coefficients. When the current coefficient does not belong to the specified number of transform coefficients, the context selection for the current coefficient is independent of the absolute coefficient level. In some embodiments, the specified number of transform coefficients corresponds to the first Nl coefficients associated with the sign prediction according to a predefined scan order in the transform block. When the current coefficient does not belong to the first Nl coefficients, the context selection is independent of the absolute coefficient level. In some embodiments, the predefined order is the entropy coding order of the sign prediction bins. In some embodiments, Nl is equal to 1, 2, 3, or 4. In some embodiments, the specified group of transform coefficients corresponds to coefficients from a transform coefficient region or a scan index range.
[0042] In some embodiments, the specified group of transform coefficients corresponds to the DC coefficient in the transform block. When the current transform coefficient is the DC coefficient, the context selection for the sign coding can depend on the absolute value of the current transform coefficient level. Otherwise, the context selection for the sign coding is independent of the absolute value of the current transform coefficient level. In some embodiments, the specified number of transform coefficients is only from the luma block. The context selection for the sign coding can depend on the absolute value of the current transform coefficient level in the luma TB and be independent of the absolute value of the current transform coefficient level in the chroma TB. In some embodiments, the specified number of transform coefficients is only associated with some specific transform block dimension, transform type, or CU coding mode.
[0043] According to another aspect of the disclosure, the context modeling for the entropy coding of the sign prediction bins for the current coefficient can be further conditioned on information about the coded sign prediction bins in the current transform block. In some embodiments, the context selection for the entropy coding of the sign prediction bins for certain coefficients can further depend on whether the first coded sign prediction or the DC sign prediction in the current transform block is correct. In some embodiments, the context selection for the entropy coding of the sign for the current coefficient can further depend on the accumulated number of sign prediction bins corresponding to the incorrect sign prediction. In some embodiments, the context selection for the entropy coding of the sign prediction bins for certain coefficients depends on whether the accumulated number of sign prediction bins corresponding to the incorrect sign prediction is greater than one or more specified thresholds. In one embodiment, the context selection for the entropy coding of the sign prediction bins for certain coefficients depends on whether the accumulated number of sign prediction bins corresponding to the incorrect sign prediction is greater than T ic where T icequal to 0, 1, 2, or 3. According to another aspect of the present application, the entropy coding of the remaining sign prediction bits can switch to bypass mode when the accumulated number of coded sign prediction bits corresponding to the erroneous sign prediction is greater than a specified threshold.
[0044] According to another aspect of the present application, the context modeling of the entropy coding of the sign prediction bits for the current coefficient can be further conditioned on the total number of sign prediction bits in the current transform block. In some embodiments, the context selection of the entropy coding of the sign prediction bits for certain coefficients in the current transform block can be further dependent on whether the total number of sign prediction bits in the current transform block is greater than one or more non-zero thresholds. According to yet another aspect of the present application, the process for a video codec can further comprise adaptively setting the values of the one or more thresholds in consideration of the coding context of the current transform block. In some embodiments, the derivation of the one or more thresholds can be further dependent on the transform block dimensions, the transform type, the color component index, the number of predicted signs, the number of non-zero coefficients, or the position of the last significant coefficient associated with the current transform block. The derivation of the one or more thresholds can be further dependent on the prediction mode of the current CU. The derivation of the one or more thresholds can be further dependent on the position or index associated with the current coefficient in the transform block. The derivation of the one or more thresholds can be further dependent on the sum of absolute values of the sign-predicted coefficients in the current transform block.
[0045] According to another aspect of the present application, the context modeling of the entropy coding of the sign prediction bits for the current coefficient can be further conditioned on information about the index or position of the current transform coefficient in the transform block, where the index of the current transform coefficient can correspond to the scan order of the coded predicted sign, or can be derived in the order of a raster scan order, a diagonal scan order (e.g., as shown in Figure 6 In some embodiments, the context selection of the entropy coding of the sign prediction bits for certain coefficients is dependent on whether the index of the current transform coefficient level is greater than or less than one or more non-zero thresholds.
[0046] In some other embodiments, the context selection for entropy encoding / decoding of symbol prediction bits for certain coefficients depends on whether the distance between the top-left block origin at position (0,0) and the current coefficient position (x,y) is greater than or less than one or more non-zero thresholds, where the distance is defined as (x+y). According to another aspect of the invention, the process for a video codec may further include adaptively setting the values of the one or more thresholds or the other one or more non-zero thresholds, taking into account the coding context of the current transform block. In some embodiments, the derivation of the one or more thresholds or the other one or more non-zero thresholds may further depend on the transform block dimension, transform type, color component index, number of predicted symbols, number of non-zero coefficients, or position of the last valid coefficient associated with the current transform block. The derivation of the one or more thresholds or the other one or more non-zero thresholds may also depend on the prediction mode of the current CU. The derivation of the one or more thresholds may further depend on the absolute level associated with the current coefficient or further depend on the sum of the absolute values of the symbol-predicted coefficients in the current transform block.
[0047] According to another aspect of the invention, the context modeling for entropy encoding and decoding of the sign prediction bits of the current coefficients in the current transform block can be further conditional on the width, height, or block size of the current transform block. In some embodiments, the context selection for entropy encoding and decoding of the sign prediction bits of certain coefficients in the current transform block depends on whether the width, height, or block size of the current transform block is greater than or less than one or more thresholds.
[0048] According to another aspect of the invention, context modeling for entropy encoding / decoding of the sign prediction bits of the current coefficients in the current transform block can be further conditional on the transform type associated with the current transform block. In some embodiments, the context selection for entropy encoding / decoding of the sign prediction bits of certain coefficients in the current transform block can further depend on the transform type associated with the current transform block. In some exemplary embodiments, when the current block transform type is a low-frequency inseparable transform (…
[0049] In the case of low-frequency non-separable transform (LFNST) or multiple transform selection (MTS), the video codec can allocate a separate set of contexts for entropy encoding and decoding of symbol prediction bits for certain transform coefficients in the current transform block.
[0050] At the decoder side, the respective cost function is evaluated for the possible hypotheses as disclosed in JVET-D0031. The decoder parses the coefficients, the signs and the sign residuals as part of its parsing process. The signs and the sign residuals are parsed at the end of a TU, at which point the decoder knows the absolute values of all coefficients. It can therefore determine which signs were predicted and, for each predicted sign, it can determine the context used to parse the sign prediction residual from the dequantized coefficient value. The knowledge of a “correct” or “incorrect” prediction is simply stored as part of the CU data of the block being parsed. At this point, the true signs of the coefficients are unknown. During reconstruction, the decoder performs similar operations to the encoder to determine the hypothesis cost. The hypothesis that achieves the minimum cost is determined as the sign prediction.
[0051] In ECM-2.0, when applying sign prediction for a current transform block, the signs of the first N sp non-zero coefficients (when available) are collected and coded according to the forward raster scan order on the top-left 4x4 sub-block in the current transform block. The number of all possible hypotheses subject to discontinuity measure is equal to 1 « N sp and increases significantly with the number of predicted signs, where “<<” is the bit-wise up-shift operator. In one proposed method, the video codec can include a constraint on the minimum absolute value of transform coefficient levels eligible for sign prediction, where the minimum absolute value can be greater than 1 for certain specified transform coefficient. Only when the absolute value of the current transform coefficient level is greater than or equal to the specified minimum value, the sign of the current coefficient under the constraint is allowed to be included in the hypotheses for sign prediction. In this way, coefficients with smaller magnitudes can be excluded from sign prediction, thereby reducing the number of predicted signs in the transform block.
[0052] In some embodiments, only when the absolute value of the current coefficient level is greater than a specified threshold T sp , the sign of the current coefficient is allowed to be included in the hypotheses for sign prediction, where T sp is greater than 0 for specified coefficients from the TB sign prediction region (the top-left 4x4 sub-block in ECM-2.0). In one embodiment, only when the absolute value of the current coefficient level is greater than 1 for certain specified coefficients, the sign of the current coefficient is allowed to be included in the hypotheses for sign prediction. The proposed method can further include adaptively determining the value of T sp . In one example, T sp may be adaptively derived according to the block width, height or size of the current transform block. In another example, T sp may be adaptively derived according to the quantization parameter selected for the current transform block. In yet another example, T may be adaptively derived according to the position or index of the current coefficient in the current block.sp In one example embodiment, T sp :
[0053] T sp equals 0 when the index of the current coefficient is less than C1;
[0054] Otherwise, T sp equals 1 when the index of the current coefficient is less than C2;
[0055] Otherwise, T sp equals 2,
[0056] where 0 < C1 < C2 < N sp .
[0057] In ECM-2.0, when sign prediction is applied to the current transform block, the signs of the first N sp non-zero coefficients (when available) are collected and coded according to the raster scan order on the top-left 4x4 sub-block. The transform coefficient region applicable to sign prediction is fixed to the top-left 4x4 sub-block of each transform block. In the proposed method, the processing of the video codec can further include adaptively setting the transform coefficient region or index range applicable to sign prediction in the current transform block according to the coding context associated with the current transform block. The video codec can include a specified scan order for collecting the signs of the first N sp coefficients in the transform block. The video codec can further include determining a maximum index applicable to sign prediction in the current transform block according to the specified scan order, wherein the signs of any transform coefficients having an index greater than the maximum index according to the specified scan order in the current transform block are not applicable to sign prediction.
[0058] Alternatively, the video codec can include determining whether a transform coefficient or sub-block is applicable to sign prediction in the current transform block according to the position (x, y) of the transform coefficient or sub-block in the current transform block and one or more specific thresholds, wherein the top-left block origin of the current transform block corresponds to the position (0, 0). In some embodiments, in certain transform blocks, the transform coefficient or sub-block at position (x, y) is applicable to sign prediction only when x is less than a first threshold and y is less than a second threshold. In some other embodiments, in certain transform blocks, the transform coefficient or sub-block at position (x, y) is applicable to sign prediction only when the distance between the transform block origin and the position of the current coefficient or sub-block (equal to (x+y)) is less than another specified threshold.
[0059] In some embodiments, the transform coefficient region or index range applicable to the sign prediction in the current transform block can be adaptively set according to the block dimension, the transform type, the color component index, the position of the last significant coefficient, or the number of non-zero coefficients related to the current transform block. In some embodiments, the video codec can reduce the transform coefficient region or index range applicable to the sign prediction in the current transform block when the width or height of the current transform block is less than a specified threshold or the size of the current transform block is less than another specified threshold. In some example embodiments, for a transform block with a block width or height less than a threshold M sp or a block size less than another threshold MN sp , the transform coefficient index range applicable to the sign prediction can be set from 0 to R1 sp , where R1 sp is the maximum index according to the specified scan order. In some embodiments, when M sp is equal to 4, 8, 16, or 32, or when MN sp is equal to 16, 64, 256, or 1028, R1 sp may be equal to 2, 3, 5, or 7. The values of R1 sp , M sp , and MN sp above are merely for illustration, and other suitable values can be selected or determined as needed.
[0060] In some embodiments, the video codec can increase the transform coefficient region or index range applicable to the sign prediction in the current transform block when the width or height of the current transform block is greater than a specified threshold or the size of the current transform block is greater than another specified threshold.
[0061] In some example embodiments, the video codec can include more than one sub-block from the low frequency transform block region when the width or height of the current transform block is greater than a specified threshold or the size of the current transform block is greater than another specified threshold. In some embodiments, the video codec can reduce the transform coefficient region or index range applicable to the sign prediction for one or more specified transform types. In some example embodiments, for one or more specified transform types, the transform coefficient index range applicable to the sign prediction is set from 0 to R2 sp , where R2 sp is the maximum index according to the specified scan order.
[0062] In some other example embodiments, the video codec can include more than one sub-block from the low frequency transform block region only when the distance (x+y) between the transform block origin and the position equal to the current coefficient is less than another specified threshold D1 sptransform coefficients at position (x, y) are applicable for sign prediction. In some embodiments, the one or more specified transform types include certain low frequency non-separable transform (LFNST) types and / or certain transform types associated with multiple transform selection (MTS). R2 sp may be equal to 2, 3, 5, and D1 sp may be equal to 1, 2, 3, 4, 5, 6, or 7. In some embodiments, for certain MTS and / or LFNST types, the transform coefficient region or index range applicable for sign prediction is reduced compared to DCT type II. In some embodiments, the video codec can extend the transform coefficient region or index range applicable for sign prediction for one or more specified transform types. In some example embodiments, the video codec can include more than one subblock from the low frequency transform block region when the transform type associated with the current residual block belongs to one or more specified transform types. In some embodiments, the one or more specified transform types include DCT type II.
[0063] The proposed method can further include signaling information about the transform coefficient region or index range applicable for sign prediction in the bitstream. In some embodiments, the video codec can signal one or more syntax elements for deriving the transform coefficient region or index range applicable for sign prediction in one or more high level parameter sets, such as sequence parameter set (SPS), picture parameter set (PPS), picture header (PH), and / or slice header (SH).
[0064] In some other embodiments, the video codec can signal more than one syntax set to derive information about more than one transform coefficient region or index range for sign prediction in one or more high level syntax sets, where each syntax set corresponds to information about deriving a particular transform coefficient region or index range applicable for sign prediction of certain specified transform blocks, and the selection of the particular transform coefficient region or index range for the current transform block depends on the original context associated with the current transform block. For example, the selection of the particular transform coefficient region or index range for the current transform block can depend on the transform block size, the transform type, the color component index, the number of prediction signs, the position of the last significant coefficient, or the number of non-zero coefficients related to the current transform block. The selection of the particular transform coefficient region or index range for the current transform block can also depend on the prediction mode of the current CU.
[0065] In ECM-2.0, the maximum number N sp of prediction signs in a transform block is signaled in the sequence parameter set (SPS) spis constrained to be less than or equal to 8. In the proposed method, the video codec can adaptively set the maximum allowed number of predictor symbols in the current transform block according to the coding context associated with the current transform block. In some embodiments, the video codec can adaptively derive the maximum allowed number of predictor symbols in the current transform block from one or more reference syntax values based on the translation context associated with the current transform block. In some embodiments, the derivation of the maximum number of allowed predictor symbols in the current transform block depends on the transform block dimension, the transform type, the color component index, the position of the last significant coefficient, or the number of non-zero coefficients associated with the current transform block. For example, when the reference syntax value N sp is set to be greater than 4, the maximum allowed number can be set to 4 for smaller transform block sizes, such as 4*4. In one example, the maximum number of allowed predictor symbols in the current transform block is set to be equal to the reference syntax value N sp multiplied by a scaling factor, where the selection of the scaling factor value depends on the current block dimension.
[0066] In some other embodiments, the video codec can signal more than one syntax set in one or more high-level parameter sets, such as the sequence parameter set (SPS), the picture parameter set (PPS), the picture header (PH), and / or the slice header (SH), where each syntax set corresponds to information about a certain parameter value for deriving the maximum number of allowed predictor symbols in a transform block for certain specified transform blocks. Furthermore, the selection of the certain parameter value for the maximum number of allowed predictor symbols in the current transform block depends on the current coding context. For example, the selection of the certain parameter value for the maximum number of allowed predictor symbols in the current transform block can depend on the transform block dimension, the transform type, the color component index, the position of the last significant coefficient, or the number of non-zero coefficients associated with the current transform block. In one embodiment, the selection of the certain parameter value for the maximum number of allowed predictor symbols in the current transform block can be determined by comparing the current block size with one or more pre-defined thresholds. In another embodiment, the selection of the certain parameter value is determined by whether a secondary transform such as LFNST is applied to the current transform block. The selection of the certain parameter value for the current transform block can also depend on the prediction mode of the current CU. In the proposed method, the processing of the video codec can also include more than one upper bound constraint, each constraining the maximum parameter value for the maximum number of allowed predictor symbols in a transform block for certain specified transform blocks. The determination of the upper bound constraint for the current transform block can depend on the transform block dimension, the transform type, the color component index, the position of the last significant coefficient, or the number of non-zero coefficients associated with the current transform block. In one embodiment, the upper bound constraint is set to be equal to 4 for transform blocks with block size less than or equal to 16.
[0067] In ECM-2.0, the coefficient sign prediction tool is disabled for the current transform block when the transform block width or height of the current transform block is greater than 128 or less than 4. According to one aspect of the present disclosure, a video codec can employ different constraints on block dimensions to enable sign prediction for different transform types. In the proposed method, the video codec includes more than one set of dimension constraints for enabling sign prediction for different transform types. The video codec further includes determining a selected set of dimension constraints for enabling sign prediction in the current transform block according to the current transform type. In some embodiments, for one or more transform types belonging to MTS and / or LFNST types, the video codec can select a set including lower maximum dimension constraint values than DCT type II.
[0068] In ECM-2.0, the signs of the first N sp non-zero coefficients are collected and coded according to the forward raster scan order on the first subblock for sign prediction. According to another aspect of the present disclosure, a video codec can determine the selected first N sp coefficients subject to sign prediction according to an order of decreasing coefficient amplitudes based on the expected distribution of transform coefficients. In one embodiment, the signs of the first N sp coefficients are collected according to a forward subblock-wise diagonal scan order starting from the DC coefficient, where the transform coefficients in each subblock are accessed in a diagonal scan order, and the subblocks in each transform block are also accessed in a diagonal scan order as shown in Figure 6 In another embodiment, the signs of the first N sp coefficients are collected and coded according to a forward subblock-wise diagonal scan order. In another embodiment, the signs of the first N sp coefficients are collected in a forward subblock-wise diagonal scan order and coded in a backward subblock-wise diagonal scan order.
[0069] In ECM-2.0, the best sign prediction hypothesis for the current transform block is determined by minimizing a cost function corresponding to a discontinuity measure through the sum of absolute differences (SAD) across block boundaries. In the proposed method, the best sign prediction hypothesis for the current transform block is determined by minimizing a cost function corresponding to a discontinuity measure through the sum of squared errors across block boundaries as follows:
[0070]
[0071] In ECM-2.0, when sign prediction is enabled for a current transform block, the top-left (lowest frequency) transform sub-block is subject to sign prediction, and the sign entropy coding of the top-left transform sub-block is skipped during the residual coding process used to entropy code the current transform block. After entropy coding all transform blocks (through the per-transform block residual coding process) and the syntax elements used to indicate the LFNST and MTS indices associated with the current CU, a separate sign coding process is used to entropy code the sign information on the transform sub-blocks of each residual block for which sign prediction is enabled in the current CU. As such, each transform block in the current CU for which sign prediction is enabled needs to be revisited after being entropy decoded through the residual decoding process for sign coding.
[0072] In the proposed method, a video coding system includes entropy coding a plurality of residual blocks in a current CU, where each of the plurality of residual blocks is coded by a residual coding process. A sign prediction bit sub of a current transform block is also entropy coded in the residual coding process used to entropy code the current transform block. Each coded sign prediction bit sub of a current coefficient indicates whether the predicted sign of the current coefficient is correct. The separate sign coding process in ECM-2.0 is removed. In some embodiments, each residual block is divided into one or more non-overlapping sub-blocks, and the residual coding process includes entropy coding individual sub-blocks in the current transform block according to a specified sub-block scan order. The residual coding process can also include one or more sub-block coding passes for entropy coding information about absolute values of transform coefficient levels in the current sub-block. The residual coding process can also include determining a set of coefficients in the current transform block for which sign prediction is enabled according to a specified rule. The residual coding process can also include entropy coding the sign prediction bit sub in the current transform block according to a specified scan order.
[0073] In some embodiments, the sign prediction is only applied to transform coefficients from the top-left transform sub-block in each transform block. The residual decoding process can further include a sub-block sign coding pass after the one or more sub-block coding passes. The one or more sub-block coding passes are used to entropy code absolute values of transform coefficient levels in the top-left sub-block, and the sub-block sign coding pass entropy decodes coefficient sign information according to a specified scan order in the top-left sub-block. In the sub-block sign coding pass, a syntax flag can be signaled for a current non-zero coefficient in the top-left sub-block. When the current non-zero coefficient is subject to sign prediction, the syntax flag signals a value of a sign prediction bit to indicate whether the sign prediction for the current coefficient is correct. Otherwise, the syntax flag signals a sign of the current non-zero coefficient. In some embodiments, the sub-block sign coding pass that entropy codes the sign information follows a reverse diagonal scan order in the top-left sub-block. In some other embodiments, the sub-block sign coding pass that entropy codes the sign information follows a forward diagonal scan order in the top-left sub-block. Context modeling of the entropy coded sign prediction bit can be further conditioned on the coded sign prediction bit. In some embodiments, the sign prediction is applied to transform coefficients from more than one sub-block in the top-left transform block region in a transform block. The proposed sub-block sign coding pass can be similarly applied to each sub-block from the top-left transform block region for sign prediction.
[0074] In ECM-2.0, as shown in Figure 5 the best sign prediction hypothesis is determined in the current transform block by minimizing a cost function that corresponds to a discontinuity measure across both the left block boundary and the top block boundary. According to another aspect of the present disclosure, a video codec can completely turn off sign prediction using only one side (left or top) of the block boundary samples to derive the predicted sign or considering the coding context associated with the current transform block. In the proposed method, a video codec includes a sign prediction mode that corresponds to deriving the predicted sign in a transform block based on a discontinuity measure across both the top block boundary and the left block boundary. The video codec further includes two additional sign prediction modes, where the first additional mode corresponds to deriving the predicted sign in a transform block based on only a discontinuity measure across the left block boundary, and the second additional mode corresponds to deriving the predicted sign in a transform block based on only a discontinuity measure across the top block boundary. In some embodiments, the cost functions corresponding to the first and second additional modes are given by equations (6) and (7) as follows, respectively:
[0075]
[0076] The video codec further includes determining a selected sign prediction mode for a current transform block. The video codec can determine the selected sign prediction mode based on the luma component only for all transform blocks in a CU. In some embodiments, the video codec can derive the selected sign prediction mode in the current transform block with consideration of a selected intra prediction direction in the current CU. In one example, the video codec can set the selected sign prediction mode to a first additional prediction mode when the selected intra prediction direction is close to a horizontal prediction direction, and set the selected sign prediction mode to a second additional prediction mode when the selected intra prediction direction is close to a vertical prediction direction.
[0077] In some other embodiments, the video codec can derive the selected sign prediction mode in the current transform block with consideration of block boundary conditions associated with the current block. For example, when a particular current block boundary overlaps with a tile or slice boundary, the video codec can determine not to use the reconstructed samples from the particular boundary to derive the predicted sign. When the reconstructed neighboring samples are not available on both the top and left block boundaries, the video codec can disable the sign prediction for the block. For another example, when it is determined from the reconstructed boundary samples that an image edge can exist on a particular current block boundary, the video codec can determine not to use the reconstructed samples from the particular boundary to derive the predicted sign.
[0078] In some other embodiments, the video codec can signal one or more syntax elements for deriving the selected sign prediction mode in the current transform block or coding unit when a residual signal exists in the current transform block or coding unit. In some embodiments, the video codec can signal the one or more syntax elements to derive the selected sign prediction mode in the current transform block only when a number of the predicted signs in the current transform block is greater than a threshold.
[0079] In some embodiments, the video codec can downsample the left block boundary and the top block boundary vertically in the cost function for discontinuity measurement with reduced computational complexity. In the proposed embodiments, the downsampling rate of the selected block boundary can be reduced by 2 times when only one of the two additional prediction modes is selected for the current transform block.
[0080] In ECM-2.0, when sign prediction is applied to a current transform block, a set of signs for up to N sp coefficients in the current transform block is jointly predicted by minimizing the cost function Eqn. (1). The number of all possible hypotheses for which discontinuity measurement needs to be performed is equal to 1 « N spand increases significantly with the number of symbols to be jointly predicted. According to another aspect of the disclosure, a set of symbols subject to sign prediction in a transform block can be predicted in subsets to reduce the number of all possible hypotheses subject to discontinuity measurement. The set of symbols jointly subject to sign prediction is referred to as a set of jointly predicted coefficient symbols (also referred to as a group of jointly predicted coefficient symbols). In the proposed method, the set of symbols subject to sign prediction in a transform block can be partitioned into one or more subsets of symbols. The prediction and coding of the set of symbols subject to sign prediction includes the prediction and coding of each subset of symbols, where the symbols in each subset are jointly predicted for all possible hypotheses of each subset by minimizing a cost function such as Eqn. (1). The symbols derived from the coded subsets can be used to reconstruct the boundary samples in the current transform block to predict the symbols of the current subset. By partitioning the set of symbols subject to sign prediction into subsets, the maximum number of predicted symbols allowed in a transform block can be further increased to 8.
[0081] In some embodiments, the video codec can employ the same subset size S sp to uniformly partition the set of symbols into one or more subsets, where the subset size is all equal to S sp , except that one subset can have a remaining size after the partitioning of the set by S sp . In some embodiments, the value of S sp may be equal to 1, 2, 3, 4, 5, 6, 7, or 8, or signaled in the bitstream. In some other embodiments, the video codec can adaptively set the subset size S sp . For example, S sp may be adaptively determined according to the transform block dimension, the transform type, the color component index, the number of non-zero coefficients, or the number of predicted symbols associated with the current transform block. S sp may also be adaptively adjusted according to the prediction mode of the current CU. In some embodiments, sign prediction can be skipped for a subset, and all symbols in the subset are coded in a bypass mode. For example, the video codec can skip the sign prediction of a subset in the current transform block, where the subset is not the first subset in the current transform block and the number of predicted symbols or non-zero coefficients in the subset is less than S sp or other specified threshold.
[0082] According to another aspect of the disclosure, the video codec can partition the set of symbols subject to sign prediction in a transform block into subsets of different sizes. For example, S (i) sp may be used to indicate the subset size of subset i. In one example, when up to 8 symbols in a transform block are subject to sign prediction, the subset size S (i) may be equal to 2, 3, 4, 5, 6, 7, or 8 for i equal to 0, 1, and 2, respectively.sp equal to 2, 2, and 4, respectively. In another example, when up to 16 symbols in a transform block are subject to sign prediction, the subset size S (i) sp equal to 2, 2, 4, and 8, respectively. In another example, when up to 16 symbols in a transform block are subject to sign prediction, the subset size S (i) sp equal to 1, 2, 3, 4, and 6, respectively.
[0083] In some embodiments, the video codec can determine the associated subset index for the sign prediction of the current transform coefficient according to the position of the current transform coefficient in the current transform block or the scan index. For example, when the scan index of the current coefficient is greater than a specified threshold T sub , the associated subset index is set equal to 1; otherwise, the associated subset index is set equal to 0. In some embodiments, T sub is equal to 0, 1, 2, 3, 4, or 5.
[0084] In some embodiments, the video codec can adaptively divide the set of symbols subject to sign prediction into subsets according to the transform block dimension, the transform type, the color component index, the number of non-zero coefficients, or the number of predicted signs associated with the current transform block. In some other embodiments, the video codec can adaptively divide the set of symbols subject to sign prediction into subsets according to the prediction mode of the current CU. In some other embodiments, the video codec can adaptively consider the absolute values of the coefficient levels in the current transform block to determine the division of the set of symbols subject to sign prediction into variable size subsets in the current transform block.
[0085] In some embodiments, the information about the division of the set of symbols subject to sign prediction into subsets can be signaled in the bitstream. In some specific embodiments, the information about the division of the set of symbols subject to sign prediction into subsets can be signaled in the sequence parameter set (SPS), the picture parameter set (PPS), the picture header (PH), and / or the slice header (SH). In some embodiments, the context modeling for the entropy coding of the sign of the current coefficient associated with the sign prediction can be further conditioned on the information resulting from the division of the set of symbols subject to sign prediction into subsets. In some embodiments, the context selection for the entropy coding of the sign of the current transform coefficient in a subset can further depend on the subset index associated with the subset.
[0086] In the present disclosure, the entropy coding of the sign of the current coefficient can refer to the entropy coding of the sign prediction bit of the current coefficient in any of the methods proposed above.
[0087] The proposed aspects, methods and related embodiments can be implemented in image and video coding systems, individually and jointly.
[0088] Any of the proposed methods can be implemented in an encoder and / or a decoder. For example, any of the proposed methods can be implemented in a coefficient coding module of an entropy encoder 122 in an encoder (e.g., the encoder 120 in FIG. 1) and / or a coefficient coding module of an entropy decoder 140 in a decoder (e.g., the decoder 130 in FIG. 1). Alternatively, any of the proposed methods can be implemented as a circuit integrated to the coefficient coding module of an encoder and / or the coefficient coding module of a decoder. Figure 1A Figure 1B Alternatively, any of the proposed methods can be implemented as a circuit integrated to the coefficient coding module of an encoder and / or the coefficient coding module of a decoder.
[0089] Figure 7 A flowchart of an exemplary video decoding system utilizing joint sign prediction according to an embodiment of the present application is shown. The steps shown in the flowchart can be implemented as program code executable on one or more processors (e.g., one or more CPUs) on the encoder side. The steps shown in the flowchart can also be implemented based on hardware, e.g., one or more electronic devices or processors arranged to perform the steps in the flowchart. At step 710, a video decoder according to the method will receive encoded transform coefficients corresponding to a residual block of a current block, wherein the encoded transform coefficients include encoded sign residues. At step 720, a maximum number allowed is determined according to a coding context related to the current block. At step 730, the encoded sign residues are decoded into decoded sign residues, wherein a total number of the set of jointly predicted coefficient signs is equal to or less than the allowed number. At step 740, a joint sign prediction of the set of jointly predicted coefficient signs is determined by selecting one of a set of hypotheses of the set of jointly predicted coefficient signs corresponding to a hypothesis that achieves a minimum cost, wherein the cost of each hypothesis in the set of hypotheses is calculated separately based on a boundary pixel of the current block and a corresponding neighboring pixel of the current block of the hypothesis, and the boundary pixel of the current block associated with each hypothesis is reconstructed using information including said each hypothesis. At step 750, the set of jointly predicted coefficient signs is reconstructed based on the decoded sign residues and the joint sign prediction.
[0090] Figure 8 A flowchart of an exemplary video coding system utilizing joint sign prediction according to embodiments of the present application is shown. In step 810, a video encoder according to the method receives transform coefficients corresponding to a residual block of a current block. In step 820, a maximum allowed number is determined according to a coding context associated with the current block. In step 830, a set of jointly predicted coefficient signs associated with a set of selected transform coefficients is determined, wherein a total number of the set of jointly predicted coefficient signs is equal to or less than the allowed maximum number. In step 840, a joint sign prediction of the set of jointly predicted coefficient signs is determined by selecting one of a set of combined hypotheses of the set of jointly predicted coefficient signs corresponding to a hypothesis that achieves a minimum cost, wherein the cost for each hypothesis in the set of hypotheses is calculated separately for the hypothesis using a boundary pixel of the current block and a corresponding neighboring pixel of the current block, and wherein the boundary pixel of the current block associated with each hypothesis is reconstructed using information comprising said each hypothesis. In step 850, a sign residual between the set of jointly predicted coefficient signs and the joint sign prediction is determined. In step 860, a context coding is applied to the sign residual to generate a coded sign residual.
[0091] The flowchart shown is intended to illustrate an example of video coding according to the present application. Those skilled in the art can modify each step, rearrange steps, split steps, or combine steps to implement the present application without departing from the spirit of the present application. In the present disclosure, specific syntax and semantics have been used to illustrate examples to implement embodiments of the present application. Those skilled in the art can practice the present application by replacing the syntax and semantics with equivalent syntax and semantics without departing from the spirit of the present application.
[0092] The above description is presented to enable any person skilled in the art to practice the present application as provided in the context of a particular application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, the present application is not intended to be limited to the particular embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features herein disclosed. In the above detailed description, various specific details are illustrated in order to provide a thorough understanding of the present application. However, those skilled in the art will understand that the present application can be practiced without these specific details.
[0093] Embodiments of the present application as described above can be implemented in various hardware, software code, or a combination of both. For example, an embodiment of the present application can be one or more circuitry integrated into a video compression chip or program code integrated into video compression software to perform the processing described herein. Embodiments of the present application can also be program code to be executed on a digital signal processor (DSP), a microprocessor, or a field programmable gate array (FPGA) to perform the processing described herein. The present application can also relate to a processing circuitry that executes specific methods carried out by a computer processor, a digital signal processor, a microprocessor or a field programmable gate array (FPGA). The processor can be configured to perform specific tasks according to the present application by executing machine-readable software code or firmware code that defines the specific methods embodied by the present application. The software code or firmware code can be developed in different programming languages and different formats or styles. The software code can also be compiled into different object code formats or styles. However, different code formats, styles, and languages of software code and other means of configuring code to perform tasks in accordance with the present application will not depart from the spirit and scope of the application.
[0094] The present application can be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The described examples are to be considered in all respects only as illustrative and not restrictive. The scope of the application is, therefore, indicated by the appended claims, rather than by the foregoing description. All changes that come within the meaning of equivalency of the claims are intended to be embraced within the scope of the claims.
Claims
1. A video decoding method, the method comprising: Receive transform coefficients of the codec for the residual block corresponding to the current block, wherein the transform coefficients of the codec include the symbol residuals of the codec associated with a set of predicted coefficient symbols; The predicted coefficient symbols are divided into one or more subsets, where each subset includes one or more jointly predicted coefficient symbols; The symbol residual of the encoding and decoding is decoded into the symbol residual of the decoding and used for one or more subsets; The joint symbol prediction of the coefficient symbols for each subset of the joint prediction is determined by selecting a hypothesis from a set of hypotheses corresponding to a combination of coefficient symbols that achieve the minimum cost. For each hypothesis, the cost between the boundary pixel of the current block and its corresponding neighboring pixel is calculated. Information including each hypothesis is used to reconstruct the boundary pixel of the current block associated with each hypothesis. The symbolic coefficients of the set of predictions are reconstructed based on the symbolic residuals from the decoding and the joint symbolic predictions.
2. The video decoding method according to claim 1, wherein, Parse one or more syntax elements from one or more high-level parameter sets to derive the coefficient symbols for that set of predictions.
3. The video decoding method according to claim 2, wherein, in, The one or more high-level parameter sets include sequence parameter sets, image parameter sets, image headers, slice headers, or combinations thereof.
4. The video decoding method according to claim 2, wherein, The maximum allowed number is determined based on the block size of the current block.
5. The video decoding method according to claim 4, wherein, When the block size is equal to 4×4, the maximum allowed number is less than or equal to 4.
6. The video decoding method according to claim 2, wherein, The maximum allowed number is determined based on the transformation type of the current block.
7. The video decoding method according to claim 1, wherein, The bits of the symbol residual associated with the predicted coefficient symbols of this encoding and decoding are entropy encoded during the residual encoding and decoding process used for the current block.
8. The video decoding method according to claim 7, wherein, The residual encoding / decoding process also includes a sub-block symbol encoding / decoding channel for entropy encoding / decoding the set of predicted coefficient symbols from one or more sub-blocks in the upper left region of the current block, after entropy encoding / decoding one or more sub-blocks of the absolute value of the transform coefficient level of the one or more sub-blocks.
9. The video decoding method according to claim 1, wherein, The cost is calculated between the boundary pixels of the current block and the corresponding adjacent pixels of the current block that are only on the top boundary of the current block or only on the left boundary of the current block.
10. The video decoding method of claim 9, wherein the selection between the top boundary of the current block and the left boundary of the current block depends on the selected intra-frame prediction direction of the current block.
11. A video decoding apparatus, the apparatus comprising one or more electronic circuits or processors, for: Receive transform coefficients of the codec for the residual block corresponding to the current block, wherein the transform coefficients of the codec include the symbol residuals of the codec associated with a set of predicted coefficient symbols; The predicted coefficient symbols are divided into one or more subsets, where each subset includes one or more jointly predicted coefficient symbols; The symbol residual of the encoding and decoding is decoded into the symbol residual of the decoding and used for one or more subsets; The joint symbol prediction of the coefficient symbols for each subset of the joint prediction is determined by selecting a hypothesis from a set of hypotheses corresponding to a combination of coefficient symbols that achieve the minimum cost. For each hypothesis, the cost between the boundary pixel of the current block and its corresponding neighboring pixel is calculated. Information including each hypothesis is used to reconstruct the boundary pixel of the current block associated with each hypothesis. The symbolic coefficients of the set of predictions are reconstructed based on the symbolic residuals from the decoding and the joint symbolic predictions.
12. A video coding method, the method comprising: Receive the transformation coefficients corresponding to the residual block of the current block; Determine the sign of a set of predicted coefficients associated with a selected set of transformation coefficients; The predicted coefficient symbols are divided into one or more subsets, where each subset includes one or more jointly predicted coefficient symbols; A joint symbol prediction for the coefficient symbols of each subset is determined by selecting a hypothesis from a set of hypotheses corresponding to a combination of coefficient symbols that achieve the minimum cost joint prediction. For each hypothesis, the cost between the boundary pixel of the current block and its corresponding neighboring pixel is calculated, and the boundary pixel of the current block associated with each hypothesis is reconstructed using information including each hypothesis. For each subset, the symbol residual between the coefficient symbols of the set of predictions and the joint symbol prediction is determined. Context encoding / decoding is applied to the symbolic residual to generate encoded / decoded symbolic residuals for each subset.
13. The video coding method of claim 12, further comprising sending one or more syntactic elements to derive a maximum allowed number of coefficient symbols for the set of predictions.
14. The video encoding method according to claim 12, wherein, The one or more syntactic elements are sent from a sequence parameter set, a picture parameter set, a picture header, a slice header, or a combination thereof.
15. The video encoding method according to claim 12, wherein, The codec context corresponds to the block size of the current block.
16. The video encoding method according to claim 12, wherein, The bits of the symbol residual associated with the predicted coefficient symbols of this encoding and decoding are entropy encoded during the residual encoding and decoding process used for the current block.
17. The video encoding method according to claim 16, wherein, The residual encoding / decoding process also includes a sub-block symbol encoding / decoding channel for entropy encoding / decoding the set of predicted coefficient symbols from one or more sub-blocks in the upper left region of the current block, after entropy encoding / decoding one or more sub-blocks of the absolute value of the transform coefficient level of the one or more sub-blocks.
18. The video encoding method according to claim 12, wherein, The cost is calculated between the boundary pixels of the current block and the corresponding adjacent pixels of the current block that are only on the top boundary of the current block or only on the left boundary of the current block.
19. The video coding method of claim 18, wherein the selection between the top boundary of the current block and the left boundary of the current block depends on the selected intra-prediction direction of the current block.