Improved residual code prediction of transform coefficients in video coding.
Patent Information
- Application Number
- JP2024514612
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-04
- Filing Date
- 2022-09-29
- Publication Date
- 2025-09-30
AI Technical Summary
Existing video coding standards, such as VVC, do not effectively predict and encode the sign of residual coefficients, leading to inefficiencies in bitrate and compression efficiency, particularly in video encoding processes.
Implement a residual code prediction method that utilizes gradient-based boundary residual prediction, multimodal residual boundary prediction, and extended residual code prediction areas, along with improved signaling techniques to enhance the prediction of residual signs in video coding.
Improves compression efficiency and reduces bitrate by accurately predicting residual signs across larger areas of transform blocks, thereby optimizing video encoding and decoding processes.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications
[0001] This PCT application claims priority to U.S. Provisional Patent Application No. 63 / 250,202, filed September 29, 2021, and to U.S. Provisional Patent Application No. 63 / 296,370, filed January 4, 2022, each of which is incorporated by reference herein. [Background technology]
[0002]
[0002] In 2020, the ITU-T Video Coding Expert Group ("ITU-T VCEG") and the ISO / IEC Moving Picture Expert Group ("ISO / IEC MPEG") Joint Video Experts Team ("JVET") published the final draft of a next-generation video codec specification, Versatile Video Coding ("VVC"). The specification further improves video coding performance over previous standards such as H.264 / AVC (Advanced Video Coding) and H.265 / HEVC (High Efficiency Video Coding). JVET continues to propose additional technologies beyond the scope of the VVC standard itself, collected under the name Enhanced Compression Model ("ECM").
[0003]
[0003] In each of the AVC, HEVC, and VVC standards, the discrete cosine transform (DCT) technique is a basic image compression technique for efficient video coding. An image is divided into coding blocks, a prediction calculation is performed on each coding block to derive a prediction unit and a residual, and then a transform calculation is performed on the residual to derive an array of coefficients. Based on the residual coefficients, a decoder may reconstruct an image based on the prediction unit and the transform coefficients, and finally output the reconstructed picture to a bitstream for transmission. By transmitting the reconstructed picture, the bitrate required for transmitting the video stream is reduced, thereby achieving rate improvement.
[0004]
[0004] Further improvements in compression efficiency have been achieved by further predicting, encoding, and decoding the residual coefficients. This is implemented in the form of a Context-Sensitive Binary Arithmetic Codec ("CABAC") in an entropy coder. However, due to challenges understood by those skilled in the art, the sign of each residual coefficient has traditionally not been predicted or encoded, but rather transmitted as "sign bits", which continue to occupy a significant portion of the transmitted bitstream and remain the target of ongoing efforts to achieve further bitrate improvements.
[0005]
[0005] At the time of writing, the latest draft of the ECM (presented as "Algorithm Description of Enhanced Compression Model 3 (ECM3)" at the 136th Meeting of the Motion Picture Experts Group ("MPEG") in October 2021) includes a proposal to predict, encode, and decode the residual coefficient code. However, it is desirable to continue to refine the code prediction proposal in order to achieve even greater bitrate improvements and compression efficiency in video coding, and so that the resulting bitrate improvements are not significantly offset by loss of prediction accuracy. Summary of the Invention
[0006]
[0006] Now referring to the accompanying drawings, in which the leftmost digit(s) of a reference number identifies the figure in which the reference number first appears, the use of the same reference number in different figures indicates similar or identical items or features. [Brief description of the drawings]
[0007] [Figure 1A] FIG. 7 is an exemplary block diagram of a video encoding process according to an exemplary embodiment of the present disclosure. [Figure 1B] FIG. 7 is an exemplary block diagram of a video decoding process according to an exemplary embodiment of the present disclosure. [Diagram 2]
[0008] FIG. 13 shows a 4×4 region of a TU of size 16×8 pixels according to the ECM3 proposal. [Diagram 3]
[0009] FIG. 13 illustrates the scan order of a 4×4 region of a TU to determine the predicted code according to the ECM3 proposal. [Figure 4A]
[0010] FIG. 13 is a diagram showing an example of residual boundary prediction of an 8×4 TU according to the ECM3 proposal. [Figure 4B]
[0010] FIG. 2 is a diagram showing an example of residual boundary prediction of an 8×4 TU according to the ECM3 proposal. [Diagram 5]
[0011] FIG. 13 shows the derived predicted boundary residuals according to the ECM3 proposal. [Figure 6A]
[0012] FIG. 1 illustrates the use of a gradient-based boundary residual code prediction process in a residual code prediction method according to an exemplary embodiment of the present disclosure. [Figure 6B]
[0012] FIG. 2 illustrates the use of a gradient-based boundary residual code prediction process in a residual code prediction method according to an exemplary embodiment of the present disclosure. [Figure 7]
[0013] FIG. 13 shows some correlation patterns between the eight nearest neighbors to the predicted pixel in the top row. [Figure 8A-C]
[0014] FIG. 13 illustrates an exemplary residual boundary prediction for TB according to the diagonal right mode of the present disclosure. [Fig. 8D-F] FIG. 2 illustrates an exemplary residual boundary prediction for TB according to the diagonal right mode of the present disclosure. [Figure 9A-C]
[0015] 13A-13C are diagrams illustrating exemplary residual boundary prediction for TB according to the diagonal left mode of the present disclosure. [Fig. 9D-F] FIG. 2 illustrates an exemplary residual boundary prediction for TB according to the diagonal left mode of the present disclosure. [Figure 10A-C]
[0016] FIG. 13 illustrates an exemplary residual boundary prediction for TB according to the first mixed mode of the present disclosure. [Fig. 10D-F] FIG. 2 illustrates an exemplary residual boundary prediction for TB according to the first mixed mode of the present disclosure. [Figure 11A-C]
[0017] FIG. 13 illustrates an exemplary residual boundary prediction for TB according to the second mixed mode of the present disclosure. [Fig. 11D-F]
[0017] FIG. 4 illustrates an exemplary residual boundary prediction for TB according to the second mixed mode of the present disclosure. [Figure 12A-C]
[0018] FIG. 13 is a diagram illustrating an example of TB residual boundary prediction according to the third mixed mode of the present disclosure. [Fig. 12D-F] FIG. 13 is a diagram illustrating an example of TB residual boundary prediction according to the third mixed mode of the present disclosure. [Figure 13A-C]
[0019] FIG. 13 is a diagram showing an example of TB residual boundary prediction according to the fourth mixed mode of the present disclosure. [Fig. 13D-F] FIG. 13 is a diagram illustrating an example of TB residual boundary prediction according to the fourth mixed mode of the present disclosure. [Figure 14A-C]
[0020] FIG. 13 is a diagram showing an example of TB residual boundary prediction according to the fifth mixed mode of the present disclosure. [Fig. 14D-F] FIG. 13 is a diagram illustrating an example of TB residual boundary prediction according to the fifth mixed mode of the present disclosure. [Figure 15A-C]
[0021] FIG. 13 is a diagram showing an example of TB residual boundary prediction according to the sixth mixed mode of the present disclosure. [Fig. 15D-F] FIG. 13 is a diagram illustrating an example of TB residual boundary prediction according to the sixth mixed mode of the present disclosure. [Figure 16]
[0022] FIG. 13 illustrates an example of a sample of reconstructed neighboring blocks compared to a template, according to an exemplary embodiment of the present disclosure. [Figure 17]
[0023] FIG. 2 illustrates an example of classifying TB codes into a first group and a second group according to an exemplary embodiment of the present disclosure. [Figure 18]
[0024] FIG. 2 illustrates an example of a TB-dependent M×N code prediction region according to an exemplary embodiment of the present disclosure. [Figure 19]
[0025] FIG. 2 illustrates an example of a residual code prediction method utilizing a sorting order according to an exemplary embodiment of the present disclosure. [Figure 20]
[0026] FIG. 1 illustrates the mismatch of quantization indexes for the same residual coefficient level between two quantizers. [Figure 21]
[0027] A diagram showing the bitstream signaling order according to the ECM3 proposal. [Figure 22]
[0028] A diagram showing the consequences of bitstream signaling ordering according to the ECM3 proposal. [Figure 23]
[0029] FIG. 1 illustrates a two-pass bitstream signaling method for a residual code prediction method utilizing a sorting order according to an exemplary embodiment of the present disclosure. [Figure 24]
[0030] FIG. 13 illustrates a zigzag scan order applied to the scan order of a 4×4 region of a TB to determine a predicted code. [Diagram 25]
[0031] FIG. 1 illustrates an example of an extended boundary for M=N=2 according to an exemplary embodiment of the present disclosure. [Figure 26]
[0032] FIG. 1 illustrates an example of discontinuity measurement with boundary expansion. [Figure 27]
[0033] FIG. 1 illustrates an example system for implementing the processes and methods described herein for implementing residual code prediction. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008]
[0034] In accordance with the VVC video coding standard ("VVC standard") and the motion prediction described therein, computer-readable instructions stored on a computer-readable storage medium are executable by one or more processors of a computing system to configure the one or more processors to perform encoder operations described by the VVC standard and decoder operations described by the VVC standard. Some of these encoder and decoder operations according to the VVC standard are described in more detail below, but these following descriptions should not be understood as being exhaustive of the encoder and decoder operations according to the VVC standard. Subsequently, a "VVC standard encoder" and a "VVC standard decoder" shall describe respective computer-readable instructions stored on a computer-readable storage medium that configure one or more processors to perform these respective operations (which may be referred to as a "reference implementation" of the encoder or decoder, by way of example).
[0009]
[0035] Furthermore, according to an exemplary embodiment of the present disclosure, the VVC standard encoder and the VVC standard decoder further include computer readable instructions stored in a computer readable storage medium executable by one or more processors of a computing system to configure the one or more processors to perform operations not specified in the VVC standard. The VVC standard encoder should not be understood as being limited to the operation of the reference implementation of the encoder, but should be understood as including further computer readable instructions to configure one or more processors of a computing system to perform further operations described herein. The VVC standard decoder should not be understood as being limited to the operation of the reference implementation of the decoder, but should be understood as including further computer readable instructions to configure one or more processors of a computing system to perform further operations described herein.
[0010]
[0036] 1A and 1B show example block diagrams of an encoding process 100 and a decoding process 150, respectively, according to an example embodiment of the present disclosure.
[0011]
[0037] In the encoding process 100, the VVC standard encoder configures one or more processors of the computing system to receive as input one or more input pictures from an image source 102. The input picture includes a number of pixels sampled by an image capture device, such as a photosensor array, and includes an uncompressed stream of multiple color channels (e.g., RGB color channels) that store color data at the original resolution of the picture, with each channel using a number of bits to store the color data for each pixel of the picture. The VVC standard encoder configures one or more processors of the computing system to store this uncompressed color data in a compressed format, with the color data stored at a resolution lower than the original resolution of the picture and encoded as a luma ("Y") channel and two chroma ("U" and "V" channels) that are lower resolution than the luma channel.
[0012]
[0038] The VVC standard encoder encodes a picture (the picture being encoded, which is called the “current picture” to distinguish it from other pictures received from the image source 102) by configuring one or more processors of the computing system to divide the original picture into units and subunits according to a partitioning structure. The VVC standard encoder configures one or more processors of the computing system to subdivide the picture into macroblocks (“MBs”), each MB having a dimension of 16×16 pixels, which may be further subdivided into partitions. The VVC standard encoder configures one or more processors of the computing system to subdivide the picture into coding tree units (“CTUs”), whose luma and chroma components may be further subdivided into coding tree blocks (“CTBs”), which may be further subdivided into coding units (“CUs”). Alternatively, the VVC standard encoder configures one or more processors of the computing system to subdivide the picture into units of N×N pixels, which may then be further subdivided into subunits. Each of these granular units of a picture may be referred to generally as a "block" for purposes of this disclosure.
[0013]
[0039] A CU is coded using one block of luma samples and two corresponding blocks of chroma samples, and the picture is not monochrome and is coded using one coding tree.
[0014]
[0040] A VVC standard encoder configures one or more processors of a computing system to subdivide a block into partitions having dimensions that are multiples of 4 x 4 pixels. For example, a partition of a block may have dimensions of 8 x 4 pixels, 4 x 8 pixels, 8 x 8 pixels, 16 x 8 pixels, or 8 x 16 pixels.
[0015]
[0041] A VVC standard encoder configures one or more processors of a computing system to encode color information of a picture at a lower resolution than an input picture by encoding color information of blocks of the picture and its subdivisions rather than color information of the pixels of the original picture at full resolution, thereby storing the color information in fewer bits than the input picture.
[0016]
[0042] Furthermore, the VVC standard encoder encodes a picture by configuring one or more processors of the computing system to perform motion prediction on blocks of the current picture. Motion prediction encoding refers to storing image data of blocks of the current picture (blocks of the original picture before encoding are called "input blocks") using motion information and a prediction unit ("PU") rather than pixel data according to intra prediction 104 or inter prediction 106.
[0017]
[0043] Motion information refers to data describing the motion of a block structure of a picture or a unit or subunit thereof, such as motion vectors and references to blocks of a current picture or a reference picture. A PU may refer to a unit or subunits corresponding to one of a plurality of block structures of a picture, such as an MB or a CTU, and the blocks are divided based on the picture data and coded according to the VVC standard. The motion information corresponding to a PU may describe a motion prediction coded by a VVC standard encoder as described herein.
[0018]
[0044] A VVC standard encoder configures one or more processors of a computing system to encode motion prediction information for each block of a picture in an inter-block encoding order, such as a raster scan order in which the first block to be decoded is the topmost and leftmost block of the picture. The block being encoded is called the "current block" to distinguish it from other blocks in the same picture.
[0019]
[0045] According to intra prediction 104, one or more processors of the computing system are configured to encode a block by referring to the motion information and PU of one or more other blocks of the same picture. According to intra predictive encoding, one or more processors of the computing system perform intra prediction 104 (also called spatial prediction) calculations by encoding the motion information of a current block based on spatially neighboring samples from spatially neighboring blocks of the current block.
[0020]
[0046] According to inter prediction 106, one or more processors of the computing system are configured to encode a block by referring to motion information and PUs of one or more other pictures. For the purpose of inter prediction encoding, one or more processors of the computing system are configured to store one or more previously encoded and decoded pictures in a reference picture buffer, where these stored pictures are called reference pictures.
[0021]
[0047] The one or more processors are configured to perform inter prediction 106 (also called temporal prediction or motion compensated prediction) calculation by encoding motion information of the current block based on samples from one or more reference pictures. The inter prediction may further be calculated according to unidirectional prediction or bidirectional prediction, where only one motion vector pointing to one reference picture is used to generate a prediction signal of the current block. In bidirectional prediction, two motion vectors, each pointing to a respective reference picture, are used to generate a prediction signal of the current block.
[0022]
[0048] The VVC standard encoder configures one or more processors of the computing system to encode the CU to include a reference index for identifying a prediction signal of a current block for reference by a VVC standard decoder. The one or more processors of the computing system can encode the CU to include an inter-prediction indicator. The inter-prediction indicator indicates List 0 prediction with reference to a first reference picture list called List 0, List 1 prediction with reference to a second reference picture list called List 1, or bidirectional prediction with reference to both reference picture lists called List 0 and List 1, respectively.
[0023]
[0049] If the inter prediction indicator indicates list 0 prediction or list 1 prediction, the one or more processors of the computing system are configured to encode the CU including a reference index that references a reference picture of a reference picture buffer referenced by list 0 or list 1, respectively. If the inter prediction indicator indicates bidirectional prediction, the one or more processors of the computing system are configured to encode the CU including a first reference index that references a first reference picture of a reference picture buffer referenced by list 0 and a second reference index that references a second reference picture of a reference picture referenced by list 1.
[0024]
[0050] The VVC standard encoder configures one or more processors of the computing system to individually encode the current block of the picture to output the respective predicted block. According to the VVC standard, the CTU can be sized to 128x128 luma samples (plus corresponding chroma samples depending on the chroma format). The CTU can be further divided into CUs according to a quadtree, a binary tree, or a ternary tree. The one or more processors of the computing system are configured to finally record the coding parameter set, such as the coding mode (intra mode or inter mode), the motion information (reference index, motion vector, etc.) of the inter-coded block, and the quantized residual coefficient, in the syntax structure of the leaf node of the partition structure.
[0025]
[0051] After the prediction block is output, the VVC standard encoder configures one or more processors of the computing system to send a set of coding parameters, such as the coding mode (i.e., intra or inter prediction), the intra or inter prediction mode, and motion information, to the entropy encoder 124 (described below).
[0026]
[0052] The VVC standard provides semantics for recording the coding parameter set of a CU. For example, with respect to the above coding parameter set, the pred_mode_flag of the CU is set to 0 for inter-coded blocks and set to 1 for intra-coded blocks, the general_merge_flag of the CU is set to indicate whether the inter prediction of the CU uses merge mode or not, the inter_affine_flag and cu_affine_type_flag of the CU are set to indicate whether the inter prediction of the CU uses affine motion compensation or not, the mvp_l0_flag and mvp_l1_flag are set to indicate the motion vector index in list 0 or list 1, respectively, and the ref_idx_l0 and ref_idx_l1 are set to indicate the reference picture index in list 0 or list 1, respectively. It should be understood that the VVC standard includes semantics for recording various other information, flags, and options that are beyond the scope of this disclosure.
[0027]
[0053] The VVC standard encoder further implements one or more mode decision and encoder control settings 108, e.g., rate control settings. The one or more processors of the computing system are configured to perform a mode decision by selecting an optimized prediction mode for a current block after intra prediction or inter prediction based on a rate-distortion optimization method.
[0028]
[0054] The rate control settings configure one or more processors of a computing system to assign different quantization parameters ("QP") to different pictures. The magnitude of the QP determines the scale at which picture information is quantized during encoding by one or more processors (as described below), and thus the extent to which encoding process 100 discards picture information from MBs of a sequence during encoding (because that information falls between steps of the scale).
[0029]
[0055] The VVC standard encoder further implements a subtractor 110. One or more processors of the computing system are configured to perform the subtraction operation by calculating the difference between the input block and the predicted block. Based on the optimized prediction mode, the predicted block is subtracted from the input block. The difference between the input block and the predicted block is called a prediction residual, or "residual" for brevity.
[0030]
[0056] Based on the prediction residual, the VVC standard encoder further implements a transform 112. One or more processors of the computing system are configured to perform a transform operation on the residual by matrix arithmetic operations to derive an array of coefficients (which may be called "residual coefficients", "transform coefficients", etc.), thereby encoding the current block as a transform block ("TB"). The transform coefficients may refer to coefficients that represent one of several spatial transformations, such as a diagonal flip, a vertical flip, or a rotation, that may be applied to a sub-block.
[0031]
[0057] It should be appreciated that the coefficients may be stored as two components, a magnitude and a sign, as will be explained in more detail below.
[0032]
[0058] The sub-blocks of a CU, such as PU and TB, can be arranged in any combination of sub-block dimensions as described above. A VVC standard encoder configures one or more processors of a computing system to subdivide a CU into a hierarchical structure of TBs, a residual quadtree ("RQT"). The RQT provides an order for motion prediction and residual coding across the sub-blocks at each level, recursively descending each level of the RQT.
[0033]
[0059] The VVC standard encoder further implements quantization 114. One or more processors of the computing system are configured to perform a quantization operation on the residual coefficients by a matrix arithmetic operation based on the quantization matrix and the above assigned QP. The residual coefficients that fall within an interval are kept, and the residual coefficients that fall outside the interval step are discarded.
[0034]
[0060] The VVC standard encoder further implements inverse quantization 116 and inverse transform 118. One or more processors of the computing system are configured to perform inverse quantization and inverse transform operations on the quantized residual coefficients by inverse matrix arithmetic operations of the quantization and transform operations described above. The inverse quantization and inverse transform operations result in a reconstructed residual.
[0035]
[0061] The VVC standard encoder further implements an adder 120. The one or more processors of the computing system are configured to perform an addition operation by adding the prediction block and the reconstructed residual to output a reconstructed block.
[0036]
[0062] The VVC standard encoder further implements a loop filter 122. The one or more processors of the computing system are configured to apply loop filters, such as a deblocking filter, a sample adaptive offset ("SAO") filter, and an adaptive loop filter ("ALF"), to the reconstructed block to output a filtered reconstructed block.
[0037]
[0063] The VVC standard encoder further configures one or more processors of the computing system to output the filtered reconstructed blocks to a decoded picture buffer (“DPB”) 200. The DPB 200 stores reconstructed pictures that are used by one or more processors of the computing system as reference pictures in encoding pictures other than the current picture, as described above with respect to inter prediction.
[0038]
[0064] The VVC standard encoder further implements an entropy encoder 124. The one or more processors of the computing system are configured to perform entropy encoding, in which symbols constituting the quantized residual coefficients are encoded by mapping into binary strings (hereinafter "bins"), which can be transmitted in an output bitstream at a compressed bit rate, according to a context-dependent binary arithmetic codec ("CABAC"). The encoded quantized residual coefficient symbols include absolute values of the residual coefficients (these absolute values are hereinafter referred to as "residual coefficient levels").
[0039]
[0065] However, while the residual coefficient levels are predicted and coded, the residual coefficient signs are signaled using bins indicating equiprobable (hereinafter "EP") states (it should be understood that coefficients with value 0 do not have signs and therefore do not need to be signaled). VVC standard encoders do not configure one or more processors of a computing system to predict the residual coefficient signs due to computational challenges (which will be understood by those skilled in the art, but need not be repeated here to understand the exemplary embodiments of the present disclosure). For these reasons, CABAC configures one or more processors of a computing system to bypass coding of the residual coefficient signs and to additionally transmit one bit per sign in the output bitstream.
[0040]
[0066] Thus, the entropy encoder configures one or more processors of the computing system to encode residual coefficient levels of the block, bypass encoding of the residual coefficient code, record the residual coefficient code together with the encoded block, record coding parameter sets, such as, for example, the coding mode, intra- or inter-prediction mode, and motion information, that are encoded within the syntax structure of the encoded block (e.g., a picture parameter set ("PPS") included in the picture header and a sequence parameter set ("SPS") included in a sequence of multiple pictures), and output the encoded block.
[0041]
[0067] The VVC standard encoder configures one or more processors of the computing system to output an encoded picture composed of encoded blocks from the entropy encoder 124. The encoded picture is output to a transmission buffer and ultimately packed into a bitstream for output from the VVC standard encoder.
[0042]
[0068] In a decoding process 150, a VVC standard decoder configures one or more processors of a computing system to receive as input one or more encoded pictures from a bitstream.
[0043]
[0069] The VVC standard decoder implements an entropy decoder 152. One or more processors of the computing system are configured to perform entropy decoding, and the bins are decoded by reversing the symbol-to-bin mapping according to CABAC, thereby recovering the entropy-coded quantized residual coefficients. The entropy decoder 152 outputs the quantized residual coefficients, outputs residual coefficient codes with bypassed coding, and also outputs syntax structures such as PPS and SPS.
[0044]
[0070] The VVC standard decoder further implements inverse quantization 154 and inverse transform 156. One or more processors of the computing system are configured to perform inverse quantization and inverse transform operations on the decoded quantized residual coefficients by inverse matrix arithmetic operations of the quantization and transform operations described above. The inverse quantization and inverse transform operations result in a reconstructed residual.
[0045]
[0071] Furthermore, based on the coding parameter set recorded by the entropy coder 124 in syntax structures such as PPS and SPS (or received by out-of-band transmission or coded within the decoder) and the coding mode included in the coding parameter set, the VVC standard decoder determines whether to apply intra prediction 156 (i.e., spatial prediction) or motion compensated prediction 158 (i.e., temporal prediction) to the reconstructed residual.
[0046]
[0072] If the coding parameter set specifies intra prediction, the VVC standard decoder configures one or more processors of the computing system to perform intra prediction 158 using the prediction information specified in the coding parameter set, which then generates a prediction signal.
[0047]
[0073] If the coding parameter set specifies inter prediction, the VVC standard decoder configures one or more processors of the computing system to perform motion compensated prediction 160 using reference pictures from the DPB 200. The motion compensated prediction 160 thereby generates a prediction signal.
[0048]
[0074] The VVC standard decoder further implements an adder 162. The adder 162 configures one or more processors of the computing system to perform an addition operation on the reconstructed residual and the prediction signal to output a reconstructed block.
[0049]
[0075] The VVC standard decoder further implements a loop filter 164. The one or more processors of the computing system are configured to apply a loop filter, such as, for example, a deblocking filter, a SAO filter, an ALF, etc., to the reconstructed block to output a filtered reconstructed block.
[0050]
[0076] The VVC standard decoder further configures one or more processors of the computing system to output the filtered reconstructed blocks to the DBP 200. As described above, the DPB 200 stores reconstructed pictures that are used by one or more processors of the computing system as reference pictures in encoding pictures other than the current picture, as described above with respect to motion compensated prediction.
[0051]
[0077] The VVC standard decoder further configures one or more processors of the computing system to output the reconstructed picture from the DPB to a display viewable by a user of the computing system, such as a television display, a personal computing monitor, a smartphone display, a tablet display, etc.
[0052]
[0078] Thus, as shown in the encoding process 100 and the decoding process 150 described above, the VVC standard encoder and the VVC standard decoder, respectively, implement motion prediction encoding according to the VVC specification. The VVC standard encoder and the VVC standard decoder, respectively, configure one or more processors of a computing system to generate a reconstructed picture based on a previous reconstructed picture of the DPB according to the motion compensation prediction described by the VVC standard, where the previous reconstructed picture serves as a reference picture in the motion compensation prediction described herein.
[0053]
[0079] As explained above, when a coded block is transmitted in a bitstream, a significant portion of the transmission bitrate is consumed by the transmission of residual coefficient codes, each code being signaled by one bit (hereinafter referred to as "code bits"), and the VVC standard does not provide any technique for further compressing these code bits to achieve further compression. In the ongoing efforts of the JVET to develop compression techniques beyond the scope of the VVC specification (presented as "Algorithm Description for Enhanced Compression Model 3 (ECM3)" at the 136th Meeting of the Moving Picture Experts Group ("MPEG") in October 2021), it is proposed to replace this operation by predicting and encoding the residual coefficient codes (hereinafter referred to as "codes" for simplicity) as well (hereinafter referred to as "residual code prediction" or "code prediction" techniques, further details of which will be described later).
[0054]
[0080] ECM3 proposes a residual code prediction technique based on the high correlation between codes that frequently occur across the boundaries of neighboring TBs (this correlation remains after the transform operation). The ECM3 proposal provides VVC standard encoders and decoders to predict the codes of rows and columns of TBs in the neighborhood of other blocks, which may save bits when encoding individual TBs, since the neighboring TBs contain redundant information. Furthermore, within a TB, regardless of its size, only codes in the topmost and leftmost 4x4 region of the TB are predicted (so that VVC standard encoders and decoders may use neighboring reconstructed blocks for prediction, as described below), while other codes of the same TB are still signaled by the EP bins.
[0055]
[0081] For the purposes of understanding the exemplary embodiments of the present disclosure, the ECM3 residual code prediction proposal is described in further detail below.
[0056]
[0082] Figure 2 shows a 4x4 region of the TB with size 16x8 pixels, and according to the ECM3 proposal, the residual code prediction technique only predicts the codes within this region.
[0057]
[0083] Furthermore, according to the residual code prediction of ECM3, up to a maximum number of codes can be predicted for each TB. Assume that the maximum number of predicted codes is represented by the variable maxNumPredSigns. The value of maxNumPredSigns is signaled towards the decoder through the SPS. Allowed values of maxNumPredSigns are between 0 and 8, inclusive. If the value of maxNumPredSigns is 0, the code prediction method is disabled for the sequence. If maxNumPredSigns is equal to 8, then at most 8 codes are predicted for the TB. If the number of codes in the top left 4x4 region is greater than the maximum allowed value (maxNumPredSigns), then only the first maxNumPredSigns codes, determined in raster scan order as described below, are predicted.
[0058]
[0084] Figure 3 shows the scanning order of the topmost and leftmost 4x4 region of the TB to determine the predicted sign according to the ECM3 proposal. The 4x4 region of the TB is first converted to a one-dimensional array by scanning the residual coefficients in raster scan order, in which the first n signs (in raster scan order) are predicted. As Figure 3 shows, the first 8 signs (assuming maxNumPredSigns==8) in raster scan order are predicted and subsequently coded, and the remaining subsequent signs are conveyed as EP bins in CABAC bypass mode. In this way, some signs are predicted and coded rather than signaled, thus saving sign bits in transmission. It should be understood that coefficients with value 0 have no sign, so no prediction or signaling is required.
[0059]
[0085] Each sign can be either positive or negative, so if n signs are predicted, 2 n There are 2 possible code value combinations. Therefore, a VVC standard encoder and decoder implemented according to ECM3 have 2 nThe predicted code can be reconstructed by performing simplified residual boundary reconstructions (hence, keeping n small avoids high computational complexity). The boundary residuals are predicted from neighboring reconstructed blocks and the current block undergoing prediction in the decoding process, and since the blocks are decoded and reconstructed in raster scan order, before the current block is decoded, the upper neighboring block and the left neighboring block have already been previously decoded and reconstructed, and the reconstructed pixels of those neighboring blocks are available to both the VVC standard encoder and the VVC standard decoder.
[0060]
[0086] For a particular boundary reconstruction, only the leftmost and topmost residual values of the block are reconstructed from inverse quantization and subsequent inverse transformation. The residual boundary reconstruction for each combination of code values is called a hypothesis. If n represents the number of predicted codes, the number of hypotheses is 2 n and the boundary residual reconstruction is 2 n It is executed times.
[0061]
[0087] Figures 4A and 4B show an example of residual boundary prediction for an 8x4 TB according to the ECM3 proposal. The residuals of the top row of the boundary are predicted from the reconstructed neighborhood of the top first row (shown in Figure 4A) or the top first two rows (shown in Figure 4B), and the residuals of the left column are predicted from the reconstructed neighborhood of the left first column (shown in Figure 4A) or the left first two columns. Here, (0,0) indicates the top left position of the TB where the code is predicted. According to Figure 4A, the boundary residuals of the top row and left column are predicted as follows: ResPred x,0 =R x,-1 -P x,0 ResPred 0,y =R -1,y -P 0,y
[0062]
[0088] According to FIG. 4B, the boundary residuals in the top row and left column are predicted as follows: ResPred x,0 =-R x,-2 +2Rx,-1 -P x,0 ResPred 0,y =-R -2,y +2R -1,y -P 0,y
[0063]
[0089] Here, ResPred x,y R represents the predicted boundary residual of the TB location (x,y) (shown in dark shading in Figures 4A and 4B). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at position (x,y).
[0064]
[0090] Figure 5 shows the predicted boundary residuals derived from the above formula according to the ECM3 proposal shown in Figure 4B. In the top row of the TB, the predicted boundary residuals are determined based on two reconstructed neighboring pixels offset vertically by 1 pixel and 2 pixels, respectively. In the leftmost row of the TB, the predicted boundary residuals are determined based on two reconstructed neighboring pixels offset horizontally by 1 pixel and 2 pixels, respectively.
[0065]
[0091] 2 n Having determined the hypotheses and the predicted boundary residuals, the encoder and decoder may test each hypothesis against the predicted boundary residuals to determine which hypothesis minimizes a cost function, where the cost function is configured such that the higher the discontinuity between neighboring blocks, the higher the cost. The cost function is defined as the sum of absolute differences ("SAD") between the predicted residual and the reconstructed residual of the hypothesis, as follows:
[0066]
number
[0067]
[0092] Here, ResPred x,yis defined above, along with (0,0) which represents the top-left position of the TB where the code is predicted, and ResRec x,y is the square root of the n are the candidates for each reconstructed residual of a hypothesis among the hypotheses, and w and h are the width and height of the TB.
[0068]
[0093] To derive each predicted code, the encoder and decoder evaluate the above cost function for each hypothesis (i.e., candidate), and the hypothesis that results in the smallest cost is selected as the predictor for the code.
[0069]
[0094] Finally, with ECM3 residual code prediction, the predicted code is no longer EP signaled in the bitstream, but is replaced by a "residual" (signaled using the associated CABAC context) that indicates whether the prediction was correct or not. In this disclosure, we use the variable errSignPred to define the code "residual." A value of errSignPred equal to 0 indicates that the predicted code is the same as the actual code, and a value of errSignPred equal to 1 indicates that the predicted code is different from the actual code.
[0070]
[0095] It should be further understood that the VVC standard encoder and decoder predict each of the n predicted symbols one by one, and refer to the errSignPred value of the previous prediction when making the subsequent prediction. According to ECM3, the predicted symbols are the first n symbols of the block in raster scan order (described in detail later with reference to FIG. 21).
[0071]
[0096] For the first predicted code, the VVC standard encoder and decoder have no previous errSignPred value. nThe cost function above is evaluated for each of the hypotheses, the hypothesis that yields the smallest cost among them is selected, and the first code of that hypothesis is designated as the predictor of the first code of the current block. The errSignPred flag of the first code is then signaled to indicate whether the prediction of the first code is correct or not.
[0072]
[0097] For the second predicted symbol, the VVC-standard encoder and decoder know the previously determined errSignPred value for the first predicted symbol and can therefore derive the actual value of the first symbol, so there are only two hypotheses to evaluate for the remaining n-1 symbols. n-1 The VVC standard encoder and decoder have only two n-1 The cost function is evaluated for each of the hypotheses, the hypothesis that results in the smallest cost is selected, and the second code of that hypothesis (i.e., the first code among the remaining n-1 codes) is designated as the predictor of the second code of the current block. Then, the errSignPred flag of the second code is signaled to indicate whether the prediction of the second code is correct or not.
[0073]
[0098] For the third predicted symbol, the VVC-standard encoder and decoder know the previously determined errSignPred value for the second predicted symbol and can therefore derive the actual value of the second symbol along with the actual value of the first symbol, so there are only two hypotheses to evaluate for the remaining n-2 symbols. n-2 The VVC standard encoder and decoder have only two n-2 The cost function is evaluated for each of the hypotheses, the hypothesis that results in the smallest cost is selected, and the third code of that hypothesis (i.e., the first code among the remaining n-2 codes) is designated as the predictor of the third code of the current block. Then, the errSignPred flag of the third code is signaled to indicate whether the prediction of the third code is correct or not.
[0074]
[0099] The VVC standard encoder and decoder repeat the above steps for each subsequent predicted code until the last predicted code. For the last predicted code, the encoder and decoder know all previously determined errSignPred values for all other predicted codes and can derive the actual values of all other codes, so there are only two hypotheses to evaluate for the remaining code. The encoder and decoder evaluate the above cost function for both hypotheses, select the hypothesis that results in the smallest cost, and designate the last code of that hypothesis (i.e., the remaining code) as the predictor of the last code of the current block. Then, the errSignPred flag of the last code is signaled to indicate whether the prediction of the last code is correct or not.
[0075]
[0100] However, as mentioned above, the residual code prediction of ECM3 still has drawbacks. Since the boundary residual is predicted from the reconstructed pixel of the upper neighborhood or the reconstructed pixel of the left neighborhood, the prediction may not be accurate when the image data of the block contains a diagonal visible border. In addition, there is still a limitation in performing the residual code prediction based only on the top-left 4×4 region of the TB, and it is desirable to predict more residual coefficient codes to achieve a larger bitrate improvement. In addition, there is also a limitation in predicting only the first n codes in the raster scan order, and it is desirable to predict more codes without excessively increasing the computational complexity.
[0076]
[0101] Furthermore, the above cost functions do not accurately measure the actual cost of each hypothesis in all cases; for example, if an image contains objects whose edges coincide with the boundaries of neighboring TBs, the assumption that there is a high correlation between residual coefficient codes across the boundaries of neighboring TBs becomes invalid, and the assumption that neighboring TBs contain redundant information becomes invalid.
[0077]
[0102] Thus, exemplary embodiments of the present disclosure provide a residual code prediction method that provides improvements over ECM3 in several respects.
[0078]
[0103] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes a gradient-based boundary residual code prediction process.
[0079]
[0104] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a residual code prediction method that utilizes multi-modal residual boundary prediction and further configures the residual boundary prediction by signaling; by calculating a predicted template; and by deriving the signaling.
[0080]
[0105] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and a VVC standard decoder that implement a residual code prediction method that utilizes multi-modal residual boundary prediction to perform two-stage residual code prediction.
[0081]
[0106] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes code prediction over the entire TB or an extended region of the TB.
[0082]
[0107] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes an extended maximum number of predicted codes.
[0083]
[0108] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that exploits a sorting order of residual coefficients.
[0084]
[0109] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes a non-raster scan order.
[0085]
[0110] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes boundary prediction for extended boundaries.
[0086]
[0111] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes an extended range of residual signaling.
[0087]
[0112] Each of the above aspects of exemplary embodiments of the present disclosure will now be described in further detail.
[0088]
[0113] 6A and 6B illustrate the utilization of a gradient-based boundary residual code prediction process in a residual code prediction method according to an exemplary embodiment of the present disclosure.
[0089]
[0114] According to the ECM3 proposal, the predicted boundary residual is determined based on the vertically offset or horizontally offset reconstructed samples of neighboring blocks. However, in the coded video bitstream, correlations between residual coefficient codes across the boundaries of neighboring TBs may occur not only vertically and horizontally, but also diagonally (i.e., when the image contains an object with a visible diagonal border across neighboring partitioned blocks), and the residual code prediction of ECM3 tends to miss these diagonal correlations.
[0090]
[0115] 6A and 6B, in accordance with an exemplary embodiment of the present disclosure, a VVC standard encoder and a VVC standard decoder each configure one or more processors of a computing system to predict a boundary residual from the nearest eight reconstructed neighboring samples of a predicted pixel. Both the case where the predicted pixel is in the top row of TB and the case where the predicted pixel is in the leftmost column of TB are shown, with the predicted pixel labeled X and its nearest neighbors labeled C-E and J-N.
[0091]
[0116] The predicted residual, ResPred, is calculated by subtracting the prediction (P) from the weighted average of the three neighbors C, D, and E as follows:
[0092]
number
[0093]
[0117] W0, W1, and W2 are weighting coefficients, calculated as the inverse of the gradient values in the three directions.
[0094]
number
[0095]
number
[0096]
[0118] It should be appreciated that, considering the distribution of the eight nearest neighbors, three general patterns of correlation between residual coefficient signs can be observed: diagonally left to right, diagonally right to left, and either vertical (top row) or horizontal (leftmost column). Figure 7 shows how these patterns can occur among the eight nearest neighbors of the predicted pixel in the top row. These patterns occur in a similar rotation for the predicted pixel in the leftmost column.
[0097]
[0119] The gradients of the three origin directions are calculated as follows: Grad0=(JC) 2 +(KD) 2 +(LE) 2 Grad1=(KC) 2 +(LD) 2 +(ME) 2 Grad2=(LC) 2 +(MD) 2 +(NE) 2
[0098]
[0120] Grad0 represents the gradient of the pattern occurring diagonally to the right, Grad1 represents the gradient of the pattern occurring vertically, and Grad2 represents the gradient of the pattern occurring diagonally to the left. A low value of Grad0 indicates that the predicted value of pixel X is more similar to neighborhood C, a low value of Grad1 indicates that the predicted value of pixel X is more similar to neighborhood D, and a low value of Grad2 indicates that the predicted value of pixel X is more similar to neighborhood E.
[0099]
[0121] The steps for computing the gradient at the predicted pixel along the leftmost column of TB can be similarly derived from the above description.
[0100]
[0122] According to the ECM3 proposal, the residual boundary prediction described above with reference to Figures 4A, 4B, and 5 is performed uniformly for all TBs, and this method does not provide optimal compression performance for all kinds of image content, especially when the image data of a block contains significant textures and borders. Therefore, an exemplary embodiment of the present disclosure provides multi-modal residual boundary prediction to improve the accuracy of the residual boundary prediction method for various image contents.
[0101]
[0123] According to an exemplary embodiment of the present disclosure, the VVC standard encoder configures one or more processors of a computing device to check each mode among a plurality of residual code prediction modes, select a mode, and signal the selected mode to a VVC standard decoder. The exemplary embodiment of the present disclosure should not be construed as being limited to a particular mode and a particular number of modes. As an example, nine code prediction modes are described below, where the above-mentioned ECM3 residual boundary prediction method is referred to as the "default mode".
[0102]
[0124] Other code prediction modes according to exemplary embodiments of the present disclosure are described as "right diagonal mode," "left diagonal mode," "first mixed mode" (including both right diagonal mode in the top row and left diagonal mode in the leftmost column), "second mixed mode" (including both left diagonal mode in the top row and right diagonal mode in the leftmost column), "third mixed mode" (including both default mode in the top row and left diagonal mode in the leftmost column), "fourth mixed mode" (including both default mode in the top row and right diagonal mode in the leftmost column), "fifth mixed mode" (including both right diagonal mode in the top row and default mode in the leftmost column), and "sixth mixed mode" (including both left diagonal mode in the top row and default mode in the leftmost column).
[0103]
[0125] 8A, 8B, 8C, 8D, 8E, and 8F show an exemplary residual boundary prediction of a current block 800 (here, TB) according to the diagonal right mode of the present disclosure. As shown in FIG. 8A, 8B, and 8C, a VVC standard encoder according to an exemplary embodiment of the present disclosure configures one or more processors of a computing device to select the diagonal right mode instead of the two neighboring rows and the two neighboring columns, and the residual of the top row 802 and the leftmost column 804 of the boundary are predicted only from the reconstructed neighbors of the first row above 806 and the first column to the left 808. Here, (0,0) indicates the top-left position of the TB where the code is predicted. Each residual of both the top row 802 and the leftmost column 804 is predicted from a different reconstructed sample to its left, for example, as follows, for example: ResPred x,0 =R x-1,-1 -P x,0 ResPred 0,y =R -1,y-1 -P 0,y
[0104]
[0126] However, as Figures 8D, 8E, and 8F show, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from reconstructed neighbors in the first two rows above 806 and 810 and the first two columns to the left 808 and 812, where (0,0) indicates the top left location of the TB from which the code is predicted. Each residual in both the top row 802 and the leftmost column 804 is predicted from two different reconstructed sample pixels respectively to its top left, for example, as follows, for example: ResPred x,0 =2R x-1,-1 -R x-2,-2 -P x,0 ResPred 0,y =2R -1,y-1 -R -2,y-2 -P 0,y
[0105]
[0127] Here (as mentioned above and repeated here for ease of reference), ResPred x,y R represents the predicted boundary residual of the TB location (x,y) (shown in dark shading in Figures 8A, 8B, 8C, 8D, 8E, and 8F). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at position (x,y).
[0106]
[0128] 9A, 9B, 9C, 9D, 9E, and 9F show an example of residual boundary prediction of a current block 800 (here, TB) according to the diagonal left mode of the present disclosure. Similar to FIG. 8A, 8B, and 8C, in FIG. 9A, 9B, and 9C, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighbors of the first row 806 above and the first column 808 to the left, where (0,0) indicates the top left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from a different reconstructed sample to its upper right, and each residual of the leftmost column 804 is predicted from a different reconstructed sample to its lower left, for example, as follows: ResPred x,0 =R x+1,-1 -P x,0 ResPred 0,y =R -1,y+1 -P 0,y
[0107]
[0129] However, similar to Figures 8D, 8E, and 8F, in Figures 9D, 9E, and 9F, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from reconstructed neighbors in the first two rows 806 and 810 above and the first two columns 808 and 812 to the left, where (0,0) indicates the top left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from two different reconstructed sample pixels to its upper right and each residual of the leftmost column 804 is predicted from two different reconstructed sample pixels to its lower left, for example, as follows, by way of example: ResPred x,0 =2R x+1,-1 -R x+2,-2 -P x,0 ResPred 0,y =2R -1,y+1 -R -2,y+2 -P 0,y
[0108]
[0130] Here (as mentioned above and repeated here for ease of reference), ResPredx,y R represents the predicted boundary residual of the TB location (x,y) (shown in dark shading in Figures 9A, 9B, 9C, 9D, 9E, and 9F). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at position (x,y).
[0109]
[0131] 10A, 10B, 10C, 10D, 10E, and 10F show an example of residual boundary prediction of a current block 800 (here, TB) according to the first mixed mode of the present disclosure. Similar to FIG. 8A, 8B, 8C, 9A, 9B, and 9C, in FIG. 10A, 10B, and 10C, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighbors of the first row 806 above and the first column 808 to the left. Here, (0,0) indicates the top-left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from a different reconstructed sample above it and each residual of the leftmost column 804 is predicted from a different reconstructed sample below it, for example, as follows: ResPred x,0 =R x-1,-1 -P x,0 ResPred 0,y =R -1,y+1 -P 0,y
[0110]
[0132] However, similar to Figures 8D, 8E, 8F, 9D, 9E, and 9F, in Figures 10D, 10E, and 10F, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from reconstructed neighbors in the first two rows 806 and 810 above and the first two columns 808 and 812 to the left, where (0,0) indicates the top left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from two different reconstructed sample pixels to its left and above, and each residual of the leftmost column 804 is predicted from two different reconstructed sample pixels to its left and below, as follows: ResPred x,0 =2R x-1,-1 -R x-2,-2 -P x,0 ResPred 0,y =2R -1,y+1 -R -2,y+2 -P 0,y
[0111]
[0133] Here (as mentioned above and repeated here for ease of reference), ResPred x,y R represents the predicted boundary residual of the TB location (x,y) (shown in dark shading in Figures 10A, 10B, 10C, 10D, 10E, and 10F). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at position (x,y).
[0112]
[0134] 11A, 11B, 11C, 11D, 11E, and 11F show an example of residual boundary prediction of a current block 800 (here, TB) according to the second mixed mode of the present disclosure. Similar to FIG. 8A, 8B, 8C, 9A, 9B, 9C, 10A, 10B, and 10C, in FIG. 11A, 11B, and 11C, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighbors of the first row 806 above and the first column 808 to the left. Here, (0,0) indicates the top left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from a different reconstructed sample to its upper right, and each residual of the leftmost column 804 is predicted from a different reconstructed sample to its upper left, for example, as follows: ResPred x,0 =R x+1,-1 -P x,0 ResPred 0,y =R -1,y-1 -P 0,y
[0113]
[0135] However, in Fig. 11D, Fig. 11E, and Fig. 11F, similar to Fig. 8D, Fig. 8E, Fig. 8F, Fig. 9D, Fig. 9E, Fig. 9F, Fig. 10D, Fig. 10E, and Fig. 10F, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighbors of the first two rows 806 and 810 above and the first two columns 808 and 812 to the left, where (0,0) indicates the top left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from two different reconstructed sample pixels to its upper right and each residual of the leftmost column 804 is predicted from two different reconstructed sample pixels to its upper left, for example, as follows, by way of example: ResPred x,0 =2R x+1,-1 -2R x+2,-2 -P x,0 ResPred 0,y =2R -1,y-1 -R -2,y-2 -P 0,y
[0114]
[0136] Here (as mentioned above and repeated here for ease of reference), ResPred x,y R represents the predicted boundary residual of the TB location (x,y) (shown in dark shading in Figures 11A, 11B, 11C, 11D, 11E, and 11F). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at position (x,y).
[0115]
[0137] 12A, 12B, 12C, 12D, 12E, and 12F show examples of residual boundary prediction of a current block 800 (here, TB) according to the third mixed mode of the present disclosure. Similar to FIG. 8A, 8B, 8C, 9A, 9B, 9C, 10A, 10B, 10C, 11A, 11B, and 11C, in FIG. 12A, 12B, and 12C, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighborhood of the first row above 806 and the first column to the left 808. Here, (0,0) indicates the top-left position of the TB where the code is predicted. Each residual in the top row 802 is predicted from a different reconstructed sample pixel above, and each residual in the leftmost column 804 is predicted from a different reconstructed sample pixel below it to the left, for example, as follows: ResPred x,0 =R x,-1 -P x,0 ResPred 0,y =R -1,y+1 -P 0,y
[0116]
[0138] However, similar to Fig. 8D, Fig. 8E, Fig. 8F, Fig. 9D, Fig. 9E, Fig. 9F, Fig. 10D, Fig. 10E, Fig. 10F, Fig. 11D, Fig. 11E, and Fig. 11F, in Fig. 12D, Fig. 12E, and Fig. 12F, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from reconstructed neighbors of the first two rows 806 and 810 above and the first two columns 808 and 812 to the left, where (0,0) indicates the top left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from two different reconstructed sample pixels above, and each residual of the leftmost column 804 is predicted from two different reconstructed sample pixels below it to the left, for example, as follows: ResPred x,0 =2R x,-1 -R x,-2 -P x,0 ResPred 0,y =2R -1,y+1 -R -2,y+2 -P0,y
[0117]
[0139] Here (as mentioned above and repeated here for ease of reference), ResPred x,y R represents the predicted boundary residual of the (x,y) location of the TB (shown in dark shading in Figures 12A, 12B, 12C, 12D, 12E, and 12F). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at the (x,y) position.
[0118]
[0140] 13A, 13B, 13C, 13D, 13E, and 13F show examples of residual boundary prediction of a current block 800 (here, TB) according to the fourth mixed mode of the present disclosure. Similar to FIG. 8A, 8B, 8C, 9A, 9B, 9C, 10A, 10B, 10C, 11A, 11B, 11C, 12A, 12B, and 12C, in FIG. 13A, 13B, and 13C, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighborhood of the first row above 806 and the first column to the left 808. Here, (0,0) indicates the top-left position of the TB where the code is predicted. Each residual in the top row 802 is predicted from a different reconstructed sample pixel above, and each residual in the leftmost column 804 is predicted from a different reconstructed sample pixel above and to the left, for example, as follows: ResPred x,0 =R x,-1 -P x,0 ResPred 0,y =R -1,y-1 -P 0,y
[0119]
[0141] However, similar to Fig. 8D, Fig. 8E, Fig. 8F, Fig. 9D, Fig. 9E, Fig. 9F, Fig. 10D, Fig. 10E, Fig. 10F, Fig. 11D, Fig. 11E, Fig. 11F, Fig. 12D, Fig. 12E, and Fig. 12F, in Fig. 13D, Fig. 13E, and Fig. 13F, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from reconstructed neighbors of the first two rows 806 and 810 above and the first two columns 808 and 812 to the left, where (0,0) indicates the top left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from two different reconstructed sample pixels above, and each residual of the leftmost column 804 is predicted from two different reconstructed sample pixels above and to the left of it, for example, as follows: ResPred x,0 =2R x,-1 -2R x,-2 -P x,0 ResPred 0,y =2R -1,y-1 -R -2,y-2 -P 0,y
[0120]
[0142] Here (as mentioned above and repeated here for ease of reference), ResPred x,y R represents the predicted boundary residual of the (x,y) location of the TB (shown in dark shading in Figures 13A, 13B, 13C, 13D, 13E, and 13F). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at the (x,y) position.
[0121]
[0143] Figures 14A, 14B, 14C, 14D, 14E, and 14F show examples of residual boundary prediction of a current block 800 (here, TB) according to the fifth mixed mode of the present disclosure. Similar to Figures 8A, 8B, 8C, 9A, 9B, 9C, 10A, 10B, 10C, 11A, 11B, 11C, 12A, 12B, 12C, 13A, 13B, and 13C, in Figures 14A, 14B, and 14C, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighborhood of the first row above and the first column to the left. Here, (0,0) indicates the top-left position of the TB where the code is predicted. Each residual in the top row 802 is predicted from a different reconstructed sample pixel to its upper left, and each residual in the leftmost column 804 is predicted from a different reconstructed sample pixel to its left, for example, as follows: ResPred x,0 =R x-1,-1 -P x,0 ResPred 0,y =R -1,y -P 0,y
[0122]
[0144] However, similar to Fig. 8D, Fig. 8E, Fig. 8F, Fig. 9D, Fig. 9E, Fig. 9F, Fig. 10D, Fig. 10E, Fig. 10F, Fig. 11D, Fig. 11E, Fig. 11F, Fig. 12D, Fig. 12E, Fig. 12F, Fig. 13D, Fig. 13E, and Fig. 13F, in Fig. 14D, Fig. 14E, and Fig. 14F, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighbors of the first two rows 806 and 810 above and the first two columns 808 and 812 to the left, where (0,0) indicates the top left position of the TB where the code is predicted. Each residual of the top row 802 is predicted from two different reconstructed sample pixels to its top left, and each residual of the leftmost column 804 is predicted from two different reconstructed sample pixels to its left, for example, as follows: ResPred x,0 =2R x-1,-1 -R x-2,-2 -P x,0 ResPred 0,y =2R -1,y -R -2,y -P 0,y
[0123]
[0145] Here (as mentioned above and repeated here for ease of reference), ResPred x,y R represents the predicted boundary residual of the (x,y) location of the TB (shown in dark shading in Figures 14A, 14B, 14C, 14D, 14E, and 14F). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at the (x,y) position.
[0124]
[0146] 15A, 15B, 15C, 15D, 15E, and 15F show examples of residual boundary prediction of a current block 800 (here, TB) according to the sixth mixed mode of the present disclosure. Similar to Fig. 8A, 8B, 8C, 9A, 9B, 9C, 10A, 10B, 10C, 11A, 11B, 11C, 12A, 12B, 12C, 13A, 13B, 13C, 14A, 14B, and 14C, in Fig. 15A, 15B, and 15C, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighborhood of the first row above 806 and the first column to the left 808. Here, (0,0) indicates the top-left position of the TB where the code is predicted. Each residual in the top row 802 is predicted from a different reconstructed sample pixel to the right and above it, and each residual in the leftmost column 804 is predicted from a different reconstructed sample pixel to the left of it, for example, as follows: ResPred x,0 =R x+1,-1 -P x,0 ResPred 0,y =R -1,y -P 0,y
[0125]
[0147] However, similar to Fig. 8D, Fig. 8E, Fig. 8F, Fig. 9D, Fig. 9E, Fig. 9F, Fig. 10D, Fig. 10E, Fig. 10F, Fig. 11D, Fig. 11E, Fig. 11F, Fig. 12D, Fig. 12E, Fig. 12F, Fig. 13D, Fig. 13E, Fig. 13F, Fig. 14D, Fig. 14E, and Fig. 14F, in Fig. 15D, Fig. 15E, and Fig. 15F, the residuals of the top row 802 and the leftmost column 804 of the boundary are predicted from the reconstructed neighbors of the first two rows 806 and 810 above and the first two columns 808 and 812 to the left, where (0,0) indicates the top left position of the TB whose code is predicted. Each residual of the top row is predicted from two different reconstructed sample pixels to its upper right and each residual of the leftmost column is predicted from two different reconstructed sample pixels to its left, for example, as follows: ResPred x,0 =2R x+1,-1 -2R x+2,-2 -P x,0 ResPred 0,y =2R -1,y -R -2,y -P 0,y
[0126]
[0148] Here (as mentioned above and repeated here for ease of reference), ResPred x,y R represents the predicted boundary residual of the (x,y) location of the TB (shown in dark shading in Figures 15A, 15B, 15C, 15D, 15E, and 15F). x,y represents the reconstructed pixels of the block in the neighborhood of the position (x, y), and P x,y represents the predicted signal of the current block at the (x,y) position.
[0127]
[0149] The nine code prediction modes described above are summarized in Table 1 below.
[0128] [Table 1]
[0129]
[0150] According to an example embodiment in which a VVC standard encoder and a VVC standard decoder implement multiple boundary prediction modes (e.g., but not limited to, the nine prediction modes shown in Table 1), the VVC standard encoder needs to signal which prediction mode is selected for final encoding. Because the VVC standard does not include multi-modal residual boundary prediction, an example embodiment of the present disclosure further provides a VVC standard encoder that implements signaling of the boundary prediction mode in the syntax structure of the coded block, and a VVC standard decoder that implements parsing of the boundary prediction mode from the syntax structure of the coded block.
[0130]
[0151] One or more processors of the computing system may be configured to signal the boundary prediction mode in a TB level syntax structure, and / or a CU level syntax structure, and / or a CTU level syntax structure, and / or a slice level syntax structure, and / or a picture level syntax structure, and / or a sequence level syntax structure. When the signaling of the boundary prediction mode is a TB level syntax structure, for each TB, the encoder selects the boundary prediction mode by minimizing the rate-distortion cost and signals it to the decoder.
[0131]
[0152] As an example, the exemplary embodiment of the present disclosure introduces two syntax elements into the block semantics according to the VVC standard: For each TB, the VVC standard encoder configures one or more processors of the computing system to signal in the syntax structure of the TB a first flag (for reference herein, referred to as default_mode_flag) to indicate whether the default mode is applied to the residual code prediction of that TB. If the value of default_mode_flag is equal to 1, the VVC standard encoder applies the default mode to the final encoding.
[0132]
[0153] If the value of default_mode_flag is equal to 0, the VVC standard encoder configures one or more processors of the computing system to signal in the syntax structure of the TB a second flag (for reference herein, referred to as borderModeIdc) that uses 3 bits (according to Table 2 below) to indicate which non-default border prediction mode the VVC standard encoder applies to the residual code prediction of the TB. The value borderModeIdc indicates the border prediction mode used for that TB. An example of a borderModeIdc value is shown in Table 2 below.
[0133] [Table 2]
[0134]
[0154] As an example, a VVC standard encoder may further implement signaling of the boundary prediction mode according to the following pseudocode: For each transformation block Signaling default_mode_flag When default_mode_flag==0 Signaling borderModeIdc
[0135]
[0155] However, since a VVC standard encoder that implements signaling of boundary prediction modes in a block syntax structure adds signaling bits to each coded picture, additional overhead is introduced into bitstream transmission and compression efficiency may be affected. To reduce the signaling overhead, the above multi-modal residual boundary prediction according to an exemplary embodiment of the present disclosure may further provide a template-based boundary mode derivation.
[0136]
[0156] According to an exemplary embodiment of the present disclosure, the VVC standard encoder and the VVC standard decoder respectively implement the derivation of the template-based boundary prediction mode without signaling by comparing the samples of the reconstructed neighboring blocks with a plurality of templates and evaluating which of the plurality of templates best fits the actual reconstructed neighboring block. The neighboring blocks are reconstructed before encoding and decoding the current block according to the encoding order, and therefore are accessible by the encoder and the decoder during the encoding and decoding process of the current block.
[0137]
[0157] FIG. 16 shows an example of reconstructed neighboring block samples compared to the template (including upper neighboring block boundary samples highlighted in white and left neighboring block samples highlighted in white). The rows of reconstructed pixels of the first three neighbors above and the columns of reconstructed pixels of the first three neighbors to the left are used to calculate the cost of each boundary prediction mode. For each boundary prediction mode described above with reference to Tables 1 and 2, the VVC standard encoder and VVC standard decoder configure one or more processors of the computing system to generate a mode template by inversely applying the respective boundary prediction mode to the corresponding samples of the current block.
[0138]
[0158] In other words, the mode template describes, for each boundary prediction mode, the ideal expected values of the reconstructed samples of neighboring blocks when that boundary prediction mode matches the corresponding samples of the current block.
[0139]
[0159] The suitability of each boundary prediction mode is evaluated by calculating a cost function between the mode template and the corresponding samples of the reconstructed neighboring blocks. For example, the SAD is calculated as the cost of each boundary prediction mode between the mode template and the corresponding samples of the reconstructed neighboring blocks as follows:
[0140]
number
[0141]
[0160] Alternatively, other cost functions such as the sum of absolute transform differences ("SATD") can be used instead of SAD.
[0142]
[0161] Here, R x,y represents the reconstructed pixel of the block near the position (x,y). (0,0) represents the position of the top-left sample of the current block, indicated by an arrow in Figure 16. x,y Let x,y denote the mode template value generated at location (x,y) based on each respective boundary prediction mode. The computation of the mode template for each respective boundary prediction mode is summarized in Table 3 below.
[0143] [Table 3]
[0144]
[0162] After calculating the cost function for each boundary prediction mode, the boundary prediction mode with the best cost function output is determined as the most suitable and selected for the current block.
[0145]
[0163] Alternatively, the exemplary embodiment of the present disclosure provides a VVC standard encoder and a VVC standard decoder that implement a residual code prediction method that utilizes multi-modal residual boundary prediction without signaling a first flag (e.g., the above-mentioned default_mode_flag) in the syntax structure to indicate whether a default mode is applied to the residual code prediction. In this way, one bit is saved in the bitstream transmission, and the signaling overhead is further reduced.
[0146]
[0164] Instead of explicit signaling, the first flag is derived from the parity bits of the sum of the transform coefficient levels. As an example, a VVC standard encoder and decoder can implement implicit signaling of the boundary prediction mode according to the following pseudocode: For each TB Calculate the sum of the absolute values of the conversion coefficient levels If the sum is even Set default_mode_flag=1 If not, Set default_mode_flag=0
[0147]
[0165] It should be appreciated that by determining the first flag that is implicitly signaled in this different manner, a VVC standard decoder can be implemented substantially similarly to the VVC standard decoder described above in which both the first flag and the second flag are signaled.
[0148]
[0166] The VVC standard encoder further configures one or more processors of the computing device to output a bitstream to satisfy the condition that if the sum of the absolute values of the transform coefficient levels is an even number, then the default mode is used, and if not, then the default mode is not used.
[0149]
[0167] Additionally, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a two-stage residual code prediction method that utilizes multi-modal residual boundary prediction.
[0150]
[0168] According to the VVC standard encoder and VVC standard decoder implementing the two-stage code prediction method, one or more processors of a computing system are configured to predict and reconstruct codes of TB in two stages. First, the codes of TB are classified into two different groups. The VVC standard encoder and VVC standard decoder apply only a default mode to predict the codes of the first group, and apply the above-mentioned multiple boundary prediction modes to predict the codes of the second group.
[0151]
[0169] FIG. 17 shows an example of classifying TB codes into a first group and a second group.
[0152]
[0170] The two-stage residual code prediction method performed on TB according to an embodiment of the present invention may include the following steps.
[0153]
[0171] In a first step, one or more processors of the computing system are configured to classify the n unknown codes between a first group and a second group. It should be understood that the proposed method is not limited to a specific classification or grouping method. For example, the first n / 2 codes may be classified into a first group and the rest into a second group. Alternatively, the classification or grouping may be performed in an interleaved manner, where the codes are classified into each group alternately.
[0154]
[0172] Then, a default mode code prediction is performed on the first group of codes, the default mode code prediction errors for these codes are calculated, and the most suitable mode (whether the default mode or another code prediction mode) for the first group of codes is selected, which is implemented according to the following steps.
[0155]
[0173] In a subsequent step, the one or more processors of the computing system apply a default code prediction mode to predict the codes of the first group and reconstruct the codes of the first group, so that the actual codes of the first group are now known.
[0156]
[0174] In a subsequent step, the one or more processors of the computing system select a second group of symbol prediction modes based on the symbol prediction errors of the first group, as follows:
[0157]
[0175] One or more processors of the computing system predict a first group of symbols for each symbol prediction mode and then calculate a respective prediction error (i.e., the number of correct symbols based on the reconstruction of the previous step) for each non-symbol prediction mode.
[0158]
[0176] The one or more processors of the computing system select, from among all the code prediction modes, the code prediction mode with the smallest prediction error as the selected code prediction mode.
[0159]
[0177] In a subsequent step, the one or more processors of the computing system apply the selected symbol prediction mode to predict the symbols of the second group and reconstruct the symbols of the second group.
[0160]
[0178] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes code prediction over the entire TB or an extended region of the TB.
[0161]
[0179] In ECM3, sign prediction is limited to the 4×4 area in the upper left of the TB. In contrast, exemplary embodiments of the present disclosure expand the sign prediction area to the entire transform block without limiting the position of the predicted sign within the current block (i.e., without limiting it to the upper left 4×4 area within the transform block). In other words, as long as the number of signs of the predicted TB is below the maximum allowable limit, the signs of the TB can be predicted. The maximum allowable limit is represented by maxNumPredSigns.
[0162]
[0180] Alternatively, the sign prediction area (M×N) depends on the width and height of the transform block. FIG. 18 shows an example of a TB-dependent M×N sign prediction area, where the value of M is the larger of 4 or one-fourth of the width of the TB, and the value of N is the larger of 4 or one-fourth of the height of the TB, and they are calculated as follows, respectively.
[0163]
Number
[0164]
[0181] One or more processors of the computing system apply the residual sign prediction described herein to predict a sign when (x<M && y<N), where (x, y) is the horizontal and vertical position of the sign to be predicted. Otherwise, one or more processors signal the sign by the EP bins without prediction.
[0165]
[0182] Alternatively, the sign prediction area (M×N) is fixed. Assume that maxM and maxN are the maximum values of M and N, respectively. The values of maxM and maxN can be fixed values (i.e., maxM = maxN = 32) and can be predefined for both the encoder and the decoder. The values of M and N can be determined as follows. M = max(maxM, width) N = max(maxN, height)
[0166]
[0183] Alternatively, the values of M and N can be determined as follows. M = min(maxM, width) N = min(maxN, height)
[0167]
[0184] One or more processors of the computing system predict a code by applying the residual code prediction described herein when (x < M && y < N), where (x, y) is the horizontal and vertical position of the predicted code. Otherwise, one or more processors signal the code by the EP bin without prediction.
[0168]
[0185] Alternatively, the VVC standard encoder signals the maximum value of M (denoted as maxM) and the maximum value of N (denoted as maxN) to the VVC standard decoder in a syntax structure such as SPS, PPS. The values of M and N can be derived as follows from maxM and maxN, and the width and height of the current TB. M = max(maxM, width) N = max(maxN, height)
[0169]
[0186] Alternatively, the values of M and N can be derived as follows. M = min(maxM, width) N = min(maxN, height)
[0170]
[0187] One or more processors of the computing system initialize NumPredSign to zero before predicting the first code of the TB, and then, when (x < M && y < N && NumPreSign < maxNumPreSigns), apply the residual code prediction described herein to predict the code, where (x, y) is the horizontal and vertical position of the predicted code, and increment NumPreSign by one. Otherwise, one or more processors signal the code by the EP bin without prediction.
[0171]
[0188] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes an extended maximum number of predicted codes.
[0172]
[0189] In ECM3, the maximum number of predicted signs (maxNumPredSigns) of a TB is signaled to the decoder, with a maximum allowed value of 8. According to an example embodiment of the present disclosure, the maximum number may be increased to 16. Alternatively, the value of the maximum number of predicted signs (maxNumPredSigns) may be changed per picture and signaled in syntax structures such as slice and / or picture headers.
[0173]
[0190] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes a code sorting order based on residual coefficients.
[0174]
[0191] Generally, the larger the absolute value of the transform coefficient level, the higher the accuracy of the sign prediction. Therefore, the exemplary embodiment of the present disclosure provides a residual sign prediction method in which the 2D transform coefficient levels are stored in a one-dimensional coefficient level array, as described above with reference to FIG.
[0175]
[0192] In a subsequent step, the one or more processors of the computing system sort the array in order of absolute values of the coefficient levels. The array is sorted in descending order of absolute values, with the residual coefficient level with the largest absolute value being placed at the beginning of the sorted array and the residual coefficient level with the smallest absolute value being placed at the end of the array. The first n symbols according to the corresponding array positions in the sorted array are predicted using a residual symbol prediction method, and the remaining symbols are signaled by EP bins.
[0176]
[0193] 19 illustrates an example of a residual sign prediction method utilizing a sorting order according to an exemplary embodiment of the present disclosure, whereby the sign of the maximum residual coefficient level in the TB is prioritized for prediction.
[0177]
[0194] Alternatively, the exemplary embodiment of the present disclosure provides a VVC standard encoder and a VVC standard decoder that implement a residual code prediction method that utilizes a code sorting order that is not based on residual coefficient levels. This is motivated by the observation that according to ECM3, since two quantizers (Q0 and Q1) are implemented, the absolute value of the level may not necessarily consistently represent the dequantized DCT transform coefficient level. Even for the same level value, the quantization index (QIdx) may be different.
[0178]
[0195] Figure 20 shows the mismatch of quantization indexes for the same residual coefficient level between two quantizers. It is shown that for the same level value, the reconstruction points can be different.
[0179]
[0196] Therefore, in order to improve the code sorting and further increase the code prediction accuracy, the code sorting is performed based on the corresponding QIdx value (instead of the residual coefficient level value). The QIdx of the transform coefficient level can be calculated as follows:
[0180]
[0197] QIdx=(abs(level)<<1)-(state&1)
[0181]
[0198] Here, the variable level represents the value of the quantized transform coefficient level, and "state" represents the dependent quantization state of the level. The quantization states are described in more detail in H. Schwarz, P. Haase, T. Nguyen, J. Pfaff, D. Marpe, T. Wiegand, "EE2-4.1: Results for dependent quantization with 8 states", JVET-V0082, April 2021. For the same quantization parameter, a larger QIdx represents a larger transform coefficient level after inverse quantization. Thus, the first n codes according to the corresponding QIdx values ordered from maximum to minimum are predicted using the residual code prediction method, and the remaining codes are signaled by the EP bins.
[0182]
[0199] The predicted sign then needs to be signaled in the bitstream for later processing by the decoder. However, as described above with reference to Figures 4A and 4B, according to the ECM3 proposal, the first maxNumPredSigns signs determined in raster scan order are predicted, so that the first 8 non-zero signs of the block in raster scan order are predicted and signaled by their errSignPred flags, and the remaining 4 non-zero signs are not predicted and signaled by the EP bins, as shown in Figure 21.
[0183]
[0200] According to ECM3, both the errSignPred flags of the predicted coefficient levels and the EP bin signaling of the non-predicted coefficients are signaled in the original raster scan order. Figure 21 shows that the errSignPred flags of the predicted codes are signaled first in the bitstream, followed by the EP bin signals of the remaining non-predicted codes.
[0184]
[0201] However, according to the above-mentioned residual code prediction method that utilizes the sorting order of the residual coefficients, the predicted codes are not necessarily the first N codes in the raster scan order, and this result is shown in Figure 22. Therefore, if the predicted codes and the non-predicted codes are coded in the bitstream in the raster scan order, they will be interleaved as shown in Figure 22.
[0185]
[0202] To solve this problem, in one or more aspects, the exemplary embodiment of the present disclosure provides a VVC standard encoder and a VVC standard decoder that implement a two-pass bitstream signaling method for a residual code prediction method that utilizes a sorting order, as shown in Figure 23. In the first pass, each predicted errSignPred flag for the code of the block is signaled in the bitstream in raster scan order or level absolute value order before signaling each non-predicted code. In the second pass, each remaining non-predicted code is signaled in raster scan order with an EP bin signal, which follows the predicted flag for the code.
[0186]
[0203] For a CABAC entropy decoder configured to process a bitstream, it is beneficial for the EP bin signals to be grouped together. While parsing the bitstream of the coded residuals of the predicted code and the bypass coded code, the entropy decoder 152 determines whether the currently parsed bin is the coded residuals of the predicted code or the bypass coded code. All the coded residuals of the predicted code are parsed before the bypass coded code. The entropy decoder 152 parses the residuals of the code and the bypass coded code and stores them in a buffer. Then, in the reconstruction process, the VVC standard decoder can perform code selection to assign the residuals of the code and correctly assign the bypass coded code to the coefficients. The entropy decoder 152 does not need to perform code selection in the parsing process and can perform code selection later in the reconstruction process.
[0187]
[0204] In one or more aspects, the exemplary embodiment of the present disclosure provides a VVC standard encoder and a VVC standard decoder that implement a residual code prediction method that utilizes a non-raster scan order for prediction. Figure 24 shows a zigzag scan order that is applied to the scan order of a 4x4 region of a TB to determine a predicted code (replacing the raster scan order shown in Figure 3).
[0188]
[0205] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes boundary prediction for extended boundaries.
[0189]
[0206] According to ECM3, all hypothesis reconstruction and boundary prediction are performed only for the top row 802 and leftmost column 804 of the TB. Thus, the exemplary embodiment of the present disclosure provides boundary prediction for the extended boundary region. In one embodiment, the boundary area is extended to the entire transform block. In another embodiment, the boundary area is extended to M top rows 814 and N leftmost columns 816 (M=N=1 represents the boundary area established by ECM3). The values of M and N can be fixed (e.g., M=N=2). The values of M and N can also be signaled to the decoder at the slice / picture level or SPS level.
[0190]
[0207] FIG. 25 illustrates an example of an extended boundary for M=N=2, according to an exemplary embodiment of the present disclosure.
[0191]
[0208] VVC standard encoders and decoders implementing the extended boundary prediction method configure one or more processors of a computing system to calculate a cost function as a sum of two sub-cost functions, each measuring discontinuities across different rows of a neighboring block and a current block, or discontinuities across different columns of a neighboring block and a current block. The first sub-cost function is R -2 , R -1 The second sub-cost function measures the discontinuity between R-1, the first row of the candidate (or the first column of the left boundary but not the top boundary) and the second row of the candidate (or the second column of the left boundary but not the top boundary).
[0192]
[0209] An example of cost calculation is shown below.
[0193]
number
[0194]
[0210] Here, ResPred x,yrefers to the predicted boundary residual (shown in dark shading in Fig. 25) of the TB location (x,y), and ResRec x,y refers to the candidate reconstructed residual of the hypothesis at position (x,y) (described above with reference to Fig. 5). (0,0) represents the top-left position of the TB whose sign is predicted, w and h are the width and height of the TB, and P x,y is the predicted sample at position (x,y).
[0195]
number
[0196]
[0211] FIG. 26 shows an example of discontinuity measurement with boundary extension.
[0197]
[0212] In one or more aspects, exemplary embodiments of the present disclosure provide a VVC standard encoder and decoder that implement a residual code prediction method that utilizes an extended range of residual signaling.
[0198]
[0213] According to ECM3, for each predicted code, the encoder is configured to signal the code residual errSignPred to the decoder. If the value of errSignPred is equal to 0, the predicted code value is the same as the actual code value. If the value of errSignPred is equal to 1, the predicted code value is not the same as the actual code value.
[0199]
[0214] There are four context variables used to code errSignPred for the luma component, and four context models are used to code errSignPred for the chroma components. The derivation of the context variables depends on whether the CU is intra- or inter-coded, and whether the predicted code is a DC code or an AC code. According to ECM3, the context index is derived as follows: For each TB: int ctxOffset=CU::isIntra(*tu.cu)?0:2; For each predicted code Assume that (x,y) is the horizontal and vertical position (from the top-left position of the TB) of the symbol to be encoded. ctxIdc=(x==0 && y==0)?0:1 ctxIdc+=ctxOffset
[0200]
[0215] However, the above does not take into account the further observation that the accuracy of the code prediction method decreases as the number of predicted codes of a transform block increases. Therefore, according to an exemplary embodiment of the present disclosure, an additional context variable is introduced into the syntax structure of the signal context when the number of predicted codes of a transform block is greater than a predefined threshold. In the proposed method, six contexts are used for luma and six for chroma. A VVC standard encoder can implement the derivation of the context index according to the following pseudocode: For each TB: int ctxOffset=CU::isIntra(*tu.cu)?0:3; For each predicted code Assume that (x,y) is the horizontal and vertical position (from the top-left position of the TB) of the symbol to be encoded. ctxIdc=(x==0 && y==0)?0:1 ctxIdc+=((x||y) && numSigns <Th)?1:0 ctxIdc+=ctxOffset
[0201]
[0216] Furthermore, depending on the number of predicted symbols of TB, one of several context models may be selected. For example, if the number of predicted symbols of TB is greater than threshold A, a first context model is used, if the number of predicted symbols of TB is less than or equal to threshold A but greater than threshold B, a second context model is used, if the number of predicted symbols of TB is less than or equal to threshold B, a third context model is used, and so on.
[0202]
[0217] Alternatively, a VVC standard encoder may implement the derivation of the context index depending on the number of predicted symbols in the TB and the current bin order according to the following pseudocode: For each TB: int ctxOffset=CU::isIntra(*tu.cu)?0:maxNumPredSigns; ctxIdc=numSigns; ctxIdc=ctxIdc+ctxOffset; For each predicted symbol: Use ctxIdc to encode or decode the current bin ctxIdc=ctxIdc+1;
[0203]
[0218] Here, maxNumPredSigns represents the maximum number of predicted signs for the current sequence, and numSigns represents the number of predicted signs in the current TB.
[0204]
[0219] The residual code prediction method according to the exemplary embodiment of the present disclosure may further perform the exclusion of boundary pixels between predictions. As mentioned above, when an image contains an object whose edge coincides with the boundary of a neighboring TB, the assumption that there is a high correlation between the residual coefficient codes across the boundary of the neighboring TB is invalid, and the assumption that the neighboring TB contains redundant information is invalid.
[0205]
[0220] To overcome this limitation of the ECM3 proposed cost function, VVC standard encoders and decoders implement a residual code prediction method that excludes some boundary pixels from the cost function evaluation, in particular those boundary pixels whose content is discontinuous with respect to nearby TBs.
[0206]
[0221] As described above with reference to FIG. 4A and FIG. 4B, the VVC standard encoder and the VVC standard decoder configure one or more processors of a computing system to predict each symbol of the n predicted symbols one by one, and for the first predicted symbol, the VVC standard encoder and the VVC standard decoder perform a 2 n One or more processors of the computing system are configured to evaluate a cost function for each of the hypotheses, the cost function evaluation including a boundary pixel set N that includes all boundary pixels.
[0207]
[0222] The VVC standard encoder and decoder configure one or more processors of the computing system to select the hypothesis that results in the smallest cost among them and designate the first code of that hypothesis as a predictor of the first code of the current block. Then, the errSignPred flag of the first code is signaled to indicate whether the prediction of the first code is correct or not.
[0208]
[0223] Subsequently, we assume that the prediction of the first sign is incorrect (i.e., the sign of the current block is predicted incorrectly compared to the selected hypothesis and the corresponding sign of the current block predicted by the selected hypothesis), and henceforth, we refer to the hypothesis designated as the predictor as A and the correct hypothesis as B.
[0209]
[0224] For each of hypotheses A and B, the VVC standard encoder and the VVC standard decoder configure one or more processors of the computing system to calculate and compare the difference between the predicted residual and the reconstructed residual of each boundary pixel. That is, for boundary pixel p, the difference between the predicted residual and the reconstructed residual of hypothesis A is calculated as diff_A(p), and the difference between the predicted residual and the reconstructed residual of hypothesis B is calculated as diff_B(p). Then, diff_A_B(p) is calculated as the difference between diff_B(p) and diff_A(p). If diff_B(p) is greater than diff_A(p) (i.e., diff_A_B(p) is greater than 0), p is excluded from the boundary pixel set N.
[0210]
[0225] As an example, each pixel that produces a larger difference between the predicted and reconstructed residuals for hypothesis B than for hypothesis A is excluded from the boundary pixel set N.
[0211]
[0226] As another example, among the pixels that result in a larger difference between the predicted and reconstructed residuals for hypothesis B than for hypothesis A, only the k pixels with the largest magnitude of difference between the difference between hypothesis A and the difference between hypothesis B (i.e., the k pixels are the maximum of diff_A_B()) are excluded from the boundary pixel set N. After excluding the boundary pixels, the remaining pixels are used as the boundary pixel set N in predicting the next bin.
[0212]
[0227] Thus, each code prediction step that results in an incorrect prediction triggers the exclusion of one or more boundary pixels from the boundary pixel set N between the previous and the next prediction, and only the remaining boundary pixels are used to predict the next code. The excluded boundary pixels can be understood as pixels whose content is discontinuous with the neighboring TBs, making the cost function less reliable. By excluding such pixels, the accuracy of the cost function and, accordingly, the prediction accuracy can be improved.
[0213]
[0228] Those skilled in the art will understand that all of the above aspects of the present disclosure may be implemented simultaneously in any combination thereof, and that all aspects of the present disclosure may be implemented in combination as yet another embodiment of the present disclosure.
[0214]
[0229] FIG. 27 illustrates an example system 2700 for implementing the above-described processes and methods for implementing residual code prediction.
[0215]
[0230] The techniques and mechanisms described herein may be implemented by multiple instances of system 2700, as well as by any other computing device, system, and / or environment. System 2700 shown in FIG. 27 is only one example of a system and is not intended to suggest any limitation as to the scope of use or functionality of any computing device utilized to perform the processes and / or procedures described above. Other well-known computing devices, systems, environments, and / or configurations that may be suitable for use with the embodiments include, but are not limited to, personal computers, server computers, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, game consoles, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, implementations using field programmable gate arrays ("FPGAs") and application specific integrated circuits ("ASICs"), and / or the like.
[0216]
[0231] The system 2700 may include one or more processors 2702 and a system memory 2704 communicatively coupled to the processors 2702. The processors 2702 may execute one or more modules and / or processes to cause the processors 2702 to perform various functions. In some embodiments, the processors 2702 may include a central processing unit ("CPU"), a graphics processing unit ("GPU"), both a CPU and a GPU, or other processing units or components known in the art. Additionally, each of the processors 2702 may possess its own local memory, which may also store program modules, program data, and / or one or more operating systems.
[0217]
[0232] Depending on the exact configuration and type of system 2700, the system memory 2704 may be volatile, such as RAM, non-volatile, such as ROM, flash memory, a mini-hard drive, a memory card, or a combination thereof. The system memory 2704 may include one or more computer-executable modules 2706 executable by the processor 2702.
[0218]
[0233] The module 2706 may include, but is not limited to, one or more of an encoder 2708 and a decoder 2710 .
[0219]
[0234] The encoder 2708 may be a VVC standard encoder that is executable by the processor 902 to implement any, some, or all aspects of the exemplary embodiments of the present disclosure described above and configure the processor 902 to perform the operations described above.
[0220]
[0235] The decoder 2710 may be a VVC standard encoder that implements any, some, or all aspects of the exemplary embodiments of the present disclosure described above and is executable by the processor 2702 to configure the processor 2702 to perform the operations described above.
[0221]
[0236] The system 2700 may further include an input / output (I / O) interface 2740 for receiving image source data and bitstream data and outputting reconstructed pictures to a reference picture buffer or DBP and / or a display buffer. The system 2700 may also include a communications module 2750 that enables the system 2700 to communicate with other devices (not shown) over a network (not shown). The network may include wired media, such as the Internet, a wired network or a direct-wired connection, and wireless media, such as acoustic, radio frequency ("RF"), infrared, and other wireless media.
[0222]
[0237] Some or all of the operations of the above-described methods may be implemented by executing computer-readable instructions stored in a computer-readable storage medium, as defined below. As used in this description and claims, the term "computer-readable instructions" includes routines, applications, application modules, program modules, programs, components, data structures, algorithms, etc. The computer-readable instructions may be implemented in a variety of system configurations, including single-processor or multi-processor systems, minicomputers, mainframe computers, personal computers, handheld computing devices, microprocessor-based programmable consumer electronics, combinations thereof, and the like.
[0223]
[0238] The computer-readable storage medium may include volatile memory (e.g., random access memory ("RAM")) and / or non-volatile memory (e.g., read only memory ("ROM"), flash memory, etc.). The computer-readable storage medium may also include additional removable and / or non-removable storage devices including, but not limited to, flash memory, magnetic storage devices, optical storage devices, and / or tape storage devices that may provide non-volatile storage of computer-readable instructions, data structures, program modules, and the like.
[0224]
[0239] A non-transitory computer-readable storage medium is an example of a computer-readable medium. A computer-readable medium includes at least two types of computer-readable media: computer-readable storage media and communication media. A computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any process or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. A computer-readable storage medium includes, but is not limited to, phase-change memory ("PRAM"), static random access memory ("SRAM"), dynamic random access memory ("DRAM"), other types of random access memory ("RAM"), read-only memory ("ROM"), electrically erasable programmable read-only memory ("EEPROM"), flash memory or other memory technology, compact disk read-only memory ("CD-ROM"), digital versatile disk ("DVD") or other optical storage device, magnetic cassette, magnetic tape, magnetic disk storage device or other magnetic storage device, or other non-transmission medium that can be used to store information for access by a computing device. In contrast, communication media may embodied computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism. Computer readable storage media as used herein should not be interpreted as transitory signals per se, such as, for example, electric waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a wave guide or other transmission medium (e.g., light pulses through a fiber optic cable), or electrical signals propagating through wires.
[0225]
[0240] Computer-readable instructions stored on one or more non-transitory computer-readable storage media that, when executed by one or more processors, may perform the operations described above with reference to Figures 1A-26. Generally, computer-readable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular abstract data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement a process.
[0226]
[0241] Although the present subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as example forms of implementing the claims.
[0227]
[0242] Exemplary embodiments of the present disclosure are further described by at least the following clauses.
[0228]
[0243] A. A method including predicting, by one or more processors (2702) of a computing system (2700), residual coefficient signs of a top row (802) of a current block (800) and a leftmost column (804) of the current block (800), where at least the residual coefficient signs of the top row (802) are predicted from the respective top-left reconstructed samples or the respective top-right reconstructed samples, or the residual coefficient signs of the leftmost column (804) are predicted from the respective top-left reconstructed samples or the respective bottom-left reconstructed samples.
[0229]
[0244] B. The method of paragraph A, wherein the residual coefficient signs in the top row (802) are predicted from their respective top-left reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective top-left reconstructed samples.
[0230]
[0245] C. The method of paragraph B, wherein the residual coefficient signs in the top row (802) are predicted from their respective top-left two reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective top-left two reconstructed samples.
[0231]
[0246] D. The method of paragraph A, wherein the residual coefficient signs in the top row (802) are predicted from their respective bottom-left reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective top-right reconstructed samples.
[0232]
[0247] E. The method of paragraph D, wherein the residual coefficient signs in the top row (802) are predicted from their respective two bottom left reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective two top right reconstructed samples.
[0233]
[0248] F. The method of paragraph A, wherein the residual coefficient signs in the top row (802) are predicted from their respective bottom-left reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective top-left reconstructed samples.
[0234]
[0249] G. The method of paragraph F, wherein the residual coefficient signs in the top row (802) are predicted from the two reconstructed samples at their respective bottom left, and the residual coefficient signs in the leftmost column (804) are predicted from the two reconstructed samples at their respective top left.
[0235]
[0250] H. The method of paragraph A, wherein the residual coefficient signs in the top row (802) are predicted from their respective top-left reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective top-right reconstructed samples.
[0236]
[0251] I. The method of paragraph H, wherein the residual coefficient signs in the top row (802) are predicted from their respective top-left two reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective top-right two reconstructed samples.
[0237]
[0252] J. The method of paragraph A, wherein the residual coefficient signs in the top row (802) are predicted from the reconstructed samples respectively to the left and the residual coefficient signs in the leftmost column (804) are predicted from the reconstructed samples respectively above.
[0238]
[0253] K. The method of paragraph J, wherein the residual coefficient signs in the top row (802) are predicted from the two reconstructed samples to the lower left of each other, and the residual coefficient signs in the leftmost column (804) are predicted from the two reconstructed samples above each other.
[0239]
[0254] L. The method of paragraph A, wherein the residual coefficient signs in the top row (802) are predicted from the reconstructed sample to the respective left and the residual coefficient signs in the leftmost column (804) are predicted from the reconstructed sample above each.
[0240]
[0255] M. The method of paragraph L, wherein the residual coefficient signs in the top row (802) are predicted from the two reconstructed samples to the upper left of each other, and the residual coefficient signs in the leftmost column (804) are predicted from the two reconstructed samples above each other.
[0241]
[0256] N. The method of paragraph A, wherein the residual coefficient signs in the top row (802) are predicted from their respective left reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective top-left reconstructed samples.
[0242]
[0257] O. The method of paragraph N, wherein the residual coefficient signs in the top row (802) are predicted from the two reconstructed samples to their respective left, and the residual coefficient signs in the leftmost column (804) are predicted from the two reconstructed samples to their respective upper left.
[0243]
[0258] P. The method of paragraph A, wherein the residual coefficient signs in the top row (802) are predicted from their respective left reconstructed samples, and the residual coefficient signs in the leftmost column (804) are predicted from their respective right-upper reconstructed samples.
[0244]
[0259] Q. The method of paragraph P, wherein the residual coefficient signs in the top row (802) are predicted from the two reconstructed samples to their left, and the residual coefficient signs in the leftmost column (804) are predicted from the two reconstructed samples to their right and to their right.
[0245]
[0260] R. The method of any one of paragraphs A-Q, further comprising one of signaling a boundary prediction mode flag in a syntax structure of the current block (800) by one or more processors (2702) or parsing the boundary prediction mode flag from the syntax structure of the current block (800) by one or more processors (2702).
[0246]
[0261] S. The method of paragraph R, further comprising one of signaling a non-default boundary prediction mode flag in a syntax structure of the current block (800) by one or more processors (2702) or parsing a non-default boundary prediction mode flag from a syntax structure of the current block (800) by one or more processors (2702).
[0247]
[0262] T. The method of paragraph R, further comprising one of: signaling a boundary prediction mode flag in a syntax structure of the current block (800) without signaling a non-default boundary prediction mode flag by one or more processors (2702); or determining, by one or more processors (2702), the non-default prediction mode flag from a parity bit of a sum of transform coefficient levels of the current block (800).
[0248]
[0263] U. The method of any one of paragraphs A-T, further comprising: generating, by one or more processors (2702), a plurality of corresponding mode templates for a plurality of boundary prediction modes, each based on a top-most row (802) and a left-most column (804) of the current block (800); calculating, by the one or more processors (2702), for each of the plurality of mode templates, a cost function between the mode template and corresponding samples of the reconstructed neighboring blocks; and selecting a boundary prediction mode having an optimal cost function output for the corresponding mode template.
[0249]
[0264] V. A method comprising: classifying, by one or more processors (2702) of a computing system (2700), a first plurality of residual coefficient codes and a second plurality of residual coefficient codes of a current block (800); predicting, by the one or more processors (2702), the first plurality of residual coefficient codes by a default code prediction mode; reconstructing, by the one or more processors (2702), the first plurality of residual coefficient codes by a plurality of code prediction modes; selecting, by the one or more processors (2702), a code prediction mode having a minimum prediction error from among each of the plurality of code prediction modes; and predicting, by the one or more processors (2702), the second plurality of residual coefficient codes by the selected code prediction mode.
[0250]
[0265] W. A method comprising predicting, by one or more processors (2702) of a computing system (2700), residual coefficient signs of a current block (800) without any restriction on the location of the predicted residual coefficient signs within the current block (800).
[0251]
[0266] X. A method comprising predicting, by one or more processors (2702) of a computing system (2700), a residual coefficient code of a current block (800), wherein the predicted residual coefficient code is confined to a region of the current block (800), the width of the region of the current block (800) being the greater of four and one-quarter the width of the current block (800), and the height of the region of the current block (800) being the greater of four and one-quarter the height of the current block (800).
[0252]
[0267] Y. A method comprising predicting, by one or more processors (2702) of a computing system (2700), a residual coefficient code of a current block (800), wherein the predicted residual coefficient code is constrained to a region of the current block (800), the width of the region of the current block (800) being the greater of a fixed width and the width of the current block (800), and the height of the region of the current block (800) being the greater of a fixed height and the height of the current block (800).
[0253]
[0268] Z. The method of paragraph Y, wherein the residual coefficient code is predicted up to a maximum tolerance limit, the maximum tolerance limit being greater than 8, and the maximum tolerance limit being signaled in a syntax structure of the current block (800).
[0254]
[0269] AA. A method comprising: sorting residual coefficient levels of a current block (800) in order of magnitude by one or more processors (2702) of a computing system (2700); and predicting, by the one or more processors (2702), a plurality of residual coefficient signs of the current block (800) corresponding to a maximum residual coefficient level of the current block (800) in the sorted order.
[0255]
[0270] AB. A method comprising: sorting, by one or more processors (2702) of a computing system (2700), quantization steps of residual coefficients of a current block (800) in order of magnitude; and predicting, by the one or more processors (2702), a plurality of residual coefficient signs of the current block (800) corresponding to a maximum quantization step of the residual coefficients of the current block (800) in the sorted order.
[0256]
[0271] AC. The method of paragraph AB or AC, further including: signaling, by one or more processors (2702), a flag for each predicted residual coefficient code in the bitstream prior to signaling the non-predicted residual coefficient code; and signaling, by one or more processors (2702), the non-predicted residual coefficient code following the signaled flag.
[0257]
[0272] AD. A method comprising predicting, by one or more processors (2702) of a computing system (2700), residual coefficient signs of a current block (800) in an order other than a raster scan order.
[0258]
[0273] AE. The method of paragraph AD, wherein the residual coefficient signs of the current block (800) are predicted in a zigzag scan order.
[0259]
[0274] AF. A method comprising predicting, by one or more processors (2702) of a computing system (2700), residual coefficient signs of a plurality of top-most rows (814) of a current block (800) and a plurality of left-most columns (816) of the current block (800).
[0260]
[0275] AG. The method of claim AF, wherein predicting, by one or more processors (2702), residual coefficient signs includes calculating, by the one or more processors (2702), a cost function as a sum of two sub-cost functions, each sub-cost function measuring across different rows of the neighboring blocks and the current block (800) or across different columns of the neighboring blocks and the current block (800).
[0261]
[0276] AH. A method comprising: predicting, by one or more processors (2702) of a computing system (2700), residual coefficient codes of a current block (800); and signaling, by the one or more processors (2702), one of a plurality of context models in a syntax structure of the current block according to the number of predicted residual coefficient codes.
[0262]
[0277] A method comprising: predicting, by one or more processors (2702) of an AI computing system (2700), one of a plurality of residual coefficient signs of a current block (800) for a boundary pixel of the current block (800) based on a boundary pixel set; evaluating, by the one or more processors (2702), a cost function according to each of a plurality of hypotheses; selecting, by the one or more processors (2702), a selected hypothesis from among the plurality of hypotheses that results in a minimum cost function output; determining, by the one or more processors (2702), that an initial sign was incorrectly predicted compared to the selected hypothesis; calculating, by the one or more processors (2702), a difference between a predicted residual and a reconstructed residual for the boundary pixel; and excluding, by the one or more processors (2702), one or more boundary pixels having a top magnitude of difference between the predicted residual and the reconstructed residual from the boundary pixel set for subsequent residual sign prediction.
Claims
1. A method for encoding a video sequence, comprising: receiving a video sequence; and The video sequence predicting a plurality of transform coefficient signs of the current block corresponding to maximum respective quantization indices of the transform coefficients of the current block in sorted order; encoding, by an entropy encoder, residuals of the predicted plurality of transform coefficient codes with a syntax structure of the current block; By encoding A method comprising:
2. The encoding step comprises: bypassing encoding of transform coefficient codes of the current block other than the predicted plurality of transform coefficient codes; outputting a coded block including the coded residuals of the predicted plurality of transform coefficient codes and equiprobable binary sequence signaling of the transform coefficient codes whose coding has been bypassed; The method of claim 1 further comprising:
3. The encoding step comprises: encoding a residual of each predicted transform coefficient code in a bitstream before encoding the bypassed transform coefficient code; encoding, in the bitstream, the bypassed transform coefficients following the signaled residual; The method of claim 2 further comprising:
4. The encoding step comprises: recording the quantization indices of the transform coefficients of the current block in raster scan order as elements of an array; sorting the elements of the array by size to generate a sorted array; The method of claim 1 further comprising:
5. 5. The method of claim 4, wherein predicting the plurality of transform coefficient codes of the current block corresponding to the largest quantization index comprises predicting each transform coefficient corresponding to the quantization index recorded in a plurality of elements at the end of the sorted array.
6. the predicted transform coefficient signs are confined to a region of the current block; the width of the region of the current block is the smaller of a fixed width and the width of the current block; the height of the region of the current block is the smaller of a fixed height and the height of the current block; The method of claim 1.
7. The method of claim 6 , wherein the encoding further comprises encoding the fixed width and the fixed height in the syntax structure of the current block.
8. The signs of the transform coefficients are predicted to the maximum extent possible, the maximum allowable limit is greater than 8; the encoding step further comprises encoding the maximum allowable limit in a syntax structure of the current block. The method of claim 1.
9. A method for decoding a bitstream, comprising: decoding the bitstream to output a video sequence; wherein said decoding comprises: decoding, by an entropy decoder, a plurality of entropy coded transform coefficients of the current block to generate decoded transform coefficients of the current block; reconstructing the current block based on residuals of predicted transform coefficient codes of the reconstructed residual of the current block corresponding to maximum respective quantization indices of the transform coefficients of the reconstructed residual of the current block in sorted order; A method comprising:
10. The decoding step comprises: parsing, by the entropy decoder, a plurality of residuals of predicted transform coefficient codes from a syntax structure of the current block; inverse quantizing and inverse transforming the decoded transform coefficients and the plurality of residuals of transform coefficient codes to generate reconstructed residuals of the current block; 10. The method of claim 9, further comprising:
11. The decryption comprises: recording quantization indices of the transform coefficients of the reconstructed residual of the current block in raster scan order as elements of an array; sorting the elements of the array by size to generate a sorted array; The method of claim 10 further comprising:
12. 12. The method of claim 11, wherein generating the reconstructed residual of the current block includes selecting a plurality of transform coefficient codes of the reconstructed residual of the current block corresponding to maximum quantization indexes, and predicting each transform coefficient corresponding to a quantization index recorded in a plurality of elements at the end of the sorted array.
13. The decryption comprises: The method of claim 11 , further comprising adding the reconstructed residual of the current block and a prediction signal including the plurality of transform coefficient signs to generate a reconstructed block.
14. the predicted transform coefficient signs are confined to a region of the current block; the width of the region of the current block is the smaller of a fixed width and the width of the current block; the height of the region of the current block is the smaller of a fixed height and the height of the current block; 10. The method of claim 9.
15. The method of claim 14 , wherein the decoding further comprises parsing the fixed width and the fixed height from a syntax structure of the current block.
16. A method for signaling a bitstream, comprising: receiving a video sequence; The video sequence predicting a plurality of transform coefficient signs of the current block corresponding to maximum respective quantization indices of the transform coefficients of the current block in sorted order; encoding, by an entropy encoder, residuals of the predicted plurality of transform coefficient codes with a syntax structure of the current block; and encoding the signaling the generated bitstream based on said encoding. A method comprising:
17. The encoding step: bypassing encoding of transform coefficient codes of the current block other than the predicted plurality of transform coefficient codes; outputting a coded block including the coded residuals of the predicted plurality of transform coefficient codes and equiprobable binary sequence signaling of the transform coefficient codes whose coding has been bypassed; 17. The method of claim 16, further comprising:
18. The encoding step: recording the quantization indices of the transform coefficients of the current block in raster scan order as elements of an array; sorting the elements of the array by size to generate a sorted array; 17. The method of claim 16, further comprising:
19. The predicted plurality of transform coefficient codes are limited to a region of the current block; the width of the region of the current block is the smaller of a fixed width and the width of the current block; the height of the region of the current block is the smaller of a fixed height and the height of the current block; 17. The method of claim 16.
20. The transform coefficient signs are predicted to a maximum allowable limit; the maximum allowable limit is greater than 8; the encoding step further comprises encoding the maximum allowable limit in a syntax structure of the current block.
17. The method of claim 16.