Transport coding techniques for conversion omission blocks
By introducing conversion omit mode and block incremental pulse decoding modulation in the video decoder, the problem of low decoding efficiency of existing video decoding systems for traditional camera capture video content is solved, and a more efficient video decoding effect is achieved.
Patent Information
- Application Number
- CN202510490218.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2020-06-24
- Publication Date
- 2025-08-15
AI Technical Summary
In efficient video decoding systems, there is a problem of low decoding efficiency for traditional cameras to capture video content, especially when using certain decoding tools.
A video decoder method is provided to control the use of decoding tools by communicating higher-order syntax elements, including conversion omission mode (TSM) and block incremental pulse coding modulation (BDPCM) to skip the conversion operation and directly quantize the residual signal, decoding the residual block using an alternative residual coding tool.
Improve video decoding efficiency, especially for screen content decoding applications, reduce conversion operations during decoding, and improve decoding efficiency and compression performance.
Smart Images

Figure CN120499376A_ABST
Abstract
Description
[0001] [Cross-reference to related applications]
[0002] This application is part of a non-provisional application claiming priority to U.S. Provisional Patent Application No. 62 / 868,830, filed June 28, 2019. The contents of the aforementioned patent application are incorporated herein by reference.
Technical field
[0003] The present disclosure relates generally to video processing and, more particularly, to a method for signaling coding a video data block. [Background Technology]
[0004] Unless otherwise indicated herein, the approaches described in this section are not prior art to the claims and are not admitted to be prior art by inclusion in this section.
[0005] In a video decoding system implementing High-Efficiency Video Coding (HEVC), an input video signal is predicted based on a reconstructed signal derived from a coded picture region. The prediction residual signal is processed using a linear transform. The transform coefficients are quantized and entropy coded along with other auxiliary information in the bitstream. After inverse transforming the dequantized transform coefficients, a reconstructed signal is generated from the prediction signal and the reconstructed residual signal. The reconstructed signal is further processed using loop filtering to remove coding artifacts. The decoded image is stored in a frame buffer and used for output and prediction of future images in the input video signal.
[0006] In HEVC, the decoded picture is divided into non-overlapping square block regions represented by associated coding tree units (CTUs). The decoded picture can be represented by a set of slices, each slice containing an integer number of CTUs. The individual CTUs in a slice are processed in raster scan order. Bi-predicted (B) slices can be decoded using intra prediction or inter prediction, where intra prediction or inter prediction uses up to two motion vectors and reference indices to predict the sample values of each block. Predicted (P) slices are decoded using intra prediction or inter prediction using up to one motion vector and reference index to predict the sample values of each block. Intra (I) slices are decoded only using intra prediction.
[0007] A CTU can be divided into multiple non-overlapping coding units (CUs) using a recursive quadtree (QT) structure to accommodate various local motion and texture characteristics. One or more prediction units (PUs) are specified for each CU. A prediction unit, together with the associated CU syntax, serves as the basic unit for communicating predictor information. The values of the relevant pixel samples within the PU are predicted using a specified prediction process. The CU can be further divided using a residual quadtree (RQT) structure to represent the associated prediction residual signal. The leaf nodes of the RQT correspond to transform units (TUs). A transform unit includes a transform block (TB) of 8x8, 16x16, or 32x32 luma samples or a transform block of four 4x4 luma samples, and two corresponding transform blocks of chroma samples for images in 4:2:0 color format. Integer transforms are applied to the transform blocks, and the level values of the quantized coefficients are entropy coded in the bitstream along with other auxiliary information.
[0008] The terms coding tree block (CTB), coding block (CB), prediction block (PB), and transform block (TB) are defined to specify a 2-D sample array of one color component associated with a CTU, CU, PU, and TU, respectively. Thus, a CTU consists of one luma CTB, two chroma CTBs, and related syntax elements. Similar relationships are valid for CUs, PUs, and TUs. Tree partitioning is usually applied to both luma and chroma, but there are exceptions when certain minimum sizes for chroma are reached. In some other coding standards, each CTU can be divided into one or more coding units (CUs) of smaller size using a quadtree with nested multiple types of trees, using binary and ternary splits. The resulting CU partitions can be square or rectangular. [Summary of the invention]
[0009] The following summary is illustrative only and is not intended to be limiting in any way. That is, the following summary is provided to introduce the concepts, key points, benefits, and advantages of the novel and non-obvious technologies described herein. Selected, but not all, implementations are further described in the detailed description that follows. Accordingly, the following summary is not intended to identify essential features of the claimed subject matter, nor is it intended to be used to determine the scope of the claimed subject matter.
[0010] Some embodiments provide a method for performing transform skip mode (TSM) in a video decoder. The video decoder receives data from a bitstream to decode it into a plurality of video pictures. The video decoder parses the bitstream for a first syntax element in a sequence parameter set (SPS) of a current sequence of video pictures. When the first syntax element indicates that transform skip mode is enabled for the current sequence of video pictures, and when a current block in a current picture of the current sequence uses transform skip mode, the video decoder reconstructs the current block by using an untransformed quantized residual signal.
[0011] When the first syntax element indicates that the current sequence of video pictures allows the switch skip mode, the video decoder parses the bitstream of the second syntax element in the SPS to indicate whether block delta pulse code modulation (BDPCM) is allowed for the current sequence of video pictures. In some embodiments, when the first syntax element indicates that the current sequence of video pictures allows the switch skip mode, the video decoder further parses the bitstream of the third syntax element to indicate whether the residual signal of the current block is entropy decoded using other residual decoding processes.
Brief Description of the Drawings
[0012] The accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated into and constitute a part of this disclosure. The accompanying drawings illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. It will be understood that the drawings are not necessarily drawn to scale, as some components may be shown out of proportion to their actual dimensions in order to clearly illustrate the concepts of the present disclosure.
[0013] Figure 1 The signaling of TSM-related signals in a high-level syntax set is conceptually illustrated.
[0014] Figure 2 An example video encoder capable of implementing a transform skip mode is described.
[0015] Figure 3 A section describing implementation of a transition skip mode in a video encoder.
[0016] Figure 4 The process of using transform skip mode during video encoding is conceptually illustrated.
[0017] Figure 5 An example video decoder capable of implementing a transition skip mode is described.
[0018] Figure 6 Describes the portion of a video decoder that implements transition skipping mode.
[0019] Figure 7 The process of using transition skip mode during video decoding is conceptually illustrated.
[0020] Figure 8 An electronic system for implementing some embodiments of the present disclosure is conceptually illustrated. [Specific implementation method]
[0021] In the detailed description that follows, many specific details are set forth by way of example to provide a thorough understanding of the relevant teachings. Any variations, derivations, and / or extensions based on the teachings described herein are within the scope of protection of this disclosure. In some cases, well-known methods, processes, components, and / or circuits related to one or more example implementations disclosed herein may be described at a relatively high level without detail to avoid unnecessarily obscuring various aspects of the teachings of this disclosure.
[0022] I. Entropy Coding of Pixel Blocks
[0023] Some embodiments of the present disclosure provide methods for controlling the use of decoding tools in a video decoding system. Some video decoding systems (e.g., Versatile Video Coding (VVC)) are developed to support various video applications. Certain decoding tools used for new decoding applications such as screen content decoding may not be suitable for decoding video content captured by traditional cameras. According to some aspects of the present invention, a video decoder can communicate one or more high-level syntax elements to control the use of certain decoding tools for a target application.
[0024] In some embodiments, a coded block flag (CBF) is used to signal whether a transform block contains any non-zero transform coefficients. When the CBF is equal to 0, no further decoding is performed on the associated transform block, and all coefficients in the current transform block are inferred to be equal to 0. Otherwise, the associated transform block contains at least one non-zero transform coefficient. Non-zero transform blocks are further divided into non-overlapping sub-blocks. A syntax element coded_sub_block_flag may be signaled to indicate whether the current sub-block contains any non-zero coefficients. When coded_sub_block_flag is equal to 0, no further coding is performed on the associated transform sub-block, and all coefficients in the current transform sub-block are inferred to be equal to 0. Otherwise, the associated transform block contains at least one non-zero transform coefficient. Multiple sub-block coding passes are used to entropy decode the transform coefficient level values in the associated sub-block. In each decoding pass, a single transform coefficient is accessed once according to a predefined scanning order.
[0025] In some embodiments, a syntax element sig_coeff_flag is signaled in the first sub-block decoding pass to indicate whether the absolute value of the current transform coefficient criterion is greater than 0. In the second decoding pass, a syntax element coeff_abs_level_greater1_flag is further signaled for the current coefficient with sig_coeff_flag equal to 1 to indicate whether the absolute value of the associated transform coefficient criterion is greater than 1. In the third decoding pass, a syntax element coeff_abs_level_greater2_flag is further signaled for the current coefficient with coeff_abs_level_greater1_flag equal to 1 to indicate whether the maximum value of the associated transform coefficient criterion is greater than 2. In the fourth and fifth sub-block decoding passes, sign information and remaining level value are further signaled via the syntax elements coeff_sign_flag and coeff_abs_level_remaining, respectively.
[0026] In some embodiments, the transform coefficients can be quantized by slave scalar quantization. The selection of one of the two quantizers is determined by a state machine having four states. The state of the current transform coefficient is determined by the state and the parity of the absolute level value of the previous transform coefficient in the scan order. The transform block is divided into non-overlapping sub-blocks. The transform coefficient calibrant in each sub-block is entropy decoded using multiple sub-block decoding channels. The syntax elements sig_coeff_flag, abs_level_gt1_flag, par_level_flag and abs_level_gt3_flag are signaled in the first sub-block decoding channel. The elements abs_level_gt1_flag and abs_level_gt3_flag indicate whether the absolute value of the current coefficient calibrant is greater than 1 and greater than 3, respectively. The syntax element par_level_flag indicates the parity bits of the absolute value of the current level. The absolute value of the partially reconstructed transform coefficient calibrant from the first channel is given by
[0027] AbsLevelPass1=sig_coeff_flag+par_level_flag+
[0028] abs_level_gt1_flag+2*abs_level_gt3_flag
[0029] The context selection for entropy coding sig_coeff_flag depends on the state of the current coefficient. Therefore, par_level_flag is signaled in the first decoding pass to derive the state of the next coefficient. The syntax elements abs_remainder and coeff_sign_flag are further signaled in subsequent sub-block decoding passes to indicate the remaining coefficient criterion value and sign, respectively. The absolute value of the fully reconstructed transformed coefficient criterion is given by
[0030] AbsLevel=AbsLevelPass1+2*abs_remainder
[0031] The conversion coefficient is given by
[0032] TransCoeffLevel=(2*AbsLevel-(QState>1?1:0))*(1-2*
[0033] coeff_sign_flag),
[0034] Among them, QState indicates the state of the current conversion coefficient.
[0035] In order to achieve high compression efficiency, the context-based adaptive binary arithmetic coding (CABAC) mode or the normal mode is used to entropy decode the values of the syntax elements in the HEVC and VVC drafts. Since the arithmetic decoder in the CABAC engine can only encode binary symbol values, the CABAC operation first needs to convert the value of the syntax element into a binary string, a process usually called binarization. In the decoding process, a probability model is gradually established based on the decoded symbols of different contexts. The selection of the modeling context for decoding the next binary symbol can be determined by the decoding information. Symbols can be decoded without a context modeling stage and an equal probability distribution is assumed (usually called bypass mode) to improve the bitstream parsing throughput.
[0036] In some embodiments, the values of the syntax elements coded_sub_block_flag, sig_coeff_flag, coeff_abs_level_greater1_flag, and coeff_abs_level_greater2_flag in the transform subblock are decoded in normal mode. The values of the syntax elements coeff_sign_flag and coeff_abs_level_remaining in the transform subblock are decoded in bypass mode. To limit the total number of normal bit bins for the entropy coded transform coefficient level in a subblock in the worst case, a maximum of eight coeff_abs_level_greater1_flag values and one coeff_abs_level_greater2_flag value can be decoded per subblock. Thus, the maximum number of normal bit bins per subblock can be limited to 25.
[0037] In some embodiments, the syntax elements sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag are signaled in the first sub-block pass. The syntax elements abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag are entropy coded in sub-block coding passes 2, 3, 4, and 5, respectively. The context modeling for sig_coeff_flag is conditioned on the sig_coeff_flag values of its two neighbors. The context modeling for abs_level_gt1_flag and par_level_flag each uses a single context. In some embodiments, the syntax element abs_level_gtx_flag[n][j], j=0..4,] specifies whether the absolute value of the transform coefficient criterion (at scan position n) is greater than (j<<1)+1 and corresponds to the syntax elements abs_level_gt1_flag, abs_level_gt3_flag, abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag, respectively. The syntax elements par_level_flag and abs_level_gtx_flag[n][j], j=0..4,] are both coded using a single context variable.
[0038] II. Transform-Skipped Blocks
[0039] In some embodiments, methods are provided for signaling transform skip (TS) mode, block-based delta pulse coded modulation (BDPCM) mode, and other tools related to transform skip. When a block is decoded using the transform skip mode, the quantized residual signal is entropy decoded without undergoing a transform operation. When a block is decoded using the BDPCM mode, the residual is quantized, and the difference between each quantized residual and its predictor (e.g., the residual of a previously decoded horizontal or vertical (depending on the BDPCM prediction direction) neighbor) is decoded.
[0040] In some embodiments, a video coder may signal multiple syntax elements in a high-level syntax (HLS) set, such as a sequence parameter set (SPS), a picture parameter set (PPS), and / or a slice header, to control the use of transform skip mode and related coding tools. In some embodiments, a high-level syntax (HLS) set represents a set of syntax for a level above the block level, such as a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, or any other set for a level above the block level. The video coder may signal one or more high-level syntax elements to indicate whether transform skip mode (TSM) is enabled in the current bitstream. When transform skip mode is enabled, the video coder may further signal one or more high-level syntax elements to indicate whether BDPCM is enabled in the current bitstream.
[0041] In some embodiments, when the transform skip mode is enabled, the video decoder may further signal one or more high-order syntax elements to indicate whether to adopt an alternative residual coding tool or process to decode the residual block in the TSM. Specifically, when a CU is decoded in transform skip mode (in other words, the transform skip mode is used for the CU), its prediction residual can be quantized and (entropy) decoded using a transform skip residual decoding process (also called an alternative residual decoding process). In some embodiments, the alternative residual decoding process is modified from the conventional transform coefficient decoding process. Specifically, the residual of the TU is decoded in units of non-overlapping sub-blocks of size 4x4, and a forward scanning order is applied to scan the sub-blocks within the transform block and the positions within the sub-blocks; the last (x, y) position (last (x, y) position) is not signaled; when all the previous flags are equal to 0, each sub-block is decoded except the last sub-block.
[0042] coded_sub_block_flag; sig_coeff_flag context modeling uses a simplified template, while the context model of sig_coeff_flag depends on the top and left neighboring values; the context model of the abs_level_gt1 flag also depends on the left and top sig_coeff_flag values; par_level_flag uses only one context model; additional flags greater than 3, 5, 7, and 9 are signaled to indicate the coefficient level, one context for each flag; the remainder value is binarized using modified parameter derivation; and the context model of the sign flag is determined based on the left and top neighboring values, and the sign flag is parsed after sig_coeff_flag to keep all context-coded bits together.
[0043] In some embodiments, the video coder may signal the SPS syntax element
[0044] sps_transform_skip_enabled_flag and PPS syntax elements pps_bdpcm_enabled_flag and pps_alternative_residual_coding_flag are used to signal whether TSM, BDPCM and alternative residual coding tools are enabled. The following provides the relevant syntax tables for SPS, PPS and transform units:
[0045]
[0046]
[0047]
[0048]
[0049]
[0050] In some embodiments, the video coder may signal the SPS syntax element
[0051] sps_transform_skip_enabled_flag, sps_bdpcm_enabled_flag and
[0052] sps_alternative_residual_coding_flag is used to signal whether TSM, BDPCM and alternative residual coding tools are enabled. The syntax tables for SPS, PPS and transform units are as follows:
[0053]
[0054]
[0055]
[0056]
[0057]
[0058] Figure 1 The signaling of TSM related signals in a high-level syntax set such as an SPS is conceptually shown. The figure shows three SPSs 110, 120, and 130. SPS 110 is applicable to a video picture sequence 115. SPS 120 is applicable to a video picture sequence 125. SPS 130 is applicable to a video picture sequence 135.
[0059] SPS 110 includes a TSM-enabling syntax element set to "false." Consequently, video sequence 115 does not allow TSM, and all blocks in sequence 115 are coded without TSM. In some embodiments, this means that each block of video sequence 125 is coded by converting spatial domain signals (e.g., prediction residuals) into transform domain signals (e.g., transform coefficients), which are then quantized and entropy coded. Furthermore, because the TSM-enabling syntax element in the SPS is set to false, there are no other TSM-related syntax elements, such as BDPCM or other residual coding.
[0060] SPS 120 includes a TSM enable syntax element set to true. Thus, TSM is enabled for video sequence 125, and some blocks in some pictures in sequence 125 are coded using TSM. For TSM-coded blocks, the spatial domain residual signal is directly quantized and entropy coded without conversion. Since the TSM enable syntax element is set to true, the SPS may include other TSM-related syntax elements, such as a BDPCM enable syntax element. In this case, the BDPCM enable syntax element is set to false, and no blocks in sequence 125 are coded using BDPCM. Although not shown, there may be a syntax element (e.g., alternate_residual_coding_flag) that enables or disables alternative residual coding for certain blocks in sequence 125.
[0061] SPS 130 includes a TSM enable syntax element set to true and a BDPCM enable flag set to true. Thus, some blocks in sequence 135 are coded using both TSM and BDPCM. For those blocks, the time-domain residual signal is coded using BDPCM before being quantized and entropy coded without being converted. Although not shown, there may be a syntax element (e.g., alternate_residual_coding_flag) that enables or disables alternate residual coding for some blocks in sequence 135.
[0062] Any of the aforementioned methods can be implemented in an encoder and / or a decoder. For example, any of the aforementioned methods can be implemented in an entropy decoding module of an encoder and / or an entropy decoding module of a decoder. Alternatively, any of the aforementioned methods can be implemented as a circuit integrated into an entropy decoding module of an encoder and / or an entropy decoding module of a decoder.
[0063] III. Video Decoder Example
[0064] Figure 2An example video encoder 200 capable of implementing a transform skip mode is illustrated. As shown, the video encoder 200 receives an input video signal from a video source 205 and encodes the signal into a bitstream 295. The video encoder 200 has several components or modules for encoding the signal from the video source 205, including at least some components selected from the following: a transform module 210, a quantization module 211, an inverse quantization module 214, an inverse transform module 215, an intra-image estimation module 220, an intra-frame prediction module 225, a motion compensation module 230, a motion estimation module 235, a loop filter 245, a reconstructed image buffer 250, an MV buffer 265, an MV prediction module 275, and an entropy encoder 290. The motion compensation module 230 and the motion estimation module 235 are part of the inter-frame prediction module 240.
[0065] In some embodiments, modules 210-290 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device or electronic apparatus. In some embodiments, modules 210-290 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic apparatus. Although modules 210-290 are illustrated as separate modules, some modules may be combined into a single module.
[0066] A video source 205 provides a raw video signal that represents the pixel data for each video frame without compression. A subtractor 208 calculates the difference between the raw video pixel data from the video source 205 and the predicted pixel data 213 from the motion compensation module 230 or the intra-frame prediction module 225. A transform module 210 converts the difference (or residual pixel data or residual signal 209) into transform coefficients 216 (e.g., by performing a discrete cosine transform or DCT). A quantization module 211 quantizes the transform coefficients 216 into quantized data (or quantized coefficients) 212, which are encoded into a bitstream 295 by an entropy encoder 290.
[0067] The inverse quantization module 214 dequantizes the quantized data (or quantized coefficients) 212 to obtain transform coefficients, and the inverse transform module 215 performs an inverse transform on the transform coefficients to generate a reconstructed residual 219. The reconstructed residual 219 is added to the predicted pixel data 213 together with the predicted pixel data 213 to generate reconstructed pixel data 217. In some embodiments, the reconstructed pixel data 217 is temporarily stored in a line buffer (not shown) for use in intra-image prediction and spatial MV prediction. The reconstructed pixels are filtered by the loop filter 245 and stored in the reconstructed image buffer 250. In some embodiments, the reconstructed image buffer 250 is a memory external to the video encoder 200. In some embodiments, the reconstructed image buffer 250 is a memory internal to the video encoder 200.
[0068] An intra-picture estimation module 220 performs intra-prediction based on the reconstructed pixel data 217 to generate intra-prediction data. The intra-prediction data is provided to an entropy encoder 290 to be encoded into a bitstream 295. The intra-prediction data is also used by an intra-prediction module 225 to generate predicted pixel data 213.
[0069] The motion estimation module 235 performs inter-frame prediction by generating MVs to refer to reference pixel data of a previously decoded frame stored in the reconstructed image buffer 250. These MVs are provided to the motion compensation module 230 to generate predicted pixel data.
[0070] Instead of encoding the complete actual MV in the bitstream, the video encoder 200 generates a predicted MV using MV prediction, and encodes the difference between the MV used for motion compensation and the predicted MV as residual motion data and stores in the bitstream 295 .
[0071] The MV prediction module 275 generates a predicted MV based on a reference MV generated for encoding a previous video frame, that is, a motion compensated MV used to perform motion compensation. The MV prediction module 275 retrieves the reference MV from the previous video frame in the MV buffer 265. The video encoder 200 stores the MV generated for the current video frame in the MV buffer 265 as a reference MV for generating the predicted MV.
[0072] The MV prediction module 275 uses the reference MV to create a predicted MV. The predicted MV can be calculated by spatial MV prediction or temporal MV prediction. The entropy encoder 290 encodes the difference (residual motion data) between the predicted MV and the motion compensated MV (MC MV) of the current frame into the bitstream 295.
[0073] The entropy encoder 290 encodes various parameters and data into a bitstream 295 by using an entropy coding technique such as context-adaptive binary arithmetic coding (CABAC) or Huffman coding. The entropy encoder 290 encodes various header elements, flags, and quantized transform coefficients 212 and residual motion data as syntax elements into the bitstream 295. The bitstream 295 is in turn stored in a storage device or sent to a decoder via a communication medium such as a network.
[0074] The loop filter 245 performs a filtering or smoothing operation on the reconstructed pixel data 217 to reduce decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering operation performed includes sample adaptive offset (SAO). In some embodiments, the filtering operation includes adaptive loop filtering (ALF).
[0075] Figure 3 A portion of the video encoder 200 implementing a transform skip mode is described. Specifically, the encoder 200 determines whether to use a skip transform operation and whether to use transform skip-related operations such as BDPCM and / or alternative residual coding for each pixel block based on whether those tools are enabled for the current picture or the current sequence including the current picture.
[0076] As illustrated, the transform module 210 performs a transform operation on the residual signal 209, and the inverse transform module 215 performs a corresponding inverse transform operation. If TSM is activated for the current block being decoded, the encoder 200 can skip the transform and inverse transform operations. When transform skip mode is used, the residual signal 209 is not processed by the transform module 210 but is directly quantized by the quantization module 211. Furthermore, when TSM is used, the output of the inverse quantization module 214 is used directly as the reconstructed residual, rather than being processed by the inverse transform module 215.
[0077] When BDPCM is enabled for the current block, the BDPCM module 311 performs BDPCM processing on the output of the quantization module 211 prior to entropy coding, and the inverse BDPCM module 314 performs corresponding BDPCM processing at the input of the inverse quantization module 214. The entropy encoder 290 receives syntax elements 390 from the decoding control module 300 and may perform either regular residual coding (RRC) processing 313 or transform-skipped residual coding (TSRC) processing 312 based on whether alternative residual coding is used.
[0078] The decoding control module 300 can control the skipping of conversion and inverse conversion operations at the conversion module 210 and the inverse conversion module 215. The decoding control module 300 can also enable or disable the corresponding BDPCM operations at the BDPCM module 311 and the inverse BDPCM module 314. The decoding control module can also enable or disable alternative residual decoding by selecting one of TSRC or RRC in the entropy encoder 290.
[0079] Depending on whether TSM, BDPCM and / or alternative residual coding is used for the current sequence of video pictures, the current picture or the current block, the decoding control module 300 can encode corresponding syntax elements, such as sps_transform_skip_enable_flag, sps_bdpcm_enable_flag and / or alternative_residual_coding_flag (for PPS or SPS or slice header) into the bitstream 295.
[0080] Figure 4A process 400 for using a transform skip mode during video encoding is conceptually illustrated. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing encoder 200 perform process 400 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing encoder 200 performs process 400.
[0081] The encoder receives (at block 410) data to be encoded as one or more video pictures in a bitstream. The encoder signals (at block 420) a TSM syntax element (e.g., sps_transform_skip_enable_flag) in the SPS of the current sequence of video pictures in the bitstream. The encoder determines (at block 425) whether TSM is allowed for the current sequence of video pictures. If TSM is allowed for the current sequence, the process proceeds to block 440. If TSM is not allowed for the current sequence, the process proceeds to block 430.
[0082] At block 430, the encoder encodes the pictures of the current sequence without using TSM. In some embodiments, when the TSM syntax element indicates that transform skip mode is not enabled for the current sequence of video pictures, all blocks in the current sequence of video pictures are encoded using quantized transform coefficients.
[0083] At block 440, the encoder signals in the bitstream a BDPCM syntax element in the SPS (e.g., sps_bdpcm_enable_flag) to indicate whether BDPCM is enabled for pictures in the current sequence. The encoder also signals (at block 450) an alternative residual coding syntax element in the bitstream (e.g., alternative_residual_coding_flag or a slice header for a PPS or SPS). The process then proceeds to block 460.
[0084] If the video picture sequence allows the use of TSM, and if the current block has TSM enabled, the encoder encodes (at block 460) the current block in the current sequence of video pictures using TSM. For example, if a flag in the bitstream indicates that TSM is active for the current block, the encoder encodes the current block using a quantized residual signal that is not converted and remains in the spatial domain.
[0085] When decoding the current block by using TSM, if BDPCM and / or alternative residual decoding are enabled for the current block, the encoding of the current block may also use BDPCM and / or alternative residual decoding mode. Specifically, when the current sequence of video images allows BDPCM and the current block enables BDPCM (for example, a flag in the bitstream indicates that BDPCM is valid for the current block), BDPCM is used to encode the current block (by using the difference between the residual signal and the previously decoded residual signal of the adjacent position to decode the residual signal at a certain position in the current block). When alternative residual decoding is enabled for the current block (for example, there is no flag in the bitstream that disables alternative residual decoding for the current slice), the residual signal of the current block is entropy encoded using alternative residual decoding (for example, TSRC), otherwise conventional residual coding (RRC) is used.
[0086] IV. Video Decoder Example
[0087] Figure 5 An example video decoder 500 capable of implementing a transform skipping mode is illustrated. As shown, video decoder 500 is an image decoding or video decoding circuit that receives a bitstream 595 and decodes the contents of the bitstream into pixel data for a video frame for display. Video decoder 500 has several components or modules for decoding bitstream 595, including some components selected from an inverse quantization module 505, an inverse transform module 510, an intra-frame prediction module 525, a motion compensation module 530, a loop filter 545, a decoded picture buffer 550, an MV buffer 565, an MV prediction module 575, and a parser 590. Motion compensation module 530 is part of inter-frame prediction module 540.
[0088] In some embodiments, modules 510-590 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device. In some embodiments, modules 510-590 are modules of hardware circuits implemented by one or more ICs of an electronic device. Although modules 510-590 are shown as separate modules, some modules may be combined into a single module.
[0089] The parser 590 (or entropy decoder) receives the bitstream 595 and performs initial parsing according to the syntax defined by the video coding or image coding standard. The parsed syntax elements include various header elements, flags, and quantized data (or quantized coefficients) 512. The parser 590 parses the various syntax elements using entropy coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman coding.
[0090] The inverse quantization module 505 dequantizes the quantized data (or quantized coefficients) 512 to obtain transform coefficients, and the inverse transform module 510 performs inverse transform on the transform coefficients 516 to generate a reconstructed residual signal 519. The reconstructed residual signal 519 is added to the predicted pixel data 513 from the intra prediction module 525 or the motion compensation module 530 to generate decoded pixel data 517. The decoded pixel data is filtered by the loop filter 545 and stored in the decoded picture buffer 550. In some embodiments, the decoded picture buffer 550 is a memory external to the video decoder 500. In some embodiments, the decoded picture buffer 550 is a memory internal to the video decoder 500.
[0091] The intra prediction module 525 receives intra prediction data from the bitstream 595 and generates predicted pixel data 513 from the decoded pixel data 517 stored in the decoded picture buffer 550. In some embodiments, the decoded pixel data 517 is also stored in a line buffer (not shown) used for intra prediction and spatial MV prediction.
[0092] In some embodiments, the contents of the decoded image buffer 550 are used for display. The display device 555 either retrieves the contents of the decoded image buffer 550 for direct display or retrieves the contents of the decoded image buffer into a display buffer. In some embodiments, the display device receives pixel values from the decoded image buffer 550 via pixel transfer.
[0093] The motion compensation module 530 generates predicted pixel data 513 from the decoded pixel data 517 stored in the decoded picture buffer 550 according to the motion compensated MV (MC MV). These motion compensated MVs are decoded by adding the residual motion data received from the bitstream 595 to the predicted MV received from the MV prediction module 575.
[0094] The MV prediction module 575 generates a predicted MV based on a reference MV generated for decoding a previous video frame (e.g., a motion compensated MV used to perform motion compensation). The MV prediction module 575 retrieves the reference MV of the previous video frame from the MV buffer 565. The video decoder 500 stores the motion compensated MV generated for decoding the current video frame in the MV buffer 565 as a reference MV for generating the predicted MV.
[0095] The loop filter 545 performs a filtering or smoothing operation on the decoded pixel data 517 to reduce decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering operation performed includes sample adaptive offset (SAO). In some embodiments, the filtering operation includes adaptive loop filtering (ALF).
[0096] Figure 6A portion illustrating implementation of a transform skip mode for video decoder 500. Specifically, decoder 500 determines whether to skip (inverse) transform operations and whether to use transform skip-related operations such as BDPCM and / or alternative residual coding for each pixel block based on whether those tools are enabled for the current picture or the current sequence including the current picture.
[0097] As shown, the inverse quantizer 514 performs an inverse quantization operation on the quantized coefficients 512 parsed by the entropy decoder 590. The output of the inverse quantizer 514 is provided to the inverse transform module 516 to be inversely transformed into a residual signal. When TSM is used, the output of the inverse quantization module 514 is directly used as the reconstructed residual, rather than being processed by the inverse transform module 515.
[0098] When BDPCM is enabled for the current block, the inverse BDPCM module 614 performs BDPCM processing at the input of the inverse quantization module 514. The entropy decoder 590 may perform a regular residual coding (RRC) process 611 or a transform skipped residual coding (TSRC) process 612 based on whether alternative residual coding is used.
[0099] The decoding control module 600 can control the skipping of the inverse transform operation at the inverse transform module 515. The decoding control module 600 can also enable or disable BDPCM operations at the inverse BDPCM module 614. The decoding control module can also enable or disable alternative residual decoding by selecting one of TSRC or RRC in the entropy decoder 590. The decoding control module 600 can generate controls corresponding to these TSM-related operations based on syntax elements 690, such as sps_transform_skip_enable_flag, sps_bdpcm_enable_flag, and / or alternate_residual_coding_flag (for PPS or SPS or slice header) parsed by the entropy decoder 590 from the bitstream 595.
[0100] Figure 7 A process 700 for using a transform skip mode during video encoding is conceptually illustrated. In some embodiments, one or more processing units (e.g., processors) of a computing device implementing decoder 500 perform process 700 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing decoder 500 performs process 700.
[0101] The decoder receives (at block 710) data from a bitstream to be decoded into one or more video pictures. The decoder parses (at block 720) the bitstream to obtain a TSM syntax element (e.g., sps_transform_skip_enable_flag) in the SPS of the current video picture sequence. The decoder determines (at block 725) whether TSM is enabled for the current video picture sequence. If TSM is enabled for the current sequence, the process proceeds to block 740. If TSM is not enabled for the current sequence, the process proceeds to block 730.
[0102] At block 730, the decoder reconstructs the pictures of the current sequence without using TSM. In some embodiments, when the TSM syntax element indicates that transform skip mode is not allowed for the current sequence of video pictures, all blocks in the current sequence of video pictures are coded by using quantized transform coefficients.
[0103] At block 740, the decoder parses the bitstream and obtains the BDPCM syntax element (e.g., sps_bdpcm_enable_flag) in the SPS to indicate whether BDPCM is enabled for pictures in the current sequence. The decoder also parses (at block 750) the bitstream and obtains the alternate residual coding syntax element (e.g., alternate_residual_coding_flag for PPS or SPS or slice header). The process then proceeds to block 760.
[0104] If TSM is allowed for the sequence of video pictures, and if TSM is enabled for the current block, the decoder reconstructs (at block 760) the current block in the current picture of the current sequence of video pictures using TSM. For example, if a flag in the bitstream indicates that TSM is active for the current block, the decoder will reconstruct the current block using the quantized residual signal that is not transformed and remains in the spatial domain.
[0105] When the current block is decoded by using TSM, if BDPCM and / or alternative residual decoding are enabled for the current block, the decoding of the current block may also use BDPCM and / or alternative residual decoding. Specifically, when the current sequence of video images allows BDPCM and the current block enables BDPCM (for example, a flag in the bitstream indicates that BDPCM is valid for the current block), BDPCM is used to decode the current block (the residual signal at a certain position in the current block is decoded by using the difference between the residual signal and the previously decoded residual signal at the adjacent position.) When alternative residual decoding is enabled for the current block (for example, there is no flag in the bitstream to disable alternative residual decoding for the current slice), the residual signal of the current block is entropy decoded using alternative residual decoding (for example, TSRC), otherwise conventional residual decoding (RRC) is used.
[0106] V. Example Electronic Systems
[0107] Many of the above features and applications are implemented as a software process designated as a set of instructions recorded on a computer-readable storage medium (also referred to as a computer-readable medium). When these instructions are executed by one or more computing or processing units (e.g., one or more processors, the core of a processor, or other processing units), they cause the processing units to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard drives, erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), etc. Computer-readable media do not include carrier waves and electronic signals transmitted wirelessly or by wired connections.
[0108] In this specification, the term "software" is intended to include firmware residing in read-only memory or applications stored in magnetic storage, which can be read into memory for processing by a processor. Likewise, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while retaining different software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of separate programs that together implement the software inventions described herein are within the scope of this disclosure. In some embodiments, the software program, when installed to run on one or more electronic systems, defines one or more specific machine implementations that implement and perform the operations of the software program.
[0109] Figure 8 An electronic system 800 is conceptually illustrated for implementing some embodiments of the present disclosure. The electronic system 800 may be a computer (e.g., a desktop computer, a personal computer, a tablet computer, etc.), a phone, a PDA, or any other type of electronic device. Such an electronic system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 800 includes a bus 805, (one or more) processing units 810, a graphics processing unit (GPU) 815, a system memory 820, a network 825, a read-only memory 830, a permanent storage device 835, an input device 840, and an output device 845.
[0110] Bus 805 collectively represents all system buses, peripheral buses, and chipset buses that communicatively connect the numerous internal devices of electronic system 800. For example, bus 805 communicatively connects processing unit(s) 810, read-only memory 830, system memory 820, and persistent storage device 835 with GPU 815.
[0111] The processing unit 810 retrieves instructions to be executed and data to be processed from these various storage units in order to perform the processes of the present disclosure. In different embodiments, the processing unit can be a single processor or a multi-core processor. Some instructions are passed to and executed by the GPU 815. The GPU 815 can offload various calculations or supplement the image processing provided by the processing unit 810.
[0112] Read-only memory (ROM) 830 stores static data and instructions used by processing unit 810 and other modules of the electronic system. On the other hand, permanent storage device 835 is a read-write storage device. This device is a non-volatile storage unit that stores instructions and data even when electronic system 800 is turned off. Some embodiments of the present disclosure use a mass storage device (such as a magnetic disk or optical disk and its corresponding disk drive) as permanent storage device 835.
[0113] Other embodiments use removable storage devices (e.g., floppy disks, flash memory devices, etc., and their corresponding disk drives) as permanent storage devices. Like permanent storage device 835, system memory 820 is a read-write storage device. However, unlike storage device 835, system memory 820 is a volatile read-write memory, such as random access memory. System memory 820 stores some instructions and data used by the processor at runtime. In some embodiments, processing according to the present disclosure is stored in system memory 820, permanent storage device 835 and / or read-only memory 830. For example, various storage units include instructions for processing multimedia clips according to some embodiments. Processing unit 810 retrieves instructions to be executed and data to be processed from these various storage units in order to perform the processing of some embodiments.
[0114] The bus 805 is also connected to input and output devices 840 and 845. The input device 840 enables a user to convey information to the electronic system and select commands. The input device 840 includes an alphanumeric keyboard and pointing device (also known as a "mouse control device"), a camera (e.g., a webcam), a microphone, or a similar device for receiving voice commands, etc. The output device 845 displays images or other output data generated by the electronic system. The output device 845 includes a printer and a display device, such as a cathode ray tube (CRT) or a liquid crystal display (LCD), as well as a speaker or similar audio output device. Some embodiments include devices that act as both input devices and output devices, such as a touch screen.
[0115] Finally, if Figure 8As shown, bus 805 also couples electronic system 800 to a network 825 via a network adapter (not shown). In this manner, the computer can be part of a computer network (e.g., a local area network ("LAN"), a wide area network ("WAN"), or an intranet, or a network (e.g., the Internet). Any or all components of electronic system 800 may be used in conjunction with the present disclosure.
[0116] Some embodiments include electronic components, such as a microprocessor, storage and memory that stores computer program instructions in a machine-readable or computer-readable medium (alternatively referred to as a computer-readable storage medium, machine-readable medium, or machine-readable storage medium). Some examples of such computer-readable media include RAM, ROM, compact disc-read only disk (CD-ROM), compact disc-recordable disk (CD-R), compact disc-rewritable disk (CD-RW), read-only digital versatile disk (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini-SD card, micro-SD card, etc.), magnetic and / or solid-state hard drives, read-only and recordable Blu-ray discs. An optical disc, an ultra-density optical disc, any other optical or magnetic medium, and a floppy disk. The computer-readable medium may store a computer program executable by at least one processing unit, and the computer program includes a set of instructions for performing various operations. Examples of computer programs or computer code include machine code such as produced by a compiler, and files including high-level code executed by a computer, electronic component, or microprocessor using an interpreter.
[0117] Although the above discussion primarily refers to microprocessors or multi-core processors executing software, many of the above features and applications are performed by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions stored on the circuits themselves. In addition, some embodiments execute software stored in programmable logic devices (PLDs), ROM, or RAM devices.
[0118] As used in this specification and any claims of this application, the terms "computer," "server," "processor," and "memory" refer to electronic or other technological devices. These terms do not include a person or group of people. For purposes of this description, the terms "display" or "displaying" refer to displaying on an electronic device. As used in this specification and any claims of this application, the terms "computer-readable medium," "computer-readable media," and "machine-readable medium" are entirely limited to tangible, physical objects that store information in a form that can be read by a computer. These terms do not include any wireless signals, wired download signals, and any other temporary signals.
[0119] Although the present disclosure has been described with reference to numerous specific details, those skilled in the art will recognize that the present disclosure may be embodied in other specific forms without departing from the spirit of the present disclosure. Figure 4 and Figure 7 ) conceptually illustrates the processes. The specific operations of these processes may not be performed in the exact order shown and described. Specific operations may not be performed in a continuous series of operations, and different specific operations may be performed in different embodiments. In addition, the process may be implemented using several sub-processes or as part of a larger macro-process. Therefore, it will be understood by those skilled in the art that the present disclosure is not limited by the foregoing illustrative details, but is defined by the appended claims.
[0120] Additional Statement
[0121] The subject matter described herein sometimes shows different components contained within or connected to other different components. It should be understood that the architectures depicted in this manner are merely exemplary, and that many other architectures that achieve the same functionality can actually be implemented. In a conceptual sense, any arrangement of components that achieve the same functionality is effectively "associated" so as to achieve the desired functionality. Therefore, any two components that are combined herein to obtain a particular functionality can be considered to be "associated" with each other to achieve the desired functionality, regardless of the architecture or intermediate components. Similarly, any two components that are associated in this manner can also be considered to be "operably connected" or "operably coupled" to each other to achieve the desired functionality, and any two components that can be associated in this manner can also be considered to be "operably couplable" to each other to achieve the desired functionality. Specific examples of "operably couplable" include, but are not limited to, components that are physically connectable and / or physically interact with each other, and / or components that are wirelessly interactive and / or wirelessly interactive, and / or components that are logically interactive and / or logically interactive.
[0122] Furthermore, with respect to the use of substantially any plural and / or singular terms herein, those skilled in the art can translate the plural to the singular, and / or the singular to the plural, as is appropriate to the context and / or application. For clarity, various singular / plural permutations are expressly set forth herein.
[0123] Those skilled in the art will understand that, in general, the terms used herein, and particularly in the appended claims (e.g., the subject matter of the appended claims), are generally intended as "open-ended" terms (e.g., the term "comprising" should be interpreted as "including but not limited to," the term "having" should be interpreted as "having at least," the term "including" should be interpreted as "including but not limited to," etc.). Those skilled in the art will also understand that if a specific number of an introduced claim recitation is intended, such intent will be explicitly stated in the claim, and in the absence of such a statement, no such intent is present. For example, to aid understanding, the appended claims may include use of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed as limiting any claim containing such introduced claim recitation with the indefinite article "a" or "an" to an invention containing only one such recitation, even if the same claim contains the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should generally be construed to mean "at least one" or "one or more"); the same applies to claim recitations introduced with definite articles. In addition, even if a specific number of introduced claim recitations is explicitly stated, those skilled in the art will recognize that such a statement should generally be construed to mean at least the stated number (e.g., the statement "two recitations" alone without other modifiers generally means at least two recitations, or two or more recitations). Furthermore, where phrases such as "at least one of A, B, and C, etc." are used, such constructions are generally intended to have the meaning of the phrase as understood by those of ordinary skill in the art (e.g., "a system having at least one of A, B, and C" would include, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Where phrases such as "at least one of A, B, or C, etc." are used, such constructions are generally intended to have the meaning of the phrase as understood by those of ordinary skill in the art (e.g., "a system having at least one of A, B, or C" would include, but is not limited to, systems having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Those of ordinary skill in the art will further understand that, whether in the specification, claims, or drawings, substantially any disjunctive word and / or phrase representing two or more alternative terms should be understood to contemplate the possibility of including one, any, or both of the terms.For example, the phrase "A or B" should be understood to include the possibilities of "A," "B," or "A and B."
[0124] It will be appreciated from the foregoing that various embodiments of the present disclosure have been described herein for illustrative purposes, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Therefore, the various embodiments disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the appended claims.
Claims
1. A video decoding method, comprising: receiving data to be encoded or decoded into a plurality of video images; signaling or parsing a first syntax element in a sequence parameter set of a current sequence of the plurality of video pictures, the first syntax element indicating whether the current sequence of the plurality of video pictures to which the current block belongs allows use of a transform skip mode; In response to the first syntax element indicating that the current sequence of the plurality of video pictures allows the use of the transform skip mode, signaling or parsing a second syntax element in a slice header of a current slice in the current sequence, the second syntax element indicating whether an alternative residual coding process is enabled for the current slice to which the current block belongs; When the first syntax element indicates that the current sequence of the multiple video pictures allows use of the transform skip mode, and when the current block in the current picture of the current sequence uses the transform skip mode, when the second syntax element indicates that the current slice has the alternative residual decoding process enabled, the current block is encoded or decoded using a quantized residual signal that is not converted according to the alternative residual decoding process.
2. An electronic device comprising: The video decoding circuit is configured to perform the following operations: receiving data to be decoded into a plurality of video images; Parsing a first syntax element in a sequence parameter set of a current sequence of the plurality of video pictures, the first syntax element indicating whether the current sequence of the plurality of video pictures to which the current block belongs allows use of a transform skip mode; In response to the first syntax element indicating that the current sequence of the video pictures allows the use of the transform skip mode, parsing a second syntax element in a slice header of a current slice in the current sequence, the second syntax element indicating whether an alternative residual decoding process is enabled for the current slice to which the current block belongs; When the first syntax element indicates that the current sequence of the multiple video pictures allows use of the transform skip mode, and when the current block in the current picture of the current sequence uses the transform skip mode, when the second syntax element indicates that the current slice has the alternative residual decoding process enabled, the current block is decoded using a quantized residual signal that is not converted according to the alternative residual decoding process.