Binarization in Transform Skip Residual Coding
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-06-24
- Publication Date
- 2026-08-12
Smart Images

Figure 112021144706122-PCT00022_ABST
Abstract
Description
Technology Field
[0001] This application is,
[0002] Claiming priority to U.S. Patent Application No. 16 / 909,892 filed on June 23, 2020, this
[0003] U.S. provisional patent application No. 62 / 865,883 filed on June 24, 2019; and
[0004] Claiming the benefit of U.S. provisional patent application No. 62 / 894,449 filed on August 30, 2019, and
[0005] The entire contents of each of these are thus integrated into a reference.
[0006] 기술 분야
[0007] The present disclosure relates to video encoding and video decoding. Background Technology
[0008] Digital video capabilities can be integrated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital aids (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite wireless phones, so-called "smartphones," video teleconferencing devices, video streaming devices, etc. Digital video devices implement video coding techniques such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions of such standards. By implementing such video coding techniques, video devices may more efficiently transmit, receive, encode, decode, and / or store digital video information.
[0009] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate redundancy inherent in video sequences. For block-based video coding, a video slice (e.g., a video picture or a part of a video picture) may be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction for reference samples in neighboring blocks of the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction for reference samples in neighboring blocks of the same picture, or temporal prediction for reference samples in other reference pictures. Pictures may be referred to as frames, and reference pictures may be referred to as reference frames. means of solving the problem
[0010] The present disclosure describes techniques related to binarization performed in transform skip residual coding. More specifically, the present disclosure describes techniques related to an entropy decoding process that converts a binary representation into quantized coefficients of a series of non-binary values. A corresponding entropy encoding process, which is the inverse process of entropy decoding, is also described herein.
[0011] According to the techniques of the present disclosure, a binarization process for coding transform skip factors is described. For an input quantization parameter (QP), the video decoder may derive a corresponding dynamic range [0, maxTsLevel] for levels of transform skip factors, where “maxTsLevel” represents the maximum possible level of transform skip factors allowed for a specific QP value, i.e., the maximum possible level of quantized residual values. The maximum possible level of transform skip factors for a block may be a function of the QP value for the block, but may also depend on the bit depth for the block. Then, the video decoder may receive an index of an interval containing levels of values for transform skip factors, or some other indication. The video decoder may additionally receive a remainder value representing the difference between the initial value of the interval containing levels of values for transform skip factors and the actual level value of the transform skip factors.
[0012] According to one example of the present disclosure, a method for decoding video data comprises: determining that a block of video data is encoded without transforming residual data for the block; determining a quantization parameter for the block of video data; determining a range for levels of quantized residual values of the block of video data based on the determined quantization parameter; dividing the range into k intervals, wherein k is an integer value; determining a level for a quantized residual value of the block based on the k intervals, wherein the step of determining a level for a quantized residual value of the block based on the k intervals comprises: receiving information indicating that the level for a quantized residual value is within a specific interval among the k intervals; receiving information indicating a difference value indicating a difference between a reference level value for the specific interval and a level for a quantized residual value of the block; and determining a level for a quantized residual value of the block based on the reference level value and the difference value. It includes the step of outputting decoded video data based on levels for quantized residual values.
[0013] According to another example of the present disclosure, a device for decoding video data comprises a memory configured to store video data and one or more processors, wherein the one or more processors determine that a block of video data is encoded without transforming residual data for the block; determine a quantization parameter for the block of video data; determine a range for levels of quantized residual values of the block of video data based on the determined quantization parameter; divide the range into k intervals, wherein k is an integer value; and determine a level for a quantized residual value of the block based on the k intervals, wherein to determine a level for a quantized residual value of the block based on the k intervals, the one or more processors receive information indicating that a level for a quantized residual value is within a specific interval among the k intervals; and receive information indicating a difference value indicating a difference between a reference level value for a specific interval and a level for a quantized residual value of the block. and is further configured to determine the level for the quantized residual value of the block based on the reference level value and the difference value; and is configured to output decoded video data based on the level for the quantized residual value.
[0014] According to another example of the present disclosure, a computer-readable storage medium for storing instructions, wherein the instructions, when executed by one or more processors, cause one or more processors to determine that a block of video data is encoded without transforming residual data for the block; cause to determine a quantization parameter for the block of video data; cause to determine a range of levels of quantized residual values of the block of video data based on the determined quantization parameter; cause to divide the range into k intervals, wherein k is an integer value; cause to determine a level of quantized residual values of the block based on the k intervals, wherein, to determine a level of quantized residual values of the block based on the k intervals, the instructions cause one or more processors to receive information indicating that a level of quantized residual values is within a specific interval among the k intervals; and cause to receive information indicating a difference value indicating a difference between a reference level value for a specific interval and a level of quantized residual values of the block. And, based on the reference level value and the difference value, the level for the quantized residual value of the block is determined; and the decoded video data is output based on the level for the quantized residual value.
[0015] According to another example, an apparatus for decoding video data comprises: means for determining that a block of video data is encoded without transforming residual data for the block; means for determining a quantization parameter for a block of video data; means for determining a range for levels of quantized residual values of a block of video data based on the determined quantization parameter; means for dividing the range into k intervals, wherein k is an integer value; means for determining a level for a quantized residual value of a block based on k intervals, wherein the means for determining a level for a quantized residual value of a block based on k intervals comprises: means for receiving information indicating that a level for a quantized residual value is within a specific interval among the k intervals; means for receiving information indicating a difference value indicating a difference between a reference level value for a specific interval and a level for a quantized residual value of a block; and means for determining a level for a quantized residual value of a block based on the reference level value and the difference value. and includes means for outputting decoded video data based on levels for quantized residual values.
[0016] According to another example of the present disclosure, a method for generating a bitstream of encoded video data comprises: determining that a block of video data is encoded without transforming residual data for the block; determining a level for a quantized residual value of the block; determining a quantization parameter for the block of video data; determining a range for levels of quantized residual values of the block of video data based on the determined quantization parameter; dividing the range into k intervals, wherein k is an integer value; determining a specific interval among the k intervals containing a level for a quantized residual value; determining a difference value representing the difference between a reference level value for the specific interval and a level for a quantized residual value of the block; and signaling a level for a quantized residual value of the block based on the k intervals, wherein the step of signaling a level for a quantized residual value of the block based on the k intervals comprises generating one or more syntax elements indicating a specific interval to be included in the bitstream of the encoded video data. The method includes the step of generating a syntax element indicating a difference value to be included in the bitstream of the encoded video data; and the step of outputting the bitstream of the encoded video data, and the step of signaling a level for the quantized residual value of the block.
[0017] According to another example of the present disclosure, a device for encoding video data comprises a memory configured to store video data and one or more processors, wherein the one or more processors determine that a block of video data is encoded without transforming residual data for the block; determine a level for a quantized residual value of the block; determine a quantization parameter for the block of video data; determine a range for levels of quantized residual values of the block of video data based on the determined quantization parameter; divide the range into k intervals, wherein k is an integer value; determine a specific interval among the k intervals containing a level for a quantized residual value; and determine a difference value representing the difference between a reference level value for the specific interval and a level for a quantized residual value of the block. Signaling a level for a quantized residual value of a block based on k intervals, wherein one or more processors are configured to signal a level for a quantized residual value of a block based on k intervals, wherein one or more processors generate one or more syntax elements indicating a specific interval to be included in a bitstream of encoded video data; generate a syntax element indicating a difference value to be included in a bitstream of encoded video data; and further configured to output a bitstream of encoded video data.
[0018] According to another example of the present disclosure, a computer-readable storage medium for storing instructions, wherein the instructions, when executed by one or more processors, cause one or more processors to determine that a block of video data is encoded without transforming residual data for the block; to determine a level for a quantized residual value of the block; to determine a quantization parameter for a block of video data; based on the determined quantization parameter, to determine a range for levels of quantized residual values of the block of video data; to divide the range into k intervals, wherein k is an integer value; to determine a specific interval among the k intervals containing a level for a quantized residual value; and to determine a difference value representing the difference between a reference level value for a specific interval and a level for a quantized residual value of the block. To signal a level for a quantized residual value of a block based on k intervals, one or more processors generate one or more syntax elements indicating a specific interval to be included in a bitstream of encoded video data; generate a syntax element indicating a difference value to be included in a bitstream of encoded video data; and further configured to output a bitstream of encoded video data, thereby signaling a level for a quantized residual value of said block.
[0019] According to another example of the present disclosure, an apparatus for generating a bitstream of encoded video data comprises: means for determining that a block of video data is encoded without transforming residual data for the block; means for determining a level for a quantized residual value of the block; means for determining a quantization parameter for a block of video data; means for determining a range for levels of quantized residual values of the block of video data based on the determined quantization parameter; means for dividing the range into k intervals, wherein k is an integer value; means for determining a specific interval among the k intervals including a level for a quantized residual value; and means for determining a difference value indicating a difference between a reference level value for a specific interval and a level for a quantized residual value of the block. Means for signaling a level for a quantized residual value of a block based on k intervals, the means for signaling a level for a quantized residual value of a block based on k intervals comprises: means for generating one or more syntax elements indicating a specific interval to be included in a bitstream of encoded video data; means for generating a syntax element indicating a difference value to be included in a bitstream of encoded video data; and means for outputting a bitstream of encoded video data.
[0020] Details of one or more examples are described in the accompanying drawings and the following description. Other features, purposes, and advantages will be apparent from the description, drawings, and claims. Brief explanation of the drawing
[0021] FIG. 1 is a block diagram illustrating an exemplary video encoding and decoding system capable of performing the techniques of the present disclosure. FIGS. 2a and 2b are conceptual diagrams illustrating an exemplary quadtree binary tree (QTBT) structure and a corresponding coding tree unit (CTU). FIG. 3 is a block diagram illustrating an exemplary video encoder capable of performing the techniques of the present disclosure. FIG. 4 is a block diagram illustrating an exemplary video decoder capable of performing the techniques of the present disclosure. Figure 5 is a flowchart illustrating an exemplary video encoding process. Figure 6 is a flowchart illustrating an exemplary video decoding process. Figure 7 is a flowchart illustrating an exemplary video encoding process. Figure 8 is a flowchart illustrating an exemplary video decoding process. Specific details for implementing the invention
[0022] Video coding (e.g., video encoding and / or video decoding) typically involves predicting a block of video data from either an already coded block of video data in the same picture (e.g., intra-prediction) or an already coded block of video data in a different picture (e.g., inter-prediction). In some cases, the video encoder also calculates residual data by comparing the predicted block with the original block. Thus, the residual data represents the difference between the predicted block and the original block. To reduce the number of bits required to signal the residual data, the video encoder may transform and quantize the residual data and signal the transformed and quantized residual data from the encoded bitstream.
[0023] The video decoder decodes residual data and adds it to the prediction block to generate a reconstructed video block that matches the original video block more closely than the prediction block alone. The compression achieved by the transform and quantization processes may be lossy, which means that the transform and quantization processes may introduce distortion into the decoded video data. Due to the loss introduced by the transform and quantization of the residual data, the first reconstructed block may have distortions or artifacts. One general type of artifact or distortion is referred to as blockiness, where the boundaries of the blocks used to code the video data are visible.
[0024] To further improve the quality of the decoded video, the video decoder may perform one or more filtering operations on the restored video blocks. Examples of these filtering operations include deblocking filtering, sample adaptive offset (SAO) filtering, and adaptive loop filtering (ALF). Parameters for these filtering operations are determined by the video encoder and may be explicitly signaled in the encoded video bitstream, or may be implicitly determined by the video decoder without the need for the parameters to be explicitly signaled in the encoded video bitstream.
[0025] In some coding scenarios, a video encoder may encode blocks of video data in a transform skip mode where the transform process described above is not performed, i.e., the transform process is skipped. Accordingly, for blocks encoded in transform skip mode, residual data is not transformed but may still be quantized. Thus, transform skip coefficients generally correspond to quantized representations of residual values, whereas transform coefficients generally correspond to residual values of the transformed block that are quantized to generate transform coefficients. As used herein, the terms coefficients may refer to transform coefficients or transform skip coefficients, and may be quantized or dequantized.
[0026] The present disclosure describes techniques related to binarization performed in transform skip residual coding. More specifically, the present disclosure describes techniques related to an entropy decoding process that converts a binary representation into quantized coefficients of a series of non-binary values. A corresponding entropy encoding process, which is the inverse process of entropy decoding, is also described herein. In the following disclosure, when a video decoder is described as receiving or parsing a syntax element, it may be assumed that a video encoder is configured to generate the same syntax element for signaling, for example, to include in a bitstream of encoded video data. Similarly, when a video encoder is described as signaling a syntax element, it may be assumed that a video decoder is configured to receive and parse the same syntax element.
[0027] According to the techniques of the present disclosure, a binarization process for coding transform skip coefficients is described. For an input quantization parameter (QP), a video decoder can derive a corresponding dynamic range [0, maxTsLevel] of levels of transform skip coefficients, where “maxTsLevel” represents the maximum level of transform skip coefficients, i.e., the maximum level of quantized residual values. In this context, “level” refers to the magnitude or absolute value of the quantized residual values.
[0028] According to one exemplary technique of the present disclosure, once the dynamic range of levels of the transformation skip factor is calculated, the range can be divided into k (inclusive) intervals as follows:
[0029] [X, t0], [t0+1, t1], [t1+1, t2], ... [t k-3 +1, t k-2 ], [t k-2 +1, maxTsLevel].
[0030] In the above example, X represents the minimum value for interval 0. As discussed later in this disclosure, in different implementations, X may be equal to 0, 1, 2, or some other value. The k intervals may have indices in the range of 0 to k-1. Again, referring to the above example, index 0 corresponds to interval [X, t0], index 1 corresponds to interval [t0+1, t1], and interval [t k-2 It continues up to index k-1 corresponding to [+1, maxTsLevel]. In the above example, t n represents the upper threshold for the nth interval, where n is in the range of 0 to k-1.
[0031] Next, the video decoder may receive an index of an interval containing the level of the value for the transform skip factor, or some other indication. The video decoder may additionally receive a remainder value representing the difference between the initial value of the interval containing the level of the value for the transform skip factor and the actual level value for the transform skip factor. As an example, if Y represents the actual level of the value for the transform skip factor and Y is within interval 2, the remainder value is equal to Y-(t1+1).
[0032] In some examples, the video decoder may receive an effective count flag indicating whether the transform skip count is equal to 0 or not equal to 0 before receiving an interval indication. In such examples, the value of X at the first interval may be equal to 1. In some examples, the video decoder may also receive a greater than 1 flag indicating whether the level of the transform skip count is equal to 1 or greater than 1 before receiving an interval indication. In such examples, the value of X at the first interval may be equal to 2. The video decoder may additionally receive a flag indicating whether the actual value for the transform skip count is negative or positive.
[0033] The distribution of values for transformation coefficients in a transformation block tends to be very different from the distribution of values for transformation skip coefficients in a non-transformed block. For example, almost all transformation coefficients in the bottom right half of a transformation block may be equal to 0. While only a few transformation coefficients near the top left corner of the block may have larger values, for example, greater than 2, some transformation coefficients between the top left corner of the transformation and the bottom left half of the block may have smaller values, for example, 1 or 2. Existing techniques for coding coefficients are generally designed to utilize the abundance of 1s and 2s, as well as the large number of 0s found in transformation blocks, which may present potential problems when coding blocks in transformation skip mode.
[0034] Unlike transform blocks, transform skip blocks have relatively few zero values, and to some extent, transform skip values have zero values, and those zero values tend not to cluster in specific regions of the transform skip blocks. Therefore, coefficient coding techniques designed to code transform blocks tend not to be efficient when coding transform skip blocks. By dividing the range of transform skip coefficient levels into k intervals and receiving one or more syntax values indicating which of the k intervals the level for the quantized residual value is within, and receiving a syntax element indicating the difference between the reference level value for the interval within which the level for the quantized residual value is within and the actual level for the quantized residual value of the block, a video decoder configured according to the techniques of the present disclosure may generate the advantage of achieving better coding efficiency when coding transform skip blocks compared to existing coefficient coding techniques. The reference level may be, for example, the lowest value included in the interval, but other reference values, such as the highest value of the interval or any other value of the interval, may also be used as the reference level.
[0035] The techniques of the present disclosure may be applied to any of the existing video codecs, such as High Efficiency Video Coding (HEVC), or may be proposed as promising coding tools for standards currently under development, such as Universal Video Coding (VVC), and other future video coding standards.
[0036] FIG. 1 is a block diagram illustrating an exemplary video encoding and decoding system (100) capable of performing the techniques of the present disclosure. The techniques of the present disclosure generally relate to coding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Accordingly, video data may include raw, unencoded video, encoded video, decoded (e.g., restored) video, and video metadata such as signaling data.
[0037] As illustrated in FIG. 1, the system (100) includes a source device (102) that provides encoded video data to be decoded and displayed by a destination device (116) in this example. In particular, the source device (102) provides the video data to the destination device (116) via a computer-readable medium (110). The source device (102) and the destination device (116) may include any of a wide range of devices, such as a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a telephone handset, such as a smartphone, a television, a camera, a display device, a digital media player, a video gaming console, a video streaming device, etc. In some cases, the source device (102) and the destination device (116) may be equipped for wireless communication according to a wireless communication standard and thus may be referred to as wireless communication devices.
[0038] In the example of FIG. 1, the source device (102) includes a video source (104), memory (106), a video encoder (200), and an output interface (108). The destination device (116) includes an input interface (122), a video decoder (300), memory (120), and a display device (118). According to the present disclosure, the video encoder (200) of the source device (102) and the video decoder (300) of the destination device (116) may be configured to apply techniques for signaling residual data for transform skip blocks.
[0039] Accordingly, the source device (102) represents an example of a video encoding device, while the destination device (116) represents an example of a video decoding device. In other examples, the source device and the destination device may include other components or arrays. For example, the source device (102) may receive video data from an external video source, such as an external camera. Similarly, the destination device (116) may interface with an external display device rather than including an integrated display device.
[0040] The system (100) as illustrated in FIG. 1 is merely one example. In general, any digital video encoding and / or decoding device may perform techniques for signaling residual data for conversion skip blocks. The source device (102) and the destination device (116) are merely examples of such coding devices in which the source device (102) generates video data coded for transmission to the destination device (116). The present disclosure refers to a "coding" device as a device that performs the coding (encoding and / or decoding) of data. Accordingly, the video encoder (200) and the video decoder (300) represent examples of coding devices, specifically a video encoder and a video decoder, respectively. In some examples, the devices (102, 116) may operate in a substantially symmetric manner such that each of the devices (102, 116) includes video encoding and decoding components. Accordingly, the system (100) may support unidirectional or bidirectional video transmission between video devices (102, 116) for, for example, video streaming, video playback, video broadcasting, or video calling.
[0041] Generally, a video source (104) represents a source of video data (i.e., raw, unencoded video data) and provides a sequential series of pictures of video data (also referred to as “frames”) to a video encoder (200) that encodes data for the pictures. The video source (104) of the source device (102) may include a video capture device such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As an additional alternative, the video source (104) may generate computer graphics-based data as source video, or as a combination of live video, archived video, and computer-generated video. In each case, the video encoder (200) encodes the captured, pre-captured, or computer-generated video data. The video encoder (200) may rearrange the pictures from the received order (sometimes referred to as “display order”) to the coding order for coding. The video encoder (200) may generate a bitstream containing encoded video data. Then, the source device (102) may output the encoded video data onto a computer-readable medium (110) through an output interface (108) for reception and / or extraction by, for example, an input interface (122) of a destination device (116).
[0042] The memory (106) of the source device (102) and the memory (120) of the destination device (116) represent general-purpose memories. In some examples, the memories (106, 120) may store raw video data, e.g., raw video from a video source (104) and raw, decoded video data from a video decoder (300). Additionally or alternatively, the memories (106, 120) may store software instructions executable by, for example, the video encoder (200) and the video decoder (300), respectively. Although the video encoder (200) and the video decoder (300) are shown separately in this example, it should be understood that the video encoder (200) and the video decoder (300) may also include internal memories for functionally similar or equivalent purposes. Furthermore, the memories (106, 120) may store encoded video data, for example, output from the video encoder (200) and input to the video decoder (300). In some examples, portions of the memories (106, 120) may be allocated as one or more video buffers to store, for example, raw, decoded, and / or encoded video data.
[0043] The computer-readable medium (110) may represent any type of medium or device capable of transmitting encoded video data from a source device (102) to a destination device (116). In one example, the computer-readable medium (110) represents a communication medium that enables the source device (102) to transmit the encoded video data directly to the destination device (116) in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, an output interface (108) may modulate a transmission signal containing the encoded video data, and an input interface (122) may demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from a source device (102) to a destination device (116).
[0044] In some examples, the source device (102) may output encoded data from the output interface (108) to the storage device (112). Similarly, the destination device (116) may access the encoded data from the storage device (112) through the input interface (122). The storage device (112) may include any of various distributed or locally accessed data storage media, such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.
[0045] In some examples, the source device (102) may output encoded video data to a file server (114) or other intermediate storage device that may store the encoded video generated by the source device (102). The destination device (116) may access the stored video data from the file server (114) via streaming or downloading. The file server (114) may be any type of server device capable of storing the encoded video data and transmitting the encoded video data to the destination device (116). The file server (114) may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a NAS (network attached storage) device. The destination device (116) may access the encoded video data from the file server (114) via any standard data connection, including an internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on the file server (114). The file server (114) and the input interface (122) may be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.
[0046] The output interface (108) and input interface (122) may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the output interface (108) and input interface (122) include wireless components, the output interface (108) and input interface (122) may be configured to transmit data, such as encoded video data, according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, etc. In some examples where the output interface (108) includes a wireless transmitter, the output interface (108) and the input interface (122) may be configured to transmit data, such as encoded video data, according to other wireless standards such as IEEE 802.11 specifications, IEEE 802.15 specifications (e.g., ZigBee™), Bluetooth™ standards, etc. In some examples, the source device (102) and / or the destination device (116) may include individual system-on-a-chip (SoC) devices. For example, the source device (102) may include an SoC device for performing functions attributed to the video encoder (200) and / or the output interface (108), and the destination device (116) may include an SoC device for performing functions attributed to the video decoder (300) and / or the input interface (122).
[0047] The techniques of the present disclosure may be applied to video coding by supporting any of various multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, internet streaming video transmissions, such as DASH (dynamic adaptive streaming over HTTP), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0048] The input interface (122) of the destination device (116) receives an encoded video bitstream from a computer-readable medium (110) (e.g., a storage device (112), a file server (114), etc.). The encoded video bitstream may include signaling information defined by a video encoder (200), which is also used by a video decoder (300), such as syntax elements having values that describe the processing and / or characteristics of video blocks or other coded units (e.g., slices, pictures, groups of pictures, sequences, etc.). A display device (118) displays decoded pictures of the decoded video data to a user. The display device (118) may represent any of various display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
[0049] Although not illustrated in FIG. 1, in some examples, the video encoder (200) and the video decoder (300) may each be integrated with an audio encoder and / or an audio decoder, and may include suitable MUX-DEMUX units or other hardware and / or software to handle multiplexed streams containing both audio and video in a common data stream. Where applicable, the MUX-DEMUX units may follow the ITU H.223 multiplexer protocol, or other protocols, such as the User Datagram Protocol (UDP).
[0050] The video encoder (200) and the video decoder (300) may each be implemented as any of various suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Where the techniques are partially implemented in software, the device may store instructions for the software on a suitable, non-transient computer-readable medium and execute those instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the video encoder (200) and the video decoder (300) may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined encoder / decoder (CODEC) in a separate device. A device including a video encoder (200) and / or a video decoder (300) may include an integrated circuit, a microprocessor, and / or a wireless communication device, such as a cellular phone.
[0051] The video encoder (200) and video decoder (300) may operate according to a video coding standard, such as ITU-T H.265, also referred to as High Efficiency Video Coding, or extensions thereof, such as multi-view and / or scalable video coding extensions. Alternatively, the video encoder (200) and video decoder (300) may operate according to other proprietary or industry standards, such as ITU-T H.266, also referred to as Versatile Video Coding (VVC), or JEM (Joint Exploration Test Model). The latest draft of the VVC standard is Bross et al., "Versatile Video Coding (Draft 5)", Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 14 th Meeting: Geneva, CH, 19-27 March 2019, JVET-N1001-v8 (hereinafter referred to as "VVC Draft 5"). However, the techniques of the present disclosure are not limited to any specific coding standard.
[0052] Generally, the video encoder (200) and the video decoder (300) may perform block-based coding of pictures. The term “block” generally refers to a structure containing data to be processed (e.g., to be encoded, to be decoded, or otherwise to be used in the encoding and / or decoding process). For example, a block may contain a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, the video encoder (200) and the video decoder (300) may code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, rather than coding red, green, and blue (RGB) data for samples of a picture, the video encoder (200) and the video decoder (300) may code luminance and chrominance components, wherein the chrominance components may include both red and blue chrominance components. In some examples, the video encoder (200) converts the received RGB-formatted data into a YUV representation prior to encoding, and the video decoder (300) converts the YUV representation into an RGB format. Alternatively, pre- and post-processing units (not shown) may perform these conversions.
[0053] The present disclosure may generally refer to the coding of pictures (e.g., encoding and decoding) to include a process of encoding or decoding data of the pictures. Similarly, the present disclosure may refer to the coding of blocks of pictures to include a process of encoding or decoding data for the blocks, e.g., prediction and / or residual coding. An encoded video bitstream generally includes a series of values for syntax elements representing coding decisions (e.g., coding modes) and partitioning of the pictures into blocks. Accordingly, references to coding a picture or block should generally be understood as coding values for the syntax elements forming the picture or block.
[0054] HEVC defines various blocks including coding units (CUs), prediction units (PUs), and transformation units (TUs). According to HEVC, a video coder (such as a video encoder (200)) partitions coding tree units (CTUs) into CUs according to a quadtree structure. That is, the video coder partitions the CTUs and CUs into four identical non-nested squares, and each node of the quadtree has either zero or four child nodes. Nodes without child nodes may be referred to as "leaf nodes," and the CUs of such leaf nodes may contain one or more PUs and / or one or more TUs. The video coder may further partition the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of the TUs. In HEVC, PUs represent inter-prediction data, while TUs represent residual data. Intra-predicted CUs include intra-predicted information such as intra-mode indications.
[0055] As another example, the video encoder (200) and video decoder (300) may be configured to operate according to JEM or VVC. According to JEM or VVC, the video encoder (e.g., video encoder (200)) partitions the picture into multiple coding tree units (CTUs). The video encoder (200) may partition the CTUs according to a tree structure such as a quadtree binary tree (QTBT) structure or a multitype tree (MTT) structure. The QTBT structure eliminates the concepts of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level partitioned according to quadtree partitioning, and a second level partitioned according to binary tree partitioning. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of binary trees correspond to coding units (CUs).
[0056] In an MTT partitioning structure, blocks may be partitioned using quadtree (QT) partitions, binary tree (BT) partitions, and one or more types of tripletree (TT) partitions. A tripletree partition is a partition in which a block is split into three sub-blocks. In some examples, a tripletree partition divides the block into three sub-blocks without splitting the original block through the center. The partitioning types in MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.
[0057] In some examples, the video encoder (200) and the video decoder (300) may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, whereas in other examples, the video encoder (200) and the video decoder (300) may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for both of the chrominance components (or two QTBT / MTT structures for individual chrominance components).
[0058] The video encoder (200) and video decoder (300) may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures per HEVC. For the purposes of explanation, the description of the techniques of the present disclosure is presented with respect to QTBT partitioning. However, it should be understood that the techniques of the present disclosure may also be applied to video coders configured to use quadtree partitioning or other types of partitioning.
[0059] Blocks (e.g., CTUs or CUs) may be grouped in various ways within a picture. For example, a brick may refer to a rectangular area of rows of CTUs within a specific tile in the picture. A tile may refer to a rectangular area of CTUs within a specific tile column or a specific tile row in the picture. A tile column refers to a rectangular area of CTUs having a height equal to the height of the picture and a width specified by syntax elements (e.g., as in a picture parameter set). A tile row refers to a rectangular area of CTUs having a height specified by syntax elements (e.g., as in a picture parameter set) and a width equal to the width of the picture.
[0060] In some examples, a tile may be partitioned into multiple bricks, each of which may contain one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick. However, a brick, which is a true subset of a tile, may not be referred to as a tile.
[0061] Bricks in a picture may also be arranged into slices. A slice may be an integer number of bricks in a picture that may be exclusively contained in a single Network Abstraction Layer (NAL) unit. In some examples, a slice includes only a number of complete tiles or a continuous sequence of complete bricks of a single tile.
[0062] The present disclosure may interchangeably use "NxN" and "N by N" to refer to the sample dimensions of a block (such as a CU or other video block) in terms of vertical and horizontal dimensions, e.g., 16x16 samples or 16 by 16 samples. Generally, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU may be arranged in rows and columns. Furthermore, CUs do not necessarily have to have the same number of samples in the horizontal direction as in the vertical direction. For example, CUs may contain N×M samples, where M does not necessarily have to be equal to N.
[0063] A video encoder (200) encodes video data for CUs and other information representing prediction and / or residual information. The prediction information indicates how the CU will be predicted to form a prediction block for the CU. The residual information generally represents sample-by-sample differences between the samples of the CU prior to encoding and the prediction block.
[0064] To predict the CU, the video encoder (200) may generally form a prediction block for the CU through inter-prediction or intra-prediction. Inter-prediction generally refers to predicting the CU from data of a previously coded picture, whereas intra-prediction generally refers to predicting the CU from data of the same picture that was previously coded. To perform inter-prediction, the video encoder (200) may generate a prediction block using one or more motion vectors. The video encoder (200) may generally perform motion search to identify a reference block that closely matches the CU, for example, in terms of the differences between the CU and the reference block. The video encoder (200) may calculate a difference metric using the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared differences (MSD), or other such difference calculations to determine whether a reference block closely matches the current CU. In some examples, the video encoder (200) may predict the current CU using unidirectional prediction or bidirectional prediction.
[0065] Some examples of JEM and VVC also provide an affine motion compensation mode that may be considered as an inter-prediction mode. In an affine motion compensation mode, the video encoder (200) may determine two or more motion vectors representing non-translational motion, such as zoom in or zoom out, rotation, perspective motion, or other irregular motion types.
[0066] To perform intra-prediction, the video encoder (200) may select an intra-prediction mode to generate prediction blocks. Some examples of JEM and VVC provide 67 intra-prediction modes, including planar mode and DC mode, as well as various directional modes. Generally, the video encoder (200) selects an intra-prediction mode that describes neighbor samples for the current block (e.g., a block of CU) to predict samples of the current block. Such samples may generally be located above, above and to the left of, or to the left of the current block in the same picture as the current block, assuming that the video encoder (200) codes CTUs and CUs in raster scan order (from left to right, from top to bottom).
[0067] The video encoder (200) encodes data indicating the prediction mode for the current block. For example, for inter-prediction modes, the video encoder (200) may encode motion information for the corresponding mode as well as data indicating which of the various available inter-prediction modes is used. For unidirectional or bidirectional inter-prediction, for example, the video encoder (200) may encode motion vectors using an Advanced Motion Vector Prediction (AMVP) or merge mode. The video encoder (200) may also encode motion vectors for an affine motion compensation mode using similar modes.
[0068] Following a prediction, such as an intra-prediction or inter-prediction of a block, the video encoder (200) may calculate residual data for the block. Residual data, such as a residual block, represents sample-by-sample differences between the block and the prediction block for the block formed using the corresponding prediction mode. The video encoder (200) may apply one or more transformations to the residual block to generate transformed data in a transformation domain instead of a sample domain. For example, the video encoder (200) may apply a Discrete Cosine Transform (DCT), an Integer Transform, a Wavelet Transform, or a conceptually similar transformation to the residual video data. Additionally, the video encoder (200) may apply a second transformation following a first transformation, such as a Mode-Dependent Non-Separable Second Transform (MDNSST), a Signal-Dependent Transform, or a Karhunen-Loeve Transform (KLT). The video encoder (200) generates transformation coefficients following the application of one or more transformations.
[0069] Although examples in which transformations are performed are described above, in some examples, the transformation may be skipped. For example, the video encoder (200) may implement a transformation skip mode in which the transformation operation is skipped. In examples where the transformation is skipped, the video encoder (200) may output coefficients corresponding to residual values instead of transformation coefficients. The coefficients corresponding to residual values may, for example, correspond to quantized residual values. In the following description, the term "coefficient" should be interpreted to include either coefficients corresponding to residual values or transformation coefficients generated from the result of the transformation.
[0070] As mentioned above, the video encoder (200) may perform quantization of the transform coefficients or residual values. Quantization generally refers to a process in which values are quantized to provide additional compression, so as to potentially reduce the amount of data used to represent those values. By performing the quantization process, the video encoder (200) may reduce the bit depth associated with some or all of the coefficients. For example, the video encoder (200) may round down n-bit values to m-bit values during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder (200) may perform a bitwise right-shift of the values to be quantized.
[0071] Following quantization, the video encoder (200) may scan the coefficients to generate a one-dimensional vector from a two-dimensional matrix containing the quantized coefficients. For transform coefficients, the scan may be designed to place higher energy (and therefore lower frequency) coefficients at the front of the vector and lower energy (and therefore higher frequency) transform coefficients at the rear of the vector. For transform skip coefficients, the same or different scans may be used. In some examples, the video encoder (200) may utilize a predefined scan order to scan the quantized coefficients to generate a serialized vector, and then entropy-encode the quantized coefficients of the vector. In other examples, the video encoder (200) may perform an adaptive scan. After scanning the quantized coefficients to form a one-dimensional vector, the video encoder (200) may entropy-encode the one-dimensional vector according to, for example, Context-Adaptive Binary Arithmetic Coding (CABAC). The video encoder (200) may also entropy-encode values for syntax elements describing metadata associated with the encoded video data for use by the video decoder (300) in decoding the video data.
[0072] To perform CABAC, the video encoder (200) may assign a context within a context model to the symbol to be transmitted. The context may, for example, be related to whether the neighbor values of the symbol are zero values. The probability determination may be based on the context assigned to the symbol.
[0073] The video encoder (200) may additionally generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, for the video decoder (300), such as picture headers, block headers, slice headers, or other syntax data, such as sequence parameter sets (SPS), picture parameter sets (PPS), or video parameter sets (VPS). Likewise, the video decoder (300) may decode such syntax data to determine how to decode the corresponding video data.
[0074] In this way, the video encoder (200) may generate a bitstream containing syntax elements describing the partitioning of the encoded video data, for example, into blocks of a picture (e.g., CUs), and prediction and / or residual information for the blocks. Ultimately, the video decoder (300) may receive the bitstream and decode the encoded video data.
[0075] Generally, the video decoder (300) performs a process opposite to that performed by the video encoder (200) to decode the encoded video data of the bitstream. For example, the video decoder (300) may decode values for the syntax elements of the bitstream using CABAC in a manner substantially similar but opposite to the CABAC encoding process of the video encoder (200). The syntax elements may define the CUs of the CTU by defining partitioning information for the picture's CTUs and partitioning of each CTU according to a corresponding partition structure, such as a QTBT structure. The syntax elements may additionally define prediction and residual information for blocks of video data (e.g., CUs).
[0076] Residual information may be represented, for example, by quantized transform coefficients or quantized transform skip coefficients. The video decoder (300) may inversely quantize the quantized coefficients of the block and, if coded in the transform block, inversely transform them to reconstruct a residual block for the block. The video decoder (300) forms a prediction block for the block using a signaled prediction mode (intra- or inter-prediction) and associated prediction information (e.g., motion information for inter-prediction). Then, the video decoder (300) may reconstruct the original block by combining the prediction block and the residual block (on a sample-by-sample basis). The video decoder (300) may perform additional processing, such as performing a deblocking process to reduce visual artifacts along the boundaries of the block.
[0077] According to the techniques of the present disclosure, a video encoder (200) may be configured to determine that a block of video data is encoded in a transform skip mode, to determine a level for a block of quantized residuals, and to determine a quantization parameter for a block of video data. Based on the determined quantization parameter, the video encoder (200) may be configured to determine a range for levels of quantized residual values of a block of video data and to divide the range into k intervals, where k represents an integer value. After determining which of the k intervals the level of quantized residual values is within, the video encoder (200) may determine a difference value representing the difference between a reference level value for the interval within which the level of quantized residual values is within and the level of quantized residual values of the block. A video encoder (200) may be configured to signal a level for a quantized residual value of a block based on k intervals by generating one or more syntax elements that indicate an interval within which a level for a quantized residual value is contained and a syntax element that indicates a difference value, in order to include it in a bitstream of encoded video data.
[0078] According to the techniques of the present disclosure, a video decoder (300) may be configured to determine that a block of video data is encoded in a transform skip mode and to determine quantization parameters for the block of video data. Based on the determined quantization parameters, the video decoder (300) may determine a range of levels of quantized residual values of the block of video data and divide the range into k intervals, where k represents an integer value. The video decoder (300) may receive one or more syntax elements indicating which of the k intervals the level of the quantized residual value is within, receive a syntax element indicating a difference value indicating a difference between a reference level value for the interval in which the level of the quantized residual value is within and a level of the quantized residual value of the block, and determine the level of the quantized residual value of the block based on the k intervals by determining the level of the quantized residual value based on the reference level value and the difference value.
[0079] By signaling the level for a quantized residual value based on which of the k intervals the level for the quantized residual value is within and a difference value indicating the difference between the level for the quantized residual value of the block and the reference level value for the interval within which the level for the quantized residual value is within, the video encoder (200) and the video decoder (300) may entropy-code the quantized residual values for blocks coded in transform skip mode using fewer bits compared to existing techniques for coding quantized residual values. By entropy-coding the quantized residual values for blocks coded in transform skip mode using fewer bits, the video encoder (200) and the video decoder (300) may improve the rate distortion trade-offs of the coded video data, as they may achieve better compression without adding any additional distortion.
[0080] The present disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" may refer to the communication of values for syntax elements and / or other data used to decode encoded video data. That is, the video encoder (200) may signal values for syntax elements in a bitstream. Generally, signaling refers to generating values in a bitstream. As mentioned above, the source device (102) may transmit the bitstream to the destination device (116) non-real-time or substantially real-time, as may occur when storing syntax elements in the storage device (112) for subsequent retrieval by the destination device (116).
[0081] FIGS. 2a and 2b are conceptual diagrams illustrating an exemplary quadtree binary tree (QTBT) structure (130) and a corresponding coding tree unit (CTU) (132). Solid lines represent quadtree splitting, and dotted lines represent binary tree splitting. At each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which splitting type (i.e., horizontal or vertical) is used, and in this example, 0 represents horizontal splitting and 1 represents vertical splitting. In the case of quadtree splitting, there is no need to indicate the splitting type, because quadtree nodes split blocks horizontally and vertically into four sub-blocks of the same size. Accordingly, the video encoder (200) may encode the syntax elements (e.g., splitting information) for the region tree levels (i.e., solid lines) of the QTBT structure (130) and the syntax elements (e.g., splitting information) for the prediction tree levels (i.e., dotted lines) of the QTBT structure (130), and the video decoder (300) may decode them. The video encoder (200) may encode the video data, such as prediction and transformation data for the CUs represented by the end leaf nodes of the QTBT structure (130), and the video decoder (300) may decode it.
[0082] Generally, the CTU (132) of FIG. 2b may be associated with parameters that define the sizes of blocks corresponding to the nodes of the QTBT structure (130) at the first and second levels. These parameters may include CTU size (indicating the size of the CTU (132) in the samples), minimum quadtree size (MinQTSize, indicating the minimum allowed quadtree leaf node size), maximum binary tree size (MaxBTSize, indicating the maximum allowed binary tree root node size), maximum binary tree depth (MaxBTDepth, indicating the maximum allowed binary tree depth), and minimum binary tree size (MinBTSize, indicating the minimum allowed binary tree leaf node size).
[0083] The root node of the QTBT structure corresponding to the CTU may have four child nodes at the first level of the QTBT structure, each of which may be partitioned according to quadtree partitioning. That is, the nodes at the first level are either leaf nodes (without child nodes) or have four child nodes. An example of the QTBT structure (130) represents such nodes as including child nodes and parent nodes with solid lines for branches. If the nodes at the first level are not larger than the maximum allowed binary tree root node size (MaxBTSize), the nodes may be further partitioned by individual binary trees. Binary tree splitting of a node may be repeated until the nodes resulting from the splitting reach the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). An example of a QTBT structure (130) represents such nodes by having dotted lines for branches. Binary tree leaf nodes are referred to as coding units (CUs) used for prediction (e.g., intra-picture or inter-picture prediction) and transformation without any additional partitioning. As discussed above, CUs may also be referred to as "video blocks" or "blocks".
[0084] In one example of a QTBT partitioning structure, the CTU size is set to 128x128 (luminance samples and two corresponding 64x64 chroma samples), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. Quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes ranging from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If a leaf quadtree node is 128x128, it will not be further split by the binary tree because its size exceeds MaxBTSize (i.e., 64x64 in this example). Otherwise, the leaf quadtree nodes will be further partitioned by the binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further splitting is not allowed. If a binary tree node has a width equal to MinBTSize (4 in this example), it implies that further horizontal splitting is not allowed. Similarly, a binary tree node with a height equal to MinBTSize implies that further vertical splitting is not allowed for that binary tree node. As mentioned above, the leaf nodes of the binary tree are referred to as CUs and are further processed according to prediction and transformation without further partitioning.
[0085] The present disclosure describes techniques related to binarization performed in transform skip residual coding, also referred to as transform skip mode. More specifically, the present disclosure describes techniques related to an entropy decoding process that transforms a binary representation into quantized coefficients of a series of non-binary values. A corresponding entropy encoding process, which is the inverse process of entropy decoding, is also described herein.
[0086] B. Bross, T. Nguyen, P. Keydel, H. Schwarz, D. Marpe, T. Wiegand, "Non-CE8: Unified Transform Type Signalling and Residual Coding for Transform Skip", JVET document JVET-M0464, Marrackech, MA, Jan 2019 (hereinafter JVET-M0464) describes techniques related to transform skip residual coding.
[0087] When implementing the transform skip residual coding described in JVET-M0464, the video decoder (300) has syntax elements sig_coeff_flag and abs_level_gtX_flag Using s, coefficient level ( CoeffLevel Decoding ), where X = 1,2,..5, par_level_flag , abs_remainder, and coeff_sign_flag is. Syntax element sig_coeff_flag Indicates whether the coefficient is not zero. Syntax element coeff_sign_flag indicates whether the coefficient is negative, and the syntax element par_level_flag indicates whether the coefficient is odd or even. Syntax element abs_level_gtX_flag s (X = 1,2,..5) is the absolute coefficient level Indicates whether it is greater than, where << represents a left shift operation. Specifically, if the coefficient is not zero (i.e., sig_coeff_flag = 0 ), video decoder (300) is flag abs_level_gt1_flag s is received and parsed, where the flag indicates whether the absolute count value is greater than 1. If the absolute count value is greater than 1, the video decoder (300) has a syntax indicating whether the absolute count value is greater than 2. abs_level_gt2_flag Receives s. Similarly, the absolute coefficient value If it is larger, the video decoder (300) is In this case, the absolute coefficient value Syntax element indicating whether it is larger abs_level_gtX_flag s is received. If the absolute count level is greater than 10, the video decoder (300) receives the difference, e.g., Syntax element displaying abs_remainder Receives and parses.
[0088] Next, the video decoder (300) may derive restored transform coefficients for non-zero coefficients as follows:
[0089]
[0090] However, the binarization techniques of JVET-M0464 do not reflect the dynamic range of absolute coefficient levels for different quantization parameters. For low QP ranges, the absolute quantized transform skip coefficients can have large values, such as greater than 20. In such cases, all coded into regular bins abs_level_gtX_flag s must be signaled by the video encoder (200), while the large remainder value ( CoeffLevel ) - 10 must also be signaled as bypass beans. Bypass coding generally refers to entropy coding that is not context-adaptive.
[0091] According to the techniques of the present disclosure, a new binarization process is proposed for transform skip residual coding. For input quantization parameters of QP, a video encoder (200) and a video decoder (300) may be configured to derive a corresponding dynamic range [0, maxTsLevel] of transform skip coefficients as follows:
[0092]
[0093] Here, QUANT_SHIFT is the quantization shift parameter (currently set to 14 in VTM5.0), qpPer is equal to QP / 6, and quantisationCoefficient is the quantizationScaler currently derived based on the lookup table quantisationLookUp in VTM5.0:
[0094] quantisationLookUp = [26214,23302,20560,18396,16384,14564].
[0095] quantizationScaler quantisationCoefficient is identical to quantisationLookUp[qpRem], where qpRem is the remainder of QP divided by 6. The value of quantisationLookUp[] is
[0096]
[0097] It is derived by.
[0098] For the purpose of explanation, the original residual can be expressed as R, and the quantized residual can be expressed as Rq. The equation used for quantization is
[0099]
[0100] and, where qStep is a function of QP:
[0101] A large value for QP corresponds to a larger value for qStep, and therefore a smaller Rq, which implies coarser quantization.
[0102] The following pseudocode illustrates an exemplary implementation of the above expression in software using an integer implementation:
[0103]
[0104] QP%6 represents the remainder when QP is divided by 6.
[0105] The value of quantisationLookUp[] is
[0106]
[0107] It can be derived by,
[0108] iQBits = 14 + qpPer, where qpPer is the quotient of QP / 6.
[0109] For a given QP value, the maximum quantized size maxTsLevel occurs for the largest residual value, which is (1 << channelBitDepth) - 1.
[0110] In one example, after the video encoder (200) and video decoder (300) calculate the dynamic range of quantized coefficients, the video encoder (200) and video decoder (300) can divide the range into k intervals:
[0111] [2, t1], [t1+1, t2], [t2+1, t3], ....... [t k-2 +1, t k-1 ], [t k-1 +1, maxTsLevel].
[0112] Syntax elements as in the current transformation skip residual coefficients described in JVET-M0464 sig_coeff_flag , coeff_sign_flag, and abs_level_gt1_flag After signaling, the video encoder (200) has syntax elements indicating the levels of absolute coefficients abs_level_gtTX_flag It may also signal. Specifically, if the absolute count level is greater than 1, the video encoder (200) has a syntax element indicating whether the absolute level is greater than t1. abs_level_gtT1_flag Signals . Similarly, for other level syntax, the absolute coefficient level is t X-1 If it is greater than, the video encoder (200) has an absolute count level t X Syntax element indicating whether it is larger abs_level_gtTX_flag Signals. Syntax element abs_level_gtTX_flag is the syntax for all coefficients of the current block abs_level_gtTX_flag a syntax element abs_level_gtT(X+1)_flag In the bitplane mode to be coded before coding, or all for a single coefficient abs_level_gtTX_flag Elements can be coded in an interleaving manner where they are coded before the next coefficient is coded. In the current transformation skip residual coding of VTM5.0 abs_level_gtX_flag Similar to coding, syntax elements abs_level_gtTX_flag It can be coded with regular beans if the number of regular coded beans (currently set as 2 * block width * block height) has not been reached. If the limit is reached, the rest of the syntax elements are bypass coded.
[0113] In the last pass sig_coeff_flag , coeff_sign_flag, abs_level_gt1_flag and abs_level_gtTX_flag After encoding, the video encoder (200) uses, for example, truncated unary coding or ly codes, syntax elements that indicate the remainder within the interval to which the absolute count level belongs in bypass mode abs_remainder It can also signal. For example, the absolute coefficient level absCoeffLevel This interval [t c-1 +1, t c If it belongs to ], syntax abs_level_gtT1_flag , abs_level_gtT1_flag, ...... abs_level_gtTc_flag will be signaled. In the last pass, the rest absCoeffLevel - ( t c-1 +1) can be signaled in the next bypass mode.
[0114] The video decoder (300) receives the syntax elements described above and may derive restored transformation coefficients for non-zero coefficients as follows:
[0115]
[0116] Because the proposed binarization process for transform-skip residual coding can better model a large dynamic range of coefficient levels, the techniques of the present disclosure can also be used with quantization parameter offsets (qp_offset), wherein QP set for the current coding unit is modified as QP - qp_offset when the proposed binarization for transform-skip residual coding is applied. qp_offset is a positive integer.
[0117] The binarization techniques described herein for transform skip residual coding may also be applied to the coded coefficients after the quantized residual DPCM (RDPCM) has been applied. However, due to residual subtraction, the maximum coefficient size is 2*maxTsLevel, not maxTsLevel. The maximum size occurs when the coefficient of value maxTsLevel is predicted to be the coefficient of value -maxTsLevel, resulting in a coefficient residual of value 2*maxTsLevel. Accordingly, the video encoder (200) and the video decoder (300) may be configured to calculate the dynamic range for the blocks coded in RDPCM as [0, 2*maxTsLevel].
[0118] According to one exemplary technique of the present disclosure, a video encoder (200) and a video decoder (300) may be configured to calculate a dynamic range of coefficients and divide the range into k (inclusive) intervals as follows:
[0119] Regarding the coefficients of the transformation skip mode:
[0120] [2, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 +1, maxTsLevel].
[0121] Regarding the coefficients of RDPCM mode:
[0122] [2, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 +1, 2*maxTsLevel].
[0123] As described in JVET-M0464, syntax elements as in the current transformation skip residual coefficients sig_coeff_flag , coeff_sign_flag and abs_level_gt1_flag After signaling, the video encoder (200) may signal the index of the interval to which the coefficient belongs and the remainder within the interval. The following are several examples of signaling interval indices.
[0124] In the first example, the absolute coefficient level is greater than 1, i.e., abs_level_gt1_flag=1 In this case, the video encoder (200) has a syntax element indicating whether the absolute level is greater than t1 abs_level_gtT1_flag Signals . Similarly, for other level syntax, the absolute coefficient level is t X-1 If it is greater than, the video encoder (200) has an absolute count level t X Syntax element indicating whether it is larger abs_level_gtTX_flag Signals. Syntax element abs_level_gtTX_flag is the syntax for all coefficients of the current block abs_level_gtTX_flag a syntax element abs_level_gtT(X+1)_flag In the bitplane mode to be coded before coding, or all for a single coefficient abs_level_gtTX_flag Elements can be coded in an interleaving manner where they are coded before the next coefficients are coded. In the current transformation skip residual coding of VTM. abs_level_gtX_flag Similar to coding, syntax elements abs_level_gtTX_flag It can be coded with regular beans if the number of regular coded beans (currently set as 2 * block width * block height) has not been reached. If the limit is reached, the rest of the syntax elements are bypass coded.
[0125] In another example, the absolute coefficient level is greater than 1, i.e. abs_level_gt1_flag=1 In this case, the video encoder (200) has coefficients for the interval [t k-1 +1, t k If it belongs to ], signal the value k. The video encoder (200) may encode the value in a bypass code, for example, by using Reiss-Golomb coding or truncated binary coding. In some examples, the video encoder (200) may encode the value into context-coded bins after binarization, for example, using unary coding, where each bin of the unary code is coded as a context-coded bin. When a regular coded bin limit is reached, such as a limit set to 2 * block width * block height, the video encoder (200) may bypass-code the remaining bins.
[0126] After signaling the interval indices, the video encoder (200) has a syntax element that indicates the remainder within the interval to which the absolute count level belongs. abs_remainder It can also signal. For example, the absolute coefficient level absCoeffLevel This interval [t c-1 +1, t c If it belongs to ], the video encoder (200) is the remainder absCoeffLevel - (t c-1+1) signals. The video encoder (200) may signal the remainder in bypass mode, for example, using truncated unary coding or lice codes. The video decoder (300) receives the syntax elements described herein and, based on the syntax elements received in the manner described above absCoeffLevel You can also determine the value for .
[0127] According to another example of the present disclosure, the video encoder (200) and the video decoder (300) may be configured to determine a dynamic range of coefficient values and divide the range into k (inclusive) intervals as follows:
[0128] Regarding the coefficients of the transformation skip mode:
[0129] [1, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 +1, maxTsLevel].
[0130] Regarding the coefficients of RDPCM mode:
[0131] [1, t1], [t1+1, t2], [t2+1, t3], ...[t k-2 +1, t k-1 ], [t k-1 +1, 2*maxTsLevel].
[0132] The video encoder (200) has syntax elements, as in the current transform skip residual coding described in JVET-M0464 sig_coeff_flag and coeff_sign_flag After signaling, the index of the interval to which the coefficient belongs and the remainder within the interval may also be signaled as described above.
[0133] FIG. 3 is a block diagram illustrating an exemplary video encoder (200) capable of performing the techniques of the present disclosure. FIG. 3 is provided for illustrative purposes and should not be considered a limitation to the techniques as broadly illustrated and described in the present disclosure. For illustrative purposes, the present disclosure describes a video encoder (200) in the context of video coding standards such as the HEVC video coding standard and the H.266 video coding standard under development. However, the techniques of the present disclosure are not limited to these video coding standards and are generally applicable to video encoding and decoding.
[0134] In the example of FIG. 3, the video encoder (200) includes a video data memory (230), a mode selection unit (202), a residual generation unit (204), a transform processing unit (206), a quantization unit (208), an inverse quantization unit (210), an inverse transform processing unit (212), a restoration unit (214), a filter unit (216), a decoded picture buffer (DPB) (218), and an entropy encoding unit (220). The entropy encoding unit (220) includes a transform skip syntax processing unit (221). Any or all of the video data memory (230), mode selection unit (202), residual generation unit (204), transform processing unit (206), quantization unit (208), inverse quantization unit (210), inverse transform processing unit (212), restoration unit (214), filter unit (216), DPB (218), and entropy encoding unit (220) may be implemented in one or more processors or in a processing circuit. Furthermore, the video encoder (200) may include additional or alternative processors or processing circuits to perform these and other functions.
[0135] The video data memory (230) may store video data to be encoded by the components of the video encoder (200). The video encoder (200) may receive video data stored in the video data memory (230) from, for example, a video source (104) (Fig. 1). The DPB (218) may act as a reference picture memory that stores reference video data for use in predicting subsequent video data by the video encoder (200). The video data memory (230) and the DPB (218) may be formed by any of the various memory devices, such as dynamic random access memory (DRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices, including synchronous DRAM (SDRAM). The video data memory (230) and the DPB (218) may be provided by the same memory device or by separate memory devices. In various examples, the video data memory (230) may be on-chip with other components of the video encoder (200) as exemplified, or off-chip with respect to those components.
[0136] In the present disclosure, a reference to the video data memory (230) should not be interpreted as being limited to memory inside the video encoder (200) unless specifically described as such, or to memory outside the video encoder (200) unless specifically described as such. Rather, a reference to the video data memory (230) should be understood as a reference memory that stores video data received by the video encoder (200) for encoding (e.g., video data for the current block to be encoded). The memory (106) of FIG. 1 may also provide temporary storage of outputs from various units of the video encoder (200).
[0137] Various units of FIG. 3 are illustrated to aid in understanding the operations performed by the video encoder (200). The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functions and are pre-configured for the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions for the operations that can be performed. For example, programmable circuits may execute software or firmware that causes the programmable circuits to operate in a manner defined by the commands of the software or firmware. Fixed-function circuits may execute software commands (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuits are generally invariant. In some examples, one or more of the units may be separate circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.
[0138] The video encoder (200) may include arithmetic logic units (ALUs), basic function units (EFUs), digital circuits, analog circuits, and / or programmable cores formed from programmable circuits. In examples where the operations of the video encoder (200) are performed using software executed by the programmable circuits, memory (106) (FI. 1) may store object code of software that the video encoder (200) receives and executes, or other memory within the video encoder (200) (not shown) may store such instructions.
[0139] The video data memory (230) is configured to store received video data. The video encoder (200) may extract a picture of video data from the video data memory (230) and provide the video data to the residual generation unit (204) and the mode selection unit (202). The video data in the video data memory (230) may be raw video data to be encoded.
[0140] The mode selection unit (202) includes a motion estimation unit (222), a motion compensation unit (224), and an intra-prediction unit (226). The mode selection unit (202) may include additional function units to perform video prediction according to different prediction modes. As examples, the mode selection unit (202) may include a palette unit, an intra-block copy unit (which may be part of the motion estimation unit (222) and / or the motion compensation unit (224)), an affine unit, a linear model (LM) unit, etc.
[0141] The mode selection unit (202) generally adjusts multiple encoding passes to test combinations of encoding parameters and the resulting rate-distortion values for such combinations. The encoding parameters may include partitioning of CTUs into CUs, prediction modes for CUs, transformation types for residual data of CUs, quantization parameters for residual data of CUs, etc. The mode selection unit (202) may ultimately select a combination of encoding parameters that has better rate-distortion values than other tested combinations.
[0142] The video encoder (200) partitions a picture extracted from the video data memory (230) into a series of CTUs and may encapsulate one or more CTUs within a slice. The mode selection unit (202) may partition the CTUs of the picture according to a tree structure, such as the quadtree structure or QTBT structure of HEVC described above. As described above, the video encoder (200) may form one or more CUs by partitioning the CTUs according to the tree structure. Such CUs may also generally be referred to as "video blocks" or "blocks."
[0143] Generally, the mode selection unit (202) also controls its components (e.g., motion estimation unit (222), motion compensation unit (224), and intra-prediction unit (226)) to generate a prediction block for the current block (e.g., the current CU, or, in HEVC, the overlapping part of PU and TU). For inter-prediction of the current block, the motion estimation unit (222) may perform motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously coded pictures stored in DPB (218). In particular, the motion estimation unit (222) may calculate a value indicating how similar a potential reference block is to the current block, for example, according to the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared differences (MSD), etc. The motion estimation unit (222) may perform these calculations using sample-by-sample differences between a reference block generally considered and a current block. The motion estimation unit (222) may also identify a reference block having the lowest value resulting from these calculations, indicating a reference block that matches most closely to the current block.
[0144] The motion estimation unit (222) may form one or more motion vectors (MVs) that define the positions of reference blocks in reference pictures for the position of the current block in the current picture. Then, the motion estimation unit (222) may provide the motion vectors to the motion compensation unit (224). For example, for unidirectional inter-prediction, the motion estimation unit (222) may provide a single motion vector, whereas for bidirectional inter-prediction, the motion estimation unit (222) may provide two motion vectors. Then, the motion compensation unit (224) may generate a prediction block using the motion vectors. For example, the motion compensation unit (224) may extract data for the reference block using the motion vectors. As another example, if the motion vector has fractional sample precision, the motion compensation unit (224) may interpolate the values for the prediction block according to one or more interpolation filters. Furthermore, for bidirectional inter-prediction, the motion compensation unit (224) may extract data for two reference blocks identified by individual motion vectors and combine the extracted data, for example, through sample-by-sample averaging or weighted averaging.
[0145] As another example, for intra-prediction or intra-prediction coding, the intra-prediction unit (226) may generate prediction blocks from samples adjacent to the current block. For example, for directional modes, the intra-prediction unit (226) may generally generate prediction blocks by mathematically combining the values of neighboring samples and populating these calculated values in a direction defined across the current block. As another example, for DC mode, the intra-prediction unit (226) may calculate the average of neighboring samples for the current block and generate prediction blocks containing this resulting average for each sample of the prediction block.
[0146] The mode selection unit (202) provides the prediction block to the residual generation unit (204). The residual generation unit (204) receives the raw, unencoded version of the current block from the video data memory (230) and the prediction block from the mode selection unit (202). The residual generation unit (204) calculates the sample-by-sample differences between the current block and the prediction block. The resulting sample-by-sample differences define the residual block for the current block. In some examples, the residual generation unit (204) may also determine the differences between sample values in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit (204) may be formed using one or more subtractor circuits that perform binary subtraction.
[0147] In the examples where the mode selection unit (202) partitions the CUs into PUs, each PU may be associated with a luminance prediction unit and a corresponding chroma prediction unit. The video encoder (200) and the video decoder (300) may support PUs of various sizes. As indicated above, the size of the CU may refer to the size of the luminance coding block of the CU, and the size of the PU may refer to the size of the luminance prediction unit of the PU. Assuming that the size of a particular CU is 2Nx2N, the video encoder (200) may support PU sizes of 2Nx2N or NxN for intra-prediction, and may support symmetric PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, etc. for inter-prediction. The video encoder (200) and video decoder (300) may also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-prediction.
[0148] In examples where the mode selection unit does not further partition the CU into PUs, each CU may be associated with a luminance coding block and a corresponding chroma coding block. As described above, the size of the CU may refer to the size of the luminance coding block of the CU. The video encoder (200) and the video decoder (300) may support CU sizes of 2Nx2N, 2NxN, or Nx2N.
[0149] For some examples, for other video coding techniques such as intra-block copy mode coding, affine mode coding, and linear model (LM) mode coding, the mode selection unit (202) generates a prediction block for the current block being encoded through individual units associated with the coding techniques. In some examples, such as palette mode coding, the mode selection unit (202) may not generate a prediction block, but instead may generate syntax elements indicating a method for restoring the block based on the selected palette. In such modes, the mode selection unit (202) may provide these syntax elements to the entropy encoding unit (220) to be encoded.
[0150] As described above, the residual generation unit (204) receives video data for the current block and the corresponding prediction block. Then, the residual generation unit (204) generates a residual block for the current block. To generate the residual block, the residual generation unit (204) calculates the sample-by-sample differences between the prediction block and the current block.
[0151] The transformation processing unit (206) applies one or more transformations to a residual block to generate a block of transformation coefficients (referred to herein as a “transformation coefficient block”). The transformation processing unit (206) may form the coefficient block by applying various transformations to the residual block. For example, the transformation processing unit (206) may apply a Discrete Cosine Transform (DCT), a Directional Transform, a Karhunen-Loeve Transform (KLT), or a conceptually similar transformation to the residual block. In some examples, the transformation processing unit (206) may perform multiple transformations on the residual block, for example, a first transformation and a second transformation, such as a rotation transformation. In some examples, the transformation processing unit (206) does not apply transformations to the residual block. For a block of video data coded in transformation skip mode, the transformation processing unit (206) may be seen as a pass-through unit that does not modify the residual block.
[0152] The quantization unit (208) may quantize the coefficients in the coefficient block to generate a quantized coefficient block. The quantization unit (208) may quantize the coefficients in the coefficient block according to the quantization parameter (QP) value associated with the current block. The video encoder (200) may adjust the degree of quantization applied to the coefficient blocks associated with the current block by adjusting the QP value associated with the CU (e.g., via the mode selection unit (202)). Quantization may introduce a loss of information, and thus, the quantized coefficients may have lower precision than the original coefficients output by the conversion processing unit (206).
[0153] The inverse quantization unit (210) and the inverse transform processing unit (212) may each apply inverse quantization and inverse transforms to the quantized coefficient block to restore the residual block from the coefficient block. For a block of video data coded in transform skip mode, the inverse transform processing unit (212) may be seen as a pass unit that does not alter the dequantized coefficient block. The restoration unit (214) may generate a restored block corresponding to the current block (although potentially having some degree of distortion) based on the prediction block generated by the mode selection unit (202) and the restored residual block. For example, the restoration unit (214) may generate the restored block by adding samples of the restored residual block to corresponding samples from the prediction block generated by the mode selection unit (202).
[0154] The filter unit (216) may perform one or more filter operations on the restored blocks. For example, the filter unit (216) may perform deblocking operations to reduce blockiness artifacts along the edges of the CUs. The operations of the filter unit (216) may be skipped in some examples.
[0155] The video encoder (200) stores the restored blocks in the DPB (218). For example, in examples where the operations of the filter unit (216) are performed, the restoration unit (214) may store the restored blocks in the DPB (218). In examples where the operations of the filter unit (216) are not performed, the filter unit (216) may store the filtered restored blocks in the DPB (218). The motion estimation unit (222) and the motion compensation unit (224) may take a reference picture from the DPB (218) formed from the restored (and potentially filtered) blocks and inter-predict blocks of subsequently encoded pictures. Additionally, the intra-prediction unit (226) may use the restored blocks in the DPB (218) of the current picture to intra-predict other blocks in the current picture.
[0156] Generally, the entropy encoding unit (220) may entropy encode syntax elements received from other functional components of the video encoder (200). For example, the entropy encoding unit (220) may entropy encode quantized coefficient blocks from the quantization unit (208). As another example, the entropy encoding unit (220) may entropy encode prediction syntax elements (e.g., motion information for inter-prediction or intra-mode information for intra-prediction) from the mode selection unit (202). The entropy encoding unit (220) may perform one or more entropy encoding operations on syntax elements, which are other examples of video data, to generate entropy-encoded data. For example, the entropy encoding unit (220) may perform a context-adaptive variable-length coding (CAVLC) operation, a CABAC operation, a V2V (variable-to-variable) length coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a probabilistic interval partitioning entropy (PIPE) coding operation, an exponential-Golomb encoding operation, or other types of entropy encoding operations on the data. In some examples, the entropy encoding unit (220) may operate in a bypass mode where syntax elements are not entropy encoded.
[0157] According to the techniques of the present disclosure, the transform skip syntax processing unit (221) of the entropy encoding unit (220) may be configured to determine a range of levels for quantized residual values of a block of video data based on quantization parameters and to signal levels for quantized residual values of the block by dividing the range into k intervals. The transform skip syntax processing unit (221) may then signal levels for quantized residual values of the block based on k intervals by generating one or more syntax elements indicating a specific interval containing levels for quantized residual values to be included in the bitstream of the encoded video data, and by generating a syntax element indicating a difference value indicating a difference between a reference level value for a specific interval and a level for quantized residual values of the block to be included in the bitstream of the encoded video data.
[0158] The video encoder (200) may output a bitstream containing entropy-encoded syntax elements necessary to restore blocks of a picture or slice. In particular, the entropy encoding unit (220) may output the bitstream.
[0159] The operations described above are described with respect to blocks. Such descriptions should be understood as operations for Luma coding blocks and / or Chroma coding blocks. As described above, in some examples, Luma coding blocks and Chroma coding blocks are Luma and Chroma components of CU. In some examples, Luma coding blocks and Chroma coding blocks are Luma and Chroma components of PU.
[0160] In some examples, operations performed on a luminal coding block do not need to be repeated for a chroma coding block. As an example, operations to identify the motion vector (MV) and reference picture for a luminal coding block do not need to be repeated to identify the MV and reference picture for chroma blocks. Rather, the MV for the luminal coding block may be scaled to determine the MV for the chroma blocks, and the reference picture may be the same. As another example, the intra-prediction process may be the same for the luminal coding block and the chroma coding block.
[0161] The video encoder (200) also represents an example of a device configured to encode video data, comprising one or more processing units implemented in a memory and circuit configured to store video data, wherein the one or more processing units are configured to determine that a block of video data is encoded in a transform skip mode; determine a quantization parameter for a block of video data; determine a range for residual values of a block of video data based on the determined quantization parameter; divide the range into k intervals, wherein k is an integer value; and determine values for one or more syntax elements based on the k intervals. The syntax elements include, for example, the abs_level_gtTX_flag and abs_remainder syntax elements described above.
[0162] FIG. 4 is a block diagram illustrating an exemplary video decoder (300) that may perform the techniques of the present disclosure. FIG. 4 is provided for illustrative purposes and is not limited to techniques as broadly illustrated and described in the present disclosure. For illustrative purposes, the present disclosure describes a video decoder (300) according to the techniques of JEM, VVC, and HEVC. However, the techniques of the present disclosure may be performed by video coding devices composed of other video coding standards.
[0163] In the example of FIG. 4, the video decoder (300) includes a coded picture buffer (CPB) memory (320), an entropy decoding unit (302), a prediction processing unit (304), an inverse quantization unit (306), an inverse transform processing unit (308), a restoration unit (310), a filter unit (312), and a decoded picture buffer (DPB) (314). The entropy decoding unit (302) includes a transform skip syntax processing unit (303). Any or all of the CPB memory (320), the entropy decoding unit (302), the prediction processing unit (304), the inverse quantization unit (306), the inverse transform processing unit (308), the restoration unit (310), the filter unit (312), and the DPB (314) may be implemented in one or more processors or in a processing circuit. Furthermore, the video decoder (300) may include additional or alternative processors or processing circuits to perform these and other functions.
[0164] The prediction processing unit (304) includes a motion compensation unit (316) and an intra-prediction unit (318). The prediction processing unit (304) may include additional units to perform predictions according to different prediction modes. As examples, the prediction processing unit (304) may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit (316)), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder (300) may include more, fewer, or different functional components.
[0165] The CPB memory (320) may store video data, such as an encoded video bitstream to be decoded by components of the video decoder (300). The video data stored in the CPB memory (320) may be obtained, for example, from a computer-readable medium (110) (Fig. 1). The CPB memory (320) may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory (320) may store video data other than syntax elements of the encoded picture, such as transient data representing outputs from various units of the video decoder (300). The DPB (314) generally stores decoded pictures that the video decoder (300) may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. CPB memory (320) and DPB (314) may be formed by any memory device among various memory devices, such as SDRAM, DRAM, MRAM, RRAM, or other types of memory devices. CPB memory (320) and DPB (314) may be provided by the same memory device or by separate memory devices. In various examples, CPB memory (320) may be on-chip with respect to other components of the video decoder (300) or off-chip with respect to those components.
[0166] Additionally or alternatively, in some examples, the video decoder (300) may retrieve coded video data from memory (120) (Fig. 1). That is, the memory (120) may store data as discussed above in CPB memory (320). Likewise, the memory (120) may store instructions to be executed by the video decoder (300) when some or all of the functions of the video decoder (300) are implemented in software to be executed by the processing circuit of the video decoder (300).
[0167] The various units illustrated in FIG. 4 are exemplified to aid in understanding the operations performed by the video decoder (300). The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to FIG. 3, fixed-function circuits refer to circuits that provide specific functions and are pre-configured for the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functions for the operations that can be performed. For example, programmable circuits may execute software or firmware that causes the programmable circuits to operate in a manner defined by instructions of the software or firmware. Fixed-function circuits may execute software instructions (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuits are generally invariant. In some examples, one or more of the units may be separate circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.
[0168] The video decoder (300) may include ALUs, EFUs, digital circuits, analog circuits, and / or programmable cores formed from programmable circuits. In examples where the operations of the video decoder (300) are performed by software running on the programmable circuits, on-chip or off-chip memory may store instructions (e.g., object code) of the software that the video decoder (300) receives and executes.
[0169] The entropy decoding unit (302) receives encoded video data from the CPB and may regenerate syntax elements by entropy decoding the video data. The prediction processing unit (304), the inverse quantization unit (306), the inverse transform processing unit (308), the restoration unit (310), and the filter unit (312) may generate decoded video data based on syntax elements extracted from the bitstream.
[0170] Generally, the video decoder (300) restores the picture on a block-by-block basis. The video decoder (300) may also perform restoration operations for each block individually (wherein the block currently being restored, i.e., being decoded, may be referred to as the “current block”).
[0171] The entropy decoding unit (302) may entropy decode not only conversion information such as quantization parameters (QP) and / or conversion mode indication(s), but also syntax elements defining the quantized coefficients of the quantized coefficient block. The inverse quantization unit (306) may use the QP associated with the quantized coefficient block to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied by the inverse quantization unit (306). The inverse quantization unit (306) may, for example, perform a bit-by-bit left-shift operation to inverse quantize the quantized coefficients. The inverse quantization unit (306) may thereby form a coefficient block containing the coefficients.
[0172] According to the techniques of the present disclosure, for a block of video data encoded in a transform skip mode, the transform skip syntax processing unit (303) of the entropy decoding unit (302) may be configured to determine a range of levels of quantized residual values of the block of video data based on determined quantization parameters and to divide the range into k intervals. The transform skip syntax processing unit (303) may then receive information indicating that a level of quantized residual values is within a specific interval among the k intervals, receive information indicating a difference value indicating a difference between a reference level value for a specific interval and a level of quantized residual values of the block, and determine a level of quantized residual values of the block based on the k intervals by determining a level of quantized residual values based on the reference level value and the difference value.
[0173] After the inverse quantization unit (306) forms the coefficient block, the inverse transform processing unit (308) may apply one or more inverse transforms to the coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit (308) may apply an inverse DCT, an inverse integer transform, an inverse KLT (Karhunen-Loeve transform), an inverse rotation transform, an inverse directional transform, or other inverse transforms to the coefficient block. For blocks coded in transform skip mode, the inverse transform processing unit (308) may be seen as a pass unit that does not modify the dequantized coefficient block.
[0174] Furthermore, the prediction processing unit (304) generates a prediction block according to the prediction information syntax elements entropy-decoded by the entropy decoding unit (302). For example, if the prediction information syntax elements indicate that the current block is inter-predicted, the motion compensation unit (316) may generate a prediction block. In this case, the prediction information syntax elements may indicate a motion vector identifying the position of the reference block in the reference picture as well as the position of the current block in the current picture, as well as the reference picture in the DPB (314) from which the reference block will be extracted. The motion compensation unit (316) may generally perform the inter-predicting process in a manner substantially similar to that described for the motion compensation unit (224) (Fig. 3).
[0175] As another example, if the prediction information syntax elements indicate that the current block is intra-predicted, the intra-predict unit (318) may generate a prediction block according to the intra-predict mode indicated by the prediction information syntax elements. Again, the intra-predict unit (318) may perform the intra-predict process in a manner substantially similar to that generally described for the intra-predict unit (226) (Fig. 3). The intra-predict unit (318) may extract data of neighbor samples for the current block from the DPB (314).
[0176] The restoration unit (310) may restore the current block using the prediction block and the residual block. For example, the restoration unit (310) may restore the current block by adding samples of the residual block to the corresponding samples of the prediction block.
[0177] The filter unit (312) may perform one or more filter operations on the restored blocks. For example, the filter unit (312) may perform deblocking operations to reduce blockiness artifacts along the edges of the restored blocks. The operations of the filter unit (312) are not necessarily performed in all examples.
[0178] The video decoder (300) may store the restored blocks in the DPB (314). As discussed above, the DPB (314) may provide reference information to the prediction processing unit (304), such as samples of the current picture for intra-prediction and previously decoded pictures for subsequent motion compensation. Additionally, the video decoder (300) may output the decoded pictures from the DPB (314) for subsequent presentation on a display device such as the display device (118) of FIG. 1.
[0179] In this way, the video decoder (300) represents an example of a video decoding device comprising one or more processing units implemented in a memory and circuit configured to store video data, the one or more processing units being configured to determine that a block of video data is encoded in a convert skip mode; to determine a quantization parameter for a block of video data; to determine a range for residual values of a block of video data based on the determined quantization parameter; to divide the range into k intervals, wherein k is an integer value; and to determine a value for a coefficient of residual data based on the k intervals. The video decoder (300) may, for example, interpret the value of one or more syntax elements based on the determined intervals.
[0180] In some implementations, the video decoder (300) may also receive a syntax element indicating that the coefficient has a value greater than 0, a syntax element indicating that the coefficient has a value greater than 1, and / or a syntax element indicating the sign of the coefficient.
[0181] The video decoder (300) receives a syntax element indicating that for a first interval among k intervals, the value for a coefficient is greater than the value included in the first interval, receives a syntax element indicating that for a second interval among k intervals, the value for a coefficient is included in the second interval, and may also receive a syntax element indicating the difference between the initial value and the value for a coefficient for the second interval. The syntax element indicating the difference may be bypass coded.
[0182] The video decoder (300) may receive flags indicating that, for individual intervals, the value for a coefficient is greater than the values included in individual intervals for flags, until it receives a flag indicating that the value for a coefficient is within the interval for flags. The video decoder (300) may then receive a syntax element indicating the difference between the initial value for the interval for flags and the value for the coefficient. The syntax element indicating the difference may, for example, be bypass coded.
[0183] The k intervals may include, for example, intervals in the range of 0 to 0, intervals in the range of 1 to 1, intervals in the range of 2 to 1 threshold, intervals in the range of 1 to 2 thresholds, and intervals in the range of 1 to 3 thresholds. The video decoder (300) may determine the first threshold, the second threshold, and the third threshold based on quantization parameters as described above. The k intervals may also include intervals in the range of 1 to 3 thresholds plus 1 maximum value.
[0184] The video decoder (300) may be configured to receive a syntax element indicating an index for one of k intervals, and the k intervals indicated by the index correspond to an interval among the k intervals containing a value for a coefficient. The video decoder (300) may then receive a syntax element indicating a difference value between a starting value for one of the k intervals and a value for a coefficient.
[0185] FIG. 5 is a flowchart illustrating an exemplary method for encoding a current block. The current block may include a current CU. Although described in relation to a video encoder (200) (Fig. 1 and Fig. 3), it should be understood that other devices may be configured to perform a method similar to that of FIG. 5.
[0186] In this example, the video encoder (200) initially predicts the current block (350). For example, the video encoder (200) may form a predicted block for the current block. The video encoder (200) may then calculate a residual block for the current block (352). To calculate the residual block, the video encoder (200) may calculate the difference between the original, unencoded block and the predicted block for the current block. Then, the video encoder (200) may transform and quantize the coefficients of the residual block (354). For blocks coded in transform skip mode, the residual block may be only quantized and not transformed at step 354. Next, the video encoder (200) may scan the quantized coefficients of the residual block (356). During or following a scan, the video encoder (200) may entropy-encode the coefficients (358). For example, the video encoder (200) may encode the coefficients using CAVLC or CABAC. The video encoder (200) may output the entropy-encoded data of the next block (360).
[0187] FIG. 6 is a flowchart illustrating an exemplary method for decoding a current block of video data. The current block may include a current CU. Although described in relation to a video decoder (300) (Fig. 1 and Fig. 4), it should be understood that other devices may be configured to perform a method similar to that of FIG. 6.
[0188] The video decoder (300) may receive entropy-coded data for the current block, such as entropy-coded data and entropy-coded prediction information for the coefficients of the residual block corresponding to the current block (370). The video decoder (300) may entropy-decode the entropy-coded data to determine prediction information for the current block and regenerate the coefficients of the residual block (372). The video decoder (300) may predict the current block using an intra- or inter-prediction mode, such as indicated by the prediction information for the current block, to calculate the prediction block for the current block (374). The video decoder (300) may then back-scan the regenerated coefficients to generate a block of quantized coefficients (376). The video decoder (300) may then back-quantize and back-transform the coefficients to generate a residual block (378). For blocks coded in transform skip mode, the video decoder (300) may simply de-quantize the block of quantized coefficients in step 378 and not inverse transform. The video decoder (300) may ultimately decode the current block by combining the prediction block and the residual block (380).
[0189] FIG. 7 is a flowchart illustrating an exemplary method for encoding a current block. The current block may include a current CU. Although described in relation to a video encoder (200) (Fig. 1 and Fig. 3), it should be understood that other devices may be configured to perform a method similar to that of FIG. 7.
[0190] The video encoder (200) determines that a block of video data is encoded without transforming the residual data for the block (400). The block may be encoded, for example, in a transform skip mode. The video encoder (200) determines a level for the quantized residual value of the block (402). To determine a level for the quantized residual value of the block, the video encoder (200) may be configured, for example, to determine a level for the residual value and to quantize the level for the residual value to determine a level for the quantized residual value.
[0191] The video encoder (200) determines quantization parameters for a block of video data (404). Based on the determined quantization parameters, the video encoder (200) determines a range of levels for the quantized residual values of the block of video data (406).
[0192] The video encoder (200) divides the range into k intervals, where k represents an integer value (408). The k intervals are, for example, the following intervals:
[0193] [X, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 It may include [+1, maxTsLevel], where maxTsLevel represents the maximum possible level for the levels of the quantized residual values of the block based on the quantization parameter for the block, and t n represents an upper threshold value for the nth interval, where n is in the range of 0 to k-1, and X represents a minimum value for the first interval, e.g., the 0th interval. As described elsewhere, X may be equal to 0, 1, 2, or some other value.
[0194] The video encoder (200) determines a specific interval among k intervals containing levels for quantized residual values (410). The video encoder (200) determines a difference value representing the difference between a reference level value for a specific interval and a level for quantized residual values of a block (412).
[0195] The video encoder (200) signals a level for the quantized residual value of the block based on k intervals (414). As part of signaling a level for the quantized residual value of the block based on k intervals, the video encoder (200) generates one or more syntax elements indicating a specific interval (416) to be included in the bitstream of the encoded video data and generates a syntax element indicating a difference value (418) to be included in the bitstream of the encoded video data. The video encoder (200) may, for example, bypass-encode the syntax element indicating the difference value.
[0196] The video encoder (200) may, for example, generate flags indicating that for individual intervals among k intervals, the level of the quantized residual value is greater than the values included in individual intervals for the flags, for inclusion in the bitstream of the encoded video data, until a flag indicating that the level of the quantized residual value is within the interval associated with the flag is generated. The video encoder (200) may additionally generate a syntax element indicating the sign of the residual value for inclusion in the bitstream of the encoded video data.
[0197] The video encoder (200) may, for example, generate a syntax element such as a 1-bit flag indicating that the level for the quantized residual value is greater than the values included in the first interval for the first interval to be included in the bitstream of the encoded video data for the first interval; and may generate a syntax element such as another 1-bit flag indicating that the level for the quantized residual value is included in the second interval for the second interval to be included in the bitstream of the encoded video data for the second interval. For a syntax element indicating a difference value, the video encoder (200) may generate a syntax element set to the difference between the reference level value for the second interval and the level for the quantized residual value of the block.
[0198] In some examples, the first interval among the k intervals may include values in the range of 1 to the first threshold, and the video encoder (200) may be configured to generate a syntax element indicating that the level of the quantized residual value is greater than 0 in order to include it in the bitstream of the encoded video data. In cases where the level for the quantized residual value is equal to 0, the video encoder (200) does not need to signal any additional information indicating the quantized residual value. That is, the video encoder does not need to generate one or more syntax elements indicating a specific interval or a syntax element indicating a difference value in order to include it in the bitstream of the encoded video data.
[0199] In some examples, the first interval among the k intervals may include values in the range of 2 to the first threshold, and the video encoder (200) may generate a syntax element indicating that the level for the quantized residual value is greater than 0 and a syntax element indicating that the level for the quantized residual value is greater than 1 in order to include it in the bitstream of the encoded video data. In cases where the level for the quantized residual value is equal to 1, the video encoder (200) does not need to signal any additional information indicating the quantized residual value. That is, the video encoder does not need to generate one or more syntax elements indicating a specific interval or a syntax element indicating a difference value in order to include it in the bitstream of the encoded video data.
[0200] As part of signaling levels for the quantized residual values of a block based on k intervals, the video encoder (200) also outputs a bitstream of encoded video data (420). The video encoder (200) may output a bitstream of encoded video, for example, by storing the bitstream in a memory device or by transmitting the bitstream of encoded video data to another device.
[0201] FIG. 8 is a flowchart illustrating an exemplary method for decoding a current block. The current block may include a current CU. Although described in relation to a video decoder (300) (Figs. 1 and 4), it should be understood that other devices may be configured to perform a method similar to that of FIG. 5.
[0202] The video decoder (300) determines that a block of video data is encoded without transforming residual data for the block (430). The block may be encoded, for example, in transform skip mode. The video decoder (300) determines quantization parameters for a block of video data (432). The video decoder (300) may, for example, receive indications of quantization parameters from the video data.
[0203] The video decoder (300) determines a range of levels for quantized residual values of a block of video data based on determined quantization parameters (434). The range of levels for quantized residual values is typically smaller than the bit depth of the video data. For example, 8-bit video data from 0 to 2 8 If the sample values have a range of -1, the quantized residual values are 0 to 2 8 It has a range of maximum values less than -1. The maximum value is a function of a specific quantization parameter used for video data. The video decoder (300) divides the range into k intervals (436).
[0204] The video decoder (300) determines the level for the quantized residual value of the block based on k intervals (438). The k intervals are, for example, the following intervals:
[0205] [X, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 It may include [+1, maxTsLevel], where maxTsLevel represents the maximum possible level for the levels of the quantized residual values of the block based on the quantization parameter for the block, and t nrepresents an upper threshold value for the nth interval, where n is in the range of 0 to k-1, and X represents a minimum value for the first interval, e.g., the 0th interval. As described elsewhere, X may be equal to 0, 1, 2, or some other value.
[0206] In examples where X is equal to 0, k intervals include intervals in the range of 0 to a first threshold, intervals in the range of a first threshold plus 1 to a second threshold, intervals in the range of a second threshold plus 1 to a third threshold, and other intervals. In examples where X is equal to 1, k intervals include intervals in the range of 1 to a first threshold, intervals in the range of a first threshold plus 1 to a second threshold, intervals in the range of a second threshold plus 1 to a third threshold, and other intervals. In examples where X is equal to 2, k intervals include intervals in the range of 2 to a first threshold, intervals in the range of a first threshold plus 1 to a second threshold, intervals in the range of a second threshold plus 1 to a third threshold, and other intervals. The k intervals also include an interval containing the maximum value for the range.
[0207] As part of determining the level for the quantized residual value of a block based on k intervals, the video decoder (300) receives information indicating that the level for the quantized residual value is within a specific interval among the k intervals (440) and receives information indicating a difference value indicating the difference between the reference level value for the specific interval and the level for the quantized residual value of the block (442). The video decoder (300) may, for example, bypass decode the syntax element indicating the difference value. The video decoder (300) then determines the level for the quantized residual value based on the reference level value and the difference value (444).
[0208] The video decoder (300) may receive flags indicating that, for individual intervals among k intervals, the level for the quantized residual value is greater than the values included in the individual intervals for the flags, until it receives a flag indicating that the level for the quantized residual value is within the interval associated with the flag.
[0209] In examples where X is equal to 1, the video decoder (300) may receive a syntax element indicating that the level for the quantized residual value is greater than 0, before receiving information indicating that the level for the quantized residual value is within a specific interval among k intervals, or information indicating a difference value. In cases where the level for the quantized residual value is equal to 0, the video decoder (300) does not need to receive any additional information indicating the quantized residual value. That is, the video decoder (300) does not need to receive information indicating that the level for the quantized residual value is within a specific interval among k intervals, or information indicating a difference value.
[0210] In examples where X is equal to 2, the video decoder (300) receives a syntax element indicating that the level for the quantized residual value is greater than 0 and a syntax element indicating that the level for the quantized residual value is greater than 1, before receiving information indicating that the level for the quantized residual value is within a specific interval among k intervals or information indicating a difference value. In cases where the level for the quantized residual value is equal to 1, the video decoder (300) does not need to receive any additional information indicating the quantized residual value. That is, the video decoder (300) does not need to receive information indicating that the level for the quantized residual value is within a specific interval among k intervals or information indicating a difference value.
[0211] The video decoder (300) outputs decoded video data based on levels for quantized residual values (446). The video decoder (300) may, for example, output decoded video data for display or output decoded video data for storage. As part of decoding video data, the video decoder (300) may, for example, dequantize levels for quantized residual values to determine levels for dequantized residual values, receive syntax elements indicating signs for dequantized residual values, and determine dequantized residual values based on levels for dequantized residual values and signs for dequantized residual values. The video decoder (300) may also determine a residual block for a block of video data, add the residual block to a prediction block for a block of video data to determine a restored block for a block of video data, and generate a picture of the decoded video data based on the restored block. The video decoder (300) may additionally perform one or more filtering operations on the restored block.
[0212] It should be recognized that, depending on the example, certain acts or events of any of the techniques described herein may be performed in different sequences, added, merged, or omitted entirely (e.g., not all described acts or events are necessary for the implementation of the techniques). Furthermore, in certain examples, acts or events may be performed simultaneously rather than sequentially, for example, through multi-threaded processing, interrupt processing, or multiple processors.
[0213] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. Where implemented in software, the functions may be stored as one or more instructions or code on a computer-readable medium or transmitted therethrough, or executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media corresponding to media of the same type as data storage media, or communication media including any medium that facilitates the transmission of a computer program from one place to another, for example, according to a communication protocol. In this way, computer-readable media may generally correspond to (1) non-transient types of computer-readable storage media or (2) communication media such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to extract instructions, code and / or data structures for the implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0214] As an example, not a limitation, these computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. Additionally, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead relate to non-transient, tangible storage media. As used herein, disks and discs include compact discs (CDs), laser discs, optical discs, digital multi-purpose discs (DVDs), floppy discs, and Blu-ray discs, wherein disks typically reproduce data magnetically, but discs reproduce data optically using lasers. The above combinations must also be included within the range of computer-readable media.
[0215] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, terms such as "processor" and "processing circuit" as used herein may refer to any of any other structures suitable for implementing the structures described above or the techniques described herein. Additionally, in some embodiments, the functions described herein may be provided within dedicated hardware and / or software modules integrated into a codec configured for or combined with encoding and decoding. Furthermore, the techniques may be fully implemented in one or more circuits or logic elements.
[0216] The techniques of the present disclosure may be implemented in a wide range of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chip sets). Various components, modules, or units are described in the present disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but implementation by different hardware units is not required. Rather, as described above, various units may be combined in a codec hardware unit or provided by a set of interoperable hardware units including one or more processors as described above, together with suitable software and / or firmware.
[0217] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
Claim 1 A method for decoding video data, comprising: determining that a block of video data is encoded without transforming residual data for said block; determining a quantization parameter for said block of video data; determining, based on the determined quantization parameter, a range for levels of quantized residual values of said block of video data, wherein the range is from 0 to a maximum possible level for said levels of quantized residual values of said block of video data; and dividing said range into k intervals, wherein k is an integer value, and each interval includes multiple levels of quantized residual values, and each interval has an associated index value between 0 and k-1. A method for decoding video data, comprising: a step of determining a level for a quantized residual value of a block of video data based on k intervals, wherein the step of determining a level for a quantized residual value of a block of video data based on k intervals comprises: receiving information indicating an index corresponding to a specific interval among the k intervals to which the level for the quantized residual value belongs; receiving information indicating a difference value indicating a difference between a reference level value for the specific interval and a level for the quantized residual value of a block of video data; and determining a level for the quantized residual value based on the reference level value and the difference value; and outputting decoded video data based on the level for the quantized residual value. Claim 2 A method for decoding video data according to claim 1, further comprising: a step of determining a level for a quantized residual value by dequantizing a level for the quantized residual value; a step of receiving a syntax element indicating a sign for the quantized residual value; and a step of determining the quantized residual value based on the level for the quantized residual value and the sign for the quantized residual value. Claim 3 A method for decoding video data according to claim 2, further comprising: a step of determining a residual block for a block of video data, wherein the residual block includes the dequantized residual value; a step of determining a restored block for a block of video data by adding the residual block to a prediction block for a block of video data; a step of generating a picture of decoded video data based on the restored block; and a step of outputting the picture of decoded video data. Claim 4 A method for decoding video data according to claim 1, wherein the first interval among the k intervals includes values in the range of 1 to 1 threshold value, and the method further comprises the step of receiving a syntax element indicating that the level for the quantized residual value is greater than 0. Claim 5 A method for decoding video data according to claim 1, wherein the first interval among the k intervals includes values in the range of 2 to 1 threshold values, and the method further comprises the steps of: receiving a syntax element indicating that the level for the quantized residual value is greater than 0; and receiving a syntax element indicating that the level for the quantized residual value is greater than 1. Claim 6 A method for decoding video data according to claim 1, further comprising: receiving a syntax element indicating that, for a first interval among k intervals, the level for the quantized residual value is greater than the values included in the first interval; and receiving a syntax element indicating that, for a second interval among k intervals, the level for the quantized residual value is included in the second interval; wherein the syntax element indicating the difference value indicates the difference between a reference level value for the second interval and the level for the quantized residual value of the block of video data. Claim 7 A method for decoding video data according to claim 1, further comprising the step of receiving, for individual intervals among k intervals, flags indicating that the level for the quantized residual value is greater than the values included in the individual intervals for the flags, until receiving a flag indicating that the level for the quantized residual value is within the interval associated with the flag. Claim 8 A method for decoding video data according to claim 1, wherein the k intervals include an interval in the range of 0 to a first threshold value, an interval in the range of a value obtained by adding 1 to the first threshold value to a second threshold value, and an interval in the range of a value obtained by adding 1 to the second threshold value to a third threshold value, and the method further comprises the step of determining the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter. Claim 9 A method for decoding video data according to claim 1, wherein the k intervals include an interval in the range of 1 to a first threshold value, an interval in the range of a value obtained by adding 1 to the first threshold value to a second threshold value, and an interval in the range of a value obtained by adding 1 to the second threshold value to a third threshold value, and the method further comprises the step of determining the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter. Claim 10 A method for decoding video data according to claim 1, wherein the k intervals include an interval in the range of 2 to a first threshold value, an interval in the range of a value obtained by adding 1 to the first threshold value to a second threshold value, and an interval in the range of a value obtained by adding 1 to the second threshold value to a third threshold value, and the method further comprises the step of determining the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter. Claim 11 A method for decoding video data according to claim 10, wherein the k intervals include intervals ranging from a value obtained by adding 1 to the third threshold value to the maximum possible level for the levels of the quantized residual values of the block of video data. Claim 12 In claim 1, the k intervals are the following intervals: [2, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 Includes [+1, maxTsLevel], where maxTsLevel represents the maximum possible level for the levels of the quantized residual values of the block of video data based on the quantization parameter for the block of video data, and t n A method for decoding video data, wherein represents an upper threshold for the nth interval, where n is in the range of 0 to k-1. Claim 13 A method for decoding video data according to claim 1, further comprising the step of bypass decoding a syntax element indicating the difference value. Claim 14 A method for generating a bitstream of encoded video data, comprising: determining that a block of video data is encoded without transforming residual data for said block; determining a level for a quantized residual value of said block of video data; determining a quantization parameter for said block of video data; determining a range for levels of quantized residual values of said block of video data based on the determined quantization parameter, wherein the range is from 0 to a maximum possible level for said levels of quantized residual values of said block of video data; dividing said range into k intervals, wherein k is an integer value, and each interval includes multiple levels for quantized residual values, and each interval has an associated index value between 0 and k-1; determining a specific interval among said k intervals that includes a level for said quantized residual value; and a reference level value for said specific interval and the quantized residual value of said block of video data A step of determining a difference value representing the difference between levels; and a step of signaling a level for the quantized residual value of the block of the video data based on k intervals, wherein the step of signaling a level for the quantized residual value of the block of the video data based on k intervals comprises: a step of generating one or more syntax elements representing an associated index value of the specific interval to be included in the bitstream of the encoded video data; and a step of generating a syntax element representing the difference value to be included in the bitstream of the encoded video data.A method for generating a bitstream of encoded video data, comprising the step of outputting the bitstream of the encoded video data and the step of signaling a level for the quantized residual value of the block of the video data. Claim 15 A method for generating a bitstream of encoded video data, wherein, in claim 14, the step of determining a level for a quantized residual value of a block of video data comprises: a step of determining a level for a residual value; and a step of quantizing the level for a residual value to determine a level for the quantized residual value. Claim 16 A method for generating a bitstream of encoded video data according to claim 15, further comprising the step of generating a syntax element indicating a sign for the residual value to be included in the bitstream of the encoded video data. Claim 17 A method for generating a bitstream of encoded video data, wherein, in claim 14, the first interval among the k intervals includes values in the range of 1 to 1 threshold value, and the method further comprises the step of generating a syntax element indicating that the level of the quantized residual value is greater than 0 in order to include it in the bitstream of the encoded video data. Claim 18 A method for generating a bitstream of encoded video data, wherein, in claim 14, the first interval among the k intervals includes values in the range of 2 to a first threshold, and the method further comprises the steps of: generating a syntax element indicating that the level for the quantized residual value is greater than 0 in order to include it in the bitstream of the encoded video data; and generating a syntax element indicating that the level for the quantized residual value is greater than 1 in order to include it in the bitstream of the encoded video data. Claim 19 In claim 14, the method further comprises: generating a syntax element indicating that, for a first interval among the k intervals, the level for the quantized residual value is greater than the values included in the first interval in order to be included in the bitstream of the encoded video data; generating a syntax element indicating that, for a second interval among the k intervals, the level for the quantized residual value is included in the second interval in order to be included in the bitstream of the encoded video data; wherein the syntax element indicating the difference value indicates the difference between the reference level value for the second interval and the level for the quantized residual value of the block of the video data. Claim 20 A method for generating a bitstream of encoded video data according to claim 14, further comprising the step of generating flags indicating that, for individual intervals among k intervals, the level for the quantized residual value is greater than the values included in the individual intervals for the flags, in order to include in the bitstream of the encoded video data, until generating a flag indicating that the level for the quantized residual value is within the interval associated with the flag, in order to include in the bitstream of the encoded video data. Claim 21 A method for generating a bitstream of encoded video data, wherein the k intervals include an interval in the range of 0 to a first threshold value, an interval in the range of a value obtained by adding 1 to the first threshold value to a second threshold value, and an interval in the range of a value obtained by adding 1 to the second threshold value to a third threshold value, and the method further comprises the step of determining the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter. Claim 22 In claim 14, the k intervals include intervals in the range of 1 to a first threshold value, intervals in the range of a value obtained by adding 1 to the first threshold value to a second threshold value, and intervals in the range of a value obtained by adding 1 to the second threshold value to a third threshold value, and the method further comprises the step of determining the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter, a method for generating a bitstream of encoded video data. Claim 23 In claim 14, the k intervals include intervals in the range of 2 to 1 threshold values, intervals in the range of 1 to 2 threshold values, and intervals in the range of 1 to 3 threshold values, and the method further comprises the step of determining the 1 threshold value, the 2 threshold value, and the 3 threshold value based on the quantization parameter, a method for generating a bitstream of encoded video data. Claim 24 A method for generating a bitstream of encoded video data, wherein the k intervals comprise intervals ranging from a value obtained by adding 1 to the third threshold value to the maximum possible level for the levels of the quantized residual values of the block of video data. Claim 25 In claim 14, the k intervals are the following intervals: [2, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 Includes [+1, maxTsLevel], where maxTsLevel represents the maximum possible level for the levels of the quantized residual values of the block of video data based on the quantization parameter for the block of video data, and t n A method for generating a bitstream of encoded video data, wherein represents an upper threshold for the nth interval, where n is in the range of 0 to k-1. Claim 26 A method for generating a bitstream of encoded video data, further comprising the step of bypass encoding a syntax element indicating the difference value in claim 14. Claim 27 A device for decoding video data comprises: a memory configured to store video data; and one or more processors, wherein the one or more processors determine that a block of video data is encoded without transforming residual data for said block; determine a quantization parameter for said block of video data; and, based on the determined quantization parameter, determine a range for levels of quantized residual values of said block of video data, wherein the range is from 0 to a maximum possible level for said levels of quantized residual values of said block of video data; and divide said range into k intervals, wherein k is an integer value, and each interval comprises multiple levels of quantized residual values, and each interval has an associated index value between 0 and k-1. A device for decoding video data, wherein, in order to determine a level for a quantized residual value of a block of video data based on k intervals, the one or more processors receive information indicating an index corresponding to a specific interval among the k intervals to which the level for the quantized residual value belongs; receive information indicating a difference value indicating a difference between a reference level value for the specific interval and a level for the quantized residual value of the block of video data; and further configured to determine a level for the quantized residual value of the block of video data based on the reference level value and the difference value; and configured to output decoded video data based on the level for the quantized residual value. Claim 28 A device for decoding video data according to claim 27, wherein the one or more processors determine a level for a quantized residual value by dequantizing a level for the quantized residual value; receive a syntax element indicating a sign for the quantized residual value; and further configured to determine the quantized residual value based on the level for the quantized residual value and the sign for the quantized residual value. Claim 29 A device for decoding video data, wherein, in claim 28, the one or more processors determine a residual block for a block of video data, wherein the residual block includes the dequantized residual value; determine a restored block for a block of video data by adding the residual block to a prediction block for a block of video data; generate a picture of decoded video data based on the restored block; and further configured to output the picture of decoded video data. Claim 30 A device for decoding video data according to claim 27, wherein the first interval among the k intervals comprises values in the range of 1 to 1 threshold value, and the one or more processors are further configured to receive a syntax element indicating that the level for the quantized residual value is greater than 0. Claim 31 A device for decoding video data according to claim 27, wherein the first interval among the k intervals comprises values in the range of 2 to 1 threshold values, and the one or more processors are further configured to receive a syntax element indicating that the level for the quantized residual value is greater than 0; and to receive a syntax element indicating that the level for the quantized residual value is greater than 1. Claim 32 In claim 27, the one or more processors are further configured to receive a syntax element indicating that, for a first interval among the k intervals, the level for the quantized residual value is greater than the values included in the first interval; and for a second interval among the k intervals, to receive a syntax element indicating that the level for the quantized residual value is included in the second interval; and a device for decoding video data, wherein the syntax element indicating the difference value indicates the difference between a reference level value for the second interval and a level for the quantized residual value of a block of video data. Claim 33 A device for decoding video data according to claim 27, wherein the one or more processors are further configured to receive, for individual intervals among the k intervals, flags indicating that the level for the quantized residual value is greater than the values included in the individual intervals for the flags, until receiving the flag indicating that the level for the quantized residual value is within the interval associated with the flag. Claim 34 A device for decoding video data according to claim 27, wherein the k intervals include intervals in the range of 0 to a first threshold value, intervals in the range of the first threshold value plus 1 to a second threshold value, and intervals in the range of the second threshold value plus 1 to a third threshold value, and the one or more processors are further configured to determine the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter. Claim 35 A device for decoding video data according to claim 27, wherein the k intervals include intervals in the range of 1 to a first threshold value, intervals in the range of a value obtained by adding 1 to the first threshold value to a second threshold value, and intervals in the range of a value obtained by adding 1 to the second threshold value to a third threshold value, and wherein the one or more processors are further configured to determine the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter. Claim 36 A device for decoding video data according to claim 27, wherein the k intervals include intervals in the range of 2 to 1 threshold values, intervals in the range of 1 to 2 threshold values, and intervals in the range of 1 to 3 threshold values, wherein the one or more processors are further configured to determine the 1 threshold value, the 2 threshold value, and the 3 threshold value based on the quantization parameter. Claim 37 A device for decoding video data according to claim 36, wherein the k intervals comprise intervals ranging from a value obtained by adding 1 to the third threshold value to the maximum possible level for the levels of the quantized residual values of the block of video data. Claim 38 In claim 27, the k intervals are the following intervals: [2, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 Includes [+1, maxTsLevel], where maxTsLevel represents the maximum possible level for the levels of the quantized residual values of the block of video data based on the quantization parameter for the block of video data, and t n A device for decoding video data, wherein represents an upper threshold value for the nth interval, where n is in the range of 0 to k-1. Claim 39 In claim 27, the device for decoding video data, wherein the one or more processors are additionally configured to bypass decode a syntax element indicating the difference value. Claim 40 A device for decoding video data according to claim 27, wherein the device comprises a wireless communication device, a receiver configured to receive encoded video data, and a display configured to display the decoded video data. Claim 41 In claim 40, the wireless communication device comprises a telephone handset, and the receiver is configured to demodulate a signal including the encoded video data according to a wireless communication standard, a device for decoding video data. Claim 42 A device for encoding video data comprises: a memory configured to store video data; and one or more processors, wherein the one or more processors determine that a block of video data is encoded without transforming residual data for said block; determine a level for a quantized residual value of said block of video data; determine a quantization parameter for said block of video data; and, based on the determined quantization parameter, determine a range for levels of quantized residual values of said block of video data, wherein the range is from 0 to a maximum possible level for said levels of quantized residual values of said block of video data; and divide said range into k intervals, wherein k is an integer value, and each interval includes multiple levels for quantized residual values, and each interval has an associated index value between 0 and k-1. Determining a specific interval among the k intervals including a level for the quantized residual value; determining a difference value representing the difference between a reference level value for the specific interval and a level for the quantized residual value of the block of video data; and signaling a level for the quantized residual value of the block of video data based on the k intervals, wherein, to signal a level for the quantized residual value of the block of video data based on the k intervals, the one or more processors generate one or more syntax elements representing an associated index value of the specific interval to be included in the bitstream of the encoded video data; and generate a syntax element representing the difference value to be included in the bitstream of the encoded video data;A device for encoding video data, further configured to output a bitstream of the encoded video data and configured to signal a level for the quantized residual value of a block of the video data. Claim 43 A device for encoding video data according to claim 42, wherein, in order to determine a level for a quantized residual value of a block of video data, the one or more processors determine a level for a residual value; and further configured to quantize a level for a residual value to determine a level for a quantized residual value. Claim 44 A device for encoding video data, wherein the one or more processors are further configured to generate a syntax element indicating a sign for the residual value to be included in the bitstream of the encoded video data. Claim 45 A device for encoding video data according to claim 42, wherein the first interval among the k intervals includes values in the range of 1 to 1 threshold value, and the one or more processors are further configured to generate a syntax element indicating that the level of the quantized residual value is greater than 0 in order to include it in the bitstream of the encoded video data. Claim 46 A device for encoding video data according to claim 42, wherein the first interval among the k intervals comprises values in the range of 2 to 1 threshold values, and the one or more processors are further configured to generate a syntax element indicating that the level for the quantized residual value is greater than 0 for inclusion in the bitstream of the encoded video data; and to generate a syntax element indicating that the level for the quantized residual value is greater than 1 for inclusion in the bitstream of the encoded video data. Claim 47 In claim 42, the one or more processors are further configured to generate a syntax element indicating that, for a first interval among the k intervals, the level for the quantized residual value is greater than the values included in the first interval for inclusion in the bitstream of the encoded video data; and for a second interval among the k intervals, to generate a syntax element indicating that the level for the quantized residual value is included in the second interval for inclusion in the bitstream of the encoded video data; and the syntax element indicating the difference value indicates the difference between the reference level value for the second interval and the level for the quantized residual value of the block of the video data, a device for encoding video data. Claim 48 A device for encoding video data according to claim 42, wherein the one or more processors are further configured to generate flags indicating that, for individual intervals among the k intervals, the level for the quantized residual value is greater than the values included in the individual intervals for the flags, for inclusion in the bitstream of the encoded video data, until the level for the quantized residual value is within the interval associated with the flag is generated. Claim 49 A device for encoding video data according to claim 42, wherein the k intervals include intervals in the range of 0 to a first threshold value, intervals in the range of a value 1 plus the first threshold value to a second threshold value, and intervals in the range of a value 1 plus the second threshold value to a third threshold value, and wherein the one or more processors are further configured to determine the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter. Claim 50 A device for encoding video data according to claim 42, wherein the k intervals include intervals in the range of 1 to a first threshold value, intervals in the range of a value obtained by adding 1 to the first threshold value to a second threshold value, and intervals in the range of a value obtained by adding 1 to the second threshold value to a third threshold value, and wherein the one or more processors are further configured to determine the first threshold value, the second threshold value, and the third threshold value based on the quantization parameter. Claim 51 A device for encoding video data according to claim 42, wherein the k intervals include intervals in the range of 2 to 1 threshold values, intervals in the range of 1 to 2 threshold values, and intervals in the range of 1 to 3 threshold values, wherein the one or more processors are further configured to determine the 1 threshold value, the 2 threshold value, and the 3 threshold value based on the quantization parameter. Claim 52 A device for encoding video data according to claim 51, wherein the k intervals comprise intervals ranging from a value obtained by adding 1 to the third threshold value to the maximum possible level for the levels of the quantized residual values of the block of video data. Claim 53 In Clause 42, the k intervals are the following intervals: [2, t1], [t1+1, t2], [t2+1, t3], ... [t k-2 +1, t k-1 ], [t k-1 Includes [+1, maxTsLevel], where maxTsLevel represents the maximum possible level for the levels of the quantized residual values of the block of video data based on the quantization parameter for the block of video data, and t n A device for encoding video data, wherein represents an upper threshold value for the nth interval, where n is in the range of 0 to k-1. Claim 54 In claim 42, the device for encoding video data, wherein the one or more processors are further configured to bypass encode a syntax element indicating the difference value. Claim 55 In claim 42, the device for encoding video data comprises a wireless communication device and further comprises a transmitter configured to transmit encoded video data. Claim 56 In claim 55, the wireless communication device comprises a telephone handset, and the transmitter is configured to modulate a signal including the encoded video data according to a wireless communication standard, a device for encoding video data. Claim 57 A computer-readable storage medium for storing instructions, wherein the instructions, when executed by one or more processors, cause the one or more processors to determine that a block of video data is encoded without transforming residual data for said block; to determine a quantization parameter for said block of video data; and, based on the determined quantization parameter, to determine a range for levels of quantized residual values of said block of video data, wherein the range is from 0 to a maximum possible level for said levels of quantized residual values of said block of video data; to divide said range into k intervals, wherein k is an integer value, and each interval includes multiple levels of quantized residual values, and each interval has an associated index value between 0 and k-1; and to determine a level for quantized residual values of said block of video data based on said k intervals, wherein said range of said block of video data based on said k intervals To determine a level for a quantized residual value, the instructions cause one or more processors to receive information indicating an index corresponding to a specific interval among the k intervals to which the level for the quantized residual value belongs; to receive information indicating a difference value indicating a difference between a reference level value for the specific interval and the level for the quantized residual value of the block of video data; and to determine a level for the quantized residual value of the block of video data, which determines the level for the quantized residual value based on the reference level value and the difference value;A computer-readable storage medium storing instructions that output decoded video data based on the level of the quantized residual value. Claim 58 A device for decoding video data, comprising: means for determining that a block of video data is encoded without transforming residual data for said block; means for determining a quantization parameter for said block of video data; means for determining a range for levels of quantized residual values of said block of video data based on said quantization parameter, wherein the range is from 0 to a maximum possible level for said levels of quantized residual values of said block of video data; means for dividing said range into k intervals, wherein k is an integer value, each interval includes multiple levels for quantized residual values, and each interval has an associated index value between 0 and k-1; means for determining a level for a quantized residual value of said block of video data based on said k intervals, wherein the means for determining a level for a quantized residual value of said block of video data based on said k intervals comprises, among said k intervals An apparatus for decoding video data, comprising: means for receiving information indicating an index corresponding to a specific interval to which a level for a quantized residual value belongs; means for receiving information indicating a difference value indicating a difference between a reference level value for the specific interval and a level for the quantized residual value of a block of video data; means for determining a level for the quantized residual value of a block of video data, comprising means for determining a level for the quantized residual value based on the reference level value and the difference value; and means for outputting decoded video data based on the level for the quantized residual value. Claim 59 In claim 1, the k intervals are the following intervals: [X, t1], [t1+1, t2], ... [t k-2 +1, t k-1 ], [t k-1 Includes [+1, maxTsLevel], where X is the minimum value for interval k=0, maxTsLevel represents the maximum possible level for the levels of the quantized residual values of the block of video data based on the quantization parameter for the block of video data, and t n A method for decoding video data, wherein represents an upper threshold for the nth interval, where n is in the range of 0 to k-1. Claim 60 In claim 27, the k intervals are the following intervals: [X, t1], [t1+1, t2], ... [t k-2 +1, t k-1 ], [t k-1 Includes [+1, maxTsLevel], where X is the minimum value for interval k=0, maxTsLevel represents the maximum possible level for the levels of the quantized residual values of the block of video data based on the quantization parameter for the block of video data, and t n A device for decoding video data, wherein represents an upper threshold for the nth interval, where n is in the range of 0 to k-1.
Citation Information
Patent Citations
Rice parameter initialization for coefficient level coding in video coding process
WO2015006602A2