Normalized bin reduction for coefficient decoding using thresholds and rice parameters

KR103003601B1Active Publication Date: 2026-08-11QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
KR1020217016652
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-05
Filing Date
2019-12-06
Publication Date
2026-08-11
Estimated Expiration
2039-12-06

Smart Images

  • Figure 112021062899034-PCT00031_ABST
    Figure 112021062899034-PCT00031_ABST
Patent Text Reader

Abstract

A video coder may be configured to determine a value for a zero parameter based on a rice parameter, and the value for the zero parameter identifies a coded value corresponding to a coefficient level of zero; receive a first coded value for a first coefficient of a second set of coefficients; and, based on the value for the zero parameter and the first coded value for the first coefficient, determine a level for the first coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] This application is:

[0002] U.S. provisional patent application No. 62 / 776,379 filed on December 6, 2018; and

[0003] Claiming the benefit of U.S. provisional patent application No. 62 / 787,681 filed on January 2, 2019,

[0004] Claiming priority to U.S. Patent Application No. 16 / 704,995 filed on December 5, 2019, and

[0005] The full contents of each of these are incorporated herein by reference.

[0006] Technology field

[0007] The present disclosure relates to video encoding and video decoding. Background Technology

[0008] background

[0009] Digital video capabilities can be integrated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite wireless phones, so-called "smartphones," video teleconferencing devices, video streaming devices, etc. Digital video devices implement video coding techniques such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC) standards, ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions of these standards. Video devices may also transmit, receive, encode, decode, and / or store digital video information more efficiently by implementing such video coding techniques.

[0010] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate redundancy inherent in video sequences. For block-based video coding, a video slice (i.e., a video picture or a part of a video picture) may be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction for reference samples in neighboring blocks within the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction for reference samples in neighboring blocks within the same picture, or temporal prediction for reference samples in other reference pictures. Pictures may be referred to as frames, and reference pictures may be referred to as reference frames.

[0011] summation

[0012] Video coding (e.g., video encoding and / or video decoding) typically involves predicting a block of video data from a block of video data already coded in the same picture (e.g., intra-prediction) or from a block of video data already coded in a different picture (e.g., inter-prediction). In some cases, the video encoder also calculates residual data by comparing the predicted block with the original block. Thus, residual data represents the difference between the predicted block of video data and the original block. To reduce the number of bits required to signal the residual data, the video encoder converts the residual data into transform coefficients, quantizes the transform coefficients, and signals the transformed and quantized coefficients in the encoded bitstream. The compression achieved by the transform and quantization processes may be lossy, which means that the transform and quantization processes may introduce distortion into the decoded video data. The present disclosure describes techniques related to transform coefficient coding.

[0013] A method for decoding video data comprises: determining a threshold number of normally coded bins for a first decoding pass; context decoding bins of syntax elements of a group of coefficients for a first set of coefficients until the threshold number of normally coded bins is reached, wherein the context-decoded bins of syntax elements include one or more significance flags, one or more parity level flags, and one or more first flags, each of the one or more significance flags indicates whether the absolute level for the corresponding coefficient is equal to zero, each of the one or more parity level flags indicates whether the absolute level for the corresponding coefficient is even or odd, and each of the one or more first flags indicates whether the absolute level for the corresponding coefficient is greater than 2; and determining values ​​for a first set of coefficients of a transform unit based on the context-decoded bins of syntax elements. A step of bypass decoding additional syntax elements for a second set of coefficients in response to reaching a threshold number of normally coded bins, wherein bypass decoding additional syntax elements comprises deriving a value for a Rice parameter for the coefficients of the second set of coefficients; and a step of determining values ​​for a second set of coefficients of a transformation unit based on additional syntax elements, wherein the step of determining values ​​for a second set of coefficients of a transformation unit based on additional syntax elements comprises a step of determining a value for a Zero parameter based on a Rice parameter, wherein the value for the Zero parameter identifies a coded value corresponding to a coefficient level of zero; a step of receiving a first coded value for a first coefficient among the second set of coefficients;and, based on the value for the zero parameter and the first coded value for the first coefficient, includes the step of determining the level for the first coefficient.;

[0014] A device for decoding video data comprises: a memory configured to store video data; and one or more processors implemented in a circuit, wherein the one or more processors determine a threshold number of normally coded bins for a first decoding pass; and perform context decoding for a first set of coefficients, wherein the bins of syntax elements of a coefficient group are context decoded until the threshold number of normally coded bins is reached, wherein the bins of context decoded syntax elements include one or more significance flags, one or more parity level flags, and one or more first flags, wherein each of the one or more significance flags indicates whether the absolute level for the corresponding coefficient is equal to zero, each of the one or more parity level flags indicates whether the absolute level for the corresponding coefficient is even or odd, and each of the one or more first flags indicates whether the absolute level for the corresponding coefficient is greater than 2; and determine values ​​for a first set of coefficients of a transformation unit based on the bins of context decoded syntax elements; In response to reaching a threshold number of normally coded bins, to bypass decode additional syntax elements for a second set of coefficients, one or more processors perform bypass decoding additional syntax elements configured to derive a value for a rice parameter for the coefficients of the second set of coefficients;And, configured to determine values ​​for a second set of coefficients of a transformation unit based on additional syntax elements, and to determine values ​​for a second set of coefficients of a transformation unit based on said additional syntax elements, said one or more processors determine a value for a zero parameter based on a rice parameter, wherein the value for the zero parameter identifies a coded value corresponding to a coefficient level of zero; receive a first coded value for a first coefficient among the second set of coefficients; and configured to determine a level for the first coefficient based on the value for the zero parameter and the first coded value for the first coefficient.;

[0015] According to one or more examples, a computer-readable storage medium stores instructions, and when the instructions are executed by one or more processors, the one or more processors determine a threshold number of normally coded bins for a first decoding pass; and for a first set of coefficients, context decoding bins of syntax elements of a group of coefficients until the threshold number of normally coded bins is reached, wherein the context-decoded bins of the context-decoded syntax elements include one or more significance flags, one or more parity level flags, and one or more first flags, each of the one or more significance flags indicates whether the absolute level for the corresponding coefficient is equal to zero, each of the one or more parity level flags indicates whether the absolute level for the corresponding coefficient is even or odd, and each of the one or more first flags indicates whether the absolute level for the corresponding coefficient is greater than 2; The instructions cause values ​​for a first set of coefficients of a transformation unit to be determined based on bins of context-decoded syntax elements; and, in response to reaching a threshold number of normally coded bins, to bypass decode additional syntax elements for a second set of coefficients, the instructions cause one or more processors to perform said bypass decoding, which derives values ​​for a rice parameter for the coefficients of the second set of coefficients.And, to determine values ​​for a second set of coefficients of a conversion unit based on additional syntax elements, and to determine values ​​for a second set of coefficients of a conversion unit based on said additional syntax elements, the instructions cause one or more processors to determine a value for a zero parameter based on a lie parameter, wherein the value for the zero parameter identifies a coded value corresponding to a coefficient level of zero; to receive a first coded value for a first coefficient among the second set of coefficients; and, based on the value for the zero parameter and the first coded value for the first coefficient, to determine a level for the first coefficient.

[0016] According to one example, an apparatus for decoding video data comprises: means for determining a threshold number of normally coded bins for a first decoding pass; means for context decoding bins of syntax elements of a group of coefficients for a first set of coefficients until a threshold number of normally coded bins is reached, wherein the bins of the context-decoded syntax elements comprise one or more significance flags, one or more parity level flags, and one or more first flags, each of the one or more significance flags indicates whether the absolute level for the corresponding coefficient is equal to zero, each of the one or more parity level flags indicates whether the absolute level for the corresponding coefficient is even or odd, and each of the one or more first flags indicates whether the absolute level for the corresponding coefficient is greater than 2; and means for determining values ​​for a first set of coefficients of a transformation unit based on the bins of the context-decoded syntax elements. In response to reaching a threshold number of normally coded bins, means for bypass decoding additional syntax elements for a second set of coefficients, wherein the means for bypass decoding additional syntax elements comprises means for deriving a value for a rice parameter for the coefficients of the second set of coefficients; and means for determining values ​​for a second set of coefficients of a transformation unit based on additional syntax elements, wherein the means for determining values ​​for a second set of coefficients of a transformation unit based on additional syntax elements comprises means for determining a value for a zero parameter based on a rice parameter, wherein the value for the zero parameter identifies a coded value corresponding to a coefficient level of zero; means for receiving a first coded value for a first coefficient among the second set of coefficients;and, based on a value for a zero parameter and a first coded value for a first coefficient, it includes means for determining a level for the first coefficient.;

[0017] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, purposes, and advantages will be apparent from the description, drawings, and claims. Brief explanation of the drawing

[0018] FIG. 1 is a block diagram illustrating an exemplary video encoding and decoding system capable of performing the techniques of the present disclosure. FIGS. 2a and 2b are conceptual diagrams illustrating an exemplary quadtree binary tree (QTBT) structure and a corresponding coding tree unit (CTU). Figure 3 illustrates an exemplary sequence of syntax elements representing absolute level values ​​for coefficients within a coding group (CG). Figure 4 shows an example of a template used to select probabilistic models. Figure 5 illustrates an example of interleaved Gt2 flags in the first pass after the Par flag. Figure 6 illustrates an example of an interleaved Gt2 flag in the first pass after the Gt1 flag. Figure 7 illustrates an example of partial coding of the last coefficient reaching the normally coded bin limit for SIG-Gt1-Par-Gt2 coding in the first coding pass. Figure 8 illustrates an example of partial coding of the last coefficient reaching the normally coded bin limit for SIG-Gt1-Gt2-Par coding in the first coding pass. FIG. 9 is a block diagram showing an exemplary video encoder capable of performing the techniques of the present disclosure. FIG. 10 is a block diagram showing an exemplary video decoder capable of performing the techniques of the present disclosure. Figures 11a and 11b are conceptual diagrams illustrating the range update process in binary arithmetic coding. Figure 12 is a conceptual diagram illustrating the output process in binary arithmetic coding. Figure 13 is a block diagram showing a context-adaptive binary arithmetic coding (CABAC) coder in a video encoder. Figure 14 is a block diagram showing a CABAC coder in a video decoder. Figure 15 is a flowchart illustrating the exemplary operation of a video encoder. Figure 16 is a flowchart illustrating the exemplary operation of a video decoder. Figure 17 is a flowchart illustrating the exemplary operation of a video decoder. Specific details for implementing the invention

[0019] details

[0020] Video coding (e.g., video encoding and / or video decoding) typically involves predicting a block of video data from a block of video data already coded in the same picture (e.g., intra-prediction) or from a block of video data already coded in a different picture (e.g., inter-prediction). In some cases, the video encoder also calculates residual data by comparing the predicted block with the original block. Thus, residual data represents the difference between the predicted block of video data and the original block. To reduce the number of bits required to signal the residual data, the video encoder transforms and quantizes the residual data and signals the transformed and quantized residual data in the encoded bitstream. The compression achieved by the transform and quantization processes may be lossy, which means that the transform and quantization processes may introduce distortion into the decoded video data.

[0021] A video decoder decodes residual data and adds it to a prediction block to generate a reconstructed video block that matches the original video block more closely than the prediction block alone. Due to losses introduced by the transformation and quantization of the residual data, the reconstructed block may have distortion or artifacts. One common type of artifact or distortion is referred to as blockiness, where the boundaries of the blocks used to code the video data are visible.

[0022] To further improve the quality of the decoded video, the video decoder may perform one or more filtering operations on the reconstructed video blocks. Examples of these filtering operations include deblocking filtering, sample adaptive offset (SAO) filtering, and adaptive loop filtering (ALF). Parameters for these filtering operations may be determined by the video encoder and may be explicitly signaled in the encoded video bitstream, or may be implicitly determined by the video decoder without the need for the parameters to be explicitly signaled in the encoded video bitstream.

[0023] As introduced above, the video encoder transforms residual data to generate transform coefficients. These transform coefficients may be further quantized. In this disclosure, the term transform coefficient, or coefficient, may refer to a quantized transform coefficient or a non-quantized transform coefficient. This disclosure describes techniques for signaling values ​​of transform coefficients, for example, quantized transform coefficients, from a video encoder to a video decoder. More specifically, this disclosure describes techniques related to an entropy decoding process that transforms a binary representation of bits into a series of non-binary values ​​of quantized transform coefficients. A corresponding entropy encoding process, which is generally the inverse process of entropy decoding, is also described in this disclosure.

[0024] In one example, the present disclosure describes techniques for determining Rice parameters used to define codes, e.g., Golomb-Rice codes or Exponential-Golomb codes, to code the residual absolute values ​​of coefficient levels for a block of coefficients used to code other indications of significance coefficients, such as coefficient levels greater than 1 and coefficient levels greater than 2, in context-adaptive binary arithmetic coding (CABAC). The coefficient levels may be levels of transformation coefficients in the case of loss coding, or levels of coefficients to which no transformation is applied in the case of lossless coding or loss coding in a transformation omission mode (i.e., residual pixel values). As described in more detail below, the coefficient level may be an absolute value for the coefficient level or a remaining level for the coefficient level.

[0025] The Rhys parameter is a tunable value used to select a set of codewords from a family of Colomb codes, e.g., Colomb-Rhys codes or exponent-Colomb codes. The codes defined by the Rhys parameter may be used to code the residual absolute value of a coefficient level for at least one coefficient in a transform unit (TU) or coefficient group (CG), that is, a block of coefficients. Each CG may be a 4×4 transform block or a 4×4 subblock of the transform block of video data. The CGs may include transform coefficients in the case of lossy coding, or may include coefficients to which no transform is applied in the case of lossless coding or lossy coding in a transform omission mode.

[0026] The present disclosure further describes techniques for determining the value of a zero parameter based on a lith parameter. The zero parameter represents a bitstream value corresponding to a zero counting level. When the probability that the counting level is zero is relatively low, a longer codeword or bitstream value may be assigned to the zero counting level so that shorter codewords may be used for non-zero values. The techniques of the present disclosure may improve video compression by improving the selection of zero parameters so that bits may be saved in the coding of counting levels.

[0027] The techniques of the present disclosure may be applied to any of the existing video codecs, such as High Efficiency Video Coding (HEVC), or may be promising coding tools for new video coding standards, such as Universal Video Coding (VVC), which is currently under development.

[0028] FIG. 1 is a block diagram illustrating an exemplary video encoding and decoding system capable of performing the techniques of the present disclosure. The techniques of the present disclosure generally relate to coding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Accordingly, video data may include raw, uncoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.

[0029] As illustrated in FIG. 1, the system (100) includes a source device (102) that provides encoded video data to be decoded and displayed by a destination device (116) in this example. In particular, the source device (102) provides video data to the destination device (116) via a computer-readable medium (110). The source device (102) and the destination device (116) may include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets, such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, etc. In some cases, the source device (102) and the destination device (116) may be equipped for wireless communication and thus may be referred to as wireless communication devices.

[0030] In the example of FIG. 1, the source device (102) includes a video source (104), memory (106), a video encoder (200), and an output interface (108). The destination device (116) includes an input interface (122), a video decoder (300), memory (120), and a display device (118). According to the present disclosure, the video encoder (200) of the source device (102) and the video decoder (300) of the destination device (116) may be configured to apply the coefficient coding techniques described herein. Thus, the source device (102) represents an example of a video encoding device, while the destination device (116) represents an example of a video decoding device. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device (102) may receive video data from an external video source, such as an external camera. Likewise, the destination device (116) may interface with an external display device rather than including an integrated display device.

[0031] The system (100) illustrated in FIG. 1 is merely one example. In general, any digital video encoding and / or decoding device may perform the counting coding techniques described herein. The source device (102) and the destination device (116) are merely examples of such coding devices in which the source device (102) generates coded video data for transmission to the destination device (116). The present disclosure refers to a "coding" device as a device that performs the coding (encoding and / or decoding) of data. Accordingly, the video encoder (200) and the video decoder (300) represent examples of coding devices, in particular a video encoder and a video decoder, respectively. In some examples, the source device (102) and the destination device (116) may operate in a substantially symmetric manner such that each of the source device (102) and the destination device (116) includes video encoding and decoding components. Accordingly, the system (100) may support unidirectional or bidirectional video transmission between a source device (102) and a destination device (116) for, for example, video streaming, video playback, video broadcasting, or video telephony.

[0032] Generally, a video source (104) represents a source of video data (i.e., raw, uncoded video data) and provides a sequential series of pictures of video data (also referred to as “frames”) to a video encoder (200) that encodes data for the pictures. The video source (104) of a source device (102) may include a video capture device such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As an additional alternative, the video source (104) may generate computer graphics-based data as source video, or as a combination of live video, archived video, and computer-generated video. In each case, the video encoder (200) encodes the captured, pre-captured, or computer-generated video data. The video encoder (200) may rearrange the pictures from the received order (sometimes referred to as the “display order”) to the coding order for coding. The video encoder (200) may generate a bitstream containing encoded video data. The source device (102) may output the encoded video data onto a computer-readable medium (110) through an output interface (108) for reception and / or extraction by, for example, the input interface (122) of the destination device (116).

[0033] The memory (106) of the source device (102) and the memory (120) of the destination device (116) represent general-purpose memories. In some examples, the memories (106, 120) may store raw video data, for example, raw video from a video source (104) and raw, decoded video data from a video decoder (300). Additionally or alternatively, the memories (106, 120) may store software instructions executable by, for example, the video encoder (200) and the video decoder (300), respectively. Although the memory (106) and memory (120) are depicted separately from the video encoder (200) and the video decoder (300) in this example, it should be understood that the video encoder (200) and the video decoder (300) may also include internal memories for functionally similar or equivalent purposes. Additionally, the memories (106, 120) may store encoded video data that is output from the video encoder (200) and input to the video decoder (300), for example. In some examples, parts of the memories (106, 120) may be allocated as one or more video buffers to store raw, decoded, and / or encoded video data, for example.

[0034] The computer-readable medium (110) may represent any type of medium or device capable of transmitting encoded video data from a source device (102) to a destination device (116). In one example, the computer-readable medium (110) represents a communication medium for enabling the source device (102) to transmit encoded video data directly to the destination device (116) in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, an output interface (108) may modulate a transmission signal containing the encoded video data, and an input interface (122) may demodulate the received transmission signal. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from a source device (102) to a destination device (116).

[0035] In some examples, the source device (102) may output encoded data from the output interface (108) to the storage device (112). Similarly, the destination device (116) may access the encoded data from the storage device (112) through the input interface (122). The storage device (112) may include any of various distributed or locally accessed data storage media, such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0036] In some examples, the source device (102) may output encoded video data to a file server (114) or other intermediate storage device that may store the encoded video generated by the source device (102). The destination device (116) may access the stored video data from the file server (114) via streaming or downloading. The file server (114) may be any type of server device capable of storing the encoded video data and transmitting the encoded video data to the destination device (116). The file server (114) may represent a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a NAS (network attached storage) device. The destination device (116) may access the encoded video data from the file server (114) via any standard data connection, including an internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., Digital Subscriber Line (DSL), cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on the file server (114). The file server (114) and the input interface (122) may be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.

[0037] The output interface (108) and input interface (122) may represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where the output interface (108) and input interface (122) include wireless components, the output interface (108) and input interface (122) may be configured to transmit data, such as encoded video data, according to cellular communication standards such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, 5G, etc. In some examples where the output interface (108) includes a wireless transmitter, the output interface (108) and the input interface (122) may be configured to transmit data, such as encoded video data, according to other wireless standards such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee™), the Bluetooth™ standard, etc. In some examples, the source device (102) and / or the destination device (116) may include individual system-on-chip (SoC) devices. For example, the source device (102) may include an SoC device for performing the functionality attributed to the video encoder (200) and / or the output interface (108), and the destination device (116) may include an SoC device for performing the functionality attributed to the video decoder (300) and / or the input interface (122).

[0038] The techniques of the present disclosure may be applied to video coding by supporting any of various multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, internet streaming video transmissions, such as DASH (dynamic adaptive streaming over HTTP), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0039] The input interface (122) of the destination device (116) receives an encoded video bitstream from a computer-readable medium (110) (e.g., a communication medium, a storage device (112), a file server (114), etc.). The encoded video bitstream may include signaling information defined by a video encoder (200), which is also used by a video decoder (300), such as syntax elements having values ​​that describe the processing and / or characteristics of video blocks or other coded units (e.g., slices, pictures, groups of pictures, sequences, etc.). A display device (118) displays decoded pictures of the decoded video data to a user. The display device (118) may include any of the various display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.

[0040] Although not illustrated in FIG. 1, in some examples, the video encoder (200) and the video decoder (300) may each be integrated with an audio encoder and / or an audio decoder, and may include suitable MUX-DEMUX units or other hardware and / or software to handle multiplexed streams containing both audio and video in a common data stream. Where applicable, the MUX-DEMUX units may follow the ITU H.223 multiplexer protocol, or other protocols, such as the User Datagram Protocol (UDP).

[0041] Each of the video encoder (200) and the video decoder (300) may be implemented as any of various suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Where the techniques are partially implemented in software, the device may store instructions for the software on a suitable non-transient computer-readable medium and execute those instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the video encoder (200) and the video decoder (300) may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined encoder / decoder (CODEC) in each device. A device including a video encoder (200) and / or a video decoder (300) may include an integrated circuit, a microprocessor, and / or a wireless communication device, such as a cellular phone.

[0042] The video encoder (22) and video decoder (300) may operate according to video coding standards such as ITU-T H.265, also referred to as High Efficiency Video Coding (HEVC), or their extensions such as multi-view and / or scalable video coding extensions. Alternatively, the video encoder (200) and video decoder (300) may operate according to other proprietary or industry standards such as ITU-T H.266, also referred to as JEM (Joint Exploration Test Model) or VVC (Versatile Video Coding). The latest draft of the VVC standard is Bross et al., “Versatile Video Coding (Draft 6),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 15 th Meeting: Gothenburg, SE, 3-12 July 2019, JVET-O2001-vE (hereinafter referred to as “VVC Draft 6”) is described. However, the techniques of the present disclosure are not limited to any specific coding standard.

[0043] Generally, the video encoder (200) and video decoder (300) may perform block-based coding of pictures. The term “block” generally refers to a structure containing data to be processed (e.g., to be encoded, to be decoded, or otherwise to be used in the encoding and / or decoding process). For example, a block may contain a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, the video encoder (200) and video decoder (300) may code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, rather than coding red, green, and blue (RGB) data for samples of a picture, the video encoder (200) and video decoder (300) may code luminance and chrominance components, wherein the chrominance components may include both red and blue chrominance components. In some examples, the video encoder (200) converts the received RGB formatted data into a YUV representation prior to encoding, and the video decoder (300) converts the YUV representation into an RGB format. Alternatively, pre- and post-processing units (not shown) may perform these conversions.

[0044] The present disclosure may generally refer to the coding of pictures (e.g., encoding and decoding), which includes a process of encoding or decoding data of the pictures. Similarly, the present disclosure may refer to the coding of blocks of pictures, which includes a process of encoding or decoding data for the blocks, such as, for example, prediction and / or residual coding. An encoded video bitstream generally includes a series of values ​​for syntax elements representing coding decisions (e.g., coding modes) and partitioning of the pictures into blocks. Accordingly, references to coding a picture or block should generally be understood as coding values ​​for the syntax elements forming the picture or block.

[0045] HEVC defines various blocks including coding units (CUs), prediction units (PUs), and transformation units (TUs). According to HEVC, a video coder (e.g., a video encoder (200)) partitions coding tree units (CTUs) into CUs according to a quadtree structure. That is, the video coder partitions the CTUs and CUs into four identical non-overlapping squares, and each node of the quadtree has either zero or four child nodes. Nodes without child nodes may be referred to as "leaf nodes," and the CUs of such leaf nodes may contain one or more PUs and / or one or more TUs. The video coder may further partition the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of the TUs. In HEVC, PUs represent inter-prediction data, while TUs represent residual data. Intra-predicted CUs include intra-predicted information such as intra-mode indication.

[0046] As another example, the video encoder (200) and video decoder (300) may be configured to operate according to JEM or VVC. According to JEM or VVC, the video encoder (e.g., video encoder (200)) partitions the picture into multiple coding tree units (CTUs). The video encoder (200) may partition the CTUs according to a tree structure such as a quadtree-binary tree (QTBT) structure or a Multi-Type Tree (MTT) structure. The QTBT structure eliminates the concepts of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level partitioned according to quadtree partitioning, and a second level partitioned according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary trees correspond to the coding units (CUs).

[0047] In the MTT partitioning structure, blocks may be partitioned using quadtree (QT) partitions, binary tree (BT) partitions, and one or more types of tripletree (TT) (also called ternary tree (TT)) partitions. A triple or ternary tree partition is a partition in which a block is split into three sub-blocks. In some examples, a triple or ternary tree partition splits the block into three sub-blocks without splitting the original block through the center. The partitioning types in MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.

[0048] In some examples, the video encoder (200) and the video decoder (300) may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video encoder (200) and the video decoder (300) may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for both chrominance components (or two QTBT / MTT structures for each chrominance component).

[0049] The video encoder (200) and video decoder (300) may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures per HEVC. For the purposes of explanation, the description of the techniques of the present disclosure is presented with respect to QTBT partitioning. However, it should be understood that the techniques of the present disclosure may also be applied to video coders configured to use quadtree partitioning or other types of partitioning.

[0050] Blocks (e.g., CTUs or CUs) may be grouped in various ways within a picture. As an example, a brick may refer to a rectangular area of ​​rows of CTUs within a specific tile in the picture. A tile may refer to a rectangular area of ​​CTUs within a specific tile row and a specific tile column in the picture. A tile column refers to a rectangular area of ​​CTUs having a width specified by syntax elements (e.g., in a picture parameter set) and a height equal to the height of the picture. A tile row refers to a rectangular area of ​​CTUs having a width equal to the width of the picture and a height specified by syntax elements (e.g., in a picture parameter set).

[0051] In some examples, a tile may be partitioned into multiple bricks, each of which may contain one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick. However, a brick, which is a true subset of a tile, may not be referred to as a tile.

[0052] Bricks in a picture may also be arranged into slices. A slice may be an integer number of bricks in a picture that may be exclusively contained in a single Network Abstraction Layer (NAL) unit. In some examples, a slice includes either multiple complete tiles or a sequence of complete bricks consisting of only one tile.

[0053] The present disclosure may interchangeably use "NxN" and "N by N," e.g., 16x16 samples or 16 by 16 samples, to refer to the sample dimensions of a block (e.g., a CU or other video block) in terms of vertical and horizontal dimensions. Generally, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU may be arranged in rows and columns. Furthermore, CUs do not necessarily have the same number of samples in the horizontal direction as in the vertical direction. For example, CUs may contain NxM samples, where M is not necessarily the same as N.

[0054] A video encoder (200) encodes video data for CUs and other information representing prediction and / or residual information. The prediction information indicates how the CU will be predicted to form a prediction block for the CU. The residual information generally represents sample-by-sample differences between the samples of the CU prior to encoding and the prediction block.

[0055] To predict the CU, the video encoder (200) may generally form a prediction block for the CU through inter-prediction or intra-prediction. Inter-prediction generally refers to predicting the CU from data of a previously coded picture, whereas intra-prediction generally refers to predicting the CU from data of the same picture previously coded. To perform inter-prediction, the video encoder (200) may generate a prediction block using one or more motion vectors. The video encoder (200) may generally perform motion search to identify a reference block that closely matches the CU in terms of the differences between the CU and the reference block. The video encoder (200) may calculate a difference metric using the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared differences (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, the video encoder (200) may predict the current CU using unidirectional prediction or bidirectional prediction.

[0056] Some examples of JEM and VVC also provide an affine motion compensation mode, which may be considered as an inter-prediction mode. In the affine motion compensation mode, the video encoder (200) may determine two or more motion vectors representing non-translational motion, such as zoom in or out, rotation, perspective motion, or other irregular motion types.

[0057] To perform intra prediction, the video encoder (200) may select an intra prediction mode to generate prediction blocks. Some examples of JEM and VVC provide 67 intra prediction modes, including planar mode and DC mode, as well as various directional modes. Generally, the video encoder (200) selects an intra prediction mode that describes neighboring samples for the current block (e.g., a block of CU) to predict samples of the current block. Such samples may generally be located above, above and to the left of, or to the left of the current block in the same picture as the current block, assuming that the video encoder (200) codes the CTUs and CUs in raster scan order (from left to right, from top to bottom).

[0058] The video encoder (200) encodes data indicating the prediction mode for the current block. For example, for inter-prediction modes, the video encoder (200) may encode motion information for the corresponding mode as well as data indicating which of the various available inter-prediction modes is used. For unidirectional or bidirectional inter-prediction, for example, the video encoder (200) may encode motion vectors using Advanced Motion Vector Prediction (AMVP) or a merge mode. The video encoder (200) may also encode motion vectors for an affine motion compensation mode using similar modes.

[0059] Following a prediction such as an intra prediction or inter prediction of a block, the video encoder (200) may calculate residual data for the block. Residual data, such as a residual block, represents sample-by-sample differences between the block and the prediction block for the block, formed using the corresponding prediction mode. The video encoder (200) may apply one or more transformations to the residual block to generate transformed data in a transformation domain instead of a sample domain. For example, the video encoder (200) may apply a discrete cosine (DCT), integer transformation, wavelet transformation, or conceptually similar transformation to the residual video data. Additionally, the video encoder (200) may apply a secondary transformation following the first transformation, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transformation, or a Karhunen-Loeve transform (KLT). The video encoder (200) generates transformation coefficients following the application of one or more transformations.

[0060] As mentioned above, following any transformations to generate transformation coefficients, the video encoder (200) may perform quantization of the transformation coefficients. Quantization generally refers to a process in which transformation coefficients are quantized to provide additional compression, such that the amount of data used to represent those transformation coefficients is reduced as much as possible. By performing the quantization process, the video encoder (200) may reduce the bit depth associated with some or all of the transformation coefficients. For example, the video encoder (200) may round down n-bit values ​​to m-bit values ​​during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder (200) may perform a bitwise right-shift of the values ​​to be quantized.

[0061] Following quantization, the video encoder (200) may scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix containing the quantized transform coefficients. The scan may be designed to place higher energy (and therefore lower frequency) transform coefficients at the front of the vector and lower energy (and therefore higher frequency) transform coefficients at the rear of the vector. In some examples, the video encoder (200) may generate a serialized vector by utilizing a predefined scan order to scan the quantized transform coefficients, and then entropy-encode the quantized transform coefficients of the vector. In other examples, the video encoder (200) may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder (200) may entropy-encode the one-dimensional vector according to, for example, context-adaptive binary arithmetic coding (CABAC). The video encoder (200) may also entropy-encode values ​​for syntax elements describing metadata associated with the encoded video data for use by the video decoder (300) in decoding the video data.

[0062] To perform CABAC, the video encoder (200) may assign a context within a context model to the symbol to be transmitted. The context may, for example, be related to whether the neighboring values ​​of the symbol are zero values. The probability determination may be based on the context assigned to the symbol.

[0063] The video encoder (200) may additionally generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, to the video decoder (300), for example, from a picture header, block header, slice header, or other syntax data, such as a sequence parameter set (SPS), picture parameter set (PPS), or video parameter set (VPS). The video decoder (300) may likewise decode such syntax data to determine how to decode the corresponding video data.

[0064] In this way, the video encoder (200) may generate a bitstream containing syntax elements describing the partitioning of the encoded video data, for example, into blocks of a picture (e.g., CUs), and prediction and / or residual information for the blocks. Ultimately, the video decoder (300) may receive the bitstream and decode the encoded video data.

[0065] Generally, the video decoder (300) performs a process opposite to that performed by the video encoder (200) to decode the encoded video data of the bitstream. For example, the video decoder (300) may decode values ​​for the syntax elements of the bitstream using CABAC in a manner substantially similar to, but opposite to, the CABAC encoding process of the video encoder (200). The syntax elements may define the CUs of the CTU by defining partitioning information for the picture's CTUs and partitioning of each CTU according to a corresponding partition structure, such as a QTBT structure. The syntax elements may additionally define prediction and residual information for blocks of video data (e.g., CUs).

[0066] Residual information may be represented, for example, by quantized transform coefficients. The video decoder (300) may inversely quantize and inversely transform the quantized transform coefficients of the block to reconstruct a residual block for the block. The video decoder (300) forms a prediction block for the block using the signaled prediction mode (intra or inter prediction) and the associated prediction information (e.g., motion information for inter prediction). The video decoder (300) may then reconstruct the original block by combining the prediction block and the residual block (on a per-sample basis). The video decoder (300) may perform additional processing, such as performing a deblocking process to reduce visual artifacts along the boundaries of the block.

[0067] The present disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to the communication of values ​​for syntax elements and / or other data used to decode encoded video data. That is, the video encoder (200) may signal values ​​for syntax elements in a bitstream. Generally, signaling refers to generating values ​​in a bitstream. As mentioned above, the source device (102) may transmit the bitstream to the destination device (116) non-real-time or substantially real-time, such as when storing syntax elements in the storage device (112) for subsequent retrieval by the destination device (116).

[0068] FIGS. 2a and 2b are conceptual diagrams illustrating an exemplary quadtree binary tree (QTBT) structure (130) and a corresponding coding tree unit (CTU) (132). Solid lines represent quadtree splitting, and dotted lines represent binary tree splitting. At each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which splitting type (i.e., horizontal or vertical) is used, where 0 indicates horizontal splitting and 1 indicates vertical splitting. For quadtree splitting, it is not necessary to indicate the splitting type, as quadtree nodes split blocks horizontally and vertically into four sub-blocks of the same size. Accordingly, the video encoder (200) may encode the syntax elements (e.g., splitting information) for the region tree levels (i.e., solid lines) of the QTBT structure (130) and the syntax elements (e.g., splitting information) for the prediction tree levels (i.e., dotted lines) of the QTBT structure (130), and the video decoder (300) may decode them. For the CUs represented by the terminal leaf nodes of the QTBT structure (130), the video encoder (200) may encode video data such as prediction and transformation data, and the video decoder (300) may decode them.

[0069] Generally, the CTU (132) of FIG. 2b may be associated with parameters that define the sizes of blocks corresponding to the nodes of the QTBT structure (130) at the first and second levels. These parameters may include CTU size (indicating the size of the CTU (132) in the samples), minimum quadtree size (MinQTSize, indicating the minimum allowed quadtree leaf node size), maximum binary tree size (MaxBTSize, indicating the maximum allowed binary tree root node size), maximum binary tree depth (MaxBTDepth, indicating the maximum allowed binary tree depth), and minimum binary tree size (MinBTSize, indicating the minimum allowed binary tree leaf node size).

[0070] The root node of the QTBT structure corresponding to the CTU may have four child nodes at the first level of the QTBT structure, each of which may be partitioned according to quadtree partitioning. That is, the nodes at the first level are either leaf nodes (no child nodes) or have four child nodes. An example of the QTBT structure (130) represents such nodes as including child nodes and parent nodes with solid lines for branches. If the nodes at the first level are not larger than the maximum allowed binary tree root node size (MaxBTSize), the nodes may be further partitioned by individual binary trees. Binary tree splitting of a node may be repeated until the nodes resulting from the split reach the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). An example of a QTBT structure (130) represents such nodes as having dotted lines for branches. Binary tree leaf nodes are referred to as coding units (CUs) used for prediction (e.g., intra-picture or inter-picture prediction) and transformation without any additional partitioning. As discussed above, CUs may also be referred to as "video blocks" or "blocks".

[0071] In one example of a QTBT partitioning structure, the CTU size is set to 128x128 (luminance samples and two corresponding 64x64 chroma samples), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. Quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes ranging from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If a leaf quadtree node is 128x128, it will not be further split by the binary tree because its size exceeds MaxBTSize (i.e., 64x64 in this example). Otherwise, the leaf quadtree nodes will be further partitioned by the binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further splitting is not allowed. When a binary tree node has a width equal to MinBTSize (4 in this example), it implies that further horizontal splitting is not allowed. Similarly, a binary tree node with a height equal to MinBTSize implies that further vertical splitting is not allowed for that binary tree node. As mentioned above, the leaf nodes of the binary tree are referred to as CUs and are further processed according to prediction and transformation without further partitioning.

[0072] TCQ (Trellis coded quantization) was proposed in H. Schwarz, T. Nguyen, D. Marpe, T. Wiegand, M. Karczewicz, M. Coban, J. Dong, “CE7: Transform coefficient coding with reduced number of regular-coded bins (tests 7.1.3a, 7.1.3b)”, JVET document JVET-L0274, Macao, CN, Oct 2018 (hereinafter, JVET-L0274). In the techniques of JVET-L0274, two scalar quantizers are used switchably for quantization / de-quantization. The scalar quantizer used for the current transformed / quantized coefficient is determined by the parity (lesser bit) of the quantized coefficient preceding the current transformed / quantized coefficient in the scanning order.

[0073] A coefficient coding scheme combined with TCQ was also proposed in JVET-L0274, wherein the context selection for decoding the quantized coefficients depends on the quantizer used. Specifically, the significance flag (SIG) of a coefficient, which indicates whether the coefficient is zero or non-zero, has three sets of context models, and the set selected for a specific SIG depends on the quantizer used for the associated coefficient. Thus, when starting to decode the SIG of the current coefficient, the entropy decoder must know the parity of the coefficient at the previous scanning position, which determines the quantizer for the current coefficient and, therefore, the context set for the SIG of that coefficient.

[0074] TU is divided into non-overlapping subblocks called coding groups (CG), and their size is typically 4x4. The decoding process described herein may sometimes be described for 4x4 CGs, but can be easily extended to any other CG sizes. The techniques of this disclosure, and thus the descriptions contained herein, relate primarily to encoding and decoding processes for absolute levels of coefficients in CGs. Other information associated with CGs, such as symbols, may be encoded or decoded in the manner described in JVET-L0274, but may also be encoded and decoded using alternative techniques.

[0075] The video encoder (200) and video decoder (300) may be configured to process syntax elements in bitstreams. For example, the following syntax elements may be used to represent the absolute level value (absLevel) for a coefficient.

[0076] · sig_coeff_flag: If absLevel is 0, this flag is equal to 0; otherwise, the flag is equal to 1.

[0077] · abs_level_gt1_flag: If sig_coeff_flag is equal to 1, the flag exists in the bitstream. If absLevel is greater than 1, it is equal to 1; otherwise, the flag is equal to 0.

[0078] · par_level_flag: If rem_abs_gt1_flag is equal to 1, the flag exists in the bitstream. If absLevel is odd, it is equal to 0, and if absLevel is even, it is equal to 1.

[0079] · abs_level_gt3_flag: If abs_level_gt1_flag is equal to 1, the flag exists in the bitstream. If absLevel is greater than 3, it is equal to 1; otherwise, the flag is equal to 0.

[0080] · abs_remainder: If abs_level_gt3_flag is equal to 1, this syntax element exists in the bitstream. It is the absolute residual value of the transform factor level coded in Golomb-Rice code.

[0081] · abs_level: This is the absolute value of the transformation coefficient level coded in Golomb-Rice.

[0082] For the sake of brevity, the syntax elements sig_coeff_flag, par_level_flag, abs_level_gt1_flag, abs_level_gt3_flag, abs_remainder, and abs_level are denoted as SIG, Par, Gt1, Gt2, remLevel, and absLevel, respectively.

[0083] The video encoder (200) and the video decoder (300) may be configured to set any of the syntax elements that are not parsed from the bitstream to a default value such as 0. Given the values ​​of the first syntax element among the five syntax elements, the value for the absolute level of the coefficient can be calculated as follows:

[0084] absoluteLevel = SIG + Gt1 + Par + (Gt2 << 1) + (remLevel << 1) (1)

[0085] Alternatively, if the coefficient is coded in a completely bypass-coded mode, absoluteLevel can be coded directly as abs_level.

[0086] Figure 3 shows an exemplary order of syntax elements representing absoluteLevels in the CG as in JVET-L0274. Other orders may also be used. As can be seen, all 5 syntax elements are parsed from the bitstream when absLevel is greater than 4.

[0087] In the example of FIG. 3, the video decoder (300) scans positions in the CG in up to four passes. In the first pass (136), the video decoder (300) parses the values ​​for SIGs, Pars, and Gt1s. Only non-zero SIGs are followed by corresponding Gt1s and Pars. That is, if the video decoder (300) determines that a SIG has a value of zero, this means that the coefficient level is equal to zero, and the video decoder (300) does not receive instances of Gt1 but receives par for that coefficient. After the first pass (136), the value for the partial absoluteLevel, denoted as absLevel1 for each position, may be reconstructed as shown in Equation (2).

[0088] absLevel1 = SIG + Par + Gt1 (2)

[0089] In some implementations, the video decoder (300) may be configured to parse up to 28 regular coded bins in the first pass (136) for 4x4 subblocks and up to 6 regular coded bins for 2x2 subblocks. Limits on the number of regular coded bins may be enforced in groups of SIG, Gt1, and Par bins, which means that each group of SIG, Gt1, and Par bins is coded as a set and switching to bypass coding in the middle of the set is not allowed.

[0090] If at least one non-zero Gt1 exists in the first pass, the video decoder (300) may be configured to scan the second pass (138). In the second pass (138), the video decoder (300) parses Gt2s for positions having non-zero Gt1s. The bins in the first pass (136) and the second pass (138) may all be normalized, which means that the probability distribution of the bins is modeled by a selected context model. If at least one non-zero Gt2 exists in the second pass (138), the video decoder (300) scans the third pass (140). During the third pass (140), the video decoder (300) parses the remLevels of positions having non-zero Gt2s. remLevel is not binary, and the video decoder (300) may bypass-code binaries of rem, which means that the binaries are assumed to be uniformly distributed and no context selection is required.

[0091] In the fourth pass (142), the video decoder (300) scans all residual coefficients that are not partially represented by the normally coded bins in the previous three passes. The coefficient levels of the additional pass (142) are coded as absolute values ​​using bypass coded bins.

[0092] The video encoder (200) and video decoder (300) may perform context modeling. The context modeling used in JVET-L0274 is also briefly introduced herein along with variations proposed by this disclosure. The context modeling discussed in more detail below generally refers to the selection of probabilistic models, also referred to as contexts, for bin-to-decode. In JVET-L0274, the syntax elements SIG, Par, Gt1, and Gt2 are coded using context modeling. The selection of contexts depends on the values ​​of absLevel1 in the local neighborhood, denoted by N. FIG. 4 shows the template of the neighborhood used. Positions within the template, but outside the current TU, may be excluded from N.

[0093] Figure 4 shows an example of a template used to select probabilistic models. The square marked "X" specifies the current scan position, and the square marked "Y" indicates the local neighborhood used.

[0094] For the current position (refer to the square with X in FIG. 4), the video decoder (300) determines the context indices of its SIG, Par, Gt1, and Gt2, denoted as ctxIdxSIG, ctxIdxPar, ctxIdxGt1, and ctxIdxGt2. To determine the context indices, the video decoder (300) may first determine three variables - numSIG, sumAbs1, and d. The variable numSIG represents the number of non-zero SIGs in N, which is expressed by Equation (3) below.

[0095] (3)

[0096] The variable sumAbs1 represents the sum of absLevel1 in N, which is expressed by Equation 4 below.

[0097] (4)

[0098] The variable d represents the diagonal measurement of the current position inside the TU, as expressed by the equation (5) below:

[0099] d = x + y (5)

[0100] Here, x and y represent the coordinates of the current position within the TU.

[0101] Given sumAbs1 and d, the video decoder (300) determines the context index for decoding SIG as follows:

[0102] For Luma, ctxIdxSIG is determined by Equation (6):

[0103] ctxIdxSIG = 18*max(0, state-1) + min( sumAbs1, 5) + (d < 2 ? 12 : ( d < 5 ? 6 : 0 )) (6)

[0104] For chroma, ctxIdxSIG is determined by Equation (7):

[0105] ctxIdxSIG = 12 * max(0, state-1) + min( sumAbs1, 5 ) + ( d < 2 ? 6 : 0 ) ) (7)

[0106] In equations (6) and (7), the variable "state" represents the current state of the state machine as defined in JVET-L0274.

[0107] Given sumSIG, sumAbs1 and d, the video decoder (300) determines the context index for decoding Par as follows:

[0108] If the current scan position is the same as the position of the last non-zero coefficient, ctxIdxPar is 0.

[0109] otherwise

[0110] o For Luma, ctxIdxPar is determined by Equation (8):

[0111] ctxIdxPar = 1 + min( sumAbs1 - numSIG, 4 ) + ( d == 0 ? 15 : ( d < 3 ? 10 : ( d < 10 ? 5 : 0 ) ) ) (8)

[0112] For chroma, ctxIdxPar is determined by (9).

[0113] ctxIdxPar = 1 + min( sumAbs1 - numSIG, 4 ) + ( d == 0 ? 5 : 0 ) (9)

[0114] ctxIdxGt1 and ctxIdxGt2 are set to the values ​​of ctxIdxPar.

[0115] The video encoder (200) and the video decoder (300) may be configured to perform RemLevel coding. The video decoder (300) derives a price parameter (pricePar) for coding the non-binary syntax elements remRemainder (remLevel) and absLevel as follows:

[0116] At the beginning of each subblock, ricePar is set to 0;

[0117] After coding the rest of the syntax elements, the rice parameter (pricePar) is modified as follows:

[0118] If ricePar is less than 3 and the final coded value of the remainder is greater than ((3 < ricePar) - 1), ricePar is incremented by 1.

[0119] To code the non-binary syntax element absLevel representing absolute quantization indices that are completely bypass-coded, the following applies:

[0120] The sum of absolute values ​​within the local template, sumAbs, is determined.

[0121] The variables ricePar and posZero are,

[0122] ricePar = riceParTable[ min( 31, sumAbs ) ]

[0123] posZero = posZeroTable[ max( 0, state - 1 ) ][ min( 31, sumAbs ) ]

[0124] Determined by a table lookup according to, where, the variable state represents the state for dependent quantization (it is equal to 0 when dependent quantization is disabled),

[0125] riceParTable[] and posZeroTable[]] are,

[0126] riceParTable

[32] = {

[0127] 0,0,0,0,0,0,0,1,1,1,1,1,1,1,2,2,2,2,2,2,2,2,2,2,2,2,2,2,3,3,3,3

[0128] };

[0129] posZeroTable[3]

[32] = {

[0130] {0,0,0,0,0,1,2,2,2,2,2,2,2,4,4,4,4,4,4,4,4,4,4,4,4,8,8,8,8,8,8,8,8,8},

[0131] {1,1,1,1,2,3,4,4,4,6,6,6,8,8,8,8,8,8,12,12,12,12,12,12,12,12,16,16,16,16,16,16}, {1,1,2,2,2,3,4,4,4,6,6,6,8,8,8,8,8,8,12,12,12,12,12,12,12,16,16,16,16,16,16,16}

[0132] };

[0133] It is given by.

[0134] The intermediate variable codeValue is derived as follows:

[0135] If absLevel is equal to 0, codeValue is set to be equal to posZero;

[0136] o Otherwise, if absLevel is less than or equal to posZero, codeValue is set to absLevel - 1;

[0137] o Otherwise (absLevel is greater than posZero), codeValue is set to be the same as absLevel.

[0138] The value of codeValue is coded using the Golom-Rice code with the rice parameter ricePar.

[0139] The video encoder (200) and video decoder (300) may be configured to perform absolute-level reconstruction. The absolute-level reconstruction may be the same as in JVET-L0274, which was discussed above regarding syntax elements in the bitstream.

[0140] The video encoder (200) and video decoder (300) may be configured to code the Gt2 flags in an interleaved manner. In some examples, instead of the described method in which the SIG, Gt1, and Par flags are coded in the first pass and the Gt2 flags are coded in the second pass, the Gt2 flags may be incorporated into the first pass after the Par flag or after the Gt1 flag, as shown in the figures below, which reduces the coding passes from 4 to 3.

[0141] FIG. 5 shows an example of an interleaved Gt2 flag in the first pass after the Par flag. With respect to FIG. 5, the video decoder (300) may determine the value for absLevel1 in the same manner as described above with respect to FIG. 3, but the order in which the various syntax elements are received is changed. For example, in FIG. 5, the video decoder (300) determines the values ​​for Gt2 as part of the first pass (162) instead of part of the second pass (e.g., the second pass (138) in FIG. 3). Accordingly, in FIG. 5, the first pass (136) and the second pass (138) of FIG. 3 are effectively combined into a single pass (the first pass (162)), and the third pass (140) and the fourth pass (142) of FIG. 3 become the second pass (164) and the third pass (166) of FIG. 5, respectively. Thus, in the example of FIG. 5, only three passes are required to transmit all syntax elements.

[0142] Figure 6 shows an example of interleaved Gt2 flags in the first pass after the Gt1 flag. In this case, absLevel1 can be calculated as follows:

[0143] absLevel1 = SIG + Par + Gt1 + (Gt2<<1)

[0144] In relation to context modeling, the equations introduced above can be used to derive the context. In relation to FIG. 6, the video decoder (300) may determine the value for absLevel1 in the same manner as described above in relation to FIG. 3, but the order in which the various syntax elements are received is changed. For example, in FIG. 6, the video decoder (300) determines the values ​​for Gt2 as part of the first pass (172) instead of part of the second pass (e.g., the second pass (138) in FIG. 3). Accordingly, in FIG. 6, the first pass (136) and the second pass (138) of FIG. 3 are effectively combined into a single pass (first pass (172)), and the third pass (140) and the fourth pass (142) of FIG. 3 become the second pass (174) and the third pass (176) of FIG. 6, respectively. Thus, in the example of FIG. 6, only three passes are required to transmit all syntax elements. In FIG. 6, the syntax elements of the first pass (172) are scanned in a different order than the syntax elements of the first pass (162) of FIG. 5, but the other passes are generally the same.

[0145] The video encoder (200) and video decoder (300) may be configured to use a partial final normal bin coded coefficient representation, wherein the values ​​for some coefficients may be partially transmitted using normal bins along with residual values ​​transmitted using bypass coding. In the coding scheme described in JVET-L0274, the last normal bin coded coefficient (e.g., Coeff K in FIG. 3), SIG, Gt1, and Par bins, to which the normal bin budget for the first coding pass is reached, are all coded as normal bins. Normal bin coding does not terminate in the middle of the SIG-Gt1-Par group. Similarly, for the SIG-Gt1-Par-Gt2 group or the SIG-Gt1-Gt2-Par group (e.g., FIG. 5 and 6), the coding for the SIG, Gt1, Par, and Gt2 flags of Coeff K is coded in normal mode. The present disclosure proposes techniques to break these constraints by allowing possible termination of normally coded beans after coding of SIG and Gt1 flags, as illustrated in FIGS. 7 and 8.

[0146] FIG. 7 illustrates an example of partial coding of the last coefficient where the normally coded bin limit for SIG-Gt1-Par-Gt2 coding is reached in the first coding pass (182). In the example of FIG. 7, the video decoder (300) scans a third pass (186) containing both remLevel values ​​and absLevel values. The value for remLevel represents the residual value between the actual value for the coefficient and the partial value determined from the first pass (182) and the second pass (184). The value for absLevel, conversely, represents the absolute value of the coefficient value.

[0147] FIG. 8 shows an example of partial coding of the last coefficient where the normally coded bin limit for SIG-Gt1-Gt2-Par coding is reached in the first coding pass (192). In FIG. 8, the syntax elements of the first pass (192) are scanned in a different order than the syntax elements of the first pass (182) of FIG. 7. The second pass (194) and the third pass (196) are generally the same as the second pass (184) and the third pass (186) of FIG. 7.

[0148] In the examples of FIGS. 7 and 8, the residual level of Coeff K is coded as remLevelFull, which is bypass-coded in the third pass (186 / 196), along with the values ​​for absLevel, which are bypass-coded. The values ​​for the coefficients are expressed as follows:

[0149] absoluteLevel = SIG + Gt1 + remLevelFull,

[0150] or

[0151] absoluteLevel = SIG + remLevelFull.

[0152] In other examples, the regular coding of the beans may end after the coding of the Par and Gt2 flags, or vice versa. In this case, the residual level of the last coefficient will be coded as half of the residual level, i.e.,

[0153] absoluteLevel = SIG +GT1 + Par + (remLevel << 1),

[0154] or

[0155] absoluteLevel = SIG +GT1 + (GT2 << 1) + (remLevel << 1).

[0156] The total number of normally coded beans may also be defined as the total number imposed on the interleaved SIG, Gt1, Gt2, and Par flags.

[0157] The video encoder (200) and the video decoder (300) may be configured to perform remLevel coding. The remLevel coding in the second coding pass may be the same as that described above for RemLevel coding. The video decoder (300) may perform rice parameter updates and derivations up to the end of Coeff K-1, where Coeff K-1 represents the second to last normally coded coefficients prior to the last normally coded coefficient (Coeff K). The video decoder (300) may decode Coeff K-1 using fully normal coding, or may decode Coeff K using fully normal coding or a combination of normal coding and bypass coding. For the remLevelFull coding of Coeff K, the video decoder (300) may update the rice parameters as follows:

[0158] riceParBypass = 2 x ricePar + lastCodedGt2Flag,

[0159] riceParBypass = riceParBypass == 1 ? riceParBypass - 1 : riceParBypass

[0160] Here, ricePar is the ricePar used for coding remLevel in the second pass, and lastCodedGt2Flag is the value of the last coded Gt2 flag in the first coding pass. Alternatively, a value for ricePar that is 2 x ricePar may be used, or a ricePar that matches the optimal coding of the residual level for Coeff K may be used.

[0161] In some examples, for coding Coeff K's remLevelFull, the video decoder (300) may update the rice parameter as follows:

[0162] 1- riceParBypass = min(ricePar>0 ? ricePar + 1 : lastCodedGt2Flag, 3)

[0163] 2- riceParBypass = min(2*ricePar + lastCodedGt2Flag, 3)

[0164] 3- riceParBypass = min(2*ricePar, 3)

[0165] For the remainder of the absLevel values ​​for the coefficients that are fully coded using bypass coding, the video decoder (300) may update riceParBypass as follows. Before coding the bypass-coded coefficients, the video decoder (300) updates riceParBypass as follows:

[0166] if (riceParBypass < 3 && absoluteLevelPrevCoeff > ((3< <riceParBypass)-1) { riceParBypass++;}

[0167] It is similar to how ricePar is updated for remLevel coding, except that the full absolute value of the previously coded coefficient (Coeff K) is used for threshold checking instead of remLevel.

[0168] The video decoder (300) may derive a posZero parameter to determine the absLevel level of any various different techniques. In one example, the video decoder (300) may derive a posZero parameter to determine the absLevel level using the following lookup table:

[0169] posZero = posZeroTableBypass[ max( 0, state - 1 ) ][ riceParBypass ]

[0170] posZeroTableBypass [3][4]={{ 1, 2, 4, 8},{ 3, 6, 12, 16},{ 4, 6, 12, 16}};

[0171] The video decoder (300) may also derive an intermediate variable codeValue to be coded as follows:

[0172] If absLevel or remLevelFull is equal to 0, codeValue is set to equal posZero;

[0173] o Otherwise, if absLevel or remLevelFull is less than or equal to posZero, codeValue is set to absLevel - 1 or remLevelFull - 1, respectively.

[0174] o Otherwise (if absLevel or remLevelFull is greater than posZero), codeValue is set to be equal to absLevel or remLevelFull, respectively.

[0175] The video decoder (300) may also code the value of codeValue using a Golom-Rice code with the rice parameter riceParBypass.

[0176] FIG. 9 is a block diagram illustrating an exemplary video encoder (200) capable of performing the techniques of the present disclosure. FIG. 9 is provided for illustrative purposes only and should not be construed as limiting the techniques as broadly illustrated and described in the present disclosure. For illustrative purposes, the present disclosure describes the video encoder (200) in the context of video coding standards such as the HEVC video coding standard and the H.266 video coding standard under development. However, the techniques of the present disclosure are not limited to these video coding standards and are generally applicable to video encoding and decoding.

[0177] In the example of FIG. 9, the video encoder (200) includes a video data memory (230), a mode selection unit (202), a residual generation unit (204), a transformation processing unit (206), a quantization unit (208), an inverse quantization unit (210), an inverse transformation processing unit (212), a reconstruction unit (214), a filter unit (216), a decoded picture buffer (DPB) (218), and an entropy encoding unit (220).

[0178] The video data memory (230) may store video data to be encoded by components of the video encoder (200). The video encoder (200) may receive video data stored in the video data memory (230) from, for example, a video source (104) (Fig. 1). The DPB (218) may act as a reference picture memory to store reference video data for use in predicting subsequent video data by the video encoder (200). The video data memory (230) and the DPB (218) may be formed by any of various memory devices, such as dynamic random access memory (DRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices, including synchronous dynamic random access memory (SDRAM). The video data memory (230) and the DPB (218) may be provided by the same memory device or by separate memory devices. In various examples, the video data memory (230) may be on-chip with other components of the video encoder (200) as illustrated, or off-chip with respect to those components.

[0179] In the present disclosure, a reference to the video data memory (230) should not be interpreted as being limited to memory inside the video encoder (200) unless specifically stated otherwise, or to memory outside the video encoder (200) unless specifically stated otherwise. Rather, a reference to the video data memory (230) should be understood as a reference memory that stores video data received by the video encoder (200) for encoding (e.g., video data for the current block to be encoded). The memory (106) of FIG. 1 may also provide temporary storage of outputs from various units of the video encoder (200).

[0180] Various units of FIG. 9 are illustrated to aid in understanding the operations performed by the video encoder (200). These units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functionality and are pre-configured for the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in the operations that can be performed. For example, programmable circuits may execute software or firmware that causes the programmable circuits to operate in a manner defined by the commands of the software or firmware. Fixed-function circuits may execute software commands (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuits are generally invariant. In some examples, one or more of the units may be separate circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.

[0181] The video encoder (200) may include arithmetic logic units (ALUs), elementary function units (EFUs), digital circuits, analog circuits, and / or programmable cores formed from programmable circuits. In examples where the operations of the video encoder (200) are performed by software executed by the programmable circuits, memory (106) (FI. 1) may store object code of the software that the video encoder (200) receives and executes, or another memory (not shown) within the video encoder (200) may store these instructions.

[0182] The video data memory (230) is configured to store received video data. The video encoder (200) may extract a picture of video data from the video data memory (230) and provide the video data to the residual generation unit (204) and the mode selection unit (202). The video data in the video data memory (230) may be raw video data to be encoded.

[0183] The mode selection unit (202) includes a motion estimation unit (222), a motion compensation unit (224), and an intra-prediction unit (226). The mode selection unit (202) may include additional function units to perform video prediction according to different prediction modes. For example, the mode selection unit (202) may include a palette unit, an intra-block copy unit (which may be part of the motion estimation unit (222) and / or the motion compensation unit (224)), an affine unit, a linear model (LM) unit, etc.

[0184] The mode selection unit (202) generally adjusts multiple encoding passes to test combinations of encoding parameters and rate-distortion values ​​of the results for such combinations. The encoding parameters may include partitioning of CTUs into CUs, prediction modes for CUs, transformation types for residual data of CUs, quantization parameters for residual data of CUs, etc. The mode selection unit (202) may ultimately select a combination of encoding parameters that has superior rate-distortion values ​​compared to other tested combinations.

[0185] The video encoder (200) partitions a picture extracted from the video data memory (230) into a series of CTUs and may encapsulate one or more CTUs within a slice. The mode selection unit (202) may partition the CTUs of the picture according to a tree structure, such as the quadtree structure or QTBT structure of HEVC described above. As described above, the video encoder (200) may form one or more CUs from partitioning the CTUs according to a tree structure. Such CUs may also generally be referred to as "video blocks" or "blocks".

[0186] Generally, the mode selection unit (202) also controls its components (e.g., motion estimation unit (222), motion compensation unit (224), and intra prediction unit (226)) to generate a prediction block for the current block (e.g., the current CU, or the overlapping part of PU and TU in HEVC). For the inter prediction of the current block, the motion estimation unit (222) may perform motion search to identify one or more closely matching reference blocks from one or more reference pictures (one or more previously coded pictures stored in DPB (218). In particular, the motion estimation unit (222) may calculate a value indicating how similar a potential reference block is to the current block, for example, based on the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared differences (MSD), etc. The motion estimation unit (222) may perform these calculations using sample-by-sample differences between a reference block generally considered and a current block. The motion estimation unit (222) may also identify a reference block having the lowest value resulting from these calculations, indicating the reference block that most closely matches the current block.

[0187] The motion estimation unit (222) may form one or more motion vectors (MV) that define the positions of reference blocks in reference pictures for the position of the current block in the current picture. The motion estimation unit (222) may then provide the motion vectors to the motion compensation unit (224). For example, for unidirectional inter-prediction, the motion estimation unit (222) may provide a single motion vector, whereas for bidirectional inter-prediction, the motion estimation unit (222) may provide two motion vectors. Then, the motion compensation unit (224) may generate a prediction block using the motion vectors. For example, the motion compensation unit (224) may extract data for the reference block using the motion vectors. As another example, if the motion vector has fractional sample precision, the motion compensation unit (224) may interpolate the values ​​for the prediction block according to one or more interpolation filters. Additionally, for bidirectional inter-prediction, the motion compensation unit (224) may extract data for two reference blocks identified by each motion vector and combine the extracted data, for example, through sample-by-sample averaging or weighted averaging.

[0188] As another example, for intra prediction or intra prediction coding, the intra prediction unit (226) may generate a prediction block from samples adjacent to the current block. For example, for directional modes, the intra prediction unit (226) may generally generate a prediction block by mathematically combining the values ​​of adjacent samples and populating these calculated values ​​in a direction defined across the current block. As another example, for DC mode, the intra prediction unit (226) may calculate the average of the adjacent samples for the current block and generate a prediction block containing the average of these results for each sample of the prediction block.

[0189] The mode selection unit (202) provides the prediction block to the residual generation unit (204). The residual generation unit (204) receives the raw, uncoded version of the current block from the video data memory (230) and the prediction block from the mode selection unit (202). The residual generation unit (204) calculates the sample-by-sample differences between the current block and the prediction block. The sample-by-sample differences of the result define the residual block for the current block. In some examples, the residual generation unit (204) may also determine the differences between sample values ​​in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit (204) may be formed using one or more subtraction circuits that perform binary subtraction.

[0190] In examples where the mode selection unit (202) partitions the CUs into PUs, each PU may be associated with a luminance prediction unit and a corresponding chroma prediction unit. The video encoder (200) and the video decoder (300) may support PUs of various sizes. As indicated above, the size of the CU may represent the size of the luminance coding block of the CU, and the size of the PU may represent the size of the luminance prediction block of the PU. Assuming the size of a specific CU is 2Nx2N, the video encoder (200) may support PU sizes of 2Nx2N or NxN for intra-prediction, and may support symmetric PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, etc. for inter-prediction. The video encoder (20) and video decoder (30) may also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-prediction.

[0191] In examples where the mode selection unit does not further partition the CU into PUs, each CU may be associated with a luminance coding block and a corresponding chroma coding block. As above, the size of the CU may refer to the size of the luminance coding block of the CU. The video encoder (200) and the video decoder (300) may support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0192] For other video coding techniques such as intra-block copy mode coding, affine mode coding, and linear model (LM) mode coding, as in some examples, the mode selection unit (202) generates a prediction block for the current block to be encoded through individual units associated with the coding technique. In some examples, such as palette mode coding, the mode selection unit (202) may not generate a prediction block, but instead generate syntax elements indicating how to reconstruct the block based on the selected palette. In these modes, the mode selection unit (202) may provide these syntax elements to the entropy encoding unit (220) to be encoded.

[0193] As described above, the residual generation unit (204) receives video data for the current block and the corresponding prediction block. The residual generation unit (204) then generates a residual block for the current block. To generate the residual block, the residual generation unit (204) calculates the sample-by-sample differences between the current block and the prediction block.

[0194] The transformation processing unit (206) applies one or more transformations to the residual block to generate a block of transformation coefficients (referred to herein as a “transformation coefficient block”). The transformation processing unit (206) may form the transformation coefficient block by applying various transformations to the residual block. For example, the transformation processing unit (206) may apply the Discrete Cosine Transform (DCT), the Directional Transform, the Karhunen-Loeve Transform (KLT), or a conceptually similar transformation to the residual block. In some examples, the transformation processing unit (206) may perform multiple transformations on the residual block, such as a first transformation and a second transformation, such as a rotation transformation. In some examples, the transformation processing unit (206) does not apply transformations to the residual block.

[0195] The quantization unit (208) may quantize the transformation coefficients in the transformation coefficient block to generate a quantized transformation coefficient block. The quantization unit (208) may quantize the transformation coefficients of the transformation coefficient block according to the quantization parameter (QP) value associated with the current block. The video encoder (202) may adjust the degree of quantization applied to the coefficient blocks associated with the current block by adjusting the QP value associated with the CU (e.g., via the mode selection unit (202)). Quantization may introduce a loss of information, and thus the quantized transformation coefficients may have lower precision than the original transformation coefficients generated by the transformation processing unit (206).

[0196] The inverse quantization unit (210) and the inverse transform processing unit (212) may each apply inverse quantization and inverse transforms to the quantized transform factor block to reconstruct a residual block from the transform factor block. The reconstruction unit (214) may generate a reconstructed block corresponding to the current block (potentially having some degree of distortion) based on the prediction block and the reconstructed residual block generated by the mode selection unit (202). For example, the reconstruction unit (214) may generate a reconstructed block by adding samples of the reconstructed residual block to corresponding samples from the prediction block generated by the mode selection unit (202).

[0197] The filter unit (216) may perform one or more filter operations on the reconstructed block. For example, the filter unit (216) may perform deblocking operations to reduce blockiness artifacts along the edges of the CUs. The operations of the filter unit (216) may be omitted in some examples.

[0198] The video encoder (200) stores the reconstructed blocks in the DPB (218). For example, in examples where the operations of the filter unit (216) are not performed, the reconstruction unit (214) may store the reconstructed blocks in the DPB (218). In examples where the operations of the filter unit (216) are performed, the filter unit (216) may store the filtered reconstructed blocks in the DPB (218). The motion estimation unit (222) and the motion compensation unit (224) may take a reference picture from the DPB (218) formed from the reconstructed (and potentially filtered) blocks and inter-predict blocks of subsequent encoded pictures. Additionally, the intra-prediction unit (226) may use the reconstructed blocks in the current picture's DPB (218) to intra-predict other blocks in the current picture.

[0199] Generally, the entropy encoding unit (220) may entropy encode syntax elements received from other functional components of the video encoder (200), including the syntax elements described above for coefficient coding. For example, the entropy encoding unit (220) may entropy encode quantized transform coefficient blocks from the quantization unit (208). As another example, the entropy encoding unit (220) may entropy encode prediction syntax elements (e.g., motion information for inter-prediction or intra-mode information for intra-prediction) from the mode selection unit (202). The entropy encoding unit (220) may perform one or more entropy encoding operations on syntax elements, which are other examples of video data, to generate entropy-encoded data. For example, the entropy encoding unit (220) may perform a context-adaptive variable-length coding (CAVLC) operation, a CABAC operation, a V2V (variable-to-variable) length coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a probability interval partitioning entropy (PIPE) coding operation, an exponential-Golomb encoding operation, or other types of entropy encoding operations on the data. In some examples, the entropy encoding unit (220) may operate in a bypass mode where syntax elements are not entropy encoded.

[0200] The video encoder (200) may output a bitstream containing entropy-encoded syntax elements necessary to reconstruct blocks of a picture or slice. In particular, the entropy encoding unit (220) may output the bitstream.

[0201] The operations described above are described with respect to blocks. These descriptions should be understood as operations for Luma coding blocks and / or Chroma coding blocks. As described above, in some examples, Luma coding blocks and Chroma coding blocks are Luma and Chroma components of CU. In some examples, Luma coding blocks and Chroma coding blocks are Luma and Chroma components of PU.

[0202] In some examples, operations performed on luminal coding blocks do not need to be repeated for chroma coding blocks. As one example, operations to identify the motion vector (MV) and reference picture for luminal coding blocks do not need to be repeated to identify the MV and reference picture for chroma blocks. Rather, the MV for luminal coding blocks may be scaled to determine the MV for chroma blocks, and the reference picture may be the same. As another example, the intra prediction process may be the same for luminal coding blocks and chroma coding blocks.

[0203] A video encoder (200) represents an example of a device configured to encode video data, comprising a memory configured to store video data, and one or more processing units implemented in a fixed function and / or programmable circuit and configured to perform the exemplary techniques described in the present disclosure.

[0204] FIG. 10 is a block diagram illustrating an exemplary video decoder (300) that may perform the techniques of the present disclosure. FIG. 10 is provided for illustrative purposes and is not limited to techniques as broadly illustrated and described in the present disclosure. For illustrative purposes, the present disclosure describes a video decoder (30) described according to the techniques of JEM and HEVC. However, the techniques of the present disclosure may be performed by video coding devices configured with other video coding standards.

[0205] In the example of FIG. 10, the video decoder (300) includes a coded picture buffer (CPB) (320), an entropy decoding unit (302), a prediction processing unit (304), an inverse quantization unit (306), an inverse transform processing unit (310), a reconstruction unit (310), a filter unit (312), and a decoded picture buffer (DPB) (314). The prediction processing unit (304) includes a motion compensation unit (316) and an intra prediction unit (318). The prediction processing unit (304) may include additional units to perform predictions according to different prediction modes. For example, the prediction processing unit (304) may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit (316)), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder (300) may include more, fewer, or different functional components.

[0206] The CPB memory (320) may store video data, such as an encoded video bitstream to be decoded by components of the video decoder (300). The video data stored in the CPB memory (320) may be obtained, for example, from a computer-readable medium (110) (Fig. 1). The CPB memory (320) may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory (320) may store video data other than syntax elements of the encoded picture, such as transient data representing outputs from various units of the video decoder (300). The DPB (314) generally stores decoded pictures that the video decoder (300) may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. CPB memory (320) and DPB (314) may be formed by any of the various memory devices, such as synchronous dynamic random access memory (SDRAM), DRAM, magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. CPB memory (320) and DPB (314) may be provided by the same memory device or by separate memory devices. In various examples, CPB memory (320) may be on-chip with respect to other components of the video decoder (300) or off-chip with respect to those components.

[0207] Additionally or alternatively, in some examples, the video decoder (300) may retrieve coded video data from memory (120) (Fig. 1). That is, the memory (120) may store data as discussed above in CPB memory (320). Likewise, the memory (120) may store instructions to be executed by the video decoder (300) when some or all of the functionality of the video decoder (300) is implemented in software executed by the processing circuit of the video decoder (300).

[0208] The various units shown in FIG. 10 are illustrated to aid in understanding the operations performed by the video encoder (300). These units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to FIG. 9, fixed-function circuits refer to circuits that provide specific functionality and are pre-configured for the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in the operations that can be performed. For example, programmable circuits may execute software or firmware that causes the programmable circuits to operate in a manner defined by the commands of the software or firmware. Fixed-function circuits may execute software commands (e.g., to receive parameters or output parameters), but the types of operations performed by the fixed-function circuits are generally invariant. In some examples, one or more of the units may be separate circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.

[0209] The video decoder (300) may include ALUs, EFUs, digital circuits, analog circuits, and / or programmable cores formed from programmable circuits. In examples where the operations of the video decoder (300) are performed by software running on the programmable circuits, on-chip or off-chip memory may store instructions (e.g., object code) of the software that the video decoder (300) receives and executes.

[0210] The entropy decoding unit (302) receives video data encoded from CPB and may entropy decode the video data to reproduce syntax elements including the syntax elements described above for coefficient coding. The prediction processing unit (304), the inverse quantization unit (306), the inverse transform processing unit (308), the reconstruction unit (310), and the filter unit (312) may generate decoded video data based on syntax elements extracted from the bitstream.

[0211] Generally, the video decoder (300) reconstructs the picture block by block. The video decoder (300) may also perform the reconstruction operation for each block individually (where the block currently being reconstructed, i.e., being decoded, may be referred to as the “current block”).

[0212] The entropy decoding unit (302) may entropy decode not only conversion information such as quantization parameters (QP) and / or conversion mode indication(s), but also syntax elements defining the quantized conversion coefficients of the quantized conversion coefficient block. The inverse quantization unit (306) may use the QP associated with the quantized conversion coefficient block to determine the degree of quantization and, similarly, the degree of inverse quantization for the inverse quantization unit (306) to apply. The inverse quantization unit (306) may, for example, perform a bit-by-bit left-shift operation to inversely quantize the quantized conversion coefficients. Thus, the inverse quantization unit (306) may form a conversion coefficient block containing the conversion coefficients.

[0213] After the inverse quantization unit (306) forms the transform coefficient block, the inverse transform processing unit (308) may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit (308) may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse directional transform, or other inverse transforms to the coefficient block.

[0214] Additionally, the prediction processing unit (304) generates a prediction block according to the prediction information syntax elements entropy-decoded by the entropy decoding unit (302). For example, if the prediction information syntax elements indicate that the current block is inter-predicted, the motion compensation unit (316) may generate a prediction block. In this case, the prediction information syntax elements may indicate a motion vector identifying the position of the reference block in the reference picture, as well as the position of the current block in the current picture, as well as the reference picture in the DPB (314) from which the reference block will be extracted. The motion compensation unit (316) may perform the inter-predict process in a manner substantially similar to that described in relation to the motion compensation unit (224) (Fig. 9).

[0215] As another example, if the prediction information syntax element indicates that the current block is being intra-predicted, the intra-prediction unit (318) may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax elements. Again, the intra-prediction unit (318) may perform the intra-prediction process in a manner substantially similar to that generally described in relation to the intra-prediction unit (226) (Fig. 9). The intra-prediction unit (318) may extract data of samples adjacent to the current block from the DPB (314).

[0216] The reconstruction unit (310) reconstructs the current block using the prediction block and the residual block. For example, the reconstruction unit (310) may reconstruct the current block by adding samples of the residual block to the corresponding samples of the prediction block.

[0217] The filter unit (312) may perform one or more filter operations on the reconstructed blocks. For example, the filter unit (312) may perform deblocking operations to reduce blockiness artifacts along the edges of the reconstructed blocks. The operations of the filter unit (312) are not necessarily performed in all examples.

[0218] The video decoder (300) may store the reconstructed blocks in the DPB (314). As discussed above, the DPB (314) may provide reference information to the prediction processing unit (304), such as samples of the current picture for intra-prediction and previously decoded pictures for subsequent motion compensation. Additionally, the video decoder (300) may output the decoded pictures from the DPB for subsequent presentation onto a display device, such as the display device (118) of FIG. 1.

[0219] In this way, the video decoder (300) represents an example of a video decoding device, the device comprising a memory configured to store video data, and one or more processing units configured to decode coefficients as implemented in the circuit and described in the present disclosure.

[0220] FIGS. 11a and FIGS. 11b illustrate examples of CABAC processes in bin n. In example (400) of FIG. 11a, in bin n, the range in bin 2 includes RangeMPS and RangeLPS, which are given by the probability (σ) of the least possible symbol (LPS) given a specific context state (σ). Example (400) shows an update of the range in bin n+1 when the value of bin n is equal to the maximum possible symbol (MPS). In this example, the row remains the same, but the value of the range in bin n+1 is reduced to the value of RangeMPS in bin n. Example (402) of FIG. 11b shows an update of the range in bin n+1 when the value of bin n is not equal to MPS (i.e., equal to LPS). In this example, the row is moved to the lower range value of RangeLPS in bin n. Also, the value of Range at empty n+1 is reduced to the value of RangeLPS at empty n.

[0221] In one example of the HEVC video coding process, the range is represented by 9 bits and the row by 10 bits. A renormalization process exists to maintain the range and row values ​​with sufficient precision. Renormalization occurs whenever the range is less than 256. Therefore, after renormalization, the range is always greater than or equal to 256. Depending on the values ​​of the range and row, the Binary Arithmetic Coder (BAC) outputs a bitstream that is '0' or '1', or updates an internal variable (called BO: bits-outstanding) to preserve future outputs. Figure 12 shows examples of range-dependent BAC outputs. For example, when the range and row exceed a specific threshold (e.g., 512), '1' is output to the bitstream. When the range and row are below a specific threshold (e.g., 512), '0' is output to the bitstream. When the range and lower bound are between specific thresholds, nothing is output to the bitstream. Instead, the BO value may be incremented, and the next bin is encoded.

[0222] In the CABAC context model of H.264 / AVC and some examples of HEVC, there are 128 states. 64 possible LPS probabilities that can be 0 to 63 (up womb σ There exists (indicated by ). Each MPS can be 0 or 1. Thus, the 128 states are 64 state probabilities for the two possible values ​​(0 or 1) for the MPS. Therefore, the states can be indexed by 7 bits.

[0223] LPS ranges ( rangeLPS σ To reduce the computation required to derive ), the results for all cases may be pre-calculated and stored as approximations in a look-up table. Thus, the LPS range can be obtained without multiplication by using a simple table lookup. Since this behavior can cause significant latency in many hardware architectures, avoiding multiplication may be important for some devices or applications.

[0224] Instead of multiplication, a 4-column pre-calculated LPS range table may be used. The range is divided into four segments. Segment indices can be derived by the query (range>>6)&3. In effect, segment indices are derived by shifting and dropping bits from the actual range. Table 1 below shows the possible ranges and their corresponding indices.

[0225] Table 1 - Range Index range 256-319 320-383 384-447 448-511 (range>>6) & 3 0 1 2 3

[0226] Subsequently, the LPS range table has 64 entries (one for each probabilistic state) multiplied by 4 (one for each range index). Each entry is the Range LPS, i.e., the value obtained by multiplying the range by the LPS probability. An example of a portion of this table is shown in Table 2 below. Table 2 represents probabilistic states 9-12. In one proposal for HEVC, the probabilistic states may be in the range of 0 to 63.

[0227] Table 2 - RangeLPS Probe status (σ) RangeLPS Index 0 Index 1 Index 2 Index 3 … … … … … 9 90 110 130 150 10 85 104 123 142 11 81 99 117 135 12 77 94 111 128 … … … … …

[0228] In each segment (i.e., range value), each probabilistic state σ The LPS range of is predefined. In other words, the probabilistic state σThe LPS range is quantized into four values ​​(i.e., one value for each range index). The specific LPS range used at a given point depends on which segment the range belongs to. The number of possible LPS ranges used in the table is a trade-off between the number of table columns (i.e., the number of possible LPS range values) and LPS range precision. Generally speaking, more columns result in smaller quantization errors for LPS range values ​​but also increase the need for more memory to store the table. Fewer columns result in larger quantization errors but reduce the memory required to store the table.

[0229] As previously explained, each LPS probability state has a corresponding probability. The probability p for each state is derived as follows:

[0230] p σ = α p σ -1

[0231] Here, the state σ ranges from 0 to 63. The constant α represents the change in probability between each context state. In one example, α = 0.9493, or more precisely, α = (0.01875 / 0.5). The probability in state σ = 0 is equal to 0.5 (i.e., p0 = 1 / 2). That is, in context state 0, LPS and MPS are equally possible. The probability in each successive state is derived by multiplying the previous state by α. Thus, the probability of LPS occurring in context state α = 1 is p0 * 0.9493 (0.5 * 0.9493 = 0.47465). In this way, as the index of state α increases, the probability of LPS occurring decreases.

[0232] CABAC is adaptive because probabilistic states are updated to follow signal statistics (i.e., the values ​​of previously coded bins). The update process is as follows: For a given probabilistic state, the update depends on the state index and the value of the encoded symbol identified as either LPS or MPS. As a result of the update process, a new probabilistic state is derived consisting of a potentially modified LPS probability estimate and, if necessary, a modified MPS value.

[0233] For empty values ​​identical to MPS, the given state index may be incremented by 1. This applies to all states except when MPS occurs at state index (62), where the LPS probability is already at its minimum (or equivalently, the maximum MPS probability is reached). In this case, the state index (62) remains fixed until LPS is shown or the last empty value is encoded (state (63) is used for the special case of the last empty value). When LPS occurs, the state index is changed by decreasing the state index by a predetermined amount as shown in the equation below. This rule generally applies to each occurrence of LPS with the following exception. Assuming LPS is encoded in a state with index σ=0, this corresponds to the equal probabilistic case, where the state index remains fixed, but the MPS value will be toggled so that the values ​​of LPS and MPS are interchanged. In all other cases, regardless of which symbol is encoded, the MPS value will not change. The derivation of transition rules for LPS probability is given LPS probability p old and its updated counterpart p new Based on the following relationship between:

[0234] p new = max( α p old , p 62 ) When MPS occurs

[0235] p new = (1- α) + α p old When LPS occurs

[0236] Regarding the actual implementation of the probability estimation process in CABAC, it is important to note that all transition rules may be realized by up to two tables, each having 63 entries of 6-bit unsigned integer values. In some examples, state transitions may be determined by a single table TransIdxLPS[σ], which determines a new updated state index TransIdxLPS[σ] when an LPS is observed for a given state index σ. MPS-driven transitions can be obtained by a simple (saturated) increment of the state index by a fixed value of 1, resulting in an updated state index min(σ+1, 62). Table 3 below is an example of a partial TransIdxLPS table.

[0237] Table 3 - TransIdxLPS Probe status (σ) New state TransIdxLPS [σ] … … 9 6 10 8 11 8 12 8 … …

[0238] The techniques described above in connection with FIGS. 11a, 11b, and 12 represent only one exemplary implementation of CABAC. It should be understood that the techniques of this disclosure are not limited to this described implementation of CABAC. For example, in older BAC approaches (e.g., the BAC approach used in H.264 / AVC), the tables RangeLPS and TransIdxLPS were tuned for low-resolution videos (i.e., Common Intermediate Format (CIF) and Quarter-CIF (QCIF) videos). With future codecs such as HEVC and VVC, a large amount of video content is HD (high definition) and, in some cases, greater than HD. Video content with HD or higher resolution tends to have different statistics than the 10-year-old QCIF sequences used to develop H.264 / AVC. As such, the RangeLPS and TransIdxLPS tables from H.264 / AVC can cause adaptation between states in a manner that is too rapid. That is, particularly when LPS occurs, the transitions between probabilistic states can be too large for the smoother, higher-resolution content of HD video. Therefore, probabilistic models used according to conventional techniques may not be accurate for HD and extra-HD content. Furthermore, as HD video content contains a larger range of pixel values, H.264 / AVC tables do not contain enough entries to account for the more extreme values ​​that may exist in HD content.

[0239] As such, for HEVC and future coding standards such as VVC, RangeLPS and TransIdxLPS tables may be modified to account for the characteristics of these new content. In particular, BAC processes for HEVC and future coding standards may use tables that allow for slower adaptation processes and may account for more extreme cases (i.e., skewed probabilities). Thus, as an example, RangeLPS and TransIdxLPS tables may be modified to achieve these purposes by including more probabilistic states and ranges than are used in BAC for H.264 / AVC or HEVC.

[0240] FIG. 13 is a block diagram of an exemplary entropy encoding unit (220) that may be configured to perform CABAC according to the techniques of the present disclosure. A syntax element (418) is input into the entropy encoding unit (220). If the syntax element is already a binary value syntax element (i.e., a syntax element having only values ​​of 0 and 1), the binarization step may be omitted. If the syntax element is a non-binary value syntax element (e.g., a syntax element represented by a number of bits such as conversion factor levels), the non-binary value syntax element is binarized by a binarizer (420). The binarizer (420) performs mapping of the non-binary value to a sequence of binary decisions. These binary decisions are often referred to as "bins". For example, for transformation factor levels, the level values ​​may be decomposed into successive bins, each bin indicating whether the absolute value of the factor level is greater than a certain value. For example, bin 0 (sometimes referred to as the significance flag) indicates whether the absolute value of the transformation factor level is greater than 0. Bin 1 indicates whether the absolute value of the transformation factor level is greater than 1, and so on. A unique mapping may be developed for each syntax element of the non-binary value.

[0241] Each bin generated by the binarizer (420) is fed to the binary arithmetic coding side of the entropy encoding unit (220). That is, for a predetermined set of syntax elements of non-binary values, each bin type (e.g., bin 0) is coded before the next bin type (e.g., bin 1). Coding may be performed in either regular mode or bypass mode. In bypass mode, the bypass coding engine (426) performs arithmetic coding using a fixed probability model, for example, using Colomb-Reis or exponential Colomb coding. Bypass mode is generally used for more predictable syntax elements.

[0242] Coding in normal mode involves performing CABAC. Normal mode CABAC is performed to code bin values ​​for which the probability of the bin value is predictable, given the values ​​of previously coded bins. The probability that a bin is LPS is determined by the context modeler (422). The context modeler (422) outputs a bin value and a context model (e.g., a probabilistic state σ). The context model may be an initial context model for a series of bins, or it may be determined based on the coded values ​​of previously encoded bins. As previously mentioned, the context modeler may update the state based on whether the previously coded bin was MPS or LPS.

[0243] After the context model and the probabilistic state σ are determined by the context modeler (422), the regular coding engine (424) performs BAC on the bin values. According to the techniques of the present disclosure, the regular coding engine (424) performs BAC using a TransIdxLPS table (430) containing more than 64 probabilistic states σ. In one example, the number of probabilistic states is 128. TransIdxLPS is used to determine which probabilistic state is used for the next bin (bin n+1) when the previous bin (bin n) is an LPS. The regular coding engine (424) may also use a RangeLPS table (128) to determine the range value for an LPS given a specific probabilistic state σ. However, according to the techniques of the present disclosure, rather than using all possible probability states σ of the TransIdxLPS table (430), probability state indices σ are mapped to grouped indices for use in the RangeLPS table. That is, each index into the RangeLPS table (428) may represent two or more of the total number of probability states. The mapping of probability state indices σ to grouped indices may be linear (e.g., by dividing by 2) or non-linear (e.g., a logarithmic function or a mapping table).

[0244] In other examples of the present disclosure, the difference between consecutive probabilistic states may be smaller by setting the parameter α to be greater than 0.9493. In one example, α = 0.9689. In another example of the present disclosure, the highest probability (p0) of LPS occurring may be set lower than 0.5. In one example, p0 may be equal to 0.493.

[0245] According to one or more techniques of the present disclosure, in contrast to using the same value of a variable (e.g., window size, scaling factor (α), and one or more of the probability update rate) used to update a probability state in a binary arithmetic coding process, the entropy encoding unit (220) may use different values ​​of the variable for different context models and / or different syntax elements. For example, the entropy encoding unit (220) may determine the value of a variable used to update a probability state in a binary arithmetic coding process for one of a plurality of context models and update the probability state based on the determined value.

[0246] FIG. 14 is a block diagram of an exemplary entropy encoding unit (302) that may be configured to perform CABAC according to the techniques of the present disclosure. The entropy decoding unit (302) of FIG. 14 performs CABAC in the inverse manner of the entropy encoding unit (220) described in FIG. 13. Coded bits from a bitstream (448) are input to the entropy decoding unit (302). Based on whether the coded bits were entropy coded using bypass mode or normal mode, the coded bits are supplied to either a context modeler (450) or a bypass decoding engine (452). If the coded bits were coded in bypass mode, the bypass decoding engine (452) may use, for example, Golom-Rice or exponential Golom decoding to extract bins of non-binary syntax elements or syntax elements of binary values.

[0247] If the coded bits are coded in normal mode, the context modeler (450) may determine a probabilistic model for the coded bits, and the normal decoding engine (454) may decode the coded bits to generate bins of syntax elements of non-binary values ​​(or the syntax elements themselves if they are binary values). After the context model and the probabilistic state σ are determined by the context modeler (450), the normal decoding engine (454) performs BAC on the bin values. According to the techniques of the present disclosure, the normal decoding engine (454) performs BAC using a TransIdxLPS table (458) containing more than 64 probabilistic states σ. In one example, the number of probabilistic states is 128, but other numbers of probabilistic states consistent with the techniques of the present disclosure may be defined. The TransIdxLPS table (458) is used to determine which probability state is used for the next bin (bin n+1) when the previous bin (bin n) is an LPS. The normalized decoding engine (454) may also use the RangeLPS table (456) to determine the range value for the LPS given a specific probability state σ. However, according to the techniques of the present disclosure, rather than using all possible probability states σ of the TransIdxLPS table (458), probability state indices σ are mapped to grouped indices for use in the RangeLPS table (456). That is, each index into the RangeLPS table (456) may represent two or more of the total number of probability states. The mapping of probability state indices σ to grouped indices may be linear (e.g., by dividing by 2) or non-linear (e.g., a logarithmic function or a mapping table).

[0248] In other examples of the present disclosure, the difference between consecutive probabilistic states may be smaller by setting the parameter α to be greater than 0.9493. In one example, α = 0.9689. In another example of the present disclosure, the highest probability (p0) of LPS occurring may be set lower than 0.5. In one example, p0 may be equal to 0.493.

[0249] After the bins are decoded by the regular decoding engine (224), the debinizer (460) may perform inverse mapping to convert the bins back into non-binary value syntax element values.

[0250] FIG. 15 is a flowchart illustrating an exemplary operation of a video encoder for encoding a current block of video data. The current block may include a current CU. Although described in relation to the video encoder (200) (Fig. 1 and Fig. 9), it should be understood that other devices may be configured to perform an operation similar to that of FIG. 15.

[0251] In this example, the video encoder (200) initially predicts the current block (550). For example, the video encoder (200) may form a predicted block for the current block. Then, the video encoder (200) may calculate a residual block for the current block (552). To calculate the residual block, the video encoder (200) may calculate the difference between the original uncoded block and the predicted block for the current block. After that, the video encoder (200) may transform and quantize the coefficients of the residual block (554). Next, the video encoder (200) may scan the quantized transformed coefficients of the residual block (556). During or after the scan, the video encoder (200) may entropy-encode the coefficients (558). For example, the video encoder (200) may encode the coefficients using CAVLC or CABAC. The video encoder (200) may output the entropy-coded data of the block (560).

[0252] FIG. 16 is a flowchart illustrating an exemplary operation of a video decoder for decoding a current block of video data. The current block may include a current CU. Although described in relation to the video decoder (300) (Figs. 1 and 3), it should be understood that other devices may be configured to perform an operation similar to that of FIG. 16.

[0253] The video decoder (300) may receive entropy-coded data for the current block, such as entropy-coded prediction information and entropy-coded data for the coefficients of the residual block corresponding to the current block (570). The video decoder (300) may entropy-decode the entropy-coded data to determine prediction information for the current block and regenerate the coefficients of the residual block (572). The video decoder (300) may predict the current block using an intra-prediction or inter-prediction mode, such as indicated by the prediction information for the current block, to calculate the prediction block for the current block (574). The video decoder (300) may then back-scan the regenerated coefficients to generate a block of quantized transform coefficients (576). After that, the video decoder (300) may generate a residual block by inverse-quantizing and inverse-transforming the coefficients (578). The video decoder (300) can also ultimately decode the current block by combining the prediction block and the residual block (580).

[0254] FIG. 17 is a flowchart illustrating the exemplary operation of a video decoder for decoding coefficient values. Although it has been described in relation to the video decoder (300) (Fig. 1 and Fig. 10), it should be understood that other devices may be configured to perform an operation similar to that of FIG. 17.

[0255] The video decoder (300) determines the threshold number of normally coded bins for the first decoding pass (602).

[0256] For a first set of coefficients, the video decoder (300) context-decodes the syntax elements of the coefficient group until a threshold number of normally coded bins is reached (604). The context-decoded bins of the syntax elements may include, for example, one or more significance flags, one or more parity level flags, and one or more first flags, as described above. Each of the one or more significance flags may indicate whether the absolute level for the coefficient is equal to 0, and each of the one or more parity level flags may indicate whether the coefficient has an even or odd absolute level. Each of the one or more first flags may indicate whether the coefficient has an absolute level greater than 2.

[0257] To context-decode the syntax elements of a group of coefficients, the video decoder (300) may perform context-adaptive binary arithmetic decoding to decode the syntax elements of a group of coefficients. In other examples, to context-decode the syntax elements of a group of coefficients until a threshold number of normally coded bins is reached, the video decoder (300) may context-decode one or more remaining syntax elements for a first set of coefficients, after determining that a threshold number of normally coded bins has been reached while coding the syntax elements for a first set of coefficients.

[0258] The video decoder (300) determines values ​​for a first set of coefficients of a transform unit based on context-decoded bins of syntax elements (606). In response to reaching a threshold number of normally coded bins, for a second set of coefficients, the video decoder (300) bypass decodes additional syntax elements (608). To bypass decode additional syntax elements, the video decoder (300) may derive values ​​for the rice parameters for the coefficients of the second set of coefficients.

[0259] The video decoder (300) determines values ​​for a second set of coefficients of a transform unit based on additional syntax elements (610). To determine values ​​for a second set of coefficients of a transform unit based on additional syntax elements, the video decoder (300) determines a value for a zero parameter based on a rice parameter (612). To determine a value for a zero parameter based on a rice parameter, the video decoder (300) may determine a value for a zero parameter, for example, based on the rice parameter and also based on the current state of the state machine. As previously mentioned, the value for a zero parameter identifies a coded value corresponding to a coefficient level of zero. The video decoder (300) may determine a value for a rice parameter, for example, from a lookup table or in some other way.

[0260] In order to determine values ​​for a second set of coefficients of a conversion unit based on additional syntax elements, the video decoder (300) also receives a first coded value for a first coefficient of the second set of coefficients (614), and determines a level for the first coefficient based on the value for the zero parameter and the first coded value for the first coefficient (616). The level for the first coefficient may be, for example, either a residual level or an absolute level.

[0261] In response to the value for the zero parameter being equal to the first coded value, the video decoder (300) may determine that the level for the first coefficient is equal to zero. In response to the first coded value being greater than the value for the zero parameter, the video decoder (300) may determine that the level for the first coefficient is equal to the first coded value. In other cases, in response to the first coded value being less than the value for the zero parameter, the video decoder (300) may determine that the level for the first coefficient is equal to the first coded value plus 1.

[0262] The video decoder (300) may also determine a decoded conversion block based on values ​​for a first set of coefficients and values ​​for a second set of coefficients; determine a reconstructed block by adding the decoded conversion block to a prediction block; determine a decoded block of video data by performing one or more filtering operations on the reconstructed block; and output a decoded picture of video data including the decoded block of video data.

[0263] It should be recognized that, depending on the example, specific actions or events of any of the techniques described herein may be performed in different sequences and may be added, merged, or removed in their entirety (e.g., not all described actions or events are essential for the execution of the techniques). Furthermore, in certain examples, actions or events may be performed simultaneously rather than sequentially, for example, through multi-threaded processing, interrupt processing, or multiple processors.

[0264] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored or transmitted as one or more instructions or code on a computer-readable medium, or may be executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium corresponding to a medium of the same type as a data storage medium, or a communication medium including any medium that facilitates the transmission of a computer program from one place to another, for example, according to a communication protocol. In this way, computer-readable media may generally correspond to (1) non-transient types of computer-readable storage media or (2) a communication medium such as a signal or carrier wave. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to extract instructions, code and / or data structures for the implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0265] As an example that is not a limitation, these computer-readable storage media may include one or more of RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other media that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer. Additionally, any connection is appropriately named as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, such coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead relate to non-transient, tangible storage media. Disks and discs, as used herein, include compact discs (CDs), laser discs, optical discs, digital multi-purpose discs (DVDs), floppy discs, and Blu-ray discs, wherein disks typically reproduce data magnetically, while discs reproduce data optically using lasers. The above combinations should also be included within the range of computer-readable media.

[0266] Instructions may be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits. Accordingly, the term “processor” may refer to any of any other structure suitable for implementing the aforementioned structure or the techniques described herein, as used herein. Additionally, in some embodiments, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or integrated into a combined codec. Furthermore, the techniques may be fully implemented in one or more circuits or logic elements.

[0267] The techniques of the present disclosure may be implemented in a wide variety of devices or apparatus, including wireless handsets, integrated circuits (ICs), or sets of ICs (e.g., chip sets). Various components, modules, or units are described in the present disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but implementation by different hardware units is not necessarily required. Rather, as described above, various units may be provided by a collection of interoperable hardware units, which may be combined in a codec hardware unit or, together with suitable software and / or firmware, include one or more processors as described above.

[0268] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

Claim 1 A method for decoding video data, the method comprises: determining a threshold number of normally coded bins for a first decoding pass; context decoding bins of syntax elements of a group of coefficients for a first set of coefficients until the threshold number of normally coded bins is reached, wherein the context-decoded bins of the syntax elements include one or more significance flags, one or more parity level flags, and one or more first flags, each of the one or more significance flags indicates whether the absolute level for the corresponding coefficient is equal to zero, each of the one or more parity level flags indicates whether the absolute level for the corresponding coefficient is even or odd, and each of the one or more first flags indicates whether the absolute level for the corresponding coefficient is greater than 2; determining values ​​for the first set of coefficients of a transformation unit based on the context-decoded bins of syntax elements; and in response to reaching the threshold number of normally coded bins, the coefficients For a second set, a step of bypass decoding additional syntax elements, wherein bypass decoding additional syntax elements comprises deriving a value for a Rice parameter for the coefficients of the second set of coefficients; and a step of determining values ​​for the second set of coefficients of the transformation unit based on the additional syntax elements, wherein the step of determining values ​​for the second set of coefficients of the transformation unit based on the additional syntax elements comprises: a step of determining a value for a Zero parameter based on the Rice parameter, wherein the value for the Zero parameter identifies a coded value corresponding to a coefficient level of zero;A method for decoding video data, comprising: receiving a first coded value for a first coefficient among a second set of the coefficients; and determining a level for the first coefficient based on the value for the zero parameter and the first coded value for the first coefficient. Claim 2 A method for decoding video data according to claim 1, wherein the level for the first coefficient includes a residual level. Claim 3 A method for decoding video data according to claim 1, wherein the level for the first coefficient includes an absolute level. Claim 4 A method for decoding video data according to claim 1, wherein the step of determining a value for the zero parameter based on the rice parameter includes the step of determining a value for the zero parameter based on the rice parameter and based on the current state of the state machine. Claim 5 A method for decoding video data according to claim 1, further comprising the step of determining that the level of the first coefficient is equal to zero in response to the value for the zero parameter being equal to the first coded value. Claim 6 A method for decoding video data according to claim 1, further comprising the step of determining that the level for the first coefficient is the same as the first coded value in response to the first coded value being greater than the value for the zero parameter. Claim 7 A method for decoding video data according to claim 1, further comprising the step of determining that the level for the first coefficient is equal to the first coded value plus 1 in response to the first coded value being smaller than the value for the zero parameter. Claim 8 A method for decoding video data according to claim 1, further comprising the step of determining a value for the rice parameter from a lookup table. Claim 9 A method for decoding video data according to claim 1, wherein the step of context-decoding the syntax elements of the coefficient group comprises the step of performing context-adaptive binary arithmetic decoding to decode the syntax elements of the coefficient group. Claim 10 A method for decoding video data according to claim 1, wherein the step of context-decoding syntax elements of the coefficient group until a threshold number of normally coded bins is reached comprises: determining that the threshold number of normally coded bins has been reached while coding syntax elements for a first set of coefficients; and context-decoding one or more residual syntax elements for a first set of coefficients. Claim 11 A method for decoding video data according to claim 1, further comprising: determining a decoded transformation block based on values ​​for a first set of coefficients and values ​​for a second set of coefficients; determining a reconstructed block by adding the decoded transformation block to a prediction block; performing one or more filtering operations on the reconstructed block to determine a decoded block of video data; and outputting a decoded picture of video data including the decoded block of video data. Claim 12 A device for decoding video data, wherein the device comprises a memory configured to store the video data; The circuit comprises one or more processors implemented therein, wherein the one or more processors: determine a threshold number of normally coded bins for a first decoding pass; perform context decoding for a first set of coefficients, wherein the bins of syntax elements of a coefficient group are context decoded until the threshold number of normally coded bins is reached, wherein the context-decoded bins of the syntax elements comprise one or more significance flags, one or more parity level flags, and one or more first flags, wherein each of the one or more significance flags indicates whether the absolute level for the corresponding coefficient is equal to zero, each of the one or more parity level flags indicates whether the absolute level for the corresponding coefficient is even or odd, and each of the one or more first flags indicates whether the absolute level for the corresponding coefficient is greater than 2; determine values ​​for the first set of coefficients of a transformation unit based on the context-decoded bins of syntax elements; and the threshold of the normally coded bins In response to reaching a number, for a second set of coefficients, bypass decoding additional syntax elements, wherein to bypass decode said additional syntax elements, said one or more processors perform bypass decoding said additional syntax elements configured to derive a value for a Rice parameter for the coefficients of the second set of said coefficients;A device for decoding video data, configured to determine values ​​for a second set of coefficients of the conversion unit based on the additional syntax elements, and to determine values ​​for a second set of coefficients of the conversion unit based on the additional syntax elements, wherein the one or more processors perform determining a value for a zero parameter based on the rice parameter, wherein the value for the zero parameter identifies a coded value corresponding to a coefficient level of zero; receive a first coded value for a first coefficient among the second set of coefficients; and, based on the value for the zero parameter and the first coded value for the first coefficient, determine a level for the first coefficient. Claim 13 A device for decoding video data, wherein the level for the first coefficient includes a residual level in claim 12. Claim 14 A device for decoding video data, wherein the level for the first coefficient in claim 12 includes an absolute level. Claim 15 A device for decoding video data according to claim 12, wherein, in order to determine a value for the zero parameter based on the rice parameter, the one or more processors are configured to determine a value for the zero parameter based on the rice parameter and based on the current state of the state machine. Claim 16 A device for decoding video data, wherein the one or more processors are also configured to determine that the level of the first coefficient is equal to zero in response to the value for the zero parameter being equal to the first coded value. Claim 17 A device for decoding video data, wherein the one or more processors are also configured to determine that the level for the first coefficient is equal to the first coded value in response to the first coded value being greater than the value for the zero parameter. Claim 18 A device for decoding video data, wherein the one or more processors are also configured to determine that the level for the first coefficient is equal to the first coded value plus 1 in response to the first coded value being smaller than the value for the zero parameter. Claim 19 In claim 12, the device for decoding video data, wherein the one or more processors are also configured to determine a value for the rice parameter from a lookup table. Claim 20 A device for decoding video data according to claim 12, wherein, in order to context decode the syntax elements of the coefficient group, the one or more processors are configured to perform context-adaptive binary arithmetic decoding to decode the syntax elements of the coefficient group. Claim 21 A device for decoding video data, wherein, in order to context decode syntax elements of a group of coefficients until a threshold number of normally coded bins is reached, the one or more processors are configured to: determine that a threshold number of normally coded bins has been reached while coding syntax elements for a first set of coefficients; and context decode one or more remaining syntax elements for a first set of coefficients. Claim 22 A device for decoding video data, wherein, in claim 12, the one or more processors also: determine a decoded transform block based on values ​​for a first set of coefficients and values ​​for a second set of coefficients; determine a reconstructed block by adding the decoded transform block to a prediction block; perform one or more filtering operations on the reconstructed block to determine a decoded block of video data; and output a decoded picture of video data including the decoded block of video data. Claim 23 In claim 12, the device for decoding video data comprises a wireless communication device and further comprises a receiver configured to receive encoded video data. Claim 24 In claim 23, the wireless communication device includes a telephone handset, and the receiver is a device for decoding video data configured to demodulate a signal including the encoded video data according to a wireless communication standard. Claim 25 A device for decoding video data according to claim 12, further comprising a display configured to display decoded video data. Claim 26 In claim 12, the device is a device for decoding video data, comprising one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box. Claim 27 A computer-readable storage medium for storing instructions, wherein the instructions, when executed by one or more processors, cause the one or more processors: to determine a threshold number of normally coded bins for a first decoding pass; to context decode bins of syntax elements of a group of coefficients for a first set of coefficients until the threshold number of normally coded bins is reached, wherein the context-decoded bins of the syntax elements include one or more significance flags, one or more parity level flags, and one or more first flags, each of the one or more significance flags indicates whether the absolute level for the corresponding coefficient is equal to zero, each of the one or more parity level flags indicates whether the absolute level for the corresponding coefficient is even or odd, and each of the one or more first flags indicates whether the absolute level for the corresponding coefficient is greater than 2, thereby performing the context decoding; and to the first set of coefficients of a transformation unit based on the context-decoded bins of syntax elements Determining values ​​for; and in response to reaching a threshold number of the normally coded bins, bypass decoding additional syntax elements for a second set of coefficients, wherein to bypass decode the additional syntax elements, the instructions cause the one or more processors to perform the bypass decoding, which derives a value for a Rice parameter for the coefficients of the second set of coefficients;A computer-readable storage medium, wherein, in order to determine values ​​for a second set of coefficients of the conversion unit based on the additional syntax elements, the instructions cause the one or more processors to: determine a value for a zero parameter based on the rice parameter, wherein the value for the zero parameter identifies a coded value corresponding to a coefficient level of zero; receive a first coded value for a first coefficient among the second set of coefficients; and determine a level for the first coefficient based on the value for the zero parameter and the first coded value for the first coefficient. Claim 28 In claim 27, a computer-readable storage medium wherein the level for the first coefficient includes a residual level. Claim 29 In claim 27, a computer-readable storage medium wherein the level for the first coefficient includes an absolute level. Claim 30 In claim 27, a computer-readable storage medium wherein, in order to determine a value for the zero parameter based on the rice parameter, the instructions cause the one or more processors to determine a value for the zero parameter based on the rice parameter and based on the current state of the state machine. Claim 31 In claim 27, the above instructions also cause the one or more processors to determine that the level of the first coefficient is equal to zero in response to the value for the zero parameter being equal to the first coded value. Claim 32 In claim 27, the instructions also cause the one or more processors to determine that the level for the first coefficient is equal to the first coded value in response to the first coded value being greater than the value for the zero parameter, a computer-readable storage medium. Claim 33 In claim 27, the instructions also cause the one or more processors to determine that the level for the first coefficient is equal to the first coded value plus 1 in response to the first coded value being smaller than the value for the zero parameter, a computer-readable storage medium. Claim 34 In claim 27, the above instructions also allow one or more processors to determine a value for the rice parameter from a lookup table, a computer-readable storage medium. Claim 35 In claim 27, a computer-readable storage medium wherein, for context decoding the syntax elements of the coefficient group, the instructions cause the one or more processors to perform context-adaptive binary arithmetic decoding to decode the syntax elements of the coefficient group. Claim 36 A computer-readable storage medium according to claim 27, wherein, for context-decoding syntax elements of a group of coefficients until a threshold number of normally coded bins is reached, the instructions cause the one or more processors: to determine that a threshold number of normally coded bins has been reached while coding syntax elements for a first set of coefficients; and to context-decode one or more remaining syntax elements for a first set of coefficients. Claim 37 In claim 27, the instructions also cause the one or more processors to: determine a decoded conversion block based on values ​​for a first set of coefficients and values ​​for a second set of coefficients; determine a reconstructed block by adding the decoded conversion block to a prediction block; perform one or more filtering operations on the reconstructed block to determine a decoded block of video data; and output a decoded picture of video data including the decoded block of video data, a computer-readable storage medium. Claim 38 A device for decoding video data, the device comprising: means for determining a threshold number of normally coded bins for a first decoding pass; means for context decoding bins of syntax elements of a group of coefficients for a first set of coefficients until the threshold number of normally coded bins is reached, wherein the context-decoded bins of the syntax elements comprise one or more significance flags, one or more parity level flags, and one or more first flags, wherein each of the one or more significance flags indicates whether the absolute level for the corresponding coefficient is equal to zero, each of the one or more parity level flags indicates whether the absolute level for the corresponding coefficient is even or odd, and each of the one or more first flags indicates whether the absolute level for the corresponding coefficient is greater than 2; means for determining values ​​for the first set of coefficients of a transformation unit based on the context-decoded bins of syntax elements; and in response to reaching the threshold number of normally coded bins, For a second set of coefficients, means for bypass decoding additional syntax elements, said means for bypass decoding additional syntax elements, said means for bypass decoding the additional syntax elements, said means for deriving a value for a Rice parameter for the coefficients of the second set of coefficients; and means for determining values ​​for the second set of coefficients of the transformation unit based on said additional syntax elements, said means for determining values ​​for the second set of coefficients of the transformation unit based on said additional syntax elements, said means for determining a value for a Zero parameter based on said Rice parameter, said means for determining a value for a Zero parameter, said value for a Zero parameter identifying a coded value corresponding to a coefficient level of zero;An apparatus for decoding video data, comprising: means for receiving a first coded value for a first coefficient among a second set of the above coefficients; and means for determining a level for the first coefficient based on a value for the zero parameter and the first coded value for the first coefficient.

Citation Information

Patent Citations

  • Context Adaptive Binary Arithmetic Coding (CABAC) with Scalable Throughput and Coding Efficiency

    US20130177069A1