Coefficient Coding for Support of Various Color Formats in Video Coding

By developing a context derivation process for CABAC that considers the color format of the picture, the technique addresses the issue of poor coding efficiency in existing video encoding and decoding methods, achieving improved compression performance.

JP7689978B2Active Publication Date: 2025-06-09QUALCOMM INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022557113
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-06
Filing Date
2021-04-07
Publication Date
2025-06-09
Estimated Expiration
2041-04-07

AI Technical Summary

Technical Problem

Existing video encoding and decoding techniques, such as those in Essential Video Coding (EVC), do not accurately account for various color formats like 4:4:4 and 4:2:2, leading to less accurate context probability selection and poor coding efficiency.

Method used

A context derivation process for Context Adaptive Binary Arithmetic Coding (CABAC) is developed that takes into account the color format of the picture, allowing for the selection of a more accurate context model for syntax elements indicating the prefixes of the x and y coordinates of the last significant transform coefficients.

Benefits of technology

This approach improves coding efficiency by selecting a context with more accurate probability, leading to better compression performance across different color formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689978000007
    Figure 0007689978000007
  • Figure 0007689978000008
    Figure 0007689978000008
  • Figure 0007689978000009
    Figure 0007689978000009
Patent Text Reader

Abstract

A method for decoding video data comprises determining, based on a color format of the picture, which context model to use from among a first context model and a second context model to determine a context increment of a syntax element indicating a prefix of an x ​​or y coordinate of a last significant transform coefficient of a color component of a block of a picture of the video data, and decoding a bin of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] This application claims priority to U.S. Patent Application No. 17 / 223,814, filed Apr. 6, 2021, and U.S. Provisional Patent Application No. 63 / 009,292, filed Apr. 13, 2020, each of which is hereby incorporated by reference in its entirety.

[0002]

[0002] This disclosure relates to video encoding and video decoding.

Background Art

[0003]

[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, so-called "smartphones", video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video coding techniques such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions of such standards. By implementing such video coding techniques, video devices can transmit, receive, encode, decode, and / or store digital video information more efficiently.

[0004]

[0004] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in a video sequence. In block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, sometimes referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in adjacent blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples in adjacent blocks in the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may sometimes be referred to as a frame, and a reference picture may sometimes be referred to as a reference frame.

Summary of the Invention

[0005]

[0005] Generally, in the present disclosure, techniques for converting coefficient coding of video data having various color formats, such as 4:2:2 and 4:4:4 in addition to 4:2:0, and for enabling coding will be described. As described herein, the Essential Video Coding Test Model 5.0 (ETM5.0) uses a context derivation process to determine the context of Context Adaptive Binary Arithmetic Coding (CABAC) of syntax elements indicating the prefixes of the x and y coordinates of the last significant transform coefficients of a block. The context specifies the probability of a symbol. The context derivation process of ETM5.0 does not take into account various color formats. This can lead to the selection of a context with less accurate probability. In the present disclosure, a context derivation process for determining the context of CABAC of syntax elements indicating the prefixes of the x and y coordinates of the last significant transform coefficients of a block will be described in terms of techniques based on the color format of the picture containing the block. This can result in the selection of a context with more accurate probability, which can ultimately lead to excellent coding efficiency.

[0006]

[0006] In one example, the present disclosure describes a method for decoding video data. The method includes determining which context model to use from a first context model and a second context model based on the color format of a picture to determine a context increment of a syntax element indicating the prefix of the x or y coordinate of the last significant transform coefficient of a color component of a block of the picture, and decoding the bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0007]

[0007] In another example, the present disclosure describes a method for encoding video data. The method includes determining which context model to use from a first context model and a second context model based on the color format of a picture in order to determine a context increment of a syntax element indicating a prefix of the x or y coordinate of the last significant transform coefficient of a color component of a block of the picture of the video data, and encoding the bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0008]

[0008] In another example, the present disclosure describes a device for decoding video data. The device includes a memory configured to store the video data and one or more processors coupled to the memory. The one or more processors are implemented in a circuit and are configured to determine which context model to use from a first context model and a second context model based on the color format of a picture in order to determine a context increment of a syntax element indicating a prefix of the x or y coordinate of the last significant transform coefficient of a color component of a block of the picture of the video data, and decode the bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0009]

[0009] In another example, the present disclosure describes a device for encoding video data. The device includes a memory configured to store video data and one or more processors coupled to the memory. The one or more processors are implemented in a circuit and are configured to determine which context model to use from a first context model and a second context model based on the color format of a picture to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data, and to encode a bin of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0010]

[0010] In another example, the present disclosure describes a device for decoding video data. The device includes means for determining which context model to use from a first context model and a second context model based on the color format of a picture to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data, and means for decoding a bin of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0011]

[0011] In another example, the present disclosure describes a device for encoding video data. The device determines, based on the color format of a picture, which context model to use between a first context model and a second context model for determining a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data, and encodes bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0012]

[0012] In another example, the present disclosure describes a computer-readable storage medium storing instructions that, when executed, cause one or more processors to determine which context model to use between a first context model and a second context model based on the color format of a picture for determining a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data, and decode bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0013]

[0013] In another example, the present disclosure describes a computer-readable storage medium storing instructions that, when executed, cause one or more processors to determine which context model to use between a first context model and a second context model based on the color format of a picture for determining a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data, and encode bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0014]

[0014] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0015]

Figure 1

[0015] A block diagram showing an exemplary video encoding and decoding system that can implement the techniques of the present disclosure.

Figure 2A

[0016] A conceptual diagram showing an exemplary quad-tree binary-tree (QTBT) structure.

Figure 2B

Figure 3

[0017] A conceptual diagram showing an adaptive transform selection for an inter-coded block.

Figure 4

[0018] A conceptual diagram showing a coefficient scanning method and a last coefficient position.

Figure 5

[0019] A table showing coding symbols for coefficients of a 16 - coefficient chunk.

Figure 6

[0020] A block diagram showing an exemplary video encoder that can implement the techniques of the present disclosure.

Figure 7

[0021] A block diagram showing an exemplary video decoder that can implement the techniques of the present disclosure.

Figure 8

[0022] A flowchart showing an exemplary method for encoding a current block.

Figure 9

[0023] A flowchart showing an exemplary method for decoding a current block of video data.

Figure 10

[0024] A flowchart showing an exemplary operation of a video encoder according to one or more techniques of the present disclosure.

Figure 11

[0025] Flowchart showing exemplary operation of a video encoder according to one or more techniques of the present disclosure.

Figure 12

[0026] Flowchart showing exemplary operation of a video decoder according to one or more techniques of the present disclosure.

Figure 13

[0027] Flowchart showing exemplary operation of a video encoder for determining a context according to one or more techniques of the present disclosure.

Figure 14

[0028] Flowchart showing exemplary operation for determining the value of a luma model variable according to one or more techniques of the present disclosure.

DETAILED DESCRIPTION

[0016]

[0029] The context model enables a video encoder to determine the context to be used in context-adaptive binary arithmetic coding (CABAC). Conventionally, in essential video coding (EVC), a video encoder (e.g., a video encoder or a video decoder) uses a first context model (e.g., a luma context model) when determining the context to be used for CABAC coding of a syntax element indicating the prefix of the coordinates of the last significant transform coefficient of the luma component of a block, and a second different context model (e.g., a chroma context model) when determining the context to be used for CABAC coding of a syntax element indicating the prefix of the coordinates of the last significant transform coefficient of the chroma component of a block. In the present disclosure, a significant transform coefficient is a non-zero transform coefficient.

[0017]

[0030] The video coder uses these two different context models in EVC because the color format of the picture is assumed to be 4:2:0. When the color format of the picture is 4:2:0, there are chroma samples that are half the number of luma samples both horizontally and vertically. Due to this difference in the number of chroma samples compared to luma samples, different statistical values can exist for the bin values of the syntax elements that indicate the prefix of the coordinates of the last significant transform coefficients for luma and chroma.

[0018]

[0031] However, other color formats are possible, such as 4:4:4 and 4:2:2. In a picture with a 4:4:4 color format, there are an equal number of luma and chroma samples both horizontally and vertically. In a picture with a 4:2:2 color format, there are chroma samples that are half the number of luma samples horizontally, and an equal number of luma and chroma samples vertically. Using the EVC chroma context model with these other color formats can lead to poor coding efficiency.

[0019]

[0032] In the present disclosure, techniques that can address this problem and thereby improve coding efficiency will be described. In one example, a video coder (e.g., a video encoder or a video decoder) may determine which context model to use from a first context model and a second context model based on the color format of a picture in order to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of a picture of video data. The video coder may code (e.g., encode or decode) a bin of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element. By determining whether to use the first context model or the second context model to determine the context increment of the syntax element, the video coder may be able to better select a context suitable for coding the bin of the syntax element. This may improve coding efficiency. In some situations, the same context model may be used for the luma component and the chroma component.

[0020]

[0033] FIG. 1 is a block diagram showing an exemplary video encoding and decoding system 100 that may implement the techniques of the present disclosure. The techniques of the present disclosure generally target coding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Thus, video data may include raw unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata such as signaling data.

[0021]

[0034] As shown in FIG. 1, system 100 includes, in this example, a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. In particular, source device 102 provides video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 may comprise any of a wide range of devices, including desktop computers, mobile devices, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, broadcast receiver devices, and the like. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and thus may be referred to as wireless communication devices.

[0022]

[0035] In the example of FIG. 1, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. According to the present disclosure, video encoder 200 of source device 102 and video decoder 300 of destination device 116 may be configured to apply techniques for coefficient coding for support of various color formats. Thus, source device 102 represents an example of a video encoding device and destination device 116 represents an example of a video decoding device. In other examples, the source device and the destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device rather than including a built-in display device.

[0023]

[0036] The system 100 shown in FIG. 1 is merely an example. Generally, any digital video encoding and / or decoding device may implement techniques for coefficient coding for support of various color formats. The source device 102 and the destination device 116 are merely examples of such coding devices that generate encoded video data for transmission from the source device 102 to the destination device 116. In the present disclosure, a "coding" device is referred to as a device that performs coding (encoding and / or decoding) of data. Thus, the video encoder 200 and the video decoder 300 represent examples of coding devices, specifically, a video encoder and a video decoder, respectively. In some examples, the source device 102 and the destination device 116 may operate substantially symmetrically such that each of the source device 102 and the destination device 116 includes video encoding and decoding components. Thus, the system 100 may support one-way or two-way video transmission between the source device 102 and the destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0024]

[0037] Generally, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a sequential series of pictures (also referred to as "frames") of video data to video encoder 200 that encodes the data of the pictures. The video source 104 of source device 102 can include a video capture device such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 104 can generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, the video encoder 200 encodes the captured video data, pre-captured video data, or computer-generated video data. The video encoder 200 can reorder the pictures from the reception order (sometimes referred to as "display order") to a coding order for coding. The video encoder 200 can generate a bitstream that includes the encoded video data. The source device 102 can then output the encoded video data onto computer-readable medium 110 via output interface 108 for reception and / or retrieval, for example, by input interface 122 of destination device 116.

[0025]

[0038] The memory 106 of the source device 102 and the memory 120 of the destination device 116 represent general-purpose memories. In some examples, the memories 106, 120 may store raw video data, such as raw video from the video source 104 and raw decoded video data from the video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions executable by, for example, the video encoder 200 and the video decoder 300, respectively. Although the memory 106 and the memory 120 are shown separately from the video encoder 200 and the video decoder 300 in this example, it should be understood that the video encoder 200 and the video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Further, the memories 106, 120 may store encoded video data, such as the output from the video encoder 200 and the input to the video decoder 300. In some examples, portions of the memories 106, 120 may be allocated as one or more video buffers, for example, to store raw decoded and / or encoded video data.

[0026]

[0039] Computer-readable medium 110 may represent any type of medium or device capable of transferring encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to directly transmit encoded video data to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. Output interface 108 may demodulate a transmission signal including the encoded video data, and input interface 122 may demodulate a received transmission signal according to a communication standard such as a wireless communication protocol. The communication medium may comprise any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be useful for facilitating communication from source device 102 to destination device 116.

[0027]

[0040] In some examples, source device 102 may output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray Disc, a DVD, a CD-ROM, a flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.

[0028]

[0041] In some examples, the source device 102 may output the encoded video data to a file server 114 or another intermediate storage device that may store the encoded video data generated by the source device 102. The destination device 116 may access the stored video data from the file server 114 via streaming or downloading. The file server 114 can be any type of server device that can store the encoded video data and transmit the encoded video data to the destination device 116. The file server 114 can represent a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a Network Attached Storage (NAS) device. The destination device 116 can access the encoded video data from the file server 114 through any standard data connection including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi® connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both that is suitable for accessing the encoded video data stored on the file server 114. The file server 114 and the input interface 122 can be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.

[0029]

[0042] The output interface 108 and the input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet® card), a wireless communication component operating according to any of various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 may be configured to transfer data such as encoded video data according to cellular communication standards such as 4G, 4G-LTE® (Long Term Evolution), LTE Advanced, 5G, and the like. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 may be configured to transfer data such as encoded video data according to other wireless standards such as the IEEE 802.11 specifications, the IEEE 802.15 specifications (e.g., ZigBee®), the Bluetooth® standards, and the like. In some examples, the source device 102 and / or the destination device 116 may each include a respective system-on-chip (SoC) device. For example, the source device 102 may include an SoC device for implementing functions attributed to the video encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device for implementing functions attributed to the video decoder 300 and / or the input interface 122.

[0030]

[0043] The techniques of the present disclosure may be applied to video coding that supports any of various multimedia applications such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as dynamic adaptive streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0031]

[0044] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, such as syntax elements having values that describe the characteristics and / or processing of video blocks or other coded units (e.g., slices, pictures, picture groups, sequences, etc.), which are also used by the video decoder 300. The display device 118 displays the decoded pictures of the decoded video data to the user. The display device 118 may represent any of various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0032]

[0045] Although not shown in FIG. 1, in some examples, the video encoder 200 and the video decoder 300 may each be integrated with an audio encoder and / or an audio decoder and may include a suitable MUX-DEMUX unit, or other hardware and / or software, to handle a multiplexed stream that includes both audio and video in a common data stream. When applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).

[0033]

[0046] The video encoder 200 and the video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques herein are implemented partially in software, the device can store software instructions on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the video encoder 200 and the video decoder 300 can be included within one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (codec) in their respective devices. Devices including the video encoder 200 and / or the video decoder 300 can comprise integrated circuits, microprocessors, and / or wireless communication devices such as cellular telephones.

[0034]

[0047] The video encoder 200 and the video decoder 300 may operate in accordance with a video coding standard such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC), or an extension thereto such as multi-view and / or scalable video coding extensions. Alternatively, the video encoder 200 and the video decoder 300 may operate in accordance with other proprietary or industry standards such as ITU-T H.266, also known as Versatile Video Coding (VVC). The most recent draft of the VVC standard is described in Bross et al., "Versatile Video Coding (Draft 8)", Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 17th Meeting: Brussels, BE, January 7-17, 2020, JVET-Q2001-vE (hereinafter referred to as "VVC Draft 8"). Alternatively, the video encoder 200 and the video decoder 300 may operate in accordance with Essential Video Coding (EVC). However, the techniques of the present disclosure are not limited to any particular coding standard.

[0035]

[0048] Generally, video encoder 200 and video decoder 300 may perform block-based coding of pictures. The term "block" generally refers to a structure that includes data to be processed (e.g., to be encoded, decoded, or used in other ways in the encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, video encoder 200 and video decoder 300 may code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, instead of coding red, green, and blue (RGB) data for samples of a picture, video encoder 200 and video decoder 300 may code a luminance component and a chrominance component, where the chrominance component may include chrominance components for both a red phase and a blue phase. In some examples, video encoder 200 converts received data in RGB format to a YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and postprocessing units (not shown) may perform these conversions.

[0036]

[0049] In the present disclosure, reference may generally be made to the coding (e.g., encoding and decoding) of pictures, including the process of encoding or decoding picture data. Similarly, in the present disclosure, reference may be made to the coding of blocks of pictures, including the process of encoding or decoding block data, e.g., prediction and / or residual coding. An encoded video bitstream generally includes a series of values for syntax elements representing coding decisions (e.g., coding modes) and the partitioning of a picture into blocks. Thus, reference to coding a picture or a block should generally be understood as coding the values of the syntax elements that form the picture or the block.

[0037]

[0050] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as video encoder 200) divides a coding tree unit (CTU) into CUs according to a quad tree structure. That is, the video coder divides the CTU and the CU into four equal non-overlapping squares, and each node of the quad tree has either zero or four child nodes. A node without child nodes may be called a "leaf node", and the CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coder may further divide the PU and the TU. For example, in HEVC, a residual quad tree (RQT) represents the division of the TU. In HEVC, the PU represents inter-prediction data, while the TU represents residual data. An intra-predicted CU includes intra-prediction information such as an intra-mode indication.

[0038]

[0051] As another example, video encoder 200 and video decoder 300 may be configured to operate according to VVC. According to VVC, a video coder (such as video encoder 200) divides a picture into a plurality of coding tree units (CTUs). Video encoder 200 may divide the CTU according to a tree structure, such as a quad tree binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concept of multiple division types, such as the separation between the CU, PU, and TU in HEVC. The QTBT structure includes two levels: a first level divided according to a quad tree division and a second level divided according to a binary tree division. The root node of the QTBT structure corresponds to the CTU. The leaf node of the binary tree corresponds to a coding unit (CU).

[0039]

[0052] In the MTT partitioning structure, blocks can be partitioned using a quad tree (QT) partition, a binary tree (BT) partition, and one or more types of triple tree (TT) (also called ternary tree (TT)) partitions. A triple tree or ternary tree partition is a partition in which a block is split into three sub-blocks. In some examples, the triple tree or ternary tree partition splits the block into three sub-blocks without splitting the original block through its center. The partition types (e.g., QT, BT, and TT) in the MTT can be symmetric or asymmetric.

[0040]

[0053] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance and chrominance components. In other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for both chrominance components (or two QTBT / MTT structures for each chrominance component).

[0041]

[0054] Video encoder 200 and video decoder 300 may be configured to use a quad tree partition, a QTBT partition, an MTT partition, or other partitioning structures according to HEVC. For the purpose of description, the description of the techniques of the present disclosure is presented with respect to the QTBT partition. However, it should be understood that the techniques of the present disclosure may also be applicable to video coders configured to use a quad tree partition or similarly other types of partitions.

[0042]

[0055] In some examples, a CTU includes a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or a CTB of samples of a picture coded using a syntax structure with three separate color planes for a monochrome picture or samples. A CTB can be an N×N block of samples for some value N such that the splitting of components to the CTB is in segments. A component is an array or a single sample from one of the three arrays (luma and two chroma) that make up a picture in 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample of an array that makes up a picture in monochrome format. In some examples, a coding block is an M×N block of samples for some values M and N such that the splitting of CTBs to the coding block is in segments.

[0043]

[0056] Blocks (e.g., CTUs or CUs) can be grouped in various ways in a picture. As an example, a brick can refer to a rectangular region of CTU rows within a particular tile in a picture. A tile can be a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile column refers to a rectangular region of CTUs having a height equal to the height of the picture and a width specified by a syntax element (such as in a picture parameter set). A tile row refers to a rectangular region of CTUs having a height specified by a syntax element (such as in a picture parameter set) and a width equal to the width of the picture.

[0044]

[0057] In some examples, a tile can be divided into a plurality of bricks, each of which can include one or more CTU rows within the tile. Tiles that are not divided into a plurality of bricks may also be referred to as bricks. However, a brick that is a true subset of a tile may not be referred to as a tile.

[0045]

[0058] The bricks in the picture can also be placed in slices. A slice can be an integer number of bricks of a picture that may be contained solely within a single network abstraction layer (NAL) unit. In some examples, a slice contains either some complete tiles, or only a continuous sequence of complete bricks of one tile.

[0046]

[0059] In the present disclosure, for example, "N×N" and "N by N" may be used interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) with respect to vertical and horizontal dimensions, such as 16×16 samples or 16 by 16 samples. Generally, a 16×16 CU has 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non - negative integer value. The samples in a CU can be arranged in rows and columns. Moreover, a CU does not necessarily have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can comprise N×M samples, where M is not necessarily equal to N.

[0047]

[0060] The video encoder 200 encodes the video data of a CU that represents prediction and / or residual information, as well as other information. The prediction information indicates how the CU should be predicted to form a predicted block of the CU. The residual information generally represents the sample - by - sample difference between the samples of the CU before encoding and the predicted block.

[0048]

[0061] To predict a CU, video encoder 200 can generally form a prediction block of the CU through inter prediction or intra prediction. Inter prediction generally refers to predicting a CU from data of a previously coded picture, while intra prediction generally refers to predicting a CU from previously coded data of the same picture. To perform inter prediction, video encoder 200 can use one or more motion vectors to generate a prediction block. Video encoder 200 can generally perform a motion search to identify a reference block that exactly matches the CU, for example, with respect to the difference between the CU and the reference block. Video encoder 200 can calculate a difference metric using sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block exactly matches the current CU. In some examples, video encoder 200 can predict the current CU using uni-directional prediction or bi-directional prediction.

[0049]

[0062] Some examples of VVC also provide an affine motion compensation mode that can be regarded as an inter prediction mode. In the affine motion compensation mode, video encoder 200 can determine two or more motion vectors representing non-translational motion, such as zoom in or out, rotation, perspective motion, or other irregular motion types.

[0050]

[0063] To perform intra prediction, video encoder 200 may select an intra prediction mode to generate a prediction block. Some examples of VVC provide 67 intra prediction modes, including various directional modes, as well as planar mode and DC mode. Generally, video encoder 200 selects an intra prediction mode that describes adjacent samples for the current block from which the samples of the current block (e.g., the block of the CU) are to be predicted. Such samples are generally above, above and to the left, or to the left of the current block in the same picture as the current block, assuming that video encoder 200 codes CTUs and CUs in raster scan order (from left to right, from top to bottom).

[0051]

[0064] Video encoder 200 encodes data representing the prediction mode for the current block. For example, in the inter prediction mode, video encoder 200 may encode which of the various available inter prediction modes is used, as well as data representing the motion information of the corresponding mode. For example, in uni - directional or bi - directional inter prediction, video encoder 200 may encode motion vectors using advanced motion vector prediction (AMVP) or merge mode. Video encoder 200 may use a similar mode to encode the motion vectors of the affine motion compensation mode.

[0052]

[0065] Following prediction such as intra prediction or inter prediction of a block, video encoder 200 may calculate residual data for the block. Residual data such as a residual block represents the per-sample difference between the block and a predicted block for the block formed using the corresponding prediction mode. Video encoder 200 may apply one or more transforms to the residual block to generate data transformed in a transform domain rather than a sample domain. For example, video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Further, video encoder 200 may apply a second transform following the first transform, such as a mode-dependent non-separable second-order transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT). Video encoder 200 generates transform coefficients following the application of one or more transforms.

[0053]

[0066] As described above, following any transform for generating transform coefficients, video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to a process of quantizing transform coefficients to reduce as much as possible the amount of data used to represent the transform coefficients and perform further compression. By performing the quantization process, video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, video encoder 200 may truncate an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 may perform a right shift of each bit of the value to be quantized.

[0054]

[0067] Following quantization, video encoder 200 may scan the transform coefficients to generate a one-dimensional vector from the two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place higher energy (and thus lower frequency) transform coefficients in front of the vector and lower energy (and thus higher frequency) transform coefficients behind the vector. In some examples, video encoder 200 may scan the quantized transform coefficients using a predefined scan order to generate a serialized vector and then entropy code the quantized transform coefficients of the vector. In other examples, video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 may entropy code the syntax elements that describe the one-dimensional vector, for example, according to context adaptive binary arithmetic coding (CABAC). Such syntax elements may include the last significant transform coefficient as described in the present disclosure. Video encoder 200 may also entropy code the values of the syntax elements that describe the metadata associated with the encoded video data that video decoder 300 uses when decoding the video data.

[0055]

[0068] To perform CABAC, video encoder 200 may assign a context within a context model to the symbol to be transmitted. The context may be related to, for example, whether the adjacent value of the symbol is a zero value. The probability determination may be based on the context assigned to the symbol.

[0056]

[0069] The video encoder 200 can further generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, for example, within other syntax data such as picture headers, block headers, slice headers, or sequence parameter sets (SPS), picture parameter sets (PPS), or video parameter sets (VPS), for the video decoder 300. The video decoder 300 can similarly decode such syntax data to determine how to decode the corresponding video data.

[0057]

[0070] In this way, the video encoder 200 can generate an encoded video bitstream, for example, including picture partitioning into blocks (e.g., CUs) and syntax elements that describe prediction and / or residual information for the blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0058]

[0071] Generally, the video decoder 300 performs a reverse process of what was performed by the video encoder 200 to decode the encoded video data in the bitstream. For example, the video decoder 300 can use CABAC in a manner that is the reverse but substantially similar to the CABAC encoding process of the video encoder 200 to decode the values of the syntax elements in the bitstream. The syntax elements can define partitioning information for partitioning the picture into CTUs to define the CUs of the CTUs and the partitioning of each CTU according to a corresponding partitioning structure such as a QTBT structure. The syntax elements can further define prediction and residual information for blocks (e.g., CUs) of the video data.

[0059]

[0072] Residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inverse quantize and inverse transform the quantized transform coefficients of a block in order to reproduce the residual block of the block. The video decoder 300 uses the signaling prediction mode (intra or inter prediction) and associated prediction information (e.g., motion information for inter prediction) to form the prediction block of the block. The video decoder 300 can then (per sample) combine the prediction block and the residual block to reproduce the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the boundaries of the block.

[0060]

[0073] As described above, the video encoder 200 and the video decoder 300 may apply CABAC encoding and decoding to the values of the syntax elements. To apply CABAC encoding to a syntax element, the video encoder 200 may binarize the value of the syntax element to form a series of one or more bits called "bins". Each bin may be associated with a corresponding bin index (binIdx). In addition, the video encoder 200 may identify a coding context. The coding context may identify the probability of a bin having a particular value. For example, the coding context may indicate a probability of 0.7 for coding a bin with a value of 0 and a probability of 0.3 for coding a bin with a value of 1. After identifying the coding context, the video encoder 200 may divide the interval into a lower sub-interval and an upper sub-interval. One of the sub-intervals may be associated with the value 0 and the other sub-interval may be associated with the value 1. The width of the sub-interval may be proportional to the probability indicated for the associated value by the identified coding context. If the bin of the syntax element has a value associated with the lower sub-interval, the encoded value may be equal to the lower boundary of the lower sub-interval. If the same bin of the syntax element has a value associated with the upper sub-interval, the encoded value may be equal to the lower boundary of the upper sub-interval. To encode the next bin of the syntax element, the video encoder 200 may repeat these steps in the interval that is the sub-interval associated with the value of the encoded bit. When the video encoder 200 repeats these steps for the next bin, the video encoder 200 may use a modified probability based on the identified coding context and the probability indicated by the actual value of the encoded bin.

[0061]

[0074] When the video decoder 300 performs CABAC decoding on the value of a syntax element, the video decoder 300 may identify a coding context. The video decoder 300 may then divide the interval into a lower sub-interval and an upper sub-interval. One of the sub-intervals may be associated with the value 0, and the other sub-interval may be associated with the value 1. The width of the sub-interval may be proportional to the probability indicated for the associated value by the identified coding context. If the encoded value is within the lower sub-interval, the video decoder 300 may decode a bin having the value associated with the lower sub-interval. If the encoded value is within the upper sub-interval, the video decoder 300 may decode a bin having the value associated with the upper sub-interval. To decode the next bin of the syntax element, the video decoder 300 may repeat these steps with the interval that is the sub-interval containing the encoded value. When the video decoder 300 repeats these steps for the next bin, the video decoder 300 may use a modified probability based on the identified coding context and the probability indicated by the decoded bin. The video decoder 300 may then inverse-binarize the bins to restore the value of the syntax element.

[0062]

[0075] In some cases, video encoder 200 may encode bins using bypass CABAC coding. Performing bypass CABAC coding on a bin may not be computationally more expensive than performing normal CABAC coding on the bin. Further, performing bypass CABAC coding may enable higher degrees of parallelization and throughput. Bins encoded using bypass CABAC coding may be referred to as “bypass bins”. Grouping bypass bins together may increase the throughput between video encoder 200 and video decoder 300. A bypass CABAC coding engine may be able to code several bins in a single cycle, while a normal CABAC coding engine may only be able to code a single bin per cycle. The bypass CABAC coding engine may be simpler because it does not select a context and assumes a probability of 1 / 2 for both symbols (0 and 1). Thus, in bypass CABAC coding, the interval is split directly in half.

[0063]

[0076] In this disclosure, generally, reference may be made to “signaling” certain information, such as a syntax element. The term “signaling” may generally refer to the communication of the values of syntax elements and / or other data used to decode the encoded video data. That is, video encoder 200 may signal the value of a syntax element in the bitstream. Generally, signaling refers to generating a value in the bitstream. As described above, source device 102 may transfer the bitstream to destination device 116 in real time or non-real time as may occur when storing syntax elements in storage device 112 for later retrieval by destination device 116.

[0064]

[0077] FIG. 2A and FIG. 2B are conceptual diagrams showing an exemplary quad-tree binary-tree (QTBT) structure 130 and a corresponding coding tree unit (CTU) 132. Solid lines represent quad-tree splitting, and dotted lines indicate binary-tree splitting. At each split (i.e., non-leaf) node of the binary tree, one flag is signaled to indicate which splitting type (i.e., horizontal or vertical) is used. Here, in this example, 0 indicates horizontal splitting and 1 indicates vertical splitting. In quad-tree splitting, since a quad-tree node splits a block horizontally and vertically into four sub-blocks of equal size, there is no need to indicate the splitting type. Thus, the video encoder 200 can encode and the video decoder 300 can decode the syntax elements (such as splitting information) at the region tree level of the QTBT structure 130 (i.e., solid lines) and the syntax elements (such as splitting information) at the prediction tree level of the QTBT structure 130 (i.e., dashed lines). The video encoder 200 can encode and the video decoder 300 can decode video data such as prediction and transform data for the CUs represented by the terminal leaf nodes of the QTBT structure 130.

[0065]

[0078] Generally, the CTU 132 in FIG. 2B can be associated with parameters that define the size of the blocks corresponding to the nodes of the QTBT structure 130 at the first and second levels. These parameters can include the CTU size (representing the size of the CTU 132 in the sample), the minimum quad-tree size (MinQTSize representing the minimum allowable quad-tree leaf node size), the maximum binary-tree size (MaxBTSize representing the maximum allowable binary-tree root node size), the maximum binary-tree depth (MaxBTDepth representing the maximum allowable binary-tree depth), and the minimum binary-tree size (MinBTSize representing the minimum allowable binary-tree leaf node size).

[0066]

[0079] The root node of the QTBT structure corresponding to the CTU may have four child nodes at the first level of the QTBT structure, and each of them may be partitioned according to the quad tree partitioning. That is, the nodes at the first level are either leaf nodes (having no child nodes) or have four child nodes. An example of the QTBT structure 130 represents nodes including a parent node and child nodes having solid lines for the branches. If the nodes at the first level are not larger than the maximum allowable binary tree root node size (MaxBTSize), the nodes may be further partitioned by respective binary trees. The binary tree splitting of one node may be repeated until the nodes resulting from the splitting reach the minimum allowable binary tree leaf node size (MinBTSize) or the maximum allowable binary tree depth (MaxBTDepth). An example of the QTBT structure 130 represents nodes having dashed lines for the branches. The binary tree leaf nodes are called coding units (CUs), and the CUs are used for prediction (e.g., intra-picture or inter-picture prediction) and transformation without further partitioning. As described above, the CUs may also be called "video blocks" or "blocks".

[0067]

[0080] In an example of the QTBT partitioning structure, the CTU size is set as 128×128 (luma samples and two corresponding 64×64 chroma samples), MinQTSize is set as 16×16, MaxBTSize is set as 64×64, MinBTSize is set as 4 (for both width and height), and MaxBTDepth is set as 4. The quad-tree partitioning is first applied to the CTU to generate quad-tree leaf nodes. The quad-tree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the quad-tree leaf node is 128×128, the leaf quad-tree node is not further split by the binary tree since its size exceeds MaxBTSize (i.e., 64×64 in this example). In other cases, the quad-tree leaf node is further partitioned by the binary tree. Thus, the quad-tree leaf node is also the root node for the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further splitting is not permitted. When a binary tree node has a width equal to MinBTSize (4 in this example), it implies that further vertical splitting is not permitted. Similarly, a binary tree node with a height equal to MinBTSize implies that further horizontal splitting is not permitted for that binary tree node. As described above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further partitioning.

[0068]

[0081] At the 124th MPEG meeting in Macau, China, requirements for a new video codec were issued. The submitted CfP responses were evaluated at the 125th MPEG meeting in Marrakech, Morocco, and the technology from proposal m46354 was selected to form the basis of the working draft and test model of the Essential Video Coding standard. The following sections of this disclosure provide an explanation of the transform coding method utilized in MPEG5 EVC and implemented in ETM5.0

[0082] The Discrete Cosine Transform (DCT2) transform is applied to the residual block between the original block and the corresponding prediction block, as is done by conventional hybrid video codecs. To support a 64×64 pipeline, the maximum allowable transform size is set to 64. If the length of the side of the CU is larger than the maximum transform size, this side is automatically split into two segments

[0069]

[0083] In addition to the normal DCT2 transform, the Adaptive Transform Selection (ATS) method can be used for both intra prediction and inter prediction cases. Table 1 shows the basis functions of the kernel design for adaptive transform selection

[0070]

Table 1

[0071]

[0084] ATS is applied to blocks where both the width and height are smaller than 32 block sizes. If the length of the width or height exceeds 32 pixels, ATS is considered not to be applied to the block

[0072]

[0085] In the case of an intracoded block, the flag is used to signal to the video decoder 300 whether ATS is applied. When the video encoder 200 selects the use of ATS in the CU as the core transform, the video encoder 200 also signals two more flags to the video decoder 300 to indicate what type of transform is used for the horizontal direction and the vertical direction, respectively. The value 0 indicates that DST-7 is used, and the value 1 indicates that DCT-8 is used.

[0073]

[0086] In the case of an inter-predicted CU with a residual (i.e., an inter-predicted CU with a residual block), the video encoder 200 may signal whether the entire residual block should be decoded or a sub-part of the residual block should be decoded. When only a sub-part of the residual block is to be coded, that sub-part of the residual block is encoded with the inferred transform type, and the other sub-parts of the residual block are zeroed out. The sub-part position information and the corresponding transform type are shown in FIG. 3. FIG. 3 is a conceptual diagram showing the adaptive transform selection for the inter-coded block 150. In the example of FIG. 3, the sub-part 152 containing the residual information is shown as A152. The sub-part containing the residual information can be 1 / 2 or 1 / 4 the size of the current CU. ATS is enabled for CUs where both the width and height are 64 or less. FIG. 3 also shows that instead of signaling the transform type performed for the intracoded CU, the transform type is derived based on the position of the sub-blocks. For example, the horizontal transform and the vertical transform for the position 0 sub-part are DCT-8 and DST-7, respectively. When at least one side of the residual TU is larger than 32, the transform is set as DCT-2.

[0074]

[0087] After the transformation is performed, scalar quantization is applied to the transformed coefficients. The range of the quantization parameter (QP) can be from 0 to 51, and the scaling factor (SF) corresponding to each QP is defined by a look-up table. The video encoder 200 can scan the transform coefficients of the coded blocks after quantization in a predefined scanning pattern and entropy code them. To further utilize the statistical properties of the transform coefficients, a bit-plane based coefficient coding method, so-called advanced coefficient coding (ADCC), is utilized in the main profile of EVC instead of the currently used run-length coding method. The ADCC method uses the following design elements.

[0075] 1. Fixed zigzag scanning pattern.

[0076] 2. Signaling the coordinates of the last non-0 transform coefficient in the scanning order.

[0077] 3. Parsing the transform coefficients in chunks of 16.

[0078] 4. Signaling the transform coefficients within each processing chunk as a sequence of significance and level flags, sign flag, and residual level.

[0079]

[0088] Among these symbols (i.e., significance and level flags, sign flag, and residual level), the bins of sigMapFlag, flagLevelA, and flagLevelB are coded with an adaptive context model, and the bins of signFlag and the binarized levelRem are coded through a bypass mode. The value of sigMapFlag can indicate whether the corresponding transform coefficient is a significant transform coefficient. The value of flagLevelA can indicate whether the level of the corresponding transform coefficient is greater than or equal to level A. The value of flagLevelB can indicate whether the level of the corresponding transform coefficient is greater than or equal to level B. The value of signFlag indicates whether the level of the transform coefficient is positive or negative. The value of levelRem indicates the residue of the level of the transform coefficient.

[0080]

[0089] To reduce the number of context-coded bins, the explicit flagLevelA and flagLevelB are binarized using Golomb codes and adaptively switched to levelRem coding, which is coded in bypass mode with equal probability. To improve throughput, the number of explicitly coded symbols flagLevelA and flagLevelB within a coded chunk is limited. Only the first N flagLevelA symbols and the first M flagLevelB symbols are coded. The explicit coding of these symbols is omitted when a specified threshold is met.

[0081]

[0090] A visualization of this method is given in FIG. 4. In particular, FIG. 4 is a conceptual diagram showing a coefficient scanning method and the last transform coefficient position 160. FIG. 4 shows the transform coefficients of block 162. The arrow indicates a zigzag scan pattern starting from the "last" significant transform coefficient with respect to the positions along the zigzag scan pattern from the DC transform coefficient 164.

[0082]

[0091] FIG. 5 is a table 170 showing coding symbols for the coefficients of a chunk of 16 coefficients. In other words, FIG. 5 gives an example of coding symbols for a chunk of 16 transform coefficients where N = 1. The symbols for which signaling is omitted are marked with an 'x' in FIG. 5.

[0083]

[0092] A part of the transform coefficient coding (i.e., a part of the transform_unit syntax structure) is defined in MPEG5 EVC as shown in Table 2 below.

[0084]

Table 2

[0085]

[0093] The residual_coding syntax structure referred to in Table 2 above can be implemented as reproduced in Table 3 below. As can be seen in Table 3, the values log2TbWidth and log2TbHeight are passed from the transform_unit syntax structure to the residual_coding syntax structure, and from the residual_coding syntax structure to the residual_coding_adv syntax structure.

[0086]

Table 3

[0087]

[0094] The residual_coding syntax structure from the transform_unit syntax structure can be implemented as shown in Table 4 below, which shows a part of the residual_coding_adv syntax structure.

[0088]

Table 4

[0089]

[0095] In Table 4 above, the syntax element last_sig_coeff_x_prefix may specify the prefix of the column (x) position of the last significant transform coefficient in the scan order within the transform block. The syntax element last_sig_coeff_y_prefix may specify the prefix of the row (y) position of the last significant transform coefficient in the scan order within the transform block. If last_sig_coeff_x_suffix does not exist, the column position (i.e., x coordinate) of the last significant transform coefficient (LastSignificantCoeffX) may be equal to the value of last_sig_coeff_x_prefix. Otherwise (if last_sig_coeff_x_suffix exists), the following may apply. LastSignificantCoeffX=(1<<((last_sig_coeff_x_prefix>>1)-1))*(2+(last_sig_coeff_x_prefix&1))+last_sig_coeff_x_suffix

[0096] Similarly, when there is no last_sig_coeff_y_suffix, the row position (i.e., y coordinate) of the last significant transform coefficient (LastSignificantCoeffY) can be equal to last_sig_coeff_y_prefix. Otherwise (when last_sig_coeff_y_suffix exists), the following may apply. LastSignificantCoeffY=(1<<((last_sig_coeff_y_prefix>>1)-1))*(2+(last_sig_coeff_y_prefix&1))+last_sig_coeff_y_suffix

[0097] A video coder (e.g., video encoder 200 or video decoder 300) may perform a process of deriving a context increment (ctxInc) for syntax elements last_sig_coeff_x_prefix and last_sig_coeff_y_prefix. The video coder may determine a context based on the context increment. In some examples, to determine a context based on the context increment, the video coder may determine a context index by adding the context increment to a context offset value (ctxIdxOffset) of a syntax element (e.g., last_sig_coeff_x_prefix and last_sig_coeff_y_prefix). The context offset value may be equal to the lowest context index value available for use with the syntax element. The following text describes the derivation process implemented in ETM5.0.

[0090] The input to this process is the variable binIdx, the color component index cIdx, and the associated transform size log2TrafoSize which is log2TbWidth for last_sig_coeff_x_prefix and log2TbHeight for last_sig_coeff_y_prefix, respectively.

[0091] The output of this process is the variable ctxInc.

[0092] The variables ctxOffset and ctxShift are derived as follows.

[0093] - If cIdx is equal to 0, the following applies.

[0094] If log2TrafoSize is less than 6, ctxOffset is set equal to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2), and ctxShift is set equal to (log2TrafoSize + 1)>>2.

[0095] - Otherwise (log2TrafoSize is 6 or greater), ctxOffset is set equal to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and ctxShift is set equal to ((log2TrafoSize + 1)>>2)<<1.

[0096] - Otherwise (cIdx is greater than 0), ctxOffset is set equal to 18, and ctxShift is set equal to Max(0,log2TrafoSize - 2)-Max(0,log2TrafoSize - 4).

[0097] The variable ctxInc is derived as follows. ctxInc=(binIdx>>ctxShift)+ctxOffset (9 - 1)

[0098] In the above text, log2TrafoSize specifies the relevant transform size, which is log2TbWidth for last_sig_coeff_x_prefix and log2TbHeight for last_sig_coeff_y_prefix, respectively. The variable log2TbWidth is equal to the base-2 logarithm of the width of the transform block. The video coder may repeat this operation for each bin of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix. The bins of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix may be the individual binary digits of the binary versions of last_sig_coeff_x_prefix and last_sig_coeff_y_prefix.

[0098]

[0099] The transform coefficient coding as specified in MPEG5 EVC does not expect coding of video signals using a color format different from 4:2:0, i.e., where the chroma component data is presented at 1 / 4 resolution compared to the luma component. Enabling support for color formats different from 4:2:0 can increase the coding efficiency for compressing videos with color formats different from 4:2:0.

[0099]

[0100] In this disclosure, techniques that may enable support for coding video signals in color formats other than the 4:2:0 color format will be described. The techniques of this disclosure may thereby increase the coding efficiency for compressing videos with color formats different from 4:2:0. The techniques of this disclosure may be used together or separately.

[0100]

[0101] According to the first technique of the present disclosure, a video coder (e.g., video encoder 200 or video decoder 300) may invoke a residual coding operation for a non-luma (cIdx not equal to 0) color component having a block size that is a function of a chroma format indicator. The following text describes the semantics in ETM5.0 of the chroma_format_idc syntax element.

[0101] chroma_format_idc specifies the chroma sampling with respect to luma sampling as specified in Clause 6.2. The value of chroma_format_idc shall be in the range of 0 to 3, inclusive.

[0102] Depending on the value of chroma_format_idc, the values of the variables SubWidthC and SubHeightC are assigned as specified in Clause 6.2, and the variable ChromaArrayType is assigned as follows.

[0103] - If chroma_format_idc is equal to 0, ChromaArrayType is set equal to 0.

[0104] - Otherwise, ChromaArrayType is set equal to chroma_format_idc.

[0105]

[0102] Furthermore, the following text from Clause 6.2 of ETM5.0 describes the color format.

[0106] The variables SubWidthC and SubHeightC are specified in Table 6-1 according to the chroma format sampling structure specified through chroma_format_idc. Other values of chroma_format_idc, SubWidthC, and SubHeightC may be specified in the future by ISO / IEC.

[0107]

Table 5

[0108] In monochrome sampling, there is only one sample array that is nominally regarded as the luma array.

[0109] In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.

[0110] In 4:2:2 sampling, each of the two chroma arrays has the same height as the luma array and half the width.

[0111] In 4:4:4 sampling, each of the two chroma arrays has the same height and width as the luma array.

[0112]

[0103] According to the first technique of the present disclosure, the text of ETM5.0 can be modified as shown in Table 5 below in order to consider different values of SubWidthC and SubHeightC in different color formats determined in Table 6-1. In the present disclosure, the proposed modification to the transform_unit syntax structure of EVC is indicated by <!---->…<!--!-->tags.

[0113]

Table 6

[0114]

[0104] Since the values of SubWidthC and SubHeightC depend on the color format and TrafoLog2Width and TrafoLog2Height are modified in Table 5 based on SubWidthC and SubHeightC, the values of TrafoLog2Width and TrafoLog2Height are modified based on the color format. Therefore, the correct values of TrafoLog2Width and TrafoLog2Height can be used in the residual_coding and residual_coding_adv syntax structures shown in Tables 3 and 4 for various color formats. Using the correct values of TrafoLog2Width and TrafoLog2Height in the residual_coding_adv syntax structure may enable correct syntax elements to be signaled for various color formats.

[0115]

[0105] Thus, in an exemplary configuration using the first technique of the present disclosure, a video coder (e.g., video encoder 200 or video decoder 300) may determine parameters of a residual coding operation based on a chroma format indicator (chroma_format_idc) applicable to a block of video data. For example, the video coder may determine one parameter of the residual coding operation as TrafoLog2Width - SubWidthC + 1 and may determine another parameter of the residual coding operation as TrafoLog2Height - SubHeightC + 1. The video coder may then perform the residual coding operation based on the determined parameters of the residual coding operation to code the residual data of the non-luma components of the block.

[0116]

[0106] The second technique of the present disclosure can improve the performance of coefficient coding for non-luma (cIdx not equal to 0) by enabling more efficient context modeling. According to the second technique of the present disclosure, a luma context model (i.e., a context model used for components in a full-resolution video signal such as luma components) can be enabled for context coding syntax elements in chroma components of a color format different from 4:2:0, i.e., with a color_format_idc not equal to 1. The proposed changes to ETM5.0 according to the second technique of the present disclosure are <!---->…<!--!-->indicated by tags.

[0117] Derivation process of ctxInc for syntax elements last_sig_coeff_x_prefix and last_sig_coeff_y_prefix The inputs to this process are the variable binIdx, the color component index cIdx, and the associated transform size log2TrafoSize which is log2TbWidth for last_sig_coeff_x_prefix and log2TbHeight for last_sig_coeff_y_prefix respectively.

[0118] The output of this process is the variable ctxInc.

[0119] <!---->The variable enableLumaModel is set to false and is modified as follows.

[0120] - When chromaArrayType is equal to 0 or 3, the variable enableLumaModel is set to true.

[0121] - Otherwise, when chromaArrayType is equal to 1 or 2 and cIdx is equal to 0, the variable enableLumaModel is set to true.

[0122] - Instead, when chromaArrayType is equal to 2, cIdx is not equal to 0, and the syntax element to be parsed is last_sig_coeff_y_prefix, the variable enableLumaModel is set equal to true. <!--!--> The variables ctxOffset and ctxShift are derived as follows.

[0123] - <!---->When enableLumaModel is equal to true <!--!-->the following applies.

[0124] - <^^>When log2TrafoSize is less than 6, ctxOffset is set equal to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2), and ctxShift is set equal to (log2TrafoSize + 1)>>2.

[0125] - Otherwise (log2TrafoSize is 6 or greater), ctxOffset is set equal to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and ctxShift is set equal to ((log2TrafoSize + 1)>>2)<<1. < / ^^> - Otherwise (<!---->when enableLumaModel is equal to false <!--!-->), ctxOffset is set equal to 18, and ctxShift is set equal to Max(0,log2TrafoSize - 2)-Max(0,log2TrafoSize - 4).

[0126] The variable ctxInc is derived as follows. ctxInc=(binIdx>>ctxShift)+ctxOffset (9 - 2)

[0107] Thus, as shown in the above text, the luma context model (i.e., the text marked with <^^>…< / ^^> tags) can be used for luma components and also, in some situations, for chroma components. Thus, according to the second technique of the present disclosure, a video coder (e.g., video encoder 200 or video decoder 300) may use a context model to derive a context increment of a first syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix). The first syntax element indicates a prefix of the x or y coordinate of the last significant transform coefficient of the luma component of a block of a picture of video data. Further, the video coder may apply CABAC to a bin of the first syntax element using the context determined based on the context increment of the first syntax element. The video coder may use the same or a different context model to derive a context increment of a second syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix) depending, in part, on the color format of the picture. The second syntax element indicates a prefix of the x or y coordinate of the last significant transform coefficient of the chroma component of the block. The video coder may apply CABAC to a bin of the second syntax element using the context determined based on the context increment of the second syntax element.

[0127]

[0108] Thus, in some examples, the video encoder 200 may determine which context model to use from a first context model and a second context model (e.g., defined by the value of the variable enableLumaModel) based on the color format of the picture to determine a context increment (ctxInc) of a syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix) indicating a prefix of the x or y coordinate of the last significant transform coefficient of the color component of a block of a picture of the video data. The video encoder 200 may encode the bin of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element. In some examples, as described above, the index of the context (with respect to a predefined set of contexts) may be determined from the sum of the context increment and an offset value that may be predefined. Similarly, the video decoder 300 may determine which context model to use from a first context model and a second context model based on the color format of the picture to determine a context increment (ctxInc) of a syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix) indicating a prefix of the x or y coordinate of the last significant transform coefficient of the color component of a block of a picture of the video data. The video decoder 300 may encode the bin of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.

[0128]

[0109] In some examples, using the first context model comprises performing one of setting the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and setting the context shift to (log2TrafoSize + 1)>>2, based on the base-2 logarithm value of the transform size of the block being less than 6, where log2TrafoSize indicates the base-2 logarithm value of the transform size. Alternatively, based on the base-2 logarithm value of the transform size being 6 or greater, the video coder may set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and set the context shift to ((log2TrafoSize + 1)>>2)<<1. The video coder may determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset, where the applicable syntax element is the first syntax element or the second syntax element. In some examples, using the second context model comprises setting the context offset ctxOffset of the second syntax element to 18 and setting the context shift ctxShift of the applicable syntax element to Max(0,log2TrafoSize - 2)-Max(0,log2TrafoSize - 4), where Max indicates the maximum value function and log2TrafoSize indicates the base-2 logarithm value of the transform size of the block. Further, using the second context model may comprise determining the context increment of the second syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the applicable syntax element.

[0129]

[0110] FIG. 6 is a block diagram showing an exemplary video encoder 200 that can implement the techniques of the present disclosure. FIG. 6 is provided for illustrative purposes and should not be considered as limiting the techniques widely exemplified and described in the present disclosure. For illustrative purposes, in the present disclosure, the video encoder 200 according to the techniques of VVC (ITU-T H.266 under development), HEVC (ITU-T H.265), and EVC will be described. However, the techniques of the present disclosure can be implemented by video encoding devices configured according to other video coding standards.

[0130]

[0111] In the example of FIG. 6, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transformation processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy encoding unit 220. Any or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transformation processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy encoding unit 220 can be implemented in one or more processors or in a processing circuit. For example, the units of the video encoder 200 can be implemented as one or more circuits or logic elements, as part of a hardware circuit, or as part of a processor of an FPGA or an ASIC. Moreover, the video encoder 200 can include additional or alternative processors or processing circuits for implementing these and other functions.

[0131]

[0112] The video data memory 230 can store video data to be encoded by components of the video encoder 200. The video encoder 200 can receive, for example, video data stored in the video data memory 230 from a video source 104 (FIG. 1). The DPB 218 can function as a reference picture memory that stores reference video data for use in prediction of subsequent video data by the video encoder 200. The video data memory 230 and the DPB 218 can be formed by any of various memory devices, such as a dynamic random access memory (DRAM) including a synchronous DRAM (SDRAM), a magnetoresistive RAM (MRAM), a resistive RAM (RRAM (registered trademark)), or other types of memory devices. The video data memory 230 and the DPB 218 can be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 can be on-chip with other components of the video encoder 200 or off-chip with respect to those components, as illustrated.

[0132]

[0113] In the present disclosure, a reference to the video data memory 230 should not be construed as being limited to the memory internal to the video encoder 200 or, unless otherwise specifically stated, limited to the memory external to the video encoder 200. Rather, a reference to the video data memory 230 is to be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., video data of the current block to be encoded). The memory 106 of FIG. 1 can also provide temporary storage of outputs from various units of the video encoder 200.

[0133] [

[0114] ]The various units of FIG. 6 are shown to assist in understanding the operations performed by video encoder 200. The units may be implemented as fixed function circuitry, programmable circuitry, or a combination thereof. Fixed function circuitry refers to circuitry that provides a specific function and is preset with respect to the operations that may be performed. Programmable circuitry refers to circuitry that may be programmed to perform various tasks and to provide a flexible function in the operations that may be performed. For example, programmable circuitry may execute software or firmware that operates the programmable circuitry in a manner defined by instructions of the software or firmware. Fixed function circuitry may execute software instructions (e.g., to receive or output parameters), but the type of operations performed by the fixed function circuitry is generally invariant. In some examples, one or more of the units may be separate circuit blocks (fixed function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0134] [

[0115] ]Video encoder 200 may include a programmable core formed from an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or programmable circuitry. In an example where the operations of video encoder 200 are performed using software executed by programmable circuitry, memory 106 (FIG. 1) may store the instructions of the software (e.g., object code) that video encoder 200 receives and executes, or another memory (not shown) within video encoder 200 may store such instructions.

[0135] [

[0116] ]Video data memory 230 is configured to store received video data. Video encoder 200 may retrieve a picture of the video data from video data memory 230 and provide the video data to residual generation unit 204 and mode selection unit 202. The video data in video data memory 230 may be raw video data to be encoded.

[0136]

[0117] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra prediction unit 226. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. By way of example, the mode selection unit 202 may include a palette unit, an intra block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, and the like.

[0137]

[0118] Generally, the mode selection unit 202 coordinates a plurality of encoding paths in order to test combinations of encoding parameters and the resulting rate distortion values for such combinations. The encoding parameters may include the partitioning of CTUs into CUs, the prediction mode for the CUs, the transform type for the residual data of the CUs, the quantization parameter for the residual data of the CUs, and the like. The mode selection unit 202 may ultimately select a combination of encoding parameters that has a rate distortion value better than other tested combinations.

[0138]

[0119] The video encoder 200 may partition a picture fetched from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 may partition the CTUs of the picture according to a tree structure, such as the HEVC QTBT structure or quad tree structure described above. As described above, the video encoder 200 may form one or more CUs from partitioning the CTUs according to the tree structure. Such CUs may also be commonly referred to as "video blocks" or "blocks".

[0139]

[0120] Generally, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for the current block (e.g., the current CU, or in HEVC, the overlapping part of the PU and TU). For the inter prediction of the current block, the motion estimation unit 222 may perform a motion search to identify one or more exactly matching reference blocks in one or more reference pictures (e.g., one or more previously coded pictures stored in the DPB 218). In particular, the motion estimation unit 222 may calculate a value representing how similar a potential reference block is to the current block, according to, for example, the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The motion estimation unit 222 may generally perform these calculations using the sample-by-sample differences between the current block and the reference block under consideration. The motion estimation unit 222 may identify the reference block having the lowest value resulting from these calculations, which indicates the reference block that most closely matches the current block.

[0140]

[0121] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in a current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, in uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, while in bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then generate a prediction block using the motion vectors. For example, the motion compensation unit 224 may retrieve data of a reference block using the motion vectors. As another example, when the motion vectors have sub-sample accuracy, the motion compensation unit 224 may interpolate values for the prediction block according to one or more interpolation filters. Further, in bi-directional inter prediction, the motion compensation unit 224 may retrieve data for two reference blocks identified by the respective motion vectors and combine the retrieved data, for example, through sample-by-sample averaging or weighted averaging.

[0141]

[0122] As another example, for intra prediction, or intra prediction coding, the intra prediction unit 226 may generate a prediction block from samples adjacent to the current block. For example, in a directional mode, the intra prediction unit 226 may generally mathematically combine the values of adjacent samples and populate these calculated values in a direction defined across the current block to generate the prediction block. As another example, in a DC mode, the intra prediction unit 226 may calculate an average of adjacent samples for the current block and generate the prediction block to include this obtained average for each sample of the prediction block.

[0142]

[0123] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the raw, unencoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block of the current block. In some examples, the residual generation unit 204 may also determine the differences between the sample values in the residual block in order to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.

[0143]

[0124] In an example where the mode selection unit 202 divides a CU into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As shown above, the size of a CU may refer to the size of the luma coding block of the CU, and the size of a PU may refer to the size of the luma prediction unit of the PU. Assuming that the size of a particular CU is 2N×2N, the video encoder 200 may support PU sizes of 2N×2N or N×N for intra prediction and symmetric PU sizes of 2N×2N, 2N×N, N×2N, N×N, or the like for inter prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitioning for PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.

[0144]

[0125] In an example where the mode selection unit 202 does not further divide the CU into PUs, each CU may be associated with a luma coding block and a corresponding chroma coding block. As described above, the size of the CU may refer to the size of the luma coding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2N×2N, 2N×N, or N×2N.

[0145]

[0126] As some examples, in the case of other video coding techniques such as intra block copy mode coding, affine mode coding, and linear model (LM) mode coding, the mode selection unit 202 generates a prediction block for the current block being coded via respective units related to the coding technique. In some examples, such as palette mode coding, the mode selection unit 202 may not generate a prediction block and instead may generate a syntax element indicating the manner in which the block should be reconstructed based on the selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to be coded to the entropy coding unit 220.

[0146]

[0127] As described above, the residual generation unit 204 receives the video data of the current block and the corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0147]

[0128] The transformation processing unit 206 applies one or more transformations to the residual block to generate a block of transformation coefficients (referred to herein as the "transformation coefficient block"). The transformation processing unit 206 may apply various transformations to the residual block to form the transformation coefficient block. For example, the transformation processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transformation processing unit 206 may perform multiple transformations, such as a primary transform and a secondary transform, such as a rotation transform, on the residual block. In some examples, the transformation processing unit 206 does not apply a transformation to the residual block.

[0148]

[0129] The quantization unit 208 may quantize the transformation coefficients in the transformation coefficient block to generate a quantized transformation coefficient block. The quantization unit 208 may quantize the transformation coefficients of the transformation coefficient block according to the quantization parameter (QP) value associated with the current block. The video encoder 200 may adjust the degree of quantization applied to the transformation coefficient block associated with the current block by adjusting the QP value associated with the CU (e.g., via the mode selection unit 202). Quantization can result in a loss of information, and thus the quantized transformation coefficients may have lower precision than the original transformation coefficients generated by the transformation processing unit 206.

[0149]

[0130] The inverse quantization unit 210 and the inverse transformation processing unit 212 may apply inverse quantization and inverse transformation, respectively, to the quantized transformation coefficient block to reconstruct the residual block from the transformation coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (potentially with some degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add the samples of the reconstructed residual block to the corresponding samples from the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0150]

[0131] The filter unit 216 may perform one or more filter operations on the reconstructed block. For example, the filter unit 216 may perform a deblocking operation to reduce blockiness artifacts along the edge of the CU. The operation of the filter unit 216 may be skipped in some examples.

[0151]

[0132] The video encoder 200 stores the reconstructed block in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 may store the reconstructed block in the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 may store the filtered reconstructed block in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve a reference picture formed from the reconstructed (and potentially filtered) blocks from the DPB 218 to inter-predict blocks of a picture to be encoded later. In addition, the intra prediction unit 226 may use the reconstructed blocks in the DPB 218 of the current picture to intra-predict other blocks in the current picture.

[0152]

[0133] Generally, the entropy encoding unit 220 may entropy-encode syntax elements received from other functional components of the video encoder 200. For example, the entropy encoding unit 220 may entropy-encode the quantized transform coefficient blocks from the quantization unit 208. As another example, the entropy encoding unit 220 may entropy-encode prediction syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from the mode selection unit 202. The entropy encoding unit 220 may perform one or more entropy encoding operations on syntax elements, which are another example of video data, to generate entropy-encoded data. For example, the entropy encoding unit 220 may perform context-adaptive variable length coding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding operations, syntax-based context-adaptive binary arithmetic coding (SBAC) operations, probability interval partitioning entropy (PIPE) coding operations, exponential Golomb coding operations, or another type of entropy encoding operation on the data. In some examples, the entropy encoding unit 220 may operate in a bypass mode in which the syntax elements are not entropy-encoded.

[0153] According to one or more techniques of the present disclosure, the entropy encoding unit 220 may determine which context model to use from a first context model and a second context model based on the color format of the picture to determine a context increment of a syntax element indicating a prefix of the x or y coordinate of the last significant transform coefficient of the color component of the block of the picture of the video data. The entropy encoding unit 220 may encode the bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element. In some examples, the entropy encoding unit 220 may use the first context model or the second context model to derive a context increment of a syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix) indicating a prefix of the x or y coordinate of the last significant transform coefficient of the color component of the block of the picture of the video data. The entropy encoding unit 220 may apply CABAC to the bins of the first syntax element using the context determined based on the context increment of the syntax element. The entropy encoding unit 220 may use the first context model when the color component is luma or when the color component is a chroma component and the color format is different from 4:2:0. When using the first context model, the entropy encoding unit 220 may set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and set the context shift to be equal to (log2TrafoSize + 1)>>2 based on the fact that the base-2 logarithm value of the transform size of the block is less than 6, where log2TrafoSize indicates the base-2 logarithm value of the transform size.Based on the fact that the base-2 logarithm value of the transform size is 6 or more, the entropy encoding unit 220 may set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and set the context shift to be equal to ((log2TrafoSize + 1)>>2)<<1. The entropy encoding unit 220 may determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset. The applicable syntax element may be the first syntax element or the second syntax element.

[0154]

[0135] The video encoder 200 may output a bitstream including entropy-encoded syntax elements necessary for reconstructing a block of a slice or a picture. For example, the entropy encoding unit 220 may output a bitstream.

[0155]

[0136] The operations described above have been described with respect to a block. Such descriptions should be understood as operations on a luma coding block and / or a chroma coding block. As described above, in some examples, the luma coding block and the chroma coding block are the luma component and the chroma component of a CU. In some examples, the luma coding block and the chroma coding block are the luma component and the chroma component of a PU.

[0156]

[0137] In some examples, the operations performed on a luma coding block need not be repeated for a chroma coding block. As an example, the operations for identifying the motion vector (MV) and reference picture of a luma coding block need not be repeated to identify the MV and reference picture of a chroma block. Instead, the MV of the luma coding block can be scaled to determine the MV of the chroma block, and the reference picture can be the same. As another example, the intra prediction process can be the same for luma coding blocks and chroma coding blocks.

[0157]

[0138] Video encoder 200 may represent an example of a device configured to encode video data, including a memory configured to store video data and one or more processing units. The one or more processing units are implemented in circuitry and configured to determine parameters of a residual coding operation based on a chroma format indicator applicable to a block of video data. The one or more processing units may perform a residual coding operation based on the determined parameters of the residual coding operation to encode residual data of a non-luma component of the block. In some examples, the one or more processing units are configured to use a context model to entropy code a prefix of the y coordinate of the last significant transform coefficient of the luma component of the block and to use the same context model to entropy code a prefix of the y coordinate of the last significant transform coefficient of the chroma component of the block based on the block having a color format of 4:2:2.

[0158]

[0139] FIG. 7 is a block diagram showing an exemplary video decoder 300 that can implement the techniques of the present disclosure. FIG. 7 is provided for illustrative purposes and is not intended to limit the techniques broadly illustrated and described in the present disclosure. For illustrative purposes, in the present disclosure, the video decoder 300 according to the techniques of VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265) will be described. However, the techniques of the present disclosure can be implemented by a video coding device configured according to other video coding standards.

[0159]

[0140] In the example of FIG. 7, the video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 can be implemented in one or more processors or in a processing circuit. For example, the units of the video decoder 300 can be implemented as one or more circuits or logic elements, as part of a hardware circuit, or as part of a processor of an FPGA or an ASIC. Moreover, the video decoder 300 can include additional or alternative processors or processing circuits for implementing these and other functions.

[0160]

[0141] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 may include additional units for performing prediction according to other prediction modes. By way of example, the prediction processing unit 304 may include a palette unit, an intra block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, and the like. In other examples, the video decoder 300 may include a greater number, a lesser number, or different functional components.

[0161]

[0142] The CPB memory 320 may store video data such as an encoded video bitstream to be decoded by the components of the video decoder 300. The video data stored in the CPB memory 320 may be obtained, for example, from the computer-readable medium 110 (FIG. 1). The CPB memory 320 may include a CPB that stores encoded video data (such as syntax elements) from the encoded video bitstream. Also, the CPB memory 320 may store video data other than the syntax elements of the coded picture, such as temporary data representing the outputs from various units of the video decoder 300. The DPB 314 generally stores decoded pictures that the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed by any of various memory devices, such as DRAM including SDRAM, MRAM, RRAM, or other types of memory devices. The CPB memory 320 and the DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with or off-chip with respect to the other components of the video decoder 300.

[0162]

[0143] Additionally or alternatively, in some examples, the video decoder 300 may retrieve the encoded video data from the memory 120 (FIG. 1). That is, the memory 120 may store data as described above using the CPB memory 320. Similarly, the memory 120 may store instructions to be executed by the video decoder 300 when some or all of the functionality of the video decoder 300 is implemented in software to be executed by the processing circuitry of the video decoder 300.

[0163]

[0144] The various units shown in FIG. 7 are shown to assist in understanding the operations performed by the video decoder 300. The units may be implemented as fixed function circuitry, programmable circuitry, or a combination thereof. Similar to FIG. 6, fixed function circuitry refers to circuitry that provides a specific function and is preset with respect to the operations that may be performed. Programmable circuitry refers to circuitry that may be programmed to perform various tasks and to provide a flexible function in the operations that may be performed. For example, programmable circuitry may execute software or firmware that causes the programmable circuitry to operate in a manner defined by instructions of the software or firmware. Fixed function circuitry may execute software instructions (e.g., to receive or output parameters), but the type of operations performed by the fixed function circuitry is generally invariant. In some examples, one or more of the units may be discrete circuit blocks (fixed function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0164]

[0145] The video decoder 300 may include a programmable core formed from an ALU, EFU, digital circuitry, analog circuitry, and / or programmable circuitry. In an example where the operations of the video decoder 300 are performed by software executed on programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software that the video decoder 300 receives and executes.

[0165]

[0146] The entropy decoding unit 302 can receive the encoded video data from the CPB and entropy-decode the video data to reproduce the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 can generate the decoded video data based on the syntax elements extracted from the bitstream.

[0166]

[0147] Generally, the video decoder 300 reconstructs pictures block by block. The video decoder 300 can perform the reconstruction operation individually for each block (where the block being currently reconstructed, i.e., decoded, may be referred to as the "current block").

[0167]

[0148] Entropy decoding unit 302 can entropy decode the quantized transform coefficients of the quantized transform coefficient block, as well as syntax elements that define transform information such as quantization parameter (QP) and / or transform mode indication. According to one or more techniques of the present disclosure, entropy decoding unit 302 determines which context model to use from a first context model and a second context model based on the color format of the picture in order to determine the context increment of a syntax element indicating the prefix of the x or y coordinate of the last significant transform coefficient of the color component of a block of the picture of the video data. Entropy decoding unit 302 can decode the bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element. In some examples, entropy decoding unit 302 can use the first context model or the second context model to derive the context increment of a syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix) indicating the prefix of the x or y coordinate of the last significant transform coefficient of the color component of a block of the picture of the video data. Entropy decoding unit 302 can apply CABAC to the bins of the first syntax element using the context determined based on the context increment of the syntax element. Entropy decoding unit 302 can use the first context model when the color component is luma, or when the color component is a chroma component and the color format is different from 4:2:0. When using the first context model, entropy decoding unit 302 can set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and set the context shift to be equal to (log2TrafoSize + 1)>>2 based on the fact that the base-2 logarithm value of the transform size of the block is less than 6, where log2TrafoSize indicates the base-2 logarithm value of the transform size.Based on the fact that the base-2 logarithm value of the transform size is 6 or more, the entropy decoding unit 302 may set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and set the context shift to be equal to ((log2TrafoSize + 1)>>2)<<1. The entropy decoding unit 302 may determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset. The applicable syntax element may be the first syntax element or the second syntax element.

[0168]

[0149] The inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine the degree of quantization and, similarly, the degree of inverse quantization to be applied by the inverse quantization unit 306. The inverse quantization unit 306 may perform, for example, a left shift operation on a bit-by-bit basis to inverse quantize the quantized transform coefficients. The inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.

[0169]

[0150] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the transform coefficient block.

[0170]

[0151] Furthermore, the prediction processing unit 304 generates a prediction block according to the prediction information syntax element entropy-decoded by the entropy decoding unit 302. For example, when the prediction information syntax element indicates that the current block is inter-predicted, the motion compensation unit 316 may generate a prediction block. In this case, the prediction information syntax element may indicate the reference picture in the DPB 314 from which the reference block should be taken, as well as the motion vector that identifies the location of the reference block in the reference picture relative to the location of the current block in the current picture. The motion compensation unit 316 may perform an inter-prediction process in a manner that is substantially the same as that described with respect to the motion compensation unit 224 (FIG. 6).

[0171]

[0152] As another example, when the prediction information syntax element indicates that the current block is intra-predicted, the intra-prediction unit 318 may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Also in this case, the intra-prediction unit 318 may perform an intra-prediction process in a manner that is substantially the same as that described with respect to the intra-prediction unit 226 (FIG. 6). The intra-prediction unit 318 may retrieve data of adjacent samples for the current block from the DPB 314.

[0172]

[0153] The reconstruction unit 310 may reconstruct the current block using the prediction block and the residual block. For example, the reconstruction unit 310 may add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.

[0173]

[0154] The filter unit 312 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operation of the filter unit 312 is not necessarily performed in all examples.

[0174]

[0155] Video decoder 300 may store the reconstructed block in DPB 314. For example, in an example where the operation of filter unit 312 is not performed, reconstruction unit 310 may store the reconstructed block in DPB 314. In an example where the operation of filter unit 312 is performed, filter unit 312 may store the filtered and reconstructed block in DPB 314. As described above, DPB 314 may provide reference information such as samples of the current picture for intra prediction and previously decoded pictures for subsequent motion compensation to prediction processing unit 304. Moreover, video decoder 300 may output the decoded picture (e.g., decoded video) from DPB 314 for subsequent presentation on a display device such as display device 118 of FIG. 1.

[0175]

[0156] In this way, video decoder 300 represents an example of a video decoding device including a memory configured to store video data and one or more processing units, and the one or more processing units are implemented in a circuit and are configured to determine parameters of a residual coding operation based on a chroma format indicator applicable to a block of video data, and to perform a residual coding operation based on the determined parameters of the residual coding operation to decode residual data of a non-luma component of the block. In some examples, the one or more processing units may be configured to use a context model to entropy code a prefix of the y coordinate of the last significant transform coefficient of the luma component of the block, and to use the same context model to entropy decode a prefix of the y coordinate of the last significant transform coefficient of the chroma component of the block based on the block having a color format of 4:2:2.

[0176]

[0157] FIG. 8 is a flowchart showing an exemplary method for encoding a current block. The current block may comprise a current CU. Although described with respect to video encoder 200 (FIGS. 1 and 6), it should be understood that other devices may be configured to perform in a manner similar to that of FIG. 8.

[0177]

[0158] In this example, video encoder 200 first predicts the current block (350). For example, video encoder 200 may form a predicted block of the current block. Video encoder 200 may then calculate a residual block of the current block (352). To calculate the residual block, video encoder 200 may calculate the difference between the original unencoded block and the predicted block for the current block. Video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (354). Next, video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or subsequent to the scan, video encoder 200 may entropy encode the transform coefficients (358). For example, video encoder 200 may encode the transform coefficients using CAVLC or CABAC. Video encoder 200 may also entropy encode the last_sig_coeff_x_prefix and last_sig_coeff_y_prefix syntax elements using a context determined according to the techniques of the present disclosure. Video encoder 200 may then output the entropy encoded data of the block (360).

[0178]

[0159] FIG. 9 is a flowchart showing an exemplary method for decoding a current block of video data. The current block may comprise a current CU. Although described with respect to video decoder 300 (FIGS. 1 and 7), it should be understood that other devices may be configured to perform in a manner similar to that of FIG. 9.

[0179]

[0160] Video decoder 300 may receive entropy-coded data of the current block, such as entropy-coded prediction information and entropy-coded data about the transform coefficients of the residual block corresponding to the current block (370). Video decoder 300 may entropy-decode the entropy-coded data to determine the prediction information of the current block and reproduce the transform coefficients of the residual block (372). Video decoder 300 may also entropy-decode the last_sig_coeff_x_prefix and last_sig_coeff_y_prefix syntax elements using the context determined according to the techniques of the present disclosure. Video decoder 300 may predict the current block, for example, using the intra or inter prediction mode indicated by the prediction information of the current block to calculate the prediction block of the current block (374). Video decoder 300 may then inverse-scan the reproduced transform coefficients to create a block of quantized transform coefficients (376). Video decoder 300 may then inverse-quantize the transform coefficients and apply an inverse transform to the transform coefficients to generate a residual block (378). Video decoder 300 may finally decode the current block by combining the prediction block and the residual block (380).

[0180] [

[0161] ]FIG. 10 is a flowchart illustrating an exemplary operation of a video coder according to one or more techniques of the present disclosure. In the example of FIG. 10, a video coder (e.g., video encoder 200 or video decoder 300) may use a context model to derive a context increment for a first syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix) (400). The video coder may use the operations of FIG. 11 to derive the context increment for the first syntax element. The first syntax element may indicate a prefix of an x or y coordinate of a last significant transform coefficient of a luma component of a block of video data. Further, the video coder may apply CABAC to a bin of the first syntax element using a context determined based on the context increment of the first syntax element (402).

[0181] [

[0162] ]The video coder may use a context model to derive a context increment for a second syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix) (404). The video coder may use the operations of FIG. 11 to derive the context increment for the second syntax element. The second syntax element may indicate a prefix of an x or y coordinate of a last significant transform coefficient of a chroma component of a block of a picture of video data. The color format of the picture may be different from 4:2:0. The video coder may apply CABAC to a bin of the second syntax element using a context determined based on the context increment of the second syntax element (406).

[0182]

[0163] FIG. 11 is a flowchart illustrating an exemplary operation of video encoder 200 in accordance with one or more techniques of the present disclosure. In the example of FIG. 11, video encoder 200 may determine which context model to use from a first context model and a second context model based on the color format of the picture to determine a context increment of a syntax element indicating a prefix of the x or y coordinate of the last significant transform coefficient of a color component (e.g., a first color component) of a block of pictures of video data (440). In some examples, video encoder 200 may use the first context model to determine a context increment of a first syntax element and may make a determination to use the second context model to determine a context increment of a second syntax element based on the color format of the pictures of the video data. Alternatively, video encoder 200 may make a determination to use the (same) first context model to determine a context increment of a second syntax element based on the color format of the pictures of the video data. In this example, the second syntax element indicates a prefix of the x or y coordinate of the last significant transform coefficient of a second color component of a block of the picture. Further, video encoder 200 may encode the bins of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element.

[0183]

[0164] In some examples, the video encoder 200 may make a determination to use a second context model to determine the context increment of a second syntax element, based on the fact that none of the following conditions are met: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, and / or ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luma component. Further, in some examples, the video encoder 200 may make a determination to use a second context model to determine the context increment based on the fact that the following condition is not met: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not indicate the prefix of the y coordinate of the last significant transform coefficient of the second color component of the block. In some examples, the video encoder 200 may make a determination to use a first context model to derive the context increment of a second syntax element based on the fact that at least one of the following conditions is met: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, or ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luma component. Further, in some examples, the video encoder 200 may make a determination to use a first context model to derive the context increment of a second syntax element based on the fact that the following condition is met: iii) the color format of the picture is 4:2:2 color format, the second color component is not a luma component, and the second syntax element indicates the prefix of the y coordinate of the last significant transform coefficient of the block. The same determination process may be applied to the first syntax element. In other words, the same conditions may be used for both the first syntax element and the second syntax element.In some examples, the first syntax element and the second syntax element may be associated with different color components. In some examples, the first syntax element and the second syntax element may be associated with different coordinates of the same color component.

[0165] Further, the video encoder 200 may encode the bins of the syntax elements by applying CABAC using the context determined based on the context increment of the syntax elements (442).

[0184]

[0166] FIG. 12 is a flowchart showing an exemplary operation of the video decoder 300 according to one or more techniques of the present disclosure. In the example of FIG. 12, the video decoder 300 may determine which context model to use from the first context model and the second context model based on the color format of the picture to determine the context increment of the syntax element indicating the prefix of the x or y coordinate of the last significant transform coefficient of the color component (e.g., the first color component) of the block of the picture of the video data (460). The video decoder 300 may make a decision to use the second context model to determine the context increment of the second syntax element based on the color format of the picture of the video data. Alternatively, the video decoder 300 may make a decision to use the (same) first context model to determine the context increment of the second syntax element based on the color format of the picture of the video data. In this example, the second syntax element indicates the prefix of the x or y coordinate of the last significant transform coefficient of the second color component of the block of the picture. Further, the video encoder 200 may encode the bins of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element.

[0185]

[0167] In some examples, the video decoder 300 may make a determination to use a second context model to determine the context increment of a second syntax element, based on the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, and / or ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the second color component is not a luma component. Further, in some examples, based on the following condition: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not indicate the prefix of the y coordinate of the last significant transform coefficient of the second color component of the block, the video decoder 300 may make a determination to use the second context model to derive the context model of the second syntax element. In some examples, based on at least one of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, or ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luma component, the video decoder 300 may make a determination to use the first context model to derive the context increment of the second syntax element. Further, in some examples, based on the following condition: iii) the color format of the picture is a 4:2:2 color format, the second color component is not a luma component, and the second syntax element indicates the prefix of the y coordinate of the last significant transform coefficient of the block, the video decoder 300 may make a determination to use the first context model to derive the context increment of the second syntax element. The same determination process may be applied to the first syntax element. In other words, the same conditions may be used for the first syntax element and the second syntax element.In some examples, the first syntax element and the second syntax element may be associated with different color components. In some examples, the first syntax element and the second syntax element may be associated with different coordinates of the same color component.

[0168] Furthermore, the video decoder 300 may decode the bins of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element (462).

[0186]

[0169] FIG. 13 is a flowchart illustrating an exemplary operation of a video coder for determining a context according to one or more techniques of the present disclosure. In the example of FIG. 13, the video coder may determine the value of a luma model variable (e.g., enableLumaModel) (500). The video coder may use the operation of FIG. 14 described below to determine the value of the luma model variable.

[0187]

[0170] Furthermore, in the example of FIG. 13, the video coder may determine whether the luma model variable is truly equal (502). In response to determining that the luma model variable is truly equal (the "YES" branch of 502), the video coder may determine whether the base-2 logarithm value of the transform size of the block (e.g., Log2TrafoSize) is less than 6 (504). Based on the base-2 logarithm value of the transform size of the block being less than 6 (the "YES" branch of 504), the video coder may set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and may set the context shift to be equal to (log2TrafoSize + 1)>>2. Log2TrafoSize indicates the base-2 logarithm value of the transform size.

[0188] In response to determining that the base-2 logarithm value of the transform size is 6 or more (the "NO" branch of 504), the video coder may set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7) and set the context shift to ((log2TrafoSize + 1)>>2)<<1 (508). log2TrafoSize indicates the base-2 logarithm value of the transform size.

[0189] However, in response to determining that the luma model variable is not truly equal (the "NO" branch of 502), the video coder may set the context offset to 18 and set the context shift to Max(0, log2TrafoSize - 2)-Max(0, log2TrafoSize - 4) (510). Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithm value of the transform size of the block.

[0190] In either case, after determining the context shift and the context offset in act 506, 508, or 510, the video coder may determine the context increment as (binIdx>>ctxShift)+ctxOffset (512). Here, binIdx is the bin index of the bin of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset. The applicable syntax element may be the first syntax element or the second syntax element described with respect to FIG. 10.

[0191]

[0174] Figure 14 is a flowchart illustrating exemplary operations for determining values of luma model variables according to one or more techniques of the present disclosure. In the example of Figure 14, a video coder (e.g., video encoder 200 or video decoder 300) may first set a luma model variable (e.g., enableLumaModel) to false (600). The video coder may then determine whether the color format is equal to 0 or 3 (602). For example, the video coder may determine whether chromaArrayType is equal to 0 or 3. As described above, a color format equal to 0 indicates monochrome, and a color format equal to 3 indicates a 4:4:4 color format. If the color format is equal to 0 or 3 (the "YES" branch of 602), the video coder may set the luma model variable to true (604).

[0192]

[0175] Otherwise, if the color format is not equal to 0 or 3 (the "NO" branch of 602), the video coder may determine whether the color format is equal to 1 or 2 and whether the current color component (e.g., cIdx) is equal to 0 (606). The current color component is the color component associated with the coded syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix). As described above, a color format equal to 1 indicates a 4:2:0 color format, and a color format equal to 2 indicates a 4:2:2 color format. If the color format is equal to 1 or 2 and the color component is equal to 0 (the "YES" branch of 606), the video coder may set the luma model variable to true (604).

[0193] Instead, if the color format is not equal to 1 or 2, or if the current color component is not equal to 0 (the "NO" branch of 606), the video coder may determine whether the color format is equal to 2, the current color component is not equal to 0, and a syntax element (e.g., last_sig_coeff_y_prefix) indicates the prefix of the y coordinate of the last significant transform coefficient of the block (608). In response to determining that the color format is equal to 2, the current color component is not equal to 0, and the syntax element indicates the prefix of the y coordinate of the last significant transform coefficient of the block (the "YES" branch of 608), the video coder may set the luma model variable to true (604). Otherwise, if the color format is not equal to 2, or the current color component is equal to 0, or the syntax element does not indicate the prefix of the y coordinate of the last significant transform coefficient of the block (the "NO" branch of 608), the video coder does not change the value of the luma model variable (610). In some examples, decision block 608 may be omitted.

[0194]

[0177] The following is a non-limiting list of aspects according to one or more techniques of the present disclosure.

[0195]

[0178] Aspect 1A. A method of coding video data, the method comprising determining parameters of a residual coding operation based on a chroma format indicator applicable to a block of the video data, and performing a residual coding operation based on the determined parameters of the residual coding operation to code residual data of a non-luma component of the block.

[0196] Aspect 2A. A method for coding video data, the method comprising using a context model to entropy code a prefix of the y coordinate of the last significant coefficient of the luma component of a block, and using the same context model to entropy code a prefix of the y coordinate of the last significant coefficient of the chroma component of the block based on the block having a color format of 4:2:2.

[0197] Aspect 3A. The method of Aspect 2A, further comprising the method of Aspect 1.

[0198] Aspect 4A. A method according to any one of Aspects 1A - 3A, wherein the coding comprises decoding.

[0199] Aspect 5A. A method according to any one of Aspects 1A - 4A, wherein the coding comprises encoding.

[0200] Aspect 6A. A device for coding video data, the device comprising one or more means for implementing the method according to any one of Aspects 1A - 5A.

[0201] Aspect 7A. The device of Aspect 6A, wherein the one or more means comprise one or more processors implemented in a circuit.

[0202] Aspect 8A. The device according to any one of Aspects 6A and 7A, further comprising a memory configured to store video data.

[0203] Aspect 9A. The device according to any one of Aspects 6A - 8A, further comprising a display configured to display the decoded video data.

[0204] Aspect 10A. A device according to any of Aspects 6A - 9A, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0205] Aspect 11A. A device according to any of Aspects 6A - 10A, wherein the device comprises a video decoder.

[0206] Aspect 12A. A device according to any of Aspects 6A - 11A, wherein the device comprises a video encoder.

[0207] Aspect 13A. A computer-readable storage medium storing instructions which, when executed, cause one or more processors to perform a method according to any of Aspects 1A - 5A.

[0208] Aspect 1B. A method for coding video data, the method comprising: using a context model to derive a context increment of a first syntax element, wherein the first syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a luma component of a block of a picture of the video data; applying context-adaptive binary arithmetic coding (CABAC) to a bin of the first syntax element using a context determined based on the context increment of the first syntax element; using a context model to derive a context increment of a second syntax element, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a chroma component of a block, and the color format of the picture is different from 4:2:0; applying CABAC to a bin of the second syntax element using a context determined based on the context increment of the second syntax element; wherein using the context model comprises performing one of: setting a context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and setting a context shift to (log2TrafoSize + 1)>>2, based on the logarithm base 2 value of the transform size of the block being less than 6, where log2TrafoSize indicates the logarithm base 2 value of the transform size; or setting a context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7) and setting a context shift to ((log2TrafoSize + 1)>>2)<<1, based on the logarithm base 2 value of the transform size being 6 or more; determining the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset, whereinA method comprising that applicable syntax elements are a first syntax element or a second syntax element.

[0209]

[0192] Aspect 2B. The method of Aspect 1B further comprises setting an enable luma model variable from false to true based on one of the following: the color format of the picture is monochrome or 4:4:4, or the color format of the picture is 4:2:0 or 4:2:2 and the color component of the applicable syntax element is luma; or using a context model to derive a context increment of a second syntax element based on that the enable luma model variable is equal to true.

[0210]

[0193] Aspect 3B. The method of Aspect 2B further comprises setting an enable luma model variable from false to true based on that the color format of the picture is 4:2:2, the color component of the applicable syntax element is a chroma component, and the applicable syntax element indicates a prefix of the y coordinate of the last significant transform coefficient of the chroma component of the block.

[0211] Aspect 4B. The context model is the first context model, the picture is the first picture, the block is the first block, the applicable syntax element is the first applicable syntax element, the context is the first context, the context shift is the first context shift, the context offset is the first context offset, and the method is based on one of: the color format of the second picture being monochrome or 4:4:4, or the color format of the second picture being 4:2:0 or 4:2:2 and the color component of the second applicable syntax element being luma, set the enable luma model variable from false to true, or based on the enable luma model variable being equal to false, use a second context model to derive the context increment of the second applicable syntax element, where using a second context model to derive the context increment of the second applicable syntax element sets the second context offset to 18 and the second context shift to Max(0, log2TrafoSize 2 - 2) - Max(0, log2TrafoSize 2 - 4), where Max denotes the maximum value function and log2TrafoSize 2 denotes the base-2 logarithm of the transform size of the second block, and determine the second context increment as (binIdx 2 >> ctxShift 2 ) + ctxOffset 2 , where binIdx 2 is the bin index of the bin of the second applicable syntax element, ctxShift 2 denotes the second context shift, and ctxOffset 2 denotes the second context offset, further comprising the method of Aspect 1B.

[0212]

[0195] Aspect 5B. The method further comprises setting an enable luma model variable from false to true based on the color format of the second picture being 4:2:2, the color components of the second applicable syntax element being the chroma components of the second block of the second picture, and the second applicable syntax element indicating a prefix of the y coordinate of the last significant transform coefficient of the chroma components of the second block, of the method of Aspect 4B.

[0213]

[0196] Aspect 6B. Applying CABAC to the bins of the first syntax element comprises decoding the bins of the first syntax element with CABAC, and applying CABAC to the bins of the second syntax element comprises decoding the bins of the second syntax element with CABAC, of the method of Aspect 1B.

[0214]

[0197] Aspect 7B. Applying CABAC to the bins of the first syntax element comprises encoding the bins of the first syntax element with CABAC, and applying CABAC to the bins of the second syntax element comprises encoding the bins of the second syntax element with CABAC, of the method of Aspect 1B.

[0215]

[0198] Aspect 8B. A device for coding video data, the device comprising a memory configured to store the video data and one or more processors coupled to the memory, the one or more processors being implemented in a circuit, using a context model to derive a context increment of a first syntax element, wherein the first syntax element indicates a prefix of an x or y coordinate of the last significant transform coefficient of the luma component of a block of the video data, applying context-adaptive binary arithmetic coding (CABAC) to a bin of the first syntax element using a context determined based on the context increment of the first syntax element, using a context model to derive a context increment of a second syntax element, wherein the second syntax element indicates a prefix of an x or y coordinate of the last significant transform coefficient of the chroma component of the block and the color format of the picture is different from 4:2:0, applying CABAC to a bin of the second syntax element using a context determined based on the context increment of the second syntax element, wherein the one or more processors, as part of using the context model, set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and set the context shift to (log2TrafoSize + 1)>>2 based on the fact that the base-2 logarithm value of the transform size of the block is less than 6, where log2TrafoSize represents the base-2 logarithm value of the transform size, set the context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7) and set the context shift to ((log2TrafoSize + 1)>>2)<<1 based on the fact that the base-2 logarithm value of the transform size is 6 or more, and determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx isA device configured to perform a bin index of bins of applicable syntax elements, where ctxShift indicates a context shift and ctxOffset indicates a context offset, and where the applicable syntax element is a first syntax element or a second syntax element.

[0216]

[0199] Aspect 9B. One or more processors are further configured to set an enable luma model variable from false to true based on one of the picture color format being monochrome or 4:4:4, or the picture color format being 4:2:0 or 4:2:2 and the color component of the applicable syntax element being luma, and the one or more processors are configured to use a context model to derive a context increment of a second syntax element based on the enable luma model variable being equal to true as part of using the context model to derive the context increment of the second syntax element, the device of Aspect 8B.

[0217]

[0200] Aspect 10B. One or more processors are further configured to set an enable luma model variable from false to true based on the picture color format being 4:2:2, the color component of the applicable syntax element being a chroma component, and the applicable syntax element indicating a prefix of the y coordinate of the last significant transform coefficient of the chroma component of a block, the device of Aspect 9B.

[0218]

[0201] Aspect 11B. One or more processors are configured to use a second context model to derive a context increment for a second applicable syntax element based on the enable luma model variable being equal to false, where the one or more processors set a second context offset to 18 and a second context shift to Max(0, log2TrafoSize 2 -2)-Max(0, log2TrafoSize 2 -4) as part of using the second context model to derive the context increment for the second applicable syntax element, where Max represents the maximum value function and log2TrafoSize 2 represents the base-2 logarithm of the transform size of the second block, and determine the second context increment as (binIdx 2 >>ctxShift 2 )+ctxOffset 2 , where binIdx 2 is the bin index of the bin of the second applicable syntax element, ctxShift 2 represents the second context shift, and ctxOffset 2 represents the second context offset, a device of Aspect 9B configured to perform the above.

[0219]

[0202] Aspect 12B. A device of Aspect 8B, where one or more processors are configured to decode the bin of the first syntax element by applying CABAC to the bin of the first syntax element and are configured to decode the bin of the second syntax element by applying CABAC to the bin of the second syntax element as part of applying CABAC to the bin of the second syntax element.

[0220] Aspect 13B. A device according to Aspect 8B, wherein one or more processors are configured to CABAC-encode bins of a first syntax element as part of applying CABAC to the bins of the first syntax element, and one or more processors are configured to CABAC-encode bins of a second syntax element as part of applying CABAC to the bins of the second syntax element.

[0221] Aspect 14B. A device according to Aspect 8B, further comprising a display configured to display decoded video data.

[0222] Aspect 15B. A device according to Aspect 8B, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0223] Aspect 16B. A device for coding video data, the device comprising means for using a context model to derive a context increment of a first syntax element, wherein the first syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a luma component of a block of video data, means for applying context-adaptive binary arithmetic coding (CABAC) to a bin of the first syntax element using a context determined based on the context increment of the first syntax element, means for using a context model to derive a context increment of a second syntax element, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a chroma component of the block and the color format of the picture is different from 4:2:0, means for applying CABAC to a bin of the second syntax element using a context determined based on the context increment of the second syntax element, wherein the means for using a context model comprises means for setting a context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and setting a context shift to (log2TrafoSize + 1)>>2 based on the fact that the base-2 logarithm value of the transform size of the block is less than 6, where log2TrafoSize indicates the base-2 logarithm value of the transform size, means for setting a context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7) and setting a context shift to ((log2TrafoSize + 1)>>2)<<1 based on the fact that the base-2 logarithm value of the transform size is 6 or more, means for determining the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset, whereinA device comprising applicable syntax elements that are either a first syntax element or a second syntax element.

[0224]

[0207] Aspect 17B. The device further comprises means for setting an enable luma model variable from false to true based on one of the following: the color format of the picture is monochrome or 4:4:4, or the color format of the picture is 4:2:0 or 4:2:2 and the color components of the applicable syntax element are luma; or means for using a context model to derive a context increment for a second syntax element, where the means for using the context model to derive a context increment for the second syntax element comprises means for using the context model to derive a context increment for the second syntax element based on the enable luma model variable being equal to true, the device of Aspect 16B.

[0225]

[0208] Aspect 18B. The device further comprises means for using a second context model to derive a context increment for a second applicable syntax element based on the enable luma model variable being equal to false, where the means for using the second context model to derive a context increment for the second applicable syntax element sets a second context offset to 18 and a second context shift to Max(0, log2TrafoSize 2 - 2) - Max(0, log2TrafoSize 2 - 4), where Max represents the maximum value function and log2TrafoSize 2 represents the base-2 logarithm of the transform size of the second block, and means for determining the second context increment as (binIdx 2 >> ctxShift 2 ) + ctxOffset 2 , where binIdx 2 is the bin index of the bin of the second applicable syntax element, ctxShift 2 represents the second context shift, and ctxOffset 2A device of aspect 17B, comprising: indicating a second context offset.

[0226] Aspect 19B. A computer-readable storage medium storing instructions that, when executed, cause one or more processors to use a context model to derive a context increment for a first syntax element, where the first syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a lum component of a block of video data, apply context-adaptive binary arithmetic coding (CABAC) to a bin of the first syntax element using a context determined based on the context increment of the first syntax element, use the context model to derive a context increment for a second syntax element, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a chroma component of a block and the color format of the picture is different from 4:2:0, and apply CABAC to a bin of the second syntax element using a context determined based on the context increment of the second syntax element, where the instructions that cause the one or more processors to use the context model, when executed, cause the one or more processors to set a context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and set a context shift to (log2TrafoSize + 1)>>2 based on the fact that the base-2 logarithm of the transform size of the block is less than 6, where log2TrafoSize indicates the base-2 logarithm of the transform size, set a context offset to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7) and set a context shift to ((log2TrafoSize + 1)>>2)<<1 based on the fact that the base-2 logarithm of the transform size is 6 or greater, and determine a context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of a bin of an applicable syntax element and ctxShift indicates the context shift.A computer-readable storage medium comprising an instruction to cause ctxOffset to indicate a context offset, where an applicable syntax element is a first syntax element or a second syntax element.

[0227]

[0210] Aspect 1C: A method for decoding video data includes determining which context model to use from a first context model and a second context model based on the color format of a picture to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture, and decoding a bin of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.

[0228]

[0211] Aspect 2C: The method of Aspect 1C, further comprising using a first context model to determine a context increment of a first syntax element where the color component is a first color component and the syntax element is a first syntax element, making a determination to use the first context model to determine a context increment of a second syntax element based on the color format of a picture of the video data, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of a block of the picture, and decoding a bin of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element.

[0229] Aspect 3C: The color component is the first color component, the syntax element is the first syntax element, and the method comprises using a first context model to determine a context increment of the first syntax element and making a determination to use a second context model to determine a context increment of a second syntax element based on a color format of a picture of video data, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of a block of the picture, and decoding a bin of the second syntax element by applying CABAC using a context determined based on the context increment of the second syntax element. The method of Aspect 1C further comprises this.

[0230] Aspect 4C: The method of Aspect 2C or 3C, where the first color component is a luma component and the second color component is a chroma component.

[0231] Aspect 5C: Making the determination to use the second context model to determine the context increment of the second syntax element comprises making a determination to use the second context model to derive the context increment of the second syntax element based on neither of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; and ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the second color component is a luma component. The method of either Aspect 3C or 4C comprises this.

[0232] Aspect 6C: Making the determination to use a second context model to determine the context increment of a second syntax element is based further on the fact that the following condition: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not satisfy the condition of indicating the y - coordinate prefix of the last significant transform coefficient of the second color component of the block, is not satisfied, of the method of Aspect 5C.

[0233] Aspect 7C: Using a first context model to derive the context increment of a first syntax element, based on the fact that the base - 2 logarithm value of the transform size of the block is less than 6, setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2), and setting the context shift ctxShift of the first syntax element to be equal to (log2TrafoSize + 1)>>2, where log2TrafoSize represents the base - 2 logarithm value of the transform size, or based on the fact that the base - 2 logarithm value of the transform size is 6 or more, setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and setting the context shift ctxShift of the first syntax element to be equal to ((log2TrafoSize + 1)>>2)<<1, and making one of these two implementations, and determining the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the first syntax element, of the method of any of Aspects 3C to 6C.

[0234] Aspect 8C: Using the second context model to determine the context increment of the second syntax element sets the context offset ctxOffset of the second syntax element to 18 and sets the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm of the transform size of the block, and determines the context increment of the second syntax element as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bin index of the bin of the second syntax element, and is a method according to any of Aspects 3C to 7C.

[0235] Aspect 9C: Determining which context model to use between the first context model and the second context model to determine the context increment of the syntax element is based on at least one of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the color component is a luma component, or iii) the color format of the picture is a 4:2:2 color format, the color component is not a luma component, and the syntax element represents the prefix of the y coordinate of the last significant transform coefficient of the color component of the block, and making a decision to use the first context model to derive the context increment of the syntax element, and is a method according to any of Aspects 1C to 8C.

[0236] Aspect 10C: A method of encoding video data includes determining which context model to use from a first context model and a second context model based on the color format of a picture of the video data to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data, and encoding a bin of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.

[0237] Aspect 11C: The color component is a first color component, the syntax element is a first syntax element, and the method further comprises using a first context model to determine a context increment of the first syntax element, making a determination to use the first context model to determine a context increment of a second syntax element based on the color format of a picture of the video data, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of a block of the picture, and encoding a bin of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element, the method of Aspect 10C.

[0238] Aspect 12C: The color component is the first color component, the syntax element is the first syntax element, and the method includes using a first context model to determine a context increment of the first syntax element, and making a determination to use a second context model to determine a context increment of a second syntax element based on a color format of a picture of video data, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of a block of the picture, and encoding a bin of the second syntax element by applying CABAC using a context determined based on the context increment of the second syntax element, which is a method of Aspect 10C.

[0239] Aspect 13C: The method of Aspect 11C or 12C, where the first color component is a luma component and the second color component is a chroma component.

[0240] Aspect 14C: Making the determination to use the second context model to determine the context increment of the second syntax element includes making a determination to use the second context model to derive the context increment of the second syntax element based on neither of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; and ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the second color component is a luma component, which is a method of either Aspect 12C or 13C.

[0241] Aspect 15C: Making the determination to use the second context model to determine the context increment of the second syntax element is further based on the fact that the following condition: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not satisfy the condition of indicating the prefix of the y coordinate of the last significant transform coefficient of the second color component of the block, and is any of the methods of Aspects 12C to 14C.

[0242] Aspect 16C: Using the first context model to derive the context increment of the first syntax element is based on the fact that the base-2 logarithm value of the transform size of the block is less than 6. Set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2), and set the context shift ctxShift of the first syntax element to be equal to (log2TrafoSize + 1)>>2, where log2TrafoSize represents the base-2 logarithm value of the transform size. Or based on the fact that the base-2 logarithm value of the transform size is 6 or more, set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and set the context shift ctxShift of the first syntax element to be equal to ((log2TrafoSize + 1)>>2)<<1. Implement one of these, and determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the first syntax element, and is any of the methods of Aspects 12C to 15C.

[0243] Aspect 17C: Using the second context model to derive the context increment of the second syntax element involves setting the context offset ctxOffset of the second syntax element to 18, setting the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm of the transform size of the block, and determining the context increment of the second syntax element as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bin index of the bin of the second syntax element, and is a method of any of Aspects 12C to 16C.

[0244] Aspect 18C: Determining which context model to use between the first context model and the second context model to determine the context increment of the syntax element is based on at least one of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the color component is a luma component, or iii) the color format of the picture is a 4:2:2 color format, the color component is not a luma component, and the syntax element represents the prefix of the y coordinate of the last significant transform coefficient of the color component of the block, and involves making a decision to use the first context model to derive the context increment of the syntax element, and is a method of any of Aspects 10C to 17C.

[0245] Aspect 19C: A device for decoding video data includes a memory configured to store the video data and one or more processors coupled to the memory, the one or more processors being implemented in a circuit and determining which context model to use from a first context model and a second context model based on the color format of a picture to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture, and decoding bins of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.

[0246] Aspect 20C: The color component is a first color component, the syntax element is a first syntax element, and the one or more processors are further configured to use the first context model to determine a context increment of the first syntax element, make a determination to use the first context model to determine a context increment of a second syntax element based on the color format of a picture of the video data, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of a block of the picture, and decode bins of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element, for the device of Aspect 19C.

[0247] Aspect 21C: The color component is the first color component, the syntax element is the first syntax element, and one or more processors use a first context model to determine a context increment for the first syntax element and make a determination to use a second context model to determine a context increment for a second syntax element based on a color format of a picture of video data, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of a block of the picture, and further configured to decode a bin of the second syntax element by applying CABAC using a context determined based on the context increment of the second syntax element, a device of either Aspect 19C or 20C.

[0248] Aspect 22C: A device of Aspect 20C or 21C, where the first color component is a luma component and the second color component is a chroma component.

[0249] Aspect 23C: A device of either Aspect 21C or 22C, where the configuration that one or more processors make a determination to use a second context model to determine a context increment for a second syntax element is based on neither of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; and ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the second color component is a luma component, and includes making a determination to use the second context model to derive the context increment for the second syntax element.

[0250] Aspect 24C: A device of Aspect 23C, wherein one or more processors are configured to make a determination to use a second context model to determine a context increment of a second syntax element based on the fact that the following condition iii) is not satisfied: the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not indicate a prefix of the y - coordinate of the last significant transform coefficient of the second color component of the block.

[0251] Aspect 25C: A device of any of Aspects 19C - 24C, wherein one or more processors are configured to, as part of using a first context model to derive a context increment of a first syntax element, set a context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and set a context shift ctxShift of the first syntax element to (log2TrafoSize + 1)>>2 based on the fact that the base - 2 logarithm value of the transform size of the block is less than 6, or set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7) and set the context shift ctxShift of the first syntax element to ((log2TrafoSize + 1)>>2)<<1 based on the fact that the base - 2 logarithm value of the transform size is 6 or more, and determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the first syntax element.

[0252] Aspect 26C: As part of using a second context model to determine the context increment of a second syntax element, one or more processors set the context offset ctxOffset of the second syntax element to 18 and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm of the transform size of the block, and determine the context increment of the second syntax element as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bin index of the bin of the second syntax element, and is configured to perform the above for any device of Aspects 21C - 25C.

[0253]

[0236] Aspect 27C: As part of determining which context model to use from a first context model and a second context model to determine the context increment of a syntax element, one or more processors determine based on at least one of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the color component is a luma component; or iii) the color format of the picture is a 4:2:2 color format, the color component is not a luma component, and the syntax element represents a prefix of the y coordinate of the last significant transform coefficient of the color component of the block, and is configured to make a decision to use the first context model to derive the context increment of the syntax element for any device of Aspects 21C - 26C.

[0254]

[0237] Aspect 28C: Any device of Aspects 19C - 27C further comprising a display configured to display the decoded video data.

[0255] Aspect 29C: A device according to any of Aspects 19C - 28C, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0256] Aspect 30C: A device for encoding video data includes a memory configured to store the video data and one or more processors coupled to the memory. The one or more processors are implemented in a circuit and are configured to determine which context model to use from a first context model and a second context model based on a color format of a picture of the video data to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture, and to encode bins of the syntax element by applying context adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.

[0257] Aspect 31C: A device according to Aspect 30C, further configured such that the color component is a first color component, the syntax element is a first syntax element, the one or more processors are configured to use the first context model to determine a context increment of the first syntax element, and to make a determination to use the first context model to determine a context increment of a second syntax element based on a color format of a picture of the video data, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of a block of the picture, and to decode bins of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element.

[0258] Aspect 32C: The color component is the first color component, the syntax element is the first syntax element, and one or more processors use a first context model to determine a context increment for the first syntax element and make a determination to use a second context model to determine a context increment for a second syntax element based on the color format of a picture of video data, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of a block of the picture, and encode a bin of the second syntax element by applying CABAC using a context determined based on the context increment of the second syntax element. The device of Aspect 30C is further configured to perform the above operations.

[0259] Aspect 33C: The device of Aspect 31C or 32C, where the first color component is a luma component and the second color component is a chroma component.

[0260] Aspect 34C: The device of any one of Aspects 32C to 33C, where the configuration that one or more processors make a determination to use a second context model to determine a context increment for a second syntax element is based on the following conditions: i) the color format of the picture is not a monochrome color format or a 4:4:4 color format, and ii) the color format of the picture is not a 4:2:0 color format or a 4:2:2 color format, and the second color component is not a luma component. Based on this, a determination is made to use the second context model to derive the context increment of the second syntax element.

[0261] Aspect 35C: A device according to Aspect 34C, wherein one or more processors are configured to make a determination to use a second context model to determine a context increment of a second syntax element based on the fact that the following condition iii) is not satisfied: the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not indicate a prefix of the y - coordinate of the last significant transform coefficient of the second color component of the block.

[0262] Aspect 36C: A device according to any of Aspects 32C - 35C, wherein one or more processors are configured to perform one of the following: based on the fact that the base - 2 logarithm of the transform size of the block is less than 6, set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2) and set the context shift ctxShift of the first syntax element to (log2TrafoSize + 1)>>2; or based on the fact that the base - 2 logarithm of the transform size is 6 or more, set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7) and set the context shift ctxShift of the first syntax element to ((log2TrafoSize + 1)>>2)<<1, where log2TrafoSize represents the base - 2 logarithm of the transform size; and determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the first syntax element.

[0263] Aspect 37C: As part of using a second context model to determine a context increment for a second syntax element, one or more processors set a context offset ctxOffset for the second syntax element to 18 and set a context shift ctxShift for the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm of the transform size of the block, and determine the context increment for the second syntax element as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bin index of the bin of the second syntax element, and is configured to perform the above of any of Devices of Aspects 32C - 36C.

[0264] Aspect 38C: As part of determining which context model to use from a first context model and a second context model to determine a context increment for a syntax element, one or more processors use the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the color component is a luma component; or iii) the color format of the picture is a 4:2:2 color format, the color component is not a luma component, and the syntax element represents a prefix of the y coordinate of the last significant transform coefficient of the color component of the block. Based on at least one of the above conditions being satisfied, it is configured to make a determination to use the first context model to derive the context increment for the syntax element, and is any of Devices of Aspects 30C - 37C.

[0265] Aspect 39C: A device according to any of Aspects 30C to 38C, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0266] Aspect 40C: A device for decoding video data includes means for determining which context model of a first context model and a second context model should be used based on a color format of a picture in order to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of a picture of the video data, and means for decoding a bin of the syntax element by applying context adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.

[0267] Aspect 41C: A device for encoding video data includes means for determining which context model of a first context model and a second context model should be used based on a color format of a picture in order to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of a picture of the video data, and means for encoding a bin of the syntax element by applying context adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.

[0268] Aspect 42C: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to determine, based on a color format of a picture, which context model to use from a first context model and a second context model to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of video data, and to decode a bin of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.

[0269] Aspect 43C: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to determine, based on a color format of a picture, which context model to use from a first context model and a second context model to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of video data, and to encode a bin of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element.

[0270] Note that, in some examples, some of the acts or events of any of the techniques described herein may be performed in a different sequence, may be added, merged, or completely excluded (e.g., not all of the described acts or events are necessary for the practice of the technique). Moreover, in some examples, the acts or events may not be performed sequentially, but may be performed simultaneously, for example, through multi-threading, interrupt processing, or multiple processors.

[0271] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or may include a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0272] By way of example and not limitation, such computer-readable storage media can comprise RAM, ROM, EEPROM (registered trademark), CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection can properly be called a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that the computer-readable storage media and data storage media are directed to non-transitory, tangible storage media rather than including connections, carrier waves, signals, or other transitory media. As used herein, disk and disc include compact disc (CD), laser disc (registered trademark), optical disc, digital versatile disc (DVD), floppy (registered trademark) disk, and Blu-ray (registered trademark) disc, where disk typically magnetically reproduces data and disc optically reproduces data with a laser. Combinations of the above should also be included within the scope of computer-readable media.

[0273]

[0256] The commands can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated logic circuits or discrete logic circuits. Thus, the terms "processor" and "processing circuit" as used herein can refer to either the above structures or any other structure suitable for implementing the techniques described herein. Further, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Also, the techniques can be fully implemented with one or more circuits or logic elements.

[0274]

[0257] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or sets of ICs (e.g., chip sets). In the present disclosure, various components, modules, or units have been described to emphasize the functional aspects of devices configured to implement the disclosed techniques, but those components, modules, or units do not necessarily have to be realized by different hardware units. Rather, as described above, the various units can be combined in a codec hardware unit, including one or more of the processors described above, along with suitable software and / or firmware, or provided by a set of interoperable hardware units.

[0275]

[0258] Various examples have been described. These and other examples fall within the scope of the following claims. The invention described in the claims of the present application at the time of initial filing is appended below. [C1] A method for decoding video data, the method comprising: determining which context model to use from a first context model and a second context model based on the color format of the picture, in order to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data; decoding a bin of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using a context determined based on the context increment of the syntax element; A method comprising the above. [C2] The color component is a first color component, and the syntax element is a first syntax element. The method further comprises: using the first context model to determine the context increment of the first syntax element; determining to use the first context model to determine a context increment of a second syntax element based on the color format of the picture of the video data, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture; decoding a bin of the second syntax element by applying CABAC using a context determined based on the context increment of the second syntax element; The method according to C1, further comprising the above. The method according to C1. [C3] The color component is a first color component, and the syntax element is a first syntax element. The method further comprises: using the first context model to determine the context increment of the first syntax element; Making a determination to use the second context model to determine a context increment of a second syntax element based on the color format of the picture of the video data, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture. Decoding a bin of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element. Further comprising. The method according to C1. [C4] The method according to C3, wherein the first color component is a luma component and the second color component is a chroma component. [C5] Making the determination to use the second context model to determine the context increment of the second syntax element is based on the following conditions: i) The color format of the picture is a monochrome color format or a 4:4:4 color format. ii) The color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luma component. The method according to C3, comprising making the determination to use the second context model to derive the context increment of the second syntax element based on that none of the above is satisfied. [C6] Making the determination to use the second context model to determine the context increment of the second syntax element is further based on the following condition: iii) The color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not indicate the prefix of the y coordinate of the last significant transform coefficient of the second color component of the block. The method according to C5. [C7] Using the first context model to derive the context increment of the first syntax element is Based on the base-2 logarithm value of the transform size of the block being less than 6, set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2), and set the context shift ctxShift of the first syntax element to be equal to (log2TrafoSize + 1)>>2, where log2TrafoSize represents the base-2 logarithm value of the transform size, or Based on the base-2 logarithm value of the transform size being 6 or more, set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and set the context shift ctxShift of the first syntax element to be equal to ((log2TrafoSize + 1)>>2)<<1 perform one of the above Determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the first syntax element The method according to C3, comprising the above. [C8] Using the second context model to determine the context increment of the second syntax element Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0,log2TrafoSize - 2)-Max(0,log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm value of the transform size of the block Determine the context increment of the second syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the second syntax element The method according to C3, comprising the above. [C9] Determining which of the first context model and the second context model should be used to determine the context increment of the syntax element is based on the following conditions: i) The color format of the picture is a monochrome color format or a 4:4:4 color format; ii) The color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the color component is a luma component, or iii) The color format of the picture is a 4:2:2 color format, the color component is not the luma component, and the syntax element indicates the prefix of the y coordinate of the last significant transform coefficient of the color component of the block The method according to C1, comprising making a determination to use the first context model to derive the context increment of the syntax element based on that at least one of the above is satisfied. [C10] A method for encoding video data, the method comprising: Determining which of a first context model and a second context model should be used based on the color format of the picture to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture; Encoding a bin of the syntax element by applying context adaptive binary arithmetic coding (CABAC) using a context determined based on the context increment of the syntax element The method comprising. [C11] The color component is a first color component, and the syntax element is a first syntax element. The method comprises: Using the first context model to determine the context increment of the first syntax element; Making a determination to use the first context model to determine a context increment of a second syntax element based on the color format of the picture of the video data, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture. Encoding the bins of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element further comprising The method according to C10 [C12] wherein the color component is a first color component and the syntax element is a first syntax element The method comprises using the first context model to determine the context increment of the first syntax element determining to use the second context model to determine the context increment of a second syntax element based on the color format of the picture of the video data, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture Encoding the bins of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element further comprising The method according to C10 [C13] The method according to C12, wherein the first color component is a luma component and the second color component is a chroma component [C14] The determining to use the second context model to determine the context increment of the second syntax element is based on the following conditions i) the color format of the picture is a monochrome color format or a 4:4:4 color format ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the second color component is a luma component The method according to C12, comprising determining to use the second context model to derive the context increment of the second syntax element based on none of the above conditions being satisfied [C15] Performing the determination of using the second context model to determine the context increment of the second syntax element is further based on the fact that the following condition: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not satisfy the condition of indicating the prefix of the y - coordinate of the last significant transform coefficient of the second color component of the block, the method according to C12. [C16] Using the first context model to derive the context increment of the first syntax element is Based on the fact that the base - 2 logarithm value of the transform size of the block is less than 6, setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2), and setting the context shift ctxShift of the first syntax element to be equal to (log2TrafoSize + 1)>>2, where log2TrafoSize represents the base - 2 logarithm value of the transform size, or Based on the fact that the base - 2 logarithm value of the transform size is 6 or more, setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and setting the context shift ctxShift of the first syntax element to be equal to ((log2TrafoSize + 1)>>2)<<1 where one of the above is performed, Determining the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the first syntax element, The method according to C12, comprising. [C17] Using the second context model to derive the context increment of the second syntax element is Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm of the transform size of the block. Determine the context increment of the second syntax element as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bin index of the bin of the second syntax element. The method according to C12, comprising the above. [C18] Determining which context model to use between the first context model and the second context model to determine the context increment of the syntax element is based on the following conditions: i) The color format of the picture is a monochrome color format or a 4:4:4 color format. ii) The color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the color component is a luma component, or iii) The color format of the picture is a 4:2:2 color format, the color component is not the luma component, and the syntax element indicates the prefix of the y coordinate of the last significant transform coefficient of the color component of the block. The method according to C10, comprising making a determination to use the first context model to derive the context increment of the syntax element based on at least one of the above being satisfied. [C19] A device for decoding video data, the device comprising: A memory configured to store the video data; One or more processors coupled to the memory, the one or more processors being implemented in a circuit. To determine the context increment of a syntax element indicating a prefix of the x or y coordinate of the last significant transform coefficient of the color component of a block of a picture of the video data, based on the color format of the picture, determine which context model to use from a first context model and a second context model, decode the bins of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element configured to perform a device. [C20] wherein the color component is a first color component and the syntax element is a first syntax element, the one or more processors use the first context model to determine the context increment of the first syntax element, make a determination to use the first context model to determine the context increment of a second syntax element based on the color format of the picture of the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last significant transform coefficient of a second color component of the block of the picture, decode the bins of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element further configured to perform the device according to C19. [C21] wherein the color component is a first color component and the syntax element is a first syntax element, the one or more processors use the first context model to determine the context increment of the first syntax element, make a determination to use the second context model to determine the context increment of a second syntax element based on the color format of the picture of the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last significant transform coefficient of a second color component of the block of the picture, Applying CABAC using the context determined based on the context increment of the second syntax element to decode the bin of the second syntax element and further configured to perform the device according to C19. [C22] The device according to C20, wherein the first color component is a luma component and the second color component is a chroma component. [C23] The determination that the one or more processors are configured to perform the determination to use the second context model to determine the context increment of the second syntax element is based on the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; and ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the second color component is a luma component The device according to C21, comprising performing the determination to use the second context model to derive the context increment of the second syntax element based on neither of the above being satisfied. [C24] The one or more processors are configured to perform the determination to use the second context model to determine the context increment of the second syntax element based on the following condition: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element does not indicate the prefix of the y coordinate of the last significant transform coefficient of the second color component of the block. The device according to C23. [C25] The one or more processors, as part of using the first context model to derive the context increment of the first syntax element, Based on the fact that the base-2 logarithm value of the transform size of the block is less than 6, set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2), and set the context shift ctxShift of the first syntax element to be equal to (log2TrafoSize + 1)>>2, where log2TrafoSize represents the base-2 logarithm value of the transform size, or Based on the fact that the base-2 logarithm value of the transform size is 6 or more, set the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and set the context shift ctxShift of the first syntax element to be equal to ((log2TrafoSize + 1)>>2)<<1 perform one of the above Determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the first syntax element The device according to C21, configured to perform the above [C26] As part of using the second context model to determine the context increment of the second syntax element, the one or more processors Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0,log2TrafoSize - 2)-Max(0,log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm value of the transform size of the block Determine the context increment of the second syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bin index of the bin of the second syntax element The device according to C21, configured to perform the above [C27] As part of determining which of the first context model and the second context model to use to determine the context increment of the syntax element, the one or more processors have the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the color component is a luma component, or iii) the color format of the picture is a 4:2:2 color format, the color component is not the luma component, and the syntax element indicates a prefix of the y coordinate of the last significant transform coefficient of the color component of the block The device according to C19, configured to make a determination to use the first context model to derive the context increment of the syntax element based on at least one of the above being satisfied. [C28] The device according to C19, further comprising a display configured to display the decoded video data. [C29] The device according to C19, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box. [C30] A device for encoding video data, the device comprising: a memory configured to store the video data; one or more processors coupled to the memory, the one or more processors being implemented in a circuit; determining which of a first context model and a second context model to use based on the color format of the picture to determine a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of a picture of the video data; encoding bins of the syntax element by applying context adaptive binary arithmetic coding (CABAC) using a context determined based on the context increment of the syntax element; and device. [C31] The color component is a first color component, and the syntax element is a first syntax element, the one or more processors are to use the first context model to determine the context increment of the first syntax element, to make a determination to use the first context model to determine the context increment of a second syntax element based on the color format of the picture of the video data, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture, to decode a bin of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element further configured to perform the device according to C30. [C32] The color component is a first color component, and the syntax element is a first syntax element, the one or more processors are to use the first context model to determine the context increment of the first syntax element, to make a determination to use the second context model to determine the context increment of a second syntax element based on the color format of the picture of the video data, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture, to encode a bin of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element further configured to perform the device according to C30. [C33] The device according to C32, wherein the first color component is a luminance component and the second color component is a chroma component. [C34] The configuration that the one or more processors make the determination to use the second context model to determine the context increment of the second syntax element is subject to the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luma component, and based on none of them being satisfied, making the determination to use the second context model to derive the context increment of the second syntax element, the device according to C32, comprising: [C35] the one or more processors are configured to make the determination to use the second context model to determine the context increment of the second syntax element based on the following condition not being satisfied: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element indicates a prefix of a y coordinate of the last significant transform coefficient of the second color component of the block, the device according to C34 [C36] as part of using the first context model to derive the context increment of the first syntax element, the one or more processors based on the base-2 logarithm value of the transform size of the block being less than 6, setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2), and setting the context shift ctxShift of the first syntax element to be equal to (log2TrafoSize + 1)>>2, where log2TrafoSize represents the base-2 logarithm value of the transform size, or based on the base-2 logarithm value of the transform size being 6 or more, setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize - 2)+((log2TrafoSize - 1)>>2)+((1<<log2TrafoSize)>>6)<<1)+((1<<log2TrafoSize)>>7), and setting the context shift ctxShift of the first syntax element to be equal to ((log2TrafoSize + 1)>>2)<<1 performing one of the above, and determine the context increment of the first syntax element as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bin index of the bin of the first syntax element The device according to C32, configured to perform [C37] As part of using the second context model to determine the context increment of the second syntax element, the one or more processors set the context offset ctxOffset of the second syntax element to 18 and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm of the transform size of the block determine the context increment of the second syntax element as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bin index of the bin of the second syntax element The device according to C32, configured to perform [C38] As part of determining which context model to use from the first context model and the second context model to determine the context increment of the syntax element, the one or more processors, the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format and the color component is a luma component, or iii) the color format of the picture is a 4:2:2 color format, the color component is not the luma component, and the syntax element indicates the prefix of the y coordinate of the last significant transform coefficient of the color component of the block The device according to C30, configured to make a determination to use the first context model to derive the context increment of the syntax element based on at least one of the following being satisfied [C39] The device according to C30, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box. [C40] A device for decoding video data, wherein the device means for determining which context model to use from a first context model and a second context model based on the color format of the picture, for determining a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data; and means for decoding a bin of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element comprising a device. [C41] A device for encoding video data, wherein the device means for determining which context model to use from a first context model and a second context model based on the color format of the picture, for determining a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data; and means for encoding a bin of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element comprising a device. [C42] A computer-readable storage medium storing instructions that, when executed, cause one or more processors to determine which context model to use from a first context model and a second context model based on the color format of the picture, for determining a context increment of a syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a color component of a block of the picture of the video data Decoding the bins of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element A computer-readable storage medium for causing the above to be performed. [C43] A computer-readable storage medium storing instructions that, when executed, cause one or more processors to Determine which context model to use from a first context model and a second context model based on the color format of the picture to determine a context increment of a syntax element indicating a prefix of the x or y coordinate of the last significant transform coefficient of the color component of a block of the picture of video data Encoding the bins of the syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the context increment of the syntax element A computer-readable storage medium for causing the above to be performed.

Claims

1. A method for decoding video data, the method comprising: determining which context model to use from a first context model and a second context model based on whether the color format of a picture of the video data is either 4:2:0 or 4:2:2; (a) the first context model for determining a first context increment of a first syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a first color component of a block of the picture; (b) the second context model for determining a second context increment of a second syntax element; wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture, and the second context increment of the second syntax element is: (i) setting a context offset ctxtOffset of the second syntax element to 18 and setting a context shift ctxtShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm value of the transform size of the block; (ii) determining the second context increment of the second syntax element as (binIdx >> ctxtShift) + ctxtOffset, where binIdx is a bin index of a bin of the second syntax element; comprising; determining the first context increment of the first syntax element using the determined first context model and determining the second context increment of the second syntax element using the determined second context model; decoding a bin of the first syntax element by applying context-adaptive binary arithmetic coding (CABAC) using a context determined based on the first context increment of the first syntax element; Applying CABAC using the context determined based on the second context increment of the second syntax element to decode the bin of the second syntax element; A method comprising. **Claim 2** The method according to claim 1, wherein the first color component is a luma component and the second color component is a chroma component. **Claim 3** The determination of using the second context model to determine the second context increment of the second syntax element is i) the color format of the picture is a monochrome color format or a 4:4:4 color format; and ii) the color format of the picture is the 4:2:0 color format or the 4:2:2 color format, and the second color component is a luma component; The method according to claim 1, comprising making the determination of using the second context model to derive the second context increment of the second syntax element based on that none of the above is satisfied. **Claim 4** The determination of using the second context model to determine the second context increment of the second syntax element is based on the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; and ii) the color format of the picture is the 4:2:0 color format or the 4:2:2 color format, and the second color component is a luma component; and iii) the color format of the picture is the 4:2:2 color format, the second color component is a chroma component, and the second syntax element indicates the prefix of the y coordinate of the last significant transform coefficient of the second color component of the block; The method according to claim 1, comprising making the determination of using the second context model to derive the second context increment of the second syntax element based on that none of the above is satisfied. **Claim 5** A method for encoding video data, the method comprising Based on whether the color format of the picture of the video data is either 4:2:0 or 4:2:2, determine which context model to use between the first context model and the second context model, and based on the color format, (a) the first context model for determining a first context increment of a first syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a first color component of a block of the picture, (b) the second context model for determining a second context increment of a second syntax element, make a determination to use, where the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture, and the second context increment of the second syntax element is (i) set the context offset ctxtOffset of the second syntax element to 18 and set the context shift ctxtShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm value of the transform size of the block, (ii) determine the second context increment of the second syntax element as (binIdx >> ctxtShift) + ctxtOffset, where binIdx is the bin index of the bin of the second syntax element, comprising, determine the first context increment of the first syntax element using the determined first context model, and determine the second context increment of the second syntax element using the determined second context model, encode the bins of the first syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the first context increment of the first syntax element, Encoding the bins of the second syntax element by applying CABAC using the context determined based on the second context increment of the second syntax element, A method comprising. **Claim 6** The method according to claim 5, wherein the first color component is a luma component and the second color component is a chroma component. **Claim 7** Said determining to use the second context model to determine the second context increment of the second syntax element is subject to the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; and ii) the color format of the picture is the 4:2:0 color format or the 4:2:2 color format, and the second color component is a luma component The method according to claim 5, comprising making said determination to use the second context model to derive the second context increment of the second syntax element based on none of the above being satisfied. **Claim 8** Said determining to use the second context model to determine the second context increment of the second syntax element is subject to the following conditions: iii) the color format of the picture is the 4:2:2 color format, the second color component is a chroma component, and the second syntax element does not indicate the prefix of the y coordinate of the last significant transform coefficient of the second color component of the block. The method according to claim 5, further based on this not being satisfied. **Claim 9** A device for decoding video data, the device comprising a memory configured to store the video data; and one or more processors coupled to the memory, the one or more processors being implemented in a circuit, Based on whether the color format of the picture of the video data is either the 4:2:0 color format or the 4:2:2 color format, determine which context model to use from among a first context model and a second context model, based on the color format, (a) The first context model for determining a first context increment of a first syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a first color component of a block of the picture; (b) The second context model for determining a second context increment of a second syntax element; determining to use, wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture, and the second context increment of the second syntax element is (i) setting a context offset ctxtOffset of the second syntax element to 18 and setting a context shift ctxtShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm value of the transform size of the block; (ii) determining the second context increment of the second syntax element as (binIdx >> ctxtShift) + ctxtOffset, where binIdx is the bin index of a bin of the second syntax element; comprising; determining the first context increment of the first syntax element using the determined first context model and determining the second context increment of the second syntax element using the determined second context model; decoding a bin of the first syntax element by applying context-adaptive binary arithmetic coding (CABAC) using a context determined based on the first context increment of the first syntax element; decoding a bin of the second syntax element by applying CABAC using a context determined based on the second context increment of the second syntax element; configured to perform; a device.

10. The device according to claim 9, wherein the first color component is a luma component and the second color component is a chroma component.

11. The one or more processors being configured to make the determination to use the second context model to determine the second context increment of the second syntax element is based on the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; and ii) the color format of the picture is the 4:2:0 color format or the 4:2:2 color format, and the second color component is a luma component none of which is satisfied, the device according to claim 9, comprising making the determination to use the second context model to derive the second context increment of the second syntax element. **Claim 12** A device for encoding video data, the device comprising a memory configured to store the video data; and one or more processors coupled to the memory, the one or more processors being implemented in circuitry, determining which context model to use from among a first context model and a second context model based on whether the color format of a picture of the video data is either the 4:2:0 color format or the 4:2:2 color format, and based on the color format (a) the first context model for determining a first context increment of a first syntax element indicating a prefix of an x or y coordinate of a last significant transform coefficient of a first color component of a block of the picture; (b) the second context model for determining a second context increment of a second syntax element; wherein the second syntax element indicates a prefix of an x or y coordinate of a last significant transform coefficient of a second color component of the block of the picture, and the second context increment of the second syntax element is (i) Set the context offset ctxtOffset of the second syntax element to 18, and set the context shift ctxtShift of the second syntax element to Max(0, log2TrafoSize - 2) - Max(0, log2TrafoSize - 4), where Max represents the maximum value function and log2TrafoSize represents the base-2 logarithm of the transform size of the block. (ii) Determine the second context increment of the second syntax element as (binIdx >> ctxtShift) + ctxtOffset, where binIdx is the bin index of the bin of the second syntax element. comprising Determine the first context increment of the first syntax element using the determined first context model, and determine the second context increment of the second syntax element using the determined second context model. Encode the bins of the first syntax element by applying context-adaptive binary arithmetic coding (CABAC) using the context determined based on the first context increment of the first syntax element. Encode the bins of the second syntax element by applying CABAC using the context determined based on the second context increment of the second syntax element. configured to perform device.

13. The device according to claim 12, wherein the first color component is a luma component and the second color component is a chroma component.

14. The determination that the one or more processors are configured to perform the determination using the second context model to determine the second context increment of the second syntax element is subject to the following conditions: i) The color format of the picture is a monochrome color format or a 4:4:4 color format. ii) The color format of the picture is the 4:2:0 color format or the 4:2:2 color format, and the second color component is a luma component. Based on none of them being satisfied, making the determination to use the second context model to derive the second context increment of the second syntax element, the device according to claim 12. **Claim 15** A computer-readable storage medium storing instructions that, when executed, cause one or more processors to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • METHOD AND APPARATUS FOR ARITHMETIC CODING OF VIDEO AND METHOD AND APPARATUS FOR ARITHMETIC CODING OF VIDEO

    JP2014533058A

  • Progressive coding of the position of the last effective coefficient

    JP2015502079A

  • Context reduction for last transform position coding

    US20120230402A1