Supporting different color formats for coefficient coding in video coding
By selecting an appropriate context model based on the color format of the video data for CABAC encoding and decoding, the problem of low encoding and decoding efficiency under different color formats is solved, thus improving the encoding and decoding efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2021-04-07
- Publication Date
- 2026-05-22
AI Technical Summary
Existing video encoding and decoding technologies fail to effectively select context models when processing video data of different color formats, resulting in low encoding and decoding efficiency.
Based on the image color format of the video data, determine which of the first and second context models to use to determine the context increment of the syntax element, and apply CABAC for encoding and decoding.
It improves the efficiency of video encoding and decoding, especially when processing images in 4:4:4 and 4:2:2 color formats, by more accurately selecting the context model, thus enhancing encoding and decoding performance.
Smart Images

Figure CN115398923B_ABST
Abstract
Description
[0001] This application claims priority to U.S. Patent Application No. 17 / 223814, filed April 6, 2021, and U.S. Provisional Patent Application No. 63 / 009292, filed April 13, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to video encoding and video decoding. Background Technology
[0003] Digital video capabilities can be integrated into a wide variety of devices, including digital televisions, digital live broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite wireless phones, so-called "smartphones," video conferencing equipment, video streaming devices, etc. Digital video devices implement video codec technologies, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Codec (AVC), ITU-T H.265 / High-Efficiency Video Codec (HEVC), and extensions to these standards. By implementing such video codec technologies, video devices can more efficiently transmit, receive, encode, decode, and / or store UI digital video information.
[0004] Video coding and decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or eliminate redundancy inherent in video sequences. For block-based video coding and decoding, video slices (e.g., video pictures or portions of video pictures) can be segmented into video blocks, which may also be referred to as codec tree units (CTUs), codec units (CUs), and / or codec nodes. Video blocks in an intra-frame coding (I) slice of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in an inter-frame coding (P or B) slice of a picture can use spatial prediction relative to reference samples in adjacent blocks within the same picture or relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention
[0005] Generally, this disclosure describes techniques for transform coefficient encoding and decoding and for encoding and decoding video data with different color formats (e.g., 4:2:2 and 4:4:4 in addition to 4:2:0). As described herein, the Basic Video Codec Test Model 5.0 (ETM 5.0) uses a context derivation process to determine the context of the syntax element prefixed with the x and y coordinates of the last valid transform coefficient of an indicator block in CABAC. The context specifies the probability of the symbol. The context derivation process of ETM 5.0 does not take into account different color formats. This may result in the selection of the context with less accurate probabilities. This disclosure describes a context derivation process for determining the context of the CABAC context of the syntax element prefixed with the x and y coordinates of the last valid transform coefficient of an indicator block, based on the color format of the image including the block. This may result in the selection of the context with more accurate probabilities, which may ultimately lead to higher encoding and decoding efficiency.
[0006] In one example, this disclosure describes a method for decoding video data, the method comprising: determining, based on the color format of an image of the video data, which of a first context model and a second context model should be used to determine a context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinates of the last valid transform coefficients of the color components of a block of the image; and decoding the binary bits of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.
[0007] In another example, this disclosure describes a method for encoding video data, the method comprising: determining, based on the color format of an image of the video data, which of a first context model and a second context model should be used to determine a context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the color components of a block of the image; and encoding the binary bits of the syntax element by applying CABAC to the context determined based on the context increment of the syntax element.
[0008] In another example, this disclosure describes an apparatus for decoding video data, the apparatus comprising: a memory configured to store video data; and one or more processors coupled to the memory, the processors being implemented in circuitry and configured to: determine, based on the color format of an image of the video data, which of a first context model and a second context model to use to determine a context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of a block of color components of the image; and decode the binary bits of the syntax element using CABAC applied to the context determined based on the context increment of the syntax element.
[0009] In another example, this disclosure describes an apparatus for encoding video data, the apparatus comprising: a memory configured to store video data; and one or more processors coupled to the memory, the processors being implemented in circuitry and configured to: determine, based on the color format of an image of the video data, which of a first context model and a second context model to determine a context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of a block of color components of the image; and encode the binary bits of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.
[0010] In another example, this disclosure describes an apparatus for decoding video data, the apparatus comprising: a component for determining, based on the color format of an image of the video data, which context model, one of a first context model and a second context model, is used to determine the context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the color components of a block of the image; and a component for decoding the binary bits of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.
[0011] In another example, this disclosure describes an apparatus for encoding video data, the apparatus comprising: a component for determining which of a first context model and a second context model to use to determine a context increment of a syntax element based on the color format of an image of the video data, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the color components of a block of the image; and a component for encoding the binary bits of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element.
[0012] In another example, this disclosure describes a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: determine, based on the color format of an image of video data, which of a first context model and a second context model should be used to determine the context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinates of the last valid transform coefficients of the color components of a block of the image; and decode the binary bits of the syntax element by using CABAC applied to the context determined based on the context increment of the syntax element.
[0013] In another example, this disclosure describes a computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to: determine, based on the color format of a picture of video data, which of a first context model and a second context model to determine the context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the color components of a block of the picture; and encode the binary bits of the syntax element by using CABAC applied to the context determined based on the context increment of the syntax element.
[0014] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will become apparent from the specification, drawings, and claims. Attached Figure Description
[0015] Figure 1 This is a block diagram illustrating an example video encoding and decoding system that can perform the techniques of this disclosure.
[0016] Figure 2A and Figure 2B This is a conceptual diagram illustrating an example quadtree binary tree (QTBT) structure and its corresponding codec tree unit (CTU).
[0017] Figure 3 This is a conceptual diagram illustrating the adaptive transform selection of inter-frame codec blocks.
[0018] Figure 4 This is a conceptual diagram illustrating the coefficient scanning method and the position of the last coefficient.
[0019] Figure 5 It is a table showing the encoding and decoding symbols of the coefficients of a block of 16 coefficients.
[0020] Figure 6 This is a block diagram illustrating an example video encoder that can perform the techniques of this disclosure.
[0021] Figure 7 This is a block diagram illustrating an example video decoder that can perform the techniques disclosed herein.
[0022] Figure 8 This is a flowchart illustrating an example method for encoding the current block.
[0023] Figure 9 This is a flowchart illustrating an example method for decoding the current block of video data.
[0024] Figure 10 This is a flowchart illustrating example operation of a video codec according to one or more technologies disclosed herein.
[0025] Figure 11 This is a flowchart illustrating example operation of a video encoder according to one or more technologies disclosed herein.
[0026] Figure 12 This is a flowchart illustrating an example operation of a video decoder according to one or more technologies disclosed herein.
[0027] Figure 13 This is a flowchart illustrating an example operation of a video codec determining context according to one or more techniques disclosed herein.
[0028] Figure 14 This is a flowchart illustrating an example operation for determining the values of brightness model variables according to one or more techniques of this disclosure. Detailed Implementation
[0029] A context model enables a video codec to determine the context used in Context-Adaptive Binary Arithmetic Coding (CABAC). Traditionally, in Essential Video Coding (EVC), a video codec (e.g., a video encoder or decoder) uses a first context model (e.g., a luma context model) to determine the context of the CABAC codec syntax element preceding the coordinates of the last valid transform coefficients for the luma component of a block, and a different second context model (e.g., a chroma context model) to determine the context of the CABAC codec syntax element preceding the coordinates of the last valid transform coefficients for the chroma component of a block. In this disclosure, valid transform coefficients are non-zero transform coefficients.
[0030] In EVC, the video codec uses these two different context models because it assumes the image color format is 4:2:0. When the image color format is 4:2:0, the chroma samples are half the number of luminance samples in both the horizontal and vertical directions. Due to this difference in the number of chroma samples compared to luminance samples, there may be different statistics regarding the values of the binary bits (bin) of the syntax elements that indicate the coordinates of the last effective transform coefficients for luminance and chroma.
[0031] However, other color formats (such as 4:4:4 and 4:2:2) are possible. In an image with a 4:4:4 color format, the number of luminance and chrominance samples are equal in both the horizontal and vertical directions. In an image with a 4:2:2 color format, the number of chrominance samples is half that of luminance samples in the horizontal direction, while the number of luminance and chrominance samples is equal in the vertical direction. Using the EVC chrominance context model with other color formats may result in poor encoding and decoding efficiency.
[0032] This disclosure describes techniques that can address this problem and thereby improve encoding / decoding efficiency. In one example, a video codec (e.g., a video encoder or decoder) can determine, based on the color format of the image in the video data, which context model—either a first or a second context model—to determine the context increment of a syntax element, which indicates a prefix of the x or y coordinates of the last valid transform coefficients of the color components of a block of image. The video codec can encode / decode (e.g., encode or decode) the bits of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element. By determining whether to use the first or second context model to determine the context increment of the syntax element, the video codec can better select the context suitable for encoding / decoding the bits of the syntax element. This can improve encoding / decoding efficiency. In some cases, the same context model can be used for both the luma and chroma components.
[0033] Figure 1 This is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. The techniques of this disclosure generally relate to encoding and / or decoding video data. Generally, video data includes any data used for processing video. Therefore, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (e.g., signaling data).
[0034] like Figure 1As shown, in this example, system 100 includes source device 102, which provides encoded video data to be decoded and displayed by destination device 116. Specifically, source device 102 provides video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can include any of a wide variety of devices, including desktop computers, mobile devices, laptops, tablets, set-top boxes, mobile phones (such as smartphones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and therefore may be referred to as wireless communication devices.
[0035] exist Figure 1 In the example, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes an input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for coefficient encoding and decoding that support different color formats. Therefore, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, the source device and destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source, such as an external camera. Similarly, destination device 116 may interface with an external display device, rather than including an integrated display device.
[0036] like Figure 1 The system 100 shown is merely an example. Generally, any digital video encoding and / or decoding device can perform techniques for encoding and decoding coefficients that support different color formats. Source device 102 and destination device 116 are merely examples of such encoding / decoding devices, where source device 102 generates encoded video data for transmission to destination device 116. This disclosure refers to an "encoding / decoding" device as a device that performs encoding and / or decoding of data. Thus, video encoder 200 and video decoder 300 represent examples of encoding / decoding devices (specifically, video encoder and video decoder). In some examples, source device 102 and destination device 116 may operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Therefore, system 100 can support one-way or two-way video transmission between source device 102 and destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0037] Generally, video source 104 represents the source of video data (i.e., raw, unencoded video data) and provides a sequence of pictures (also called “frames”) of video data to video encoder 200, which encodes the data of the pictures. Video source 104 of source device 102 may include video capture devices such as cameras, video archives containing previously captured raw video, and / or video feed interfaces that receive video from video content providers. Alternatively, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from the received order (sometimes referred to as “display order”) into an encoding / decoding order for encoding and decoding. Video encoder 200 may generate a bitstream comprising the encoded video data. The source device 102 can then output the encoded video data to a computer-readable medium 110 via the output interface 108 for reception and / or retrieval, for example, by the input interface 122 of the destination device 116.
[0038] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data, such as raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, memories 106 and 120 may store software instructions executable by, for example, video encoder 200 and video decoder 300. Although memories 106 and 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store (e.g., output from video encoder 200 and input to video decoder 300) encoded video data. In some examples, portions of memories 106 and 120 may be allocated as one or more video buffers, for example, to store raw decoded and / or encoded video data.
[0039] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to transmit encoded video data directly to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, output interface 108 can demodulate the transmitted signal including the encoded video data, and input interface 122 can demodulate the received transmitted signal. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network, such as the Internet. The communication medium can include a router, switch, base station, or any other device that can be used to facilitate communication from source device 102 to destination device 116.
[0040] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 can include any of a wide variety of distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.
[0041] In some examples, source device 102 may output encoded video data to file server 114 or another intermediate storage device that may store the encoded video data generated by source device 102. Destination device 116 may access the stored video data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing and sending encoded video data to destination device 116. File server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 may access the encoded video data from file server 114 via any standard data connection, including an internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming protocol, a download protocol, or a combination thereof.
[0042] Output interface 108 and input interface 122 can represent a wireless transmitter / receiver, a modem, a wired network component (e.g., an Ethernet card), a wireless communication component operating according to any of a variety of IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 can be configured to transmit data (e.g., encoded video data) according to cellular communication standards (such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to operate according to other wireless standards (such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee)). TM Bluetooth TM The source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include an SoC device for performing functions attributed to video encoder 200 and / or output interface 108, and destination device 116 may include an SoC device for performing functions attributed to video decoder 300 and / or input interface 122.
[0043] The technology disclosed herein can be applied to video encoding and decoding that supports any of a wide variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto a data storage medium, decoding digital video stored on a data storage medium, or other applications.
[0044] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200 (which is also used by the video decoder 300), such as syntax elements having values describing the characteristics and / or processing of video blocks or other encoding / decoding units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays decoded images of the decoded video data to the user. The display device 118 may represent any of a wide variety of display devices, such as liquid crystal displays (LCDs), plasma displays, organic light-emitting diode (OLED) displays, or other types of display devices.
[0045] although Figure 1 Not shown, but in some examples, the video encoder 200 and video decoder 300 may each be integrated with the audio encoder and / or audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams that include both audio and video in a common data stream. Where applicable, the MUX-DEMUX units may conform to the ITU H.223 multiplexer protocol or other protocols (such as User Datagram Protocol (UDP)).
[0046] The video encoder 200 and video decoder 300 can each be implemented as any of a wide variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When these techniques are implemented in part in software, the device may store software instructions in a suitable non-transitory computer-readable medium and execute these instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, and either the video encoder 200 or the video decoder 300 may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices (such as cellular phones).
[0047] The video encoder 200 and video decoder 300 can operate according to video codec standards such as ITU-T H.265 (also known as High Efficiency Video Codec (HEVC)) or its extensions (such as Multi-View and / or Scalable Video Codec Extensions)). Alternatively, the video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards such as ITU-T H.266 (also known as Universal Video Codec (VVC)). The latest draft of the VVC standard is described in the following literature: “Versatile Video Coding (Draft 8)” by Bross et al., Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC1 / SC 29 / WG 11, 17th meeting: Brussels, Belgium, January 7-17, 2020, JVET-Q2001-vE (hereinafter referred to as “VVC Draft 8”). Alternatively, the video encoder 200 and video decoder 300 can operate according to Basic Video Codec (EVC). However, the techniques disclosed herein are not limited to any particular codec standard.
[0048] Generally, video encoder 200 and video decoder 300 can perform block-based encoding and decoding of images. The term "block" generally refers to a structure that includes the data to be processed (e.g., data encoded, decoded, or otherwise used during encoding and / or decoding). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, video encoder 200 and video decoder 300 can encode and decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, video encoder 200 and video decoder 300 can encode and decode luminance and chrominance components, rather than encoding and decoding the red, green, and blue (RGB) data of the image samples, where the chrominance components may include both red and blue chrominance components. In some examples, video encoder 200 converts the received RGB format data to a YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and postprocessing units (not shown) may perform these conversions.
[0049] This disclosure generally relates to the encoding and decoding of images (e.g., encoding and decoding), including processes for encoding or decoding image data. Similarly, this disclosure can relate to the encoding and decoding of blocks of images, including processes for encoding or decoding block data (e.g., prediction and / or residual encoding and decoding). Encoded video bitstreams typically include a series of values for representing encoding / decoding decisions (e.g., encoding / decoding modes) and syntax elements that segment images into blocks. Therefore, references to encoding and decoding images or blocks should generally be understood as encoding and decoding the values of syntax elements that form images or blocks.
[0050] HEVC defines various blocks, including codec units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video codec (such as a video encoder 200) partitions a codec tree unit (CTU) into CUs according to a quadtree structure. That is, the video codec splits the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and a CU of such a leaf node can include one or more PUs and / or one or more TUs. The video codec can further partition PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a partition of a TU. In HEVC, a PU represents inter-frame prediction data, while a TU represents residual data. Intra-predicted CUs contain intra-frame prediction information, such as intra-frame mode indications.
[0051] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, the video codec (such as video encoder 200) segments the image into multiple codec tree units (CTUs). Video encoder 200 can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple segmentation types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level of segmentation based on quadtree segmentation and a second level of segmentation based on binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to codec units (CUs).
[0052] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more ternary tree (TT) partitioning methods (also known as triplet (TT)). A ternary tree or triplet partitioning is a partition that divides a block into three sub-blocks. In some examples, a ternary tree or triplet partitioning splits a block into three sub-blocks without splitting the original block through a center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.
[0053] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).
[0054] The video encoder 200 and video decoder 300 can be configured to use per-HEVC quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structures. For illustrative purposes, the description of the techniques disclosed herein is presented in relation to QTBT segmentation. However, it should be understood that the techniques disclosed herein can also be applied to video codecs configured to use quadtree segmentation or other types of segmentation.
[0055] In some examples, a CTU includes a code-decode tree block (CTB) of luma samples, two corresponding CTBs of chroma samples from an image with three sample arrays, or a CTB of samples from a monochrome image or an image encoded using three separate color planes and a syntax structure for encoding and decoding samples. For some value of N, a CTB can be an N×N sample block such that splitting the components into CTBs is a partition. A component is an array or a single sample from one of the three arrays (one luma and two chroma) that make up an image in a 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample from an array or array that makes up an image in a monochrome format. In some examples, for some values of M and N, the code-decode block is an M×N sample block such that splitting the CTB into code blocks is a partition.
[0056] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a brick can refer to a rectangular area of a CTU row within a specific slice in an image. A slice can be a rectangular area of CTUs within a specific slice column and a specific slice row in an image. A slice column is a rectangular area of CTUs whose height is equal to the image height and whose width is specified by a syntax element (e.g., as in an image parameter set). A slice row is a rectangular area of CTUs whose width is equal to the image width and whose height is specified by a syntax element (e.g., as in an image parameter set).
[0057] In some examples, a slice can be divided into multiple bricks, each brick potentially including one or more CTU rows within the slice. A slice that is not divided into multiple bricks can also be referred to as a brick. However, bricks that are a true subset of a slice may not be referred to as a slice.
[0058] The bricks in an image can also be arranged into slices. A slice can be an integer number of bricks in the image, which can be exclusively contained in a single Network Abstraction Layer (NAL) unit. In some examples, a slice consists of multiple complete slices or a continuous sequence of complete bricks that consists of only one slice.
[0059] This disclosure uses “N×N” and “N multiplied by N” interchangeably to refer to the sample size of a block (such as a CU or other video block) in the vertical and horizontal dimensions, for example, 16×16 samples or 16 by 16 samples. Generally, a 16×16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×NCU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as it does in the vertical direction. For example, a CU can include N×M samples, where M is not necessarily equal to N.
[0060] The video encoder 200 encodes video data containing representation prediction and / or residual information of the CU, as well as other information. The prediction information indicates how the CU will be predicted to form a prediction block of the CU. The residual information typically represents the sample-by-sample difference between the samples of the CU before encoding and the prediction block.
[0061] To predict the Cubic Frame (CU), the video encoder 200 typically forms a prediction block of the CU through inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU based on previously encoded data of a frame, while intra-frame prediction generally refers to predicting the CU based on previously encoded data of the same frame. To perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate prediction blocks. The video encoder 200 can typically perform motion search to identify, for example, a reference block that closely matches the CU in terms of the difference between the CU and a reference block. The video encoder 200 can use the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such differences to calculate a difference metric to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.
[0062] Some examples of VVC also provide an affine motion compensation mode, which can be viewed as an inter-frame prediction mode. In the affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular motion types).
[0063] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of VVC provide 67 intra-frame prediction modes, including various directional modes as well as planar and DC modes. Generally, the video encoder 200 selects the intra-frame prediction mode that describes the adjacent samples of the current block (e.g., the block of the CU) upon which the samples of the current block are based. Assuming that the video encoder 200 encodes and decodes the CTU and CU in raster scan order (from left to right, from top to bottom), such samples are typically located above, to the upper left, or to the left of the current block in the same frame as the current block.
[0064] The video encoder 200 encodes data representing the prediction mode of the current block. For inter-frame prediction modes, for example, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used, along with the corresponding motion information. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may use Advanced Motion Vector Prediction (AMVP) or merging modes to encode motion vectors. The video encoder 200 may use similar modes to encode motion vectors for affine motion compensation modes.
[0065] After prediction (such as intra-frame or inter-frame prediction for a block), the video encoder 200 can compute residual data for that block. Residual data (such as residual blocks) represents the sample-by-sample difference between predicted blocks formed using the corresponding prediction modes. The video encoder 200 can apply one or more transforms to the residual blocks to generate transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply Discrete Cosine Transform (DCT), integer transform, wavelet transform, or conceptually similar transforms to the residual video data. Additionally, the video encoder 200 can apply a second transform after the first transform, such as Mode Correlated Inseparable Quadratic Transform (MDNSST), Signal Correlation Transform, Karhunen-Loeve Transform (KLT), etc. The video encoder 200 generates transform coefficients after applying one or more transforms.
[0066] As described above, after any transform that generates the transform coefficients, the video encoder 200 can perform quantization on the transform coefficients. Quantization generally refers to a process in which transform coefficients are quantized to minimize the amount of data used to represent them, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 can round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bit-by-bit right shift of the value to be quantized.
[0067] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place transform coefficients with higher energy (and therefore lower frequency) before the vector and transform coefficients with lower energy (and therefore higher frequency) after the vector. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to generate a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy encode the syntax elements describing the one-dimensional vector, for example, according to Context Adaptive Binary Arithmetic Codec (CABAC). Such syntax elements can include the last valid transform coefficient as described in this disclosure. The video encoder 200 can also entropy encode the values of syntax elements used to describe metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.
[0068] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. The context may involve, for example, whether the neighboring values of a symbol are zero. Probability determination can be based on the context assigned to the symbol.
[0069] The video encoder 200 can also generate, for example, syntax data (e.g., block-based syntax data, image-based syntax data, and sequence-based syntax data) or other syntax data (e.g., sequence parameter sets (SPS), image parameter sets (PPS), or video parameter sets (VPS)) destined for the video decoder 300 in image headers, block headers, and slice headers. The video decoder 300 can similarly decode this syntax data to determine how to decode the corresponding video data.
[0070] In this way, the video encoder 200 can generate a bitstream that includes encoded video data (e.g., syntax elements describing the segmentation of an image into blocks (e.g., CUs) and prediction and / or residual information for the blocks). Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.
[0071] Generally, the video decoder 300 performs a process reciprocal to that performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC to decode the values of the syntax elements of the bitstream in a manner substantially similar to, but reciprocal to, the CABAC encoding process of the video encoder 200. Syntax elements can define segmentation information for segmenting images into CTUs and for segmenting each CTU according to a corresponding segmentation structure (such as a QTBT structure) to define the CUs of the CTU. Syntax elements can also define prediction and residual information for blocks (e.g., CUs) of the video data.
[0072] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inverse transform the quantized transform coefficients of the block to reproduce the residual block of the block. The video decoder 300 uses the signaled prediction mode (intra-frame or inter-frame prediction) and associated prediction information (e.g., motion information from inter-frame prediction) to form the prediction block of the block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along block boundaries.
[0073] As described above, the video encoder 200 and video decoder 300 can apply CABAC encoding and decoding to the values of syntax elements. To apply CABAC encoding to syntax elements, the video encoder 200 can binarize the values of the syntax elements to form a sequence of one or more bits, referred to as a "bin". Each binary bit can be associated with a corresponding binary bit index (binIdx). Furthermore, the video encoder 200 can identify the encoding / decoding context. The encoding / decoding context can identify the probability that a binary bit has a specific value. For example, the encoding / decoding context can indicate a probability of 0.7 for encoding / decoding a binary bit with a value of 0, and a probability of 0.3 for encoding / decoding a binary bit with a value of 1. After identifying the encoding / decoding context, the video encoder 200 can split an interval into a lower sub-interval and an upper sub-interval. One sub-interval can be associated with a value of 0, and the other sub-interval can be associated with a value of 1. The width of the sub-interval can be proportional to the probability indicated by the identified encoding / decoding context for the associated value. If a bit of a syntax element has a value associated with a lower sub-interval, the encoded value can be equal to the lower boundary of the lower sub-interval. If the same bit of a syntax element has a value associated with an upper sub-interval, the encoded value can be equal to the lower boundary of the upper sub-interval. To encode the next bit of a syntax element, the video encoder 200 can repeat these steps if the interval is a sub-interval associated with the value of the encoded bit. When the video encoder 200 repeats these steps for the next bit, it can use a modified probability based on the probability indicated by the recognized encoding / decoding context and the actual value of the encoded bit.
[0074] When the video decoder 300 performs CABAC decoding on the value of a syntax element, it can recognize the encoding / decoding context. The video decoder 300 can then split an interval into a lower sub-interval and an upper sub-interval. One sub-interval can be associated with a value of 0, and the other with a value of 1. The width of the sub-interval can be proportional to the probability indicated by the recognized encoding / decoding context for the associated value. If the encoded value is within the lower sub-interval, the video decoder 300 can decode the bits containing the value associated with the lower sub-interval. If the encoded value is within the upper sub-interval, the video decoder 300 can decode the bits containing the value associated with the upper sub-interval. To decode the next bit of the syntax element, the video decoder 300 can repeat these steps if the interval is a sub-interval that includes the encoded value. When the video decoder 300 repeats these steps for the next bit, it can use a modified probability based on the probability indicated by the recognized encoding / decoding context and the decoded bits. The video decoder 300 can then binarize the binary bits to recover the values of the syntax elements.
[0075] In some cases, the video encoder 200 can use bypass CABAC encoding / decoding to encode and decode binary bits. Performing bypass CABAC encoding / decoding on binary bits may be less computationally expensive than performing regular CABAC encoding / decoding. Furthermore, performing bypass CABAC encoding / decoding allows for a higher degree of parallelization and throughput. Binary bits encoded using bypass CABAC encoding / decoding can be referred to as "bypass bits." Grouping bypass bits together can increase the throughput of both the video encoder 200 and the video decoder 300. A bypass CABAC encoding / decoding engine can encode and decode several binary bits in a single cycle, while a regular CABAC encoding / decoding engine can only encode and decode a single binary bit in a single cycle. A bypass CABAC encoding / decoding engine may be simpler because it does not select a context and can assume a 1 / 2 probability for two symbols (0 and 1). Therefore, in bypass CABAC encoding / decoding, the interval is directly split in half.
[0076] This disclosure generally relates to "signaling notification" of certain information, such as syntax elements. The term "signaling notification" can generally refer to communication relating to the value of a syntax element and / or other data used for decoding encoded video data. That is, the video encoder 200 can signal the value of a syntax element within the bitstream. Generally, signaling notification refers to generating a value within the bitstream. As described above, the source device 102 can transmit the bitstream to the destination device 116 substantially in real-time or non-real-time (such as when syntax elements are stored in storage device 112 for later retrieval by the destination device 116).
[0077] Figure 2A and Figure 2B This is a conceptual diagram illustrating an example Quadtree Binary Tree (QTBT) structure 130 and a corresponding Code-Decoder Tree Unit (CTU) 132. Solid lines represent quadtree splits, and dashed lines represent binary tree splits. In each split (i.e., non-leaf) node of the binary tree, a signal is sent to indicate which split type (i.e., horizontal or vertical) is used, where in this example, 0 indicates a horizontal split and 1 indicates a vertical split. For quadtree splits, since a quadtree node splits a block horizontally and vertically into four equal-sized sub-blocks, it is not necessary to indicate the split type. Accordingly, the video encoder 200 can encode syntax elements (such as split information) at the region tree level (i.e., solid lines) and the prediction tree level (i.e., dashed lines) of the QTBT structure 130, and the video decoder 300 can decode these syntax elements. The video encoder 200 can encode video data (such as prediction and transform data) of the CU represented by the terminal leaf nodes of the QTBT structure 130, and the video decoder 300 can decode this video data.
[0078] Generally speaking, Figure 2B The CTU 132 can be associated with parameters that define the block size corresponding to the nodes at the first and second levels of the QTBT structure 130. These parameters may include the CTU size (the size of the CTU 132 in sample points), the minimum quadtree size (MinQTSize, which represents the minimum allowed quadtree leaf node size), the maximum binary tree size (MaxBTSize, which represents the maximum allowed binary tree root node size), the maximum binary tree depth (MaxBTDepth, which represents the maximum allowed binary tree depth), and the minimum binary tree size (MinBTSize, which represents the minimum allowed binary tree leaf node size).
[0079] The root node corresponding to a CTU in a QTBT structure can have four child nodes at the first level of the QTBT structure, each of which can be segmented according to a quadtree partition. That is, a first-level node is a leaf node (without child nodes) or has four child nodes. An example of QTBT structure 130 represents such a node as including a parent node and child nodes with solid-line branches. If a first-level node is not larger than the maximum allowed binary tree root node size (MaxBTSize), the node can be further partitioned by the corresponding binary tree. The binary tree partitioning of a node can be iterated until the nodes generated from the partitioning reach the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). An example of QTBT structure 130 represents such a node as having dashed-line branches. The binary tree leaf nodes are called codec units (CUs), which are used for prediction (e.g., intra-picture or inter-picture prediction) and transformation without any further partitioning. As discussed above, CUs can also be referred to as “video chunks” or “blocks”.
[0080] In one example of a QTBT partitioning structure, the CTU size is set to 128×128 (luminance samples and two corresponding 64×64 chrominance samples), MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. Quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If the quadtree leaf node is 128×128, it will not be further partitioned by the binary tree because its size exceeds MaxBTSize (i.e., 64×64 in this example). Otherwise, the quadtree leaf node will be further partitioned by the binary tree. Therefore, the quadtree leaf node is also the root node of the binary tree, and its binary tree depth is 0. When the binary tree depth reaches MaxBTDepth (4 in this example), further partitioning is not allowed. When the width of a binary tree node equals MinBTSize (4 in this example), it means that further vertical splitting is not allowed. Similarly, a binary tree node with a height equal to MinBTSize means that further horizontal splitting is not allowed for that binary tree node. As mentioned above, the leaf nodes of the binary tree are called CUs and are further processed according to the prediction and transformation without further splitting.
[0081] At the 124th MPEG meeting held in Macau, China, requirements for new video codecs were released. At the 125th MPEG meeting held in Marrakech, Morocco, submitted CfP responses were evaluated, and the technology from proposal m46354 was selected as the basis for the working draft and test model of the basic video codec standard. The following sections of this disclosure provide a description of the transform codec approach used in MPEG5 EVC and implemented in ETM5.0.
[0082] Similar to traditional hybrid video codecs, a Discrete Cosine Transform (DCT2) is applied to the residual block between the original block and the corresponding predicted block. To support a 64×64 pipeline, the maximum allowed transform size is set to 64. If the length of an edge in a CU exceeds the maximum transform size, that edge is automatically split into two partitions.
[0083] In addition to the normal DCT2 transform, the Adaptive Transform Selection (ATS) method can be used for both intra-frame and inter-frame prediction. Table 1 shows the basis functions of the core design of the Adaptive Transform Selection method.
[0084] Table 1. Transform basis functions for DCT-8 and DST-7 with N-point input
[0085]
[0086] ATS is applied to blocks that are smaller than 32 pixels in both width and height. If the width or height exceeds 32 pixels in length, ATS is not considered for that block.
[0087] For blocks that have undergone intra-frame encoding / decoding, a flag is used to signal to the video decoder 300 whether to apply ATS. If the video encoder 200 selects to use ATS as the core transform in the CU, the video encoder 200 signals to the video decoder 300 to notify two additional flags to indicate which type of transform is used for the horizontal and vertical directions, respectively. A value of 0 indicates that DST-7 is used, while a value of 1 indicates that DCT-8 is used.
[0088] For a CU with residuals and inter-frame prediction (i.e., a CU with residual blocks and inter-frame prediction), the video encoder 200 can signal whether the entire residual block or a sub-part of the residual block should be decoded. When only a sub-part of the residual block should be encoded / decoded, that sub-part of the residual block is encoded using the inferred transform type, and the other sub-parts of the residual block are set to zero. The sub-part location information and the corresponding transform type are as follows: Figure 3 As shown. Figure 3 This is a conceptual diagram illustrating adaptive transform selection for block 150 after inter-frame encoding and decoding. Figure 3In the example, the sub-part 152 containing residual information is indicated as 152. The sub-part containing residual information can be half or a quarter of the current CU size. ATS is allowed for CUs with a width and height not exceeding 64. Figure 3 It also demonstrates that the transform type is derived based on the position of the sub-block, rather than signaling the transform type as is done for intra-frame encoded CUs. For example, the horizontal and vertical transforms for the position 0 sub-block are DCT-8 and DST-7, respectively. The transform is set to DCT-2 when at least one side of the residual TU is greater than 32.
[0089] After the transform, the transformed coefficients are scalar quantized. The quantization parameter (QP) ranges from 0 to 51, and the scaling factor (SF) corresponding to each QP is defined by a lookup table. The video encoder 200 can scan the transform coefficients of the encoded / decoded block in a predefined scan pattern after quantization and entropy encode them. To further utilize the statistical properties of the transform coefficients, a bit-plane-based coefficient encoding / decoding method, the so-called Advanced Coefficient Codec (ADCC), is used in the main EVC profile instead of the currently used run-length encoding / decoding method. The ADCC method utilizes the following design elements:
[0090] 1. Fixed zigzag scanning pattern.
[0091] 2. Send a signal to notify the coordinates of the last non-zero transform coefficient in the scanning sequence.
[0092] 3. Analyze the transformation coefficients in the 16 blocks.
[0093] 4. Signal the transformation coefficients within each processing block as a sequence of importance and level flags, symbol flags, and remaining levels.
[0094] Of these symbols (i.e., importance and level flags, symbol flags, and remaining levels), the binary bits of sigMapFlag, flagLevelA, and flagLevelB are encoded using an adaptive context model; the binarized binary bits of signFlag and levelRem are encoded using a bypass mode. The value of sigMapFlag indicates whether the corresponding transform coefficient is a valid transform coefficient. The value of flagLevelA indicates whether the level of the corresponding transform coefficient is greater than or equal to level A. The value of flagLevelB indicates whether the level of the corresponding transform coefficient is greater than or equal to level B. The value of levelRem indicates the remaining portion of the transform coefficient's level.
[0095] To reduce the number of bits involved in context encoding / decoding, explicit flagLevelA and flagLevelB can be adaptively switched to levelRem encoding / decoding. LevelRem encoding / decoding uses Golomb codes for binarization and employs a bypass mode with equal probability. To improve throughput, the number of flagLevelA and flagLevelB symbols explicitly encoded / decoded within the encoded / decoded chunk is limited. Only the first N flagLevelA and the first M flagLevelB symbols are encoded / decoded. Explicit encoding / decoding of these symbols is omitted when a specified threshold is met.
[0096] Figure 4 A visualization of the method is provided. Specifically, Figure 4 This is a conceptual diagram illustrating the coefficient scanning method and the final transformation coefficient position 160. Figure 4 The transform coefficients of block 162 are shown. The arrows indicate the position of the zigzag scan pattern starting from the "last" valid transform coefficient, with respect to the position along the zigzag scan pattern starting from the DC transform coefficient 164.
[0097] Figure 5 Table 170 shows the encoded and decoded symbols of the coefficients of a block of 16 coefficients. In other words, Figure 5 An example of encoded and decoded symbols for blocks of 16 transform coefficients is given, where N = 1. In Figure 5 In the text, the symbol with omitted signaling is marked with an X.
[0098] As shown in Table 2 below, a portion of the transform coefficient encoding and decoding is defined in MPEG5 EVC (i.e., a portion of the transform_unit syntax structure).
[0099] Table 2
[0100]
[0101] As reproduced in Table 3 below, the residual_coding syntax structure referenced in Table 2 above can be implemented. As shown in Table 3, the values log2TbWidth and log2TbHeight are passed from the transform_unit syntax structure to the residual_coding syntax structure, and from the residual_coding syntax structure to the residual_coding_adv syntax structure.
[0102] Table 3
[0103]
[0104] The residual_coding syntax structure derived from the transform_unit syntax structure can be implemented as shown in Table 4 below, which shows a portion of the residual_coding_adv syntax structure.
[0105] Table 4
[0106]
[0107]
[0108]
[0109]
[0110] In Table 4 above, the syntax element `last_sig_coeff_x_prefix` can specify a prefix for the column (x) position of the last valid transform coefficient in the scan order within the transform block. The syntax element `last_sig_coeff_y_prefix` can specify a prefix for the row (y) position of the last valid transform coefficient in the scan order within the transform block. If `last_sig_coeff_x_suffix` does not exist, the column position (i.e., x-coordinate) of the last valid transform coefficient (LastSignificantCoeffX) can be equal to the value of `last_sig_coeff_x_prefix`. Otherwise (if `last_sig_coeff_x_suffix` exists), the following applies:
[0111] LastSignificantCoeffX=(1<<((last_sig_coeff_x_prefix>>1)-1))*(2+(last_sig_coeff_x_prefix&1))+last_sig_coeff_x_suffix
[0112] Similarly, if `last_sig_coeff_y_suffix` does not exist, the row position (i.e., y-coordinate) of the last valid transformation coefficient (LastSignificantCoeffY) can be equal to `last_sig_coeff_y_prefix`. Otherwise (if `last_sig_coeff_y_suffix` exists), the following applies:
[0113] LastSignificantCoeffY=(1<<((last_sig_coeff_y_prefix>>1)-1))*(2+(last_sig_coeff_y_prefix&1))+last_sig_coeff_y_suffix
[0114] Video codecs (e.g., video encoder 200 or video decoder 300) can perform a derivation of the context increment (ctxInc) for the syntax elements last_sig_coeff_x_prefix and last_sig_coeff_y_prefix. The video codec can determine the context based on the context increment. In some examples, to determine the context based on the context increment, the video codec can determine the context index by adding the context increment to the context offset value (ctxIdxOffset) of the syntax elements (e.g., last_sig_coeff_x_prefix and last_sig_coeff_y_prefix). The context offset value can be equal to the lowest context index value that can be used with the syntax element. The derivation process implemented in ETM 5.0 is described below.
[0115] The inputs to this process are the variable binIdx, the color component index cIdx, and the associated transform size log2TrafoSize, which is log2TbWidth for last_sig_coeff_x_prefix and log2TbhHeight for last_sig_coeff_y_prefix.
[0116] The output of this process is the variable ctxInc.
[0117] The variables ctxOffset and ctxShift are derived as follows:
[0118] – If cIdx equals 0, then the following applies:
[0119] – If log2TrafoSize is less than 6, then ctxOffset is set to equal to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2), and ctxShift is set to equal to (log2TrafoSize+1)>>2.
[0120] Otherwise (log2TrafoSize is greater than or equal to 6), ctxOffset is set to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and ctxShift is set to equal to ((log2TrafoSize+1)>>2)<<1.
[0121] Otherwise (cIdx is greater than 0), ctxOffset is set to 18 and ctxShift is set to Max(0, log2TrafoSize–2)–Max(0, log2TrafoSize–4).
[0122] The variable ctxInc is derived as follows:
[0123] ctxInc=(binIdx>>ctxShift)+ctxOffset (9-1)
[0124] In the above text, `log2TrafoSize` specifies the associated transform size, which is `log2TbWidth` for `last_sig_coeff_x_prefix` and `log2TbhHeight` for `last_sig_coeff_y_prefix`. The variable `log2TbWidth` is equal to the base-2 logarithm of the transform block width. The video codec can repeat this operation for each bit of `last_sig_coeff_x_prefix` and `last_sig_coeff_y_prefix`. The bits of `last_sig_coeff_x_prefix` and `last_sig_coeff_y_prefix` can be individual binary numbers of the binarized versions of `last_sig_coeff_x_prefix` and `last_sig_coeff_y_prefix`.
[0125] Transform coefficient encoding and decoding as specified in MPEG5 EVC is not intended for encoding and decoding video signals with color formats other than 4:2:0, i.e., when the chroma component data is presented at one-quarter resolution compared to the luminance component. Implementing support for color formats other than 4:2:0 can improve the encoding and decoding efficiency of compressing video with color formats other than 4:2:0.
[0126] This disclosure describes techniques for enabling the encoding and decoding of video signals with color formats other than 4:2:0. Therefore, the techniques of this disclosure can improve the encoding and decoding efficiency for compressing video with color formats other than 4:2:0. The techniques of this disclosure can be used together or separately.
[0127] According to the first technique of this disclosure, a video codec (e.g., video encoder 200 or video decoder 300) can invoke residual encoding / decoding operations for non-luminance (cIdx not equal to 0) color components, the block size of which is a function of the chroma format indicator. The semantics of the chroma_format_idc syntax element in ETM 5.0 are described below.
[0128] chroma_format_idc specifies the chroma sampling relative to the luminance sampling as defined in Section 6.2. The value of chroma_format_idc should be in the range of 0 to 3 (inclusive).
[0129] Based on the value of chroma_format_idc, the values of variables SubWidthC and SubHeightC are assigned as specified in Section 6.2, and the variable ChromaArrayType is assigned as follows:
[0130] – If chroma_format_idc equals 0, then ChromaArrayType is set to equal to 0.
[0131] Otherwise, ChromaArrayType is set to equal chroma_format_idc.
[0132] In addition, the following text from Article 6.2 of ETM 5.0 describes the color format.
[0133] Table 6-1 specifies the variables SubWidthC and SubHeightC, which depend on the chroma format sampling structure specified by chroma_format_idc. ISO / IEC may specify other values for chroma_format_idc, SubWidthC, and SubHeightC in the future.
[0134] Table 6-1 shows the SubWidthC and SubHeightC values derived from chroma_format_idc.
[0135] chroma_format_idc Color format SubWidthC SubHeightC 0 Monochrome 1 1 1 4:2:0 2 2 2 4:2:2 2 1 3 4:4:4 1 1
[0136] In monochrome sampling, there is only one sample array, which is usually considered to be the brightness array.
[0137] In 4:2:0 sampling, the height and width of each chroma array in the two chroma arrays are half that of the luminance array.
[0138] In 4:2:2 sampling, the height and width of each chroma array in the two chroma arrays are half that of the luminance array.
[0139] In 4:4:4 sampling, the height and width of each chroma array in the two chroma arrays are the same as those of the luminance array.
[0140] According to the first technique of this disclosure, the text of ETM 5.0 can be modified as shown in Table 5 below to take into account the different values of SubWidthC and SubHeightC in different color formats as determined in Table 6-1. In this disclosure, the proposed modification to the transform_unit syntax structure of EVC uses <!>...< / !> Use labels to indicate.
[0141] Table 5
[0142]
[0143] Since the values of SubWidthC and SubHeightC depend on the color format, and the values of TrafoLog2Width and TrafoLog2Height are modified based on SubWidthC and SubHeightC in Table 5, the values of TrafoLog2Width and TrafoLog2Height are indeed modified based on the color format. Therefore, for different color formats, the correct values of TrafoLog2Width and TrafoLog2Height can be used in the residual_coding and residual_coding_adv syntax structures shown in Tables 3 and 4. Using the correct values of TrafoLog2Width and TrafoLog2Height in the residual_coding_adv syntax structure allows for signaling the correct syntax elements for different color formats.
[0144] Therefore, in the example configuration of the first technique disclosed herein, the video codec (e.g., video encoder 200 or video decoder 300) can determine the parameters of the residual encoding / decoding operation based on the chroma format indicator (chroma_format_idc) applicable to the video data block. For example, the video codec can determine one parameter of the residual encoding / decoding operation as TrafoLog2Width – SubWidthC+1, and another parameter of the residual encoding / decoding operation as TrafoLog2Height – SubHeightC+1. The video codec can then perform the residual encoding / decoding operation based on the determined parameters of the residual encoding / decoding operation to encode and decode the residual data of the non-luminance components of the block.
[0145] The second technique of this disclosure can improve the performance of non-luminance (cIdx not equal to 0) coefficient encoding and decoding by allowing more efficient context modeling. According to the second technique of this disclosure, for color formats other than 4:2:0 (i.e., color_format_idc not equal to 1), a luminance context model (i.e., a context model for components (such as luminance components) of a full-resolution video signal) can be enabled for context-coded syntax elements in the chroma components. The proposed changes to ETM 5.0 according to the second technique of this disclosure...< / !> Use labels to indicate.
[0146] The derivation of ctxInc for the syntax elements last_sig_coeff_x_prefix and last_sig_coeff_y_prefix
[0147] The inputs to this process are the variable binIdx, the color component index cIdx, and the associated transform size log2TrafoSize, which is log2TbWidth for last_sig_coeff_x_prefix and log2TbHeight for last_sig_coeff_y_prefix, respectively.
[0148] The output of this process is the variable ctxInc.
[0149] The variable enableLumaModel is set to FALSE and modified as follows:
[0150] – If chromaArrayType is equal to 0 or 3, the variable enableLumaMode is set to TRUE.
[0151] Otherwise, if chromaArrayType equals 1 or 2 and cIdx equals 0, the variable enableLumaModel is set to TRUE.
[0152] Otherwise, if chromaArrayType equals 2, cIdx is not equal to 0, and the syntax element to be parsed is last_sig_coeff_y_prefix, then the variable enableLumaModel is set to TRUE.< / !>
[0153] The variables ctxOffset and ctxShift are derived as follows:
[0154] – If <!> enableLumaModel equals TRUE< / !> Then the following situations apply:
[0155] If log2TrafoSize is less than 6, then ctxOffset is set to equal to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2), and ctxShift is set to equal to (log2TrafoSize+1)>>2.
[0156] Otherwise (log2TrafoSize is greater than or equal to 6), ctxOffset is set to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and ctxShift is set to equal to ((log2TrafoSize+1)>>2)<<1.< / ^^>
[0157] Otherwise (<!> enableLumaModel equals FALSE)< / !> ), ctxOffset is set to equal to 18, and ctxShift is set to equal to Max(0,log2TrafoSize–2)–Max(0,log2TrafoSize–4).
[0158] The variable ctxInc is derived as follows:
[0159] ctxInc=(binIdx>>ctxShift)+ctxOffset (9-2)
[0160] Therefore, as shown above, in some cases, the luminance context model (i.e., using <^^>...)< / ^^> The text marked with a tag can be used for the luminance component and also for the chrominance component. Therefore, according to the second technique of this disclosure, a video codec (e.g., video encoder 200 or video decoder 300) can use a context model to derive the context increment of a first syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix). The first syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the luminance component of a block of video data. Additionally, the video codec can apply CABAC to the binary bits of the first syntax element using the context determined based on the context increment of the first syntax element. The video codec can use the same or different context models to derive the context increment of a second syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix), depending in part on the color format of the image. The second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the chrominance component of a block. The video codec can apply CABAC to the binary bits of the second syntax element using a context determined based on the context increment of the second syntax element.
[0161] Therefore, in some examples, the video encoder 200 can determine, based on the color format of the image in the video data, which context model (e.g., defined by the value of the variable `enableLumaModel`) to determine the context increment (ctxInc) of a syntax element (e.g., `last_sig_coeff_x_prefix` or `last_sig_coeff_y_prefix`), which indicates the prefix of the x or y coordinate of the last valid transform coefficient of the color component of the image block. The video encoder 200 can encode the binary bits of the syntax element by applying CABAC to the context determined based on the context increment of the syntax element. In some examples, as described above, the index of the context (with respect to a predefined set of contexts) can be determined based on the sum of the context increment and a predefined offset value. Similarly, the video decoder 300 can determine, based on the color format of the image in the video data, which context model—either a first or a second—to determine the context increment (ctxInc) of a syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix), that indicates the prefix of the x or y coordinates of the last valid transform coefficient of the color components of a block of the image. The video decoder 300 can encode the binary bits of the syntax element using CABAC applied to the context determined based on the context increment of the syntax element.
[0162] In some examples, using the first context model includes one of the following: setting the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) based on the base-2 logarithm of the block's transform size being less than 6; and setting the context shift to be equal to (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size. Alternatively, if the base-2 logarithm of the transform size is greater than or equal to 6, the video codec may set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift to equal to ((log2TrafoSize+1)>>2)<<1. The video codec can determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bit index of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset, where the applicable syntax element is either the first syntax element or the second syntax element. In some examples, using the second context model includes: setting the context offset ctxOffset of the second syntax element to 18; and setting the context shift ctxShift of the applicable syntax element to Max(0,log2TrafoSize–2)–Max(0,log2TrafoSize–4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the block's transform size. In addition, using a second context model may include determining the context increment of the second syntax element as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the applicable syntax element.
[0163] Figure 6 This is a block diagram illustrating an example video encoder 200 that can perform the techniques disclosed herein. Figure 6 This disclosure is provided for illustrative purposes and should not be construed as limiting the techniques extensively illustrated and described herein. For illustrative purposes, this disclosure describes a video encoder 200 based on VVC (ITU-T H.266, under development), HEVC (ITU-T H.265), and EVC technologies. However, the techniques of this disclosure can be implemented by video encoding devices configured to conform to other video codec standards.
[0164] exist Figure 6In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filtering unit 216, a decoded picture buffer (DPB) 218, and an entropy encoding unit 220. Any one or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filtering unit 216, DPB 218, and entropy encoding unit 220 can be implemented in one or more processors or in processing circuitry. For example, units of the video encoder 200 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0165] The video data storage device 230 can store video data that will be encoded by the components of the video encoder 200. The video encoder 200 can obtain data from, for example, a video source 104 (…). Figure 1 The video encoder 200 receives video data stored in video data memory 230. DPB 218 can act as a reference picture memory, storing reference video data for use by the video encoder 200 when predicting subsequent video data. Video data memory 230 and DPB 218 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, video data memory 230 can be on-chip (as shown) with other components of the video encoder 200, or off-chip relative to those components.
[0166] In this disclosure, references to video data memory 230 should not be construed as being limited to memory inside video encoder 200 (unless specifically described therein) or memory outside video encoder 200 (unless specifically described therein). Rather, references to video data memory 230 should be understood as a reference memory storing video data received by video encoder 200 for encoding (e.g., video data of the current block to be encoded). Figure 1 The memory 106 can also provide temporary storage for the outputs from various units of the video encoder 200.
[0167] It shows Figure 6 Various units help understand the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits are circuits that provide a specific function and are pre-configured regarding the operations that can be performed. Programmable circuits are circuits that can be programmed to perform various tasks and provide flexible functionality in terms of the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units can be integrated circuits.
[0168] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core, formed according to programmable circuitry. In an example where the operation of the video encoder 200 is performed using software executed by programmable circuitry, memory 106 ( Figure 1 The video encoder 200 may store instructions (e.g., object code) of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.
[0169] The video data storage unit 230 is configured to store received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be the raw video data to be encoded.
[0170] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units to perform video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-block copying unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.
[0171] Mode selection unit 202 typically coordinates multiple coding passes to test combinations of coding parameters and the resulting rate-distortion values for these combinations. Coding parameters may include segmenting the CTU into CUs, the prediction mode of the CUs, the transformation type of the CU residual data, and the quantization parameters of the CU residual data. Mode selection unit 202 can ultimately select a combination of coding parameters that yields a better rate-distortion value than other test combinations.
[0172] The video encoder 200 can segment images retrieved from the video data storage 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can segment the CTUs of the image according to a tree structure (such as the QTBT structure described above or the quadtree structure of HEVC). As mentioned above, the video encoder 200 can form one or more CUs by segmenting CTUs according to a tree structure. Such CUs can also generally be referred to as "video blocks" or "blocks".
[0173] Generally, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate predicted blocks for the current block (e.g., the current CU, or the overlapping portion of PU and TU in HEVC). For inter-frame prediction of the current block, motion estimation unit 222 may perform motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously encoded / decoded pictures stored in DPB 218). Specifically, motion estimation unit 222 may calculate values representing the similarity between the potential reference block and the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. Motion estimation unit 222 may typically perform these calculations using sample-by-sample differences between the current block and the reference block being considered. Motion estimation unit 222 may identify the reference block with the lowest value obtained from these calculations, indicating the reference block that most closely matches the current block.
[0174] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the position of a reference block in a reference image relative to the position of a current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data for reference blocks. As another example, if the motion vectors have fractional sample precision, motion compensation unit 224 can interpolate the values of prediction blocks according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by the corresponding motion vectors and combine the retrieved data, for example, by sample-by-sample averaging or weighted averaging.
[0175] As another example, for intra-prediction or intra-prediction codec, intra-prediction unit 226 can generate a prediction block based on samples adjacent to the current block. For example, in directional mode, intra-prediction unit 226 can typically mathematically combine the values of adjacent samples and fill these calculated values across the current block in defined directions to generate a prediction block. As another example, in DC mode, intra-prediction unit 226 can calculate the average of the adjacent samples of the current block and generate a prediction block to include this resulting average for each sample of the prediction block.
[0176] Mode selection unit 202 provides the prediction block to residual generation unit 204. Residual generation unit 204 receives the original, uncoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block of the current block. In some examples, residual generation unit 204 can also determine the differences between sample values in the residual block to generate the residual block using Residual Differential Pulse Code Modulation (RDPCM). In some examples, residual generation unit 204 can be formed using one or more subtractor circuits performing binary subtraction.
[0177] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As mentioned above, the size of a CU can refer to the size of its luma codec block, and the size of a PU can refer to the size of the luma prediction unit of the PU. Assuming a specific CU size is 2N×2N, video encoder 200 can support PU sizes of 2N×2N or N×N for intra-frame prediction, and symmetrical PU sizes of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric segmentation for PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.
[0178] In the example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luma codec block and a corresponding chroma codec block. As mentioned above, the size of the CU can refer to the size of the luma codec block of the CU. The video encoder 200 and the video decoder 300 can support CU sizes of 2N×2N, 2N×N, or N×2N.
[0179] For other video codec techniques (such as intra-block copy mode codec, affine mode codec, and linear model (LM) mode codec), mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the codec technique. In some examples (such as palette mode codec), mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block should be reconstructed based on a selected palette. In this mode, mode selection unit 202 can provide these syntax elements to entropy coding unit 220 for encoding.
[0180] As described above, the residual generation unit 204 receives video data of the current block and the corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.
[0181] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a discrete cosine transform (DCT), direction transformation, Karhunen-Loeve transform (KLT), or conceptually similar transformations to the residual block. In some examples, transform processing unit 206 may perform multiple transformations on the residual block, such as primary and secondary transformations (e.g., rotation transformations). In some examples, transform processing unit 206 does not apply any transformations to the residual block.
[0182] Quantization unit 208 can quantize the transform coefficients in the transform coefficient block to generate a quantized transform coefficient block. Quantization unit 208 can quantize the transform coefficients of the transform coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and therefore, the quantized transform coefficients may have lower precision than the original transform coefficients generated by transform processing unit 206.
[0183] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform to the quantized transform coefficient block, respectively, to reconstruct the residual block based on the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although potentially with some degree of distortion). For example, the reconstruction unit 214 can add the samples of the reconstructed residual block to the corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.
[0184] Filtering unit 216 can perform one or more filtering operations on the reconstructed block. For example, filtering unit 216 can perform a deblocking operation to reduce block artifacts along the edges of the CU. In some examples, the operations of filtering unit 216 can be skipped.
[0185] The video encoder 200 stores reconstructed blocks in the DPB 218. For example, in an example where the operation of the filtering unit 216 is not required, the reconstructed unit 214 can store the reconstructed blocks in the DPB 218. In an example where the operation of the filtering unit 216 is required, the filtering unit 216 can store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve reference images formed based on the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction of blocks in subsequent encoded images. Additionally, the intra-frame prediction unit 226 can use the reconstructed blocks of the current image in the DPB 218 to perform intra-frame prediction of other blocks in the current image.
[0186] Generally, entropy coding unit 220 can entropy-encode syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy-encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy-encode predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, another example of video data, to generate entropy-encoded data. For example, entropy coding unit 220 can perform context-adaptive variable-length codec (CAVLC), CABAC, variable-to-variable (V2V) length codec, syntax-based context-adaptive binary arithmetic codec (SBAC), probabilistic interval segmented entropy (pipeline) codec, exponential Golomb coding, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in a bypass mode in which entropy coding of syntax elements is not performed.
[0187] According to one or more techniques of this disclosure, entropy coding unit 220 can determine which context model, either a first context model or a second context model, to use to determine the context increment of a syntax element, which indicates a prefix of the x or y coordinates of the last valid transform coefficient of the color component of a block of the image, based on the color format of the image data. Entropy coding unit 220 can encode the binary bits of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element. In some examples, entropy coding unit 220 can use either the first or second context model to derive the context increment of a syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix), which indicates a prefix of the x or y coordinates of the last valid transform coefficient of the color component of a block of the image data. Entropy coding unit 220 can apply CABAC to the binary bits of the first syntax element using the context determined based on the context increment of the syntax element. When the color component is luma or when the color component is chroma and the color format is different from 4:2:0, entropy coding unit 220 can use the first context model. When using the first context model, the entropy coding unit 220 can set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) based on the base-2 logarithm of the block transform size being less than 6, and set the context shift to be equal to (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size. The entropy coding unit 220 can also set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift to equal ((log2TrafoSize+1)>>2)<<1. The entropy coding unit 220 can determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bit index of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset. The applicable syntax element can be either the first syntax element or the second syntax element.
[0188] Video encoder 200 can output a bitstream containing entropy-encoded syntax elements required for reconstructing slices or images. For example, entropy coding unit 220 can output a bitstream.
[0189] The operations described above are relative to blocks. This description should be understood as operations on luma codec blocks and / or chroma codec blocks. As mentioned above, in some examples, the luma codec block and chroma codec block are the luma and chroma components of the CU. In some examples, the luma codec block and chroma codec block are the luma and chroma components of the PU.
[0190] In some examples, for the chroma codec block, it is not necessary to repeat the operations performed for the luma codec block. As an example, it is not necessary to repeat the operations used to identify the motion vector (MV) and reference image of the luma codec block to identify the MV and reference image of the chroma block. Instead, the MV of the luma codec block can be scaled to determine the MV of the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma codec blocks.
[0191] Video encoder 200 can represent an example of a device configured to encode video data, the device including: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to determine parameters for a residual encoding / decoding operation based on a chroma format indicator applicable to the video data block. The one or more processing units can perform the residual encoding / decoding operation based on the determined parameters to encode the residual data of the non-luminance components of the block. In some examples, the one or more processing units can be configured to: use a context model to entropy encode the prefix of the y-coordinate of the last valid transform coefficient of the luminance component of the block; and use the same context model to entropy encode the prefix of the y-coordinate of the last valid transform coefficient of the chroma component of the block, based on the block's color format being 4:2:2.
[0192] Figure 7 This is a block diagram illustrating an example video decoder 300 that can perform the techniques disclosed herein. Figure 7 This disclosure is provided for illustrative purposes and does not limit the techniques broadly exemplified and described herein. For illustrative purposes, this disclosure describes a video decoder 300 based on VVC (ITU-T H.266, under development) and HEVC (ITU-T H.265) technologies. However, the techniques of this disclosure can be implemented by video codec devices configured to conform to other video codec standards.
[0193] exist Figure 7In the example, the video decoder 300 includes a codec picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filtering unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filtering unit 312, and DPB 314 can be implemented in one or more processors or in processing circuitry. For example, units of the video decoder 300 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, FPGA, or ASIC. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0194] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-frame prediction unit 318. The prediction processing unit 304 may include additional units to perform predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.
[0195] CPB memory 320 can store video data, such as encoded video bitstreams, that will be decoded by components of video decoder 300. It can be stored from, for example, computer-readable medium 110 (…). Figure 1 The video data stored in the CPB memory 320 is obtained from the encoded video bitstream. The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory 320 may store video data other than the syntax elements of the encoded pictures, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores decoded pictures, which the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures from the encoded video bitstream. The CPB memory 320 and DPB 314 may be formed from any of a wide variety of memory devices, such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices. The CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.
[0196] Additionally or alternatively, in some examples, the video decoder 300 can be drawn from the memory 120 ( Figure 1 The encoded video data is retrieved from the memory. That is, memory 120 can store data as discussed above regarding CPB memory 320. Similarly, when some or all of the functions of video decoder 300 are implemented in software executed by the processing circuitry of video decoder 300, memory 120 can store instructions executed by video decoder 300.
[0197] It shows Figure 7 The various units shown aid in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to... Figure 6 Fixed-function circuits are circuits that provide a specific function and are pre-configured regarding the operations that can be performed. Programmable circuits are circuits that can be programmed to perform various tasks and provide flexible functionality in terms of the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.
[0198] The video decoder 300 may include an ALU, an EFU, digital circuitry, analog circuitry, and / or a programmable core formed according to programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executed on programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.
[0199] Entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to reproduce the syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filtering unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.
[0200] Generally, the video decoder 300 reconstructs the image on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").
[0201] Entropy decoding unit 302 can perform entropy decoding on syntax elements defining quantized transform coefficients of a block of quantized transform coefficients, as well as transform information such as quantization parameters (QP) and / or transform mode indications. According to one or more techniques of this disclosure, entropy decoding unit 302 can determine which context model, either a first or a second context model, to use to determine the context increment of a syntax element that indicates a prefix of the x or y coordinate of the last valid transform coefficient of the color component of the block of the image, based on the color format of the image in the video data. Entropy decoding unit 302 can decode the binary bits of the syntax element by applying CABAC using the context determined based on the context increment of the syntax element. In some examples, entropy decoding unit 302 can use either the first or the second context model to derive the context increment of a syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix) that indicates a prefix of the x or y coordinate of the last valid transform coefficient of the color component of the block of the image in the video data. Entropy decoding unit 302 can apply CABAC to the binary bits of the first syntax element using a context determined based on the context increment of the syntax element. When the color component is luminance or when the color component is chrominance and the color format is different from 4:2:0, entropy decoding unit 302 can use a first context model. When using the first context model, entropy decoding unit 302 can set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) based on the base-2 logarithm of the block transform size being less than 6, and set the context shift to equal (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size. Entropy decoding unit 302 can also set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift to equal ((log2TrafoSize+1)>>2)<<1. The entropy decoding unit 302 can determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the binary index of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset. The applicable syntax element can be either the first syntax element or the second syntax element.
[0202] The inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the quantization level, and similarly, determine the inverse quantization level to be applied by the inverse quantization unit 306. The inverse quantization unit 306 can, for example, perform a bit-by-bit left shift operation to inverse quantize the quantized transform coefficients. The inverse quantization unit 306 can thereby form a transform coefficient block including the transform coefficients.
[0203] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the transform coefficient block.
[0204] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is predicted inter-frame, the motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements can indicate the reference picture from which the reference block is to be retrieved in the DPB 314 and the motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can typically be configured with respect to the motion compensation unit 224 ( Figure 6 The method described is basically similar to the way the inter-frame prediction process is performed.
[0205] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, then intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Furthermore, intra-prediction unit 318 can typically be configured with respect to intra-prediction unit 226 ( Figure 6 The intra-prediction process is performed in a manner that is essentially similar to that described above. The intra-prediction unit 318 can retrieve data from neighboring samples of the current block from the DPB 314.
[0206] Reconstruction unit 310 can reconstruct the current block using prediction blocks and residual blocks. For example, reconstruction unit 310 can add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.
[0207] Filtering unit 312 can perform one or more filtering operations on the reconstructed block. For example, filtering unit 312 can perform a deblocking operation to reduce block artifacts along the edges of the reconstructed block. The operation of filtering unit 312 is not necessarily performed in all examples.
[0208] The video decoder 300 can store reconstructed blocks in the DPB 314. For example, in an example where the filtering unit 312 is not operated, the reconstructed unit 310 can store the reconstructed blocks in the DPB 314. In an example where the filtering unit 312 is operated, the filtering unit 312 can store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 can provide reference information (such as samples of the current image for intra-frame prediction and samples of previously decoded images for subsequent motion compensation) to the prediction processing unit 304. Furthermore, the video decoder 300 can output the decoded image (e.g., decoded video) from the DPB 314 for subsequent display on a display device (e.g., [unclear text - likely a device name]). Figure 1 It is displayed on the display device 118.
[0209] In this manner, video decoder 300 represents an example of a video decoding device, which includes: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: determine parameters for a residual encoding / decoding operation based on a chroma format indicator applicable to the video data block; and perform the residual encoding / decoding operation based on the determined parameters to decode the residual data of the non-luminance components of the block. In some examples, the one or more processing units may be configured to: entropy encode the prefix of the y-coordinate of the last valid transform coefficient of the luminance component of the block using a context model; and entropy decode the prefix of the y-coordinate of the last valid transform coefficient of the chroma component of the block using the same context model, based on the block's color format being 4:2:2.
[0210] Figure 8 This is a flowchart illustrating an example method for encoding the current block. The current block may include the current CU. Although regarding video encoder 200 ( Figure 1 and Figure 6 The description is provided, but it should be understood that other devices can be configured to perform similar actions. Figure 8 The method.
[0211] In this example, the video encoder 200 initially predicts the current block (350). For example, the video encoder 200 may form a predicted block for the current block. The video encoder 200 may then compute a residual block for the current block (352). To compute the residual block, the video encoder 200 may compute the difference between the original unencoded block and the predicted block for the current block. The video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (354). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (358). For example, the video encoder 200 may use CAVLC or CABAC to encode the transform coefficients. The video encoder 200 may also entropy encode the last_sig_coeff_x_prefix and last_sig_coeff_y_prefix syntax elements using a context determined according to the techniques of this disclosure. The video encoder 200 can then output entropy-encoded block data (360).
[0212] Figure 9 This is a flowchart illustrating an example method for decoding the current block of video data. The current block may include the current CU. Although regarding video decoder 300 ( Figure 1 and Figure 7 The description is provided, but it should be understood that other devices can be configured to perform similar actions. Figure 9 The method.
[0213] The video decoder 300 can receive entropy-encoded data of the current block, such as entropy-encoded prediction information and entropy-encoded data of the transform coefficients of the residual block corresponding to the current block (370). The video decoder 300 can entropy decode the entropy-encoded data to determine the prediction information of the current block and reproduce the transform coefficients of the residual block (372). The video decoder 300 can also entropy decode the last_sig_coeff_x_prefix and last_sig_coeff_y_prefix syntax elements using a context determined according to the techniques of this disclosure. The video decoder 300 can, for example, use an intra-frame or inter-frame prediction mode indicated by the prediction information of the current block to predict the current block (374). The video decoder 300 can then perform an inverse scan on the reproduced transform coefficients (376) to create a quantized transform coefficient block. The video decoder 300 can then inverse quantize the transform coefficients and apply the inverse transform to the transform coefficients to generate the residual block (378). Ultimately, the video decoder 300 can decode the current block by combining the predicted block and the residual block (380).
[0214] Figure 10 This is a flowchart illustrating example operation of a video encoder according to one or more technologies disclosed herein. Figure 10 In the example, the video codec (e.g., video encoder 200 or video decoder 300) can use a context model to deduce the context increment (400) of the first syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix). The video codec can use Figure 11 The operation is used to derive the context increment of the first syntax element. The first syntax element may indicate a prefix of the x or y coordinates of the last valid transform coefficient of the luminance component of the video data block. Furthermore, the video codec can use the context determined based on the context increment of the first syntax element to apply CABAC to the binary bits (402) of the first syntax element.
[0215] Video codecs can use a context model to deduce the context increment (404) of a second syntax element (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix). Figure 11 The operation is used to derive the context increment of the second syntax element. The second syntax element can be a prefix of the x or y coordinates of the last valid transform coefficient of the chroma component of the image block of the video data. The color format of the image can be different from 4:2:0. The video codec can use the context determined based on the context increment of the second syntax element to apply CABAC to the binary bits (406) of the second syntax element.
[0216] Figure 11 This is a flowchart illustrating example operation of a video encoder 200 according to one or more technologies disclosed herein. Figure 11In the example, video encoder 200 can determine, based on the color format of the image in the video data, which context model, either a first or a second, should be used to determine the context increment of a syntax element that indicates a prefix (440) of the x or y coordinate of the last valid transform coefficient of the color component (e.g., the first color component) of the image block. In some examples, video encoder 200 can use the first context model to determine the context increment of the first syntax element and can determine, based on the color format of the image in the video data, use the second context model to determine the context increment of the second syntax element. Alternatively, video encoder 200 can determine, based on the color format of the image in the video data, use the (same) first context model to determine the context increment of the second syntax element. In this example, the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block. Furthermore, video encoder 200 can encode the binary bits of the second syntax element by applying CABAC using the context determined based on the context increment of the second syntax element.
[0217] In some examples, the video encoder 200 may determine to use a second context model to derive the context increment of a second syntax element based on the following conditions not being met: i) the image's color format is a monochrome color format or a 4:4:4 color format, and / or ii) the image's color format is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luma component. Additionally, in some examples, the video encoder 200 may determine to use a second context model to determine the context increment based on the following conditions not being met: iii) the image's color format is 4:2:2, the second color component is a chroma component, and the second syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block. In some examples, the video encoder 200 may determine to use a first context model to derive the context increment of a second syntax element based on at least one of the following conditions: i) the image's color format is a monochrome color format or a 4:4:4 color format, or ii) the image's color format is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luma component. Additionally, in some examples, the video encoder 200 can determine to derive the context increment of the second syntax element using the first context model based on the following conditions: iii) the image color format is 4:2:2 color format, and the second color component is not a luminance component, and the second syntax element is a prefix of the y-coordinate of the last valid transform coefficient of the block. The same determination process can be applied to the first syntax element. In other words, the same conditions can be used for both the first and second syntax elements. In some examples, the first and second syntax elements can be associated with different color components. In some examples, the first and second syntax elements can be associated with different coordinates of the same color component.
[0218] Additionally, the video encoder 200 can encode the binary bits of a syntax element by using CABAC with a context determined based on the context increment of the syntax element (442).
[0219] Figure 12 This is a flowchart illustrating example operation of a video decoder 300 according to one or more technologies disclosed herein. Figure 12In the example, video decoder 300 can determine, based on the color format of the image in the video data, which context model, either a first or a second, to determine the context increment of a syntax element that indicates a prefix (460) of the x or y coordinate of the last valid transform coefficient of the color component (e.g., the first color component) of the image block. Video decoder 300 can also determine, based on the color format of the image in the video data, to use the second context model to determine the context increment of a second syntax element. Alternatively, video decoder 300 can determine, based on the color format of the image in the video data, to use the (same) first context model to determine the context increment of a second syntax element. In this example, the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block. Furthermore, video encoder 200 can encode the binary bits of the second syntax element by applying CABAC to the context determined based on the context increment of the second syntax element.
[0220] In some examples, the video decoder 300 may determine to use a second context model to derive the context increment of a second syntax element based on the following conditions not being met: i) the image's color format is a monochrome color format or a 4:4:4 color format, and / or ii) the image's color format is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luminance component. Additionally, in some examples, the video decoder 300 may determine to use a second context model to derive the context model of a second syntax element based on the following conditions not being met: iii) the image's color format is 4:2:2, the second color component is a chrominance component, and the second syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block. In some examples, the video decoder 300 can determine to derive the context increment of the second syntax element using the first context model based on at least one of the following conditions: i) the image's color format is a monochrome color format or a 4:4:4 color format, or ii) the image's color format is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luminance component. Additionally, in some examples, the video decoder 300 can determine to derive the context increment of the second syntax element using the first context model based on the following conditions: iii) the image's color format is a 4:2:2 color format, and the second color component is not a luminance component, and the second syntax element is a prefix of the y-coordinate of the last valid transform coefficient of the indicator block. The same determination process can be applied to the first syntax element. In other words, the same conditions can be used for both the first and second syntax elements. In some examples, the first and second syntax elements can be associated with different color components. In some examples, the first and second syntax elements can be associated with different coordinates of the same color component.
[0221] Additionally, the video decoder 300 can decode the binary bits of a syntax element by applying CABAC using a context determined based on the context increment of the syntax element (462).
[0222] Figure 13 This is a flowchart illustrating example operations of a video codec determining context according to one or more techniques according to this disclosure. Figure 13 In the example, the video codec can determine the value (500) of the luminance model variable (e.g., enableLumaModel). The video codec can use the values described below. Figure 14 The operation determines the value of the brightness model variable.
[0223] In addition, Figure 13In the example, the video codec can determine whether the luminance model variable is true (502). In response to determining that the luminance model variable is true (the "yes" branch of 502), the video codec can determine whether the base-2 logarithm of the block's transform size (e.g., Log2TrafoSize) is less than 6 (504). Based on the fact that the base-2 logarithm of the block's transform size is less than 6 (the "yes" branch of 504), the video codec can set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) and can set the context shift to equal (log2TrafoSize+1)>>2.
[0224] In response to a base-2 logarithm value for determining the transform size being greater than or equal to 6 (the "No" branch of 504), the video codec can set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift to equal to ((log2TrafoSize+1)>>2)<<1(508).
[0225] However, in response to the determination that the luminance model variable is not equal to true (the "No" branch of 502), the video codec can set the context offset to 18 and the context shift to Max(0,log2TrafoSize–2)–Max(0,log2TrafoSize–4) (510). Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the block's transform size.
[0226] In any case, after determining the context shift and context offset in actions 506, 508, or 510, the video codec can determine the context increment as (binIdx >> CTX shift) + CTX offset (512). Here, binIdx is the bit index of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset. The applicable syntax element can be about... Figure 10 The first or second syntax element described.
[0227] Figure 14 This is a flowchart illustrating an example operation for determining the values of brightness model variables according to one or more techniques disclosed herein. Figure 14In the example, the video codec (e.g., video encoder 200 or video decoder 300) may initially set the luminance model variable (e.g., enableLumaModel) to false (600). The video encoder may then determine whether the color format is equal to 0 or 3 (602). For example, the video codec may determine whether chromaArrayType is equal to 0 or 3. As mentioned above, a color format of 0 indicates monochrome, and a color format of 3 indicates a 4:4:4 color format. If the color format is equal to 0 or 3 (the "yes" branch of 602), the video codec may set the luminance model variable to true (604).
[0228] Otherwise, if the color format is not equal to 0 or 3 (the "No" branch of 602), the video codec can determine whether the color format is equal to 1 or 2 and the current color component (e.g., cIdx) is equal to 0 (606). The current color component is the color component associated with the syntax element being encoded (e.g., last_sig_coeff_x_prefix or last_sig_coeff_y_prefix). As mentioned above, a color format equal to 1 indicates a 4:2:0 color format, and a color format equal to 2 indicates a 4:2:2 color format. If the color format is equal to 1 or 2 and the color component is equal to 0 (the "Yes" branch of 606), the video codec can set the luminance model variable to true (604).
[0229] Otherwise, if the color format is not equal to 1 or 2 or the current color component is not equal to 0 (the "No" branch of 606), the video codec can determine whether the color format is equal to 2 and the current color component is not equal to 0 and the syntax element (e.g., last_sig_coeff_y_prefix) indicates the prefix of the y-coordinate of the last valid transform coefficient of the block (608). In response to determining that the color format is equal to 2 and the current color component is not equal to 0 and the syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the block (the "Yes" branch of 608), the video codec can set the luminance model variable to true (604). Otherwise, if the color format is not equal to 2 or the current color component is equal to 0 or the syntax element does not indicate the prefix of the y-coordinate of the last valid transform coefficient of the block (the "No" branch of 608), the video codec does not change the value of the luminance model variable (610). In some examples, decision box 608 may be omitted.
[0230] The following is a non-limiting list of aspects of one or more technologies based on this disclosure.
[0231] Aspect 1A: A method for encoding and decoding video data, the method comprising: determining parameters for a residual encoding and decoding operation based on a chroma format indicator applicable to a block of video data; and performing the residual encoding and decoding operation to encode and decode residual data of a non-luminance component of the block based on the determined parameters for the residual encoding and decoding operation.
[0232] Aspect 2A: A method for encoding and decoding video data, the method comprising: using a context model to entropy encode and decode the prefix of the y-coordinate of the last valid coefficient of the luminance component of a block; and using the same context model to entropy encode and decode the prefix of the y-coordinate of the last valid coefficient of the chrominance component of the block, based on the color format of the block being 4:2:2.
[0233] Aspect 3A: The method according to aspect 2A further includes the method according to aspect 1.
[0234] Aspect 4A: The method according to any one of Aspects 1A-3A, wherein encoding / decoding includes decoding.
[0235] Aspect 5A: The method according to any one of Aspects 1A-4A, wherein encoding / decoding includes encoding.
[0236] Aspect 6A: An apparatus for encoding and decoding video data, the apparatus comprising one or more components for performing a method according to any one of aspects 1A-5A.
[0237] Aspect 7A: The device according to aspect 6A, wherein the one or more components include one or more processors implemented in a circuit.
[0238] Aspect 8A: The device according to any one of aspects 6A and 7A further includes a memory configured to store video data.
[0239] Aspect 9A: The device according to any one of aspects 6A-8A further includes a display configured to display decoded video data.
[0240] Aspect 10A: The device according to any one of aspects 6A-9A, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.
[0241] Aspect 11A: The device according to any one of aspects 6A-10A, wherein the device includes a video decoder.
[0242] Aspect 12A: The device according to any one of aspects 6A-11A, wherein the device includes a video encoder.
[0243] Aspect 13A: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the method according to any one of aspects 1A-5A.
[0244] Aspect 1B: A method for encoding and decoding video data, the method comprising: using a context model to derive a context increment of a first syntax element, wherein the first syntax element indicates a prefix of the x or y coordinates of the last valid transform coefficients of the luminance component of a block of picture data; applying context-adaptive binary arithmetic codec (CABAC) to the binary bits of the first syntax element using a context determined based on the context increment of the first syntax element; using a context model to derive a context increment of a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinates of the last valid transform coefficients of the chrominance component of a block, and the color format of the picture is different from 4:2:0; and applying CABAC to the binary bits of the first syntax element using a context determined based on the context increment of the second syntax element. BAC is applied to the binary bits of the second syntax element, where using the context model includes: doing one of the following: if the base-2 logarithm of the block transform size is less than 6, setting the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) and setting the context shift to equal to (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size; or if the base-2 logarithm of the transform size is greater than or equal to 6, setting the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift to equal to ((log2TrafoSize+1)>>2)<<1; and determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the binary index of the binary bits of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset, where the applicable syntax element is the first syntax element or the second syntax element.
[0245] Aspect 2B: The method according to aspect 1B further includes setting the enable brightness model variable from false to true based on one of the following: the color format of the image is monochrome or 4:4:4, or the color format of the image is 4:2:0 or 4:2:2, and the color component of the applicable syntax element is brightness, or wherein using the context model to derive the context increment of the second syntax element includes using the context model to derive the context increment of the second syntax element based on the enable brightness model variable being true.
[0246] Aspect 3B: The method according to aspect 2B further includes enabling the luminance model variable from false to true based on the color format of the image being 4:2:2, the color component of the applicable syntax element being the chromaticity component, and the prefix of the y-coordinate of the last valid transform coefficient of the chromaticity component of the applicable syntax element indicating the block.
[0247] Aspect 4B: According to the method described in Aspect 1B, wherein the context model is a first context model, the image is a first image, the block is a first block, the applicable syntax element is a first applicable syntax element, the context is a first context, the context shift is a first context shift, and the context offset is a first context offset, and wherein the method further includes: based on the second image's color format being monochrome or 4:4:4, or the second image's color format being 4:2:0 or 4:2:2 and the second applicable syntax element's color component being luminance, or based on the enable luminance model variable being equal to false, setting the enable luminance model variable from false to true, using the second context model to derive the context increment of the second applicable syntax element, wherein the second context model is used to derive the second applicable... The context increment of the syntax element includes: setting the second context offset to 18 and setting the second context shift to Max(0,log2TrafoSize-2–2)–Max(0,log2TrafoSize2–4), where Max indicates the maximum value function and log2TrafoSize2 indicates the base-2 logarithmic value of the transform size of the second block; and setting the second context increment to (binIdx2>>ctxShift2)+ctxOffset2, where binIdx2 is the binary index of the binary bits of the second applicable syntax element, ctxShift2 indicates the second context shift, and ctxOffset2 indicates the second context offset.
[0248] Aspect 5B: According to the method of aspect 4B, wherein the method further includes enabling a luminance model variable from false to true based on the color format of the second image being 4:2:2, the color component of the second applicable syntax element being the chromaticity component of the second block of the second image, and the second applicable syntax element indicating a prefix of the y-coordinate of the last valid transform coefficient of the chromaticity component of the second block.
[0249] Aspect 6B: According to the method of aspect 1B, applying CABAC to the binary bits of the first syntax element includes CABAC decoding of the binary bits of the first syntax element, and applying CABAC to the binary bits of the second syntax element includes CABAC decoding of the binary bits of the second syntax element.
[0250] Aspect 7B: According to the method of aspect 1B, applying CABAC to the binary bits of the first syntax element includes CABAC encoding the binary bits of the first syntax element, and applying CABAC to the binary bits of the second syntax element includes CABAC encoding the binary bits of the second syntax element.
[0251] Aspect 8B: An apparatus for encoding and decoding video data, the apparatus comprising: a memory configured to store video data; and one or more processors coupled to the memory, the one or more processors being implemented in circuitry and configured to: derive a context increment of a first syntax element using a context model, wherein the first syntax element indicates a prefix of the x or y coordinates of the last valid transform coefficient of the luminance component of a block of video data; apply context-adaptive binary arithmetic codec (CABAC) to the binary bits of the first syntax element using a context determined based on the context increment of the first syntax element; derive a context increment of a second syntax element using a context model, wherein the second syntax element indicates a prefix of the x or y coordinates of the last valid transform coefficient of the chrominance component of a block, and the color format of the image is different from 4:2:0; and use... The context determined based on the context increment of the second syntax element applies CABAC to the binary bits of the second syntax element, wherein the one or more processors are configured as part of the context model: if the base-2 logarithm of the block transform size is less than 6, the context offset is set to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2), and the context shift is set to equal to (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size; if the base-2 logarithm of the transform size is greater than or equal to 6, the context offset is set to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift to equal to ((log2TrafoSize+1)>>2)<<1; and determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the binary index of the binary bits of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset, where the applicable syntax element is the first syntax element or the second syntax element.
[0252] Aspect 9B: The apparatus according to aspect 8B, wherein the one or more processors are further configured to set an enable luminance model variable from false to true based on one of the following: the color format of the image is monochrome or 4:4:4, or the color format of the image is 4:2:0 or 4:2:2 and the color component of the applicable syntax element is luminance; and wherein the one or more processors are configured to, as part of using the context model to derive the context increment of the second syntax element based on the enable luminance model variable being true, use the context model to derive the context increment of the second syntax element.
[0253] Aspect 10B: The device according to aspect 9B, wherein the one or more processors are further configured to enable the luminance model variable from false to true based on the color format of the image being 4:2:2, the color component of the applicable syntax element is a chromaticity component, and the applicable syntax element is a prefix of the y-coordinate of the last valid transform coefficient of the chromaticity component of the block.
[0254] Aspect 11B: The device according to Aspect 9B, wherein the one or more processors are configured to derive the context increment of the second applicable syntax element using a second context model based on enabling a luminance model variable equal to false, wherein the one or more processors are configured as part of deriving the context increment of the second applicable syntax element using the second context model: setting the second context offset to 18, and setting the second context shift to Max(0,log2TrafoSize-2–2)–Max(0,log2TrafoSize2–4), where Max indicates the maximum value function, and log2TrafoSize2 indicates the base-2 logarithmic value of the transform size of the second block; and determining the second context increment as (binIdx2>>ctxShift2)+ctxOffset2, where binIdx2 is the bit index of the binary bits of the second applicable syntax element, ctxShift2 indicates the second context shift, and ctxOffset2 indicates the second context offset.
[0255] Aspect 12B: The apparatus according to aspect 8B, wherein the one or more processors are configured to perform CABAC decoding on the binary bits of a first syntax element as part of applying CABAC to the binary bits of a first syntax element, and wherein the one or more processors are configured to perform CABAC decoding on the binary bits of a second syntax element as part of applying CABAC to the binary bits of a second syntax element.
[0256] Aspect 13B: The apparatus according to aspect 8B, wherein the one or more processors are configured to CABAC encode the binary bits of the first syntax element as part of applying CABAC to the binary bits of the first syntax element, and wherein the one or more processors are configured to CABAC encode the binary bits of the second syntax element as part of applying CABAC to the binary bits of the second syntax element.
[0257] Aspect 14B: The apparatus according to aspect 8B further includes a display configured to display decoded video data.
[0258] Aspect 15B: The device according to aspect 8B, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.
[0259] Aspect 16B: An apparatus for encoding and decoding video data, the apparatus comprising: means for deriving a context increment of a first syntax element using a context model, wherein the first syntax element indicates a prefix of the x or y coordinates of the last valid transform coefficient of the luminance component of a block of video data; means for applying context-adaptive binary arithmetic codec (CABAC) to the bits of the first syntax element using a context determined based on the context increment of the first syntax element; means for deriving a context increment of a second syntax element using a context model, wherein the second syntax element indicates a prefix of the x or y coordinates of the last valid transform coefficient of the chrominance component of a block, and the color format of the image is different from 4:2:0; and means for applying context-adaptive binary arithmetic codec (CABAC) to the bits of the first syntax element using a context determined based on the context increment of the second syntax element. The context is defined by the component that applies CABAC to the binary bits of the second syntax element. The component for using the context model includes: a component for setting the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) based on a base-2 logarithm of the block transform size less than 6, and a component for setting the context shift to equal to (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size; and a component for setting the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7) and a component that sets the context shift to be equal to ((log2TrafoSize+1)>>2)<<1; and a device for determining the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bit index of the binary bits of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset, where the applicable syntax element is either the first syntax element or the second syntax element.
[0260] Aspect 17B: The apparatus according to aspect 16B, wherein the apparatus further includes means for enabling a luminance model variable from false to true based on the color format of the image being monochrome or 4:4:4 or the color format of the image being 4:2:0 or 4:2:2 and the color component of the applicable syntax element being luminance, or wherein the means for using a context model to derive a context increment of a second syntax element includes means for using a context model to derive a context increment of a second syntax element based on enabling a luminance model variable to be true.
[0261] Aspect 18B: The device according to aspect 17B further includes means for deriving the context increment of the second applicable syntax element using a second context model based on the enable brightness model variable being equal to a false second context model, wherein the means for deriving the context increment of the second applicable syntax element using the second context model includes: means for setting the second context offset to 18 and setting the second context shift to Max(0,log2TrafoSize2–2)–Max(0,log2TrafoSize2–4), wherein Max indicates the maximum value function and log2TrafoSize2 indicates the base-2 logarithmic value of the transformation size of the second block; and means for determining the second context increment as (binIdx2>>ctxShift2)+ctxOffset2, wherein binIdx2 is the bit index of the binary bits of the second applicable syntax element, ctxShift2 indicates the second context shift, and ctxOffset2 indicates the second context offset.
[0262] Aspect 19B: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: derive a context increment for a first syntax element using a context model, wherein the first syntax element indicates a prefix of the x or y coordinates of the last valid transform coefficient of the luminance component of a block of video data; apply context-adaptive binary arithmetic codec (CABAC) to the bits of the first syntax element using a context determined based on the context increment of the first syntax element; derive a context increment for a second syntax element using a context model, wherein the second syntax element indicates a prefix of the x or y coordinates of the last valid transform coefficient of the chrominance component of a block, and the color format of the picture is different from 4:2:0; and apply CABAC using a context determined based on the context increment of the second syntax element. In the binary bits of the second syntax element, the instructions that cause one or more processors to use the context model include instructions that, when executed, cause one or more processors to perform the following operations: Based on a base-2 logarithm of the block transform size being less than 6, set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) and set the context shift to equal to (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size; Based on a base-2 logarithm of the transform size being greater than or equal to 6, set the context offset to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift to equal to ((log2TrafoSize+1)>>2)<<1; and determine the context increment as (binIdx>>ctxShift)+ctxOffset, where binIdx is the binary index of the binary bits of the applicable syntax element, ctxShift indicates the context shift, and ctxOffset indicates the context offset, where the applicable syntax element is the first syntax element or the second syntax element.
[0263] Aspect 1C: A method for decoding video data, comprising: determining, based on the color format of an image of the video data, which of a first context model and a second context model to use to determine a context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinates of the last valid transform coefficients of the color components of a block of the image; and decoding the binary bits of the syntax element using context-adaptive binary arithmetic codec (CABAC) with the context determined based on the context increment of the syntax element.
[0264] Aspect 2C: The method according to aspect 1C, wherein: the color component is a first color component, and the syntax element is a first syntax element, and the method further includes: determining a context increment of the first syntax element using a first context model; determining a context increment of a second syntax element based on the color format of the image in the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block; and decoding the binary bits of the second syntax element by using CABAC with the context determined based on the context increment of the second syntax element.
[0265] Aspect 3C: According to the method of aspect 1C, wherein: the color component is a first color component, and the syntax element is a first syntax element, and the method further includes: determining a context increment of the first syntax element using a first context model; determining a context increment of the second syntax element using a second context model based on the color format of the image in the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block; and decoding the binary bits of the second syntax element by using CABAC with the context determined based on the context increment of the second syntax element.
[0266] Aspect 4C: The method described according to aspect 2C or 3C, wherein the first color component is a luminance component and the second color component is a chromaticity component.
[0267] Aspect 5C: The method according to any one of Aspects 3C and 4C, wherein determining the use of a second context model to determine the context increment of the second syntax element comprises determining the use of the second context model to derive the context increment of the second syntax element based on the following conditions not being met: i) the color format of the image is a monochrome color format or a 4:4:4 color format, and ii) the color format of the image is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luminance component.
[0268] Aspect 6C: According to the method described in aspect 5C, wherein the context increment of the second syntax element is determined by using a second context model based on the following condition not being met: iii) the color format of the image is 4:2:2, the second color component is a chroma component, and the second syntax element is a prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block.
[0269] Aspect 7C: The method according to any one of Aspects 3C to 6C, wherein deriving the context increment of the first syntax element using a first context model comprises: performing one of the following operations: setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) based on the base-2 logarithm of the block transform size being less than 6, and setting the context shift ctxShift of the first syntax element to equal to (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size; or setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift ctxShift of the first syntax element to equal ((log2TrafoSize+1)>>2)<<1; and determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the binary bit index of the binary bits of the first syntax element.
[0270] Aspect 8C: The method according to any one of Aspects 3C to 7C, wherein determining the context increment of the second syntax element using a second context model comprises: setting the context offset ctxOffset of the second syntax element to 18, and setting the context shift ctxShift of the second syntax element to Max(0,log2TrafoSize-2)-Max(0,log2TrafoSize-4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the transform size of the block; and determining the context increment of the second syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element.
[0271] Aspect 9C: The method according to any one of Aspects 1C to 8C, wherein determining which context model, one of the first context model and the second context model, is used to determine the context increment of the syntax element comprises determining that the first context model is used to derive the context increment of the syntax element based on at least one of the following conditions: i) the color format of the image is a monochrome color format or a 4:4:4 color format; ii) the color format of the image is a 4:2:0 color format or a 4:2:2 color format, and the color component is a luminance component; or iii) the color format of the image is a 4:2:2 color format, and the color component is not a luminance component, and the syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the color component of the block.
[0272] Aspect 10C: A method for encoding video data, comprising: determining, based on the color format of an image of the video data, which of a first context model and a second context model to use to determine a context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of a block of color components of the image; and encoding the binary bits of the syntax element by applying context-adaptive binary arithmetic codec (CABAC) using the context determined based on the context increment of the syntax element.
[0273] Aspect 11C: The method according to aspect 10C, wherein: the color component is a first color component, and the syntax element is a first syntax element, and the method further includes: determining a context increment of the first syntax element using a first context model; determining a context increment of a second syntax element using the first context model based on the color format of the image in the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block; and encoding the binary bits of the second syntax element by using CABAC applied to the context determined based on the context increment of the second syntax element.
[0274] Aspect 12C: The method according to aspect 10C, wherein: the color component is a first color component, and the syntax element is a first syntax element, and the method further includes: determining a context increment of the first syntax element using a first context model; determining a context increment of the second syntax element using a second context model based on the color format of the image in the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block; and encoding the binary bits of the second syntax element by using CABAC applied to the context determined based on the context increment of the second syntax element.
[0275] Aspect 13C: The method according to aspect 11C or 12C, wherein the first color component is a luminance component and the second color component is a chromaticity component.
[0276] Aspect 14C: The method according to any one of Aspects 12C and 13C, wherein determining the use of a second context model to determine the context increment of the second syntax element comprises determining the use of the second context model to derive the context increment of the second syntax element based on the following conditions not being met: i) the color format of the image is a monochrome color format or a 4:4:4 color format, and ii) the color format of the image is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luminance component.
[0277] Aspect 15C: The method according to any one of Aspects 12C to 14C, wherein the context increment of the second syntax element is determined by using a second context model based on the condition that the following conditions are not met: iii) the color format of the picture is 4:2:2, the second color component is a chromaticity component, and the second syntax element is a prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block.
[0278] Aspect 16C: The method according to any one of Aspects 12C to 15C, wherein deriving the context increment of the first syntax element using a first context model comprises: performing one of the following operations: setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) based on the base-2 logarithm of the block transform size being less than 6, and setting the context shift ctxShift of the first syntax element to equal (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size; or setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift ctxShift of the first syntax element to equal ((log2TrafoSize+1)>>2)<<1; and determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the binary bit index of the binary bits of the first syntax element.
[0279] Aspect 17C: The method according to any one of Aspects 12C to 16C, wherein deriving the context increment of the second syntax element using a second context model comprises: setting the context offset ctxOffset of the second syntax element to 18, and setting the context shift ctxShift of the second syntax element to Max(0,log2TrafoSize-2)-Max(0,log2TrafoSize-4), where Max indicates a maximum value function, and log2TrafoSize indicates a base-2 logarithmic value of the transform size of the block; and determining the context increment of the second syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element.
[0280] Aspect 18C: The method according to any one of Aspects 10C to 17C, wherein determining which context model, one of a first context model and a second context model, is used to determine the context increment of a syntax element comprises determining that the first context model is used to derive the context increment of the syntax element based on at least one of the following conditions: i) the color format of the image is a monochrome color format or a 4:4:4 color format; ii) the color format of the image is a 4:2:0 color format or a 4:2:2 color format, and the color component is a luminance component; or iii) the color format of the image is a 4:2:2 color format, and the color component is not a luminance component, and the syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the color component of the block.
[0281] Aspect 19C: An apparatus for decoding video data includes: a memory configured to store video data; and one or more processors coupled to the memory, the one or more processors being implemented in a circuit and configured to: determine, based on the color format of an image of the video data, which of a first context model and a second context model to determine a context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of a block of color components of the image; and decode the binary bits of the syntax element using context-adaptive binary arithmetic codec (CABAC) determined by the context increment of the syntax element.
[0282] Aspect 20C: The apparatus according to aspect 19C, wherein: the color component is a first color component, and the syntax element is a first syntax element, and one or more processors are further configured to: determine a context increment of the first syntax element using a first context model; determine a context increment of a second syntax element using the first context model based on the color format of the image of the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block; and decode the binary bits of the second syntax element by using CABAC with the context determined based on the context increment of the second syntax element.
[0283] Aspect 21C: The apparatus according to any one of Aspects 19C and 20C, wherein: the color component is a first color component, and the syntax element is a first syntax element, and one or more processors are further configured to: determine a context increment of the first syntax element using a first context model; determine a context increment of the second syntax element using a second context model based on the color format of the image of the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block; and decode the binary bits of the second syntax element by using CABAC with the context determined based on the context increment of the second syntax element.
[0284] Aspect 22C: The apparatus according to aspect 20C or 21C, wherein the first color component is a luminance component and the second color component is a chromaticity component.
[0285] Aspect 23C: The device according to any one of aspects 21C and 22C, wherein the one or more processors are configured to determine the context increment of the second syntax element using the second context model includes determining the context increment of the second syntax element using the second context model based on the following conditions not being met: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, and ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luminance component.
[0286] Aspect 24C: The device according to aspect 23C, wherein the one or more processors are configured to determine the context increment of the second syntax element using a second context model based on the following condition not being met: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element is a prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block.
[0287] Aspect 25C: The apparatus according to any one of aspects 19C to 24C, wherein the one or more processors are configured to, as part of deriving a context increment of a first syntax element using a first context model, perform one of the following operations: setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) based on a base-2 logarithm of the block transform size being less than 6, and setting the context shift ctxShift of the first syntax element to equal (log2TrafoSize+1)>>2, where log2TrafoSize indicates a base-2 logarithm of the transform size; or setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift ctxShift of the first syntax element to equal ((log2TrafoSize+1)>>2)<<1; and determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the binary bit index of the binary bits of the first syntax element.
[0288] Aspect 26C: The device according to any one of Aspects 21C to 25C, wherein one or more processors are configured as part of determining the context increment of a second syntax element using a second context model: setting the context offset ctxOffset of the second syntax element to 18, and setting the context shift ctxShift of the second syntax element to Max(0,log2TrafoSize-2)-Max(0,log2TrafoSize-4), where Max indicates a maximum value function and log2TrafoSize indicates a base-2 logarithmic value of the transform size of the block; and determining the context increment of the second syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element.
[0289] Aspect 27C: The apparatus according to any one of aspects 21C to 26C, wherein the one or more processors are configured, as part of determining which of the first and second context models to use to determine the context increment of a syntax element, to determine to use the first context model to derive the context increment of the syntax element based on at least one of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the color component is a luminance component; or iii) the color format of the picture is a 4:2:2 color format, and the color component is not a luminance component, and the syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the color component of the block.
[0290] Aspect 28C: The device according to any one of aspects 19C to 27C further includes a display configured to display decoded video data.
[0291] Aspect 29C: The device according to any one of aspects 19C to 28C, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.
[0292] Aspect 30C: An apparatus for encoding video data includes: a memory configured to store video data; and one or more processors coupled to the memory, the one or more processors being implemented in a circuit and configured to: determine, based on the color format of an image of the video data, which of a first context model and a second context model is used to determine a context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of a block of color components of the image; and encode the binary bits of the syntax element using context-adaptive binary arithmetic codec (CABAC) with the context determined based on the context increment of the syntax element.
[0293] Aspect 31C: The apparatus according to aspect 30C, wherein: the color component is a first color component, and the syntax element is a first syntax element, and one or more processors are further configured to: determine a context increment of the first syntax element using a first context model; determine a context increment of a second syntax element using the first context model based on the color format of the image of the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block; and decode the binary bits of the second syntax element by using CABAC with the context determined based on the context increment of the second syntax element.
[0294] Aspect 32C: The apparatus according to aspect 30C, wherein: the color component is a first color component, and the syntax element is a first syntax element, and one or more processors are further configured to: determine a context increment of the first syntax element using a first context model; determine a context increment of the second syntax element using a second context model based on the color format of the image of the video data, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block; and encode the binary bits of the second syntax element by using CABAC with the context determined based on the context increment of the second syntax element.
[0295] Aspect 33C: The apparatus according to aspect 31C or 32C, wherein the first color component is a luminance component and the second color component is a chromaticity component.
[0296] Aspect 34C: The device according to any one of aspects 32C to 33C, wherein the one or more processors are configured to determine the context increment of the second syntax element using the second context model includes determining the context increment of the second syntax element using the second context model based on the following conditions not being met: i) the color format of the picture is a monochrome color format or a 4:4:4 color format, and ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the second color component is a luminance component.
[0297] Aspect 35C: The device according to aspect 34C, wherein the one or more processors are configured to determine the context increment of the second syntax element using a second context model based on the following condition not being met: iii) the color format of the picture is 4:2:2, the second color component is a chroma component, and the second syntax element is a prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block.
[0298] Aspect 36C: The device according to any one of aspects 32C to 35C, wherein the one or more processors are configured as part of deriving a context increment of a first syntax element using a first context model: performing one of the following operations: setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2) based on the base-2 logarithm of the block's transform size being less than 6, and setting the context shift ctxShift of the first syntax element to equal (log2TrafoSize+1)>>2, where log2TrafoSize indicates the base-2 logarithm of the transform size; or setting the context offset ctxOffset of the first syntax element to 3*(log2TrafoSize-2)+((log2TrafoSize-1)>>2)+((1<<log2TrafoSize)> >6)<<1)+((1<<log2TrafoSize)> >7), and set the context shift ctxShift of the first syntax element to equal ((log2TrafoSize+1)>>2)<<1; and determine the context increment of the first syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the binary bit index of the binary bits of the first syntax element.
[0299] Aspect 37C: The device according to any one of aspects 32C to 36C, wherein the one or more processors are configured as part of determining the context increment of a second syntax element using a second context model: setting the context offset ctxOffset of the second syntax element to 18, and setting the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize-2)-Max(0, log2TrafoSize-4), where Max indicates a maximum value function, and log2TrafoSize indicates a base-2 logarithmic value of the transform size of the block; and determining the context increment of the second syntax element as (binIdx>>ctxShift)+ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element.
[0300] Aspect 38C: The apparatus according to any one of aspects 30C to 37C, wherein the one or more processors are configured, as part of determining which of the first and second context models to use to determine the context increment of a syntax element, to determine to use the first context model to derive the context increment of the syntax element based on at least one of the following conditions: i) the color format of the picture is a monochrome color format or a 4:4:4 color format; ii) the color format of the picture is a 4:2:0 color format or a 4:2:2 color format, and the color component is a luminance component; or iii) the color format of the picture is a 4:2:2 color format, and the color component is not a luminance component, and the syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the color component of the block.
[0301] Aspect 39C: The device according to any one of aspects 30C to 38C, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.
[0302] Aspect 40C: An apparatus for decoding video data, comprising: means for determining which context model, one of a first context model and a second context model, is used to determine a context increment of a syntax element based on the color format of an image of the video data, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the color components of a block of the image; and means for decoding the binary bits of the syntax element using context-adaptive binary arithmetic codec (CABAC) with the context determined based on the context increment of the syntax element.
[0303] Aspect 41C: An apparatus for encoding video data, comprising: means for determining which context model, one of a first context model and a second context model, is used to determine a context increment of a syntax element based on the color format of an image of the video data, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of a block of color components of the image; and means for encoding the binary bits of the syntax element by applying context adaptive binary arithmetic codec (CABAC) using the context determined based on the context increment of the syntax element.
[0304] Aspect 42C: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: determine, based on the color format of a picture of video data, which of a first context model and a second context model to determine the context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the color components of a block of the picture; and decode the binary bits of the syntax element using context-adaptive binary arithmetic codec (CABAC) with the context determined based on the context increment of the syntax element.
[0305] Aspect 43C: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: determine, based on the color format of a picture of video data, which of a first context model and a second context model to determine the context increment of a syntax element, the syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the color component of a block of the picture; and to encode the binary bits of the syntax element using context-adaptive binary arithmetic codec (CABAC) with the context determined based on the context increment of the syntax element.
[0306] It should be recognized that, depending on the example, certain actions or events of any technique described herein may be performed in a different sequence, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technique). Furthermore, in some examples, actions or events may be performed concurrently (e.g., through multithreading, interrupt handling, or multiprocessor) rather than sequentially.
[0307] In one or more examples, the described functionality may be implemented using hardware, software, firmware, or any combination thereof. If implemented in software, these functions may be stored on or transmitted thereon as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium (such as a data storage medium) or a communication medium (including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol). In this way, a computer-readable medium may generally correspond to (1) a tangible, non-transitory computer-readable storage medium, or (2) a communication medium (such as a signal or carrier wave). A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0308] By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store required program code in the form of instructions or data structures and that can be accessed by a computer. Furthermore, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. Disks and optical discs as used herein include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0309] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Accordingly, the terms "processor" and "processing circuit" as used herein can refer to any of the foregoing structures or any other structures suitable for implementing the techniques described herein. Furthermore, in some aspects, the functions described herein can be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Moreover, these techniques can be implemented entirely within one or more circuit or logic elements.
[0310] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be combined in a codec hardware unit with suitable software and / or firmware, or provided by a collection of interoperable hardware units (including one or more processors as described above).
[0311] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method for decoding video data, the method comprising: Based on the color format of the images in the video data, determine which context model to use, the first or the second context model, and then determine which one to use based on the color format: (a) The first context model determines a first context increment of a first syntax element, the first syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the first color component of the block of the image; (b) The second context model determines a second context increment for a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block, wherein the second context increment of the second syntax element includes: (i) Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize – 2) – Max(0, log2TrafoSize – 4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the width or height of the transform block; and (ii) The second context increment of the second syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element; The binary bits of the first syntax element are decoded by using context-adaptive binary arithmetic codec (CABAC) determined based on the first context increment of the first syntax element. as well as The binary bits of the second syntax element are decoded by applying the CABAC using the context determined based on the second context increment of the second syntax element.
2. The method according to claim 1, wherein, The first color component is a luminance component, and the second color component is a chromaticity component.
3. The method according to claim 1, wherein, Determining to use the second context model to determine the second context increment of the second syntax element includes determining to use the second context model to determine the context increment of the second syntax element based on the following condition not being met: i) The color format of the image is either monochrome or 4:4:4 color format, and ii) The color format of the image is 4:2:0 or 4:2:2, and the second color component is a luminance component.
4. The method according to claim 3, wherein, The second context increment of the second syntax element is determined based on the following condition: iii) the color format of the image is the 4:2:2 color format, the second color component is a chroma component, and the second syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block.
5. The method according to claim 1, wherein, Determining the first context increment of the first syntax element using the first context model includes: Perform one of the following operations: Based on the base-2 logarithm of the width or height of the transform block being less than 6, set the context offset ctxOffset of the first syntax element to 3 * (log2TrafoSize – 2) + ((log2TrafoSize – 1) >> 2), and set the context shift ctxShift of the first syntax element to equal (log2TrafoSize + 1) >> 2, where log2TrafoSize indicates the base-2 logarithm of the width or height of the transform block; or based on the base-2 logarithm of the width or height of the transform block being greater than or equal to 6, set the context offset ctxOffset of the first syntax element to 3 * (log2TrafoSize – 2) + ((log2TrafoSize – 1) >> 2) + (((1 << log2TrafoSize) >> 6) << 1) + ((1 << log2TrafoSize) >> 7), and set the context shift ctxShift of the first syntax element to equal (( log2TrafoSize + 1) >> 2) << 1; and The first context increment of the first syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the first syntax element.
6. A method for encoding video data, the method comprising: Based on the color format of the images in the video data, determine which context model to use, the first or the second context model, and then determine which one to use based on the color format: (a) The first context model determines a first context increment of a first syntax element, the first syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the first color component of the block of the image; (b) The second context model determines a second context increment for a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block, wherein the second context increment of the second syntax element includes: (i) Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize – 2) – Max(0, log2TrafoSize – 4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the width or height of the transform block; and (ii) The second context increment of the second syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element; The binary bits of the first syntax element are encoded by using context-adaptive binary arithmetic codec (CABAC) determined based on the first context increment of the first syntax element. as well as The binary bits of the second syntax element are encoded by applying the CABAC using the context determined based on the second context increment of the second syntax element.
7. The method according to claim 6, wherein, The first color component is a luminance component, and the second color component is a chromaticity component.
8. The method according to claim 6, wherein, Determining to use the second context model to determine the second context increment of the second syntax element includes determining to use the second context model to determine the second context increment of the second syntax element based on the following condition not being met: i) The color format of the image is either monochrome or 4:4:4 color format, and ii) The color format of the image is 4:2:0 or 4:2:2, and the second color component is a luminance component.
9. The method according to claim 6, wherein, The second context increment of the second syntax element is determined based on the following condition: iii) the color format of the image is the 4:2:2 color format, the second color component is a chroma component, and the second syntax element indicates the prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block.
10. The method according to claim 6, wherein, Determining the first context increment of the first syntax element using the first context model includes: Perform one of the following operations: Based on the base-2 logarithm of the width or height of the transform block being less than 6, set the context offset ctxOffset of the first syntax element to 3 * (log2TrafoSize – 2) + ((log2TrafoSize – 1) >> 2), and set the context shift ctxShift of the first syntax element to equal (log2TrafoSize + 1) >> 2, where log2TrafoSize indicates the base-2 logarithm of the width or height of the transform block; or based on the base-2 logarithm of the width or height of the transform block being greater than or equal to 6, set the context offset ctxOffset of the first syntax element to 3 * (log2TrafoSize – 2) + ((log2TrafoSize – 1) >> 2) + (((1 << log2TrafoSize) >> 6) << 1) + ((1 << log2TrafoSize) >> 7), and set the context shift ctxShift of the first syntax element to equal (( log2TrafoSize + 1) >> 2) << 1; and The first context increment of the first syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the first syntax element.
11. An apparatus for decoding video data, the apparatus comprising: The memory is configured to store the video data; as well as One or more processors coupled to the memory, the one or more processors being implemented in a circuit and configured to: Based on the color format of the images in the video data, determine which context model to use, the first or the second context model, and then determine which one to use based on the color format: (a) The first context model determines a first context increment of a first syntax element, the first syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the first color component of the block of the image; (b) The second context model determines a second context increment for a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block, wherein the second context increment of the second syntax element includes: (i) Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize – 2) – Max(0, log2TrafoSize – 4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the width or height of the transform block; and (ii) The second context increment of the second syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element; The binary bits of the first syntax element are decoded by using context-adaptive binary arithmetic codec (CABAC) determined based on the first context increment of the first syntax element. as well as The binary bits of the second syntax element are decoded by applying the CABAC using the context determined based on the second context increment of the second syntax element.
12. The device according to claim 11, wherein, The first color component is a luminance component, and the second color component is a chromaticity component.
13. The device according to claim 11, wherein, The one or more processors are configured to determine the second context increment of the second syntax element using the second context model based on the following condition not being met: i) The color format of the image is either monochrome or 4:4:4 color format, and ii) The color format of the image is 4:2:0 or 4:2:2, and the second color component is a luminance component.
14. The device according to claim 13, wherein, The one or more processors are configured to further determine the context increment of the second syntax element based on the following condition: iii) the color format of the image is the 4:2:2 color format, the second color component is a chroma component, and the second syntax element indicates a prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block.
15. The device according to claim 11, wherein, The one or more processors are configured as part of the first context increment used to determine the first syntax element: Perform one of the following operations: Based on the base-2 logarithm of the width or height of the transform block being less than 6, set the context offset ctxOffset of the first syntax element to 3 * (log2TrafoSize – 2) + ((log2TrafoSize – 1) >> 2), and set the context shift ctxShift of the first syntax element to equal (log2TrafoSize + 1) >> 2, where log2TrafoSize indicates the base-2 logarithm of the width or height of the transform block; or based on the base-2 logarithm of the width or height of the transform block being greater than or equal to 6, set the context offset ctxOffset of the first syntax element to 3 * (log2TrafoSize – 2) + ((log2TrafoSize – 1) >> 2) + (((1 << log2TrafoSize) >> 6) << 1) + ((1 << log2TrafoSize) >> 7), and set the context shift ctxShift of the first syntax element to equal (( log2TrafoSize + 1 ) >> 2) << 1; as well as The first context increment of the first syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bit of the first syntax element.
16. The device of claim 11, further comprising a display configured to display decoded video data.
17. The device according to claim 11, wherein, The device includes one or more of a mobile device and a broadcast receiver device.
18. An apparatus for encoding video data, the apparatus comprising: The memory is configured to store the video data; as well as One or more processors coupled to the memory, the one or more processors being implemented in a circuit and configured to: Based on the color format of the images in the video data, determine which context model to use, the first or the second context model, and then determine which one to use based on the color format: (a) The first context model determines a first context increment of a first syntax element, the first syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the first color component of the block of the image; (b) The second context model determines a second context increment for a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block, wherein the second context increment of the second syntax element includes: (i) Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize – 2) – Max(0, log2TrafoSize – 4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the width or height of the transform block; and (ii) The second context increment of the second syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element; The binary bits of the first syntax element are encoded by using context-adaptive binary arithmetic codec (CABAC) determined based on the first context increment of the first syntax element. as well as The binary bits of the second syntax element are encoded by applying the CABAC using the context determined based on the second context increment of the second syntax element.
19. The device according to claim 18, wherein, The first color component is a luminance component, and the second color component is a chromaticity component.
20. The device according to claim 18, wherein, The one or more processors are configured to determine the second context increment of the second syntax element using the second context model based on the following condition not being met: i) The color format of the image is either monochrome or 4:4:4 color format, and ii) The color format of the image is 4:2:0 or 4:2:2, and the second color component is a luminance component.
21. The device according to claim 20, wherein, The one or more processors are configured to further determine the second context increment of the second syntax element based on the following condition: iii) the color format of the image is the 4:2:2 color format, the second color component is a chroma component, and the second syntax element indicates a prefix of the y-coordinate of the last valid transform coefficient of the second color component of the block.
22. The device according to claim 18, wherein, The one or more processors are configured as part of the first context increment used to determine the first syntax element: Perform one of the following operations: Based on the base-2 logarithm of the width or height of the transform block being less than 6, set the context offset ctxOffset of the first syntax element to 3 * (log2TrafoSize – 2) + ((log2TrafoSize – 1) >> 2), and set the context shift ctxShift of the first syntax element to equal (log2TrafoSize + 1) >> 2, where log2TrafoSize indicates the base-2 logarithm of the width or height of the transform block; or based on the base-2 logarithm of the width or height of the transform block being greater than or equal to 6, set the context offset ctxOffset of the first syntax element to 3 * (log2TrafoSize – 2) + ((log2TrafoSize – 1) >> 2) + (((1 << log2TrafoSize) >> 6) << 1) + ((1 << log2TrafoSize) >> 7), and set the context shift ctxShift of the first syntax element to equal (( log2TrafoSize + 1 ) >> 2) << 1; as well as The first context increment of the first syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bit of the first syntax element.
23. The device according to claim 18, wherein, The device includes one or more of a mobile device and a broadcast receiver device.
24. An apparatus for decoding video data, the apparatus comprising: The component used to determine which context model, the first or the second, to use for the color format of the image based on the video data, and to determine the following based on the color format: using (a) The first context model determines a first context increment of a first syntax element, the first syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the first color component of the block of the image; (b) The second context model determines a second context increment for a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block, wherein the second context increment of the second syntax element includes: (i) Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize – 2) – Max(0, log2TrafoSize – 4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the width or height of the transform block; and (ii) The second context increment of the second syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element; A component for decoding the binary bits of the first syntax element by applying context-adaptive binary arithmetic codec (CABAC) based on the context increment determined by the first context of the first syntax element; and A component for decoding the binary bits of the second syntax element by applying the CABAC using a context determined based on the second context increment of the second syntax element.
25. An apparatus for encoding video data, the apparatus comprising: The component used to determine which context model, the first or the second, to use for the color format of the image based on the video data, and to determine the following based on the color format: using (a) The first context model determines a first context increment of a first syntax element, the first syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the first color component of the block of the image; (b) The second context model determines a second context increment for a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block, wherein the second context increment of the second syntax element includes: (i) Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize – 2) – Max(0, log2TrafoSize – 4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the width or height of the transform block; and (ii) The second context increment of the second syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element; A component for encoding the binary bits of the first syntax element by applying context-adaptive binary arithmetic codec (CABAC) based on the first context increment of the first syntax element; and A component for encoding the binary bits of the second syntax element by applying the CABAC using a context determined based on the second context increment of the second syntax element.
26. A computer-readable storage medium having instructions stored thereon, the instructions causing one or more processors to: Based on the color format of the image from the video data, determine which context model to use, either the first or second context model, and then determine the appropriate one based on the color format: (a) The first context model determines a first context increment of a first syntax element, the first syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the first color component of the block of the image; (b) The second context model determines a second context increment for a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block, wherein the second context increment of the second syntax element includes: (i) Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize – 2) – Max(0, log2TrafoSize – 4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the width or height of the transform block; and (ii) The second context increment of the second syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element; The binary bits of the first syntax element are decoded by using context-adaptive binary arithmetic codec (CABAC) determined based on the first context increment of the first syntax element. as well as The binary bits of the second syntax element are decoded by applying the CABAC using the context determined based on the second context increment of the second syntax element.
27. A computer-readable storage medium having instructions stored thereon, the instructions causing one or more processors, when executed, to: Based on the color format of the image from the video data, determine which context model to use, either the first or second context model, and then determine the appropriate one based on the color format: (a) The first context model determines a first context increment of a first syntax element, the first syntax element indicating a prefix of the x or y coordinate of the last valid transform coefficient of the first color component of the block of the image; (b) The second context model determines a second context increment for a second syntax element, wherein the second syntax element indicates a prefix of the x or y coordinate of the last valid transform coefficient of the second color component of the image block, wherein the second context increment of the second syntax element includes: (i) Set the context offset ctxOffset of the second syntax element to 18, and set the context shift ctxShift of the second syntax element to Max(0, log2TrafoSize – 2) – Max(0, log2TrafoSize – 4), where Max indicates the maximum value function, and log2TrafoSize indicates the base-2 logarithmic value of the width or height of the transform block; and (ii) The second context increment of the second syntax element is determined as (binIdx >> ctxShift) + ctxOffset, where binIdx is the bit index of the binary bits of the second syntax element; The binary bits of the first syntax element are encoded by using context-adaptive binary arithmetic codec (CABAC) determined based on the first context increment of the first syntax element. as well as The binary bits of the second syntax element are encoded by applying the CABAC using the context determined based on the second context increment of the second syntax element.