Chroma component coding
By employing single-tree block segmentation and intra-frame chroma copying in video encoding, the problem of limited chroma decoding modes in existing technologies is solved, thereby improving encoding efficiency and quality and reducing bandwidth requirements.
Patent Information
- Application Number
- CN202480039838.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-03
- Filing Date
- 2024-06-04
- Publication Date
- 2026-01-13
AI Technical Summary
In existing video coding standards, chroma decoding mode is only enabled when the block is split into two trees, which limits the use of more efficient decoding modes and leads to insufficient coding efficiency and quality.
By employing single-tree block segmentation technology and applying intra-frame chroma copy mode to encode or decode chroma blocks of video data, a more efficient decoding mode can be used in the case of single-tree segmentation.
It improves the efficiency of video decoding, reduces power consumption, reduces the bandwidth for transmitting encoded video data, and improves the quality of video decoding.
Smart Images

Figure CN121336397A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Patent Application No. 18 / 732,460, filed June 3, 2024, and U.S. Provisional Patent Application No. 63 / 584,403, filed September 21, 2023, and U.S. Provisional Patent Application No. 63 / 509,732, filed June 22, 2023, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 18 / 732,460, filed June 3, 2024, claims the benefit of U.S. Provisional Patent Application No. 63 / 584,403, filed September 21, 2023, and U.S. Provisional Patent Application No. 63 / 509,732, filed June 22, 2023. TECHNICAL FIELD
[0002] This disclosure relates to video encoding and video decoding. BACKGROUND
[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio telephones, so-called “smart phones,” video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video coding techniques, such as those described in the standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile Video Coding (VVC), and extensions of such standards, as well as proprietary video codecs / formats, such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media, and the like. By implementing such video coding techniques, video devices can more efficiently send, receive, encode, decode, and / or store digital video information.
[0004] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy of the video sequence. For block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) can be partitioned into video blocks, which can also be referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in a neighboring block within the same picture. Video blocks in an inter-coded (P or B) slice of a picture can use spatial or temporal prediction with respect to reference samples in neighboring blocks within the same picture or reference pictures, respectively. Pictures can be referred to as frames, and reference pictures can be referred to as reference frames. SUMMARY
[0005] In general, this disclosure describes techniques for chroma component coding. In some video standards and implementations, some chroma coding modes are only enabled when a block is partitioned with dual tree partitioning (e.g., the chroma has a separate partition tree that can be different from the partition tree of the corresponding luma block). Such limitations can prevent the use of more efficient coding modes. This disclosure describes techniques that include using such chroma coding modes when a block can be partitioned with single tree partitioning. Such techniques can improve the efficiency of video coding, thereby reducing power consumption, reducing the bandwidth needed to transmit encoded video data, and / or improving video coding quality.
[0006] In one example, a method includes determining to use single tree block partitioning on a block of video data; applying an intra chroma copy mode to a chroma block of the block; and encoding or decoding the chroma block based on the single tree block partitioning and the intra chroma copy mode.
[0007] In another example, a device includes one or more memories configured to store video data and one or more processors in communication with the one or more memories and configured to determine to use single tree block partitioning on a block of the video data; apply an intra chroma copy mode to a chroma block of the block; and encode or decode the chroma block based on the single tree block partitioning and the intra chroma copy mode.
[0008] In another example, a device includes means for determining to use single tree block partitioning on a block of video data; means for applying an intra chroma copy mode to a chroma block of the block; and means for encoding or decoding the chroma block based on the single tree block partitioning and the intra chroma copy mode.
[0009] In another example, a computer-readable storage medium is encoded with instructions that, when executed, cause one or more programmable processors to determine to use single tree block partitioning on a block of video data; apply an intra chroma copy mode to a chroma block of the block; and encode or decode the chroma block based on the single tree block partitioning and the intra chroma copy mode.
[0010] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF DRAWINGS
[0011] Figure 1 is a block diagram illustrating an example video encoding and decoding system that can perform the techniques of this disclosure.
[0012] Figure 2 is a conceptual diagram illustrating an example of intra prediction modes.
[0013] Figure 3 is a conceptual diagram illustrating example reference samples for wide-angle intra prediction.
[0014] Figure 4 is a conceptual diagram illustrating an example of discontinuity in directions beyond 45 degrees.
[0015] Figure 5 is a conceptual diagram illustrating example positions of samples used to derive a and b.
[0016] Figure 6 is a conceptual diagram illustrating an example spatial portion of a convolution filter.
[0017] Figure 7 is a conceptual diagram illustrating an example reference region (with padding) used to derive filter coefficients.
[0018] Figure 8 is a conceptual diagram illustrating example spatial samples of a GL-CCCM.
[0019] Figure 9 is a conceptual diagram illustrating an example of 5 positions in a reconstructed luma sample.
[0020] Figure 10 is a conceptual diagram illustrating an example reference region of a BVG-CCCM.
[0021] Figure 11 is a flowchart illustrating example chroma component coding techniques in accordance with one or more aspects of this disclosure.
[0022] Figure 12 is a block diagram illustrating an example video encoder that can perform the techniques of this disclosure.
[0023] Figure 13 This is a block diagram illustrating an example video decoder that can perform the techniques of this disclosure.
[0024] Figure 14 This is a flowchart illustrating an example method for encoding the current block according to the technology of this disclosure.
[0025] Figure 15 This is a flowchart illustrating an example method for decoding the current block according to the technology of this disclosure. Detailed Implementation
[0026] Typically, this disclosure describes techniques for chroma component decoding. In some video standards and specific implementations, certain chroma decoding modes are enabled only when blocks are partitioned into two-tree partitions (e.g., chroma has a separate partition tree that may differ from the corresponding luma block partition tree). Such limitations may prevent the use of more efficient decoding modes. This disclosure describes techniques that include using such chroma decoding modes when blocks can be partitioned using single-tree partitioning. The techniques disclosed herein can save processing power, reduce the bandwidth used to transmit encoded video data, and / or improve decoding performance (e.g., improve the image quality of both encoded and decoded video data).
[0027] Figure 1 This is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. The techniques of this disclosure generally involve decoding (encoding and / or decoding) video data. Typically, video data includes any data used for processing video. Therefore, video data may include unencoded raw video, encoded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.
[0028] like Figure 1 As shown, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, source device 102 provides the video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can be or may include any of a wide range of devices, such as desktop computers, laptop computers, mobile devices, tablet computers, set-top boxes, mobile phones such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, etc. In some cases, source device 102 and destination device 116 may be configured for wireless communication and are therefore referred to as wireless communication devices.
[0029] exist Figure 1In the example, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes an input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for chroma component decoding. Therefore, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, the source device and destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device, rather than including an integrated display device.
[0030] like Figure 1 The system 100 shown is merely an example. Typically, any digital video encoding and / or decoding device can perform techniques for chroma component decoding. Source device 102 and destination device 116 are merely examples of such decoding devices, where source device 102 generates decoded video data for transmission to destination device 116. This disclosure refers to a “decoding” device as a device that performs the decoding (e.g., encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices, specifically, a video encoder and a video decoder, respectively. In some examples, source device 102 and destination device 116 may operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Therefore, system 100 may support one-way or two-way video transmission between source device 102 and destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0031] In general, video source 104 represents a source of video data (i.e., raw, uncoded video data) and provides a sequential series of pictures (also referred to as “frames”) of the video data to video encoder 200, which encodes data for the pictures. Video source 104 of source device 102 can include a video capture device, such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface to receive video from a video content provider. As a further alternative, video source 104 can generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured, pre-captured, or computer-generated video data. Video encoder 200 can rearrange the pictures from the received order (sometimes referred to as “display order”) into the coding order for coding. Video encoder 200 can generate a bitstream including encoded video data. Source device 102 can then output the encoded video data via output interface 108 onto computer- readable medium 110 for reception and / or retrieval by, for example, input interface 122 of destination device 116.
[0032] Memory 106 of source device 102 and memory 120 of destination device 116 represent general purpose memories. In some examples, memories 106, 120 can store raw video data, e.g., raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, memories 106, 120 can store software instructions capable of execution by, e.g., video encoder 200 and video decoder 300, respectively. While memories 106 and 120 are shown as separate from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 can also include internal memories for similar or equivalent purposes. Furthermore, memories 106, 120 can store encoded video data, e.g., from an output of video encoder 200 and an input of video decoder 300. In some examples, portions of memories 106, 120 can be allocated as one or more video buffers, e.g., to store raw decoded and / or encoded video data.
[0033] Computer-readable medium 110 can represent any type of medium or device capable of transporting encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium to enable source device 102 to transmit encoded video data directly to destination device 116 in real-time, e.g., via a radio frequency network or computer-based network. Output interface 108 can modulate the transmission signal including the encoded video data, and input interface 122 can demodulate the received transmission signal, according to a communication standard, such as a wireless communication protocol. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from source device 102 to destination device 116.
[0034] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.
[0035] In some examples, source device 102 can output encoded video data to file server 114, or another intermediate storage device, which can store the encoded video data generated by source device 102. Destination device 116 can access stored video data from file server 114 via streaming or download.
[0036] The file server 114 can be any type of server device capable of storing encoded video data and transmitting that encoded video data to the destination device 116. The file server 114 can represent a web server (e.g., for a website), a server configured to provide file delivery protocol services (such as the File Delivery Protocol (FTP) or the File Delivery over Unidirectional Transport (FLUTE) protocol), a content delivery network (CDN) device, a hypertext transfer protocol (HTTP) server, a multimedia broadcast multicast service (MBMS) or enhanced MBMS (eMBMS) server, and / or a network attached storage (NAS) device. The file server 114 can additionally or alternatively implement one or more HTTP streaming protocols, such as Dynamic Adaptive Streaming over HTTP (DASH), HTTP Live Streaming (HLS), Real Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, and the like.
[0037] The destination device 116 can access encoded video data from the file server 114 through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both that is suitable for accessing encoded video data stored on the file server 114. The input interface 122 can be configured to operate according to any one or more of various protocols discussed above for retrieving or receiving media data from the file server 114, or other such protocols for retrieving media data.
[0038] The output interface 108 and the input interface 122 can represent wireless transmitters / receivers, modems, wired networking components (e.g., Ethernet cards), wireless communication components operating according to any of a variety of IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 comprise wireless components, the output interface 108 and the input interface 122 can be configured to transfer data, such as encoded video data, according to a cellular communication standard, such as 4G, 4G-LTE (Long-Term Evolution), LTE Advanced, 5G, or the like. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 can be configured to transfer data, such as encoded video data, according to other wireless standards ™ ™ Standard, etc. In some examples, source device 102 and / or destination device 116 can comprise respective system on a chip (SoC) devices. For example, source device 102 can comprise a SoC device to perform the functionality attributable to video encoder 200 and / or output interface 108, and destination device 116 can comprise a SoC device to perform the functionality attributable to video decoder 300 and / or input interface 122.
[0039] The techniques of this disclosure can be applied to video coding in support of any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions, digital video that is encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0040] Input interface 122 of destination device 116 receives an encoded video bitstream from computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, or the like). The encoded video bitstream can include signaling information defined by video encoder 200 and also used by video decoder 300, such as syntax elements having values
[0041] Although in Figure 1Although not shown, in some examples, video encoder 200 and video decoder 300 can each be integrated with an audio encoder and / or audio decoder (e.g., an audio CODEC), and can include appropriate MUX-DEMUX units or other hardware and / or software to handle multiplexed streams including both audio and video in a common data stream. Example audio CODECs can include AAC, AC-3, AC-4, ALAC, ALS, AMBE, AMR, AMR-WB (G.722.2), AMR-WB+, aptx (various versions), ATRAC, BroadVoice (BV16, BV32), CELT, Enhanced AC-3 (E-AC-3), EVS, FLAC, G.711, G.722, G.722.1, G.722.2 (AMR-WB), G.723.1, G.726, G.728, G.729, G.729.1, GSM-FR, HE-AAC, iLBC, iSAC, LA Lyra, Monkey's Audio, MP1, MP2 (MPEG-1, 2 Audio Layer II), MP3, Musepack, Nellymoser Asao, OptimFROG, Opus, Sac, Satin, SBC, SILK, Siren 7, Speex, SVOPC, True Audio (TTA), TwinVQ, USAC, Vorbis (Ogg), WavPack, and Windows Media Audio.
[0042] Video encoder 200 and video decoder 300 each can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware or any combinations thereof. When the techniques are implemented partially in software, a device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of video encoder 200 and video decoder 300 can be included in one or more encoders or decoders, any of which alone can be a part of a combined encoder / decoder (CODEC) in a respective device. A device including video encoder 200 and / or video decoder 300 can implement video encoder 200 and / or video decoder 300 in a processing circuitry, such as an integrated circuit and / or a microprocessor. Such a device can be a wireless communication device, such as a cellular phone or any other type of device described herein.
[0043] Video encoder 200 and video decoder 300 can operate according to video coding standards, such as ITU-T H.265, also referred to as High Efficiency Video Coding (HEVC), or extensions thereof, such as the multi-view and / or scalable video coding extensions. Alternatively, video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards, such as ITU-T H.266, also referred to as Versatile Video Coding (VVC). In other examples, video encoder 200 and video decoder 300 can operate according to proprietary video codecs / formats, such as AOMedia Video 1 (AV1), extensions of AV1, and / or subsequent versions of AV1 (e.g., AV2). In other examples, video encoder 200 and video decoder 300 can operate according to other proprietary formats or industry standards. The techniques of this disclosure, however, are not limited to any particular coding standard or format. Generally, video encoder 200 and video decoder 300 can be configured to perform the techniques of this disclosure in connection with any video coding technology that uses chroma component coding.
[0044] Generally, video encoder 200 and video decoder 300 can perform block-based coding of pictures. The term “block” generally refers to a structure containing data to be processed (e.g., encoded, decoded, or otherwise used) during encoding and / or decoding. For example, a block can include a two-dimensional matrix of samples of luma and / or chroma data. Generally, video encoder 200 and video decoder 300 can code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, rather than coding red, green, and blue (RGB) data for samples of a picture, video encoder 200 and video decoder 300 can code luma and chroma components, where the chroma components can include both red hue and blue hue chroma components. In some examples, video encoder 200 converts received RGB format data to a YUV representation prior to encoding, and video decoder 300 converts the YUV representation to the RGB format. Alternatively, pre- and post-processing units (not shown) can perform these conversions.
[0045] This disclosure can generally relate to coding (e.g., encoding and decoding) of pictures to include processes of encoding or decoding data of pictures. Similarly, this disclosure can relate to coding of blocks of pictures to include processes of encoding or decoding (e.g., prediction and / or residual coding) data for blocks. A coded video bitstream generally includes a series of values for syntax elements representing coding decisions (e.g., coding modes) as well as partitioning of pictures into blocks. Accordingly, references to coding of pictures or blocks generally should be understood to refer to coding values of syntax elements forming the pictures or blocks.
[0046] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as video encoder 200) partitions a coding tree unit (CTU) into CUs according to a quad tree structure. That is, the video coder partitions a CTU and CUs into four equal, non overlapping squares, and each node of the quad tree has either zero or four child nodes. Nodes with zero child nodes can be referred to as“leaf nodes,” and CUs of such leaf nodes can include one or more PUs and / or one or more TUs. The video coder can further partition PUs and TUs. For example, in HEVC, a residual quad tree (RQT) represents partitioning of TUs. In HEVC, PUs represent inter prediction data, while TUs represent residual data. Intra predicted CUs include intra prediction information, such as an intra mode indication.
[0047] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, a video coder (such as video encoder 200) partitions a picture into CTUs. Video encoder 200 can partition a CTU according to a tree structure, such as a quad-tree binary tree (QTBT) structure or Multi-Type Tree (MTT) structure. The QTBT structure removes the concepts of multiple partition types, such as the separation between CUs, PUs, and TUs of HEVC. The QTBT structure includes two levels: a first level partitioned according to quad-tree splitting, and a second level partitioned according to binary tree splitting. The root node of the QTBT structure corresponds to a CTU. Leaf nodes of the binary tree correspond to CUs.
[0048] In the MTT partition structure, blocks can be partitioned using quad tree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) (also referred to as tri-tree (TT)) partitioning. A ternary tree or tri-tree partition is a partition in which a block is split into three sub-blocks. In some examples, a ternary tree or tri-tree partition divides a block into three sub-blocks without dividing the original block through the center. The partition types (e.g., QT, BT, and TT) in the MTT can be symmetric or asymmetric.
[0049] When operating according to the AV1 codec, video encoder 200 and video decoder 300 can be configured to code video data in units of blocks. In AV1, the largest coding block that can be processed is referred to as a superblock. In AV1, a superblock can be 128x128 luma samples or 64x64 luma samples. However, in subsequent video coding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luma sample sizes. In some examples, a superblock is the top level of a block quad tree. Video encoder 200 can further partition a superblock into smaller coding blocks. Video encoder 200 can partition superblocks and other coding blocks into smaller blocks using square or non-square partitions. Non-square blocks can include N / 2xN blocks, NxN / 2 blocks, N / 4xN blocks, and NxN / 4 blocks. Video encoder 200 and video decoder 300 can perform separate prediction and transform processing for each coding block.
[0050] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that can be coded independently of other tiles. That is, video encoder 200 and video decoder 300 can encode and decode coding blocks within a tile without using video data from other tiles. However, video encoder 200 and video decoder 300 can perform filtering across tile boundaries. The size of a tile can be uniform or non-uniform. Tile-based coding can enable parallel processing and / or multi-threading of encoder and decoder implementations.
[0051] In some examples, video encoder 200 and video decoder 300 can use a single QTBT or MTT structure to represent each of luma and chroma components, while in other examples, video encoder 200 and video decoder 300 can use two or more QTBT or MTT structures, such as one QTBT / MTT structure for luma components and another QTBT / MTT structure for two chroma components (or two QTBT / MTT structures for respective chroma components).
[0052] Video encoder 200 and video decoder 300 can be configured to use quad tree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partition structures.
[0053] In some examples, a CTU includes a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples of a picture having three sample arrays, or a monochrome picture or a CTB of samples of a picture coded using three separate color planes and syntax structures for coding samples. A CTB can be an NxN block of samples of some N value, such that one partitioning is to divide components into CTBs. A component is an array or a single sample from one of the three arrays (luma and two chroma) that make up a 4:2:0, 4:2:2, or 4:4:4 color format picture, or an array or a single sample of an array that makes up a monochrome format picture. In some examples, a coding block is an MxN block of samples of some M value and N value, such that one partitioning is to divide CTBs into coding blocks.
[0054] Blocks (e.g., CTUs or CUs) can be grouped in pictures in various ways. As one example, a tile can refer to a rectangular region of CTU rows within a particular tile in a picture. A tile column can refer to a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile row can refer to a rectangular region of CTUs having a height equal to a height of the picture and a width specified by a syntax element (e.g., such as in a picture parameter set) and a tile column can refer to a rectangular region of CTUs having a width equal to a width of the picture and a height specified by a syntax element (e.g., such as in a picture parameter set).
[0055] In some examples, a tile can be divided into multiple bricks, each of which can include one or more CTU rows within the tile. A tile that is not divided into multiple bricks can also be referred to as a brick. However, a brick that is a true subset of a tile can not be referred to as a tile. Bricks in a picture can also be arranged in slices. A slice can be an integer number of tiles of a picture, which can be uniquely contained in a single network abstraction layer (NAL) unit. In some examples, a slice includes multiple complete tiles or only a contiguous sequence of complete bricks of a single tile.
[0056] The present disclosure can use “NxN” and “N by N” interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in the vertical and horizontal dimensions, e.g., 16x16 samples or 16 by 16 samples. In general, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Likewise, an NxN CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a nonnegative integer value. The samples in a CU can be arranged in rows and columns. Moreover, a CU need not necessarily have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can comprise NxM samples, where M need not necessarily equal N.
[0057] Video encoder 200 encodes video data representing prediction and / or residual information for a CU, as well as other information. Prediction information indicates how to predict the CU in order to form a prediction block for the CU. Residual information generally represents sample-by-sample differences between the CU prior to encoding and the prediction block.
[0058] To predict a CU, video encoder 200 generally forms a prediction block for the CU through inter prediction or intra prediction. Inter prediction generally refers to predicting the CU from data of a previously coded picture, whereas intra prediction generally refers to predicting the CU from previously coded data of the same picture. To perform inter prediction, video encoder 200 can use one or more motion vectors to generate the prediction block. Video encoder 200 can generally perform a motion search to identify a reference block that closely matches the CU, e.g., according to a difference between the CU and the reference block. Video encoder 200 can calculate the difference metric using a sum of absolute difference (SAD), sum of squared difference (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether a reference block closely matches the current CU. In some examples, video encoder 200 can use uni -prediction or bi-prediction to predict the current CU.
[0059] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter prediction mode. Under the affine motion compensation mode, video encoder 200 can determine two or more motion vectors that represent non-translational motion, such as scaling or zooming, rotation, perspective motion, or other irregular types of motion.
[0060] To perform intra prediction, video encoder 200 can select an intra prediction mode to generate the prediction block. Some examples of VVC provide sixty-seven intra prediction modes, including various directional modes, as well as a planar mode and a DC mode. Generally, video encoder 200 selects an intra prediction mode that describes neighboring samples of the current block (e.g., of a CU) from which to predict samples of the current block. Such samples can generally be located above, above and to the left, or to the left of the current block in the same picture, assuming video encoder 200 is coding CTUs and CUs in a raster scan order (left-to-right, top-to-bottom).
[0061] Video encoder 200 encodes data representing the prediction mode of the current block. For example, for inter prediction modes, video encoder 200 can encode data indicating which of the various available inter prediction modes to use, as well as motion information for the corresponding mode. For example, for uni - or bi-prediction, video encoder 200 can encode motion vectors using advanced motion vector prediction (AMVP) or merge mode. Video encoder 200 can use similar modes to encode motion vectors for the affine motion compensation mode.
[0062] AV1 includes two general techniques for encoding and decoding blocks of video data. The two general techniques are intra prediction (e.g., intra prediction or spatial prediction) and inter prediction (e.g., inter prediction or temporal prediction). In the context of AV1, when using an intra prediction mode to predict a block of a current frame of video data, video encoder 200 and video decoder 300 do not use video data from other frames of the video data. For most intra prediction modes, video encoder 200 encodes the block of the current frame based on differences between sample values in the current block and prediction values generated from reference samples in the same frame. Video encoder 200 determines the prediction values generated from the reference samples based on the intra prediction mode.
[0063] Following prediction, such as intra prediction or inter prediction of a block, video encoder 200 can calculate residual data for the block. The residual data, such as a residual block, represents sample-by-sample differences between the block and a prediction block formed using the corresponding prediction mode. Video encoder 200 can apply one or more transforms to the residual block to produce transformed data in a transform domain rather than in the sample domain. For example, video encoder 200 can apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform. In addition, video encoder 200 can apply a secondary transform following the first transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), or the like. Video encoder 200 produces transform coefficients following application of the one or more transforms.
[0064] As noted above, following any transforms that produce transform coefficients, video encoder 200 can perform quantization of the transform coefficients. Quantization generally refers to a process in which transform coefficients are quantized to possibly reduce the amount of data used to represent the transform coefficients, providing further compression. By performing the quantization process, video encoder 200 can reduce the bit depth of some or all of the transform coefficients. For example, video encoder 200 can round n-bit values down to m-bit values during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 can perform a bit shift to the right on values to be quantized.
[0065] After quantization, video encoder 200 can scan the transform coefficients, producing a one-dimensional vector from the two-dimensional matrix comprising the quantized transform coefficients. The scan can be designed to place higher energy (and hence less frequent) transform coefficients earlier in the vector and lower energy (and hence more frequent) transform coefficients later in the vector. In some examples, video encoder 200 can utilize a pre-defined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy encode the quantized transform coefficients of the vector. In other examples, video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 can entropy encode the one-dimensional vector, e.g., according to context adaptive binary arithmetic coding (CABAC). Video encoder 200 can also entropy encode values for syntax elements that describe metadata associated with encoded video data used by video decoder 300 in decoding the video data.
[0066] To perform CABAC, video encoder 200 can assign a context within a context model to a symbol to be transmitted. The context can relate to, for example, whether neighboring values of the symbol are zero-valued or not. The probability determination can be based on the context assigned to the symbol.
[0067] Video encoder 200 can further generate syntax data, such as block-based, picture-based, and sequence-based syntax data, or other syntax data such as a sequence parameter set (SPS), picture parameter set (PPS), or video parameter set (VPS), to video decoder 300, e.g., in picture headers, block headers, slice headers, or the like. Video decoder 300 can likewise decode such syntax data to determine how to decode corresponding video data.
[0068] In this way, video encoder 200 can generate a bitstream including encoded video data, e.g., syntax elements describing partitioning of a picture into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Ultimately, video decoder 300 can receive the bitstream and decode the encoded video data.
[0069] In general, video decoder 300 performs a reciprocal process to that performed by video encoder 200 to decode the encoded video data of the bitstream. For example, video decoder 300 can decode values for syntax elements of the bitstream using CABAC in substantially a reciprocal manner, although in reverse order, to the CABAC encoding process of video encoder 200. The syntax elements can define partitioning information for partitioning a picture into CTUs, and partitioning each CTU according to a corresponding partition structure such as a QTBT structure to define CUs of the CTU. The syntax elements can further define prediction and residual information for blocks (e.g., CUs) of video data.
[0070] The residual information can be represented by, for example, quantized transform coefficients. Video decoder 300 can inverse quantize and inverse transform the quantized transform coefficients of the block to reproduce a residual block for the block. Video decoder 300 forms a prediction block for the block using the signaled prediction mode (intra prediction or inter prediction) and related prediction information (e.g., motion information for inter prediction). Video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reproduce the original block. Video decoder 300 can perform additional processing such as performing a deblocking process to reduce visual artifacts along boundaries of the blocks.
[0071] This disclosure can generally relate to “signaling” certain information, such as syntax elements. The term “signaling” can generally refer to the communication of values for syntax elements and / or other data used for decoding encoded video data. That is, video encoder 200 can signal values for syntax elements in a bitstream. Generally, signaling refers to generating values in a bitstream. As noted above, source device 102 can transmit the bitstream to destination device 116 in real time or not in real time, such as can occur when syntax elements are stored to storage device 112 for later retrieval by destination device 116.
[0072] According to techniques of this disclosure, a method includes determining to use single tree block partitioning on a block of video data; applying an intra chroma copy mode to a chroma block of the block; and encoding or decoding the chroma block based on the single tree block partitioning and the intra chroma copy mode.
[0073] Figure 2 is a conceptual diagram illustrating examples of intra prediction modes. Now describing chroma intra prediction in the Versatile Video Coding (VVC) standard, including intra mode coding with 67 intra prediction modes 290. To capture arbitrary edge directions present in natural video, the number of directional intra modes in VVC is extended from 33 used in HEVC to 65. The new directional modes not in HEVC are depicted in Figure 2 with dashed lines with arrows, and the planar and DC modes remain unchanged. Such denser directional intra prediction modes apply to all block sizes and for both luma and chroma intra prediction.
[0074] In VVC, for non-square blocks, several regular angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. Wide-angle intra prediction will be described later in this disclosure.
[0075] In HEVC, each intra coded block has a square shape, and each side of the square shape has a length of a power of 2. Thus, generating an intra predictor using the DC mode does not require a division operation. In VVC, a block can have a rectangular shape, such that in general case a division operation becomes necessary according to the block. To avoid the division operation for DC prediction, only the longer side of the rectangular shape is used to calculate the average value of the non-square block.
[0076] Wide-angle intra prediction for non-square blocks is now described. The regular angular intra prediction directions are defined as 45 degrees to -135 degrees in the clockwise direction. In VVC, for non-square blocks, several regular angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. For example, the video encoder 200 can signal the replaced modes to the video decoder 300. The total number of intra prediction modes is unchanged, i.e., 67, and the intra mode coding method is unchanged.
[0077] Figure 3 is a conceptual diagram illustrating example reference samples for wide-angle intra prediction. To support these prediction directions, the top reference of length 2W+1 and the left reference of length 2H+1 are defined as shown in Figure 3 For example, the top reference 360 has a length of two times the samples of the width of the rectangular block 362 (8 samples) plus one sample (2 x 8 samples + 1 sample = 17 samples). Similarly, the top reference 370 has a length of two times the samples of the width of the rectangular block 372 (4 samples) plus one sample (2 x 4 samples + 1 sample = 9 samples). For example, the left reference 364 has a length of two times the samples of the height of the rectangular block 362 (4 samples) plus one sample (2 x 4 samples + 1 sample = 9 samples). Similarly, the left reference 374 has a length of two times the samples of the height of the rectangular block 372 (8 samples) plus one sample (2 x 8 samples + 1 sample = 17 samples).
[0078] The number of replaced modes in the wide-angle direction modes depends on the aspect ratio (e.g., the ratio of the width / height) of the block. The replaced intra prediction modes are listed in Table 1.
[0079] Table 1: Intra prediction modes replaced by wide-angle modes
[0080]
[0081] Figure 4 is a conceptual diagram illustrating an example of discontinuity in the case of directions beyond 45 degrees. As Figure 4As shown, in the case of wide-angle intra-frame prediction, two vertically adjacent predicted samples of predicted sample 450 can use two non-adjacent reference samples of reference sample 452. Therefore, for example, if the wide-angle mode represents a non-fractional offset, low-pass reference sample filtering and side smoothing filtering can be applied to wide-angle prediction to reduce the negative impact of the increased gap ∆pα. For example, video encoder 200 or video decoder 300 can apply such filtering and smoothing filtering. There are eight wide-angle modes that meet this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When using any of these modes to predict blocks, the samples in the reference buffer are directly copied without any interpolation. This modification reduces the number of samples that need to be smoothed. Furthermore, this technique makes the design of non-fractional modes consistent with both regular prediction modes and wide-angle modes.
[0082] In VVC, 4:2:2 and 4:4:4 chroma formats, as well as 4:2:0, are supported. The chroma export mode (DM) export table for the 4:2:2 chroma format was originally ported from HEVC, with the number of entries expanded from 35 to 67 to match the expansion of intra-prediction modes. Since the HEVC specification does not support prediction angles below −135 degrees or above 45 degrees, luma intra-prediction modes in the range of 2 to 5 can be mapped to 2. Therefore, the 4:2:2 chroma DM export table updates the chroma format by replacing some values in the mapping table entries to more accurately convert the prediction angle of chroma blocks.
[0083] The cross-component linear model prediction is now described. To reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode can be used in VVC. For this mode, chromaticity samples are predicted based on reconstructed luminance samples from the same CU using a linear model as follows:
[0084]
[0085] in This represents the predicted chromaticity samples in the CU, and This represents the downsampled and reconstructed luminance samples from the same CU.
[0086] CCLM parameters ( and The CCLM parameters are derived from adjacent chroma samples and their corresponding downsampled luminance samples. For example, a video encoder 200 or a video decoder 300 can derive CCLM parameters.
[0087] Figure 5 This is a conceptual diagram illustrating example locations used to derive α and β. Figure 5An example of the position of the left and above samples is shown for the samples of the current block involved in the CCLM mode. For example, for the chroma block 550, the corresponding block 560 of reconstructed luma samples is shown. The video encoder 200 or the video decoder 300 can use the samples shown in the circles to determine a and b.
[0088] The above template and the left template can be used together to calculate the linear model coefficients. The above template and / or the left template can also be used alternatively for the other two linear model (LM) modes, referred to as the LM A and LM L modes.
[0089] In the LM T mode, only the above template is used to calculate the linear model coefficients. To get more samples, the above template can be extended to (W+H) samples. In the LM L mode, only the left template is used to calculate the linear model coefficients. To get more samples, the left template can be extended to (H+W) samples. In the LM LT mode, the left template and the above template are used to calculate the linear model coefficients. For example, the video encoder 200 or the video decoder 300 can use the above template, the left template, or both the above template and the left template described above.
[0090] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a downsampling ratio of 2 to 1 in the horizontal and vertical directions. The selection of the downsampling filter is specified by an SPS level flag. For example, the video encoder 200 can signal an SPS level flag that indicates which downsampling filter the video decoder 300 should apply.
[0091] Chroma intra mode coding is now described. The video encoder 200 or the video decoder 300 can use chroma intra mode coding. For chroma intra mode coding, the chroma intra mode coding allows a total of 8 intra modes. Those modes include five traditional intra modes and three cross-component linear model modes (CCLM, LM A, and LM L). The chroma mode signaling and derivation process is shown in Table 2. The chroma mode coding can directly depend on the intra prediction mode of the corresponding luma block. Since separate block partitioning structures for the luma component and the chroma component are enabled in I slices, one chroma block can correspond to multiple luma blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luma block that covers the center position of the current chroma block can be directly inherited.
[0092] Table 2: Chroma prediction mode derivation from luma mode when cclm is enabled.
[0093]
[0094] Regardless of the value of sps_cclm_enabled_flag, a single binarization table is used, as shown in Table 3.
[0095] Table 3: Unified binarization table for chroma prediction modes
[0096]
[0097] In Table 3, the first bin indicates whether the mode is regular (0) or LM mode (1). If the mode is LM mode, the next bin indicates whether the mode is LM_CHROMA (0). If the mode is not LM_CHROMA, the next bin indicates whether the mode is LM_L (0) or LM_A (1). For such cases, when sps_cclm_enabled_flag is 0, the first bin of the binarization table for intra_chroma_pred_mode can be discarded before entropy coding. Or in other words, the first bin is inferred to be 0 and thus not coded. This single binarization table is used for both cases when sps_cclm_enabled_flag is equal to 0 and 1. The first two bins in Table 3 can be context coded with their own context models, and the rest of the bins can be bypass coded.
[0098] Intra block copy is now described. Intra block copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding efficiency for screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder (e.g., video encoder 200) to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed inside the current picture.
[0099] IBC in VVC is now described. The luma block vector of an IBC coded CU is integer-precision. The chroma block vector is also rounded to integer precision. When used in combination with advanced motion vector resolution (AMVR), the IBC mode can switch between 1-pel motion vector precision and 4-pel motion vector precision.
[0100] The video encoder can signal the IBC mode as follows. For example, at the CU level, the IBC mode can be signaled with a flag and can be signaled as an IBC advanced motion vector prediction (AMVP) mode or an IBC skip / merge mode as follows: 1) IBC skip / merge mode: a merge candidate index is used to indicate which block vector from a list of neighboring candidate IBC coded blocks is used to predict the current block. The merge list can include spatial, history-based motion vector predictor (HMVP), and / or paired candidates. 2) IBC AMVP mode: block vector difference is coded in the same way as motion vector difference. Block vector prediction technique uses two candidates as predictors, one from the left neighbor and one from the above neighbor (if IBC coded). When either neighbor is not available, a default block vector can be used as a predictor. A flag is signaled to indicate the block vector predictor index.
[0101] The interaction between IBC mode and other inter coding tools in VVC, such as paired merge candidates, HMVP, combined intra / inter prediction mode (CIIP), merge mode with motion vector difference (MMVD), and geometric partition mode (GPM), is as follows: 1) IBC can be used together with paired merge candidates and HMVP. A new paired IBC merge candidate can be generated by averaging the two IBC merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference. 2) IBC cannot be used in combination with the following interworking tools: affine motion, CIIP, MMVD, and GPM. 3) When using DUAL_TREE partitioning, IBC is not allowed for chroma coding blocks.
[0102] Unlike the HEVC screen content coding extension, the current picture is no longer included as one of the reference pictures in the reference picture list 0 for IBC prediction. The derivation process of the motion vector for IBC mode excludes all neighboring blocks in the inter mode, and vice versa. The following IBC design aspects are applied: 1) IBC shares the same process as in the regular MV merge mode, including pair-wise merge candidates and history-based motion predictor, but does not allow temporal motion vector prediction (TMVP) and zero vector, as they are not valid for IBC mode; 2) separate HMVP buffers (5 candidates each) for regular motion vectors (MVs) and IBC; 3) block vector constraints are implemented in the form of bitstream conformance constraints, the video encoder 200 ensures that no invalid vectors exist in the bitstream, and the merge mode should not be used if the merge candidate is invalid (e.g., out of range or zero); such bitstream conformance constraints are expressed in terms of virtual buffers, as described below; 4) for deblocking, IBC is treated as an inter mode; 6) if the IBC prediction mode is used to code the current block, AMVR does not use quarter-pel; instead, AMVR is signaled to indicate only whether the MV is inter-pel or 4 integer-pel; and / or 7) the number of IBC merge candidates can be signaled separately from the number of regular, subblock, and geometric merge candidates in the slice header.
[0103] IntraTMP in ECM is now described. In the enhanced compression model (ECM), in addition to IBC, there is yet another intra block copy mode, called intra template matching prediction (IntraTMP).
[0104] Intra template matching prediction is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the video encoder 200 searches for the template in the reconstructed part of the current frame that is most similar to the current template, and uses the corresponding block as the prediction block. The video encoder 200 then signals the use of this mode, and the same prediction operation is performed at the video decoder 300.
[0105] The prediction signal is generated by matching the L-shaped causal neighbor of the current block with another block in a predefined search area. The sum of absolute differences (SAD) is used as the cost function.
[0106] Within each zone, the video decoder 300 searches for the template with the smallest SAD with respect to the current template, and uses the corresponding block of the template with the smallest SAD as the prediction block.
[0107] The intra template matching tool is enabled for CUs with width and height size smaller than or equal to 64. The maximum CU size for intra template matching is configurable.
[0108] When the decoder-side intra-mode export (DIMD) is not used for the current CU, a special flag is used to signal the intra-template matching prediction mode at the CU level. For example, video encoder 200 can signal such a special flag to video decoder 300.
[0109] The new chroma intra-prediction modes in the Enhanced Compression Model (ECM) are now described. First, the Multi-Model Linear Model (MMLM) is described. The Cross-Component Linear Model (CCLM) included in the VVC is extended by adding three MMLM modes. See, for example, Zhang et al., “Enhanced Cross-component Linear Model Intra-prediction,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 4th Meeting: Chengdu, CN, 15-21 Oct. 2016, JVET-D0110. In each MMLM mode, reconstructed neighboring samples are divided into two classes using a threshold, which is the average of the luminance reconstructed neighboring samples. The linear model for each class is derived using the Least Mean Square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. The video encoder 200 or the video decoder 300 can use either the CCLM mode or the MMLM mode.
[0110] The slope adjustment for CCLM is now described. Slope adjustment is applied to both CCLM and MMLM prediction. For example, a video encoder 200 or a video decoder 300 can perform slope adjustment. The adjustment includes a tilt linear function that maps luminance values to chrominance values relative to a center point determined by the average luminance values of a reference sample.
[0111] CCLM uses a two-parameter model to map luminance values to chrominance values. The slope parameter is... "and bias parameters" The mapping is defined as follows:
[0112]
[0113] The slope parameter "u" is signaled to update the model to the following form:
[0114]
[0115] in
[0116]
[0117]
[0118] By this selection, the mapping function is tilted or rotated around the point with luminance value The average of the reference luminance samples used in the model creation is to provide a meaningful modification to the model.
[0119] The video encoder 200 or the video decoder 300 can provide a slope adjustment parameter. The slope adjustment parameter is provided as an integer between -4 and 4, inclusive, and is signaled in the bitstream (e.g., by the video encoder 200). The unit of the slope adjustment parameter is 1 / 8 of the chroma sample value per one luminance sample value (for 10-bit content).
[0120] The adjustment is applicable to CCLM models using reference samples above and to the left of the block, but not to "single-sided" modes (LM_A, LM_L). Such selection is based on a trade-off consideration of coding efficiency and complexity.
[0121] When applying slope adjustment to the multi-mode CCLM model, both models can be adjusted, and thus up to two slope updates can be signaled for a single chroma block. For example, the video encoder 200 can determine and signal up to two slope updates to the video decoder 300.
[0122] A convolutional cross-component model is now described. In Astola, et al. “AHG12: Convolutional cross-component model (CCCM) for intra prediction,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 26th Meeting: by Teleconference, 20-29 Apr. 2022, JVET-Z0064 and Astola, et al., “EE2-1.1a: Convolutional cross-component intra prediction model,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 27th Meeting: by Teleconference, 13-22 Jul. 2022, JVET-AA0057, a convolutional cross-component model (CCCM) is proposed to predict chroma samples from reconstructed luma samples in a similar spirit as performed by the CCLM mode. As with CCLM, when chroma sub-sampling is used, reconstructed luma samples are down-sampled to match the lower resolution chroma grid. Video encoder 200 or video decoder 300 can utilize the CCCM technique and / or the CCLM technique.
[0123] Further, similar to CCLM, there is an option to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one derived for samples above the average luma reference value and the other for the rest of the samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.
[0124] Figure 6 is a conceptual diagram illustrating an example spatial part of a convolutional filter. The convolutional 7-tap filter consists of a 5-tap plus sign shaped spatial component (herein referred to as “spatial 5-tap component 600”), a non-linear term, and a bias term. The input to the spatial 5-tap component 600 of the filter consists of the center (C) luma sample co-located with the chroma sample to be predicted and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, as shown in Figure 6 Video encoder 200 or video decoder 300 can apply the convolutional 7-tap filter.
[0125] The non-linear term P is the square of the central luma sample C, and is scaled to the sample value range of the content:
[0126] P = (C*C + midVal) » bitDepth
[0127] That is, for 10-bit content, it is computed as:
[0128] P = (C*C + 512) » 10
[0129] The bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM), and is set to the mid chroma value (512 for 10-bit content).
[0130] The output of the filter is computed as the convolution between the filter coefficients c i and the input values, and is clipped to the range of valid chroma samples:
[0131] predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B
[0132] Figure 7 is an illustrative conceptual diagram of the example reference region (with padding) used to derive the filter coefficients. The filter coefficients ci are computed by minimizing the mean squared error (MSE) between the predicted and reconstructed chroma samples in the reference region 702 (represented by the hatched padding box) of the PU 700. Figure 7 The reference region 702 is illustrated as including 6 rows of chroma samples above and to the left of the PU 700. The reference region 702 is extended one PU width to the right, and one PU height below the PU boundary of the PU 700 (via the extension 704 shown in the box filled with dashed lines). The reference region 702 can be adjusted to include only available samples. The extension 704 of the reference region 702 is used to support the “side samples” of the positive form spatial filter, and is padded when in an unavailable region.
[0133] The MSE minimization is performed by computing the autocorrelation matrix of the luma inputs, and the cross-correlation vector between the luma inputs and the chroma outputs. The autocorrelation matrix is LDL-decomposed, and the final filter coefficients are computed using back-substitution. This process roughly follows the computation of the adaptive linear filter (ALF) filter coefficients in ECM, but chooses LDL-decomposition instead of Cholesky-decomposition to avoid using square root operations. The proposed method only uses integer arithmetic.
[0134] A gradient and location-based CCCM (GL-CCCM) is now described.Figure 8 is a conceptual diagram illustrating example spatial samples of GL-CCCM. When performing the GL-CCCM technique as shown Figure 8 , video encoder 200 or video decoder 300 can utilize spatial samples 800.
[0135] The GL-CCCM technique uses gradient and position information, rather than the 4 spatial neighboring samples in the CCCM filter. The GL-CCCM filter for prediction is:
[0136] predChromaVal = c0C + c1G y + c2G x + c3Y + c4X + c5P + c6B
[0137] where G y and G x are the vertical and horizontal gradients, respectively, and are calculated as:
[0138] G y = (2N + NW + NE) - (2S + SW + SE)
[0139] G x = (2W + NW + SW) - (2E + NE + SE)
[0140] Furthermore, in some examples, the Y and X parameters are the vertical and horizontal positions of the center luma sample.
[0141] The remaining parameters can be the same as the CCCM tool. The reference region for parameter calculation is the same as the CCCM technique.
[0142] Gradient linear model is now described. For YUV 4:2:0 color format, a gradient linear model (GLM) technique can be used to predict chroma samples from luma sample gradients. Two modes are supported: a two-parameter GLM mode and a three-parameter GLM mode. Video encoder 200 or video decoder 300 can utilize the GLM technique.
[0143] Compared to the CCLM mode, the two-parameter GLM mode utilizes luma sample gradients instead of down-sampled luma values to derive the linear model. Specifically, when the two-parameter GLM mode is applied, the input to the CCLM process, i.e., the down-sampled luma sample , is replaced by the luma sample gradient . Other parts of the CCLM mode (e.g., parameter derivation, prediction sample linear conversion) remain unchanged. Thus, in the two-parameter GLM mode case, video encoder 200 or video decoder 300 can use the luma sample gradient , slope parameter , and a bias parameter to determine a prediction value of a chroma sample as follows:
[0144]
[0145] In the three-parameter GLM mode, the chroma samples can be predicted based on the luma sample gradient and down-sampled luma values with different parameters. The model parameters of the three-parameter GLM mode are derived from 6 columns and 6 rows of neighboring samples via the LDL decomposition based MSE minimization technique used in the CCCM. For example, in the three-parameter GLM mode case, the video encoder 200 or the video decoder 300 can determine a prediction value of a chroma sample as follows:
[0146]
[0147] wherein is the prediction value of the chroma sample, is the luma sample gradient, is the down-sampled luma sample, is a first slope parameter, is a second slope parameter, is a third slope parameter, and is a bias parameter.
[0148] A direct block vector mode for chroma prediction is now described. The video encoder 200 or the video decoder 300 can utilize the direct block vector mode for chroma prediction. In Huo, et al., “EE2-3.1: Direct block vector mode for chroma prediction,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 29th Meeting: by Teleconference, 11-20 Jan. 2023, JVET-AC0071, a direct block vector (DBV) is described to improve coding efficiency of chroma components when bi-tree is activated in an intra slice. This technique describes or includes two specific implementations.
[0149] Figure 9 is a conceptual diagram illustrating an example of 5 positions in the reconstructed luma sample. When chroma bi-tree is activated in an intra slice, for a chroma CU 900 coded with DBV mode, if Figure 9If one of the five positions (i.e., top-left 902, top-right 904, center 906, bottom-left 908, or bottom-right 910) is coded, then the luma block vector bvL is used to derive the chroma block vector bvC.
[0150] A block vector guided CCCM (BVG-CCCM) is now described. Video encoder 200 or video decoder 300 can utilize the BVG-CCCM.
[0151] In Youvalari, et al., “AHG12: Block vector guided CCCM,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 30th Meeting: Antalya, TR, 21-28 Apr. 2023, JVET-AD0100, a BVG-CCCM technique is proposed to improve coding efficiency of ECM. The BVG-CCCM technique uses the block vector of a collocated luma block coded in IBC or intraTMP mode to determine the reference region used to compute the CCCM parameters. Then, the reference region in luma and the corresponding region in chroma channels are used to compute the CCCM parameters. The prediction uses the computed model parameters and the collocated luma samples to make the CCCM prediction.
[0152] Figure 10 FIG. 10 is a conceptual diagram illustrating an example reference region of the BVG-CCCM. For example, a collocated luma block 1002 can be collocated with a current block 1000. The current block 1000 can be a chroma block corresponding to the collocated luma block 1002. Video encoder 200 or video decoder 300 can use the block vector 1004 of the collocated luma block 1002 coded in IBC or intraTMP mode to determine the reference region (e.g., reference chroma 1010) used to compute the CCCM parameters. For example, video encoder 200 or video decoder 300 can determine the block vector 1008 corresponding to the block vector 1004 to determine the reference region (e.g., reference chroma 1010). Then, the reference region in luma (e.g., reference luma 1006) and the corresponding region in chroma channels (e.g., reference chroma 1010) can be used to compute the CCCM parameters. When performing the CCCM prediction, the prediction uses the computed model parameters and the collocated luma samples.
[0153] Similar to the direct block vector (DBV) mode in ECM-8.0, five positions in the co-located luma block region (e.g., of co-located luma block 1002) are scanned to determine the block vector to be used in the BVG-CCCM technique. The use of the BVG-CCCM mode is signaled with a PU-level flag that is CABAC coded. For example, video encoder 200 can signal the use of the BVG-CCCM mode to video decoder 300 via the BVG-CCCM flag at the PU level. The BVG-CCCM flag is signaled if the co-located luma block 1002 is coded with IBC or intra TMP mode and the cross-component index is LM_CHROMA IDX or MM_LM_CHROMA IDX.
[0154] In ECM, the DBV mode is only enabled when the block partitioning uses dual tree (e.g., chroma has separate partition tree, which can be different from the corresponding luma block partition tree). In the case of single tree block partitioning, when the luma block is coded using Intra TMP, the chroma coding mode can be chosen from regular intra prediction modes and cross component prediction modes. For example, when the corresponding luma block is coded using Intra TMP, video encoder 200 or video decoder 300 can choose between regular intra prediction modes and cross component prediction modes when coding the chroma block with single tree partitioning. However, in such cases, the video coder (e.g., video encoder 200 or video decoder 300) does not have the option of intra block copy (IBC) mode. The cross component prediction modes include CCLM, MMLM, CCCM, GLM, and other modes that predict chroma samples from luma samples.
[0155] According to the techniques of this disclosure, intra chroma copy modes are enabled in single tree block partitioning, such as the DBV mode. For example, intra chroma copy modes in single tree block partitioning, in one example the DBV mode, are enabled. In one example, such modes can be enabled when the corresponding luma block has a block vector, which can occur when, for example, the corresponding luma block is coded using Intra TMP or IBC. Video encoder 200 or video decoder 300 can utilize intra chroma copy modes, such as the DBV mode, while using single tree block partitioning.
[0156] In one example, the DBV mode is indicated by the same flag as in the dual tree partitioning case and is signaled as a separate mode. A separate context from the dual tree DBV case can be used to signal the DBV flag in the single tree case. In one example, the context used to signal the DBV flag can be shared for the single tree and dual tree cases. Video encoder 200 can signal the flag and video decoder 300 can parse the flag to determine whether to use the DBV mode.
[0157] In another example, when chroma coding utilizes DM mode, DBV mode is implicitly applied for single tree partition case. In some examples, video encoder 200 can signal a flag to video decoder 300 to indicate that chroma is coded with DM mode. If the luma block has a block vector, e.g., when the block is coded by IntraTMP or IBC, DM mode indicates DBV mode, otherwise, DM mode indicates regular intra prediction mode. In such cases, video encoder 200 can not signal DBV flag and video decoder 300 can implicitly determine whether to apply DBV mode instead of parsing DBV flag.
[0158] In yet another example, if a luma block is coded by IntraTMP or IBC mode in single tree partition, chroma mode coding is forced to DBV mode. For example, if video encoder 200 or video decoder 300 codes a luma block using IntraTMP mode or IBC mode with single tree partition, video encoder 200 or video decoder 300 will code the corresponding chroma block using DBV mode. Therefore, syntax elements to indicate other modes (regular intra prediction mode and cross component prediction mode) will not be signaled in the bitstream as those syntax elements are not needed. In such examples, video encoder 200 can not signal syntax elements to indicate other modes and video decoder 300 can not parse such elements.
[0159] In yet another example, if a luma block has a block vector, e.g., if the block is coded by IntraTMP or IBC mode in single tree partition, chroma mode coding is forced to one of DBV mode or cross component prediction mode and regular intra prediction mode can be disabled in such cases. Therefore, corresponding syntax elements to indicate regular intra prediction mode can be saved so that they can not be signaled in the bitstream. For example, when a luma block in single tree partition has a block vector, video encoder 200 does not signal corresponding syntax elements to indicate regular intra prediction mode and video decoder 300 can not parse corresponding syntax elements.
[0160] In yet another example, if a luma block has a block vector, e.g., if the block is coded by IntraTMP or IBC mode in single tree partition, whether chroma mode coding is forced to DBV mode is controlled by a high level syntax (e.g., SPS level flag, picture header or slice header). For example, video encoder 200 can signal a syntax element with high level syntax to indicate whether chroma mode coding is forced to DBV mode. Video decoder 300 can parse the syntax element to determine whether chroma mode coding is forced to DBV mode for such blocks for which the syntax element applies.
[0161] Intra TMP fusion modes in which the luma intra predictor is generated by blending multiple reference blocks, such reference blocks being from multiple BVs of an intra template matching process, are disclosed in Zhang, et al., “Non-EE2: Intra Template-Matching Prediction Fusion,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 29th Meeting: by Teleconference, 11-20 Jan. 2023, JVET-AC0069; Huo, et al., “Non-EE2: A Fusion method of Intra Template Matching Prediction (Intra TMP),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 29th Meeting: by Teleconference, 11-20 Jan. 2023, JVET-AC0110; Zhang, et al., “EE2-1.11: Intra Template-Matching Prediction Fusion,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 30th Meeting: Antalya, TR, 21-28 Apr. 2023, JVET-AD0072; and Huo, et al., “EE2-1.16: A Fusion method of Intra Template Matching Prediction (Intra TMP),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 30th Meeting: Antalya, TR, 21-28 Apr. 2023, JVET-AD0116.Huo, et al., “EE2-1.15a: Intra template matching (Intra TMP) based on linear filter model,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 30th Meeting: Antalya, TR, 21-28 Apr. 2023, JVET-AD0112, discloses applying linear filtering to luma Intra TMP predictors, where a 6-tap linear filter includes 5 spatial luma samples in the reference block and a bias term. The filter coefficients can be derived by minimizing the MSE on samples between the reference template (obtained by block vectors) and the current template. In Li, et al., “EE2-1.12: Intra TMP with sub-pel precision,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 30th Meeting: Antalya, TR, 21-28 Apr. 2023, JVET-AD0125, sub-pixel precision Intra TMP is disclosed, supporting three sub-pixel precisions, including half-pixel, quarter-pixel, and three quarter-pixels, with eight directions around the integer-pixel position. In Kidani, et al., “Non-EE2: Bi-predictive IBC for natural and screen content,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 30th Meeting: Antalya, TR, 21-28 Apr. 2023, JVET-AD0134, bi-predictive IBC is disclosed. However, the behavior of chroma DBV mode (which borrows block vectors from some co-located luma block) is undefined. This disclosure also describes different techniques for chroma DBV mode.
[0162] In one example, when a luma block uses Intra TMP fusion, a video encoder 200 or video decoder 300 that uses chroma DBV mode also generates a predictor by blending multiple reference blocks obtained from the same set of BVs used in luma Intra TMP fusion.
[0163] In another example, when the luma block uses linear filtering, the same filter is applied to the chroma DBV predictor. For example, video encoder 200 or video decoder 300 can apply the same filter for linear filtering to the chroma DBV predictor.
[0164] In yet another example, when the luma block uses linear filtering, a linear filter is applied to the chroma DBV predictor. The linear filter is derived by minimizing the MSE on the chroma samples between the reference template identified by the block vector and the current template. For example, video encoder 200 or video decoder 300 can derive a linear filter by minimizing the MSE on the chroma samples between the reference template identified by the block vector and the current template and applying the derived linear filter to the chroma BVD predictor. In yet another example, if the luma block has multiple block vectors (e.g., IBC for bi-prediction and Intra TMP fusion), only one block vector is used for chroma block copy. Thus, no weighted prediction or fusion is applied in DBV mode to limit the computational complexity of the mode. For example, video encoder 200 or video decoder 300 can not apply weighted prediction or fusion in DBV mode. In one example, the selected block vector is one stored in List 0. In another example, the selected block vector is the first block vector in luma Intra TMP (with the lowest template matching cost).
[0165] In yet another example, if a filtering process is applied to the corresponding luma block, the filtering process is not applied to the chroma coding to reduce complexity. For example, if a filtering technique is applied to the corresponding luma block, video encoder 200 or video decoder 300 can not apply the filtering technique to the chroma block.
[0166] In yet another example, if the luma block is coded using sub-pixel accuracy Intra TMP mode, the block copy for the chroma block skips the sub-pixel parameters, e.g., only the integer part of the luma BV derived by Intra TMP is used. For example, if the corresponding luma block is coded using sub-pixel accuracy Intra TMP mode, video encoder 200 or video decoder 300 can skip the sub-pixel parameters for the chroma block copy and only use the integer part of the luma BV derived by Intra TMP.
[0167] In yet another example, if a sub-pixel precision Intra TMP mode is used to code a luma block, then the sub-pixel parameters for block copy of chroma blocks are skipped if the block vector points to an invalid region (e.g., a region outside the picture boundary, a region that has not been coded, etc.). Then only the integer part of the luma BV derived by Intra TMP is used. For example, if a block vector points to an invalid region, then if a sub-pixel precision Intra TMP mode is used to code a corresponding luma block, video encoder 200 or video decoder 300 can skip the sub-pixel parameters for chroma block copy.
[0168] Figure 11 FIG. 11 is a flowchart illustrating an example chroma component coding technique in accordance with one or more aspects of the disclosure. Video encoder 200 or video decoder 300 can determine that single tree block partitioning is used on a block of video data (1100). For example, video encoder 200 can determine that single tree block partitioning is utilized on a block of video data. For example, video decoder 300 can determine that single tree block partitioning is used on the block because video encoder 200 used single tree block partitioning on the block.
[0169] Video encoder 200 or video decoder 300 can apply an intra chroma copy mode to a chroma block of the block (1102). For example, video encoder 200 or video decoder 300 can use an intra chroma copy mode on the chroma block. The intra chroma copy mode can be a block copy mode similar to intra block copy. For example, the intra chroma copy mode can be a DBV mode. For example, when applying the intra chroma copy mode, video encoder 200 or video decoder 300 can use a luma block vector from a corresponding luma block of the block of video data for the chroma block.
[0170] Video encoder 200 or video decoder 300 can encode or decode the chroma block based on the single tree block partitioning and the intra chroma copy mode (1104). For example, video encoder 200 can encode the chroma block using the single tree partitioning and the intra chroma copy mode. Video decoder 300 can decode the chroma block using the single tree partitioning and the intra chroma copy mode.
[0171] In some examples, the intra chroma copy mode is a DBV mode. In some examples, prior to applying the intra chroma copy mode to the chroma block, video encoder 200 or video decoder 300 can determine that a luma block of the block has a block vector that is collocated with the chroma block. In some examples, applying the intra chroma copy mode to the chroma block is based at least in part on the luma block having the block vector. In some examples, video encoder 200 can signal or video decoder 300 can parse a high-level syntax element that indicates whether to use the intra chroma copy mode. As used herein, parsing includes determining a value of a syntax element. In such cases, the value of the syntax element indicates whether to use the intra chroma copy mode.
[0172] In some examples, as part of determining that a luma block has a block vector, video encoder 200 or video decoder 300 can determine whether the luma block is decoded using Intra-Template Matching Prediction (IntraTMP) mode or Intra-Block Copy (IBC) mode. In some examples, video encoder 200 can signal or video decoder 300 can parse a DBV mode flag indicating that DBV mode is being used.
[0173] In some examples, before applying the intra-frame chroma copy mode to the chroma block, the video encoder 200 or video decoder 300 may determine that the chroma decoding mode of the chroma block includes a chroma derived mode (DM) mode. In such examples, applying the intra-frame chroma copy mode to the chroma block is based at least in part on a chroma decoding mode that includes the DM mode.
[0174] In some examples, the block is the first block, and the chroma block is the first chroma block. In some examples, the video encoder 200 or video decoder 300 may determine to use single-tree block segmentation on the second block of video data. In some examples, the video encoder 200 or video decoder 300 may determine that the luma block of the second block has a block vector, which co-addresses with the second chroma block of the second block. In some examples, based on the use of single-tree block segmentation on the second block and the luma block having a block vector, the video encoder 200 or video decoder 300 may disable the regular intra-frame prediction mode. In some examples, the video encoder 200 or video decoder 300 may apply either a DBV mode or a cross-component prediction mode to the second chroma block. The cross-component prediction mode may include a mode for predicting chroma samples based on corresponding luma samples. For example, the corresponding luma sample may be a luma sample from the luma block that corresponds to a chroma sample from the second chroma block.
[0175] In some examples, video encoder 200 or video decoder 300 may determine to encode or decode luma blocks of video data using an IntraTMP fusion mode. In some examples, video encoder 200 or video decoder 300 may, based on the determination of the luma blocks to be encoded or decoded using the IntraTMP fusion mode, blend the same set of block vectors for the corresponding chroma blocks of the video data in a direct block vector mode to generate a blended block vector. In some examples, video encoder 200 or video decoder 300 may encode or decode the corresponding chroma blocks based on the blended block vector.
[0176] In some examples, the video encoder 200 or the video decoder 300 may determine a linear filter to be applied to the luminance blocks of the video data. In some examples, the video encoder 200 or the video decoder 300 may apply a linear filter to the direct block vector of the corresponding chroma block.
[0177] In some examples, the video encoder 200 or the video decoder 300 may determine to apply a first linear filter to the luminance block of the video data. In some examples, the video encoder 200 or the video decoder 300 may determine a second linear filter to apply to the direct block vector of the corresponding chroma block. In such examples, determining the second linear filter includes determining the minimum mean square error of the chroma samples between the reference template identified by the direct block vector and the current template.
[0178] In some examples, the video encoder 200 or video decoder 300 may determine which luminance blocks of video data to apply a filter to. In such examples, the video encoder 200 or video decoder 300 may, based on the determination of which luminance blocks to apply a filter to, avoid applying a filter to the corresponding chroma blocks of video data.
[0179] In some examples, the video encoder 200 or the video decoder 300 may determine that a luma block of video data has multiple block vectors. In some examples, the video encoder 200 or the video decoder 300 may select a single block vector for decoding the corresponding chroma block based on the fact that a luma block has multiple block vectors. In some examples, the single block vector includes block vectors stored in list 0 or block vectors used for the luma block with the lowest template matching cost.
[0180] In some examples, the video encoder 200 or video decoder 300 may determine to use a subpixel precision IntraTMP mode to encode or decode the luminance blocks of the video data. In some examples, the video encoder 200 or video decoder 300 may, based on the determination to use a subpixel precision IntraTMP mode to encode or decode the luminance blocks, only use the integer part of the luminance block vector derived by the subpixel precision IntraTMP mode for the corresponding chroma blocks of the video data.
[0181] In some examples, video encoder 200 or video decoder 300 may determine to encode or decode luma blocks of video data using a subpixel precision intra template matching prediction (IntraTMP) mode. In some examples, video encoder 200 or video decoder 300 may determine that the block vector derived by the IntraTMP mode points to an invalid region. In some examples, based on the determination that luma blocks are encoded or decoded using a subpixel precision IntraTMP mode and that the block vector points to an invalid region, video encoder 200 or video decoder 300 may use only the integer portion of the luma block vector derived by the IntraTMP mode for the corresponding chroma blocks of the video data.
[0182] Figure 12 This is a block diagram illustrating an example video encoder 200 that can perform the techniques of this disclosure. Figure 12 This disclosure is provided for illustrative purposes and should not be construed as a limitation on the techniques extensively illustrated and described herein. For illustrative purposes, this disclosure describes the video encoder 200 in accordance with the techniques of VVC and HEVC. However, the techniques of this disclosure can be performed by video encoding devices configured for other video decoding standards and video decoding formats, such as AV1 and subsequent formats of AV1 video decoding.
[0183] exist Figure 12 In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy coding unit 220 may be implemented in one or more processors or in processing circuitry. For example, the units of the video encoder 200 may be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0184] Video data storage 230 can store video data to be encoded by components of video encoder 200. Video encoder 200 can receive data from, for example, video source 104 (…). Figure 1The video data memory 230 receives video data stored in the video data memory 230. The DPB 218 can act as a reference picture memory, storing reference video data for use when the video encoder 200 predicts subsequent video data. The video data memory 230 and DPB 218 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 can be on-chip (as illustrated) with other components of the video encoder 200, or off-chip relative to those components.
[0185] In this disclosure, references to video data memory 230 should not be construed as limited to memory inside video encoder 200 (unless specifically described) or memory outside video encoder 200 (unless specifically described). Rather, references to video data memory 230 should be understood as a reference memory that stores video data received by video encoder 200 for encoding (e.g., video data for the current block to be encoded). Figure 1 The memory 106 can also provide temporary storage for the outputs from various units of the video encoder 200.
[0186] Examples Figure 12 Various units help understand the operations performed by the video encoder 200. Units can be implemented as fixed-function circuits, programmable circuits, or combinations thereof. Fixed-function circuits are circuits that provide specific functionality and are pre-defined for the operations that can be performed. Programmable circuits are circuits that can be programmed to perform various tasks and provide flexible functionality for the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more units in the unit may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units in the unit may be integrated circuits.
[0187] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core, all formed by programmable circuitry. In an example where the operation of the video encoder 200 is performed using software executed by programmable circuitry, memory 106 ( Figure 1The video encoder 200 may store instructions (e.g., target code) of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.
[0188] The video data storage unit 230 is configured to store received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be raw video data to be encoded.
[0189] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units for performing video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-frame block copying unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.
[0190] Mode selection unit 202 typically coordinates multiple coding channels to test combinations of coding parameters and the resulting rate-distortion values for such combinations. Coding parameters may include the CTU-CU partitioning, the prediction mode for the CU, the transformation type of the residual data for the CU, the quantization parameters of the residual data for the CU, etc. Mode selection unit 202 can ultimately select a combination of coding parameters that has a better rate-distortion value compared to other tested combinations.
[0191] The video encoder 200 can divide images retrieved from the video data storage 230 into a series of CTUs, and encapsulate one or more CTUs within slices. The mode selection unit 202 can divide the image's CTUs according to the tree structure described above (such as an MTT structure, a QTBT structure, a superblock structure, or the quadtree structure described above). As described above, the video encoder 200 can form one or more CUs by dividing CTUs according to a tree structure. Such CUs are also commonly referred to as "video blocks" or "blocks".
[0192] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate prediction blocks for the current block (e.g., the current CU, or, in HEVC, the overlapping portion of PU and TU). To perform inter-frame prediction for the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in DPB 218). Specifically, motion estimation unit 222 may calculate values representing the similarity between a potential reference block and the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. Motion estimation unit 222 may typically perform these calculations using sample-by-sample differences between the current block and the reference blocks under consideration. Motion estimation unit 222 may identify reference blocks with the lowest values produced by these calculations to indicate the reference block that best matches the current block.
[0193] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the location of a reference block in a reference image relative to the location of the current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate a prediction block. For example, motion compensation unit 224 can use the motion vectors to retrieve data for the reference block. As another example, where the motion vectors have fractional sample precision, motion compensation unit 224 can interpolate the values of the prediction block according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by corresponding motion vectors and combine the retrieved data, for example, by per-sample averaging or weighted averaging.
[0194] When operating according to the AV1 video decoding format, the motion estimation unit 222 and the motion compensation unit 224 can be configured to encode the decoded blocks of video data (e.g., both luma and chroma decoded blocks) using translational motion compensation, affine motion compensation, overlap block motion compensation (OBMC), and / or composite inter-intra-frame prediction.
[0195] As another example, for intra-prediction or intra-prediction decoding, intra-prediction unit 226 may generate a prediction block from samples adjacent to the current block. For example, for directional mode, intra-prediction unit 226 may typically mathematically combine the values of adjacent samples and fill these calculated values across the current block in a defined direction to produce a prediction block. As another example, for DC mode, intra-prediction unit 226 may calculate the average of the adjacent samples of the current block and generate a prediction block to include the resulting average for each sample of the prediction block.
[0196] When operating according to the AV1 video decoding format, the intra-frame prediction unit 226 can be configured to encode decoded blocks of video data (e.g., both luma and chroma decoded blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, luma-chroma (CFL) prediction, intra-block copying (IBC), and / or palette modes. The mode selection unit 202 may include additional functional units for performing video prediction based on other prediction modes.
[0197] Mode selection unit 202 provides a prediction block to residual generation unit 204. Residual generation unit 204 receives an uncoded raw version of the current block from video data memory 230 and a prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, residual generation unit 204 may also determine the differences between sample values in the residual block to generate the residual block using residual differential pulse decoding modulation (RDPCM). In some examples, residual generation unit 204 may be formed using one or more subtractor circuits performing binary subtraction.
[0198] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As noted above, the size of a CU can refer to the size of the luma decoding block of the CU, while the size of a PU can refer to the size of the luma prediction unit of the PU. Assuming a particular CU size is 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra-frame prediction, and symmetric PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, or similar for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-frame prediction.
[0199] In an example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luminance decoding block and a corresponding chrominance decoding block. As mentioned above, the size of the CU can refer to the size of the luminance decoding block of the CU. The video encoder 200 and the video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.
[0200] For other video decoding techniques, such as intra-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding, as some examples, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples (such as palette mode decoding), mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block is reconstructed based on a selected palette. In such modes, mode selection unit 202 may provide these syntax elements to entropy coding unit 220 for encoding.
[0201] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.
[0202] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a discrete cosine transform (DCT), direction transform, Karhunen-Loeve transform (KLT), or conceptually similar transformations to the residual block. In some examples, transform processing unit 206 may perform multiple transformations on the residual block, such as primary and secondary transformations (e.g., rotation transformations). In some examples, transform processing unit 206 does not apply any transformations to the residual block.
[0203] When operating according to AV1, transform processing unit 206 may apply one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a combination of horizontal / vertical transforms, which may include the Discrete Cosine Transform (DCT), the Asymmetric Discrete Sine Transform (ADST), the Reversed ADST (e.g., ADST in reverse order), and the Identity Transform (IDTX). When using the Identity Transform, the transform is skipped in either the vertical or horizontal direction. In some examples, transform processing may be skipped entirely.
[0204] Quantization unit 208 quantizes the transform coefficients in a transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 quantizes the transform coefficients of the transform coefficient block according to the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may cause information loss, and therefore, the quantized transform coefficients may have lower accuracy compared to the original transform coefficients produced by transform processing unit 206.
[0205] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct the residual block based on the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block (although potentially with some degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 can add samples of the reconstructed residual block to corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.
[0206] Filter unit 216 may perform one or more filtering operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce block artifacts along the edges of the CU. In some examples, the operation of filter unit 216 may be skipped.
[0207] When operating according to AV1, filter unit 216 may perform one or more filtering operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce block artifacts along the edges of the CU. In other examples, filter unit 216 may apply a constrained direction enhancement filter (CDEF) after deblocking and may include the application of a non-separable, nonlinear, low-pass directional filter based on the estimated edge direction. Filter unit 216 may also include a loop recovery filter applied after CDEF and may include a separable symmetric normalized Wiener filter or a dual-guided filter.
[0208] The video encoder 200 stores reconstructed blocks in the DPB 218. For example, in an example where the filter unit 216 is not operated, the reconstruction unit 214 may store reconstructed blocks in the DPB 218. In an example where the filter unit 216 is operated, the filter unit 216 may store filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve reference images formed by the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction for blocks of subsequent encoded images. Additionally, the intra-frame prediction unit 226 may use the reconstructed blocks of the current image in the DPB 218 to perform intra-frame prediction for other blocks in the current image.
[0209] Typically, entropy coding unit 220 can entropy-encode syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy-encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy-encode predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements (another example of video data) to generate entropy-coded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probability interval partitioning entropy (PIPE) decoding, exponential Golomb coding, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in a bypass mode where syntax elements are not entropy-encoded.
[0210] The video encoder 200 can output a bitstream that includes the entropy coding syntax elements required to reconstruct slices or blocks of images. Specifically, the entropy coding unit 220 can output a bitstream.
[0211] According to AV1, entropy coding unit 220 can be configured as a symbol-to-symbol adaptive multi-symbol arithmetic decoder. The syntax elements in AV1 consist of an N-element alphabet, and the context (e.g., a probability model) consists of a set of N probabilities. Entropy coding unit 220 can store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). Entropy coding unit 220 can perform recursive scaling to update the context using an update factor based on the alphabet size.
[0212] The operations described above are relative to blocks. This description should be understood as operations applied to luma decoding blocks and / or chroma decoding blocks. As described above, in some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the CU. In some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the PU.
[0213] In some examples, it is not necessary to repeat the operations performed relative to the luma decoding block for the chroma decoding block. As an example, the operations for identifying the MV and reference image of the luma decoding block do not need to be repeated for identifying the MV and reference image of the chroma block. Instead, the MV used for the luma decoding block can be scaled to determine the MV used for the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma decoding blocks.
[0214] Video encoder 200 represents an example of a device configured to encode video data, the device including a memory configured to store the video data, and one or more processing units implemented in circuitry and configured to determine, use single-tree block segmentation on a block of video data; apply an intra-frame chroma copy mode to the chroma block of the block; and encode the chroma block based on the single-tree block segmentation and the intra-frame chroma copy mode.
[0215] Figure 13 This is a block diagram illustrating an example video decoder 300 that can perform the techniques of this disclosure. Figure 13 This disclosure is provided for illustrative purposes and not for limiting the techniques extensively illustrated and described herein. For illustrative purposes, the video decoder 300 is described in accordance with VVC and HEVC techniques. However, the techniques of this disclosure can be implemented by video decoding devices configured for other video decoding standards.
[0216] exist Figure 13In the example, the video decoder 300 includes a decoded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a DPB 314. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or in processing circuitry. For example, the units of the video decoder 300 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0217] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-prediction unit 318. The prediction processing unit 304 may include additional units that perform predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copying unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.
[0218] When operating according to AV1, motion compensation unit 316 can be configured to decode video data blocks (e.g., both luma and chroma blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite inter-intra-frame prediction, as described above. Intra-frame prediction unit 318 can be configured to decode video data blocks (e.g., both luma and chroma blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, CFL, IBC, and / or palette mode, as described above.
[0219] CPB memory 320 can store video data to be decoded by components of video decoder 300, such as encoded video bitstreams. For example, it can be retrieved from computer-readable medium 110 ( Figure 1The video data stored in the CPB memory 320 is obtained. The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Furthermore, the CPB memory 320 may store video data other than the syntax elements of the decoded picture, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores a decoded picture that the video decoder 300 may output, and / or uses the decoded picture as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed from any of a variety of memory devices, such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices. The CPB memory 320 and the DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.
[0220] Additionally or alternatively, in some examples, the video decoder 300 may be located from the memory 120 ( Figure 1 The decoded video data can be retrieved from the memory. In other words, memory 120 can utilize CPB memory 320 to store data as discussed above. Similarly, when some or all of the functionality of video decoder 300 is implemented in software to be executed by the processing circuitry of video decoder 300, memory 120 can store instructions to be executed by video decoder 300.
[0221] Examples Figure 13 The various units shown help to understand the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to... Figure 12 Fixed-function circuits are circuits that provide specific functionality and are predefined for the operations they can perform. Programmable circuits are circuits that can be programmed to perform various tasks and provide flexible functionality for the operations they can perform. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is usually immutable. In some examples, one or more units in a cell can be different circuit blocks (fixed-function or programmable), and in some examples, one or more units in a cell can be integrated circuits.
[0222] The video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executed on programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.
[0223] The entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to reproduce the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.
[0224] Typically, the video decoder 300 reconstructs the image block by block. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").
[0225] Entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients defining the quantized transform coefficient block, as well as transform information (such as quantization parameters (QP) and / or transform mode indications). Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization to be applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bit-by-bit left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 can thereby form a transform coefficient block including the transform coefficients.
[0226] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the transform coefficient block.
[0227] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is an inter-frame prediction, the motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements may indicate the reference picture from which the reference block is to be retrieved in the DPB 314, and a motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can generally be configured according to the parameters relative to the motion compensation unit 224 ( Figure 12 The method described is essentially the same as the method used to perform the inter-frame prediction process.
[0228] As another example, when the prediction information syntax element indicates that the current block is intra-predictive, the intra-predictive unit 318 can generate a prediction block according to the intra-predictive mode indicated by the prediction information syntax element. Similarly, the intra-predictive unit 318 can generally be associated with the intra-predictive unit 226 ( Figure 12 The intra-prediction process is performed in a manner substantially similar to that described above. The intra-prediction unit 318 can retrieve data of neighboring samples of the current block from the DPB 314.
[0229] Reconstruction unit 310 can use the prediction block and the residual block to reconstruct the current block. For example, reconstruction unit 310 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block.
[0230] Filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce block artifacts along the edges of the reconstructed block. The operation of filter unit 312 is not necessarily performed in all examples.
[0231] The video decoder 300 can store reconstructed blocks in the DPB 314. For example, in an example where the operation of the filter unit 312 is not performed, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the operation of the filter unit 312 is performed, the filter unit 312 can store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 can provide reference information (such as samples of the current image for intra-frame prediction and samples of previously decoded images for subsequent motion compensation) to the prediction processing unit 304. Furthermore, the video decoder 300 can output decoded images (e.g., decoded video) from the DPB 314 for use in applications such as... Figure 1 The subsequent presentation on display devices such as display device 118.
[0232] In this manner, video decoder 300 represents an example of a video decoding device, the device including a memory configured to store video data; and one or more processing units implemented in circuitry and configured to encode video data, the one or more processing units including the memory configured to store video data; and one or more processing units implemented in circuitry and configured to determine single-tree block segmentation on a block of video data; apply an intra-frame chroma copy mode to the chroma block of the block; and decode the chroma block based on the single-tree block segmentation and the intra-frame chroma copy mode.
[0233] Figure 14This is a flowchart illustrating an example method for encoding the current block according to the technology of this disclosure. The current block may be or may include the current CU. Although this relates to video encoder 200 ( Figure 1 and Figure 12 This is a description, but it should be understood that other devices can be configured to perform the same actions. Figure 14 Similar to the method.
[0234] In this example, the video encoder 200 initially predicts the current block (400). For example, the video encoder 200 may form a prediction block for the current block. The video encoder 200 may then compute a residual block for the current block (402). To compute the residual block, the video encoder 200 may compute the difference between the unencoded original block for the current block and the prediction block. The video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (404). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (406). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (408). For example, the video encoder 200 may use CAVLC or CABAC to encode the transform coefficients. The video encoder 200 may then output the entropy-encoded data of the block (410).
[0235] Figure 15 This is a flowchart illustrating an example method for decoding a current block of video data according to the technology of this disclosure. The current block may be or may include the current CU. Although relative to video decoder 300 ( Figure 1 and Figure 13 This is a description, but it should be understood that other devices can be configured to perform similar actions. Figure 15 Similar to the method.
[0236] The video decoder 300 can receive entropy-coded data for the current block, such as entropy-coded prediction information and entropy-coded data for the transform coefficients of the residual block corresponding to the current block (500). The video decoder 300 can entropy decode the entropy-coded data to determine the prediction information for the current block and reproduce the transform coefficients of the residual block (502). The video decoder 300 can predict the current block, for example, using an intra-frame prediction mode or an inter-frame prediction mode indicated by the prediction information of the current block (504), to compute the prediction block of the current block. The video decoder 300 can then perform an inverse scan on the reproduced transform coefficients (506) to create a block of quantized transform coefficients. The video decoder 300 can then inverse quantize the transform coefficients and apply the inverse transform to the transform coefficients to produce the residual block (508). The video decoder 300 can finally decode the current block by combining the prediction block and the residual block (510).
[0237] The following numbered clauses illustrate one or more aspects of the devices and technologies described in this disclosure.
[0238] Clause 1A. A method for decoding video data, the method comprising: determining a block of the video data to be segmented using a single-tree block; applying an intra-frame chroma copy mode to a chroma block of the block; and decoding the chroma block based on the single-tree block segmentation and the intra-frame chroma copy mode.
[0239] Clause 2A. The method described in accordance with Clause 1A, wherein the intra-frame chroma copy mode is a direct block vector (DBV) mode.
[0240] Clause 3A. The method according to Clause 1A or Clause 2A, the method further comprising determining, before applying the intra-frame chroma copy mode to the chroma block, that the luma block of the block has a block vector, wherein applying the intra-frame chroma copy mode to the chroma block is at least in part based on the luma block of the block having the block vector.
[0241] Clause 4A. The method described in Clause 3A further includes a higher-order syntax element that signals or parses an indication of whether an intra-frame chroma copy mode is used.
[0242] Clause 5A. The method according to Clause 3A or Clause 4A, wherein determining that the luma block has a block vector includes determining that the luma block is decoded using IntraTemplate Matching Prediction (IntraTMP) mode or IntraBlock Copy (IBC) mode.
[0243] Clause 6A. The method according to any one of Clauses 2A to 5A, the method further comprising signaling or parsing a DBV mode flag indicating that a DBV mode is being used.
[0244] Clause 7A. The method according to any one of Clauses 2A to 5A, the method further comprising determining, before applying the intra-frame chroma copy mode to the chroma block, that the chroma decoding mode includes a chroma-derived mode (DM) mode, wherein applying the intra-frame chroma copy mode to the chroma block is based at least in part on the chroma decoding mode including the DM mode.
[0245] Clause 8A. The method according to Clause 2A, the method further comprising, before applying the intra-chroma copy mode to the chroma block, determining to decode the luma block of the block using an intra-template matching prediction (IntraTMP) mode or an intra-block copy (IBC) mode, wherein applying the intra-chroma copy mode to the chroma block is based on decoding the luma block using the IntraTMP mode or the IBC mode.
[0246] Clause 9A. The method according to Clause 1A or Clause 2A, wherein the block is a first block and the chroma block is a first chroma block, the method further comprising: determining that a single-tree block segmentation is applied to a second block of the video data; determining that a luma block of the second block has a block vector; disabling a regular intra-frame prediction mode based on the application of single-tree block segmentation to the second block and the luma block having a block vector; and applying either a DBV mode or a cross-component prediction mode to the second chroma block of the second block.
[0247] Clause 10A. The method according to Clause 9A, wherein the cross-component prediction mode includes a mode for predicting chromaticity samples based on corresponding luminance samples.
[0248] Clause 11A. A method for decoding video data, the method comprising: determining a luma block of the video data to be decoded using an IntraTMP fusion mode; based on the determination to decode the luma block using the IntraTMP fusion mode, mixing the same set of block vectors for the luma block in a direct block vector mode for a corresponding chroma block of the video data; and decoding the corresponding chroma block based on the mixed block vectors.
[0249] Clause 12A. The method according to any one of Clauses 1A to 11A, the method further comprising: determining a linear filter for a luminance block of the video data; and applying the linear filter to a direct block vector of a corresponding chroma block.
[0250] Clause 13A. The method according to any one of Clauses 1A to 11A, the method further comprising: determining a first linear filter to be applied to a luminance block of the video data; and determining a second linear filter to be applied to a direct block vector of the corresponding chroma block, wherein determining the second linear filter includes determining the minimum mean square error of chroma samples between a reference template identified by the direct block vector and a current template.
[0251] Clause 14A. The method according to any one of Clauses 1A to 11A, the method further comprising: determining to apply a filter to a luminance block of the video data; and, based on the determination to apply the filter to the luminance block, avoiding applying the filter to a corresponding chroma block of the video data.
[0252] Clause 15A. The method according to any one of Clauses 1A to 14A, the method further comprising: determining that the luminance block of the video data has a plurality of block vectors; and selecting a single block vector for a corresponding chrominance block based on the luminance block having the plurality of block vectors.
[0253] Clause 16A. The method according to Clause 15A, wherein the single block vector includes a block vector stored in list 0 or a block vector for the luma block having the lowest template matching cost.
[0254] Clause 17A. The method according to any one of Clauses 1A to 16A, the method further comprising: determining a luminance block of the video data to be decoded using a subpixel precision intra template matching prediction (IntraTMP) mode; and, based on the determination to decode the luminance block using the subpixel precision IntraTMP mode, using only the integer portion of the block vector derived by the IntraTMP mode for the corresponding chroma block of the video data.
[0255] Clause 18A. The method according to any one of Clauses 1A to 16A, the method further comprising: determining a luma block of the video data to be decoded using a subpixel precision intra template matching prediction (IntraTMP) mode; determining that a block vector derived by the IntraTMP mode points to an invalid region; and, based on the determination that the luma block is decoded using the subpixel precision IntraTMP mode and that the block vector points to an invalid region, using only the integer portion of the block vector derived by the IntraTMP mode for the corresponding chroma block of the video data.
[0256] Clause 19A. The method according to any one of Clauses 1A to 18A, wherein decoding includes decoding.
[0257] Clause 20A. The method according to any one of Clauses 1A to 19A, wherein decoding includes encoding.
[0258] Clause 21A. An apparatus for decoding video data, the apparatus comprising one or more components for performing the method according to any one of Clauses 1A to 20A.
[0259] Clause 22A. The device pursuant to Clause 21A, wherein the one or more components include one or more processors implemented in a circuit.
[0260] Clause 23A. The device according to any one of Clauses 21A or 22A, the device further includes a memory for storing the video data.
[0261] Clause 24A. The device according to any one of Clauses 21A to 23A, the device further comprising a display configured to display decoded video data.
[0262] Clause 25A. The device pursuant to any one of Clauses 21A to 24A, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.
[0263] Clause 26A. The device pursuant to any one of Clauses 21A to 25A, wherein the device includes a video decoder.
[0264] Clause 27A. The device pursuant to any one of Clauses 21A to 26A, wherein the device includes a video encoder.
[0265] Clause 28A. A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the method according to any one of Clauses 1A to 20A.
[0266] Clause 1B. A method for encoding or decoding video data, the method comprising: determining a block of the video data to be segmented using a single-tree block; applying an intra-frame chroma copy mode to a chroma block of the block; and encoding or decoding the chroma block based on the single-tree block segmentation and the intra-frame chroma copy mode.
[0267] Clause 2B. The method described in accordance with Clause 1B, wherein the intra-frame chroma copy mode is a direct block vector (DBV) mode.
[0268] Clause 3B. The method according to Clause 1B or Clause 2B, the method further comprising, before applying the intra-frame chroma copy mode to the chroma block, determining that the luma block of the block has a block vector, the luma block being co-located with the chroma block, wherein applying the intra-frame chroma copy mode to the chroma block is at least in part based on the luma block having the block vector.
[0269] Clause 4B. The method described in Clause 3B further includes a higher-order syntax element that signals or parses an indication of whether an intra-frame chroma copy mode is used.
[0270] Clause 5B. The method according to Clause 3B or Clause 4B, wherein determining that the luma block has a block vector includes determining that the luma block is decoded using IntraTemplate Matching Prediction (IntraTMP) mode or IntraBlock Copy (IBC) mode.
[0271] Clause 6B. The method according to any one of Clauses 2B to 5B, the method further comprising signaling or resolving a DBV mode flag indicating that a DBV mode is being used.
[0272] Clause 7B. The method according to any one of Clauses 2B to 5B, the method further comprising, before applying the intra-frame chroma copy mode to the chroma block, determining that the chroma decoding mode for the chroma block includes a chroma-derived mode (DM) mode, wherein applying the intra-frame chroma copy mode to the chroma block is based at least in part on the chroma decoding mode including the DM mode.
[0273] Clause 8B. The method according to Clause 1B or Clause 2B, wherein the block is a first block and the chroma block is a first chroma block, the method further comprising: determining that a single-tree block segmentation is applied to a second block of the video data; determining that a luma block of the second block has a block vector, the luma block being co-located with a second chroma block of the second block; disabling a conventional intra-frame prediction mode based on the application of single-tree block segmentation to the second block and the luma block having a block vector; and applying one of a DBV mode or a cross-component prediction mode to the second chroma block, wherein the cross-component prediction mode includes a mode for predicting chroma samples based on corresponding luma samples.
[0274] Clause 9B. The method according to Clause 1B or Clause 2B, wherein the method further comprises: determining a luma block of the video data to be encoded or decoded using an IntraTMP fusion mode; based on the determination to encode or decode the luma block using the IntraTMP fusion mode, mixing the same set of block vectors used for the luma block in a direct block vector mode for the corresponding chroma block of the video data to generate a mixed block vector; and encoding or decoding the corresponding chroma block based on the mixed block vector.
[0275] Clause 10B. The method according to any one of Clauses 1B to 9B, the method further comprising: determining a linear filter for a luminance block of the video data; and applying the linear filter to a direct block vector of a corresponding chroma block.
[0276] Clause 11B. The method according to any one of Clauses 1B to 9B, the method further comprising: determining a first linear filter to be applied to a luminance block of the video data; and determining a second linear filter to be applied to a direct block vector of the corresponding chroma block, wherein determining the second linear filter includes determining the minimum mean square error of chroma samples between a reference template identified by the direct block vector and a current template.
[0277] Clause 12B. The method according to any one of Clauses 1B to 9B, the method further comprising: determining to apply a filter to a luminance block of the video data; and, based on the determination to apply the filter to the luminance block, avoiding applying the filter to a corresponding chroma block of the video data.
[0278] Clause 13B. The method according to any one of Clauses 1B to 12B, the method further comprising: determining that the luminance block of the video data has a plurality of block vectors; and selecting a single block vector for decoding the corresponding chrominance block based on the luminance block having the plurality of block vectors.
[0279] Clause 14B. The method according to Clause 13B, wherein the single block vector includes a block vector stored in list 0 or a block vector for the luma block having the lowest template matching cost.
[0280] Clause 15B. The method according to any one of Clauses 1B to 14B, the method further comprising: determining a luminance block of the video data to be encoded or decoded using a subpixel precision intra template matching prediction (IntraTMP) mode; and, based on the determination to encode or decode the luminance block using the subpixel precision IntraTMP mode, using only the integer portion of the luminance block vector derived by the subpixel precision IntraTMP mode for the corresponding chroma block of the video data.
[0281] Clause 16B. The method according to any one of Clauses 1B to 14B, the method further comprising: determining a luma block of the video data to be encoded or decoded using a subpixel precision intra template matching prediction (IntraTMP) mode; determining that a block vector derived by the IntraTMP mode points to an invalid region; and, based on the determination that the luma block is encoded or decoded using the subpixel precision IntraTMP mode and that the block vector points to an invalid region, using only the integer portion of the luma block vector derived by the IntraTMP mode for the corresponding chroma block of the video data.
[0282] Clause 17B. An apparatus for encoding or decoding video data, the apparatus comprising: one or more memories configured to store the video data; and one or more processors communicating with the one or more memories and configured to: determine a single-tree block segmentation on a block of the video data; apply an intra-frame chroma copy mode to a chroma block of the block; and encode or decode the chroma block based on the single-tree block segmentation and the intra-frame chroma copy mode.
[0283] Clause 18B. An apparatus for encoding or decoding video data as described in Clause 17B, wherein the apparatus further comprises: a camera configured to capture the video data; and a video encoder wherein one or more processors are configured to encode the chroma blocks.
[0284] Clause 19B. An apparatus for encoding or decoding video data as described in Clause 17B or claim 18B, wherein the apparatus further comprises: a display configured to display the decoded video data; and a video decoder, wherein the one or more processors are configured to decode the chroma blocks.
[0285] Clause 20B. An apparatus for encoding or decoding video data, the apparatus comprising: means for determining single-tree block segmentation on a block of the video data; means for applying an intra-frame chroma copy mode to a chroma block of the block; and means for encoding or decoding the chroma block based on the single-tree block segmentation and the intra-frame chroma copy mode.
[0286] Clause 21B. A computer-readable storage medium encoded with instructions that, when executed, cause one or more programmable processors to determine to use single-tree block segmentation on a block of video data; apply an intra-frame chroma copy mode to the chroma blocks of the block; and encode or decode the chroma blocks based on the single-tree block segmentation and the intra-frame chroma copy mode.
[0287] It should be recognized that, based on the examples, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all actions or events described are necessary for implementing the technique). Furthermore, in some examples, actions or events may be performed concurrently (e.g., through multithreading, interrupt handling, or multiple processors) rather than sequentially.
[0288] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. Thus, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.
[0289] By way of example, and not limitation, such computer-readable storage media may include one or more of RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs use lasers to reproduce data optically. Combinations of these should also be included within the scope of computer-readable media.
[0290] Instructions can be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuit" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.
[0291] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but implementation by different hardware units is not necessarily required. Rather, as described above, various units may be combined in a codec hardware unit, or various units may be provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.
[0292] Various examples have been described. These and other examples are within the scope of the following claims.
Claims
1. A method of encoding or decoding video data, the method comprising: determining that single tree block partitioning is used on a block of the video data; applying an intra chroma copy mode to a chroma block of the block; and encoding or decoding the chroma block based on the single tree block partitioning and the intra chroma copy mode.
2. The method of claim 1, wherein the intra chroma copy mode is a direct block vector (DBV) mode.
3. The method of claim 1, the method further comprising, prior to applying the intra chroma copy mode to the chroma block, determining that a luma block of the block has a block vector, the luma block being collocated with the chroma block, wherein applying the intra chroma copy mode to the chroma block is based at least in part on the luma block having the block vector.
4. The method of claim 3, the method further comprising signaling or parsing a high level syntax element indicating whether an intra chroma copy mode is used.
5. The method of claim 3, wherein determining that the luma block has a block vector comprises determining that the luma block is coded using an intra template matching prediction (IntraTMP) mode or an intra block copy (IBC) mode.
6. The method of claim 2, the method further comprising signaling or parsing a DBV mode flag indicating that a DBV mode is being used.
7. The method of claim 2, the method further comprising, prior to applying the intra chroma copy mode to the chroma block, determining that a chroma coding mode for the chroma block includes a chroma derived mode (DM) mode, wherein applying the intra chroma copy mode to the chroma block is based at least in part on the chroma coding mode including the DM mode.
8. The method of claim 1, wherein the block is a first block and the chroma block is a first chroma block, the method further comprising: determining that single tree block partitioning is used on a second block of the video data; determining that a luma block of the second block has a block vector, the luma block being collocated with a second chroma block of the second block; based on the single tree block partitioning being used on the second block and the luma block having the block vector, disabling regular intra prediction modes; and applying one of a DBV mode or a cross component prediction mode to the second chroma block, wherein the cross component prediction mode includes a mode to predict chroma samples based on corresponding luma samples.
9. The method of claim 1, wherein method further comprises: determining that an intra template matching prediction (IntraTMP) fusion mode is used to encode or decode a luma block of the video data; based on the determination that the IntraTMP fusion mode is used to encode or decode the luma block, for a corresponding chroma block of the video data, blending a same set of block vectors used for the luma block in a direct block vector mode to generate a blended block vector; and encoding or decoding the corresponding chroma block based on the blended block vector.
10. The method of claim 1, the method further comprising: determining a linear filter applied to a luma block of the video data; and applying the linear filter to a direct block vector of a corresponding chroma block.
11. The method of claim 1, the method further comprising: determining to apply a first linear filter to a luma block of the video data; and determining a second linear filter applied to a direct block vector of a corresponding chroma block, wherein determining the second linear filter comprises determining a minimum mean squared error of chroma samples between a reference template and a current template identified by the direct block vector.
12. The method of claim 1, the method further comprising: determining to apply a filter to a luma block of the video data; and based on the determination to apply a filter to the luma block, refraining from applying a filter to a corresponding chroma block of the video data.
13. The method of claim 1, the method further comprising: determining that a luma block of the video data has multiple block vectors; and based on the luma block having the multiple block vectors, selecting a single block vector for coding a corresponding chroma block.
14. The method of claim 13, wherein the single block vector comprises a block vector stored in List 0 or a block vector for the luma block having a lowest template matching cost.
15. The method of claim 1, the method further comprising: determining that a luma block of the video data is encoded or decoded using a sub-pixel precision intra template matching prediction (Intra TMP) mode; and based on the determination that the luma block is encoded or decoded using the sub-pixel precision Intra TMP mode, using only an integer portion of a luma block vector derived by the sub-pixel precision Intra TMP mode for a corresponding chroma block of the video data.
16. The method of claim 1, the method further comprising: determining that a luma block of the video data is encoded or decoded using a sub-pixel precision intra template matching prediction (Intra TMP) mode; determining that a block vector derived by the Intra TMP mode points to an invalid region; and based on the determination that the luma block is encoded or decoded using the sub-pixel precision Intra TMP mode and that the block vector points to an invalid region, using only an integer portion of a luma block vector derived by the Intra TMP mode for a corresponding chroma block of the video data.
17. A device for encoding or decoding video data, the device comprising: one or more memories configured to store the video data; and one or more processors in communication with the one or more memories and configured to: determine to use single tree block partitioning on a block of the video data; apply an intra chroma copy mode to a chroma block of the block; and encode or decode the chroma block based on single tree block partitioning and the intra chroma copy mode.
18. The device for encoding or decoding video data of claim 17, wherein the device further comprises: a camera configured to capture the video data; and a video encoder, wherein the one or more processors are configured to encode the chroma block.
19. The apparatus for encoding or decoding video data according to claim 17, wherein the apparatus further comprises: a display configured to display decoded video data. and a video decoder, wherein the one or more processors are configured to decode the chroma block.
20. An apparatus for encoding or decoding video data, the apparatus comprising: means for determining that single tree block partitioning is used on a block of the video data; means for applying an intra chroma copy mode to chroma blocks of the block; and means for encoding or decoding the chroma blocks based on single tree block partitioning and the intra chroma copy mode.