Joint chrominance residual video coding

By using lossless Joint Chromatic Residual Decoding (JCCR) technology, a joint residual block is constructed and a lossless resJointC' is generated during decoding. This solves the problem that it is difficult to achieve lossless decoding in the joint chroma mode in the existing technology, improves video quality and reduces the amount of information sent.

CN114731418BActive Publication Date: 2026-02-24QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080081232.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-08
Filing Date
2020-12-09
Publication Date
2026-02-24
Estimated Expiration
2040-12-09

AI Technical Summary

Technical Problem

Existing video decoding standards struggle to combine lossless decoding modes when implementing joint chroma mode, resulting in lossy decoding and impacting video quality.

Method used

The lossless Joint Chromatic Residual Decoding (JCCR) technique is adopted. By constructing a joint residual block (resJointC), the amount of information sent during encoding is reduced, and a lossless resJointC' is generated during decoding to reconstruct the chroma block.

Benefits of technology

This achieves lossless decoding, improving video quality while reducing the amount of information transmitted and increasing video compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114731418B_ABST
    Figure CN114731418B_ABST
Patent Text Reader

Abstract

A video decoder can be configured to receive, for a first chroma component of a block of video data, information for first residual samples corresponding to a difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component, determine intermediate reconstructed samples based on the first residual samples, receive, for a second chroma component of the block of video data, information for second residual samples corresponding to a difference between a second chroma block of the second chroma component and the intermediate reconstructed samples, reconstruct the first chroma block based on the first residual samples and the first prediction block, reconstruct the second chroma block based on the second residual samples and the intermediate reconstructed samples, and output decoded video data including the reconstructed first chroma block and the reconstructed second chroma block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Application No. 17 / 115,544, filed December 8, 2020, which claims the benefit of U.S. Provisional Patent Application No. 62 / 945,750, filed December 9, 2019, the entire contents of each of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to video encoding and video decoding. Background Technology

[0003] Digital video capabilities can be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones (so-called "smartphones"), video conferencing equipment, video streaming devices, and more. Digital video devices implement video decoding technologies, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Decoding (AVC)), ITU-T H.265 / High Efficiency Video Decoding (HEVC), and extensions to such standards. By implementing such video decoding technologies, video devices can more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video decoding, video slices (e.g., video pictures or portions of video pictures) can be segmented into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in a slice of a picture that has been intra-decoded (I) are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in a slice of a picture that has been inter-decoded (P or B) can use spatial prediction relative to reference samples in adjacent blocks within the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005] In summary, this disclosure describes techniques for joint chroma residual decoding (JCCR). In JCCR, residual chroma blocks (e.g., the difference between a chroma block and its corresponding chroma prediction block) are decoded together (e.g., a residual block is decoded against the residual chroma block). In some examples, JCCR may result in lossy decoding, and therefore JCCR may be disabled when lossless decoding is enabled. However, by utilizing lossless decoding or at least less lossy decoding, video quality can be improved. This disclosure describes examples of techniques for JCCR with lossless or less lossy decoding, thereby allowing the benefits of less lossy decoding while achieving at least some of the benefits of JCCR (e.g., a reduction in the amount of residual information transmitted with the signal).

[0006] According to one example, a method for decoding video data includes: receiving information about a first residual sample for a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block in the first chroma component and a first prediction block of the first chroma component; determining intermediate reconstruction samples based on the first residual sample; receiving information about a second residual sample for a second chroma component of a block of video data, wherein the second residual sample corresponds to the difference between a second chroma block in the second chroma component and the intermediate reconstruction samples; reconstructing a first chroma block based on the first residual sample and the first prediction block; reconstructing a second chroma block based on the second residual sample and the intermediate reconstruction samples; and outputting decoded video data including the reconstructed first chroma block and the reconstructed second chroma block.

[0007] According to another example, a method includes: determining a first residual sample of a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; determining intermediate reconstructed samples based on the first residual sample and a second prediction block of a second chroma component; determining a second residual sample of a second chroma component of a block of video data, wherein the second residual sample corresponds to the difference between a second chroma block of the second chroma component and an intermediate reconstructed sample; and outputting syntax elements representing the first residual sample and the second residual sample in a bitstream of encoded video data.

[0008] According to another example, an apparatus for decoding video data includes: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: receive information about a first residual sample for a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; determine intermediate reconstructed samples based on the first residual sample; receive information about a second residual sample for a second chroma component of a block of video data, wherein the second residual sample corresponds to the difference between a second chroma block of the second chroma component and the intermediate reconstructed samples; reconstruct the first chroma block based on the first residual sample and the first prediction block; reconstruct the second chroma block based on the second residual sample and the intermediate reconstructed samples; and output decoded video data including the reconstructed first chroma block and the reconstructed second chroma block.

[0009] According to another example, an apparatus for encoding video data includes: a memory configured to store video data; and one or more processors implemented in circuitry and configured to: determine a first residual sample of a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; determine intermediate reconstructed samples based on the first residual sample and a second prediction block of a second chroma component; determine a second residual sample of a second chroma component of a block of video data, wherein the second residual sample corresponds to the difference between a second chroma block of the second chroma component and an intermediate reconstructed sample; and output syntax elements representing the first residual sample and the second residual sample in a bitstream of the encoded video data.

[0010] According to another example, a computer-readable storage medium stores instructions that, when executed by one or more processors, cause the one or more processors to: receive information about a first residual sample for a block of video data of a first chroma component, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; determine intermediate reconstructed samples based on the first residual sample; receive information about a second residual sample for a block of video data of a second chroma component, wherein the second residual sample corresponds to the difference between a second chroma block of the second chroma component and the intermediate reconstructed samples; reconstruct the first chroma block based on the first residual sample and the first prediction block; reconstruct the second chroma block based on the second residual sample and the intermediate reconstructed samples; and output decoded video data including the reconstructed first chroma block and the reconstructed second chroma block.

[0011] According to another example, a computer-readable storage medium stores instructions that, when executed by one or more processors, cause the one or more processors to: determine a first residual sample of a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; determine an intermediate reconstructed sample based on the first residual sample and a second prediction block of a second chroma component; determine a second residual sample of a second chroma component of a block of video data, wherein the second residual sample corresponds to the difference between a second chroma block of the second chroma component and an intermediate reconstructed sample; and output syntax elements representing the first residual sample and the second residual sample in a bitstream of encoded video data.

[0012] According to another example, an apparatus for decoding video data includes: a unit for receiving information of a first residual sample of a first chroma component for a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; a unit for determining intermediate reconstructed samples based on the first residual sample; a unit for receiving information of a second residual sample of a second chroma component for a block of video data, wherein the second residual sample corresponds to the difference between a second chroma block of the second chroma component and the intermediate reconstructed samples; a unit for reconstructing a first chroma block based on the first residual sample and the first prediction block; a unit for reconstructing a second chroma block based on the second residual sample and the intermediate reconstructed samples; and a unit for outputting decoded video data including the reconstructed first chroma block and the reconstructed second chroma block.

[0013] According to another example, an apparatus for encoding video data includes: a unit for determining a first residual sample of a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; a unit for determining intermediate reconstructed samples based on the first residual sample and a second prediction block of a second chroma component; a unit for determining a second residual sample of a second chroma component of a block of video data, wherein the second residual sample corresponds to the difference between a second chroma block of the second chroma component and an intermediate reconstructed sample; and a unit for outputting syntax elements representing the first residual sample and the second residual sample in a bitstream of encoded video data.

[0014] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. Attached Figure Description

[0015] Figure 1This is a block diagram illustrating an example video encoding and decoding system that can perform the techniques described in this disclosure.

[0016] Figure 2A and Figure 2B This is a conceptual diagram showing an example quadtree binary tree (QTBT) structure and its corresponding decoding tree unit (CTU).

[0017] Figure 3 This is a block diagram illustrating an example video encoder that can perform the techniques described in this disclosure.

[0018] Figure 4 This is a block diagram illustrating an example video decoder that can perform the techniques described in this disclosure.

[0019] Figure 5 This is a flowchart illustrating an example of encoding video data.

[0020] Figure 6 This is a flowchart illustrating an example of decoding video data.

[0021] Figure 7 This is a flowchart illustrating an example of encoding video data.

[0022] Figure 8 This is a flowchart illustrating an example of decoding video data. Detailed Implementation

[0023] Video decoding (e.g., video encoding and / or video decoding) typically involves predicting blocks of video data from already decoded blocks of video data in the same frame (e.g., intra-frame prediction) or from already decoded blocks of video data in different frames (e.g., inter-frame prediction). In some cases, the video encoder also computes residual data by comparing the predicted blocks with the original blocks. Therefore, the residual data represents the difference between the predicted and original blocks. To reduce the number of bits required to signal the residual data, the video encoder transforms and quantizes the residual data, and signals the transformed and quantized residual data in the encoded bitstream. Compression achieved through the transformation and quantization process can be lossy, meaning that the transformation and quantization process may introduce distortion into the decoded video data.

[0024] The video decoder decodes the residual data and adds it to the prediction block to produce a reconstructed video block that matches the original video block more closely than the individual prediction blocks. Due to the loss introduced by transforming and quantizing the residual data, the first reconstructed block may have distortion or artifacts. A common type of artifact or distortion is called block artifacts, where the boundaries of the blocks used to decode the video data are visible.

[0025] To further improve the quality of the decoded video, the video decoder can perform one or more filtering operations on the reconstructed video blocks. Examples of these filtering operations include deblocking filtering, Sample Adaptive Offset (SAO) filtering, and Adaptive Loop Filtering (ALF). The parameters of these filtering operations can be determined by the video encoder and explicitly signaled in the encoded video bitstream, or they can be implicitly determined by the video decoder without needing to be explicitly signaled in the encoded video bitstream.

[0026] As will be explained in more detail below, some video decoding standards support lossless decoding modes. In lossless decoding mode, the decoded video data matches the encoded video data without errors or distortion. One technique used to achieve lossless video decoding is to skip the quantization and inverse quantization processes described above. As will be explained in more detail below, video data is often decoded in luminance sample blocks with two corresponding chroma sample blocks. Video data can be decoded in a joint chroma mode (also known as joint chroma residual decoding (JCCR) mode), where a single chroma residual block is encoded for two corresponding chroma sample blocks. Existing video standards do not currently implement lossless decoding modes that can also be used in conjunction with joint chroma modes.

[0027] This disclosure describes techniques for implementing JCCR modes that also support lossless decoding. As described in this disclosure, for some decoding scenarios, implementing lossless decoding using JCCR modes can advantageously lead to improved compression compared to other available lossless decoding techniques.

[0028] In video decoding, for YUV (e.g., Y, Cb, Cr) formats, there may be two chroma components (e.g., Cb and Cr). For convenience, this paper describes an example with Cb and Cr blocks, but these techniques can be extended to other types of chroma blocks and other video formats (e.g., RGB). For video encoding, the video encoder determines the residual block (referred to as resCb) of the Cb block and the residual block (referred to as resCr) of the Cr block. ResCb and resCr can represent the difference between the corresponding chroma prediction block and the original chroma block. That is, resCb represents the difference between the original Cb block and the Cb prediction block, and resCr represents the difference between the original Cr block and the Cr prediction block. In some techniques, for example, after one or more of transformation to the transform domain or frequency domain, quantization, and context-based decoding or bypass decoding, the video encoder signals information for resCb and signals information for resCr separately. The video decoder receives information for resCb' and resCr' separately and adds resCb and resCr to the corresponding Cb prediction blocks and Cr prediction blocks to reconstruct the Cb and Cr blocks. In lossy decoding, due to losses caused by the quantization and dequantization processes, resCb' and resCr' may not be exactly equal to resCb and resCr. In some implementations of lossless decoding, resCb' and resCr' may be equal to resCb and resCr.

[0029] In JCCR mode, the video encoder constructs a joint residual block (called resJointC) based on resCb and resCr, instead of decoding resCb and resCr separately. The video encoder signals the information for resJointC, which reduces the amount of information that needs to be signaled compared to signaling the information for resCb and resCr separately. The video decoder receives the information for resJointC, generates resJointC (called resJointC', indicating that resJointC' is located on the video decoder side) based on the received information, and then generates resCb' and resCr'.

[0030] In some cases, using JCCR may lead to lossy decoding. For example, resCb' (e.g., the residual Cb block on the video decoder side) may differ from resCb (e.g., the residual Cb block on the video encoder side), and resCr' (e.g., the residual Cr block on the video decoder side) may differ from resCr (e.g., the residual Cr block on the video decoder side). The use of JCCR may lead to lossy decoding due to the transform and quantization operations performed by the video encoder. However, in some cases, the operations used to generate the joint residual chroma block (e.g., resJointC) may themselves be lossy. Techniques that lead to lossy decoding using JCCR can be referred to as Type I JCCR.

[0031] For higher video quality, there may be cases where lossless video decoding is preferred. Because JCCR is lossy in some examples, lossless video decoding may not be available. Furthermore, in cases where lossless video decoding is feasible, the benefits of JCCR (e.g., reducing the amount of residual information transmitted as a signal) may not be available.

[0032] This disclosure describes examples of techniques for a second type of JCCR. A second type of JCCR may be lossless or may be less lossy compared to a first type of JCCR. In some examples, a second type of JCCR may be used where the transform and / or quantization are bypassed or where the transform and / or quantization process is lossless. However, examples of a second type of JCCR should not be construed as limiting oneself to lossless examples or to examples where the transform and / or quantization are bypassed or the transform and / or quantization process is lossless.

[0033] According to the technology of this disclosure, a video decoder can be configured to receive information about a first residual sample for a first chroma component of a block of video data, and to reconstruct a first chroma block based on the first residual sample and a first prediction block. The first residual sample can be determined in such a way that adding the first residual sample to a sample of the first prediction block yields the original first chroma block, thus resulting in lossless decoding of the first chroma block.

[0034] The video decoder can also determine intermediate reconstructed samples of the second chroma component of a block of video data based on the first residual sample and the second predicted block. The video decoder can receive information for the second residual sample and reconstruct the second chroma block by adding the second residual sample to the intermediate reconstructed sample. The second residual sample can be determined in a manner that allows adding the second residual sample to the intermediate reconstructed sample to obtain the original second chroma block, thus leading to lossless decoding of the second chroma block. Furthermore, since the intermediate reconstructed sample is typically very close to the sample of the original second chroma block, the second residual sample can usually be a relatively small value, for example, with a large number of 0s and 1s, and therefore can be decoded efficiently using relatively few bits. Therefore, by reconstructing the second chroma block based on the second residual sample and the intermediate reconstructed sample, the video coding and decoding system can advantageously achieve lossless decoding while still benefiting from the bit savings offered by JCCR for some decoding scenarios.

[0035] Figure 1 This is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. In summary, the techniques of this disclosure are directed to decoding (encoding and / or decoding) video data. Typically, video data includes any data used for processing video. Therefore, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).

[0036] like Figure 1 As shown, in this example, system 100 includes source device 102, which provides encoded video data to be decoded and displayed by destination device 116. Specifically, source device 102 provides the video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can include any of a wide range of devices, including desktop computers, notebook computers (i.e., laptop computers), tablet computers, set-top boxes, mobile phones such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and may therefore be referred to as wireless communication devices.

[0037] exist Figure 1In the example, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for Joint Chromatic Residual Decoding (JCCR). Therefore, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, the source device and destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source, such as an external camera. Similarly, destination device 116 may interface with an external display device, rather than including an integrated display device. In this disclosure, operations described as being performed on the video encoder side can be performed, for example, by the video encoder 200 of system 100, and operations described as being performed on the video decoder side can be performed, for example, by the video decoder 300 of system 100.

[0038] like Figure 1 The system 100 shown is merely an example. Typically, any digital video encoding and / or decoding device can perform the techniques used for JCCR. Source device 102 and destination device 116 are merely examples of such decoding devices, where source device 102 generates decoded video data for transmission to destination device 116. In this disclosure, "decoding device" refers to a device that performs the decoding (e.g., encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices (specifically, video encoder and video decoder). In some examples, source device 102 and destination device 116 may operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Therefore, system 100 can support one-way or two-way video transmission between source device 102 and destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0039] Typically, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a sequential series of pictures (also referred to as “frames”) of the video data to video encoder 200, which encodes the data used for the pictures. Video source 104 of source device 102 may include video capture devices such as cameras, video archives containing previously captured raw video, and / or video feed interfaces for receiving video from video content providers. Alternatively, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from their received order (sometimes referred to as “display order”) to a decoding order for decoding. Video encoder 200 may generate a bitstream comprising the encoded video data. Then, the source device 102 can output encoded video data to a computer-readable medium 110 via the output interface 108 for reception and / or retrieval by, for example, the input interface 122 of the destination device 116.

[0040] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data, such as raw video from video source 104 and raw decoded video data from video decoder 300. Alternatively, memories 106 and 120 may store software instructions executable by, for example, video encoder 200 and video decoder 300. Although memories 106 and 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store, for example, encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106 and 120 may be allocated as one or more video buffers, for example, to store raw decoded and / or encoded video data.

[0041] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to directly transmit encoded video data to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, output interface 108 can modulate the transmitted signal including the encoded video data, and input interface 122 can demodulate the received transmitted signal. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, switch, base station, or any other device that may be useful for facilitating communication from source device 102 to destination device 116.

[0042] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessible data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0043] In some examples, source device 102 may output encoded video data to file server 114 or another intermediate storage device that may store the encoded video generated by source device 102. Destination device 116 may access the stored video data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing and sending encoded video data to destination device 116. File server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 may access the encoded video data from file server 114 via any standard data connection (including an Internet connection). This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to streaming protocols, download protocols, or combinations thereof.

[0044] Output interface 108 and input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data (such as encoded video data) according to cellular communication standards (such as 4G, 4G-LTE (Long Term Evolution), improved LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to operate according to other wireless standards (such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee)). TM ),Bluetooth TM The source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include an SoC device for performing the functions assigned to video encoder 200 and / or output interface 108, and destination device 116 may include an SoC device for performing the functions assigned to video decoder 300 and / or input interface 122.

[0045] The technology disclosed herein can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding of digital video stored on data storage media, or other applications.

[0046] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, such as syntax elements (which are also used by the video decoder 300), which have values ​​describing the characteristics and / or processing of video blocks or other decoding units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays a decoded picture of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device.

[0047] Despite Figure 1 Not shown, but in some examples, the video encoder 200 and video decoder 300 may each be integrated with the audio encoder and / or audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams including both audio and video in a common data stream. Where applicable, the MUX-DEMUX unit may comply with the ITU H.223 multiplexer protocol or other protocols (such as User Datagram Protocol (UDP)).

[0048] The video encoder 200 and video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices (such as cellular phones).

[0049] Examples of video decoding standards include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual (MPEG-4 Part 2), ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Decoding (SVC) and Multi-View Video Decoding (MVC) extensions), and ITU-T H.265 (also known as ISO / IEC MPEG-4 HEVC) with its extensions.

[0050] The video encoder 200 and video decoder 300 can operate according to video decoding standards such as ITU-T H.265 (also known as High Efficiency Video Coding (HEVC)) or extensions thereof (such as Multi-View and / or Scalable Video Coding Extensions)). Alternatively, the video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards such as the Joint Exploratory Test Model (JEM) or the ITU-T H.266 standard (also known as Universal Video Coding (VVC)). The latest draft of the VVC standard is described in the following document: Bross et al., “Versatile Video Coding (Draft 7)”, Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC29 / WG 11, 16th meeting: Geneva, Switzerland, 1-11 October 2019, JVET-P2001-v14 (hereinafter referred to as “VVC Draft 7”). However, the technology of this disclosure is not limited to any particular decoding standard.

[0051] Typically, video encoder 200 and video decoder 300 can perform block-based decoding of images. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during encoding and / or decoding). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, video encoder 200 and video decoder 300 can decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, while video encoder 200 and video decoder 300 can decode luminance and chrominance components, not the red, green, and blue (RGB) data samples used for images, where the chrominance components may include both red and blue hue chrominance components. In some examples, video encoder 200 converts the received RGB-formatted data to a YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and post-processing units (not shown) can perform these conversions.

[0052] In summary, this disclosure may relate to the decoding (e.g., encoding and decoding) of images to include the process of encoding or decoding the data of an image. Similarly, this disclosure may relate to the decoding of blocks of an image to include the process of encoding or decoding the data used for the blocks (e.g., prediction and / or residual decoding). Encoded video bitstreams typically include a series of values ​​for representing decoding decisions (e.g., decoding modes) and syntax elements that segment the image into blocks. Therefore, references to decoding images or blocks should generally be understood as decoding the values ​​of the syntax elements used to form images or blocks.

[0053] HEVC defines various blocks, including decoding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (such as a video encoder 200) partitions the decoding tree unit (CTU) into CUs based on a quadtree structure. That is, the video decoder partitions the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and the CU of such leaf nodes can include one or more PUs and / or one or more TUs. The video decoder can further partition the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a partition of the TU. In HEVC, PUs represent inter-frame prediction data, while TUs represent residual data. CUs with intra-frame prediction include intra-frame prediction information, such as intra-frame mode indication.

[0054] As another example, the video encoder 200 and video decoder 300 can be configured to operate according to JEM or VVC. According to JEM or VVC, the video decoder (such as the video encoder 200) segments the image into multiple decoding tree units (CTUs). The video encoder 200 can segment the CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple segmentation types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level segmented according to quadtree segmentation and a second level segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to the decoding units (CUs).

[0055] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning (also known as triplet tree (TT)) partitioning. A ternary tree or triplet tree partitioning is a partition in which a block is divided into three sub-blocks. In some examples, a ternary tree or triplet tree partitioning divides a block into three sub-blocks without partitioning the original block through a center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0056] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).

[0057] The video encoder 200 and video decoder 300 can be configured to use per-HEVC quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structures. For illustrative purposes, a description of the techniques of this disclosure is given with respect to QTBT segmentation. However, it should be understood that the techniques of this disclosure can also be applied to video decoders configured to use quadtree segmentation or other types of segmentation.

[0058] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a brick can refer to a rectangular area of ​​a row of CTUs within a specific tile in an image. A tile can be a rectangular area of ​​a CTU within a specific tile column or a specific tile row in an image. A tile column refers to a rectangular area of ​​a CTU with a height equal to the height of the image and a width specified by a syntax element (e.g., as in an image parameter set). A tile row refers to a rectangular area of ​​a CTU with a height specified by a syntax element (e.g., as in an image parameter set) and a width equal to the width of the image.

[0059] In some examples, a tile may be divided into multiple bricks, where each brick may include one or more CTU rows within the tile. A tile that is not divided into multiple bricks may also be referred to as a brick. However, a brick that is a true subset of a tile may not be referred to as a tile.

[0060] The bricks in an image can also be arranged as slices. A slice can be an integer number of bricks in the image, and it can be uniquely contained within a single Network Abstraction Layer (NAL) unit. In some examples, a slice consists of a continuous sequence of several complete tiles or a single complete tile.

[0061] This disclosure uses "NxN" and "N by N" interchangeably to refer to the sample size of a block (such as a CU or other video block) in the vertical and horizontal dimensions, for example, 16x16 samples or 16 by 16 samples. Typically, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as it does in the vertical direction. For example, a CU can include NxM samples, where M is not necessarily equal to N.

[0062] The video encoder 200 encodes video data for use in predicting and / or residual information, as well as other information, for the CU. The prediction information indicates how the CU will be predicted to form a prediction block for the CU. The residual information typically represents the sample-by-sample difference between a sample of the CU before encoding and the prediction block.

[0063] To predict the Cubic Frame (CU), the video encoder 200 typically forms prediction blocks for the CU using either inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU based on data from previously decoded images, while intra-frame prediction generally refers to predicting the CU based on data from previously decoded images of the same frame. To perform inter-frame prediction, the video encoder 200 can generate prediction blocks using one or more motion vectors. The video encoder 200 can typically perform motion search to identify, for example, reference blocks that closely match the CU in terms of the difference between the CU and a reference block. The video encoder 200 can calculate difference metrics using sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.

[0064] Some examples of JEM and VVC also provide affine motion compensation modes, which can be considered as inter-frame prediction modes. In affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular types of motion).

[0065] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of JEM and VVC provide sixty-seven intra-frame prediction modes, including various directional modes, as well as planar and DC modes. Typically, the video encoder 200 selects an intra-frame prediction mode that describes the samples of the current block (e.g., a block of a CU) to be predicted based on, which are the neighboring samples of the current block. Assuming the video encoder 200 decodes the CTU and CU in raster scan order (from left to right, from top to bottom), such samples can typically be located above, to the upper left, or to the left of the current block within the same image.

[0066] The video encoder 200 encodes data representing the prediction mode used for the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used, as well as motion information for the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may use Advanced Motion Vector Prediction (AMVP) or merging modes to encode motion vectors. The video encoder 200 may use similar modes to encode motion vectors used for affine motion compensation modes.

[0067] Following a prediction, such as intra-frame or inter-frame prediction of a block, the video encoder 200 can compute residual data for that block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and the prediction block used to form the block using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply a Discrete Cosine Transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 can apply a second transform after the first transform, such as a Mode-dependent Inseparable Quadratic Transform (MDNSST), a Signal-dependent Transform, a Karhunen-Loeve Transform (KLT), etc. The video encoder 200 produces transform coefficients after applying one or more transforms.

[0068] As described above, after any transformation to produce transform coefficients, the video encoder 200 can perform quantization of the transform coefficients. Quantization generally refers to the process in which the transform coefficients are quantized to potentially reduce the amount of data used to represent them, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 can round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bitwise right shift of the values ​​to be quantized.

[0069] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place higher-energy (and therefore lower-frequency) transform coefficients before the vector and lower-energy (and therefore higher-frequency) transform coefficients after the vector. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy-encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy-encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic decoding (CABAC). The video encoder 200 can also entropy-encode the values ​​of syntax elements used to describe metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.

[0070] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. The context may involve, for example, whether the neighboring values ​​of a symbol are zero. Probability determination can be based on the context assigned to the symbol.

[0071] The video encoder 200 can also generate, for example, syntax data (such as block-based syntax data, image-based syntax data, and sequence-based syntax data) or other syntax data (such as sequence parameter sets (SPS), image parameter sets (PPS), or video parameter sets (VPS)) destined for the video decoder 300 in image headers, block headers, and slice headers. Similarly, the video decoder 300 can decode such syntax data to determine how to decode the corresponding video data.

[0072] In this way, the video encoder 200 can generate a bitstream that includes encoded video data, such as syntax elements describing the segmentation of an image into blocks (e.g., CUs) and prediction and / or residual information for those blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0073] Typically, the video decoder 300 performs a process reciprocal to that performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 may use CABAC to decode the values ​​of syntax elements used for the bitstream in a manner substantially similar to, but reciprocal to, the CABAC encoding process of the video encoder 200. Syntax elements may define segmentation information for segmenting images into CTUs, and for segmenting each CTU according to a corresponding segmentation structure (such as a QTBT structure) to define the CUs of the CTU. Syntax elements may also define prediction and residual information for blocks (e.g., CUs) of the video data.

[0074] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inversely transform the quantized transform coefficients of the block to regenerate a residual block for that block. The video decoder 300 uses the transmitted prediction mode (intra-frame prediction or inter-frame prediction) and associated prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for that block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to regenerate the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the block boundaries.

[0075] According to the technology of this disclosure, video encoder 200 and video decoder 300 can be configured to perform a technique based on chroma residual joint decoding (JCCR). As described above and in more detail below, in JCCR, residual information for two chroma components can be combined to reduce the amount of information transmitted with the signal. However, some JCCR techniques may not be used for lossless decoding. In some examples of this disclosure, video encoder 200 and video decoder 300 may utilize JCCR techniques, but in a manner that is lossless or at least less lossy compared to other JCCR techniques.

[0076] In some examples of this disclosure, the video decoder 300 may perform the following operations: receiving information for a first residual sample for a first chroma component, wherein the first residual sample is based on the difference between a first chroma block and a first prediction block of the first chroma component; determining intermediate samples based on the first residual sample; receiving information for a second residual sample for a second chroma component, wherein the second residual sample is based on the difference between a second chroma block and an intermediate sample of the second chroma component, wherein the second chroma block and the first chroma block are associated together (e.g., part of the same CU); reconstructing the first chroma block based on the first residual sample and the first prediction block; and reconstructing the second chroma block based on the second residual sample and the intermediate sample.

[0077] In some examples, the video decoder 300 may perform the following operations: receiving information for a set of values, wherein the set of values ​​is generated based on a first residual sample set of a first chroma block of a first chroma component and a second residual sample set of a second chroma block of a second chroma component, wherein the first chroma block and the second chroma block are associated; determining a first intermediate sample based on the set of values; determining a second intermediate sample based on the set of values; receiving information for the first residual sample, wherein the first residual sample is based on the difference between the first chroma block of the first chroma component and the first intermediate sample; receiving information for the second residual sample, wherein the second residual sample is based on the difference between the second chroma block of the second chroma component and the second intermediate sample; reconstructing the first chroma block based on the first residual sample and the first intermediate sample; and reconstructing the second chroma block based on the second residual sample and the second intermediate sample.

[0078] In some examples, the video encoder 200 may perform the following operations: signaling information for a first residual sample for a first chroma component, wherein the first residual sample is based on the difference between a first chroma block and a first prediction block of the first chroma component; determining intermediate samples based on the first residual sample; and signaling information for a second residual sample for a second chroma component, wherein the second residual sample is based on the difference between a second chroma block and an intermediate sample of the second chroma component, and wherein the second chroma block and the first chroma block are associated together (e.g., part of the same CU).

[0079] In some examples, the video encoder 200 may perform the following operations: generating a value set based on a first residual sample set of a first chroma component and a second residual sample set of a second chroma component, wherein the first residual sample set is based on the difference between a first chroma block and a first prediction block in the first chroma component, and the second residual sample set is based on the difference between a second chroma block and a second prediction block in the second chroma component, wherein the first chroma block and the second chroma block are associated; signaling information for the value set; determining a first intermediate sample based on the value set; determining a second intermediate sample based on the value set; signaling information for the first residual sample, wherein the first residual sample is based on the difference between a first chroma block and a first intermediate sample in the first chroma component; and signaling information for the second residual sample, wherein the second residual sample is based on the difference between a second chroma block and a second intermediate sample in the second chroma component.

[0080] In summary, this disclosure may involve "signaling" certain information (such as syntax elements). The term "signaling" can generally refer to the transmission of values ​​for syntax elements and / or other data for decoding encoded video data. That is, video encoder 200 can signal values ​​for syntax elements in the bitstream. Generally, signaling refers to generating values ​​in the bitstream. As described above, source device 102 can transmit the bitstream to destination device 116 substantially in real time or non-real time (such as when syntax elements are stored in storage device 112 for later retrieval by destination device 116).

[0081] Figure 2A and Figure 2B This is a conceptual diagram illustrating an example Quadtree Binary Tree (QTBT) structure 130 and its corresponding Decoding Tree Unit (CTU) 132. Solid lines represent quadtree splits, while dashed lines indicate binary tree splits. In each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which split type (i.e., horizontal or vertical) is used, where, in this example, 0 indicates a horizontal split and 1 indicates a vertical split. For quadtree splits, since the quadtree node splits the block horizontally and vertically into four sub-blocks of equal size, there is no need to indicate the split type. Therefore, the video encoder 200 can encode the following, and the video decoder 300 can decode the following: syntax elements (such as split information) for the region tree level (i.e., solid lines) of the QTBT structure 130, and syntax elements (such as split information) for the prediction tree level (i.e., dashed lines) of the QTBT structure 130. The video encoder 200 can encode video data (such as prediction and transform data) for a CU represented by the terminal leaf nodes of the QTBT structure 130, while the video decoder 300 can decode the video data.

[0082] generally, Figure 2B The CTU 132 can be associated with parameters that define the size of the block corresponding to the node at the first and second levels of the QTBT structure 130. These parameters can include the CTU size (representing the size of the CTU 132 in the sample), the minimum quadtree size (MinQTSize, which represents the minimum allowed quadtree leaf node size), the maximum binary tree size (MaxBTSize, which represents the maximum allowed binary tree root node size), the maximum binary tree depth (MaxBTDepth, which represents the maximum allowed binary tree depth), and the minimum binary tree size (MinBTSize, which represents the minimum allowed binary tree leaf node size).

[0083] The root node corresponding to the CTU in a QTBT structure can have four child nodes at the first level of the QTBT structure, where each child node can be partitioned according to a quadtree. That is, the nodes at the first level are leaf nodes (without child nodes) or have four child nodes. An example of QTBT structure 130 represents such a node as including a parent node and child nodes with solid-line branches. If the nodes at the first level are not larger than the maximum allowed binary tree root node size (MaxBTSize), these nodes can be further partitioned by the corresponding binary tree. The binary tree split of a node can be iterated until the nodes generated from the split reach the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). An example of QTBT structure 130 represents such a node as having dashed-line branches. The binary tree leaf nodes are called decoding units (CUs), which are used for prediction (e.g., intra-image or inter-image prediction) and transformation without any further partitioning. As discussed above, CUs can also be referred to as “video chunks” or “blocks”.

[0084] In one example of a QTBT segmentation structure, the CTU size is set to 128x128 (luminance samples and two corresponding 64x64 chrominance samples), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. First, a quadtree segmentation is applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes ranging from 16x16 (i.e., MinQTSize) to 128x128 (i.e., the CTU size). If a leaf quadtree node is 128x128, it will not be further split by the binary tree because this size exceeds MaxBTSize (i.e., 64x64 in this example). Otherwise, the leaf quadtree node will be further split by the binary tree. Therefore, the quadtree leaf node also serves as the root node of the binary tree and has a binary tree depth of 0. When the depth of a binary tree reaches MaxBTDepth (4 in this example), further splitting is not allowed. A binary tree node with a width equal to MinBTSize (4 in this example) means that further horizontal splitting is not allowed. Similarly, a binary tree node with a height equal to MinBTSize means that further vertical splitting is not allowed for that binary tree node. As mentioned above, the leaf nodes of the binary tree are called CUs and are further processed according to predictions and transformations without further splitting.

[0085] “Versatile Video Coding (Draft 6)”, ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC29 / WG 11 Joint Video Experts Group (JVET), 15th Meeting: Gothenburg, Sweden, July 3-12, 2019. JVET-O2001-vD (hereinafter referred to as “VVC Draft 6”) supports a mode for joint decoding of chroma residuals. The use (activation) of the joint chroma decoding mode is indicated by the TU level flag tu_joint_cbcr_residual_flag, and the selected mode is implicitly indicated by the chroma CBF. The flag tu_joint_cbcr_residual_flag exists if any one or both chroma CBFs of the TU are equal to 1. In the Picture Parameter Set (PPS) and slice header, the chroma QP offset value is signaled for the joint chroma residual decoding mode, to distinguish it from the usual chroma QP offset value signaled for the conventional chroma residual decoding mode. These chroma QP offset values ​​are used to derive the chroma QP values ​​for those blocks decoded using the joint chroma residual decoding mode. When the corresponding joint chroma decoding mode (Mode 2 in Table 1 below) is active in the TU, this chroma QP offset is added to the applied luminance-derived chroma QP during quantization and decoding of that TU. For other modes (Mode 1 and Mode 3 in Table 1 below), the chroma QP is derived in the same manner as for regular Cb or Cr blocks. Table 1 depicts the process of reconstructing the chroma residuals (resCb and resCr) from the transmitted transform blocks.When this mode is activated, a single joint chroma residual block (resJointC'[x][y] in Table 1) is sent by signal, and the residual blocks of Cb (resCb) and Cr (resCr) are derived taking into account information such as tu_cbf_cb, tu_cbf_rr, and CSign, where CSign is the sign value specified in the image header, as described in the following documents: Helmrich et al., “CE7: Joint chroma residual coding with multiple modes (tests CE7-2.1, CE7-2.2)”, Joint Video Experts Group (JVET) of ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 15th Meeting: Gothenburg, Sweden, July 3-12, 2019, JVET-O0105; and Helmrich et al., “Alternative configuration for joint chroma residual coding”, ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11. Joint Video Experts Group (JVET) of 1 / SC 29 / WG 11, 15th Meeting: Gothenburg, Sweden, July 3-12, 2019, JVET-O0543.

[0086] On the encoder side, the joint chroma components are derived, as explained below. Depending on the mode, the video encoder 200 generates resJointC, as shown in Table 1. Table 2 shows the reconstruction of the chroma residuals for different modes performed by the video decoder 300.

[0087] For all modes, the original residuals on the encoder side are calculated as: resCb[x][y] = origCb[x][y] – predCb[x][y], and resCr[x][y] = origCr[x][y] – predCr[x][y]. Due to transform and quantization, and due to joint representation, the residuals res'Cb[x][y] and res'Cr[x][y] derived on the decoder side are generally different from resCb[x][y] and resCr[x][y].

[0088] Furthermore, on the decoder side, due to transform / quantization, the joint residual resJointC' may not be the same as the joint residual resJointC generated by the video encoder 200, unless it bypasses transform and quantization. Even if resJointC is transmitted losslessly (e.g., transform-quantization bypass or transform-skipped, where QP=4 for bit depth 8), lossless reconstruction of the chroma residual may not be guaranteed due to the joint representation of the residuals.

[0089] Table 1: Generation of Joint Chromaticity Residues

[0090]

[0091] Table 2: Reconstruction of chromaticity residuals

[0092]

[0093] For all modes, the decoder-side reconstruction (e.g., by video decoder 300) is as follows: recCb[x][y] = predCb[x][y] + res'Cb[x][y], and recCr[x][y] = predCr[x][y] + res'Cr[x][y]. The three joint chroma decoding modes described above can be supported in intra-frame prediction CUs. For non-intra-frame CUs, only mode 2 can be supported. Therefore, in non-intra-frame CUs, the syntax element tu_joint_cbcr_residual_flag is only possible when both chroma cbf values ​​are 1.

[0094] In Xu et al.'s "CE8-related: A SPS Level Flag for BDPCM and JCCR", ITU-T SG16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 Joint Video Experts Group (JVET), 15th Meeting: Gothenburg, Sweden, July 3-12, 2019, JVET-O0376, a Sequence Parameter Set (SPS) level flag has been added to control the enabling / disabling of joint Cb-Cr for each video sequence. Therefore, by setting this SPS flag to zero, an explicit way to disable JCCR can be implemented in VVC.

[0095] The lossless decoding capabilities of the VVC codec are described below. VVC also plans to support lossless decoding, which is currently being studied in the Ad-hoc Group (AHG) and core experiments. During the 16th session of JVET, AHG18 produced a lossless VTM anchor, which is based on HEVC transform quantization bypass and extended to a new VVC tool with some canonical and non-canonical changes. Subsequently, in "AHG 18: Enabling lossless coding with minimal impact on VVC design", ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 Joint Video Experts Group (JVET), 16th session: Geneva, Switzerland, October 1-11, 2019, JVET-P0606, an improved design for achieving lossless decoding with minimal impact on VVC design was proposed.

[0096] However, using JCCR in situations requiring lossless decoding (or at least less lossy decoding compared to previous JCCR implementations) can present problems. For JVET activities currently enabled for lossless decoding, disable the JCCR tool at the SPS level. This is because, for both the Cb and Cr components, JCCR only sends a residual with the signal, and typically, this method may not be able to perform lossless decoding on the residual, and therefore, lossless decoding may be unavailable.

[0097] This disclosure describes several techniques for implementing lossless decoding using existing JCCR patterns with minimal modification. The example techniques can be used together or separately. For convenience, this disclosure describes examples of techniques one and two. However, these techniques should not be construed as being limited to only two techniques.

[0098] The current JCCR method generates the joint chromaticity residual on the encoder side through a linear combination of the Cb and Cr residuals. Typically, the joint residual calculated in this way is not a lossless version of the Cb and Cr residuals.

[0099] In Technique 1, two residuals (resJointC and resJointC2) are sent as signals to achieve lossless decoding (for the original JCCR, only resJointC is sent as a signal). Table 3 shows the procedures on the encoder side (e.g., the operation of video encoder 200) and the decoder side (e.g., the operation of video decoder 300). For convenience, components Cb and Cr are defined as the first and second components. For modes 1 and 2, Cb and Cr are the first and second components, respectively, and for mode 3, Cr and Cb are the first and second components, respectively. As shown in Table 3, video encoder 200 calculates resJointC in such a way that lossless reconstruction of the first component is available and guaranteed in some cases. That is, video encoder 200 and video decoder 300 use a lossless version of resCb or resCr (depending on the mode) as resJointC. The clipping boundary is (0, 2). bitdepth –1), where bitdepth is the bit depth of the internal (processing) codec operation.

[0100] In the second step, an intermediate step for reconstructing the second component (rec1Cr / rec1Cb) is performed according to the decoding process of the original JCCR method with clipping. Subsequently, resJointC2 is calculated by subtracting this intermediate reconstruction from the original second component. Here, the clipping stage ensures that the intermediate reconstruction value is within the dynamic range, and subsequently, clipping also ensures that resJointC2 will have the same dynamic range as other residuals (e.g., resJointC, etc.).

[0101] To perform lossless decoding, resJointC and resJointC2 should be decoded in a lossless manner. This can be achieved by bypassing the transform and quantization stages (as in HEVC lossless mode), or by using a transform skip of QP = 4 + 2 * (8-bit depth), which makes the quantization step size = 1, and since there is no transform, lossless decoding is feasible. In some examples, the QP offset of the JCCR mode also needs to be zero.

[0102] Table 3: Encoding and Decoding Processes for Technology 1

[0103]

[0104] The following is an example with a 2x2 residual and an internal bit depth of 10 bits. In this example, CSign = +1 and mode 1 are used.

[0105] The original Cb chromaticity block (origCb) is And the Cb prediction block (predCb) is Therefore, the video encoder 200 can determine the Cb residual block (resCb) as The original Cr chromaticity block (origCr) is And the Cr prediction block (predCb) is Therefore, the video encoder 200 can determine the Cr residual block (resCr) as The video encoder 200 can determine the resJointC to be transmitted as resCb = origCb – predCb, which is equal to

[0106] The video decoder 300 can reconstruct the Cb chroma block (recCb) into resJointC+predCb.

[0107] Both the video encoder 200 and the video decoder 300 can define rec1Cr as This produces

[0108] The video encoder 200 can determine the resJointC2 to be sent as the signal. The video decoder 300 can reconstruct the Cr chroma block (recCr) into resJointC2+rec1Cr, which is equal to origCr.

[0109] Although Method 1 is based on the original JCCR method (where the scaling factor of the second component relative to the first component is 1 / 2), such restrictions can be removed for more generalized methods with scaling factors of k1 or k2 (in the range of (0,1), excluding values ​​0 and 1), as shown in Table 3.1. Values ​​k1 and k2 can be signaled at the SPS, picture header, CTU level, or TU level. The value for b is a positive integer, which can be predetermined, fixed, or signaled at the SPS, picture header, or CTU level.

[0110] Table 3.1: Generalized Version of Technique 1

[0111]

[0112] To implement the techniques in Tables 3 and 3.1 above, the video encoder 200 can, for example, be configured to determine a first residual sample (e.g., resCb[x][y] = resJointC[x][y]) of a first chroma component of a block of video data, where the first residual sample corresponds to the difference between a first chroma block (e.g., recCb[x][y]) of the first chroma component and a first prediction block (e.g., predCb[x][y]) of the first chroma component. The video encoder 200 determines an intermediate reconstruction sample (e.g., rec1Cr[x][y]) based on the first residual sample (e.g., resJointC[x][y]) and a second prediction block (e.g., predCr[x][y]) of the second chroma component, and determines a second residual sample (e.g., resJointC2[x][y]) of the second chroma component of a block of video data, where the second residual sample corresponds to the difference between a second chroma block (e.g., recCr[x][y]) of the second chroma component and an intermediate reconstruction sample (e.g., rec1Cr[x][y]). The video encoder 200 can output syntax elements representing the first residual sample and the second residual sample from the bit stream of encoded video data.

[0113] To implement the techniques in Tables 3 and 3.1 above, the video decoder 300 may, for example, be configured to receive information of a first residual sample (e.g., resCb[x][y] = resJointC[x][y]) for a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block (e.g., recCb[x][y]) and a first prediction block (e.g., predCb[x][y]) of the first chroma component. The video decoder 300 may determine an intermediate reconstruction sample (e.g., rec1Cr[x][y]) based on the first residual sample (e.g., resJointC[x][y]) and receive information of a second residual sample (e.g., resJointC2[x][y]) for a block of video data, wherein the second residual sample corresponds to the difference between a second chroma block (e.g., recCr[x][y]) and an intermediate reconstruction sample (e.g., rec1Cr[x][y]) of the second chroma component of the second chroma component. The video decoder 300 can reconstruct a first chroma block based on a first residual sample and a first predicted block (e.g., recCb[x][y] = predCb[x][y] + resCb[x][y]), and reconstruct a second chroma block based on a second residual sample and intermediate reconstructed samples (e.g., recCr[x][y] = rec1Cr[x][y] + resJointC2[x][y]). It should be understood that the above examples are described with respect to modes 1 and 2, where the first chroma component is a Cb component and the second chroma component is a Cr component. For mode 3, the first chroma component can be a C3 component, and the second chroma component can be a Cb component.

[0114] In Example Technique 1, two residuals, resJointC and resJointC2, are signaled. In Example Technique 2, three residuals can be signaled. For example, in Example Technique 2, three residuals, resJointC', resJointC1, and resJointC2, are signaled to achieve lossless decoding (as mentioned above, for the original JCCR, only resJointC' is signaled, and for Technique 1, resJointC and resJointC2 are signaled).

[0115] The encoder-side (e.g., performed by video encoder 200) and decoder-side (e.g., performed by video decoder 300) processes are shown in Tables 4 and 5, respectively. For all modes, the original residuals on the encoder side are calculated as: resCb[x][y] = origCb[x][y] – predCb[x][y], and resCr[x][y] = origCr[x][y] – predCr[x][y]. In some examples, resJointC (e.g., the uncompressed residuals) is calculated in the same manner as the original JCCR process previously shown in Table 1. However, reconstructing the Cb and Cr components using only the joint residuals resJointC' (which is equal to resJointC for lossless decoding) may not provide a lossless representation. Therefore, additional minor residuals for Cb and Cr used for lossless decoding are derived as resJointC1 and resJointC2, as shown in Table 4. In some examples, resJointC' does not necessarily have to be the same as resJointC, so lossy decoding (transformation and / or quantization) of resJointC can even be performed.

[0116] Compared to Technique 1, Technique 2 uses an additional residual (three in total), but the advantage is that resJointC' can be a lossy version of resJointC (e.g., transformation and / or quantization can be performed), which may be better in terms of compression.

[0117] Table 4: Generation of Joint Chromaticity Residues for Technique 2

[0118]

[0119] Table 5: Reconstruction of chromaticity components for Technique 2

[0120]

[0121] Similar to Technique 1, Technique 2 can be extended to a more generalized method for scaling factors of k1 or k2 (within the range of (0,1), excluding values ​​0 and 1), as shown in Tables 6 and 7. Values ​​k1 and k2 can be signaled at the SPS or picture header, CTU level, or TU level.

[0122] Table 6: Generalized version of Technology 2 (encoder)

[0123]

[0124] Table 7: Generalized version of technology 2 (decoder)

[0125]

[0126] The following describes Transform Unit (TU) level and HLS (High-Level Syntax) level controls. Technique 1 or Technique 2 can be used for lossy or lossless decoding. Therefore, the TU level flag can indicate whether the JCCR uses one residual (the original JCCR) or more than one residual (a modified JCCR, which can be either Technique 1 or Technique 2). In some examples, additional signaling can be used to specify communication between Technique 1 and Technique 2. SPS / Picture header level controls can also be provided to disable the option to use a modified JCCR, thus avoiding the use of the TU level flag for the modified JCCR. In some examples, a modified JCCR can be used to replace the original JCCR.

[0127] Figure 3 This is a block diagram illustrating an example video encoder 200 that can perform the techniques described in this disclosure. Figure 3 This disclosure is provided for illustrative purposes and should not be construed as limiting the techniques as extensively illustrated and described herein. For illustrative purposes, this disclosure describes the video encoder 200 in the context of video decoding standards such as the HEVC video decoding standard and the H.266 (VVC) video decoding standard under development. However, the techniques of this disclosure are not limited to these video decoding standards and are generally applicable to video encoding and decoding.

[0128] exist Figure 3 In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy coding unit 220 can be implemented in one or more processors or in processing circuitry. For example, the units of the video encoder 200 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0129] The video data storage device 230 can store video data to be encoded by the components of the video encoder 200. The video encoder 200 can obtain data from, for example, a video source 104 (…). Figure 1The video data memory 230 receives video data stored in the video data memory 230. The DPB 218 can act as a reference picture memory, storing reference video data for use when the video encoder 200 predicts subsequent video data. The video data memory 230 and DPB 218 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 can be on-chip (as shown) with other components of the video encoder 200, or off-chip relative to those components.

[0130] In this disclosure, references to video data memory 230 should not be construed as limited to memory internal to video encoder 200 (unless so specifically described) or memory external to video encoder 200 (unless so specifically described). Rather, references to video data memory 230 should be understood as reference memory storing video data received by video encoder 200 for encoding (e.g., video data for the current block to be encoded). Figure 1 The memory 106 can also provide temporary storage for the outputs from the various units of the video encoder 200.

[0131] It shows Figure 3 The various units help to understand the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide a specific function and are pre-configured in terms of the operations they can perform. Programmable circuits refer to circuits that can be programmed to perform a variety of tasks and provide flexible functionality in terms of the operations they can perform. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by instructions in the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units can be integrated circuits.

[0132] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core, formed according to programmable circuitry. In an example where the operation of the video encoder 200 is performed using software executed by programmable circuitry, memory 106 ( Figure 1 The video encoder 200 may store software instructions (e.g., object code) received and executed by the video encoder 200, or such instructions may be stored in another memory (not shown) within the video encoder 200.

[0133] The video data storage unit 230 is configured to store the received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be the raw video data to be encoded.

[0134] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units that perform video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0135] Mode selection unit 202 typically coordinates multiple coding passes to test combinations of coding parameters and the rate-distortion values ​​obtained for such combinations. Coding parameters may include segmenting the CTU into CUs, the prediction mode for the CUs, the transformation type of the residual data for the CUs, and the quantization parameters of the residual data for the CUs. Mode selection unit 202 can ultimately select a combination of coding parameters that yields a better rate-distortion value than other tested combinations.

[0136] The video encoder 200 can segment images retrieved from the video data storage 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can segment the CTUs of the image according to a tree structure (such as the QTBT structure or quadtree structure of HEVC described above). As described above, the video encoder 200 can segment CTUs according to a tree structure to form one or more CUs. Such CUs can also be referred to as "video blocks" or "blocks".

[0137] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate prediction blocks for the current block (e.g., the current CU, or the overlapping portion of PU and TU in HEVC). To perform inter-frame prediction for the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in DPB 218). Specifically, motion estimation unit 222 may calculate values ​​representing the similarity between the potential reference block and the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared error (MSD), etc. Motion estimation unit 222 may typically perform these calculations using sample-by-sample differences between the current block and the reference blocks being considered. Motion estimation unit 222 may identify the reference block with the lowest value obtained from these calculations, indicating the reference block that most closely matches the current block.

[0138] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the position of a reference block in a reference image relative to the position of the current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data for the reference blocks. As another example, if the motion vectors have fractional-sample precision, motion compensation unit 224 can interpolate the values ​​used for the prediction blocks according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by corresponding motion vectors and combine the retrieved data, for example, by per-sample averaging or weighted averaging.

[0139] As another example, for intra-prediction or intra-prediction decoding, intra-prediction unit 226 can generate a prediction block based on samples adjacent to the current block. For example, in directional mode, intra-prediction unit 226 can typically mathematically combine the values ​​of adjacent samples and fill these calculated values ​​across the current block in a defined direction to generate a prediction block. As another example, in DC mode, intra-prediction unit 226 can calculate the average of the adjacent samples of the current block and generate a prediction block to include the obtained average for each sample of the prediction block.

[0140] Mode selection unit 202 provides a prediction block to residual generation unit 204. Residual generation unit 204 receives the original, uncoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, residual generation unit 204 may also determine the difference between sample values ​​in the residual block to generate the residual block using residual differential pulse decode-modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form residual generation unit 204.

[0141] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As noted above, the size of a CU can refer to the size of the luma decoding block of the CU, while the size of a PU can refer to the size of the luma prediction unit of the PU. Assuming a particular CU size is 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra-frame prediction, and 2Nx2N, 2NxN, Nx2N, NxN, or similar symmetrical PU sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric segmentation for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-frame prediction.

[0142] In the example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luma decoding block and a corresponding chroma decoding block. As mentioned above, the size of the CU can refer to the size of the luma decoding block of the CU. Video encoder 200 and video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0143] For other video decoding techniques (such as block-based copy mode decoding, affine mode decoding, and linear model (LM) mode decoding, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples (such as palette mode decoding), mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block should be reconstructed based on the selected palette. In such a mode, mode selection unit 202 can provide these syntax elements to entropy coding unit 220 for encoding.

[0144] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0145] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 can apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 can apply a discrete cosine transform (DCT), direction transform, Karhunen-Loeve transform (KLT), or conceptually similar transformations to the residual block. In some examples, transform processing unit 206 can perform multiple transformations on the residual block, such as primary and secondary transformations (e.g., rotation transformations). In some examples, transform processing unit 206 does not apply any transformations to the residual block.

[0146] Quantization unit 208 can quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 can quantize the transform coefficients of the transform coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may cause information loss, and therefore, the quantized transform coefficients may have lower accuracy compared to the original transform coefficients generated by transform processing unit 206.

[0147] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although potentially with some degree of distortion). For example, the reconstruction unit 214 can add samples from the reconstructed residual block to corresponding samples from the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0148] Filter unit 216 can perform one or more filtering operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce block artifacts along the edges of the CU. In some examples, the operation of filter unit 216 can be skipped.

[0149] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 can store the reconstructed blocks in the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 can store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve a reference image formed from the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction of blocks in subsequent encoded images. Additionally, the intra-frame prediction unit 226 can use the reconstructed blocks of the current image in the DPB 218 to perform intra-frame prediction of other blocks in the current image.

[0150] Typically, entropy coding unit 220 can entropy code syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy code quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy code predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, another example of video data, to generate entropy-coded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, exponential-Golomb coding, or another type of entropy coding operation on the data. In some examples, the entropy coding unit 220 can operate in a bypass mode where syntax elements are not entropy encoded.

[0151] The video encoder 200 can output a bitstream that includes entropy-encoded syntax elements required for reconstructing slices or blocks of images. Specifically, the entropy coding unit 220 can output a bitstream.

[0152] The block description refers to the operations described above. This description should be understood as referring to the operations used for the luma decoding block and / or chroma decoding block. As mentioned above, in some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the CU. In some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the PU.

[0153] In some examples, it is not necessary to repeat the operations performed for the luma decoding block for the chroma decoding block. As an example, it is not necessary to repeat the operations used to identify the motion vector (MV) and reference image for the luma decoding block to identify the MV and reference image for the chroma block. Specifically, the MV for the luma decoding block can be scaled to determine the MV for the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma decoding blocks.

[0154] Video encoder 200 represents an example of a device configured to encode video data, the device including: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: signal information for a first residual sample for a first chroma component, wherein the first residual sample is based on the difference between a first chroma block and a first prediction block of the first chroma component; determine intermediate samples based on the first residual sample; and signal information for a second residual sample for a second chroma component, wherein the second residual sample is based on the difference between a second chroma block and an intermediate sample of the second chroma component, wherein the second chroma block and the first chroma block are associated together.

[0155] One or more processing units of the video encoder 200 may also be configured to: generate a value set (e.g., resJointC) based on a first residual sample set (e.g., resCb) of a first chroma component and a second residual sample set (e.g., resCr) of a second chroma component, wherein the first residual sample set is based on the difference between a first chroma block and a first prediction block in the first chroma component, and the second residual sample set is based on the difference between a second chroma block and a second prediction block in the second chroma component, wherein the first chroma block and the second chroma block are associated (e.g., part of the same CU); and transmit a signal for... Information about the value set; information about the value set is transmitted by signal; a first intermediate sample (e.g., rec1Cb) is determined based on the value set; a second intermediate sample (e.g., rec1Cr) is determined based on the value set; information about a first residual sample (e.g., resJoint1) is transmitted by signal, wherein the first residual sample is based on the difference between the first chromaticity block of the first chromaticity component and the first intermediate sample; and information about a second residual sample (e.g., resJoint2) is transmitted by signal, wherein the second residual sample is based on the difference between the second chromaticity block of the second chromaticity component and the second intermediate sample.

[0156] Figure 4 This is a block diagram illustrating an example video decoder 300 capable of performing the techniques described herein. Figure 4This disclosure is provided for illustrative purposes and does not limit the scope of the techniques illustrated and described extensively in this disclosure. For illustrative purposes, this disclosure describes a video decoder 300 based on the techniques of JEM, VVC, and HEVC. However, the techniques of this disclosure can be implemented by video decoding devices configured for other video decoding standards.

[0157] exist Figure 4 In the example, the video decoder 300 includes a decoded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 134. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 134 can be implemented in one or more processors or in processing circuitry. For example, units of the video decoder 300 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0158] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-frame prediction unit 318. The prediction processing unit 304 may include an addition unit that performs predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0159] CPB memory 320 can store video data to be decoded by components of video decoder 300, such as encoded video bitstreams. For example, it can be stored from computer-readable medium 110 ( Figure 1The video data stored in the CPB memory 320 is obtained. The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory 320 may store video data other than the syntax elements of the decoded picture, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores decoded pictures that the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and DPB 314 may be formed from any of a variety of memory devices, such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices. The CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.

[0160] Alternatively or concurrently, in some examples, the video decoder 300 can be derived from the memory 120 ( Figure 1 The decoded video data is retrieved. In other words, memory 120 can utilize CPB memory 320 to store data, as discussed above. Similarly, when some or all of the functions of video decoder 300 are implemented using software to be executed by the processing circuitry of video decoder 300, memory 120 can store instructions to be executed by video decoder 300.

[0161] It shows Figure 4 The various units shown aid in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to... Figure 3 Fixed-function circuits refer to circuits that provide a specific function and are pre-configured in terms of the operations they can perform. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in terms of the operations they can perform. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by instructions in the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.

[0162] The video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or a programmable core formed from programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executing on the programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0163] Entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to regenerate syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on syntax elements extracted from the bitstream.

[0164] Typically, the video decoder 300 reconstructs the image on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").

[0165] Entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients defining the quantized transform coefficient block, as well as transform information (such as quantization parameters (QP) and / or transform mode indications). Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the quantization level, and similarly, determine the inverse quantization level to be applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bitwise left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 can thus form a transform coefficient block including the transform coefficients.

[0166] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply the inverse DCT, inverse integer transform, inverse Karhunen-Loeve transform (KLT), inverse rotation transform, inverse direction transform, or another inverse transform to the transform coefficient block.

[0167] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-frame predicted, the motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements may indicate the reference picture from which the reference block is to be retrieved in the DPB 314, and a motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can typically be configured with respect to the motion compensation unit 224 ( Figure 3The method described is basically similar to the way the inter-frame prediction process is performed.

[0168] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, then intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, intra-prediction unit 318 can typically be configured with respect to intra-prediction unit 226 ( Figure 3 The intra-prediction process is performed in a manner substantially similar to that described above. The intra-prediction unit 318 can retrieve data from neighboring samples of the current block from the DPB 314.

[0169] Reconstruction unit 310 can reconstruct the current block using the prediction block and the residual block. For example, reconstruction unit 310 can reconstruct the current block by adding the samples of the residual block to the corresponding samples of the prediction block.

[0170] Filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce block artifacts along the edges of the reconstructed block. The operation of filter unit 312 is not necessarily performed in all examples.

[0171] The video decoder 300 can store the reconstructed blocks in the DPB 314. For example, in an example where the filter unit 312 is not operated, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the filter unit 312 is operated, the filter unit 312 can store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 can provide reference information (such as samples of the current image for intra-frame prediction and samples of previously decoded images for subsequent motion compensation) to the prediction processing unit 304. Furthermore, the video decoder 300 can output decoded images (e.g., decoded video) from the DPB 314 for use in applications such as... Figure 1 The subsequent presentation on the display device 118.

[0172] In this manner, video decoder 300 represents an example of a video decoding device, which includes: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: receive information for a first residual sample for a first chroma component, wherein the first residual sample is based on the difference between a first chroma block and a first prediction block of the first chroma component; determine intermediate samples based on the first residual sample; receive information for a second residual sample for a second chroma component, wherein the second residual sample is based on the difference between a second chroma block and an intermediate sample of the second chroma component, and wherein the second chroma block and the first chroma block are associated together; reconstruct the first chroma block based on the first residual sample and the first prediction block; and reconstruct the second chroma block based on the second residual sample and the intermediate sample.

[0173] The processing unit of the video decoder 300 can also be configured to: receive information for a value set, wherein the value set is generated based on a first residual sample set (e.g., resCb) of a first chroma block of a first chroma component and a second residual sample set (e.g., resCr) of a second chroma block of a second chroma component, wherein the first chroma block and the second chroma block are associated (e.g., part of the same CU); determine a first intermediate sample (e.g., rec1Cb) based on the value set; determine a second intermediate sample (e.g., rec1Cr) based on the value set; receive information for a first residual sample (e.g., resJoint1), wherein the first residual sample is based on the difference between the first chroma block of the first chroma component and the first intermediate sample; receive information for a second residual sample (e.g., resJoint2), wherein the second residual sample is based on the difference between the second chroma block of the second chroma component and the second intermediate sample; reconstruct a first chroma block based on the first residual sample and the first intermediate sample; and reconstruct a second chroma block based on the second residual sample and the second intermediate sample.

[0174] Figure 5 This is a flowchart illustrating an example method for encoding the current block. The current block may include the current CU. Although regarding video encoder 200 ( Figure 1 and Figure 3 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 5 Similar to the method.

[0175] In this example, the video encoder 200 initially predicts the current block (350). For example, the video encoder 200 may form a predicted block for the current block. Then, the video encoder 200 may calculate residual information for the current block (352). To calculate the residual information, the video encoder 200 may calculate resJointC and resJointC2 for technique one as described above (e.g., based on determining rec1Cr or rec1Cb), or the video encoder 200 may calculate resJointC, resJointC1, and resJointC2 for technique two as described above (e.g., based on determining rec1Cr and rec1Cb).

[0176] Optionally (e.g., when lossless decoding is not required), the video encoder 200 can then transform and quantize the residual information (354). The video encoder 200 can scan the residual information (if arranged in blocks) (356). During or after scanning, the video encoder 200 can entropy encode the residual information (358). For example, the video encoder 200 can use CAVLC or CABAC to encode the residual information. The video encoder 200 can then output (e.g., transmit as a signal) the entropy-encoded residual information (360).

[0177] Figure 6 This is a flowchart illustrating an example method for decoding the current block of video data. The current block may include the current CU. Although regarding video decoder 300 ( Figure 1 and Figure 4 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 6 Similar to the method.

[0178] The video decoder 300 can receive entropy-encoded residual information, such as resJointC and resJointC2 for Technique 1 or resJointC, resJointC1, and resJointC2 for Technique 2 (370). The video decoder 300 can determine intermediate values ​​(e.g., rec1Cb and rec1Cr for Technique 1 and Technique 2 as described above) (372). The video decoder 300 can predict the current block, for example, using an intra-frame or inter-frame prediction mode indicated by the prediction information of the current block (374), to compute a prediction block for the current block. The video decoder 300 can then perform an inverse scan of the residual information (if the residual information is in block form) (376) to create a block of residual information. Where lossless decoding is not required and is therefore shown in dashed lines, the video decoder 300 can then perform inverse quantization and inverse transform on the transform coefficients to produce residual information (378). Finally, the video decoder 300 can decode the current block by combining the prediction block and the residual information (380).

[0179] Figure 7 This is a flowchart illustrating an example method for encoding the current block. The current block can be the current CU, which includes a first chroma component, a second chroma component, and potentially other components (such as a luma component). For example, the current block can be decoded in one or both of a lossless decoding mode or a chroma residual joint decoding mode. Although regarding video encoder 200 ( Figure 1 and Figure 3 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 7 Similar to the method.

[0180] For a block of video data with respect to its first chroma component, the video decoder 300 determines a first residual sample corresponding to the difference between the first chroma block and the first predicted block of the first chroma component (400). The video decoder 300 determines intermediate reconstructed samples based on the first residual sample and the second predicted block of the second chroma component (402). The video decoder 300 determines a second residual sample for the second chroma component of the block of video data, wherein the second residual sample corresponds to the difference between the second chroma block and the intermediate reconstructed sample of the second chroma component (404). The video decoder 300 outputs syntax elements representing the first and second residual samples in the bitstream of the encoded video data (406). For example, the video decoder 300 may output a bitstream for storage or transmission.

[0181] Figure 8This is a flowchart illustrating an example method for decoding the current block of video data. The current block can be the current CU, which includes a first chroma component, a second chroma component, and potentially other components (such as a luma component). For example, the current block can be decoded in one or both of a lossless decoding mode or a chroma residual joint decoding mode. Although regarding video decoder 300 ( Figure 1 and Figure 4 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 8 Similar to the method.

[0182] For a block of video data with respect to its first chroma component, the video decoder 300 receives information for a first residual sample, the first residual sample corresponding to the difference (420) between the first chroma block and the first predicted block of the first chroma component. The video decoder 300 may use intra-frame prediction, inter-frame prediction, or some other such prediction mode to obtain the first predicted block.

[0183] The video decoder 300 determines intermediate reconstructed samples (422) based on the first residual sample. To determine the intermediate reconstructed samples, the video decoder 300 may, for example, determine the intermediate reconstructed samples based on the first residual sample and the second prediction block of the second chroma component. To determine the intermediate reconstructed samples, the video decoder 300 may, for example, determine the intermediate reconstructed samples according to the patterns in Tables 3, 3.1 or 4-6 above, for example, rec1Cr[x][y] as described above.

[0184] For the second chroma component of a block of video data, the video decoder 300 receives information for a second residual sample, which corresponds to the difference (424) between the second chroma block in the second chroma component and an intermediate reconstructed sample. Since the intermediate reconstructed sample can usually be relatively close to the actual value of the second chroma block, the second residual sample can usually be decoded using relatively few bits.

[0185] The video decoder 300 reconstructs a first chroma block (426) based on a first residual sample and a first prediction block. To reconstruct the first chroma block based on the first residual sample and the first prediction block, the video decoder 300 may, for example, add the first residual sample to the first prediction block. The video decoder 300 reconstructs a second chroma block (428) based on a second residual sample and an intermediate reconstructed sample. To reconstruct the second chroma block based on the second residual sample and the intermediate reconstructed sample, the video decoder 300 may, for example, add the value of the second residual sample to the corresponding value of the intermediate reconstructed sample. In the case where the video decoder 300 implements lossless decoding, the reconstructed first chroma block and the reconstructed second chroma block can accurately match the corresponding original first chroma block and second chroma block without any kind of filtering operation.

[0186] The video decoder 300 outputs decoded video data (430) including a reconstructed first chroma block and a reconstructed second chroma block. For example, the video decoder 300 may output the decoded video for display, output the decoded video data for long-term storage in a non-volatile storage medium, or output the decoded video data to a buffer or other short-term memory for use in decoding subsequent images of the video.

[0187] The following terms describe aspects of the video encoder 200 and / or the video decoder 300, as well as the techniques that can be implemented by the video encoder 200 and / or the video decoder 300.

[0188] Aspect 1: A method for decoding video data includes: receiving information for a first residual sample for a first chroma component, wherein the first residual sample is based on the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; determining an intermediate sample based on the first residual sample; receiving information for a second residual sample for a second chroma component, wherein the second residual sample is based on the difference between a second chroma block of the second chroma component and the intermediate sample, wherein the second chroma block and the first chroma block are associated together; reconstructing the first chroma block based on the first residual sample and the first prediction block; and reconstructing the second chroma block based on the second residual sample and the intermediate sample.

[0189] Aspect 2: According to the method of aspect 1, wherein the first residual sample represents the difference between the first chromaticity block of the first chromaticity component and the first prediction block of the first chromaticity component.

[0190] Aspect 3: The method according to any one of Aspects 1 and 2, wherein determining the intermediate sample comprises: determining the intermediate sample based on the first residual sample and the second prediction block of the second chromaticity component.

[0191] Aspect 4: The method according to any one of Aspects 1-3, wherein the first chroma block and the second chroma block have the same decoding unit (CU).

[0192] Aspect 5: The method according to any one of Aspects 1-4, wherein determining the intermediate sample includes one of the following: for the first mode, rec1Cr[x][y] = Clip(predCr[x][y] + CSign * resJointC[x][y] >> 1) or Clip(predCr[x][y] + CSign * k1 * resJointC[x][y]), where rec1Cr[x][y] represents the intermediate sample, Clip is a clipping operation, predCr[x][y] represents the sample of the second prediction block of the second chromaticity component, Csign is the sign value, resJointC[x][y] represents the sample of the first residual sample, >> is a right shift or division by 2 operation, and k1 is a pre-stored or signal-transmitted parameter; for the second mode, rec1Cr[x][y] = Clip(predCr[x][y] + CSign * resJointC[x][y]), where, rec1Cr[x][y] represents the intermediate sample, Clip is a clipping operation, predCr[x][y] represents the sample of the second prediction block of the second chromaticity component, CSign is the sign value, and resJointC[x][y] represents the sample of the first residual sample; or for the third mode, rec1Cb[x][y] = Clip(predCb[x][y] + CSign * resJointC[x][y] >> 1) or Clip(predCb[x][y] + CSign * k2 * resJointC[x][y]), where rec1Cr[x][y] represents the intermediate sample, Clip is a clipping operation, predCb[x][y] represents the sample of the second prediction block of the second chromaticity component, CSign is the sign value, resJointC[x][y] represents the sample of the first residual sample, >> is a right shift or division by two operation, and k2 is a pre-stored or signaled parameter.

[0193] Aspect 6: A method for encoding video data includes: signaling information for a first residual sample for a first chroma component, wherein the first residual sample is based on the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; determining an intermediate sample based on the first residual sample; and signaling information for a second residual sample for a second chroma component, wherein the second residual sample is based on the difference between a second chroma block of the second chroma component and the intermediate sample, and wherein the second chroma block and the first chroma block are associated together.

[0194] Aspect 7: According to the method of aspect 6, wherein the first residual sample represents the difference between the first chromaticity block of the first chromaticity component and the first prediction block of the first chromaticity component.

[0195] Aspect 8: The method according to any one of Aspects 6 and 7, wherein determining the intermediate sample comprises: determining the intermediate sample based on a second prediction block of the first residual sample and the second chromaticity component.

[0196] Aspect 9: The method according to any one of Aspects 6-8, wherein the first chroma block and the second chroma block have the same decoding unit (CU).

[0197] Aspect 10: The method according to any one of Aspects 6-9, wherein determining the intermediate sample includes one of the following: for a first mode, rec1Cr[x][y] = Clip(predCr[x][y] + CSign * resJointC[x][y] >> 1) or Clip(predCr[x][y] + CSign * k1 * resJointC[x][y]), where rec1Cr[x][y] represents the intermediate sample, Clip is a clipping operation, predCr[x][y] represents a sample of the second prediction block of the second chromaticity component, Csign is the sign value, resJointC[x][y] represents a sample of the first residual sample, >> is a right shift or division by 2 operation, and k1 is a pre-stored or signal-transmitted parameter; for a second mode, rec1Cr[x][y] = Clip(predCr[x][y] + CSign * resJointC[x][y]), where, rec1Cr[x][y] represents the intermediate sample, Clip is a clipping operation, predCr[x][y] represents the sample of the second prediction block of the second chromaticity component, CSign is the sign value, and resJointC[x][y] represents the sample of the first residual sample; or for the third mode, rec1Cb[x][y] = Clip(predCb[x][y] + CSign * resJointC[x][y] >> 1) or Clip(predCb[x][y] + CSign * k2 * resJointC[x][y]), where rec1Cr[x][y] represents the intermediate sample, Clip is a clipping operation, predCb[x][y] represents the sample of the second prediction block of the second chromaticity component, CSign is the sign value, resJointC[x][y] represents the sample of the first residual sample, >> is a right shift or division by two operation, and k2 is a pre-stored or signaled parameter.

[0198] Aspect 11: A method for decoding video data includes: receiving information for a set of values, wherein the set of values ​​is generated based on a first residual sample set of a first chroma block of a first chroma component and a second residual sample set of a second chroma block of a second chroma component, wherein the first chroma block and the second chroma block are associated; determining a first intermediate sample based on the set of values; determining a second intermediate sample based on the set of values; receiving information for the first residual sample, wherein the first residual sample is based on the difference between the first chroma block of the first chroma component and the first intermediate sample; receiving information for the second residual sample, wherein the second residual sample is based on the difference between the second chroma block of the second chroma component and the second intermediate sample; reconstructing the first chroma block based on the first residual sample and the first intermediate sample; and reconstructing the second chroma block based on the second residual sample and the second intermediate sample.

[0199] Aspect 12: According to the method of aspect 11, determining the first intermediate sample includes: determining the first intermediate sample based on the value set and a first prediction block of the first color component, and wherein determining the second intermediate sample includes: determining the second intermediate sample based on the value set and a second prediction block of the second color component.

[0200] Aspect 13: The method according to any one of aspects 11 and 12, wherein the first chroma block and the second chroma block have the same decoding unit (CU).

[0201] Aspect 14: The method according to any one of Aspects 11-13, wherein determining the first intermediate sample based on the set of values ​​and determining the second intermediate sample based on the set of values ​​comprises one of the following: for the first mode, rec1Cb[x][y] = Clip(predCb[x][y] + resJointC[x][y]) and rec1Cr[x][y] = Clip(predCr[x][y] + CSign * resJointC[x][y] >> 1) or rec1Cr[x][y] = Clip(predCr[x][y] + CSign * k1 * resJointC[x][y]), wherein rec 1Cb[x][y] represents the first intermediate sample, Clip is the clipping operation, predCb[x][y] represents the first prediction block, resJointC[x][y] represents the set of values, rec1Cr[x][y] represents the second intermediate sample, predCr[x][y] represents the second prediction block, CSign is the sign value, >> is the right shift or division by 2 operation, and k1 is a pre-stored or signaled parameter; for the second mode, rec1Cb[x][y] = Clip(predCb[x][y] + resJointC[x][y]) and rec1Cr[x][y] = Clip(predCr[x][y] + CSign n*resJointC[x][y]), where rec1Cb[x][y] represents the first intermediate sample, Clip is the clipping operation, predCb[x][y] represents the first prediction block, resJointC[x][y] represents the set of values, rec1Cr[x][y] represents the second intermediate sample, predCr[x][y] represents the second prediction block, and CSign is the sign value; or for the third mode, rec1Cb[x][y] = Clip(predCb[x][y] + CSign * resJointC[x][y] >> 1) and rec1Cr[x][y] = Clip(predCr[x][y] * resJointC[x][y] >> 1) and rec1Cr[x][y] = Clip(predCr[x][y] * resJointC[x][y] >> 1) rec1Cb[x][y] = Clip(predCb[x][y] + CSign * resJointC[x][y]) or rec1Cb[x][y] = Clip(predCb[x][y] + CSign * k2 * resJointC[x][y]), where rec1Cb[x][y] represents the first intermediate sample, Clip is a clipping operation, predCb[x][y] represents the first prediction block, resJointC[x][y] represents the set of values, rec1Cr[x][y] represents the second intermediate sample, predCr[x][y] represents the second prediction block, CSign is the sign value, >> is a right shift or division by 2 operation, and k2 is a pre-stored or signaled parameter.

[0202] Aspect 15: A method for encoding video data includes: generating a value set based on a first residual sample set of a first chroma component and a second residual sample set of a second chroma component, wherein the first residual sample set is based on the difference between a first chroma block and a first prediction block of the first chroma component, and the second residual sample set is based on the difference between a second chroma block and a second prediction block of the second chroma component, wherein the first chroma block and the second chroma block are associated; signaling information for the value set; determining a first intermediate sample based on the value set; determining a second intermediate sample based on the value set; signaling information for the first residual sample, wherein the first residual sample is based on the difference between the first chroma block and the first intermediate sample of the first chroma component; and signaling information for the second residual sample, wherein the second residual sample is based on the difference between the second chroma block and the second intermediate sample of the second chroma component.

[0203] Aspect 16: According to the method of aspect 15, determining the first intermediate sample includes: determining the first intermediate sample based on the value set and a first prediction block of the first color component, and wherein determining the second intermediate sample includes: determining the second intermediate sample based on the value set and a second prediction block of the second color component.

[0204] Aspect 17: The method according to any one of Aspects 15 and 16, wherein the first chroma block and the second chroma block have the same decoding unit (CU).

[0205] Aspect 18: The method according to any one of Aspects 15-17, wherein generating the set of values ​​comprises one of the following: for a first mode, resJointC[x][y] = (4*resCb[x][y] + 2*CSign*resCr[x][y]) / 5 or resJointC[x][y] = (4*resCb[x][y] + 4*k1*CSign*resCr[x][y]) / 5, wherein resJointC[x][y] represents the set of values, resCb[x][y] represents the first set of residual samples, CSign is a sign value, resCr[x][y] represents the second set of residual samples, and k1 is a pre-stored or signaled value; for a second mode, resJointC[x][y] = (resCb[x][y] + CSign*resCr[x][y]) / 2 or resJointC[x][y] = (resCb[x][y] + CSign*resCr[x][y]) / 2. y]+CSign*resCr[x][y]) / 2, where resJointC[x][y] represents the set of values, resCb[x][y] represents the first residual sample set, CSign is the sign value, resCr[x][y] represents the second residual sample set, and k1 is a pre-stored or signaled value; or for the third mode, resJointC[x][y]=(4*resCr[x][y]+2*CSign*resCb[x][y]) / 5 or resJointC[x][y]=(4*resCr[x][y]+4*k2*CSign*resCb[x][y]) / 5, where resJointC[x][y] represents the set of values, resCr[x][y] represents the first residual sample set, CSign is the sign value, resCb[x][y] represents the second residual sample set, and k2 is a pre-stored or signaled value.

[0206] Aspect 19: The method according to any one of Aspects 15-18, wherein determining the first intermediate sample based on the set of values ​​and determining the second intermediate sample based on the set of values ​​comprises one of the following: for the first mode, rec1Cb[x][y] = Clip(predCb[x][y] + resJointC[x][y]) and rec1Cr[x][y] = Clip(predCr[x][y] + CSign * resJointC[x][y] >> 1) or rec1Cr[x][y] = Clip(predCr[x][y] + CSign * k1 * resJointC[x][y]), wherein rec 1Cb[x][y] represents the first intermediate sample, Clip is the clipping operation, predCb[x][y] represents the first prediction block, resJointC[x][y] represents the set of values, rec1Cr[x][y] represents the second intermediate sample, predCr[x][y] represents the second prediction block, CSign is the sign value, >> is the right shift or division by 2 operation, and k1 is a pre-stored or signaled parameter; for the second mode, rec1Cb[x][y] = Clip(predCb[x][y] + resJointC[x][y]) and rec1Cr[x][y] = Clip(predCr[x][y] + CSign n*resJointC[x][y]), where rec1Cb[x][y] represents the first intermediate sample, Clip is the clipping operation, predCb[x][y] represents the first prediction block, resJointC[x][y] represents the set of values, rec1Cr[x][y] represents the second intermediate sample, predCr[x][y] represents the second prediction block, and CSign is the sign value; or for the third mode, rec1Cb[x][y] = Clip(predCb[x][y] + CSign * resJointC[x][y] >> 1) and rec1Cr[x][y] = Clip(predCr[x][y] * resJointC[x][y] >> 1) and rec1Cr[x][y] = Clip(predCr[x][y] * resJointC[x][y] >> 1) rec1Cb[x][y] = Clip(predCb[x][y] + CSign * resJointC[x][y]) or rec1Cb[x][y] = Clip(predCb[x][y] + CSign * k2 * resJointC[x][y]), where rec1Cb[x][y] represents the first intermediate sample, Clip is a clipping operation, predCb[x][y] represents the first prediction block, resJointC[x][y] represents the set of values, rec1Cr[x][y] represents the second intermediate sample, predCr[x][y] represents the second prediction block, CSign is the sign value, >> is a right shift or division by 2 operation, and k2 is a pre-stored or signaled parameter.

[0207] Aspect 20: An apparatus for decoding video data includes: a memory configured to store video data; and processing circuitry coupled to the memory and configured to perform the method of any one of claims 1-5 and 11-14.

[0208] Aspect 21: An apparatus for encoding video data includes: a memory configured to store video data; and processing circuitry coupled to the memory and configured to perform the method of any one of claims 6-10 and 15-19.

[0209] Aspect 22: The device according to any one of aspects 20 or 21 further includes: a display configured to display decoded video data.

[0210] Aspect 23: The device according to any one of Aspects 20-22, wherein the device includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0211] Aspect 24: A computer-readable storage medium having instructions thereon stored thereon, the instructions, when executed, causing one or more processors to perform any one of aspects 1-5 or 11-14.

[0212] Aspect 25: An apparatus for decoding video data, the apparatus comprising a unit for performing the method of any one of aspects 1-5 or 11-14.

[0213] It should be recognized that, based on the examples, certain actions or events of any of the techniques described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for implementing the techniques). Furthermore, in some examples, actions or events may be performed concurrently rather than sequentially, for example, through multithreaded processing, interrupt handling, or multiple processors.

[0214] In one or more examples, the described functionality can be implemented using hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium can include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures used to implement the techniques described in this disclosure. Computer program products can include computer-readable media.

[0215] For example, rather than limiting, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures, and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave). However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically copy data, while optical discs use lasers to optically copy data. The combination of the above items should also be included within the scope of computer-readable media.

[0216] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuitry" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, the techniques can be fully implemented in one or more circuit or logic elements.

[0217] The technologies disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed technologies, but implementation by different hardware units is not necessarily required. Specifically, as described above, various units may be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.

[0218] Various examples have been described. These and other examples are within the scope of the appended claims.

Claims

1. A method for decoding video data, the method comprising: Information is received for a first residual sample of a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; The second prediction block of the second chromaticity component of the block is determined using at least one of intra-frame prediction or inter-frame prediction. An intermediate reconstruction sample is determined based on the first residual sample and the second prediction block, wherein the intermediate reconstruction sample corresponds to the second prediction block of the corresponding scaled sample added to the first residual sample; Information is received for a second residual sample of the second chroma component of the block of video data, wherein the second residual sample corresponds to the difference between the second chroma block of the second chroma component and the intermediate reconstructed sample; The first chromaticity block is reconstructed based on the first residual sample and the first prediction block; The second chroma block is reconstructed based on the second residual sample and the intermediate reconstructed sample; and The output includes decoded video data consisting of the reconstructed first chroma block and the reconstructed second chroma block.

2. The method according to claim 1, wherein, The video data blocks are decoded in a lossless decoding mode.

3. The method according to claim 1, wherein, The corresponding scaled sample of the first residual sample corresponds to the first residual sample multiplied by the sign value and the scaling factor.

4. The method according to claim 1, wherein, The first chroma block and the second chroma block have the same decoding unit (CU).

5. The method according to claim 1, wherein, Reconstructing the second chroma block based on the second residual sample and the intermediate reconstruction sample includes adding the value of the second residual sample to the corresponding value of the intermediate reconstruction sample.

6. The method according to claim 1, further comprising: It is determined that the block of the video data is encoded using a chroma residual joint decoding mode.

7. The method according to claim 1, wherein, Reconstructing the first chroma block based on the first residual sample and the first prediction block includes: adding the first reconstructed chroma block to the first prediction block of the first chroma component to generate the first chroma block; Reconstructing the second chroma block based on the second residual sample and the intermediate reconstruction sample includes adding the value of the second residual sample to the corresponding value of the intermediate reconstruction sample.

8. The method according to claim 1, wherein, Determining the intermediate reconstruction sample includes determining the intermediate reconstruction sample according to one of the following equations: rec1Cr[x][y]=Clip(predCr[x][y]+CSign*resJointC[x][y]>>1) or Clip(predCr[x][y]+CSign*k1*resJointC[x][y]), where, rec1Cr[x][y] represents the intermediate reconstructed sample. Clip is a clipping operation. predCr[x][y] represents the sample of the second prediction block of the second chromaticity component. Csign is a sign value. resJointC[x][y] represents the sample of the first residual sample. >> is a right shift or division by 2 operation, and k1 is a parameter that is either pre-stored or sent by a signal.

9. The method according to claim 1, wherein, Determining the intermediate reconstruction samples includes determining the intermediate reconstruction samples according to the following equation: rec1Cr[x][y]=Clip(predCr[x][y]+CSign*resJointC[x][y]), where, rec1Cr[x][y] represents the intermediate reconstructed sample. Clip is a clipping operation. predCr[x][y] represents the sample of the second prediction block of the second chromaticity component. Csign is the sign value, and resJointC[x][y] represents the sample of the first residual sample.

10. The method according to claim 1, wherein, Determining the intermediate reconstruction sample includes determining the intermediate reconstruction sample according to one of the following equations: rec1Cb[x][y]=Clip(predCb[x][y]+CSign*resJointC[x][y]>>1) or Clip(predCb[x][y]+CSign*k2*resJointC[x][y]), where, rec1Cr[x][y] represents the intermediate reconstructed sample. Clip is a clipping operation. predCb[x][y] represents a sample of the second prediction block for the second chromaticity component. CSign is the sign value. resJointC[x][y] represents the sample of the first residual sample. >> is a right shift or division by two operation, and k2 is a parameter that is either pre-stored or sent by a signal.

11. A method for encoding video data, the method comprising: Determine a first residual sample of a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; The second prediction block of the second chromaticity component of the block is determined using at least one of intra-frame prediction or inter-frame prediction. An intermediate reconstruction sample is determined based on the second prediction block of the first residual sample and the second chromaticity component, wherein the intermediate reconstruction sample corresponds to the second prediction block of the corresponding scaled sample added to the first residual sample; Determine a second residual sample of the second chroma component of the block of video data, wherein the second residual sample corresponds to the difference between the second chroma block of the second chroma component and the intermediate reconstructed sample; and Output syntax elements representing the first residual sample and the second residual sample in the bitstream of encoded video data.

12. The method according to claim 11, wherein, The video data blocks are decoded in a lossless decoding mode.

13. The method according to claim 11, wherein, The first chroma block and the second chroma block have the same decoding unit (CU).

14. The method according to claim 11, wherein, Determining the intermediate reconstruction sample based on the second prediction block of the first residual sample and the second chromaticity component includes adding the value of the first residual sample to the value of the sample of the second prediction block.

15. The method of claim 11, further comprising: It is determined that the block of the video data is encoded using a chroma residual joint decoding mode.

16. An apparatus for decoding video data, the apparatus comprising: A memory configured to store the video data; One or more processors, which are implemented in a circuit and configured to: Information is received for a first residual sample of a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; The second prediction block of the second chromaticity component of the block is determined using at least one of intra-frame prediction or inter-frame prediction. An intermediate reconstruction sample is determined based on the first residual sample and the second prediction block, wherein the intermediate reconstruction sample corresponds to the second prediction block of the corresponding scaled sample added to the first residual sample; Receive information about a second residual sample for the second chroma component of the block of video data, wherein the second residual sample corresponds to the difference between the second chroma block of the second chroma component and the intermediate reconstructed sample; The first chromaticity block is reconstructed based on the first residual sample and the first prediction block; The second chroma block is reconstructed based on the second residual sample and the intermediate reconstructed sample; and The output includes decoded video data consisting of the reconstructed first chroma block and the reconstructed second chroma block.

17. The device according to claim 16, wherein, The video data blocks are decoded in a lossless decoding mode.

18. The device according to claim 16, wherein, The corresponding scaled sample of the first residual sample corresponds to the first residual sample multiplied by the sign value and the scaling factor.

19. The device according to claim 16, wherein, The first chroma block and the second chroma block have the same decoding unit (CU).

20. The device according to claim 16, wherein, In order to reconstruct the second chroma block based on the second residual sample and the intermediate reconstruction sample, the one or more processors are further configured to add the value of the second residual sample to the corresponding value of the intermediate reconstruction sample.

21. The device according to claim 16, wherein, The one or more processors are further configured to: It is determined that the block of the video data is encoded using a chroma residual joint decoding mode.

22. The device according to claim 16, wherein, In order to reconstruct the first chroma block based on the first residual sample and the first prediction block, the one or more processors are further configured to: add the first reconstructed chroma block to the first prediction block of the first chroma component to generate the first chroma block; In order to reconstruct the second chroma block based on the second residual sample and the intermediate reconstruction sample, the one or more processors are further configured to add the value of the second residual sample to the corresponding value of the intermediate reconstruction sample.

23. The device according to claim 16, wherein, To determine the intermediate reconstructed samples, the one or more processors are further configured to determine the intermediate reconstructed samples according to one of the following equations: rec1Cr[x][y]=Clip(predCr[x][y]+CSign*resJointC[x][y]>>1) or Clip(predCr[x][y]+CSign*k1*resJointC[x][y]), where, rec1Cr[x][y] represents the intermediate reconstructed sample. Clip is a clipping operation. predCr[x][y] represents the sample of the second prediction block of the second chromaticity component. Csign is a sign value. resJointC[x][y] represents the sample of the first residual sample. >> is a right shift or division by 2 operation, and k1 is a parameter that is either pre-stored or sent by a signal.

24. The device according to claim 16, wherein, To determine the intermediate reconstructed samples, the one or more processors are further configured to determine the intermediate reconstructed samples according to one of the following equations: rec1Cr[x][y]=Clip(predCr[x][y]+CSign*resJointC[x][y]), where, rec1Cr[x][y] represents the intermediate reconstructed sample. Clip is a clipping operation. predCr[x][y] represents the sample of the second prediction block of the second chromaticity component. Csign is the sign value, and resJointC[x][y] represents the sample of the first residual sample.

25. The device according to claim 16, wherein, To determine the intermediate reconstructed samples, the one or more processors are further configured to determine the intermediate reconstructed samples according to one of the following equations: rec1Cb[x][y]=Clip(predCb[x][y]+CSign*resJointC[x][y]>>1) or Clip(predCb[x][y]+CSign*k2*resJointC[x][y]), where, rec1Cr[x][y] represents the intermediate reconstructed sample. Clip is a clipping operation. predCb[x][y] represents a sample of the second prediction block for the second chromaticity component. CSign is the signal value. resJointC[x][y] represents the sample of the first residual sample. >> is a right shift or division by two operation, and k2 is a parameter that is either pre-stored or sent by a signal.

26. The device according to claim 16, further comprising: A display configured to show decoded video data.

27. The device according to claim 16, wherein, The device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.

28. An apparatus for encoding video data, the apparatus comprising: A memory configured to store the video data; One or more processors, which are implemented in a circuit and configured to: Determine a first residual sample of a first chroma component of a block of video data, wherein the first residual sample corresponds to the difference between a first chroma block of the first chroma component and a first prediction block of the first chroma component; The second prediction block of the second chromaticity component of the block is determined using at least one of intra-frame prediction or inter-frame prediction. An intermediate reconstruction sample is determined based on the second prediction block of the first residual sample and the second chromaticity component, wherein the intermediate reconstruction sample corresponds to the second prediction block of the corresponding scaled sample added to the first residual sample; Determine a second residual sample of the second chroma component of the block of video data, wherein the second residual sample corresponds to the difference between the second chroma block of the second chroma component and the intermediate reconstructed sample; and Output syntax elements representing the first residual sample and the second residual sample in the bitstream of encoded video data.

29. The device according to claim 28, wherein, The video data blocks are decoded in a lossless decoding mode.

30. The device according to claim 28, wherein, In order to determine the intermediate reconstructed sample based on the first residual sample and the second prediction block of the second chromaticity component, the one or more processors are further configured to add the value of the first residual sample to the value of the sample of the second prediction block.

Citation Information

Patent Citations

  • Image processing device and method

    WO2019054200A1