Method and apparatus for video coding in 4:4:4 chroma format

The method for decoding video data in the 4:4:4 chroma format improves encoding efficiency by using an adaptive color space transform to reduce redundancy, addressing the challenge of high redundancy in the 4:4:4 format.

JP7684448B2Active Publication Date: 2025-05-27BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024000363
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-23
Filing Date
2024-01-04
Publication Date
2025-05-27
Estimated Expiration
2040-09-23

AI Technical Summary

Technical Problem

The challenge in video encoding is to improve the encoding efficiency of video data in the 4:4:4 chroma format, which has higher redundancy compared to other chroma formats like 4:2:0 and 4:2:2, thereby affecting compression efficiency.

Method used

The proposed solution involves a method for decoding video data that includes receiving video data corresponding to an encoding unit from a bitstream, determining if the encoding unit has a non-zero residual, and if so, checking if it is encoded using an adaptive color space transform (ACT). Based on this, the method decides whether to perform an inverse ACT on the video data.

Benefits of technology

This approach enhances the encoding efficiency of video data in the 4:4:4 chroma format by effectively utilizing the adaptive color space transform to reduce redundancy among color components, thereby improving compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007684448000009
    Figure 0007684448000009
  • Figure 0007684448000010
    Figure 0007684448000010
  • Figure 0007684448000011
    Figure 0007684448000011
Patent Text Reader

Abstract

To provide a system and a method for improving the efficiency of coding a video which is coded in a specific saturation format.SOLUTION: The method includes the steps of: receiving, from a bit stream, video data which corresponds to a coded unit coded in an inter-prediction mode or an intra-block copy mode; receiving a second syntax element showing whether the coded unit is coded by using an adaptive color space conversion (ACT) from the video data according to that the first syntax element received by the video data shows whether the coded unit has a residual other than 0 and that the first syntax element has a value other than 0; allocating the value of 0 to the second syntax element according to the determination that the first syntax element has the value of 0; and determining whether to execute reverse adaptive color space conversion (ACT) on the video data of the coded unit according to the value of the second syntax element.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 904,539, entitled "METHODS AND APPARATUS OF VIDEO CODING IN 4:4:4 CHROMA FORMAT", filed on September 23, 2019, which is incorporated herein by reference in its entirety.

[0002] This application generally relates to the encoding and compression of video data, and more particularly, to methods and systems for improving the encoding efficiency of video encoded in 4:4:4 chroma format.

Background Art

[0003] Digital video is supported by various electronic devices such as digital TVs, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. The electronic devices implement video compression / decompression standards such as those defined by the MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) standards to transmit, receive, encode, decode, and / or store digital video data. Video compression typically involves performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in the video data. In block-based video coding, a video frame is divided into one or more slices, and each slice has a plurality of video blocks, which may also be referred to as coding tree units (CTUs). Each CTU may contain one coding unit (CU) or may be recursively divided into smaller CUs until a predetermined minimum CU size is reached. Each CU (also called a leaf CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be encoded in either intra, inter, or IBC mode. Video blocks in an intra-coded (I) slice of a video frame are encoded using spatial prediction with respect to reference samples in adjacent blocks within the same video frame. Video blocks in an inter-coded (P or B) slice of a video frame may use spatial prediction with respect to reference samples in adjacent blocks within the same video frame or temporal prediction with respect to reference samples in other previous and / or subsequent reference video frames.

[0004] Spatial or temporal prediction based on previously encoded reference blocks, such as neighboring blocks, results in a prediction block for the current video block to be encoded. The process of finding the reference blocks may be achieved by a block - matching algorithm. Residual data representing the pixel difference between the current block to be encoded and the prediction block is referred to as a residual block or prediction error. Inter - encoded blocks are encoded according to a motion vector indicating a reference block in a reference frame forming the prediction block and the residual block. The process of determining the motion vector is typically referred to as motion estimation. Intra - encoded blocks are encoded according to an intra - prediction mode and the residual block. For further compression, the residual block may be transformed from the pixel domain to a transform domain, such as the frequency domain, resulting in residual transform coefficients, which may then be quantized. The quantized transform coefficients, initially arranged in a two - dimensional array, are scanned to produce a one - dimensional vector of transform coefficients, which may then be entropy - encoded into a video bitstream to achieve further compression.

[0005] The encoded video bitstream is then stored in a computer - readable storage medium (e.g., flash memory) so as to be accessed by another electronic device having digital video capabilities or transmitted directly to the electronic device, either wired or wirelessly. The electronic device then performs video restoration (a process opposite to the above - mentioned video compression) by, for example, parsing the encoded video bitstream to obtain syntax elements from the bitstream and reconstructing digital video data from the encoded video bitstream to its original form based at least in part on the syntax elements obtained from the bitstream, and rendering the reconstructed digital video data on a display of the electronic device.

[0006] As digital video quality increases from high definition to 4K×2K or even 8K×4K, the amount of video data to be encoded / decoded increases exponentially. This always poses a challenge regarding how video data can be encoded / decoded more efficiently while maintaining the image quality of the decoded video data.

[0007] Certain video content, such as screen content video, is encoded in a 4:4:4 chroma format where all three components (one luminance component and two chroma components) have the same resolution. The 4:4:4 chroma format contains more redundancy (which is disadvantageous for achieving good compression efficiency) compared to the 4:2:0 chroma format and the 4:2:2 chroma format. However, the 4:4:4 chroma format is still a preferred encoding format for many applications that require high fidelity to maintain color information such as sharp edges in the decoded video. Considering the redundancy present in 4:4:4 chroma format video, there is a basis to believe that significant encoding improvement can be achieved by exploiting the correlation relationships between the three color components (e.g., Y, Cb, and Cr in the YCbCr domain, or G, B, and R in the RGB domain) of 4:4:4 video. Due to these correlation relationships, during the development of the Screen Content Coding (SCC) extension of HEVC, an Adaptive Color Space Transform (ACT) tool is used to exploit the correlation relationships between the three color components.

SUMMARY OF THE INVENTION

PROBLEMS TO BE SOLVED BY THE INVENTION

[0008] This application describes implementations related to encoding and decoding of video data, and more particularly, systems and methods for improving the encoding efficiency of video encoded in a specific chroma format.

MEANS FOR SOLVING THE PROBLEMS

[0009] According to a first aspect of the present application, a method for decoding video data includes receiving video data corresponding to an encoding unit from a bitstream, where the encoding unit is encoded in an inter prediction mode or an intra block copy mode, receiving a first syntax element from the video data, where the first syntax element indicates whether the encoding unit has a non-zero residual, receiving a second syntax element from the video data according to a determination that the first syntax element has a non-zero value, where the second syntax element indicates whether the encoding unit is encoded using an adaptive color space transform (ACT), assigning a value of zero to the second syntax element according to a determination that the first syntax element has a value of zero, and determining whether to perform an inverse ACT on the video data of the encoding unit according to the value of the second syntax element.

[0010] According to a second aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory. The programs, when executed by the one or more processing units, cause the electronic device to execute the method for decoding video data described above.

[0011] According to a third aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by an electronic device having one or more processing units. The programs, when executed by the one or more processing units, cause the electronic device to execute the method for decoding video data described above.

[0012] Included to provide a further understanding of the implementation, the accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate the described implementations and together with the description serve to clarify the underlying principles. Like reference numerals refer to corresponding parts.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 4E

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

[0014] Here, specific implementations will be referred to in detail, and examples thereof are shown in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternative examples may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0015] FIG. 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a destination device 14. Source device 12 and destination device 14 may include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0016] In some implementations, the destination device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may include any wireless or wired communication medium such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be useful in facilitating communication from the source device 12 to the destination device 14.

[0017] In some other implementations, the encoded video data may be transmitted from the output interface 22 to the storage device 32. Thereafter, the encoded video data in the storage device 32 may be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessible data storage media such as a hard drive, a Blu-ray disk, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that may hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data via streaming or download from the storage device 32. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include a web server (e.g., for a website), an FTP server, a network-attached storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data through any standard data connection including a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored in the file server. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0018] As shown in FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include sources such as a video capture device such as a video camera, a video archive including previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, when the video source 18 is a video camera of a security monitoring system, the source device 12 and the destination device 14 may form a camera phone or a video phone. However, the implementations described in this application may generally be applicable to video encoding and may be applied to wireless and / or wired applications.

[0019] Video that is captured, pre-captured, or computer-generated may be encoded by the video encoder 20. The encoded video data may be transmitted directly to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may further (or alternatively) be stored in the storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.

[0020] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem and receive encoded video data via link 16. The encoded video data communicated via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.

[0021] In some implementations, the destination device 14 may include a display device 34 that can be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 may display the decoded video data to the user and may include any of various display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0022] The video encoder 20 and the video decoder 30 may operate according to proprietary or industry standards such as VVC, HEVC, MPEG-4, Part 10, Advanced Video Coding (AVC), or extensions of such standards. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally assumed that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is also generally assumed that the video decoder 30 of the destination device 14 may be configured to decode video data according to any of these current or future standards.

[0023] The video encoder 20 and the video decoder 30 may each be implemented as any of various suitable encoder circuits such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, an electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and execute those instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in their respective devices.

[0024] FIG. 2 is a block diagram showing an exemplary video encoder 20 according to some implementations described in this application. The video encoder 20 may perform intra and inter prediction encoding of video blocks within a video frame. Intra prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence.

[0025] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A deblocking filter (not shown) may be disposed between the adder 62 and the DPB 64 to filter the block boundaries to remove block distortion artifacts from the reconstructed video. In addition to the deblocking filter, a loop filter (not shown) may be used to filter the output of the adder 62. The video encoder 20 may take the form of fixed or programmable hardware units, or may be divided among one or more of the illustrated fixed or programmable hardware units.

[0026] The video data memory 40 may store video data encoded by the components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from the video source 18. The DPB 64 is a buffer that stores reference video data for use when encoding video data by the video encoder 20 (e.g., in an intra or inter prediction coding mode). The video data memory 40 and the DPB 64 may be formed by any of various memory devices. In various examples, the video data memory 40 may be on the same chip as other components of the video encoder 20 or off-chip with respect to those components.

[0027] As shown in FIG. 2, after receiving video data, the partitioning unit 45 in the prediction processing unit 41 partitions the video data into video blocks. This partitioning may include partitioning the video frame into slices, tiles, or other larger coding units (CUs) according to a predefined partitioning structure such as a quadtree structure associated with the video data. The video frame may be divided into a plurality of video blocks (or a set of video blocks referred to as tiles). The prediction processing unit 41 may select one of a plurality of possible prediction coding modes for the current video block, such as one of a plurality of intra prediction coding modes or one of a plurality of inter prediction coding modes, based on error results (e.g., coding rate and distortion level). The prediction processing unit 41 may provide the resulting intra- or inter-prediction coded block to the adder 50 to generate a residual block, and also to the adder 62 to reconstruct the coded block for later use as part of a reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra mode indicators, partitioning information, and other such syntax information to the entropy coding unit 56.

[0028] To select an appropriate intra prediction coding mode for the current video block, the intra prediction processing unit 46 in the prediction processing unit 41 may perform intra prediction coding of the current video block with respect to one or more adjacent blocks in the same frame as the current block to be coded in order to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 in the prediction processing unit 41 perform inter prediction coding of the current video block with respect to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 may perform a plurality of coding paths, for example, to select an appropriate coding mode for each block of the video data.

[0029] In some implementations, the motion estimation unit 42 determines an inter prediction mode for the current video frame by generating a motion vector indicating the displacement of a prediction unit (PU) of a video block in the current video frame relative to a prediction block in a reference video frame according to a predetermined pattern within a sequence of video frames. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion of a video block. The motion vector may indicate, for example, the displacement of the PU of a video block in the current video frame or picture relative to a prediction block in a reference frame (or other coding unit) relative to the current block being coded within the current frame (or other coding unit). The predetermined pattern may specify video frames in the sequence as P-frames or B-frames. The intra BC unit 48 may determine a vector, such as a block vector, for intra BC coding in a similar manner to the determination of the motion vector by the motion estimation unit 42 for inter prediction, or may utilize the motion estimation unit 42 to determine the block vector.

[0030] The prediction block is a block in the reference frame that is considered to closely match the PU of the video block to be coded with respect to the pixel difference that can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some implementations, the video encoder 20 may calculate values for sub-pixel positions of the reference frames stored in the DPB 64. For example, the video encoder 20 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Accordingly, the motion estimation unit 42 may perform motion search for both full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0031] The motion estimation unit 42 calculates a motion vector for a prediction unit (PU) of a video block in an inter-predicted coded frame by comparing the position of a prediction block of a reference frame selected from the first reference frame list (list 0) or the second reference frame list (list 1), each of which identifies one or more reference frames stored in the DPB 64, with the position of the PU. The motion estimation unit 42 transmits the calculated motion vector to the motion compensation unit 44 and then to the entropy coding unit 56.

[0032] Motion compensation performed by the motion compensation unit 44 may involve fetching or generating a prediction block based on the motion vector determined by the motion estimation unit 42. Upon receiving the motion vector for the PU of the current video block, the motion compensation unit 44 may identify the position of the prediction block indicated by the motion vector in one of the reference frame lists, obtain the prediction block from the DPB 64, and transfer the prediction block to the adder 50. The adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by the motion compensation unit 44 from the pixel values of the currently coded video block. The pixel difference values forming the residual video block may include the difference components of luminance or chrominance or both. The motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use by the video decoder 30 when decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector used to identify the prediction block, any flag indicating the prediction mode, or any other syntax information described herein. It should be noted that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes.

[0033] In some implementations, the intra BC unit 48 may generate vectors and fetch prediction blocks in a manner similar to that described above in relation to the motion estimation unit 42 and the motion compensation unit 44, where the prediction blocks are in the same frame as the currently encoded block, and the vectors are referred to as block vectors as opposed to motion vectors. In particular, the intra BC unit 48 may determine the intra prediction mode to be used for encoding the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, between separate encoding paths, and test their performance by rate-distortion analysis. Next, the intra BC unit 48 may select the appropriate intra prediction mode to be used from among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values using rate-distortion analysis for the various tested intra prediction modes and select the intra prediction mode having the best rate-distortion characteristics among the test modes as the appropriate intra prediction mode to be used. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to create the encoded block, along with the bitrate (i.e., the number of bits) used to create the encoded block. The intra BC unit 48 may calculate the ratio from the distortion and rate for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0034] In other examples, the intra BC unit 48 may use all or part of the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra BC prediction, according to the implementations described herein. In any case, for intra block copy, the predicted block may be a block that is considered to closely match the block to be coded, with respect to the pixel differences that may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics, and the identification of the predicted block may include the calculation of values for sub-pixel positions.

[0035] Regardless of whether the predicted block is from the same frame by intra prediction or from a different frame by inter prediction, the video encoder 20 may form a residual video block by subtracting the pixel values of the predicted block from the pixel values of the currently coded video block, thereby forming pixel difference values. The pixel difference values that form the residual video block may include differences in both the luminance and chrominance components.

[0036] As described above, the intra prediction processing unit 46 may perform intra prediction on the current video block as an alternative to inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or intra block copy prediction performed by the intra BC unit 48. In particular, the intra prediction processing unit 46 may determine the intra prediction mode to be used for encoding the current block. To do this, the intra prediction processing unit 46 may, for example, encode the current block using various intra prediction modes between separate encoding paths, and the intra prediction processing unit 46 (or in some examples, the mode selection unit) may select an appropriate intra prediction mode to be used from the tested intra prediction modes. The intra prediction processing unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra prediction mode in the bitstream.

[0037] After the prediction processing unit 41 determines a prediction block for the current video block via either inter prediction or intra prediction, the adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more transform units (TUs) and provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0038] The conversion processing unit 52 may send the resulting conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients in order to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The quantization degree may be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 may then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy encoding unit 56 may perform this scan.

[0039] Following quantization, the entropy encoding unit 56 entropy encodes the quantized conversion coefficients into a video bit stream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding or another entropy coding method or technique. The encoded bit stream may then be transmitted to the video decoder 30 or archived in the storage device 32 for later transmission to the video decoder 30 or acquisition by the video decoder 30. The entropy encoding unit 56 may also entropy encode the motion vectors and other syntax elements for the current video frame being encoded.

[0040] The inverse quantization unit 58 and the inverse conversion processing unit 60 each apply inverse quantization and inverse conversion to reconstruct the residual video block in the pixel domain to generate a reference block for the prediction of other video blocks. As described above, the motion compensation unit 44 may generate a motion-compensated prediction block from one or more reference blocks of the frames stored in the DPB 64. The motion compensation unit 44 may apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.

[0041] The adder 62 adds the reconstructed residual block to the motion-compensated prediction block created by the motion compensation unit 44 in order to create a reference block for storage in the DPB 64. The reference block may then be used as a prediction block by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 to inter-predict another video block in a subsequent video frame.

[0042] FIG. 3 is a block diagram showing an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process generally opposite to the encoding process described above with respect to the video encoder 20 in connection with FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on the motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 may generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.

[0043] In some examples, a unit of the video decoder 30 may be tasked with performing an implementation of the present application. Also, in some examples, an implementation of the present disclosure may be divided among one or more units of the video decoder 30. For example, the intra BC unit 85 may perform an implementation of the present application alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.

[0044] The video data memory 79 may store video data such as an encoded video bit stream that is decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (such as a flash drive or a hard disk). The video data memory 79 may include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bit stream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use when decoding video data by the video decoder 30 (e.g., in an intra or inter prediction encoding mode). The video data memory 79 and the DPB 92 may be formed by any of various memory devices such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive random access memory (RRAM), or other types of memory devices. For purposes of illustration, the video data memory 79 and the DPB 92 are shown as two separate components of the video decoder 30 in FIG. 3. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or by separate memory devices. In some examples, the video data memory 79 may be on the same chip as other components of the video decoder 30 or off-chip with respect to those components.

[0045] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. Video decoder 30 may receive syntax elements at the video frame level and / or at the video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then transfers the motion vectors and other syntax elements to prediction processing unit 81.

[0046] When the video frame is encoded as an intra-predicted (I) frame or for intra-encoded prediction blocks in other types of frames, intra prediction processing unit 84 of prediction processing unit 81 may generate prediction data for video blocks of the current video frame based on the signaled intra prediction mode and reference data from previously decoded blocks of the current frame.

[0047] When the video frame is encoded as an inter-predicted (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 creates one or more prediction blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks may be created from a reference frame in one of the reference frame lists. Video decoder 30 may configure reference frame lists, list 0 and list 1, using default configuration techniques based on the reference frames stored in DPB 92.

[0048] In some examples, when a video block is encoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 creates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within the reconstructed area of the same picture as the current video block defined by the video encoder 20.

[0049] The motion compensation unit 82 and / or the intra BC unit 85 parse the motion vector and other syntax elements to determine prediction information for the video blocks of the current video frame, and then use the prediction information to create a prediction block for the currently decoded video block. For example, the motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode (e.g., intra or inter prediction) used to encode the video blocks of the video frame, an inter prediction frame type (e.g., B or P), configuration information for one or more of the reference frame lists for the frame, a motion vector for each video block encoded with inter prediction of the frame, an inter prediction status for each video block encoded with inter prediction of the frame, and other information for decoding the video blocks in the current video frame.

[0050] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as flags, to determine that the current video block is predicted using the intra BC mode, configuration information about which video blocks of the frame are within the reconstructed area and should be stored in the DPB 92, a block vector for each video block predicted with intra BC of the frame, an intra BC prediction status for each video block predicted with intra BC of the frame, and other information for decoding the video blocks in the current video frame.

[0051] The motion compensation unit 82 may also perform interpolation using an interpolation filter to calculate interpolated values for sub-integer pixels of a reference block, as used by the video encoder 20 during the encoding of video blocks. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and create a prediction block using the interpolation filter.

[0052] The inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80 using the same quantization parameter calculated by the video encoder 20 for each video block in the video frame to determine the quantization degree. The inverse transform processing unit 88 applies an inverse transform, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to reconstruct the residual block in the pixel domain.

[0053] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 adds the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85 to reconstruct the decoded video block for the current video block. A loop filter (not shown) may be arranged between the adder 90 and the DPB 92 to further process the decoded video block. The decoded video block in a given frame is then stored in the DPB 92 that stores the reference frames used for subsequent motion compensation of the next video block. The DPB 92, or a memory device separate from the DPB 92, may store the decoded video for later presentation on a display device such as the display device 34 of FIG. 1.

[0054] In a typical video encoding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three sample arrays denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other cases, a frame may be monochromatic and thus include only one two-dimensional array of luminance samples.

[0055] As shown in FIG. 4A, a video encoder 20 (or more specifically, a partitioning unit 45) first generates an encoded representation of a frame by partitioning the frame into a set of coding tree units (CTUs). A video frame may include consecutive ordered integers of CTUs in a left-to-right, top-to-bottom raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTUs are signaled by the video encoder 20 in the sequence parameter set such that all CTUs in the video sequence have the same size of either 128×128, 64×64, 32×32, or 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As shown in FIG. 4B, each CTU may include one coding tree block (CTB) of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements used to encode the samples of the coding tree blocks. The syntax elements include characteristics of different types of units of the encoded pixel blocks, including inter or intra prediction, intra prediction mode, motion vectors, and other parameters, and describe how the video sequence can be reconstructed in the video decoder 30. In a monochromatic picture, or a picture having three separate color planes, a CTU may include a single coding tree block and syntax elements used to encode the samples of the coding tree block. The coding tree block may be an N×N block of samples.

[0056] To achieve better performance, the video encoder 20 may recursively perform a quadtree partition, a ternary tree partition, a binary tree partition, or a combination of both on the coding tree blocks of the CTU, and divide the CTU into smaller coding units (CUs). As shown in FIG. 4C, first, a 64×64 CTU 400 is divided into four smaller CUs each having a block size of 32×32. Of the four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs 430 and 440 are each further divided into four CUs with a block size of 8×8. FIG. 4D illustrates a quadtree data structure showing the final result of the partitioning process of the CTU 400 as shown in FIG. 4C, and each leaf node of the quadtree corresponds to one CU with a size ranging from 32×32 to 8×8. Similar to the CTU shown in FIG. 4B, each CU may include an encoded block (CB) of luminance samples and two corresponding encoded blocks of chrominance samples of a frame of the same size, and syntax elements used to encode the samples of the encoded blocks. In a monochrome picture, or a picture having three separate color planes, the CU may include a single encoded block and a syntax structure used to encode the samples of the encoded block. The quadtree partition shown in FIGS. 4C and 4D is for illustrative purposes only, and it should be noted that one CTU can be divided into CUs to adapt to various local characteristics based on quadtree / ternary tree / binary tree partitions. In a multi-tree structure, one CTU can be partitioned by a quadtree structure, and each leaf CU of the quadtree can be further partitioned by binary tree and ternary tree structures. As shown in FIG. 4E, there are five partition types, namely quaternary partition, horizontal binary partition, vertical binary partition, horizontal ternary partition, and vertical ternary partition.

[0057] In some implementations, video encoder 20 may further divide the coded block of the CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same inter or intra prediction is applied. The prediction unit (PU) of the CU may include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements used to predict the prediction block. In a monochrome picture, or a picture having three separate color planes, the PU may include a single prediction block and a syntax structure used to predict the prediction block. Video encoder 20 may generate prediction luma, Cb, and Cr blocks for the luma, Cb, and Cr prediction blocks of each PU of the CU.

[0058] Video encoder 20 may generate a prediction block for the PU using intra prediction or inter prediction. When video encoder 20 generates a prediction block for the PU using intra prediction, video encoder 20 may generate the prediction block of the PU based on the decoded samples of the frame associated with the PU. When video encoder 20 generates a prediction block for the PU using inter prediction, video encoder 20 may generate the prediction block of the PU based on the decoded samples of one or more frames other than the frame associated with the PU.

[0059] After the video encoder 20 generates the predicted luminance, Cb, and Cr blocks for one or more PUs of a CU, the video encoder 20 may generate a luminance residual block for the CU by subtracting the predicted luminance block of the CU from its original luminance encoded block such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample in one of the predicted luminance blocks of the CU and the corresponding sample in the original luminance encoded block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and the corresponding sample in the original Cb encoded block of the CU, and each sample in the Cr residual block of the CU indicates the difference between a Cr sample in one of the predicted Cr blocks of the CU and the corresponding sample in the original Cr encoded block of the CU.

[0060] Further, as illustrated in FIG. 4C, the video encoder 20 may decompose the luminance, Cb, and Cr residual blocks of the CU into one or more luminance, Cb, and Cr transform blocks using a quadtree partitioning. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A transform unit (TU) of the CU may include a transform block of luminance samples, two corresponding transform blocks of chroma samples, and syntax elements used to transform the transform block samples. Thus, each TU of the CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with the TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture, or a picture having three separate color planes, the TU may include a single transform block and a syntax structure used to transform the samples of the transform block.

[0061] The video encoder 20 may apply one or more transforms to the luminance transform blocks of the TUs to generate luminance coefficient blocks for the TUs. The coefficient blocks may be two-dimensional arrays of transform coefficients. The transform coefficients may be scalar quantities. The video encoder 20 may apply one or more transforms to the Cb transform blocks of the TUs to generate Cb coefficient blocks for the TUs. The video encoder 20 may apply one or more transforms to the Cr transform blocks of the TUs to generate Cr coefficient blocks for the TUs.

[0062] After generating the coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), the video encoder 20 may quantize the coefficient blocks. Quantization generally refers to the process by which transform coefficients are quantized to provide further compression by reducing the amount of data used to represent the transform coefficients when possible. After the video encoder 20 quantizes the coefficient blocks, the video encoder 20 may entropy encode the syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream including a bit sequence that forms a representation of the encoded frame and associated data to be stored in the storage device 32 or transmitted to the destination device 14.

[0063] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the opposite of the encoding process performed by the video encoder 20. For example, the video decoder 30 may inverse-transform the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 may also reconstruct the encoded block of the current CU by adding samples of the prediction block for the current CU's PU to the corresponding samples of the transform block of the current CU's TU. After reconstructing the encoded blocks for each CU of the frame, the video decoder 30 may reconstruct the frame.

[0064] As described above, video encoding mainly uses two modes, namely intra-frame prediction (or intra prediction) and inter-frame prediction (or inter prediction), to achieve video compression. Palette-based encoding is another encoding method adopted by many video encoding standards. In palette-based encoding, which may be particularly suitable for encoding screen-generated content, a video coder (e.g., the video encoder 20 or the video decoder 30) forms a palette table of colors representing the video data of a given block. The palette table contains the most dominant (e.g., frequently used) pixel values in a given block. Pixel values that do not frequently appear in the video data of a given block are either not included in the palette table or are included in the palette table as an escape color.

[0065] Each entry in the palette table includes an index for the corresponding pixel value in the palette table. The palette index for a sample in a block may be encoded to indicate which entry from the palette table should be used to predict or reconstruct which sample. This palette mode begins with a process that generates a palette predictor for the first block of a classification of a picture, slice, tile, or other such video block. As described below, palette predictors for subsequent video blocks are typically generated by updating previously used palette predictors. For purposes of illustration, it is assumed that the palette predictor is defined at the picture level. In other words, a picture may include a plurality of encoded blocks each having its own palette table, but there is one palette predictor for the entire picture.

[0066] To reduce the bits necessary to signal palette entries in a video bitstream, a video decoder may utilize a palette predictor to determine new palette entries in the palette table used to reconstruct a video block. For example, the palette predictor may include palette entries from a previously used palette table, or may be initialized with the most recently used palette table by including all entries of the most recently used palette table. In some implementations, the palette predictor may include fewer entries than all entries from the most recently used palette table, and at this time, some entries from other previously used palette tables may be incorporated. The palette predictor may have the same size as the palette table used to encode different blocks, or may be larger or smaller than the palette table used to encode different blocks. In one example, the palette predictor is implemented as a first-in first-out (FIFO) table that includes 64 palette entries.

[0067] To generate a palette table for a block of video data from a palette predictor, a video decoder may receive a 1-bit flag for each entry of the palette predictor from an encoded video bitstream. The 1-bit flag may have a first value (e.g., binary 1) indicating that the associated entry of the palette predictor should be included in the palette table, or a second value (e.g., binary 0) indicating that the associated entry of the palette predictor should not be included in the palette table. If the size of the palette predictor is larger than the palette table used for a block of video data, the video decoder may stop receiving further flags when the maximum size for the palette table is reached.

[0068] In some implementations, some entries in the palette table may be signaled directly in the encoded video bitstream instead of being determined using the palette predictor. For such entries, the video decoder may receive three separate m-bit values from the encoded video bitstream that indicate pixel values for the luminance and two chrominance components associated with the entry, where m represents the bit depth of the video data. Compared to the multiple m-bit values required for a directly signaled palette entry, a palette entry derived from the palette predictor requires only a 1-bit flag. Thus, signaling some or all of the palette entries using the palette predictor can significantly reduce the number of bits required to signal new palette table entries, thereby improving the overall encoding efficiency of palette mode encoding.

[0069] In many cases, the palette predictor for a block is determined based on the palette table used to encode one or more previously encoded blocks. However, when encoding the first coding tree unit in a picture, slice, or tile, the palette table of previously encoded blocks may not be available. Therefore, the palette predictor cannot be generated using the entries of the previously used palette table. In such cases, a series of palette predictor initializers may be signaled in the sequence parameter set (SPS) and / or picture parameter set (PPS), which are values used to generate the palette predictor when the previously used palette table is not available. The SPS generally refers to the syntax structure of the syntax elements applied to a series of consecutive encoded video pictures called the coded video sequence (CVS), which is determined by the content of the syntax elements found in the PPS, which is referred to by the syntax elements found in each slice segment header. The PPS generally refers to the syntax structure of the syntax elements applied to one or more individual pictures within the CVS, which is determined by the syntax elements found in each slice segment header. Thus, the SPS is generally considered a higher-level syntax structure than the PPS, which means that the syntax elements included in the SPS generally change less frequently and apply to a larger portion of the video data compared to the syntax elements included in the PPS.

[0070] 5A to 5B are block diagrams showing examples of applying the adaptive color space transform (ACT) technique to convert the residual between the RGB color space and the YCgCo color space according to some implementations of the present disclosure.

[0071] In the HEVC screen content coding extension, ACT is applied to adaptively convert the residual from one color space (e.g., RGB) to another color space (e.g., YCgCo). Thus, the correlation (e.g., redundancy) among the three color components (e.g., R, G, and B) is significantly reduced in the YCgCo color space. Further, in the existing ACT design, the adaptation of different color spaces is implemented at the transform unit (TU) level by signaling one flag, tu_act_enabled_flag, for each TU. When the flag tu_act_enabled_flag is equal to 1, this indicates that the residual of the current TU is encoded in the YCgCo space. Otherwise (i.e., the flag is equal to 0), this indicates that the residual of the current TU is encoded in the original color space (i.e., without color space conversion). Additionally, different color space conversion formulas are applied depending on whether the current TU is encoded in the lossless mode or the lossy mode. Specifically, the forward and inverse color space conversion formulas between the RGB color space and the YCgCo color space for the lossy mode are defined in FIG. 5A.

[0072] In the case of the lossless mode, a reversible RGB - YCgCo conversion (also known as YCgCo - LS) is used. The reversible RGB - YCgCo conversion is implemented based on the lifting operations shown in FIG. 5B and the related description.

[0073] As shown in FIG. 5A, the forward and inverse color conversion matrices used in the lossy mode are not normalized. Thus, the magnitude of the YCgCo signal is smaller than the magnitude of the original signal after the color conversion is applied. To compensate for the magnitude reduction caused by the forward color conversion, an adjusted quantization parameter is applied to the residual in the YCgCo domain. Specifically, when the color space conversion is applied, the QP value QP Y , QP Cg , and QP CoThey are set to QP-5, QP-5, and QP-3 respectively, where QP is the quantization parameter used in the original color space.

[0074] FIG. 6 is a block diagram applying a technique of luminance mapping by chroma scaling (LMCS) in an exemplary video data decoding process according to some implementations of the present disclosure.

[0075] In VVC, LMCS is used as a new coding tool applied before in-loop filters (e.g., deblocking filter, SAO, and ALF). Generally, LMCS has two main modules: 1) in-loop mapping of the luminance component based on an adaptive piecewise linear model, and 2) chroma residual scaling depending on luminance. FIG. 6 shows a modified decoding process to which LMCS is applied. In FIG. 6, the decoding modules performed in the mapped domain include an entropy decoding module, an inverse quantization module, an inverse transform module, a luminance intra prediction module, and a luminance sample reconstruction module (i.e., addition of luminance prediction samples and luminance residual samples). The decoding modules performed in the original (i.e., unmapped) domain include a motion compensation prediction module, a chroma intra prediction module, a chroma sample reconstruction module (i.e., addition of chroma prediction samples and chroma residual samples), and all in-loop filter modules such as a deblocking module, an SAO module, and an ALF module. The new operation modules introduced by LMCS include a forward mapping module 610 for luminance samples, a backward mapping module 620 for luminance samples, and a chroma residual scaling module 630.

[0076] The in-loop mapping of LMCS can adjust the dynamic range of the input signal to improve the coding efficiency. The in-loop mapping of the luminance samples in the existing LMCS design is constructed based on two mapping functions: one forward mapping function FwdMap and one corresponding inverse mapping function InvMap. The forward mapping function uses a piecewise linear model containing 16 equally sized pieces to transmit signals from the encoder to the decoder. The inverse mapping function can be directly derived from the forward mapping function and thus does not need to be transmitted.

[0077] The parameters of the luminance mapping model are signaled at the slice level. First, a presence flag is signaled to indicate whether the luminance mapping model should be signaled for the current slice. If the luminance mapping model exists within the current slice, the corresponding piecewise linear model parameters are further signaled. In addition, at the slice level, another LMCS control flag is signaled to enable / disable the LMCS for the slice.

[0078] The chroma residual scaling module 630 is designed to compensate for the interaction of quantization accuracy between the luminance signal and the corresponding chroma signal when in-loop mapping is applied to the luminance signal. Whether chroma residual scaling is enabled or disabled for the current slice is also signaled in the slice header. When the luminance mapping is enabled, an additional flag is signaled to indicate whether luminance-dependent chroma residual scaling is applied. When the luminance mapping is not used, the luminance-dependent chroma residual scaling is always disabled and no additional flag is required. In addition, the chroma residual scaling is always disabled for CUs containing four or fewer chroma samples.

[0079] FIG. 7 is a block diagram showing an exemplary video decoding process in which a video decoder implements the technique of inverse adaptive color space transform (ACT) according to some implementations of the present disclosure.

[0080] Similar to the ACT design in the SCC of HEVC, the ACT in VVC converts the intra / inter prediction residual of one CU in the 4:4:4 chroma format from the original color space (e.g., RGB color space) to the YCgCo color space. As a result, the redundancy between the three color components can be reduced for better coding efficiency. FIG. 7 shows a flowchart of the decoding in which inverse ACT is applied in the VVC framework by adding an inverse ACT module 710. When processing a CU encoded with ACT enabled, entropy decoding, inverse quantization, and transformation based on inverse DCT / DST are first applied to the CU. Then, as shown in FIG. 7, inverse ACT is activated to convert the decoded residual from the YCgCo color space to the original color space (e.g., RGB and YCbCr). In addition, since ACT is not normalized in the lossy mode, a QP adjustment of (-5, -5, -3) is applied to the Y, Cg, and Co components to compensate for the change in the magnitude of the transformed residual.

[0081] In some embodiments, the ACT method reuses the same HEVC ACT core transform to perform color conversion between different color spaces. Specifically, two different color conversions are applied depending on whether the current CU is encoded in a lossy or lossless manner. In the lossy case, the forward and inverse color conversions use the irreversible YCgCo transformation matrix shown in FIG. 5A. In the lossless case, the reversible color conversion YCgCo-LS shown in FIG. 5B is applied. Furthermore, different from the existing ACT design, the following changes are introduced to the proposed ACT scheme to cope with the interaction with other coding tools in the VVC standard.

[0082] For example, since the residual of one CU in HEVC can be divided into multiple TUs, an ACT control flag is signaled separately for each TU to indicate whether color space conversion needs to be applied. However, as described above in connection with FIG. 4E, in place of the multiple-partition type concept and thus to remove the separate CU, PU, and TU partitions in HEVC, in VVC, one quadtree nested by binary and ternary structures is applied. This means that, in most cases, as long as the corresponding maximum conversion size is smaller than the width or height of one component of the CU, the leaf node of one CU is also used as a unit for prediction and conversion processing without further partitioning. Based on such a partitioning structure, in the present disclosure, it is proposed to adaptively enable and disable ACT at the CU level. Specifically, one flag cu_act_enabled_flag is signaled for each CU to select between the original color space and the YCgCo color space for encoding the residual of the CU. When the flag is equal to 1, this indicates that the residuals of all TUs within the CU are encoded in the YCgCo color space. Otherwise, when the flag cu_act_enabled_flag is equal to 0, all residuals of the CU are encoded in the original color space.

[0083] FIG. 8 is a flowchart 800 showing an exemplary process by which a video decoder decodes video data by conditionally executing techniques of inverse adaptive color space conversion (ACT) according to some implementations of the present disclosure.

[0084] As shown in FIGS. 5A and 5B, the ACT can affect the decoded residual only when the current CU contains at least one non-zero coefficient. If all the coefficients obtained from entropy decoding are 0, the reconstructed residual remains 0 regardless of whether inverse ACT is applied. In the case of the inter-mode and intra-block copy (IBC) mode, the information on whether a CU contains non-zero coefficients is indicated by the CU root coding block flag (CBF), i.e., cu_cbf. When the flag is equal to 1, this means that there are residual syntax elements in the video bitstream for the current CU. Otherwise (i.e., the flag is equal to 0), this means that the residual syntax elements of the current CU are not signaled in the video bitstream, or in other words, it is assumed that all the residuals of the CU are 0. Therefore, in some embodiments, it is proposed that the flag cu_act_enabled_flag is signaled only when the root CBF flag cu_cbf of the current CU is equal to 1 for the inter and IBC modes. Otherwise (i.e., the flag cu_cbf is equal to 0), the flag cu_act_enabled_flag is not signaled and the ACT is disabled from decoding the residual of the current CU. On the other hand, different from the inter and IBC modes, the root CBF flag is not signaled for the intra-mode, i.e., the flag of cu_cbf is not used to condition the existence of the flag cu_act_enabled_flag for the intra CU. Conversely, when the ACT is applied to an intra CU, it is proposed to use the ACT flag to conditionally enable / disable the signaling of the CBF of the luminance component. For example, when an intra CU uses the ACT, the decoder assumes that at least one component contains non-zero coefficients. Therefore, when the ACT is enabled for an intra CU and there are no non-zero residuals in its transform block except for the last transform block, it is assumed that the CBF for that last transform block is 1 without signaling.For an intra CU containing exactly one TU, when the CBF for its two chroma components (indicated by tu_cbf_cb and tu_cbf_cr) is 0, without signal transmission, it is presumed that the CBF flag of the last component (i.e., tu_cbf_luma) is always 1. In one embodiment, such a presumption rule for the luminance CBF is only enabled for an intra CU containing only one single TU for residual encoding.

[0085] To conditionally execute inverse ACT in the coding unit, the video decoder first receives video data corresponding to the coding unit (e.g., encoded in 4:4:4 format) from the bitstream, and the coding unit is encoded in an inter prediction mode or an intra block copy mode (810).

[0086] Next, the video decoder receives a first syntax element (e.g., the CU root coding block flag cu_cbf) from the video data, and the first syntax element indicates whether the coding unit has a non-zero residual (820).

[0087] If the first syntax element has a non-zero value (e.g., 1 indicating that there is a residual syntax element in the bitstream for the coding unit) (830), the video decoder then receives a second syntax element (e.g., cu_act_enabled_flag) from the video data, and the second syntax element indicates whether the coding unit is encoded using an adaptive color space transform (ACT) (830-1).

[0088] On the other hand, if the first syntax element has a value of 0 (e.g., 0 indicating that there is no residual syntax element in the bitstream for the coding unit) (840), the video decoder assigns a value of 0 to the second syntax element (e.g., sets cu_act_enabled_flag to 0) (840-1).

[0089] The video decoder then determines whether to perform inverse ACT on the video data of the coding unit according to the value of the second syntax element (for example, if the second syntax element has a value of 0, the execution of inverse ACT is stopped, and if the second syntax element has a value other than 0, inverse ACT is executed. The value of the second syntax element may be received from the video data or may be assigned based on the logic described above) (850).

[0090] In some embodiments, the coding unit is coded in a 4:4:4 chroma format, and each of the components (for example, one luminance and two chromas) has the same sample rate.

[0091] In some embodiments, the first syntax element having a value of 0 indicates that there are no residual syntax elements in the bitstream for the coding unit, and the first syntax element having a value other than 0 indicates that there are residual syntax elements in the bitstream for the coding unit.

[0092] In some embodiments, the first syntax element includes a cu_cbf flag, and the second syntax element includes a cu_act_enabled flag.

[0093] In some embodiments, when the coding unit is coded in an intra prediction mode, the video decoder conditionally receives a syntax element (for example, tu_cbf_y) for decoding the luminance component of the coding unit.

[0094] To conditionally receive the syntax element for decoding the luminance component, the video decoder first receives the video data corresponding to the coding unit from the bitstream. The coding unit is coded in an intra prediction mode, and the coding unit includes a first chroma component, a second chroma component, and one luminance component. In some embodiments, the coding unit includes only one transform unit.

[0095] Next, the video decoder receives from the video data a first syntax element (e.g., cu_act_enabled_flag) indicating whether the coding unit is coded using ACT. For example, cu_act_enabled_flag being equal to "1" indicates that the coding unit is coded using ACT, and cu_act_enabled_flag being equal to "0" indicates that the coding unit is not coded using ACT (e.g., thus inverse ACT does not need to be performed).

[0096] After receiving the first syntax element from the video data, the video decoder receives from the video data a second syntax element (e.g., tu_cbf_cb) and a third syntax element (e.g., tu_cbf_cr), where the second syntax element indicates whether the first chroma component has a non-zero residual, and the third syntax element indicates whether the second chroma component has a non-zero residual. For example, tu_cbf_cb or tu_cbf_cr being equal to "1" indicates that the first chroma component or the second chroma component respectively has at least one non-zero residual, and tu_cbf_cb or tu_cbf_cr being equal to "0" indicates that the first chroma component or the second chroma component respectively does not have a non-zero residual.

[0097] If the first syntax element has a value other than 0 (e.g., 1 indicating that inverse ACT should be performed) and at least one of the two chroma components contains a non-zero residual (e.g., tu_cbf_cb == 1 or tu_cbf_cr == 1), the video decoder receives from the video data a fourth syntax element (e.g., tu_cbf_y), and the fourth syntax element indicates whether the luminance component has a non-zero residual.

[0098] On the other hand, when the first syntax element has a value other than 0 and the residuals of both chroma components have only 0 (e.g., tu_cbf_cb == 0 and tu_cbf_cr == 0), the video decoder assigns a default value (e.g., a value other than 0) indicating that the luminance component has a residual other than 0 to the fourth syntax element. As a result, the video decoder does not receive the value for the fourth syntax element from the video data.

[0099] After determining the value for the fourth syntax element (e.g., by receiving the value from the video data or by assigning a default value other than 0 to the fourth syntax element), the video decoder determines whether to reconstruct the coding unit from the video data according to the fourth syntax element.

[0100] In some embodiments, the coding unit includes only one transform unit (TU).

[0101] In some embodiments, determining whether to reconstruct the coding unit from the video data according to the fourth syntax element includes reconstructing the residual of the luminance component according to the determination that the fourth syntax element has a value other than 0, and canceling the reconstruction of the residual of the luminance component according to the determination that the fourth syntax element has a value of 0.

[0102] Considering the strong correlation among the three components of 4:4:4 video, the intra modes used to predict the luminance component and the chrominance component are often the same for a given coded block. Therefore, in order to reduce the ACT signaling overhead, it is proposed to enable ACT for one intra CU only when its chrominance component uses the same intra prediction mode (i.e., DM mode) as the luminance component. In some embodiments, there are two ways to conditionally signal the ACT enable / disable flag and the chrominance intra prediction mode. In one embodiment of the present disclosure, it is proposed to signal the ACT enable / disable flag before signaling the intra prediction mode of one intra CU. In other words, when the ACT flag (i.e., cu_act_enabled_flag) is equal to 1, the intra prediction mode of the chrominance component is not signaled but is assumed to be the DM mode (i.e., reusing the same intra prediction mode as the luminance component). Otherwise (i.e., cu_act_enabled_flag is 0), the intra prediction mode of the chrominance component is still signaled. In another embodiment of the present disclosure, it is proposed to signal the ACT enable / disable flag after signaling the intra prediction mode. In this case, the ACT flag cu_act_enabled_flag needs to be signaled only when the value of the parsed chrominance intra prediction mode is the DM mode. Otherwise (i.e., the chrominance intra prediction mode is not equal to DM), the flag cu_act_enabled_flag need not be signaled and is assumed to be 0. In yet another embodiment, it is proposed to enable ACT for all possible chrominance intra modes. When such a method is applied, the flag cu_act_enabled_flag is always signaled regardless of the chrominance intra prediction mode.

[0103] To conditionally signal the ACT activation / inactivation flag and the chroma intra prediction mode, the video decoder first receives video data corresponding to the coding unit from the bitstream, and the coding unit is coded in the intra prediction mode, and the coding unit includes two chroma components and one luma component.

[0104] The video decoder then receives from the video data a first syntax element (e.g., cu_act_enabled_flag) indicating that the coding unit is coded using ACT.

[0105] The video decoder then receives from the video data a second syntax element, which represents the intra prediction parameter of the luma component of the coding unit (e.g., representing one of 67 intra prediction directions).

[0106] If the second syntax element has a non-zero value indicating that the coding unit is coded using ACT, the video decoder reconstructs the two chroma components of the coding unit by applying the same intra prediction parameter as the luma component of the coding unit to the two chroma components of the coding unit.

[0107] In some embodiments, the intra prediction parameter indicates the intra prediction direction applied to generate the intra prediction samples of the coding unit.

[0108] When ACT is enabled for one CU, ACT needs to access the residuals of all three components to perform the color space conversion. However, as described above, the VVC design cannot guarantee that each CU always contains information of all three components. In some embodiments of the present disclosure, when the CU does not contain information of all three components, ACT should be disabled.

[0109] First, when the individual tree (also known as the "dual tree") partitioning structure is applied, the luminance samples and chrominance samples within one CTU are partitioned into CUs based on separate partitioning structures. As a result, the CUs within the luminance partition tree contain only the coding information of the luminance component, and the CUs within the chrominance partition tree contain only the coding information of the two chrominance components. The switching between the single tree partitioning structure and the individual tree partitioning structure is performed at various levels, such as the sequence level, picture level, slice level, and coding unit group level, etc. Therefore, when it is found that the individual tree is applied to one region, the ACT presumes that all CUs (both luminance CUs and chrominance CUs) within the region are invalidated, does not signal the ACT flag, and instead the ACT flag is presumed to be 0.

[0110] Second, when the ISP mode is enabled, the TU partition is applied only to the luminance samples, and the chrominance samples are encoded but not further divided into multiple TUs. Assuming that N is the number of ISP sub - partitions (i.e., TUs) for one intra - CU, according to the current ISP design, only the last TU contains both the luminance and chrominance components, and the first N - 1 ISP TUs are composed of only the luminance component. According to an embodiment of the present disclosure, ACT is disabled in the ISP mode. There are two ways to disable ACT for the ISP mode. In the first method, it is proposed to signal an ACT enable / disable flag (i.e., cu_act_enabled_flag) before signaling the ISP mode syntax. In such a case, when the flag cu_act_enabled_flag is equal to 1, the ISP mode is not signaled in the bit - stream and is assumed to be always 0 (i.e., switched off). In the second method, it is proposed to bypass the signaling of the ACT flag using the signaling of the ISP mode. Specifically, in this method, the ISP mode is signaled before the flag cu_act_enabled_flag. When the ISP mode is selected, the flag cu_act_enabled_flag is not signaled and is assumed to be 0. Otherwise (when the ISP mode is not selected), the flag cu_act_enabled_flag is signaled to adaptively select the color space for the residual encoding of the CU.

[0111] In addition to invalidating ACT for CUs where the luminance partitioning structure and the chroma partitioning structure do not match, in the present disclosure, it is also proposed to invalidate LMCS for CUs to which ACT is applied. In one embodiment, when one CU selects the YCgCo color space to encode its residual (i.e., ACT is 1), it is proposed to invalidate both luminance mapping and chroma residual scaling. In another embodiment, when ACT is enabled for one CU, it is proposed to invalidate only chroma residual scaling, and luminance mapping can still be applied to adjust the dynamic range of the output luminance samples. In the last embodiment, it is proposed to enable both luminance mapping and chroma residual scaling for the CU to which ACT is applied to encode its residual.

[0112] To invalidate ACT signaling by the binary tree partitioning structure, the video decoder obtains from the bitstream information indicating whether the coding units in the video data are coded by a single tree partitioning or by a binary tree partitioning.

[0113] If the coding unit is coded using a single tree partitioning and each coding unit contains both a luminance component and a chroma component, the video decoder receives a second syntax element (e.g., cu_act_enabled_flag) from the video data, and the value of the first syntax element indicates whether to perform inverse adaptive color space transform (ACT) for each coding unit.

[0114] On the other hand, if the coding unit is coded using a binary tree partitioning, the coding units in the luminance partitioning tree of the binary tree partitioning contain only the coding information related to the luminance component of the coding unit, and the coding units in the chroma partitioning tree of the binary tree partitioning contain only the coding information related to the chroma component of the coding unit, the video decoder assigns a value of 0 to the second syntax element.

[0115] The video decoder then determines whether to perform an inverse ACT on each coded unit within the coded tree unit according to a second syntax element.

[0116] In some embodiments, determining whether to perform an inverse ACT on each coded unit within the coded tree unit according to a second syntax element includes performing an inverse ACT on each coded unit according to a determination that the second syntax element has a value other than 0, and refraining from performing an inverse ACT on each coded unit according to a determination that the second syntax element has a value of 0.

[0117] In some embodiments, to invalidate the ISP mode by ACT signaling, the video decoder first receives video data corresponding to the coding unit from the bitstream. Next, the video decoder receives a first syntax element (e.g., cu_act_enabled_flag) from the video data, and the first syntax element indicates whether the coding unit is coded using ACT. If the first syntax element has a value of 0, the video decoder receives a second syntax element from the video data, and the second syntax element indicates whether the coding unit is coded using the ISP mode. If the first syntax element has a value other than 0, the video decoder assigns a value of 0 to the second syntax element, indicating that the coding unit is not coded using the ISP mode. The video decoder then determines whether to reconstruct the coding unit from the video data using the ISP mode according to the second syntax element. In the current VVC, when the ISP mode is enabled, the TU partition is applied only to the luminance samples, and the chrominance samples are coded but not further divided into multiple TUs. According to one embodiment of the present disclosure, since there is rich texture information in the chrominance plane, it is also proposed to enable the ISP mode for chrominance coding in 4:4:4 video. Based on this embodiment, different methods may be used. In one method, one additional ISP index is signaled and shared by two chrominance components. In another method, it is proposed to signal two additional ISP indexes separately, one for Cb / B and the other for Cr / R. In a third method, it is proposed to reuse the ISP index used for the luminance component for the ISP prediction of the two chrominance components.

[0118] The row-column weighted intra prediction (MIP) method is an intra prediction technique. To predict samples of a rectangular block with width W and height H, MIP takes as input one line consisting of H reconstructed neighboring boundary samples to the left of the block and one line consisting of W reconstructed neighboring boundary samples above the block. If the reconstructed samples are not available, they are generated in the same way as conventional intra prediction. The generation of the prediction signal is performed based on three steps: averaging, matrix-vector multiplication, and linear interpolation, as shown in FIG. 10.

[0119] In the current VVC, the MIP mode is enabled only for the luma component. For the same reason as enabling the ISP mode for the chroma component, in one embodiment, it is proposed to enable MIP for the chroma components of 4:4:4 video. Two signaling methods can be applied. In the first method, it is proposed to signal two MIP modes separately, using one for the luma component and the other for the two chroma components. In the second method, it is proposed to signal only one single MIP mode shared by the luma and chroma components.

[0120] To enable MIP for the chroma components in the 4:4:4 chroma format, the video decoder receives video data corresponding to an encoding unit from the bitstream, where the encoding unit is encoded in the intra prediction mode, the encoding unit includes two chroma components and one luminance component, and the chroma components and the luminance component have the same resolution. Next, the video decoder receives a first syntax element (e.g., intra_mip_flag) from the video data that indicates that the luminance component of the encoding unit is encoded using the MIP tool. If the first syntax element has a value other than 0 indicating that the luminance component of the encoding unit is encoded using the MIP tool, the video decoder receives a second syntax element (e.g., intra_mip_mode) from the video data that indicates the MIP mode applied to the luminance component of the encoding unit, and reconstructs the two chroma components of the encoding unit by applying the MIP mode of the luminance component of the encoding unit to the two chroma components of the encoding unit. The following table shows the syntax design specifications for decoding video data using ACT in VVC.

[0121] First, to indicate whether ACT is enabled or not at the sequence level, one additional syntax element, e.g., sps_act_enabled_flag, is added to the sequence parameter set (SPS). In some embodiments, when a color space conversion is applied to video content where the luminance component and the chroma components have the same resolution, one bitstream compliance requirement needs to be added so that ACT can be enabled only for the 4:4:4 chroma format. Table 1 shows the modified SPS syntax table with the above syntax added.

Table 1

[0122] The fact that the flag sps_act_enabled_flag is equal to 1 indicates that the adaptive color space conversion is enabled. The fact that the flag sps_act_enabled_flag is equal to 0 indicates that the adaptive color space conversion is disabled, and for the CU referring to this SPS, the flag cu_act_enabled_flag is not signaled and is presumed to be 0. When ChromaArrayType is not equal to 3, it is a requirement compliant with the bitstream that the value of sps_act_enabled_flag be equal to 0.

Table 2

[0123] The fact that the flag cu_act_enabled_flag is equal to 1 indicates that the residual of the coding unit is coded in the YCgCo color space. The fact that the flag cu_act_enabled_flag is equal to 0 indicates that the residual of the coding unit is coded in the original color space. If the flag cu_act_enabled_flag does not exist, it is presumed to be equal to 0.

Table 3

[0124] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a tangible medium such as a data storage medium or a computer-readable storage medium corresponding to a communication medium that includes any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementation of the implementations described in this application. A computer program product may include a computer-readable medium.

[0125] The terms used in the description of the implementations in this specification are for the purpose of describing particular implementations only and are not intended to limit the scope of the claims. When used in the description of the implementations and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, the term “and / or” as used in this specification is to be understood to refer to and encompass any and all possible combinations of one or more of the associated listed items. Further, the terms “comprise” and / or “comprising” as used in this specification are to be understood to specify the presence of the stated features, elements, and / or components but do not preclude the presence or addition of one or more other features, elements, components, and / or groups thereof.

[0126] Also, terms such as first, second, etc. may be used in this specification to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, unless departing from the scope of the implementation, the first electrode may be referred to as the second electrode, and similarly, the second electrode may be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but they are not the same electrode.

[0127] The description of this application is presented for purposes of illustration and description and is not intended to be exhaustive or limited to the invention in the disclosed form. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art who benefit from the teachings presented in the foregoing description and related drawings. The embodiments are selected and described in order to best clarify the principles of the invention, its practical application, enable others skilled in the art to understand the invention for various implementations, and best utilize the fundamental principles and various implementations with various modifications to suit the particular uses contemplated. Accordingly, it should be understood that the scope of the claims should not be limited to the specific examples of the disclosed implementations, but is intended to include modifications and other implementations within the scope of the appended claims.

Claims

1. 1. A method for decoding video data, comprising the steps of: receiving video data corresponding to a coding unit from a bitstream, the coding unit being coded in an intra prediction mode, the coding unit having a first chroma component, a second chroma component, and a luma component; determining a first syntax element from the video data; determining a second syntax element and a third syntax element from the video data, the second syntax element indicating whether a first transform block for the first chroma component has non-zero residual data, and the third syntax element indicating whether a second transform block for the second chroma component has non-zero residual data; in response to the first syntax element having a non-zero value and at least one of the first transform block and the second transform block having non-zero residual data, determining a fourth syntax element from the video data, the fourth syntax element indicating whether a third transform block for the luma component has non-zero residual data; reconstructing the coding unit from the video data according to the fourth syntax element; The method includes:

2. in response to the first syntax element having the non-zero value and both the first transform block and the second transform block including only zero residual data, The method of claim 1 , further comprising assigning a non-zero value to the fourth syntax element indicating that the third transform block has non-zero residual data.

3. The method of claim 1 , wherein the first syntax element indicates whether adaptive color space transformation (ACT) has been used to apply residual data of the coding unit.

4. The method of claim 1 , wherein the coding unit includes only one transform unit (TU).

5. reconstructing the coding unit according to the fourth syntax element, reconstructing a luma residual in accordance with a determination that the fourth syntax element has a value other than zero; and The method of claim 1 , comprising: refraining from reconstructing the luma residual in accordance with a determination that the fourth syntax element has a value of zero.

6. 1. A method for encoding a coding unit in a video frame, comprising the steps of: performing intra prediction on the coding unit, the coding unit having a first chroma component, a second chroma component, and a luma component; Determining a value of a first syntax element; determining a value of a second syntax element and a value of a third syntax element, the second syntax element indicating whether a first transform block for the first chroma component has non-zero residual data and the third syntax element indicating whether a second transform block for the second chroma component has non-zero residual data; in response to the first syntax element having a non-zero value and at least one of the first transform block and the second transform block having non-zero residual data, determining a value of a fourth syntax element, the fourth syntax element indicating whether a third transform block for the luma component has non-zero residual data; reconstructing the coding unit according to the fourth syntax element; A method comprising:

7. in response to the first syntax element having the non-zero value and both the first transform block and the second transform block including only zero residual data, The method of claim 6 , further comprising assigning a non-zero value to the fourth syntax element indicating that the third transform block has non-zero residual data.

8. The method of claim 6 , wherein the first syntax element indicates whether adaptive color space transformation (ACT) is used to apply residual data of the coding unit.

9. The method of claim 6 , wherein the coding unit includes only one transform unit (TU).

10. reconstructing the coding unit according to the fourth syntax element, reconstructing a luma residual in accordance with a determination that the fourth syntax element has a value other than zero; and The method of claim 6 , comprising: abandoning reconstruction of the luma residual in accordance with a determination that the fourth syntax element has a value of zero.

11. one or more processing units; a memory coupled to the one or more processing units; a plurality of programs stored in said memory which, when executed by said one or more processing units, cause said one or more processing units to perform a method according to any one of claims 1 to 10; An electronic device comprising:

12. 11. A non-transitory computer-readable storage medium storing a plurality of programs for execution by an electronic device having one or more processing units, the plurality of programs, when executed by the one or more processing units, causing the one or more processing units to perform a method according to any one of claims 1 to 10.

13. A method for storing a bitstream, comprising the steps of: A method, wherein the bitstream is decoded by a method for decoding video data according to any of claims 1 to 5.

14. A method for storing a bitstream, comprising the steps of: A method, wherein the bitstream is generated by a method for encoding coding units in a video frame according to any of claims 6 to 10.

Citation Information

Patent Citations

  • System and method for rgb video coding enhancement

    JP2017513335A

  • QP Derivation and Offset for Adaptive Color Conversion in Video Coding

    JP2017531395A