A method, apparatus, and medium for video decoding
By introducing adaptive color transformation and chroma residual scaling technology in video encoding and decoding, the problem of low encoding and decoding efficiency of 4:4:4 chroma format video in the existing standards is solved, and more efficient encoding and decoding performance is achieved, especially in the high-frequency information processing of chroma components.
Patent Information
- Application Number
- CN202410413838.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-12
- Filing Date
- 2021-06-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-06-14
AI Technical Summary
When processing 4:4:4 chroma format video, the existing video encoding and decoding standards fail to fully utilize the correlation between Cb/B components and Cr/R components, resulting in low encoding and decoding efficiency. Especially in applications with high fidelity requirements, existing tools do not handle chroma components inadequately.
Adaptive color transformation (ACT) process is adopted to conditionally apply the chromaticity residual scaling technology. By deciding whether to perform residual encoding and decoding in the YCgCo color space at the encoding and decoding unit level, and combining brightness mapping and chroma residual scaling, the encoding and decoding process of video blocks is optimized.
It improves the encoding and decoding efficiency of 4:4:4 video, reduces redundancy, and improves the encoding and decoding performance of high-fidelity video, especially in the high-frequency information processing of chroma components.
Smart Images

Figure CN118118672B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese Patent Application No. 202180042158.3, which is the national phase application in China of the international patent application PCT / US2021 / 037197 filed on June 14, 2021, and this international patent application claims the priority of the US Patent Application No. 63 / 038,692 filed on June 12, 2020. Technical Field
[0002] This application generally relates to video data encoding, decoding, and compression, and particularly relates to methods and systems for conditionally applying a chrominance residual scaling process according to an adaptive color transform (ACT) process. Background
[0003] A variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc., support digital video. Electronic devices send, receive, encode, decode, and / or store digital video data by implementing video compression / decompression standards defined by, for example, the MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) standards. Video compression typically includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove the redundancy inherent in video data. For block-based video encoding and decoding, a video frame is divided into one or more strips, each strip having a plurality of video blocks, which may also be referred to as coding tree units (CTUs). Each CTU may contain a coding unit (CU) or be recursively divided into smaller CUs until a preset minimum CU size is reached. Each CU (also referred to as a leaf CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be encoded and decoded in intra-frame, inter-frame, or intra-block copy (IBC) mode. Intra-frame encoded (I) strips of video blocks in a video frame are encoded using spatial prediction relative to reference samples in adjacent blocks within the same video frame. Video blocks in inter-frame encoded (P (forward predicted picture) or B (bi-directionally predicted picture)) strips of a video frame can be encoded using spatial prediction relative to reference samples in adjacent blocks within the same video frame or using temporal prediction relative to reference samples in other previous and / or future reference video frames.
[0004] Spatial and temporal prediction based on previously encoded reference blocks (e.g., neighboring blocks) generates a predicted block for the current video block to be coded / decoded. The process of finding the reference blocks can be accomplished by a block matching algorithm. Residual data representing the pixel difference between the current block to be coded / decoded and the predicted block is referred to as a residual block or prediction error. An inter-coded block is coded based on a motion vector pointing to a reference block in a reference frame forming the predicted block, and the residual block. The process of determining the motion vector is typically referred to as motion estimation. An intra-coded / decoded block is coded based on an intra-prediction mode and the residual block. For further compression, the residual block is transformed from the pixel domain to a transform domain, such as the frequency domain, to produce residual transform coefficients, which can then be quantized. The quantized transform coefficients, initially arranged as a two-dimensional array, can be scanned to produce a one-dimensional vector of transform coefficients, which is then entropy-coded into a video bitstream to achieve more compression.
[0005] The encoded video bitstream is then saved in a computer-readable storage medium (e.g., flash memory) to be accessed by another electronic device having digital video capabilities, or directly transmitted to the electronic device in a wired or wireless manner. The electronic device then performs video decompression (which is a process opposite to the video compression described above) by, for example, parsing the encoded video bitstream to obtain syntax elements from the bitstream and reconstructing the digital video data from the encoded video bitstream into its original format at least in part based on the syntax elements obtained from the bitstream, and renders the reconstructed digital video data on a display of the electronic device.
[0006] As digital video quality goes from high definition to 4K×2K or even 8K×4K, the amount of video data to be encoded / decoded grows exponentially. There have been challenges in how to more efficiently encode / decode video data while maintaining the image quality of the decoded video data.
[0007] Some video content (e.g., screen content video) is encoded in the 4:4:4 chroma format, where all three components (the luminance component and the two chroma components) have the same resolution. Although the 4:4:4 chroma format includes more redundancy compared to the 4:2:0 and 4:2:2 chroma formats (which is not conducive to achieving good compression efficiency), the 4:4:4 chroma format is still the preferred encoding format for many applications that require high fidelity to preserve color information (such as sharp edges) in the decoded video. Given the redundancy present in 4:4:4 chroma format videos, there is evidence that significant codec improvements can be achieved by exploiting the correlation between the three color components (e.g., Y, Cb, and Cr in the YCbCr domain; or G, B, and R in the RGB domain) of 4:4:4 videos. Due to these correlations, during the development of the HEVC Screen Content Coding (SCC) extension, an Adaptive Color Transform (ACT) tool was adopted to exploit the correlation between the three color components. Summary of the Invention
[0008] This application describes embodiments related to the encoding and decoding of video data, and more particularly, to methods and systems for conditionally applying a chroma residual scaling process according to an Adaptive Color Transform (ACT) process.
[0009] For a video signal initially captured in the 4:4:4 color format, if the decoded video signal requires high fidelity and there is abundant information redundancy in the original color space (e.g., RGB video), it is preferable to encode the video in the original space. Although some inter-component codec tools (e.g., Cross-Component Linear Model Prediction (CCLM)) in the current VVC standard can improve the codec efficiency of 4:4:4 video encoding and decoding, the redundancy between these three components is not completely eliminated. This is because only the Y / G component is used to predict the Cb / B component and the Cr / R component, without considering the correlation between the Cb / B component and the Cr / R component. Accordingly, further decorrelating the three color components can improve the codec performance of 4:4:4 video encoding and decoding.
[0010] In the current VVC standard, the design of existing inter and intra tools mainly focuses on videos captured in 4:2:0 chroma format. Therefore, to achieve a better complexity / performance trade-off, most of these codec tools are only applicable to the luma component, while being disabled for the chroma components (e.g., Position-Dependent Intra Prediction Combination (PDPC), Multiple Reference Lines (MRL), and Sub-Division Prediction (ISP)), or different operations are used for the luma and chroma components (e.g., interpolation filters applied to motion compensated prediction). However, compared to 4:2:0 videos, video signals in 4:4:4 chroma format exhibit very different characteristics. For example, the Cb / B and Cr / R components of 4:4:4 YCbCr and RGB videos exhibit richer color information and more high-frequency information (e.g., edges and textures) than the chroma components in 4:2:0 videos. Considering this, using the same design of some existing codec tools in VVC may not always be optimal for 4:2:0 and 4:4:4 videos.
[0011] According to a first aspect of the present application, a method for decoding a video block encoded and decoded using chroma residual scaling includes: receiving, from a bitstream, a plurality of syntax elements associated with a coding and decoding unit, wherein the syntax elements include a first coding block flag (CBF) of residual samples of a first chroma component of the coding and decoding unit, a second CBF of residual samples of a second chroma component of the coding and decoding unit, and a third syntax element indicating whether an Adaptive Color Transform (ACT) is applied to the coding and decoding unit; determining, based on the first CBF, the second CBF, and the third syntax element, whether to perform the chroma residual scaling on the residual samples of the first chroma component and the second chroma component; scaling, based on corresponding scaling parameters, the residual samples of at least one of the first chroma component and the second chroma component according to a determination to perform the chroma residual scaling on the residual samples of at least one of the first chroma component and the second chroma component; and reconstructing samples of the coding and decoding unit using the luma and scaled chroma residual samples.
[0012] In some embodiments, determining whether to perform the chroma residual scaling on the residual samples of the first chroma component and the second chroma component based on the first CBF, the second CBF, and the third syntax element includes: in response to determining from the third syntax element that the ACT is applied to the coding and decoding unit: applying an inverse ACT to the luma and chroma residual samples of the coding and decoding unit; and determining to perform the chroma residual scaling on the residual samples of the first chroma component and the second chroma component after the inverse ACT regardless of the first CBF and the second CBF.
[0013] According to a second aspect of the present application, a method for decoding a video block encoded and decoded using chrominance residual scaling includes: receiving, from a bitstream, a plurality of syntax elements associated with a coding and decoding unit, where the syntax elements include a first coding block flag (CBF) of residual samples of a first chrominance component of the coding and decoding unit, a second CBF of residual samples of a second chrominance component of the coding and decoding unit, and a third syntax element indicating whether an adaptive color transform (ACT) is applied to the coding and decoding unit; determining, based on the first CBF and the second CBF, whether to perform the chrominance residual scaling on the residual samples of the first chrominance component and the second chrominance component; based on determining to perform the chrominance residual scaling on the residual samples of at least one of the first chrominance component and the second chrominance component, scaling the residual samples of at least one of the first chrominance component and the second chrominance component based on corresponding scaling parameters; and based on determining from the third syntax element that the ACT is applied to the coding and decoding unit, applying an inverse ACT to the luminance and chrominance residual samples of the coding and decoding unit after scaling.
[0014] According to a third aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory. When executed by the one or more processing units, the programs cause the electronic device to perform the method for decoding video data as described above.
[0015] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by an electronic device having one or more processing units. When executed by the one or more processing units, the programs cause the electronic device to perform the method for decoding video data as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification, illustrate the described embodiments and, together with the specification, serve to explain the basic principles. Like reference numerals refer to corresponding parts.
[0017] Figure 1 is a block diagram illustrating an exemplary video encoding and decoding system according to some embodiments of the present disclosure.
[0018] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.
[0019] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.
[0020] Figures 4A to 4Eis a block diagram illustrating how a frame is recursively partitioned into multiple video blocks having different sizes and shapes according to some embodiments of the present disclosure.
[0021] Figure 5A and Figure 5B is a block diagram illustrating an example of applying an adaptive color space transform (ACT) technique to transform residuals between an RGB color space and a YCgCo color space according to some embodiments of the present disclosure.
[0022] Figure 6 is a block diagram of applying a luminance mapping with chroma scaling (LMCS) technique during an exemplary video data decoding process according to some embodiments of the present disclosure.
[0023] Figure 7 is a block diagram illustrating an exemplary video decoding process according to some embodiments of the present disclosure, by which a video decoder implements an inverse adaptive color space transform (ACT) technique.
[0024] Figure 8A and Figure 8B is a block diagram illustrating an exemplary video decoding process according to some embodiments of the present disclosure, by which a video decoder implements an inverse adaptive color space transform (ACT) and a chroma residual scaling technique.
[0025] Figure 9 is a flowchart of an exemplary process according to some embodiments of the present disclosure, by which a video decoder decodes video data by conditionally performing a chroma residual scaling operation on residuals of a coding / decoding unit. DETAILED DESCRIPTION
[0026] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that various alternative solutions may be used and the subject matter may be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.
[0027] In some embodiments, methods are provided for improving the coding / decoding efficiency of the VVC standard for 4:4:4 video. Generally, the main features of the techniques in the present disclosure are summarized as follows.
[0028] In some embodiments, these methods are implemented to improve existing ACT designs that implement adaptive color space conversion in the residual domain. In particular, special consideration is given to handling the interaction of ACT with some existing coding and decoding tools in VVC.
[0029] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. As Figure 1 shown, system 10 includes a source device 12 that generates and encodes video data to be decoded by a destination device 14 at a later time. Source device 12 and destination device 14 may include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0030] In some embodiments, destination device 14 may receive encoded video data to be decoded via a link 16. Link 16 may include any type of communication medium or device capable of moving the encoded video data from source device 12 to destination device 14. In one example, link 16 may include a communication medium that enables source device 12 to directly transmit the encoded video data to destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (such as a local area network, a wide area network, or a global network (such as the Internet)). The communication medium may include routers, switches, base stations, or any other device that may be used to facilitate communication from source device 12 to destination device 14.
[0031] In some other embodiments, the encoded video data may be transmitted from the output interface 22 to the storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, flash memory, volatile memory, or non-volatile memory, or any other suitable digital storage media for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0032] As Figure 1 shown, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include sources such as a video capture device, e.g., a camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of these sources. As an example, if the video source 18 is a camera of a security surveillance system, the source device 12 and the destination device 14 may form a camera phone or a video phone. However, the embodiments described in this application can generally be applied to video encoding and decoding and can be applied to wireless and / or wired applications.
[0033] The captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be transmitted directly from the output interface 22 of the source device 12 to the destination device 14. The encoded video data may also (or alternatively) be stored on the storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.
[0034] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem and receives the encoded video data via a link 16. The encoded video data transmitted via the link 16 or provided on a storage device 32 may include various syntax elements generated by the video encoder 20 for use by the video decoder 30 when decoding the video data. Such syntax elements may be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored in a file server.
[0035] In some embodiments, the destination device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data to a user and may include any one of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0036] The video encoder 20 and the video decoder 30 may operate according to proprietary or industry standards, such as VVC, HEVC, MPEG-4 Part 10, Advanced Video Coding (AVC), AVS, or extensions of such standards. It should be understood that the present application is not limited to a particular video coding / decoding standard and may be applicable to other video coding / decoding standards. Generally, it is contemplated that the video encoder 20 of the source device 12 may be configured to encode video data according to any one of these current or future standards. Similarly, it is generally also contemplated that the video decoder 30 of the destination device 14 may be configured to decode video data according to any one of these current or future standards.
[0037] The video encoder 20 and the video decoder 30 may each be implemented as any one of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and any one of the one or more encoders or decoders may be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.
[0038] Figure 2FIG. 0 is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described in the present application. The video encoder 20 may perform intra prediction coding and decoding and inter prediction coding and decoding on video blocks within a video frame. Intra prediction coding and decoding relies on spatial prediction to reduce or remove spatial redundancy of video data within a given video frame or picture. Inter prediction coding and decoding relies on temporal prediction to reduce or remove temporal redundancy of video data within adjacent video frames or pictures of a video sequence.
[0039] As Figure 2 shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In some embodiments, the video encoder 20 further includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter, such as a deblocking filter (not shown), may be located between the adder 62 and the DPB 64 to filter block boundaries to remove block effect artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (not shown) may be used to filter the output of the adder 62. Further loop filtering, such as sample adaptive offset (SAO) and adaptive loop filter (ALF), may be applied to the reconstructed CU before placing the reconstructed CU in the reference image storage and using it as a reference for encoding future video blocks. The video encoder 20 may take the form of fixed or programmable hardware units, or may be partitioned among one or more of the illustrated fixed or programmable hardware units.
[0040] The video data memory 40 may store video data to be encoded by components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from a video source 18. The DPB 64 is a buffer that stores reference video data for use by the video encoder 20 when encoding video data (e.g., in an intra prediction coding and decoding mode or an inter prediction coding and decoding mode). The video data memory 40 and the DPB 64 may be formed of any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20 or off-chip relative to those components.
[0041] As Figure 2As shown, after receiving the video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles, or other larger coding units (CUs) according to a predefined segmentation structure, such as a quadtree structure associated with the video data. The video frame may be divided into a plurality of video blocks (or a set of video blocks referred to as tiles). The prediction processing unit 41 may select one of a plurality of possible prediction coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one of a plurality of intra prediction coding modes or one of a plurality of inter prediction coding modes. The prediction processing unit 41 may provide the resulting intra prediction coding block or inter prediction coding block to the adder 50 to generate a residual block, and to the adder 62 to reconstruct the coded block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra mode indicators, segmentation information, and other such syntax information to the entropy coding unit 56.
[0042] To select an appropriate intra prediction coding mode for the current video block, the intra prediction processing unit 46 within the prediction processing unit 41 may perform intra prediction coding of the current video block relative to one or more adjacent blocks in the same frame as the current block to be coded and decoded, to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter prediction coding of the current video block relative to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 may execute multiple coding and decoding channels, for example, in order to select an appropriate coding and decoding mode for each block of the video data.
[0043] In some embodiments, the motion estimation unit 42 determines the inter prediction mode of the current video frame by generating a motion vector according to a predetermined pattern within the video frame sequence, where the motion vector indicates the displacement of the prediction unit (PU) of the video block within the current video frame relative to the prediction block within the reference video frame. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector, which estimates the motion of the video block. The motion vector may, for example, indicate the displacement of the PU of the video block within the current video frame or picture relative to the prediction block (or other coded units) within the reference frame, where the prediction block is relative to the current block (or other coded units) coded within the current frame. The predetermined pattern may designate the video frames in the sequence as P frames or B frames. The intra BC unit 48 may determine the vector for performing intra BC coding in a manner similar to the way the motion estimation unit 42 determines the motion vector for inter prediction, for example, a block vector, or may utilize the motion estimation unit 42 to determine the block vector.
[0044] A predicted block is a block in a reference frame that is considered to closely match a PU of a video block to be coded or decoded in terms of pixel differences, which can be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some embodiments, video encoder 20 may calculate values at sub-integer pixel positions of the reference frames stored in DPB 64. For example, video encoder 20 may insert values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frames. Thus, motion estimation unit 42 may perform a motion search relative to full-pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy.
[0045] Motion estimation unit 42 calculates a motion vector of a PU of a video block in an inter-predicted coded frame by comparing the position of the PU with the position of a predicted block of a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1), each of the lists identifying one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44, and then to entropy coding unit 56.
[0046] Motion compensation performed by motion compensation unit 44 may involve obtaining or generating a predicted block based on the motion vector determined by motion estimation unit 42. After receiving the motion vector of the PU of the current video block, motion compensation unit 44 may locate the predicted block pointed to by the motion vector in one of the reference frame lists, obtain the predicted block from DPB 64 and forward the predicted block to adder 50. Then, adder 50 forms a residual video block with pixel differences by subtracting the pixel values of the predicted block provided by motion compensation unit 44 from the pixel values of the coded current video block. The pixel differences forming the residual video block may include a luminance difference component or a chrominance difference component or both. Motion compensation unit 44 may also generate syntax elements associated with the video blocks of a video frame for use by video decoder 30 when decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vectors for identifying the predicted blocks, any flags indicating the prediction mode, or any other syntax information described herein. Note that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are illustrated separately for conceptual purposes.
[0047] In some embodiments, the intra BC unit 48 may generate vectors and obtain prediction blocks in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, but where the prediction block is in the same frame as the current block being coded, and where the vector is referred to as a block vector relative to the motion vector. In particular, the intra BC unit 48 may determine an intra prediction mode for encoding the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during a separate encoding pass, and test its performance through rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values using rate-distortion analysis for the various tested intra prediction modes and select the intra prediction mode with the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original uncoded block (which was encoded to produce the encoded block) and the bit rate (i.e., the number of bits) used to produce the encoded block. The intra BC unit 48 may calculate a ratio based on the distortion and rate of each encoded block to determine which intra prediction mode exhibits the best rate-distortion value for the block.
[0048] In other examples, the intra BC unit 48 may use the motion estimation unit 42 and the motion compensation unit 44, in whole or in part, to perform such functions for intra BC prediction in accordance with the embodiments described herein. In either case, for intra block copy, the prediction block may be a block that is considered to closely match the block to be coded in terms of pixel differences, which may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics, and the identification of the prediction block may include calculating values at sub-integer pixel positions.
[0049] Regardless of whether the prediction block is from the same frame according to intra prediction or from a different frame according to inter prediction, the video encoder 20 may form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being coded, thereby forming pixel differences. The pixel differences for forming the residual video block may include luminance component differences and chrominance component differences.
[0050] As described above, the intra prediction processing unit 46 may perform intra prediction on the current video block as an alternative to inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or intra block copy prediction performed by the intra BC unit 48. In particular, the intra prediction processing unit 46 may determine an intra prediction mode for encoding the current block. To this end, the intra prediction processing unit 46 may, for example, encode the current block using various intra prediction modes during a separate encoding pass, and the intra prediction processing unit 46 (or, in some examples, a mode selection unit) may select an appropriate intra prediction mode from the tested intra prediction modes to use. The intra prediction processing unit 46 may provide information indicating the selected intra prediction mode of the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode in the bitstream.
[0051] After the prediction processing unit 41 determines a predicted block of the current video block via inter prediction or intra prediction, the summer 50 forms a residual video block by subtracting the predicted block from the current video block. The residual video data in the residual block may be included in one or more transform units (TUs) and provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0052] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0053] After quantization, the entropy coding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding / decoding (CAVLC), context adaptive binary arithmetic coding / decoding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding / decoding, or other entropy coding methods or techniques. The encoded bitstream may then be transmitted to the video decoder 30 or archived in the storage device 32 for later transmission to or retrieval by the video decoder 30. The entropy coding unit 56 may also entropy encode the motion vectors and other syntax elements of the current video frame being coded / decoded.
[0054] The inverse quantization unit 58 and the inverse transform processing unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual video block in the pixel domain to generate a reference block for predicting other video blocks. As described above, the motion compensation unit 44 can generate a motion-compensated prediction block from one or more reference blocks of the frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0055] The adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by the motion compensation unit 44 to generate a reference block for storage in the DPB 64. The reference block can then be used as a prediction block by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 to perform inter prediction on another video block in a subsequent video frame.
[0056] Figure 3 FIG. is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 can perform a decoding process opposite to the encoding process described above in connection with Figure 2 the video encoder 20. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra prediction unit 84 can generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.
[0057] In some examples, the units of the video decoder 30 can be assigned to perform embodiments of the present application. Similarly, in some examples, the embodiments of the present disclosure can be divided among one or more units of the video decoder 30. For example, the intra BC unit 85 can perform embodiments of the present application alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 can be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.
[0058] The video data memory 79 can store video data to be decoded by other components of the video decoder 30, such as an encoded video bitstream. For example, the video data stored in the video data memory 79 can be obtained from a storage device 32, a local video source (such as a camera), via wired or wireless network transmission of the video data or by accessing a physical data storage medium (e.g., a flash drive or a hard disk). The video data memory 79 can include a codec picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use by the video decoder 30 when decoding video data (e.g., in an intra prediction codec mode or an inter prediction codec mode). The video data memory 79 and the DPB 92 can be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, in Figure 3 the video data memory 79 and the DPB 92 are depicted as two different components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 can be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 can be on-chip with other components of the video decoder 30 or off-chip relative to those components.
[0059] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 can receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantized coefficients, motion vectors, or intra prediction mode indicators and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors and other syntax elements to the prediction processing unit 81.
[0060] When a video frame is decoded as an intra prediction codec (I) frame or for an intra codec prediction block in other types of frames, the intra prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for video blocks of the current video frame based on the signal-sent intra prediction mode and reference data from previously decoded blocks of the current frame.
[0061] When a video frame is coded / decoded as an inter-predicted coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each prediction block may be generated from a reference frame within one of the reference frame lists. The video decoder 30 may construct the reference frame lists: list 0 and list 1, using a default construction technique based on the reference frames stored in the DPB 92.
[0062] In some examples, when a video block is coded / decoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block defined by the video encoder 20.
[0063] The motion compensation unit 82 and / or the intra BC unit 85 determine the prediction information for the video blocks of the current video frame by parsing the motion vectors and other syntax elements, and then use the prediction information to generate the prediction blocks for the decoded current video blocks. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) for coding / decoding the video blocks of the video frame, the inter-predicted frame type (e.g., B or P), the construction information of one or more of the reference frame lists in the reference frame list of the frame, the motion vectors of each inter-predicted coded video block of the frame, the inter-predicted state of each inter-predicted coded video block of the frame, and other information for decoding the video blocks in the current video frame.
[0064] Similarly, the intra BC unit 85 may use some of the received syntax elements (e.g., flags) to determine that the current video block is predicted using: the intra BC mode, the construction information that the video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, the block vectors of each intra BC predicted video block of the frame, the intra BC prediction state of each intra BC predicted video block of the frame, and other information for decoding the video blocks in the current video frame.
[0065] The motion compensation unit 82 may also perform interpolation using an interpolation filter as used by the video encoder 20 during encoding of the video block to calculate the interpolation values of sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and use the interpolation filter to generate the prediction blocks.
[0066] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80, using the same quantization parameter for determining the degree of quantization calculated by the video encoder 20 for each video block in the video frame. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.
[0067] After the motion compensation unit 82 or the intra BC unit 85 generates a predicted block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block of the current video block by summing the residual block from the inverse transform processing unit 88 and the corresponding predicted block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter (not shown) may be positioned between the adder 90 and the DPB 92 to further process the decoded video block. Loop filtering such as deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) may be applied to the reconstructed CU before placing the reconstructed CU in the reference image memory. Then the decoded video blocks in a given frame are stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video blocks. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later presentation on a display device such as Figure 1 display device 34.
[0068] In a typical video coding / decoding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three arrays of samples, denoted as SL, SCb, and SCr, respectively. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other instances, a frame may be monochrome and thus include only one two-dimensional array of luminance samples.
[0069] As Figure 4A shown, the video encoder 20 (or more specifically, the partitioning unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of coding tree units (CTUs). A video frame may include an integer number of CTUs serially ordered in raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in the sequence parameter set such that all CTUs in the video sequence have the same size, i.e., one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As Figure 4BAs shown, each CTU may include one Coding Tree Block (CTB) of luminance samples, two corresponding Coding Tree Blocks of chrominance samples, and syntax elements for coding and decoding the samples of the Coding Tree Blocks. The syntax elements describe the attributes of different types of units of the coded block of pixels and how the video sequence can be reconstructed at the video decoder 30. The syntax elements include inter prediction or intra prediction, intra prediction mode, motion vector, and other parameters. In a monochrome picture or a picture with three separate color planes, the CTU may include a single Coding Tree Block and syntax elements for coding and decoding the samples of the Coding Tree Block. The Coding Tree Block may be an N×N sample block.
[0070] To achieve better performance, the video encoder 20 may recursively perform tree splitting (such as binary tree splitting, ternary tree splitting, quadtree splitting, or a combination of both) on the Coding Tree Blocks of the CTU and divide the CTU into smaller Coding Units (CUs). As Figure 4C depicted, first, the 64×64 CTU 400 is divided into four smaller CUs, each with a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four 16×16 CUs according to the block size. The two 16×16 CUs 430 and 440 are each further divided into four 8×8 CUs according to the block size. Figure 4D depicts a quadtree data structure that illustrates the final result of the division process of the CTU 400 as depicted in Figure 4C . Each leaf node of the quadtree corresponds to a CU with a corresponding size in the range of 32×32 to 8×8. Similar to the Figure 4B depicted CTU, each CU may include a coded block (CB) of luminance samples and two corresponding coded blocks of chrominance samples of the same size frame, as well as syntax elements for coding and decoding the samples of the coded blocks. In a monochrome picture or a picture with three separate color planes, the CU may include a single coded block and a syntax structure for coding and decoding the samples of the coded block. It should be noted that Figure 4C and Figure 4D the quadtree splitting depicted in is only for illustrative purposes, and a CTU can be divided into multiple CUs to adapt to different local characteristics based on quadtree / ternary tree / binary tree splitting. In a multi-type tree structure, a CTU is split by a quadtree structure, and each quadtree leaf CU can be further split by a binary tree structure or a ternary tree structure. As Figure 4E shown, there are five splitting types, namely quaternary splitting, horizontal binary splitting, vertical binary splitting, horizontal ternary splitting, and vertical ternary splitting.
[0071] In some embodiments, video encoder 20 may further divide the coded / decoded block of a CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same prediction (inter-frame or intra-frame) is applied. The prediction unit (PU) of a CU may include a prediction block for luma samples, two corresponding prediction blocks for chroma samples, and syntax elements for predicting the prediction block. In a monochrome picture or a picture with three separate color planes, the PU may include a single prediction block and syntax structures for predicting the prediction block. Video encoder 20 may generate predicted luma, Cb, and Cr blocks for the luma, Cb, and Cr prediction blocks of each PU of the CU.
[0072] Video encoder 20 may use intra-frame prediction or inter-frame prediction to generate the prediction blocks of the PU. If video encoder 20 uses intra-frame prediction to generate the prediction blocks of the PU, video encoder 20 may generate the prediction blocks of the PU based on the decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate the prediction blocks of the PU, video encoder 20 may generate the prediction blocks of the PU based on the decoded samples of one or more frames other than the frame associated with the PU.
[0073] After video encoder 20 generates the predicted luma, Cb, and Cr blocks of one or more PUs of a CU, video encoder 20 may generate a luma residual block of the CU by subtracting the predicted luma block of the CU from its original luma coded / decoded block, such that each sample in the luma residual block of the CU indicates the difference between the luma sample in one of the predicted luma blocks of the CU and the corresponding sample in the original luma coded / decoded block of the CU. Similarly, video encoder 20 may generate a Cb residual block and a Cr residual block of the CU respectively, such that each sample in the Cb residual block of the CU indicates the difference between the Cb sample in one of the predicted Cb blocks of the CU and the corresponding sample in the original Cb coded / decoded block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between the Cr sample in one of the predicted Cr blocks of the CU and the corresponding sample in the original Cr coded / decoded block of the CU.
[0074] In addition, as Figure 4CAs illustrated, video encoder 20 can use quadtree partitioning to decompose the luminance, Cb, and Cr residual blocks of a CU into one or more luminance, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. The transform unit (TU) of a CU can include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the samples of the transform block. Thus, each TU of a CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU can be a sub-block of the luminance residual block of the CU. The Cb transform block can be a sub-block of the Cb residual block of the CU. The Cr transform block can be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU can include a single transform block and a syntax structure for transforming the samples of the transform block.
[0075] Video encoder 20 can apply one or more transforms to the luminance transform block of a TU to generate a luminance coefficient block of the TU. A coefficient block can be a two-dimensional array of transform coefficients. The transform coefficients can be scalars. Video encoder 20 can apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block of the TU. Video encoder 20 can apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block of the TU.
[0076] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 can quantize the coefficient block. Quantization generally refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 can entropy code the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream including a bit sequence forming a representation of the encoded and decoded frame and associated data, which is stored in storage device 32 or transmitted to destination device 14.
[0077] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct a frame of video data at least in part based on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on a coefficient block associated with a TU of a current CU to reconstruct a residual block associated with the TU of the current CU. Video decoder 30 also reconstructs the coded block of the current CU by adding the samples of the prediction block of the PU of the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the coded block of each CU in a frame, video decoder 30 may reconstruct the frame.
[0078] As described above, video coding and decoding mainly use two modes to achieve video compression, namely, intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). Palette-based coding and decoding is another coding and decoding scheme adopted by many video coding and decoding standards. In palette-based coding and decoding, which may be particularly applicable to screen-generated content coding and decoding, a video codec (e.g., video encoder 20 or video decoder 30) forms a palette table representing the colors of the video data of a given block. The palette table includes the most dominant (e.g., frequently used) pixel values in the given block. Pixel values that are not frequently represented in the video data of the given block are not included in the palette table or are included in the palette table as escape colors.
[0079] Each entry in the palette table includes an index of the corresponding pixel value in the palette table. The palette indices of the samples in the block may be coded and decoded to indicate which entry in the palette table is to be used to predict or reconstruct which sample. The palette mode begins with the process of generating a palette prediction value for the first block of a picture, slice, tile, or other such grouping of video blocks. As will be explained below, the palette prediction values for subsequent video blocks are typically generated by updating the previously used palette prediction value. For illustrative purposes, it is assumed that the palette prediction value is defined at the picture level. In other words, a picture may include multiple coded blocks, each coded block having its own palette table, but there is one palette prediction value for the entire picture.
[0080] To reduce the bits needed to signal palette entries in a video bitstream, a video decoder can utilize palette prediction values to determine new palette entries in a palette table for reconstructing video blocks. For example, the palette prediction values can include palette entries from a previously used palette table, or can even be initialized with a recently used palette table by including all entries of the recently used palette table. In some embodiments, the palette prediction values can include less than all entries from the recently used palette table, and then combined with some entries from other previously used palette tables. The palette prediction values can have the same size as the palette tables used for coding / decoding different blocks, or can be larger or smaller than the palette tables used for coding / decoding different blocks. In one example, the palette prediction value is implemented as a first-in-first-out (FIFO) table including 64 palette entries.
[0081] To generate a palette table for a video data block from the palette prediction values, the video decoder can receive a one-bit flag for each entry of the palette prediction values from the encoded video bitstream. The one-bit flag can have a first value (e.g., binary one) indicating that the associated entry of the palette prediction value will be included in the palette table or a second value (e.g., binary zero) indicating that the associated entry of the palette prediction value will not be included in the palette table. If the size of the palette prediction value is larger than the palette table for the video data block, the video decoder can stop receiving more flags once the maximum size of the palette table is reached.
[0082] In some embodiments, some entries in the palette table can be signaled directly in the encoded video bitstream rather than determined using palette prediction values. For such entries, the video decoder can receive three separate m-bit values from the encoded video bitstream, where the m-bit values indicate the pixel values of the luminance component and two chrominance components associated with the entry, and where m represents the bit depth of the video data. Compared to the multiple m-bit values required for directly signaled palette entries, those palette entries obtained from the palette prediction values only require a one-bit flag. Thus, using palette prediction values to signal some or all palette entries can significantly reduce the number of bits needed to transmit the entries of a new palette table, thereby improving the overall coding / decoding efficiency of the palette mode coding / decoding.
[0083] In many instances, a palette prediction value for a block is determined based on a palette table used to encode or decode one or more previously encoded or decoded blocks. However, when encoding or decoding the first coding tree unit in a picture, slice, or tile, the palette table of the previously encoded blocks may not be available. Therefore, the entries of the previously used palette table cannot be used to generate the palette prediction value. In such cases, a series of initial palette prediction values can be signaled in the sequence parameter set (SPS) and / or picture parameter set (PPS). The initial values are the values used to generate the palette prediction value when the previously used palette table is not available. The SPS generally refers to the syntax structure of the syntax elements applied to a series of consecutive coded video pictures called the coded video sequence (CVS), as determined by the content of the syntax elements found in the PPS, which refers to the syntax elements found in each slice segment header. The PPS generally refers to the syntax structure of the syntax elements applied to one or more individual pictures within the CVS, as determined by the syntax elements found in each slice segment header. Therefore, the SPS is generally considered a higher-level syntax structure than the PPS, which means that compared to the syntax elements included in the PPS, the syntax elements included in the SPS generally change less frequently and apply to a larger portion of the video data.
[0084] Figures 5A to 5B is a block diagram illustrating an example of applying an adaptive color space transform (ACT) technique to transform residuals between an RGB color space and a YCgCo color space according to some embodiments of the present disclosure.
[0085] In the HEVC screen content coding extension, ACT is applied to adaptively transform the residuals from one color space (e.g., RGB) to another color space (e.g., YCgCo) such that the correlation (e.g., redundancy) between the three color components (e.g., R, G, and B) is significantly reduced in the YCgCo color space. Further, in the existing ACT design, adaptive execution of different color spaces is performed at the transform unit (TU) level by signaling a flag tu_act_enabled_flag for each TU. When the flag tu_act_enabled_flag is equal to one, it indicates that the residuals of the current TU are encoded in the YCgCo space; otherwise (i.e., the flag is equal to 0), it indicates that the residuals of the current TU are encoded in the original color space (i.e., no color space conversion is performed). In addition, different color space transform formulas are applied depending on whether the current TU is encoded in a lossless mode or a lossy mode. Specifically, the forward and inverse color space transform formulas between the RGB color space and the YCgCo color space in the lossy mode are defined in Figure 5A
[0086] For the lossless mode, a reversible version of the RGB - YCgCo transform (also referred to as YCgCo - LS) is used. The reversible version of the RGB - YCgCo transform is implemented based on the lifting operations and related descriptions depicted in Figure 5B The forward and inverse color transform matrices used in the lossy mode are unnormalized as shown in
[0087] As Figure 5A shown, the amplitude of the YCgCo signal is less than the amplitude of the original signal after applying the color transform. To compensate for the amplitude decrease caused by the forward color transform, an adjusted quantization parameter is applied to the residuals in the YCgCo domain. Specifically, when applying the color space transform, the QP values QP Y QP Cg and QP Co are set to QP - 5, QP - 5, and QP - 3 respectively, where QP is the quantization parameter used in the original color space.
[0088] Figure 6 is a block diagram of applying the luminance mapping with chroma scaling (LMCS) technique in an exemplary video data decoding process according to some embodiments of the present disclosure.
[0089] In VVC, LMCS is used as a new in - loop decoding tool applied before in - loop filters (e.g., deblocking filter, SAO, and ALF). Generally, LMCS has two main modules: 1) loop mapping of the luminance component based on an adaptive piece - wise linear model; 2) luminance - related chroma residual scaling. Figure 6 shows the modified decoding process in the case of applying LMCS. In Figure 6 , the decoding modules executed in the mapping domain include an entropy decoding block, an inverse quantization block, an inverse transform block, a luminance intra - prediction module, and a luminance sample reconstruction module (i.e., adding luminance predicted samples and luminance residual samples). The decoding modules executed in the original (i.e., non - mapped) domain include a motion - compensated prediction module, a chroma intra - prediction module, a chroma sample reconstruction module (i.e., chroma predicted samples plus chroma residual samples), and all in - loop filter modules such as a deblock module, an SAO module, and an ALF module. The new operation modules introduced by LMCS include a forward mapping module 610 for luminance samples, an inverse mapping module 620 for luminance samples, and a chroma residual scaling module 630.
[0090] The loop mapping of LMCS can adjust the dynamic range of the input signal to improve the coding and decoding efficiency. The loop mapping of the luminance samples in the existing LMCS design is based on two mapping functions: a forward mapping function FwdMap and a corresponding inverse mapping function InvMap. The forward mapping function uses a piecewise linear model with sixteen equally sized segments to transmit the signal from the encoder to the decoder. The inverse mapping function can be directly obtained from the forward mapping function and thus does not need to be signaled.
[0091] The parameters of the luminance mapping model are signaled at the slice level. First, the presence flag is signaled to indicate whether the luminance mapping model of the current slice is to be signaled. If the luminance mapping model exists in the current slice, the corresponding piecewise linear model parameters are further signaled. In addition, at the slice level, another LMCS control flag is signaled to enable / disable LMCS for the slice.
[0092] The chrominance residual scaling module 630 is designed to compensate for the interaction of quantization accuracy between the luminance signal and its corresponding chrominance signal when applying loop mapping to the luminance signal. It is also signaled in the slice header whether chrominance residual scaling is enabled or disabled for the current slice. When luminance mapping is enabled, an additional flag is signaled to indicate whether luminance-related chrominance residual scaling is applied. When luminance mapping is not used, luminance-related chrominance residual scaling is always disabled and no additional flag is required. In addition, for a CU containing less than or equal to four chrominance samples, chrominance residual scaling is always disabled.
[0093] Figure 7 FIG. is a block diagram illustrating an exemplary video decoding process according to some embodiments of the present disclosure, by which a video decoder implements an inverse adaptive color space transform (ACT) technique.
[0094] Similar to the ACT design in HEVC SCC, ACT in VVC transforms the intra / inter prediction residuals of a CU in 4:4:4 chroma format from the original color space (e.g., RGB color space) to the YCgCo color space. As a result, the redundancy between the three color components can be reduced to obtain better coding and decoding efficiency. Figure 7 FIG. depicts a decoding flowchart of how to apply inverse ACT in the VVC framework by adding an inverse ACT module 710. When processing a CU encoded with ACT enabled, entropy decoding, inverse quantization, and DCT / DST-based inverse transform are first applied to the CU. After that, as Figure 7As depicted, the inverse ACT is invoked to convert the decoded residuals from the YCgCo color space to the original color space (e.g., RGB and YCbCr). Additionally, since the ACT in the lossy mode is not normalized, a QP adjustment of (-5, -5, -3) is applied to the Y, Cg, and Co components to compensate for the variation amplitude of the transform residuals.
[0095] In some embodiments, the ACT method reuses the same ACT core transform of HEVC for color conversion between different color spaces. Specifically, two different versions of color transforms are applied depending on whether the current CU is encoded / decoded in a lossy or lossless manner. The forward and inverse color transforms for the lossy case use the invertible YCgCo transform matrix as depicted in Figure 5A For the lossless case, the invertible color transform YCgCo-LS as shown in Figure 5B is applied. Additionally, different from existing ACT designs, the ACT scheme introduces the following changes to handle its interaction with other coding tools in the VVC standard.
[0096] For example, since the residuals of a CU in HEVC can be partitioned into multiple TUs, an ACT control flag is signaled separately for each TU to indicate whether color space conversion needs to be applied. However, as described above in conjunction with Figure 4E a quadtree with binary and ternary partition structures nested is applied in VVC to replace the multi-partition type concept, thus removing the separate CU, PU, and TU partitions in HEVC. This means that in most cases, a CU leaf node is also used as the unit for prediction and transform processing without further partitioning, unless the maximum supported transform size is less than the width or height of one component of the CU. Based on such a partition structure, ACT at the CU level is enabled and disabled adaptively. Specifically, a flag cu_act_enabled_flag is signaled for each CU to select between the original color space and the YCgCo color space for encoding / decoding the residuals of the CU. If the flag equals 1, it indicates that the residuals of all TUs within the CU are encoded in the YCgCo color space. Otherwise, if the flag cu_act_enabled_flag equals 0, all residuals of the CU are encoded in the original color space.
[0097] In some embodiments, there are different scenarios for disabling ACT. When ACT is enabled for a CU, access to the residuals of all three components is required for color space conversion. However, the VVC design cannot guarantee that each CU always contains information of all three components. According to embodiments of the present disclosure, in those cases where a CU does not contain information of all three components, ACT should be forced to be disabled.
[0098] First, in some embodiments, when applying the split tree partitioning structure, the luma and chroma samples within a CTU are partitioned into multiple CUs based on the split partitioning structure. As a result, the CUs in the luma split tree contain only the coding and decoding information of the luma component, while the CUs in the chroma split tree contain only the coding and decoding information of the two chroma components. According to the current VVC, the switch between the single tree partitioning structure and the split tree partitioning structure is performed at the slice level. Thus, according to an embodiment of the present disclosure, when it is found that the split tree is applied to a slice, for all CUs (including luma CUs and chroma CUs) within the slice, ACT will always be disabled without signaling the ACT flag that is inferred to be zero.
[0099] Second, in some embodiments, when the ISP mode (described further below) is enabled, TU partitioning is applied only to luma samples while the chroma samples are coded and decoded without further partitioning into multiple TUs. Assume that N is the number of ISP sub-partitions (i.e., TUs) of an intra-frame CU. According to the current ISP design, only the last TU contains both the luma component and the chroma component, while the first N - 1 ISP TUs are composed of only the luma component. According to an embodiment of the present disclosure, ACT is disabled in the ISP mode. There are two ways to disable ACT for the ISP mode. In the first method, the ACT enable / disable flag (i.e., cu_act_enabled_flag) is signaled before signaling the syntax of the ISP mode. In such a case, when the flag cu_act_enabled_flag is equal to one, the ISP mode will not be signaled in the bitstream but will always be inferred to be zero (i.e., turned off). In the second method, the ISP mode signaling is used to bypass the signaling of the ACT flag. Specifically, in this method, the ISP mode is signaled before the flag cu_act_enabled_flag. When the ISP mode is selected, the flag cu_act_enabled_flag is not signaled and is inferred to be zero. Otherwise (when the ISP mode is not selected), the flag cu_act_enabled_flag will still be signaled to adaptively select the color space for residual coding of the CU.
[0100] In some embodiments, in addition to forcibly disabling ACT for CUs with misaligned luminance and chrominance partitioning structures, LMCS is also disabled for CUs to which ACT is applied. In one embodiment, when a CU selects the YCgCo color space to encode and decode its residuals (i.e., ACT is one), both luminance mapping and chrominance residual scaling are disabled. In another embodiment, when ACT is enabled for a CU, only chrominance residual scaling is disabled, and luminance mapping can still be applied to adjust the dynamic range of the output luminance samples. In the last embodiment, for CUs that apply ACT to encode and decode their residuals, both luminance mapping and chrominance residual scaling are enabled. There may be multiple ways to enable chrominance residual scaling for CUs to which ACT is applied. In one method, chrominance residual scaling is applied before inverse ACT during decoding. By this method, it means that when ACT is applied, chrominance residual scaling is applied to the chrominance residuals (i.e., Cg and Co residuals) in the YCgCo domain. In another method, chrominance residual scaling is applied after inverse ACT. Specifically, by the second method, chrominance scaling is applied to the residuals in the original color space. Assuming the input video is captured in RGB format, this means that chrominance residual scaling is applied to the residuals of the B and R components.
[0101] In some embodiments, a syntax element (e.g., sps_act_enabled_flag) is added to the sequence parameter set (SPS) to indicate whether ACT is enabled at the sequence level. Additionally, since the color space conversion is applied to video content with the same resolution for luminance and chrominance components (e.g., 4:4:4 chroma format 4:4:4), a bitstream compliance requirement needs to be added such that ACT can only be enabled for the 4:4:4 chroma format. Table 1 illustrates the modified SPS syntax table with the above-mentioned syntax added.
[0102]
[0103] Table 1 Modified SPS Syntax Table
[0104] Specifically, sps_act_enabled_flag equal to 1 indicates that ACT is enabled and sps_act_enabled_flag equal to 0 indicates that ACT is disabled, such that the flag cu_act_enabled_flag is not signaled for CUs that reference the SPS but is inferred to be 0. When ChromaArrayType is not equal to 3, the bitstream compliance requirement is that the value of sps_act_enabled_flag should be equal to 0.
[0105] In another embodiment, the sps_act_enabed_flag is not always signaled, but the signaling of this flag is conditional on the chroma type of the input signal. Specifically, since ACT can only be applied when the luma and chroma components are at the same resolution, the flag sps_act_enabled_flag is only signaled when the input video is captured in the 4:4:4 chroma format. With such a change, the modified SPS syntax table is as follows:
[0106]
[0107] Table 2 Modified SPS Syntax Table with Signaling Conditions
[0108] In some embodiments, the syntax design specifications for decoding video data using ACT are illustrated in the following table.
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117] Table 3 Signaling ACT Mode Specifications
[0118] The flag cu_act_enabled_flag being equal to 1 indicates that the residual of the coding / decoding unit is coded / decoded in the YCgCo color space, and the flag cu_act_enabled_flag being equal to 0 indicates that the residual of the coding / decoding unit is coded / decoded in the original color space (e.g., RGB or YCbCr). When the flag cu_act_enabled_flag does not exist, it is inferred to be equal to 0.
[0119] In some embodiments, the ACT signaling is conditional on the coding / decoding block flag (CBF). As Figure 5A and Figure 5BAs indicated, the ACT can only affect the decoded residual if the current CU contains at least one non - zero coefficient. If all the coefficients obtained from entropy decoding are zero, the reconstructed residual will be the same whether the inverse ACT is applied or not. For inter - frame mode and intra - block copy (IBC) mode, the information on whether a CU contains non - zero coefficients is indicated by the CU root codec block flag (CBF) (i.e., cu_cbf). When this flag is equal to one, it means that the residual syntax elements exist in the bitstream of the current CU. Otherwise (i.e., the flag is equal to 0), it means that the residual syntax elements of the current CU will not be signaled, and all the residuals of the inferred CU are zero. Therefore, for inter - frame mode and IBC mode, it is proposed to signal only the flag cu_act_enabled_flag when the root CBF flag cu_cbf of the current CU is equal to one. Otherwise (i.e., the flag cu_cbf is equal to 0), the flag cu_act_enabled_flag will not be signaled, and the ACT will always be disabled for decoding the residuals of the current CU. On the other hand, different from inter - frame and IBC modes, the root CBF flag is not signaled for intra - frame mode, i.e., there is no flag of cu_cbf that can be used to regulate the presence of the flag cu_act_enabled_flag for intra - frame CUs.
[0120] In some embodiments, an ACT flag is used to conditionally enable / disable the CBF signaling of the luminance component when applying the ACT to an intra - frame CU. Specifically, given an intra - frame CU using the ACT, the decoder always assumes that at least one component contains non - zero coefficients. Therefore, when the ACT is enabled for an intra - frame CU and there are no non - zero residuals in its transform blocks (except its last transform block), the CBF of its last transform block is inferred to be one without signaling. For an intra - frame CU containing only one TU, it means that if the CBFs of its two chrominance components (as indicated by tu_cbf_cb and tu_cbf_cr) are zero, the CBF flag of the last component (i.e., tu_cbf_luma) is always inferred to be one without signaling. In one embodiment, this inference rule for the luminance CBF is only enabled for intra - frame CUs containing only a single TU for residual codec.
[0121] Figure 8A and Figure 8B are block diagrams illustrating an exemplary video decoding process according to some embodiments of the present disclosure. The video decoder implements the inverse adaptive color space transform (ACT) and chrominance residual scaling techniques through the process. In some embodiments, the ACT (e.g., Figure 7 the inverse ACT 710 in Figure 6Both chrominance residual scaling 630) in are used to encode and decode the video bitstream. In some other embodiments, chrominance residual scaling is used but the inverse ACT 710 is not used simultaneously to encode and decode the video bitstream.
[0122] More specifically, Figure 8A An embodiment is depicted where the video codec performs chrominance residual scaling 630 before the inverse ACT 710. As a result, the video codec performs luminance mapping and chrominance residual scaling 630 in the color space transform domain. For example, assuming the input video is captured in RGB format and transformed into the YCgCo color space, the video codec performs chrominance residual scaling 630 on the chrominance residuals Cg and Co based on the luminance residual Y in the YCgCo color space.
[0123] In some embodiments, when chrominance residual scaling is applied in the YCgCo domain as Figure 8A shown, the corresponding chrominance residual samples fed into the chrominance residual scaling module are in the YCgCo domain. Correspondingly, the chrominance CBF flags of the current block (i.e., tu_cb_cbf and tu_cr_cbf) can be used to indicate whether there are any non-zero chrominance residual samples that need to be scaled. In this case, to avoid unnecessary chrominance scaling at the decoder, an additional check condition regarding the chrominance CBF flags can be added to ensure that chrominance residual scaling is invoked only when at least one of the two chrominance CBF flags is not zero.
[0124] Figure 8B An alternative embodiment is depicted where the video codec performs chrominance residual scaling 630 after the inverse ACT 710. As a result, the video codec performs luminance mapping and chrominance residual scaling 630 in the original color space domain. For example, assuming the input video is captured in RGB format, the video codec applies chrominance residual scaling to the B and R components.
[0125] In some embodiments, when chrominance residual scaling is applied in the RGB domain as Figure 8B shown, the corresponding residual samples fed into the chrominance residual scaling module are in the RGB domain. In this case, the chrominance CBF flags cannot indicate whether the corresponding B and R residual samples are all zero. Therefore, in this method, when ACT is applied to a CU, the two chrominance CBF flags cannot be used to decide whether chrominance residual scaling should be bypassed. When ACT is not applied to a CU, the two chrominance CBF flags in the YCgCo space can still be used to decide whether chrominance residual scaling can be bypassed.
[0126] In some embodiments, the following shows changes to the current VVC specification when the implemented signaling conditions regarding checking the CU-level ACT flag are applied to chroma residual scaling at the decoder:
[0127] 8.7.5.3 Picture reconstruction for chroma samples using the luminance-dependent chroma residual scaling process
[0128] - Set recSamples[xCurr+i][yCurr+j] to be equal to Clip1(predSamples[i][j]+resSamples[i][j]) if one or more of the following conditions are true:
[0129] - ph_chroma_residual_scale_flag is equal to 0.
[0130] - sh_lmcs_used_flag is equal to 0.
[0131] - nCurrSw*nCurrSh is less than or equal to 4.
[0132] - tu_cb_coded_flag[xCurr][yCurr] is equal to 0, and tu_cr_coded_flag[xCurr][yCurr] is equal to 0, and cu_act_enabled_flag[xCurr*SubWidthC][yCurr*SubHeightC] is equal to 0.
[0133] - Otherwise, the following applies:
[0134] - The current luminance position (xCurrY, yCurrY) is obtained as follows:
[0135] (xCurrY, yCurrY) = (xCurr*SubWidthC, yCurr*SubHeightC) (1234)
[0136] - Designate the luminance position (xCuCb, yCuCb) as the top-left luminance sample position of the luminance samples of the coding unit included in (xCurrY / sizeY*sizeY, yCurrY / sizeY*sizeY).
[0137] - The variables availL and availT are obtained as follows:
[0138] - Call the obtaining process of the neighboring block availability specified in Clause 6.4.4 with the position (xCurr, yCurr) set to be equal to (xCuCb, yCuCb), the neighboring luma position (xNbY, yNbY) set to be equal to (xCuCb - 1, yCuCb), checkPredModeY set to be equal to FALSE (false), and cIdx set to be equal to 0 as inputs, and assign the output to availL.
[0139] - Call the obtaining process of the neighboring block availability specified in Clause 6.4.4 with the position (xCurr, yCurr) set to be equal to (xCuCb, yCuCb), the neighboring luma position (xNbY, yNbY) set to be equal to (xCuCb, yCuCb - 1), checkPredModeY set to be equal to FALSE (false), and cIdx set to be equal to 0 as inputs, and assign the output to availT.
[0140] - The variable currPic specifies an array of reconstructed luma samples in the current picture.
[0141] - To obtain the variable varScale, the following ordered steps apply:
[0142] 1. The variable invAvgLuma is obtained as follows:
[0143] - The array recLuma[i] (where i = 0..(2 * sizeY - 1)) and the variable cnt are obtained as follows:
[0144] - Set the variable cnt to be equal to 0.
[0145] - When availL is equal to TRUE (true), set the array recLuma[i] (where i = 0..sizeY - 1) to be equal to currPic[xCuCb - 1][Min(yCuCb + i, pps_pic_height_in_luma_samples - 1)] (where i = 0..sizeY - 1), and set cnt to be equal to sizeY.
[0146] - When availT is equal to TRUE (true), set the array recLuma[cnt + i] (where i = 0..sizeY - 1) to be equal to currPic[Min(xCuCb + i, pps_pic_width_in_luma_samples - 1)][yCuCb - 1] (where i = 0..sizeY - 1), and set cnt to be equal to (cnt + sizeY).
[0147] - The variable invAvgLuma is obtained as follows:
[0148] - If cnt is greater than 0, the following applies:
[0149]
[0150] - Otherwise (cnt equals 0), the following applies:
[0151] invAvgLuma = 1 << (BitDepth - 1) (1236)
[0152] 2. The variable idxYInv is obtained by calling the identification of the piecewise function index process of the luminance sample specified in Article 8.8.2.3 with the variable lumaSample set to be equal to invAvgLuma as the input and idxYInv as the output.
[0153] 3. The variable varScale is obtained as follows:
[0154] varScale = ChromaScaleCoeff[idxYInv] (1237)
[0155] Figure 9 FIG. 900 is a flowchart illustrating an exemplary process according to some embodiments of the present disclosure, by which a video decoder (e.g., video decoder 30) decodes video data by conditionally performing a chrominance residual scaling operation on the residual of a coded unit.
[0156] The video decoder 30 receives a plurality of syntax elements associated with a coded unit from a bitstream, wherein the syntax elements include a first coded block flag (CBF) of the residual samples of the first chrominance component of the coded unit, a second CBF of the residual samples of the second chrominance component of the coded unit, and a third syntax element indicating whether an adaptive color transform (ACT) is applied to the coded unit (910).
[0157] The video decoder 30 then determines whether to perform chrominance residual scaling on the residual samples of the first chrominance component and the second chrominance component according to the first CBF, the second CBF, and the third syntax element (920).
[0158] According to the determination to perform the chrominance residual scaling on the residual samples of at least one of the first chrominance component and the second chrominance component, the video decoder 30 further scales the residual samples of at least one of the first chrominance component and the second chrominance component based on the corresponding scaling (930).
[0159] Video decoder 30 further uses the luminance residual samples and the scaled chrominance residual samples to reconstruct the samples (940) of the coding / decoding unit.
[0160] In some embodiments, determining whether to perform the chrominance residual scaling (920) on the residual samples of the first chrominance component and the second chrominance component according to the first CBF, the second CBF, and the third syntax element includes: determining that the ACT is applied to the coding / decoding unit from the third syntax element to: apply the inverse ACT to the luminance and chrominance residual samples of the coding / decoding unit; and determine to perform the chrominance residual scaling on the residual samples of the first chrominance component and the second chrominance component after the inverse ACT regardless of the first CBF and the second CBF.
[0161] In some embodiments, an inverse transform is applied to the residual samples of the coding / decoding unit before applying the inverse ACT.
[0162] In some embodiments, dequantization is applied to the residual samples of the coding / decoding unit before applying the inverse transform.
[0163] In some embodiments, determining whether to perform the chrominance residual scaling (920) on the residual samples of the first chrominance component and the second chrominance component according to the first CBF, the second CBF, and the third syntax element includes: determining that the ACT is not applied to the coding / decoding unit from the third syntax element to: when the CBF associated with the chrominance component is not zero, determine to perform the chrominance residual scaling on the residual samples of the chrominance component of the coding / decoding unit; or when the CBF associated with the chrominance component is zero, determine to bypass the chrominance residual scaling of the residual samples of the chrominance component of the coding / decoding unit.
[0164] In some embodiments, the first CBF is 0 when non-zero chrominance residual samples do not exist in the residual samples of the first chrominance component. In some embodiments, the second CBF is 0 when non-zero chrominance residual samples do not exist in the residual samples of the second chrominance component.
[0165] In some embodiments, the corresponding scaling parameter is obtained from the reconstructed luminance samples at the co-located position.
[0166] In some embodiments, the input of the inverse ACT is in the YCbCr space.
[0167] In some embodiments, the output of the inverse ACT is in the RGB space.
[0168] In some embodiments, a method for decoding a video block encoded and decoded using chrominance residual scaling includes: receiving, from a bitstream, a plurality of syntax elements associated with an encoding and decoding unit, wherein the syntax elements include a first coding block flag (CBF) of residual samples of a first chrominance component of the encoding and decoding unit, a second CBF of residual samples of a second chrominance component of the encoding and decoding unit, and a third syntax element indicating whether an adaptive color transform (ACT) is applied to the encoding and decoding unit; determining, based on the first CBF and the second CBF, whether to perform the chrominance residual scaling on the residual samples of the first chrominance component and the second chrominance component; based on determining to perform the chrominance residual scaling on the residual samples of at least one of the first chrominance component and the second chrominance component, scaling the residual samples of at least one of the first chrominance component and the second chrominance component based on corresponding scaling parameters; and based on determining from the third syntax element that the ACT is applied to the encoding and decoding unit, applying an inverse ACT to the luminance and chrominance residual samples of the encoding and decoding unit after scaling. In some embodiments, an inverse transform is applied to the residual samples of the encoding and decoding unit before performing the chrominance residual scaling. In some embodiments, an inverse quantization is applied to the residual samples of the encoding and decoding unit before applying the inverse transform.
[0169] In some embodiments, determining, based on the first CBF and the second CBF, whether to perform the chrominance residual scaling on the residual samples of the first chrominance component and the second chrominance component includes: determining to perform the chrominance residual scaling on the residual samples of the chrominance component of the encoding and decoding unit when the CBF associated with the chrominance component is not zero; and determining to bypass the chrominance residual scaling on the residual samples of the chrominance component of the encoding and decoding unit when the CBF associated with the chrominance component is zero.
[0170] Further embodiments also include combining or otherwise rearranging various subsets of the above embodiments in various other embodiments.
[0171] In one or more instances, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.
[0172] The terms used in the description of the embodiments herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or groups thereof.
[0173] It should also be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the embodiments, the first electrode may be referred to as the second electrode, and similarly, the second electrode may be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but the first electrode and the second electrode are not the same electrode.
[0174] The description of the present application has been presented for purposes of illustration and description, and the description is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art from the foregoing description and the teachings presented in the associated drawings. The embodiments are selected and described in order to best explain the principles of the invention, the practical application, and to enable others of ordinary skill in the art to understand the various embodiments of the invention and to best utilize the basic principles and the various embodiments with various modifications suitable for the particular purposes contemplated. Accordingly, it is to be understood that the scope of the claims should not be limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the appended claims.
Claims
1. A method for video coding, the method comprising: Determining whether to perform chrominance residual scaling on residual samples of at least one of a first chrominance component and a second chrominance component of a coding and decoding unit according to at least one of the following conditions, the conditions including: whether the quantized transform coefficients of the first chrominance component include non-zero chrominance quantized transform coefficients, whether the quantized transform coefficients of the second chrominance component include non-zero chrominance quantized transform coefficients, and whether an adaptive color transform (ACT) is applied to the coding and decoding unit, wherein the coding and decoding unit is associated with a predefined partitioning method, and wherein the predefined partitioning method includes quadtree partitioning, horizontal ternary partitioning, vertical ternary partitioning, horizontal binary partitioning, or vertical binary partitioning; Performing the chrominance residual scaling on residual samples of at least one of the first chrominance component and the second chrominance component according to the determination; Scaling the residual samples of at least one of the first chrominance component and the second chrominance component based on corresponding scaling parameters; and Reconstructing chrominance samples of the coding and decoding unit using the scaled chrominance residual samples; and coding at least one of a first coding block flag (CBF) for the residual samples of the first chrominance component of the coding and decoding unit, a second CBF for the residual samples of the second chrominance component of the coding and decoding unit, and a third syntax element indicating whether the ACT is applied to the coding and decoding unit into a video bitstream; Wherein, according to the determination that the ACT is applied to the coding and decoding unit, applying an inverse ACT to the luminance residual samples and chrominance residual samples of the coding and decoding unit; and Wherein determining whether to perform chrominance residual scaling on residual samples of at least one of the first chrominance component and the second chrominance component of the coding and decoding unit includes: According to the determination that the ACT is applied to the coding and decoding unit, determining to perform the chrominance residual scaling on the residual samples of the first chrominance component and the second chrominance component after the inverse ACT regardless of whether the quantized transform coefficients of the first chrominance component include non-zero chrominance quantized transform coefficients and whether the quantized transform coefficients of the second chrominance component include non-zero chrominance quantized transform coefficients.
2. The method according to claim 1, wherein Determining whether to perform chrominance residual scaling on residual samples of at least one of the first chrominance component and the second chrominance component of the coding and decoding unit includes: According to the determination that the ACT is not applied to the coding and decoding unit: When the quantized transform coefficients of one of the first chrominance component and the second chrominance component of the coding and decoding unit include non-zero chrominance quantized transform coefficients, determining to perform the chrominance residual scaling on the residual samples of the one chrominance component of the coding and decoding unit; and When the quantized transform coefficients of the one chrominance component do not include any non-zero chrominance quantized transform coefficients, determining not to perform the chrominance residual scaling on the residual samples of the one chrominance component of the coding and decoding unit.
3. The method according to claim 1, further comprising: Apply an inverse transform to the transform coefficients of the codec unit before applying the inverse ACT.
4. The method according to claim 3, further comprising: Apply inverse quantization to the quantized transform coefficients of the codec unit before applying the inverse transform.
5. The method according to claim 1, wherein when no non-zero chrominance quantized transform coefficients are present in the quantized transform coefficients of the first chrominance component, the first codec block flag CBF is zero; and when no non-zero chrominance quantized transform coefficients are present in the quantized transform coefficients of the second chrominance component, the second CBF is zero.
6. The method according to claim 1, wherein The corresponding scaling parameter is obtained from the reconstructed luma samples.
7. The method according to claim 1, wherein The input of the inverse ACT is in the YCgCo color space.
8. The method according to claim 1, wherein, The output of the inverse ACT is in the RGB color space.
9. The method according to claim 1, wherein When the third syntax element is not present in the video bitstream, it is inferred to be zero.
10. An electronic device, comprising: one or more processing units; a memory coupled to the one or more processing units; and a plurality of programs stored in the memory, which when executed by the one or more processing units cause the electronic device to perform the method according to any one of claims 1 to 9.
11. A non-transitory computer-readable storage medium storing a plurality of programs for execution by an electronic device having one or more processing units, wherein, The plurality of programs, when executed by the electronic device, cause the one or more processing units to perform the method according to any one of claims 1 to 9, and the non-transitory computer-readable storage medium stores a bitstream generated by the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing a bitstream generated by the method according to any one of claims 1 to 9.
13. A computer program product comprising instructions that, when executed by a processor, implement the method according to any one of claims 1 to 9.
14. A method for transmitting a bitstream, wherein, The bitstream is generated by the method according to any one of claims 1 to 9.