A video encoding method, device and medium
By performing interception operations and applying inverse ACT before adaptive color space transformation (ACT), the problem of low efficiency of 4:4:4 video encoding and decoding in the prior art is solved, and more efficient video data encoding and decoding is achieved.
Patent Information
- Application Number
- CN202410103810.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-25
- Filing Date
- 2021-01-05
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-01-05
AI Technical Summary
When processing 4:4:4 chroma format video, existing video encoding and decoding technology fails to fully utilize the correlation between the three color components, resulting in low encoding and decoding efficiency.
Prior to adaptive color space transformation (ACT), an interception operation is performed to limit the dynamic range of the residuals and an inverse ACT is applied to improve the encoding and decoding performance of the video data.
By reducing the redundancy between color components, the encoding and decoding efficiency of 4:4:4 video is improved and the image quality retention ability is enhanced.
Smart Images

Figure CN117939152B_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese patent application No. 202180018865.9, which is a Chinese national phase application of international patent application PCT / US2021 / 012165 filed on January 5, 2021, and which claims priority to U.S. patent application No. 62 / 957,273 filed on January 5, 2020 and U.S. patent application No. 62 / 965,859 filed on January 25, 2020. Technical Field
[0002] The present application relates generally to video data encoding and compression, and more particularly to methods and systems for performing truncation operations prior to adaptive color space transform (ACT). Background Art
[0003] Various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. support digital video. Electronic devices send, receive, encode, decode and / or store digital video data by implementing video compression / decompression standards defined by MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Codec (AVC), High Efficiency Video Codec (HEVC) and Versatile Video Codec (VVC) standards. Video compression typically includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data. For block-based video codecs, a video frame is divided into one or more slices, each of which has multiple video blocks, which may also be referred to as coding tree units (CTUs). Each CTU may contain a coding unit (CU) or be recursively divided into smaller CUs until a predefined minimum CU size is reached. Each CU (also called leaf-CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be coded or decoded in intra, inter, or intra block copy (IBC) mode. Video blocks in intra-coded (I) slices of a video frame are encoded using spatial prediction relative to reference samples in neighboring blocks within the same video frame. Video blocks in inter-coded (P (forward predicted picture) or B (bidirectional predicted picture)) slices of a video frame can use spatial prediction relative to reference samples in neighboring blocks within the same video frame or use temporal prediction relative to reference samples in other previous and / or future reference video frames.
[0004] A prediction block for the current video block to be coded is generated based on spatial or temporal prediction of a previously coded reference block (e.g., a neighboring block). The process of finding a reference block can be accomplished by a block matching algorithm. The residual data representing the pixel difference between the current block to be coded and the prediction block is called a residual block or a prediction error. The inter-coded block is encoded according to a motion vector pointing to a reference block in a reference frame forming the prediction block, and a residual block. The process of determining a motion vector is typically referred to as motion estimation. The intra-coded block is encoded according to an intra-frame prediction mode and a residual block. For further compression, the residual block is transformed from a pixel domain to a transform domain, such as a frequency domain, thereby generating residual transform coefficients, which can then be quantized. The quantized transform coefficients, which are initially arranged as a two-dimensional array, can be scanned to generate a one-dimensional vector of transform coefficients, and then entropy encoded into a video bitstream to achieve more compression.
[0005] The encoded video bitstream is then stored in a computer-readable storage medium (e.g., a flash memory) to be accessed by another electronic device with digital video capabilities, or directly transmitted to the electronic device in a wired or wireless manner. The electronic device then performs video decompression (which is the reverse process of the video compression described above) by, for example, parsing the encoded video bitstream to obtain syntax elements from the bitstream and reconstructing digital video data from the encoded video bitstream to its original format based at least in part on the syntax elements obtained from the bitstream, and rendering the reconstructed digital video data on a display of the electronic device.
[0006] As digital video quality increases from HD to 4Kx2K or even 8Kx4K, the amount of video data to be encoded / decoded increases exponentially. There has always been a challenge on how to encode / decode video data more efficiently while maintaining the image quality of the decoded video data.
[0007] Some video content (e.g., screen content video) is encoded in a 4:4:4 chroma format, where all three components (luminance component and two chroma components) have the same resolution. Although the 4:4:4 chroma format includes more redundancy than the 4:2:0 chroma format and the 4:2:2 chroma format (which is not friendly to achieving good compression efficiency), the 4:4:4 chroma format is still the preferred encoding format for many applications that require high fidelity to preserve color information (such as sharp edges) in the decoded video. In view of the redundancy present in 4:4:4 chroma format video, there is evidence that significant coding and decoding improvements can be achieved by exploiting the correlation between the three color components of 4:4:4 video (e.g., Y, Cb, and Cr in the YCbCr domain; or G, B, and R in the RGB domain). Due to these correlations, during the development of the HEVC Screen Content Codec (SCC) extension, the Adaptive Color Space Transform (ACT) tool was used to exploit the correlation between the three color components. Summary of the invention
[0008] The present application describes embodiments related to video data encoding and decoding, and more particularly to methods and systems for performing a truncation operation prior to adaptive color space transform (ACT).
[0009] For video signals originally captured in a 4:4:4 color format, if the decoded video signal requires high fidelity and there is a large amount of information redundancy in the original color space (e.g., RGB video), the video is preferably encoded in the original space. Although some inter-component coding tools in the current VVC standard (e.g., cross-component linear model prediction (CCLM)) can improve the efficiency of 4:4:4 video coding, the redundancy between the three components is not completely eliminated. This is because only the Y / G component is used to predict the Cb / B component and the Cr / R component, without considering the correlation between the Cb / B component and the Cr / R component. Accordingly, further decorrelation of the three color components can improve the coding and decoding performance for 4:4:4 video coding.
[0010] In the current VVC standard, the design of existing inter-frame tools and intra-frame tools is mainly focused on videos captured in 4:2:0 chroma format. Therefore, in order to achieve a better complexity / performance trade-off, most of these codec tools are only applicable to luma components, while being disabled for chroma components (e.g., position-dependent intra prediction combination (PDPC), multiple reference lines (MRL), and sub-partition prediction (ISP)), or different operations are used for luma components and chroma components (e.g., interpolation filters applied to motion compensated prediction). However, compared with 4:2:0 video, video signals in 4:4:4 chroma format exhibit very different characteristics. For example, the Cb / B components and Cr / R components of 4:4:4YCbCr and RGB videos exhibit richer color information and have more high-frequency information (e.g., edges and textures) than the chroma components in 4:2:0 video. Taking this into account, it may always be best to use the same design of some existing codec tools in VVC for both 4:2:0 and 4:4:4 videos.
[0011] According to the first aspect of the present application, a method for decoding video data includes: receiving video data corresponding to a coding unit from a bitstream, wherein the coding unit is encoded and decoded by an intra-frame prediction mode or an inter-frame prediction mode; receiving a first syntax element from the video data, wherein the first syntax element indicates whether the coding unit has been encoded and decoded using an adaptive color space transform (ACT); processing the video data to generate a residual of the coding unit; performing a clipping operation on the residual of the coding unit based on a determination that the coding unit has been encoded and decoded using the ACT based on the first syntax element; and applying an inverse ACT to the residual of the coding unit after the clipping operation.
[0012] In some embodiments, the clipping operation limits the dynamic range of the residual of the coding unit to a predefined range for processing by the inverse ACT.
[0013] According to a second aspect of the present application, an electronic device comprises one or more processing units, a memory and a plurality of programs stored in the memory. When the programs are executed by the one or more processing units, the electronic device executes the method for decoding video data as described above.
[0014] According to a third aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by an electronic device having one or more processing units. When the programs are executed by the one or more processing units, the electronic device executes the method for decoding video data as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of the specification, illustrate the described embodiments and together with the description serve to explain the basic principles. Like reference numerals refer to corresponding parts.
[0016] Figure 1 is a block diagram illustrating an exemplary video encoding and decoding system according to some embodiments of the present disclosure.
[0017] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.
[0018] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.
[0019] FIG. 4A to FIG. 4E is a block diagram illustrating how a frame may be recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure.
[0020] Figure 5A and 5B is a block diagram illustrating an example of applying an adaptive color space conversion (ACT) technique to convert a residual between an RGB color space and a YCgCo color space according to some embodiments of the present disclosure.
[0021] Figure 6 is a block diagram of a technique for applying Luma Mapping with Chroma Scaling (LMCS) in an exemplary video data decoding process according to some embodiments of the present disclosure.
[0022] Figure 7 is a block diagram illustrating an exemplary video decoding process by which a video decoder implements an inverse adaptive color space transform (ACT) technique according to some embodiments of the present disclosure.
[0023] Fig. 8A and Figure 8B is a block diagram illustrating an exemplary video decoding process by which a video decoder implements techniques for inverse adaptive color space transform (ACT) and luma mapping with chroma scaling (LMCS) according to some embodiments of the present disclosure.
[0024] Fig. 9 is a block diagram illustrating exemplary decoding logic for switching between performing adaptive color space transform (ACT) and block differential pulse code modulation (BDPCM) according to some embodiments of the present disclosure.
[0025] Fig.10is a decoding flow chart for applying different quantization parameter (QP) offsets for different components when luma and chroma internal bit depths are different according to some embodiments of the present disclosure.
[0026] Fig.11A and Fig. 11B is a block diagram illustrating an exemplary video decoding process according to some embodiments of the present disclosure, by which a video decoder implements a clipping technique to limit the dynamic range of the residual of a coding unit to a predefined range for processing by inverse ACT.
[0027] Fig.12 is a flowchart illustrating an exemplary process according to some embodiments of the present disclosure, by which a video decoder decodes video data by performing a truncation operation that limits the dynamic range of the residual of a coding unit to a predefined range for processing by an inverse ACT. DETAILED DESCRIPTION
[0028] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that various alternatives may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.
[0029] In some embodiments, the method is provided to improve the encoding and decoding efficiency of the VVC standard for 4:4:4 video. In general, the main features of the technology in the present disclosure are summarized as follows.
[0030] In some embodiments, the method is implemented to improve existing ACT designs that can perform adaptive color space conversion in the residual domain. In particular, special consideration is given to handling the interaction between ACT and some existing codec tools in VVC.
[0031] In some embodiments, the method is implemented to improve the efficiency of some existing inter-frame and intra-frame codec tools in the VVC standard for 4:4:4 video, including: 1) enabling 8-tap interpolation filter for chroma components; 2) enabling PDPC for intra-frame prediction of chroma components; 3) enabling MRL for intra-frame prediction of chroma components; 4) enabling ISP partitioning for chroma components.
[0032] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. Figure 1 As shown, system 10 includes a source device 12 that generates and encodes video data to be decoded at a later time by a destination device 14. Source device 12 and destination device 14 may include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0033] In some embodiments, the destination device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may include a communication medium for enabling the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. The encoded video data may be modulated and transmitted to the destination device 14 according to a communication standard such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form a portion of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be used to facilitate communication from the source device 12 to the destination device 14.
[0034] In some other embodiments, the encoded video data may be transferred from the output interface 22 to a storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile memory or a non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include a web server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data through any standard data connection, including a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on a file server. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0035] like Figure 1 As shown, source device 12 includes video source 18, video encoder 20 and output interface 22. Video source 18 may include a source such as a video capture device, such as a camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, if video source 18 is a camera of a security monitoring system, source device 12 and destination device 14 may form a camera phone or a video phone. However, the embodiments described in the present application may be generally applicable to video encoding and decoding and may be applied to wireless and / or wired applications.
[0036] Captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for later access by destination device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.
[0037] Destination device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 may include a receiver and / or a modem and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored in a file server.
[0038] In some implementations, destination device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0039] The video encoder 20 and the video decoder 30 may operate in accordance with a proprietary or industry standard such as VVC, HEVC, MPEG-4 Part 10, Advanced Video Codec (AVC), or an extension of such a standard. It should be understood that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that the video encoder 20 of the source device 12 may be configured to encode video data in accordance with any of these current or future standards. Similarly, it is also generally contemplated that the video decoder 30 of the destination device 14 may be configured to decode video data in accordance with any of these current or future standards.
[0040] The video encoder 20 and the video decoder 30 can each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.
[0041] Figure 22 is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described herein. The video encoder 20 may perform intra-frame prediction codecs and inter-frame prediction codecs on video blocks within a video frame. Intra-frame prediction codecs rely on spatial predictions to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame prediction codecs rely on temporal predictions to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence.
[0042] like Figure 2 As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partition unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A deblocking filter (not shown) can be located between the adder 62 and the DPB 64 to filter the block boundaries to remove blocking artifacts from the reconstructed video. In addition to the deblocking filter, a loop filter (not shown) can also be used to filter the output of the adder 62. The video encoder 20 can take the form of a fixed or programmable hardware unit, or can be divided into one or more of the fixed or programmable hardware units shown.
[0043] The video data memory 40 can store video data to be encoded by the components of the video encoder 20. The video data in the video data memory 40 can be obtained, for example, from the video source 18. The DPB 64 is a buffer that stores reference video data for use by the video encoder 20 when encoding the video data (e.g., in an intra-frame prediction codec mode or an inter-frame prediction codec mode). The video data memory 40 and the DPB 64 can be formed by any of a variety of memory devices. In various examples, the video data memory 40 can be on-chip with other components of the video encoder 20, or off-chip relative to those components.
[0044] like Figure 2As shown, after receiving the video data, the partition unit 45 within the prediction processing unit 41 partitions the video data into video blocks. The partitioning may also include partitioning the video frame into strips, tiles, or other larger coding units (CUs) according to a predefined partitioning structure (such as a quadtree structure associated with the video data). The video frame may be divided into a plurality of video blocks (or sets of video blocks referred to as tiles). The prediction processing unit 41 may select one of a plurality of possible prediction codec modes for the current video block based on error results (e.g., coding rate and distortion level), such as one of a plurality of intra-frame prediction codec modes or one of a plurality of inter-frame prediction codec modes. The prediction processing unit 41 may provide the resulting intra-frame prediction coding block or inter-frame prediction coding block to the adder 50 to generate a residual block, and to the adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra-frame mode indicators, partition information, and other such syntax information to the entropy coding unit 56.
[0045] In order to select an appropriate intra-frame prediction codec mode for the current video block, the intra-frame prediction processing unit 46 within the prediction processing unit 41 can perform intra-frame prediction codec on the current video block relative to one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter-frame prediction codec on the current video block relative to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple encoding passes, for example, to select an appropriate encoding mode for each block of video data.
[0046] In some embodiments, the motion estimation unit 42 determines the inter-frame prediction mode of the current video frame by generating a motion vector according to a predetermined pattern within the sequence of video frames, the motion vector indicating the displacement of a prediction unit (PU) of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by the motion estimation unit 42 is the process of generating motion vectors, which estimates the motion of video blocks. The motion vector may, for example, indicate the displacement of a PU of a video block within the current video frame or picture relative to a prediction block (or other coded unit) within a reference frame, the prediction block (or other coded unit) relative to the current block (or other coded unit) encoded and decoded within the current frame. The predetermined pattern may designate the video frames in the sequence as P frames or B frames. The intra BC unit 48 may determine a vector, such as a block vector, for intra BC encoding and decoding in a manner similar to the manner in which the motion estimation unit 42 determines the motion vector for inter-frame prediction, or the motion estimation unit 42 may be utilized to determine the block vector.
[0047] A prediction block is a block of a reference frame that is considered to closely match the PU of the video block to be encoded in terms of pixel difference, which can be determined by the sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics. In some embodiments, the video encoder 20 can calculate the values of sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 can insert the values of the quarter pixel positions, one-eighth pixel positions, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 can perform a motion search relative to the full pixel positions and the fractional pixel positions and output the motion vector with fractional pixel accuracy.
[0048] Motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-prediction coded frame by comparing the position of the PU to the position of a prediction block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.
[0049] The motion compensation performed by the motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by the motion estimation unit 42. After receiving the motion vector of the PU of the current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, obtain the prediction block from the DPB 64 and forward the prediction block to the adder 50. The adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by the motion compensation unit 44 from the pixel values of the current video block being coded. The pixel difference values forming the residual video block may include a luminance difference component or a chrominance difference component or both. The motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by the video decoder 30 when decoding the video block of the video frame. The syntax elements may include, for example, syntax elements defining a motion vector for describing the prediction block, any flag indicating a prediction mode, or any other syntax information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are illustrated separately for conceptual purposes.
[0050] In some embodiments, the intra BC unit 48 may generate a vector and obtain a prediction block in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but wherein the prediction block is in the same frame as the current block being encoded and decoded, and wherein the vector is referred to as a block vector relative to the motion vector. In particular, the intra BC unit 48 may determine an intra prediction mode for encoding the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during a separate encoding pass, and test its performance by rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate a rate-distortion value using rate-distortion analysis for the various tested intra prediction modes and select an intra prediction mode with the best rate-distortion characteristic among the tested modes as the appropriate intra prediction mode to be used. The rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original, uncoded block (coded to produce the coded block) and the bit rate (i.e., the number of bits) used to produce the coded block. Intra BC unit 48 may calculate a ratio based on the distortion and rate for each coded block to determine which intra-prediction mode exhibits the best rate-distortion value for the block.
[0051] In other examples, intra BC unit 48 may use motion estimation unit 42 and motion compensation unit 44 in whole or in part to perform such functions for intra BC prediction according to embodiments described herein. In either case, for intra block copying, the prediction block may be a block that is considered to closely match the block to be encoded in terms of pixel difference, which may be determined by sum of absolute difference (SAD), sum of square difference (SSD), or other difference metrics, and identification of the prediction block may include calculating values for sub-integer pixel positions.
[0052] Regardless of whether the prediction block is from the same frame according to intra-frame prediction or from different frames according to inter-frame prediction, the video encoder 20 can form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded and decoded to form pixel difference values. The pixel difference values forming the residual video block may include a luminance component difference value and a chrominance component difference value.
[0053] As described above, the intra-prediction processing unit 46 can perform intra-prediction on the current video block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or the intra-block copy prediction performed by the intra BC unit 48. In particular, the intra-prediction processing unit 46 can determine the intra-prediction mode used to encode the current block. To this end, the intra-prediction processing unit 46 can encode the current block using various intra-prediction modes, for example during a separate encoding pass, and the intra-prediction processing unit 46 (or in some examples, a mode selection unit) can select an appropriate intra-prediction mode from the tested intra-prediction modes to use. The intra-prediction processing unit 46 can provide information indicating the selected intra-prediction mode of the block to the entropy encoding unit 56. The entropy encoding unit 56 can encode the information indicating the selected intra-prediction mode in the bitstream.
[0054] After prediction processing unit 41 determines a prediction block for the current video block via inter-prediction or intra-prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more transform units (TUs) and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.
[0055] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan on the matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform the scan.
[0056] After quantization, entropy coding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or other entropy coding methods or techniques. The encoded bitstream may then be transmitted to video decoder 30 or archived in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.
[0057] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain to generate a reference block for predicting other video blocks. As described above, motion compensation unit 44 may generate a motion compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0058] Adder 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block may then be used as a prediction block by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 to inter-predict another video block in a subsequent video frame.
[0059] Figure 3 1 is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction processing unit 84, and an intra-frame BC unit 85. The video decoder 30 may perform operations generally in conjunction with the above. Figure 2 The decoding process is the reverse of the encoding process described with respect to video encoder 20. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, and intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from entropy decoding unit 80.
[0060] In some examples, units of the video decoder 30 may be assigned to perform embodiments of the present application. Likewise, in some examples, embodiments of the present disclosure may be divided between one or more units of the video decoder 30. For example, the intra BC unit 85 may perform embodiments of the present application alone or in combination with other units of the video decoder 30 (such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 (such as the motion compensation unit 82).
[0061] The video data memory 79 may store video data to be decoded by other components of the video decoder 30, such as an encoded video bitstream. For example, the video data stored in the video data memory 79 may be obtained from the storage device 32, a local video source (such as a camera) via a wired or wireless network transmission of the video data or by accessing a physical data storage medium (e.g., a flash drive or a hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use when the video decoder 30 decodes the video data (e.g., in an intra-frame prediction codec mode or an inter-frame prediction codec mode). The video data memory 79 and the DPB 92 may be formed by any of a variety of memory devices, such as a dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are shown in FIG. Figure 3 92 are depicted as two distinct components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or by separate memory devices. In some examples, the video data memory 79 may be on-chip with other components of the video decoder 30, or off-chip relative to those components.
[0062] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks and associated syntax elements of an encoded video frame. The video decoder 30 may receive syntax elements at the video frame level and / or at the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra-frame prediction mode indicators and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors and other syntax elements to the prediction processing unit 81.
[0063] When a video frame is encoded and decoded as an intra-frame prediction codec (I) frame or an intra-frame codec prediction block in other types of frames, the intra-frame prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-frame prediction mode transmitted by the signal and the reference data from the previously decoded block of the current frame.
[0064] When the video frame is encoded as an inter-frame prediction codec (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks of the video block of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each prediction block can be generated from a reference frame in one of the reference frame lists. The video decoder 30 can construct the reference frame lists: List 0 and List 1 based on the reference frames stored in the DPB 92 using a default construction technique.
[0065] In some examples, when a video block is encoded or decoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block defined by video encoder 20.
[0066] The motion compensation unit 82 and / or the intra BC unit 85 determine the prediction information of the video block of the current video frame by parsing the motion vector and other syntax elements, and then use the prediction information to generate the prediction block of the decoded current video block. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to encode and decode the video block of the video frame, the inter prediction frame type (e.g., B or P), the construction information of one or more reference frame lists in the reference frame list of the frame, the motion vector of each inter prediction coded video block of the frame, the inter prediction state of each inter prediction coded decoded video block of the frame, and other information used to decode the video block in the current video frame.
[0067] Similarly, the intra BC unit 85 may use some of the received syntax elements (e.g., flags) to determine that the current video block is predicted using the intra BC mode, construction information that the video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding the video blocks in the current video frame.
[0068] Motion compensation unit 82 may also perform interpolation using interpolation filters to calculate interpolated values for sub-integer pixels of a reference block, as used by video encoder 20 during encoding of the video block. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from the received syntax elements and use the interpolation filters to produce the prediction block.
[0069] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80, using the same quantization parameters calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to reconstruct the residual block in the pixel domain.
[0070] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 reconstructs the decoded video block of the current video block by summing the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter (not shown) may be positioned between the adder 90 and the DPB 92 to further process the decoded video block. The decoded video blocks in a given frame are then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later presentation on a display such as a video player. Figure 1 On a display device such as a display device 34.
[0071] In a typical video encoding and decoding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other examples, a frame may be monochrome and therefore include only a two-dimensional array of luma samples.
[0072] like Figure 4A As shown, the video encoder 20 (or more specifically, the partition unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of coding tree units (CTUs). A video frame may include an integer number of CTUs sequentially ordered from left to right and from top to bottom in a raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, i.e., one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As Figure 4BAs shown, each CTU may include one coding tree block (CTB) of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements for encoding and decoding samples of the coding tree blocks. The syntax elements describe the properties of different types of units of the coding blocks of pixels and how the video sequence can be reconstructed at the video decoder 30. The syntax elements include inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vector and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding and decoding samples of the coding tree block. The coding tree block may be an NxN sample block.
[0073] To achieve better performance, the video encoder 20 may recursively perform tree partitioning (such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination of both) on the coding tree blocks of the CTU and partition the CTU into smaller coding units (CUs). Figure 4C As shown, a 64x64 CTU 400 is first divided into four smaller CUs, each with a block size of 32x32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four 16x16 CUs by block size. Two 16x16 CUs 430 and 440 are each further divided into four 8x8 CUs by block size. Figure 4D Depicted diagram Figure 4C The quadtree data structure of the final result of the partitioning process of the CTU 400 depicted in FIG. 4 is shown in FIG. 4 , with each leaf node of the quadtree corresponding to a CU with a corresponding size ranging from 32x32 to 8x8. Figure 4B In the depicted CTU, each CU may include a coding block (CB) of luma samples and two corresponding coding blocks of chroma samples of the same size frame, as well as syntax elements for encoding and decoding the samples of the coding blocks. In monochrome pictures or pictures with three separate color planes, a CU may include a single coding block and syntax structures for encoding and decoding the samples of the coding block. It should be noted that Figure 4C and Figure 4D The quadtree partitioning depicted in FIG is for illustration purposes only, and a CTU can be partitioned into multiple CUs to accommodate different local characteristics based on quadtree / ternary tree / binary tree partitioning. In the multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU can be further partitioned by a binary tree structure or a ternary tree structure. Figure 4E As shown, there are five types of partitions, namely, quadruple partition, horizontal binary partition, vertical binary partition, horizontal ternary partition, and vertical ternary partition.
[0074] In some embodiments, the video encoder 20 may further partition the coding block of the CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. The prediction unit (PU) of a CU may include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements for predicting the prediction blocks. In a monochrome picture or a picture with three separate color planes, a PU may include a single prediction block and a syntax structure for predicting the prediction block. The video encoder 20 may generate predicted luma, Cb, and Cr blocks for the luma, Cb, and Cr prediction blocks of each PU of the CU.
[0075] Video encoder 20 may use intra prediction or inter prediction to generate a prediction block for a PU. If video encoder 20 uses intra prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0076] After the video encoder 20 generates the predicted luma, Cb, and Cr blocks of one or more PUs of a CU, the video encoder 20 may generate a luma residual block of the CU by subtracting the predicted luma block of the CU from its original luma coding block, so that each sample in the luma residual block of the CU indicates the difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block of the CU, respectively, so that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0077] In addition, if Figure 4CAs illustrated, the video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A transform unit (TU) of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements for transforming the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.
[0078] The video encoder 20 may apply one or more transforms to the luma transform block of the TU to generate a luma coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficient may be a scalar. The video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block of the TU. The video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block of the TU.
[0079] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to the process of quantizing transform coefficients to possibly reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode the syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Ultimately, the video encoder 20 may output a bitstream including a sequence of bits forming a representation of a coded frame and associated data, which is stored in the storage device 32 or transmitted to the destination device 14.
[0080] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 can reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the opposite of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on a coefficient block associated with a TU of the current CU to reconstruct a residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coding block of the current CU by adding the samples of the prediction block of the PU of the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the coding block of each CU of the frame, the video decoder 30 can reconstruct the frame.
[0081] As described above, video codecs mainly use two modes, i.e., intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction) to achieve video compression. Palette-based codecs are another codec scheme adopted by many video codec standards. In palette-based codecs that may be particularly suitable for screen-generated content codecs, a video codec (e.g., video encoder 20 or video decoder 30) forms a palette table representing the color of a given block of video data. The palette table includes the most dominant (e.g., frequently used) pixel values in a given block. Pixel values that are not frequently represented in the video data of a given block are not included in the palette table, or are included in the palette table as escape colors.
[0082] Each entry in the palette table includes an index to the corresponding pixel value in the palette table. The palette index of the sample in the block can be encoded and decoded to indicate which entry in the palette table is to be used to predict or reconstruct which sample. The palette mode begins with a process of generating a palette prediction value (predictor) for the first block of a picture, strip, tile, or other such grouping of video blocks. As will be explained below, the palette prediction value for subsequent video blocks is typically generated by updating the previously used palette prediction value. For the purpose of illustration, it is assumed that the palette prediction value is defined at the picture level. In other words, a picture may include multiple coding blocks, each coding block has its own palette table, but there is one palette prediction value for the entire picture.
[0083] In order to reduce the bits required to signal palette entries in a video bitstream, a video decoder can utilize a palette prediction value to determine a new palette entry in a palette table to reconstruct a video block. For example, the palette prediction value may include palette entries from a previously used palette table, or may even be initialized with a most recently used palette table by including all entries of the most recently used palette table. In some embodiments, the palette prediction value may include less than all entries from the most recently used palette table, and then combine some entries from other previously used palette tables. The palette prediction value may have the same size as the palette table used to encode and decode a different block, or may be larger or smaller than the palette table used to encode and decode a different block. In one example, the palette prediction value is implemented as a first-in, first-out (FIFO) table including 64 palette entries.
[0084] To generate a palette table for a block of video data from palette prediction values, a video decoder may receive a one-bit flag for each entry of a palette prediction value from an encoded video bitstream. The one-bit flag may have a first value (e.g., a binary one) indicating that an associated entry of the palette prediction value is to be included in the palette table or a second value (e.g., a binary zero) indicating that an associated entry of the palette prediction value is not to be included in the palette table. If the size of the palette prediction value is larger than the palette table for the block of video data, the video decoder may stop receiving more flags once a maximum size of the palette table is reached.
[0085] In some embodiments, some entries in the palette table can be directly signaled in the encoded video bitstream instead of being determined using palette prediction values. For such entries, the video decoder can receive three separate m-bit values from the encoded video bitstream, indicating the pixel values of the luminance component and two chrominance components associated with the entry, where m represents the bit depth of the video data. Compared with the multiple m-bit values required for palette entries that are directly signaled, those palette entries obtained from the palette prediction values only require a one-bit flag. Therefore, signaling some or all palette entries using palette prediction values can significantly reduce the number of bits required to signal new palette table entries, thereby improving the overall codec efficiency of palette mode codecs.
[0086] In many instances, the palette prediction value of a block is determined based on the palette table used to encode and decode one or more previously encoded blocks. However, when encoding and decoding the first coding tree unit in a picture, slice, or tile, the palette table of the previously encoded block may not be available. Therefore, the entries of the previously used palette table cannot be used to generate the palette prediction value. In this case, a series of palette prediction value initial values can be signaled in a sequence parameter set (SPS) and / or a picture parameter set (PPS), which are the values used to generate the palette prediction value when the previously used palette table is not available. SPS generally refers to the syntax structure of syntax elements applied to a series of consecutive encoded video pictures called a coded video sequence (CVS), as determined by the content of the syntax elements found in the PPS, which refers to the syntax elements found in each slice segment header. PPS generally refers to the syntax structure of syntax elements applied to one or more individual pictures within a CVS, as determined by the syntax elements found in each slice segment header. Therefore, the SPS is generally considered to be a higher-level syntax structure than the PPS, which means that the syntax elements included in the SPS generally change less frequently and apply to a larger portion of the video data than the syntax elements included in the PPS.
[0087] Figures 5A to 5B is a block diagram illustrating an example of applying an adaptive color space conversion (ACT) technique to convert a residual between an RGB color space and a YCgCo color space according to some embodiments of the present disclosure.
[0088] In the HEVC screen content codec extension, ACT is applied to adaptively transform the residual from one color space (e.g., RGB) to another color space (e.g., YCgCo), so that in the YCgCo color space, the correlation (e.g., redundancy) between the three color components (e.g., R, G, and B) is significantly reduced. Further, in the existing ACT design, the adaptation of different color spaces is performed at the transform unit (TU) level by signaling a flag tu_act_enabked_flag for each TU. When the flag tu_act_enabked_flag is equal to one, it indicates that the residual of the current TU is encoded and decoded in the YCgCo space; otherwise (i.e., the flag is equal to 0), it indicates that the residual of the current TU is encoded and decoded in the original color space (i.e., no color space conversion is performed). In addition, different color space transformation formulas are applied depending on whether the current TU is encoded and decoded in lossless mode or lossy mode. Specifically, in Figure 5A The forward and inverse color space transformation formulas between RGB color space and YCgCo color space for lossy mode are defined in .
[0089] For lossless mode, a reversible version of the RGB-YCgCo transform (also called YCgCo-LS) is used. The reversible version of the RGB-YCgCo transform is based on Figure 5B and the lifting operation described in the related description.
[0090] like Figure 5A As shown, the forward and inverse color transform matrices used in lossy mode are not normalized. Therefore, the amplitude of the YCgCo signal is smaller than the amplitude of the original signal after the color transform is applied. In order to compensate for the amplitude drop caused by the forward color transform, the adjusted quantization parameter is applied to the residual in the YCgCo domain. Specifically, when the color space transform is applied, the QP values QPY, QPCg and QPCo used to quantize the YCgCo domain residual are set to QP-5, QP-5 and QP-3, respectively, where QP is the quantization parameter used in the original color space.
[0091] Figure 6 is a block diagram of a technique for applying Luma Mapping with Chroma Scaling (LMCS) in an exemplary video data decoding process according to some embodiments of the present disclosure.
[0092] In VVC, LMCS is applied as a new codec tool before the loop filter (e.g., deblocking filter, SAO, and ALF). Generally, LMCS has two main modules: 1) loop mapping of the luminance component based on an adaptive piecewise linear model; 2) luminance-related chrominance residual scaling. Figure 6 A modified decoding process applying LMCS is shown. Figure 6 In the LMCS, the decoding modules performed in the mapped domain include an entropy decoding module, an inverse quantization module, an inverse transform module, a luma intra prediction module, and a luma sample reconstruction module (i.e., additional luma prediction samples and luma residual samples). The decoding modules performed in the original (i.e., non-mapped) domain include a motion compensation prediction module, a chroma intra prediction module, a chroma sample reconstruction module (i.e., additional chroma prediction samples and chroma residual samples), and all loop filter modules such as a deblocking module, a SAO module, and an ALF module. The new operation modules introduced by LMCS include a luma sample forward mapping module 610, a luma sample inverse mapping module 620, and a chroma residual scaling module 630.
[0093] The loop mapping of LMCS can adjust the dynamic range of the input signal to improve the encoding and decoding efficiency. The loop mapping of the luminance samples in the existing LMCS design is constructed based on two mapping functions: a forward mapping function FwdMap and a corresponding inverse mapping function InvMap. The forward mapping function uses a piecewise linear model with sixteen equal-sized segments to transmit signals from the encoder to the decoder. The inverse mapping function can be obtained directly from the forward mapping function, so it does not need to be transmitted by signal.
[0094] The parameters of the brightness mapping model are signaled at the slice level. First, a presence flag is signaled to indicate whether the brightness mapping model of the current slice is to be signaled. If the brightness mapping model exists in the current slice, the corresponding piecewise linear model parameters are further signaled. In addition, at the slice level, another LMCS control flag is signaled to enable / disable LMCS for the slice.
[0095] The chroma residual scaling module 630 is designed to compensate for the interaction of quantization precision between the luma signal and its corresponding chroma signal when loop mapping is applied to the luma signal. Whether chroma residual scaling is enabled or disabled for the current slice is also signaled in the slice header. When luma mapping is enabled, an additional flag is signaled to indicate whether luma-dependent chroma residual scaling is applied. When luma mapping is not used, luma-dependent chroma residual scaling is always disabled and no additional flag is required. In addition, chroma residual scaling is always disabled for CUs containing less than or equal to four chroma samples.
[0096] Figure 7 is a block diagram illustrating an exemplary video decoding process by which a video decoder implements an inverse adaptive color space transform (ACT) technique according to some embodiments of the present disclosure.
[0097] Similar to the ACT design in HEVC SCC, ACT in VVC converts the intra / inter prediction residual of a CU in 4:4:4 chroma format from the original color space (e.g., RGB color space) to the YCgCo color space. As a result, the redundancy between the three color components can be reduced to obtain better coding efficiency. Figure 7 A decoding flow chart is depicted on how to apply inverse ACT in the VVC framework by adding an inverse ACT module 710. When processing a CU with ACT coding enabled, entropy decoding, inverse quantization, and inverse DCT / DST-based transform are first applied to the CU. Afterwards, Figure 7As depicted, the inverse ACT is called to convert the decoded residual from the YCgCo color space to the original color space (e.g., RGB and YCbCr). In addition, since the ACT in lossy mode is not normalized, a QP adjustment of (-5,-5,-3) is applied to the Y, Cg, and Co components to compensate for the varying magnitude of the transform residual.
[0098] In some embodiments, the ACT method reuses the same ACT core transform of HEVC to perform color conversion between different color spaces. Specifically, two different versions of the color transform are applied depending on whether the current CU is encoded in a lossy or lossless manner. The forward and reverse color transforms for the lossy case use the irreversible YCgCo transform matrix, such as Figure 5A For the lossless case, a reversible color transform YCgCo-LS is applied, as Figure 5B In addition, unlike the existing ACT design, the ACT scheme introduces the following changes to handle its interaction with other codec tools in the VVC standard.
[0099] For example, since the residual of a CU in HEVC may be partitioned into multiple TUs, an ACT control flag is signaled separately for each TU to indicate whether color space conversion needs to be applied. Figure 4E As described, a quadtree nested with binary and ternary partition structures is applied in VVC to replace the concept of multiple partition types, thereby removing the separate CU, PU and TU partitions in HEVC. This means that in most cases, a CU leaf node is also used as a unit for prediction and transform processing without further partitioning, unless the maximum supported transform size is less than the width or height of a component of the CU. Based on such a partition structure, ACT is adaptively enabled and disabled at the CU level. Specifically, a flag cu_act_enables_flag is transmitted with a signal for each CU to select between the original color space and the YCgCo color space for encoding and decoding the residual of the CU. If the flag is equal to 1, it indicates that the residuals of all TUs within the CU are encoded and decoded in the YCgCo color space. Otherwise, if the flag cu_act_enables_flag is equal to 0, all residuals of the CU are encoded and decoded in the original color space.
[0100] In some embodiments, there are different scenarios for disabling ACT. When ACT is enabled for a CU, the residuals of all three components need to be accessed for color space conversion. However, the VVC design cannot guarantee that each CU always contains information of the three components. According to an embodiment of the present disclosure, in those cases where the CU does not contain information of all three components, ACT should be forcibly disabled.
[0101] First, in some embodiments, when a separate tree partition structure is applied, the luma and chroma samples within a CTU are partitioned into CUs based on the separate partition structure. As a result, the CU in the luma partition tree contains only the coding information of the luma component, while the CU in the chroma partition tree contains only the coding information of the two chroma components. According to the current VVC, the switching between the single tree partition structure and the separate tree partition structure is performed at the slice level. Therefore, according to an embodiment of the present disclosure, when it is found that a separate tree is applied to a slice, ACT will always be disabled for all CUs within the slice (including luma CUs and chroma CUs), and the ACT flag will not be transmitted by signal, and it will be inferred to be zero instead.
[0102] Secondly, in some embodiments, when the ISP mode is enabled (further described below), the TU partition is applied only to luma samples, while the chroma samples are encoded and decoded without further partitioning into multiple TUs. Assuming that N is the number of ISP sub-partitions (i.e., TUs) of an intra-frame CU, according to the current ISP design, only the last TU contains both luma and chroma components, while the first N-1 ISP TU consists only of luma components. According to an embodiment of the present disclosure, ACT is disabled in ISP mode. There are two methods to disable ACT for ISP mode. In the first method, the ACT enable / disable flag (i.e., cu_act_enables_flag) is signaled before the syntax of the ISP mode is signaled. In this case, when the flag cu_act_enables_flag is equal to one, the ISP mode is not signaled in the bitstream, but it is always inferred to be zero (i.e., turned off). In the second method, ISP mode signaling is used to bypass the signaling of the ACT flag. Specifically, in this method, the ISP mode is signaled before the flag cu_act_enables_flag. When ISP mode is selected, the flag cu_act_enables_flag is not signaled and is inferred to be zero. Otherwise (ISP mode is not selected), the flag cu_act_enabled_flag is still signaled to adaptively select the color space for residual coding of the CU.
[0103] In some embodiments, in addition to forcibly disabling ACT for CUs whose luma and chroma partition structures are not aligned, LMCS is also disabled for CUs to which ACT is applied. In one embodiment, when a CU selects the YCgCo color space to encode and decode its residual (ie, ACT is one), both luma mapping and chroma residual scaling are disabled. In another embodiment, when ACT is enabled for a CU, only chroma residual scaling is disabled, while luma mapping can still be applied to adjust the dynamic range of the output luma samples. In the last embodiment, for CUs to which ACT is applied to encode and decode their residuals, luma mapping and chroma residual scaling are both enabled. There may be multiple ways to enable chroma residual scaling for CUs to which ACT is applied. In one method, chroma residual scaling is applied before inverse ACT during decoding. By this method, it means that when ACT is applied, chroma residual scaling is applied to chroma residuals (ie, Cg and Co residuals) in the YCgCo domain. In another method, chroma residual scaling is applied after inverse ACT. Specifically, by the second method, chroma scaling is applied to residuals in the original color space. Assuming the input video is captured in RGB format, this means that chroma residual scaling is applied to the residuals of the B and R components.
[0104] In some embodiments, a syntax element (e.g., sps_act_enabled_flag) is added to a sequence parameter set (SPS) to indicate whether ACT is enabled at the sequence level. In addition, since color space conversion is applied to video content with the same resolution for luma and chroma components (e.g., 4:4:4 chroma format 4:4:4), a bitstream conformance requirement needs to be added so that ACT can only be enabled for 4:4:4 chroma format. Table 1 illustrates a modified SPS syntax table with the above syntax added.
[0105]
[0106] Table 1 Modified SPS syntax table
[0107] Specifically, sps_act_enabled_flag equal to 1 indicates that ACT is enabled, and sps_act_enabled_flag equal to 0 indicates that ACT is disabled, so that for CUs referring to SPS, the flag cu_act_enabled_flag is not signaled but is inferred to be 0. When ChromaArrayType is not equal to 3, a requirement for bitstream conformance is that the value of sps_act_enabled_flag shall be equal to 0.
[0108] In another embodiment, instead of always signaling sps_act_enabed_flag, the signaling of the flag depends on the chroma type of the input signal. Specifically, given that ACT can only be applied when luma and chroma components are at the same resolution, the flag sps_act_enabled_falg is signaled only when the input video is captured in 4:4:4 chroma format. With such a change, the modified SPS syntax table is:
[0109]
[0110] Table 2 Modified SPS Syntax Table with Signaling Conditions In some embodiments, syntax design specifications for decoding video data using ACT are described in the following table.
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120] Table 3 Signaling ACT mode specifications
[0121] The flag cu_act_enabled_flag equal to 1 indicates that the residual of the coding unit is coded in the YCgCo color space, and the flag cu_act_enabled_flag equal to 0 indicates that the residual of the coding unit is coded in the original color space (eg, RGB or YCbCr). When the flag cu_act_enabled_flag is not present, it is inferred to be equal to 0.
[0122] In the current VVC working draft, when the input video is captured in 4:4:4 chroma format, the transform skip mode can be applied to both the luminance component and the chroma component. Based on such a design, in some embodiments, three methods are used below to handle the interaction between ACT and transform skip.
[0123] In one method, when transform skip mode is enabled for an ACT CU, transform skip mode is applied only to luma components and not to chroma components. In some embodiments, the syntax design specification for this method is described in the following table.
[0124]
[0125] Table 4 Syntax specification when transform skip mode is applied only to luma component
[0126] In another approach, transform skip mode is applied to both luma and chroma components. In some embodiments, the syntax design specification for this approach is described in the following table.
[0127]
[0128]
[0129] Table 5 Syntax specification when transform skip mode is applied to both luma and chroma components
[0130] In yet another approach, when ACT is enabled for a CU, transform skip mode is always disabled. In some embodiments, the syntax design specifications for this approach are described in the following table.
[0131]
[0132]
[0133] Table 6 Syntax specification when transform skip mode is always disabled
[0134] Fig. 8A and Figure 8B is a block diagram illustrating an exemplary video decoding process by which a video decoder implements techniques for inverse adaptive color space transform (ACT) and luma mapping with chroma scaling according to some embodiments of the present disclosure. In some embodiments, using ACT (e.g., Figure 7 Inverse ACT 710 in ) and chroma residual scaling (e.g., Figure 6 In some other embodiments, the video bitstream is encoded and decoded using chroma residual scaling 630 in ACT. In other embodiments, the video bitstream is encoded and decoded using chroma residual scaling but not ACT at the same time, so inverse ACT 710 is not required.
[0135] More specifically, Fig. 8AAn embodiment is depicted in which the video codec performs chroma residual scaling 630 prior to the inverse ACT 710. As a result, the video codec performs luma mapping in the color space transform domain as well as chroma residual scaling 630. For example, assuming that the input video is captured in RGB format and transformed into the YCgCo color space, the video codec performs chroma residual scaling 630 on the chroma residuals Cg and Co according to the luma residual Y in the YCgCo color space.
[0136] Figure 8B An alternative embodiment is depicted in which the video codec performs chroma residual scaling 630 after the inverse ACT 710. As a result, the video codec performs luma mapping in the original color space domain as well as chroma residual scaling 630. For example, assuming the input video is captured in RGB format, the video codec applies chroma residual scaling to the B and R components.
[0137] Fig. 9 is a block diagram illustrating exemplary decoding logic for switching between performing adaptive color space transform (ACT) and block differential pulse code modulation (BDPCM) according to some embodiments of the present disclosure.
[0138] BDPCM is a codec tool for screen content encoding and decoding. In some embodiments, the BDPCM enable flag is signaled at the sequence level in the SPS. The BDPCM enable flag is signaled only when the transform skip mode is enabled in the SPS.
[0139] When BDPCM is enabled, if the CU size is less than or equal to MaxTsSize×MaxTsSize (in luma samples), and if the CU is intra-coded, a flag is transmitted at the CU level, where MaxTsSize is the maximum block size allowed for transform skip mode. The flag indicates whether conventional intra-coding or BDPCM is used. If BDPCM is used, another BDPCM prediction direction flag is further transmitted to indicate whether the prediction is horizontal or vertical. The block is then predicted using the unfiltered reference samples using the conventional horizontal or vertical intra prediction process. The residual is quantized, and the difference between each quantized residual and its predicted value, i.e., the previously coded residual at the adjacent position horizontally or vertically (depending on the BDPCM prediction direction), is coded.
[0140] For a block of size M (height) × N (width), let r i,j (0≤i≤M-1, 0≤j≤N-1) is the prediction residual. Let Q(r i,j )(0≤i≤M-1, 0≤j≤N-1) represents the quantized version of the residual ri,j. Applying BDPCM to the quantized residual value yields a value with the elements The modified M×N array in, is predicted from its adjacent quantized residual values. For the vertical BDPCM prediction mode, for 0≤j≤(N-1), the following formula is used to obtain
[0141]
[0142] For the horizontal BDPCM prediction mode, for 0≤i≤(M-1), the following formula is used to obtain
[0143]
[0144] At the decoder side, the above process is reversed to calculate Q(r i,j ), 0≤i≤M-1, 0≤j≤N-1, as follows:
[0145]
[0146]
[0147] The inverse quantized residual Q -1 (Q(r i,j )) is added to the intra block prediction value to produce the reconstructed sample value.
[0148] The predicted quantized residual value is converted to Sent to the decoder. For the MPM mode of future intra-mode codecs, if the BDPCM prediction direction is horizontal or vertical, respectively, the horizontal or vertical prediction mode is stored for the BDPCM-coded CU. For deblocking, if the blocks on the block boundary side are all coded using BDPCM, then that particular block boundary will not be deblocked. According to the latest VVC working draft, when the input video is in 4:4:4 chroma format, BDPCM can be applied to both luma and chroma components by signaling two separate flags (i.e., intra_bdpcm_chroma_flag and intra_bdpcm_chroma_flag) for luma and chroma channels at the CU level.
[0149] In some embodiments, the video codec performs different logic to handle the interaction between ACT and BDPCM. For example, when ACT is applied to an intra-frame CU, BDPCM is enabled for the luma component but disabled for the chroma component (910). In some embodiments, when ACT is applied to an intra-frame CU, BDPCM is enabled for the luma component, but BDPCM signaling for the chroma component is disabled. When bypassing the signaling of chroma BDPCM, in one embodiment, the values of intra_bdpcm_chroma_flag and intra_bdpcm_chroma_dir are equal to the values of the luma component (i.e., intra_bdpcm_flag and intra_bdpcm_dir_flag), (i.e., for chroma BDPCM, the same BDPCM direction as luma is used). In another embodiment, when intra_bdpcm_chroma_flag and intrra_bdpcm_chroma_dir_flag are signaled for ACT mode, their values are set to zero (i.e., for chroma components, chroma BDPCM mode is disabled). The corresponding coding unit modification syntax table is as follows:
[0150]
[0151]
[0152]
[0153] Table 7 Enable BDPCM only for the luminance component
[0154] In some embodiments, when ACT is applied to an intra CU, BDPCM (920) is enabled for both the luma component and the chroma component. The corresponding coding unit modification syntax table is as follows:
[0155]
[0156]
[0157]
[0158] Table 8 enables BDPCM for both luminance and chrominance components
[0159] In some embodiments, when ACT is applied to an intra CU, BDPCM is disabled for both the luma component and the chroma component (930). In this case, no BDPCM-related syntax elements need to be signaled. The corresponding coding unit modification syntax table is as follows:
[0160]
[0161]
[0162]
[0163] Table 9 disables BDPCM for both luminance and chrominance components
[0164] In some embodiments, a constrained chroma BDPCM signaling method is used for the ACT mode. Specifically, the signaling of the chroma BDPCM enable / disable flag (i.e., intra_bdpcm_flag) depends on the presence of luma BDPCM (i.e., intra_bdpcm_flag) when ACT is applied. The flag intra_bdpcm_chroma_flag is signaled only when the flag intra_bdpcm_flag is equal to one (i.e., luma BDPCM mode is enabled). Otherwise, the flag intra_bdpcm_chroma_flag is inferred to be zero (i.e., chroma BDPCM is disabled). When the flag intra_bdpcm_chroma_flag is equal to one (i.e., chroma BDPCM is enabled), the BDPCM direction applied to the chroma component is always set equal to the luma BDPCM direction, i.e., the value of the flag intra_bdpcm_chroma_dir_flag is always set equal to the value of intra_bdpcm_flag. The corresponding coding unit modified syntax table is as follows:
[0165]
[0166]
[0167]
[0168] Table 10 Determines the signaling of the chroma BDPCM enable / disable flag depending on the presence of luma BDPCM
[0169] In some embodiments, when ACT is applied, the presence of the flag intra_bdpcm_flag does not depend on the value of intra_bdpcm_flag, but rather conditionally enables the signaling of the chroma BDPCM mode when the intra-frame prediction mode of the luma component is horizontal or vertical. Specifically, with this approach, the flag intra_bdpcm_chroma_flag is signaled only when the luma intra-frame prediction direction is purely horizontal or vertical. Otherwise, the flag intra_bdpcm_chroma_flag is inferred to be 0 (which means disabling chroma BDPCM). When the flag intra_bdpcm_flag is equal to one (i.e., chroma BDPCM is enabled), the BDPCM direction applied to the chroma component is always set equal to the luma intra-frame prediction direction. In the table below, the numbers 18 and 50 represent the current intra-frame prediction indexes for horizontal and vertical intra-frame predictions in the current VVC draft. The corresponding syntax table modified by the coding unit is as follows:
[0170]
[0171]
[0172]
[0173] Table 11 Signaling of the chroma BDPCM enable / disable flag when the intra prediction mode of the luma component is horizontal or vertical
[0174] In some embodiments, context modeling for luma / chroma BDPCM mode is implemented. In the current BDPCM design in VVC, the same context modeling is reused for BDPCM signaling of luma and chroma components. Specifically, a single context is shared by the luma BDPCM enable / disable flag (i.e., intra_bdpcm_flag) and the chroma BDPCM enable / disable flag (i.e., intra_bdpcm_chroma_flag), and another single context is shared by the luma BDPCM direction flag (i.e., intra_bdpcm_dir_flag) and the chroma BDPCM direction flag (i.e., intra_bdpcm_chroma_dir_flag).
[0175] In some embodiments, to improve codec efficiency, in one method, separate contexts are used to signal BDPCM enable / disable for luma and chroma components. In another embodiment, separate contexts are used for signaling BDPCM direction flags for luma and chroma components. In yet another embodiment, two additional contexts are used to encode and decode the chroma BDPCM enable / disable flag, wherein the first context is used to signal intra_bdpcm_chroma_flag when luma BDPCM mode is enabled, and the second context is used to signal intra_bdpcm_chroma_flag when luma BDPCM mode is disabled.
[0176] In some embodiments, ACT is processed with lossless codec. In the HEVC standard, the lossless mode of a CU is indicated by signaling a CU-level flag cu_transquant_bypass_flag as one. However, in the ongoing VVC standardization process, a different lossless enabling method is applied. Specifically, when a CU is encoded and decoded in lossless mode, it is only necessary to skip the transform and use a quantization step of one. This can be achieved by signaling a CU-level QP value of one and signaling a TU-level transform_skip_flag of one. Therefore, in one embodiment of the present disclosure, the lossy ACT transform and the lossless ACT transform are switched for a CU / TU according to the value of transform_skip_flag and the QP value. When the flag transform_skip_flag is equal to one and the QP value is equal to 4, the lossless ACT transform is applied; otherwise, the lossy version of the ACT transform is applied, as shown below.
[0177] If transform_skip_flag is equal to 1 and QP is equal to 4, the residual sample r Y 、r Cb and r Cr The (nTbW)×(nTbH) array is modified as follows, where x=0..nTbW-1, y=0..nTbH-1:
[0178] tmp=r Y [x][y]-(r Cb [x][y]>>1)
[0179] r Y [x][y]=tmp+r Cb [x][y]
[0180] r Cb [x][y]=tmp-(r Cr [x][y]>>1)
[0181] r Cr [x][y]=r Cb [x][y]+rCr[x][y]
[0182] Otherwise, the residual sample points rY, r Cb and r Cr The (nTbW)×(nTbH) array is modified as follows, where x=0..nTbW-1, y=0..nTbH-1:
[0183] tmp=r Y [x][y]-r Cb [x][y]
[0184] r Y [x][y]=rY[x][y]+r Cb [x][y]
[0185] r Cb [x][y]=tmp-r cr [x][y]
[0186] r Cr [x][y]=tmp+r Cr [x][y]
[0187] In the above description, different ACT transform matrices are used for lossy codec and lossless codec. In order to achieve a more unified design, the lossless ACT transform matrix is used for both lossy codec and lossless codec. In addition, since the lossless ACT transform increases the dynamic range of the Cg and Co components by 1 bit, an additional 1-bit right shift is applied to the Cg and Co components after the forward ACT transform, and a 1-bit left shift is applied to the Cg and Co components before the inverse ACT transform. As described below,
[0188] If transform_skip_flag is equal to 0 or QP is not equal to 4, the residual sample r Cb and r Cr The (nTbW)×(nTbH) array is modified as follows, where x=0..nTbW-1, y=0..nTbH-1:
[0189] r Cb [x][y]=r Cb [x][y]<<1
[0190] r Cr [x][y]=r Cr [x][y]<<1
[0191] Then the residual sample array rY, r Cband r Cr The (nTbW)×(nTbH) is modified as follows, where x=0..nTbW-1, y=0..nTbH-1:
[0192] tmp=r Y [x][y]-(r Cb [x][y]>>1)
[0193] rY[x][y]=tmp+r Cb [x][y]
[0194] r Cb [x][y]=tmp-(r Cr [x][y]>>1)
[0195] r Cr [x][y]=r Cb [x][y]+r Cr [x][y]
[0196] In addition, it can be seen from the above that when ACT is applied, QP offsets (-5, -5, -3) are applied to the Y, Cg, and Co components. Therefore, for small input QP values (e.g., <5), negative QP will be used for quantization / dequantization of undefined ACT transform coefficients. To solve this problem, a clipping operation is added after the QP adjustment of ACT so that the applied QP value is always equal to or greater than zero, that is, QP'=max(QP org -QP offset , 0), where QP is the original QP, QP offset is the ACT QP offset, and QP' is the adjusted QP value.
[0197] In the above method, although the same ACT transformation matrix (ie, lossless ACT transformation matrix) is used for both lossy coding and lossless coding, the following two problems can still be identified:
[0198] Different inverse ACT operations are still applied depending on whether the current CU is a lossy CU or a lossless CU. Specifically, for a lossless CU, an inverse ACT transform is applied; for a lossy CU, an additional right shift needs to be applied before the inverse ACT transform. In addition, the decoder needs to know whether the current CU is encoded or decoded in lossy mode or lossless mode. This is inconsistent with the current WC lossless design. Specifically, unlike the HEVC lossless design (in which the lossless mode of a CU is indicated by signaling a cu_transquant_bypass_flag), the lossless encoding and decoding in VVC is performed in a purely non-normative way, i.e., skipping the transform of the prediction residual (enabling transform skip mode for luma and chroma components), selecting an appropriate QP value (i.e., 4), and explicitly disabling codec tools (such as loop filters) that prevent lossless encoding and decoding.
[0199] The QP offset used for normalizing the ACT transform is now fixed. However, the choice of the best QP offset in terms of codec efficiency may depend on the content itself. Therefore, it may be more beneficial to allow flexible QP offset signaling when the ACT tool is enabled in order to maximize its codec gain.
[0200] Based on the above considerations, a unified ACT design is implemented as follows. First, lossless ACT forward transform and inverse transform are applied to CUs encoded and decoded in lossy and lossless modes. Second, instead of using fixed QP offsets, the QP offsets applied to the ACT CU (i.e., three QP offsets applied to the Y, Cg, and Co components) are explicitly signaled in the bitstream. Third, in order to prevent possible overflow problems of the QP applied to the ACT CU, a clipping operation is applied to the resulting QP of each ACT CU to the valid QP range. It can be seen that, based on the above method, the selection between lossy codec and lossless codec can be achieved by pure encoder-only modification (i.e., using different encoder settings). The decoding operations of lossy and lossless codecs for ACT CUs are the same. Specifically, in order to enable lossless codec, in addition to the existing encoder-side lossless configuration, the encoder only needs to signal the values of the three QP offsets to be zero. On the other hand, in order to enable lossy codec, the encoder can signal a non-zero QP offset. For example, in one embodiment, to compensate for the dynamic range change caused by the lossless ACT transform in a lossy codec, QP offsets (-5, 1, 3) can be signaled for the Y, Cg, and Co components when ACT is applied. On the other hand, ACT QP offsets can be signaled at different codec levels, such as sequence parameter set (SPS), picture parameter set (PPS), picture header, coding block group level, etc., so that different QP adaptations can be provided at different granularities. The following table gives an example of performing QP offset signaling in SPS.
[0201]
[0202] Table 12 Syntax specification for QP offset signaling in SPS
[0203] In another embodiment, an advanced control flag (e.g., pictureheader_act_qp_offset_present_flag) is added at the SPS or PPS. When the flag is equal to zero, it means that the QP offset signaled in the SPS or PPS will be applied to all CUs encoded and decoded in ACT mode. Otherwise, when the flag is equal to one, additional QP offset syntax (e.g., picture_header_y_qp_offset_plus5, picture_header_cg_qp_offset_minus 1, and picture_header_co_qp_offset_minus3) can be further signaled in the picture header to individually control the QP value applied to the ACT CU in a specific picture.
[0204] On the other hand, the signaled QP offset should also be applied to limit the final ACT QP value to the valid dynamic range. In addition, different clipping ranges can be applied to CUs coded and decoded with transform and without transform. For example, when no transform is applied, the final QP should be no less than 4. Assuming that the ACT QP offset is signaled at the SPS level, the corresponding QP value of the ACT CU can be described as follows:
[0205] QpY=((qPY_PRED+CuQpDeltaVal+64+2*QpBdOffset+
[0206] sps_act_y_qp_offset)%(64+QpBdOffset))-QpBdOffset
[0207] Qp′ Cb =Clip3(-QpBdOffset,63,qP Cb +pps_cb_qp_offset+slice_cb_qp_offset+CuQpOff
[0208] set Cb +sps_act_cg_offset)+QpBdOffset
[0209] Qp′ Cr =Clip3(-QpBdOffset,63,qP Cr+pps_cr_qp_offset+slice_cr_qp_offset+CuQpOffset
[0210] Cr+sps_act_co_offset)+QpBdOffset
[0211] Qp′ CbCr =Clip3(-QpBdOffset,63,qP CbCr +pps_joint_cbcr_qp_offset+
[0212] slice_joint_cbcr_qp_offset+CuQpOffset CbCr +sps_act_cg_offset)+QpBdOffset
[0213] In another embodiment, the ACT enable / disable flag is signaled at the SPS level, while the ACT QP offset is signaled at the PPS level to allow greater flexibility for the encoder to adjust the QP offset applied to the ACT CU in order to improve encoding and decoding efficiency. Specifically, the SPS and PPS syntax tables with the described changes are shown in the following table.
[0214]
[0215]
[0216]
[0217] Table 13 Syntax specification for signaling ACT enable / disable flag at SPS level and ACT QP offset at PPS level
[0218] pps_act_qp_offset_present_flag equal to 1 specifies that pps_act_y_qp_offset_plus5, pps_act_cg_qp_offset_minus1, and pps_act_co_qp_offset_minus3 are present in the bitstream. When pps_act_qp_offset_present_flag is equal to 0, the syntax elements pps_act_y_qp_offset_plus5, pps_act_cg_qp_offset_minus1, and pps_act_co_qp_offset_minus3 are not present in the bitstream. Bitstream conformance is that the value of pps_act_qp_offset_present_flag shall be 0 when sps_act_enabled_flag is equal to 0.
[0219] pps_act_y_qp_offset_plus5, pps_act_cg_qp_offset_minus1, and pps_act_co_qp_offset_minus3 are used to determine the offsets applied to the quantization parameter values used for the luma and chroma components of coding blocks whose cu_act_enabled_flag is equal to 1. When not present, the values of pps_act_y_qp_offset_plus5, pps_act_cg_qp_offset_minus1, and pps_act_cr_qp_offset_minus3 are inferred to be equal to 0.
[0220] In the above PPS signaling, the same QP offset value is applied to the ACTCU when the joint coding and decoding of chroma residual (JCCR) mode is applied or not. Given that only the residual of the chroma component of a signal is encoded and decoded in the JCCR mode, such a design may not be optimal. Therefore, in order to achieve a better codec gain, when the JCCR mode is applied to an ACT CU, a different QP offset can be applied to encode and decode the residual of the chroma component. Based on this consideration, a separate QP offset signaling is added to the JCCR mode in the PPS, as follows.
[0221]
[0222] Table 14 Syntax specification for adding separate QP offset signaling for JCCR mode in PPS
[0223] pps_joint_cbcr_qp_offset is used to determine the offset applied to the quantization parameter value used for the chroma residual of a coded block to which joint chroma residual coding is applied. When not present, the value of pps_joint_cbcr_qp_offset is inferred to be equal to zero.
[0224] In some embodiments, Fig.10 A method for processing ACT when the internal luma bit depth and chroma bit depth are different is shown. Specifically, Fig.10 is a decoding flow chart for applying different QP offsets for different components when luma and chroma internal bit depths are different according to some embodiments of the present disclosure.
[0225] According to the existing VVC specification, the luminance component and the chrominance component are allowed to use different internal bit depths (expressed as BitDepth Y and BitDepth C ) for encoding and decoding. However, existing ACT designs always assume that the internal luma bit depth and chroma bit depth are the same. In this section, methods are implemented to improve ACT designs when BitDepthy is not equal to BitDepthc.
[0226] In the first method, the ACT tool is always disabled when the internal luma bit depth is not equal to the bit depth of the chroma components.
[0227] In the second approach, the bit depth alignment of luma and chroma components is implemented in the second solution by left-shifting the component with smaller bit depth to match the bit depth of the other component; the scaled components are then rescaled to the original bit depth by right-shifting after color conversion.
[0228] Similar to HEVC, the quantization step size increases by approximately 2 with each increment of QP. 1 / 6 times, exactly doubling every 6 increments. Based on such a design, in the second method, in order to compensate for the internal bit depth between luma and chroma, the QP value for the component with a smaller internal bit depth is increased by 6Δ, where Δ is the difference between the luma internal bit depth and the chroma internal bit depth. Then, by applying a Δ bit right shift, the residual of the component is shifted back to the original dynamic range. Figure 10 illustrates the corresponding decoding process when the above method is applied. For example, assuming that the input QP value is qp, the default QP values applied to the Y, Cg, and Co components are equal to qp-5, qp-5, and qp-3. Further, assuming that the luma internal bit depth is higher than the chroma bit depth, that is, Δ=BitDepthY-BitDepthC. Thus, the final QP values applied to the luma component and the chroma component are equal to qp-5, qp-5+6Δ, and qp-3+6Δ.
[0229] In some embodiments, encoder acceleration logic is implemented. To select a color space for residual codec of a CU, the most straightforward approach is to have the encoder check each codec mode (e.g., intra-codec mode, inter-codec mode, and IBC mode) twice, once with ACT enabled and once with ACT disabled. This may roughly double the encoding complexity. To further reduce the encoding complexity of ACT, the present disclosure implements the following encoder acceleration logic:
[0230] First, since the YCgCo space is more compact than the RGB space, when the input video is in RGB format, the rate-distortion (RD) cost of the ACT tool is checked first, and then the RD cost of the ACT tool is checked when the input video is in RGB format. In addition, the calculation of the RD cost of the color space transform disabled is performed only when there is at least one non-zero coefficient when ACT is enabled. Alternatively, when the input video is in YCbCr format, the RD cost of ACT disabled is checked after the RD check of ACT enabled. The second RD check (i.e., ACT enabled) is performed only when there is at least one non-zero coefficient when ACT is disabled.
[0231] Secondly, in order to reduce the number of codec modes tested, the same codec mode is used for both color spaces. More specifically, for intra mode, the intra prediction mode selected for full RD cost comparison is shared between the two color spaces; for inter mode, the selected motion vectors, reference pictures, motion vector prediction values, and merge indexes (for inter merge mode) are shared between the two color spaces; for IBC mode, the selected block vectors and block vector prediction values and merge indexes (for IBC merge mode) are shared between the two color spaces.
[0232] Third, due to the quad / bin / ternary tree partition structure used in VVC, the same block partition can be obtained through different partition combinations. In order to speed up the selection of color space, an ACT enable / disable decision is implemented when the same block is achieved through different partition paths. Specifically, when the CU is encoded and decoded for the first time, the selected color space used to encode and decode the residual of a specific CU is stored. Then, when the same CU is obtained through another partition path, no choice is made between the two spaces, but the stored color space decision is directly reused.
[0233] Fourth, considering the strong correlation between a CU and its spatial neighbors, the implementation uses the color space selection information of its spatial neighboring blocks to determine how many color spaces need to be checked for the residual encoding and decoding of the current CU. For example, if a sufficient number of spatial neighboring blocks choose the YCgCo space to encode and decode their residuals, then it can be reasonably inferred that the current CU is likely to choose the same color space. Accordingly, the RD check for the encoding and decoding of the residual of the current CU in the original color space can be skipped. If there are enough spatial neighbors that choose the original color space, the RD check for the residual encoding and decoding in the YCgCo domain can be bypassed. Otherwise, both color spaces need to be tested.
[0234] Fifth, given the strong correlation between CUs in the same region, a CU can choose the same color space as its parent CU to encode and decode its residual. Or the child CU can obtain the color space from the information of its parent CU, such as the selected color space and the RD cost of each color space. Therefore, for simple coding complexity, if the residual of a CU's parent CU is encoded in the YCgCo domain, the check of the RD cost of residual encoding and decoding in the RGB domain will be skipped for the CU; in addition, if the residual of its parent CU is encoded in the RGB domain, the check of the RD cost of residual encoding and decoding in the YCgCo domain will be skipped. Another conservative approach is to use the RD cost of its parent CU in both color spaces if two color spaces are tested in the encoding of its parent CU. If its parent CU selects the YCgCo color space and the RD cost of YCgCo is much smaller than the RD cost of RGB, the RGB color space will be skipped, and vice versa.
[0235] In some embodiments, 4:4:4 video codec efficiency is improved by enabling only luma codec tools for chroma components. Because the main focus of VVC design is for videos captured in 4:2:0 chroma format, most existing inter / intra codec tools are enabled only for luma components and disabled for chroma components. But as mentioned earlier, video signals in 4:4:4 chroma format exhibit very different characteristics when compared to 4:2:0 video signals. For example, similar to the luma component, the Cb / B and Cr / R components of 4:4:4YCbCr / RGB video typically contain useful high-frequency texture and edge information. This is different from the chroma components in 4:2:0 video, which are typically very smooth and contain much less information than the luma component. Based on such analysis, when the input video is in 4:4:4 chroma format, the following method is implemented to extend some of the current luma-only codec tools in VVC to chroma components.
[0236] First, enable the luma interpolation filter for the chroma component. Like HEVC, the VVC standard uses motion compensated prediction technology to exploit redundancy between temporally adjacent pictures, which supports motion vectors accurate to sixteen pixels for the Y component and thirty-two pixels for the Cb and Cr components. The fractional samples are interpolated using a set of separable 8-tap filters. The fractional interpolation of the Cb and Cr components is essentially the same as the fractional interpolation of the Y component, except that a separable 4-tap filter is used in the case of the 4:2:0 video format. This is because for 4:2:0 video, the Cb and Cr components contain much less information than the Y component, and compared to using an 8-tap interpolation filter, a 4-tap interpolation filter can reduce the complexity of the fractional interpolation filtering without affecting the efficiency of motion compensated prediction of the Cb and Cr components.
[0237] As previously described, existing 4-tap chroma interpolation filters may not be able to efficiently interpolate fractional samples for motion compensated prediction of chroma components in 4:4:4 video. Therefore, in one embodiment of the present disclosure, the same set of 8-tap interpolation filters (used for luma components in 4:2:0 video) is used for fractional sample interpolation of both luma and chroma components in 4:4:4 video. In another embodiment, in order to achieve a better trade-off between codec efficiency and complexity, adaptive interpolation filter selection is enabled for chroma samples in 4:4:4 video. For example, an interpolation filter selection flag may be signaled at the SPS, PPS, and / or slice levels to indicate whether to use an 8-tap interpolation filter (or other interpolation filter) or a default 4-tap interpolation filter for chroma components at various codec levels.
[0238] Second, enable PDPC and MRL for the chroma components.
[0239] The position-dependent intra prediction combination (PDPC) tool in VVC extends the above idea by employing a weighted combination of intra prediction samples and unfiltered reference samples. In the current VVC working draft, PDPC is enabled for the following intra modes without signaling: planar, DC, horizontal (i.e., mode 18), vertical (i.e., mode 50), angular direction close to the diagonal direction of the lower left corner (i.e., modes 2, 3, 4, .., 10), and angular direction close to the diagonal direction of the upper right corner (i.e., modes 58, 59, 60, ..., 66). Assume that the prediction sample located at coordinates (x, y) is pred(x, y), and its corresponding value after PDPC is calculated as follows
[0240] pred(x,y)=(wL×R-1,y+wT×Rx,-l-wTL×R-1,-1+(64-wL-wT+wTL)×
[0241] pred(x,y)+32)>>6
[0242] Where Rx,-1, R-1,y represent the reference samples located above and to the left of the current sample (x, y), respectively, and R-1,-1 represents the reference sample located at the upper left corner of the current block. The weights wL, wT, and wTL in the above equations are adaptively selected according to the prediction mode and the sample position, as described below, assuming that the size of the current coding block is W×H:
[0243] For DC mode,
[0244] wT=32>>((y<<1)>>shift), wL=32>>((x<<1)>>shift), wTL=(wL>>4)+(
[0245] wT>>4)
[0246] For planar mode,
[0247] wT=32>>((y<<1)>>shift), wL=32>>((x<<1)>>shift), wTL=0
[0248] For horizontal mode:
[0249] wT=32>>((y<<1)>>shift), wL=32>>((x<<1)>>shift), wTL=wT
[0250] For vertical mode:
[0251] wT=32>>((y<<1)>>shift), wL=32>>((x<<1)>>shift), wTL=wL
[0252] For the lower left diagonal direction:
[0253] wT=16>>((y<<1)>>shift), wL=16>>((x<<1)>>shift), wTL=0
[0254] For the upper right diagonal direction:
[0255] wT=16>>((y<<1)>>shift), wL=16>>((x<<1)>>shift), wTL=0
[0256] Among them, shift=(log2(W)-2+log2(H)-2+2)>>2.
[0257] Unlike HEVC, which only uses the nearest reconstructed sample row / column as reference, multiple reference lines (MRL) are introduced in VVC, where two additional rows / columns are used for intra prediction. The index of the selected reference row / column is signaled from the encoder to the decoder. When non-nearest rows / columns are selected, planar mode and DC mode are excluded from the set of intra modes that can be used to predict the current block.
[0258] In the current VVC design, the PDPC tool is only used by the luma component to reduce / eliminate discontinuities between intra-prediction samples and their reference samples (obtained from reconstructed neighboring samples). However, as mentioned earlier, there may be rich texture information in the chroma blocks in a video signal in a 4:4:4 chroma format. Therefore, tools like PDPC that use a weighted average of unfiltered reference samples and intra-prediction samples to improve prediction quality should also be beneficial for improving the chroma encoding and decoding efficiency of 4:4:4 video. Based on such considerations, in one embodiment of the present disclosure, the PDPC process is enabled for intra-prediction of chroma components in 4:4:4 video.
[0259] The same considerations can also be extended to the MRL tool. In the current VVC, MRL cannot be applied to chroma components. Based on an embodiment of the present disclosure, MRL is enabled for chroma components of 4:4:4 video by signaling an MRL index for a chroma component of an intra CU. Based on this embodiment, different methods can be used. In one method, an additional MRL index can be signaled and shared by the Cb / B and Cr / R components. In another method, two MRL indexes are signaled, one for each chroma component. In a third method, the luma MRL index is reused for intra prediction of the chroma components, so that no additional MRL signaling is required to enable MRL for the chroma components.
[0260] Third, enable ISP for the chroma components.
[0261] In some embodiments, a coding tool called sub-partition prediction (ISP) is introduced into VVC to further improve the efficiency of intra-frame coding and decoding. The traditional intra-frame mode only uses the reconstructed samples adjacent to a CU to generate the intra-frame prediction samples of the block. Based on such a design, the spatial correlation between the prediction samples and the reference samples is roughly proportional to the distance between them. Therefore, the internal samples (especially the samples located in the lower right corner of the block) usually have worse prediction quality than the samples close to the block boundary. ISP divides the current CU into 2 or 4 sub-blocks in the horizontal or vertical direction according to the block size, and each sub-block contains at least 16 samples. The reconstructed samples in a sub-block can be used as a reference for predicting the samples in the next sub-block. Repeat the above process until all sub-blocks in the current CU are encoded and decoded. In addition, in order to reduce signaling overhead, all sub-blocks in an ISP CU share the same intra-frame mode. In addition, according to the existing ISP design, sub-block partitioning is only applicable to the luminance component. Specifically, only the luma samples of one ISP CU can be further divided into multiple sub-blocks (or TUs), and each luma sub-block is encoded and decoded separately. However, the chroma samples of the ISP CU are not divided. In other words, for the chroma component, the CU is used as a processing unit for intra prediction, transformation, quantization, and entropy encoding and decoding without further partitioning.
[0262] In the current VVC, when the ISP mode is enabled, TU partitioning is applied only to luma samples, while chroma samples are encoded and decoded without further splitting into multiple TUs. According to an embodiment of the present disclosure, due to the rich texture information in the chroma plane, the ISP mode is also enabled for chroma encoding and decoding in 4:4:4 video. Based on this embodiment, different methods can be used. In one method, an additional ISP index is transmitted with a signal and shared by two chroma components. In another method, two additional ISP indexes are transmitted separately with a signal, one for Cb / B and the other for Cr / R. In a third method, the ISP index used for the luma component is reused for ISP prediction of the two chroma components.
[0263] Fourth, matrix-based intra prediction (MIP) is enabled as a new intra prediction technique for chroma components.
[0264] To predict samples of a rectangular block of width W and height H, MIP takes as input a row of H reconstructed neighboring boundary samples to the left of the block and a row of W reconstructed neighboring boundary samples above the block. If reconstructed samples are not available, they are generated as in conventional intra prediction.
[0265] In some embodiments, the MIP mode is enabled only for the luma component. For the same reason that the ISP mode is enabled for the chroma components, in one embodiment, the MIP is enabled for the chroma components of the 444 video. Two signaling methods can be applied. In the first method, two MIP modes are signaled separately, one for the luma component and the other for the two chroma components. In the second method, only a single MIP mode shared by the luma component and the chroma components is signaled.
[0266] Fifth, enable Multiple Transform Selection (MTS) for chroma components.
[0267] In addition to DCT-II adopted in HEVC, the MTS scheme is also used for residual coding and decoding of both inter-coded blocks and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII.
[0268] In the current VVC, the MTS tool is enabled only for the luma component. In one embodiment of the present disclosure, MIP is enabled for the chroma components of the 444 video. Two signaling methods can be applied. In the first method, when MTS is enabled for a CU, two transform indexes are signaled separately, one for the luma component and the other for the two chroma components. In the second method, when MTS is enabled, a single transform index shared by the luma component and the chroma components is signaled.
[0269] In some embodiments, unlike the HEVC standard that uses a fixed lookup table to derive the used quantization parameter (QP) from the chroma component based on the luma QP, in the VVC standard, a luma-to-chroma mapping table is transmitted from the encoder to the decoder, and the luma-to-chroma mapping table is defined by several pivot points of a piecewise linear function. Specifically, the syntax elements and reconstruction process of the luma-to-chroma mapping table are described as follows:
[0270]
[0271]
[0272] Table 15 Syntax elements and reconstruction process of luma to chroma mapping table
[0273] same_qp_table_for_chroma equal to 1 specifies that only one chroma QP map table is signaled and that it applies to both Cb and Cr residuals and also to the joint Cb-Cr residual when sps_joint_cbcr_enabled_flag is equal to 1. same_qp_table_for_chroma equal to 0 specifies that the chroma QP map tables are signaled in the SPS, two for Cb and Cr and one for the joint Cb-Cr when sps_joint_cbcr_enabled_flag is equal to 1. When same_qp_table_for_chroma is not present in the bitstream, the value of same_qp_table_for_chroma is inferred to be equal to 1.
[0274] qp_table_start_minus26[i] plus 26 specifies the starting luma and chroma QPs used to describe the i-th chroma QP map. The value of qp_table_start_minus26[i] shall be in the range of -26 - QpBdOffset to 36 (inclusive). When qp_table_start_minus26[i] is not present in the bitstream, the value of qp_table_start_minus26[i] is inferred to be equal to 0.
[0275] num_points_in_qp_table_minus1[i] plus 1 specifies the number of points used to describe the i-th chroma QP map table. The value of num_points_in_qp_table_minus1[i] shall be in the range of 0 to 63+QpBdOffset (inclusive). When num_points_in_qp_table_minus1[0] is not present in the bitstream, the value of num_points_in_qp_table_minus1[0] is inferred to be equal to 0.
[0276] delta_qp_in_val_minus1[i][j] specifies the delta value used to derive the input coordinates of the jth pivot point of the i-th chroma QP map. When delta_qp_in_val_minus1[0][j] is not present in the bitstream, the value of delta_qp_in_val_minus1[0][j] is inferred to be equal to 0.
[0277] delta_qp_diff_val[i][j] specifies the delta value used to obtain the output coordinates of the j-th pivot point of the i-th chroma QP map.
[0278] The i-th chroma QP mapping table ChromaQpTable[i] is obtained as follows for i = 0..numQpTables-1:
[0279] qpInVal[i][0] = -qp_table_start-_minus26[i] + 26
[0280] qpOutVal[i][0] = qpInVal[i][0]
[0281] for (j = 0; j <= num_points_in_qp_table_minus1[i]; j++) {
[0282] qpInVal[i][j + 1] = qpInVal[i][j] + delta_qp_in_val_minus1[i][j] + 1
[0283] qpOutVal[i][j + 1] = qpOutVal[i][j] +
[0284] (delta_qp_in_val_minus1[i][j] ^ delta_qp_diff_val[i][j])
[0285] }
[0286] ChromaQpTable[i][qpInVal[i][0]] = qpOutVal[i][0]
[0287] for (k = qpInVal[i][0] - 1; k >= -QpBdOffset; k--)
[0288] ChromaQpTable[i][k] = Clip3(-QpBdOffset, 63,
[0289] ChromaQpTable[i][k + 1] - 1)
[0290] for (j = 0; j < +num_points_in_qp_table_minus1[i]; j++) {
[0291] sh = (delta_qp_in_val_minus1[i][j] + 1) >> 1
[0292] for (k = qpInVal[i][j] + 1, m = 1; k <= qpInval[i][j + 1]; k++, m++)
[0293] ChromaQpTable[i][k]=ChromaQpTable[i][qpInVal[i][j]]+
[0294] ((qpOutVa1[i][j+1]-qpOutVal[i][j])*m+sh) /
[0295] (delta_qp_in_val_minus1[i][j]+1)
[0296] }
[0297] for(k=qpInVal[i][num_points_in_qp_table_minusl[i]+1]+1;k<==63;k++)
[0298] ChromaQpTable[i][k]=Clip3(-QpBdOffset, 63,
[0299] ChromaQpTable[i][k-1]+1)
[0300] In some embodiments, disclosed herein is an improved luma-to-chroma mapping function for RGB video.
[0301] In some embodiments, when the input video is in RGB format, a two-segment linear function is transmitted from the encoder to the decoder to map the luma QP to the chroma QP. This is accomplished by setting the syntax elements same_qp_table_for_chroma=1, qp_table_start_minus26[0]=0, num_points_in_qp_table__minus1[0]=0, delta_qp_in_val_minus1[0][0]=0, and delta_qp_diff_val[0][0]=0. Specifically, the corresponding luma-to-chroma QP mapping function is defined as
[0302]
[0303] Assuming the intra codec bit depth is 10 bits, Table 16 illustrates the luma to chroma QP mapping function applied to the RGB codec.
[0304]
[0305] Table 16 Luma to Chroma QP Mapping Table for RGB Codec
[0306] As shown in (5), when the luma QP is greater than 26, unequal QP values are used to encode and decode the luma component and the chroma component. In some embodiments, given that weighted chroma distortion is used when calculating the RD cost for mode decisions (see the following equation for details), unequal QP values have an impact not only on the quantization / dequantization process, but also on the decisions made during rate-distortion (RD) optimization.
[0307] J mode =(SSE luma +w chroma ·SSE chromd )+λ mode ·R mode (6)
[0308] Among them, SSE luma and SSE chroma are the distortion of the luminance component and the chrominance component respectively; R mode is the number of digits; mode is the Lagrange multiplier; w chroma is the weighting parameter of chroma distortion, which is calculated as
[0309]
[0310] However, there is a stronger correlation between the three channels of RGB video than YCbCr / YUV video. Therefore, when the video content is captured in RGB format (i.e., the information in R, G, and B is equally important), there is usually strong texture and high-frequency information in all three components. Therefore, in one embodiment of the present disclosure, equal QP values are applied to all three channels for RGB encoding and decoding. This can be accomplished by setting the corresponding luma to chroma QP mapping syntax elements to same_qp_table_for_chroma=1, qp_table_start_minus26[0]=0, num_points_in_qp_table_minus1[0]=0, delta_qp_in_val_minus1[0][0]=0, and delta_qp_diff_val[0][0]=1. Accordingly, using the method disclosed herein, the RGB luma to chroma mapping function is illustrated by equation (8) and Table 17.
[0311] QP c =QP L (8)
[0312]
[0313] Table 17 Luma to Chroma QP Mapping Table for RGB Codec
[0314] Fig.11A and Fig. 11B is a block diagram illustrating an exemplary video decoding process according to some embodiments of the present disclosure, by which a video decoder implements a truncation technique to limit the dynamic range of a residual of a coding unit to a predefined range for processing by inverse ACT.
[0315] More specifically, Fig.11A An example of applying an intercept operation 1102 to the input of an inverse ACT 1104 is shown. Fig. 11B An example is shown where intercept operations 1102 and 1106 are applied to both the input and output of inverse ACT 1104. In some embodiments, intercept operations 1102 and 1106 are the same. In some embodiments, intercept operations 1102 and 1106 are different.
[0316] In some embodiments, the bit depth control method is implemented with ACT. According to the existing ACT design, the input of the inverse ACT process at the decoder is the output residual from other residual decoding processes (e.g., inverse transform, inverse BDPCM, and inverse JCCR). In the current implementation, those residual samples can reach the maximum value of a 16-bit signed integer. With such a design, the inverse ACT cannot be implemented through a 16-bit implementation, which is very expensive for hardware implementation. To solve this problem, a truncation operation is applied to the input residual of the inverse ACT process. In one embodiment, the following truncation operation is applied to the input residual of the inverse ACT process:
[0317] Clip input =Clip(-(2 Bitdepth -1), 2 Bitdepth -1, M)
[0318] Among them, Bitdepth is the internal codec bit depth.
[0319] In another embodiment, another truncation operation is applied to the input residual of the inverse ACT process as follows:
[0320] Clip input =Clip(-(2 15 -1), 2 15 -1, M)
[0321] In addition, in another embodiment, a truncation operation is applied to the output residual of the inverse ACT. In one embodiment, the following truncation operation is applied to the output of the inverse ACT:
[0322] Clip output =Clip(-(2 Bitdepth -1), 2 Bitdepth -1, M)
[0323] In another embodiment, the following truncation operation is applied to the output of the inverse ACT:
[0324] Clip output =Clip(-(2 15 -1), 2 15 -1, M)
[0325] like Figure 5B As shown in the lifting operation depicted in , when the reversible YCgCo transform is applied, the dynamic range of the Cg and Co components will increase by 1 bit due to the lifting operation. Therefore, in order to maintain the accuracy of the residual samples output from the reversible YCgCo transform, in one embodiment, the input residual of the inverse ACT is truncated based on the following equation:
[0326] Clip input =Clip(-2 Bitdepth+1 , 2 Bitdepth+1 -1, M)
[0327] In another embodiment, different clipping operations are applied to the input Y, Cg and Co residuals of the inverse ACT. Specifically, since the bit depth of the Y component remains unchanged before and after the reversible ACT transform, the input luminance residual of the inverse ACT is clipped by the following operation:
[0328] Clip input =Clip(-2 Bitdepth , 2 Bitdepth -1, M)
[0329] For the Cg and Co components, due to the increase in bit depth, the corresponding input residuals of the inverse ACT are truncated by the following operation:
[0330] Clip input =Clip(-2 Bitdepth+1 , 2 Bitdepth+1 -1, M)
[0331] In another embodiment, the input residual of the inverse ACT is truncated by the following operation:
[0332] Clip input =Clip(-2 C , 2 C -1, M)
[0333] Among them, C is a fixed number.
[0334] Fig.121200 is a flowchart illustrating an exemplary process according to some embodiments of the present disclosure, by which a video decoder (e.g., video decoder 30) decodes video data by performing a truncation operation that limits the dynamic range of the residual of a coding unit to within a predefined range for processing by inverse ACT.
[0335] The video decoder 30 receives video data corresponding to a coding unit from a bitstream, wherein the coding unit is coded or decoded by an intra prediction mode or an inter prediction mode (1210).
[0336] Video decoder 30 then receives a first syntax element from the video data. The first syntax element indicates whether the coding unit has been coded using adaptive color space transform (ACT) (1220).
[0337] Then, the video decoder 30 processes the video data to generate a residual of the coding unit (1230). Based on determining that the coding unit has been coded using ACT based on the first syntax element, the video decoder 30 performs a truncation operation on the residual of the coding unit (1240). The video decoder 30 applies an inverse ACT to the residual of the coding unit after the truncation operation (1250).
[0338] In some embodiments, the clipping operation limits the dynamic range of the residual of the coding unit to a predefined range for processing by the inverse ACT.
[0339] In some embodiments, the intercept operation is defined as
[0340] Clip input =Clip(-2 Bitdepth+1 , 2 Bitdepth+1 -1, M)
[0341] Among them, M is the input of the clipping operation, Bitdepth is the internal codec bit depth, Clip input is the output of the interception operation, ranging from -2 Bitdepth+1 to(2 Bitdepth+1 -1) In some embodiments, if the input M is less than -2 Bitdepth+1 , the output of the intercept operation is set to -2 Bitdepth+1 In some embodiments, if the input M is greater than (2 Bitdepth+1 -1), the output of the intercept operation is set to (2 Bitdepth+1 -1).
[0342] In some embodiments, the intercept operation is defined as
[0343] Clip input =Clip(-2 Bitdepth , 2Bitdepth -1, M)
[0344] Among them, M is the input of the clipping operation, Bitdepth is the internal codec bit depth, Clip input is the output of the interception operation, ranging from -2 Bitdepth to(2 Bitdepth -1) In some embodiments, if the input M is less than -2 Bitdepth , the output of the intercept operation is set to -2 Bitdepth In some embodiments, if the input M is greater than (2 Bitdepth -1), the output of the intercept operation is set to ( 2Bitdepth -1).
[0345] In some embodiments, before performing the truncation operation, video decoder 30 applies an inverse transform to the residual of the coding unit.
[0346] In some embodiments, after applying the inverse ACT to the residual of the coding unit, video decoder 30 applies a second truncation operation to the residual of the coding unit.
[0347] In some embodiments, the clipping operation adjusts the dynamic range of the residual of the coding unit to within the fixed intra-codec bit depth implemented by inverse ACT.
[0348] In some embodiments, a fixed intra-codec bit depth of 15 is implemented.
[0349] In some embodiments, after receiving the first syntax element from the video data, video decoder 30 receives a second syntax element from the video data, wherein the second syntax element indicates a variable intra-codec bit depth used in the truncation operation.
[0350] In some embodiments, after receiving the first syntax element from the video data, the video decoder 30 receives a second syntax element from the video data, wherein the second syntax element indicates a first intra-codec bit depth, and the second intra-codec bit depth used in the truncation operation is the first intra-codec bit depth plus 1.
[0351] In some embodiments, the clipping operation further includes a first clipping operation applied to a luma component of the residual of the coding unit, and a second clipping operation applied to a chroma component of the residual of the coding unit.
[0352] In some embodiments, the first clipping operation limits the dynamic range of the luminance component of the residual of the coding unit to within the range of the first internal codec bit depth plus 1, and the second clipping operation limits the dynamic range of the chrominance component of the residual of the coding unit to within the range of the second internal codec bit depth plus 1, and the second internal codec bit depth is the first internal codec bit depth plus 1.
[0353] In one or more examples, the described functions can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted through a computer-readable medium as one or more instructions or codes and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media corresponding to tangible media such as data storage media or communication media including any media that facilitates, for example, transferring a computer program from one place to another according to a communication protocol. In this way, computer-readable media can generally correspond to (1) non-transient tangible computer-readable storage media or (2) communication media such as signals or carrier waves. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to obtain instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product can include a computer-readable medium.
[0354] The terms used in the description of the embodiments herein are for the purpose of describing specific embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more items in the associated enumerated items. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of stated features, elements, and / or parts, but do not exclude the presence or addition of one or more other features, elements, parts, and / or groups thereof.
[0355] It should also be understood that although the terms first, second, etc. can be used to describe various elements in this article, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the scope of the embodiment, the first electrode can be referred to as the second electrode, and similarly, the second electrode can be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but the first electrode and the second electrode are not the same electrode.
[0356] The description of the present application has been presented for the purpose of illustration and description, and the description is not intended to be exhaustive or limited to the invention in the form disclosed. With the benefit of the teachings presented in the foregoing description and the associated drawings, many modifications, variations and alternative embodiments will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to best explain the principles of the invention, the practical application, and to enable other persons skilled in the art to understand the various embodiments of the invention and to best utilize the basic principles and various embodiments with various modifications suitable for the intended specific use. Therefore, it should be understood that the scope of the claims should not be limited to the specific examples of the disclosed embodiments, and modifications and other embodiments are intended to be included within the scope of the appended claims.
Claims
1. A video encoding method, comprising: Obtain multiple coding units segmented from a video picture; For each of the coding units: generating a first syntax element, the first syntax element indicating whether the coding unit has been encoded using an adaptive color space transform ACT; The first syntax element is transmitted to a decoding side, so that the decoding side performs an operation when determining, according to the first syntax element, that the coding unit has been encoded using the adaptive color space transform ACT, the operation comprising: Performing a truncation operation on the residual of the coding unit; applying an inverse ACT to the residual of the coding unit after the truncation operation; and reconstructing the coding unit based on the residual after applying the inverse ACT, The clipping operation further includes a first clipping operation applied to a luminance component of the residual of the coding unit, and a second clipping operation applied to a chrominance component of the residual of the coding unit.
2. The method according to claim 1, wherein: The clipping operation limits the dynamic range of the residual of the coding unit to a predefined range for processing by the inverse ACT.
3. The method according to claim 1, wherein: One or more of the first interception operation and the second interception operation is defined as Clip input =Clip(-2 Bitdepth+1 ,2 Bitdepth+1 -1,M) Where M is the input of the clipping operation, Bitdepth is the internal codec bit depth, Clip input is the output of the intercept operation, the output of the intercept operation is in the range -2 Bitdepth+1 to (2 Bitdepth+1 -1).
4. The method according to claim 1, wherein: The operations further include applying an inverse transform to generate the residual of the coding unit before performing the truncation operation.
5. The method according to claim 1, wherein: The operations further include applying a third truncation operation to the residual of the coding unit after applying the inverse ACT to the residual of the coding unit.
6. The method according to claim 1, wherein: The clipping operation adjusts a dynamic range of the residual of the coding unit within a fixed intra-codec bit depth range to be processed by the inverse ACT, wherein the fixed intra-codec bit depth range is determined based on a fixed intra-codec bit depth.
7. The method according to claim 6, wherein: The fixed intra-codec bit depth is 15.
8. The method according to claim 1, further comprising: After generating the first syntax element, a second syntax element is generated, wherein the second syntax element indicates a variable intra-codec bit depth used in the truncation operation.
9. The method according to claim 1, further comprising: After generating the first syntax element, a second syntax element is generated, wherein the second syntax element indicates a first intra codec bit depth, and a second intra codec bit depth used in the truncation operation is the first intra codec bit depth plus 1.
10. The method according to claim 1, wherein: The first clipping operation limits the dynamic range of the luminance component of the residual of the coding unit to a range of a first internal codec bit depth plus 1, and the second clipping operation limits the dynamic range of the chrominance component of the residual of the coding unit to a range of a second internal codec bit depth plus 1, and the second internal codec bit depth is the first internal codec bit depth plus 1.
11. An electronic device comprising: one or more processing units; a memory coupled to the one or more processing units; as well as A plurality of programs stored in the memory, when the plurality of programs are executed by the one or more processing units, enable the electronic device to perform the method as claimed in any one of claims 1 to 10 to generate a video bitstream and store it in the memory.
12. A non-transitory computer-readable storage medium storing a plurality of programs and bitstreams for execution by an electronic device having one or more processing units, wherein: When the plurality of programs are executed by the one or more processing units, the electronic device executes the method according to any one of claims 1 to 10 to generate the bitstream.
13. A computer program product comprising instructions for execution by a computing device having one or more processors, wherein: When the instructions are executed by the one or more processors, the computing device performs the method of any one of claims 1 to 10 to generate a video bitstream.
Citation Information
Patent Citations
Inter-component de-correlation for video coding
CN107079157A
Video encoding method and system using adaptive color conversion
JP2017005688A