Method and apparatus for video encoding using a palette mode

By using the palette mode in video encoding and decoding, the problem of efficient processing of high-quality video data is solved, and more efficient video encoding and decoding is achieved.

CN118741091BActive Publication Date: 2025-06-17BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411015769.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-01-11
Filing Date
2021-01-11
Publication Date
2025-06-17
Estimated Expiration
2041-01-11

AI Technical Summary

Technical Problem

With the improvement of digital video quality, the amount of video data to be encoded/decoded has increased exponentially. How to encode/decode more efficiently while maintaining image quality has become a challenge.

Method used

The video encoding and decoding is performed using the palette mode, and the minimum palette mode block size is determined by receiving syntax elements associated with the encoding unit from the bitstream, and the flag decoding encoding unit is enabled according to the palette mode.

Benefits of technology

Improve the efficiency of video encoding and decoding, reduce the demand for video bitstreams, and enhance the encoding performance under high-quality video conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118741091B_ABST
    Figure CN118741091B_ABST
Patent Text Reader

Abstract

A method for an electronic device to perform decoding of video data. The method includes: receiving, from a bitstream, a plurality of syntax elements associated with a coding unit, wherein the plurality of syntax elements indicate a size of the coding unit and a coding tree type of the coding unit; determining, according to the coding tree type of the coding unit, a minimum palette mode block size for the coding unit; according to determining that the size of the coding unit is greater than the minimum palette mode block size: receiving, from the bitstream, a palette mode enable flag associated with the coding unit; and decoding the coding unit from the bitstream according to the palette mode enable flag.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese Patent Application No. 202180019765.8, which is the national phase application in China of the international patent application PCT / US2021 / 012964 filed on January 11, 2021, and the international patent application claims the priority of the US patent application No. 62 / 959,913 filed on January 11, 2020. Technical Field

[0002] This application generally relates to video data encoding, decoding, and compression, and specifically, to methods and systems for video encoding and decoding using a palette mode. Background Art

[0003] Various electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc., support digital video. The electronic devices send, receive, encode, decode, and / or store digital video data by implementing video compression / decompression standards defined by MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) standards. Video compression typically includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove the redundancy inherent in the video data. For block-based video encoding and decoding, a video frame is divided into one or more slices, each slice having a plurality of video blocks, which may also be referred to as coding tree units (CTUs). Each CTU may contain one coding unit (CU) or be recursively divided into smaller CUs until a preset minimum CU size is reached. Each CU (also referred to as a leaf CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be encoded and decoded in intra-frame, inter-frame, or IBC mode. Intra-frame encoding (I) of video blocks in a slice of a video frame is performed using spatial prediction relative to reference samples in adjacent blocks within the same video frame. Inter-frame encoding (P or B) of video blocks in a slice of a video frame can use spatial prediction relative to reference samples in adjacent blocks within the same video frame or temporal prediction relative to reference samples in other previous and / or future reference video frames.

[0004] Spatial or temporal prediction based on previously encoded reference blocks (e.g., neighboring blocks) generates a prediction block for a current video block to be coded / decoded. The process of finding the reference blocks can be done through a block matching algorithm. Residual data representing the pixel difference between the current block to be coded / decoded and the prediction block is referred to as a residual block or prediction error. An inter-coded block is coded based on a motion vector pointing to a reference block in a reference frame forming the prediction block, and the residual block. The process of determining the motion vector is typically referred to as motion estimation. An intra-coded block is coded based on an intra-prediction mode and the residual block. For further compression, the residual block is transformed from the pixel domain to a transform domain, such as the frequency domain, to generate residual transform coefficients, which can then be quantized. The quantized transform coefficients, initially arranged as a two-dimensional array, can be scanned to produce a one-dimensional vector of transform coefficients, which is then entropy coded into a video bitstream for more compression.

[0005] Then, the encoded video bitstream is saved in a computer-readable storage medium (e.g., flash memory) to be accessed by another electronic device having digital video capabilities, or directly transmitted to the electronic device in a wired or wireless manner. Then, the electronic device performs video decompression (which is a process opposite to the video compression described above) by, for example, parsing the encoded video bitstream to obtain syntax elements from the bitstream and reconstructing the digital video data from the encoded video bitstream into its original format at least partially based on the syntax elements obtained from the bitstream, and renders the reconstructed digital video data on a display of the electronic device.

[0006] As digital video quality goes from high definition to 4K×2K or even 8K×4K, the amount of video data to be coded / decoded grows exponentially. There has been a challenge in encoding / decoding video data more efficiently while maintaining the image quality of the decoded video data. Summary of the Invention

[0007] This application describes embodiments related to video data encoding and decoding, and more specifically, describes embodiments related to systems and methods for video encoding and decoding using a palette mode.

[0008] According to a first aspect of the present application, a method for decoding video data includes: receiving, from a bitstream, a plurality of syntax elements associated with a coding unit, where the plurality of syntax elements indicate a size of the coding unit and a coding tree type of the coding unit; determining, based on the coding tree type of the coding unit, a minimum palette mode block size for the coding unit; determining that the size of the coding unit is greater than the minimum palette mode block size: receiving, from the bitstream, a palette mode enable flag associated with the coding unit; and decoding the coding unit from the bitstream according to the palette mode enable flag.

[0009] According to a second aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory. When executed by the one or more processing units, the programs cause the electronic device to perform the method of decoding video data as described above.

[0010] According to a third aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by an electronic device having one or more processing units. When executed by the one or more processing units, the programs cause the electronic device to perform the method of decoding video data as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification, illustrate the described embodiments and together with the description serve to explain the basic principles. Like reference numerals refer to corresponding parts.

[0012] Figure 1 is a block diagram illustrating an exemplary video encoding and decoding system according to some embodiments of the present disclosure.

[0013] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0014] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0015] Figures 4A to 4E is a block diagram illustrating how a frame is recursively partitioned into a plurality of video blocks having different sizes and shapes according to some embodiments of the present disclosure.

[0016] Figures 5A to 5D is a block diagram illustrating an example of using a palette table to encode and decode video data according to some embodiments of the present disclosure.

[0017] Figure 6 is a flowchart illustrating an exemplary process by which a video decoder implements the technique of decoding video data according to some embodiments of the present disclosure.

[0018] Figure 7 is a block diagram illustrating an example context adaptive binary arithmetic coding / decoding (CABAC) engine according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0019] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that various alternative solutions can be used without departing from the scope of the claims, and the subject matter can be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein can be implemented on many types of electronic devices having digital video capabilities.

[0020] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. As Figure 1 shown, system 10 includes a source device 12 that generates and encodes video data to be decoded by a destination device 14 at a later time. The source device 12 and the destination device 14 can include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some embodiments, the source device 12 and the destination device 14 are equipped with wireless communication capabilities.

[0021] In some embodiments, the destination device 14 can receive the encoded video data to be decoded via a link 16. The link 16 can include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 can include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data can be modulated and transmitted to the destination device 14 according to a communication standard such as a wireless communication protocol. The communication medium can include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network (such as the Internet). The communication medium can include routers, switches, base stations, or any other device that can be used to facilitate communication from the source device 12 to the destination device 14.

[0022] In some other embodiments, the encoded video data may be transmitted from the output interface 22 to the storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of a variety of distributed or local access data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile memory, a non-volatile memory, or any other suitable digital storage media for storing the encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device 14 may access the encoded video data through any standard data connection, which includes a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0023] As Figure 1 shown, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include sources such as a video capture device, e.g., a camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As an example, if the video source 18 is a camera of a security surveillance system, the source device 12 and the destination device 14 may form a camera phone or a video phone. However, the embodiments described in the present application can generally be applicable to video coding and decoding and can be applied to wireless and / or wired applications.

[0024] The captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be directly transmitted to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may also (or alternatively) be stored on the storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.

[0025] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 can include a receiver and / or a modem and receives encoded video data via a link 16. The encoded video data transmitted via the link 16 or provided on a storage device 32 can include various syntax elements generated by the video encoder 20 for use by the video decoder 30 to decode the video data. Such syntax elements can be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored in a file server.

[0026] In some embodiments, the destination device 14 can include a display device 34, which can be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data to a user and can include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0027] The video encoder 20 and the video decoder 30 can operate according to proprietary or industry standards, such as VVC, HEVC, MPEG-4 Part 10, Advanced Video Coding (AVC), or extensions of such standards. It should be understood that the present application is not limited to a particular video coding / decoding standard and can be applicable to other video coding / decoding standards. In general, it is contemplated that the video encoder 20 of the source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is generally also contemplated that the video decoder 30 of the destination device 14 can be configured to decode video data according to any of these current or future standards.

[0028] The video encoder 20 and the video decoder 30 can each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, and any of the one or more encoders or decoders can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0029] Figure 2FIG. is a block diagram illustrating an exemplary video encoder 20 in accordance with some embodiments described in the present application. The video encoder 20 may perform intra prediction coding and inter prediction coding of video blocks within a video frame. Intra prediction coding relies on spatial prediction to reduce or remove spatial redundancy of video data within a given video frame or picture. Inter prediction coding relies on temporal prediction to reduce or remove temporal redundancy of video data within adjacent video frames or pictures of a video sequence.

[0030] As Figure 2 shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In some embodiments, the video encoder 20 further includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A deblocking filter (not shown) may be located between the adder 62 and the DPB 64 to filter block boundaries to remove block effect artifacts from the reconstructed video. A loop filter (not shown) other than the deblocking filter may also be used to filter the output of the adder 62. The video encoder 20 may be in the form of fixed or programmable hardware units, or may be partitioned among one or more of the illustrated fixed or programmable hardware units.

[0031] The video data memory 40 may store video data to be encoded by components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from a video source 18. The DPB 64 is a buffer that stores reference video data for encoding video data by the video encoder 20 (e.g., in intra prediction coding mode or inter prediction coding mode). The video data memory 40 and the DPB 64 may be formed of any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20, or off-chip relative to those components.

[0032] As Figure 2As shown, after receiving the video data, the partitioning unit 45 in the prediction processing unit 41 partitions the video data into video blocks. This partitioning may also include partitioning the video frame into strips, tiles, or other larger coding units (CUs) according to a preset partitioning structure, such as a quadtree structure associated with the video data. The video frame may be partitioned into multiple video blocks (or a set of video blocks referred to as tiles). The prediction processing unit 41 may select one of multiple possible prediction coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one of multiple intra prediction coding modes or one of multiple inter prediction coding modes. The prediction processing unit 41 may provide the resulting intra prediction coded block or inter prediction coded block to the adder 50 to generate a residual block, and provide it to the adder 62 to reconstruct the coded block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra mode indicators, partitioning information, and other such syntax information to the entropy coding unit 56.

[0033] To select an appropriate intra prediction coding mode for the current video block, the intra prediction processing unit 46 in the prediction processing unit 41 may perform intra prediction coding of the current video block relative to one or more adjacent blocks in the same frame as the current block to be coded and decoded to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 in the prediction processing unit 41 perform inter prediction coding of the current video block relative to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 may execute multiple coding channels, for example, to select an appropriate coding mode for each block of the video data.

[0034] In some embodiments, the motion estimation unit 42 determines the inter prediction mode of the current video frame by generating a motion vector according to a predetermined pattern within the video frame sequence, where the motion vector indicates the displacement of the prediction unit (PU) of the video block within the current video frame relative to the prediction block within the reference video frame. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion of the video block. The motion vector may, for example, indicate the displacement of the PU of the video block within the current video frame or picture relative to the prediction block (or other coded unit) within the reference frame, where the prediction block is relative to the current block (or other coded unit) coded within the current frame. The predetermined pattern may specify the video frames in the sequence as P frames or B frames. The intra BC unit 48 may determine the vector for performing intra BC coding in a manner similar to the way the motion estimation unit 42 determines the motion vector for inter prediction, e.g., a block vector, or may utilize the motion estimation unit 42 to determine the block vector.

[0035] A prediction block is a block in a reference frame that is considered to closely match a PU of a video block to be coded or decoded, and the pixel difference can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some embodiments, the video encoder 20 may calculate values at sub-integer pixel positions of the reference frames stored in the DPB 64. For example, the video encoder 20 may insert values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frames. Thus, the motion estimation unit 42 may perform a motion search with respect to full-pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy.

[0036] The motion estimation unit 42 calculates a motion vector of a PU of a video block in an inter-predicted coded frame by comparing the position of the PU with the position of a prediction block of a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1), and each of the lists identifies one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44, and then to the entropy coding unit 56.

[0037] The motion compensation performed by the motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by the motion estimation unit 42. After receiving the motion vector of the PU of the current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, obtain the prediction block from the DPB 64 and forward the prediction block to the adder 50. Then, the adder 50 forms a residual video block with pixel differences by subtracting the pixel values of the prediction block provided by the motion compensation unit 44 from the pixel values of the coded current video block. The pixel differences forming the residual video block may include luminance differences or chrominance differences or both. The motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use by the video decoder 30 when decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vectors for identifying the prediction blocks, any flags indicating the prediction mode, or any other syntax information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes.

[0038] In some embodiments, the intra BC unit 48 may generate vectors and obtain prediction blocks in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, but where the prediction block is in the same frame as the current block being encoded, and where, relative to a motion vector, the vector is referred to as a block vector. Specifically, the intra BC unit 48 may determine an intra prediction mode for encoding the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during a separate encoding pass, and test its performance via rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode among the various tested intra prediction modes to use and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values using rate-distortion analysis for the various tested intra prediction modes and select the intra prediction mode having the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original, unencoded block (which was encoded to produce the encoded block) and the bit rate (i.e., number of bits) used to produce the encoded block. The intra BC unit 48 may calculate a ratio from the distortion and rate of each encoded block to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0039] In other examples, the intra BC unit 48 may use the motion estimation unit 42 and the motion compensation unit 44, in whole or in part, to perform such functions for intra BC prediction in accordance with the embodiments described herein. In either case, for intra block copy, the prediction block may be a block that is considered to closely match the block to be coded and decoded in terms of pixel differences, which may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics, and the identification of the prediction block may include calculating values at sub-integer pixel positions.

[0040] Regardless of whether the prediction block is from the same frame according to intra prediction or from a different frame according to inter prediction, the video encoder 20 may form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming pixel difference values. The pixel difference values for forming the residual video block may include luminance component differences and chrominance component differences.

[0041] As described above, the intra prediction processing unit 46 may perform intra prediction on the current video block as an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or the intra block copy prediction performed by the intra BC unit 48. Specifically, the intra prediction processing unit 46 may determine an intra prediction mode for encoding the current block. To this end, the intra prediction processing unit 46 may, for example, encode the current block using various intra prediction modes during a separate encoding pass, and the intra prediction processing unit 46 (or in some examples, the mode selection unit) may select an appropriate intra prediction mode from the tested intra prediction modes for use. The intra prediction processing unit 46 may provide information indicating the selected intra prediction mode of the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra prediction mode in the bitstream.

[0042] After the prediction processing unit 41 determines the predicted block of the current video block via inter prediction or intra prediction, the adder 50 forms a residual video block by subtracting the predicted block from the current video block. The residual video data in the residual block may be included in one or more transform units (TUs) and provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0043] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy encoding unit 56 may perform the scan.

[0044] After quantization, the entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques. The encoded bitstream may then be transmitted to the video decoder 30 or archived in the storage device 32 for later transmission to the video decoder 30 or retrieval by the video decoder. The entropy encoding unit 56 may also entropy encode the motion vectors and other syntax elements of the encoded current video frame.

[0045] The dequantization unit 58 and the inverse transform processing unit 60 respectively apply dequantization and inverse transform to reconstruct the residual video block in the pixel domain to generate a reference block for predicting other video blocks. As described above, the motion compensation unit 44 can generate a motion-compensated prediction block from one or more reference blocks of the frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for motion estimation.

[0046] The adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by the motion compensation unit 44 to generate a reference block for storage in the DPB 64. The reference block can then be used by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 as a prediction block for inter-frame prediction of another video block in a subsequent video frame.

[0047] Figure 3 FIG. is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, a dequantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 can perform a decoding process generally opposite to the encoding process described above in connection with Figure 2 the video encoder 20. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra prediction unit 84 can generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.

[0048] In some examples, tasks can be assigned to the units of the video decoder 30 to implement the embodiments of the present application. Similarly, in some examples, the embodiments of the present disclosure can be divided among one or more units of the video decoder 30. For example, the intra BC unit 85 can implement the embodiments of the present application alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 can be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.

[0049] The video data memory 79 can store video data to be decoded by other components of the video decoder 30, such as an encoded video bitstream. For example, the video data stored in the video data memory 79 can be obtained from a storage device 32, a local video source (such as a camera), via wired or wireless network transmission of the video data or by accessing a physical data storage medium (e.g., a flash drive or a hard disk). The video data memory 79 can include a coded picture buffer (CPB) that stores coded video data from the encoded video bitstream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use by the video decoder 30 in decoding video data (e.g., in an intra prediction coding mode or an inter prediction coding mode). The video data memory 79 and the DPB 92 can be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are depicted in Figure 3 as two different components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 can be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 can be on-chip with other components of the video decoder 30 or off-chip relative to those components.

[0050] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 can receive syntax elements at the video frame level and / or at the video block level. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantized coefficients, motion vectors, or intra prediction mode indicators and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors and other syntax elements to the prediction processing unit 81.

[0051] When a video frame is encoded as an intra prediction coding (I) frame or for intra-coded prediction blocks in other types of frames, the intra prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for video blocks of the current video frame based on the signal-transmitted intra prediction mode and reference data from previously decoded blocks of the current frame.

[0052] When a video frame is encoded as an inter-predicted (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks of video blocks of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each prediction block may be generated from a reference frame within one of the reference frame lists. The video decoder 30 may construct the reference frame lists: list 0 and list 1 using a default construction technique based on the reference frames stored in the DPB 92.

[0053] In some examples, when a video block is encoded and decoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block defined by the video encoder 20.

[0054] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for video blocks of the current video frame by parsing the motion vectors and other syntax elements, and then use the prediction information to generate prediction blocks for the decoded current video blocks. For example, the motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) for encoding video blocks of the video frame, an inter-predicted frame type (e.g., B or P), construction information of one or more of the reference frame lists in the reference frame list of the frame, a motion vector for each inter-predicted encoded video block of the frame, an inter-predicted state of each inter-predicted encoded video block of the frame, and other information for decoding video blocks in the current video frame.

[0055] Similarly, the intra BC unit 85 may use some of the received syntax elements (e.g., flags) to determine that the current video block is predicted using: the intra BC mode, construction information that the video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, a block vector for each intra BC predicted video block of the frame, an intra BC prediction state of each intra BC predicted video block of the frame, and other information for decoding video blocks in the current video frame.

[0056] The motion compensation unit 82 may also perform interpolation using an interpolation filter as used by the video encoder 20 during encoding of the video block to calculate the interpolation values of sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and use the interpolation filter to generate the prediction block.

[0057] The inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80, using the same quantization parameter for determining the quantization degree calculated by the video encoder 20 for each video block in the video frame. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.

[0058] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vectors and other syntax elements, the adder 90 reconstructs the decoded video block of the current video block by summing the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter (not shown) may be positioned between the adder 90 and the DPB 92 to further process the decoded video block. Then the decoded video blocks in a given frame are stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video blocks. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later presentation on a display device such as Figure 1 the display device 34 as shown.

[0059] In a typical video coding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three arrays of samples, denoted as SL, SCb, and SCr, respectively. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other instances, a frame may be monochrome and thus include only one two-dimensional array of luminance samples.

[0060] As Figure 4A shown, the video encoder 20 (or more specifically, the partitioning unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of coding tree units (CTUs). A video frame may include an integer number of CTUs sorted consecutively in raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in the sequence parameter set such that all CTUs in the video sequence have the same size, i.e., one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As Figure 4BAs shown, each CTU may include one coding tree block (CTB) of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements for encoding and decoding the samples of the coding tree blocks. The syntax elements describe the attributes of different types of units of the coded blocks of pixels and how the video sequence can be reconstructed at the video decoder 30. The syntax elements include inter-prediction or intra-prediction, intra-prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding and decoding the samples of the coding tree block. The coding tree block may be an N×N sample block.

[0061] To achieve better performance, the video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination of both, on the coding tree block of the CTU, and divide the CTU into smaller coding units (CUs). As Figure 4C depicted, first, the 64×64 CTU 400 is divided into four smaller CUs, and the block size of each CU is 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four 16×16 CUs according to the block size. The two 16×16 CUs 430 and 440 are each further divided into four 8×8 CUs according to the block size. Figure 4D depicts a quadtree data structure that illustrates the final result of the partitioning process of the CTU 400 as depicted in Figure 4C . Each leaf node of the quadtree corresponds to a CU with a corresponding size in the range of 32×32 to 8×8. Similar to the Figure 4B depicted CTU, each CU may include a coded block (CB) of luminance samples and two corresponding coded blocks of chrominance samples of a frame of the same size, and syntax elements for encoding and decoding the samples of the coded blocks. In a monochrome picture or a picture with three separate color planes, the CU may include a single coded block and a syntax structure for encoding and decoding the samples of the coded block. It should be noted that Figure 4C and Figure 4D the quadtree partitioning depicted in is for illustrative purposes only, and a CTU may be divided into multiple CUs to adapt to different local characteristics based on quadtree / ternary tree / binary tree partitioning. In a multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU may be further partitioned by a binary tree structure or a ternary tree structure. As Figure 4E shown, there are five partitioning types, namely quaternary partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.

[0062] In some embodiments, the video encoder 20 may further divide the coded block of the CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. The prediction unit (PU) of the CU may include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements for predicting the prediction block. In a monochrome picture or a picture with three separate color planes, the PU may include a single prediction block and a syntax structure for predicting the prediction block. The video encoder 20 may generate prediction luma, Cb, and Cr blocks for the luma, Cb, and Cr prediction blocks of each PU of the CU.

[0063] The video encoder 20 may use intra prediction or inter prediction to generate the prediction blocks of the PU. If the video encoder 20 uses intra prediction to generate the prediction blocks of the PU, the video encoder 20 may generate the prediction blocks of the PU based on the decoded samples of the frame associated with the PU. If the video encoder 20 uses inter prediction to generate the prediction blocks of the PU, the video encoder 20 may generate the prediction blocks of the PU based on the decoded samples of one or more frames other than the frame associated with the PU.

[0064] After the video encoder 20 generates the prediction luma, Cb, and Cr blocks of one or more PUs of the CU, the video encoder 20 may generate a luma residual block of the CU by subtracting the prediction luma block of the CU from its original luma coded block, such that each sample in the luma residual block of the CU indicates the difference between the luma sample in one of the prediction luma blocks of the CU and the corresponding sample in the original luma coded block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block of the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between the Cb sample in one of the prediction Cb blocks of the CU and the corresponding sample in the original Cb coded block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between the Cr sample in one of the prediction Cr blocks of the CU and the corresponding sample in the original Cr coded block of the CU.

[0065] In addition, as Figure 4CAs illustrated, video encoder 20 may use quadtree partitioning to decompose the luminance, Cb, and Cr residual blocks of a CU into one or more luminance, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. The transform unit (TU) of a CU may include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the transform block samples. Thus, each TU of a CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and syntax structures for transforming the samples of the transform block.

[0066] Video encoder 20 may apply one or more transforms to the luminance transform block of a TU to generate a luminance coefficient block of the TU. A coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalars. Video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block of the TU. Video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block of the TU.

[0067] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to the process of quantizing transform coefficients to possibly reduce the amount of data used to represent the transform coefficients to provide further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 may entropy code the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream including a bit sequence representing an encoded frame and associated data, which is stored in storage device 32 or transmitted to destination device 14.

[0068] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct a frame of video data at least partially based on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on a coefficient block associated with a TU of a current CU to reconstruct a residual block associated with the TU of the current CU. Video decoder 30 also reconstructs the coded block of the current CU by adding the samples of the prediction block of the PU of the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the coded block of each CU in a frame, video decoder 30 may reconstruct the frame.

[0069] As described above, video coding and decoding mainly use two modes, namely, intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction) to achieve video compression. Palette-based coding and decoding is another coding and decoding scheme adopted by many video coding and decoding standards. In palette-based coding and decoding, which may be particularly applicable to content coding and decoding for screen generation, a video codec (e.g., video encoder 20 or video decoder 30) forms a palette table representing the colors of the video data of a given block. The palette table includes the most dominant (e.g., frequently used) pixel values in the given block. Pixel values that are not frequently represented in the video data of the given block are not included in the palette table or are included in the palette table as an escape color.

[0070] Each entry in the palette table includes an index of the corresponding pixel value in the palette table. The palette indices of the samples in the block may be coded and decoded to indicate which entry in the palette table is to be used to predict or reconstruct which sample. The palette mode begins with the process of generating a palette prediction value for the first block of a picture, slice, tile, or other such grouping of video blocks. As will be explained below, the palette prediction values for subsequent video blocks are typically generated by updating the previously used palette prediction value. For illustrative purposes, it is assumed that the palette prediction value is defined at the picture level. In other words, a picture may include multiple coded blocks, each coded block having its own palette table, but there is one palette prediction value for the entire picture.

[0071] To reduce the bits required to signal palette entries in a video bitstream, a video decoder can utilize a palette prediction value to determine new palette entries in a palette table for reconstructing video blocks. For example, the palette prediction value can include palette entries from a previously used palette table, or can even be initialized with the most recently used palette table by including all entries of the most recently used palette table. In some embodiments, the palette prediction value can include fewer than all entries from the most recently used palette table and then combine some entries from other previously used palette tables. The palette prediction value can have the same size as the palette table used for coding / decoding different blocks, or can be larger or smaller than the palette table used for coding / decoding different blocks. In one example, the palette prediction value is implemented as a first-in-first-out (FIFO) table including 64 palette entries.

[0072] To generate a palette table for a video data block from the palette prediction value, the video decoder can receive a one-bit flag for each entry of the palette prediction value from the encoded video bitstream. The one-bit flag can have a first value (e.g., binary one) indicating that the associated entry of the palette prediction value will be included in the palette table or a second value (e.g., binary zero) indicating that the associated entry of the palette prediction value will not be included in the palette table. If the size of the palette prediction value is greater than the palette table for the video data block, the video decoder can stop receiving more flags once the maximum size of the palette table is reached.

[0073] In some embodiments, some entries in the palette table can be signaled directly in the encoded video bitstream instead of using the palette prediction value to determine. For such entries, the video decoder can receive three separate m-bit values from the encoded video bitstream, where the m-bit values indicate the pixel values of the luminance and two chrominance components associated with the entry, and m represents the bit depth of the video data. Compared with the multiple m-bit values required for the palette entries signaled directly, those palette entries obtained from the palette prediction value only require a one-bit flag. Thus, using the palette prediction value to signal some or all palette entries can significantly reduce the bits required to signal the entries of the new palette table, thereby improving the overall coding / decoding efficiency of the palette mode coding / decoding.

[0074] In many instances, the palette prediction value of a block is determined based on the palette table used to encode and decode one or more previously encoded and decoded blocks. However, when encoding and decoding the first coding tree unit in a picture, slice, or tile, the palette table of the previously encoded and decoded blocks may not be available. Therefore, it is not possible to use the entries of the previously used palette table to generate the palette prediction value. In such cases, an initial value sequence of the palette prediction value can be signaled in the sequence parameter set (SPS) and / or picture parameter set (PPS). This initial value is the value used to generate the palette prediction value when the previously used palette table is not available. The SPS generally refers to the syntax structure of the syntax elements applied to a series of consecutive encoded video pictures called the coded video sequence (CVS), as determined by the content of the syntax elements found in the PPS, which is referenced by the syntax elements found in each slice segment header. The PPS generally refers to the syntax structure of the syntax elements applied to one or more individual pictures within the CVS, as determined by the syntax elements found in each slice segment header. Therefore, the SPS is generally considered a higher-level syntax structure than the PPS, which means that compared to the syntax elements included in the PPS, the syntax elements included in the SPS generally change less frequently and are applied to a larger portion of the video data.

[0075] Figures 5A to 5B is a block diagram illustrating an example of using a palette table to encode and decode video data according to some embodiments of the present disclosure.

[0076] For palette (PLT) mode signaling, the palette mode is encoded as the prediction mode for the coding unit, i.e., the prediction mode of the coding unit can be MODE_INTRA, MODE_INTER, MODE_IBC, and MODE_PLT. If the palette mode is used, the pixel values in the CU are represented by a small set of representative colors. This set is called the palette. For pixels with values close to the palette colors, the palette index is signaled. For pixels with values outside the palette, these pixels are represented by an escape symbol, and the quantized pixel values are directly signaled. The syntax and related semantics of the palette mode in the current VVC draft specification are illustrated in Tables 1 and 2 below.

[0077] To decode a palette mode coded block, the decoder needs to decode the palette colors and indices from the bitstream. The palette colors are defined by the palette table and encoded by the palette table coding syntax (e.g., palette_predictor_run, num_signaled_palette_entries, new_palette_entries). The escape flag palette_escape_val_present_flag is signaled for each CU to indicate whether there is an escape symbol in the current CU. If there is an escape symbol, an entry is added to the palette table and the last index is assigned to the escape mode. The palette indices for all pixels in the CU form the palette index map and are encoded by the palette index map coding syntax (e.g., num_palette_indices_minus1, palette_idx_idc, copy_above_indices_for_final_run_flag, palette_transpose_flag, copy_above_palette_indices_flag, palette_run_prefix, palette_run_suffix). An example of a palette mode coded CU is shown in Figure 5A where the palette size is 4. The first 3 samples in the CU are reconstructed using palette entries 2, 0, and 3 respectively. The "x" samples in the CU represent escape symbols. The CU-level flag palette_escaoe_val_present_flag indicates whether there are any escape symbols in the CU. If there is an escape symbol, the palette size is incremented and the last index is used to indicate the escape symbol. Thus, in Figure 5A index 4 is assigned to the escape symbol.

[0078] If the palette index (e.g., index 4 in Figure 5A ) corresponds to an escape symbol, additional overhead is signaled to indicate the corresponding color of the sample.

[0079] In some embodiments, on the encoder side, it is necessary to derive an appropriate palette to use with the CU. For lossy coding palette derivation, a modified k-means clustering algorithm is used. The first sample of the block is added to the palette. Then, for each subsequent sample in the block, the sum of absolute differences (SAD) between the sample and each current palette color is calculated. If the distortion of each component is less than the threshold of the palette entry corresponding to the minimum SAD, the sample is added to the cluster belonging to the palette entry. Otherwise, the sample is added as a new palette entry. When the number of samples mapped to a cluster exceeds the threshold, the centroid of the cluster is updated and becomes the palette entry for that cluster.

[0080] In the next step, the clusters are sorted in descending order of usage. Then, the palette entries corresponding to each entry are updated. Typically, the centroid of the cluster is used as the palette entry. However, when considering the encoding and decoding cost of the palette entry, a rate-distortion analysis is performed to analyze whether any entry from the palette prediction value is more suitable to be used as the updated palette entry instead of the centroid. This process continues until all clusters have been processed or the maximum palette size is reached. Finally, if a cluster has only one sample and the corresponding palette entry is not in the palette prediction value, the sample is converted to an escape symbol. Additionally, duplicate palette entries are removed and their clusters are merged.

[0081] After palette derivation, each sample in the block is assigned the index of the nearest (in terms of SAD) palette entry. Then, the sample is assigned to either the 'INDEX' or 'COPY_ABOVE' mode. For each sample where both the 'INDEX' and 'COPY_ABOVE' modes are possible, the run of each mode is determined. Then, the cost of encoding and decoding that mode is calculated. The mode with the lower cost is selected.

[0082] For the encoding and decoding of the palette table, a palette prediction value is maintained. The maximum size of the palette and the maximum size of the palette prediction value can both be signaled in the SPS (or other encoding levels, such as PPS, slice header, etc.). The palette prediction value is initialized at the start of each slice where the palette prediction value is reset to 0. For each entry in the palette prediction value, a reuse flag is signaled to indicate whether it is part of the current palette. As Figure 5B shown, the reuse flag palette_predictor_run is sent. After that, the number of new palette entries is signaled using the exponential Golomb code of order 0 via the syntax num_signaled_palette_entries. Finally, the component values of the new palette entries new_palette_entries[] are signaled. After encoding and decoding the current CU, the current palette is used to update the palette prediction value, and the entries from the previous palette prediction value that are not reused in the current palette are added to the end of the new palette prediction value until the maximum allowed size is reached.

[0083] To encode and decode the palette index map, the index is encoded using a horizontal or vertical traversal scan, as Figure 5C shown. The scan order is explicitly signaled in the bitstream using the palette_transpose_flag.

[0084] The palette indices are encoded and decoded using two main palette sample modes: 'INDEX' and 'COPY_ABOVE'. In the 'INDEX' mode, the palette index is explicitly signaled. In the 'COPY_ABOVE' mode, the palette index of the sample in the previous row is copied. For both the 'INDEX' and 'COPY_ABOVE' modes, the run value, which specifies the number of pixels encoded using the same mode, is signaled. The mode is signaled using a flag except when it is the top row when using horizontal scanning, or the first column when using vertical scanning, or when the previous mode was 'COPY_ABOVE'.

[0085] In some embodiments, the encoding and decoding order of the index map is as follows: First, the number of index values of the CU is signaled using the syntax num_palette_indices_minus1, and then the actual index values of the entire CU are signaled using the syntax palette_idx_idc. Both the number of indices and the index values are encoded and decoded in the bypass mode. This groups the bits of the bypass encoding and decoding related to the indices together. Then the palette modes (INDEX or COPY_ABOVE) and runs are signaled in an interleaved manner using the syntax copy_above_palette_indices_flag, palette_run_prefix, and palette_run_suffix. The copy_above_palette_indices_flag is a context encoding flag (only one bit), the codeword of palette_run_prefix is determined by the process described in Table 3 below, and the first 5 bits are context encoded. The palette_run_suffix is decoded as bypass bits. Finally, the component escape values corresponding to the escape samples of the entire CU are grouped together and decoded in the bypass mode. After signaling the index values, the additional syntax element copy_above_indices_for_final_run_flag is signaled. This syntax element combined with the number of indices eliminates the need to signal the run value corresponding to the last run in the block.

[0086] In the reference software of VVC (VTM), the dual tree is enabled for I slices, which separates the coding units of the luminance component and the chrominance components. Therefore, the palette is applied separately to the luminance (Y component) and chrominance (Cb and Cr components). If the dual tree is disabled, the palette will be applied jointly to the Y, Cb, Cr components.

[0087] Table 1 Palette Encoding and Decoding Syntax

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094] Table 2 Palette Coding / Decoding Semantics

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101] Table 3 Binary Codewords and CABAC Context Selections for the Syntax palette_run_prefix

[0102]

[0103]

[0104]

[0105] At the 15th JVET meeting, a row-based CG (document number JVET-O0120, accessible at http: / / phenix.int-evry.fr / jvet / was proposed to simplify the buffer usage and syntax in the palette mode of VTM6.0. As the coefficient group (CG) used in transform coefficient coding, the CU is divided into multiple row-based coefficient groups, each consisting of m samples. For each CG, the index run, palette index value, and quantized color of the escape mode are sequentially encoded / parsed. Therefore, the pixels in the row-based CG can be reconstructed after parsing the syntax elements (e.g., the index run, palette index value, and escape quantized color of the CG), which greatly reduces the buffer requirements in the palette mode of VTM6.0, where the syntax elements of the entire CU must be parsed (and stored) before reconstruction.

[0106] In this application, each CU in the palette mode is divided into multiple segments of m samples (m = 8 in this test) based on a traversal scan pattern, as Figure 5D shown.

[0107] The encoding order of palette run encoding and decoding in each segment is as follows: For each pixel, a context - encoded binary bit run_copy_flag = 0 is signaled, which indicates that the pixel has the same mode as the previous pixel, i.e., both the previous scanned pixel and the current pixel have run type COPY_ABOVE or both the previous scanned pixel and the current pixel have run type INDEX and the same index value. Otherwise, run_copy_flag = 1 is signaled.

[0108] If the current pixel and the previous pixel are in different modes, a context - encoded binary bit copy_above_palette_indices_flag indicating the run type of the pixel (i.e., INDEX or COPY_ABOVE) is signaled. In this case, if the sample is in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not have to parse the run type because the INDEX mode is used by default. If the previously parsed run type is COPY_ABOVE, the decoder also does not have to parse the run type.

[0109] After palette run encoding and decoding the pixels in a segment, the index value (in INDEX mode) and the quantized escape color are encoded as bypass binary bits and grouped separately from the encoding / parsing of the context - encoded binary bits to improve the throughput within each row - based CG. Since the index value is now encoded / parsed after run encoding and decoding, the encoder does not have to signal the number of index values num_palette_indices_minusl and the final run type copy_above_indices_for_final_run_flag. The syntax of the CG palette mode is shown in Table 4.

[0110] Table 4 Palette Encoding and Decoding Syntax

[0111]

[0112]

[0113]

[0114]

[0115]

[0116] Figure 6 FIG. 600 is a flowchart illustrating an exemplary process according to some embodiments of the present disclosure, by which a video decoder (e.g., video decoder 30) implements techniques for decoding video data.

[0117] For the palette mode in VVC, the palette mode can be applied to CUs equal to or smaller than 64×64 pixels. In some embodiments, a minimum palette mode block size is proposed to reduce complexity, such that the palette mode is disabled for coding units smaller than the minimum palette mode block size. For example, it is proposed to disable the palette mode for all blocks smaller than a certain threshold (e.g., 16 samples). Since there are different chroma formats (e.g., 4:4:4, 4:2:2, 4:2:0) and different coding tree types (e.g., SINGLE TREE, DUAL_TREE_LUMA, and DUAL_TREE_CHROMA), this threshold may vary. Note that "SINGLE_TREE" indicates that the luminance and chrominance components of an image are partitioned in the same way, such that the two components share the same palette table and palette prediction values in the palette mode. In contrast, "DUAL_TREE" indicates that the luminance and chrominance components of an image are partitioned separately, such that the two components have different palette tables and palette prediction values in the palette mode. For example, for the YUV 4:2:0 format with the "DUAL_TREE" type, i.e., considering different components separately, the palette mode for the chrominance component of CUs smaller than 16 samples should be disabled to reduce complexity. Table 5 below gives an example of the proposed syntax.

[0118] Table 5 - Palette mode enable flags for different coding tree types and chroma formats

[0119]

[0120] In Table 5, pred_mode_plt_flag specifies whether the palette mode is enabled (e.g., value 1) or disabled (e.g., value 0) for the coding unit. Parameters such as SubWidthC and SubHeightC are associated with the chroma format of the coding unit as follows:

[0121] Chrominance format SubWidthC SubHeightC Monochrome 1 1 4:4:4 1 1 4:2:2 2 1 4:2:0 2 2

[0122] In monochrome sampling, there is only one sample array that is nominally considered a luminance array. In 4:2:0 sampling, each of the two chrominance arrays has half the height and half the width of the luminance array. In 4:2:2 sampling, each of the two chrominance arrays has the same height as the luminance array and half the width. In 4:4:4 sampling, each of the two chrominance arrays has the same height and width as the luminance array.

[0123] In another embodiment, for the single-tree case, the palette mode is disabled for small-sized blocks depending on the luminance block size. In one example of the YUV 420 format, in the single-tree case, the palette mode is disabled for CUs smaller than 16 pixels depending on the luminance block size. In a specific example, the palette mode can be enabled for an 8×4 CU that includes 8×4 luminance samples and two 4×2 chrominance samples because palette enabling depends on the size of the luminance samples and does not consider the chrominance size.

[0124] During decoding of the bitstream, the video decoder 30 first receives a plurality of syntax elements (610) associated with the coding unit from the bitstream. The plurality of syntax elements indicate the size of the coding unit and the coding tree type of the coding unit. For example, the coding tree type of the coding unit can be one of SINGLE_TREE, DUAL_TREE_LUMA, or DUAL_TREE_CHROMA. The video decoder 30 then determines the minimum palette mode block size of the coding unit according to the coding tree type of the coding unit (620). For example, as shown in Table 5 above, when the coding tree type of the coding unit is SINGLE_TREE or DUAL_TREE_LUMA, the video decoder 30 sets the minimum palette mode block size to 16 samples. When the coding tree type of the coding unit is DUAL_TREE_CHROMA, the video decoder 30 first determines the chrominance format of the coding unit, and then sets the minimum palette mode block size according to the chrominance format as shown in the table above. For example, when the chrominance format is 4:4:4, the minimum palette mode block size is 16 samples; when the chrominance format is 4:2:2, the minimum palette mode block size is 32 samples; and when the chrominance format is 4:2:0, the minimum palette mode block size is 64 samples.

[0125] Based on determining that the size of the coding unit is greater than the minimum palette mode block size (630), the video decoder 30 receives a palette mode enable flag (640) associated with the coding unit from the bitstream, and then decodes the coding unit from the bitstream according to the palette mode enable flag (650). In some embodiments, when the palette mode enable flag indicates that the palette mode is enabled for the coding unit, the video decoder 30 generates a palette table for the current unit from the bitstream (670), and then decodes the coding unit from the bitstream using the generated palette table (680), as described above in connection with Figures 5A to 5D as described.

[0126] Figure 7 is a block diagram illustrating an exemplary context adaptive binary arithmetic coding (CABAC) engine according to some embodiments of the present disclosure.

[0127] Context adaptive binary arithmetic coding (CABAC) is a form of entropy coding used in many video coding standards, such as H.264 / MPEG-4 AVC, High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC). CABAC is based on arithmetic coding, with some innovations and changes to adapt it to the needs of video coding standards. For example, CABAC encodes and decodes binary symbols, thus maintaining low complexity and allowing probability modeling of the more common bits of any symbol. The probability model is adaptively selected based on local context, thus allowing better probability modeling, since coding modes are usually well correlated locally. Finally, CABAC uses multiplication-free range division by using quantized probability ranges and probability states.

[0128] CABAC has multiple probability modes for different contexts. CABAC first converts all non-binary symbols to binary. Then, for each binary bit (or bit), the codec selects the probability model to use and then uses information from nearby elements to optimize the probability estimate. Finally, arithmetic coding is applied to compress the data.

[0129] Context modeling provides an estimate of the conditional probability of the coded / decoded symbols. With a suitable context model, the given inter-symbol redundancy can be exploited by switching between different probability models based on the coded / decoded symbols near the current symbol to be coded. Coding / decoding the data symbols involves the following stages.

[0130] Binarization: CABAC uses binary arithmetic coding and decoding, which means that only binary decisions (1 or 0) are encoded. Before arithmetic coding and decoding, non-binary value symbols (such as transform coefficients or motion vectors) are "binarized" or converted into binary codes. This process is similar to the process of converting data symbols into variable-length codes, but the binary codes are further encoded (by the arithmetic coder) before transmission. The phase is repeated for each binary bit (or "bit") of the binarized symbol.

[0131] Context model selection: A "context model" is a probability model for one or more binary bits of a binarized symbol. This model can be selected from the available model selection based on the statistics of the most recently coded data symbols. The context model stores the probabilities of each binary bit being "1" or "0".

[0132] Arithmetic coding: The arithmetic coder encodes each binary bit according to the selected probability model. Note that each binary bit has only two sub-ranges (corresponding to "0" and "1").

[0133] Probability update: The selected context model is updated based on the actually coded and decoded values (for example, if the binary bit value is "1", the frequency count for "1" is increased).

[0134] By decomposing each non-binary syntax element value into a sequence of binary bits, the further processing of each binary bit value in CABAC depends on the associated coding / decoding mode decision, which can be selected as either the normal mode or the bypass mode. The latter is chosen for binary bits that are assumed to be uniformly distributed, and thus, the entire normal binary arithmetic coding (and decoding) process is simply bypassed. In the normal coding mode, each binary bit value is encoded using a normal binary arithmetic coding / decoding engine, where the associated probability model is either determined by a fixed selection based on the type of the syntax element and the bit position or bit index (binIdx) in the binarized representation of the syntax element, or adaptively selected from two or more probability models according to the relevant side information (e.g., the spatial neighbors, components, depth, or size of the CU / PU / TU, or the position within the TU). The selection of the probability model is referred to as context modeling. As an important design decision, the latter case is typically only applied to the most frequently observed binary bits, while other less frequently observed binary bits in general are processed using a combined (usually zero-order) probability model. In this way, CABAC is able to perform selective adaptive probability modeling at the sub-symbol level, and thus, provides an efficient tool for exploiting the redundancy between symbols while significantly reducing the overall modeling or learning cost. Note that for both the fixed and adaptive cases, in principle, the switching from one probability model to another can occur between any two consecutive normally coded / decoded binary bits. In summary, the design of the context model in CABAC reflects the aim of finding a good compromise between the conflicting goals of avoiding unnecessary modeling cost overhead and exploiting statistical dependencies to a large extent.

[0135] The parameters of the probability model in CABAC are adaptive, which means that the adaptation of the model probabilities to the statistical variations of the binary bit source is performed in a backward-adaptive and synchronized manner on a bit-by-bit basis in both the encoder and the decoder; this process is called probability estimation. For this purpose, each probability model in CABAC can take one of 126 different states, where the associated model probability value p ranges within the interval [0:01875; 0:98125]. These two parameters of each probability model are stored in the context memory as 7-bit entries: 6 bits for each of the 63 probability states, representing the model probability pLPS of the least probable symbol (LPS); 1 bit representing nMPS, i.e., the value of the most probable symbol (MPS).

[0136] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to tangible media such as data storage media or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.

[0137] The terms used in the description of the embodiments herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or groups thereof.

[0138] It should also be understood that although terms such as first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the embodiments, a first electrode may be referred to as a second electrode, and similarly, a second electrode may be referred to as a first electrode. The first electrode and the second electrode are both electrodes, but the first electrode and the second electrode are not the same electrode.

[0139] The description of the present application has been presented for purposes of illustration and description, and the description is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications, variations and alternative embodiments will be apparent to those of ordinary skill in the art from the foregoing description and the teachings presented in the associated drawings. The embodiments are chosen and described in order to best explain the principles of the invention, the practical application, and to enable others of ordinary skill in the art to understand the various embodiments of the invention and best utilize the basic principles and the various embodiments with various modifications suitable for the particular purposes contemplated. Accordingly, it is to be understood that the scope of the claims should not be limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the appended claims.

Claims

1. A method for encoding video data, comprising: Determine a plurality of variables associated with a coding unit, where the plurality of variables indicate a luminance coding block size and a coding tree type, where the coding unit is associated with a predefined partitioning method, and where the predefined partitioning method includes quaternary partitioning, horizontal ternary partitioning, vertical ternary partitioning, horizontal binary partitioning, or vertical binary partitioning; Determine a palette mode block size threshold for the coding unit according to the coding tree type; Determine that at least the luminance coding block size is greater than the palette mode block size threshold; Determine a palette mode enable flag associated with the coding unit; and Encode the coding unit according to the palette mode enable flag, where determining the palette mode block size threshold for the coding unit according to the coding tree type includes: When the coding tree type is SINGLE_TREE or DUAL_TREE_LUMA: Set the palette mode block size threshold to 16; and When the coding tree type is DUAL_TREE_CHROMA: Set the palette mode block size threshold according to the chroma format, where the coding tree type being SINGLE_TREE indicates that the luminance component and the chrominance component of an image are partitioned in the same way, the coding tree type being DUAL_TREE_LUMA indicates that the luminance component and the chrominance component of an image are partitioned separately and the luminance component of the image is currently being processed, and the coding tree type being DUAL_TREE_CHROMA indicates that the luminance component and the chrominance component of an image are partitioned separately and the chrominance component of the image is currently being processed.

2. The method according to claim 1, wherein, Encoding the coding unit according to the palette mode enable flag includes: When the palette mode enable flag indicates that the palette mode is enabled for the coding unit: Determine a palette table for the coding unit; and Encode the coding unit using the palette table.

3. The method according to claim 1, wherein, When the chroma format is 4:4:4, the palette mode block size threshold is 16.

4. The method according to claim 1, wherein, When the chroma format is 4:2:2, the palette mode block size threshold is 32.

5. The method according to claim 1, wherein, When the chroma format is 4:2:0, the palette mode block size threshold is 64.

6. An electronic device, comprising: One or more processing units; A memory coupled to the one or more processing units; And A plurality of programs stored in the memory, which when executed by the one or more processing units cause the electronic device to perform the method according to any one of claims 1 to 5.

7. A non-transitory computer-readable storage medium storing a plurality of programs for execution by an electronic device having one or more processing units, wherein, The plurality of programs, when executed by the electronic device, cause the one or more processing units to perform the method according to any one of claims 1 to 5.

8. A computer program product comprising instructions for execution by a computing device having one or more processors, wherein when the instructions are executed by the one or more processors, the computing device stores a bitstream generated by the method according to any one of claims 1 to 5.

9. A method for sending a bitstream, wherein, The bitstream is generated by the method according to any one of claims 1 to 5.