Method and apparatus for performing rate-distortion analysis for palette mode
By using the accuracy of internal bit depth in video encoding for rate distortion analysis of palette mode, updating the palette table and encoding the bitstream, the problem of efficient encoding and image quality is solved, and more efficient video data encoding and decoding is achieved.
Patent Information
- Application Number
- CN202080055525.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-09-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-09-24
AI Technical Summary
With the improvement of digital video quality, the amount of video data is increasing exponentially. How to encode/decode video data more effectively while maintaining the image quality of the decoded video data is an ongoing challenge.
Rate distortion analysis for palette mode is performed using the accuracy of internal bit depth, the palette table is updated by performing rate distortion analysis on the encoding blocks, and the updated palette table and the corresponding palette index map are encoded into a bitstream.
It improves the efficiency and quality of video data encoding, reduces distortion during the encoding process, and maintains high-quality decoded video data.
Smart Images

Figure CN114208172B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 905,346, filed on September 24, 2019, entitled “Video Codec Using Palette Mode,” the entire contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] Embodiments of the present invention generally relate to video data encoding, decoding and compression, and more particularly to methods and systems for performing rate-distortion analysis for palette modes using precision of internal bit depth. Background Art
[0004] Various electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit, receive, encode, decode and / or store digital video data by implementing video compression / decompression standards defined by MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC) standards. Video compression typically includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data. For block-based video coding, a video frame is divided into one or more slices, each slice having multiple video blocks, which may also be referred to as coding tree units (CTUs). Each CTU may contain a coding unit (CU) or may be recursively split into smaller CUs until a predefined minimum CU size is reached. Each CU (also called a leaf-CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be encoded in intra, inter, or IBC mode. Video blocks in intra-coded (I) slices of a video frame are encoded using spatial prediction relative to reference samples in neighboring blocks in the same video frame. Video blocks in inter-coded (P or B) slices of a video frame can use spatial prediction relative to reference samples in neighboring blocks in the same video frame or temporal prediction relative to reference samples in other previous and / or future reference videos.
[0005] A prediction block for the current video block to be encoded is generated based on spatial or temporal prediction of previously encoded reference blocks (e.g., neighboring blocks). The process of finding a reference block can be accomplished by a block matching algorithm. The residual data representing the pixel differences between the current block to be encoded and the prediction block is called a residual block or prediction error. Inter-coded blocks are encoded based on motion vectors and residual blocks pointing to reference blocks in a reference frame that form the prediction block. The process of determining motion vectors is generally referred to as motion estimation. Intra-coded blocks are encoded based on intra-frame prediction modes and residual blocks. For further compression, the residual block is transformed from the pixel domain to a transform domain, such as the frequency domain, producing residual transform coefficients, which can then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, can be scanned to produce a one-dimensional vector of transform coefficients, which can then be entropy encoded into the video bitstream to achieve even more compression.
[0006] The encoded video bitstream is then stored in a computer-readable storage medium (e.g., flash memory) to be accessed by another electronic device with digital video capabilities or directly transmitted to the electronic device in a wired or wireless manner. The electronic device then performs video decompression (which is the reverse process of the above-mentioned video compression) by, for example, parsing the encoded video bitstream to obtain syntax elements from the bitstream and reconstructing digital video data from the encoded video bitstream to its original format based at least in part on the syntax elements obtained from the bitstream, and presents the reconstructed digital video data on a display of the electronic device.
[0007] As the quality of digital video increases from HD to 4K×2K and even 8K×4K, the amount of video data to be encoded / decoded increases exponentially. How to encode / decode video data more efficiently while maintaining the image quality of the decoded video data is an ongoing challenge. Summary of the invention
[0008] The present application describes embodiments related to video data encoding and decoding, and more particularly, to systems and methods for performing rate-distortion analysis for palette modes using precision of internal bit depth.
[0009] According to the first aspect of the present application, a method for encoding video data includes: identifying a coding block for palette mode encoding; determining a palette table for the coding block; updating the palette table by performing rate-distortion analysis on the coding block, wherein the rate calculation and distortion calculation of the coding block are set to use an internal coding bit depth of a reference sample for the coding block; and encoding the updated palette table and the corresponding palette index map of the coding block into a bitstream.
[0010] According to a second aspect of the present invention, an electronic device comprises one or more processing units, a memory and a plurality of programs stored in the memory. When executed by the one or more processing units, the programs enable the electronic device to perform the method for decoding video data as described above.
[0011] According to a third aspect of the present invention, a non-transitory computer-readable storage medium stores a plurality of programs executed by an electronic device having one or more processing units. When executed by the processing unit, the programs cause the electronic device to perform the method for decoding video data as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings, which are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification, illustrate the described embodiments and together with the description serve to explain the basic principles. The same reference numerals refer to corresponding parts.
[0013] Figure 1 is a block diagram illustrating an exemplary video encoding and decoding system according to some embodiments of the present invention.
[0014] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present invention.
[0015] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present invention.
[0016] Figures 4A to 4E is a block diagram showing how a frame is recursively quadtree partitioned into multiple video blocks of different sizes according to some embodiments of the present application.
[0017] Figure 5 is a block diagram illustrating an example of determining and using a palette table for encoding and decoding video data according to some embodiments of the present invention.
[0018] Figure 6 is a flow chart illustrating an exemplary process for a video encoder implementing a technique for performing rate-distortion analysis for palette mode using precision of internal bit depth according to some embodiments of the present invention. DETAILED DESCRIPTION
[0019] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that various alternatives may be used without departing from the scope of the claims, and that the subject matter may be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein may be implemented on a variety of types of electronic devices having digital video capabilities.
[0020] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present invention. Figure 1 As shown, system 10 includes a source device 12 that generates and encodes video data that is subsequently decoded by a destination device 14. Source device 12 and destination device 14 may include any of a variety of electronic devices, including a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a video camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.
[0021] In some embodiments, the target device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the target device 14. In one example, the link 16 may include a communication medium so that the source device 12 can directly transmit the encoded video data to the target device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the target device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form a part of a packet-based network, such as a local area network, a wide area network, or a global network, such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may help facilitate communication from the source device 12 to the target device 14.
[0022] In some other embodiments, the encoded video data can be transferred from the output interface 22 to the storage device 32. Subsequently, the target device 14 can access the encoded video data in the storage device 32 through the input interface 28. The storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, the storage device 32 may correspond to a file server or another intermediate storage device that can save the encoded video data generated by the source device 12. The target device 14 can access the stored video data from the storage device 32 by streaming or downloading. The file server can be any type of computer capable of storing encoded video data and transmitting the encoded video data to the target device 14. Exemplary file servers include network servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The target device 14 can access the encoded video data through any standard data connection, including a wireless channel suitable for accessing the encoded video data stored on the file server (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of the two. The transmission of the encoded video data from the storage device 32 can be a streaming transmission, a download transmission, or a combination of the two.
[0023] like Figure 1 As shown, source device 12 includes video source 18, video encoder 20 and output interface 22. Video source 18 may include a source such as a video capture device, such as a camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of these sources. As an example, if video source 18 is a camera of a security monitoring system, source device 12 and target device 14 may form a camera phone or a video phone. However, the embodiments described in this application are generally applicable to video encoding and decoding, and may be applicable to wireless and / or wired applications.
[0024] Captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to target device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored on storage device 32 for subsequent access by target device 14 or other devices for decoding and / or playback. Output interface 22 may also include a modem and / or a transmitter.
[0025] Target device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 may include a receiver and / or a modem and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include a variety of syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. These syntax elements may be included within the encoded video data transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0026] In some implementations, the target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with the target device 14. The display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0027] The video encoder 20 and the video decoder 30 may operate according to proprietary or industry standards, such as VVC, HEVC, MPEG-4 Part 10, Advanced Video Coding (AVC), or extensions of such standards. It should be understood that the present application is not limited to a specific video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally envisioned that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is also generally envisioned that the video decoder 30 of the target device 14 may be configured to decode video data according to any of these current or future standards.
[0028] The video encoder 20 and the video decoder 30 can each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for the software in an appropriate non-transitory computer-readable medium and use one or more processors to execute these instructions in hardware to perform the video encoding / decoding operations disclosed in the present invention. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in a corresponding device.
[0029] Figure 2is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described herein. The video encoder 20 may perform intra-frame and inter-frame prediction encoding of video blocks within a video frame. Intra-frame prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence.
[0030] like Figure 2 As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 also includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copy (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A deblocking filter (not shown) may be located between the adder 62 and the DPB 64 to filter the block boundaries, thereby removing block artifacts from the reconstructed video. In addition to the deblocking filter, a loop filter (not shown) may also be used to filter the output of the adder 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided in one or more of the fixed or programmable hardware units shown.
[0031] The video data memory 40 may store video data encoded by the components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from the video source 18. The DPB 64 is a buffer that stores reference video data used by the video encoder 20 when encoding the video data (e.g., in an intra-frame or inter-frame prediction coding mode). The video data memory 40 and the DPB 64 may be formed by any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20, or off-chip relative to these components.
[0032] like Figure 2As shown, after receiving the video data, the segmentation unit 45 within the prediction processing unit 41 divides the video data into video blocks. The division may also include dividing the video frame into slices, tiles, or other larger coding units (CUs) according to a predefined division structure, such as a quadtree structure associated with the video data. The video frame may be divided into a plurality of video blocks (or video block groups referred to as tiles). The prediction processing unit 41 may select a prediction coding mode from a plurality of possible prediction coding modes for the current video block based on error results (such as coding rate and distortion level), such as one of one or more inter-frame prediction coding modes in a plurality of intra-frame prediction coding modes. The prediction processing unit 41 may provide the resulting intra-frame or inter-frame prediction coding block to the adder 50 to generate a residual block, and to the adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra-frame mode indicators, segmentation information, and other such syntax information to the entropy coding unit 56.
[0033] In order to select an appropriate intra-frame prediction coding mode for the current video block, intra-frame prediction processing unit 46 within prediction processing unit 41 may perform intra-frame prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-frame prediction coding of the current video block relative to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes, for example, to select an appropriate coding mode for each block of video data.
[0034] In some embodiments, motion estimation unit 42 determines the inter-prediction mode for a current video frame according to a predetermined pattern for a sequence of video frames by generating a motion vector that indicates the displacement of a prediction unit (PU) of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of a video block. A motion vector, for example, may indicate the displacement of a PU of a video block within a current video frame or picture relative to a prediction block within a reference frame (or other coded unit) relative to a current block (or other coded unit) that is encoded within the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra BC unit 48 may determine a vector, such as a block vector, for intra BC coding in a manner similar to the manner in which motion estimation unit 42 determines motion vectors for inter prediction, or may utilize motion estimation unit 42 to determine the block vector.
[0035] A prediction block is a block of a reference frame that is considered to closely match the PU of the video block to be encoded in terms of pixel difference, which can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some embodiments, the video encoder 20 may calculate values for sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 may interpolate quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, the motion estimation unit 42 may perform a motion search relative to full pixel positions and fractional pixel positions and output a motion vector with fractional pixel precision.
[0036] Motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-prediction coded frame by comparing the position of the PU to the position of a prediction block of a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.
[0037] The motion compensation performed by the motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by the motion estimation unit 42. After receiving the motion vector of the PU of the current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from the DPB 64, and forward the prediction block to the adder 50. The adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by the motion compensation unit 44 from the pixel values of the current video block being encoded. These pixel difference values forming the residual video block may include a luminance difference component or a chrominance difference component or both. The motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by the video decoder 30 when decoding the video block of the video frame. These syntax elements may include syntax elements such as defining a motion vector for identifying the predictive block, any flag indicating the prediction mode, or any other syntax information described herein. It should be noted that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, but are illustrated separately for conceptual purposes.
[0038] In some embodiments, the intra BC unit 48 may generate vectors and obtain prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but these prediction blocks are located in the same frame as the current block being encoded and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine an intra prediction mode for encoding the current block. In some examples, the intra BC unit 48 may, for example, encode the current block using various intra prediction modes during separate encoding passes, and test their performance through rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode from among the various tested intra prediction modes to use and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values using rate-distortion analysis for various tested intra prediction modes, and select an intra prediction mode with the best rate-distortion characteristics from among the tested modes as an appropriate intra prediction mode to use. The rate-distortion analysis typically determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra BC unit 48 may calculate ratios from the distortions and rates for the various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.
[0039] In other examples, intra BC unit 48 may use, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction in accordance with embodiments described herein. In either case, for intra block copying, a prediction block may be a block that is considered to closely match the block to be encoded, in terms of pixel differences, which may be determined by a sum of absolute differences (SAD), sum of squares (SSD), or other difference metric, and identification of the prediction block may include calculation of values for sub-integer pixel positions.
[0040] Regardless of whether the prediction block is from the same frame according to intra-frame prediction or from different frames according to inter-frame prediction, video encoder 20 can form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming pixel difference values. These pixel difference values forming the residual video block may include luma and chroma component differences.
[0041] As described above, the intra-prediction processing unit 46 may perform intra-prediction on the current video block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44 or the intra-block copy prediction performed by the intra BC unit 48. Specifically, the intra-prediction processing unit 46 may determine an intra-prediction mode for encoding the current block. To this end, the intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, for example, during separate encoding passes, and the intra-prediction processing unit 46 (or the mode selection unit in some examples) may select an appropriate intra-prediction mode to use from the tested intra-prediction modes. The intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra-prediction mode in the bitstream.
[0042] After prediction processing unit 41 determines a prediction block for the current video block by inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more transform units (TUs) and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as discrete cosine transform (DCT) or a conceptually similar transform.
[0043] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then scan the matrix containing the quantized transform coefficients. Alternatively, entropy encoding unit 56 may perform such a scan.
[0044] After quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques. This encoded bitstream may then be transmitted to video decoder 30, or archived in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 may also entropy encode the motion vectors and other syntax elements for the current video frame being encoded.
[0045] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain to generate a reference block for predicting other video blocks. As described above, motion compensation unit 44 may generate a motion compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for motion estimation.
[0046] Adder 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block stored in DPB 64. The reference block may then be used as a prediction block by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 to inter-predict another video block in a subsequent video frame.
[0047] Figure 3 is a block diagram showing an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 also includes a motion compensation unit 82, an intra-frame prediction processing unit 84, and an intra-frame block copy unit (intra-frame BC unit) 85. The video decoder 30 performs a decoding process, which is generally combined with Figure 2 The encoding process described with respect to video encoder 20 is reversed. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, and intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from entropy decoding unit 80.
[0048] In some examples, units of video decoder 30 may be assigned tasks to perform embodiments of the present invention. Furthermore, in some examples, embodiments of the present invention may be divided between one or more units of video decoder 30. For example, intra BC unit 85 may perform embodiments of the present invention alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction processing unit 84, and entropy decoding unit 80. In some examples, video decoder 30 may not include intra BC unit 85 and the functions of intra BC unit 85 may be performed by other components of prediction processing unit 81, such as motion compensation unit 82.
[0049] The video data memory 79 may store video data, such as an encoded video bitstream, to be decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32 via a wired or wireless network communication of the video data, from a local video source (such as a camera), or may be obtained by accessing a physical data storage medium (such as a flash drive or hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use when the video decoder 30 decodes the video data (such as in an intra-frame or inter-frame prediction coding mode). The video data memory 79 and the DPB 92 may be formed by any of a variety of memory devices, such as a dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. For illustrative purposes, in Figure 3 92 are shown as two different components of the video decoder 30. However, it is obvious to those skilled in the art that the video data memory 79 and the DPB 92 can be provided by the same memory device or a separate memory device. In some examples, the video data memory 79 can be on-chip with other components of the video decoder 30, or off-chip relative to these components.
[0050] During the decoding process, the video decoder 30 receives a coded video bitstream representing video blocks of a coded video frame and associated syntax elements. The video decoder 30 may receive these syntax elements at the video frame level and / or at the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantization coefficients, motion vectors or intra-frame prediction mode indicators and other syntax elements. The entropy decoding unit 80 then forwards these motion vectors and other syntax elements to the prediction processing unit 81.
[0051] When the video frame is encoded as an intra-frame prediction coded (I) frame or for intra-frame coded prediction blocks in other types of frames, the intra-frame prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-frame prediction mode sent by the signal and the reference data from the previously decoded block of the current frame.
[0052] When the video frame is encoded as an inter-frame prediction coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the video block of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each of these prediction blocks may be generated from a reference frame in one of these reference frame lists. The video decoder 30 may construct the reference frame lists, i.e., List 0 and List 1, based on the reference frames stored in the DPB 92 using a default construction technique.
[0053] In some examples, when encoding the video block according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates prediction blocks for the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. These prediction blocks may be within the same reconstruction region of the picture as the current video block defined by video encoder 20.
[0054] The motion compensation unit 82 and / or the intra BC unit 85 determine the prediction information for the video block of the current video frame by parsing these motion vectors and other syntax elements, and then use the prediction information to generate a prediction frame for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (such as intra-frame or inter-frame prediction) used to encode the video block of the video frame, the inter-frame prediction frame type (such as B or P), the construction information of one or more reference frame lists for the frame, the motion vector of each inter-frame prediction encoded video block of the frame, the inter-frame prediction state of each inter-frame prediction encoded video block of the frame, and other information for decoding these video blocks in the current video frame.
[0055] Similarly, the intra BC unit 85 may use some of the received syntax elements (e.g., an identifier) to determine construction information of which video blocks of the frame were previously predicted using the intra BC mode and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction states for each intra BC predicted video block of the frame, and other information for decoding a framework of these video blocks in the current video frame.
[0056] Motion compensation unit 82 may also use the interpolation filters to interpolate during encoding of the video blocks to calculate interpolated values for sub-integer pixels of reference blocks as did video encoder 20. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from the received syntax elements and use the interpolation filters to produce a prediction block.
[0057] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80 using the same quantization parameter calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to reconstruct the residual block in the pixel domain.
[0058] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on these vectors and other syntax elements, the adder 90 reconstructs the encoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. An in-loop filter (not shown) may be located between the adder 90 and the DPB 92 to further process the decoded video block. The decoded video blocks in a given frame are then stored in the DPB 92, which stores reference frames for subsequent motion compensation of later video blocks. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later display on a display device (such as Figure 1 is presented on a display device 34).
[0059] In a typical video encoding process, a video sequence usually includes a set of ordered frames or pictures. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other cases, a frame may be monochrome and therefore include only one two-dimensional array of luma samples.
[0060] like Figure 4A As shown, the video encoder 20 (or more specifically, the segmentation unit 45) generates an encoded representation of the frame by first dividing the frame into a set of coding tree units (CTUs). A video frame may include an integer number of CTUs that are consecutively ordered in a raster scan order from left to right and from top to bottom. Each CTU is the largest logical coding unit and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, i.e., one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. Figure 4BAs shown, each CTU may include one coding tree block (CTB) of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements for encoding samples of these coding tree blocks. These syntax elements describe the characteristics of different types of units of coded blocks of coded pixel blocks and how to reconstruct the video sequence at the video decoder 30, including inter-frame or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding samples of the coding tree block. A coding tree block may be an N×N block of samples.
[0061] To achieve better performance, the video encoder 20 may recursively perform tree partitioning on the coding tree blocks of the CTU, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination of the two, and partition the CTU into smaller coding units (CUs). Figure 4C As shown, the 64×64 CTU 400 is first divided into four smaller CUs, each of which has a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are both divided into four 16×16 CUs according to the block size. The two 16×16 CUs 430 and 440 are further divided into four 8×8 CUs according to the block size. Figure 4D A quadtree data structure is shown, and the figure shows Figure 4C The final result of the partitioning process of CTU 400 is shown in FIG. 4 , where each leaf node of the quadtree corresponds to a CU with a size ranging from 32×32 to 8×8. Figure 4B Each CU may include a coding block (CB) of luma samples and two corresponding coding blocks of chroma samples of the same size frame and syntax elements for encoding the samples of the coding blocks, as shown in the CTU shown in FIG. 1 . In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures for encoding the samples of the coding block. It should be noted that in Figure 4C and 4D The quadtree partitioning shown in FIG is for illustration purposes only, and a CTU can be split into CUs to accommodate different local characteristics based on quadtree / ternary / binary tree partitioning. In the multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU can be further partitioned by a binary tree and a ternary tree structure. Figure 4E As shown, there are five types of segmentation, namely, four-pronged segmentation, horizontal two-pronged segmentation, vertical two-pronged segmentation, horizontal three-pronged segmentation, and vertical three-pronged segmentation.
[0062] In some embodiments, the video encoder 20 may further partition the coding block of the CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples on which the same (inter-frame or intra-frame) prediction is applied. The prediction unit (PU) of a CU may include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements for prediction of these prediction blocks. In a monochrome picture or a picture with three separate color planes, a PU may include a single prediction block and a syntax structure for predicting the prediction block. The video encoder 20 may generate predicted luma, Cb, and Cr blocks for the luma, Cb, and Cr prediction blocks of each PU of the CU.
[0063] Video encoder 20 may use intra prediction or inter prediction to generate a prediction block for a PU. If video encoder 20 uses intra prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate a prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0064] After the video encoder 20 generates predicted luma, Cb, and Cr blocks for one or more PUs of a CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predicted luma block of the CU from its original luma coding block, so that each sample in the luma residual block of the CU indicates the difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, so that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0065] In addition, if Figure 4CAs shown, the video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A transform unit (TU) of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements for transforming these transform block samples. Therefore, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some examples, the luma transform block associated with the TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.
[0066] The video encoder 20 may apply one or more transforms to the luma transform block of the TU to generate a luma coefficient block for the TU. The coefficient block may be a two-dimensional array of multiple transform coefficients. The transform coefficient may be a scalar. The video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.
[0067] After generating a coefficient block (such as a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to the process of quantizing transform coefficients to possibly reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode the syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform context adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream including a sequence of bits that form a representation of the encoded frame and related data, which is stored in the storage device 32 or transmitted to the target device 14.
[0068] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct frames of the video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally mutual with the encoding process performed by the video encoder 20. For example, the video decoder 30 may inversely transform the coefficient blocks associated with the TUs of the current CU to reconstruct the residual blocks associated with the TUs of the current CU. The video decoder 30 may also reconstruct the coding blocks of the current CU by adding the samples of the prediction blocks of the PUs of the current CU to the samples of the transform blocks of the TUs of the current CU. After reconstructing the coding blocks for each CU of the frame, the video decoder 30 may reconstruct the frame.
[0069] As described above, video codecs mainly use two modes to achieve video compression, namely intra-frame prediction and inter-frame prediction. Palette-based codecs are another coding scheme adopted by many video coding standards. In palette-based codecs that may be particularly suitable for screen-generated content codecs, a video codec (e.g., video encoder 20 or video decoder 30) forms a palette table representing the color of the video data of a given block. The palette table includes the most important (e.g., frequently used) pixel values in the given block. Pixel values that are not frequently represented in the video data of the given block are either not included in the palette table or included in the palette table as escape colors.
[0070] Each entry in the palette table includes an index to the corresponding pixel value in the palette table. The palette index for the sample in the block can be encoded to indicate which entry from the palette table will be used to predict or reconstruct which sample. The palette mode begins with the process of generating a palette predictor for the first block of a picture, slice, tile, or other such video block grouping. As described below, the palette predictor for subsequent video blocks is typically generated by updating the previously used palette predictor. For purposes of illustration, it is assumed that the palette predictor is defined at the picture level. In other words, a picture may include multiple coding blocks, each with its own palette table, but there is only one palette predictor for the entire picture.
[0071] In order to reduce the bits required to signal palette entries in a video bitstream, a video decoder can utilize a palette predictor to determine new palette entries in a palette table for reconstructing a video block. For example, the palette predictor can include palette entries from a previously used palette table, or even be initialized with a most recently used palette table by including all entries of the most recently used palette table. In some embodiments, the palette predictor can include less than all entries of the most recently used palette table, and then merge some entries from other previously used palette tables. The palette predictor can have the same size as a palette table used to encode a different block, or can be larger or smaller than a palette table used to encode a different block. In one example, the palette predictor is implemented as a first-in, first-out (FIFO) table that includes 64 palette entries.
[0072] To generate a palette table for a block of video data from the palette predictor, a video decoder may receive a one-bit flag for each entry of the palette predictor from an encoded video bitstream. The one-bit flag may have a first value (e.g., binary 1) indicating that an associated entry of the palette predictor is to be included in the palette table or a second value (e.g., binary 0) indicating that an associated entry of the palette predictor is not included in the palette table. If the size of the palette predictor is larger than the palette table for the block of video data, the video decoder may stop receiving more flags once a maximum size of the palette table is reached.
[0073] In some embodiments, some entries in the palette table can be directly signaled in the encoded video bitstream rather than determined using the palette predictor. For these entries, the video decoder can receive three separate m-bit values from the encoded video bitstream that indicate pixel values for the luminance and two chrominance components associated with the entry, where m represents the bit depth of the video data. The palette entries derived from the palette predictor require only a one-bit flag, compared to the multiple m-bit values required for palette entries sent directly by signal. Therefore, using the palette predictor to signal some or all palette entries can significantly reduce the number of bits required to signal new palette table entries, thereby improving the overall encoding efficiency of palette mode encoding.
[0074] In many cases, the palette predictor for a block is determined based on the palette table used to encode one or more previously encoded blocks. However, when encoding the first coding tree unit in a picture, slice, or tile, the palette table of the previously encoded block may not be available. Therefore, the palette predictor cannot be generated using the entries of the previously used palette table. In this case, a sequence of initial values of the palette predictor can be signaled in a sequence parameter set (SPS) and / or a picture parameter set (PPS), which are the values used to generate the palette predictor when the previously used palette table is not available. SPS generally refers to the syntax structure of syntax elements applied to a series of consecutively encoded video pictures called a coded video sequence (CVS), which is determined by the content of the syntax elements found in the PPS, and the syntax elements found in each slice header reference the syntax elements found in the PPS. PPS generally refers to the syntax structure of syntax elements applied to one or more individual pictures within a CVS, and the one or more individual pictures are determined by the syntax elements found in each slice header. Therefore, the SPS is generally considered to be a higher-level syntax structure than the PPS, which means that the syntax elements included in the SPS generally change less and apply to a larger portion of the video data than the syntax elements included in the PPS.
[0075] Figure 5 is a block diagram illustrating an example of determining and using a palette table for encoding and decoding video data in a picture 500 according to some embodiments of the present invention. The picture 500 includes a first block 510 associated with a first palette table 520 and a second block 530 associated with a second palette table 540. Since the second block 530 is located to the right of the first block 510, the second palette table 540 can be related to the first palette table 520. A palette predictor 550 is associated with the picture 500 and is used to store zero or more palette entries from earlier generated palette tables including the first palette table 520, and add some of them to the second palette table 540. It should be noted that Figure 5 The various blocks shown in may correspond to CTUs, CUs, PUs, or TUs as described above, which are not limited to the block structure of any particular codec standard and may be compatible with future block-based codec standards.
[0076] Generally speaking, the palette table includes several pixel values for the block currently being encoded and decoded (such as Figure 5510 or 530 in the block). In some examples, a video codec (such as video encoder 20 or video decoder 30) can encode and decode a palette table for each color component of the block. For example, the video encoder 20 can encode a palette table for the luminance component of the block, another palette table for the chrominance Cb component of the block, and another palette table for the chrominance Cr component of the block. In this case, the first palette table 520 and the second palette table 540 can each become multiple palette tables. In other examples, the video encoder 20 can encode a single palette table for all color components of the block. In this case, the i-th entry in the palette table is a triple value of (Yi, Cbi, Cri), each of which corresponds to a component of a sample in the encoded block. Therefore, the representation of the first palette table 520 and the second palette table 540 is merely an example and is not intended to be limiting.
[0077] As described herein, a video codec (such as video encoder 20 or video decoder 30) may use the indexes I1, ... I N The pixels of the first block 510 are encoded and decoded instead of directly encoding and decoding the actual pixel values of the first block 510. For example, for each sample in the first block 510, the video encoder 20 can encode the index value of the sample, where the index value has an associated pixel value in the first palette table 520. The video encoder 20 encodes the first palette table 520 into an encoded video data bitstream and sends the bitstream to the video decoder 30 for palette-based decoding at the decoder side. In general, one or more palette tables can be transmitted for each block, or one or more palette tables can be shared between different blocks. The video decoder 30 can obtain the index value from the video bitstream generated by the video encoder 20, and reconstruct the pixel value using the corresponding pixel values of these index values in the first palette table 520. In other words, for each corresponding index value of the block, the video decoder 30 can determine the entry in the first palette table 520 based on the index value. Then, the video decoder 30 replaces the corresponding index value in the block with the pixel value specified by the entry in the first palette table 520.
[0078] In some embodiments, a video codec (such as video encoder 20 or video decoder 30) determines a second palette table 540 based at least in part on a palette predictor 550 associated with the picture 500. The palette predictor 550 may include some or all of the entries of the first palette table 520, and may also include entries from other previously generated palette tables. In some examples, the palette predictor 550 is implemented using a first-in-first-out table, where when entries of the first palette table 520 are added to the palette predictor 550, the oldest entries currently in the palette predictor 550 are deleted to keep the size of the palette predictor 550 at or below a maximum. In other examples, different techniques may be used to update and / or maintain the palette predictor 550.
[0079] In one example, the video encoder 20 may encode a pred_palette_flag for each block (e.g., the second block 530) to indicate whether the palette table 540 for the second block 530 is predicted from one or more palette tables associated with one or more other blocks (e.g., the neighboring block 510). For example, when the value of such a flag is binary "1", the video decoder 30 may determine that the palette table 540 for the second block 530 is predicted from one or more palette tables previously decoded, and therefore a new palette table for the second block 540 is not included in the video bitstream containing the pred_palette_flag. When such a flag is binary "0", the video decoder 30 may determine that the second palette table 540 for the second block 530 is included in the video bitstream as a new palette table. In some examples, the pred_palette_flag may be encoded separately for each different color component of the block (e.g., for a video block in a YCbCr space, three flags, one for Y, one for Cb, and one for Cr). In other examples, a single pred_palette_flag may be encoded for all color components of a block.
[0080] In the above example, pred_palette_flag is signaled for each block to indicate each entry of the palette table predicted for the current block. This means that the second palette table 540 is identical to the first palette table 520, and no additional information is signaled. In other examples, one or more syntax elements may be signaled on a per-entry basis. That is, a flag may be signaled for each entry of the previous palette table to indicate whether the entry exists in the current palette table. If a palette entry is not predicted, the palette entry may be explicitly signaled. In other examples, these two methods may be used in combination.
[0081] When predicting the second palette table 540 based on the first palette table 520, the video encoder 20 and / or the video decoder 30 may locate the block from which the predicted palette table is determined. The predicted palette table may be associated with one or more neighboring blocks (i.e., the second block 530) of the block currently being encoded. Figure 5 As shown, when determining the predicted palette table for the second block 530, the video encoder 20 and / or the video decoder 30 may locate the left neighboring block, namely the first block 510. In other examples, the video encoder 20 and / or the video decoder 30 may locate one or more blocks at other positions relative to the second block 530, such as the upper block in the picture 500. In another example, the palette table for the last block in the scan order using the palette mode may be used as the predicted palette table for the second block 530.
[0082] The video encoder 20 and / or the video decoder 30 may determine a block for palette prediction according to a predetermined order of block positions. For example, the video encoder 20 and / or the video decoder 30 may initially identify a left neighbor block for palette prediction, i.e., the first block 510. If the left neighbor block is not available for prediction (e.g., the left neighbor block is encoded using a mode other than a palette-based coding mode, such as an intra-frame prediction mode or an inter-frame prediction mode, or the left neighbor block is located at the leftmost edge of a picture or a fragment), the video encoder 20 and / or the video decoder 30 may identify an upper neighbor block in the picture 500. The video encoder 20 and / or the video decoder 30 may continue to search for available blocks according to a predetermined order of block positions until a block having a palette table that can be used for palette prediction is located. In some examples, the video encoder 20 and / or the video decoder 30 may determine a prediction palette based on reconstructed samples of multiple blocks and / or adjacent blocks by applying one or more formulas, functions, rules, etc. to generate a prediction palette table based on a palette table of one or a combination of multiple adjacent blocks (spatially or in a scan order). In one example, a predicted palette table including palette entries from one or more previously encoded neighboring blocks includes several (N) entries. In this case, the video encoder 20 first sends a binary vector V having the same size as the predicted palette table (i.e., size N) to the video decoder 30. Each entry in the binary vector indicates whether the corresponding entry in the predicted palette table will be reused or copied to the palette table for the current block. For example, V(i)=1 indicates that the i-th entry in the predicted palette table for the neighboring block will be reused or copied to the palette table for the current block, and the entry may have a different index in the current block.
[0083] In other examples, video encoder 20 and / or video decoder 30 may construct a candidate list including several potential candidates for palette prediction. In these examples, video encoder 20 may encode an index into the candidate list to indicate a candidate block in the list from which the current block for palette prediction is selected. Video decoder 30 may construct the candidate list in the same manner, decode the index, and use the decoded index to select the palette of the corresponding block for use with the current block. In another example, the palette table of the candidate block indicated in the list may be used as a predicted palette table for each prediction of the palette table of the current block.
[0084] In some embodiments, one or more syntax elements may indicate whether a palette table (e.g., the second palette table 540) is predicted entirely from a predicted palette (e.g., the first palette table 520, which may be composed of entries from one or more previously encoded blocks), or whether specific entries of the second palette table 540 are predicted. For example, an initial syntax element may indicate whether all entries in the second palette table 540 are predicted. If the initial syntax element indicates that not all entries are predicted (e.g., a flag having a binary value of zero), one or more additional syntax elements may indicate which entries of the second palette table 540 are predicted from the predicted palette table.
[0085] In some embodiments, the size of the palette table (eg, the number of pixel values contained in the palette table) may be fixed or may be signaled using one or more syntax elements in the coded bitstream.
[0086] In some embodiments, the video encoder 20 can encode the pixels of the block without precisely matching the pixel values in the palette table with the actual pixel values in the corresponding block of video data. For example, the video encoder 20 and the video decoder 30 can merge or combine (i.e., quantize) different entries in the palette table when the pixel values of these entries are within a predetermined range of each other. In other words, if there is already an existing pixel value within the error range of the new pixel value, the new pixel value is not added to the palette table, but the index of the existing pixel value is used to encode the sample in the block corresponding to the new pixel value. It should be noted that this lossy encoding and decoding process has no effect on the operation of the video decoder 30, and the video decoder 30 can decode the pixel values in the same manner regardless of whether the particular palette table is lossless or lossy.
[0087] In some embodiments, the video encoder 20 may select an entry in the palette table as a predicted pixel value for encoding a pixel value in the block. The video encoder 20 may then determine the difference between the actual pixel value and the selected entry as a residual and encode the residual. The video encoder 20 may generate a residual block that includes residual values for pixels in the block predicted by the entry in the palette table, and then the video encoder 20 applies a transform and quantization (as described above in conjunction with the above) to the residual block. Figure 2 As described above). Video encoder 20 can generate quantized residual transform coefficients in this manner. In another example, the residual block can be encoded and decoded losslessly (without transform and quantization) or without transform. Video decoder 30 can inverse transform and inverse quantize the transform coefficients to reproduce the residual block, and then reconstruct the pixel value using the predicted palette entry value for the pixel value and the residual value.
[0088] In some embodiments, the video encoder 20 may determine an error threshold for constructing a palette table, which is referred to as a differential value. For example, if the actual pixel value at a position in a block produces an absolute difference between the actual pixel value and an existing pixel value entry in the palette table that is less than or equal to the differential value, the video encoder 20 may send an index value to identify the corresponding index of the pixel value entry in the palette table for reconstructing the actual pixel value at the position. If the actual pixel value at a position in a block produces an absolute difference between the actual pixel value and an existing pixel value entry in the palette table that is greater than the differential value, the video encoder 20 may send the actual pixel value and add the actual pixel value as a new entry to the palette table. To construct the palette table, the video decoder 30 may use a differential value signaled by the encoder, rely on a fixed or known differential value, or infer or derive a differential value.
[0089] As described above, the video encoder 20 and / or the video decoder 30 may use a codec mode including an intra prediction mode, an inter prediction mode, a lossless codec palette mode, and a lossy codec palette mode when encoding and decoding video data. The video encoder 20 and the video decoder 30 may encode and decode one or more syntax elements indicating whether palette-based codec is enabled. For example, at each block, the video encoder 20 may encode a syntax element indicating whether a palette-based codec mode will be used for the block (e.g., CU or PU). For example, the syntax element may be signaled at a block level (e.g., CU level) in an encoded video bitstream, and then the video decoder 30 receives the syntax element when decoding the encoded video bitstream.
[0090] In some embodiments, these syntax elements may be transmitted at a higher level than the block level. For example, the video encoder 20 may signal these syntax elements at a slice level, a tile level, a PPS level, or an SPS level. In this case, a value equal to 1 indicates that the palette mode is used to encode all blocks at or below this level, so that no additional mode information (such as the palette mode or other modes) is signaled at the block level. A value equal to 0 indicates that the palette mode is not used to encode blocks at or below this level.
[0091] In some embodiments, the fact that a higher-level syntax element enables the palette mode does not mean that the palette mode must be used to encode every block at or below the higher level. On the contrary, another syntax element at the CU level or even the TU level may still be needed to indicate whether the CU-level or TU-level block is encoded or decoded using the palette mode, and if so, the corresponding palette table is to be constructed. In some embodiments, a video codec (such as video encoder 20 and video decoder 30) selects a threshold (such as 32) for the minimum block size based on the number of samples in the block, so that the palette mode is not allowed to be used for blocks with a block size below the threshold. In this case, there is no signaling of any syntax element for such a block. It should be noted that the threshold for the minimum block size can be explicitly signaled in the bitstream, or the threshold can be implicitly set to a default value that both the video encoder 20 and the video decoder 30 comply with.
[0092] The pixel value of one position of a block may be the same as (or within a differential value range of) the pixel value of another position of the block. For example, adjacent pixel positions of a block typically have the same pixel value, or may be mapped to the same index value in the palette table. Therefore, the video encoder 20 may encode one or more syntax elements that indicate a plurality of consecutive pixels or index values in a given scan order that have the same pixel value or index value. A string of similarly valued pixel or index values may be referred to herein as a "run". For example, if two consecutive pixels or indices in a given scan order have different values, the run is equal to 0. If two consecutive pixels or indices in a given scan order have the same value but the third pixel or index in the scan order has a different value, the run is equal to 1. For three consecutive indices or pixels with the same value, the run is 2, and so on. The video decoder 30 may obtain the syntax element indicating the run from the encoded bitstream and use the data to determine the number of consecutive positions with the same pixel or index value.
[0093] Figure 6is a flow chart illustrating an exemplary process 600 for a video encoder implementing a technique for performing rate-distortion analysis for palette mode using precision of internal bit depth according to some embodiments of the invention.
[0094] In some embodiments, when performing rate-distortion analysis in this palette mode, the video encoder 20 uses the precision of the internal bit depth. During the encoding process, the rate-distortion analysis provides a better choice for the trade-off between bit rate and distortion in the form of an analytical expression. The rate-distortion analysis may include the calculation of distortion and bit rate. The bit rate is defined as the number of bits per data sample to be stored or transmitted. The distortion is defined as a numerical value corresponding to the difference between the input and output signals, which is typically modeled based on human perception and aesthetics. In one example, a distortion calculation, such as the sum of absolute differences (SAD) or the sum of squared differences (SSD), is used to select the index of the nearest palette entry using the internal bit depth precision.
[0095] In another example, the video encoder performs a rate-distortion analysis for exporting a palette in a palette mode based on internal bit depth precision. The cluster centroids are typically used as palette entries when exporting the palette at the encoder side. The video encoder performs a rate-distortion analysis to analyze whether an entry from the palette predictor is more suitable than the centroids for use as updated palette entries when considering the cost of encoding and decoding the palette entries.
[0096] When encoding video data, video encoder 20 first identifies a coding block for palette mode encoding (eg, Figure 5 Next, the video encoder 20 determines a palette table (such as Figure 5 The video encoder updates the palette table 520 in (620) by performing a rate-distortion analysis on the coding block (630). The rate calculation and distortion calculation of the coding block are set to use the internal bit depth of the reference samples for the coding block.
[0097] After performing the rate-distortion analysis, the video encoder encodes the updated palette table and the corresponding palette index map of the coding block into a bitstream (640).
[0098] In some embodiments, the video encoder uses different weights for the luma component cost and the chroma component cost in the rate-distortion analysis for the palette mode. Such rate-distortion analysis can be used to derive palette mode related parameters (such as index coding mode, copy-over mode, or index run mode, etc.) for a given sample value in a CU. In one example, the chroma component distortion can be weighted in the total cost calculation for the rate-distortion analysis of the palette mode. For example, the distortion of the chroma component (such as the sum of absolute differences (SAD) and / or the sum of square differences (SSD)) is multiplied by a constant less than 1, such as 0.8 or 0.7. In another example, in the total cost calculation for the rate-distortion analysis of the palette mode, the cost of the chroma rate-distortion is multiplied by a constant less than 1, such as 0.8 or 0.7. In some embodiments, the coding block includes a luma component and a chroma component, and different weights are applied to the luma component and the chroma component for rate calculation and / or distortion calculation. In some embodiments, for distortion calculation, the weight applied to the luma component is higher than the weight applied to the chroma component. In some embodiments, the video encoder determines the palette table by grouping the samples in the coding block into a plurality of clusters based on their respective pixel values; determining a cluster centroid for each cluster in the plurality of clusters that exceeds a predefined threshold; and inserting the cluster centroid as a new entry in the palette table. In some embodiments, the new entry replaces an existing entry in the palette table corresponding to the sample of the cluster. In some embodiments, as described above, for clusters that exceed the predefined threshold, the video encoder replaces the centroid of the cluster with an entry in a palette predictor associated with the coding block based on a cost result of the rate-distortion analysis.
[0099] In some embodiments, for a cluster of only one sample, the video encoder converts the sample to an escape symbol when a corresponding entry in the palette table does not exist in the palette predictor associated with the coding block.
[0100] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on or transmitted through a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which corresponds to tangible media such as data storage media, or communication media including any media that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) non-temporary tangible computer-readable storage media or (2) communication media such as signals or carrier waves. Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures to implement the embodiments described in the present application. A computer program product may include computer-readable media.
[0101] In some embodiments, quantized residual differential pulse codec modulation (RDPCM) is provided for screen content encoding and decoding. When RDPCM is enabled, if the size of the CU is less than or equal to 32×32 luminance samples and if the CU is intra-coded, a flag is sent at the CU level. This flag indicates whether conventional intra-coding or RDPCM is used. If RDPCM is used, the RDPCM prediction direction flag is transmitted to indicate whether the prediction is horizontal or vertical. The coding block is then predicted using a conventional horizontal or vertical intra prediction process and unfiltered reference samples. The residual is quantized, and the difference between each quantized residual and its predictor, that is, the residual of the previous codec at a horizontal or vertical adjacent position (depending on the RDPCM prediction direction), is encoded and decoded.
[0102] For a block of size M (height) × N (width), let r i,j , 0≤i≤M-1, 0≤j≤N-1 is the prediction residual. Let Q(r i,j ), 0≤i≤M-1, 0≤j≤N-1 represents the residual r i,j RDPCM is applied to these quantized residuals, resulting in a product with elements The improved M×N array in is predicted from its adjacent quantized residuals. For vertical RDPCM prediction mode, for 0≤j≤(N-1), the following formula is used to derive
[0103]
[0104] For the horizontal RDPCM prediction mode, for 0≤i≤(M-1), the following formula is used to derive
[0105] On the decoder side, the above process is reversed to calculate Q(r i,j ), 0≤i≤M-1, 0≤j≤N-1, as shown below:
[0106] If using vertical RDPCM;
[0107] If using horizontal RDPCM.
[0108] The inverse quantized residual, Q -1 (Q(r i,j )), added to the intra block prediction value to produce the reconstructed sample value. The predicted quantized residual value is converted to the reconstructed sample value using the same residual coding process as in the transform skip mode residual coding. is sent to the decoder. As for the MPM modes for future intra-mode encoding, since there is no luma intra-mode encoded and decoded for RDPCM-coded CUs, the first MPM intra-mode is associated with the current CU and is used for intra-mode encoding and decoding of chroma blocks of the current CU and subsequent CUs. For deblocking, if two blocks on both sides of a block boundary are encoded and decoded using RDPCM, this particular block boundary will not be deblocked.
[0109] In existing quantization designs for RDPCM samples, given a quantization parameter (or QP), when the size of the transform block is a power of 4, the quantization scale used for RDPCM samples is the same as the scale used in conventional quantization of samples in other encoding tools. In some embodiments, the quantization process is unified by using the same shift and / or offset operations from the RDPCM quantization. For example, the RDPCM quantization process is used for quantization of escape samples in palette mode. Therefore, the quantization design for escape samples is the same as the quantization process for samples in RDPCM mode.
[0110] In some embodiments, the QP value is remapped for the quantization process of the escape sample. For the escape sample quantization, the range of QP can be defined differently from the quantization in other coding modes. For example, for the escape sample quantization, the minimum allowed QP can be defined as 4, because when the QP is equal to 4, the quantization step becomes 1. In another example, for the escape sample quantization, the maximum allowed QP can be defined as 61. In some embodiments, the remapping process can be derived by a specific equation. In one example, the QP used for the escape sample quantization process is derived as:
[0111] QP esca =MIN(((MAX(4,QP cu)–2) / 6)*6+4,61);
[0112] Where QPcu is the QP for the given CU. If the CU is coded in palette mode, the actual QP used when quantizing the escape samples of the CU is QP esca .
[0113] In some embodiments, fixed length binarization is used for escape samples. The length of the codeword of the fixed length binarization may depend on certain parameters of a given block. For example, the length of the fixed length binarization depends on the QP value and the internal bit depth. For example, the length len of the fixed length binarization can be derived as:
[0114] len=(bitDepth–(floor(QP–4) / 6)));
[0115] Wherein, the QP in the above equation is set to the actual QP value for the current block encoded in palette mode.
[0116] In some embodiments, the kth order Exp-Golomb binarization is used for escape sample encoding and decoding. Depending on certain parameters of a given block, the value of k may be different. For example, the value of k for the Exp-Golomb binarization used for escape sample encoding depends on the value of QP and the internal bit depth. For example, the value of k can be derived as:
[0117] k = (a – floor(QP / b));
[0118] Wherein, the QP in the above equation is set to the actual QP value for the current block encoded in the palette mode, and a and b are constants, for example, a=6, b=10.
[0119] In some embodiments, a truncated binary codeword is used for binarization of escape samples. The maximum value of the truncated binary codeword depends on certain parameters of a given block. For example, the maximum value of the truncated binary codeword depends on the QP value and the internal bit depth.
[0120] The terms used in the description of the embodiments herein are only used for the purpose of describing specific embodiments and are not intended to limit the scope of the claims. The singular forms "one" and "the / said" used in the description of the embodiments and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and covers any and all possible combinations of one or more associated listed items. It will be further understood that when used in this specification, the term "comprising" specifies the presence of stated features, elements and / or components, but does not exclude the presence or addition of other one or more features, elements, components and / or their groups.
[0121] It should also be understood that although the terms first, second, etc. may be used to describe various elements in this article, these elements should not be limited to these terms. These terms are only used to distinguish one element from another element. For example, without departing from the scope of the embodiment, the first electrode may be referred to as the second electrode, and similarly, the second electrode may be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but not the same electrode.
[0122] The description of the present application is presented for the purpose of illustration and description and is not intended to be exhaustive or to limit the invention in the disclosed form. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art with the benefit of the teachings presented in the foregoing description and the associated drawings. The embodiments are selected and described in order to best explain the principles of the invention, the practical application, and to enable other persons of ordinary skill in the art to understand the various implementations of the invention and to best utilize the basic principles and various implementations with various modifications, as applicable to the intended specific use. Therefore, it should be understood that the scope of the claims is not limited to the specific examples of the disclosed embodiments, and modifications and other embodiments are intended to be included within the scope of the appended claims.
Claims
1. A method for encoding video data, the method comprising: Identify encoding blocks for palette mode encoding; Determining a palette table for the coding block; updating the palette table by performing a rate-distortion analysis on the coding block, wherein the rate calculation and the distortion calculation of the coding block are arranged to update an entry in the palette table using a precision of an intra-coding bit depth for a reference sample of the coding block, the entry comprising an index of a corresponding pixel value in the palette table; as well as The updated palette table and the corresponding palette index map of the coding block are encoded into a bitstream.
2. The method according to claim 1, wherein: The coding block includes a luma component and a chroma component, and different weights are applied to the luma component and the chroma component for the rate calculation and / or the distortion calculation.
3. The method according to claim 2, wherein: For the distortion calculation, a higher weight is applied to the luma component than to the chroma components.
4. The method of claim 1, wherein determining the palette table comprises: grouping the samples in the coding block into a plurality of clusters based on respective pixel values of the samples; for each cluster in the plurality of clusters that exceeds a predefined threshold, determining a cluster centroid; as well as The cluster centroid is inserted into the palette table as a new entry in the palette table.
5. The method according to claim 4, wherein: The new entry replaces an existing entry in the palette table corresponding to a sample of a cluster.
6. The method of claim 4, wherein updating the palette table comprises: For clusters exceeding the predefined threshold, the centroid of the cluster is replaced with an entry in a palette predictor associated with the coding block according to a cost result of the rate-distortion analysis.
7. The method according to claim 4, wherein: Updating the palette table includes: For a cluster with only one sample, the sample is converted to an escape symbol when there is no corresponding entry in the palette table in the palette predictor associated with the coding block.
8. A computing device comprising: one or more processors; a memory coupled to the one or more processors; as well as A plurality of programs stored in the memory, when executed by the one or more processors, cause the computing device to perform the method of any one of claims 1-7.
9. A non-transitory computer-readable storage medium storing a plurality of programs executed by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, causes the computing device to perform the method of any one of claims 1-7.
Citation Information
Patent Citations
Encoding strategies for adaptive switching of color spaces, color sampling rates and / or bit depths
US20160261885A1
Two-dimensional palette coding for screen content coding
US20180307457A1
Fractional Quantization Parameter Offset In Video Compression
US20190020875A1
Palette coding for screen content coding
US20190158854A1