Video decoding method, electronic device, program, bitstream storage method, bitstream reception method

The palette mode in video encoding and decoding addresses the challenge of efficiently processing high-definition video data by using a palette table to represent pixel values, enhancing encoding efficiency and maintaining image quality.

JP7708928B2Active Publication Date: 2025-07-15BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024078545
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-08-15
Filing Date
2024-05-14
Publication Date
2025-07-15
Estimated Expiration
2040-08-14

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges in efficiently encoding and decoding high-definition to 4K×2K or 8K×4K video data while maintaining image quality, as the amount of data to be processed exponentially increases.

Method used

Implementing a palette mode for video encoding and decoding, which involves using a palette table to represent dominant pixel values and encoding pixel values as indexes, with lossless or lossy methods for escape samples based on quantization parameter thresholds.

Benefits of technology

Enhances encoding efficiency by reducing the bits required for signaling palette entries, thereby improving the overall encoding efficiency and maintaining image quality during decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708928000009
    Figure 0007708928000009
  • Figure 0007708928000010
    Figure 0007708928000010
  • Figure 0007708928000011
    Figure 0007708928000011
Patent Text Reader

Abstract

To provide a video decoding method.SOLUTION: A video decoding method includes: obtaining, from a bitstream, a palette-mode coded block, the coded block including both escape samples and non-escape samples; determining a quantization parameter from information included in a parameter set associated with the coded block; identifying escape samples in the coded block; when the quantization parameter is greater than a threshold value, using a predetermined formula to obtain values of reconstructed samples on the basis of values of the escape samples; and, when the quantization parameter is equal to the threshold value, setting values of reconstructed samples to values of the escape samples.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to decoding and compressing video data, and more particularly, to a method and system for video decoding using a palette mode.

Background Art

[0002] Digital video is supported by various electronic devices such as digital TVs, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. The electronic device transmits, receives, encodes, decodes, and / or stores digital video data by implementing video compression / decompression standards defined by MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding) standards. Video compression typically includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in the video data. In the case of block-based video coding, a video frame is divided into one or more slices, and each slice has a plurality of video blocks that may also be referred to as coding tree units (CTUs). Each CTU can contain one coding unit (CU) or be recursively divided into smaller CUs until a predetermined minimum coding unit (CU) size is reached. Each CU (also referred to as a leaf CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be encoded in either intra mode, inter mode, or IBC mode. Video blocks within an intra-coded (I) slice of a video frame are encoded using spatial prediction with respect to reference samples in adjacent blocks within the same video frame. Video blocks within an inter-coded (P or B) slice of a video frame can use spatial prediction with respect to reference samples in adjacent blocks within the same video frame or temporal prediction with respect to reference samples in other previous and / or future reference video frames.

[0003] Previous encoded reference blocks, e.g., spatial or temporal prediction based on neighboring blocks, result in a prediction block for the current video block to be encoded. The process of finding the reference blocks can be achieved by a block matching algorithm. Residual data representing the pixel difference between the current block to be encoded and the prediction block is called a residual block or prediction error. An inter-coded block is encoded according to a reference block within a reference frame forming the prediction block and a motion vector indicating the residual block. The process of determining the motion vector is typically called motion estimation. An intra-coded block is encoded according to an intra-prediction mode and a residual block. For further compression, the residual block is transformed from the pixel domain to a transform domain, e.g., a frequency domain, resulting in residual transform coefficients, which can then be quantized. The quantized transform coefficients are first arranged in a two-dimensional array, scanned to generate a one-dimensional vector of the transform coefficients, and then entropy encoded into the video bitstream to achieve further compression.

[0004] The encoded video bitstream is then stored in a computer-readable storage medium (e.g., flash memory) so that it can be accessed by another electronic device having digital video capabilities or transmitted directly to the electronic device, either wired or wirelessly. The electronic device then parses the encoded video bitstream to obtain syntax elements from the bitstream and reconstructs the digital video data from the encoded video bitstream back to its original format based at least in part on the syntax elements obtained from the bitstream, thereby performing video decompression (a process opposite to the above-described video compression), and rendering the reconstructed digital video data on the display of the electronic device.

[0005] In digital video quality ranging from high definition to 4K×2K or 8K×4K, the amount of video data to be encoded / decoded increases exponentially. A method for more efficiently encoding / decoding video data while maintaining the image quality of the decoded video data has always been a problem.

Summary of the Invention

Problems to be Solved by the Invention

[0006] This application relates to implementations related to the encoding and decoding of video data, and more particularly, to a system and method for encoding and decoding video using a palette mode.

Means for Solving the Problems

[0007] According to a first aspect of the present application, a method for decoding video data receives, from a bitstream, video data corresponding to a palette mode encoded block, determines a quantization parameter value from information included in a parameter set related to the palette mode encoded block, identifies quantization escape samples of the palette mode encoded block, and performs inverse quantization on the quantized escape samples according to a predetermined formula to obtain a reconstructed escape sample value according to a determination that the quantization parameter value is greater than a threshold value, and sets the reconstructed escape sample to the quantization escape sample value according to a determination that the quantization parameter value is less than or equal to the threshold value.

[0008] According to a second aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory. When the programs are executed by the one or more processing units, the electronic device is caused to execute a method for decoding video data as described above.

[0009] According to a third aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs to be executed by an electronic device having one or more processing units. When the programs are executed by the one or more processing units, the electronic device is caused to execute a method for decrypting video data as described above.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 4E

Figure 5

Figure 6

[0011] The accompanying drawings are included to provide a further understanding of the embodiments, are incorporated in and constitute a part of this specification, illustrate the described embodiments, and together with the description serve to explain the underlying principles. Like reference numerals refer to corresponding elements.

[0012] Reference will now be made in detail to specific examples, which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to facilitate understanding of the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that various alternative forms may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0013] FIG. 1 is a block diagram showing an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a destination device 14. Source device 12 and destination device 14 can each comprise any of a variety of electronic devices, including a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a video game console, a video streaming device, etc. In some implementations, source device 12 and destination device 14 comprise wireless communication capabilities.

[0014] In one implementation, the destination device 14 can receive the encoded video data that is decoded via the link 16. The link 16 can include any type of communication medium or device that can move the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may comprise a communication medium to enable the source device 12 to transmit the encoded video data directly in real time to the destination device 14. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum, or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, a switch, a base station, or any other device that may be useful to facilitate communication from the source device 12 to the destination device 14.

[0015] In some other implementations, the encoded video data may be sent from the output interface 22 to the storage device 32. Subsequently, the encoded video data within the storage device 32 can be accessed by the destination device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or local access data storage media, such as a hard drive, Blu-ray disk, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data. In a further example, the storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 can access the video data stored in the storage device 32 via streaming or downloading. The file server can be any type of computer that can store the encoded video data and send the encoded video data to the destination device 14. Exemplary file servers include web servers, FTP servers, NAS (network attached storage) devices, or local disk drives. The destination device 14 can access the encoded video data via any standard data connection, including a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 can be a streaming transmission, a download transmission, or a combination of both.

[0016] As shown in FIG. 1, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 can include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, when the video source 18 is a video camera of a security monitoring system, the source device 12 and the destination device 14 can form a camera phone or a video phone. However, the implementations described in the present application are generally applicable to video encoding and are applicable to wireless and / or wired applications.

[0017] Captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data may be directly transmitted to the destination device 14 via the output interface 22 of the source device 12. The encoded video data can also be stored in the storage device 32 so that it can be accessed later by the destination device 14 or other devices for decoding and / or playback. The output interface 22 can further include a modem and / or a transmitter.

[0018] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 includes a receiver and / or a modem and can receive encoded video data via link 16. The encoded video data communicated via link 16 or provided on storage device 32 can include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements may be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.

[0019] In some implementations, display device 34, which can be an integrated display device of destination device 14, and an external display device configured to communicate with destination device 14 can be included. Display device 34 displays the decoded video data to the user and can include any of various display devices such as a liquid crystal display, a plasma display, an organic light emitting diode, or another type of display device.

[0020] Video encoder 20 and video decoder 30 can operate according to proprietary specifications or industry standards such as VVC, HEVC, MPEG-4, Part 10, AVC (Advanced Video Coding), or extensions of such standards. It should be understood that the present application is not limited to specific video encoding / decoding standards and is applicable to other video encoding / decoding standards. In general, it is contemplated that video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, in general, it is also contemplated that video decoder 30 of destination device 14 can be configured to decode video data according to any of these current or future standards.

[0021] The video encoder 20 and the video decoder 30 can each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device can store software instructions in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations of the present disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and any of them may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.

[0022] FIG. 2 is a block diagram showing an exemplary video encoder 20 according to some implementations described in the present application. The video encoder 20 can perform intra prediction encoding and inter prediction encoding of video blocks within a video frame. Intra prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence.

[0023] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a conversion processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra prediction processing unit 46, and an intra block copy unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58, an inverse conversion processing unit 60, and an adder 62 for video block reconstruction. A deblocking filter (not shown) can be arranged between the adder 62 and the DPB 64 to filter the block boundary and remove block noise artifacts from the reconstructed video. An in-loop filter (not shown) may be used to filter the output of the adder 62 in addition to the deblocking filter. The video encoder 20 can be in the form of a fixed or programmable hardware unit, or can be divided among one or more of the fixed or programmable hardware units shown.

[0024] The video data memory 40 can store video data encoded by the components of the video encoder 20. The video data in the video data memory 40 can be obtained, for example, from the video source 18. The DPB 64 is a buffer that stores reference video data for use in encoding video data by the video encoder 20 (e.g., in an intra prediction encoding mode or an inter prediction encoding mode). The video data memory 40 and the DPB 64 can be formed by any of various memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20, or off-chip with respect to these components.

[0025] As shown in FIG. 2, after receiving the video data, the splitting unit 45 in the prediction processing unit 41 splits the video data into video blocks. This splitting can also include splitting the video frame into slices, tiles, or other larger coding units (CUs) according to a predefined splitting structure such as a quadtree structure associated with the video data. The video frame can be split into a plurality of video blocks (or a set of video blocks called tiles). The prediction processing unit 41 can select one of a plurality of possible prediction coding modes, such as one of a plurality of intra prediction coding modes or one of a plurality of inter prediction coding modes, for the current video block based on error results (e.g., coding rate and distortion level). The prediction processing unit 41 can provide the resulting intra or inter prediction coding block to the adder 50 to generate a residual block and then provide it to the adder 62 that reconstructs the coding block for use as part of the reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra mode indicators, splitting information, and other such syntax information to the entropy coding unit 56.

[0026] To select an appropriate intra prediction coding mode for the current video block, the intra prediction processing unit 46 in the prediction processing unit 41 can perform intra prediction coding of the current video block on one or more adjacent blocks within the same frame as the current block to be coded in order to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 in the prediction processing unit 41 perform inter prediction coding of the current video block with respect to one or more prediction blocks within one or more reference frames in order to provide temporal prediction. The video encoder 20 can execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of the video data.

[0027] In some implementations, the motion estimation unit 42 determines the inter - prediction mode of the current video frame by generating a motion vector indicating the displacement of a prediction unit (PU) of a video block in the current video frame relative to a prediction block in a reference video frame according to a predetermined pattern within a sequence of video frames. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector, and the motion vector estimates the motion of a video block. The motion vector can indicate, for example, the displacement of the PU of a video block in the current video frame or picture relative to a prediction block in a reference frame (or other coded unit) with respect to the current block being coded within the current frame (or other coded unit). The predetermined pattern can specify video frames within the sequence as P - frames or B - frames. The intra - BC unit 48 can determine a vector for intra - BC coding, such as a block vector, in a manner similar to the determination of the motion vector by the motion estimation unit 42 for inter - prediction, or the motion estimation unit 42 can be utilized to determine the block vector.

[0028] The prediction block is a block in the reference frame that is considered to closely match the PU of the video block to be coded with respect to pixel differences, which can be determined by SAD (sum of absolute difference), SSD (sum of square difference), or other difference metrics. In some implementations, the video encoder 20 can calculate the values of sub - integer pixel positions of the reference frames stored in the DPB 64. For example, the video encoder 20 can interpolate the values of 1 / 4 - pixel positions, 1 / 8 - pixel positions, or other fractional - pixel positions of the reference frames. Thus, the motion estimation unit 42 can perform motion search for both full - pixel positions and fractional - pixel positions and output a motion vector with fractional - pixel accuracy.

[0029] The motion estimation unit 42 calculates the motion vector of the PU of the video block in the inter-predicted coded frame by comparing the position of the PU with the position of the predicted block of the reference frame selected from the first reference frame list (List0) or the second reference frame list (List1) that each identifies one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44 and then to the entropy coding unit 56.

[0030] The motion compensation executed by the motion compensation unit 44 can include fetching or generating a predicted block based on the motion vector determined by the motion estimation unit 42. Upon receiving the motion vector for the PU of the current video block, the motion compensation unit 44 can search for the predicted block pointed to by the motion vector in one of the reference frame lists, retrieve the predicted block from the DPB 64, and transmit the predicted block to the adder 50. The adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the predicted block provided by the motion compensation unit 44 from the pixel values of the currently encoded video block. The pixel difference values forming the residual video block can include luminance or chrominance components or both. The motion compensation unit 44 can also generate syntax elements related to the video block of the video frame for use by the video decoder 30 when decoding the video block of the video frame. The syntax elements can include, for example, syntax elements defining the motion vector used to identify the predicted block, any flags indicating the prediction mode, or any other syntax information described herein. It should be noted that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes.

[0031] In some implementations, the intra BC unit 48 can generate vectors in a similar manner as described above in relation to the motion estimation unit 42 and the motion compensation unit 44, and can fetch prediction blocks. However, the prediction blocks are in the same frame as the currently encoded block, and the vectors are called block vectors as opposed to motion vectors. In particular, the intra BC unit 48 can determine the intra prediction mode used to encode the current block. In some examples, the intra BC unit 48 can encode the current block using various intra prediction modes, for example, between separate encoding paths, and test their performance through rate-distortion analysis. Next, the intra BC unit 48 can use the appropriate intra prediction mode among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 can calculate rate-distortion values using rate-distortion analysis for the various tested intra prediction modes, and select, as the appropriate intra prediction mode to use, the intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to generate the encoded block, as well as the bit rate (i.e., the number of bits) used to generate the encoded block. The intra BC unit 48 can calculate a ratio from the distortion and rate for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0032] In other examples, the intra BC unit 48 can use the motion estimation unit 42 and the motion compensation unit 44, either wholly or in part, to perform such functions for intra BC prediction according to the implementation forms described herein. In any case, in the case of intra block copy, the prediction block can be a block that is considered to closely match the block to be encoded, with respect to the pixel difference, which can be determined by SAD (sum of absolute difference), SSD (sum of squared difference), or other difference metrics, and the identification of the prediction block can include the calculation of sub-integer pixel position values.

[0033] Regardless of whether the prediction block is from the same frame according to intra prediction or from a different frame according to inter prediction, the video encoder 20 can form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded to form pixel difference values. The pixel difference values for forming the residual video block can include both the luminance component difference and the chrominance difference.

[0034] As described above, the intra prediction processing unit 46 can perform an intra prediction of the current video block in place of the inter prediction executed by the motion estimation unit 42 and the motion compensation unit 44, or the intra block copy prediction executed by the intra BC unit 48. In particular, the intra prediction processing unit 46 can determine the intra prediction mode to be used for encoding the current block. To do so, the intra prediction processing unit 46 can, for example, encode the current block using various intra prediction modes during separate encoding paths, and the intra prediction processing unit 46 (or, in some examples, the mode selection unit) can select an appropriate intra prediction mode to use from the tested intra prediction modes. The intra prediction processing unit 46 can provide information indicating the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 can encode the information indicating the selected intra prediction mode into the bitstream.

[0035] After the prediction processing unit 41 determines the predicted block of the current video block via either inter prediction or intra prediction, the adder 50 forms a residual video block by subtracting the predicted block from the current video block. The residual video data in the residual block can be included in one or more transform units (TUs) and supplied to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0036] The conversion processing unit 52 can send the obtained conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. Also, the quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 can then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0037] Following quantization, the entropy coding unit 56 entropy-codes the quantized conversion coefficients into the video bit stream, for example, using context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bit stream may then be sent to the video decoder 30 or may be archived in the storage device 32 for transmission or retrieval to a future video decoder 30. The entropy coding unit 56 can also entropy-code the motion vectors and other syntax elements for the current video frame being encoded.

[0038] The inverse quantization unit 58 and the inverse conversion processing unit 60 apply inverse quantization and inverse conversion, respectively, to reconstruct the residual video block within the pixel region for generating a reference block for prediction of other video blocks. As described above, the motion compensation unit 44 can generate a motion compensation prediction block from one or more reference blocks of the frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-pixel values for use in motion estimation.

[0039] The adder 62 adds the residual block reconstructed into the motion compensation prediction block generated by the motion compensation unit 44 to generate a reference block for storage in the DPB 64. The reference block can then be used by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 as a prediction block for predicting another video block within a subsequent video frame.

[0040] FIG. 3 is a block diagram showing an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further has a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 can perform a decoding process that is substantially the reverse of the encoding process described above with respect to the video encoder 20 in relation to FIG. 2. For example, the motion compensation unit 82 can generate prediction data based on the motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 can generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.

[0041] In some examples, tasks may be assigned to the units of the video decoder 30 to perform the implementation of the present application. Also, in some examples, the implementation of the present disclosure can be divided among one or more of the units of the video decoder 30. For example, the intra BC unit 85 can perform the implementation of the present application alone or in combination with other units of the video decoder 30 such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 such as the motion compensation unit 82.

[0042] Since the video data memory 79 is to be decoded by other components of the video decoder 30, it can store video data such as an encoded video bitstream. The video data stored in the video data memory 79 can be obtained, for example, from the storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or a hard disk). The video data memory 79 can include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use when the video decoder 30 decodes video data (e.g., in an intra or inter prediction coding mode). The video data memory 79 and the DPB 92 can be formed by any of various memory devices such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM (registered trademark)), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are shown as two separate components of the video decoder 30 in FIG. 3. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip with other components of the video decoder 30 or off-chip with respect to those components.

[0043] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. Video decoder 30 can receive syntax elements at the video frame level and / or at the video block level. The entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Next, the entropy decoding unit 80 transfers the motion vectors and other syntax elements to the prediction processing unit 81.

[0044] When the video frame is encoded as an intra prediction coded (I) frame or for an intra coded prediction block within another type of frame, the intra prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video blocks of the current video frame based on the signaled intra prediction mode and reference data from previously decoded blocks of the current frame.

[0045] When the video frame is encoded as an inter prediction coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the video blocks of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame within one of the reference frame lists. Video decoder 30 can configure the reference frame lists, list 0 and list 1, using a default configuration technique based on the reference frames stored in DPB 92.

[0046] In some examples, when a video block is encoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within the reconstructed region of the same picture as the current video block defined by the video encoder 20.

[0047] The motion compensation unit 82 and / or the intra BC unit 85 determines prediction information for a video block of the current video frame by parsing the motion vector and other syntax elements, and then uses the prediction information to generate a prediction block for the currently decoded video block. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used to encode the video block of the video frame, the inter prediction frame type (e.g., B or P), the configuration information for one or more of the reference frame lists for the frame, the motion vector for each inter prediction encoded video block of the frame, the inter prediction status for each inter prediction encoded video block of the frame, and other information for decoding the video block in the current video frame.

[0048] Similarly, the intra BC unit 85 can use some of the received syntax elements, such as flags, to determine that the current video block is predicted using the intra BC mode, which video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, the block vector for each intra BC predicted video block of the frame, the intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding the video block within the current video frame.

[0049] Also, the motion compensation unit 82 may perform interpolation using an interpolation filter as used by the video encoder 20 during the encoding of video blocks, and calculate interpolation values for sub-integer pixels of the reference blocks. In this case, the motion compensation unit 82 can determine the interpolation filter used by the video encoder 20 from the received syntax elements, and generate a prediction block using the interpolation filter.

[0050] The inverse quantization unit 86 inverse quantizes the bitstream decoded by the entropy decoding unit 80 and the quantized transform coefficients provided to the entropy for each video block in the video frame using the same quantization parameter calculated by the video encoder 20, to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to reconstruct the residual block in the pixel domain.

[0051] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. An in-loop filter (not shown) can be arranged between the adder 90 and the DPB 92 to further process the decoded video block. The decoded video blocks within a given frame are stored in the DPB 92 that stores the reference frames used for subsequent motion compensation of the next video blocks. The DPB 92, or a memory device separate from the DPB 92, can also store the decoded video for later presentation on a display device such as the display device 34 of FIG. 1.

[0052] In a typical video encoding process, a video sequence typically includes an ordered set of frames or pictures. Each frame can include three sample arrays denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other examples, a frame may be monochromatic and thus include only one two-dimensional array of luminance samples.

[0053] As shown in FIG. 4A, video encoder 20 (or more specifically, splitting unit 45) first generates an encoded representation of a frame by splitting the frame into a set of coding tree units (CTUs). A video frame can include an integral number of CTUs that are sequentially ordered in a raster scan order from left to right and from top to bottom. Each CTU is the largest logical encoding unit, and the width and height of the CTU are signaled to video encoder 20 by the sequence parameter set such that all CTUs in the video sequence have the same size as one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As shown in FIG. 4B, each CTU can comprise one coding tree block (CTB) of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements used to encode the samples of the coding tree block. The syntax elements include properties of different types of units of an encoded block of pixels, such as inter or intra prediction, intra prediction mode, motion vectors, and other parameters, and describe how the video sequence can be reconstructed at video decoder 30. In a monochrome picture or a picture having three separate color planes, a CTU can comprise a single coding tree block and syntax elements used to encode the samples of the coding tree block. The coding tree block may be an N×N block of samples.

[0054] To achieve better performance, the video encoder 20 can recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination of both, on the coding tree blocks of the CTU, and divide the CTU into smaller coding units (CUs). As shown in FIG. 4C, the 64×64 CTU 400 is first divided into four smaller CUs, each having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs of 16×16 by the block size. The two 16×16 CUs 430 and 440 are each further divided into four CUs of 8×8 by the block size. FIG. 4D shows a quadtree data structure indicating the final result of the partitioning process of the CTU 400 as shown in FIG. 4C, and each leaf node of the quadtree corresponds to one CU of each size in the range from 32×32 to 8×8. Similar to the CTU shown in FIG. 4B, each CU can include a coding block (CB) of luminance samples, two corresponding coding blocks of chrominance samples of a frame of the same size, and syntax elements used to code the samples of the coding blocks. For a monochrome picture or a picture having three separate color planes, the CU can include a single coding block and a syntax structure used to code the samples of the coding block. The quadtree partitioning shown in FIGS. 4C and 4D is for illustrative purposes only, and it should be noted that one CTU can be divided into CUs to adapt to various local characteristics based on quadtree / ternary tree / binary tree partitioning. In a multi-type tree structure, one CTU is divided by a quadtree structure, and each quadtree leaf CU can be further divided by a binary tree structure and a ternary tree structure. As shown in FIG. 4E, there are five partitioning types, namely, quadtree partitioning, horizontal binary tree partitioning, vertical binary tree partitioning, horizontal ternary tree partitioning, and vertical ternary tree partitioning.

[0055] In some implementations, video encoder 20 can further divide the coding block of a CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same prediction, inter or intra, is applied. The prediction unit (PU) of a CU can include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements used to predict the prediction block. For a monochrome picture or a picture having three separate color planes, the PU can include a single prediction block and a syntax structure used to predict the prediction block. Video encoder 20 can generate prediction luma, Cb, and Cr blocks for the luma, Cb, and Cr prediction blocks of each PU of the CU.

[0056] Video encoder 20 may use intra prediction or inter prediction to generate a prediction block for a PU. When video encoder 20 uses intra prediction to generate a prediction block of a PU, video encoder 20 can generate the prediction block of the PU based on the decoded samples of the frame related to the PU. When video encoder 20 uses inter prediction to generate a prediction block of a PU, video encoder 20 can generate the prediction block of the PU based on the decoded samples of one or more frames other than the frame related to the PU.

[0057] After the video encoder 20 generates prediction luminance, Cb, and Cr blocks for one or more PUs of a CU, the video encoder 20 can generate a luminance residual block for the CU by subtracting the prediction luminance block of the CU from the original luminance encoding block of the CU such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample within one of the prediction luminance blocks of the CU and the corresponding sample in the original luminance encoding block of the CU. Similarly, the video encoder 20 can generate a Cb residual block for the CU such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample within one of the prediction Cb blocks of the CU and the corresponding sample in the original Cb encoding block of the CU, and can generate a Cr residual block for the CU such that each sample in the Cr residual block of the CU indicates the difference between a Cr sample within one of the prediction Cr blocks of the CU and the corresponding sample in the original Cr encoding block of the CU.

[0058] Furthermore, as shown in FIG. 4C, the video encoder 20 may use quadtree partitioning to decompose the luminance, Cb, and Cr residual blocks of the CU into one or more luminance, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A transform unit (TU) of a CU may include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. For a monochrome picture or a picture having three separate color planes, a TU can include a single transform block and a syntax structure used to transform the samples of the transform block.

[0059] Video encoder 20 can apply one or more transforms to the luminance transform block of a TU to generate a luminance coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 can apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block of the TU. Video encoder 20 can apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0060] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 can quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized, perhaps reducing the amount of data used to represent the transform coefficients and providing further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 can entropy code the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream containing a bit sequence forming a representation of the encoded frame and associated data, which can either be stored in storage device 32 or transmitted to destination device 14.

[0061] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can analyze the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process executed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the encoded block of the current CU by adding the samples of the prediction block for the PU of the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the encoded block for each CU of the frame, the video decoder 30 can reconstruct the frame.

[0062] As described above, video encoding mainly uses two modes, namely, intra-frame prediction (or intra prediction) and inter-frame prediction (or inter prediction), to achieve video compression. Palette-based encoding is another encoding method adopted by many video encoding standards. In palette-based encoding, which may be particularly suitable for screen generation content encoding, a video coder (e.g., the video encoder 20 or the video decoder 30) forms a palette table of colors representing the video data of a given block. The palette table includes the most dominant (e.g., frequently used) pixel values within the given block. Pixel values that are not frequently represented in the video data of the specified block are either not included in the palette table or are included in the palette table as an escape color.

[0063] Each entry in the palette table includes the index of the corresponding pixel value in the palette table. The palette index for samples within a block may be encoded to indicate which entry from the palette table is used to predict or reconstruct which sample. This palette mode begins with the process of generating a palette predictor for the first block of a grouping of pictures, slices, tiles, or other video blocks. As described below, the palette predictors for subsequent video blocks are typically generated by updating previously used palette predictors. For purposes of explanation, it is assumed that the palette predictors are defined at the picture level. In other words, a picture can include multiple encoded blocks each having its own palette table, but there is one palette predictor for the entire picture.

[0064] To reduce the bits required for signaling palette entries in the video bitstream, the video decoder can utilize a palette predictor to determine new palette entries in the palette table used for reconstruction of video blocks. For example, the palette predictor can include palette entries from a previously used palette table, or even start with the last used palette table by including all entries of the last used palette table. In some implementations, the palette predictor includes fewer entries than all entries from the last used palette table and can then incorporate some entries from other previously used palette tables. The palette predictor may have the same size as the palette table used to encode different blocks, or may be larger or smaller than the palette table used to encode different blocks. In one example, the palette predictor is implemented as a first-in first-out (FIFO) table that includes 64 palette entries.

[0065] To generate a palette table for a block of video data from a palette predictor, the video decoder can receive a 1-bit flag for each entry of the palette predictor from the encoded video bitstream. The 1-bit flag can have a first value (e.g., binary 1) indicating that the associated entry of the palette predictor is included in the palette table, or a second value (e.g., binary 0) indicating that the associated entry of the palette predictor is not included in the palette table. If the size of the palette predictor is larger than the palette table used for a block of video data, the video decoder may stop receiving more flags when the maximum size of the palette table is reached.

[0066] In some implementations, instead of some entries of the palette table being determined using the palette predictor, they may be directly signaled within the encoded video bitstream. For such entries, the video decoder can receive three separate m-bit values (where m represents the bit depth of the video data) from the encoded video bitstream that indicate the pixel values of the luminance and chrominance components associated with the entry. Compared to the multiple m-bit values required for directly signaled palette entries, those palette entries derived from the palette predictor require only a 1-bit flag. Thus, signaling some or all palette entries using the palette predictor can significantly reduce the number of bits required to signal new entries of the palette table, thereby improving the overall encoding efficiency of palette mode encoding.

[0067] In many cases, the palette predictor of a block is determined based on a palette table used to encode one or more previously encoded blocks. However, when encoding the first coded tree unit in a picture, slice, or tile, the palette table of previously encoded blocks may not be available. Therefore, it is not possible to generate a palette predictor using entries of a previously used palette table. In such cases, the sequence of palette predictor initializers may be signaled in a sequence parameter set (SPS) and / or a picture parameter set (PPS), which are values used to generate a palette predictor when a previously used palette table is not available. The SPS generally refers to the syntax structure of syntax elements applied to a series of consecutive coded video pictures called a coded video sequence (CVS), which is determined by the content of the syntax elements found in the PPS that are referenced by syntax elements found in each slice segment header. The PPS generally refers to the syntax structure of syntax elements applied to one or more individual pictures in the CVS, as determined by syntax elements found in each slice segment header. Therefore, the SPS is generally considered to be at a higher level of syntax structure than the PPS, and the syntax elements included in the SPS generally mean that they are not changed as frequently as the syntax elements included in the PPS and are applied to a larger portion of the video data.

[0068] FIG. 5 is a block diagram illustrating an example of determining and using a palette table for encoding video data within picture 500 according to some implementations of the present disclosure. Picture 500 includes a first block 510 associated with a first palette table 520 and a second block 530 associated with a second palette table 540. Since the second block 530 is to the right of the first block 510, the second palette table 540 may be determined based on the first palette table 520. Palette predictor 550 is associated with picture 500 and is used to collect zero or more palette entries from the first palette table 520 and construct zero or more palette entries within the second palette table 540. The various blocks shown in FIG. 5 may correspond to CTUs, CUs, PUs, or TUs as described above, and it should be noted that the blocks are not limited to the block structure of any particular encoding standard and may be compatible with future block-based encoding standards.

[0069] Generally, a palette table contains a dominant and / or representative number of pixel values in the currently encoded block (e.g., block 510 or 530 in FIG. 5). In some examples, a video coder (e.g., video encoder 20 or video decoder 30) can encode a palette table separately for each chroma of the block. For example, video encoder 20 can encode a palette table for the luminance component of the block, a separate palette table for the chroma Cb component of the block, and yet another separate palette table for the chroma Cr component of the block. In this case, the first palette table 520 and the second palette table 540 may each be a plurality of palette tables. In other examples, video encoder 20 can encode a single palette table for all chromas of the block. In this case, the i-th entry of the palette table is three values (Yi, Cbi, Cri), where each value corresponds to one component of a pixel. Thus, the representation of the first palette table 520 and the second palette table 540 is merely an example and is not limiting.

[0070] As described herein, rather than directly encoding the actual pixel values of the first block 510, a video coder (such as video encoder 20 or video decoder 30) can use a palette-based encoding scheme to encode the pixels of the first block 510 using indexes I1, …, IN. For example, for each pixel within the first block 510, the video encoder 20 can encode the index value of the pixel, and the index value is associated with the pixel value within the first palette table 520. The video encoder 20 can encode the first palette table 520 and transmit it in the encoded video data bitstream for use by the video decoder 30 for palette-based decoding on the decoder side. Generally, one or more palette tables may be transmitted per block or shared among different blocks. The video decoder 30 can obtain the index values from the video bitstream generated by the video encoder 20 and can reconstruct the pixel values using the corresponding pixel values of the index values within the first palette table 520. In other words, for each index value for a block, the video decoder 30 may determine an entry within the first palette table 520. The video decoder 30 then replaces each index value within the block with the pixel value specified by the determined entry within the first palette table 520.

[0071] In some implementations, a video coder (e.g., video encoder 20 or video decoder 30) determines a second palette table 540 based at least in part on a palette predictor 550 associated with a picture 500. The palette predictor 550 includes some or all of the entries of a first palette table 520 and may also include entries from other palette tables. In some examples, the palette predictor 550 is implemented using a first-in first-out table, where adding an entry of the first palette table 520 to the palette predictor 550 causes the currently oldest entry within the palette predictor 550 to be evicted to keep the palette predictor 550 below its maximum size. In other examples, different techniques may be used to update and / or maintain the palette predictor 550.

[0072] In one example, the video encoder 20 encodes a pred_palette_flag for each block (e.g., second block 530) to indicate whether the palette table for the block is predicted from one or more other palette tables associated with one or more other blocks such as adjacent block 510. For example, if the value of such a flag is binary, the video decoder 30 may determine that the second palette table 540 for the second block 530 is predicted from one or more previously decoded palette tables and thus that a new palette table for the second block 540 is not included in the video bitstream containing the pred_palette_flag. If such a flag is binary zero, the video decoder 30 may determine that the second palette table 540 for the second block 530 is included in the video bitstream as a new palette table. In some examples, the pred_palette_flag may be encoded separately for different chroma of the block (e.g., for a video block in the YCbCr space, three flags, one for Y, one for Cb, and one for Cr). In other examples, a single pred_palette_flag may be encoded for all chroma of the block.

[0073] In the above example, pred_palette_flag is signaled block by block, indicating that all entries in the palette table of the current block are predicted. This means that the second palette table 540 is identical to the first palette table 520 and no additional information is signaled. In other examples, one or more syntax elements may be signaled entry by entry. That is, a flag can be signaled for each entry in the previous palette table to indicate whether the entry exists in the current palette table. If a palette entry is not predicted, the palette entry may be signaled explicitly. In other examples, these two methods can be combined.

[0074] When predicting the second palette table 540 according to the first palette table 520, the video encoder 20 and / or the video decoder 30 can identify the block for which the predicted palette table is determined. The predicted palette table may be associated with one or more adjacent blocks of the currently encoded block, i.e., the second block 530. As shown in FIG. 5, when determining the predicted palette table of the second block 530, the video encoder 20 and / or the video decoder 30 can identify the left adjacent block, i.e., the first block 510. In other examples, the video encoder 20 and / or the video decoder 30 can place one or more blocks at other positions relative to the second block 530, such as an upper block within the picture 500. In another example, the palette table of the last block in the scan order using the palette mode can be used as the predicted palette table of the second block 530.

[0075] Video encoder 20 and / or video decoder 30 can determine blocks for palette prediction according to a predetermined order of block positions. For example, video encoder 20 and / or video decoder 30 can first identify the left adjacent block, i.e., the first block 510, for palette prediction. If the left adjacent block is not available for prediction (e.g., the left adjacent block is encoded in a mode other than a palette-based coding mode such as an intra prediction mode or an inter prediction mode, or is located at the leftmost end of a picture or slice), video encoder 20 and / or video decoder 30 can identify the upper adjacent block within picture 500. Video encoder 20 and / or video decoder 30 can continue to search for available blocks according to the predetermined order of block positions until it finds a block having a palette table available for palette prediction. In some examples, video encoder 20 and / or video decoder 30 applies one or more formulas, functions, rules, etc. to generate a predicted palette table based on the palette tables of one or more adjacent blocks or based on a combination of a plurality of adjacent blocks (spatially or in scan order), so as to determine a predicted palette based on a plurality of blocks of adjacent blocks and / or reconstructed samples. In one example, the predicted palette table includes palette entries from one or more previously encoded adjacent blocks and includes the number of entries, N. In this case, video encoder 20 first transmits a binary vector V having the same size as the predicted palette table, i.e., size N, to video decoder 30. Each entry of the binary vector indicates whether the corresponding entry of the predicted palette table is reused or copied to the palette table of the current block. For example, V(i)=1 means that the i-th entry of the predicted palette table of the adjacent block is reused or copied to the palette table of the current block. There may be a different index for the current block.

[0076] In yet other examples, video encoder 20 and / or video decoder 30 can construct a candidate list that includes a number of potential candidates for palette prediction. In such examples, the video encoder 20 can encode an index into the candidate list to indicate a candidate block within the list from which the current block used for palette prediction is selected. The video decoder 30 can construct the candidate list in the same way, decode the index, and use the decoded index to select the palette of the corresponding block for use with the current block. In another example, a palette table of the indicated candidate blocks in the list can be used as a prediction palette table for each entry of the palette table of the current block.

[0077] In some implementations, one or more syntax elements can indicate whether a palette table, such as a second palette table 540, is fully predicted from a prediction palette (e.g., a first palette table 520 that may be composed of entries from one or more previously encoded blocks), or whether a particular entry of the second palette table 540 is predicted. For example, an initial syntax element can indicate whether all entries within the second palette table 540 are predicted. If the initial syntax element indicates that not all entries are predicted (e.g., a flag having a binary zero value), one or more additional syntax elements can indicate which entries of the second palette table 540 are predicted from the prediction palette table.

[0078] In some implementations, the size of the palette table, e.g., the number of pixel values included in the palette table, may be fixed or may be signaled using one or more syntax elements within the encoded bitstream.

[0079] In some implementations, the video encoder 20 can encode the pixels of a block without exactly matching the pixel values in the palette table to the actual pixel values in the corresponding block of the video data. For example, the video encoder 20 and the video decoder 30 can merge or combine (i.e., quantize) different entries in the palette table when the pixel values of the entries are within a predetermined range of each other. In other words, if an existing pixel value within the error margin of the new pixel value already exists, the new pixel value is not added to the palette table, and the samples within the block corresponding to the new pixel value are encoded with the index of the existing pixel value. Note that this process of lossy encoding does not affect the operation of the video decoder 30. The video decoder can decode the pixel values in the same way regardless of whether a particular palette table is lossless or lossy.

[0080] In some implementations, the video encoder 20 can select an entry in the palette table as a predicted pixel value for encoding the pixel values within a block. The next video encoder 20 can then determine the difference between the actual pixel value and the selected entry as a residual and encode the residual. The video encoder 20 can generate a residual block that includes the residual values for the pixels within the block predicted by the entry in the palette table, and then apply transformation and quantization (as described above in connection with FIG. 2) to the residual block. In this way, the video encoder 20 can generate quantized residual transform coefficients. In another example, the residual block may be encoded losslessly (without transformation and quantization) or without transformation. The video decoder 30 can inverse-transform and inverse-quantize the transform coefficients to reproduce the residual block, and then reconstruct the pixel values using the predicted palette input value and the residual values for the pixel values.

[0081] In some implementations, the video encoder 20 can determine an error threshold called a delta value to build a palette table. For example, when the actual pixel value at a position within a block generates an absolute difference between the actual pixel value below the delta value and an existing pixel value entry in the palette table, the video encoder 20 can send an index value to identify the corresponding index of the pixel value entry in the palette table for use when reconstructing the actual pixel value at that position. When the actual pixel value at a position within a block generates an absolute difference value between the actual pixel value in the palette table and an existing pixel value entry that is greater than the delta value, the video encoder 20 can send the actual pixel value and add the actual pixel value as a new entry to the palette table. To configure the palette table, the video decoder 30 can use the delta value signaled by the encoder, depend on a fixed or known delta value, or infer or derive the delta value.

[0082] As described above, the video encoder 20 and / or the video decoder 30 can use encoding modes including an intra prediction mode, an inter prediction mode, a lossless encoding palette mode, and a lossy encoding palette mode when encoding video data. The video encoder 20 and the video decoder 30 can encode one or more syntax elements indicating whether palette-based encoding is enabled. For example, in each block, the video encoder 20 can encode a syntax element indicating whether a palette-based encoding mode is used for the block (e.g., a CU or a PU). For example, this syntax element can be signaled in the encoded video bitstream at the block level (e.g., the CU level) and then received by the video decoder 30 when decoding the encoded video bitstream.

[0083] In some implementations, the above-mentioned syntax elements may be sent at a level higher than the block level. For example, the video encoder 20 can signal such syntax elements at the slice level, tile level, PPS level, or SPS level. In this case, a value equal to 1 indicates that all blocks below this level are encoded using the palette mode, and no additional mode information, such as the palette mode or other modes, is signaled at the block level. A value equal to zero indicates that blocks below this level are not encoded using the palette mode.

[0084] In some implementations, the fact that a higher-level syntax element enables the palette mode does not mean that each block at this higher or lower level must be encoded in the palette mode. Rather, another CU-level or TU-level syntax element is still needed to indicate whether a block at the CU level or TU level is encoded in the palette mode, and if so, the corresponding palette table should be constructed. In some embodiments, the video coder (e.g., the video encoder 20 and the video decoder 30) selects a threshold (e.g., 32) for the number of samples in a block of the minimum block size so that the palette mode is not permitted for blocks whose block size is less than the threshold. In this case, no signaling of syntax elements for such blocks is performed. Note that the threshold for the minimum block size can be explicitly signaled in the bitstream or implicitly set to a default value compiled by both the video encoder 20 and the video decoder 30.

[0085] The pixel value at one position of a block may be the same as (or within the range of a delta value) the pixel values at other positions of the block. For example, it is common for adjacent pixel positions of a block to have the same pixel value or to be mapped to the same index value within a palette table. Thus, video encoder 20 can encode one or more syntax elements indicating the number of consecutive pixels or index values in a given scan order having the same pixel value or index value. A string of pixels or index values with similar values may be referred to herein as a "run". For example, if two consecutive pixels or indexes in a given scan order have different values, the run is equal to zero. If two consecutive pixels or indexes in a given scan order have the same value, but a third pixel or index in the scan order has a different value, the run is equal to 1. For three consecutive indexes or pixels having the same value, the run is 2, and so on. Video decoder 30 can obtain the syntax element indicating the run from the encoded bitstream and use that data to determine the number of consecutive positions having the same pixel or index value.

[0086] FIG. 6 is a flowchart showing an exemplary process 600 for implementing a technique by which video decoder 30 decodes video data using a palette-based scheme according to some implementations of the present disclosure. Specifically, video decoder 30 uses information related to quantization parameters to determine whether escape samples (e.g., pixels that are not represented by palette entries in a palette table as described above in connection with FIG. 5 and that require additional overhead in the bitstream for signaling) within a palette mode encoded block are decoded in a lossless manner (i.e., the reconstructed pixel value of the escape sample is the same as the original pixel value) or in a lossy manner (i.e., the reconstructed pixel value of the escape sample is different from the original pixel value due to quantization).

[0087] To implement the palette-based method, the video decoder 30 receives, from the bitstream, video data corresponding to a palette mode encoded block (610). For example, the palette mode encoded block includes both escape samples and non-escape samples (e.g., pixel values represented by palette entries in a palette table).

[0088] Next, the video decoder 30 determines a quantization parameter value from information included in a parameter set associated with the palette mode encoded block (620). For example, the quantization parameter value may be associated with one of a plurality of coding levels and may be obtained, for example, from a sequence parameter set, a picture parameter set, a group header corresponding to a group of tiles, a tile header, etc.

[0089] In some embodiments, for the quantization design for escape samples, the scale of quantization for escape samples is the same as the normal quantization used for samples (e.g., samples under a transform skip case and / or a transform case) with other coding tools when a specific quantization parameter is given, but the actual operation for escape sample quantization is defined to be different from the normal quantization operation. For example, the quantization operation for escape samples includes different shift and / or offset operations from regular quantization.

[0090] In some embodiments, the quantization processes for escape samples and non-escape samples are the same. In one example, a standard quantization process is used for quantization of escape samples in palette mode. As a result, the quantization design for escape samples becomes the same as the quantization process for samples in transform skip mode and / or transform mode. The following equations relate to the corresponding quantization and inverse quantization processes applied in an encoder and a decoder when using quantization / inverse quantization in transform skip mode for encoding palette escape colors.

Equation

[0091] pResi and pResi’ are the original and reconstructed residual coefficients, pLevel is the quantized value, transformShift is the bit shift used to compensate for the dynamic range increase by 2D transform, which is equal to 15 - bitDepth - (log2(W) + log2(H)) / 2, where W and H are the width and height of the current transform unit, and bitDepth is the encoding bit depth. scale[] and descale[] are quantization and inverse quantization look-up tables with 14-bit and 6-bit precision, defined as follows. [Table 1] Table 1

[0092] If the size of the transform block is not a power of 4, another look-up table is defined as follows. [Table 2] Table 2

[0093] Next, the video decoder 30 identifies (630) the quantized escape samples within the palette mode encoded block. For example, the quantized escape samples can be associated with additional overhead in the bitstream to distinguish them from non-escape samples represented by palette entries in the palette table.

[0094] When video decoder 30 determines an escape sample within a palette mode coded block, video decoder 30 determines a particular method for decoding the quantized escape sample based on information from the quantization parameter value. In accordance with the determination that the quantization parameter value is greater than a threshold (e.g., the value of quantization parameter (QP) is greater than 4) (640), video decoder 30 performs inverse quantization on the quantized escape sample according to a predefined formula (e.g., based on formula (2) and Table 1 or Table 2 described above) to obtain a reconstructed escape sample value (640-1). Since the reconstructed escape sample value may be different from the original sample value for quantization, this decoding process is typically a lossy process. In accordance with the determination that the quantization parameter value is less than or equal to the threshold (e.g., the value of QP is equal to 4 or less than 4) (650), video decoder 30 sets the reconstructed escape sample to the quantized escape sample value (650-1). In this case, since the reconstructed escape sample value is the same as the original sample value, the decoding process is a lossless process. In some embodiments, video decoder 30 performs steps 640 and 650 while determining the quantization parameter value in parallel and from information included in a parameter set associated with the palette mode coded block (e.g., while performing step 620).

[0095] In some embodiments, the video decoder 30 performs inverse quantization on all quantization escape samples of all CUs within or below the level at which the quantization parameter value is signaled. For example, if the quantization parameter value is obtained from the picture parameter set, all quantized escape samples of the CUs within the picture are decoded according to a predetermined formula. In other words, if the value of QP is a specific value, for example, 4 or less, all CUs below the level at which the quantization parameter information is signaled may be encoded in lossless palette mode, and if the value of QP is a specific value, for example, greater than 4, it means that the lossless palette mode cannot be used to encode any CU below the level at which this QP information is signaled.

[0096] In some embodiments, the video decoder 30 first determines the delta quantization parameter from the information included in the parameter set associated with the palette mode encoded block, and then determines the quantization parameter value by adding the delta quantization parameter to the reference quantization parameter value as the quantization parameter value. An exemplary code for signaling the delta quantization parameter value of the palette mode encoded CU is shown below.

Table 3

[0097] In some embodiments, the delta quantization parameter is always signaled for the palette mode encoded block. An example of the code for always signaling the delta quantization parameter of the palette mode encoded block is shown below.

Table 4

[0098] In some embodiments, the delta quantization parameters for the luminance component and the chrominance component are always signaled separately for the palette mode coded blocks. Exemplary code for signaling the delta quantization parameters separately to the luminance component and the chrominance component of the palette mode coded blocks is shown below.

Table 5

Table 6

[0099] In some embodiments, the quantization coefficients corresponding to the escape samples of the palette mode coded blocks are binary coded using fixed-length binarization based on the quantization parameter value and the bit depth. The fixed-length binary length is determined according to one of the following formulas. Binary length = (bit depth - quantization parameter value / 6) (3) Binary length = (bit depth - (floor(quantization parameter - 4) / 6)) (4)

[0100] In some embodiments, the binarization method is adaptively switched at different coding levels, such as the sequence parameter set (SPS), the picture parameter set (PPS), the slice, and / or the group of coded blocks. In such a case, the encoder has the flexibility to dynamically select the binarization method to signal information in the bitstream.

[0101] In some embodiments, when the palette mode is lossless, the fixed-length binarization process is used for the escape samples. In one example, the escape samples are directly coded based on their binary format values, and each bit is coded as a CABAC bypass bin when the palette mode is lossless.

[0102] In some embodiments, the parameter set associated with the pallet mode encoded block includes the following.A first syntax element (e.g., palette_max_size) for specifying the maximum allowable palette size of a palette mode coding block, a second syntax element (e.g., palette_max_area) for specifying the maximum allowable palette area of a palette mode coding block, a third syntax element (e.g., palette_max_predictor_size) for specifying the maximum allowable palette predictor size of a palette mode coding block, a fourth syntax element (e.g., delta_palette_max_predictor_size) for specifying the difference between the maximum allowable palette predictor size and the maximum allowable palette size of a palette mode coding block, a fifth syntax element (e.g., palette_predictor_initializer_present_flag) for initializing a sequence palette predictor for a palette mode coding block, a sixth syntax element (e.g., num_palette_predictor_initializer_minus1) for specifying the number of entries of a palette predictor initializer minus 1 of a palette mode coding block, a seventh syntax element (e.g., palette_predictor_initializers[component][i]) for specifying the value of a component of the i-th palette entry used to initialize an array of predictor palette entries, an eighth syntax element (e.g., luma_bit_depth_entry_minus8_initializers) for specifying the bit depth value of the luminance component of an entry of a palette predictor initializer minus 8, a ninth syntax element (e.g., chroma_bit_depth_entry_minus8_initializers) for specifying the bit depth value of the chroma component of an entry of a palette predictor initializer minus 8, a tenth syntax element (e.g., luma_bit_depth_entry_minus8) for specifying the bit depth value of the luminance component of an entry of a palette minus 8, an eleventh syntax element (e.g., chroma_bit_depth_entry_minus8_initializers) for specifying the bit depth value of the chroma component of an entry of a palette minus 8.

[0103] In some embodiments, determining a quantization parameter value from information included in a parameter set associated with a palette mode coded block includes the following. In accordance with a determination that the quantization coefficients of the palette mode coded block do not correspond to escape samples of the palette mode coded block (e.g., the coefficients are coded with a palette), set the quantization parameter value equal to the quantization parameter value (QPcu) included in the parameter set corresponding to the current CU, and calculate the quantization parameter value according to the formula MIN(((MAX(4, QPcu) - 2) / 6)*6 + 4, 61) in accordance with a determination that the quantization coefficients of the palette mode coded block correspond to escape samples of the palette mode coded block.

[0104] In some embodiments, video decoder 30 determines a quantization parameter value from information included in a parameter set associated with a palette mode coded block in accordance with a determination that the quantization coefficients of the palette mode coded block correspond to escape samples of the palette mode coded block, and restricts the quantization parameter value according to the formula MAX(4, QPcu).

[0105] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally can correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for the implementations described herein. A computer program product may include a computer-readable medium.

[0106] The terms used in the description of the embodiments herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms “a,” “an,” etc. are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, the term “and / or” as used herein refers to any and all possible combinations of one or more of the associated listed items and is to be understood to be inclusive. The term “including,” etc. as used herein, while specifying the presence of the stated features, elements, and / or components, further is understood not to preclude the presence or addition of one or more other features, elements, components, and / or groups thereof.

[0107] Also, terms such as first, second, etc. may be used in this specification to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of implementation, the first electrode can be called the second electrode, and similarly, the second electrode can be called the first electrode. The first electrode and the second electrode are both electrodes, but they are not the same electrode.

[0108] The description of the present application is presented for purposes of illustration and explanation and is not intended to be exhaustive or to limit the invention to the disclosed form. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art who benefit from the teachings presented in the foregoing description and the related drawings. This embodiment is selected and described in order to best explain the principles of the invention, its practical application, and to enable others skilled in the art to understand the invention for various implementations and to best utilize the underlying principles and various implementations with various modifications suitable for the particular uses contemplated. Accordingly, it should be understood that the claims are not to be limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the appended claims.

Claims

1. Obtain an encoded block in palette mode, wherein the encoded block includes escape samples and non-escape samples, determine quantization parameters from the signaled information of the sequence parameter set related to the encoded block, identify the escape samples of the encoded block, when the quantization parameter is greater than a threshold, obtain the value of the reconstructed sample based on the value of the escape sample using a predetermined formula, when the quantization parameter is equal to the threshold, set the value of the reconstructed sample to the value of the escape sample, determining the quantization parameter includes determining the quantization parameter according to the function: MAX(4, QPcu), wherein the QPcu represents a quantization parameter value corresponding to the signaled information of the sequence parameter set related to the encoded block, A video decoding method.

2. According to the determination that the quantization parameter is greater than a threshold, obtain the reconstructed sample by a loss process, According to the determination that the quantization parameter is equal to the threshold, obtain the reconstructed sample by a lossless process, The video decoding method according to claim 1.

3. The predetermined formula is related to a look-up table defined as The video decoding method according to claim 1.

4. The escape sample 【Table 1】 is related to additional overhead of the bitstream, The video decoding method according to claim 1.

5. The threshold is equal to 4, The video decoding method according to claim 1.

6. An electronic device, comprising one or more processing units, a memory connected to one or more of the processing units, and a plurality of programs stored in the memory, which, when executed by one or more of the processing units, cause the electronic device to execute the video decoding method according to any one of claims 1 to 5. An electronic device.

7. A program executed by an electronic device including one or more processors, which, when executed by one or more of the processors, causes the electronic device to execute the video decoding method according to any one of claims 1 to 5. A program. ​ ​ ​ ​ ​ ​ ​

8. A bitstream storage method, wherein the bitstream is decoded by the video decoding method according to any one of Claims 1 to 5. A bitstream storage method.

9. A bitstream reception method, wherein the bitstream is decoded by the video decoding method according to any one of Claims 1 to 5. A bitstream reception method.

Citation Information

Patent Citations

  • Coding escape pixels for palette mode coding

    US20160227225A1

  • Escape color coding for palette coding mode

    US20160227231A1

  • Improved encoding process using a palette mode

    US20160316214A1

  • Coding escape pixels for palette coding

    WO2016123513A1

  • Robust encoding / decoding of escape-coded pixels in palette mode

    WO2016197314A1