Video encoding method, electronic device, program, bitstream storage method, bitstream transmission method

The use of a palette mode for video decoding, with lossless or lossy encoding of escape samples based on a quantization parameter threshold, addresses the challenge of efficiently encoding and decoding high-definition video data, enhancing encoding efficiency and maintaining picture quality.

JP7708929B2Active Publication Date: 2025-07-15BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024078546
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-08-15
Filing Date
2024-05-14
Publication Date
2025-07-15
Estimated Expiration
2040-08-14

AI Technical Summary

Technical Problem

Existing video encoding technologies face challenges in efficiently encoding and decoding high-definition to 4K×2K or 8K×4K video data while maintaining picture quality, as the amount of data to be processed exponentially increases.

Method used

A method for video decoding using a palette mode, involving the use of a palette table to represent dominant pixel values and encoding pixel indices, with lossless or lossy encoding of escape samples based on a quantization parameter threshold, to enhance encoding efficiency.

Benefits of technology

Improves encoding efficiency by reducing the bits required for signaling palette entries, maintaining picture quality through lossless or lossy decoding of escape samples, and optimizing data transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708929000009
    Figure 0007708929000009
  • Figure 0007708929000010
    Figure 0007708929000010
  • Figure 0007708929000011
    Figure 0007708929000011
Patent Text Reader

Abstract

To provide a video coding method.SOLUTION: A video coding method includes: dividing a video frame into a plurality of blocks; determining a quantization parameter of a current block of the plurality of blocks, the current block including both escape samples and non-escape samples; identifying escape samples in the current block; when the quantization parameter is greater than a threshold value, using a predetermined formula to obtain values of reconstructed samples on the basis of values of the escape samples; when the quantization parameter is equal to the threshold value, determining values of the reconstructed samples to values of the escape samples; and coding information associated with the quantization parameter into a bitstream.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to the encoding and compression of video data, and more particularly, to a method and system for video decoding using a palette mode.

Background Art

[0002] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices implement video compression / decompression standards defined by MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding) standards to transmit, receive, encode, decode, and / or store digital video data. Video compression usually includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data. In the case of block-based video coding, a video frame is divided into one or more slices, and each slice has a plurality of video blocks that can also be called coding tree units (CTUs). Each CTU can contain one CU or be recursively divided into smaller CUs until a predetermined minimum coding unit (CU) size is reached. Each CU (also called a leaf CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be encoded in either an intra mode, an inter mode, or an IBC mode. Video blocks within an intra-coded (I) slice of a video frame are encoded using spatial prediction with respect to reference samples in adjacent blocks within the same video frame. Video blocks within an inter-coded (P or B) slice of a video frame can use spatial prediction with respect to reference samples in adjacent blocks within the same video frame or temporal prediction with respect to reference samples in other previous and / or future reference video frames.

[0003] Prior encoded reference blocks, e.g., spatial or temporal prediction based on neighboring blocks, result in a prediction block for the current video block to be encoded. The process of finding the reference block can be achieved by a block matching algorithm. Residual data representing the pixel difference between the current block to be encoded and the prediction block is called the residual block or prediction error. An inter-encoded block is encoded according to a reference block within a reference frame forming the prediction block and a motion vector indicating the residual block. The process of determining the motion vector is typically called motion estimation. An intra-encoded block is encoded according to an intra prediction mode and a residual block. For further compression, the residual block is transformed from the pixel domain to a transform domain, e.g., a frequency domain, resulting in residual transform coefficients, which can then be quantized. The quantized transform coefficients are first arranged in a two-dimensional array, scanned to generate a one-dimensional vector of the transform coefficients, and then entropy encoded into a video bitstream to achieve further compression.

[0004] The encoded video bitstream is then stored in a computer-readable storage medium (e.g., flash memory) so that it can be accessed by another electronic device having a digital video function or transmitted directly to the electronic device, either wired or wirelessly. The electronic device then parses the encoded video bitstream to obtain syntax elements from the bitstream and reconstructs the digital video data from the encoded video bitstream into its original format based at least in part on the syntax elements obtained from the bitstream, thereby performing video decompression (a process opposite to the above-described video compression), and rendering the reconstructed digital video data on a display of the electronic device.

[0005] In digital video quality ranging from high definition to 4K×2K or 8K×4K, the amount of video data to be encoded / decoded increases exponentially. There has always been a problem of how to encode / decrypt video data more efficiently while maintaining the picture quality of the decoded video data.

Summary of the Invention

Problems to be Solved by the Invention

[0006] This application relates to implementations related to the encoding and decoding of video data, and more particularly, to a system and method for encoding and decoding video using a palette mode.

Means for Solving the Problems

[0007] According to a first aspect of the present application, a method for decoding video data receives video data corresponding to a palette mode encoded block from a bitstream, determines a quantization parameter value from information included in a parameter set related to the palette mode encoded block, identifies quantization escape samples of the palette mode encoded block, and performs inverse quantization on the quantized escape samples according to a predetermined formula to obtain a reconstructed escape sample value according to a determination that the quantization parameter value is greater than a threshold value, and sets the reconstructed escape sample to the quantization escape sample value according to a determination that the quantization parameter value is less than or equal to the threshold value.

[0008] According to a second aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory. When the programs are executed by the one or more processing units, the electronic device is caused to execute a method for decoding video data as described above.

[0009] According to a third aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs to be executed by an electronic device having one or more processing units. When the programs are executed by the one or more processing units, the electronic device is caused to execute a method for decoding video data as described above.

Brief Description of the Drawings

[0010]

Fig. 1

Fig. 2

Fig. 3

Fig. 4A

Fig. 4B

Fig. 4C

Fig. 4D

Fig. 4E

Fig. 5

Fig. 6

[0011] The accompanying drawings are included to provide a further understanding of the embodiments, are incorporated in and constitute a part of this specification, show the described embodiments, and together with the description serve to explain the underlying principles. Like reference numerals refer to corresponding elements.

[0012] Reference is made in detail to specific examples, which are illustrated in the accompanying drawings. In the following detailed description, numerous non - limiting specific details are set forth in order to assist in an understanding of the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternative forms can be used without departing from the scope of the claims, and the subject matter can be practiced without these specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented on many types of electronic devices having digital video capabilities.

[0013] FIG. 1 is a block diagram showing an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a destination device 14. The source device 12 and the destination device 14 can comprise any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set - top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some implementations, the source device 12 and the destination device 14 include wireless communication capabilities.

[0014] In one implementation, the destination device 14 can receive the encoded video data that is decoded via the link 16. The link 16 can include any type of communication medium or device that can move the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 may comprise a communication medium to enable the source device 12 to transmit the encoded video data directly and in real time to the destination device 14. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can comprise any wireless or wired communication medium, such as the radio frequency (RF) spectrum, or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, a switch, a base station, or any other device that may be useful to facilitate communication from the source device 12 to the destination device 14.

[0015] In some other implementations, the encoded video data may be sent from the output interface 22 to the storage device 32. Subsequently, the encoded video data within the storage device 32 can be accessed by the destination device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or local access data storage media, such as a hard drive, Blu-ray disk, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data. In a further example, the storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 can access the video data stored in the storage device 32 via streaming or downloading. The file server can be any type of computer that can store the encoded video data and send the encoded video data to the destination device 14. Exemplary file servers include web servers, FTP servers, NAS (network attached storage) devices, or local disk drives. The destination device 14 can access the encoded video data via any standard data connection, including a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0016] As shown in FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 can include a video capture device, such as a video camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As an example, when the video source 18 is a video camera of a security monitoring system, the source device 12 and the destination device 14 can form a camera phone or a video phone. However, the implementations described in this application are generally applicable to video encoding and are applicable to wireless and / or wired applications.

[0017] Captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data may be directly transmitted to the destination device 14 via the output interface 22 of the source device 12. The encoded video data can also be stored in the storage device 32 so that it can be accessed later by the destination device 14 or other devices for decoding and / or playback. The output interface 22 can further include a modem and / or a transmitter.

[0018] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 includes a receiver and / or a modem and can receive encoded video data via link 16. The encoded video data communicated via link 16 or provided on storage device 32 can include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.

[0019] In some implementations, display device 34, which can be an integrated display device of destination device 14, can include an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to the user and can include any of various display devices such as a liquid crystal display, a plasma display, an organic light emitting diode, or another type of display device.

[0020] Video encoder 20 and video decoder 30 can operate according to proprietary specifications or industry standards such as VVC, HEVC, MPEG-4, Part 10, AVC (Advanced Video Coding), or extensions of such standards. It should be understood that the present application is not limited to specific video encoding / decoding standards and is applicable to other video encoding / decoding standards. Generally, it is contemplated that video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, generally, it is also contemplated that video decoder 30 of destination device 14 can be configured to decode video data according to any of these current or future standards.

[0021] The video encoder 20 and the video decoder 30 can each be implemented as any of various suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device can store software instructions in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations of the present disclosure. Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and any of them may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.

[0022] FIG. 2 is a block diagram showing an exemplary video encoder 20 according to some implementations described in the present application. The video encoder 20 can perform intra prediction encoding and inter prediction encoding of video blocks within a video frame. Intra prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence.

[0023] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a splitting unit 45, an intra prediction processing unit 46, and an intra block copy unit 48. In some implementations, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A deblocking filter (not shown) can be arranged between the adder 62 and the DPB 64 to filter the block boundary and remove block noise artifacts from the reconstructed video. An in-loop filter (not shown) may be used to filter the output of the adder 62 in addition to the deblocking filter. The video encoder 20 can be in the form of a fixed or programmable hardware unit, or can be divided among one or more of the illustrated fixed or programmable hardware units.

[0024] The video data memory 40 can store video data encoded by the components of the video encoder 20. The video data in the video data memory 40 can be obtained, for example, from the video source 18. The DPB 64 is a buffer that stores reference video data for use in encoding video data by the video encoder 20 (e.g., in the intra prediction encoding mode or the inter prediction encoding mode). The video data memory 40 and the DPB 64 can be formed by any of various memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20, or off-chip with respect to these components.

[0025] As shown in FIG. 2, after receiving the video data, the splitting unit 45 in the prediction processing unit 41 splits the video data into video blocks. This splitting can also include splitting the video frame into slices, tiles, or other larger coding units (CUs) according to a predefined splitting structure such as a quadtree structure associated with the video data. The video frame can be split into a plurality of video blocks (or a set of video blocks called tiles). The prediction processing unit 41 can select, based on error results (e.g., coding rate and distortion level), one of a plurality of possible prediction coding modes for the current video block, such as one of a plurality of intra prediction coding modes or one of a plurality of inter prediction coding modes. The prediction processing unit 41 can provide the resulting intra or inter prediction coding block to the adder 50 to generate a residual block and subsequently provide it to the adder 62 that reconstructs the coding block for use as part of the reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra mode indicators, splitting information, and other such syntax information to the entropy coding unit 56.

[0026] To select an appropriate intra prediction coding mode for the current video block, the intra prediction processing unit 46 in the prediction processing unit 41 can perform intra prediction coding of the current video block with respect to one or more adjacent blocks within the same frame as the current block to be coded to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 in the prediction processing unit 41 perform inter prediction coding of the current video block with respect to one or more prediction blocks within one or more reference frames to provide temporal prediction. The video encoder 20 can execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of the video data.

[0027] In some implementations, the motion estimation unit 42 determines the inter prediction mode of the current video frame by generating a motion vector indicating the displacement of the prediction unit (PU) of the video block in the current video frame with respect to the prediction block in the reference video frame according to a predetermined pattern within the sequence of video frames. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector, and the motion vector estimates the motion of the video block. The motion vector can indicate, for example, the displacement of the PU of the video block in the current video frame or picture with respect to the prediction block in the reference frame (or other coded unit) for the current block being coded in the current frame (or other coded unit). The predetermined pattern can specify the video frames in the sequence as P frames or B frames. The intra BC unit 48 can determine a vector for intra BC coding, such as a block vector, in a manner similar to the determination of the motion vector by the motion estimation unit 42 for inter prediction, or the motion estimation unit 42 can be utilized to determine the block vector.

[0028] The prediction block is a block in the reference frame that is considered to closely match the PU of the video block to be coded with respect to pixel differences, which can be determined by SAD (sum of absolute difference), SSD (sum of square difference), or other difference metrics. In some implementations, the video encoder 20 can calculate the values of the sub - integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 can interpolate the values of the 1 / 4 - pixel position, 1 / 8 - pixel position, or other fractional pixel positions of the reference frame. Accordingly, the motion estimation unit 42 can perform motion search for both integer pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0029] The motion estimation unit 42 calculates the motion vector of the PU of the video block in the inter-predicted coded frame by comparing the position of the PU with the position of the predicted block of the reference frame selected from the first reference frame list (List0) or the second reference frame list (List1) that identifies one or more reference frames each stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44 and then to the entropy coding unit 56.

[0030] The motion compensation performed by the motion compensation unit 44 can include fetching or generating a predicted block based on the motion vector determined by the motion estimation unit 42. Upon receiving the motion vector for the PU of the current video block, the motion compensation unit 44 can search for the predicted block pointed to by the motion vector in one of the reference frame lists, retrieve the predicted block from the DPB 64, and transmit the predicted block to the adder 50. The adder 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the predicted block provided by the motion compensation unit 44 from the pixel values of the currently encoded video block. The pixel difference values forming the residual video block can include the luminance or chrominance components or both. The motion compensation unit 44 can also generate syntax elements related to the video blocks of the video frame for use by the video decoder 30 when decoding the video blocks of the video frame. The syntax elements can include, for example, syntax elements that define the motion vector used to identify the predicted block, any flags indicating the prediction mode, or any other syntax information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes.

[0031] In some implementations, the intra BC unit 48 can generate vectors and fetch prediction blocks in a manner similar to that described above in relation to the motion estimation unit 42 and the motion compensation unit 44, where the prediction blocks are in the same frame as the currently encoded block and the vectors are called block vectors as opposed to motion vectors. In particular, the intra BC unit 48 can determine the intra prediction mode to be used for encoding the current block. In some examples, the intra BC unit 48 can encode the current block using various intra prediction modes, for example, between separate encoding paths, and test their performance through rate-distortion analysis. Next, the intra BC unit 48 can use the appropriate intra prediction mode among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 can calculate rate-distortion values using rate-distortion analysis for the various tested intra prediction modes and select, as the appropriate intra prediction mode to use, the intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to generate the encoded block, as well as the bitrate (i.e., the number of bits) used to generate the encoded block. The intra BC unit 48 can calculate a ratio from the distortion and rate for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0032] In other examples, the intra BC unit 48 can use all or part of the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra BC prediction according to the implementation forms described in this specification. In any case, in the case of intra block copy, the prediction block can be a block that is considered to closely match the block to be encoded with respect to the pixel difference, which can be determined by SAD (sum of absolute difference), SSD (sum of squared difference), or other difference metrics. The identification of the prediction block can include the calculation of the values of sub-pixel positions.

[0033] Regardless of whether the prediction block is from the same frame according to intra prediction or from different frames according to inter prediction, the video encoder 20 can form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded to form pixel difference values. The pixel difference values for forming the residual video block can include both the luminance component difference and the chrominance difference.

[0034] As described above, the intra prediction processing unit 46 can perform intra prediction on the current video block as an alternative to the inter prediction executed by the motion estimation unit 42 and the motion compensation unit 44, or the intra block copy prediction executed by the intra BC unit 48. In particular, the intra prediction processing unit 46 can determine the intra prediction mode to be used for encoding the current block. To do so, the intra prediction processing unit 46 can, for example, encode the current block using various intra prediction modes during separate encoding paths, and the intra prediction processing unit 46 (or, in some examples, the mode selection unit) can select an appropriate intra prediction mode for use from among the tested intra prediction modes. The intra prediction processing unit 46 can provide information indicating the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 can encode the information indicating the selected intra prediction mode into the bitstream.

[0035] After the prediction processing unit 41 determines the predicted block of the current video block via either inter prediction or intra prediction, the adder 50 forms a residual video block by subtracting the predicted block from the current video block. The residual video data within the residual block can be included in one or more transform units (TUs) and supplied to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0036] The conversion processing unit 52 can send the obtained conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients to further reduce the bit rate. Also, the quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 can then perform a scan of the matrix containing the quantized conversion coefficients. Alternatively, the entropy encoding unit 56 may perform the scan.

[0037] Following quantization, the entropy encoding unit 56 entropy-encodes the quantized conversion coefficients into the video bitstream, for example, using context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream may then be sent to the video decoder 30, or may be archived in the storage device 32 for transmission or retrieval to a future video decoder 30. The entropy encoding unit 56 can also entropy-encode the motion vectors and other syntax elements for the current video frame being encoded.

[0038] The inverse quantization unit 58 and the inverse conversion processing unit 60 respectively apply inverse quantization and inverse conversion to reconstruct the residual video block within the pixel region for generating a reference block for prediction of other video blocks. As described above, the motion compensation unit 44 can generate a motion compensation prediction block from one or more reference blocks of the frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.

[0039] The adder 62 adds the residual block reconstructed into the motion compensation prediction block generated by the motion compensation unit 44 to generate a reference block for storing in the DPB 64. The reference block can then be used by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 as a prediction block for predicting another video block within a subsequent video frame.

[0040] FIG. 3 is a block diagram showing an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further has a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 can perform a decoding process that is substantially the reverse of the encoding process described above with respect to the video encoder 20 in relation to FIG. 2. For example, the motion compensation unit 82 can generate prediction data based on the motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 can generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.

[0041] In some examples, tasks may be assigned to the units of the video decoder 30 to perform the implementation of the present application. Also, in some examples, the implementation of the present disclosure can be divided among one or more of the units of the video decoder 30. For example, the intra BC unit 85 can perform the implementation of the present application alone or in combination with other units of the video decoder 30 such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 such as the motion compensation unit 82.

[0042] Since the video data memory 79 is to be decoded by other components of the video decoder 30, it can store video data such as an encoded video bitstream. The video data stored in the video data memory 79 can be obtained, for example, from the storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or a hard disk). The video data memory 79 can include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use when the video decoder 30 decodes video data (e.g., in an intra or inter prediction coding mode). The video data memory 79 and the DPB 92 can be formed by any of various memory devices such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM (registered trademark)), or other types of memory devices. For the sake of explanation, the video data memory 79 and the DPB 92 are shown as two separate components of the video decoder 30 in FIG. 3. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip with other components of the video decoder 30 or off-chip with respect to those components.

[0043] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. Video decoder 30 can receive syntax elements at the video frame level and / or at the video block level. The entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Next, the entropy decoding unit 80 transfers the motion vectors and other syntax elements to the prediction processing unit 81.

[0044] When a video frame is encoded as an intra prediction encoded (I) frame or for an intra encoded prediction block in another type of frame, the intra prediction processing unit 84 of the prediction processing unit 81 may generate prediction data for a video block of the current video frame based on the signaled intra prediction mode and reference data from previously decoded blocks of the current frame.

[0045] When a video frame is encoded as an inter prediction encoded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for a video block of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame in one of the reference frame lists. Video decoder 30 can configure the reference frame lists, list 0 and list 1, using default construction techniques based on the reference frames stored in DPB 92.

[0046] In some examples, when a video block is encoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within the reconstructed region of the same picture as the current video block defined by the video encoder 20.

[0047] The motion compensation unit 82 and / or the intra BC unit 85 determines prediction information for the video blocks of the current video frame by parsing the motion vectors and other syntax elements, and then uses the prediction information to generate a prediction block for the currently decoded video block. For example, the motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) used to encode the video blocks of the video frame, an inter prediction frame type (e.g., B or P), configuration information for one or more of the reference frame lists for the frame, a motion vector for each inter prediction encoded video block of the frame, an inter prediction status for each inter prediction encoded video block of the frame, and other information for decoding the video blocks in the current video frame.

[0048] Similarly, the intra BC unit 85 can use some of the received syntax elements, such as flags, to determine that the current video block was predicted using the intra BC mode, which video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, a block vector for each intra BC predicted video block of the frame, an intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding the video blocks within the current video frame.

[0049] Also, the motion compensation unit 82 may perform interpolation using an interpolation filter such as that used by the video encoder 20 during the encoding of video blocks, and calculate interpolation values for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filter used by the video encoder 20 from the received syntax elements, and generate a prediction block using the interpolation filter.

[0050] The inverse quantization unit 86 inverse quantizes the bitstream decoded by the entropy decoding unit 80 and the quantized transform coefficients provided to the entropy for each video block in the video frame using the same quantization parameter calculated by the video encoder 20, and determines the degree of quantization. The inverse transform processing unit 88 applies an inverse transform, such as an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to reconstruct the residual block in the pixel domain.

[0051] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. An in-loop filter (not shown) can be arranged between the adder 90 and the DPB 92 to further process the decoded video block. The decoded video blocks within a given frame are stored in the DPB 92 that stores the reference frames used for subsequent motion compensation of the next video block. The DPB 92, or a memory device separate from the DPB 92, can also store the decoded video for later presentation on a display device such as the display device 34 of FIG. 1.

[0052] In a typical video encoding process, a video sequence typically includes an ordered set of frames or pictures. Each frame can include three sample arrays denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other examples, a frame may be monochromatic and thus include only one two-dimensional array of luminance samples.

[0053] As shown in FIG. 4A, a video encoder 20 (or more specifically, a splitting unit 45) first generates an encoded representation of a frame by splitting the frame into a set of coding tree units (CTUs). A video frame can include an integer number of CTUs that are continuously ordered in a raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled to the video encoder 20 in a sequence parameter set such that all CTUs in the video sequence have the same size as one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As shown in FIG. 4B, each CTU can comprise one coding tree block (CTB) of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements used to encode the samples of the coding tree block. The syntax elements include properties of different types of units of an encoded block of pixels, such as inter or intra prediction, intra prediction mode, motion vectors, and other parameters, and describe how the video sequence can be reconstructed in a video decoder 30. In a monochrome picture or a picture having three separate color planes, a CTU can comprise a single coding tree block and syntax elements used to encode the samples of the coding tree block. The coding tree block may be an N×N block of samples.

[0054] To achieve better performance, the video encoder 20 can recursively perform tree splitting, such as binary tree splitting, ternary tree splitting, quadtree splitting, or a combination of both, on the coding tree blocks of the CTU, and split the CTU into smaller coding units (CUs). As shown in FIG. 4C, the 64×64 CTU 400 is first split into four smaller CUs, each having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each split into four CUs of 16×16 by the block size. The two 16×16 CUs 430 and 440 are each further split into four CUs of 8×8 by the block size. FIG. 4D shows a quadtree data structure showing the final result of the splitting process of the CTU 400 as shown in FIG. 4C, and each leaf node of the quadtree corresponds to one CU of each size in the range from 32×32 to 8×8. Similar to the CTU shown in FIG. 4B, each CU can include a coding block (CB) of luminance samples, two corresponding coding blocks of chrominance samples of a frame of the same size, and syntax elements used to code the samples of the coding block. In a monochrome picture or a picture having three separate color planes, the CU can include a single coding block and a syntax structure used to code the samples of the coding block. The quadtree splitting shown in FIGS. 4C and 4D is for illustrative purposes only, and it should be noted that one CTU can be split into CUs and adapted to various local characteristics based on quadtree / ternary tree / binary tree splitting. In a multi-type tree structure, one CTU is split by a quadtree structure, and each quadtree leaf CU can be further split by a binary tree structure and a ternary tree structure. As shown in FIG. 4E, there are five splitting types, namely, quadtree splitting, horizontal binary tree splitting, vertical binary tree splitting, horizontal ternary tree splitting, and vertical ternary tree splitting.

[0055] In some implementations, the video encoder 20 may further divide the coding block of the CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same prediction, inter or intra, is applied. The prediction unit (PU) of the CU may comprise a prediction block of luminance samples, two corresponding prediction blocks of chrominance samples, and syntax elements used to predict the prediction block. For a monochrome picture or a picture having three separate color planes, the PU can comprise a single prediction block and a syntax structure used to predict the prediction block. The video encoder 20 can generate prediction luminance, Cb, and Cr blocks for the luminance, Cb, and Cr prediction blocks of each PU of the CU.

[0056] The video encoder 20 may use intra prediction or inter prediction to generate a prediction block for the PU. When the video encoder 20 uses intra prediction to generate the prediction block of the PU, the video encoder 20 may generate the prediction block of the PU based on the decoded samples of the frame associated with the PU. When the video encoder 20 uses inter prediction to generate the prediction block of the PU, the video encoder 20 may generate the prediction block of the PU based on the decoded samples of one or more frames other than the frame associated with the PU.

[0057] After the video encoder 20 generates prediction luminance, Cb, and Cr blocks for one or more PUs of a CU, the video encoder 20 may generate a luminance residual block for the CU by subtracting the prediction luminance block of the CU from the original luminance encoded block such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample within one of the prediction luminance blocks of the CU and the corresponding sample in the original luminance encoded block of the CU. Similarly, the video encoder 20 may generate a Cb residual block for the CU such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample within one of the prediction Cb blocks of the CU and the corresponding sample in the original Cb encoded block of the CU, and may generate a Cr residual block for the CU such that each sample in the Cr residual block of the CU indicates the difference between a Cr sample within one of the prediction Cr blocks of the CU and the corresponding sample in the original Cr encoded block of the CU.

[0058] Further, as shown in FIG. 4C, the video encoder 20 may use quadtree partitioning to decompose the luminance, Cb, and Cr residual blocks of the CU into one or more luminance, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A transform unit (TU) of a CU may comprise a transform block of luminance samples, two corresponding transform blocks of chroma samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture having three separate color planes, a TU can comprise a single transform block and a syntax structure used to transform the samples of the transform block.

[0059] Video encoder 20 can apply one or more transforms to the luminance transform block of a TU to generate a luminance coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar values. Video encoder 20 can apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block of the TU. Video encoder 20 can apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0060] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 can quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized, likely reducing the amount of data used to represent the transform coefficients and providing further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 can entropy encode the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream including a bit sequence forming a representation of the encoded frame and associated data, which can either be stored in storage device 32 or transmitted to destination device 14.

[0061] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can analyze the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 can also reconstruct the encoded block of the current CU by adding the samples of the prediction block for the PU of the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the encoded block for each CU of the frame, the video decoder 30 can reconstruct the frame.

[0062] As described above, video encoding mainly uses two modes, namely, intra-frame prediction (or intra prediction) and inter-frame prediction (or inter prediction), to achieve video compression. Palette-based encoding is another encoding method adopted by many video encoding standards. In palette-based encoding, which may be particularly suitable for screen generation content encoding, a video coder (e.g., the video encoder 20 or the video decoder 30) forms a palette table of colors representing the video data of a given block. The palette table contains the most dominant (e.g., frequently used) pixel values within the given block. Pixel values that are not frequently represented in the video data of the specified block are either not included in the palette table or are included in the palette table as an escape color.

[0063] Each entry in the palette table contains an index of the corresponding pixel value in the palette table. The palette index for samples within a block may be encoded to indicate which entry from the palette table is used to predict or reconstruct which sample. This palette mode begins with the process of generating a palette predictor for the first block of a grouping of pictures, slices, tiles, or other video blocks. As will be described below, palette predictors for subsequent video blocks are typically generated by updating previously used palette predictors. For purposes of explanation, it is assumed that the palette predictor is defined at the picture level. In other words, a picture can contain multiple coded blocks each having its own palette table, but there is one palette predictor for the entire picture.

[0064] To reduce the bits required for signaling palette entries in the video bitstream, the video decoder can utilize a palette predictor to determine new palette entries within the palette table used for reconstruction of video blocks. For example, the palette predictor can include palette entries from a previously used palette table, or even start with the last used palette table by including all entries of the last used palette table. In some implementations, the palette predictor includes fewer entries than all entries from the last used palette table and can then incorporate some entries from other previously used palette tables. The palette predictor may have the same size as the palette table used to encode different blocks, or may be larger or smaller than the palette table used to encode different blocks. In one example, the palette predictor is implemented as a first-in first-out (FIFO) table containing 64 palette entries.

[0065] To generate a palette table for a block of video data from a palette predictor, a video decoder can receive a 1-bit flag for each entry of the palette predictor from an encoded video bitstream. The 1-bit flag can have a first value (e.g., binary 1) indicating that the associated entry of the palette predictor is included in the palette table, or a second value (e.g., binary 0) indicating that the associated entry of the palette predictor is not included in the palette table. If the size of the palette predictor is larger than the palette table used for a block of video data, the video decoder may stop receiving more flags when the maximum size of the palette table is reached.

[0066] In some implementations, instead of some entries of the palette table being determined using the palette predictor, they may be directly signaled within the encoded video bitstream. For such entries, the video decoder can receive three separate m-bit values from the encoded video bitstream, indicating the pixel values of the luminance and chrominance components associated with the entry (where m represents the bit depth of the video data). Compared to the multiple m-bit values required for directly signaled palette entries, those palette entries derived from the palette predictor require only a 1-bit flag. Thus, signaling some or all palette entries using the palette predictor can significantly reduce the number of bits required to signal new entries of the palette table, thereby improving the overall encoding efficiency of palette mode encoding.

[0067] In many cases, the palette predictor for a block is determined based on a palette table used to encode one or more previously encoded blocks. However, when encoding the first coding tree unit in a picture, slice, or tile, the palette table of previously encoded blocks may not be available. Therefore, it is not possible to generate a palette predictor using the entries of the previously used palette table. In such cases, the sequence of palette predictor initializers may be signaled in the sequence parameter set (SPS) and / or picture parameter set (PPS), which are values used to generate a palette predictor when the previously used palette table is not available. The SPS generally refers to the syntax structure of the syntax elements applied to a series of consecutive encoded video pictures called the coded video sequence (CVS), which is determined by the content of the syntax elements found in the PPS, which is referred to by the syntax elements found in each slice segment header. The PPS generally refers to the syntax structure of the syntax elements applied to one or more individual pictures in the CVS, as determined by the syntax elements found in each slice segment header. Therefore, the SPS is generally regarded as a higher level of syntax structure than the PPS, and the syntax elements included in the SPS generally mean that they are not changed as frequently as the syntax elements included in the PPS and are applied to a larger part of the video data.

[0068] FIG. 5 is a block diagram illustrating an example of determining and using a palette table for encoding video data within picture 500 according to some implementations of the present disclosure. Picture 500 includes a first block 510 associated with a first palette table 520 and a second block 530 associated with a second palette table 540. Since the second block 530 is to the right of the first block 510, the second palette table 540 may be determined based on the first palette table 520. Palette predictor 550 is associated with picture 500 and is used to collect zero or more palette entries from the first palette table 520 and construct zero or more palette entries within the second palette table 540. The various blocks shown in FIG. 5 can correspond to CTUs, CUs, PUs, or TUs as described above, and it should be noted that the blocks are not limited to the block structure of any particular encoding standard and can be compatible with future block-based encoding standards.

[0069] Generally, a palette table contains a dominant and / or representative number of pixel values in the currently encoded block (e.g., block 510 or 530 in FIG. 5). In some examples, a video coder (e.g., video encoder 20 or video decoder 30) can encode a palette table separately for each chroma of the block. For example, video encoder 20 can encode a palette table for the luminance component of the block, a separate palette table for the chroma Cb component of the block, and yet another separate palette table for the chroma Cr component of the block. In this case, the first palette table 520 and the second palette table 540 may each be a plurality of palette tables. In other examples, video encoder 20 can encode a single palette table for all chromas of the block. In this case, the i-th entry of the palette table is three values (Yi, Cbi, Cri), each value corresponding to one component of a pixel. Thus, the representations of the first palette table 520 and the second palette table 540 are merely examples and not limiting.

[0070] As described herein, rather than directly encoding the actual pixel values of the first block 510, a video coder (such as video encoder 20 or video decoder 30) can use a palette-based encoding scheme to encode the pixels of the first block 510 using indices I1, …, IN. For example, for each pixel within the first block 510, the video encoder 20 can encode the index value of the pixel, and the index value is associated with a pixel value within the first palette table 520. The video encoder 20 can encode the first palette table 520 and transmit it in the encoded video data bitstream for use by the video decoder 30 for palette-based decoding on the decoder side. In general, one or more palette tables may be transmitted per block or may be shared across different blocks. The video decoder 30 can obtain the index values from the video bitstream generated by the video encoder 20 and can reconstruct the pixel values using the corresponding pixel values of the index values within the first palette table 520. In other words, for each index value for the block, the video decoder 30 may determine an entry within the first palette table 520. The video decoder 30 then replaces each index value within the block with the pixel value specified by the determined entry within the first palette table 520.

[0071] In some implementations, a video coder (e.g., video encoder 20 or video decoder 30) determines a second palette table 540 based at least in part on a palette predictor 550 associated with a picture 500. The palette predictor 550 includes some or all of the entries of a first palette table 520 and may also include entries from other palette tables. In some examples, the palette predictor 550 is implemented using a first-in first-out table, where adding an entry of the first palette table 520 to the palette predictor 550 causes the currently oldest entry within the palette predictor 550 to be evicted to keep the palette predictor 550 below its maximum size. In other examples, different techniques may be used to update and / or maintain the palette predictor 550.

[0072] In one example, the video encoder 20 encodes a pred_palette_flag for each block (e.g., second block 530) to indicate whether the palette table for the block is predicted from one or more other palette tables associated with one or more other blocks such as an adjacent block 510. For example, if the value of such a flag is binary, the video decoder 30 may determine that the second palette table 540 for the second block 530 is predicted from one or more previously decoded palette tables and thus that a new palette table for the second block 540 is not included in the video bitstream containing the pred_palette_flag. If such a flag is binary zero, the video decoder 30 may determine that the second palette table 540 for the second block 530 is included in the video bitstream as a new palette table. In some examples, the pred_palette_flag may be encoded separately for different chroma of a block (e.g., for a video block in the YCbCr space, three flags, one for Y, one for Cb, and one for Cr). In other examples, a single pred_palette_flag may be encoded for all chroma of a block.

[0073] In the above example, pred_palette_flag is signaled block by block, indicating that all entries in the palette table of the current block are predicted. This means that the second palette table 540 is identical to the first palette table 520 and no additional information is signaled. In other examples, one or more syntax elements may be signaled entry by entry. That is, the flag can be signaled for each entry in the previous palette table to indicate whether the entry exists in the current palette table. If a palette entry is not predicted, the palette entry may be signaled explicitly. In other examples, these two methods can be combined.

[0074] When predicting the second palette table 540 according to the first palette table 520, the video encoder 20 and / or the video decoder 30 can identify the block for which the predicted palette table is determined. The predicted palette table may be associated with one or more adjacent blocks of the currently encoded block, i.e., the second block 530. As shown in FIG. 5, when determining the predicted palette table of the second block 530, the video encoder 20 and / or the video decoder 30 can identify the left adjacent block, i.e., the first block 510. In other examples, the video encoder 20 and / or the video decoder 30 can place one or more blocks at other positions relative to the second block 530, such as the upper blocks within the picture 500. In another example, the palette table of the last block in the scan order using the palette mode can be used as the predicted palette table of the second block 530.

[0075] Video encoder 20 and / or video decoder 30 can determine blocks for palette prediction according to a predetermined order of block positions. For example, video encoder 20 and / or video decoder 30 can first identify the left adjacent block, i.e., the first block 510, for palette prediction. If the left adjacent block is not available for prediction (e.g., the left adjacent block is encoded in a mode other than a palette-based coding mode such as an intra prediction mode or an inter prediction mode, or is located at the leftmost end of a picture or slice), video encoder 20 and / or video decoder 30 can identify the upper adjacent block within picture 500. Video encoder 20 and / or video decoder 30 can continue to search for available blocks according to the predetermined order of block positions until it finds a block having a palette table available for palette prediction. In some examples, video encoder 20 and / or video decoder 30 applies one or more formulas, functions, rules, etc. to generate a predicted palette table based on the palette tables of one or more adjacent blocks or based on a combination of multiple adjacent blocks (spatially or in scan order), so as to determine a predicted palette based on multiple blocks of adjacent blocks and / or reconstructed samples. In one example, a predicted palette table including palette entries from one or more previously encoded adjacent blocks includes the number of entries, N. In this case, video encoder 20 first transmits to video decoder 30 a binary vector V having the same size as the predicted palette table, i.e., size N. Each entry of the binary vector indicates whether the corresponding entry of the predicted palette table is reused or copied to the palette table of the current block. For example, V(i)=1 means that the i-th entry of the predicted palette table of the adjacent block is reused or copied to the palette table of the current block. There may be another index for the current block.

[0076] In yet other examples, the video encoder 20 and / or the video decoder 30 can construct a candidate list that includes a number of potential candidates for palette prediction. In such examples, the video encoder 20 can encode an index into the candidate list to indicate a candidate block within the list from which the current block used for palette prediction is selected. The video decoder 30 can construct the candidate list in the same way, decode the index, and use the decoded index to select the palette of the corresponding block for use with the current block. In another example, a palette table of the indicated candidate block in the list can be used as a prediction palette table for each entry of the palette table of the current block.

[0077] In some implementations, one or more syntax elements can indicate whether a palette table, such as the second palette table 540, is fully predicted from a prediction palette (e.g., a first palette table 520 that may be composed of entries from one or more previously encoded blocks), or whether a particular entry of the second palette table 540 is predicted. For example, an initial syntax element can indicate whether all entries within the second palette table 540 are predicted. If the initial syntax element indicates that not all entries are predicted (e.g., a flag having a binary zero value), one or more additional syntax elements can indicate which entries of the second palette table 540 are predicted from the prediction palette table.

[0078] In some implementations, the size of the palette table, e.g., the number of pixel values included in the palette table, may be fixed or may be signaled using one or more syntax elements within the encoded bitstream.

[0079] In some implementations, the video encoder 20 can encode the pixels of a block without exactly matching the pixel values in the palette table to the actual pixel values within the corresponding block of the video data. For example, the video encoder 20 and the video decoder 30 can merge or combine (i.e., quantize) different entries in the palette table when the entry pixel values are within a predetermined range of each other. In other words, if an existing pixel value within the error margin of the new pixel value already exists, the new pixel value is not added to the palette table, and the samples within the block corresponding to the new pixel value are encoded with the index of the existing pixel value. Note that this process of lossy encoding does not affect the operation of the video decoder 30. The video decoder can decode the pixel values in the same way regardless of whether a particular palette table is lossless or lossy.

[0080] In some implementations, the video encoder 20 can select an entry in the palette table as a predicted pixel value for encoding the pixel values within a block. The next video encoder 20 can then determine the difference between the actual pixel value and the selected entry as a residual and encode the residual. The video encoder 20 can generate a residual block that includes the residual values for the pixels within the block predicted by an entry in the palette table, and then apply transformation and quantization (as described above in connection with FIG. 2) to the residual block. In this way, the video encoder 20 can generate quantized residual transform coefficients. In another example, the residual block may be encoded losslessly (without transformation and quantization) or without transformation. The video decoder 30 can inverse-transform and inverse-quantize the transform coefficients to reproduce the residual block, and then reconstruct the pixel values using the predicted palette input value and the residual value for the pixel values.

[0081] In some implementations, the video encoder 20 can determine an error threshold called a delta value to build a palette table. For example, when the actual pixel value at a position within a block generates an absolute difference between the actual pixel value below the delta value and an existing pixel value entry in the palette table, the video encoder 20 can send an index value to identify the corresponding index of the pixel value entry in the palette table for use when reconstructing the actual pixel value at that position. When the actual pixel value at a position within a block generates an absolute difference value between the actual pixel value in the palette table and an existing pixel value entry that is greater than the delta value, the video encoder 20 can send the actual pixel value and add the actual pixel value as a new entry to the palette table. To configure the palette table, the video decoder 30 can use the delta value signaled by the encoder, depend on a fixed or known delta value, or infer or derive the delta value.

[0082] As described above, the video encoder 20 and / or the video decoder 30 can use encoding modes including an intra prediction mode, an inter prediction mode, a lossless encoding palette mode, and a lossy encoding palette mode when encoding video data. The video encoder 20 and the video decoder 30 can encode one or more syntax elements indicating whether palette-based encoding is enabled. For example, in each block, the video encoder 20 can encode a syntax element indicating whether a palette-based encoding mode is used for the block (e.g., a CU or a PU). For example, this syntax element can be signaled in the encoded video bitstream at the block level (e.g., the CU level) and then received by the video decoder 30 when decoding the encoded video bitstream.

[0083] In some implementations, the above syntax elements may be sent at a level higher than the block level. For example, the video encoder 20 may signal such syntax elements at the slice level, tile level, PPS level, or SPS level. In this case, a value equal to 1 indicates that all blocks below this level are encoded using the palette mode and that additional mode information, such as the palette mode or other modes, is not signaled at the block level. A value equal to zero indicates that blocks below this level are not encoded using the palette mode.

[0084] In some implementations, the fact that a higher-level syntax element enables the palette mode does not mean that each block at this higher or lower level must be encoded in the palette mode. Rather, another CU-level or TU-level syntax element is still needed to indicate whether a block at the CU level or TU level is encoded in the palette mode, and if so, the corresponding palette table should be constructed. In some implementations, the video coder (e.g., video encoder 20 and video decoder 30) selects a threshold (e.g., 32) for the number of samples in a block of the minimum block size such that the palette mode is not permitted for blocks whose block size is less than the threshold. In this case, signaling of syntax elements for such blocks is not performed. Note that the threshold for the minimum block size can be signaled explicitly in the bitstream or implicitly set to a default value that is compiled by both the video encoder 20 and the video decoder 30.

[0085] The pixel value at one position of a block may be the same as (or within the range of the delta value) the pixel values at other positions of the block. For example, it is common for adjacent pixel positions of a block to have the same pixel value or to be mapped to the same index value within a palette table. Thus, video encoder 20 can encode one or more syntax elements indicating the number of consecutive pixels or index values in a given scan order having the same pixel value or index value. A string of pixels or index values with similar values may be referred to herein as a "run". For example, if two consecutive pixels or indexes in a given scan order have different values, the run is equal to zero. If two consecutive pixels or indexes in a given scan order have the same value, but a third pixel or index in the scan order has a different value, the run is equal to 1. For three consecutive indexes or pixels having the same value, the run is 2, and so on. Video decoder 30 can obtain the syntax element indicating the run from the encoded bitstream and use that data to determine the number of consecutive positions having the same pixel or index value.

[0086] FIG. 6 is a flowchart showing an exemplary process 600 for implementing a technique by which video decoder 30 decodes video data using a palette-based scheme according to some implementations of the present disclosure. Specifically, video decoder 30 uses information related to the quantization parameter to determine whether escape samples (e.g., pixels that are not represented by palette entries in the palette table as described above in connection with FIG. 5 and that require additional overhead in the bitstream for signaling) within a palette mode encoded block are decoded in a lossless manner (i.e., the reconstructed pixel value of the escape sample is the same as the original pixel value) or in a lossy manner (i.e., the reconstructed pixel value of the escape sample is different from the original pixel value due to quantization).

[0087] To implement the palette-based method, video decoder 30 receives, from the bitstream, video data corresponding to a palette mode coded block (610). For example, the palette mode coded block includes both escape samples and non-escape samples (e.g., pixel values represented by palette entries in a palette table).

[0088] Next, video decoder 30 determines a quantization parameter value from information included in a parameter set associated with the palette mode coded block (620). For example, the quantization parameter value may be associated with one of a plurality of coding levels and may be obtained, for example, from a sequence parameter set, a picture parameter set, a group header corresponding to a group of tiles, a tile header, etc.

[0089] In some embodiments, for the quantization design for escape samples, the scale of quantization for escape samples is the same as the normal quantization used for samples (e.g., samples under a transform skip case and / or a transform case) with other coding tools when a specific quantization parameter is given, but the actual operation for escape sample quantization is defined to be different from the normal quantization operation. For example, the quantization operation for escape samples includes different shift and / or offset operations from regular quantization.

[0090] In some embodiments, the quantization process for escape samples and non-escape samples is the same. In one example, a standard quantization process is used for quantization of escape samples in palette mode. As a result, the quantization design for escape samples becomes the same as the quantization process for samples in transform skip mode and / or transform mode. The following equations relate to the corresponding quantization and inverse quantization processes applied in an encoder and a decoder when using quantization / inverse quantization in transform skip mode for coding palette escape colors.

Equation

[0091] pResi and pResi’ are the original and reconstructed residual coefficients, pLevel is the quantized value, transformShift is the bit shift used to compensate for the dynamic range increase due to the 2D transform, which is equal to 15 - bitDepth - (log2(W)+log2(H)) / 2, where W and H are the width and height of the current transform unit, and bitDepth is the coding bit depth. scale[] and descale[] are quantization and inverse quantization look-up tables with 14-bit and 6-bit precision, defined as follows. [Table 1] Table 1

[0092] If the size of the transform block is not a power of 4, another look-up table is defined as follows. [Table 2] Table 2

[0093] Next, the video decoder 30 identifies (630) the quantized escape samples within the palette mode coded block. For example, the quantized escape samples can be associated with additional overhead in the bitstream to distinguish them from non-escape samples represented by palette entries in the palette table.

[0094] When video decoder 30 determines an escape sample within a palette mode coded block, video decoder 30 determines a particular method for decoding the quantized escape sample based on information from the quantization parameter value. In accordance with the determination that the quantization parameter value is greater than a threshold (e.g., the value of quantization parameter (QP) is greater than 4) (640), video decoder 30 performs inverse quantization on the quantized escape sample according to a predefined formula (e.g., based on formula (2) and Table 1 or Table 2 described above) to obtain a reconstructed escape sample value (640-1). Since the reconstructed escape sample value may be different from the original sample value for quantization, this decoding process is typically a lossy process. In accordance with the determination that the quantization parameter value is less than or equal to the threshold (e.g., the value of QP is equal to 4 or less than 4) (650), video decoder 30 sets the reconstructed escape sample to the quantized escape sample value (650-1). In this case, since the reconstructed escape sample value is the same as the original sample value, the decoding process is a lossless process. In some embodiments, video decoder 30 performs step 640 and step 650 while determining the quantization parameter value from information included in a parameter set associated with the palette mode coded block in parallel (e.g., while performing step 620).

[0095] In some embodiments, the video decoder 30 performs inverse quantization on all quantization escape samples of all CUs within or below the level at which the quantization parameter value is signaled. For example, if the quantization parameter value is obtained from the picture parameter set, all quantized escape samples of the CUs within the picture are decoded according to a predetermined formula. In other words, if the value of QP is a specific value, for example, 4 or less, all CUs below the level at which the quantization parameter information is signaled may be encoded in lossless palette mode. If the value of QP is a specific value, for example, greater than 4, it means that the lossless palette mode cannot be used to encode any CU below the level at which this QP information is signaled.

[0096] In some embodiments, the video decoder 30 first determines the delta quantization parameter from the information included in the parameter set associated with the palette mode encoded block, and then determines the quantization parameter value by adding the delta quantization parameter to the reference quantization parameter value as the quantization parameter value. An exemplary code for signaling the delta quantization parameter value of the palette mode encoded CU is shown below. [Table 3] Table 3

[0097] In some embodiments, the delta quantization parameter is always signaled for the palette mode encoded block. An example of the code for always signaling the delta quantization parameter of the palette mode encoded block is shown below. [Table 4] Table 4

[0098] In some embodiments, the delta quantization parameters for the luminance component and the chroma component are always signaled separately for the palette mode coded block. Exemplary code for signaling the delta quantization parameters separately to the luminance component and the chroma component of the palette mode coded block is shown below. [Table 5] [Table 6] Table 5

[0099] In some embodiments, the quantization coefficients corresponding to the escape samples of the palette mode coded block are binary coded using fixed-length binarization based on the quantization parameter value and the bit depth. The fixed-length binary length is determined according to either of the following formulas. Binary length = (bit depth - quantization parameter value / 6) (3) Binary length = (bit depth - (floor(quantization parameter - 4) / 6)) (4)

[0100] In some embodiments, the binarization method is adaptively switched at different coding levels, such as a sequence parameter set (SPS), a picture parameter set (PPS), a slice, and / or a group of coded blocks. In such cases, the encoder has the flexibility to dynamically select the binarization method to signal information in the bitstream.

[0101] In some embodiments, when the palette mode is lossless, the fixed-length binarization process is used for the escape samples. In one example, the escape samples are directly coded based on their binary format values, and each bit is coded as a CABAC bypass bin when the palette mode is lossless.

[0102] In some embodiments, the parameter set associated with the pallet mode coded block includes the following.A first syntax element (e.g., palette_max_size) for specifying the maximum allowable palette size of a palette mode coding block, a second syntax element (e.g., palette_max_area) for specifying the maximum allowable palette area of a palette mode coding block, a third syntax element (e.g., palette_max_predictor_size) for specifying the maximum allowable palette predictor size of a palette mode coding block, a fourth syntax element (e.g., delta_palette_max_predictor_size) for specifying the difference between the maximum allowable palette predictor size and the maximum allowable palette size of a palette mode coding block, a fifth syntax element (e.g., palette_predictor_initializer_present_flag) for initializing a sequence palette predictor for a palette mode coding block, a sixth syntax element (e.g., num_palette_predictor_initializer_minus1) for identifying the number of entries of a palette predictor initializer minus 1 of a palette mode coding block, a seventh syntax element (e.g., palette_predictor_initializers[component][i]) for specifying the value of the component of the i-th palette entry used to initialize an array of predictor palette entries, an eighth syntax element (e.g., luma_bit_depth_entry_minus8_initializers) for identifying the bit depth value of the luminance component of the entries of a palette predictor initializer minus 8, a ninth syntax element (e.g., chroma_bit_depth_entry_minus8_initializers) for identifying the bit depth value of the chroma component of the entries of a palette predictor initializer minus 8, a tenth syntax element (e.g., luma_bit_depth_entry_minus8) for identifying the bit depth value of the luminance component of the entries of a palette minus 8, an eleventh syntax element (e.g., chroma_bit_depth_entry_minus8_initializers) for identifying the bit depth value of the chroma component of the entries of a palette minus 8.

[0103] In some embodiments, determining a quantization parameter value from information included in a parameter set associated with a palette mode coded block includes the following. In accordance with a determination that the quantization coefficients of the palette mode coded block do not correspond to escape samples of the palette mode coded block (e.g., the coefficients are coded with a palette), set the quantization parameter value equal to the quantization parameter value (QPcu) included in the parameter set corresponding to the current CU, and in accordance with a determination that the quantization coefficients of the palette mode coded block correspond to escape samples of the palette mode coded block, calculate the quantization parameter value according to the formula MIN(((MAX(4, QPcu) - 2) / 6)*6 + 4, 61).

[0104] In some embodiments, video decoder 30 determines a quantization parameter value from information included in a parameter set associated with a palette mode coded block in accordance with a determination that the quantization coefficients of the palette mode coded block correspond to escape samples of the palette mode coded block, and restricts the quantization parameter value according to the formula MAX(4, QPcu).

[0105] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium generally can correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to search for instructions, code, and / or data structures for implementations described herein. A computer program product may include a computer-readable medium.

[0106] The terms used in the description of the embodiments herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms such as "a" are intended to include the plural forms as well, unless the context clearly dictates otherwise. Also, the term "and / or" as used herein refers to any and all possible combinations of one or more of the associated listed items and is understood to encompass them. When the terms "including" and the like are used herein, they specify the presence of the described features, elements, and / or components, but it is further understood that they do not preclude the presence or addition of one or more other features, elements, components, and / or groups thereof.

[0107] Also, terms such as first, second, etc. may be used in this specification to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the implementation, the first electrode can be called the second electrode, and similarly, the second electrode can be called the first electrode. The first electrode and the second electrode are both electrodes, but they are not the same electrode.

[0108] The description of the present application is presented for purposes of illustration and explanation and is not intended to be exhaustive or limited to the invention in the disclosed form. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art who benefit from the teachings presented in the foregoing description and the associated drawings. This embodiment is selected and described in order to best explain the principles of the invention, its practical application, and to enable other skilled artisans to understand the invention for various implementations and to best utilize the underlying principles and various implementations with various modifications suitable for the particular uses contemplated. Accordingly, it should be understood that the claims are not intended to be limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the appended claims.

Claims

1. Divide a video frame into a plurality of blocks, Determine the quantization parameter of the current block among the plurality of blocks, The current block includes escape samples and non-escape samples, Identify the escape samples of the current block, When the quantization parameter is greater than a threshold value, Using a predetermined formula, based on the values of the escape samples, obtain the values of the reconstructed samples, When the quantization parameter is equal to the threshold value, Determine the value of the reconstructed sample as the value of the escape sample, Encode information related to the quantization parameter into a bitstream, Encoding the information related to the quantization parameter into the bitstream includes encoding the information related to the quantization parameter into the sequence parameter set of the bitstream, Determining the quantization parameter includes determining the quantization parameter according to the function: MAX(4, QPcu), QPcu represents the quantization parameter value corresponding to the current block, Video encoding method.

2. According to the determination that the quantization parameter is greater than a threshold value, Obtain the reconstructed sample by a loss process, According to the determination that the quantization parameter is equal to the threshold value, Obtain the reconstructed sample by a lossless process, The video encoding method according to claim 1.

3. The predetermined formula 【Table 1】 is related to a look-up table defined as The video encoding method according to claim 1.

4. Encode additional overhead related to the escape samples into the bitstream, The video encoding method according to claim 1, further comprising this.

5. The threshold value is equal to 4, The video encoding method according to claim 1.

6. Transmit the bitstream, The video encoding method according to claim 1.

7. An electronic device, One or more processing units, A memory connected to one or more of the processing units, When executed by one or more of the processing units, a plurality of programs stored in the memory that cause the electronic device to execute the video encoding method according to any one of claims 1 to 6, including, Electronic device.

8. A program executed by an electronic device including one or more processors, When executed by one or more of the processors, cause the electronic device to execute the video encoding method according to any one of claims 1 to 6. Program. **Claim 9** The bitstream is generated by the video encoding method according to any one of claims 1 to 6. Bitstream storage method. **Claim 10** The bitstream is generated by the video encoding method according to any one of claims 1 to 6. Bitstream transmission method.

Citation Information

Patent Citations

  • Coding escape pixels for palette mode coding

    US20160227225A1

  • Escape color coding for palette coding mode

    US20160227231A1

  • Improved encoding process using a palette mode

    US20160316214A1

  • Coding escape pixels for palette coding

    WO2016123513A1

  • Robust encoding / decoding of escape-coded pixels in palette mode

    WO2016197314A1