Method and apparatus for coding video data in palette mode
By determining joint or separate encoding for luma and chroma components and setting maximum palette and predictor sizes, the encoding efficiency in video coding standards is improved, addressing challenges in advanced formats like VVC/H.266.
Patent Information
- Application Number
- JP2025131223
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-30
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2040-11-17
AI Technical Summary
Existing video coding standards face challenges in optimizing the palette table size and predictor size for luma and chroma components in palette mode, which affects coding efficiency in advanced video coding formats like VVC/H.266.
Determine whether the luma and chroma components of a coding unit are jointly or separately encoded in palette mode, and set a maximum palette table size and predictor size based on this determination to improve encoding efficiency.
Enhances coding efficiency by optimizing palette table and predictor sizes, aligning with the goal of achieving higher compression performance in video coding standards like VVC/H.266.
Smart Images

Figure 2025169949000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 954,843, filed December 30, 2019, which is incorporated herein by reference in its entirety.
[0002] Technical Field
[0002] The present disclosure relates generally to video processing, and more particularly to methods and apparatus for signaling and determining maximum palette table size and maximum palette predictor size based on a coding tree structure for luma and chroma components in palette mode. [Background technology]
[0003] background
[0003] A video is a series of still pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video may be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC) / H.265 standard, the Versatile Video Coding (VVC) / H.266 standard, and the AVS standard, that specify specific video coding formats are developed by standardization organizations. As increasingly advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes increasingly higher. Summary of the Invention [Means for solving the problem]
[0004] Disclosure Overview
[0004] In some embodiments, an exemplary palette encoding method includes determining whether the luma component of a coding unit (CU) and the chroma components of the CU are jointly encoded in palette mode or separately encoded, and in response to the luma component and the chroma component being jointly encoded in palette mode, determining a first maximum palette table size for the CU, determining a first maximum palette predictor size for the CU, and predicting the CU based on the first maximum palette table size and the first maximum palette predictor size.
[0005] In some embodiments, an exemplary video processing device includes at least one memory for storing instructions and at least one processor configured to execute the instructions to cause the device to determine whether a luma component of a CU and a chroma component of the CU are jointly or separately coded in a palette mode, and, in response to the luma component and the chroma component being jointly coded in the palette mode, determine a first maximum palette table size for the CU, determine a first maximum palette predictor size for the CU, and predict the CU based on the first maximum palette table size and the first maximum palette predictor size.
[0006] In some embodiments, an exemplary non-transitory computer-readable storage medium stores a set of instructions executable by one or more processing devices to cause a video processing apparatus to: determine whether a luma component of a CU and a chroma component of the CU are jointly or separately coded in a palette mode; and, in response to the luma component and the chroma component being jointly coded in the palette mode, determine a first maximum palette table size for the CU, determine a first maximum palette predictor size for the CU, and predict the CU based on the first maximum palette table size and the first maximum palette predictor size.
[0007] BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments and various aspects of the present disclosure are set forth in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]
[0008] [Figure 1]
[0008] FIG. 1 is a schematic diagram illustrating the structure of an example video sequence, according to some embodiments of the present disclosure. [Figure 2A]
[0009] FIG. 2 is a schematic diagram illustrating an example encoding process of a hybrid video coding system consistent with embodiments of the present disclosure. [Figure 2B]
[0010] FIG. 2 is a schematic diagram illustrating another example encoding process of a hybrid video coding system, consistent with embodiments of the present disclosure. [Figure 3A]
[0011] FIG. 2 is a schematic diagram illustrating an example decoding process for a hybrid video coding system, consistent with embodiments of the present disclosure. [Figure 3B]
[0012] FIG. 2 is a schematic diagram illustrating another example decoding process for a hybrid video coding system, consistent with embodiments of the present disclosure. [Figure 4]
[0013] 1 is a block diagram of an exemplary apparatus for encoding or decoding video, consistent with some embodiments of the present disclosure. [Figure 5]
[0014] 1 shows a schematic diagram of an exemplary block coded in palette mode, according to some embodiments of the present disclosure. [Figure 6]
[0015] 1 illustrates a schematic diagram of an example process for updating a palette predictor after encoding a coding unit, according to some embodiments of the present disclosure. [Figure 7]
[0016] 1 shows exemplary Table 1 illustrating exemplary uniform maximum predictor sizes and maximum palette sizes according to some embodiments of the present disclosure. [Figure 8]
[0017] 1 shows exemplary Table 2 illustrating exemplary maximum predictor sizes and maximum palette sizes according to some embodiments of the present disclosure. [Figure 9]
[0018] 1 shows an example Table 3 illustrating an example decoding process using a predefined maximum palette predictor size and maximum palette size according to some embodiments of the present disclosure. [Figure 10]
[0019] 10 shows exemplary Table 4 illustrating an exemplary palette encoding syntax table for using a predefined maximum palette predictor size and maximum palette size according to some embodiments of the present disclosure. [Figure 11]
[0020] 10 shows exemplary Table 5 illustrating an exemplary derivation of the maximum palette size and maximum palette predictor size for an individual palette, according to some embodiments of the present disclosure. [Figure 12]
[0021] 10 shows exemplary Table 6 illustrating another exemplary derivation of the maximum palette size and maximum palette predictor size for an individual palette, according to some embodiments of the present disclosure. [Figure 13]
[0022] 1 shows exemplary Table 7 illustrating an exemplary sequence parameter set (SPS) syntax table according to some embodiments of the present disclosure. [Figure 14]
[0023] 10 shows exemplary Table 8 illustrating another exemplary derivation of the maximum palette size and maximum palette predictor size for an individual palette, according to some embodiments of the present disclosure. [Figure 15]
[0024] 10 shows exemplary Table 9 illustrating another exemplary derivation of the maximum palette size and maximum palette predictor size for an individual palette, according to some embodiments of the present disclosure. [Figure 16]
[0025] 1 shows exemplary Table 10 illustrating an exemplary picture header (PH) syntax according to some embodiments of the present disclosure. [Figure 17]
[0026] 10 shows exemplary Table 11 illustrating exemplary derivations of maximum palette size and maximum palette predictor size for I-slices, P-slices, and B-slices according to some embodiments of the present disclosure. [Figure 18]
[0027] 12 shows an example Table 12 illustrating an example slice header (SH) syntax according to some embodiments of the present disclosure. [Figure 19]
[0028] 1 illustrates a flowchart of an exemplary palette encoding method according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0009] Detailed Description
[0029] Reference will now be made in detail to the exemplary embodiments illustrated in the accompanying drawings. The following description will refer to the accompanying drawings in which like numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations set forth in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Instead, they are merely examples of apparatus and methods consistent with aspects related to the present invention as set forth in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with incorporated terms and / or definitions, the terms and definitions provided herein will control.
[0010]
[0030] The ITU-T Video Coding Expert Group (VCEG) and the ISO / IEC Moving Picture Expert Group (MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC) / H.266 standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC) / H.265 standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 but with half the bandwidth.
[0011]
[0031] To achieve the same subjective quality as HEVC / H.265 at half the bandwidth, JVET has developed technology that exceeds HEVC using the JEM (joint exploration model) reference software. As the coding technology has been incorporated into JEM, JEM has achieved significantly higher coding performance than HEVC.
[0012]
[0032] The VVC standard is a recent development and continues to add more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0013]
[0033] Video is a series of still pictures (or "frames") arranged in time sequence to preserve visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with a display capability) can be used to display such pictures in time sequence. In some applications, the video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for surveillance, conferencing, or live broadcasting.
[0014]
[0034] To reduce the storage space and transmission bandwidth required in such applications, video may be compressed before storage and transmission and decompressed before display. Compression and decompression may be performed by software executed by a processor (e.g., a processor in a general-purpose computer) or by dedicated hardware. A compression module is commonly referred to as an "encoder," and a decompression module is commonly referred to as a "decoder." Encoders and decoders may be collectively referred to as a "codec." Encoders and decoders may be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed on a computer-readable medium. Video compression and decompression may be performed by a variety of algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x family, etc. In some applications, a codec may decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard, in which case the codec is sometimes called a "transcoder."
[0015]
[0035] A video encoding process can identify and retain useful information that can be used for picture reconstruction and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such an encoding process may be called "lossy." Otherwise, it may be called "lossless." Most encoding processes are lossy; this is a tradeoff to reduce the required storage space and transmission bandwidth.
[0016]
[0036] Useful information about the picture being encoded (called the "current picture") includes changes relative to a reference picture (e.g., a previously encoded and reconstructed picture). Such changes may include pixel position changes, luminance changes, or color changes, of which position changes are the most important. Position changes of pixels representing an object may reflect the object's motion between the reference picture and the current picture.
[0017]
[0037] A picture that is coded without referencing another picture (i.e., it is its own reference picture) is called an "I-picture." A picture that is coded using a previous picture as a reference picture is called a "P-picture." A picture that is coded using both a previous picture and a future picture as reference pictures (i.e., the referencing is "bidirectional") is called a "B-picture."
[0018]
[0038] 1 illustrates the structure of an example video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be actual video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., actual video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files saved on a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.
[0019]
[0039] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture for a picture (e.g., picture 104) may not be immediately preceding or following that picture. For example, picture 104's reference picture may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiment to the example shown in FIG.
[0020]
[0040] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, they may divide a picture into basic segments and encode or decode the picture segment by segment. Such basic segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, a basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or a "coding tree unit" (CTU) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit may have a variable size of picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any shape and size of pixels. The size and shape of the basic processing unit may be selected for each picture based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit.
[0021]
[0041] A basic processing unit may be a logical unit that may include a collection of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic lightness information, one or more chroma components (e.g., Cb and Cr) representing color information, and related syntax elements (where the luma and chroma components may have the same size basic processing unit). The luma and chroma components are sometimes referred to as "coding tree blocks" (CTBs) in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit can be repeated for each of its luma and chroma components.
[0022]
[0042] Video coding has multiple stages of operation, examples of which are shown in FIGS. 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still be too large to process and therefore may be further divided into segments referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunits may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" (CUs) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits may have the same or smaller size as the basic processing units. Similar to basic processing units, basic processing subunits are also logical units that may contain a collection of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated on each of its luma and chroma components. Note that such division can be performed to further levels depending on the processing needs. Note also that different stages can use different schemes to divide the basic processing units.
[0023]
[0043] For example, in a mode decision stage (an example of which is shown in FIG. 2B ), the encoder may decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder may divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in the case of H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing sub-unit.
[0024]
[0044] As another example, in the prediction stage (an example of which is shown in FIGS. 2A-2B), the encoder can perform prediction operations at the level of basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and perform prediction operations at the level of the segments.
[0025]
[0045] As another example, in the transform stage (an example of which is shown in FIGS. 2A-2B), the encoder can perform transform operations on residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and perform transform operations at the segment level. Note that the division scheme of the same basic processing subunit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0026]
[0046] 1, the basic processing unit 112 is further divided into 3x3 basic processing sub-units, the boundaries of which are indicated by dotted lines. Different basic processing units of the same picture may be divided into basic processing sub-units in different schemes.
[0027]
[0047] In some implementations, to provide parallel processing capabilities and error resilience for video encoding and decoding, a picture can be divided into multiple regions for processing, allowing the encoding or decoding process to process a region of the picture without relying on information from any other region of the picture. That is, each region of the picture can be processed independently. This allows a codec to process different regions of a picture in parallel, thus improving coding efficiency. Also, if data for one region is corrupted during processing or lost during network transmission, the codec can accurately encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two region types: "slice" and "tile." It should also be noted that different pictures in video sequence 100 may have different partition schemes for dividing the picture into regions.
[0028]
[0048] 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing subunits, and regions of structure 110 in FIG. 1 are merely examples, and the present disclosure is not limited to these embodiments.
[0029]
[0049] FIG. 2A illustrates a schematic diagram of an example encoding process 200A consistent with embodiments of the present disclosure. For example, encoding process 200A can be performed by an encoder. As shown in FIG. 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 of FIG. 1, video sequence 202 can include a set of pictures (referred to as "original pictures") arranged in a temporal order. Similar to structure 110 of FIG. 1, each original picture in video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can perform process 200A at the level of basic processing units for each original picture in video sequence 202. For example, the encoder can perform process 200A in an iterative manner, in which case the encoder can encode one basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for a region of each original picture in video sequence 202 (eg, regions 114-118).
[0030]
[0050] In FIG. 2A , an encoder may send a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may generate a residual BPU 210 by subtracting the prediction BPU 208 from the original BPU. The encoder may send the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may send the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the “forward path.” During process 200A, after quantization stage 214, the encoder may send quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may generate a prediction reference 224 to be used in prediction stage 204 for the next iteration of process 200A by adding reconstructed residual BPU 222 to prediction BPU 208. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0031]
[0051] The encoder may perform process 200A iteratively to encode each original BPU of the original picture (in the forward path) and to generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.
[0032]
[0052] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any manner of inputting data.
[0033]
[0053] In the prediction stage 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224 and perform a prediction operation to generate predicted data 206 and a predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting predicted data 206, which can be used to reconstruct the original BPU as a predicted BPU 208 from the predicted data 206 and the prediction reference 224.
[0034]
[0054] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 typically differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder can subtract it from the original BPU to generate residual BPU 210. For example, the encoder can subtract pixel values (e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values of the original BPU. Each pixel of residual BPU 210 may have a residual value as a result of such subtraction between corresponding pixels of the original BPU and predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 may have fewer bits, but can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.
[0035]
[0055] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns may have the same size (e.g., the size of the residual BPU 210). Each basis pattern may represent a variation frequency (e.g., frequency of brightness variation) component of the residual BPU 210. No basis pattern can be reconstructed from any combination (e.g., linear combination) of the other basis patterns. That is, this decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is analogous to a discrete Fourier transform of a function, where the basis patterns are analogous to basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are analogous to the coefficients associated with the basis functions.
[0036]
[0056] Different transform algorithms can use different basis patterns. For example, various transform algorithms, such as a discrete cosine transform or a discrete sine transform, can be used in transform stage 212. The transform in transform stage 212 is reversible. That is, the encoder can reconstruct residual BPU 210 by inversely operating the transform (called an "inverse transform"). For example, to reconstruct pixels of residual BPU 210, the inverse transform may be multiplying the values of corresponding pixels of the basis pattern by their associated coefficients and adding these products to generate a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Compared to residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.
[0037]
[0057] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns may represent different variation frequencies (e.g., brightness variation frequencies). Because the human eye is generally good at recognizing low-frequency variations, the encoder can ignore high-frequency variation information without significant quality degradation in decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of high-frequency basis patterns may be converted to zero, and transform coefficients of low-frequency basis patterns may be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed into transform coefficients through the inverse operation of quantization (called "dequantization").
[0038]
[0058] Because the encoder ignores such division remainders in rounding operations, the quantization stage 214 may be lossy. In general, the quantization stage 214 may contribute the most information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values of the quantization parameter or other parameters of the quantization process.
[0039]
[0059] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, parameters of the prediction operation, the transform type in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), or encoder control parameters (e.g., bitrate control parameters). The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0040]
[0060] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may generate reconstructed transform coefficients by performing inverse quantization on the quantized transform coefficients 216. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may generate a prediction reference 224 to be used in the next iteration of process 200A by adding the reconstructed residual BPU 222 to a prediction BPU 208.
[0041]
[0061] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed by an encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages of FIG. 2A.
[0042]
[0062] 2B shows a schematic diagram of another example encoding process 200B consistent with embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x family). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.
[0043]
[0063] In general, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can predict a current BPU by using pixels from one or more already-encoded neighboring BPUs within the same picture. That is, the prediction reference 224 in spatial prediction may include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can predict a current BPU by using regions from one or more already-encoded pictures. That is, the prediction reference 224 in temporal prediction may include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0044]
[0064] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. With respect to the original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs encoded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, or polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, for example, by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the used video coding standard. In the case of intra prediction, the prediction data 206 may include, for example, the locations (e.g., coordinates) of the used neighboring BPUs, the sizes of the used neighboring BPUs, parameters of the extrapolation, or the orientations of the used neighboring BPUs relative to the original BPU.
[0045]
[0065] As another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. With respect to the original BPU of the current picture, the prediction reference 224 may include one or more pictures (called "reference pictures") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be encoded and reconstructed for each BPU. For example, the encoder may generate a reconstructed BPU by adding the reconstructed residual BPU 222 to the predicted BPU 208. Once all the reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (called a "search window") of the reference picture. The location of the search window in the reference picture may be determined based on the location of the original BPU in the current picture. For example, the search window may be centered at a location in the reference picture that has the same coordinates as the original BPU of the current picture, or may extend outward by a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., using a pel-recursive algorithm or a block-matching algorithm), the encoder can determine such a region as a matching region. The matching region may have different dimensions (e.g., smaller, equal, larger, or a different shape) than the original BPU. Because the reference picture and the current picture are temporally separated in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the location of the original BPU over time. The encoder may record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., like picture 106 in FIG. 1), the encoder can search for a matching region and determine its associated motion vector for each reference picture. In some embodiments, the encoder can assign weights to pixel values of the matching region in each matching reference picture.
[0046]
[0066] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, or zooming. In the case of inter prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, or weights associated with the reference pictures.
[0047]
[0067] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder can move the matching region of the reference picture according to the motion vector, in which case the encoder can predict the original BPU of the current picture. When multiple reference pictures are used (e.g., as in picture 106 of FIG. 1), the encoder can move the matching region of the reference picture according to each motion vector and average the pixel values of the matching region. In some embodiments, if the encoder assigns weights to the pixel values of the matching region of each matching reference picture, the encoder can add a weighted sum of the pixel values of the moved matching region.
[0048]
[0068] In some embodiments, inter-prediction may be unidirectional or bidirectional. Unidirectional inter-prediction may use one or more reference pictures in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which a reference picture (e.g., picture 102) precedes picture 104. Bidirectional inter-prediction may use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture in which reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.
[0049]
[0069] Referring further to the forward path of process 200B, after spatial prediction 2042 and temporal prediction stage 2044, in mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique in which the encoder can select a prediction mode to minimize the value of a cost function depending on the bitrates of candidate prediction modes and the distortion of reconstructed reference pictures under such candidate prediction modes. Depending on the selected prediction mode, the encoder can generate a corresponding predicted BPU 208 and predicted data 206.
[0050]
[0070] In the reconstruction path of process 200B, if an intra prediction mode was selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU encoded and reconstructed within the current picture), the encoder can send the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If an inter prediction mode was selected in the forward path, after generating the prediction reference 224 (e.g., the current picture with all BPUs encoded and reconstructed), the encoder can send the prediction reference 224 to a loop filter stage 232 where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by inter prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, or adaptive loop filtering. The loop-filtered reference picture may be stored in a buffer 234 (or a "decoded picture buffer") for later use (e.g., to be used as an inter-prediction reference picture for a future picture in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in the binary encoding stage 226 along with the quantized transform coefficients 216, the prediction data 206, and other information.
[0051]
[0071] FIG. 3A shows a schematic diagram of an example decoding process 300A consistent with embodiments of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder can follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, a decoder can perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder can decode one fundamental processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for a region (e.g., region 114-118) of each picture encoded in video bitstream 228.
[0052]
[0072] In FIG. 3A , a decoder may send a portion of a video bitstream 228 associated with an encoded picture's fundamental processing unit (referred to as an "encoding BPU") to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may send the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may send the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may generate a prediction reference 224 by adding the reconstructed residual BPU 222 to the prediction BPU 208. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may send the prediction reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.
[0053]
[0073] The decoder may iteratively perform process 300A to decode each encoded BPU of the encoded picture and generate a prediction reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all encoded BPUs of the encoded picture, the decoder may output the picture to video stream 304 for display and proceed to decoding the next encoded picture in video bitstream 228.
[0054]
[0074] In binary decoding stage 302, the decoder may perform the inverse operation of the binary encoding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, a transform type, parameters of the quantization process (e.g., quantization parameters), or encoder control parameters (e.g., bitrate control parameters). In some embodiments, if video bitstream 228 is transmitted in packets over a network, the decoder may depacketize video bitstream 228 before sending it to binary decoding stage 302.
[0055]
[0075] 3B shows a schematic diagram of another example decoding process 300B consistent with embodiments of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x family). Compared to process 300A, process 300B further divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and further includes loop filter stage 232 and buffer 234.
[0056]
[0076] In process 300B, for an encoding basic processing unit (referred to as the “current BPU”) of an encoded picture being decoded (referred to as the “current picture”), prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra prediction was used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, or the direction of the neighboring BPUs relative to the original BPU. As another example, if inter prediction was used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for the inter-prediction operation may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions in each reference picture, or one or more motion vectors associated with each of the matching regions.
[0057]
[0077] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in temporal prediction stage 2044. Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder can generate a prediction BPU 208. The decoder can generate a prediction reference 224 by adding the prediction BPU 208 and the reconstructed residual BPU 222, as shown in FIG. 3A.
[0058]
[0078] In process 300B, the decoder may send the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may send the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may send the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner shown in FIG. 2B . The loop filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as an inter-prediction reference picture for a future encoded picture in the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include parameters of the loop filter (e.g., loop filter strength).
[0059]
[0079] FIG. 4 is a block diagram of an example apparatus 400 for encoding or decoding video, according to an embodiment of the present disclosure. As shown in FIG. 4, the apparatus 400 may include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a dedicated machine for video encoding or decoding. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, the processor 402 may include any combination of several central processing units (i.e., "CPUs"), graphics processing units (i.e., "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems-on-chips (SoCs), or application-specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0060]
[0080] The device 400 may also include memory 404 configured to store data (e.g., an instruction set, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing stages of processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and the data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, CompactFlash (CF) cards, etc. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped as a single logical component.
[0061]
[0081] Bus 410 may be a communication device that transfers data between components within apparatus 400, such as an internal bus (e.g., a CPU memory bus) or an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port).
[0062]
[0082] For the sake of clarity and simplicity, in this disclosure, the processor 402 and other data processing circuitry will be collectively referred to as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single, independent module, or may be fully or partially integrated with any other components of the device 400.
[0063]
[0083] The device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, or a mobile communications network, etc.) In some embodiments, the network interface 406 may include any combination of several network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0064]
[0084] In some embodiments, apparatus 400 may optionally further include a peripheral interface 408 to provide connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), or a video input device (e.g., a camera, or an input interface coupled to a video archive), etc.
[0065]
[0085] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within apparatus 400. For example, some or all stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules of apparatus 400, such as program instructions that may be loaded into memory 404. As another example, some or all stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of apparatus 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, or NPUs).
[0066]
[0086] In the quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 in FIG. 2A or 2B, inverse quantization 218 in FIG. 3A or 3B), a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used for coding a picture or slice may be signaled at a high level, for example, using the init_qp_minus26 syntax element in the picture parameter set (PPS) and the slice_qp_delta syntax element in the slice header. Furthermore, the QP value may be adapted at a local level per CU using a delta QP value signaled at the granularity of the quantization group.
[0067]
[0087] In VVC, palette mode can be used with 4:4:4 color format. When palette mode is enabled, if the CU size is 64x64 or less, a flag is sent at the CU level indicating whether palette mode is used or not.
[0068]
[0088] FIG. 5 shows a schematic diagram of an example block 500 coded in palette mode, according to some embodiments of the present disclosure. As shown in FIG. 5, when palette mode is utilized to code a current CU (e.g., block 500), sample values at each position in the CU (e.g., position 501, position 502, position 503, or position 504) are represented by a small set of representative color values. This set is called a "palette" or "palette table" (e.g., palette 510). For sample positions with values close to palette colors, a corresponding palette index (e.g., index 0, index 1, index 2, or index 3) is signaled. According to some disclosed embodiments, color values outside the palette table can be specified by signaling an escape index (e.g., index 4). Then, for all positions in the CU that use the escape color index, the (quantized) color component values are signaled for each of these positions.
[0069]
[0089] To encode the palette table, a palette predictor is maintained. The palette predictor is initialized to 0 (e.g., empty) at the beginning of each slice in the non-wavefront case and at the beginning of each CTU row in the wavefront case. In some cases, the palette predictor can also be initialized to 0 at the beginning of a tile. FIG. 6 shows a schematic diagram of an example process for updating the palette predictor after encoding and decoding a coding unit according to some embodiments of the present disclosure. As shown in FIG. 6, for each entry in the palette predictor, a reuse flag is signaled indicating whether it is included in the current palette table of the current CU. The reuse flag is transmitted using run-length coding of zeros, followed by the number of new palette entries and the component values of the new palette entries. After encoding and / or decoding a palette-encoded CU, the palette predictor is updated using the current palette table, and entries from the previous palette predictor that are not reused in the current palette table are added to the end of the new palette predictor until it reaches the maximum allowed size.
[0070]
[0090] In some embodiments, for each CU, an escape flag is signaled to indicate whether an escape symbol exists in the current CU. If an escape symbol exists, the palette table is extended by one (as shown in FIG. 5) and the last index is assigned the escape symbol.
[0071]
[0091] Referring to Figure 5, the palette indices of samples in a CU form a palette index map. The index map is coded using horizontal or vertical transverse scanning. The scanning order is explicitly signaled in the bitstream using the syntax element "palette_transpose_flag". The palette index map is coded using index execution mode or index copy mode.
[0072]
[0092] According to some embodiments, the tree structure of an I-slice is signaled by the syntax element "qtbtt_dual_tree_intra_flag" in the Sequence Parameter Set (SPS) syntax. The syntax element qtbtt_dual_tree_intra_flag equal to 1 indicates that two separate coding_tree syntax structures are used for the luma and chroma components of the I-slice, respectively. The syntax element qtbtt_dual_tree_intra_flag equal to 0 indicates that separate coding_tree syntax structures are not used for the luma and chroma components of the I-slice. Furthermore, P slices and B slices are always coded as single-tree slices. Consistent with the disclosed embodiments, an I-picture is an intra-coded picture that does not reference other pictures in the encoding / decoding process. Both P pictures and B pictures are inter-coded pictures, which are decoded with reference to other pictures. The difference between P pictures and B pictures is that each block in a P picture can only refer to a maximum of one block in each reference picture, whereas each block in a B picture can refer to a maximum of two blocks in each reference picture.
[0073]
[0093] According to some embodiments, for slices with dual luma / chroma trees, different palettes (e.g., different palette tables) are applied separately to the luma (Y component) and chroma (Cb and Cr components). For dual-tree slices (e.g., dual luma / chroma trees), each entry in the luma palette table includes only a Y value, and each entry in the chroma palette table includes both a Cb value and a Cr value. For single-tree slices, the palette is applied jointly to the Y, Cb, and Cr components (e.g., each entry in the palette table includes a Y, Cb, and Cr value). Also, for certain color formats such as 4:2:0 and 4:2:2 color formats, coding units (CUs) in a single-tree slice may have separate luma and chroma trees due to limitations on the allowable minimum chroma coding block size. Therefore, for these color formats, a CU in a single-tree slice may have a local dual-tree structure (e.g., a single tree at the slice level but a dual tree at the CU level).
[0074]
[0094] Therefore, a single-tree slice coding unit may have separate luma and chroma trees in the case of a non-inter-SCIPU (smallest chroma intra prediction unit), since chroma is not allowed to be further split, but luma is. In single-tree coding, a SCIPU is defined as a coding tree node whose chroma block size is 16 chroma samples or more and has at least one child luma block with less than 64 luma samples. As mentioned above, the separate tree associated with a SCIPU is called a local dual tree.
[0075]
[0095] Based on the tree type of a slice (e.g., single tree or dual tree), two types of palette tables ("joint palette" and "separate palette") can be used for the slice. A single-tree slice can be palette coded using a joint palette table. Each entry in the joint palette table includes a Y color component, a Cb color component, and a Cr color component, and all color components of a coding unit (CU) in a single-tree slice (except for the local dual tree mentioned above) are jointly coded using the joint palette table. In contrast, a dual-tree slice is palette coded using two separate palettes. The luma and chroma components of a dual-tree slice require different palette tables and are coded separately. Therefore, for a dual-tree slice, two index maps (one for the luma component and one for the chroma component) are signaled in the bitstream.
[0076]
[0096] 7 shows exemplary Table 1 illustrating exemplary uniform maximum predictor sizes and maximum palette sizes according to some embodiments of the present disclosure. As shown in Table 1, the maximum palette predictor size for both the joint and individual palettes is uniformly set to 63, and the maximum palette size for both the joint and individual palette tables is uniformly set to 31. However, as noted above, for a dual-tree slice / CU with separate luma and chroma trees, two separate palette tables are required, while for a single-tree slice / CU with a joint luma-chroma tree, only one joint palette table is required. Thus, the complexity of generating a separate palette table for a dual-tree slice / CU is approximately twice the complexity of generating a joint palette table for a single-tree slice / CU.
[0077]
[0097] Consistent with some disclosed embodiments, to address the computational complexity and time imbalance associated with palette coding of dual-tree slices / CUs versus palette coding of single-tree slices / CUs, the maximum predictor size of the individual luma and chroma trees can be set smaller than the maximum predictor size of a single (e.g., joint) luma-chroma tree. Alternatively or additionally, the maximum palette size (e.g., maximum palette table size) of the individual luma and chroma trees can be set smaller than the maximum palette size of a single (e.g., joint) luma-chroma tree.
[0078]
[0098] In some disclosed embodiments, the following six variables are defined to represent maximum predictor sizes and maximum palette sizes: In particular, the variable "max_plt_predictor_size_joint" represents the maximum predictor size for a joint palette. The variable "max_plt_predictor_size_luma" represents the maximum predictor size for the luma component when a separate palette is used. The variable "max_plt_predictor_size_chroma" represents the maximum predictor size for the chroma components when a separate palette is used. The variable "max_plt_size_joint" represents the maximum palette size for a joint palette. The variable "max_plt_size_luma" represents the maximum palette size for the luma component when a separate palette is used. The variable "max_plt_size_chroma" represents the maximum palette size for the chroma components when a separate palette is used.
[0079]
[0099] In some embodiments, the maximum palette predictor size and maximum palette size are a predefined set of fixed values and do not need to be signaled to the video decoder. Figure 8 shows example Table 2 illustrating example maximum predictor sizes and maximum palette sizes according to some embodiments of this disclosure.
[0080]
[0100] In some embodiments, the maximum palette predictor size and maximum palette size for a dual-tree slice are set to half of the maximum palette predictor size and maximum palette size for a single-tree slice. As shown in Table 2, the maximum palette predictor size and maximum palette size for a joint palette (i.e., for a single-tree slice) are defined as 63 and 31, respectively. The maximum palette predictor size and maximum palette size for separate palettes for both the luma and chroma components (i.e., for a dual-tree slice) are defined as 31 and 15, respectively.
[0081]
[0101] 9 shows exemplary Table 3 illustrating an exemplary decoding process using a predefined maximum palette predictor size and maximum palette size according to some embodiments of the present disclosure. As shown in Table 3, changes to the palette mode decoding process currently proposed in VVC Draft 7 are highlighted and italicized in boxes 901-906, and content to be deleted from the palette mode decoding process currently proposed in VVC Draft 7 is shown in boxes 905-906, struck through, and italicized. In this embodiment, if a CU is coded as a local dual tree (e.g., separate luma / chroma local trees for a single tree slice), the maximum predictor size for coding the local dual tree is set to the maximum predictor size of the joint palette.
[0082]
[0102] 10 shows exemplary Table 4, which illustrates an exemplary palette encoding syntax table for using a predefined maximum palette predictor size and maximum palette size, according to some embodiments of the present disclosure. Changes to the syntax, as compared to the syntax used to implement a uniform maximum predictor size and maximum palette size shown in Table 1, are highlighted and italicized in boxes 1001-1003 in Table 4, and syntax elements that are deleted from the syntax are shown in boxes 1002-1003, struck through, and italicized in Table 4.
[0083]
[0103] In some embodiments, the maximum palette size of a joint palette and the difference between the maximum palette size of a joint palette and the maximum predictor size of a joint palette are signaled to the decoder via SPS syntax. Exemplary semantics consistent with this embodiment are described as follows: The syntax element "sps_max_plt_size_joint_minus1" specifies the maximum allowed palette size of a joint palette table minus 1. The value of the syntax element sps_max_plt_size_joint_minus1 is in the range of 0 to 63, inclusive. If the syntax element sps_max_plt_size_joint_minus1 is not present, its value is inferred to be 0. Additionally, the syntax element "sps_delta_max_plt_predictor_size_joint" specifies the difference between the maximum allowed palette predictor size and the maximum allowed palette size of a joint palette. The value of the syntax element sps_delta_max_plt_predictor_size_joint is in the range of 0 to 63, inclusive. If the syntax element sps_delta_max_plt_predictor_size_joint is not present, its value is inferred to be 0.
[0084]
[0104] The maximum palette size and maximum palette predictor size of the individual luma / chroma palettes are not signaled. Instead, they are derived from the maximum palette size and maximum palette predictor size of the joint palette. Figure 11 shows an example Table 5 illustrating an example derivation of the maximum palette size and maximum palette predictor size of the individual palettes according to some embodiments of the present disclosure.
[0085]
[0105] In the example shown in Table 5, when separate luma / chroma palettes are used, the syntax element max_plt_size_joint is distributed equally between the luma and chroma components. Consistent with this disclosure, it is also possible for the maximum palette size of a joint palette to be distributed unevenly between the luma and chroma components. Figure 12 shows exemplary Table 6 illustrating another example derivation of the maximum palette size and maximum palette predictor size for separate palettes, according to some embodiments of this disclosure. Table 6 shows an example of uneven distribution.
[0086]
[0106] 13 shows exemplary Table 7 illustrating an exemplary sequence parameter set (SPS) syntax table according to some embodiments of the present disclosure. Changes to the syntax, as compared to the syntax used to implement the uniform maximum predictor size and maximum palette size shown in Table 1, are highlighted in box 1301 and in italics in Table 7. Although not shown in Table 7, it is contemplated that the maximum palette size and maximum palette predictor size of the individual luma / chroma palettes may also be signaled in the SPS, along with the maximum palette size and maximum palette predictor size of the joint palette.
[0087]
[0107] In some embodiments, maximum palette size and maximum palette predictor size related syntax are transmitted through the picture header (PH). Exemplary semantics consistent with this embodiment are described as follows: The syntax element "pic_max_plt_size_joint_minus1" specifies the maximum allowed palette size of the joint palette table minus 1 for slices associated with the PH. The value of the syntax element pic_max_plt_size_joint_minus1 is in the range of 0 to 63, inclusive. The syntax element "pic_delta_max_plt_predictor_size_joint" specifies the difference between the maximum allowed palette predictor size and the maximum allowed palette size for the joint palette for slices associated with the PH. The maximum allowed value of the syntax element pic_delta_max_plt_predictor_size_joint is 63. If the syntax element pic_delta_max_plt_predictor_size_joint is not present, its value is inferred to be 0.
[0088]
[0108] The maximum palette size and maximum palette predictor size of the individual luma / chroma palettes are not signaled. Instead, they are derived from the maximum palette size and maximum palette predictor size of the joint palette. Figure 14 shows example Table 8 illustrating another example derivation of the maximum palette size and maximum palette predictor size of the individual palettes according to some embodiments of the present disclosure.
[0089]
[0109] In the example shown in Table 8, the syntax element max_plt_size_joint is distributed equally between the luma and chroma components when separate luma / chroma palettes are used. Consistent with this disclosure, it is also possible for the maximum palette size of a joint table to be distributed unevenly between the luma and chroma components. Figure 15 shows exemplary Table 9 illustrating another example derivation of the maximum palette size and maximum palette predictor size for separate palettes, according to some embodiments of this disclosure. Table 9 shows an example of uneven distribution.
[0090]
[0110] 16 shows exemplary Table 10 illustrating exemplary PH syntax, according to some embodiments of the present disclosure. Changes to the syntax, as compared to the syntax used to implement the uniform maximum predictor size and maximum palette size shown in Table 1, are highlighted in box 1601 and in italics in Table 10. Although not shown in Table 10, it is contemplated that the maximum palette size and maximum palette predictor size of the separate luma / chroma palettes may also be signaled in the picture header, along with the maximum palette size and maximum palette predictor size of the joint palette.
[0091]
[0111] In some embodiments, syntax related to the maximum palette size and maximum palette predictor size is signaled in each slice by the slice header. Exemplary semantics consistent with this embodiment are described as follows:
[0092]
[0112] Specifically, the syntax elements "slice_max_plt_size_joint_minus1" and "slice_delta_max_plt_predictor_size_joint" are conditionally signaled if the slice is coded as a single-tree slice. The syntax element slice_max_plt_size_joint_minus1 specifies the maximum allowed palette size of the joint palette table minus 1 for a single-tree slice. It is a bitstream conformance requirement that the maximum value of the syntax element slice_max_plt_size_joint is 63. The syntax element slice_delta_max_plt_predictor_size_joint specifies the difference between the maximum allowed palette predictor size and the maximum allowed palette size of the joint palette for a single-tree slice. The maximum allowed value of the syntax element slice_delta_max_plt_predictor_size_joint is 63. If the syntax element slice_delta_max_plt_predictor_size_joint is not present, its value is inferred to be 0.
[0093]
[0113] The syntax elements "slice_max_plt_size_luma_minus1" and "slice_delta_max_plt_predictor_size_luma" are conditionally signaled if the slice is coded as a dual-tree slice. The syntax element slice_max_plt_size_luma_minus1 specifies the maximum allowed palette size of the luma palette table minus 1 for a dual-tree slice. If the syntax element slice_max_plt_size_luma is not present, its value is inferred to be 0. It is a bitstream conformance requirement that the maximum value of the syntax element slice_max_plt_size_luma_minus1 be 63. The syntax element slice_delta_max_plt_predictor_size_luma specifies the difference between the maximum allowed palette predictor size and the maximum allowed palette size of the luma palette for a dual-tree slice. The maximum allowed value of the syntax element slice_delta_max_plt_predictor_size_luma is 63. If the syntax element slice_delta_max_plt_predictor_size_luma is not present, its value is inferred to be 0.
[0094]
[0114] FIG. 17 shows an example Table 11 illustrating an example derivation of maximum palette size and maximum palette predictor size for I-slices, P-slices, and B-slices according to some embodiments of the present disclosure.
[0095]
[0115] 18 shows exemplary Table 12 illustrating exemplary SH syntax according to some embodiments of the present disclosure. Changes to that syntax, compared to the syntax used to implement uniform maximum predictor size and maximum palette size shown in Table 1, are highlighted in box 1801 and in italics in Table 12. The prediction update procedure consistent with this embodiment is the same as that shown in Table 3, and the palette encoding syntax consistent with this embodiment is the same as that shown in Table 4.
[0096]
[0116] 19 shows a flowchart of an exemplary palette encoding method 1900 according to some embodiments of the present disclosure. Method 1900 may be performed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may perform method 1900. In some embodiments, method 1900 may be implemented by a computer program product embodied in a computer-readable medium including computer-executable instructions, such as program code, executed by a computer (e.g., apparatus 400 of FIG. 4).
[0097]
[0117] In step 1901, a decision can be made as to whether the luma component of a CU and the chroma components of the CU are jointly or separately coded in palette mode. For example, a variable treeType can be used to indicate whether the luma component of a CU and the chroma components of the CU are jointly or separately coded in palette mode (e.g., as shown in Table 3 of FIG. 9 or Table 4 of FIG. 10).
[0098]
[0118] In step 1903, a first maximum palette table size for the CU may be determined in response to the luma and chroma components being jointly coded in palette mode. In some embodiments, the first maximum palette table size for the CU may be determined based on the value of a first syntax element signaled in the video bitstream (e.g., the syntax element sps_max_plt_size_joint_minus1 shown in Table 7 of FIG. 13 or the syntax element pic_max_plt_size_joint_minus1 shown in Table 10 of FIG. 16).
[0099]
[0119] At step 1905, a first maximum palette predictor size for the CU may be determined in response to the luma and chroma components being jointly coded in palette mode. In some embodiments, the first maximum palette predictor size for the CU may be determined based on the value of a first syntax element and the value of a second syntax element signaled in the video bitstream (e.g., the syntax element sps_delta_max_plt_predictor_size_joint shown in Table 7 of FIG. 13 or the syntax element pic_delta_max_plt_predictor_size_joint shown in Table 10 of FIG. 16). For example, the first maximum palette predictor size for the CU may be determined to be the sum of the value of the first syntax element and the value of the second syntax element (e.g., as shown in Table 5 of FIG. 11, Table 6 of FIG. 12, Table 8 of FIG. 14, or Table 9 of FIG. 15). In some embodiments, the first syntax element and the second syntax element are signaled in an SPS associated with the CU (e.g., as shown in Table 7 of FIG. 13) or in a PH associated with the CU (e.g., as shown in Table 10 of FIG. 16).
[0100]
[0120] In step 1907, in response to the luma and chroma components being jointly coded in palette mode, a CU may be predicted based on a first maximum palette table size and a first maximum palette predictor size. For example, a CU may be predicted as shown in Table 3 of FIG. 9.
[0101]
[0121] In some embodiments, method 1900 may include, in response to the luma component and the chroma component being coded separately in palette mode, determining a second maximum palette table size for the CU based on the first maximum palette table size, determining a second maximum palette predictor size for the CU based on the first maximum palette predictor size, and predicting the CU based on the second maximum palette table size and the second maximum palette predictor size. The second maximum palette table size or the second maximum palette predictor size is for the luma component or the chroma component. For example, the maximum palette table size or the maximum palette predictor size for the luma component or the chroma component can be determined based on Table 5 of FIG. 11 , Table 6 of FIG. 12 , Table 8 of FIG. 14 , or Table 9 of FIG. 15 .
[0102]
[0122] In some embodiments, method 1900 may include determining a first maximum palette table size for the CU to be a first predetermined value. Method 1900 may also include determining a first maximum palette predictor size for the CU to be a second predetermined value. For example, as shown in Table 2 of FIG. 8, the maximum palette table size for the joint palette may be 31, and the maximum palette predictor size for the joint palette may be 63. In some embodiments, method 1900 may include determining a third maximum palette table size for the CU to be a third predetermined value, in response to the luma and chroma components being coded separately in the palette mode, and predicting the CU based on the third maximum palette table size. The third predetermined value is smaller than the first predetermined value. For example, as shown in Table 2 of FIG. 8, the maximum palette table size for the luma or chroma palette may be 15.
[0103]
[0123] In some embodiments, the method 1900 may include predicting the CU based on a first maximum palette predictor size (e.g., as shown in Table 3 of FIG. 9 ) in response to the luma and chroma components being coded separately in palette mode and the CU being part of a single tree slice.
[0104]
[0124] In some embodiments, method 1900 may include, in response to the luma and chroma components being coded separately in palette mode and the CU not being part of a single tree slice, determining a third maximum palette predictor size for the CU to be a fourth predetermined value and predicting the CU based on the third maximum palette predictor size. The fourth predetermined value is less than the second predetermined value. For example, as shown in Table 2 of FIG. 8, the maximum palette predictor size for a separate palette may be 31. The CU may be predicted as shown in Table 3 of FIG. 9.
[0105]
[0125] In some embodiments, the method 1900 may include determining whether a picture slice including a CU is a single-tree slice or a dual-tree slice, and, in response to the picture slice being a single-tree slice, determining a first maximum palette table size for the CU in the picture slice based on a value of a third syntax element signaled in a slice header of the picture slice, and determining a first maximum palette predictor size for the CU based on the value of the third syntax element and the value of a fourth syntax element signaled in the slice header. The first maximum palette predictor size for the CU may be determined to be the sum of the value of the third syntax element and the value of the fourth syntax element. For example, as shown in Table 11 of Figure 17, in response to a picture slice being a single-tree slice (e.g., slice_type != I | | qtbtt_dual_tree_intra_flag == 0), the maximum palette table size for the joint palette can be determined based on the value of the syntax element slice_max_plt_size_joint_minus1 signaled in the slice header (e.g., SH as shown in Table 12 of Figure 18), and the maximum palette predictor size for the joint palette can be determined to be the sum of the value of the syntax element slice_max_plt_size_joint_minus1 and the value of the syntax element slice_delta_max_plt_predictor_size_joint signaled in the slice header (e.g., SH as shown in Table 12 of Figure 18).
[0106]
[0126] In some embodiments, the method 1900 may include, in response to the picture slice being a dual-tree slice, determining a fourth maximum palette table size for the CU based on a value of a fifth syntax element signaled in the slice header, determining a fourth maximum palette predictor size for the CU based on the value of the fifth syntax element and the value of a sixth syntax element signaled in the slice header, and predicting the CU based on the fourth maximum palette table size and the fourth maximum palette predictor size. The fourth maximum palette predictor size for the CU may be determined to be the sum of the value of the fifth syntax element and the value of the sixth syntax element. For example, as shown in Table 11 of Figure 17, in response to a picture slice being a dual tree slice, the maximum palette table size for the luma or chroma palette can be determined based on the value of the syntax element slice_max_plt_size_luma_minus1 signaled in the slice header (e.g., SH as shown in Table 12 of Figure 18), and the maximum palette predictor size for the luma or chroma palette can be determined to be the sum of the value of the syntax element slice_max_plt_size_luma_minus1 and the value of the syntax element slice_delta_max_plt_predictor_size_luma signaled in the slice header (e.g., SH as shown in Table 12 of Figure 18).
[0107]
[0127] The embodiments can be further described using the following clauses. 1. Determining whether the luma component of a coding unit (CU) and the chroma component of the CU are jointly coded in palette mode or coded separately; In response to the luma and chroma components being jointly encoded in palette mode, determining a first maximum palette table size for the CU; determining a first maximum palette predictor size for the CU; and predicting a CU based on a first maximum palette table size and a first maximum palette predictor size; A palette encoding method comprising: 2. Determining a first maximum palette table size for a CU includes: 10. The method of claim 1, comprising determining a first maximum palette table size for a CU based on a value of a first syntax element signaled in a video bitstream. 3. Determining a first maximum palette predictor size for a CU 3. The method of claim 2, comprising determining a first maximum palette predictor size for a CU based on a value of a first syntax element and a value of a second syntax element signaled in the video bitstream. 4. Determining a first maximum palette predictor size for a CU 4. The method of clause 3, comprising determining a first maximum palette predictor size for the CU to be a sum of a value of a first syntax element and a value of a second syntax element. 5. The method of clause 3 or 4, wherein the first syntax element and the second syntax element are signaled in a sequence parameter set (SPS) associated with the CU. 6. The method of clause 3 or 4, wherein the first syntax element and the second syntax element are signaled in a picture header (PH) associated with the CU. 7. In response to the luma and chroma components being separately coded in palette mode, determining a second maximum palette table size for the CU based on the first maximum palette table size; determining a second maximum palette predictor size for the CU based on the first maximum palette predictor size; and predicting a CU based on a second maximum palette table size and a second maximum palette predictor size; 7. The method of any one of clauses 1 to 6, further comprising: 8. The method of clause 7, wherein the second maximum palette table size or the second maximum palette predictor size is for the luma component or the chroma component. 9. Determining a first maximum palette table size for a CU includes: 10. The method of claim 1, comprising determining a first maximum palette table size for the CU to be a first predetermined value. 10. Determining a first maximum palette predictor size for a CU 10. The method of any one of clauses 1 to 9, comprising determining a first maximum palette predictor size for the CU to be a second predetermined value. 11. In response to the luma component and the chroma component being separately coded in palette mode, determining a third maximum palette table size for the CU to be a third predetermined value; and predicting the CU based on a third maximum palette table size; further comprising 11. The method of clause 9 or 10, wherein the third predetermined value is less than the first predetermined value. 12. The method of clause 11, further comprising predicting a CU based on a first maximum palette predictor size in response to the luma component and chroma component being coded separately in palette mode and the CU being part of a single tree slice. 13. In response to the luma and chroma components being coded separately in palette mode and the CU not being part of a single tree slice, determining a third maximum palette predictor size for the CU to be a fourth predetermined value; and predicting a CU based on a third maximum palette predictor size; further comprising 13. The method of any one of clauses 10 to 12, wherein the fourth predetermined value is less than the second predetermined value. 14. Determining whether the picture slice containing the CU is a single-tree slice or a dual-tree slice; In response to the picture slice being a single tree slice, determining a first maximum palette table size for a CU in the picture slice based on a value of a third syntax element signaled in a slice header of the picture slice; and determining a first maximum palette predictor size for the CU based on the value of the third syntax element and the value of the fourth syntax element signaled in the slice header; 2. The method of clause 1, further comprising: 15. Determining a first maximum palette predictor size for a CU includes: 15. The method of clause 14, comprising determining a first maximum palette predictor size for the CU to be the sum of a value of a third syntax element and a value of a fourth syntax element. 16. In response to the picture slice being a dual tree slice, determining a fourth maximum palette table size for the CU based on a value of a fifth syntax element signaled in the slice header; determining a fourth maximum palette predictor size for the CU based on the value of the fifth syntax element and the value of the sixth syntax element signaled in the slice header; and predicting a CU based on a fourth maximum palette table size and a fourth maximum palette predictor size; 16. The method of clause 14 or 15, further comprising: 17. Determining a fourth maximum palette predictor size for a CU 17. The method of clause 16, comprising determining a fourth maximum palette predictor size for the CU to be the sum of a value of a fifth syntax element and a value of a sixth syntax element. 18. A video processing device, at least one memory for storing instructions; and at least one processor, wherein the at least one processor: determining whether a luma component of a coding unit (CU) and a chroma component of the CU are jointly coded in palette mode or separately coded; In response to the luma and chroma components being jointly encoded in palette mode, determining a first maximum palette table size for the CU; determining a first maximum palette predictor size for the CU; and predicting a CU based on a first maximum palette table size and a first maximum palette predictor size; 11. A video processing device configured to execute instructions to cause the device to: 19. At least one processor: determining a first maximum palette table size for the CU based on a value of a first syntax element signaled in the video bitstream; 19. An apparatus as described in clause 18, configured to execute instructions to cause the apparatus to: 20. At least one processor: determining a first maximum palette predictor size for the CU based on the value of the first syntax element and the value of a second syntax element signaled in the video bitstream; 20. An apparatus as described in clause 19, configured to execute instructions to cause the apparatus to: 21. At least one processor: determining a first maximum palette predictor size for the CU to be the sum of the value of the first syntax element and the value of the second syntax element; 21. An apparatus as described in clause 20, configured to execute instructions to cause the apparatus to: 22. The apparatus of clause 20 or 21, wherein the first syntax element and the second syntax element are signaled in a sequence parameter set (SPS) associated with the CU. 23. The apparatus of clause 20 or 21, wherein the first syntax element and the second syntax element are signaled in a picture header (PH) associated with the CU. 24. At least one processor: In response to the luma and chroma components being separately coded in palette mode, determining a second maximum palette table size for the CU based on the first maximum palette table size; determining a second maximum palette predictor size for the CU based on the first maximum palette predictor size; and predicting a CU based on a second maximum palette table size and a second maximum palette predictor size; 24. An apparatus according to any one of clauses 18 to 23, configured to execute instructions to cause the apparatus to: 25. The apparatus of clause 24, wherein the second maximum palette table size or the second maximum palette predictor size is for a luma component or a chroma component. 26. At least one processor: determining a first maximum palette table size for the CU to be a first predetermined value; 19. An apparatus as described in clause 18, configured to execute instructions to cause the apparatus to: 27. At least one processor: determining a first maximum palette predictor size for the CU to be a second predetermined value; 27. An apparatus according to clause 18 or 26, configured to execute instructions to cause the apparatus to: 28. At least one processor: In response to the luma and chroma components being separately coded in palette mode, determining a third maximum palette table size for the CU to be a third predetermined value; and predicting the CU based on a third maximum palette table size; configured to execute instructions to cause the device to 28. The apparatus of clause 26 or 27, wherein the third predetermined value is less than the first predetermined value. 29. At least one processor: predicting a CU based on a first maximum palette predictor size in response to the luma component and the chroma component being separately coded in a palette mode and the CU being part of a single tree slice; 29. An apparatus as described in clause 28, configured to execute instructions to cause the apparatus to: 30. At least one processor: In response to the luma and chroma components being coded separately in palette mode and the CU not being part of a single tree slice, determining a third maximum palette predictor size for the CU to be a fourth predetermined value; and predicting a CU based on a third maximum palette predictor size; configured to execute instructions to cause the device to 30. The apparatus of any one of clauses 27 to 29, wherein the fourth predetermined value is less than the second predetermined value. 31. At least one processor: determining whether a picture slice containing a CU is a single-tree slice or a dual-tree slice; In response to the picture slice being a single tree slice, determining a first maximum palette table size for a CU in the picture slice based on a value of a third syntax element signaled in a slice header of the picture slice; and determining a first maximum palette predictor size for the CU based on the value of the third syntax element and the value of the fourth syntax element signaled in the slice header; 19. An apparatus as described in clause 18, configured to execute instructions to cause the apparatus to: 32. At least one processor: determining a first maximum palette predictor size for the CU to be the sum of the value of a third syntax element and the value of a fourth syntax element; 32. The apparatus of clause 31, configured to execute instructions to cause the apparatus to: 33. At least one processor: In response to the picture slice being a dual tree slice, determining a fourth maximum palette table size for the CU based on a value of a fifth syntax element signaled in the slice header; determining a fourth maximum palette predictor size for the CU based on the value of the fifth syntax element and the value of the sixth syntax element signaled in the slice header; and predicting a CU based on a fourth maximum palette table size and a fourth maximum palette predictor size; 33. An apparatus according to clause 31 or 32, configured to execute instructions to cause the apparatus to: 34. At least one processor: determining a fourth maximum palette predictor size for the CU to be the sum of the value of the fifth syntax element and the value of the sixth syntax element; 34. An apparatus as described in clause 33, configured to execute instructions to cause the apparatus to: 35. A non-transitory computer-readable storage medium having stored thereon an instruction set, the instruction set comprising: determining whether a luma component of a coding unit (CU) and a chroma component of the CU are jointly coded in palette mode or separately coded; In response to the luma and chroma components being jointly encoded in palette mode, determining a first maximum palette table size for the CU; determining a first maximum palette predictor size for the CU; and predicting a CU based on a first maximum palette table size and a first maximum palette predictor size; A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including: 36. The instruction set is determining a first maximum palette table size for the CU based on a value of a first syntax element signaled in the video bitstream; 36. A non-transitory computer-readable storage medium as recited in clause 35, executable by one or more processing devices to cause a video processing device to perform 37. The instruction set is determining a first maximum palette predictor size for the CU based on the value of the first syntax element and the value of a second syntax element signaled in the video bitstream; 37. The non-transitory computer-readable storage medium of clause 36, executable by one or more processing devices to cause a video processing device to perform 38. The instruction set is determining a first maximum palette predictor size for the CU to be the sum of the value of the first syntax element and the value of the second syntax element; 38. A non-transitory computer-readable storage medium as recited in clause 37, executable by one or more processing devices to cause a video processing device to perform 39. The non-transitory computer-readable storage medium of clause 37 or 38, wherein the first syntax element and the second syntax element are signaled in a sequence parameter set (SPS) associated with the CU. 40. The non-transitory computer-readable storage medium of clause 37 or 38, wherein the first syntax element and the second syntax element are signaled in a picture header (PH) associated with the CU. 41.The instruction set is In response to the luma and chroma components being separately coded in palette mode, determining a second maximum palette table size for the CU based on the first maximum palette table size; determining a second maximum palette predictor size for the CU based on the first maximum palette predictor size; and predicting a CU based on a second maximum palette table size and a second maximum palette predictor size; A non-transitory computer-readable storage medium according to any one of clauses 35 to 40, executable by one or more processing devices to cause a video processing device to perform the above. 42. The non-transitory computer-readable storage medium of clause 41, wherein the second maximum palette table size or the second maximum palette predictor size is for a luma component or a chroma component. 43.The instruction set is determining a first maximum palette table size for the CU to be a first predetermined value; 36. A non-transitory computer-readable storage medium as recited in clause 35, executable by one or more processing devices to cause a video processing device to perform 44.The instruction set is determining a first maximum palette predictor size for the CU to be a second predetermined value; 44. A non-transitory computer-readable storage medium as described in clause 35 or 43, executable by one or more processing devices to cause a video processing device to perform the steps of: 45.The instruction set is In response to the luma and chroma components being separately coded in palette mode, determining a third maximum palette table size for the CU to be a third predetermined value; and predicting the CU based on a third maximum palette table size; executable by one or more processing devices to cause a video processing device to perform 45. The non-transitory computer-readable storage medium of clause 43 or 44, wherein the third predetermined value is less than the first predetermined value. 46. The instruction set is predicting a CU based on a first maximum palette predictor size in response to the luma component and the chroma component being separately coded in a palette mode and the CU being part of a single tree slice; 46. A non-transitory computer-readable storage medium as recited in clause 45, executable by one or more processing devices to cause a video processing device to perform 47.The instruction set is In response to the luma and chroma components being coded separately in palette mode and the CU not being part of a single tree slice, determining a third maximum palette predictor size for the CU to be a fourth predetermined value; and predicting a CU based on a third maximum palette predictor size; executable by one or more processing devices to cause a video processing device to perform 47. The non-transitory computer-readable storage medium of any one of clauses 44 to 46, wherein the fourth predetermined value is less than the second predetermined value. 48. The instruction set is determining whether a picture slice containing a CU is a single-tree slice or a dual-tree slice; In response to the picture slice being a single tree slice, determining a first maximum palette table size for a CU in the picture slice based on a value of a third syntax element signaled in a slice header of the picture slice; and determining a first maximum palette predictor size for the CU based on the value of the third syntax element and the value of the fourth syntax element signaled in the slice header; 36. A non-transitory computer-readable storage medium as recited in clause 35, executable by one or more processing devices to cause a video processing device to perform 49. The instruction set is determining a first maximum palette predictor size for the CU to be the sum of the value of a third syntax element and the value of a fourth syntax element; 49. A non-transitory computer-readable storage medium as recited in Clause 48, executable by one or more processing devices to cause a video processing device to perform the steps of: 50.The instruction set is In response to the picture slice being a dual tree slice, determining a fourth maximum palette table size for the CU based on a value of a fifth syntax element signaled in the slice header; determining a fourth maximum palette predictor size for the CU based on the value of the fifth syntax element and the value of the sixth syntax element signaled in the slice header; and predicting a CU based on a fourth maximum palette table size and a fourth maximum palette predictor size; 50. A non-transitory computer-readable storage medium as described in clause 48 or 49, executable by one or more processing devices to cause a video processing device to perform the steps of: 51.The instruction set is determining a fourth maximum palette predictor size for the CU to be the sum of the value of the fifth syntax element and the value of the sixth syntax element; 51. A non-transitory computer-readable storage medium as described in clause 50, executable by one or more processing devices to cause a video processing device to perform the steps of:
[0108]
[0128] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which may be executed by a device (such as the disclosed encoders and decoders) to perform the methods described above. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or other magnetic data storage media, CD-ROMs, other optical data storage media, any physical media with a pattern of holes, RAM, PROMs, and EPROMs, FLASH-EPROMs or other flash memories, NVRAMs, caches, registers, other memory chips or cartridges, and networked versions of the above. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0109]
[0129] It should be noted that relational terms herein, such as "first" and "second," are used only to distinguish one entity or operation from another, and do not require or imply an actual relationship or ordering between those entities or operations. Also, the words "comprising," "having," "containing," and "including," and other similar forms, are intended to be equivalent in meaning and to be open-ended in that the term or terms following any one of these terms is not an exhaustive list of such term or terms, or limited to only the listed term or terms.
[0110]
[0130] As used herein, unless specifically stated otherwise, the term "or" encompasses all possible combinations unless impracticable. For example, if it is stated that a database may include A or B, then the database may include A, or B, or A and B, unless specifically stated otherwise or impracticable. As a second example, if it is stated that a database may include A, B, or C, then the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless specifically stated otherwise or impracticable.
[0111]
[0131] It is understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it may be stored in the above computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art will also understand that more than one of the above modules / units can be integrated into one module / unit, and that each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0112]
[0132] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the above specification and examples be considered exemplary only, with the true scope and spirit of the invention being indicated by the following claims. Additionally, the order of steps depicted in the figures is intended for illustrative purposes only and is not intended to be limited to any particular order of steps. Thus, one skilled in the art will recognize that these steps may be performed in different orders while performing the same method.
[0113]
[0133] In the drawings and specification, illustrative embodiments have been disclosed. However, many variations and modifications to these embodiments may be made. Accordingly, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. determining whether a luma component of a coding unit (CU) and a chroma component of the CU are jointly coded in palette mode or separately coded; In response to the luma component and the chroma component being jointly encoded in the palette mode, determining a first maximum palette table size for the CU; determining a first maximum palette predictor size for the CU; and predicting the CU based on the first maximum palette table size and the first maximum palette predictor size; A palette encoding method comprising:
2. determining the first maximum palette table size for the CU, determining the first maximum palette table size for the CU based on a value of a first syntax element signaled in a video bitstream; determining the first maximum palette predictor size for the CU based on the value of the first syntax element and a value of a second syntax element signaled in the video bitstream; The method of claim 1 , comprising:
3. determining the first maximum palette predictor size for the CU, 3. The method of claim 2, comprising determining the first maximum palette predictor size for the CU to be the sum of the value of the first syntax element and the value of the second syntax element.
4. The method of claim 2 , wherein the first syntax element and the second syntax element are signaled in a sequence parameter set (SPS) or a picture header (PH) associated with the CU.
5. In response to the luma component and the chroma component being separately encoded in the palette mode, determining a second maximum palette table size for the CU based on the first maximum palette table size; determining a second maximum palette predictor size for the CU based on the first maximum palette predictor size; and predicting the CU based on the second maximum palette table size and the second maximum palette predictor size; The method of claim 1 further comprising:
6. determining the first maximum palette table size for the CU, determining the first maximum palette table size for the CU to be a first predetermined value; determining the first maximum palette predictor size for the CU to be a second predetermined value; The method of claim 1 , comprising:
7. the first predetermined value is 31; and 7. The method of claim 6, wherein the second predetermined value is 63.
8. In response to the luma component and the chroma component being separately encoded in the palette mode, determining a third maximum palette table size for the CU to be a third predetermined value; and predicting the CU based on the third maximum palette table size; further comprising The method of claim 6 , wherein the third predetermined value is less than the first predetermined value.
9. The method of claim 8 , wherein the third predetermined value is 15.
10. 9. The method of claim 8, further comprising: predicting the CU based on the first maximum palette predictor size in response to the luma component and the chroma component being separately coded in the palette mode and the CU being part of a single tree slice.
11. In response to the luma component and the chroma component being separately coded in the palette mode and the CU not being part of a single tree slice, determining a third maximum palette predictor size for the CU to be a fourth predetermined value; and predicting the CU based on the third maximum palette predictor size; further comprising The method of claim 8 , wherein the fourth predetermined value is less than the second predetermined value.
12. The method of claim 11 , wherein the fourth predetermined value is 31.
13. determining whether a picture slice containing the CU is a single-tree slice or a dual-tree slice; In response to the picture slice being a single tree slice, determining the first maximum palette table size for the CU in the picture slice based on a value of a third syntax element signaled in a slice header of the picture slice; and determining the first maximum palette predictor size for the CU based on the value of the third syntax element and a value of a fourth syntax element signaled in the slice header; The method of claim 1 further comprising:
14. In response to the picture slice being a dual tree slice, determining a fourth maximum palette table size for the CU based on a value of a fifth syntax element signaled in the slice header; determining a fourth maximum palette predictor size for the CU based on the value of the fifth syntax element and a value of a sixth syntax element signaled in the slice header; and predicting the CU based on the fourth maximum palette table size and the fourth maximum palette predictor size; 14. The method of claim 13, further comprising:
15. A video processing device, at least one memory for storing instructions; and at least one processor, wherein the at least one processor: determining whether a luma component of a coding unit (CU) and a chroma component of the CU are jointly coded in palette mode or separately coded; In response to the luma component and the chroma component being jointly encoded in the palette mode, determining a first maximum palette table size for the CU; determining a first maximum palette predictor size for the CU; and predicting the CU based on the first maximum palette table size and the first maximum palette predictor size; a video processing device configured to execute the instructions to cause the device to:
16. 1. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising: determining whether a luma component of a coding unit (CU) and a chroma component of the CU are jointly coded in palette mode or separately coded; In response to the luma component and the chroma component being jointly encoded in the palette mode, determining a first maximum palette table size for the CU; determining a first maximum palette predictor size for the CU; and predicting the CU based on the first maximum palette table size and the first maximum palette predictor size; A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including:
17. The instruction set comprises: determining the first maximum palette table size for the CU based on a value of a first syntax element signaled in a video bitstream; determining the first maximum palette predictor size for the CU based on the value of the first syntax element and a value of a second syntax element signaled in the video bitstream; 20. The non-transitory computer-readable storage medium of claim 16, executable by the one or more processing devices to cause the video processing device to:
18. The instruction set comprises: In response to the luma component and the chroma component being separately encoded in the palette mode, determining a second maximum palette table size for the CU based on the first maximum palette table size; determining a second maximum palette predictor size for the CU based on the first maximum palette predictor size; and predicting the CU based on the second maximum palette table size and the second maximum palette predictor size; 20. The non-transitory computer-readable storage medium of claim 16, executable by the one or more processing devices to cause the video processing device to:
19. The instruction set comprises: determining the first maximum palette table size for the CU to be a first predetermined value; determining the first maximum palette predictor size for the CU to be a second predetermined value; 20. The non-transitory computer-readable storage medium of claim 16, executable by the one or more processing devices to cause the video processing device to:
20. The instruction set comprises: In response to the luma component and the chroma component being separately encoded in the palette mode, determining a third maximum palette table size for the CU to be a third predetermined value; and predicting the CU based on the third maximum palette table size; executable by the one or more processing devices to cause the video processing device to perform 20. The non-transitory computer-readable storage medium of claim 19, wherein the third predetermined value is less than the first predetermined value.
Citation Information
Patent Citations
Maximum palette parameters in palette-based video coding
JP2017520162A
Methods of Palette Based Prediction for Non-444 Color Format in Video and Image Coding
US20170374372A1