Methods for palette prediction

Optimizing palette predictor updates and deblocking filters in video coding systems by simplifying the palette predictor process and treating palette mode as an independent coding mode improves processing efficiency and reduces resource consumption.

JP2025113279AActive Publication Date: 2025-08-01HFI INNOVATION INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025080012
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-04-04
Filing Date
2025-05-12
Publication Date
2025-08-01
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

Existing video coding standards face inefficiencies in palette predictor update processes and deblocking filter strength determination for palette mode, leading to hardware resource constraints and suboptimal performance.

Method used

Simplify the palette predictor update process by setting a fixed number of reuse flags and initialize the palette predictor to a predefined value, and treat palette mode as an independent coding mode for deblocking filter strength determination.

Benefits of technology

Enhances processing efficiency and reduces resource consumption in hardware implementations by optimizing palette predictor updates and deblocking filters in video coding systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113279000001_ABST
    Figure 2025113279000001_ABST
Patent Text Reader

Abstract

To provide a computer-implemented method for encoding video.SOLUTION: The method includes: receiving a video frame for processing; generating one or more coding units of the video frame; and processing one or more coding units using one or more palette predictors having palette entries, where each palette entry of the one or more palette predictors has a corresponding reuse flag, and where the number of reuse flags for each palette predictor is set to a fixed number for a corresponding coding unit.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications

[0001] This disclosure claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 005,305, filed Apr. 4, 2020, and U.S. Provisional Patent Application No. 63 / 002,594, filed Mar. 31, 2020, both of which are hereby incorporated by reference in their entirety.

[0002] Technical Field

[0002] This disclosure generally relates to video processing, and more particularly, to the use of palette modes in video encoding and decoding.

Background Art

[0003] Background

[0003] Video is a set of still pictures (or “frames”) that capture visual information. To reduce memory storage and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly techniques based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Video coding standards such as the HEVC (High Efficiency Video Coding) (HEVC / H.265) standard, the VVC (Versatile Video Coding) (VVC / H.266) standard, and the AVS standard, which specify particular video coding formats, have been developed by standardization organizations. As more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is becoming higher.

Summary of the Invention

Means for Solving the Problems

[0004] Summary of the Disclosure

[0004] Embodiments of the present disclosure provide a computer-implemented method for a palette predictor. In some embodiments, the method includes receiving a video frame for processing, generating one or more coded units of the video frame, and processing the one or more coded units using one or more palette predictors having palette entries, wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and the number of reuse flags of each palette predictor is set to a fixed number for the corresponding coded unit.

[0005]

[0005] Embodiments of the present disclosure provide an apparatus. In some embodiments, the apparatus includes a memory configured to store instructions and a processor coupled to the memory, the processor being configured to execute the instructions to cause the apparatus to receive a video frame for processing, generate one or more coded units of the video frame, and process the one or more coded units using one or more palette predictors having palette entries, wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and the number of reuse flags of each palette predictor is set to a fixed number for the corresponding coded unit.

[0006]

[0006] Embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing. In some embodiments, the method includes receiving a video frame for processing, generating one or more coded units of the video frame, and processing the one or more coded units using one or more palette predictors having palette entries, wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and the number of reuse flags of each palette predictor is set to a fixed number for the corresponding coded unit.

[0007]

[0007] Embodiments of the present disclosure provide a computer-implemented method for deblocking a palette-mode filter. In some embodiments, the method includes receiving a video frame for processing, generating one or more coded units of the video frame, each coded unit having one or more coded blocks, and setting a boundary filter strength to 1 in response to at least a first coded block of two adjacent coded blocks being coded in palette mode and a second coded block of the two adjacent coded blocks having a coding mode different from the palette mode.

[0008]

[0008] Embodiments of the present disclosure provide an apparatus. In some embodiments, the apparatus includes a memory configured to store instructions and a processor coupled to the memory. The processor is configured to execute the instructions to cause the apparatus to receive a video frame for processing, generate one or more coded units of the video frame, each coded unit having one or more coded blocks, and set a boundary filter strength to 1 in response to at least a first coded block of two adjacent coded blocks being coded in palette mode and a second coded block of the two adjacent coded blocks having a coding mode different from the palette mode.

[0009]

[0009] Embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to initiate a method for performing video data processing. In some embodiments, the method includes receiving a video frame for processing, generating one or more coding units of the video frame, each coding unit having one or more coding blocks, and setting a boundary filter strength to 1 in response to at least a first coding block of two adjacent coding blocks being coded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode.

[0010] Brief Description of the Drawings

[0010] Embodiments and various aspects of the present disclosure are shown in the following detailed description and the accompanying drawings. The various features shown in the figures are not drawn to scale.

Brief Description of the Drawings

[0011]

Figure 1

[0011] It is a schematic diagram illustrating the structure of an example of a video sequence according to some embodiments of the present disclosure.

Figure 2A

[0012] It is a schematic diagram illustrating an exemplary encoding process of a hybrid video coding system conforming to an embodiment of the present disclosure.

Figure 2B

[0013] It is a schematic diagram illustrating another exemplary encoding process of a hybrid video coding system conforming to an embodiment of the present disclosure.

Figure 3A

[0014] It is a schematic diagram illustrating an exemplary decoding process of a hybrid video coding system conforming to an embodiment of the present disclosure.

Figure 3B

[0015] It is a schematic diagram illustrating another exemplary decoding process of a hybrid video coding system that conforms to an embodiment of the present disclosure.

Figure 4

[0016] It is a block diagram of an exemplary apparatus for encoding or decoding video according to some embodiments of the present disclosure.

Figure 5

[0017] It illustrates a block encoded in palette mode according to some embodiments of the present disclosure.

Figure 6

[0018] An example of a palette predictor update process is shown.

Figure 7

[0019] An example of a palette coding syntax is shown.

Figure 8

[0020] It shows a flowchart of a palette predictor update process according to some embodiments of the present disclosure. [[ID=<<MASK_BEGIN>>10<<MASK_END>>]]

Figure 9

[0021] An example of a palette predictor update process according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>11<<MASK_END>>]]

Figure 10

[0022] An example of a decoding process in palette mode is shown. [[ID=<<MASK_BEGIN>>12<<MASK_END>>]]

Figure 11

[0023] An example of a decoding process in palette mode according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>13<<MASK_END>>]]

Figure 12

[0024] It shows a flowchart of another palette predictor update process according to some embodiments of the present disclosure. [[ID=<<MASK_BEGIN>>14<<MASK_END>>]]

Figure 13

[0025] An example of another palette predictor update process according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>15<<MASK_END>>]]

Figure 14

[0026] An example of a palette coding syntax according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>16<<MASK_END>>]]

Figure 15

[0027] An example of a decoding process in palette mode according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>17<<MASK_END>>]]

Figure 16

[0028] Hardware design of a decoder for a palette mode that implements a part of a palette predictor update process according to some embodiments of the present disclosure is illustrated. [[ID=<<MASK_BEGIN>>18<<MASK_END>>]]

Figure 17

[0029] An example of a decoding process in palette mode according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>19<<MASK_END>>]]

Figure 18

[0030] An example of an initialization process in palette mode is shown. [[ID=<<MASK_BEGIN>>20<<MASK_END>>]]

Figure 19

[0031] An example of an initialization process in palette mode according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>21<<MASK_END>>]]

Figure 20

[0032] An example of palette predictor update and run - length encoding of corresponding reuse flags is shown. [[ID=<<MASK_BEGIN>>22<<MASK_END>>]]

Figure 21

[0033] An example of palette predictor update and run - length encoding of corresponding reuse flags according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>23<<MASK_END>>]]

Figure 22

[0034] An example of palette predictor update and run - length encoding of corresponding reuse flags according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>24<<MASK_END>>]]

Figure 23

[0035] An example of a palette coding syntax according to some embodiments of the present disclosure is shown. [[ID=<<MASK_BEGIN>>25<<MASK_END>>]]

Figure 24

[0036] An example of palette coding semantics is shown. [[ID=<<MASK_BEGIN>>26<<MASK_END>>]]

Figure 25

[0037] An example of palette coding semantics according to some embodiments of the present disclosure is shown.

Best Mode for Carrying Out the Invention

[0012] Detailed Description

[0038] Here, exemplary embodiments will be described in detail, and these examples are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise specified, the same numbers in different drawings represent the same or similar elements. The implementation forms described in the following description of the exemplary embodiments do not represent all implementation forms that conform to the present invention. Rather, they are merely examples of devices and methods that conform to the aspects related to the present invention enumerated in the appended claims. Specific aspects of the present disclosure will be described in more detail below. If the terms and definitions described herein conflict with the terms and / or definitions incorporated by reference, the description herein shall prevail.

[0013]

[0039] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC) (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor standard, the High Efficiency Video Coding (HEVC) (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0014]

[0040] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has developed technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. Since the coding technology was incorporated into JEM, JEM has achieved significantly improved coding performance compared to HEVC.

[0015]

[0041] The VVC standard has been recently developed and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system as those used in the latest video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, and so on.

[0016]

[0042] Video is a set of still pictures (or "frames") arranged in chronological order to store visual information. To capture and store those pictures in chronological order, a video capture device (e.g., a camera) can be used, and to display such pictures in chronological order, a video playback device (e.g., a TV, computer, smartphone, tablet computer, video player, or any end-user terminal with a display function) can be used. Also, in some applications, for surveillance, meetings, or live broadcasts, etc., the video capture device can transmit the captured video to a video playback device (e.g., a computer with a monitor) in real time.

[0017]

[0043] To reduce the memory space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is generally called an "encoder", and the module for decompression is generally called a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented as any of various suitable hardware, software, or combinations thereof. For example, hardware implementation forms of the encoder and decoder can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementation forms of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed on a computer-readable medium. The compression and decompression of video can be implemented by various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard. In this case, the codec can be called a "transcoder".

[0018]

[0044] The video encoding process can identify and retain useful information that can be used to reconstruct the picture and ignore information that is not important for reconstruction. If the ignored unimportant information cannot be fully reconstructed, such an encoding process can be called "irreversible". Otherwise, the encoding process can be called "reversible". Most encoding processes are irreversible, which is a trade-off for reducing the required memory space and transmission bandwidth.

[0019]

[0045] The useful information of the symbolized picture (referred to as the "current picture") includes changes with respect to a reference picture (e.g., a previously symbolized and reconstructed picture). Such changes can include changes in pixel position, luminance, or color, among which the position change is the most important. The position change of the pixel group representing an object can reflect the movement of the object between the reference picture and the current picture.

[0020]

[0046] A picture that is coded without referring to another picture (i.e., the picture is its own reference picture) is called an "I picture". When some or all of the blocks in a picture (e.g., blocks generally referring to parts of a video picture) are predicted using intra prediction or inter prediction with one reference picture (e.g., uni - directional prediction), the picture is called a "P picture". When at least one block in a picture is predicted using two reference pictures (e.g., bi - directional prediction), the picture is called a "B picture".

[0021]

[0047] FIG. 1 shows an example structure of a video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live video or a captured and archived video. The video 100 can be a real - world video, a video generated by a computer (e.g., a computer game video), or a combination thereof (e.g., a real - world video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., a video file stored in a storage device), or a video supply interface (e.g., a video broadcast transceiver) for receiving videos from a video content provider.

[0022]

[0048] As shown in FIG. 1, the video sequence 100 can include a series of pictures temporally arranged along a timeline, including pictures 102, 104, 106, and 108. Pictures 102 to 106 are consecutive, and there are more pictures between picture 106 and picture 108. In FIG. 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and as indicated by the arrow, its reference picture is picture 102. Picture 106 is a B picture, and as indicated by the arrow, its reference pictures are pictures 104 and 108. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately before or after that picture. For example, the reference picture of picture 104 can be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102 to 106 are merely examples, and the present disclosure is not limited to the examples shown in FIG. 1 for the embodiments of the reference pictures.

[0023]

[0049] Typically, a video codec does not encode or decode an entire picture at once because the calculation of such a task is complex. More precisely, a video codec can divide a picture into basic segments and encode or decode the picture segment by segment. Such a basic segment is called a basic processing unit (BPU) in the present disclosure. For example, the structure 110 in FIG. 1 shows an example of the structure of a picture (for example, any one of pictures 102 to 108) of the video sequence 100. In the structure 110, the picture is divided into 4×4 basic processing units, and its boundaries are indicated by dashed lines. In some embodiments, the basic processing unit can be called a "macroblock" in some video coding standards (for example, the MPEG family, H.261, H.263, or H.264 / AVC), or a "coding tree unit" (CTU) in some other video coding standards (for example, H.265 / HEVC or H.266 / VVC). The basic processing unit can have variable sizes such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, etc., or pixels of any arbitrary shape and size within the picture. The size and shape of the basic processing unit can be selected for each picture based on the balance between the coding efficiency and the detail maintained within the basic processing unit.

[0024]

[0050] The basic processing unit can be a logical unit that may include various types of video data groups stored in a computer memory (for example, within a video frame buffer). For example, the basic processing unit of a color picture can include a luma component (Y) representing achromatic luminance information, one or more chroma components (for example, Cb and Cr) representing color information, and related syntax elements that may have basic processing units of the same size for the luma and chroma components. In some video coding standards (for example, H.265 / HEVC or H.266 / VVC), the luma and chroma components can be referred to as "coding tree blocks" ("CTB:Coding Tree Block"). Any operation performed on the basic processing unit can be repeatedly performed on each of its luma and chroma components.

[0025]

[0051] The encoding of the video has multiple operation stages, examples of which are shown in FIGS. 2A-2B and FIGS. 3A-3B. At each stage, since the size of the basic processing unit may still be too large to be processed, in the present disclosure, it can be further divided into segments called "basic processing subunits". In some embodiments, the basic processing subunit may be called a "block" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or may be called a "coding unit" ("CU: Coding Unit") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit can have the same size as or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit and can include a group of various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in a computer memory (e.g., within a video frame buffer). Any operation performed on the basic processing subunit can be repeatedly performed on each of its luma and chroma components. Note that such division can be performed at further levels according to processing requirements. Also note that the various stages can divide the basic processing unit in various ways.

[0026]

[0052] For example, in the mode determination stage (an example of which is shown in FIG. 2B), the encoder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, but such a determination may be made when the basic processing unit is too large. The encoder can divide the basic processing unit into a plurality of basic processing subunits (e.g., CUs as in H.265 / HEVC or H.266 / VVC) and can determine the type of prediction for each individual basic processing subunit.

[0027]

[0053] In another example, during the prediction stage (an example of which is shown in FIGS. 2A-2B), the coder can perform a prediction operation at the level of a basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform the prediction operation at that level.

[0028]

[0054] In another example, during the transform stage (an example of which is shown in FIGS. 2A-2B), the coder can perform a transform operation on a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. It should be noted that the same basic processing subunit division method may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0029]

[0055] In the structure 110 of FIG. 1, the basic processing unit 112 is further divided into 3×3 basic processing subunits, and its boundaries are indicated by dotted lines. Different basic processing units of the same picture can be divided into basic processing subunits in different ways.

[0030]

[0056] In some implementations, to provide parallel processing and error resilience capabilities for video encoding and decoding, a picture can be divided into regions for processing, such that the encoding or decoding process can be made independent of the information in any other region of the picture for a given region of the picture. In other words, each region of the picture can be processed independently. By doing so, the codec can process different regions of the picture in parallel, increasing the coding efficiency. Also, if the data of a region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the damaged or lost data, providing an error resilience function. In some video coding standards, a picture can be divided into various types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions, namely, "slices" and "tiles". It should also be noted that the various pictures of video sequence 100 may have various partitioning schemes for dividing the picture into regions.

[0031]

[0057] For example, in FIG. 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown by solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in FIG. 1 are merely examples, and the present disclosure does not limit its embodiments.

[0032]

[0058] FIG. 2A shows a schematic diagram of an example of an encoding process 200A that conforms to an embodiment of the present disclosure. For example, the encoding process 200A can be executed by an encoder. As shown in FIG. 2A, the encoder can encode the video sequence 202 into the video bitstream 228 according to the process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 can include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can execute the process 200A at the level of the basic processing unit for each original picture of the video sequence 202. For example, the encoder can execute the process 200A repeatedly, in which case the encoder can encode the basic processing unit in one iteration of the process 200A. In some embodiments, the encoder can execute the process 200A in parallel for the regions (e.g., regions 114-118) of each original picture of the video sequence 202.

[0033]

[0059] In FIG. 2A, the encoder can supply the basic processing unit of the original picture of video sequence 202 (referred to as the "original BPU") to prediction stage 204 to generate prediction data 206 and predicted BPU 208. The encoder can subtract the predicted BPU 208 from the original BPU to generate residual BPU 210. The encoder can supply the residual BPU 210 to transform stage 212 and quantization stage 214 to generate quantized transform coefficients 216. The encoder can supply the prediction data 206 and the quantized transform coefficients 216 to binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226 and 228 can be referred to as the "forward path". During process 200A, after quantization stage 214, the encoder can supply the quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222 and 224 of process 200A can be referred to as the "reconstruction path". The reconstruction path can be used by both the encoder and the decoder to ensure that the same reference data is used for prediction.

[0034]

[0060] The encoder can repeatedly execute process 200A to encode each original BPU of the original picture (in the forward path) and generate prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder can proceed to encode the next picture in video sequence 202.

[0035]

[0061] Referring to process 200A, the coder can receive video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive (s)" can refer to receiving, inputting, obtaining, extracting, acquiring, reading, accessing, or any act in any way for inputting data.

[0036]

[0062] In the prediction stage 204, in the current iteration, the coder can receive the original BPU and prediction criterion 224, and perform a prediction operation to generate prediction data 206 and predicted BPU 208. Prediction criterion 224 can be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction stage 204 is to reduce the redundancy of information by extracting prediction data 206, and prediction data 206 can be used to reconstruct the original BPU by extracting it as predicted BPU 208 from prediction data 206 and prediction criterion 224.

[0037]

[0063] Ideally, predicted BPU 208 can be the same as the original BPU. However, due to the non-ideal prediction operation and reconstruction operation, predicted BPU 208 generally differs slightly from the original BPU. To record such a difference, after generating predicted BPU 208, the coder can subtract it from the original BPU to generate residual BPU 210. For example, the coder can subtract the pixel value (e.g., grayscale value or RGB value) of predicted BPU 208 from the corresponding pixel value of the original BPU. Each pixel of residual BPU 210 can have a residual value as a result of such subtraction between the corresponding pixel of the original BPU and predicted BPU 208. Compared with the original BPU, prediction data 206 and residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. In this way, the original BPU is compressed.

[0038]

[0064] To further compress the residual BPU 210, in the transformation stage 212, the coder can reduce the spatial redundancy of the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", where each basis pattern is related to a "transformation coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent the varying frequency (e.g., luminance varying frequency) component of the residual BPU 210. None of the basis patterns can be reproduced from any combination (e.g., linear combination) of the other basis patterns. In other words, the decomposition can decompose the variations of the residual BPU 210 into the frequency domain. Such a decomposition is similar to the discrete Fourier transform of a function, the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients related to the basis functions.

[0039]

[0065] If the transformation algorithm is different, the basis patterns used may also be different. In the transformation stage 212, various transformation algorithms can be used, such as the discrete cosine transform, the discrete sine transform, etc. The transformation in the transformation stage 212 can be performed inversely. That is, the coder can restore the residual BPU 210 by the inverse operation of the transformation (referred to as "inverse transformation"). For example, to restore the pixels of the residual BPU 210, the inverse transformation can multiply the values of the corresponding pixels of the basis patterns by their respective associated coefficients and add the products to produce a weighted sum. In a video coding standard, both the coder and the decoder can use the same transformation algorithm (and thus the same basis patterns). Therefore, the coder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basis patterns from the coder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits and can be used to reconstruct the residual BPU 210 without significant quality degradation. In this way, the residual BPU 210 is further compressed.

[0040]

[0066] The coder can further compress the transform coefficients in the quantization stage 214. In the transformation process, various basis patterns can represent various fluctuation frequencies (e.g., luminance fluctuation frequencies). Since the human eye can generally recognize low-frequency fluctuations better, the coder can ignore the information of high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the coder can generate the quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as the "quantization scale factor") and rounding the quotient to the nearest integer. After such an operation, some transform coefficients of the high-frequency basis pattern can be converted to zero, and the transform coefficients of the low-frequency basis pattern can be converted to smaller integers. The coder can ignore the quantized transform coefficients 216 with zero values, thereby further compressing the transform coefficients. The quantization process can also be performed in reverse. In that case, the quantized transform coefficients 216 can be reconstructed into transform coefficients by the inverse operation of quantization (referred to as "inverse quantization").

[0041]

[0067] Since the coder ignores the remainder of such division in the rounding operation, the quantization stage 214 can be irreversible. Usually, the quantization stage 214 may be the cause of the largest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain various levels of information loss, the coder can use quantization parameters of various values or any other parameters of the quantization process.

[0042]

[0068] In the binary encoding stage 226, the coder can encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy encoding, variable-length encoding, arithmetic encoding, Huffman encoding, context-adaptive binary arithmetic encoding, or any other reversible or irreversible compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the coder can encode other information in the binary encoding stage 226, such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), the coder control parameters (e.g., bitrate control parameters), and so on. The coder can generate the video bitstream 228 using the output data of the binary encoding stage 226. In some embodiments, the video bitstream 228 can be further packetized for network transmission.

[0043]

[0069] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the coder can perform inverse quantization on the quantized transform coefficients 216 to generate the reconstructed transform coefficients. In the inverse transform stage 220, the coder can generate the reconstructed residual BPU 222 based on the reconstructed transform coefficients. The coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the prediction reference 224 used in the next iteration of process 200A.

[0044]

[0070] Note that it is possible to encode the video sequence 202 using other variants of the process 200A. In some embodiments, each step of the process 200A can be executed in a different order by an encoder. In some embodiments, one or more steps of the process 200A can be combined into a single step. In some embodiments, a single step of the process 200A can be divided into multiple steps. For example, the transform step 212 and the quantization step 214 can be combined into a single step. In some embodiments, the process 200A can include additional steps. In some embodiments, the process 200A can omit one or more steps of FIG. 2A.

[0045]

[0071] FIG. 2B shows a schematic diagram of another example of an encoding process 200B that conforms to an embodiment of the present disclosure. The process 200B can be modified from the process 200A. For example, the process 200B can be used by an encoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared with the process 200A, the forward path of the process 200B additionally includes a mode determination step 230, and divides the prediction step 204 into a spatial prediction step 2042 and a temporal prediction step 2044. The reconstruction path of the process 200B additionally includes a loop filter step 232 and a buffer 234.

[0046]

[0072] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can predict the current BPU using pixels from one or more adjacent BPUs that have already been coded within the same picture. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can predict the current BPU using regions from one or more pictures that have already been coded. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the temporal redundancy inherent in a picture.

[0047]

[0073] Referring to process 200B, in the forward path, the coder performs prediction operations at the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, at the spatial prediction stage 2042, the coder can perform intra prediction. In the original BPU of the coded picture, the prediction reference 224 can include one or more adjacent BPUs that are coded (in the forward path) and reconstructed (in the reconstructed path) within the same picture. The coder can generate the predicted BPU 208 by extrapolating the adjacent BPUs. The extrapolation technique can include, for example, linear extrapolation or linear interpolation, polynomial extrapolation or polynomial interpolation, etc. In some embodiments, the coder can perform extrapolation at the pixel level, such as by extrapolating the values of the corresponding pixels for each pixel of the predicted BPU 208. The adjacent BPUs used for extrapolation can be located relative to the original BPU in various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom - left, bottom - right, top - left, or top - right of the original BPU), or in any direction defined by the video coding standard being used. In intra prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the parameters of the extrapolation, the direction of the adjacent BPUs used relative to the original BPU, etc.

[0048]

[0074] In another example, in the temporal prediction stage 2044, the coder can perform inter prediction. In the original BPU of the current picture, the prediction reference 224 can include one or more pictures (referred to as "reference pictures") that are encoded (in the forward path) and reconstructed (in the reconstructed path). In some embodiments, the reference pictures can be encoded and reconstructed for each BPU. For example, the coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same picture are generated, the coder can generate the picture reconstructed as a reference picture. The coder can perform an operation of "motion estimation" to search for a matching region in a certain range (referred to as a "search window") within the reference picture. The position of the search window within the reference picture can be determined based on the position of the original BPU within the current picture. For example, the search window can be centered at a position having the same coordinates as the original BPU within the current picture within the reference picture and can be expanded over a predetermined distance. When the coder identifies a region similar to the original BPU within the search window (for example, by using a pixel recursive algorithm, a block matching algorithm, etc.), the coder can determine such a region as a matching region. The matching region can have dimensions different from those of the original BPU (for example, smaller, equal, larger or different in shape). Since the reference picture and the current picture are temporally separated within the timeline (as shown in FIG. 1, for example), it can be considered that the matching region "moves" to the position of the original BPU as time passes. The coder can record such a direction and distance of motion as a "motion vector". When multiple reference pictures are used (such as picture 106 in FIG. 1, for example), the coder can search for a matching region for each reference picture and obtain its associated motion vector. In some embodiments, the coder can assign weights to the pixel values of the matching regions of each matching reference picture.

[0049]

[0075] Using motion estimation, various types of motion such as, for example, translation, rotation, scaling, etc. can be identified. In inter prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the matching region, the motion vector related to the matching region, the number of reference pictures, the weights related to the reference pictures, etc.

[0050]

[0076] To generate the predicted BPU 208, the encoder can perform an operation of "motion compensation". Using motion compensation, the predicted BPU 208 can be reconstructed based on the prediction data 206 (e.g., motion vector) and the prediction reference 224. For example, the encoder can move the matching region of the reference picture according to the motion vector, in which case the encoder can predict the original BPU of the current picture. (e.g., like picture 106 in FIG. 1) When multiple reference pictures are used, the encoder can move the matching regions of the reference pictures according to their respective motion vectors and average the pixel values of the matching regions. In some embodiments, when the encoder assigns weights to the pixel values of the matching regions of each matching reference picture, the encoder can add the weighted sum of the pixel values of the moved matching regions.

[0051]

[0077] In some embodiments, inter prediction can be one-way or two-way. One-way inter prediction can use one or more reference pictures that are in the same temporal direction with respect to the current picture. For example, picture 104 in FIG. 1 is a one-way inter-predicted picture, and the reference picture (e.g., picture 102) precedes picture 104. Two-way inter prediction can use one or more reference pictures that are in both temporal directions with respect to the current picture. For example, picture 106 in FIG. 1 is a two-way inter-predicted picture, and the reference pictures (e.g., pictures 104 and 108) are in both temporal directions with respect to picture 104.

[0052]

[0078] Continuing to refer to the forward path of process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, at the mode determination stage 230, the coder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the coder can perform rate-distortion optimization techniques, in which case the coder can select a prediction mode according to the bitrate of the candidate prediction modes and the distortion of the reconstructed reference pictures under the candidate prediction modes to minimize the value of the cost function. According to the selected prediction mode, the coder can generate the corresponding predicted BPU 208 and the predicted data 206.

[0053]

[0079] In the reconstruction path of process 200B, when the intra prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current BPU that is encoded and reconstructed within the current picture), the coder can directly supply the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). The coder can supply the prediction reference 224 to the loop filter stage 232, where the coder can apply a loop filter to the prediction reference 224 to reduce or remove the distortion (e.g., blocking artifacts) caused during the coding of the prediction reference 224. The coder can apply various loop filter techniques at the loop filter stage 232, such as, for example, deblocking, sample adaptive offset, adaptive loop filter, etc. The loop filtered reference picture can be stored in buffer 234 (or "decoded picture buffer") for later use (e.g., for use as an inter prediction reference picture for subsequent pictures of video sequence 202). The coder can store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the coder can encode loop filter parameters (e.g., loop filter strength) at the binary coding stage 226 along with the quantized transform coefficients 216, prediction data 206, and other information.

[0054]

[0080] FIG. 3A shows a schematic diagram of an example of a decoding process 300A that conforms to an embodiment of the present disclosure. Process 300A can be a decompression process corresponding to the compression process 200A in FIG. 2A. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 in FIGS. 2A-2B), generally, the video stream 304 is not identical to the video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, the decoder can execute process 300A at the level of a basic processing unit (BPU) for each picture encoded within the video bitstream 228. For example, the decoder can execute process 300A repeatedly, in which case the decoder can decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder can execute process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded within the video bitstream 228.

[0055]

[0081] In FIG. 3A, the decoder can supply a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode that portion into prediction data 206 and quantized transform coefficients 216. The decoder can supply the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can supply the prediction data 206 to the prediction stage 204 to generate a predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can supply the prediction reference 224 to the prediction stage 204 for performing a prediction operation in the next iteration of process 300A.

[0056]

[0082] The decoder can repeatedly execute process 300A to decode each encoded BPU of the encoded picture and generate a prediction reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all the encoded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.

[0057]

[0083] In binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context adaptive binary arithmetic coding, or any other reversible compression algorithm) used by the encoder. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as, for example, the prediction mode, the parameters of the prediction operation, the type of transform, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bit rate control parameters), and so on. In some embodiments, when the video bitstream 228 is transmitted in packet units over a network, the decoder can depacketize the video bitstream 228 and then supply it to the binary decoding stage 302.

[0058]

[0084] FIG. 3B shows a schematic diagram of another example of a decoding process 300B that conforms to an embodiment of the present disclosure. The process 300B can be modified from the process 300A. For example, the process 300B can be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared with the process 300A, the process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.

[0059]

[0085] In process 300B, in the encoded basic processing unit (referred to as the "current BPU") of the encoded picture being decoded (referred to as the "current picture"), the predicted data 206 decoded from the binary decoding stage 302 by the decoder can include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra prediction is used by the encoder to encode the current BPU, the predicted data 206 can include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation can include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as a reference, the sizes of the adjacent BPUs, extrapolation parameters, the directions of the adjacent BPUs with respect to the original BPU, etc. In another example, if inter prediction is used by the encoder to encode the current BPU, the predicted data 206 can include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. The parameters of the inter prediction operation can include, for example, the number of reference pictures related to the current BPU, the weights respectively related to the reference pictures, the positions (e.g., coordinates) of one or more matching regions in each reference picture, one or more motion vectors respectively related to the matching regions, etc.

[0060]

[0086] Based on the prediction mode indicator, the decoder can determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of the execution of such spatial or temporal prediction are described in Figure 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder can generate the predicted BPU 208. As described in Figure 3A, the decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate the prediction reference 224.

[0061]

[0087] In process 300B, the decoder can supply the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, when the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly supply the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). When the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the decoder can supply the prediction reference 224 to the loop filter stage 232 to reduce or remove distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop-filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as an inter prediction reference picture for the subsequent encoded picture of the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data can further include loop filter parameters (e.g., the strength of the loop filter). In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter prediction is used to encode the current BPU, the prediction data includes loop filter parameters.

[0062]

[0088] FIG. 4 is a block diagram of an example of an apparatus 400 for encoding or decoding video that conforms to an embodiment of the present disclosure. As shown in FIG. 4, the apparatus 400 can include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a dedicated device for video encoding or decoding. The processor 402 can be any type of circuitry capable of handling or processing information. For example, the processor 402 can include any combination of any number of central processing units (i.e., "CPUs"), graphics processing units (i.e., "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), system on chips (SoCs), application specific integrated circuits (ASICs), etc. In some embodiments, the processor 402 can also be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, the processor 402 can include a plurality of processors including a processor 402a, a processor 402b, and a processor 402n.

[0063]

[0089] Device 400 can also include a memory 404 configured to store a set of data (e.g., instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data can include program instructions (e.g., program instructions for implementing steps in processes 200A, 200B, 300A, or 300B) and processing data (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and the processing data (e.g., via bus 410) and execute the program instructions to perform operations or handling on the processing data. The memory 404 can include a high-speed random access memory device or a non-volatile memory device. In some embodiments, the memory 404 can include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (registered trademark) (CF) cards, etc. The memory 404 can also be a group of memories grouped as a single logical component (not shown in FIG. 4).

[0064]

[0090] The bus 410 can be a communication device that transfers data between components within the device 400, such as an internal bus (e.g., a CPU memory bus), an external bus (e.g., a USB (Universal Serial Bus) port, a PCI (Peripheral Component Interconnect) Express port), etc.

[0065]

[0091] To make the description understandable without causing ambiguity, in the present disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuits can be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Additionally, the data processing circuits can be a single independent module, or can be fully or partially integrated with any other component of the device 400.

[0066]

[0092] The device 400 can further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 can include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth (registered trademark) adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.

[0067]

[0093] In some embodiments, optionally, the device 400 can further include a peripheral device interface 408 to provide connection to one or more peripheral devices. As shown in FIG. 4, the peripheral devices can include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), etc.

[0068]

[0094] Note that the video codec (e.g., the codec that executes processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. In another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).

[0069]

[0095] FIG. 5 illustrates a CU coded in palette mode. In VVC draft 8, the palette mode can be used in monochrome, 4:2:0, 4:2:2, and 4:4:4 color formats. When the palette mode is enabled and the CU size is 64x64 or less and there are more than 16 samples indicating whether the palette mode is used, a flag is sent at the CU level. When the palette mode is used for coding of the (current) CU 500, the sample value at each position within the CU is represented by a small set of representative color values. This set is called the palette 510. For sample positions having values approximating the palette colors 501, 502, 503, the corresponding palette index is signaled. It is also possible to specify color values outside the palette by signaling an escape index 504. Next, for all positions within the CU that use the escape color index, the (quantized) color component values are signaled for each of those respective positions.

[0070]

[0096] For palette coding, a palette predictor is maintained. FIG. 6 shows an exemplary process for updating the palette predictor after each coding unit 600 is coded. The predictor is initialized to 0 (i.e., empty) at the start of each slice for non-wavefront cases and at the start of each CTU row for wavefront cases. For each entry of the palette predictor, a reuse flag is signaled to indicate whether to include it in the current palette of the current CU. The reuse flag is sent using zero run-length coding. Thereafter, the number of new palette entries and the component values of the new palette entries are signaled. After coding the palette-coded CU, the palette predictor will be updated using the current palette, and entries from the previous palette predictor not reused in the current palette are added at the end of the new palette predictor until the maximum allowed size is reached.

[0071]

[0097] An escape flag is signaled for each CU to indicate whether an escape symbol exists within the current CU. If the escape symbol exists, the palette table is extended by the last index assigned to be the escape symbol. As shown in the example of FIG. 5, the palette indices of the samples within the CU form a palette index map. The index map is coded using a horizontal cross scan or a vertical cross scan. The scan order is explicitly signaled in the bitstream using the palette_transpose_flag. The palette index map is coded using an index run mode or an index copy mode.

[0072]

[0098] In VVC Draft 8, the deblocking filter process includes defining the block boundary, deriving the boundary filter strength based on the coding modes of two adjacent blocks along the defined block boundary, deriving the number of samples to be filtered, and applying the deblocking filter to the samples. When an edge is at the boundary of a coding unit, coding sub-block unit, or transform unit, that edge is defined as the block boundary. Next, the boundary filter strength is calculated based on the coding modes of two adjacent blocks according to the following six rules. (1) When both of the two coded blocks are coded in the BDPCM mode, the boundary filter strength is set to 0. (2) Otherwise, when one of the coded blocks is coded in the intra mode, the boundary filter strength is set to 2. (3) Otherwise, when one of the coded blocks is coded in the CIIP mode, the boundary filter strength is set to 2. (4) Otherwise, when one of the coded blocks contains one or more non-zero coefficient levels, the boundary filter strength is set to 1. (5) Otherwise, when one of the blocks is coded in the IBC mode and the other block is coded in the inter mode, the boundary filter strength is set to 1. (6) Otherwise (when both of the two blocks are coded in the IBC mode or the inter mode), the boundary filter strength is derived using the reference pictures and motion vectors of the two blocks.

[0073]

[0099] In VVC Draft 8, the process of calculating the boundary filter strength is described in more detail. Specifically, in VVC Draft 8, eight sequential and comprehensive scenarios are given. In Scenario 1, when cIdx is equal to 0 and both samples p0 and q0 are in a coded block where intra_bdpcm_luma_flag is equal to 1, bS[xDi][yDj] is set equal to 0. Otherwise, in Scenario 2, when cIdx is greater than 0 and both samples p0 and q0 are in a coded block where intra_bdpcm_chroma_flag is equal to 1, bS[xDi][yDj] is set equal to 0. Otherwise, in Scenario 3, when sample p0 or q0 is in a coded block of a coded unit coded using an intra prediction mode, bS[xDi][yDj] is set equal to 2. Otherwise, in Scenario 4, when the block edge is also the edge of the coded block and sample p0 or q0 is in a coded block where ciip_flag is equal to 1, bS[xDi][yDj] is set equal to 2. Otherwise, in Scenario 5, when the block edge is also the edge of the transform block and sample p0 or q0 is in a transform block containing one or more non-zero transform coefficient levels, bS[xDi][yDj] is set equal to 1. Otherwise, in Scenario 6, when the prediction mode of the coded sub-block containing sample p0 is different from the prediction mode of the coded sub-block containing sample q0 (i.e., one of the coded sub-blocks is coded in the IBC prediction mode and the other is coded in the inter prediction mode), bS[xDi][yDj] is set equal to 1.

[0074]

[0100] Otherwise, in Scenario 7, when cIdx is equal to 0, edgeFlags[xDi][yDj] is equal to 2, and one or more of the following conditions are true, bS[xDi][yDj] is set equal to 1.

[0075]

[0101] Condition (1): Both the coded sub-block containing sample p0 and the coded sub-block containing sample q0 are coded in the IBC prediction mode, and the absolute difference between the horizontal or vertical components of the block vectors used for prediction of the two coded sub-blocks is 8 or more in 1 / 16 luma sample units.

[0076]

[0102] Condition (2): For the prediction of the coded sub-block containing sample p0, a different reference picture or a different number of motion vectors is used compared to the prediction of the coded sub-block containing sample q0. Regarding Condition (2), it should be noted that the determination of whether the reference pictures used for the two coded sub-blocks are the same or different is based only on which picture is being referenced, regardless of whether the prediction is formed using an index to reference picture list 0 or an index to reference picture list 1, and regardless of whether the index positions within the reference picture list are different. Also, regarding Condition (2), it should be noted that the number of motion vectors used for prediction of the coded sub-block whose top-left sample covers (xSb, ySb) is equal to PredFlagL0[xSb][ySb] + PredFlagL1[xSb][ySb].

[0077]

[0103] Condition (3): One motion vector is used to predict the coded sub-block containing sample p0, and one motion vector is used to predict the coded sub-block containing sample q0, and the absolute difference between the horizontal or vertical components of the motion vectors used is 8 or more in 1 / 16 luma sample units.

[0078]

[0104] Condition (4): Two motion vectors and two different reference pictures are used to predict the coded sub-block containing sample p0, and the two motion vectors of the same two reference pictures are used to predict the coded sub-block containing sample q0. The absolute difference between the horizontal or vertical components of the two motion vectors used in the prediction of the two coded sub-blocks of the same reference picture is 8 or more in 1 / 16 luma sample units.

[0079]

[0105] Condition (5): Two motion vectors of the same reference picture are used to predict the coded sub-block containing sample p0, and the two motion vectors of the same reference picture are used to predict the coded sub-block containing sample q0. The following two conditions are both true. Condition (5.1): The absolute difference between the horizontal or vertical components of the motion vectors in list 0 used in the prediction of the two coded sub-blocks is 8 or more in 1 / 16 luma sample units, or the absolute difference between the horizontal or vertical components of the motion vectors in list 1 used in the prediction of the two coded sub-blocks is 8 or more in 1 / 16 luma sample units. Condition (5.2): The absolute difference between the horizontal or vertical components of the motion vector in list 0 used in the prediction of the coded sub-block containing sample p0 and the motion vector in list 1 used in the prediction of the coded sub-block containing sample q0 is 8 or more in 1 / 16 luma sample units, or the absolute difference between the horizontal or vertical components of the motion vector in list 1 used in the prediction of the coded sub-block containing sample p0 and the motion vector in list 0 used in the prediction of the coded sub-block containing sample q0 is 8 or more in 1 / 16 luma sample units.

[0080]

[0106] Otherwise, i.e., if none of the above seven scenarios are satisfied, in scenario 8, the variable bS[xDi][yDj] is set equal to 0. After deriving the boundary filter strength, the number of samples to be filtered is derived and the deblocking filter is applied to the samples. Note that when a block is coded in palette mode, the number of samples is set equal to 0. This means that the deblocking filter is not applied to blocks coded in palette mode.

[0081]

[0107] As described above, the structure of the current palette of a block includes two parts. First, the entries in the current palette can be predicted from the palette predictor. For each entry of the palette predictor, a reuse flag is signaled to indicate whether this entry is included in the current palette. Second, the component values of the current palette entries can be signaled directly. After obtaining the current palette, the palette predictor is updated using the current palette. FIG. 7 shows a part of section 7.3.10.6 (“Palette Coding Syntax”) of VVC draft 8. When parsing the syntax of a palette-coded block, the reuse flag (i.e., “palette_predictor_run” 701 in FIG. 7) is decoded first, followed by the component values of the new palette entries (i.e., “num_signalled_palette_entries” 702 and “new_palette_entries” 703 in FIG. 7). In VVC draft 8, it is necessary to know the number of entries of the palette predictor (i.e., “PredictorPaletteSize[startComp]” 704 in FIG. 7) before parsing the reuse flag. This means that when two adjacent blocks are both coded in palette mode, the syntax of the second block cannot be parsed until the update process of the palette predictor of the first block is completed.

[0082]

[0108] However, in the conventional palette predictor update process, it is necessary to check the reuse flag for each entry. In a worse scenario, up to 63 checks may be required. These conventional designs may not be hardware-friendly for at least the following two reasons. First, context-based adaptive binary arithmetic coding (CABAC) syntax parsing has to wait until the palette predictor is fully updated. Usually, CABAC syntax parsing is the slowest module in hardware. This may reduce the processing capacity of CABAC. Second, implementing the palette predictor update process at the CABAC syntax parsing stage may impose a burden on the hardware.

[0083]

[0109] In addition, another problem with these conventional designs is that the boundary filter strength is not defined for the palette mode. When one of the adjacent blocks is coded in the palette mode and the other adjacent block is coded in the IBC mode or the inter mode, the boundary filter strength is not defined.

[0084]

[0110] Embodiments of the present disclosure provide implementations for addressing one or more of the above-described problems. These embodiments may improve the palette predictor update process and enhance the efficiency, speed, and resource consumption of the system implementing the above-described palette prediction process or a similar process.

[0085]

[0111] In some embodiments, the palette predictor update process is simplified and its complexity is reduced so that hardware resources can be made available for other purposes. FIG. 8 shows a flowchart of a palette predictor update process 800 according to an embodiment of the present disclosure. The method 800 can be executed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B) or by a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B), or can be executed by one or more software or hardware components of a device (e.g., device 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can execute the method 800. In some embodiments, the method 800 can be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions such as program code executed by a computer (e.g., device 400 of FIG. 4). Referring to FIG. 8, the method 800 may include the following steps 802-804.

[0086]

[0112] In step 802, when the palette predictor is updated, all palette entries of the current palette are added as a first set of entries before the new palette predictor. In step 804, all palette entries from the previous palette predictor are added to the end of the new palette predictor, regardless of whether the entry is reused in the current palette as a second set of entries, and the second set of entries is after the first set of entries.

[0087]

[0113] For example, FIG. 9 presents a simplified palette predictor update process that matches the process described in FIG. 8. The current palette is generated by processes 901 and 902. The current palette includes entries reused from the previous palette predictor and signaled new palette entries. In process 903, all palette entries of the current palette are added as a set of first entries of the new palette predictor (corresponding to step 802). In process 904, regardless of whether an entry is reused in the current palette, palette entries from the previous palette predictor are added as a set of second entries of the new palette predictor (corresponding to step 804).

[0088]

[0114] The benefit is that the size of the new palette predictor is calculated by adding the size of the current palette and the size of the previous palette predictor without inspecting the value of the reuse flag. The update process is much simpler than the conventional design of VVC draft 8 (as shown in FIG. 10), which may enhance the processing ability of CABAC.

[0089]

[0115] For example, FIG. 11 shows an example of a palette mode decoding process that matches an embodiment of the present disclosure. The differences including the deleted part 1101 between the conventional design of FIG. 10 and the disclosed design of FIG. 11 are emphasized in the text with a strikethrough.

[0090]

[0116] In some embodiments, although the palette predictor update process is simplified, the palette predictor may have two or more identical entries (i.e., redundancy of the palette predictor) because the reuse flag is not inspected. Even though it is improved, this may mean that the prediction effect of the palette predictor may decrease.

[0091]

[0117] Figure 12 shows a flowchart of another palette predictor update process 1200 according to an embodiment of the present disclosure. The method 1200 can be executed by an encoder (e.g., by process 200A of FIG. 2A or 200B of FIG. 2B) or by a decoder (e.g., by process 300A of FIG. 3A or 300B of FIG. 3B), or can be executed by one or more software or hardware components of a device (e.g., device 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can execute the method 1200. In some embodiments, the method 1200 can be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions such as program code executed by a computer (e.g., device 400 of FIG. 4). Referring to FIG. 12, the method 1200 may include the following steps 1202-1206.

[0092]

[0118] In this exemplary embodiment, in order to eliminate the redundancy of the palette predictor and simplify the palette predictor update process, the entries of the previous palette predictor between the first reuse entry and the last reuse entry are directly discarded. Only the entries from the first entry to the first reuse entry, and the entries from the last reuse entry to the last entry of the previous palette predictor are added to the new palette predictor. In summary, the palette predictor update process is modified as follows. In step 1202, each entry of the current palette is added to the new palette predictor as a set of first entries. In step 1204, each entry from the first entry to the first reuse entry of the previous palette predictor is added to the new palette predictor as a set of second entries after the set of first entries. In step 1206, each entry from the last reuse entry to the last entry of the previous palette predictor is added to the new palette predictor as a set of third entries after the set of second entries.

[0093]

[0119] For example, FIG. 13 presents a simplified palette predictor update process that matches the process described in FIG. 12. In process 1301, all palette entries of the current palette are added to the palette predictor as a first set of entries of the new palette predictor (corresponding to step 1202). In process 1302, each entry from the first entry to the first reuse entry of the previous palette predictor is added to the new palette predictor as a second set of entries after the first set of entries (corresponding to step 1204). In process 1303, each entry from the last reuse entry to the last entry of the previous palette predictor is added to the new palette predictor as a third set of entries after the second set of entries (corresponding to step 1206). In some embodiments, the third set of entries is after the first set of entries, and the second set of entries is after the third set of entries.

[0094]

[0120] Note that the first reuse entry and the last entry can be derived when parsing the reuse flag. There is no need to check the value of the reuse flag. The size of the new palette predictor is calculated by adding the size of the current palette, the size from 0 to the first reuse entry, and the size from the last reuse entry to the last entry. FIG. 14 shows an example of a palette coding syntax that matches an embodiment of the present disclosure. The changes including the added portion 1401 between the conventional design of FIG. 7 and the disclosed design of FIG. 14 are enclosed by a dashed line. FIG. 15 shows an example of a palette mode decoding process that matches an embodiment of the present disclosure. The changes including the deleted portion 1501 between the conventional design of FIG. 10 and the disclosed design of FIG. 15 are emphasized with text struck through, and the added portion 1502 is enclosed by a dashed line.

[0095]

[0121] In the previously disclosed embodiments, the palette predictor update process is simplified, but this process needs to be implemented in the CABAC syntax analysis stage of the hardware design. In some embodiments, the number of reuse flags is set to a fixed value. Therefore, CABAC can continue syntax analysis without waiting for the palette predictor update process. Furthermore, the palette predictor update process can be implemented in different pipeline stages other than the CABAC syntax analysis stage, which makes it possible to enhance the flexibility of the hardware design.

[0096]

[0122] FIG. 16 illustrates a decoder hardware design in palette mode. Data structure 1601 and data structure 1602 are exemplarily illustrated for CABAC and decoded palette pixels. The decoder hardware design 1603 includes a predictor update module 1631 and a CABAC syntax analysis module 1632, and the predictor update process and the CABAC syntax analysis process are in parallel. Therefore, CABAC can continue syntax analysis without waiting for the palette update process.

[0097]

[0123] In order to set the number of reuse flags to a fixed value for each coded block, in some embodiments, the size of the palette predictor is initialized to a predefined value at the beginning of each slice for non-wavefront cases and at the beginning of each CTU row for wavefront cases. The predefined value is set to the maximum size of the palette predictor. In one example, the number of reuse flags is 31 or 63 depending on the slice type and the dual-tree mode setting. When the slice type is an I slice and the dual-tree mode is on (referred to as case 1), the number of reuse flags is set to 31. Otherwise (when the slice type is a B slice / P slice, or when the slice type is an I slice and the dual-tree mode is off (referred to as case 2)), the number of reuse flags is set to 63. Numbers of reuse flags other than the two cases can be used, and it may be beneficial to maintain a relationship where the number of reuse flags in case 1 is twice that in case 2. In addition, when initializing the palette predictor, the value of each entry and each component is set to 0 or (1<<(sequence bit depth - 1)).

[0098]

[0124] For example, FIG. 17 shows an example of a palette mode decoding process that conforms to an embodiment of the present disclosure. The changes including the deleted portion 1701 between the conventional design of FIG. 10 and the disclosed design of FIG. 17 are emphasized with text struck through, and the added portion 1702 is surrounded by a dashed line. Similarly, FIG. 19 shows an example of an initialization process that conforms to an embodiment of the present disclosure. The changes including the deleted portion 1901 between the conventional design of FIG. 18 and the disclosed design of FIG. 19 are emphasized with text struck through, and the added portion 1902 is surrounded by a dashed line.

[0099]

[0125] In some embodiments, when the reuse flag is set to a fixed value, or when the embodiment of FIG. 9 is implemented, some redundancy may occur in the palette predictor. In some embodiments, this may not be a problem because those redundant entries are not selected for prediction on the encoder side. Examples of palette predictor updates are shown in FIGS. 20, 21, and 22. FIG. 20 shows an example according to the procedure specified in VVC Draft 8, FIG. 21 shows an example according to some implementations (referred to as the first embodiment) of the method proposed in FIG. 9, and FIG. 22 shows another example according to some implementations of the embodiment (referred to as the third embodiment) in which the reuse flag is set to a fixed value. As shown by the comparison of FIGS. 20, 21, and 22, assuming there are no newly signaled palette entries, the method proposed in the first embodiment (FIG. 21) can increase the number of bits for signaling the reuse flag 2101, while the method proposed in the third embodiment (FIG. 22) maintains the same number of bits for signaling the reuse flag 2201 as the method (FIG. 20) in the VVC Draft 8 design for signaling the reuse flag 2001.

[0100]

[0126] In some embodiments, the palette predictor is initially initialized to a fixed value. This means that there may be redundant entries in the palette predictor. As described above, redundant entries may not be a problem, but the design of this embodiment does not prevent the encoder from using those redundant entries. If the encoder accidentally selects one of these redundant entries, the coding performance of the palette predictor, and thus the performance of the palette mode, may also decrease. To prevent this case, in some embodiments, bitstream compliance is added when signaling the reuse flag. In some embodiments, bitstream compliance means that when signaling the reuse flag, the value of the size of the palette predictor is equal to the maximum size of the palette predictor. More specifically, range constraints are added to the binary value of the reuse flag.

[0101]

[0127] For example, FIG. 23 shows an example of a palette coding syntax that conforms to an embodiment of the present disclosure. The changes including the deleted portion 2301 between the conventional design of FIG. 7 and the disclosed design of FIG. 23 are highlighted in strikethrough text, and the added portion 2302 is enclosed by a dashed line. Similarly, FIG. 25 shows an example of a palette coding semantics that conforms to an embodiment of the present disclosure. The changes including the deleted portion 2501 between the conventional design of FIG. 24 and the disclosed design of FIG. 25 are highlighted in strikethrough text, and the added portion 2502 is enclosed by a dashed line.

[0102]

[0128] Further still, the present disclosure provides the following method to address the issues of the palette mode deblocking filter.

[0103]

[0129] In some embodiments, the palette mode is treated as a subset of the intra prediction mode. Thus, when one of the adjacent blocks is coded in the palette mode, the boundary filter strength is set equal to 2. More specifically, the scenario 3 is described as follows by changing the portion of the VVC draft 8 specification that details the process of calculating the boundary filter strength: "When sample p0 or q0 is in the coded block of a coding unit coded using the intra prediction mode or the palette mode, bS[xDi][yDj] is set equal to 2". The added language is underlined.

[0104]

[0130] In some embodiments, the palette mode is treated as an independent coding mode. When the coding modes of two adjacent blocks are different, the boundary filter strength is set equal to 1. More specifically, by changing the part of the VVC draft 8 specification that details the process of calculating the boundary filter strength, scenario 6 is described as follows: "When the prediction mode of the coded sub-block containing sample p0 is different from the prediction mode of the coded sub-block containing sample q0, bS[xDi][yDj] is set equal to 1". The language to be deleted is struck through.

[0105]

[0131] In some embodiments, the palette mode is treated as an independent coding mode. Similar to the BDPCM mode setting, when one of the coded blocks is in the palette mode, the boundary filter strength is set equal to 0. More specifically, by changing the part of the VVC draft 8 specification that details the process of calculating the boundary filter strength, a new scenario is inserted between scenario 3 and scenario 4 and is described as follows: "Otherwise, if the block edge is also the edge of the coded block and sample p0 or q0 is in a coded block where pred_mode_plt_flag is equal to 1, bS[xDi][yDj] is set equal to 0". Using this modified process, scenarios 4 to 8 are renumbered to scenarios 5 to 9, resulting in 9 scenarios. It should be noted that a block coded in the palette mode is equivalent to a pred_mode_plt_flag equal to 1.

[0106]

[0132] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions may be executed by a device (such as the disclosed encoder and decoder) for performing the above method. Common forms of non-transitory media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROM, and EPROM, FLASH (registered trademark)-EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and networked versions of these. The device can include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.

[0107]

[0133] Note that in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. Moreover, "comprising", "having", "containing", "including", and other similar forms of words are intended to be equivalent in meaning, and it is not intended that the singular or plural items following any one of these words be an exhaustive listing of such singular or plural items, nor is it intended to be limited to only the singular or plural items enumerated. It is intended to be non-limiting in the sense that it is not limited to only the singular or plural items enumerated.

[0108]

[0134] As used herein, unless otherwise specified, the term "or" includes all possible combinations, except when the combination is infeasible. For example, if it is said that a certain database may include A or B, then unless otherwise specified or infeasible, the database may include A, B, A and B. As a second example, if it is said that a certain database may include A, B, or C, then unless otherwise specified or infeasible, the database may include A, B, C, A and B, A and C, B and C, A and B and C.

[0109]

[0135] It will be understood that the above embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. When implemented by software, the software can be stored in the above computer-readable medium. When the software is executed by a processor, it can execute the disclosed method. The computing units and other functional units described in the present disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that a plurality of the above modules / units can be combined into one module / unit, and the above modules / units can also be further divided into a plurality of sub-modules / sub-units respectively.

[0110]

[0136] The following clauses can be used to further describe the embodiments. 1. A video data processing method including processing one or more coding units using one or more palette predictors, wherein a palette predictor among the one or more palette predictors adds all palette entries of the current palette as a first set of entries of the palette predictor, and Add the palette entries from the previous palette predictor as a set of second entries of the palette predictor, after the set of first entries of the palette predictor, regardless of the value of the reuse flag of the palette entries of the previous palette predictor. A video data processing method updated by 2. The method according to clause 1, wherein each palette entry of the first set and the second set of palette predictors includes a reuse flag. 3. The method according to clause 1, further comprising receiving a video frame for processing and generating one or more coding units of the video frame. The method according to clause 1, further comprising 4. The method according to any one of clauses 1 to 3, wherein each palette entry of the palette predictor has a corresponding reuse flag, and the number of reuse flags of the palette predictor is set to a fixed number for the corresponding coding unit. 5. A video data processing method including processing one or more coding units using one or more palette predictors, wherein the palette predictor among the one or more palette predictors adds all palette entries of the current palette as a set of first entries of the palette predictor, adds one or more palette entries of the previous palette predictor within a first range that starts from the first palette entry of the previous palette predictor and ends at the first palette entry of the previous palette predictor having a set of reuse flags, as a set of one or more second entries of the palette predictor to the palette predictor, and adds one or more palette entries of the previous palette predictor within a second range that starts from the last palette entry of the previous palette predictor having a set of reuse flags and ends at the last palette entry of the previous palette predictor, as a set of one or more third entries of the palette predictor to the palette predictor. A video data processing method updated by, where a second set of entries and a third set of entries are after a first set of entries. 6. The method according to clause 5, wherein each palette entry of the first set of entries, the second set of entries, and the third set of entries of the palette predictor includes a reuse flag. 7. Receiving a video frame for processing, Generating one or more coded units of the video frame, The method according to clause 5, further comprising. 8. Each palette entry of one or more palette predictors has a corresponding reuse flag, The method according to any one of clauses 5 to 7, wherein the number of reuse flags of each palette predictor is set to a fixed number for the corresponding coded unit. 9. A video data processing method, Receiving a video frame for processing, Generating one or more coded units of the video frame, Processing one or more coded units using one or more palette predictors having palette entries, Including, Each palette entry of one or more palette predictors has a corresponding reuse flag, A video data processing method, wherein the number of reuse flags of each palette predictor is set to a fixed number for the corresponding coded unit. 10. The palette predictor, Adding all palette entries of the current palette as a first set of entries of the palette predictor, and, Adding entries from a previous palette predictor not reused in the current palette as a second set of entries of the palette predictor, where the second set of entries is after the first set of entries, The method according to clause 9, updated by. 11. The method according to clause 9, wherein the fixed number is set based on a slice type and a dual-tree mode setting. 12. The method according to clause 9, wherein the size of one of the one or more palette predictors is initialized to a predefined value at the beginning of the slice in the case of non-wavefront. 13. The method according to clause 9, wherein the size of one of the one or more palette predictors is initialized to a predefined value at the beginning of the coding unit row in the case of wavefront. 14. The method according to clause 12 or 13, further comprising adding bitstream compliance that the value of the size of the palette predictor is equal to the maximum size of the palette predictor when signaling the reuse flag. 15. The method according to any one of clauses 12 to 14, wherein when initializing one or more palette predictors, the value of each entry and each component are set to 0 or (1 << (sequence bit depth - 1)). 16. The method according to any one of clauses 9 to 15, further comprising applying a range constraint to the binarized value of the reuse flag. 17. An apparatus for performing video data processing, a memory configured to store instructions, and a processor coupled to the memory, the processor executing instructions to cause the apparatus to process one or more coding units using one or more palette predictors is configured to be executed, wherein the palette predictor among the one or more palette predictors adds all palette entries of the current palette as a set of first entries of the palette predictor, and adds the palette entries from the previous palette predictor as a set of second entries of the palette predictor, which is a set of second entries after the set of first entries, regardless of the value of the reuse flag of the palette entries of the previous palette predictor, is updated by, the apparatus. 18. The apparatus according to clause 17, wherein each palette entry of the first set of entries and the second set of entries of the palette predictor includes a reuse flag. 19. The processor executes instructions to cause the apparatus to receive a video frame for processing, and generate one or more coded units of the video frame The apparatus according to clause 17, further configured to perform. 20. Each palette entry of the palette predictor has a corresponding reuse flag, The number of reuse flags of the palette predictor is set to a fixed number for the corresponding coded unit. The apparatus according to any one of clauses 17 to 19. 21. An apparatus for performing video data processing, a memory configured to store instructions, and a processor coupled to the memory, the processor executing instructions to cause the apparatus to process one or more coded units using one or more palette predictors configured to perform, wherein the palette predictor among the one or more palette predictors adds all palette entries of the current palette as a first set of entries of the palette predictor; adds one or more palette entries of the previous palette predictor as a second set of one or more entries of the palette predictor within a first range starting from the first palette entry of the previous palette predictor and ending at the first palette entry of the previous palette predictor having a set of reuse flags, and adds one or more palette entries of the previous palette predictor as a third set of one or more entries of the palette predictor within a second range starting from the last palette entry of the previous palette predictor having a set of reuse flags and ending at the last palette entry of the previous palette predictor. An apparatus updated by 22. The apparatus according to clause 21, wherein each palette entry of the first set of palette predictors and the second set of palette predictors includes a reuse flag. 23. The processor executes instructions to cause the apparatus to receive a video frame for processing, and generate one or more coded units of the video frame The apparatus according to clause 21, further configured to perform. 24. Each palette entry of the palette predictor has a corresponding reuse flag, The number of reuse flags of the palette predictor is set to a fixed number for the corresponding coded unit. The apparatus according to any one of clauses 21 to 23. 25. An apparatus for performing video data processing, a memory configured to store instructions, a processor coupled to the memory, the processor executing instructions to cause the apparatus to receive a video frame for processing, generate one or more coded units of the video frame, and process one or more coded units using one or more palette predictors having palette entries configured to perform, each palette entry of the one or more palette predictors has a corresponding reuse flag, The number of reuse flags of each palette predictor is set to a fixed number for the corresponding coded unit. An apparatus. 26. The processor executes instructions to cause the apparatus to add all palette entries of the current palette as the first set of palette predictor entries, and Add an entry from a palette predictor before being reused in the current palette as a set of second entries of the palette predictor, where the set of second entries is after the set of first entries. The apparatus according to clause 25, further configured to cause the palette predictor to be updated by . 27. The apparatus according to clause 25, wherein a fixed number is set based on a slice type and a dual-tree mode setting. 28. The apparatus according to clause 25, wherein the size of one of the one or more palette predictors is initialized to a predefined value at the beginning of a slice in the case of a non-wavefront. 29. The apparatus according to clause 25, wherein the size of one of the one or more palette predictors is initialized to a predefined value at the beginning of a coded unit row in the case of a wavefront. 30. The processor executes instructions to cause the apparatus to Add bitstream compliance that the value of the size of the palette predictor is equal to the maximum size of the palette predictor when signaling a reuse flag. The apparatus according to clause 28 or 29, further configured to cause the processor to execute the above. 31. When initializing one or more palette predictors, the value of each entry and each component are set to 0 or (1 << (sequence bit depth - 1)). The apparatus according to any one of clauses 28 to 30. 32. The processor executes instructions to cause the apparatus to Apply a range constraint to the binary value of the reuse flag. The apparatus according to any one of clauses 25 to 31, further configured to cause the processor to execute the above. 33. A non-transitory computer-readable medium storing a set of instructions, executable by one or more processors of the apparatus to cause the apparatus to start a method for performing video data processing, the method including: Processing one or more coded units using one or more palette predictors. A palette predictor among one or more palette predictors adds all palette entries of the current palette as a set of first entries of the palette predictor, and adds palette entries from the previous palette predictor as a set of second entries of the palette predictor, which is a set of second entries after the set of first entries, regardless of the value of the reuse flag of the palette entries of the previous palette predictor A non - transitory computer - readable medium updated by the above. 34. The non - transitory computer - readable medium according to clause 33, wherein each palette entry of the set of first entries and the set of second entries of the palette predictor includes a reuse flag. 35. The method includes receiving a video frame for processing, and generating one or more coded units of the video frame, The non - transitory computer - readable medium according to clause 33, which further includes the above. 36. Each palette entry of the palette predictor has a corresponding reuse flag, The non - transitory computer - readable medium according to any one of clauses 33 - 35, wherein the number of reuse flags of the palette predictor is set to a fixed number for the corresponding coded unit. 37. A non - transitory computer - readable medium storing a set of instructions, which is executable by one or more processors of a device to start a method for performing video data processing on the device, and the method includes processing one or more coded units using one or more palette predictors, A palette predictor among one or more palette predictors adds all palette entries of the current palette as a set of first entries of the palette predictor, In a first range, one or more palette entries of a previous palette predictor that are within the first range starting from the first palette entry of the previous palette predictor and ending at the first palette entry of the previous palette predictor having a reuse flag set are added to the palette predictor as a set of one or more second entries of the palette predictor, and In a second range, one or more palette entries of the previous palette predictor that are within the second range starting from the last palette entry of the previous palette predictor having a reuse flag set and ending at the last palette entry of the previous palette predictor are added to the palette predictor as a set of one or more third entries of the palette predictor, A non - transitory computer - readable medium updated by and having the set of second entries and the set of third entries after the set of first entries. 38. The non - transitory computer - readable medium according to clause 37, wherein each palette entry of the set of first entries and the set of second entries of the palette predictor includes a reuse flag. 39. The method is Receiving a video frame for processing, Generating one or more coded units of the video frame, The non - transitory computer - readable medium according to clause 37, further comprising. 40. Each palette entry of the palette predictor has a corresponding reuse flag, The non - transitory computer - readable medium according to any one of clauses 37 - 39, wherein the number of reuse flags of the palette predictor is set to a fixed number for a corresponding coded unit. 41. A non - transitory computer - readable medium storing a set of instructions, executable by one or more processors of a device to start a method for performing video data processing on the device, the method comprising Receiving a video frame for processing, Generating one or more coded units of the video frame, Processing one or more coded units using one or more palette predictors having palette entries comprising each palette entry of the one or more palette predictors having a corresponding reuse flag a non - transitory computer - readable medium, wherein the number of reuse flags for each palette predictor is set to a fixed number for a corresponding coded unit 42. The palette predictor is updated by adding all palette entries of the current palette as a set of first entries of the palette predictor, and adding entries from a previous palette predictor not reused in the current palette as a set of second entries of the palette predictor, the set of second entries being after the set of first entries The non - transitory computer - readable medium according to clause 41 43. The non - transitory computer - readable medium according to clause 41, wherein the fixed number is set based on a slice type and a dual - tree mode setting 44. The non - transitory computer - readable medium according to clause 41, wherein the size of one of the one or more palette predictors is initialized to a predefined value at the beginning of a slice in the non - wavefront case 45. The non - transitory computer - readable medium according to clause 41, wherein the size of one of the one or more palette predictors is initialized to a predefined value at the beginning of a coded unit row in the wavefront case 46. Adding bit - stream compliance that the value of the size of the palette predictor is equal to the maximum size of the palette predictor when signaling the reuse flag The non - transitory computer - readable medium according to clause 44 or 45, further comprising 47. When initializing one or more palette predictors, the value of each entry and each component is set to 0 or (1<<(sequence bit depth - 1)) 48. Applying a range constraint to the binary value of the reuse flag The non-transitory computer-readable medium according to any one of clauses 41 to 47, further comprising 49. A method for deblocking a palette mode filter, comprising: Receiving a video frame for processing; Generating one or more coded units of the video frame, wherein each of the one or more coded units has one or more coding blocks; Setting a boundary filter strength to 2 in response to at least a first coding block of two adjacent coding blocks being coded in a palette mode; The method comprising. 50. A method for deblocking a palette mode filter, comprising: Receiving a video frame for processing; Generating one or more coded units of the video frame, wherein each of the one or more coded units has one or more coding blocks; Setting a boundary filter strength to 1 in response to at least a first coding block of two adjacent coding blocks being coded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode; The method comprising. 51. An apparatus for performing video data processing, comprising: A memory configured to store instructions; A processor coupled to the memory, the processor executing the instructions to cause the apparatus to: Receive a video frame for processing; Generate one or more coded units of the video frame, wherein each of the one or more coded units has one or more coding blocks; and Setting the boundary filter strength to 1 in response to at least a first coded block of two adjacent coded blocks being coded in palette mode and a second coded block of the two adjacent coded blocks having a coding mode different from the palette mode An apparatus configured to cause the execution of the above. A non-transitory computer-readable medium storing a set of instructions, executable by one or more processors of an apparatus to cause the apparatus to begin a method of video data processing, the method comprising: Receiving a video frame for processing; Generating one or more coding units of a video frame, each of the one or more coding units having one or more coded blocks; Setting the boundary filter strength to 1 in response to at least a first coded block of two adjacent coded blocks being coded in palette mode and a second coded block of the two adjacent coded blocks having a coding mode different from the palette mode; A non-transitory computer-readable medium including the above.

[0111]

[0137] In the above specification, embodiments have been described with respect to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications can be made to the described embodiments. By studying this specification and practicing the invention disclosed herein, other embodiments may become apparent to those skilled in the art. This specification and the examples are regarded as merely illustrative, and it is intended that the true scope and spirit of the invention be indicated by the following claims. It is also intended that the order of steps shown in the figures be for illustrative purposes only and not be limited to any particular order of steps. Thus, those skilled in the art can understand that these steps can be executed in different orders while implementing the same method.

[0112]

[0138] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications can be made to these embodiments. Accordingly, specific terms are used, but they are used only in a general and illustrative sense and not for purposes of limitation.

Claims

1. A method for processing video data, comprising: receiving a video frame for processing; generating one or more coding units of the video frame; processing the one or more coding units using one or more palette predictors having palette entries; wherein each palette entry of the one or more palette predictors has a corresponding reuse flag; and the number of reuse flags of each palette predictor is set to a fixed number for the corresponding coding unit. A method for processing video data.

2. The method according to claim 1, wherein the palette predictor is updated by: adding all palette entries of the current palette as a set of first entries of the palette predictor; and adding entries from a previous palette predictor not reused in the current palette as a set of second entries of the palette predictor, the set of second entries being after the set of first entries. The method according to claim 1, wherein the palette predictor is updated by:

3. adding all palette entries of the current palette as a set of first entries of the palette predictor; and adding palette entries from a previous palette predictor as a set of second entries of the palette predictor, the set of second entries being after the set of first entries, regardless of the value of the reuse flag of the palette entries of the previous palette predictor. The method according to claim 1, wherein the palette predictor is updated by: [[ID=I7]]adding all palette entries of the current palette as a set of first entries of the palette predictor;

4. adding one or more palette entries of the previous palette predictor within a first range starting from a first palette entry of the previous palette predictor and ending at the first palette entry of the previous palette predictor having a set of reuse flags as a set of one or more second entries of the palette predictor to the palette predictor; and ​ ​ In a second range, starting from the last palette entry having a reuse flag set of the previous palette predictor and ending at the last palette entry of the previous palette predictor, one or more palette entries of the previous palette predictor within the second range are added to the palette predictor as a set of third one or more entries of the palette predictor The method according to claim 1, wherein the set of second entries and the set of third entries are after the set of first entries, and the palette predictor is updated thereby **Claim 5** The method according to claim 1, wherein the fixed number is set based on a slice type and a dual tree mode setting **Claim 6** The method according to claim 1, wherein the size of one of the one or more palette predictors is initialized to a predefined value at the start of a slice in the case of a non-wavefront **Claim 7** An apparatus for performing video data processing, comprising a memory configured to store instructions, and one or more processors, the one or more processors executing the instructions to cause the apparatus to receive a video frame for processing, generate one or more coded units of the video frame, and process the one or more coded units using one or more palette predictors having palette entries wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and the number of reuse flags for each palette predictor is set to a fixed number for a corresponding coded unit **Claim 8** The one or more processors further execute the instructions to cause the apparatus to update the palette predictor by adding all palette entries of the current palette as a set of first entries of the palette predictor, and adding entries from the previous palette predictor not reused in the current palette as a set of second entries of the palette predictor, the set of second entries being after the set of first entries The apparatus according to claim 7, further configured to perform the above **Claim 9** The one or more processors further execute the instructions to cause the apparatus to ​ adding all palette entries of the current palette as a set of first entries of the palette predictor; and adding palette entries from a previous palette predictor as a second set of entries of the palette predictor, the second set of entries being after the first set of entries, regardless of the value of a reuse flag of the palette entries of the previous palette predictor. updating the palette predictor by The apparatus of claim 7 , further configured to cause the execution of:

10. The one or more processors execute the instructions to cause the device to: adding all palette entries of the current palette as a set of first entries of said palette predictor; adding one or more palette entries of the previous palette predictor within a first range, the first range starting from a first palette entry of the previous palette predictor and ending with a first palette entry of the previous palette predictor that has a reuse flag set, to the palette predictor as a second set of one or more entries of the palette predictor; and adding one or more palette entries of the previous palette predictor that are within a second range, the second range starting from the last palette entry of the previous palette predictor that has a reuse flag set and ending with the last palette entry of the previous palette predictor, to the palette predictor as a third set of one or more entries of the palette predictor. updating the palette predictor by 8. The apparatus of claim 7, further configured to cause execution of: wherein the second set of entries and the third set of entries are after the first set of entries.

11. The apparatus of claim 7 , wherein the fixed number is set based on a slice type and a dual-tree mode setting.

12. 8. The apparatus of claim 7, wherein a size of one of the one or more palette predictors is initialized to a predefined value at the beginning of a slice for a non-wavefront case.

13. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method comprising: Receiving a video frame for processing; Generating one or more coded units of the video frame; Processing one or more coded units using one or more palette predictors having palette entries, wherein each palette entry of the one or more palette predictors has a corresponding reuse flag; a non-transitory computer-readable medium, wherein the number of reuse flags for each palette predictor is set to a fixed number for a corresponding coded unit. **Claim 14** The palette predictor is updated by adding all palette entries of the current palette as a first set of entries of the palette predictor, and adding entries from the previous palette predictor that are not reused in the current palette as a second set of entries of the palette predictor, the second set of entries being after the first set of entries; A non-transitory computer-readable medium according to claim 13. **Claim 15** The palette predictor is updated by adding all palette entries of the current palette as a first set of entries of the palette predictor, and adding palette entries from the previous palette predictor as a second set of entries of the palette predictor, the second set of entries being after the first set of entries, regardless of the value of the reuse flag of the palette entries of the previous palette predictor; A non-transitory computer-readable medium according to claim 13. **Claim 16** The palette predictor is updated by adding all palette entries of the current palette as a first set of entries of the palette predictor, adding one or more palette entries of the previous palette predictor within a first range, the first range starting from the first palette entry of the previous palette predictor and ending at the first palette entry of the previous palette predictor having a set of reuse flags, as a second set of one or more entries of the palette predictor to the palette predictor, and ​ ​ In a second range, starting from the last palette entry having a reuse flag set of the previous palette predictor and ending at the last palette entry of the previous palette predictor, adding one or more palette entries of the previous palette predictor that are within the second range as a set of third one or more entries of the palette predictor to the palette predictor The non-transitory computer-readable medium according to claim 13, updated by, wherein the set of second entries and the set of third entries are after the set of first entries. **Claim 17** The non-transitory computer-readable medium according to claim 13, wherein the fixed number is set based on a slice type and a dual-tree mode setting. **Claim 18** A method for deblocking a palette mode filter, comprising: Receiving a video frame for processing; Generating one or more coded units of the video frame, wherein each of the one or more coded units has one or more coding blocks; Setting a boundary filter strength to 1 in response to at least a first coding block of two adjacent coding blocks being coded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode; The method including the above. **Claim 19** An apparatus for performing video data processing, comprising: A memory configured to store instructions; and One or more processors, wherein the one or more processors execute the instructions to cause the apparatus to: Receive a video frame for processing; Generate one or more coded units of the video frame, wherein each of the one or more coded units has one or more coding blocks; and Set a boundary filter strength to 1 in response to at least a first coding block of two adjacent coding blocks being coded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode. The apparatus is configured to perform the above. **Claim 20** A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method comprising: Receiving a video frame for processing; One or more coded units of the video frame, each of the one or more coded units having one or more coding blocks; Setting a boundary filter strength to 1 in response to at least a first coding block of two adjacent coding blocks being coded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode; A non-transitory computer-readable medium comprising the above.

Citation Information

Patent Citations

  • Determining application of deblocking filtering to palette coded blocks in video coding

    WO2015191834A1

  • An encoder, a decoder and corresponding methods of boundary strength derivation of deblocking filter

    WO2020114513A1

  • Condition dependent coding of palette mode usage indication

    WO2020169105A1

  • Encoding device, decoding device, encoding method, and decoding method

    WO2021161914A1