Palette prediction method
By simplifying the palette predictor update process and fixedly setting the number of reuse flags, the problem of inefficient hardware implementation in palette mode is solved, encoding efficiency is improved, and the boundary filter strength is reasonably defined, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202180026652.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-04
- Filing Date
- 2021-03-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-03-31
AI Technical Summary
The existing video encoding technology has a high complexity in the palette predictor update process in the palette mode, resulting in low hardware implementation efficiency, and the boundary filter intensity is undefined in different encoding modes, affecting the encoding efficiency.
Simplify the process of updating the palette predictor, by fixedly setting the number of reuse flags, and implementing CABAC parsing in parallel in hardware design, defining the palette mode as an independent encoding mode to determine the boundary filter intensity.
The hardware implementation efficiency and coding efficiency of the palette prediction process are improved, the delay of CABAC throughput is reduced, and the reasonable calculation of boundary filter strength is ensured.
Smart Images

Figure CN115428455B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This disclosure claims priority to U.S. Provisional Application No. 63 / 005,305, filed Apr. 4, 2020, and U.S. Provisional Application No. 63 / 002,594, filed Mar. 31, 2020, which are hereby incorporated by reference in their entirety. Technical field
[0003] This disclosure generally relates to video processing, and more particularly, to the use of palette modes in video encoding and decoding. Background art
[0004] Video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is typically referred to as encoding, and the decompression process is typically referred to as decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and in - loop filtering. Standardization organizations have developed video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify particular video coding formats. As more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is getting higher and higher. Summary of the invention
[0005] Embodiments of this disclosure provide a computer - implemented method for a palette predictor. In some embodiments, the method includes: receiving a video frame for processing; generating one or more coding units of the video frame; and processing the one or more coding units using one or more palette predictors having a plurality of palette entries, wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number.
[0006] Embodiments of the present disclosure provide an apparatus. In some embodiments, the apparatus includes: a memory configured to store instructions; and a processor coupled to the memory and configured to execute the instructions to cause the apparatus to perform: receiving a video frame for processing; generating one or more coding units of the video frame; and processing one or more coding units using one or more palette predictors having a plurality of palette entries, wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number.
[0007] Embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing a set of instructions that can be executed by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing. In some embodiments, the method includes: receiving a video frame for processing; generating one or more coding units of the video frame; and processing one or more coding units using one or more palette predictors having a plurality of palette entries, wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number.
[0008] Embodiments of the present disclosure provide a computer-implemented method for a deblocking filter in a palette mode. In some embodiments, the method includes: receiving a video frame for processing; generating one or more coding units for the video frame, wherein each coding unit of the one or more coding units has one or more coding blocks; and setting a boundary filter strength to 1 in response to at least a first coding block of two adjacent coding blocks being encoded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode.
[0009] Embodiments of the present disclosure provide an apparatus. In some embodiments, the apparatus includes: a memory configured to store instructions; and a processor coupled to the memory and configured to execute the instructions to cause the apparatus to perform: receiving a video frame for processing; generating one or more coding units for the video frame, wherein each coding unit of the one or more coding units has one or more coding blocks; and setting a boundary filter strength to 1 in response to at least a first coding block of two adjacent coding blocks being encoded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode.
[0010] Embodiments of the present disclosure provide a non - transitory computer - readable storage medium storing an instruction set, which can be executed by one or more processors of a device to cause the device to initiate a method for performing video data processing. In some embodiments, the method includes: receiving a video frame for processing; generating one or more coding units for the video frame, wherein each of the one or more coding units has one or more coding blocks; and in response to at least a first coding block of two adjacent coding blocks being encoded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode, setting a boundary filter strength to 1. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Embodiments and various aspects of the present disclosure are shown in the following detailed description and the drawings. The various features shown in the drawings are not drawn to scale.
[0012] Figure 1 is a schematic structural diagram of an exemplary video sequence according to some embodiments of the present application.
[0013] Figure 2A is a schematic diagram showing an exemplary encoding process of a hybrid video coding system consistent with embodiments of the present application.
[0014] Figure 2B is a schematic diagram showing another exemplary encoding process of a hybrid video coding system consistent with embodiments of the present application.
[0015] Figure 3A is a schematic diagram showing an exemplary decoding process of a hybrid video coding system consistent with embodiments of the present application.
[0016] Figure 3B is a schematic diagram showing another exemplary decoding process of a hybrid video coding system consistent with embodiments of the present application.
[0017] Figure 4 is a block diagram of an exemplary device for encoding or decoding a video according to some embodiments of the present application.
[0018] Figure 5 shows an illustration of a block encoded in a palette mode according to some embodiments of the present application.
[0019] Figure 6 shows an example of a palette predictor update process.
[0020] Figure 7 shows an example palette coding syntax.
[0021] Figure 8Shows a flowchart of a palette predictor update process according to some embodiments of the present application.
[0022] Figure 9 Shows an example of a palette predictor update process according to some embodiments of the present application.
[0023] Figure 10 Shows an exemplary decoding process for the palette mode.
[0024] Figure 11 Shows an exemplary decoding process for the palette mode according to some embodiments of the present application.
[0025] Figure 12 Shows a flowchart of another palette predictor update process according to some embodiments of the present application.
[0026] Figure 13 Shows an example of another palette predictor update process according to some embodiments of the present application.
[0027] Figure 14 Shows an example of a palette coding syntax according to some embodiments of the present application.
[0028] Figure 15 Shows an exemplary decoding process for the palette mode according to some embodiments of the present application.
[0029] Figure 16 Shows a schematic diagram of a decoder hardware design for the palette mode for implementing a part of the palette predictor update process according to some embodiments of the present application.
[0030] Figure 17 Shows an exemplary decoding process for the palette mode according to some embodiments of the present application.
[0031] Figure 18 Shows an exemplary initialization process for the palette mode.
[0032] Figure 19 Shows an exemplary initialization process for the palette mode according to some embodiments of the present application.
[0033] Figure 20 Shows the correspondence of the palette predictor update and the reuse flag Run-length encoding Example of the code.
[0034] Figure 21 Shows an example of the corresponding run - length encoding of the palette predictor update and the reuse flag according to some embodiments of the present disclosure.
[0035] Figure 22Shows an example of the corresponding run - length encoding of palette predictor updates and reuse flags according to some embodiments of the present disclosure.
[0036] Figure 23 Shows an exemplary palette coding syntax according to some embodiments of the present application.
[0037] Figure 24 Shows exemplary palette coding semantics.
[0038] Figure 25 Shows exemplary palette coding semantics according to some embodiments of the present application. Detailed Description
[0039] Now, reference will be made in detail to the exemplary embodiments, which are illustrated in the accompanying drawings. The following description refers to the drawings, unless otherwise specified, where the same numbers in different drawings represent the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatuses and methods consistent with aspects related to the present disclosure as set forth in the appended claims. Specific aspects of the present disclosure are described in more detail below. If there is a conflict with the terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0040] The Joint Video Exploration Team (JVET) of the ITU - T Video Coding Experts Group (ITU - T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0041] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been using the Joint Exploration Model (JEM) reference software to explore technologies beyond HEVC. As coding techniques are incorporated into JEM, JEM has achieved higher coding performance than HEVC.
[0042] The VVC standard has been recently developed and continues to include more coding techniques that provide better compression performance. VVC is based on the hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0043] A video is a sequence of still images (or "frames") arranged in chronological order to store visual information. These images can be acquired and stored in chronological order using a video capture device (e.g., a camera), and such images in the time series can be displayed using a video playback device (e.g., a television, computer, smartphone, tablet computer, video player, or any end-user terminal with a display function). Additionally, in some applications, the video capture device can send the captured video in real time to a video playback device (e.g., a computer with a monitor), such as for surveillance, conferencing, or live broadcasting.
[0044] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or by dedicated hardware. The module for compression is generally referred to as an "encoder", and the module for decompression is generally referred to as a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, a hardware implementation of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. A software implementation of the encoder and decoder can include program code fixed in a computer-readable medium, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process. Video compression and decompression can be achieved through various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress a video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be referred to as a "transcoder".
[0045] The video encoding process can identify and retain useful information that can be used to reconstruct an image and ignore unimportant reconstruction information. If the unimportant information cannot be fully reconstructed when ignored, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off to reduce the required storage space and transmission bandwidth.
[0046] The useful information of an encoded image (referred to as the "current image") includes the changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in the position of pixels, changes in brightness, or changes in color, with the change in position being the most concerned. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference image and the current image.
[0047] An image encoded without reference to another image (i.e., it is its own reference image) is called an "I-image". An image encoded using a previous image as a reference image is called a "P-image", and an image encoded using both a previous image and a future image as reference images is called a "B-image" (the reference is "bidirectional").
[0048] Figure 1 Fig. 5 shows the structure of an example video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live video or a video that has been captured and archived. Video 100 can be a real-life video, a computer-generated video (e.g., a computer game video), or a combination of both (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) that receives video from a video content provider.
[0049] As Figure 1 shown, the video sequence 100 can include a series of images arranged in time along a timeline, including images 102, 104, 106, and 108. Images 102-106 are consecutive, and there are more images between images 106 and 108. In Figure 1 , image 102 is an I-image, and its reference image is image 102 itself. Image 104 is a P-image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B-image, and its reference images are images 104 and 108, as indicated by the arrows. In some embodiments, the reference image of an image (e.g., image 104) may not be immediately before or after the image. For example, the reference image of image 104 can be an image before image 102. It should be noted that the reference images of images 102-106 are merely examples, and the present disclosure does not limit the embodiments of the reference images as Figure 1 shown.
[0050] Generally, due to the computational complexity of the encoding and decoding tasks, video codecs do not encode or decode an entire image at once. Instead, they can divide the image into basic segments and encode or decode the image segments one by one. In the present disclosure, such a basic segment is called a basic processing unit ("BPU"). For example, Figure 1The structure 110 therein shows an example structure of an image (e.g., any of the images 102 - 108) of the video sequence 100. In the structure 110, the image is divided into 4×4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit may have a variable size in the image, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing unit can be selected for the image based on a balance of coding efficiency and the level of detail to be maintained in the basic processing unit. The CTU is the largest block unit and can include up to 128×128 luminance samples (plus corresponding chrominance samples depending on the chrominance format). The CTU can be further divided into coding units (CUs) using a quadtree, binary tree, ternary tree, or a combination thereof.
[0051] The basic processing unit can be a logical unit, which can include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, the basic processing unit of a color image can include a luminance component (Y) representing achromatic luminance information, one or more chrominance components representing color information (e.g., Cb and Cr), and associated syntax elements, where the luminance and chrominance components can have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components can be referred to as "coding tree blocks" ("CTB"). Any operation performed on the basic processing unit can be repeated for each of its luminance and chrominance components.
[0052] Video coding has multiple operation stages, examples of which are Figures 2A - 2B and Figures 3A - 3BAs shown. For each stage, the size of the basic processing unit may still be too large for processing, so it can be further divided into segments called "basic processing subunits" in the present disclosure. In some embodiments, the basic processing subunit may be called a "block" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or an "encoding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same size as the basic processing unit or a smaller size than the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to a further level according to processing needs. It should also be noted that different stages may use different schemes to divide the basic processing unit.
[0053] For example, in the mode decision stage (an example of which is shown in Figure 2B ), the encoder can decide what prediction mode (e.g., intra prediction or inter prediction) to use for the basic processing unit, which may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC), and decide the prediction type for each individual basic processing subunit.
[0054] For another example, in the prediction stage (an example of which is shown in Figures 2A - 2B ), the encoder can perform prediction operations at the level of the basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level the prediction operations can be performed.
[0055] For another example, in the transform stage (an example of which is shown in Figures 2A - 2BAs shown (in [reference], etc.), the encoder can perform a transformation operation on a residual basic processing unit (e.g., a CU). However, in some cases, the basic processing unit may still be too large to process. The encoder can further divide the basic processing unit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level the transformation operation can be performed. It should be noted that the partitioning scheme of the same basic processing unit can be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU can have different sizes and numbers.
[0056] In Figure 1 the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing subunits, the boundaries of which are shown as dashed lines. Different basic processing units of the same image can be divided into basic processing subunits in different schemes.
[0057] In some embodiments, in order to provide the ability for parallel processing and the fault tolerance ability for video encoding and decoding, an image can be divided into regions for processing, such that for a region of the image, the encoding or decoding process can be independent of information from any other region of the image. In other words, each region of the image can be processed separately. By doing so, the codec can process different regions of the image in parallel, thus improving the encoding efficiency. In addition, when the data of a region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the damaged or lost data, thus providing fault tolerance. In some video coding standards, an image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different images of the video sequence 100 can have different partitioning schemes for dividing the image into regions.
[0058] For example, in Figure 1 the structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that Figure 1 the basic processing units, basic processing subunits, and structural regions in 110 are only examples, and the present disclosure does not limit its embodiments.
[0059] Figure 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A can be executed by an encoder. As Figure 2AAs shown, the encoder may encode video sequence 202 into video bitstream 228 according to process 200A. Similar to Figure 1 the video sequence 100 in Figure 1 , the video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to Figure 1 the structure 110 in
[0060] , each original image of the video sequence 202 may be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder may execute process 200A at the level of basic processing units for each original image of the video sequence 202. For example, the encoder may execute process 200A in an iterative manner, where the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may execute process 200A in parallel for regions (e.g., regions 114-118) of each original image of the video sequence 202.
[0060] Referring to Figure 2A , the encoder may feed a basic processing unit of an original image of the video sequence 202 (referred to as "original BPU") to prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate residual BPU 210. The encoder may feed the residual BPU 210 to transform stage 212 and quantization stage 214 to 216 generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after the quantization stage 214, the encoder may feed the quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path". The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0061] The encoder may iteratively execute process 200A to encode each original BPU (in the forward path) of the encoded original image and generate prediction reference 224 for the next original BPU (in the reconstruction path) of the encoded original image. After encoding all the original BPUs of the original image, the encoder may continue to encode the next image in the video sequence 202.
[0062] Referring to process 200A, the encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). The term "receive" as used herein may refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or using for inputting data in any manner.
[0063] In the prediction stage 204, at the current iteration, the encoder may receive the original BPU and the prediction reference 224, and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of a previous iteration of process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from the prediction data 206 and the prediction reference 224 that can be used to reconstruct the original BPU into the predicted BPU 208.
[0064] Ideally, the predicted BPU 208 may be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To record these differences, when generating the predicted BPU 208, the encoder may subtract it from the original BPU to generate a residual BPU 210. For example, the encoder may subtract the value of the corresponding pixel of the predicted BPU 208 (e.g., grayscale value or RGB value) from the value of the pixel of the original BPU. Each pixel of the residual BPU 210 may have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared with the original BPU, the prediction data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.
[0065] To further compress the residual BPU 210, in the transform stage 212, the encoder may reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "base patterns". Each base pattern is associated with a "transformation coefficient". The base patterns may have the same size (e.g., the size of the residual BPU 210), and each base pattern may represent a frequency component of the variation of the residual BPU 210 (e.g., the frequency of brightness variation). None of the base patterns can be reproduced from any combination (e.g., linear combination) of any other base patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the base images are similar to the basic functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basic functions.
[0066] Different transformation algorithms can use different basic patterns. Various transformation algorithms can be used at transformation stage 212, such as, for example, discrete cosine transform, discrete sine transform, etc. The transformation at transformation stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transformation (referred to as "inverse transformation"). For example, to recover the pixels of the residual BPU 210, the inverse transformation can be multiplying the values of the corresponding pixels of the basic pattern by the corresponding correlation coefficients and summing the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transformation algorithm (and thus have the same basic pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 therefrom without receiving the basic pattern from the encoder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Thus, the residual BPU 210 is further compressed.
[0067] The encoder can further compress the transformation coefficients at quantization stage 214. During the transformation process, different basic patterns can represent different change frequencies (e.g., luminance change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore the information of high-frequency changes without causing significant quality degradation in decoding. For example, at quantization stage 214, the encoder can generate the quantized transformation coefficients 216 by dividing each transformation coefficient by an integer value (referred to as "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transformation coefficients of the high-frequency basic pattern can be converted to zero, and the transformation coefficients of the low-frequency basic pattern can be converted to smaller integers. The encoder can ignore the quantized transformation coefficients 216 with zero values, whereby the transformation coefficients are further compressed. This quantization process is also reversible, where the quantized transformation coefficients 216 can be reconstructed as transformation coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0068] Since the encoder ignores the remainder of the division in the rounding operation, the quantization stage 214 can be lossy. Generally, the quantization stage 214 can contribute the most information loss in process 200A. The greater the information loss, the fewer the number of bits required for the quantized transformation coefficients 216. To obtain different levels of information loss, the encoder can use different quantization parameter values or any other parameters of the quantization process.
[0069] In the binary encoding stage 226, the encoder can encode the prediction data 206 and the quantized transform coefficients 216 using binary encoding techniques, such as entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context - adaptive binary arithmetic encoding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform at the transform stage 212, the parameters of the quantization process (e.g., quantization parameter), the encoder control parameters (e.g., bit - rate control parameter), etc. The encoder can use the output data of the binary encoding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 can be further packetized for network transmission.
[0070] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate the reconstructed transform coefficients. In the inverse transform stage 220, the encoder can generate the reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate the prediction reference 224 that will be used in the next iteration of process 200A.
[0071] It should be noted that other variants of process 200A can be used to encode the video sequence 202. In some embodiments, the stages of process 200A can be executed by the encoder in a different order. In some embodiments, one or more stages of process 200A can be combined into a single stage. In some embodiments, a single stage of process 200A can be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be combined into a single stage. In some embodiments, process 200A can include additional stages. In some embodiments, process 200A can omit Figure 2A one or more of the
[0072] Figure 2B FIG. shows a schematic diagram of another example encoding process 200B according to an embodiment of the present disclosure. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B further includes a loop filter stage 232 and a buffer 234.
[0073] Generally, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra prediction") can use pixels from one or more already-encoded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter prediction") can use regions from one or more already-encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce the temporal redundancy inherent in the image.
[0074] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra prediction. For the original BPU of the encoded image, the prediction reference 224 can include one or more adjacent BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder can generate the predicted BPU 208 by interpolating the adjacent BPUs. The interpolation techniques can include, for example, linear interpolation or interpolation, polynomial interpolation or interpolation, etc. In some embodiments, the encoder can perform interpolation at the pixel level, e.g., by interpolating the values of the corresponding pixels of each pixel of the predicted BPU 208. The adjacent BPUs used for interpolation can be located in various directions relative to the original BPU, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., at the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 can include, for example, the positions (e.g., coordinates) of the adjacent BPUs used, the sizes of the adjacent BPUs used, the parameters of the interpolation, the direction of the adjacent BPUs used relative to the original BPU, etc.
[0075] For another example, in the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed on a per-BPU basis. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same image have been generated, the encoder may generate a reconstructed image as a reference image. The encoder may perform an operation of "motion estimation" to search for a matching region within the range of the reference image (referred to as the "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a position in the reference image that has the same coordinates as the original BPU in the current image and may extend outward a predetermined distance. When the encoder identifies (e.g., by using a pel-recursive algorithm, a block-matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder may determine such a region as the matching region. The matching region may have a different size (e.g., smaller than, equal to, larger than, or a different shape) from the original BPU. Since the reference image and the current image are temporally separated on the timeline (e.g., as Figure 1 shown), the matching region may be considered to "move" over time to the position of the original BPU. The encoder may record the direction and distance of this motion as a "motion vector". When multiple reference images are used (e.g., as in Figure 1 image 106), the encoder may search for the matching region and determine its associated motion vector for each reference image. In some embodiments, the encoder may assign weights to the pixel values of the matching regions of the respective matching reference images.
[0076] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.
[0077] To generate the predicted BPU 208, the encoder may perform an operation of "motion compensation". Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., as in Figure 1For the image 106) therein, the encoder can move the matching region of the reference image according to the respective motion vectors and average pixel values of the matching regions. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of the respective matching reference images, the encoder can add the weighted sums of the pixel values of the moved matching regions.
[0078] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same time direction relative to the current image. For example, Figure 1 the image 104 therein is an unidirectional inter-frame prediction image, where the reference image (i.e., image 102) is before image 04. Bidirectional inter-frame prediction can use one or more reference images in two time directions relative to the current image. For example, Figure 1 the image 106 therein is a bidirectional inter-frame prediction image, where the reference images (i.e., images 104 and 08) are in two time directions relative to image 104.
[0079] Still referring to the forward path of process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform rate-distortion optimization techniques, where the encoder can select a prediction mode to minimize the value of a cost function according to the bit rate of the candidate prediction mode and the distortion of the reconstructed reference image under the candidate prediction mode. According to the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.
[0080] In the reconstruction path of process 200B, if an intra prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), the encoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If an inter prediction mode has been selected in the forward path, after generating the prediction reference 224 (e.g., the current image in which all BPUs have been encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filter stage 232. At this stage, the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion introduced by inter prediction (e.g., blocking artifacts). The encoder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive compensation, adaptive loop filter, etc. The loop-filtered reference image can be stored in the buffer 234 (or "decoded image buffer") for later use (e.g., as an inter prediction reference image for future images of the video sequence 202). The encoder can store one or more reference images in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode the parameters of the loop filter (e.g., loop filter strength) as well as the quantized transform coefficients 216, prediction data 206, and other information at the binary coding stage 226.
[0081] Figure 3A FIG. shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A can be a decompression process corresponding to Figure 2A the compression process 200A therein. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss during the compression and decompression processes (e.g., Figures 2A - 2B the quantization stage 214 therein), generally, the video stream 304 is different from the video sequence 202. Similar to Figures 2A - 2B processes 200A and 200B therein, the decoder can perform process 300A on each image encoded in the video bitstream 228 at the basic processing unit (BPU) level. For example, the decoder can perform process 300A in an iterative manner, where the decoder can decode the basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for each region (e.g., regions 114 - 118) of each image encoded in the video bitstream 228.
[0082] As Figure 3AAs shown, the decoder can feed a portion of the video bitstream 228 associated with the basic processing unit of the encoded image (referred to as the "encoded BPU") into the binary decoding stage 302, where the decoder can decode this portion into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 into the inverse quantization stage 218 and the inverse transform stage 220 to generate the reconstructed residual BPU 222. The decoder can feed the prediction data 206 into the prediction stage 204 to generate the predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., the decoded image buffer in a computer memory). The decoder can feed the prediction reference 224 into the prediction stage 204 for performing prediction operations in the next iteration of process 300A.
[0083] The decoder can iteratively execute process 300A to decode each encoded BPU of the encoded image and generate the prediction reference 224 for the next encoded BPU of the encoded image. After decoding all the encoded BPUs of the encoded image, the decoder can output the image to the video stream 304 for display and continue to decode the next encoded image in the video bitstream 228.
[0084] In the binary decoding stage 302, the decoder can perform the inverse operations of the binary encoding techniques used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context - adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, the parameters of the prediction operation, the transform type, the parameters of the quantization process (e.g., quantization parameter), the encoder control parameters (e.g., bitrate control parameter), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over a network, the decoder can unpack the video bitstream 228 before feeding it into the binary decoding stage 302.
[0085] Figure 3B A schematic diagram of another example decoding process 300B according to an embodiment of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filtering stage 232 and a buffer 2344.
[0086] In process 300B, for an encoded basic processing unit (referred to as the "current BPU") of a decoded encoded image (referred to as the "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data, depending on what prediction mode the encoder used to encode the current BPU. For example, if the encoder uses intra prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation can include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as references, the sizes of the adjacent BPUs, interpolation parameters, the orientation of the adjacent BPUs relative to the original BPU, etc. For another example, if the encoder uses inter prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. The parameters of the inter prediction operation can include, for example, the number of reference images associated with the current BPU, the weights respectively associated with the reference images, the positions (e.g., coordinates) of one or more matching regions in the corresponding reference images, one or more motion vectors respectively associated with the matching regions, etc.
[0087] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of performing such spatial prediction or temporal prediction are described in Figure 2B and will not be repeated hereinafter. After performing such spatial prediction or temporal prediction, the decoder can generate a predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in Figure 3A
[0088] In process 300B, the decoder can feed the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference image in which all BPUs are decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortions (e.g., blocking artifacts). The decoder can, as Figure 2B Apply the loop filter to the prediction reference 224 in the manner shown. The reference image for loop filtering can be stored in buffer 234 (e.g., the decoded picture buffer in computer memory) for later use (e.g., as the inter-prediction reference picture for future encoded pictures in video bitstream 228). The decoder can store one or more reference pictures in buffer 234 for use at the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction is used to encode the current BPU, the prediction data may further include the parameters of the loop filter (e.g., loop filter strength).
[0089] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding video according to an embodiment of the present disclosure. As Figure 4 shown, the apparatus 400 may include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 may become a dedicated machine for video encoding or decoding. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), a field programmable gate array (FPGA), a system on chip (SoC), an application specific integrated circuit (ASIC), etc. in any combination. In some embodiments, the processor 402 may also be a group of processors grouped as a single logical component. For example, as Figure 4 shown, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0090] The apparatus 400 may further include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, as Figure 4As shown, the stored data may include program instructions (e.g., for implementing the stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. The memory 404 may include high-speed random access storage devices or non-volatile storage devices. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical discs, magnetic disks, hard disk drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a group of memories grouped as a single logical component ( Figure 4 not shown in
[0091] Bus 410 may be a communication device for transferring data between components inside the device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., universal serial bus port, peripheral component interconnect express port), or the like.
[0092] For ease of explanation without ambiguity, in this disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuits may be implemented entirely in hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuits may be a single separate module or may be fully or partially incorporated into any other component of the device 400.
[0093] The device 400 may further include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0094] In some embodiments, optionally, the device 400 may further include a peripheral interface 408 to provide connections to one or more peripheral devices. As Figure 4 shown, the peripheral devices may include, but are not limited to, cursor control devices (e.g., mouse, touchpad, or touchscreen), keyboards, displays (e.g., cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces coupled to video archives), etc.
[0095] It should be noted that a video codec (e.g., the codec performing processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instances that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0096] Figure 5 A schematic diagram of a CU encoded in palette mode is shown. In VVC draft 8, the palette mode can be used in monochrome, 4:2:0, 4:2:2, and 4:4:4 color formats. When the palette mode is enabled, if the CU size is less than or equal to 64x64 and greater than 16 samples, a flag is sent at the CU level to indicate whether the palette mode is used. If the (current) CU 500 is encoded using the palette mode, the sample value at each position in the CU is represented by a small set of representative color values. This set is called the palette 510. For sample positions with values close to palette colors 501, 502, 503, the corresponding palette index is signaled. A color value outside the palette can also be specified by signaling an escape index 504. Then, for all positions in the CU that use the escape color index, the (quantized) color component values for each of these positions are signaled.
[0097] For the encoding of the palette, the palette predictor is maintained. Figure 6 An exemplary process for updating the palette predictor after encoding each coding unit 600 is shown. At the start of each slice for the non-wavefront case and at the start of each CTU row for the wavefront case, the predictor is initialized to 0 (i.e., empty). For each entry in the palette predictor, a reuse flag is signaled to indicate whether it will be included in the current palette of the current CU. The reuse flag is sent using run-length encoding of zeros. After that, the number of new palette entries and the component values of the new palette entries are signaled. After encoding the palette for the CU, the palette predictor is updated using the current palette, and the entries from the previous palette predictor that are not reused in the current palette are added at the end of the new palette predictor until the allowed maximum size is reached.
[0098] Signal an escape flag for each CU to indicate whether there is an escape symbol in the current CU. If there is an escape symbol, increment the palette table by one and assign the last index to the escape symbol. The palette indices of the samples in the CU form a palette index map as shown in the example of Figure 5 . The index map is encoded using a horizontal or vertical traversal scan. The scan order is explicitly signaled in the bitstream using the palette_transpose_flag. The palette index map is encoded using the index run mode or the index copy mode.
[0099] In VVC Draft 8, the deblocking filter process includes defining block boundaries, deriving boundary filter strengths based on the coding modes of two adjacent blocks along the defined block boundaries, deriving the number of samples to be filtered, and applying the deblocking filter to the samples. When the edge is a coding unit, coding sub-block unit, or transform unit boundary, the edge is defined as a block boundary. Then, the boundary filter strength is calculated based on the coding modes of the two adjacent blocks according to the following six rules. (1) If both coding blocks are coded in the BDPCM mode, the boundary filter strength is set to 0. (2) Otherwise, if one of the coding blocks is coded in the intra mode, the boundary filter strength is set to 2. (3) Otherwise, if one of the coding blocks is coded in the CIIP mode, the boundary filter strength is set to 2. (4) Otherwise, if one of the coding blocks contains one or more non-zero coefficient levels, the boundary filter strength is set to 1. (5) Otherwise, if one block is coded in the IBC mode and the other block is coded in the inter mode, the boundary filter strength is set to 1. (6) Otherwise (both blocks are coded in the IBC or inter mode), the boundary filter strength is derived using the reference images and motion vectors of the two blocks.
[0100] VVC Draft 8 gives a relatively detailed description of the calculation process of the boundary filtering strength. Specifically, VVC Draft 8 presents eight consecutive and detailed scenarios. In Scenario 1, if cIdx is equal to 0 and both samples p0 and q0 are in the coding block where intra_bdpcm_luma_flag is equal to 1, then bS[xDi][yDj] is set to be equal to 0. Otherwise, in Scenario 2, if cIdx is greater than 0 and both samples p0 and q0 are in the coding block where intra_bdpcm_chroma_flag is equal to 1, then bS[xDi][yDj] is set to be equal to 0. Otherwise, in Scenario 3, if sample p0 or q0 is in the coding block of the coding unit encoded in the intra prediction mode, then bS[xDi][yDj] is set to be equal to 2. Otherwise, in Scenario 4, if the block edge is also the coding block edge and sample p0 or q0 is in the coding block where ciip_flag is equal to 1, then bS[xDi][yDj] is set to be equal to 2. Otherwise, in Scenario 5, if the block edge is also the transform block edge and sample p0 or q0 is in the transform block containing one or more non-zero transform coefficient levels, then bS[xDi][yDj] is set to be equal to 1. Otherwise, in Scenario 6, if the prediction mode of the coding sub-block containing sample p0 is different from the prediction mode of the coding sub-block containing sample q0 (i.e., one of the coding sub-blocks is encoded in the IBC prediction mode and the other is encoded in the inter prediction mode), bS[xDi][yDj] is set to be equal to 1.
[0101] Otherwise, in Scenario 7, if cIdx is equal to 0, edgeFlags[xDi][yDj] is equal to 2, and one or more of the following conditions are true, then bS[xDi][yDj] is set to be equal to 1.
[0102] Condition (1): Both the coding sub-block containing sample p0 and the coding sub-block containing sample q0 are encoded using the IBC prediction mode, and the absolute difference between the horizontal or vertical components of the block vectors used in the predictions of the two coding sub-blocks is greater than or equal to 8 in units of 1 / 16 luma samples.
[0103] Condition (2): For the prediction of the coded block containing sample p0, a different reference image or a different number of motion vectors is used compared to the prediction of the coded block containing sample q0. For Condition (2), note that the determination of whether the reference images used for the two coded blocks are the same or different is based only on which images are referenced, regardless of whether Index 0 or Index 1 of the reference image list is used to form the prediction. Also, the index positions in the reference image list are not considered. Similarly, for Condition (2), note that the number of motion vectors used to predict the coded block with the top-left sample covering (xSb, ySb) is equal to PredFlagL0[xSb][ySb] + PredFlagL1[xSb][ySb].
[0104] Condition (3): One motion vector is used to predict the coded block containing sample p0, one motion vector is used to predict the coded block containing sample q0, and the absolute difference between the horizontal or vertical components of the motion vectors used is greater than or equal to 8 in units of 1 / 16 luma samples.
[0105] Condition (4): Two motion vectors and two different reference images are used to predict the coded block containing sample p0, two motion vectors for the same two reference images are used to predict the coded block containing sample q0, and the absolute difference between the horizontal or vertical components of the two motion vectors used in the prediction of the two coded blocks for the same reference image is greater than or equal to 8 in units of 1 / 16 luma samples.
[0106] Condition (5): Two motion vectors for the same reference image are used to predict the coded block containing sample p0, two motion vectors for the same reference image are used to predict the coded block containing sample q0, and both of the following conditions hold. Condition (5.1): The absolute difference between the horizontal or vertical components of the List 0 motion vectors used in the prediction of the two coded blocks is greater than or equal to 8 in units of 1 / 16 luma samples, or the absolute difference between the horizontal or vertical components of the List 1 motion vectors used in the prediction of the two coded blocks is greater than or equal to 8 in units of 1 / 16 luma samples. Condition (5.2): The absolute difference between the horizontal or vertical components of the List 0 motion vector used in the prediction of the coded block containing sample p0 and the List 1 motion vector used in the prediction of the coded block containing sample q0 is greater than or equal to 8 in units of 1 / 16 luma samples, or the absolute difference between the horizontal or vertical components of the List 1 motion vector used in the prediction of the coded block containing sample p0 and the List 0 motion vector used in the prediction of the coded block containing sample q0 is greater than or equal to 8 in units of 1 / 16 luma samples.
[0107] Otherwise, if none of the previous seven scenarios are satisfied, in Scenario 8, the variable bS[xDi][yDj] is set to be equal to 0. After deriving the boundary filter strength, the number of samples to be filtered is obtained, and deblocking filtering is applied to the samples. Note that when a block is encoded in palette mode, the number of samples is set to be equal to 0. This means that deblocking filtering is not applied to blocks encoded in palette mode.
[0108] As previously mentioned, the construction of the current palette of a block consists of two parts. First, the entries in the current palette can be predicted from the palette predictor. For each entry in the palette predictor, a reuse flag is issued to indicate whether the entry is included in the current palette. Second, the component values of the current palette entries can be signalled directly. After obtaining the current palette, the palette predictor is updated using the current palette. Figure 7 Part of Section 7.3.10.6 ("Palette Coding Syntax") of VVC Draft 8 is shown. When parsing the syntax of a palette-coded block, the reuse flag (i.e., Figure 7 "palette_predictor_run" in 701) is decoded first, followed by the component values of the new palette entries (i.e., Figure 7 "num_signalled_palette_entries" 702 and "new_palette_entries" 703) in. In VVC Draft 8, the number of palette predictor entries (i.e., Figure 7 "PredictorPaletteSize[startComp]" 704) in needs to be known before parsing the reuse flag. This means that when two adjacent blocks are both encoded in palette mode, the syntax of the second block cannot be parsed until the palette predictor update process of the first block has been completed.
[0109] However, during the conventional palette predictor update process, the reuse flag of each entry needs to be checked. In the worst case, up to 63 checks are required. These traditional designs may not be suitable for hardware for at least two reasons. First, context-based adaptive binary arithmetic coding (CABAC) parsing needs to wait until the palette predictor is fully updated. Generally, CABAC parsing is the slowest module in hardware. This may reduce the CABAC throughput. Second, when implementing the palette predictor update process during the CABAC parsing stage, it may cause trouble to the hardware.
[0110] In addition, another problem in these conventional designs is that the boundary filter strength is not defined for the palette mode. When one of the adjacent blocks is encoded in palette mode and the other adjacent block is encoded in IBC or inter prediction mode, the boundary filter strength is undefined.
[0111] Embodiments of the present disclosure provide implementations that address one or more of the above problems. These implementations can improve the palette predictor update process, thereby enhancing the efficiency, speed, and resource consumption of systems implementing the above palette prediction process or similar processes.
[0112] In some embodiments, the palette predictor update process is simplified to reduce its complexity, thereby freeing up hardware resources. Figure 8 A flowchart of a palette predictor update process 800 according to an embodiment of the present disclosure is shown. Method 800 may be performed by one or more software or hardware components of an encoder (e.g., via process 200A of Figure 2A or process 200B of Figure 2B ), a decoder (e.g., via process 300A of Figure 3A or process 300B of Figure 3B ), or a device (e.g., device 400 of Figure 4 ). For example, a processor (e.g., processor 402 of Figure 4 ) may execute method 800. In some embodiments, method 800 may be implemented by a computer program product included in a computer-readable medium, the computer program product including computer-executable instructions, such as program code, executed by a computer (e.g., device 400 of Figure 4 ). As Figure 8 shown, method 800 may include the following steps 802 - 804.
[0113] In step 802, when updating the palette predictor, all palette entries of the current palette are added to the front of the new palette predictor as a first set of entries. In step 804, all palette entries from the previous palette predictor are added to the end of the new palette predictor, regardless of whether these entries are reused in the current palette, as a second set of entries, and the second set of entries follows the first set of entries.
[0114] For example, Figure 9 provides a simplified palette predictor update process consistent with the process described in Figure 8 . The current palette is generated by process 901 and process 902. The current palette includes entries reused from the previous palette predictor and newly signaled palette entries. In process 903, all palette entries of the current palette are the first set of entries of the new palette predictor (corresponding to step 802). In process 904, the palette entries from the previous palette predictor are the second set of entries of the new palette predictor, regardless of whether these entries are reused in the current palette (corresponding to step 804).
[0115] The advantage is that, without checking the value of the reuse flag, the size of the new palette predictor is calculated by adding the size of the current palette and the size of the previous palette predictor. Since the update process is much simpler than the conventional design in VVC Draft 8 (such as Figure 10 as shown), the CABAC throughput can be increased.
[0116] For example, Figure 11 shows an exemplary decoding process for the palette mode consistent with an embodiment of the present disclosure. Figure 10 the conventional design of Figure 11 and the changes between the disclosed design include the removed part 1101 highlighted by the deleted text.
[0117] In some embodiments, although the palette predictor update process is simplified, due to the lack of checking the reuse flag, the palette predictor may have two or more identical entries (i.e., redundancy in the palette predictor). Although still an improvement, this may mean that the prediction efficiency of the palette predictor may be reduced.
[0118] Figure 12 shows a flowchart of another palette predictor update process 1200 according to an embodiment of the present disclosure. Method 1200 may be performed by an encoder (e.g., through Figure 2A process 200A or Figure 2B process 200B), a decoder (e.g., through Figure 3A process 300A or Figure 3B process 300B) or one or more software or hardware components of a device (e.g., Figure 4 device 400). For example, a processor (e.g., Figure 4 processor 402) may perform method 1200. In some embodiments, method 1200 may be implemented by a computer program product included in a computer-readable medium, the computer program product including computer-executable instructions executed by a computer (e.g., Figure 4 device 400), such as program code. As Figure 12 shown, method 1200 may include the following steps 1202 - 1206.
[0119] In this exemplary embodiment, to remove redundancy in the palette predictor and keep the palette predictor update process simple, the previous palette predictor entries between the first reuse entry and the last reuse entry are directly discarded. Only the entries from the first entry to the first reuse entry and from the last reuse entry of the previous palette predictor to the last entry of the previous palette predictor are added to the new palette predictor. In summary, the palette predictor update process is modified as follows: In step 1202, each entry of the current palette is added to the new palette predictor as the first set of entries. In step 1204, each entry from the first to the first reuse entry of the previous palette predictor is added to the new palette predictor as the second set of entries after the first set of entries. In step 1206, each entry from the last reuse entry of the previous palette predictor to the last entry of the previous palette predictor is added to the new palette predictor as the third set of entries after the second set of entries.
[0120] For example, Figure 13 a simplified palette predictor update process consistent with the Figure 12 process described is provided. In process 1301, all palette entries of the current palette are added to the palette predictor as the first set of entries of the new palette predictor (corresponding to step 1202). In process 1302, each entry from the first to the first reuse entry of the previous palette predictor is added to the new palette predictor as the second set of entries after the first set of entries (corresponding to step 1204). In process 1303, each entry from the last reuse entry of the previous palette predictor to the last entry of the previous palette predictor is added to the new palette predictor as the third set of entries after the second set of entries (corresponding to step 1206). In some embodiments, the third set of entries is after the first set of entries, and the second set of entries is after the third set of entries.
[0121] Note that when parsing the reuse flag, the first reuse entry and the last entry can be derived. There is no need to check the value of the reuse flag. The size of the new palette predictor is calculated by adding the size of the current palette, the size from 0 to the first reuse entry, and the size from the last reuse entry to the last entry. Figure 14 An exemplary palette coding syntax consistent with an embodiment of the present application is shown. Figure 7 The conventional design of Figure 14 and the design of the present disclosure in Figure 15 include an additional part marked by 1401. An exemplary decoding process for the palette mode according to an embodiment of the present application is shown. Figure 10 The conventional design of Figure 15Changes between the designs of the present disclosure in [ ] include removal portions 1501 with deleted text highlighting and added portions 1502 with markings.
[0122] In previously disclosed embodiments, although the palette predictor update process is simplified, it needs to be implemented in the CABAC parsing stage in the hardware design. In some embodiments, the number of reuse flags is set to a fixed value. Thus, CABAC can continue parsing without waiting for the palette predictor update process. Additionally, the palette predictor update process can be implemented in different pipeline stages outside the CABAC parsing stage, which provides greater flexibility for the hardware design.
[0123] Figure 16 A schematic diagram of a decoder hardware design for the palette mode is shown. Exemplarily shown are data structures 1601 and data structures 1602 for CABAC and decoding palette pixels. The decoder hardware design 1603 includes a predictor update module 1631 and a CABAC parsing module 1632, where the predictor update process and the CABAC parsing process are parallel. Thus, CABAC can continue parsing without waiting for the palette update process.
[0124] For each coding block, to set the number of reuse flags to a fixed value, in some embodiments, at the start of each stripe for the non-wavefront case and at the start of each CTU row for the wavefront case, the size of the palette predictor is initialized to a predefined value. The predefined value is set to the maximum size of the palette predictor. In one example, depending on the stripe type and the dual-tree mode setting, the number of reuse flags is 31 or 63. When the stripe type is an I-stripe and the dual-tree mode is enabled (referred to as case 1), the number of reuse flags is set to 31. Otherwise (the stripe type is a B- / P-stripe or the stripe type is an I-stripe and the dual-tree mode is not enabled, (referred to as case 2)), the number of reuse flags is set to 63. The reuse flags in the two cases can also use other numbers, and it may be beneficial to maintain a factor-of-2 relationship between the number of reuse flags in case 1 and the number of reuse flags in case 2. Additionally, when initializing the palette predictor, the value of each entry and each component is set to 0 or (1<<(sequence bit depth - 1)).
[0125] For example, Figure 17 An exemplary decoding process for the palette mode according to an embodiment of the present disclosure is shown. Figure 10 The conventional design of [ ] and Figure 17 Changes between the designs of the present disclosure in [ ] include removal portions 1701 with deleted text highlighting and added portions 1702 with markings. Similarly, Figure 19 An exemplary initialization process consistent with an embodiment of the present disclosure is shown.Figure 18 between the conventional design of Figure 19 and the design of the present disclosure of
[0126] In some embodiments, when the reuse flag is set to a fixed value, or in an implementation Figure 9 of Figure 20 , some redundancy can be introduced into the palette predictor. In some embodiments, this may not be a problem because those redundant entries are never selected for prediction on the encoder side. An example of the update of the palette predictor is as Figure 21 and Figure 22 shown. Figure 20 shows an example according to the process specified in VVC draft 8. Figure 21 shows an example of some embodiments according to the method proposed in Figure 9 (referred to as the first embodiment). And Figure 22 shows another example of some implementations according to an embodiment, where the reuse flag is set to a fixed value (referred to as the third embodiment). By Figure 20 and Figure 21 comparison with 22 shows that assuming no new palette entries are signaled, the method proposed in the first embodiment ( Figure 21 ) can increase the number of bits of the reuse flag 2101 used for signaling. While in the third embodiment ( Figure 22 ), the method proposed keeps the number of bits of the reuse flag 2201 used for signaling the same as the number of bits of the reuse flag 2001 in the method in the VVC draft 8 design ( Figure 20 ).
[0127] In some embodiments, the palette predictor is first initialized to a fixed value. This means that there may be redundant entries in the palette predictor. Although redundant entries may not be a problem as described above, the design in this embodiment does not prevent the encoder from using those redundant entries. If the encoder selects one of these redundant entries, the encoding performance of the palette predictor and thus the encoding performance of the palette mode may be reduced. To prevent this, in some embodiments, bitstream consistency is added when signaling the reuse flag. In some embodiments, when signaling the reuse flag, the bitstream consistency having a value of the size of the palette predictor is equal to the maximum size of the palette predictor. More specifically, range constraints are added to the binarized value of the reuse flag.
[0128] For example, Figure 23 shows an exemplary palette encoding syntax consistent with the embodiments of the present disclosure. Figure 7 between the conventional design ofFigure 23 Changes between the designs of the present disclosure include the removed part 2301 highlighted with deleted text and the added part 2302 marked. Similarly, Figure 25 illustrates example palette coding semantics consistent with embodiments of the present disclosure. Figure 24 The conventional design of Figure 25 Changes between the designs of the present disclosure include the removed part 2501 highlighted with deleted text and the added part 2502 marked.
[0129] In addition, the present disclosure provides the following method to solve the problem of the deblocking filter in the palette mode.
[0130] In some embodiments, the palette mode is regarded as a subset of the intra prediction mode. Therefore, if one of the adjacent blocks is encoded in the palette mode, the boundary filter strength is set to be equal to 2. More specifically, the part of the VVC draft 8 specification that details the process of calculating the boundary filtering strength is changed, so that scenario 3 now states: "If sample p0 or q0 is in the coding block of the coding unit encoded in the intra prediction mode or the palette mode, bS[xDi][yDj] is set to be equal to 2." The added statement is underlined.
[0131] In some embodiments, the palette mode is regarded as an independent coding mode. When the coding modes of two adjacent blocks are different, the boundary filter strength is set to be equal to 1. More specifically, the part of the VVC draft 8 specification that details the process of calculating the boundary filtering strength has been changed, so that scenario 6 now states: "The prediction mode of the coding sub-block containing sample p0 is different from the prediction mode of the coding sub-block containing sample q0 bS[xDi][yDj] is set to be equal to 1." The deleted statement is deleted.
[0132] In some embodiments, the palette mode is regarded as an independent coding mode. Similar to the BDPCM mode setting, when one of the coding blocks is in the palette mode, the boundary filter strength is set to be equal to 0. More specifically, the part of the VVC draft 8 specification that details the process of calculating the boundary filtering strength is changed to insert a new scenario between scenario 3 and scenario 4, which states: "Otherwise, if the block edge is also the coding block edge, and sample p0 or q0 is in the coding block where pred_mode_plt_flag is equal to 1, then bS[xDi][yDj] is set to be equal to 0." Using this modified process, there will be 9 scenarios, and scenarios 4 - 8 are renumbered as scenarios 5 - 9. It should be noted that a block encoded in the palette mode is equivalent to pred_mode_plt_flag being equal to 1.
[0133] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) to perform the above method. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAM, PROM, and EPROM, FLASH-EPROM or any other flash memory, NVRAM, caches, registers, any other storage chip or cartridge storage, and networked versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0134] It should be noted that relational terms such as "first" and "second" herein are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "comprising", "having", "including", and "include" and other similar forms are equivalent in meaning and are open-ended, because one or more items following any of these words do not mean an exhaustive list of the one or more items or are limited to the listed one or more items.
[0135] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations, unless infeasible. For example, if it is stated that a database may include A or B, then the database may include A, or B, or A and B, unless otherwise explicitly stated or infeasible. As a second example, if it is stated that a database may include A, B, or C, then the database may include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise explicitly stated or infeasible.
[0136] It should be understood that the above embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. The software can perform the disclosed method when executed by a processor. The computing units and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that the above multiple modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0137] The embodiments can be further described using the following terms:
[0138] 1. A video data processing method, comprising:
[0139] Processing one or more coding units using one or more palette predictors, wherein a palette predictor of the one or more palette predictors is updated by:
[0140] Adding all palette entries of the current palette as a first set of entries of the palette predictor; and
[0141] Adding a plurality of palette entries from a previous palette predictor as a second set of entries of the palette predictor, regardless of the value of the reuse flag of the palette entries in the previous palette predictor, wherein the second set of entries follows the first set of entries.
[0142] 2. The method according to clause 1, wherein each palette entry in the first set of entries and the second set of entries of the palette predictor includes a reuse flag.
[0143] 3. The method according to clause 1, further comprising:
[0144] Receiving a video frame for processing, and
[0145] Generating the one or more coding units of the video frame.
[0146] 4. The method according to any one of clauses 1 to 3, wherein each palette entry of the palette predictor has a corresponding reuse flag, and
[0147] wherein, for a corresponding coding unit, the number of reuse flags used for the palette predictor is set to a fixed number.
[0148] 5. A video data processing method, comprising:
[0149] Processing one or more coding units using one or more palette predictors, wherein a palette predictor of the one or more palette predictors is updated by:
[0150] Adding all palette entries of the current palette as a first set of entries of the palette predictor;
[0151] Add one or more palette entries of a previous palette predictor within a first range to the palette predictor as a second set of one or more entries of the palette predictor, where the first range starts at a first palette entry of the previous palette predictor and ends at a first palette entry of the previous palette predictor having a reuse flag group, and add one or more palette entries of the previous palette predictor within a second range to the palette predictor as a third set of one or more entries of the palette predictor, where the second range starts at a last palette entry of the previous palette predictor having a reuse flag group and ends at a last palette entry of the previous palette predictor, and the second set of entries and the third set of entries are after the first set of entries.
[0152] 6. The method according to clause 5, wherein each palette entry of the first set of entries, the second set of entries, and the third set of entries of the palette predictor includes a reuse flag.
[0153] 7. The method according to clause 5, further comprising:
[0154] Receiving a video frame for processing, and
[0155] Generating one or more coding units of the video frame.
[0156] 8. The method according to any one of clauses 5 to 7, wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and
[0157] wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number.
[0158] 9. A video data processing method, comprising:
[0159] Receiving a video frame for processing;
[0160] Generating one or more coding units of the video frame; and
[0161] Processing one or more coding units using one or more palette predictors having a plurality of palette entries,
[0162] wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and
[0163] wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number.
[0164] 10. The method according to clause 9, wherein the palette predictor is updated as follows:
[0165] Adding all the palette entries of the current palette as the first set of entries of the palette predictor; and
[0166] Adding the entries from the previous palette predictor that are not reused in the current palette as the second set of entries of the palette predictor, where the second set of entries follows the first set of entries.
[0167] 11. The method according to clause 9, wherein the fixed number is set based on the slice type and the dual-tree mode setting.
[0168] 12. The method according to clause 9, wherein the size of the one or more palette predictors is initialized to a predetermined value at the start of a slice for a non-wavefront case.
[0169] 13. The method according to clause 9, wherein the size of the one or more palette predictors is initialized to a predetermined value at the start of a coding unit row for a wavefront case.
[0170] 14. The method according to clause 12 or 13, further comprising: adding bitstream consistency with the value of the palette predictor size, which is equal to the maximum size of the palette predictor when the reuse flag is signaled.
[0171] 15. The method according to any one of clauses 12 to 14, wherein when initializing the one or more palette predictors, the value of each entry and each component is set to 0 or (1 << (sequence bit depth - 1)).
[0172] 16. The method according to any one of clauses 9 to 15, further comprising adding range constraints to the binarized value of the reuse flag.
[0173] 17. An apparatus for video data processing, the apparatus comprising:
[0174] A memory for storing instructions; and
[0175] A processor coupled to the memory and configured to execute the instructions to cause the apparatus to perform the following operations:
[0176] Processing one or more coding units using one or more palette predictors, wherein a palette predictor of the one or more palette predictors is updated as follows:
[0177] Adding all the palette entries of the current palette as the first set of entries of the palette predictor; and
[0178] Add a plurality of palette entries from a previous palette predictor as a second set of entries of the palette predictor, regardless of the value of the reuse flag of the palette entries in the previous palette predictor, wherein the second set of entries follows the first set of entries.
[0179] 18. The apparatus according to clause 17, wherein each palette entry in the first set of entries and the second set of entries of the palette predictor includes a reuse flag.
[0180] 19. The apparatus according to clause 17, wherein the processor is further configured to execute the instructions to cause the apparatus to perform:
[0181] Receive a video frame for processing, and
[0182] Generate one or more coding units of the video frame.
[0183] 20. The apparatus according to any one of clauses 17 to 19, wherein each palette entry of the palette predictor has a corresponding reuse flag, and
[0184] wherein, for a corresponding coding unit, the number of reuse flags for the palette predictor is set to a fixed number.
[0185] 21. An apparatus for video data processing, the apparatus comprising:
[0186] A memory for storing instructions; and
[0187] A processor coupled to the memory and configured to execute the instructions to cause the apparatus to perform the following operations:
[0188] Process one or more coding units using one or more palette predictors, wherein one of the one or more palette predictors is updated by:
[0189] Add all palette entries of the current palette as a first set of entries of the palette predictor;
[0190] Add one or more palette entries of a previous palette predictor within a first range to the palette predictor as a second set of one or more entries of the palette predictor, wherein the first range starts at a first palette entry of the previous palette predictor and ends at a first palette entry of the previous palette predictor having a group of reuse flags, and
[0191] Add one or more palette entries of a previous palette predictor within a second range to the palette predictor as a third set of one or more entries of the palette predictor, where the second range starts and ends at the last palette entry of the previous palette predictor having a reuse flag group, and the second set of entries and the third set of entries are after the first set of entries.
[0192] 22. The apparatus according to clause 21, wherein each palette entry in the first set of entries and the second set of entries of the palette predictor includes a reuse flag.
[0193] 23. The apparatus according to clause 21, wherein the processor is further configured to execute the instructions to cause the apparatus to perform:
[0194] Receive a video frame for processing, and
[0195] Generate one or more coding units of the video frame.
[0196] 24. The apparatus according to any one of clauses 21 to 23, wherein each palette entry of the palette predictor has a corresponding reuse flag, and
[0197] wherein, for a corresponding coding unit, the number of reuse flags for the palette predictor is set to a fixed number.
[0198] 25. An apparatus for video data processing, the apparatus comprising:
[0199] A memory for storing instructions; and
[0200] A processor coupled to the memory and configured to execute the instructions to cause the apparatus to perform the following operations:
[0201] Receive a video frame for processing;
[0202] Generate one or more coding units of the video frame; and
[0203] Process one or more coding units using one or more palette predictors having a plurality of palette entries,
[0204] wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and
[0205] wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number.
[0206] 26. The apparatus according to clause 25, wherein the processor is further configured to execute the instructions to cause the apparatus to perform:
[0207] Update the palette predictor by:
[0208] Adding all the palette entries of the current palette as a first set of entries of the palette predictor; and
[0209] Adding the entries from the previous palette predictor that are not reused in the current palette as a second set of entries of the palette predictor, wherein the second set of entries follows the first set of entries.
[0210] 27. The apparatus according to clause 25, wherein the fixed quantity is set based on the slice type and the dual-tree mode setting.
[0211] 28. The apparatus according to clause 25, wherein the size of the one or more palette predictors is initialized to a predetermined value at the start of a slice for a non-wavefront case.
[0212] 29. The apparatus according to clause 25, wherein the size of the one or more palette predictors is initialized to a predetermined value at the start of a coding unit row for a wavefront case.
[0213] 30. The apparatus according to clause 28 or 29, wherein the processor is further configured to execute the instructions to cause the apparatus to perform: adding bitstream consistency with a value of the size of the palette predictor, which is equal to the maximum size of the palette predictor when the reuse flag is signaled.
[0214] 31. The apparatus according to any one of clauses 28 to 30, wherein when initializing the one or more palette predictors, the value of each entry and each component is set to 0 or (1 << (sequence bit depth - 1)).
[0215] 32. The apparatus according to any one of clauses 25 to 31, wherein the processor is further configured to execute the instructions to cause the apparatus to perform: adding range constraints to the binarized value of the reuse flag.
[0216] 33. A non-transitory computer-readable medium storing a set of instructions that can be executed by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising:
[0217] Processing one or more coding units using one or more palette predictors, wherein one of the one or more palette predictors is updated by:
[0218] Add all the palette entries of the current palette as the first set of entries of the palette predictor; and
[0219] Add a plurality of palette entries from a previous palette predictor as the second set of entries of the palette predictor, regardless of the value of the reuse flag of the palette entries in the previous palette predictor, wherein the second set of entries is after the first set of entries.
[0220] 34. The non-transitory computer-readable medium according to clause 33, wherein each palette entry in the first set of entries and the second set of entries of the palette predictor includes a reuse flag.
[0221] 35. The non-transitory computer-readable medium according to clause 33, wherein the method further comprises:
[0222] Receiving a video frame for processing, and
[0223] Generating one or more coding units of the video frame.
[0224] 36. The non-transitory computer-readable medium according to any one of clauses 33 to 35, wherein each palette entry of the palette predictor has a corresponding reuse flag, and
[0225] wherein, for a corresponding coding unit, the number of reuse flags for the palette predictor is set to a fixed number.
[0226] 37. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising:
[0227] Processing one or more coding units using one or more palette predictors, wherein one of the one or more palette predictors is updated by:
[0228] Adding all the palette entries of the current palette as the first set of entries of the palette predictor;
[0229] Adding one or more palette entries of a previous palette predictor within a first range to the palette predictor as the second set of one or more entries of the palette predictor, wherein the first range starts at the first palette entry of the previous palette predictor and ends at the first palette entry of the previous palette predictor having a group of reuse flags, and
[0230] Adding one or more palette entries of a previous palette predictor within a second range to the palette predictor as one or more entries of a third set of the palette predictor, wherein the second range starts at the last palette entry of the previous palette predictor having a reuse flag group and ends at the last palette entry of the previous palette predictor, and the second set of entries and the third set of entries are after the first set of entries.
[0231] 38. The non-transitory computer-readable medium according to clause 37, wherein each palette entry in the first set of entries and the second set of entries of the palette predictor includes a reuse flag.
[0232] 39. The non-transitory computer-readable medium according to clause 37, wherein the method further comprises:
[0233] Receiving a video frame for processing, and
[0234] Generating one or more coding units of the video frame.
[0235] 40. The non-transitory computer-readable medium according to any one of clauses 37 to 39, wherein each palette entry of the palette predictor has a corresponding reuse flag, and
[0236] wherein, for a corresponding coding unit, the number of reuse flags for the palette predictor is set to a fixed number.
[0237] 41. A non-transitory computer-readable medium storing a set of instructions executable by one or more processors of a device to cause the device to initiate a method for performing video data processing, the method comprising:
[0238] Receiving a video frame for processing;
[0239] Generating one or more coding units of the video frame; and
[0240] Processing one or more coding units using one or more palette predictors having palette entries,
[0241] wherein each palette entry of the one or more palette predictors has a corresponding reuse flag, and
[0242] wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number.
[0243] 42. The non-transitory computer-readable medium according to clause 41, wherein the palette predictor is updated by:
[0244] Add all the palette entries of the current palette as the first set of entries of the palette predictor; and
[0245] Add the entries from the previous palette predictor that are not reused in the current palette as the second set of entries of the palette predictor, where the second set of entries follows the first set of entries.
[0246] 43. The non-transitory computer-readable medium according to clause 41, wherein the fixed number is set based on the stripe type and the dual-tree mode setting.
[0247] 44. The non-transitory computer-readable medium according to clause 41, wherein the size of the one or more palette predictors is initialized to a predetermined value at the start of a stripe for a non-wavefront case.
[0248] 45. The non-transitory computer-readable medium according to clause 41, wherein the size of the one or more palette predictors is initialized to a predetermined value at the start of a coding unit row for a wavefront case.
[0249] 46. The non-transitory computer-readable medium according to clause 44 or 45, further comprising adding bitstream consistency with the value of the palette predictor size, which is equal to the maximum size of the palette predictor when the reuse flag is signaled.
[0250] 47. The non-transitory computer-readable medium according to any one of clauses 44 to 46, wherein when initializing the one or more palette predictors, the value of each entry and each component is set to 0 or (1<<(sequence bit depth - 1)).
[0251] 48. The non-transitory computer-readable medium according to any one of clauses 41 to 47, further comprising adding range constraints to the binarized value of the reuse flag.
[0252] 49. A method for a deblocking filter in palette mode, comprising:
[0253] Receiving a video frame for processing;
[0254] Generating one or more coding units for the video frame, wherein each coding unit in the one or more coding units has one or more coding blocks; and
[0255] In response to at least the first coding block of two adjacent coding blocks being encoded in palette mode, setting the boundary filter strength to 2.
[0256] 50. A method for a deblocking filter in palette mode, comprising:
[0257] Receive a video frame for processing;
[0258] Generate one or more coding units for the video frame, wherein each coding unit of the one or more coding units has one or more coding blocks; and
[0259] In response to at least a first coding block of two adjacent coding blocks being coded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode, set a boundary filter strength to 1.
[0260] 51. An apparatus for processing video data, the apparatus comprising:
[0261] A memory for storing instructions; and
[0262] A processor coupled to the memory and configured to execute the instructions to cause the apparatus to perform the following operations:
[0263] Receive a video frame for processing;
[0264] Generate one or more coding units for the video frame, wherein each coding unit of the one or more coding units has one or more coding blocks; and
[0265] In response to at least a first coding block of two adjacent coding blocks being coded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode, set a boundary filter strength to 1.
[0266] 52. A non-transitory computer-readable medium storing a set of instructions that can be executed by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising:
[0267] Receive a video frame for processing;
[0268] Generate one or more coding units for the video frame, wherein each coding unit of the one or more coding units has one or more coding blocks; and
[0269] In response to at least a first coding block of two adjacent coding blocks being coded in a palette mode and a second coding block of the two adjacent coding blocks having a coding mode different from the palette mode, set a boundary filter strength to 1.
[0270] In the foregoing specification, embodiments have been described with reference to numerous specific details, which may vary according to implementation. Certain modifications and changes may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. The specification and examples are considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims. The sequence of steps shown in the drawings is for illustrative purposes only and is not intended to be limited to any particular sequence of steps. Accordingly, those skilled in the art will appreciate that these steps may be performed in a different order while implementing the same method.
[0271] In the drawings and specification, exemplary embodiments have been disclosed. However, many variations and modifications may be made to these embodiments. Accordingly, although specific terms have been employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. A video data processing method, comprising: Receiving a video frame for processing; Generating one or more coding units of the video frame; and Processing one or more coding units using one or more palette predictors having a plurality of palette entries, Among them, Each palette entry of the one or more palette predictors having a corresponding reuse flag, and Wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number, and the fixed number of the reuse flags is set to 31 or 63 based on the slice type and the dual-tree mode setting.
2. The method according to claim 1, wherein The palette predictor is updated by: Adding all palette entries of the current palette as a first set of entries of the palette predictor; and Adding entries from a previous palette predictor that are not reused in the current palette as a second set of entries of the palette predictor, wherein the second set of entries is after the first set of entries.
3. The method according to claim 1, wherein The palette predictor is updated by: Adding all palette entries of the current palette as a first set of entries of the palette predictor; and Adding palette entries from a previous palette predictor as a second set of entries of the palette predictor, regardless of the value of the reuse flag of the palette entries in the previous palette predictor, wherein the second set of entries is after the first set of entries.
4. The method according to claim 1, wherein The palette predictor is updated by: Adding all palette entries of the current palette as a first set of entries of the palette predictor; Adding one or more palette entries of a previous palette predictor within a first range to the palette predictor as a second set of one or more entries of the palette predictor, wherein the first range starts at the first palette entry of the previous palette predictor and ends at the first palette entry of the previous palette predictor having a group of reuse flags, and Adding one or more palette entries of a previous palette predictor within a second range to the palette predictor as a third set of one or more entries of the palette predictor, wherein the second range starts at the last palette entry of the previous palette predictor having a group of reuse flags and ends at the last palette entry of the previous palette predictor, and the second set of entries and the third set of entries are after the first set of entries.
5. The method according to claim 1, wherein Setting the fixed number to 31 or 63 based on the slice type and the dual-tree mode setting further includes: When the slice type is an I-slice and the dual-tree mode is on, setting the fixed number to 31; Or, when the slice type is a B- / P-slice or the slice type is an I-slice and the dual-tree mode is off, setting the fixed number to 63.
6. The method according to claim 1, characterized in that, The size of the one or more palette predictors is initialized to a predetermined value at the start of a slice for a non-wavefront case.
7. An apparatus for performing video data processing, the apparatus comprising: A memory for storing instructions; and One or more processors configured to execute the instructions to cause the apparatus to perform the following operations: Receive a video frame for processing; Generate one or more coding units of the video frame; and Process one or more coding units using one or more palette predictors having a plurality of palette entries, Among them, Each palette entry of the one or more palette predictors having a corresponding reuse flag, and Wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number, and the fixed number of reuse flags is set to 31 or 63 based on the slice type and the dual-tree mode setting.
8. The device according to claim 7, characterized in that, The one or more processors are further configured to execute the instructions to cause the apparatus to perform: Update the palette predictor by: Adding all palette entries of the current palette as a first set of entries of the palette predictor; and Adding entries from a previous palette predictor that are not reused in the current palette as a second set of entries of the palette predictor, wherein the second set of entries is after the first set of entries.
9. The device according to claim 7, characterized in that, The one or more processors are further configured to execute the instructions to cause the apparatus to perform: Update the palette predictor by: Adding all palette entries of the current palette as a first set of entries of the palette predictor; and Adding palette entries from a previous palette predictor as a second set of entries of the palette predictor regardless of the value of the reuse flag of the palette entries in the previous palette predictor, wherein the second set of entries is after the first set of entries.
10. The device according to claim 7, characterized in that, The one or more processors are further configured to execute the instructions to cause the apparatus to perform: Update the palette predictor by: Adding all palette entries of the current palette as a first set of entries of the palette predictor; Adding one or more palette entries of a previous palette predictor within a first range to the palette predictor as a second set of one or more entries of the palette predictor, wherein the first range starts at a first palette entry of the previous palette predictor and ends at a first palette entry of the previous palette predictor having a group of reuse flags, and Adding one or more palette entries of a previous palette predictor within a second range to the palette predictor as a third set of one or more entries of the palette predictor, wherein the second range starts at a last palette entry of the previous palette predictor having a group of reuse flags and ends at a last palette entry of the previous palette predictor, and the second set of entries and the third set of entries are after the first set of entries.
11. The device according to claim 7, wherein, Setting the fixed number to 31 or 63 based on the slice type and the dual-tree mode setting further includes: When the slice type is an I-slice and the dual-tree mode is enabled, setting the fixed number to 31; Alternatively, when the slice type is a B- / P-slice or the slice type is an I-slice and the dual-tree mode is off, the fixed number is set to 63.
12. The device according to claim 7, characterized in that, The size of the one or more palette predictors is initialized to a predefined value at the start of the slice in the non-wavefront case.
13. A non-transitory computer-readable medium storing a video bitstream, the bitstream being for processing according to the following method, the method comprising: generating one or more coding units of a video frame; and processing the one or more coding units using one or more palette predictors having a plurality of palette entries, Among them, each palette entry of the one or more palette predictors having a corresponding reuse flag, and wherein, for a corresponding coding unit, the number of reuse flags for each palette predictor is set to a fixed number, and the fixed number of reuse flags is set to 31 or 63 based on the slice type and dual-tree mode setting.
14. The non-transitory computer-readable medium according to claim 13, wherein, The palette predictor is updated by: adding all palette entries of the current palette as a first set of entries of the palette predictor; and adding entries from a previous palette predictor that are not reused in the current palette as a second set of entries of the palette predictor, wherein the second set of entries follows the first set of entries.
15. The non-transitory computer-readable medium according to claim 13, wherein, The palette predictor is updated by: adding all palette entries of the current palette as a first set of entries of the palette predictor; and adding palette entries from a previous palette predictor regardless of the value of the reuse flag of the palette entries in the previous palette predictor, wherein the second set of entries follows the first set of entries.
16. The non-transitory computer-readable medium according to claim 13, wherein, The palette predictor is updated by: adding all palette entries of the current palette as a first set of entries of the palette predictor; adding one or more palette entries of a previous palette predictor within a first range to the palette predictor as a second set of one or more entries of the palette predictor, wherein the first range starts at the first palette entry of the previous palette predictor and ends at the first palette entry of the previous palette predictor having a group of reuse flags, and adding one or more palette entries of a previous palette predictor within a second range to the palette predictor as a third set of one or more entries of the palette predictor, wherein the second range starts at the last palette entry of the previous palette predictor having a group of reuse flags and ends at the last palette entry of the previous palette predictor, and the second set of entries and the third set of entries follow the first set of entries.
17. The non-transitory computer-readable medium according to claim 13, wherein, Setting the fixed number to 31 or 63 based on the slice type and dual-tree mode setting further comprises: when the slice type is an I-slice and the dual-tree mode is on, setting the fixed number to 31; alternatively, when the slice type is a B- / P-slice or the slice type is an I-slice and the dual-tree mode is off, setting the fixed number to 63.
18. A method for a deblocking filter in palette mode, the method being executed in an encoder, comprising: Receiving a video frame for processing; Generate two or more coding units for the video frame, wherein, Each of the two or more coding units has one or more coding blocks; Determining whether at least one of a first coding block in two adjacent coding blocks and a second coding block in the two adjacent coding blocks is encoded in an intra prediction mode; In response to the first coding block and the second coding block not being encoded in an intra prediction mode, determining whether at least one of the first coding block and the second coding block is encoded in a CIIP (Combined Inter and Intra Prediction) mode; In response to the first coding block and the second coding block not being encoded in a CIIP mode, determining whether at least one of the first coding block and the second coding block is encoded in a palette mode; And In response to the first coding block being encoded in a palette mode and the second coding block having a coding mode different from the palette mode, setting the boundary filter strength to 1.
19. An apparatus for video data processing, applied to an encoder, the apparatus comprising: A memory for storing instructions; and One or more processors configured to execute the instructions to cause the apparatus to perform the following operations: Receiving a video frame for processing; Generate two or more coding units for the video frame, wherein, Each of the two or more coding units has one or more coding blocks; Determining whether at least one of a first coding block in two adjacent coding blocks and a second coding block in the two adjacent coding blocks is encoded in an intra prediction mode; In response to the first coding block and the second coding block not being encoded in an intra prediction mode, determining whether at least one of the first coding block and the second coding block is encoded in a CIIP (Combined Inter and Intra Prediction) mode; In response to the first coding block and the second coding block not being encoded in a CIIP mode, determining whether at least one of the first coding block and the second coding block is encoded in a palette mode; And In response to the first coding block being encoded in a palette mode and the second coding block having a coding mode different from the palette mode, setting the boundary filter strength to 1.
20. A non - transitory computer - readable medium storing a video bitstream, the bitstream being for processing according to the following method, the method comprising: Generate two or more coding units for video frames, wherein, Each of the two or more coding units has one or more coding blocks; Determining whether at least one of a first coding block in two adjacent coding blocks and a second coding block in the two adjacent coding blocks is encoded in an intra prediction mode; In response to the first coding block and the second coding block not being encoded in an intra prediction mode, determining whether at least one of the first coding block and the second coding block is encoded in a CIIP (Combined Inter and Intra Prediction) mode; In response to the first coding block and the second coding block not being encoded in a CIIP mode, determining whether at least one of the first coding block and the second coding block is encoded in a palette mode; And In response to the first coding block being coded in palette mode and the second coding block having a coding mode different from the palette mode, set the boundary filter strength to 1.
Citation Information
Patent Citations
Method and device for decoding with palette mode
US20200092546A1