Method, apparatus, storage medium, and program product for processing video data
By disabling the combination of LFNST and ACT in the VVC standard, using only one of LFNST or ACT for video encoding, solving the encoding process complexity and delay issues, achieving more efficient video compression.
Patent Information
- Application Number
- CN202180029638.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-09
- Filing Date
- 2021-06-09
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-06-09
AI Technical Summary
Existing video encoding technology In high-efficiency video encoding standards such as VVC/H.266, the combination of low-frequency inseparable transformation (LFNST) and adaptive color transformation (ACT) leads to an increase in the complexity of the encoding/decoding process, affecting efficiency and delay.
At the encoding layer video sequence (CLVS) and encoding unit (CU) levels, encoding is optimized by conditional enablement and bitstream consistency constraints by disabling the combination of LFNST and ACT.
Reduces the complexity of the encoding/decoding process, improves coding efficiency, reduces the delay of the decoding pipeline, and improves compression performance.
Smart Images

Figure CN115443655B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This disclosure claims priority to U.S. Provisional Application No. 63 / 036,499, filed on June 9, 2020, the entire content of which is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to video processing, and more particularly, to methods, apparatuses, storage media, and program products for processing video data. Background Art
[0004] Video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats using standardized video coding techniques, most commonly based on prediction, transformation, quantization, entropy coding, and loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the Audio Video Coding Standard (AVS) standard, etc., specify specific video coding formats and are formulated by standardization organizations. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is also getting higher and higher. Summary of the Invention
[0005] Embodiments of this disclosure provide a method for processing video data. The method includes: receiving one or more video sequences for processing; and encoding the one or more video sequences using only one of the Low - Frequency Non - Separable Transform (LFNST) and the Adaptive Color Transform (ACT).
[0006] Embodiments of this disclosure provide an apparatus for processing video data. The apparatus includes: a memory configured to store instructions; and a processor coupled to the memory and configured to execute the instructions to cause the apparatus to perform: receiving one or more video sequences for processing; and encoding the one or more video sequences using only one of the Low - Frequency Non - Separable Transform (LFNST) and the Adaptive Color Transform (ACT).
[0007] Embodiments of this disclosure provide a non - transitory computer - readable medium storing a set of instructions that can be executed by one or more processors of a device to cause the device to initiate a method for performing video data processing. The method includes: receiving one or more video sequences for processing; and encoding the one or more video sequences using only one of the Low - Frequency Non - Separable Transform (LFNST) and the Adaptive Color Transform (ACT).
[0008] Embodiments of the present disclosure provide an apparatus for processing video data, the apparatus comprising: a first receiving module configured to receive one or more video sequences for processing; and a first processing module configured to encode the one or more video sequences using only one of a low-frequency non-separable transform (LFNST) and an adaptive color transform (ACT).
[0009] Embodiments of the present disclosure provide a computer program product comprising a computer program which, when executed by a processor, implements any of the methods described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.
[0011] Figure 1 is a schematic diagram showing the structure of an example video sequence according to some embodiments of the present disclosure.
[0012] Figure 2A is a schematic diagram showing an exemplary encoding process of a hybrid video coding system consistent with embodiments of the present disclosure.
[0013] Figure 2B is a schematic diagram showing another exemplary encoding process of a hybrid video coding system consistent with embodiments of the present disclosure.
[0014] Figure 3A is a schematic diagram showing an exemplary decoding process of a hybrid video coding system consistent with embodiments of the present disclosure.
[0015] Figure 3B is a schematic diagram showing another exemplary decoding process of a hybrid video coding system consistent with embodiments of the present disclosure.
[0016] Figure 4 is a block diagram of an exemplary apparatus for encoding or decoding video according to some embodiments of the present disclosure.
[0017] Figure 5 shows an exemplary processing flow of an encoding process.
[0018] Figure 6 shows an exemplary processing flow of a decoding process.
[0019] Figure 7 shows an exemplary flowchart of an encoding method according to some embodiments of the present disclosure.
[0020] Figure 8 shows another exemplary flowchart of an encoding method according to some embodiments of the present disclosure.
[0021] Figure 9 Shows an exemplary SPS syntax according to some embodiments of the present disclosure.
[0022] Figure 10 Shows the exemplary semantics of the updated syntax element sps_act_enabled_flag according to some embodiments of the present disclosure.
[0023] Figure 11 Shows the exemplary semantics of the updated syntax element sps_act_enabled_flag according to some embodiments of the present disclosure.
[0024] Figure 12 Shows an exemplary flowchart of an encoding method according to some embodiments of the present disclosure.
[0025] Figure 13 Shows an exemplary SPS syntax with an updated sps_lfnst_enabled_flag according to some embodiments of the present disclosure.
[0026] Figure 14 Shows the exemplary semantics of the updated syntax element sps_lfnst_enabled_flag according to some embodiments of the present disclosure.
[0027] Figure 15 Shows the exemplary semantics of the updated syntax element sps_lfnst_enabled_flag according to some embodiments of the present disclosure.
[0028] Figure 16 Shows an exemplary flowchart of an encoding method for LFNST transformation and ACT transformation according to some embodiments of the present disclosure.
[0029] Figure 17 Shows an exemplary syntax including the syntax element cu_act_enabled_flag according to some embodiments of the present disclosure.
[0030] Figure 18 Shows the exemplary semantics of the updated syntax element cu_act_enabled_flag according to some embodiments of the present disclosure.
[0031] Figure 19 Shows the exemplary semantics of the updated syntax element lfnst_idx according to some embodiments of the present disclosure.
[0032] Figure 20 Shows an exemplary flowchart of an encoding method for LFNST and ACT according to some embodiments of the present disclosure.
[0033] Figure 21Illustrates the exemplary semantics of the updated variable ApplyLfnstFlag according to some embodiments of the present disclosure. Detailed Description
[0034] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, and like numbers in different drawings represent the same or similar elements unless otherwise noted. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with aspects of the present disclosure as set forth in the appended claims. Specific aspects of the present disclosure are described in more detail below. If there is a conflict with the terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0035] The Joint Video Exploration Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0036] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been using the Joint Exploration Model (JEM) reference software to explore technologies beyond HEVC. As coding techniques are incorporated into JEM, JEM has achieved higher coding performance than HEVC.
[0037] The VVC standard has recently been developed and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0038] Video is a set of static images (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store these images in chronological order, and a video playback device (e.g., a television, computer, smartphone, tablet, video player, or any end-user terminal with a display function) can be used to display such images in chronological order. In addition, in some applications, a video capture device can transmit the captured video to a video playback device (e.g., a computer with a monitor) in real time, such as for monitoring, conferencing, or live streaming, etc.
[0039] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is typically called an "encoder", and the module for decompression is typically called a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be achieved through various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard. In this case, the codec can be referred to as a "transcoder".
[0040] The video encoding process can identify and retain useful information that can be used to reconstruct the image and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off made to reduce the required storage space and transmission bandwidth.
[0041] The useful information of the image being encoded (referred to as the "current image") includes the changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in the position of pixels, changes in brightness, or changes in color, with the most concern being changes in position. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference image and the current image.
[0042] An image encoded without referring to another image (i.e., it is its own reference image) is called an "I-image". If some or all of the blocks in the image (e.g., blocks that typically refer to parts of a video image) are predicted using intra-frame prediction or inter-frame prediction (e.g., single prediction) using one reference image, the image is called a "P-image". If at least one block in the image is predicted using two reference images (e.g., bidirectional prediction), the image is called a "B-image".
[0043] Figure 1Shows the structure of an example video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live video or a captured and archived video. The video 100 can be a real video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., a video file stored in a storage device), or a video feed interface for receiving videos from a video content provider (e.g., a video broadcast transceiver).
[0044] As Figure 1 shown, the video sequence 100 can include a series of images arranged in time along a time axis, including images 102, 104, 106, and 108. Images 102 - 106 are consecutive, and there are more images between images 106 and 108. In Figure 1 this, image 102 is an I-image, and its reference image is image 102 itself. Image 104 is a P-image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B-image, and its reference images are images 104 and 108, as indicated by the arrows. In some embodiments, the reference image of an image (e.g., image 104) may not be immediately before or after the image. For example, the reference image of image 104 can be an image before image 102. It should be noted that the reference images of images 102 - 106 are only examples, and the present disclosure does not limit the embodiments of reference images to Figure 1 the examples shown in
[0045] Generally, due to the computational complexity of the encoding and decoding tasks, video codecs do not encode or decode an entire image at once. Instead, they can divide the image into basic segments and encode or decode the image segment by segment. Such basic segments are referred to as basic processing units ("BPUs") in the present disclosure. For example, Figure 1The structure 110 therein shows an example structure of an image (e.g., any one of images 102 - 108) of the video sequence 100. In the structure 110, the image is divided into 4×4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in the image can have variable sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for the image based on a balance of coding efficiency and the level of detail to be maintained in the basic processing unit.
[0046] The basic processing unit can be a logical unit, which can include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, the basic processing unit of a color image can include a luminance component (Y) representing achromatic brightness information, one or more chrominance components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chrominance components can have the same basic processing unit size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components can be referred to as "coding tree blocks" ("CTB"). Any operation performed on the basic processing unit can be repeated for each of its luminance and chrominance components.
[0047] Video coding has multiple operation stages, examples of which are in Figures 2A - 2B and Figures 3A - 3BShown in. For each stage, the size of the basic processing unit may still be too large to process, so it can be further divided into segments called "basic processing subunits" in the present disclosure. In some embodiments, the basic processing subunit may be called a "block" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The size of the basic processing subunit can be the same as or smaller than that of the basic processing unit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be carried out to a further level according to processing needs. It should also be noted that different stages can use different schemes to divide the basic processing unit.
[0048] For example, in the mode decision stage (an example of which is shown in Figure 2B ), the encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing subunit.
[0049] For another example, in the prediction stage (an example of which is shown in Figures 2A - 2B ), the encoder can perform prediction operations at the basic processing subunit (e.g., CU) level. However, in some cases, the basic processing subunit may still be too large to process. The encoder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and prediction operations can be performed at the level of these segments.
[0050] For another example, in the transform stage (an example of which is shown in Figures 2A - 2BAs shown (in [description]), the encoder may perform a transformation operation on a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The encoder may further divide the basic processing subunit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and the transformation operation can be performed at the level of these segments. It should be noted that the partitioning scheme for the same basic processing subunit may be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU may have different sizes and numbers.
[0051] In Figure 1 the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing subunits, and their boundaries are shown as dashed lines. Different basic processing units of the same image can be divided into basic processing subunits in different schemes.
[0052] In some implementations, to provide parallel processing capabilities and fault tolerance for video encoding and decoding, an image can be divided into multiple regions for processing, such that for one region of the image, the encoding or decoding process can be independent of information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving the encoding efficiency. Additionally, when the data of one region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the damaged or lost data, thereby providing fault tolerance. In some video coding standards, an image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different images of the video sequence 100 can have different partitioning schemes for dividing the image into multiple regions.
[0053] For example, in Figure 1 the structure 110 is divided into three regions 114, 116, and 118, and their boundaries are shown as solid lines within the structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that Figure 1 the basic processing units, basic processing subunits, and regions of the structure 110 in [description] are only examples, and the present disclosure does not limit its embodiments.
[0054] Figure 2A shows a schematic diagram of an example encoding process 200A consistent with an embodiment of the present disclosure. For example, the encoding process 200A can be performed by an encoder. As Figure 2AAs shown, an encoder may encode video sequence 202 into video bitstream 228 according to process 200A. Similar to Figure 1 video sequence 100 in Figure 1 , video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to
[0055] structure 110 in Figure 2A , each original image of video sequence 202 may be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder may perform process 200A on each original image of video sequence 202 at the basic processing unit level. For example, the encoder may perform process 200A in an iterative manner, where the encoder may encode one basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel on regions (e.g., regions 114-118) of each original image of video sequence 202.
[0056] In Figure 2A , the encoder may feed the basic processing units of the original images of video sequence 202 (referred to as "original BPUs") to prediction stage 204 to generate prediction data 206 and predicted BPU 208. The encoder may subtract predicted BPU 208 from the original BPU to generate residual BPU 210. The encoder may feed residual BPU 210 to transform stage 212 and quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed prediction data 206 and quantized transform coefficients 216 to binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after quantization stage 214, the encoder may feed quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add reconstructed residual BPU 222 to predicted BPU 208 to generate prediction reference 224, which is used in prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path". The reconstruction path may be used to ensure that the encoder and decoder use the same reference data for prediction.
[0056] The encoder may iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate prediction reference 224 (in the reconstruction path) for encoding the next original BPU of the original image. After encoding all original BPUs of the original image, the encoder may continue to encode the next image in video sequence 202.
[0057] Referring to process 200A, an encoder may receive video sequence 202 generated by a video capture device (e.g., a camera). The term "receive" as used herein may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or otherwise inputting data.
[0058] In prediction stage 204, in the current iteration, the encoder may receive an original BPU and prediction reference 224, and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, which can be used to reconstruct the original BPU into the predicted BPU 208 from the prediction data 206 and the prediction reference 224.
[0059] Ideally, the predicted BPU 208 may be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To record this difference, after generating the predicted BPU 208, the encoder may subtract it from the original BPU to generate a residual BPU 210. For example, the encoder may subtract the value of a pixel (e.g., a grayscale value or an RGB value) of the predicted BPU 208 from the value of the corresponding pixel of the original BPU. Each pixel of the residual BPU 210 may have a residual value that is the result of such subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the prediction data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Thus, the original BPU is compressed.
[0060] To further compress the residual BPU 210, in transform stage 212, the encoder may reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", each basis pattern being associated with a "transformation coefficient". The basis patterns may have the same size (e.g., the size of the residual BPU 210). Each basis pattern may represent a frequency component (e.g., the frequency of brightness change) of the variation of the residual BPU 210. No basis pattern can be reproduced from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition may decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transformation coefficients are similar to the coefficients associated with the basis functions.
[0061] Different transformation algorithms can use different basic patterns. Various transformation algorithms can be used in the transformation stage 212, such as the discrete cosine transform, the discrete sine transform, etc. The transformation in the transformation stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transformation (referred to as "inverse transformation"). For example, in order to recover the pixels of the residual BPU 210, the inverse transformation can be to multiply the values of the corresponding pixels of the basic pattern by the corresponding correlation coefficients and sum the products to generate a weighted sum. For video coding standards, both the encoder and the decoder can use the same transformation algorithm (and thus the same basic pattern). Therefore, the encoder may only record the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from these coefficients without receiving the basic pattern from the encoder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significantly reducing the quality. Therefore, the residual BPU 210 is further compressed.
[0062] The encoder can further compress the transformation coefficients in the quantization stage 214. During the transformation process, different basic patterns can represent different change frequencies (e.g., luminance change frequencies). Since the human eye is usually better at recognizing low-frequency changes, the encoder can ignore the information of high-frequency changes without significantly degrading the decoding quality. For example, in the quantization stage 214, the encoder can generate the quantized transformation coefficients 216 by dividing each transformation coefficient by an integer value (referred to as "quantization scale factor") and rounding the quotient to its nearest integer. After such an operation, some transformation coefficients of the high-frequency basic pattern can be converted to zero, and the transformation coefficients of the low-frequency basic pattern can be converted to smaller integers. The encoder can ignore the quantized transformation coefficients 216 with zero values, further compressing the transformation coefficients through this operation. The quantization process is also reversible, where the quantized transformation coefficients 216 can be reconstructed as transformation coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0063] Because the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 may be lossy. Generally, the quantization stage 214 may cause the most information loss in the process 200A. The greater the information loss, the fewer bits the quantized transformation coefficients 216 may require. To obtain different degrees of information loss, the encoder can use different values of the quantization parameter or any other parameter of the quantization process.
[0064] In the binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using binary coding techniques (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform in the transform stage 212, the parameters of the quantization process (e.g., quantization parameter), the encoder control parameters (e.g., bitrate control parameter), etc. The encoder may use the output data of the binary coding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0065] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate the reconstructed transform coefficients. In the inverse transform stage 220, the encoder may generate the reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the prediction reference 224 that will be used in the next iteration of process 200A.
[0066] It should be noted that other variants of process 200A may be used to encode the video sequence 202. In some embodiments, the various stages of process 200A may be executed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit Figure 2A one or more of the
[0067] Figure 2B FIG. shows a schematic diagram of another exemplary coding process 200B consistent with embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared with process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0068] Generally, prediction techniques can be classified into two categories: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-frame prediction") can use pixels from one or more encoded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the spatial redundancy inherent in the picture. Temporal prediction (e.g., inter-picture prediction or "inter-frame prediction") can use regions from one or more encoded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include the encoded pictures. Temporal prediction can reduce the temporal redundancy inherent in the picture.
[0069] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra-frame prediction. For the original BPU of the picture being encoded, the prediction reference 224 can include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder can generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques can include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder can perform extrapolation at the pixel level, such as by extrapolating the values of the corresponding pixels for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation can be in various directions relative to the original BPU, such as in the vertical direction (e.g., on top of the original BPU), horizontal direction (e.g., to the left of the original BPU), diagonal direction (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction defined in the video coding standard being used. For intra-frame prediction, the prediction data 206 can include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, the extrapolation parameters, the direction of the neighboring BPUs used relative to the original BPU, etc.
[0070] For another example, during the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed on a per-BPU basis. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the reconstructed BPU. When all the reconstructed BPUs of the same image are generated, the encoder may generate the reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching region within the range of the reference images (referred to as the "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at the position in the reference image that has the same coordinates as the original BPU in the current image and may extend outward a predetermined distance. When the encoder (e.g., by using a pixel recursive algorithm, a block matching algorithm, etc.) identifies a region in the search window that is similar to the original BPU, the encoder may determine such a region as the matching region. The matching region may have a different size (e.g., smaller, equal to, larger, or different in shape) from the original BPU. Since the reference image and the current image are temporally separated on the time axis (e.g., as Figure 1 shown), the matching region may be considered to "move" over time to the position of the original BPU. The encoder may record the direction and distance of this motion as a "motion vector". When using multiple reference images (e.g., as in Figure 1 image 106), the encoder may search for the matching region for each reference image and determine its associated motion vector. In some embodiments, the encoder may assign weights to the pixel values of the matching regions of the respective matching reference images.
[0071] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.
[0072] To generate the predicted BPU 208, the encoder may perform an operation of "motion compensation". Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When using multiple reference images (e.g., as in Figure 1In the image 106), the encoder can move the matching region of the reference image based on the respective motion vectors and average pixel values of the matching regions. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of the respective matching reference images, the encoder can sum the weighted sums of the pixel values of the moved matching regions.
[0073] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 The image 104 in is a unidirectional inter-frame prediction image, where the reference image (e.g., image 102) is before image 104. Bidirectional inter-frame prediction can use one or more reference images in two temporal directions relative to the current image. For example, Figure 1 The image 106 in is a bidirectional inter-frame prediction image, where the reference images (e.g., images 104 and 108) are in two temporal directions relative to image 106.
[0074] Still referring to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform rate-distortion optimization techniques, where the encoder can select a prediction mode based on the bit rate of the candidate prediction modes and the distortion of the reconstructed reference images under the candidate prediction modes to minimize the value of the cost function. According to the selected prediction mode, the encoder can generate the corresponding predicted BPU 208 and predicted data 206.
[0075] In the reconstruction path of process 200B, if an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the currently encoded and reconstructed current BPU in the current image), the encoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current image). The encoder can feed the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion (e.g., block artifacts) introduced during the encoding of the prediction reference 224. The encoder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The reference image for loop filtering can be stored in the buffer 234 (or "decoded picture buffer") for later use (e.g., as an inter prediction reference image for future images of the video sequence 202). The encoder can store one or more reference images in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder can encode the parameters of the loop filter (e.g., loop filter strength) as well as the quantized transform coefficients 216, prediction data 206, and other information at the binary coding stage 226.
[0076] Figure 3A FIG. shows a schematic diagram of an example decoding process 300A consistent with embodiments of the present disclosure. Process 300A can be a decompression process corresponding to Figure 2A the compression process 200A in. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss during the compression and decompression processes (e.g., Figure 2A and Figure 2B the quantization stage 214 in), generally, the video stream 304 is different from the video sequence 202. Similar to Figure 2A and Figure 2B processes 200A and 200B in, the decoder can perform process 300A on each image encoded in the video bitstream 228 at the basic processing unit (BPU) level. For example, the decoder can perform process 300A in an iterative manner, where the decoder can decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for each region (e.g., regions 114 - 118) of each image encoded in the video bitstream 228.
[0077] In Figure 3AIn [description], the decoder may feed a portion of the video bitstream 228 associated with a basic processing unit of the encoded image (referred to as an "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder may decode this portion into predicted data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the predicted data 206 to the prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded image buffer in a computer memory). The decoder may feed the prediction reference 224 to the prediction stage 204 to perform a prediction operation in the next iteration of process 300A.
[0078] The decoder may iteratively execute process 300A to decode each encoded BPU of the encoded image and generate a prediction reference 224 for decoding the next encoded BPU of the encoded image. After decoding all the encoded BPUs of the encoded image, the decoder may output the image to the video stream 304 for display and continue to decode the next encoded image in the video bitstream 228.
[0079] In the binary decoding stage 302, the decoder may perform the inverse operations of the binary encoding techniques used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context - adaptive binary arithmetic encoding, or any other lossless compression algorithm). In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as the prediction mode, the parameters of the prediction operation, the transform type, the parameters of the quantization process (e.g., quantization parameter), the encoder control parameters (e.g., bit - rate control parameter), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over a network, the decoder may unpack the video bitstream 228 before feeding it to the binary decoding stage 302.
[0080] Figure 3B A schematic diagram of another exemplary decoding process 300B consistent with embodiments of the present disclosure is shown. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filter stage 232 and a buffer 234.
[0081] In process 300B, for an encoded basic processing unit (referred to as "current BPU") of an encoded image being decoded (referred to as "current image"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data, depending on what prediction mode the encoder used to encode the current BPU. For example, if the encoder uses intra prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, and the like. The parameters of the intra prediction operation can include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as references, the sizes of the adjacent BPUs, extrapolation parameters, the directions of the adjacent BPUs relative to the original BPU, and the like. For another example, if the encoder uses inter prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, and the like. The parameters of the inter prediction operation can include, for example, the number of reference images associated with the current BPU, the weights respectively associated with the reference images, the positions (e.g., coordinates) of one or more matching regions in each reference image, one or more motion vectors respectively associated with the matching regions, and the like.
[0082] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Figure 2B Details of performing such spatial prediction or temporal prediction are described in [reference], and will not be repeated below. After performing such spatial prediction or temporal prediction, the decoder can generate a predicted BPU 208. As Figure 3A described, the decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.
[0083] In process 300B, the decoder may feed prediction reference 224 to spatial prediction stage 2042 or temporal prediction stage 2044 to perform prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in spatial prediction stage 2042, after generating prediction reference 224 (e.g., the decoded current BPU), the decoder may directly feed prediction reference 224 to spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current image). If the current BPU is decoded using inter prediction in temporal prediction stage 2044, after generating prediction reference 224 (e.g., the reference image in which all BPUs have been decoded), the decoder may feed prediction reference 224 to loop filter stage 232 to reduce or eliminate distortion (e.g., block artifacts). The decoder may apply the loop filter to prediction reference 224 in the manner as Figure 2B described. The loop-filtered reference image may be stored in buffer 234 (e.g., a decoded image buffer in a computer memory) for later use (e.g., as an inter prediction reference image for future encoded images of video bitstream 228). The decoder may store one or more reference images in buffer 234 for use in temporal prediction stage 2044. In some embodiments, the prediction data may also include parameters of the loop filter (e.g., loop filter strength). In some embodiments, when the prediction mode indicator of prediction data 206 indicates that inter prediction is used to encode the current BPU, the prediction data includes parameters of the loop filter.
[0084] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding video consistent with embodiments of the present disclosure. As Figure 4 shown, apparatus 400 may include a processor 402. When processor 402 executes the instructions described herein, apparatus 400 may become a special-purpose machine for video encoding or decoding. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), system on a chip (SoC), application specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may also be a group of processors grouped as a single logic component. For example, as Figure 4As shown, the processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0085] The apparatus 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as Figure 4 shown, the stored data may include program instructions (e.g., program instructions for implementing the stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and the data for processing (e.g., via bus 410) and execute the program instructions to perform operations or control on the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical discs, magnetic disks, hard disk drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a set of memories grouped as a single logical component ( Figure 4 not shown in the figure).
[0086] The bus 410 may be a communication device for transmitting data between components within the apparatus 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), etc.
[0087] For ease of explanation without causing ambiguity, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits" in this disclosure. The data processing circuits may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. In addition, the data processing circuits may be a single independent module, or may be fully or partially combined into any other component of the apparatus 400.
[0088] The apparatus 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0089] In some embodiments, optionally, the apparatus 400 may also include a peripheral interface 408 to provide connections to one or more peripheral devices. AsFigure 4 As shown, the peripheral device may include but is not limited to a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), etc.
[0090] It should be noted that the video codec (e.g., the codec that executes processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).
[0091] Figure 5 An exemplary processing flow of an encoding process is shown. In some embodiments, the encoding process 500 may be applied to the VVC standard. As Figure 5 shown, in the encoding process 500, the transform stage 212 may include an Adaptive Color Transform (ACT) 212A, which is also applied to reduce the redundancy between the three color components in the 4:4:4 chroma format. The ACT 212A performs a loop color space conversion in the prediction residual domain by adaptively converting the residual from the input color space (e.g., typically in the RGB color space) to the YCgCo space. Two color spaces are adaptively selected by indicating an ACT flag at the coding unit (CU) level. When the ACT flag is equal to 1, the residual of the CU is encoded in the YCgCo space. When the ACT flag is not equal to 1 (e.g., the ACT flag is equal to 0), the residual of the CU is encoded in the original color space (e.g., typically in the RGB color space). The encoding process 500 may also include a primary transform 212B, which may be applied to the output of the ACT 212A, and a Low-Frequency Non-Separable Transform (LFNST) 212C, which may be applied to the output of the primary transform 212B. Quantization 214 may receive the output of the LFNST 212C.
[0092] Figure 6 An exemplary processing flow of a decoding process is shown. In some embodiments, the decoding process 600 may be applied to the VVC standard. As Figure 6As shown, in the decoding process 600, the inverse transform stage 220 may include an inverse LFNST 220C and an inverse primary transform 220B, and the inverse LFNST 220C is applied between the inverse quantization 218 and the inverse primary transform 220B. The decoding process may further include an inverse ACT 220A, which may be applied to the output of the inverse primary transform 220B. The reconstructed residual BPU 222 may receive the output of the inverse ACT 220A. In the LFNST 212C, a 4×4 non-separable transform or an 8×8 non-separable transform is applied according to the block size. Four transform sets are used in the LFNST 212C. For each transform set, the selected non-separable quadratic transform candidate is further specified by an explicitly indicated LFNST index. The index is indicated once in the bitstream of each intra-coded unit (CU) after the transform coefficients.
[0093] In VVC, the discrete cosine transform type II (DCT-II) is used as the primary transform. In addition to DCT-II, a multi-transform selection (MTS) scheme is also used, in which the primary transform is selected from multiple selected transforms of discrete cosine transform 8 (DCT8) / discrete sine transform 7 (DST7).
[0094] As Figure 5 shown, the transform 212 can be a very complex unit. For example, due to matrix multiplication, the LFNST 212C is a complex process. As a result, cascading three inverse transforms (e.g., the inverse LFNST 220C as Figure 6 shown is followed by the inverse primary transform 220B, and then the inverse ACT 220A) requires more clock cycles and increases the latency of the decoding pipeline. In addition, the LFNST matrix set is trained using the residual signal in the YUV color space. They may not be optimized for the ACT-coded blocks because the ACT residuals are typically in the YCoCg color space. Therefore, in terms of compression performance, the combination of the LFNST 212A and the ACT 212C may not be beneficial.
[0095] Embodiments of the present disclosure provide an updated encoding and decoding scheme to solve the problems listed above, such as improving compression performance and reducing latency. Figure 7 An exemplary flowchart of an encoding method 700 for LFNST and ACT according to some embodiments of the present disclosure is shown. As Figure 7 shown, the method 700 may be executed by an encoder (e.g., through Figure 2A the process 200A or Figure 2B the process 200B), or by one or more software or hardware components of a device (e.g., Figure 4 the device 400). For example, one or more processors (e.g., Figure 4The processor 402) may execute method 700. In some embodiments, method 700 may be implemented by a computer program product that is included in a computer-readable medium and includes computer-executable instructions such as program code that is executed by a computer (e.g., Figure 4 device 400). Method 700 may include steps 702 and 704 below.
[0096] In step 702, one or more video sequences are received for processing. In step 704, the one or more video sequences are encoded using only one of a low-frequency non-separable transform (LFNST) and an adaptive color transform (ACT). For example, in transform stage 212 (see Figure 5 ), LFNST 212C or ACT 212B may be used, but a combination of LFNST 212C and ACT 212B is not allowed. If LFNST 212C is used, ACT 212B is disabled. If ACT 212B is used, LFNST 212C is disabled. Thus, the efficiency of the encoding / decoding process is improved. In some embodiments, method 700 is used at the coded layer video sequence (CLVS) level. In some embodiments, method 700 is used at the CU level.
[0097] At the CLVS level, in some embodiments, ACT 212A is conditionally processed based on LFNST 212C. Figure 8 An exemplary flowchart of an encoding method 800 for LFNST and ACT according to some embodiments of the present disclosure is shown. It can be understood that method 800 may be Figure 7 part of step 704 in method 700. In step 802, ACT is enabled based on the state of LFNST. For example, if LFNST is enabled in the CLVS, there is no need to determine ACT, such that ACT cannot be used in the CLVS. ACT can be enabled in the CLVS only when LFNST is not applied in the CLVS. Thus, if LFNST is enabled, ACT cannot be enabled.
[0098] In VVC (e.g., VVC Draft 9), there is a sequence parameter set (SPS) syntax element sps_act_enabled_flag to enable or disable the ACT of CLVS. The syntax element sps_act_enabled_flag being equal to 1 may indicate that the ACT may be used for encoding / decoding images of CLVS. The syntax element sps_act_enabled_flag being equal to 0 may mean that the ACT is not used for encoding / decoding images of CLVS. In some embodiments, another SPS flag sps_lfnst_enabled_flag may be used to indicate whether LFNST is enabled or disabled in CLVS.
[0099] Figure 9 An exemplary SPS syntax 900 according to some embodiments of the present disclosure is shown. The SPS syntax structure 900 may be used in method 800. As Figure 9 shown, changes to the previous VVC are shown in italics, and the proposed deleted syntax is further shown with a strikethrough. As Figure 9 shown, in VVC (e.g., VVC Draft 9), based on the value of syntax element 902 (e.g., sps_lfnst_enabled_flag), syntax element 901 (e.g., sps_act_enabled_flag) is conditionally indicated. If LFNST is enabled in CLVS (e.g., syntax element 902 is equal to 1), then syntax element 901 is not indicated and is inferred to be equal to 0. Since the indication of syntax element 901 depends on the value of syntax element 902, syntax element 902 is sent before syntax element 901. As Figure 9 shown, if ChromaArrayType (chroma type) is equal to 3 (e.g., 4:4:4 chroma subsampled video), the syntax element sps_max_luma_transform_size_64_flag is equal to 0, and syntax element 902 is equal to 0 (reference box 903), then syntax element 901 is indicated. As a result, if syntax element 902 is not enabled (e.g., sps_lfnst_enabled_flag is equal to 0), then syntax element 901 is indicated. Therefore, the combination of LFNST and ACT is not allowed for encoding or decoding, thus reducing the complexity of the encoding / decoding process.
[0100] Furthermore, the semantics of the syntax element sps_act_enabled_flag may be updated. Figure 10 An exemplary semantics of the updated syntax element sps_act_enabled_flag according to some embodiments of the present disclosure is shown. AsFigure 10 As shown, changes to the previous VVC are shown in italics. Portion 1001 is added and is inferred to be equal to 0 when sps_act_enabled_flag does not exist. Thus, the semantic definition of the syntax element sps_act_enabled_flag defines the value of sps_act_enabled_flag when sps_act_enabled_flag does not exist in the encoded / decoded bitstream.
[0101] In some embodiments, bitstream consistency constraints may be imposed on sps_act_enabled_flag and the semantics of the syntax element sps_act_enabled_flag may be updated. Figure 11 An exemplary semantics of the updated syntax element sps_act_enabled_flag according to some embodiments of the present disclosure is shown. As Figure 11 shown, changes to the previous VVC are shown in italics, with reference box 1101. As Figure 11 shown, bitstream consistency constraint 1101 is added such that when sps_lfnst_enabled_flag (sps_lfnst enable flag) is enabled (e.g., the value of sps_lfnst_enabled_flag is equal to 1), sps_act_enabled_flag is not enabled (e.g., the value of sps_act_enabled_flag is equal to 0). Thus, the combination of LFNST and ACT is not allowed for encoding or decoding, thereby reducing the complexity of the encoding / decoding process.
[0102] In some embodiments, LFNST is conditionally processed based on ACT. Figure 12 An exemplary flowchart of an encoding method 1200 for LFNST and ACT according to some embodiments of the present disclosure is shown. It can be understood that method 1200 may be Figure 7 part of step 704 in method 700. At step 1202, LFNST is enabled based on the state of ACT. For example, if ACT is applied in CLVS, there is no need to determine LFNST such that LFNST cannot be used in CLVS. Only when ACT is not applied in CLVS can LFNST be enabled in CLVS. Thus, if ACT is enabled in CLVS, LFNST cannot be enabled.
[0103] Figure 13 An exemplary SPS syntax with an updated sps_lfnst_enabled_flag according to some embodiments of the present disclosure is shown. SPS syntax structure 1300 may be used in method 1200. In Figure 13Among them, the changes from the previous VVC are shown in italics as shown in box 1302. As Figure 13 shown, in VVC (e.g., VVC Draft 9), based on the value of sps_act_enabled_flag, the syntax element 1301 (e.g., sps_lfnst_enabled_flag) is conditionally indicated. If the value of sps_act_enabled_flag is equal to 0 (i.e., ACT is not enabled), the syntax element 1301 is indicated. If the value of sps_act_enabled_flag is equal to 1 (i.e., ACT is enabled), the syntax element 1301 is not indicated and is inferred to be equal to 0. Therefore, the combination of LFNST and ACT is not allowed for encoding or decoding, thus reducing the complexity of the encoding / decoding process.
[0104] In addition, the semantics of the syntax element sps_lfnst_enabled_flag can be updated. Figure 14 An exemplary semantics of the updated syntax element sps_lfnst_enabled_flag according to some embodiments of the present disclosure is shown. As Figure 14 shown, the changes to the previous VVC are shown in italics. Part 1401 is added and is inferred to be equal to 0 when sps_lfnst_enabled_flag does not exist. Therefore, the semantics of the syntax element sps_lfnst_enabled_flag is more sound.
[0105] Figure 15 Another exemplary semantics of the updated syntax element sps_lfnst_enabled_flag according to some embodiments of the present disclosure is shown. As Figure 15 shown, the changes to the previous VVC are shown in italics as shown in box 1501. As Figure 15 shown, a consistency constraint 1501 is added such that when sps_act_enabled_flag is enabled (e.g., the value of sps_act_enabled_flag is equal to 1), sps_lfnst_enabled_flag is not enabled (e.g., the value of sps_lfnst_enabled_flag is equal to 0). Therefore, the combination of LFNST and ACT is not allowed for encoding or decoding, thus reducing the complexity of the encoding / decoding process.
[0106] In some embodiments, bitstream consistency constraints are imposed such that the values of sps_act_enabled_flag and sps_lfnst_enabled_flag cannot both be equal to 1 at the same time. For example, the following constraint can be applied: "The requirement for bitstream consistency is that in the bitstream of a single CLVS, the values of sps_act_enabled_flag and sps_lfnst_enabled_flag should not both be equal to 1".
[0107] In the embodiments described above, the combination of LFNST and ACT is disabled at the CLVS level. In some embodiments, LFNST and ACT are at the CU level. Figure 16 An exemplary flowchart of an encoding method 1600 for LFNST transform and ACT transform according to some embodiments of the present disclosure is shown. It can be understood that method 1600 can be Figure 7 a part of step 704 in method 700. In step 1602, LFNST is enabled based on the state of ACT in the current coding unit. For example, when ACT is not enabled in the current CU, it is determined whether LFNST is used in the current CU. If ACT is enabled in the current CU, there is no need to determine whether to use LFNST, such that LFNST cannot be used in the current CU. LFNST can only be used in the current CU when ACT is not enabled in the current CU. Therefore, the combination of LFNST and ACT is not allowed in the current CU.
[0108] In VVC (e.g., VVC draft 9), there is a CU-level syntax element cu_act_enabled_flag (cu_act enable flag) to indicate whether ACT is applied in the current CU. If the value of cu_act_enabled_flag is equal to 1 (i.e., ACT is enabled in the current CU), then ACT is applied in the current CU. Another CU-level syntax element lfnst_idx (lfnst index) is used, and the syntax element lfnst_index specifies whether and which one of two LFNST kernels in a selected transform set is used. Generally, four transform sets are predefined in LFNST, and each transform set has two inseparable transforms (e.g., two LFNST kernels). The syntax element lfnst_index being equal to 0 specifies that LFNST is not used in the current CU.
[0109] To prohibit the combination of LFNST and ACT at the CU level, if the value of cu_act_enabled_flag is equal to 1 (i.e., ACT is not enabled in the current CU), then lfnst_index is not indicated and is inferred to be 0.
[0110] In some embodiments, the syntax element lfnst_index can be indicated based on the cu_act_enabled_flag. Figure 17 An exemplary syntax including the syntax element cu_act_enabled_flag according to some embodiments of the present disclosure is shown. The syntax structure 1700 can be used in the method 1600. As Figure 17 shown, the changes to the previous VVC are shown in italics, reference box 1701. As Figure 17 shown, if the cu_act_enabled_flag is equal to 1, the syntax element 1702 (e.g., lfnst_idx) is not indicated. The combination of LFNST and ACT is not allowed for encoding or decoding, thus reducing the complexity of the encoding / decoding process.
[0111] In some embodiments, a bitstream consistency constraint can be imposed such that when the value of the syntax element 1702 (e.g., lfnst_idx) is not equal to 0, the value of the cu_act_enabled_flag is equal to 0. The semantics of the syntax element cu_act_enabled_flag can be updated. Figure 18 An exemplary semantics of the updated syntax element cu_act_enabled_flag according to some embodiments of the present disclosure is shown. As Figure 18 shown, the changes to the previous VVC are shown in italics, reference box 1801. As Figure 18 shown, the bitstream consistency constraint 1801 is added such that when lfnst_idx is not equal to 0 (i.e., LFNST is used in the current CU), the cu_act_enabled_flag is not enabled (e.g., the value of the cu_act_enabled_flag is equal to 0). Therefore, the combination of LFNST and ACT is not allowed for encoding or decoding, thereby reducing the complexity of the encoding / decoding process.
[0112] In some embodiments, a bitstream consistency constraint can be imposed such that when the value of the cu_act_enabled_flag is equal to 1, the value of the lfnst_index is equal to 0. The semantics of the syntax element lfnst_idx can be updated. Figure 19 An exemplary semantics of the updated syntax element lfnst_idx according to some embodiments of the present disclosure is shown. As Figure 19 shown, the changes to the previous VVC are shown in italics, reference box 1901. As Figure 19As shown, the bitstream consistency constraint 1901 is added such that when cu_act_enabled_flag is enabled (i.e., the value of cu_act_enabled_flag is equal to 1), lfnst_idx is not enabled (e.g., the value of lfnst_idx is equal to 0). Thus, the combination of LFNST and ACT is not allowed for encoding or decoding, thereby reducing the complexity of the encoding / decoding process.
[0113] Figure 20 FIG. shows an exemplary flowchart of an encoding method 2000 for LFNST and ACT according to some embodiments of the present disclosure. It can be understood that method 2000 can be Figure 7 a part of step 704 in method 700. In step 2002, a variable for indicating the application of LFNST in the current coding unit is determined based on ACT not being enabled in the current coding unit. If ACT is enabled in the current CU, the variable is not derived, such that when ACT is enabled, LFNST cannot be used in the current CU. The variable is only derived when LFNST is used in the current CU and ACT is not enabled, to ensure that the combination of LFNST and ACT is not allowed. Figure 21 FIG. shows an exemplary semantics of the updated variable ApplyLfnstFlag according to some embodiments of the present disclosure. The semantics can be used in method 2000. As Figure 21 shown, the changes to the previous VVC are shown in italics. As Figure 21 shown, the variable 2101 (e.g., ApplyLfnstFlag) can be derived only when ACT is not enabled in the current CU (e.g., the value of the syntax element 2103 (e.g., cu_act_enabled_flag) is equal to 0) and LFNST is used (e.g., the value of the syntax element 2102 (e.g., lfnst_idx) is greater than 0). Thus, the combination of LFNST and ACT is not allowed for encoding or decoding, thereby reducing the complexity of the encoding / decoding process.
[0114] It can be understood that although the present disclosure relates to providing various syntax elements for inference based on values equal to 0 or 1, the values can be configured in any way (e.g., 1 or 0) to provide appropriate inferences.
[0115] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) to perform the above method. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a hole pattern, RAM, PROM, and EPROM, FLASH-EPROM, or any other flash memory, NVRAM, caches, registers, any other storage chip or cartridge, and network versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0116] It should be noted that relational terms in this document, such as "first" and "second", etc., are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. Additionally, the words "comprising", "having", "including", and other similar forms are intended to have the same meaning and are open-ended, because one or more items following any of these words do not mean an exhaustive list of these one or more items, or are limited to the listed one or more items.
[0117] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations, unless infeasible. For example, if it is stated that a database may include A or B, then the database may include A, or B, or A and B, unless otherwise clearly stated or infeasible. As a second example, if it is stated that a database may include A, B, or C, then the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C, unless otherwise clearly stated or infeasible.
[0118] It can be understood that the above embodiments can be implemented by hardware or software (program code) or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. When the software is executed by a processor, it can perform the disclosed method. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that multiple of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0119] The embodiments can be further described using the following clauses:
[0120] 1. A computer-implemented video encoding method, comprising:
[0121] Receive one or more video sequences for processing; and
[0122] Encode the one or more video sequences using only one of a low-frequency non-separable transform (LFNST) and an adaptive color transform (ACT).
[0123] 2. The method according to clause 1, wherein the LFNST and the ACT are at the level of a coded layer video sequence (CLVS).
[0124] 3. The method according to clause 2, wherein encoding the one or more video sequences using only one of the LFNST and the ACT further comprises:
[0125] Enable the ACT based on the state of the LFNST.
[0126] 4. The method according to clause 3, wherein enabling the ACT based on the state of the LFNST further comprises:
[0127] Determine whether to enable the ACT based on the state of the LFNST; and
[0128] Enable the ACT based on determining that the LFNST is not enabled.
[0129] 5. The method according to clause 4, wherein determining whether to enable the ACT further comprises:
[0130] Based on bitstream consistency constraints, in response to a second flag being indicated, determine that a first flag is not indicated, the first flag enabling the ACT and the second flag enabling the LFNST.
[0131] 6. The method according to clause 2, wherein encoding the one or more video sequences using only one of the LFNST and the ACT further comprises:
[0132] Enable the LFNST based on the state of the ACT.
[0133] 7. The method according to clause 6, wherein enabling the LFNST based on the state of the ACT further comprises:
[0134] Determine whether to enable the LFNST based on the state of the ACT; and
[0135] Enable the LFNST based on determining that the ACT is not enabled.
[0136] 8. The method according to clause 7, wherein determining whether to enable the LFNST further comprises:
[0137] Based on bitstream consistency constraints, in response to a second flag being indicated, determine that a first flag is not indicated, the first flag enabling the LFNST and the second flag enabling the ACT.
[0138] 9. A method according to claim 1, wherein LFNST and ACT are at the coding unit (CU) level.
[0139] 10. The method according to claim 9, wherein encoding the one or more video sequences using only one of LFNST and ACT further comprises:
[0140] Enabling LFNST based on the state that ACT is not used in the current coding unit.
[0141] 11. The method according to claim 9, wherein encoding the one or more video sequences using only one of LFNST and ACT further comprises:
[0142] Based on bitstream consistency constraints, in response to LFNST being used in the current coding unit, determining that a first flag is not indicated, the first flag enabling ACT in the current coding unit.
[0143] 12. The method according to claim 9, wherein encoding the one or more video sequences using only one of LFNST and ACT further comprises:
[0144] Based on bitstream consistency constraints, in response to the first flag being indicated in the current coding unit, determining that LFNST is not used in the current coding unit, the first flag enabling ACT in the current coding unit.
[0145] 13. The method according to claim 9, wherein encoding the one or more video sequences using only one of LFNST and ACT further comprises:
[0146] Based on ACT not being enabled in the current coding unit, determining a variable for indicating that LFNST is applied in the current coding unit.
[0147] 14. An apparatus for performing video data processing, the apparatus comprising: [[ID=3,0]]
[0148] A memory configured to store instructions; and
[0149] One or more processors configured to execute the instructions to cause the apparatus to perform:
[0150] Receiving one or more video sequences for processing; and
[0151] Encoding the one or more video sequences using only one of low-frequency non-separable transform (LFNST) and adaptive color transform (ACT).
[0152] 15. The apparatus according to claim 14, wherein LFNST and ACT are at the coding layer video sequence (CLVS) level.
[0153] 16. The apparatus according to claim 15, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0154] Enable ACT based on the state of LFNST.
[0155] 17. The apparatus according to claim 16, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0156] Determine whether to enable ACT based on the state of LFNST; and
[0157] Enable ACT based on determining that LFNST is not enabled.
[0158] 18. The apparatus according to claim 17, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0159] Based on the bitstream consistency constraint, in response to the second flag being indicated, determine that the first flag is not indicated, the first flag enables ACT and the second flag enables LFNST.
[0160] 19. The apparatus according to claim 15, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0161] Enable LFNST based on the state of ACT.
[0162] 20. The apparatus according to claim 19, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0163] Determine whether to enable LFNST based on the state of ACT; and
[0164] Enable LFNST based on determining that ACT is not enabled.
[0165] 21. The apparatus according to claim 20, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0166] Based on the bitstream consistency constraint, in response to the second flag being indicated, determine that the first flag is not indicated, the first flag enables LFNST and the second flag enables ACT.
[0167] 22. The apparatus according to claim 14, wherein LFNST and ACT are at the coding unit (CU) level.
[0168] 23. The apparatus according to claim 22, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0169] Enable LFNST based on the state that ACT is not used in the current coding unit.
[0170] 24. The apparatus according to Article 22, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0171] Based on the bitstream consistency constraint, in response to the LFNST being used in the current coding unit, determine that the first flag is not indicated, and the first flag enables the ACT in the current coding unit.
[0172] 25. The apparatus according to Article 22, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0173] Based on the bitstream consistency constraint, in response to the first flag being indicated in the current coding unit, determine that the LFNST is not used in the current coding unit, and the first flag enables the ACT in the current coding unit.
[0174] 26. The apparatus according to Article 22, wherein the processor is further configured to execute instructions to cause the apparatus to perform:
[0175] Based on the ACT not being enabled in the current coding unit, determine a variable for indicating that the LFNST is applied in the current coding unit.
[0176] 27. A non-transitory computer-readable medium storing an instruction set, the instruction set being executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising:
[0177] Receiving one or more video sequences for processing; and
[0178] Encoding the one or more video sequences using only one of a low-frequency non-separable transform (LFNST) and an adaptive color transform (ACT).
[0179] 28. The non-transitory computer-readable medium according to Article 27, wherein the LFNST and the ACT are at the coded layer video sequence (CLVS) level.
[0180] 29. The non-transitory computer-readable medium according to Article 28, wherein the method further comprises:
[0181] Enabling the ACT based on the state of the LFNST.
[0182] 30. The non-transitory computer-readable medium according to Article 29, wherein the method further comprises:
[0183] Determining whether to enable the ACT based on the state of the LFNST; and
[0184] Enabling the ACT based on determining that the LFNST is not enabled.
[0185] 31. The non - transitory computer - readable medium according to Article 30, wherein the method further comprises:
[0186] Based on the bit - stream consistency constraint, in response to the second flag being indicated, determining that the first flag is not indicated, the first flag enabling ACT and the second flag enabling LFNST.
[0187] 32. The non - transitory computer - readable medium according to Article 28, wherein the method further comprises:
[0188] Enabling LFNST based on the state of ACT.
[0189] 33. The non - transitory computer - readable medium according to Article 32, wherein the method further comprises:
[0190] Determining whether to enable LFNST based on the state of ACT; and
[0191] Enabling LFNST based on determining that ACT is not enabled.
[0192] 34. The non - transitory computer - readable medium according to Article 33, wherein the method further comprises:
[0193] Based on the bit - stream consistency constraint, in response to the second flag being indicated, determining that the first flag is not indicated, the first flag enabling LFNST and the second flag enabling ACT.
[0194] 35. The non - transitory computer - readable medium according to Article 27, wherein LFNST and ACT are at the coding unit (CU) level.
[0195] 36. The non - transitory computer - readable medium according to Article 35, wherein the method further comprises:
[0196] Enabling LFNST based on the state that ACT is not used in the current coding unit.
[0197] 37. The non - transitory computer - readable medium according to Article 35, wherein the method further comprises:
[0198] Based on the bit - stream consistency constraint, in response to LFNST being used in the current coding unit, determining that the first flag is not indicated, the first flag enabling ACT in the current coding unit.
[0199] 38. The non - transitory computer - readable medium according to Article 35, wherein the method further comprises:
[0200] Based on the bit - stream consistency constraint, in response to the first flag being indicated in the current coding unit, determining that LFNST is not used in the current coding unit, the first flag enabling ACT in the current coding unit.
[0201] A non-transitory computer-readable medium according to Article 35, wherein the method further comprises:
[0202] Based on that ACT is not enabled in the current coding unit, determining a variable for indicating that LFNST is applied in the current coding unit.
[0203] The present disclosure also provides an apparatus for processing video data, comprising:
[0204] A first receiving module, configured to receive one or more video sequences for processing; and
[0205] A first processing module, configured to encode the one or more video sequences using only one of a low-frequency non-separable transform (LFNST) and an adaptive color transform (ACT).
[0206] The present disclosure also provides a computer program product, comprising a computer program which, when executed by a processor, implements any of the methods described above.
[0207] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary with the implementation. Certain adjustments and modifications may be made to the described embodiments. Considering the specification and practice of the present disclosure disclosed herein, other embodiments may be apparent to those skilled in the art. The specification and examples are only to be regarded as exemplary, and the true scope and spirit of the present disclosure are indicated by the following claims. The step order shown in the figures is also for illustrative purposes only and is not intended to be limited to any particular step order. Therefore, those skilled in the art can understand that these steps can be executed in a different order while implementing the same method.
[0208] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are used, they are used only in a general and descriptive sense and not for purposes of limitation.
Claims
1. A method for processing video data, wherein, Comprising: Receiving one or more video sequences for processing; And Encoding the one or more video sequences using only one of a low-frequency non-separable transform (LFNST) and an adaptive color transform (ACT); the LFNST and the ACT are at the encoded layer video sequence (CLVS) level or the coding unit (CU) level; Wherein, at the CLVS level, when the LFNST is enabled, the ACT is not enabled; when the ACT is enabled, the LFNST is not enabled; At the CU level, when the ACT in the current coding unit is not enabled, the LFNST is enabled; when the ACT in the current coding unit is enabled, the LFNST is not enabled.
2. The method according to claim 1, wherein, When at the CLVS level, encoding the one or more video sequences using only one of the LFNST and the ACT further comprises: Enabling the ACT based on the state of the LFNST.
3. The method according to claim 2, wherein Enabling the ACT based on the state of the LFNST further comprises: Determining whether to enable the ACT based on the state of the LFNST; and Enabling the ACT based on determining that the LFNST is not enabled.
4. The method according to claim 1, wherein, When at the CLVS level, encoding the one or more video sequences using only one of the LFNST and the ACT further comprises: Enabling the LFNST based on the state of the ACT.
5. The method according to claim 4, wherein, Enabling the LFNST based on the state of the ACT further comprises: Determining whether to enable the LFNST based on the state of the ACT; and Enabling the LFNST based on determining that the ACT is not enabled.
6. The method according to claim 1, wherein, When at the CU level, encoding the one or more video sequences using only one of the LFNST and the ACT further comprises: Enabling the LFNST based on the state that the ACT is not used in the current coding unit.
7. An apparatus for processing video data, wherein, The apparatus comprises: A memory configured to store instructions; and One or more processors configured to execute the instructions to cause the apparatus to perform: Receiving one or more video sequences for processing; and Encoding the one or more video sequences using only one of a low-frequency non-separable transform (LFNST) and an adaptive color transform (ACT); the LFNST and the ACT are at the encoded layer video sequence (CLVS) level or the coding unit (CU) level; Wherein, at the CLVS level, when the LFNST is enabled, the ACT is not enabled; when the ACT is enabled, the LFNST is not enabled; At the CU level, when the ACT in the current coding unit is not enabled, the LFNST is enabled; when the ACT in the current coding unit is enabled, the LFNST is not enabled.
8. The apparatus according to claim 7, wherein, When at the CLVS level, the processor is further configured to execute instructions to cause the apparatus to perform: Enabling the ACT based on the state of the LFNST.
9. The apparatus according to claim 7, wherein When at the CLVS level, the processor is further configured to execute instructions to cause the apparatus to perform: Enable the LFNST based on the state of the ACT.
10. The apparatus according to claim 7, wherein At the CU level, the processor is further configured to execute instructions to cause the device to perform: Enable the LFNST based on the state that the ACT is not used in the current coding unit.
11. A non - transitory computer - readable medium storing an instruction set, wherein, The instruction set can be executed by one or more processors of the device to cause the device to initiate a method for performing video data processing, the method including: Receiving one or more video sequences for processing; and Encoding the one or more video sequences using only one of a low-frequency non-separable transform (LFNST) and an adaptive color transform (ACT); the LFNST and the ACT are at the coded layer video sequence (CLVS) level or the coding unit (CU) level; wherein, at the CLVS level, when the LFNST is enabled, the ACT is not enabled; when the ACT is enabled, the LFNST is not enabled; at the CU level, when the ACT in the current coding unit is not enabled, the LFNST is enabled; when the ACT in the current coding unit is enabled, the LFNST is not enabled.
12. The non-transitory computer-readable medium according to claim 11, wherein, At the CLVS level, the method further includes: Enable the ACT based on the state of the LFNST.
13. The non-transitory computer-readable medium according to claim 11, wherein, At the CLVS level, the method further includes: Enable the LFNST based on the state of the ACT.
14. The non-transitory computer-readable medium according to claim 11, wherein, At the CU level, the method further includes: Enable the LFNST based on the state that the ACT is not used in the current coding unit.
Citation Information
Patent Citations
QP derivation and offset for adaptive color transform in video coding
US20160100167A1
Method and apparatus for processing image signal
WO2020067694A1