Method and apparatus for processing video content
By disabling CCALF, ALF, DQ, and SDH processes under certain conditions, the method addresses the challenge of optimizing video coding efficiency and complexity in advanced standards like VVC/H.266, achieving improved compression and reduced computational demands.
Patent Information
- Application Number
- JP2025081931
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-22
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-07
AI Technical Summary
Existing video coding standards face challenges in achieving optimal compression efficiency and reducing computational complexity while maintaining video quality, particularly with the adoption of advanced techniques like VVC/H.266.
The method involves disabling cross-component adaptive loop filter (CCALF) and chroma adaptive loop filter (ALF) processes, as well as dependent quantization (DQ) and sign data hiding (SDH) for specific conditions in video content processing, to optimize video encoding and decoding.
This approach enhances compression efficiency and reduces computational requirements, aligning with the goals of VVC/H.266 by improving coding performance and maintaining video quality.
Smart Images

Figure 2025116018000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 63 / 028,615, filed May 22, 2020, the entire contents of which are incorporated herein by reference.
[0002] Technical Field FIELD OF THE DISCLOSURE
[0002] The present disclosure relates generally to video processing, and more particularly to methods and apparatus for processing video content using VVC high-level syntax cleanup. [Background technology]
[0003] background
[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (e.g., HEVC / H.265) standard, the Versatile Video Coding (e.g., VVC / H.266) standard, and the AVS standard, which specify particular video coding formats, are developed by standardization organizations. As more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]
[0004] overview
[0004] Embodiments of the present disclosure provide a method and apparatus for processing video content, which may include receiving a bitstream including the video content, determining whether a first signal related to the video content satisfies a given condition, and disabling both a cross-component adaptive loop filter (CCALF) process and a chroma adaptive loop filter (ALF) process in response to determining that the first signal satisfies the given condition.
[0005]
[0005] The device may include a memory that stores a set of instructions and one or more processors, the one or more processors configured to execute the set of instructions to cause the device to receive a bitstream including video content, determine whether a first signal associated with the video content satisfies a given condition, and in response to determining that the first signal satisfies the given condition, disable both a cross-component adaptive loop filter (CCALF) process and a chroma adaptive loop filter (ALF) process.
[0006]
[0006] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer to cause the computer to perform a method for processing video content, the method including receiving a bitstream including the video content, determining whether a first signal associated with the video content satisfies a given condition, and in response to determining that the first signal satisfies the given condition, disabling both a cross-component adaptive loop filter (CCALF) process and a chroma adaptive loop filter (ALF) process.
[0007]
[0007] Embodiments of the present disclosure also provide a method and apparatus for processing video content. The method may include receiving a bitstream including the video content, determining whether a first signal related to the video content satisfies a given condition, and disabling dependent quantization (DQ) and sign data hiding (SDH) for at least one slice in response to determining that the first signal satisfies the given condition.
[0008]
[0008] The device may include a memory that stores a set of instructions and one or more processors, the one or more processors configured to execute the set of instructions to cause the device to receive a bitstream including video content, determine whether a first signal associated with the video content satisfies a given condition, and in response to determining that the first signal satisfies the given condition, disable dependent quantization (DQ) and sign data hiding (SDH) for at least one slice.
[0009]
[0009] An embodiment of the present disclosure further provides a non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer to cause the computer to perform a method for processing video content, the method including receiving a bitstream including the video content, determining whether a first signal associated with the video content satisfies a given condition, and in response to determining that the first signal satisfies the given condition, disabling dependent quantization (DQ) and sign data hiding (SDH) for at least one slice.
[0010] BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]
[0011] [Figure 1]
[0011] FIG. 1 shows a schematic diagram illustrating an example structure of a video sequence, consistent with some embodiments of the present disclosure. [Figure 2A]
[0012] 1 shows a schematic diagram illustrating an example encoding process of a hybrid video coding system consistent with certain embodiments of the present disclosure. [Figure 2B]
[0013] 1 shows a schematic diagram illustrating another exemplary encoding process of a hybrid video coding system, consistent with certain embodiments of the present disclosure. [Figure 3A]
[0014] 1 shows a schematic diagram illustrating an example decoding process for a hybrid video coding system consistent with certain embodiments of the present disclosure. [Figure 3B]
[0015] 1 shows a schematic diagram illustrating another exemplary decoding process for a hybrid video coding system, consistent with certain embodiments of the present disclosure. [Figure 4]
[0016] 1 shows a block diagram of an exemplary device for encoding or decoding video consistent with some embodiments of the present disclosure. [Figure 5]
[0017] 1 shows a flowchart of an exemplary computer-implemented method for processing video content, consistent with certain embodiments of the present disclosure. [Figure 6]
[0018] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 7]
[0019] 1 illustrates an example picture header syntax structure consistent with certain embodiments of this disclosure. [Figure 8]
[0020] 1 illustrates an example slice header syntax structure consistent with certain embodiments of the present disclosure. [Figure 9]
[0021] 1 illustrates an example picture parameter set syntax structure consistent with certain embodiments of this disclosure. [Figure 10]
[0022] 1 illustrates an example picture header syntax structure consistent with certain embodiments of this disclosure. [Figure 11]
[0023] 1 illustrates an example slice header syntax structure consistent with certain embodiments of the present disclosure. [Figure 12]
[0024] 1 shows a flowchart of an exemplary computer-implemented method for processing video content, consistent with certain embodiments of the present disclosure. [Figure 13]
[0025] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 14]
[0026] 1 illustrates an example slice header syntax structure consistent with certain embodiments of the present disclosure. [Figure 15]
[0027] 1 illustrates an exemplary general constraint syntax structure consistent with certain embodiments of the present disclosure. [Figure 16A]
[0028] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 16B] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 16C] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 16D] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 16E] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 16F] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 16G]1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 16H] 1 illustrates an example sequence parameter set syntax structure consistent with certain embodiments of the present disclosure. [Figure 17A]
[0029] 1 illustrates an example picture header syntax structure consistent with certain embodiments of this disclosure. [Figure 17B] 1 illustrates an exemplary picture header syntax structure consistent with certain embodiments of the present disclosure. [Figure 17C] 1 illustrates an exemplary picture header syntax structure consistent with certain embodiments of the present disclosure. [Figure 17D] 1 illustrates an exemplary picture header syntax structure consistent with certain embodiments of the present disclosure. [Figure 17E] 1 illustrates an exemplary picture header syntax structure consistent with certain embodiments of the present disclosure. [Figure 17F] 1 illustrates an exemplary picture header syntax structure consistent with certain embodiments of the present disclosure. [Figure 18A]
[0030] 1 illustrates an example slice header syntax structure consistent with certain embodiments of the present disclosure. [Figure 18B] 1 illustrates an example slice header syntax structure consistent with certain embodiments of the present disclosure. [Figure 18C] 1 illustrates an example slice header syntax structure consistent with certain embodiments of the present disclosure. [Figure 18D] 1 illustrates an example slice header syntax structure consistent with certain embodiments of the present disclosure. [Figure 18E] 1 illustrates an example slice header syntax structure consistent with certain embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0012] Detailed Description
[0031] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description will refer to the accompanying drawings in which, unless otherwise indicated, like numerals in different drawings represent the same or similar elements. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with the present invention. Rather, they are merely examples of devices and methods consistent with aspects related to the present invention as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein will control.
[0013]
[0032] The ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.
[0014]
[0033] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing a technology that goes beyond HEVC using the joint exploration model (JEM) reference software. Because the coding technology has been incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.
[0015]
[0034] The VVC standard is a recent development and continues to include more coding techniques that result in better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, and H.263.
[0016]
[0035] A video is a set of still pictures (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display these pictures in chronological order. Furthermore, in some applications, a video capture device can transmit captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as surveillance, conferencing, or live broadcasting.
[0017]
[0036] To reduce the storage space and transmission bandwidth required by such applications, video can be compressed before storage and transmission and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or dedicated hardware. A module for compression is generally referred to as an "encoder," and a module for decompression is generally referred to as a "decoder." Encoders and decoders can be collectively referred to as a "codec." Encoders and decoders can be implemented as various suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression may be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec may decompress video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec may be called a "transcoder."
[0018]
[0037] A video coding process can identify and retain useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, then such a coding process can be called "lossy." Otherwise, such a coding process can be called "lossless." Most coding processes are lossy; this is a tradeoff to reduce the required storage space and transmission bandwidth.
[0019]
[0038] Useful information about the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include pixel position changes, luminance changes, or color changes, of which position changes are the most important. Position changes of pixels representing an object can reflect the object's movement between the reference picture and the current picture.
[0020]
[0039] A picture that is coded without reference to another picture (i.e., the picture is its own reference picture) is called an "I-picture." A picture is called a "P-picture" if some or all of the blocks in the picture (e.g., blocks that generally refer to portions of a video picture) are predicted using intra- or inter-prediction with one reference picture (e.g., unidirectional prediction). A picture is called a "B-picture" if at least one block in the picture is predicted using two reference pictures (e.g., bidirectional prediction).
[0021]
[0040] The present disclosure is directed to methods and apparatus for processing video content that conforms to the above-mentioned video coding standards.
[0022]
[0041] 1 illustrates the structure of an example video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.
[0023]
[0042] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture and its reference picture is picture 102 itself. Picture 104 is a P-picture and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) need not immediately precede or follow that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the examples shown in FIG.
[0024]
[0043] Typically, video codecs do not encode or decode an entire picture at once because such a task is computationally complex. Rather, video codecs may divide a picture into elementary segments and encode or decode the picture segment by segment. In this disclosure, such elementary segments are referred to as basic processing units ("BPUs"). For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) and as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing unit can have variable sizes within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of a basic processing unit can be selected for a picture based on a balance between coding efficiency and the level of detail desired to be preserved within the basic processing unit. A CTU is the largest block unit and can contain as many as 128x128 luma samples (plus corresponding chroma samples depending on the chroma format). A CTU can be further partitioned into coding units (CUs) using a quadtree, binary tree, ternary tree, or a combination thereof.
[0025]
[0044] A basic processing unit may be a logical unit that may include various types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements of the basic processing unit, where the luma and chroma components may have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit can be repeated for each of its luma and chroma components.
[0026]
[0045] Video coding has multiple operational stages, examples of which are shown in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to process and therefore may be further divided into segments referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-units may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-units may have the same or smaller size than the basic processing units. Similar to basic processing units, basic processing sub-units are also logical units that may contain various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated for each of its luma and chroma components. It should be noted that such division can be performed to further levels depending on the processing needs. It should also be noted that different stages can divide the basic processing unit using different schemes.
[0027]
[0046] For example, in a mode decision stage (an example of which is shown in FIG. 2B ), an encoder may decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder may divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing sub-unit.
[0028]
[0047] In another example, in the prediction stage (one example of which is shown in FIGS. 2A-2B), the encoder can perform prediction operations at the level of elementary processing sub-units (e.g., CUs). However, in some cases, elementary processing sub-units may still be too large to process. The encoder can further divide the elementary processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.
[0029]
[0048] In another example, in the transform stage (one example of which is shown in FIGS. 2A-2B), the encoder may perform a transform operation on a residual elementary processing sub-unit (e.g., a CU). However, in some cases, the elementary processing sub-unit may still be too large to process. The encoder may further divide the elementary processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. It should be noted that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0030]
[0049] In structure 110 of Figure 1, basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown by dotted lines. Different basic processing units of the same picture can be divided into basic processing sub-units in different ways.
[0031]
[0050] In some implementations, to provide video encoding and decoding with parallel processing and error resilience, a picture can be divided into regions for processing, thereby enabling the encoding or decoding process for a region of a picture to not depend on information from any other region of the picture. In other words, each region of a picture can be processed independently. This allows a codec to process different regions of a picture in parallel, thus increasing coding efficiency. Furthermore, if data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience. Some video coding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that various pictures in video sequence 100 may have different partitioning schemes for dividing the picture into regions.
[0032]
[0051] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.
[0033]
[0052] FIG. 2A shows a schematic diagram of an example encoding process 200A consistent with embodiments of the present disclosure. For example, encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 of FIG. 1, video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to structure 110 of FIG. 1, each original picture of video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of a basic processing unit for each original picture of video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for each original picture region of video sequence 202 (eg, regions 114-118).
[0034]
[0053] 2A , an encoder may feed a basic processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder may feed the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may feed quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to predicted BPU 208 to generate a prediction reference 224 used in prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0035]
[0054] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in video sequence 202.
[0036]
[0055] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any manner of inputting data.
[0037]
[0056] In the prediction stage 204, in the current iteration, the encoder receives the original BPU and a prediction reference 224 and can perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 can be generated from the reconstruction path of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting prediction data 206, which can be used to reconstruct the original BPU from the prediction data 206 and the prediction reference 224 as a predicted BPU 208.
[0038]
[0057] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, an encoder may generate predicted BPU 208 and then subtract it from the original BPU to generate residual BPU 210. For example, the encoder may subtract pixel values (e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values of the original BPU. As a result of such subtraction between corresponding pixels of the original BPU and predicted BPU 208, each pixel of residual BPU 210 may have a residual value. Compared to the original BPU, prediction data 206 and residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant loss of quality. Thus, the original BPU is compressed.
[0039]
[0058] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency (e.g., luminance variation frequency) component of the residual BPU 210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, the basis patterns are similar to basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.
[0040]
[0059] Different transform algorithms can use different basis patterns. For example, various transform algorithms can be used in transform stage 212, such as a discrete cosine transform, a discrete sine transform, etc. The transform in transform stage 212 is reversible. That is, the encoder can reconstruct residual BPU 210 by inversely operating the transform (referred to as an "inverse transform"). For example, to reconstruct a pixel of residual BPU 210, the inverse transform can multiply the value of the corresponding pixel in the basis pattern by the associated respective coefficient and add the products to obtain a weighted sum. In video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Although the transform coefficients may have fewer bits compared to residual BPU 210, they can be used to reconstruct residual BPU 210 without significant loss of quality. Therefore, the residual BPU 210 is further compressed.
[0041]
[0060] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore high-frequency fluctuation information without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization scale factor") and rounding the quotient to its nearest neighbor. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also reversible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called "dequantization").
[0042]
[0061] Quantization stage 214 may be lossy because the encoder ignores the remainder of such a division in a rounding operation. Typically, quantization stage 214 may contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values of the quantization parameter or any other parameter of the quantization process.
[0043]
[0062] In the binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary coding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transform in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder may generate a video bitstream 228 using the output data of the binary coding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0044]
[0063] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.
[0045]
[0064] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, an encoder can perform the stages of process 200A in a different order. In some embodiments, one or more stages of process 200A can be combined into a single stage. In some embodiments, a single stage of process 200A can be separated into multiple stages. For example, transform stage 212 and quantization stage 214 can be combined into a single stage. In some embodiments, process 200A can include additional stages. In some embodiments, process 200A can omit one or more stages in FIG. 2A .
[0046]
[0065] 2B shows a schematic diagram of another example encoding process 200B consistent with embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and separates prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0047]
[0066] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels of one or more already coded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions of one or more already coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0048]
[0067] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. With respect to an original BPU of a picture being coded, the prediction reference 224 may include one or more neighboring BPUs coded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction specified within the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.
[0049]
[0068] In another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. With respect to the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, a reference picture may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered at a location of the reference picture that has the same coordinates as the original BPU in the current picture and may extend over a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pel recursion algorithm, a block matching algorithm, etc.), the encoder can determine that region as a matching region. The matching region may have different dimensions (e.g., smaller, equal, larger, or a different shape) than the original BPU. Because the reference picture and the current picture are separated in time in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., picture 106 in FIG. 1), the encoder can find the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the matching region in each matching reference picture.
[0050]
[0069] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, scaling, etc. In inter prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, etc.
[0051]
[0070] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift matching regions of a reference picture according to a motion vector, within which the encoder may predict the original BPU of the current picture. If multiple reference pictures are used (e.g., picture 106 of FIG. 1), the encoder may shift matching regions of the reference pictures according to their respective motion vectors and average pixel values of the matching regions. In some embodiments, if the encoder assigns weights to pixel values of the matching regions of the respective matching reference pictures, the encoder may add a weighted sum of pixel values of the shifted matching regions.
[0052]
[0071] In some embodiments, inter-prediction can be unidirectional or bidirectional. Unidirectional inter-prediction can use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which a reference picture (e.g., picture 102) precedes picture 104. Bidirectional inter-prediction can use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture in which reference pictures (e.g., pictures 104 and 108) are in both temporal directions relative to picture 104.
[0053]
[0072] Continuing with reference to the forward path of process 200B, after spatial prediction step 2042 and temporal prediction step 2044, in mode decision step 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrates of the candidate prediction modes and the distortion of the reconstructed reference picture under the candidate prediction modes. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.
[0054]
[0073] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating prediction reference 224 (e.g., the current BPU being encoded and reconstructed in the current picture), the encoder may feed prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). The encoder may feed prediction reference 224 to loop filter stage 232, where the encoder may apply a loop filter to prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) introduced during encoding of prediction reference 224. For example, the encoder may apply various loop filter techniques in loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference picture may be stored in buffer 234 (or a “decoded picture buffer”) for later use (e.g., for use as an inter-prediction reference picture for future pictures of video sequence 202). The encoder may store one or more reference pictures in a buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in a binary coding stage 226.
[0055]
[0074] FIG. 3A shows a schematic diagram of an example decoding process 300A consistent with embodiments of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for a region (e.g., regions 114-118) of each picture encoded in video bitstream 228.
[0056]
[0075] In FIG. 3A , a decoder may feed a portion of a video bitstream 228 associated with a basic processing unit of a coded picture (referred to as a “coded BPU”) to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may feed the prediction reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.
[0057]
[0076] The decoder may iteratively perform process 300A to decode each coded BPU of a coded picture and generate a prediction reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of a coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.
[0058]
[0077] In binary decoding stage 302, the decoder may perform an inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if video bitstream 228 is transmitted in packets over a network, the decoder may depacketize video bitstream 228 before feeding it to binary decoding stage 302.
[0059]
[0078] 3B shows a schematic diagram of another example decoding process 300B consistent with embodiments of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.
[0060]
[0079] In process 300B, for a coded basic processing unit (referred to as a "current BPU") of a coded picture being decoded (referred to as a "current picture"), prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of extrapolation, directions of the neighboring BPUs relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter-prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, etc.
[0061]
[0080] Based on the prediction mode indicator, the decoder may decide whether to perform spatial prediction (e.g., intra prediction) in a spatial prediction step 2042 or temporal prediction (e.g., inter prediction) in a temporal prediction step 2044. Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. As described in FIG. 3A, the decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.
[0062]
[0081] In process 300B, the decoder may feed a prediction reference 224 to a spatial prediction stage 2042 or a temporal prediction stage 2044 for performing a prediction operation within a next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the decoder may feed the prediction reference 224 to a loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B . The loop filtered reference pictures may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as inter-prediction reference pictures for future coded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data may further include loop filter parameters (e.g., loop filter strength). In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data includes loop filter parameters.
[0063]
[0082] FIG. 4 is a block diagram of an example device 400 for encoding or decoding video consistent with embodiments of the present disclosure. As shown in FIG. 4, device 400 may include a processor 402. When processor 402 executes the instructions described herein, device 400 may become a dedicated machine for encoding or decoding video. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (“CPUs”), graphics processing units (“GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may be a set of processors grouped together as a single logical entity. For example, as shown in Figure 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0064]
[0083] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to operate on or process the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped together as a single logical entity.
[0065]
[0084] Bus 410, such as an internal bus (e.g., a CPU memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like, may be a communication device that transfers data between components within device 400.
[0066]
[0085] For ease of explanation and to avoid ambiguity, this disclosure will collectively refer to the processor 402 and other data processing circuitry as "data processing circuitry." The data processing circuitry may be implemented entirely in hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, independent module or may be fully or partially combined within any other component of the device 400.
[0067]
[0086] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communications network, etc.) In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0068]
[0087] In some embodiments, device 400 may optionally further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or input interface coupled to a video archive), etc.
[0069]
[0088] It should be noted that a video codec (e.g., a codec that performs process 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).
[0070]
[0089] In VVC, an adaptive loop filter (ALF) with block-based filter adaptation is applied. For the luma component, one of 25 filters is selected for each 4x4 block based on local gradient direction and activity. In addition to the ALF, VVC Draft 9 also uses a cross-component adaptive loop filter (CCALF). The CCALF filter is designed to work in parallel with the luma ALF.
[0071]
[0090] VVC supports three types of Adaptation Parameter Set (APS) Network Abstraction Layer (NAL) units: ALF_APS, LMCS_APS, and SCALING_APS. Filter coefficients for the ALF process and CCALF process are signaled within the ALF_APS. Up to 25 sets of luma ALF filter coefficients and clipping value indices and up to 8 sets of chroma ALF filter coefficients and clipping value indices can be signaled within one ALF APS. In addition, up to 4 sets of CCALF filter coefficients for the Cb component and up to 4 sets of CCALF filter coefficients for the Cr component are signaled.
[0072]
[0091] In VVC Draft 9, there are two residual coding methods, including (a) the regular residual coding (RRC) method and (b) the transform skip residual coding (TSRC) method. These residual coding methods are designated herein as "residual_coding" and "residual_ts_coding." In VVC Draft 9, both transform skip (TS) and block differential pulse code modulation (BDPCM) blocks are allowed to select either the RRC method or the TSRC method under certain conditions. If the value of the slice-level flag slice_ts_residual_coding_disabled_flag for a slice is equal to 0, the TS and BDPCM-coded blocks of that slice select TSRC. If the value of the slice-level flag slice_ts_residual_coding_disabled_flag for a slice is equal to 1, the TS and BDPCM-coded blocks of that slice select RRC.
[0073]
[0092] The current design of VVC has shortcomings directed to the no_aps_constraint_flag syntax element, the no_tsrc_constraint_flag syntax element, and the order of syntax. This disclosure provides suggested methods, such as updating the definitions of syntax elements and updating syntax tables, to address these shortcomings.
[0074]
[0093] The first drawback concerns the no_aps_constraint_flag syntax element. The current design of VVC has several constraint flags signaled within the SPS. Certain coding tools can use those constraint flags (e.g., the no_aps_constraint_flag syntax element) to define profiles that are deactivated. The no_aps_constraint_flag syntax element specifies whether there are APS NAL units in the bitstream. The following are the semantics of no_aps_constraint_flag:
[0075]
[0094] The no_aps_constraint_flag syntax element equal to 1 specifies that there cannot be any NAL units in OlsInScope with nuh_unit_type equal to PREFIX_APS_NUT or SUFFIX_APS_NUT, and that sps_lmcs_enabled_flag and sps_scaling_list_enabled_flag may both be equal to 0. The no_aps_constraint_flag syntax element equal to 0 imposes no such constraint.
[0076]
[0095] As mentioned above, VVC supports three types of APS NAL units: ALF_APS, LMCS_APS, and SCALING_APS. The above semantic definition of no_aps_constraint_flag implies that if the no_aps_constraint_flag syntax element is equal to 1, then the coding tools LMCS and scaling list are disabled by setting sps_lmcs_enabled_flag and sps_scaling_list_enabled_flag equal to 0. However, in the current VVC design, ALF and CCALF can still be enabled even if there is no APS signaled (i.e., no_aps_constraint_flag is equal to 1).
[0077]
[0096] Since VVC allows a fixed / default set of filters for luma ALF, the current VVC design can enable the ALF process for the luma component without transmitting an APS. However, there is no such fixed / default set of filters for the chroma ALF and CCALF processes. Therefore, it is asserted that the chroma ALF and CCALF processes cannot be performed without signaling the filter sets by APS.
[0078]
[0097] In the proposed method, both chroma ALF and CCALF processes are disabled when APS is not signaled (i.e., no_aps_constraint_flag is equal to 1). The method for disabling chroma ALF and CCALF processes involves updating the definition of the no_aps_constraint_flag syntax element and introducing a new flag. This method addresses the above drawbacks and provides better code efficiency.
[0079]
[0098] Figure 5 shows a flowchart of an exemplary computer-implemented method for processing video content consistent with some embodiments of the present disclosure. The method may be performed by a decoder (e.g., by process 300A of Figure 3A or process 300B of Figure 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of Figure 4). For example, a processor (e.g., processor 402 of Figure 4) may perform the method of Figure 5. In some embodiments, the method may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, for execution by a computer (e.g., apparatus 400 of Figure 4).
[0080]
[0099] The method of FIG. 5 may include the following steps.
[0081]
[0100] In step 501, a bitstream containing video content is received by a decoder (eg, by process 300A of FIG. 3A or process 300B of FIG. 3B).
[0082]
[0101] In step 502, it is determined whether a first signal associated with video content satisfies a given condition. In some embodiments, this determination includes determining whether the first signal indicates that Adaptation Parameter Set (APS) Network Abstraction Layer (NAL) units are present in the received bitstream. For example, it is determined whether the no_aps_constraint_flag syntax element is equal to 1, since a no_aps_constraint_flag syntax element equal to 1 indicates the absence of APS NAL units.
[0083]
[0102] In step 503, in response to determining that the first signal satisfies a given condition, both the cross-component adaptive loop filter (CCALF) process and the chroma adaptive loop filter (ALF) process are disabled. As described above, if the no_aps_constraint_flag syntax element is equal to 1, both the CCALF process and the chroma ALF process are disabled.
[0084]
[0103] In some embodiments, the CCALF process can be disabled at the sequence level. The definition of the no_aps_constraint_flag syntax element can be updated by including sps_ccalf_enabled_flag. A no_aps_constraint_flag syntax element equal to 1 specifies that sps_ccalf_enabled_flag can be equal to 0. A sps_ccalf_enabled_flag syntax element equal to 0 specifies that the CCALF process is disabled and not applied when decoding pictures in CLVS that are associated with video content in the received bitstream. In some embodiments, when the no_aps_constraint_flag syntax element is equal to 1, the value of sps_ccalf_enabled_flag can be equal to 0. When the no_aps_constraint_flag syntax element is equal to 0, such constraints cannot be imposed.
[0085]
[0104] In some embodiments, disabling chroma ALF processing at the slice level may be implemented by updating the definition of the no_aps_constraint_flag syntax element. The definition of the no_aps_constraint_flag syntax element may be updated by including sh_alf_cb_flag, sh_alf_cr_flag, and sh_num_alf_aps_ids_luma. For example, a no_aps_constraint_flag syntax element equal to 1 specifies that the values of sh_alf_cb_flag, sh_alf_cr_flag, and sh_num_alf_aps_ids_luma for all slices in OlsInScope may be equal to 0. To use a fixed / default set of filters for luma ALF processing, the value of sh_num_alf_aps_ids_luma is set to 0. In some embodiments, if the no_aps_constraint_flag syntax element is equal to 1, then the sh_alf_cb_flag syntax element, the sh_alf_cr_flag syntax element, and the sh_num_alf_aps_ids_luma syntax element are equal to 0. If the no_aps_constraint_flag syntax element is equal to 0, then no such constraints can be imposed.
[0086]
[0105] In some embodiments, disabling chroma ALF processing can occur at the picture level instead of the slice level. The definition of the no_aps_constraint_flag syntax element can be updated by including ph_alf_cb_flag, ph_alf_cr_flag, and ph_num_alf_aps_ids_luma. For example, a no_aps_constraint_flag syntax element equal to 1 may specify that the values of ph_alf_cb_flag, ph_alf_cr_flag, and ph_num_alf_aps_ids_luma for all slices in OlsInScope can be equal to 0. To use a fixed / default set of filters for luma ALF processing, the value of ph_num_alf_aps_ids_luma is set to 0. In some embodiments, if the no_aps_constraint_flag syntax element is equal to 1, then the ph_alf_cb_flag syntax element, the ph_alf_cr_flag syntax element, and the ph_num_alf_aps_ids_luma syntax element are equal to 0. If the no_aps_constraint_flag syntax element is equal to 0, then no such constraints can be imposed.
[0087]
[0106] As explained above, the chroma ALF process can be disabled at the slice header level by updating the definition of the no_aps_constraint_flag syntax element. In some embodiments, the chroma ALF process can be controlled by providing a signal in the bitstream.
[0088]
[0107] In some embodiments, the method may further include providing a second signal for controlling the chroma ALF process at a picture parameter set (PPS) level, a sequence parameter set (SPS) level, a picture header (PH) level, or a slice header (SH) level. The second signal may be an SPS-level flag or a PPS-level flag.
[0089]
[0108] The picture header carries information about a particular picture and contains information common to all slices belonging to the same picture. The SPS contains sequence-level information shared by all pictures in a complete coded layered video sequence (CLVS) and can provide an overall picture of what the bitstream contains and how the information in the bitstream can be used. The SPS is at a higher level than the PH and SH, and the SPS level flags can be used to control the chroma ALF process at the SPS level, the PH level, or the SH level.
[0090]
[0109] Similarly, PPS contains syntax elements for CLVS, PPS is at a higher level than PH and SH, and PPS level flags can be used to control the chroma ALF process at the PPS level, PH level, or SH level.
[0091]
[0110] In some embodiments, an SPS-level flag (e.g., sps_chroma_alf_enabled_flag) is used to control the ALF process for chroma components. For example, an sps_chroma_alf_enabled_flag syntax element equal to 0 may specify that an adaptive loop filter is disabled and not applied when decoding chroma components of a picture in CLVS. An sps_chroma_alf_enabled_flag syntax element equal to 1 may specify that an adaptive loop filter may be enabled and applied when decoding chroma components of a picture in CLVS. If absent, the value of sps_chroma_alf_enabled_flag is inferred to be equal to 0. It is understood that this example gives different results based on the element being 0 or 1, but it will be understood that the indication of 0 and 1 is a design choice and that the results can be reversed with respect to this and other syntax elements (e.g., an sps_chroma_alf_enabled_flag syntax element equal to 0 may specify that an adaptive loop filter is enabled).
[0092]
[0111] 6-8 show examples of the sequence parameter set (SPS), picture header (PH), and slice header (SH) syntax tables of the proposed method, respectively. The modifications proposed below can be made to the VVC standard or implemented within other video coding technologies.
[0093]
[0112] 6, the proposed flag sps_chroma_alf_enabled_flag (e.g., element 601) is conditionally signaled if sps_alf_enabled_flag equals 1 and ChromaArrayType != 0. In addition, the flag sps_ccalf_enabled_flag (e.g., element 602) is also conditionally signaled if sps_alf_enabled_flag equals 1 and ChromaArrayType != 0.
[0094]
[0113] Figure 7 shows the ph_alf_aps_id_luma[i] syntax element, which specifies the aps_adaptation_parameter_set_id of the i-th ALF APS to which the luma component of the slice associated with the PH refers. If the ph_alf_aps_id_luma[i] syntax element is present, the following applies: The value of alf_luma_filter_signal_flag may be equal to 1 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i]. - The TemporalId of an APS NAL unit with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i] may be less than or equal to the TemporalId of the picture associated with the PH. If ChromaArrayType is equal to 0, the value of aps_chroma_present_flag of an APS NAL unit with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i] may be equal to 0. - If sps_ccalf_enabled_flag is equal to 0, the values of alf_cc_cb_filter_signal_flag and alf_cc_cr_filter_signal_flag may be equal to 0 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i]. Furthermore, in some embodiments of the present disclosure, as shown in element 701 of FIG. 7, if sps_chroma_alf_enabled_flag is equal to 0, the value of alf_chroma_filter_signal_flag may be equal to 0 for APS NAL units having aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i].
[0095]
[0114] Figure 8 shows the sh_alf_aps_id_luma[i] syntax element, which specifies the aps_adaptation_parameter_set_id of the i-th ALF APS to which the luma component of a slice refers. If sh_alf_enabled_flag is equal to 1 and sh_alf_aps_id_luma[i] is not present, the value of sh_alf_aps_id_luma[i] is inferred to be equal to the value of ph_alf_aps_id_luma[i]. If sh_alf_aps_id_luma[i] is present, the following applies: - The TemporalId of an APS NAL unit with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i] may be less than or equal to the TemporalId of a coded slice NAL unit. The value of alf_luma_filter_signal_flag may be equal to 1 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i]. If ChromaArrayType is equal to 0, the value of aps_chroma_present_flag may be equal to 0 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i]. - If sps_ccalf_enabled_flag is equal to 0, the values of alf_cc_cb_filter_signal_flag and alf_cc_cr_filter_signal_flag may be equal to 0 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i]. Furthermore, in some embodiments of the present disclosure, as shown in element 801 of FIG. 8, if sps_chroma_alf_enabled_flag is equal to 0, the value of alf_chroma_filter_signal_flag may be equal to 0 for APS NAL units having aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i].
[0096]
[0115] In some embodiments, the definition of no_aps_constraint_flag can be updated by introducing a new SPS level flag, sps_chroma_alf_enabled_flag syntax element in Figures 6-8. For example, the no_aps_constraint_flag syntax element equal to 1 specifies that there cannot be any NAL units in OlsInScope with nuh_unit_type equal to PREFIX_APS_NUT or SUFFIX_APS_NUT, and that sps_chroma_alf_enabled_flag (introduced in elements 601, 701, and 801 in Figures 6-8), sps_ccalf_enabled_flag, sps_lmcs_enabled_flag, and sps_scaling_list_enabled_flag (introduced in elements 602, 702, and 802 in Figures 6-8) can all be equal to 0. In some embodiments, if the no_aps_constraint_flag syntax element is equal to 1, the value of sh_num_alf_aps_ids_luma for all slices in OlsInScope may be equal to 0. If the no_aps_constraint_flag syntax element is equal to 0, no such constraints can be imposed.
[0097]
[0116] In some embodiments, a second signal (e.g., a PPS level flag) may also be introduced to control the chroma ALF process at the PPS level, the PH level, or the SH level. For example, a PPS level flag (e.g., pps_chroma_alf_enabled_flag) may be used to control the ALF process of a chroma component. A pps_chroma_alf_enabled_flag syntax element equal to 0 may specify that the adaptive loop filter is disabled and not applied when decoding a chroma component of a picture with reference to a PPS. A pps_chroma_alf_enabled_flag syntax element equal to 1 may specify that the adaptive loop filter may be enabled and applied when decoding a chroma component of a picture with reference to a PPS. If absent, the value of the pps_chroma_alf_enabled_flag syntax element is inferred to be equal to 0. It is understood that this example gives different results based on the element being 0 or 1, but the 0 and 1 designations are design choices and that the results can be reversed for this and other syntax elements (e.g., a pps_chroma_alf_enabled_flag syntax element equal to 0 can specify that the adaptive loop filter is enabled).
[0098]
[0117] 9-11 show example picture parameter set (PPS), picture header (PH), and slice header (SH) syntax tables of the proposed method, respectively. The modifications proposed below can be made to the VVC standard or implemented within other video coding technologies.
[0099]
[0118] As shown in element 901 of FIG. 9, the proposed flag pps_chroma_alf_enabled_flag is conditionally signaled if pps_chroma_tool_offsets_present_flag is equal to 1.
[0100]
[0119] Figure 10 shows the ph_alf_aps_id_luma[i] syntax element, which specifies the aps_adaptation_parameter_set_id of the i-th ALF APS to which the luma component of the slice associated with the PH refers. If ph_alf_aps_id_luma[i] is present, the following applies: The value of alf_luma_filter_signal_flag may be equal to 1 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i]. - The TemporalId of an APS NAL unit with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i] may be less than or equal to the TemporalId of the picture associated with the PH. If ChromaArrayType is equal to 0, the value of aps_chroma_present_flag of an APS NAL unit with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i] may be equal to 0. - If sps_ccalf_enabled_flag is equal to 0, the values of alf_cc_cb_filter_signal_flag and alf_cc_cr_filter_signal_flag may be equal to 0 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i]. Furthermore, in some embodiments of the present disclosure, as shown in element 1001 of FIG. 10, if pps_chroma_alf_enabled_flag is equal to 0, the value of alf_chroma_filter_signal_flag may be equal to 0 for APS NAL units having aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to ph_alf_aps_id_luma[i].
[0101]
[0120] Figure 11 shows the ph_alf_aps_id_luma[i] syntax element, which specifies the aps_adaptation_parameter_set_id of the i-th ALF APS to which the luma component of the slice associated with the PH refers. If ph_alf_aps_id_luma[i] is present, the following applies: sh_alf_aps_id_luma[i] syntax element, which specifies the aps_adaptation_parameter_set_id of the i-th ALF APS to which the luma component of the slice refers. If sh_alf_enabled_flag is equal to 1 and sh_alf_aps_id_luma[i] is absent, the value of sh_alf_aps_id_luma[i] is inferred to be equal to the value of ph_alf_aps_id_luma[i]. If sh_alf_aps_id_luma[i] is present, the following applies: - The TemporalId of an APS NAL unit with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i] may be less than or equal to the TemporalId of a coded slice NAL unit. The value of alf_luma_filter_signal_flag may be equal to 1 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i]. If ChromaArrayType is equal to 0, the value of aps_chroma_present_flag may be equal to 0 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i]. - If sps_ccalf_enabled_flag is equal to 0, the values of alf_cc_cb_filter_signal_flag and alf_cc_cr_filter_signal_flag may be equal to 0 for APS NAL units with aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i]. Furthermore, in some embodiments of the present disclosure, as shown in element 1101 of FIG. 11, if pps_chroma_alf_enabled_flag is equal to 0, the value of alf_chroma_filter_signal_flag may be equal to 0 for APS NAL units having aps_params_type equal to ALF_APS and aps_adaptation_parameter_set_id equal to sh_alf_aps_id_luma[i].
[0102]
[0121] In some embodiments, the definition of no_aps_constraint_flag can be updated by introducing a new PPS level flag pps_chroma_alf_enabled_flag syntax element in Figures 9-11. For example, the no_aps_constraint_flag syntax element equal to 1 specifies that there cannot be any NAL units in OlsInScope with nuh_unit_type equal to PREFIX_APS_NUT or SUFFIX_APS_NUT, and that pps_chroma_alf_enabled_flag (introduced by elements 901, 1001, and 1101 in Figures 9-11), sps_ccalf_enabled_flag, sps_lmcs_enabled_flag, and sps_scaling_list_enabled_flag (introduced by elements 902, 1002, and 1102 in Figures 9-11) can all be equal to 0. In some embodiments, if the no_aps_constraint_flag syntax element is equal to 1, the value of sh_num_alf_aps_ids_luma for all slices in OlsInScope may be equal to 0. If the no_aps_constraint_flag syntax element is equal to 0, no such constraints can be imposed.
[0103]
[0122] In some embodiments, a semantic constraint is applied to no_aps_constraint_flag to disable both ALF and CCALF. For example, a no_aps_constraint_flag syntax element equal to 1 may specify that there cannot be NAL units in OlsInScope with nuh_unit_type equal to PREFIX_APS_NUT or SUFFIX_APS_NUT, and that sps_alf_enabled_flag, sps_ccalf_enabled_flag, sps_lmcs_enabled_flag, and sps_scaling_list_enabled_flag may all be equal to 0. In some embodiments, when the no_aps_constraint_flag syntax element is equal to 1, the values of sps_alf_enabled_flag, sps_ccalf_enabled_flag, sps_lmcs_enabled_flag, and sps_scaling_list_enabled_flag may all be 0. If the no_aps_constraint_flag syntax element is equal to 0, no such constraints can be imposed.
[0104]
[0123] A second drawback of the current design of VVC relates to the defeat of transform skip residual coding (TSRC). The defeat of TSRC is prevented under either of the following two scenarios:
[0105]
[0124] In the first scenario, enabling dependent quantization (DQ) or sign data hiding (SDH) of a slice may cause the disabling of the TSRC to be prevented.
[0106]
[0125] In the second scenario, a flag related to the constraints of the TSRC may be used to prevent the TSRC from being overridden.
[0107]
[0126] Both of the above scenarios should be addressed so that TSRC can be disabled if necessary.
[0108]
[0127] In the first scenario, enabling DQ or SDH for a slice may prevent disabling of TSRC. Specifically, according to the slice header syntax of the current VVC design, when a slice-level DQ flag (i.e., sh_dep_quant_used_flag) or a slice-level SDH flag (i.e., sh_sign_data_hiding_used_flag) is equal to 1, the TSRC disabled flag (i.e., sh_ts_residual_coding_disabled_flag) is set to 0.
[0109]
[0128] Setting the sh_dep_quant_used_flag syntax element to 0 may disable TSRC. According to the definition of the sh_dep_quant_used_flag syntax element, an sh_ts_residual_coding_disabled_flag syntax element equal to 0 specifies that the residual_ts_coding() syntax structure is used to parse residual samples of transform skip blocks of the current slice. In addition, an sh_dep_quant_used_flag syntax element equal to 1 specifies that dependent quantization is used for the current slice. An sh_sign_data_hiding_used_flag syntax element equal to 1 specifies that sign bit hiding is used for the current slice. For TSRC to be disabled, both the sh_dep_quant_used_flag syntax element and the sh_sign_data_hiding_used_flag syntax element are required to be equal to 0. Enabling DQ or SDH for a slice may cause TSRC to be prevented from being disabled.
[0110]
[0129] In the second scenario, the no_tsrc_constraint_flag syntax element can be used to not allow TSRC disabling by setting the TSRC disabled flag to 0. In particular, a no_tsrc_constraint_flag syntax element equal to 1 specifies that sh_ts_residual_coding_disabled_flag can be equal to 0. A no_tsrc_constraint_flag syntax element equal to 0 does not impose such a constraint. As noted, a sh_ts_residual_coding_disabled_flag syntax element equal to 0 specifies that the residual_ts_coding() syntax structure is used to parse residual samples of transform skip blocks of the current slice. When the sh_ts_residual_coding_disabled_flag syntax element is set to 0, TSRC is enabled.
[0111]
[0130] In a conventional VVC video coding design, TSRC can be disabled for a given slice (e.g., sh_ts_residual_coding_disabled_flag is 1) if all of the following conditions are met: - Dependent Quantization (DQ) is disabled (sh_dep_quant_used_flag == 0) AND - Sign Data Hiding (SDH) is disabled (sh_sign_data_hiding_used_flag == 0) Disabling both DQ and SDH may disable TSRC, while enabling either DQ or SDH may prevent the disabling of TSRC.
[0112]
[0131] In some embodiments of the present disclosure, constraints are applied to the no_tsrc_constraint_flag syntax element to disable certain features such as TSRC. For example, the following embodiment is proposed to disallow the following two combinations: - no_tsrc_constraint_flag equals 1 and DQ is enabled - no_tsrc_constraint_flag equals 1 and SDH is enabled
[0113]
[0132] These proposed methods are provided to address situations where TSRC override is prevented. Some embodiments of the present disclosure apply constraints to the no_tsrc_constraint_flag syntax element and DQ and SDH related syntax elements to override TSRC. Current syntax elements are updated to provide TSRC override features with better code efficiency.
[0114]
[0133] For example, Figure 12 shows a flowchart of an exemplary computer-implemented method for processing video content consistent with some embodiments of the present disclosure. The method may be performed by a decoder (e.g., by process 300A of Figure 3A or process 300B of Figure 3B) or by one or more software or hardware components of an apparatus (e.g., apparatus 400 of Figure 4). For example, a processor (e.g., processor 402 of Figure 4) may perform the method. In some embodiments, the method may be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions, such as program code, for execution by a computer (e.g., apparatus 400 of Figure 4). The method may include the following steps:
[0115]
[0134] In step 1201, a bitstream containing video content is received, for example by a decoder.
[0116]
[0135] In step 1202, it is determined whether a first signal associated with the video content satisfies a given condition. The first signal may be a no_tsrc_constraint_flag syntax element. In some embodiments, it is determined whether the no_tsrc_constraint_flag syntax element is equal to 1.
[0117]
[0136] In step 1203, in response to determining that the first signal satisfies a given condition, Dependent Quantization (DQ) and Sign Data Hiding (SDH) are disabled for at least one slice. For example, if the no_tsrc_constraint_flag syntax element is determined to be equal to 1, DQ and SDH are disabled for at least one slice.
[0118]
[0137] In some embodiments, the method may further include disabling transform skip residual coding (TSRC) for at least one slice in response to determining that the first signal satisfies a given condition. For example, if a no_tsrc_constraint_flag syntax element is determined to be equal to 1, the TSRC disable flag is set to 1.
[0119]
[0138] In some embodiments, DQ and SDH may be disabled at the SPS level or at the SH level, hi some embodiments, DQ and SDH may be disabled for all slices.
[0120]
[0139] At the SPS level, DQ and SDH disablement applies to all slices. In some embodiments, constraints related to SPS level DQ and SDH syntax elements can be used to update the definition of the no_tsrc_constraint_flag syntax element. In some embodiments, constraints related to the no_tsrc_constraint_flag syntax element can be used to update the definition of both SPS level DQ and SDH syntax elements. In some embodiments, conditional signaling between the no_tsrc_constraint_flag syntax element and SPS level DQ and SDH syntax elements can be used to update the SPS syntax.
[0121]
[0140] At the SH level, the definition of the no_tsrc_constraint_flag syntax element can be updated with constraints related to the SH level DQ and SDH syntax elements for all slices. In some embodiments, the definition of both the SH level DQ and SDH syntax elements for a slice can be updated with constraints related to the no_tsrc_constraint_flag syntax element, and DQ and SDH disabling applies to the slice. In some embodiments, conditional signaling between the no_tsrc_constraint_flag syntax element and the SH level DQ and SDH syntax elements for a slice can be used to update the SH syntax, and DQ and SDH disabling applies to the slice. At the SH level, DQ and SDH disabling can apply to all slices or to a given slice, as described above.
[0122]
[0141] The modifications proposed below can be made to the VVC standard or can be implemented within other video coding technologies.
[0123]
[0142] In some embodiments, a semantic constraint is applied to no_tsrc_constraint_flag to disable DQ and SDH at the SPS level. Thus, the following proposed changes to the semantics of no_tsrc_constraint_flag may be made to the VVC standard or implemented within other video coding technologies. For example, a no_tsrc_constraint_flag syntax element equal to 1 may specify that sh_ts_residual_coding_disabled_flag may be equal to 1, and that sps_dep_quant_enabled_flag and sps_sign_data_hiding_enabled_flag may both be equal to 0. In some embodiments, if the no_tsrc_constraint_flag syntax element is equal to 1, then the sh_ts_residual_coding_disabled_flag syntax element, sps_dep_quant_enabled_flag syntax element, and sps_sign_data_hiding_enabled_flag syntax elements are equal to 0. If the no_tsrc_constraint_flag syntax element is equal to 0, then no such constraints can be imposed. At the SPS level, DQ and SDH disabling applies to all slices.
[0124]
[0143] In some embodiments, a semantic constraint is applied to no_tsrc_constraint_flag to disable DQ and SDH in all of the slices of CLVS. Therefore, the following proposed changes to the semantics of no_tsrc_constraint_flag can be made to the VVC standard or implemented in other video coding technologies. For example, a no_tsrc_constraint_flag syntax element equal to 1 may specify that for all slices, sh_ts_residual_coding_disabled_flag may be equal to 1, and sh_dep_quant_used_flag and sh_sign_data_hiding_used_flag may both be equal to 0. In some embodiments, when the no_tsrc_constraint_flag syntax element is equal to 1, the sh_ts_residual_coding_disabled_flag syntax element, the sh_dep_quant_used_flag syntax element, and the sh_sign_data_hiding_used_flag syntax element are equal to 0. When the no_tsrc_constraint_flag syntax element is equal to 0, no such constraints can be imposed.
[0125]
[0144] In some embodiments, we apply semantic constraints to both sps_dep_quant_enabled_flag and sps_sign_data_hiding_enabled_flag to disable DQ and SDH at the SPS level. The following proposed changes to the semantics of sps_dep_quant_enabled_flag and sps_sign_data_hiding_enabled_flag may be made to the VVC standard or implemented within other video coding technologies. For example, an sps_dep_quant_enabled_flag syntax element equal to 0 may specify that dependent quantization is disabled and not used for pictures that reference an SPS. An sps_dep_quant_enabled_flag syntax element equal to 1 may specify that dependent quantization is enabled and may be used for pictures that reference an SPS. In some embodiments, if the value of no_tsrc_constraint_flag is equal to 1, the value of sps_dep_quant_enabled_flag may be equal to 0. For example, an sps_sign_data_hiding_enabled_flag syntax element equal to 0 specifies that sign bit hiding is disabled and not used for pictures that reference an SPS. An sps_sign_data_hiding_enabled_flag syntax element equal to 1 specifies that sign bit hiding may be enabled and used for pictures that reference an SPS. If sps_sign_data_hiding_enabled_flag is absent, sps_sign_data_hiding_enabled_flag is inferred to be equal to 0. In some embodiments, if the value of no_tsrc_constraint_flag is equal to 1, the value of sps_sign_data_hiding_enabled_flag may be equal to 0. According to the updated definitions of the sps_dep_quant_enabled_flag and sps_sign_data_hiding_enabled_flag syntax elements, DQ and SDH can be disabled at the SPS level for all slices.
[0126]
[0145] In some embodiments, semantic constraints are applied to slice-level DQ flags (e.g., sh_dep_quant_used_flag) and SDH flags (e.g., sh_sign_data_hiding_used_flag). For example, an sh_dep_quant_used_flag syntax element equal to 0 may specify that dependent quantization is not used for the current slice. An sh_dep_quant_used_flag syntax element equal to 1 may specify that dependent quantization is used for the current slice. If sh_dep_quant_used_flag is absent, sh_dep_quant_used_flag is inferred to be equal to 0. In some embodiments, if the value of no_tsrc_constraint_flag is equal to 1, the value of sh_dep_quant_used_flag may be equal to 0. For example, an sh_sign_data_hiding_used_flag syntax element equal to 0 may specify that sign bit hiding is not used for the current slice. The sh_sign_data_hiding_used_flag syntax element equal to 1 may specify that sign bit hiding is used for the current slice. If sh_sign_data_hiding_used_flag is absent, sh_sign_data_hiding_used_flag is inferred to be equal to 0. In some embodiments, if the value of no_tsrc_constraint_flag is equal to 1, the value of sh_sign_data_hiding_used_flag may be equal to 0. For the current slice, DQ and SDH may be disabled at the SH level.
[0127]
[0146] In the following embodiments of the present disclosure, conditional signaling can be used to disable the TSRC to update the SPS syntax table. In some embodiments, conditional signaling can be used to disable the TSRC to update the SH syntax table.
[0128]
[0147] At the SPS level or the SH level, SDH may be enabled when DQ is disabled according to the SPS syntax table and the SH syntax table. To prevent DQ enablement for the disabled DQ, a first signal is used to disable SDH after DQ is disabled. In some embodiments, in response to determining that the first signal satisfies a given condition, it is determined whether DQ is disabled for at least one slice, and in response to DQ being disabled, SDH is disabled for the at least one slice. For example, if it is determined that the no_tsrc_constraint_flag syntax element is equal to 1, it is determined whether DQ is disabled. If it is determined that DQ is disabled, SDH is disabled.
[0129]
[0148] In some embodiments, the SPS syntax table is updated to disable TSRC for all slices. Figure 13 shows the SPS syntax table of the proposed method. For example, if no_tsrc_constraint_flag is equal to 0, then sps_dep_quant_enabled_flag and sps_sign_data_hiding_enabled_flag are conditionally signaled. The SPS syntax table modifications proposed below can be added to the VVC standard or implemented in other video coding technologies. For example, as shown in elements 1301 and 1302 of Figure 13, the syntax of "!no_tsrc_constraint_flag" is added. If the no_tsrc_constraint_flag syntax element is equal to 1, then the value of "!no_tsrc_constraint_flag" is equal to 0 and the value of sps_dep_quant_enabled_flag is set to 0. An sps_dep_quant_enabled_flag syntax element equal to 0 specifies that dependent quantization is disabled at the SPS level for all slices. Within element 1302, if the no_tsrc_constraint_flag syntax element is equal to 1, the value of "!no_tsrc_constraint_flag" is equal to 0 and the value of sps_sign_data_hiding_enabled_flag is still set to 0 even if "!sps_dep_quant_enabled_flag" is equal to 1. An sps_sign_data_hiding_enabled_flag syntax element equal to 0 specifies that sign data hiding is disabled at the SPS level for all slices.
[0130]
[0149] As noted, if both DQ and SDH are disabled for a slice, the TSRC for the slice may be disabled. Since both DQ and SDH are disabled for all slices, TSRC is disabled for all slices. In some embodiments, the SH syntax table is updated to disable the TSRC for the slice. Figure 14 shows the SH syntax table of the proposed method. For example, if no_tsrc_constraint_flag is equal to 0, the slice-level DQ sh_dep_quant_used_flag and SDH sh_sign_data_hiding_used_flag are conditionally signaled. As shown in elements 1401 and 1402 of Figure 14, "!no_tsrc_constraint_flag" is added in the syntax. If the no_tsrc_constraint_flag syntax element is equal to 1, the sh_dep_quant_used_flag syntax element is set to 0. As shown in element 1402, if the no_tsrc_constraint_flag syntax element is equal to 1, the sh_sign_data_hiding_used_flag syntax element is still set to 0 even if the value of "!sh_sign_data_hiding_used_flag" is 1. The sh_ts_residual_coding_disabled_flag syntax element is set to 1 because sh_ts_residual_coding_disabled_flag is conditionally signaled when both the sh_dep_quant_used_flag syntax element and the sh_sign_data_hiding_used_flag syntax element are equal to 0.
[0131]
[0150] In some embodiments, a semantic constraint is applied to no_tsrc_constraint_flag to disable transform skip mode. For example, a no_tsrc_constraint_flag syntax element equal to 1 may specify that sps_transform_skip_enabled_flag and sh_ts_residual_coding_disabled_flag may be equal to 0. In some embodiments, when the no_tsrc_constraint_flag syntax element is equal to 1, the sps_transform_skip_enabled_flag syntax element is equal to 0. When the no_tsrc_constraint_flag syntax element is equal to 0, no such constraint can be imposed.
[0132]
[0151] In some embodiments, the general constraint syntax is updated by removing the no_tsrc_constraint_flag syntax element. It should be noted that disabling DQ or SDH may affect coding efficiency. If transform skip is disabled at the SPS level, TSRC may be disabled. The constraint flag "no_transform_skip_constraint_flag" can be used to disable transform skip mode. The no_transform_skip_constraint_flag syntax element equal to 1 may specify that transform skip mode is disabled. When the no_transform_skip_constraint_flag syntax element is equal to 1, there is no transform skip mode, so TSRC is implicitly disabled. Therefore, the functionality of no_tsrc_constraint_flag overlaps with that of no_transform_skip_constraint_flag. To disable TSRC, the no_tsrc_constraint_flag syntax element can be removed from the example general constraint syntax shown in FIG. 15, and the no_transform_skip_constraint_flag syntax element can be used without the need to disable DQ and SDH.
[0133]
[0152] The third shortcoming of the current VVC design concerns the ordering of syntax. For example, in VVC Draft 9, the ordering of syntax is not modular. In the SPS syntax table, the syntax related to transformations is scattered in different places in the syntax table. In the picture header (PH) and slice header (SH), some of the syntax elements for in-loop filters are at the top of the syntax table, while others are at the bottom.
[0134]
[0153] In some embodiments, as shown in Figures 16A-H, syntax elements in an SPS syntax table are ordered so that all syntax elements related to transforms are placed together. Syntax elements related to maximum transform size (e.g., element 1601 in Figure 16C), transform skipping (e.g., element 1602 in Figure 16D), multiple transform sets (MTS) (e.g., element 1603 in Figure 16D), and low-frequency non-separable transforms (LFNSTs) (e.g., element 1604 in Figure 16D) are signaled consecutively in the example SPS syntax table. Similarly, sps_lmcs_enabled_flag (e.g., element 1605 in Figure 16D) is signaled immediately after the signaling of sps_ccalf_enabled_flag. The syntax elements marked with strikethrough are where the syntax elements were located before reordering.
[0135]
[0154] In some embodiments, the picture header (PH) syntax is reordered as shown in Figures 17A-F. Picture header (PH) syntax elements related to the SAO and deblocking processes are signaled immediately after the signaling of syntax related to ALF and LMCS, for example as shown in elements 17011 and 17012 in Figures 17B-C. The syntax elements marked with strikethrough are where the syntax elements were located before the reordering.
[0136]
[0155] In some embodiments, the slice header (SH) syntax is reordered as shown in Figures 18A-E. For example, slice header syntax elements related to SAO and deblocking processes (e.g., element 1801 in Figure 18B) are signaled immediately after syntax related to ALF and LMCS. The syntax elements marked with strikethrough are where the syntax elements were located before reordering. It will be understood that the above embodiments can be combined during implementation.
[0137]
[0156] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which may be executed by an apparatus (such as the disclosed encoders and decoders) to perform the above-described methods. Common non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. An apparatus may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0138]
[0157] It should be noted that relational terms such as "first" and "second" herein are used merely to distinguish one entity or operation from another and do not require or imply any actual relationship or order between those entities or operations. Furthermore, terms such as "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and are open-ended in that the items following any one of these terms are not intended to be an exhaustive list of such items or to be limited only to the items they list.
[0139]
[0158] As used herein, unless otherwise specified, the word "or" includes all possible combinations unless impracticable. For example, if a database is stated to include A or B, the database can include A, B, A and B, unless otherwise specified or impracticable. As a second example, if a database is stated to include A, B or C, the database can include A, B, C, A and B, A and C, B and C, A and B and C, unless otherwise specified or impracticable.
[0140]
[0159] It will be understood that the above-described embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computational units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.
[0141]
[0160] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications to the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims. The order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to the particular order of steps. As such, one skilled in the art will recognize that steps can be performed in different orders while implementing the same method.
[0142]
[0161] Although illustrative embodiments have been disclosed in the drawings and herein, many variations and modifications to those embodiments may be made. Accordingly, although specific terms have been employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A computer-implemented method for processing video content, comprising: receiving a bitstream containing video content; determining whether a first signal associated with the video content satisfies a given condition; and and disabling both a cross-component adaptive loop filter (CCALF) process and a chroma adaptive loop filter (ALF) process in response to determining that the first signal satisfies the given condition. A method comprising:
2. Determining whether the first signal satisfies the given condition includes: determining whether the first signal indicates that an Adaptation Parameter Set (APS) Network Abstraction Layer (NAL) unit is present in the received bitstream; The method of claim 1 , comprising:
3. Providing a second signal for controlling the chroma ALF process at a picture parameter set level, a sequence parameter set level, a picture header level, or a slice header level. The method of claim 1 further comprising:
4. the second signal indicates whether the ALF process is enabled when decoding chroma components of pictures in a coded layer video sequence (CLVS) associated with the video content; The method of claim 3.
5. the second signal indicates whether the ALF process is enabled when decoding chroma components of pictures that reference a Picture Parameter Set (PPS) associated with the video content. The method of claim 3.
6. 1. A computer-implemented method for processing video content, comprising: receiving a bitstream containing video content; determining whether a first signal associated with the video content satisfies a given condition; and disabling dependent quantization (DQ) and sign data hiding (SDH) for at least one slice in response to determining that the first signal satisfies the given condition. A method comprising:
7. disabling transform skip residual coding (TSRC) for at least one slice in response to determining that the first signal satisfies the given condition. The method of claim 6 further comprising:
8. in response to the DQ and the SDH being disabled for the at least one slice, disabling the TSRC for the at least one slice; The method of claim 6.
9. disabling the DQ and the SDH for the at least one slice in response to determining that the first signal satisfies the given condition; and disabling the DQ and the SDH for the at least one slice at a slice header level, a picture header level, or a sequence parameter set level in response to the determining that the first signal satisfies the given condition. The method of claim 6, comprising:
10. disabling the DQ and the SDH for the at least one slice in response to determining that the first signal satisfies the given condition; determining whether the DQ is disabled for the at least one slice in response to determining that the first signal satisfies the given condition; and disabling the SDH for the at least one slice in response to the DQ being disabled. The method of claim 6, comprising:
11. An apparatus for processing video content, a memory for storing a set of instructions; and one or more processors, wherein the one or more processors: receiving a bitstream containing video content; determining whether a first signal associated with the video content satisfies a given condition; and and disabling both a cross-component adaptive loop filter (CCALF) process and a chroma adaptive loop filter (ALF) process in response to determining that the first signal satisfies the given condition. an apparatus configured to execute the set of instructions to cause the apparatus to perform
12. Determining whether the first signal satisfies the given condition includes: determining whether the first signal indicates that an Adaptation Parameter Set (APS) Network Abstraction Layer (NAL) unit is present in the received bitstream; 12. The device of claim 11, comprising:
13. the one or more processors Providing a second signal for controlling the chroma ALF process at a picture parameter set level, a sequence parameter set level, a picture header level, or a slice header level.
12. The device of claim 11, configured to execute the set of instructions to further cause the device to:
14. the second signal indicates whether the ALF process is enabled when decoding chroma components of pictures in a coded layer video sequence (CLVS) associated with the video content; 14. The device of claim 13.
15. the second signal indicates whether the ALF process is enabled when decoding chroma components of pictures that reference a Picture Parameter Set (PPS) associated with the video content.
14. The device of claim 13.
16. 1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer to cause the computer to perform a method for processing video content, the method comprising: receiving a bitstream containing video content; determining whether a first signal associated with the video content satisfies a given condition; and disabling dependent quantization (DQ) and sign data hiding (SDH) for at least one slice in response to determining that the first signal satisfies the given condition.
1. A non-transitory computer-readable medium comprising:
17. The at least one processor disabling transform skip residual coding (TSRC) for at least one slice in response to determining that the first signal satisfies the given condition.
20. The non-transitory computer-readable medium of claim 16, configured to execute the set of instructions to cause the computer to further:
18. in response to the DQ and the SDH being disabled for the at least one slice, disabling the TSRC for the at least one slice; 17. The non-transitory computer-readable medium of claim 16.
19. disabling the DQ and the SDH for the at least one slice in response to determining that the first signal satisfies the given condition; and disabling the DQ and the SDH for the at least one slice at a slice header level, a picture header level, or a sequence parameter set level in response to the determining that the first signal satisfies the given condition.
20. The non-transitory computer-readable medium of claim 16, comprising:
20. disabling the DQ and the SDH for the at least one slice in response to determining that the first signal satisfies the given condition; determining whether the DQ is disabled for the at least one slice in response to determining that the first signal satisfies the given condition; and disabling the SDH for the at least one slice in response to the DQ being disabled.
20. The non-transitory computer-readable medium of claim 16, comprising:
Citation Information
Patent Citations
Header parameter set for video coding
WO2020097232A1
High level bitstream syntax for quantization parameters
WO2021180165A1