Method and system for processing luma and chroma signals
Luma mapping with chroma scaling enhances video coding efficiency by optimizing luma and chroma processing, achieving reduced bandwidth usage while maintaining quality in video processing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2026-03-06
AI Technical Summary
Existing video coding techniques face challenges in achieving high coding efficiency, particularly in handling luma and chroma components, leading to suboptimal bandwidth utilization and quality preservation in video processing.
Implementing luma mapping with chroma scaling, which includes processing luma samples to determine a chroma scale factor for chroma samples, enhancing coding efficiency by better utilizing the range of luma code values and improving subjective quality.
The method and system achieve improved coding efficiency, allowing for the same subjective quality with reduced bandwidth, particularly in standard and high dynamic range video signals.
Smart Images

Figure 0007825687000007 
Figure 0007825687000008 
Figure 0007825687000009
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 865,815, filed June 24, 2019, which is incorporated by reference in its entirety.
[0002] Technical Field
[0002] The present disclosure relates generally to video processing, and more particularly to methods and systems for luma mapping with chroma scaling. [Background technology]
[0003]
[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify specific video coding formats, are developed by standardization organizations. As more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes higher. Summary of the Invention [Means for solving the problem]
[0004] Disclosure Overview
[0004] Embodiments of the present disclosure provide methods and systems for performing in-loop luma mapping with chroma scaling and cross-component linear models.
[0005]
[0005] In one exemplary embodiment, the method includes receiving data representing a first block and a second block in a picture, the data including a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determining an average value of the plurality of luma samples associated with the second block; determining a chroma scale factor for the first block based on the average value; and processing the plurality of chroma samples associated with the first block using the chroma scale factor.
[0006]
[0006] In some embodiments, the system includes a memory for storing a set of instructions and at least one processor, the at least one processor configured to execute the set of instructions to cause the system to receive data representing a first block and a second block in a picture, the data including a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block, determine an average value of the plurality of luma samples associated with the second block, determine a chroma scale factor for the first block based on the average value, and process the plurality of chroma samples associated with the first block using the chroma scale factor.
[0007] BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]
[0008] [Figure 1] 8 illustrates an example structure of a video sequence according to some embodiments of the present disclosure. [Figure 2A]
[0009] 1 illustrates a schematic diagram of an example encoding process according to some embodiments of the present disclosure. [Figure 2B]
[0010] 10 shows a schematic diagram of another example of an encoding process according to some embodiments of the present disclosure. [Figure 3A]
[0011] 1 illustrates a schematic diagram of an example of a decoding process according to some embodiments of the present disclosure. [Figure 3B]
[0012] 10 shows a schematic diagram of another example of a decoding process according to some embodiments of the present disclosure. [Figure 4]
[0013] 1 illustrates a block diagram of an example of a device for encoding or decoding video, according to some embodiments of the present disclosure. [Figure 5]
[0014] 1 shows a schematic diagram of an exemplary luma mapping with chroma scaling (LMCS) process, according to some embodiments of the present disclosure. [Figure 6]
[0015] 10 is a tile group level syntax table for LMCS, according to some embodiments of the present disclosure. [Figure 7]
[0016] 10 is another tile group level syntax table for LMCS, according to some embodiments of the present disclosure. [Figure 8]
[0017] 1 is a slice-level syntax table for LMCS, according to some embodiments of the present disclosure. [Figure 9]
[0018] 1 is a syntax table for an LMCS piecewise linear model, according to some embodiments of the present disclosure. [Figure 10]
[0019] 10 illustrates an example of sample locations used to derive α and β, according to some embodiments of the present disclosure. [Figure 11]
[0020] 10 is a table for deriving a chroma prediction mode from a luma mode when CCLM is enabled, according to some embodiments of the present disclosure. [Figure 12]
[0021] 1 is an exemplary coding tree unit syntax structure according to some embodiments of the present disclosure. [Figure 13]
[0022] 1 is an exemplary dual tree split syntax structure according to some embodiments of the present disclosure. [Figure 14]
[0023] 1 is an exemplary coding tree unit syntax structure according to some embodiments of the present disclosure. [Figure 15]
[0024] 1 is an exemplary dual tree split syntax structure according to some embodiments of the present disclosure. [Figure 16A]
[0025] 1 illustrates an exemplary chroma tree partitioning according to some embodiments of the present disclosure. [Figure 16B]
[0026] 1 illustrates an example luma tree partitioning according to some embodiments of the present disclosure. [Figure 17]
[0027] 10 illustrates an exemplary simplification of an averaging operation according to some embodiments of the present disclosure. [Figure 18]
[0028] 10A-10C illustrate example samples used in an average calculation to derive a chroma scale factor, according to some embodiments of the present disclosure. [Figure 19]
[0029] 10 illustrates an example of deriving a chroma scale factor for a block at the right or bottom boundary of a picture, according to some embodiments of the present disclosure. [Figure 20]
[0030] 1 is an exemplary coding tree unit syntax structure according to some embodiments of the present disclosure. [Figure 21]
[0031] 10 is another exemplary coding tree unit syntax structure according to some embodiments of the present disclosure. [Figure 22]
[0032] 10 is an example modified signaling of an LMCS piecewise linear model at the slice level, in accordance with some embodiments of the present disclosure. [Figure 23]
[0033] 1 is a flowchart of an exemplary method for processing video content, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0009] Detailed Description
[0034] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise indicated, like numerals in different figures represent the same or similar elements. The implementations described in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with aspects related to the present invention as recited in the appended claims. Unless otherwise specified, the term "or" encompasses all possible combinations, unless impracticable. For example, if a component is stated to include A or B, that component can include A or B, or A and B, unless otherwise specified or impracticable. As a second example, if a component is stated to include A, B, or C, that component can include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise specified or impracticable.
[0010]
[0035] A video is a set of still pictures (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in chronological order, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display these pictures in chronological order. Furthermore, in some applications, a video capture device can transmit captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as surveillance, conferencing, or live broadcasting.
[0011]
[0036] To reduce the storage space and transmission bandwidth required by such applications, video can be compressed before storage and transmission and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or dedicated hardware. A module for compression is generally referred to as an "encoder," and a module for decompression is generally referred to as a "decoder." Encoders and decoders can be collectively referred to as a "codec." Encoders and decoders can be implemented as various suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be called a "transcoder."
[0012]
[0037] A video coding process can identify and retain useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, then such a coding process can be called "lossy." Otherwise, such a coding process can be called "lossless." Most coding processes are lossy; this is a tradeoff to reduce the required storage space and transmission bandwidth.
[0013]
[0038] Useful information about the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include pixel position changes, luminance changes, or color changes, of which position changes are the most relevant. Position changes of pixels representing an object can reflect the object's movement between the reference picture and the current picture.
[0014]
[0039] A picture that is coded without reference to another picture (i.e., the picture is its own reference picture) is called an "I-picture." A picture that is coded using a past picture as a reference picture is called a "P-picture." A picture that is coded using both past and future pictures as reference pictures (i.e., the referencing is "bidirectional") is called a "B-picture."
[0015]
[0040] As mentioned above, one of the goals in developing new video coding techniques is to improve coding efficiency, i.e., to use less coded data to represent the same picture quality. This disclosure provides a method and system for luma mapping with chroma scaling. Luma mapping is a process for mapping luma samples for use in the loop filter, and chroma scaling is a luma-dependent process for scaling chroma residual values. The same subjective quality as HEVC / H.265 using half the bandwidth, LMCS has two main components: 1) a process for mapping input luma code values to a new set of code values for use in the coding loop, and 2) a luma-dependent process for scaling chroma residual values. The luma mapping process improves coding efficiency for standard and high dynamic range video signals by better utilizing the range of luma code values allowed at a specified bit depth.
[0016]
[0041] 1 illustrates the structure of an example video sequence 100 using video coding according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.
[0017]
[0042] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) need not immediately precede or follow that picture. For example, the reference picture of picture 104 may be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the examples shown in FIG.
[0018]
[0043] Typically, video codecs do not encode or decode an entire picture at once because such a task is computationally complex. Rather, video codecs may divide a picture into elementary segments and encode or decode the picture segment by segment. In this disclosure, such elementary segments are referred to as basic processing units ("BPUs"). For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) and as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Basic processing units can have variable sizes within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for a picture based on a balance between coding efficiency and the level of detail one wishes to preserve within the basic processing unit.
[0019]
[0044] A basic processing unit may be a logical unit that may include various types of video data stored in computer memory (e.g., a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements of the basic processing unit, where the luma and chroma components may have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit can be repeated for each of its luma and chroma components.
[0020]
[0045] Video coding involves multiple operational stages, examples of which are detailed in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to process and therefore may be further divided into segments referred to as "basic processing sub-units" in this disclosure. In some embodiments, the basic processing sub-units may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-units may have the same or smaller size than the basic processing units. Similar to basic processing units, basic processing sub-units are also logical units that may contain various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated for each of its luma and chroma components. It should be noted that such division can be performed to further levels depending on the processing needs. It should also be noted that different stages can divide the basic processing unit using different schemes.
[0021]
[0046] For example, in a mode decision stage (one example of which is detailed in FIG. 2B ), the encoder may decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large to make such a decision. The encoder may divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing sub-unit.
[0022]
[0047] In another example, in the prediction stage (one example of which is detailed in FIG. 2A), the encoder can perform prediction operations at the level of elementary processing sub-units (e.g., CUs). However, in some cases, elementary processing sub-units may still be too large to process. The encoder can further divide the elementary processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.
[0023]
[0048] In another example, in the transform stage (one example of which is detailed in FIG. 2A ), the encoder may perform a transform operation on a residual elementary processing sub-unit (e.g., a CU). However, in some cases, the elementary processing sub-unit may still be too large to process. The encoder may further divide the elementary processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. It should be noted that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0024]
[0049] In structure 110 of Figure 1, basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown by dotted lines. Different basic processing units of the same picture can be divided into basic processing sub-units in different ways.
[0025]
[0050] In some implementations, to provide parallel processing and error resilience for video encoding and decoding, a picture can be divided into regions for processing, thereby allowing the encoding or decoding process for a region of a picture to not depend on information from any other region of the picture. In other words, each region of a picture can be processed independently. This allows a codec to process different regions of a picture in parallel, thus increasing coding efficiency. Furthermore, if data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience. Some video coding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that various pictures in video sequence 100 may have different partitioning schemes for dividing the picture into regions.
[0026]
[0051] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.
[0027]
[0052] FIG. 2A shows a schematic diagram of an example encoding process 200A according to some embodiments of the present disclosure. An encoder may follow process 200A to encode a video sequence 202 into a video bitstream 228. Similar to video sequence 100 of FIG. 1, video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to structure 110 of FIG. 1, each original picture of video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of the basic processing units for each original picture of video sequence 202. For example, the encoder may perform process 200A in an iterative manner, where the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions (e.g., regions 114-118) of each original picture of video sequence 202.
[0028]
[0053] In FIG. 2A , an encoder may feed a basic processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder may feed the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may feed quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to predicted BPU 208 to generate a prediction reference 224 used in prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0029]
[0054] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a predicted reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.
[0030]
[0055] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any action in any manner to input data.
[0031]
[0056] In the prediction stage 204, in the current iteration, the encoder receives the original BPU and a prediction reference 224 and can perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 can be generated from the reconstruction path of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from the prediction data 206 and the prediction reference 224 that can be used to reconstruct the original BPU as a predicted BPU 208.
[0032]
[0057] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder may subtract it from the original BPU to generate residual BPU 210. For example, the encoder may subtract pixel values (e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values of the original BPU. As a result of such subtraction between corresponding pixels of the original BPU and predicted BPU 208, each pixel of residual BPU 210 may have a residual value. Compared to the original BPU, prediction data 206 and residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant loss of quality.
[0033]
[0058] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency (e.g., luminance variation frequency) component of the residual BPU 210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, the basis patterns are similar to basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.
[0034]
[0059] Different transform algorithms can use different basis patterns. For example, various transform algorithms can be used in transform stage 212, such as a discrete cosine transform, a discrete sine transform, etc. The transform in transform stage 212 is reversible. That is, the encoder can reconstruct residual BPU 210 by inversely operating the transform (referred to as an "inverse transform"). For example, to reconstruct a pixel of residual BPU 210, the inverse transform can multiply the value of the corresponding pixel in the basis pattern by the associated respective coefficient and add the products to obtain a weighted sum. In video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Although the transform coefficients may have fewer bits compared to residual BPU 210, they can be used to reconstruct residual BPU 210 without significant loss of quality. Therefore, the residual BPU 210 is further compressed.
[0035]
[0060] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore high-frequency fluctuation information without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest neighbor. After such an operation, some transform coefficients of the high-frequency basis patterns can be converted to zero, and some transform coefficients of the low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also reversible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients by the inverse operation of quantization (referred to as "dequantization").
[0036]
[0061] Quantization stage 214 may be lossy because the encoder ignores any remainder of such division in a rounding operation. Typically, quantization stage 214 may contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values of the quantization parameter or any other parameter of the quantization process.
[0037]
[0062] In the binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary coding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, parameters of the prediction operation, the type of transform in the transform stage 212, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder may generate a video bitstream 228 using the output data of the binary coding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0038]
[0063] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.
[0039]
[0064] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, an encoder can perform the stages of process 200A in a different order. In some embodiments, one or more stages of process 200A can be combined into a single stage. In some embodiments, a single stage of process 200A can be separated into multiple stages. For example, transform stage 212 and quantization stage 214 can be combined into a single stage. In some embodiments, process 200A can include additional stages. In some embodiments, process 200A can omit one or more stages of FIG. 2A.
[0040]
[0065] 2B shows a schematic diagram of another example encoding process 200B according to some embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder conforming to a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and separates prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0041]
[0066] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels of one or more already coded neighboring BPUs in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions of one or more already coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0042]
[0067] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra prediction. With respect to an original BPU of a picture being coded, the prediction reference 224 may include one or more neighboring BPUs coded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating the value of a corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction specified within the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.
[0043]
[0068] In another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. With respect to the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, a reference picture may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all the reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region for a range (referred to as a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered at a location having the same coordinates in the reference picture as the original BPU in the current picture, and may extend over a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pel recursion algorithm, a block matching algorithm, etc.), the encoder can determine that region as a matching region. The matching region may have different dimensions (e.g., smaller, equal, larger, or different shape) than the original BPU. Because the reference picture and the current picture are separated in time in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., picture 106 in FIG. 1), the encoder can find the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the matching region in each matching reference picture.
[0044]
[0069] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, scaling, etc. In inter prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, etc.
[0045]
[0070] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference picture according to the motion vector, and thus, the encoder may predict the original BPU of the current picture. When multiple reference pictures are used (e.g., picture 106 of FIG. 1), the encoder may move the matching region of the reference picture according to each motion vector and average the pixel values of the matching region. In some embodiments, if the encoder assigns weights to the pixel values of the matching region of each matching reference picture, the encoder may add a weighted sum of the pixel values of the moved matching region.
[0046]
[0071] In some embodiments, inter-prediction can be unidirectional or bidirectional. Unidirectional inter-prediction can use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter-prediction can use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture in which the reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.
[0047]
[0072] Continuing with reference to the forward path of process 200B, after spatial prediction step 2042 and temporal prediction step 2044, in mode decision step 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrates of the candidate prediction modes and the distortion of the reconstructed reference picture under the candidate prediction modes. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.
[0048]
[0073] In the reconstruction path of process 200B, if an intra-prediction mode is selected in the forward path, after generating a prediction reference 224 (e.g., a current BPU that has been coded and reconstructed in a current picture), the encoder can directly feed the prediction reference 224 to a spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If an inter-prediction mode is selected in the forward path, after generating a prediction reference 224 (e.g., a current picture in which all BPUs have been coded and reconstructed), the encoder can feed the prediction reference 224 to a loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) caused by inter-prediction. For example, the encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop filtered reference pictures may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., for use as inter-predicted reference pictures for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in a binary coding stage 226 along with the quantized transform coefficients 216, the prediction data 206, and other information.
[0049]
[0074] FIG. 3A shows a schematic diagram of an example decoding process 300A according to some embodiments of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for a region (e.g., regions 114-118) of each picture encoded in video bitstream 228.
[0050]
[0075] In FIG. 3A , a decoder may feed a portion of a video bitstream 228 associated with a basic processing unit of a coded picture (referred to as a “coded BPU”) to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may feed the predicted reference 224 to the prediction stage 204 for performing a prediction operation in a next iteration of the process 300A.
[0051]
[0076] The decoder may iteratively perform process 300A to decode each coded BPU of the coded picture and generate a predicted reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.
[0052]
[0077] In binary decoding stage 302, the decoder may perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as, for example, a prediction mode, parameters of the prediction operation, type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if video bitstream 228 is transmitted over a network in packets, the decoder may depacketize video bitstream 228 before feeding it to binary decoding stage 302.
[0053]
[0078] 3B shows a schematic diagram of another example decoding process 300B according to some embodiments of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.
[0054]
[0079] In process 300B, prediction data 206 decoded by the decoder from binary decoding stage 302 for a coded basic processing unit (referred to as a "current BPU") of a coded picture to be decoded (referred to as a "current picture") may include various types of data, depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of extrapolation, directions of the neighboring BPUs relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter-prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, etc.
[0055]
[0080] Based on the prediction mode indicator, the decoder may decide whether to perform spatial prediction (e.g., intra prediction) in a spatial prediction step 2042 or temporal prediction (e.g., inter prediction) in a temporal prediction step 2044. Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. As described in FIG. 3A, the decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.
[0056]
[0081] In process 300B, the decoder may feed the predicted reference 224 to a spatial prediction stage 2042 or a temporal prediction stage 2044 for performing a prediction operation within a next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may feed the prediction reference 224 to a loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B . The loop filtered reference pictures may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as inter-prediction reference pictures for future coded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include loop filter parameters (e.g., loop filter strength).
[0057]
[0082] FIG. 4 is a block diagram of an example device 400 for encoding or decoding video, in accordance with some embodiments of the present disclosure. As shown in FIG. 4, device 400 may include a processor 402. When processor 402 executes the instructions described herein, device 400 may become a dedicated machine for encoding or decoding video. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (“CPUs”), graphics processing units (“GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, integrated circuit (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may be a set of processors grouped together as a single logical entity. For example, as shown in Figure 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0058]
[0083] Device 400 may also include memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to operate on or process the data for processing. Memory 404 may include high-speed random access storage or non-volatile storage. In some embodiments, memory 404 may include any combination of any number of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped together as a single logical entity.
[0059]
[0084] Bus 410, such as an internal bus (e.g., a CPU memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like, may be a communication device that transfers data between components within device 400.
[0060]
[0085] For ease of explanation and to avoid ambiguity, this disclosure will collectively refer to the processor 402 and other data processing circuitry as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuitry may be a single, independent module or may be fully or partially combined within any other component of the device 400.
[0061]
[0086] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.) In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0062]
[0087] In some embodiments, device 400 may optionally further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or input interface coupled to a video archive), etc.
[0063]
[0088] It should be noted that a video codec (e.g., a codec that executes process 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).
[0064]
[0089] 5 shows a schematic diagram of an exemplary luma mapping with chroma scaling (LMCS) process 500 according to some embodiments of the present disclosure. For example, the process 500 may be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). The LMCS is a new processing block that is applied before the loop filter 232 in FIG. 2B. The LMCS may also be referred to as a reshaper.
[0065]
[0090] The LMCS process 500 may include an in-loop mapping of luma component values based on an adaptive piecewise linear model and luma-dependent chroma residual scaling of chroma components.
[0066]
[0091] 5, the in-loop mapping of luma component values based on an adaptive piecewise linear model may include a forward mapping stage 518 and an inverse mapping stage 508. The luma-dependent chroma residual scaling of the chroma components may include chroma scaling 520.
[0067]
[0092] The sample values before mapping or after reverse mapping may be referred to as samples in the original domain, and the sample values after mapping and before reverse mapping may be referred to as samples in the mapped domain. When LMCS is enabled, some stages in the process 500 may be performed in the mapped domain rather than the original domain. It will be appreciated that the forward mapping stage 518 and the reverse mapping stage 508 may be enabled / disabled at the sequence level using the SPS flag.
[0068]
[0093] As shown in Figure 5, Q -1 &T -1 Steps 504, reconstruction 506, and intra prediction 514 are performed within the map domain. -1 &T -1 Stage 504 may include inverse quantization and inverse transform, reconstruction 506 may include summing luma prediction and luma residual, and intra prediction 508 may include luma intra prediction.
[0069]
[0094] The loop filter 510, motion compensation stages 516 and 530, intra prediction stage 528, reconstruction stage 522, and decoded picture buffers (DPBs) 512 and 526 are performed in the original (i.e., unmapped) domain. In some embodiments, the loop filter 510 may include deblocking, an adaptive loop filter (ALF), and a sample adaptive offset (SAO), the reconstruction stage 522 may include chroma prediction and addition of chroma residual, and the DPBs 512 and 526 may store decoded pictures as reference pictures.
[0070]
[0095] In some embodiments, a method for processing video content using luma mapping according to a piecewise linear model may be applied.
[0071]
[0096] In-loop mapping of the luma component can adjust the signal statistics of the input video by redistributing codewords across the dynamic range to improve compression efficiency. Luma mapping utilizes a forward mapping function "FwdMap" and a corresponding inverse mapping function "InvMap". The "FwdMap" function is signaled using a piecewise linear model with 16 equal sections. The "InvMap" function does not need to be signaled, but instead is derived from the "FwdMap" function.
[0072]
[0097] The signaling of the piecewise linear model is shown in Table 1 of FIG. 6 and Table 2 of FIG. 7. Later, in Draft 5 of VVC, the signaling of the piecewise linear model was changed to Table 3 of FIG. 8 and Table 4 of FIG. 9. Tables 1 and 3 show the syntax structure of the tile group header and slice header. A reshaper model parameters present flag may be signaled first to indicate whether a luma mapping model is present in the target tile group or target slice. If a luma mapping model is present in the current tile group / slice, the piecewise linear model parameters corresponding to the target tile group or target slice may be signaled in tile_group_reshaper_model() / lmcs_data() using the syntax elements shown in Table 2 of FIG. 7 and Table 4 of FIG. 9. The piecewise linear model divides the dynamic range of the input signal into 16 equal partitions. For each partition, the number of codewords assigned to the partition can be used to represent the linear mapping parameters of the partition. In an example of a 10-bit input, each of the 16 partitions of the input may have 64 codewords assigned to that partition by default. The number of signaled codewords can be used to calculate a scale factor and adjust the mapping function appropriately for that partition. Table 2 of FIG. 7 and Table 4 of FIG. 9 also inclusively define minimum and maximum indices over the number of signaled codewords, such as "reshaper_model_min_bin_idx" and "reshaper_model_delta_max_bin_idx" found in Table 2, and "lmcs_min_bin_idx" and "lmcs_delta_max_bin_idx" found in Table 4. If the partition index is less than "reshaper_model_min_bin_idx" or "lmcs_min_bin_idx," or greater than "15-reshaper_model_max_bin_idx" or "15-lmcs_delta_max_bin_idx," the number of codewords for that partition is not signaled and is inferred to be zero. In other words, no codewords are assigned to the partition and no mapping / scaling is applied.
[0073]
[0098] Another reshaper enable flag (e.g., "tile_group_reshaper_enable_flag" or "slice_lmcs_enabled_flag") may be signaled at the tile group header level or slice header level to indicate whether the LMCS process shown in FIG. 5 is applied to the target tile group or target slice. If the reshaper is enabled for the target tile group or target slice and the target tile group or target slice does not use dual-tree partitioning, an additional chroma scaling enable flag may be signaled to indicate whether chroma scaling is enabled for the target tile group or target slice. It will be understood that dual-tree partitioning may also be referred to as chroma-separate tree. In the following, the present disclosure describes dual-tree partitioning in more detail.
[0074]
[0099] A piecewise linear model can be constructed based on the signaled syntax elements in Table 2 or Table 4 as follows: The i-th piece (i=0...15) of the "FwdMap" piecewise linear model can be defined by two input pivot points InputPivot[ ] and two mapped pivot points MappedPivot[ ]. The mapped pivot points MappedPivot[ ] can be the output of the "FwdMap" piecewise linear model. Assuming the bit depth of an exemplary input video is 10 bits, InputPivot[ ] and MappedPivot[ ] can be calculated based on the signaled syntax as follows: It will be understood that the bit depth can be different from 10 bits. a) Using the syntax elements in Table 2: 1)OrgCW=64 2) For i=0:16, InputPivot[i]=i*OrgCW 3)i=reshaper_model_min_bin_idx: in reshaper_model_max_bin_idx SignaledCW[i]=OrgCW+(1¬2*reshape_model_bin_delta_sign_CW[i])*reshape_model_bin_delta_abs_CW[i]; 4) For i = 0:16, calculate MappedPivot[i] as follows: MappedPivot[0]=0; (i=0; i<16; i++) MappedPivot[i+1]=MappedPivot[i]+SignaledCW[i] b) Using the syntax elements in Table 4: 1)OrgCW=64 2) For i=0:16, InputPivot[i]=i*OrgCW 3)i=lmcs_min_bin_idx: For lmcsl_max_bin_idx, SignaledCW[i]=OrgCW+(1¬2*lmcs_bin_delta_sign_CW[i])*lmcsl_bin_delta_abs_CW[i]; 4) For i = 0:16, calculate MappedPivot[i] as follows: MappedPivot[0]=0; (i=0; i<16; i++) MappedPivot[i+1]=MappedPivot[i]+SignaledCW[i]
[0075]
[0100] The inverse mapping function "InvMap" is defined by InputPivot[] and MappedPivot[]. Unlike "FwdMap", in the "InvMap" piecewise linear model, the two input pivot points for each piece are defined by MappedPivot[] and the two output pivot points are defined by InputPivot[]. In this way, the inputs of "FwdMap" are divided into equal pieces, but the inputs of "InvMap" are not guaranteed to be divided into equal pieces.
[0076]
[0101] As shown in Figure 5, for inter-coded blocks, motion compensation prediction can be performed in the map domain. In other words, after motion compensation prediction 516, Y pred and apply a “FwdMap” function 518 to map the luma prediction blocks in the original region to the map region (Y′ pred =FwdMap(Y pred )). For intra-coded blocks, the "FwdMap" function is not applied because the reference samples used in intra prediction are already in the map region. After the reconstructed block 506, Y r An "InvMap" function 508 can be applied to convert the reconstructed luma values in the map domain back to reconstructed luma values in the original domain (
number
[0077]
[0102] The luma mapping process (forward or inverse mapping) can be performed using a look-up table (LUT) or using on-the-fly calculations. If a LUT is used, the "FwdMapLUT[]" and "InvMapLUT[]" tables can be pre-calculated and pre-stored for use at the tile group or slice level, and the forward and inverse mappings can be simply calculated as FwdMap(Y pred )=FwdMapLUT[Y pred ] and InvMap(Y r )=InvMapLUT[Y r ] can be implemented as follows.
[0078]
[0103] Instead, on-the-fly calculations can be used. Take the forward mapping function "FwdMap" as an example. To identify the partition to which a luma sample belongs, the sample value can be right-shifted by 6 bits (corresponding to 16 equal partitions, assuming 10-bit video) to obtain the partition index. The linear model parameters for that partition are then taken and applied on-the-fly to calculate the mapped luma value. The "FwdMap" function is evaluated as follows: Y'pred=FwdMap(Y pred )=((b2-b1) / (a2-a1))*(Y pred -a1)+b1 where "i" is the partition index, a1 is InputPivot[i], a2 is InputPivot[i+1], b1 is MappedPivot[i], and b2 is MappedPivot[i+1].
[0079]
[0104] The "InvMap" function can be computed on the fly in a similar way, except that the partitions in the map region are not guaranteed to be of equal size, so a conditional check must be applied instead of a simple right bit shift when finding the partition to which a sample value belongs.
[0080]
[0105] In some embodiments, a method may be provided for processing video content using luma-dependent chroma residual scaling.
[0081]
[0106] Chroma residual scaling can be used to compensate for interactions between luma signals and chroma signals corresponding to the luma signal. Whether chroma residual scaling is enabled can also be signaled at the tile group level or slice level. As shown in Table 1 of FIG. 6 and Table 3 of FIG. 8, when luma mapping is enabled and when dual-tree partitioning is not applied to the current tile group, an additional flag (e.g., "tile_group_reshaper_chroma_residual_scale_flag" or "slice_chroma_residual_scale_flag") can be signaled to indicate whether luma-dependent chroma residual scaling is enabled. When luma mapping is not used or dual-tree partitioning is used within the target tile group (or target slice), luma-dependent chroma residual scaling can be disabled accordingly. Furthermore, luma-dependent chroma residual scaling can be disabled for chroma blocks whose area is 4 or less.
[0082]
[0107] Chroma residual scaling depends on the average value of the luma prediction blocks (for both intra-coded and inter-coded blocks) corresponding to the chroma signal. The average "avgY'" of the luma prediction blocks can be determined using the following formula:
number
[0083]
[0108] Use the following steps to find the chroma scale factor value "C" for chroma residual scaling: ScaleInv " can be obtained. 1) Index Y of the piecewise linear model to which avgY' belongs Idx is found based on the InvMap function. 2) C ScaleInv =cScaleInv[Y Idx ] holds, where cScaleInv[] is a pre-computed LUT with, for example, 16 sections.
[0084]
[0109] In some embodiments, in the LMCS method, a pre-computed LUT, cScaleInv[i], where i is in the range of 0 to 15, can be derived based on the 64-entry static LUT ChromaResidualScaleLut and the value of the signaled codeword SignaledCW[i] as follows: ChromaResidualScaleLut
[64] ={ 16384, 16384, 16384, 16384, 16384, 16384, 16384, 8192, 8192, 8192, 8192, 5461, 5461, 5461, 5461, 4096, 4096, 4096, 4096, 3277, 3277, 3277, 3277, 2731, 2731, 2731, 2731, 2341, 2341, 2341, 2048, 2048, 2048, 1820, 1820, 1820, 1638, 1638, 1638, 1638, 1489, 1489, 1489, 1489, 1365, 1365, 1365, 1260, 1260, 1260, 1260, 1170, 1170, 1170, 1092, 1092, 1092, 1024, 1024, 1024, 1024}; shiftC=11 - If (SignaledCW[i]==0) holds, cScaleInv[i]=(1< <shiftC) - Otherwise, cScaleInv[i]=ChromaResidualScaleLut[(SignaledCW[i]>>1)-1]
[0085]
[0110] As an example, assume the input is 10-bit, the static LUT "ChromaResidualScaleLut[]" contains 64 entries, and the signaled codeword "SignaledCW[]" is in the range [0,128]. Therefore, to construct the chroma scale factor LUT "cScaleInv[]", we use a division by 2 (or a right shift by 1). The LUT "cScaleInv[]" can be constructed at the tile group (or slice level).
[0086]
[0111] If the current block can be coded using intra, CIIP, or intra block copy (IBC) mode, avgY' can be determined as the average of the intra, CIIP, or IBC predicted luma values. Otherwise, avgY' is determined as the average of the forward-mapped inter-predicted luma values (i.e., Y' in FIG. 3). pred ) is calculated as the average of the current picture reference (CPR) mode. Unlike luma mapping, which is performed sample-based, IBC uses the "C ScaleInv」 is a constant value for all chroma blocks. ScaleInv」 Using , chroma residual scaling can be applied at the decoder side as follows:
number
[0087]
[0112] where:
number
[0088]
[0113] In some embodiments, a method may be provided for processing video content using cross-component linear model prediction.
[0089]
[0114] To reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode can be used. In CCLM, chroma samples are predicted based on the reconstructed luma samples of the same coded unit (CU) by using a linear model as follows: pred C (i,j)=α·rec L '(i,j)+β
[0090]
[0115] where pred C (i,j) represents the predicted chroma sample in the CU, and rec L (i,j) represents the downsampled reconstructed luma sample of the same CU.
[0091]
[0116] The linear model parameters α and β are derived based on the relationship between the luma and chroma values from two sample positions. The two sample positions may include a first luma sample position having the maximum luma sample value and a second luma sample position having the minimum luma sample value among a set of downsampled adjacent luma samples, and their corresponding chroma samples. The linear model parameters α and β are obtained according to the following equations:
number
[0092]
[0117] where Y a and X a represent the luma and chroma values of the first luma sample position, respectively. b and Y b represent the luma and chroma values of the second luma sample position, respectively.
[0093]
[0118] FIG. 10 illustrates an example of sample positions involved in CCLM mode according to some embodiments of the present disclosure.
[0094]
[0119] The calculation of the parameter α can be implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are expressed in exponential notation. For example, diff is approximated by a 4-bit significant part and an exponent. As a result, the table for 1 / diff is reduced to 16 elements for 16 values of the significant part as follows: DivTable [] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}
[0095]
[0120] The "DivTable[]" table also reduces the computational complexity and reduces the memory size required to store the required tables.
[0096]
[0121] In addition to being able to use the top and left positions together to calculate the linear model coefficients, they can also be used alternatively in two other LM modes, called the LM_A mode and the LM_L mode.
[0097]
[0122] In LM_A mode, only samples in the top position are used to calculate the linear model coefficients. To get more samples, the top position can be extended to include (W+H) samples. In LM_L mode, only samples in the left position are used to calculate the linear model coefficients. To get more samples, the left position can be extended to include (H+W) samples.
[0098]
[0123] For non-square blocks, the top template is extended to W+W and the left template is extended to H+H.
[0099]
[0124] To match the chroma sample positions for 4:2:0 video sequences, two types of downsampling filters can be applied to the luma samples to achieve a 2:1 downsampling ratio both horizontally and vertically. The choice of downsampling filter can be specified by the SPS level flag. The two downsampling filters corresponding to "Type 0" and "Type 2" content, respectively, are as follows:
number
[0100]
[0125] It will be appreciated that if the top reference line is at the boundary of a CTU, only one luma line (a common line buffer in intra prediction) is used to calculate the downsampled luma sample.
[0101]
[0126] The calculation of this parameter can be done as part of the decoding process, rather than simply as a search operation in the encoder. As a result, no syntax is used to communicate the values of α and β to the decoder. The α and β parameters are calculated separately for each of the chroma components.
[0102]
[0127] A total of eight intra modes may be allowed for chroma intra mode coding. These modes include five conventional intra modes and three cross-component linear model modes (e.g., CCLM, LM_A, and LM_L). The process for signaling and deriving chroma modes when CCLM is enabled is shown in Table 5 of Figure 9. Chroma mode coding of a chroma block may depend on the intra prediction mode of the luma block corresponding to the chroma block. Because separate block partitioning structures for luma and chroma components are enabled within an I slice (described below), one chroma block may correspond to multiple luma blocks. Therefore, the chroma derivation mode (DM) inherits the intra prediction mode of the corresponding luma block that covers the center position of the current chroma block.
[0103]
[0128] In some embodiments, a method may be provided for processing video content using dual-tree partitioning.
[0104]
[0129] In the VVC draft, the coding tree scheme supports the ability for luma and chroma to have separate block tree partitions. This is also called dual-tree partitioning. In the VVC draft, the signaling of dual-tree partitioning is shown in Table 6 of Figure 12 and Table 7 of Figure 13. In the later VVC draft 5, dual-tree partitioning is signaled as in Table 8 of Figure 14 and Table 9 of Figure 15. If a sequence-level control flag (e.g., "qtbtt_dual_tree_intra_flag") signaled in the SPS is turned on and the target tile group (or target slice) is intra-coded, block partition information can be signaled separately, first for luma and then for chroma. Dual-tree partitioning is not allowed for inter-coded tile groups / slices (e.g., P and B tile groups / slices). When the separate block tree mode is applied, the luma coding tree blocks (CTBs) are divided into CUs by a first coding tree structure, and the chroma CTBs are divided into chroma CUs by a second coding tree structure, as shown in Table 7 of Figure 13.
[0105]
[0130] If luma and chroma blocks are allowed to have different partitions, problems may arise for coding tools with dependencies between various color components. For example, when LMCS is applied, the average value of the luma blocks corresponding to the target chroma block can be used to determine the scale factor to be applied to the target chroma block. When dual-tree partitioning is used, determining the average value of the luma blocks may incur latency for the entire CTU. For example, if the luma blocks of a CTU are partitioned once vertically and the chroma blocks of the CTU are partitioned once horizontally, both luma blocks of the CTU are decoded to calculate the average value before the first chroma block of the CTU can be decoded. In VVC, a CTU can be as large as 128x128 in units of luma samples, causing a significant increase in the latency of decoding a chroma block. Therefore, VVC Draft 4 and Draft 5 may prohibit the combination of dual-tree partitioning and luma-dependent chroma scaling. When dual-tree partitioning is enabled for a target tile group (or target slice), chroma scaling may be forced off. Note that the luma mapping part of the LMCS is still allowed in dual-tree splitting, since it only affects the luma component and does not have dependency issues across color components.
[0106]
[0131] Another example of a coding tool that relies on dependencies between color components to achieve better coding efficiency is called the cross-component linear model (CCLM), discussed above. In CCLM, neighboring luma and chroma reconstructed samples can be used to derive cross-component parameters. The cross-component parameters can be applied to corresponding reconstructed luma samples of a target chroma block to derive predictors for the chroma components. When dual-tree partitioning is used, it is not guaranteed that the luma and chroma partitions are aligned. Therefore, CCLM cannot begin for a chroma block until all of the corresponding luma blocks containing samples used for CCLM have been reconstructed.
[0107]
[0132] 16A-16B illustrate exemplary chroma tree and luma tree partitioning according to some embodiments of the present disclosure. FIG. 16A illustrates an exemplary partitioning structure for a chroma block 1600. FIG. 16B illustrates an exemplary partitioning structure for a luma block 1610 corresponding to the chroma block 1600 of FIG. 16. In FIG. 16A, the chroma block 1600 is quadranted into four sub-blocks, with the bottom-left sub-block further quadranted into four sub-blocks. The block with the grid pattern is the current block to be predicted. In FIG. 16B, the luma block 1610 is bisected horizontally into two sub-blocks, with the region with the grid pattern corresponding to the target chroma block to be predicted. To derive CCLM parameters, values of neighboring reconstructed samples, represented by solid circles, are required. Therefore, prediction of the target chroma block cannot begin until reconstruction of the lower luma block is complete, which results in significant latency.
[0108]
[0133] In some embodiments, a method may be provided for processing video content using virtual pipeline data units.
[0109]
[0134] The VVC standard introduces the concept of a virtual pipeline data unit (VPDU) for a more hardware-friendly implementation. A VPDU is defined as a non-overlapping MxM luma (L) / NxN chroma (C) unit within a picture. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. Different stages process different VPDUs simultaneously. The size of a VPDU is roughly proportional to the buffer size in most pipeline stages, so keeping the VPDU size small is important. In VVC, the VPDU size is set to 64x64 samples. Therefore, all coding tools employed in VVC cannot violate the VPDU constraints. For example, the maximum transform size can only be 64x64 because all transform blocks must operate in the same pipeline stage. Due to the VPDU constraint, intra-prediction blocks should also be limited to 64x64. Therefore, in an intra-coded tile group / slice (e.g., an I tile group / slice), the CTU is forced to be split into four 64x64 blocks (if the CTU is larger than 64x64), and each 64x64 block can be further split by the dual tree structure. Thus, when the dual tree is enabled, the common root of the luma splitting tree and the chroma splitting tree is at the 64x64 block size.
[0110]
[0135] There are several problems with the current design of the LMCS and CCLM.
[0111]
[0136] First, the derivation of the tile group level chroma scale factor LUT "cScaleInv[]" cannot be easily extended. The derivation process currently relies on a constant chroma LUT "ChromaResidualScaleLut" with 64 entries. For 10-bit video with 16 partitions, an additional step of divide by 2 must be applied. If the number of partitions changes (e.g., if 8 partitions are used instead of 16 partitions), the derivation process must be modified to apply a divide by 4 instead of 2. This additional step can cause a loss of precision.
[0112]
[0137] Second, the Y used to obtain the chroma scale factor, e.g. Idx To calculate the luma mean, the average value of all luma blocks is used. Considering the maximum CTU size of 128x128, the average luma value may be calculated based on 16,384 (128x128) luma samples, and such calculation is complex. Furthermore, when a 128x128 luma block division is selected by the encoder, the block is likely to contain homogeneous content. Therefore, a subset of luma samples in a block may be sufficient to calculate the luma mean.
[0113]
[0138] Third, during dual-tree splitting, chroma scaling is set off to avoid potential pipeline issues in the hardware decoder. However, this dependency can be avoided if explicit signaling is used to indicate the chroma scale factor to be applied (instead of using the corresponding luma sample to derive the chroma scale factor to be applied). Enabling chroma scaling within intra-coded tile groups / slices can further improve coding efficiency.
[0114]
[0139] Fourth, conventionally, delta codeword values are signaled for each of the 16 partitions. In many cases, it is recognized that only a limited number of different codewords are used for the 16 partitions. Thus, signaling overhead can be further reduced.
[0115]
[0140] Fifthly, the CCLM parameters are derived using luma and chroma reconstructed samples from blocks that are causally close to the target chroma block. In dual-tree splitting, luma block splitting and chroma block splitting are not necessarily aligned. Therefore, a luma block having a plurality of luma blocks or an area larger than the target chroma block may correspond to the target chroma block. To derive the CCLM parameters of the target chroma block, as shown in FIGS. 16A to 16B, all corresponding luma blocks must be reconstructed. Such reconstruction causes latency in pipeline implementation and reduces the throughput of the hardware decoder.
[0116]
[0141] To address the above problems, embodiments of the present disclosure are shown as follows.
[0117]
[0142] Embodiments of the present disclosure provide a method for processing video content by removing a chroma scaling LUT.
[0118]
[0143] As described above, a 64-entry chroma LUT is not easily extensible and may cause problems when other piecewise linear models (e.g., 8-piece, 4-piece, 64-piece, etc.) are used. To achieve the same coding efficiency, the chroma scale factor can be set to the same as the corresponding piecewise luma scale factor, so such a LUT is also unnecessary. In some embodiments of the present disclosure, the piecewise index of the current chroma block is represented by Y Idx and the chroma scale factor is determined using the following steps: ·Y Idx >reshaper_model_max_bin_idx or Y Idx <reshaper_model_min_bin_idx holds or SignaledCW[Y Idx]=0, set chroma_scaling to the default, chroma_scaling=1.0, i.e. no scaling is applied. Otherwise, set chroma_scaling to SignaledCW[Y Idx ] / OrgCW.
[0119]
[0144] The chroma scale factors derived above have fractional precision. A fixed-point approximation can be applied to avoid dependencies on hardware / software platforms. Furthermore, the inverse chroma scaling needs to be done on the decoder side. Such division can be implemented by fixed-point arithmetic using a multiplication followed by a right shift. We denote the number of bits in the fixed-point approximation by CSCALE_FP_PREC. The following can be used to determine the inverse chroma scale factor in fixed-point precision: inverse_chroma_scaling[Y Idx ]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(SignaledCW[Y Idx ]>>1)) / SignaledCW[Y Idx ]; where luma_bit_depth is the luma bit depth and TOTAL_NUMBER_PIECES is the total number of pieces in the piecewise linear model, which is set to 16 in VVC Draft 4. Note that the value of inverse_chroma_scaling may only need to be calculated once per tile group / slice, and the division above is an integer division operation.
[0120]
[0145] Further quantization can be applied to derive the chroma scale factors and inverse scale factors. For example, an inverse chroma scale factor can be calculated for every even (2×m) value of SignaledCW, and odd (2×m+1) values of SignaledCW reuse the chroma scale factor of the neighboring even-valued scale factor. In other words, the following can be used: for(i=reshaper_model_min_bin_idx; i<=reshaper_model_max_bin_idx; i++) { tempCW=SignaledCW[i]>>1)<<1; inverse_chroma_scaling[i]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(tempCW>>1)) / tempCW; }
[0121]
[0146] The above embodiment of quantizing the chroma scale factors can be further generalized, for example, to calculate an inverse chroma scale factor for every nth value of SignaledCW, while all other adjacent values share the same chroma scale factor. For example, "n" can be set to 4, meaning that every four adjacent codeword values share the same inverse chroma scale factor value. The value of "n" is preferably a power of 2, which allows for the use of shifts to calculate divisions. By representing the value of log2(n) as LOG2_n, the above formula can be modified as follows: tempCW=SignaledCW[i]>>LOG2_n)< <LOG2_n。
[0122]
[0147] Finally, the value of LOG2_n can be a function of the number of pieces used in the piecewise linear model. If fewer pieces are used, it is beneficial to use a larger LOG2_n. For example, if TOTAL_NUMBER_PIECES is 16 or less, LOG2_n can be set to 1 + (4 - log2(TOTAL_NUMBER_PIECES)). If TOTAL_NUMBER_PIECES is greater than 16, LOG2_n can be set to 0.
[0123]
[0148] SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a method for processing video content by simplifying the averaging of luma prediction blocks.
[0124]
[0149] As discussed above, the current chroma block's division index "Y Idx To determine ", the average value of the corresponding luma block can be used. However, for large block sizes, the averaging process may involve a large number of luma samples. In the worst case, 128x128 luma samples may be involved in the averaging process.
[0125]
[0150] Embodiments of the present disclosure provide a simplified averaging process to reduce the worst case to using only NxN luma samples (N is a power of 2).
[0126]
[0151] In some embodiments, if both dimensions of a 2D luma block are less than or equal to a preset threshold N (in other words, at least one of the two dimensions is greater than N), then "downsampling" can be applied to use only N positions within that dimension. Without loss of generality, take the horizontal dimension as an example. If the width is greater than N, then only samples at position x (x=i×(width>>log2(N)), i=0,...N-1) are used for averaging.
[0127]
[0152] 17 shows an exemplary simplification of the averaging operation according to some embodiments of the present disclosure. In this example, N is set to 4, and only 16 luma samples (shaded samples) in the block are used for averaging. It will be understood that the value of N is not limited to 4. For example, N can be set to any value that is a power of 2. In other words, N can be 1, 2, 4, 8, etc.
[0128]
[0153] In some embodiments, different values of N may be applied to the horizontal and vertical dimensions. In other words, the worst case of the averaging operation may use NxM samples. In some embodiments, the number of samples may be limited within the averaging process without considering the dimensions. For example, a maximum of 16 samples may be used. These 16 samples may be distributed within the horizontal or vertical dimensions in a 1x16, 16x1, 2x8, 8x2, 4x4 format, or in a format that fits the shape of the target block. For example, if the block is vertical, 2x8 is used; if the block is horizontal, 8x2 is used; and if the block is square, 4x4 is used.
[0129]
[0154] Such simplification may cause the average value to differ from the true average of all luma blocks, but any such difference is likely to be small, since when large block sizes are chosen, the content within the blocks tends to be more homogeneous.
[0130]
[0155] Furthermore, the decoder-side motion vector refinement (DMVR) mode is a complex process within the VVC standard, especially for the decoder, because DMVR requires the decoder to perform motion estimation to derive motion vectors before motion compensation can be applied. The bidirectional optical flow (BDOF) mode within the VVC standard can further complicate this situation, because BDOF is an additional sequential process that must be applied after DMVR to obtain the luma prediction block. Because chroma scaling requires the average value of the corresponding luma prediction block, DMVR and BDOF may be applied before the average value can be calculated, causing latency issues.
[0131]
[0156] To solve this latency issue, some embodiments of the present disclosure use luma prediction blocks before DMVR and BDOF to calculate an average luma value, and then use the average luma value to obtain a chroma scale factor, which allows chroma scaling to be applied in parallel with the DMVR and BDOF process, thus significantly reducing latency.
[0132]
[0157] Consistent with this disclosure, variations in latency reduction are contemplated. In some embodiments, this latency reduction can be combined with the simplified averaging process described above, which uses only a portion of the luma prediction block to calculate the average luma value. In some embodiments, the luma prediction block can be used after the DMVR process and before the BDOF process to calculate the average luma value. The average luma value is then used to obtain the chroma scale factor. This design allows chroma scaling to be applied in parallel with the BDOF process while maintaining accuracy in determining the chroma scale factor. The DMVR process may refine the motion vector, and therefore, using the prediction sample with the refined motion vector after the DMVR process may be more accurate than using the prediction sample with the motion vector before the DMVR process.
[0133]
[0158] Furthermore, in the VVC standard, a CU syntax structure (e.g., coding_unit()) may include a syntax element "cu_cbf" to indicate whether there are any non-zero residual coefficients in the target CU. At the TU level, the TU syntax structure transform_unit() includes syntax elements tu_cbf_cb and tu_cbf_cr to indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the target TU. In VVC Draft 4, when chroma scaling is enabled at the tile group level or slice level, averaging of the corresponding luma block can be invoked. The present disclosure also provides a method for bypassing the luma averaging process. Consistent with the disclosed embodiments, since the chroma scaling process is applied to residual chroma coefficients, the luma averaging process can be bypassed if there are no non-zero chroma coefficients. This can be determined based on the following conditions: Condition 1: cu_cbf is equal to 0 Condition 2: tu_cbf_cr and tu_cbf_cb are both equal to 0
[0134]
[0159] If condition 1 or condition 2 is met, the luma averaging process can be bypassed.
[0135]
[0160] In the above embodiment, only NxN samples of the prediction block are used to derive the average value, which simplifies the averaging process. For example, if N is equal to 1, only the top-left sample of the prediction block may be used. However, even in this simplified case, the prediction block needs to be generated first, which causes latency. Therefore, in some embodiments, we consider that the reference luma samples can be directly used to derive the chroma scale factors. This allows the decoder to derive the chroma scale factors in parallel with the luma prediction process, thereby reducing latency. In other words, intra prediction and inter prediction are processed separately.
[0136]
[0161] For intra prediction, already decoded neighboring samples in the same picture can be used as reference samples for generating the prediction block. These reference samples include a sample at the top of the target block, a sample to the left of the target block, and a sample at the top left of the target block. An average of all these reference samples can be used to derive the chroma scale factor. Alternatively, an average of only a portion of these reference samples can be used. For example, only the M reference samples (e.g., M=3) closest to the top left position of the target block can be averaged.
[0137]
[0162] As another example, the M reference samples averaged to derive the chroma scale factor are not closest to the top left position, but are distributed along the top and left boundaries of the target block, as shown in FIG. 18. FIG. 18 shows exemplary samples used in the average calculation to derive the chroma scale factor. As shown in FIG. 18, the exemplary samples are represented by dotted rectangles. In each of the exemplary rectangles 1801-1807 shown in FIG. 18, 1, 2, 3, 4, 5, 6, and 8 samples are averaged. The average calculation of the present disclosure can be replaced with a weighted average, in which different samples can have different weights in the average calculation. For example, to avoid a division operation in the average calculation, the sum of the weights can be a power of 2.
[0138]
[0163] In the case of inter prediction, reference samples from temporal reference pictures can be used to generate the prediction block. These reference samples are identified by a reference picture index and a motion vector. If the motion vector has fractional precision, interpolation can be applied. To calculate the average of the reference samples, either the interpolated reference samples or the pre-interpolated reference samples (i.e., the motion vector clipped to integer precision) can be used. Consistent with the disclosed embodiments, all of the reference samples can be used to calculate the average. Alternatively, only a portion of the reference samples (e.g., the reference sample corresponding to the top-left position of the target block) can be used to calculate the average.
[0139]
[0164] As shown in Figure 5, inter prediction is performed in the original region, while intra prediction is performed in the reshaped region. Thus, in inter prediction, forward mapping is applied to the prediction block, and the luma prediction block after forward mapping is used to calculate the average value of the luma block. To reduce latency, the average value can be calculated using the luma prediction block before forward mapping. For example, the entire luma block before forward mapping, an NxN portion of the luma block before forward mapping, or the top-left sample of the luma block before forward mapping can be used.
[0140]
[0165] Embodiments of the present disclosure further provide a method for processing video content with chroma scaling for dual-tree partitioning.
[0141]
[0166] Because the dependency on luma blocks can cause hardware design complexity, chroma scaling can be turned off for intra-coded tile groups / slices, which allows for dual-tree partitioning, but this restriction can cause a loss of coding efficiency.
[0142]
[0167] Because the CTU is the common root of both the luma coding tree and the chroma coding tree, deriving the chroma scale factor at the CTU level can eliminate the dependency between chroma and luma in the dual-tree division. For example, to derive the chroma scale factor, reconstructed luma or chroma samples adjacent to the CTU are used. This chroma scale factor can be used for all chroma samples in the CTU. In this example, the above-mentioned reference sample averaging method can be applied to average the reconstructed samples adjacent to the CTU. The average of all these reference samples can be used to derive the chroma scale factor. Alternatively, the average of only a portion of these reference samples can be used. For example, only the M reference samples (e.g., M=4, 8, 16, 32, or 64) closest to the top-left position of the target block can be averaged.
[0143]
[0168] However, for a CTU on the bottom or right boundary of a picture, such as the gray CTU in Figure 19, not all samples of the CTU may be within the picture boundary. In this case, only neighboring reconstructed samples on the boundary of the CTU within the picture boundary (gray samples in Figure 19) can be used to derive the chroma scale factor. However, a variable number of samples in the average calculation requires undesirable division operations in hardware implementation. Therefore, embodiments of the present disclosure provide a method for padding picture boundary samples to a constant that is a power of two so that division operations in the average calculation can be avoided. For example, as shown in Figure 19, a padded sample outside the bottom boundary of the picture is generated from sample 1905, which is the closest sample to the padded sample among all samples on the bottom boundary of the picture. In addition to deriving chroma scale factors at the CTU level, chroma scale factors can also be derived on a fixed grid. Considering a virtual pipeline data unit (VPDU), which is defined as the data unit processed by the pipeline stages, the chroma scale factor can be derived at the VPDU level. In VVC Draft 5, a VPDU is defined as a 64x64 block on the luma sample grid. Therefore, embodiments of the present disclosure derive chroma scale factors at the granularity of 64x64 blocks. In VVC Draft 6, a VPDU is defined as an MxM block on the luma sample grid, where M is the smaller of the CTU size and 64. The CTU-level derivation method described above can also be used at the VPDU level.
[0144]
[0169] In some embodiments, in addition to deriving the chroma scale factor on a fixed grid smaller than the CTU, the factor is derived only once per CTU and used for all grid units (e.g., VPDUs) within the CTU. For example, derive the chroma scale factor on the first VPDU of the CTU and use that factor for all VPDUs within the CTU. It will be understood that the method at the VPDU level, which uses a limited number of adjacent samples during derivation (e.g., only the adjacent samples corresponding to the first VPDU within the CTU), is equivalent to the derivation at the CTU level.
[0145]
[0170] Average the sample values of the corresponding luma blocks to calculate avgY’ at the CTU level at the VPDU level or any other arbitrary fixed-size block unit level, and determine the segmentation index Y Idx and obtain the chroma scale factor inverse_chroma_scaling[Y Idx , alternatively, to avoid the dependency on luma in the case of dual-tree splitting, the chroma scale factor can also be explicitly signaled in the bitstream.
[0146]
[0171] The chroma scale index can be signaled at multiple levels. For example, as shown in Table 10 of FIG. 20 and Table 11 of FIG. 21, the chroma scale index can be signaled at the coding unit (CU) level together with the chroma prediction mode. To determine the chroma scale factor of the target chroma block, the syntax element lmcs_scaling_factor_idx (element 2002 in FIG. 20 and element 2102 in FIG. 21) is used. If there is no lmcs_scaling_factor_idx, the chroma scale factor of the target chroma block is inferred to be 1.0 with floating-point precision or equally (1<<CSCALE_FP_PREC) with fixed-point precision. The range of allowable values of lmcs_chroma_scaling_idx can be determined at the tile group level or the slice level, which will be described later.
[0147]
[0172] Depending on the possible values of lmcs_chroma_scaling_idx, its signaling cost may be too high, especially for small blocks. Therefore, in some embodiments of the present disclosure, the signaling conditions in Table 10 of FIG. 20 may additionally include a condition for block size. For example, this "lmcs_chroma_scaling_idx" syntax element (element 2002 of FIG. 20) is signaled only if the target block contains N chroma samples or less, or if the target block has a width greater than a given width W and / or a height greater than a given height H. For smaller blocks, if lmcs_chroma_scaling_idx is not signaled, the decoder can determine its chroma scale factor. As an example, the chroma scale factor can be set to 1.0 with floating-point precision. In some embodiments, a default lmcs_chroma_scaling_idx value can be added at the tile group header or slice header level (Table 1 of FIG. 10). Blocks that do not have lmcs_chroma_scaling_idx signaled (e.g., small blocks) can use this tile group / slice level default index to derive the chroma scale factor corresponding to the block. In some embodiments, the chroma scale factor of a small block can be inherited from its neighbors (e.g., top or left neighbors) that explicitly signal a scale factor.
[0148]
[0173] In addition to signaling this "lmcs_chroma_scaling_idx" syntax element at the CU level, this syntax element can also be signaled at the CTU level. However, when the maximum CTU size in VVC is 128x128, the chroma scaling by the "lmcs_chroma_scaling_idx" syntax element signaled at the CTU level may be too coarse. Therefore, in some embodiments of the present disclosure, this "lmcs_chroma_scaling_idx" syntax element may be signaled using a fixed granularity. For example, one lmcs_chroma_scaling_idx may be signaled per region of 16x16 samples (or a region of 64x64 samples in a VPDU) and applied to samples within the region of 16x16 samples (or a region of 64x64 samples).
[0149]
[0174] The range of lmcs_chroma_scaling_idx for a target tile group / slice depends on the number of chroma scale factor values allowed within the target tile group / slice. The range of lmcs_chroma_scaling_idx may be determined by existing methods within VVC that rely on a 64-entry chroma LUT, as discussed above. Alternatively, the range of lmcs_chroma_scaling_idx may also be determined using the chroma scale factor calculations discussed above.
[0150]
[0175] As an example, in the "quantization" method described above, the value of LOG2_n is set to 2 (i.e., "n" is set to 4), and the codeword assignment for each segment in the piecewise linear model for the target tile group / slice is set as follows: {0, 65, 66, 64, 67, 62, 62, 64, 64, 64, 67, 64, 64, 62, 61, 0}. Therefore, any codeword value between 64 and 67 can have the same scale factor value (e.g., 1.0 in decimal precision), and any codeword value between 60 and 63 can have the same scale factor value (e.g., 60 / 64 = 0.9375 in decimal precision), so there are only two possible scale factor values for the entire tile group. In the two end segments with no codewords assigned, the chroma scale factor can be set to 1.0 by default. Therefore, in this example, one bit is sufficient to signal lmcs_chroma_scaling_idx for a block in the target slice. The block may contain CU, CTU or fixed regions depending on the signaling level of the chroma scale factor.
[0151]
[0176] Besides using a piecewise linear model to derive possible chroma scale factor values, in some embodiments the encoder can signal a set of chroma scale factor values in the tile group / slice header, and then at the block level, the value of the chroma scale factor can be determined using this set and the value of lmcs_chroma_scaling_idx for that block.
[0152]
[0177] Alternatively, to reduce the signaling cost, the chroma scale factor can be predicted from a neighboring block. For example, a flag can be used to indicate that the chroma scale factor of the target block is equal to that of the neighboring block of the target block. The neighboring block can be an upper or left neighboring block. Thus, up to two bits can be signaled for the target block. For example, the first bit of the two bits can indicate whether the chroma scale factor of the target block is equal to that of the left neighbor of the target block, and the second bit can indicate whether the chroma scale factor of the target block is equal to that of the top neighbor of the target block. If the value of neither bit indicates that the chroma scale factor of the target block is equal to that of the top neighbor or the left neighbor, the lmcs_chroma_scaling_idx syntax can be signaled.
[0153]
[0178] Depending on the different possible values of "lmcs_chroma_scaling_idx", variable length codewords can be used for "code lmcs_chroma_scaling_idx" to reduce the average code length.
[0154]
[0179] Context-based adaptive binary arithmetic coding (CABAC) can be applied to code "lmcs_chroma_scaling_idx". The CABAC context associated with a target block may depend on the "lmcs_chroma_scaling_idx" of the target block's neighboring blocks. For example, the left neighboring block or the top neighboring block can be used to form the CABAC context. Regarding the binarization of "lmcs_chroma_scaling_idx", truncated rice binarization can be used to binarize "lmcs_chroma_scaling_idx".
[0155]
[0180] By signaling "lmcs_chroma_scaling_idx", the encoder can select lmcs_chroma_scaling_idx adaptively with respect to rate-distortion cost. Thus, rate-distortion optimization can be used to select lmcs_chroma_scaling_idx to improve coding efficiency, which may help offset the increased signaling cost.
[0156]
[0181] Embodiments of the present disclosure further provide a method for processing video content that involves signaling an LMCS piecewise linear model.
[0157]
[0182] The LMCS method in VVC Draft 4 uses a piecewise linear model with 16 partitions, but the number of unique values of SignaledCW[i] within a tile group / slice tends to be much less than 16. For example, some of the 16 partitions may use the default number of codewords, "OrgCW," and some of the 16 partitions may have the same number of codewords as each other. Therefore, when signaling an LMCS piecewise linear model, the number of unique codewords may be signaled in the form of "listUniqueCW[]," and an index of listUniqueCW[] may be sent for each partition of the LMCS piecewise linear model to select the codeword for the target partition.
[0158]
[0183] The modified syntax table is shown in Table 12 of FIG. 22, where the syntax elements 2202 and 2204 shown in italics have been revised in accordance with this embodiment.
[0159]
[0184] The semantics of the disclosed signaling method are as follows, with changes underlined: reshaper_model_min_bin_idx specifies the minimum bin (or partition) index used in the reshaper construction process. The value of reshaper_model_min_bin_idx shall be in the range 0 to MaxBinIdx. The value of MaxBinIdx shall be equal to 15. reshaper_model_delta_max_bin_idx specifies the maximum allowed bin (or partition) index MaxBinIdx minus the maximum bin index used in the reshaper construction process. The value of reshaper_model_max_bin_idx is set equal to MaxBinIdx-reshape_model_delta_max_bin_idx. reshaper_model_bin_delta_abs_cw_prec_minus1 plus 1 specifies the number of bits used to represent the syntax reshape_model_bin_delta_abs_CW[i]. reshaper_model_bin_num_unique_cw_minus1 plus 1 specifies the size of the codeword array listUniqueCW. reshaper_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value for the ith bin. reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshaper_model_bin_delta_abs_CW[i] as follows: If reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is a positive value. Otherwise (reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0) the corresponding variable RspDeltaCW[i] is negative. If reshape_model_bin_delta_sign_CW_flag[i] is missing, the corresponding variable RspDeltaCW[i] is inferred to be equal to 0. The variable RspDeltaCW[i] is derived as RspDeltaCW[i]=(1-2*reshape_model_bin_delta_sign_CW[i])*reshape_model_bin_delta_abs_CW[i]. The variable listUniqueCW[0] is set equal to OrgCW. The variables listUniqueCW[i] for i=1... reshaper_model_bin_num_unique_cw_minus1 are set equal to is derived as follows: - Set the variable OrgCW to (1< <BitDepth Y) / (MaxBinIdx+1). - listUniqueCW[i] =OrgCW+RspDeltaCW[i-1] reshaper_model_bin_cw_idx[i] specifies the index into the array listUniqueCW[] used to derive RspCW[i]. The value of reshaper_model_bin_cw_idx[i] shall be in the range 0 to (reshaper_model_bin_num_unique_cw_minus1+1). RspCW[i] is derived as follows: - if reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx holds, RspCW[i]= listUniqueCW[reshaper_model_bin_cw_idx[i]]. - Otherwise, RspCW[i]=0.
[0160]
[0185] BitDepth Y If the value of is equal to 10, the value of RspCW[i] can be in the range of 32 to 2*OrgCW-1.
[0161]
[0186] SUMMARY Embodiments of the present disclosure provide a method for processing video content with conditional chroma scaling at the block level.
[0162]
[0187] As shown in Table 1 of FIG. 6, whether chroma scaling is applied can be determined by the tile_group_reshaper_chroma_residual_scale_flag signaled at the tile group / slice level. However, it may be beneficial to determine whether to apply chroma scaling at the block level. For example, in some embodiments, a CU-level flag can be signaled to indicate whether chroma scaling is applied to the target block. The presence of the CU-level flag can be conditioned based on the tile group-level flag "tile_group_reshaper_chroma_residual_scale_flag." In other words, the CU-level flag can be signaled only if chroma scaling is allowed at the tile group / slice level. The CU-level flag may allow the encoder to decide whether to use chroma scaling based on whether chroma scaling is beneficial for the target block, but it may also incur signaling overhead.
[0163]
[0188] Consistent with the disclosed embodiments, to avoid the above signaling overhead, whether chroma scaling is applied to a block can be conditioned based on the prediction mode of the target block. For example, if the target block is inter-predicted, the prediction signal tends to be good, especially if its reference picture is close in terms of temporal distance. In this case, chroma scaling can be bypassed because the residual is expected to be very small. For example, pictures in higher temporal levels tend to have reference pictures that are close in terms of temporal distance. For a block, chroma scaling can be disabled in these pictures that use nearby reference pictures. To determine whether this condition is met, the difference in picture order count (POC) between the target picture and the reference picture of the target block can be used.
[0164]
[0189] In some embodiments, chroma scaling may be disabled for all inter-coded blocks, in some embodiments, chroma scaling may be disabled for all intra-coded blocks, and in some embodiments, chroma scaling may be disabled for combined intra-inter prediction (CIIP) modes defined in the VVC standard.
[0165]
[0190] In the VVC standard, the CU syntax structure "coding_unit()" may include a syntax element "cu_cbf" to indicate whether there are any non-zero residual coefficients in the target CU. At the TU level, the TU syntax structure "transform_unit()" may include syntax elements "tu_cbf_cb" and "tu_cbf_cr" to indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the target TU. Based on these flags, the chroma scaling process can be conditioned. As described above, if there are no non-zero residual coefficients, the averaging of the corresponding luma chroma scaling process can be invoked, and then the chroma scaling process can be bypassed. The present disclosure provides a method for bypassing the luma averaging process.
[0166]
[0191] Embodiments of the present disclosure provide a method for processing video content that involves the derivation of CCLM parameters.
[0167]
[0192] As mentioned above, in VVC5, CCLM parameters for predicting a target chroma block are derived using luma and chroma reconstructed samples from neighboring blocks. In the case of dual trees, the luma block division and the chroma block division may not be aligned. In other words, to derive CCLM parameters for one NxM chroma block, multiple neighboring luma blocks or luma blocks with a size larger than 2Nx2M (for color format 4:2:0) are reconstructed, which may introduce latency.
[0168]
[0193] To reduce latency, CCLM parameters are derived at the CTU / VPDU level, for example. To derive the CCLM parameters, reconstructed luma and chroma samples from adjacent CTUs / VPDUs can be used. The derived parameters can be applied to all blocks within a CTU / VPDU. For example, the parameters can be derived using the formulas described in Cross-Component Linear Model Prediction, where X a and Y a are the luma and chroma values of the luma sample position with the maximum luma sample value of the CTU / VPDU adjacent luma samples, respectively. b and Y b and represent the luma and chroma values of the luma sample position with the smallest luma sample of the CTU / VPDU adjacent luma samples, respectively. For those skilled in the art, any other derivation process can be used in combination with the concept of derivation of CTU / VPDU-level parameters proposed in this specification.
[0169]
[0194] In addition to deriving CCLM parameters at the CTU / VPDU level, such derivation process can be performed on a fixed luma grid. In VVC Draft 5, if dual-tree partitioning is used, separate luma and chroma partitioning can start from a 64x64 luma grid. In other words, the partitioning of 128x128 CTUs into 64x64 CUs can be performed jointly, rather than separately, for luma and chroma. Thus, as another example, CCLM parameters can be derived on a 64x64 luma grid. Adjacent reconstructed luma and chroma samples of a 64x64 grid unit can be used to derive CCLM parameters for all chroma blocks within a 64x64 grid unit. Compared to a CTU-level derivation, which can reach 128x128 in luma samples, a 64x64 unit-level derivation can be more accurate and still does not have the pipeline latency issues present in the current VVC Draft 5. In addition to this example, the derivation of CCLM parameters can be further simplified by skipping the derivation for some grids. For example, CCLM parameters are derived only on the first 64x64 block in a CTU, and the derivation for subsequent 64x64 blocks in the same CTU is skipped. The parameters derived based on the first 64x64 block can be used for all blocks in the CTU.
[0170]
[0195] FIG. 23 shows a flowchart of an example method 2300 for processing video content according to some embodiments of the present disclosure. In some embodiments, method 2300 may be performed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec may be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding or converting a video sequence into another code. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence may be a surveillance video sequence that may be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence may include multiple pictures. The device may perform method 2300 at the picture level. For example, the device may process pictures one by one in method 2300. In another example, the device may process multiple pictures at a time in method 2300. Method 2300 may include the following steps.
[0171]
[0196] In step 2302, data representing a first block and a second block in a picture may be received. The plurality of blocks may include the first block and the second block. In some embodiments, the first block may be a target chroma block (e.g., chroma block 1600 in FIG. 16A ), and the second block may be a coded tree block (CTB), a transform unit (TU), or a virtual pipeline data unit (VPDU). A virtual pipeline data unit is a non-overlapping unit in a picture that has a size less than or equal to the size of the coded tree unit of the picture. For example, if the size of a CTU is 128x128 pixels, the VPDU may have a size smaller than the size of the CTU, and the size of the VPDU (e.g., 64x64 pixels) may be proportional to the buffer size in most pipeline stages of hardware (e.g., a hardware decoder).
[0172]
[0197] In some embodiments, the coding tree block may be a luma block corresponding to a target chroma block (e.g., luma block 1610 of FIG. 16B). Thus, the data may include a plurality of chroma samples associated with a first block and a plurality of luma samples associated with a second block. The plurality of chroma samples associated with the first block includes a plurality of chroma residual samples within the first block.
[0173]
[0198] In step 2304, an average value of a plurality of luma samples associated with the second block may be determined. The plurality of luma samples may include the samples described with respect to Figures 18-19. As an example, as shown in Figure 19, the plurality of luma samples may include a plurality of reconstructed luma samples (e.g., shaded samples 1905 and padded samples 1903) on a left boundary 1901 of the second block (e.g., 1902) or on a top boundary of the second block. It will be understood that the plurality of reconstructed luma samples may belong to an adjacent reconstructed luma block (e.g., 1904).
[0174]
[0199] The method 2300 may further include determining whether a first luma sample of the plurality of luma samples associated with the second block is outside a picture boundary, and, in response to determining that the first luma sample is outside the picture boundary, setting a value of the first luma sample to a value of a second luma sample of the plurality of luma samples that is within the picture boundary. The picture boundary may include one of a right picture boundary and a bottom picture boundary. For example, it may be determined that the padded sample 1903 is outside the bottom picture boundary, and thus the value of the padded sample 1903 is set to be the value of the shaded sample 1905, which is the closest sample to the padded sample 1903 of all samples on the bottom picture boundary.
[0175]
[0200] It will be appreciated that if the second block (e.g., 1902) crosses a picture boundary, a padded sample (e.g., padded sample 1903) can be created such that the number of luma samples can be a constant, typically a power of two, to avoid division operations.
[0176]
[0201] In step 2306, a chroma scale factor for the first block may be determined based on the average value. As discussed above with respect to Figure 18, in intra prediction, decoded samples in neighboring blocks of the same picture may be used as reference samples to generate a predictive block. For example, the average value of the samples in the neighboring blocks may be used as the luma average value for determining the chroma scale factor of a target block (e.g., the first block in this example), and the chroma scale factor for the first block may be determined using the luma average value of the second block.
[0177]
[0202] In step 2308, the chroma scale factors can be used to process the chroma samples associated with the first block. As discussed above with respect to Figure 5, the chroma scale factors can be applied at the decoder side to construct a chroma scale factor LUT at the tile group level and to the reconstructed chroma residual of the target block. Similarly, the chroma scale factors can also be applied at the encoder side.
[0178]
[0203] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which may be executed by an apparatus (such as the disclosed encoders and decoders) to perform the above-described methods. Common non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, flash EPROMs or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. An apparatus may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0179]
[0204] Embodiments may be further described using the following clauses: 1. A computer-implemented method for processing video content, comprising: receiving data representing a first block and a second block in a picture, the data including a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determining an average value of a plurality of luma samples associated with the second block; determining a chroma scale factor for the first block based on the average value; and Processing a plurality of chroma samples associated with the first block using a chroma scale factor. A computer-implemented method comprising: 2. The method of clause 1, wherein the plurality of luma samples associated with the second block includes a plurality of reconstructed luma samples on a left boundary of the second block or on a top boundary of the second block. 3. Determining whether a first luma sample of a plurality of luma samples associated with a second block is outside a picture boundary; and in response to determining that the first luma sample is outside the picture boundary, setting a value of the first luma sample to a value of a second luma sample of the plurality of luma samples that is within the picture boundary. 3. The method of clause 2, further comprising: 4. Determining whether a first luma sample of a plurality of luma samples associated with a second block is outside a picture boundary; and In response to determining that the first luma sample is outside the picture boundary, setting a value of the first luma sample to a value of a second luma sample of the plurality of luma samples that is on the picture boundary. 4. The method of clause 3, further comprising: 5. The method according to clause 4, wherein the picture boundary is one of a right boundary of the picture and a bottom boundary of the picture. 6. The method of any one of clauses 1 to 5, wherein the second block is a coding tree block, a transform unit, or a virtual pipeline data unit, and the size of the virtual pipeline data unit is less than or equal to the size of the coding tree unit of the picture. 7. The method of clause 6, wherein the virtual pipeline data unit is a non-overlapping unit within a picture. 8. The method of any one of clauses 1-7, wherein the plurality of chroma samples associated with the first block comprises a plurality of chroma residual samples within the first block. 9. The method of any one of clauses 1-8, wherein the first block is a target chroma block and the second block is a luma block corresponding to the target chroma block. 10. A system for processing video content, comprising: a memory for storing a set of instructions; and at least one processor, the at least one processor providing the system with: receiving data representing a first block and a second block in a picture, the data including a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determining an average value of a plurality of luma samples associated with the second block; determining a chroma scale factor for the first block based on the average value; and Processing a plurality of chroma samples associated with the first block using a chroma scale factor. configured to execute a set of instructions to cause system. 11. The system of clause 10, wherein the plurality of luma samples associated with the second block includes a plurality of reconstructed luma samples on a left boundary of the second block or on a top boundary of the second block. 12. At least one processor in the system determining whether a first luma sample of the plurality of luma samples associated with the second block is outside a picture boundary; and in response to determining that the first luma sample is outside the picture boundary, setting a value of the first luma sample to a value of a second luma sample of the plurality of luma samples that is within the picture boundary. 12. The system of claim 11, configured to execute a set of instructions to further cause: 13. The system of clause 12, wherein the second luma sample is on a picture boundary. 14. The system of clause 13, wherein the picture boundary is one of a right picture boundary and a bottom picture boundary. 15. A system described in any one of clauses 10 to 14, wherein the second block is a coding tree block, a transform unit, or a virtual pipeline data unit, and the size of the virtual pipeline data unit is less than or equal to the size of the coding tree unit of the picture. 16. The system of clause 15, wherein the virtual pipeline data unit is a non-overlapping unit within a picture. 17. The system of any one of clauses 10-16, wherein the plurality of chroma samples associated with the first block comprises a plurality of chroma residual samples within the first block. 18. The system of any one of clauses 10-17, wherein the first block is a target chroma block and the second block is a luma block corresponding to the target chroma block. 19. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising: receiving data representing a first block and a second block in a picture, the data including a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determining an average value of a plurality of luma samples associated with the second block; determining a chroma scale factor for the first block based on the average value; and Processing a plurality of chroma samples associated with the first block using a chroma scale factor. 1. A non-transitory computer-readable medium comprising: 20. The non-transitory computer-readable medium of clause 19, wherein the plurality of luma samples associated with the second block includes a plurality of reconstructed luma samples on a left boundary of the second block or on a top boundary of the second block.
[0180]
[0205] It should be noted that relational terms such as "first" and "second" herein are used merely to distinguish one entity or operation from another and do not require or imply any actual relationship or order between those entities or operations. Furthermore, terms such as "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and are open-ended in that the items following any one of these terms are not intended to be an exhaustive list of such items or to be limited only to the items they list.
[0181]
[0206] As used herein, unless otherwise specified, the word "or" includes all possible combinations unless impracticable. For example, if a database is stated to include A or B, the database can include A or B, or A and B, unless otherwise specified or impracticable. As a second example, if a database is stated to include A, B, or C, the database can include A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C, unless otherwise specified or impracticable.
[0182]
[0207] It will be understood that the above-described embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the disclosed methods. The computational units and other functional units described in this disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.
[0183]
[0208] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications to the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims. The order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to the particular order of steps. As such, one skilled in the art will recognize that steps can be performed in different orders while implementing the same method.
[0184]
[0209] Although illustrative embodiments have been disclosed in the drawings and herein, many variations and modifications to those embodiments may be made. Accordingly, although specific terms have been employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A method for decoding a bitstream associated with a video sequence, comprising: reconstructing chroma-coded blocks based on syntax elements encoded in the bitstream; determining an inverse chroma scale factor associated with the reconstructed chroma-coded block based on the encoded syntax elements in the bitstream; and performing inverse chroma residual scaling of the reconstructed chroma-coded block using the inverse chroma scale factor. wherein the inverse chroma scale factor is determined based on a bit depth of the reconstructed chroma-coded block, a variable signaled in the bitstream, and a number of bits used in a fixed-point approximation.
2. The method of claim 1 , wherein the reconstructed chroma-coded blocks belong to a tile group, the method further comprising applying the inverse chroma scale factor to all chroma-coded blocks in the tile group.
3. The method of claim 1 , wherein the reconstructed chroma-coded blocks belong to a slice, the method further comprising applying the inverse chroma scale factor to all chroma-coded blocks in the slice.
4. The method of claim 1 , wherein the variable signaled in the bitstream relates to a partition index of the reconstructed chroma-coded block.
5. The method of claim 1 , wherein the variable signaled in the bitstream indicates the number of codewords used in a piecewise linear model.
6. 1. A method for encoding a video sequence into a bitstream, comprising: encoding variables for chroma scaling of target chroma-coded blocks into the bitstream; determining a chroma scale factor based on the variable, a bit-width of the target chroma-coded block, and a number of bits used in a fixed-point approximation; performing chroma residual scaling of the target chroma-coded block using the chroma scale factor; and encoding one or more syntax elements signaling the scaled target chroma-coded blocks into the bitstream. A method comprising:
7. The method of claim 6 , wherein the target chroma-coded block belongs to a tile group, the method further comprising applying the chroma scale factor to all chroma-coded blocks in the tile group.
8. The method of claim 6 , wherein the target chroma-coded block belongs to a slice, the method further comprising applying the chroma scale factor to all chroma-coded blocks in the slice.
9. The method of claim 6 , wherein the variable signaled in the bitstream relates to a partition index of the target chroma-coded block.
10. The method of claim 6 , wherein the variable indicates the number of codewords used in a piecewise linear model.
Citation Information
Patent Citations
Interactions between in-loop reshaping and inter coding tools
WO2020156526A1
Signaling of in-loop reshaping information using parameter sets
WO2020156529A1
Methods for cross component dependency reduction
WO2020216246A1