Method and system for processing luminance and chrominance signals

Through the chromaticity scaling brightness mapping and cross-component linear model, the problem of low encoding efficiency of high dynamic range video signals is solved, and more efficient video encoding and transmission is achieved.

CN114375582BActive Publication Date: 2025-08-08ALIBABA GROUP HOLDING LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080043294.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-24
Filing Date
2020-05-29
Publication Date
2025-08-08
Estimated Expiration
2040-05-29

AI Technical Summary

Technical Problem

When existing video encoding technologies deal with high dynamic range video signals, their encoding efficiency is low and it is difficult to effectively utilize the bit depth range, resulting in high storage and transmission bandwidth requirements.

Method used

The chromaticity mapping method of chromaticity scaling is used to determine the chromaticity scaling factor and process the chromaticity samples of the image block to improve encoding efficiency. Combining the in-loop brightness mapping and cross component linear model, the video encoding process is optimized.

Benefits of technology

Improves video encoding efficiency, reduces the bandwidth required for storage and transmission, while maintaining or improving video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114375582B_ABST
    Figure CN114375582B_ABST
Patent Text Reader

Abstract

The present disclosure provides a system and method for processing video content. The method may include: receiving data representing a first block and a second block in an image, the data including a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determining an average value of the plurality of luma samples associated with the second block; determining a chroma scaling factor for the first block based on the average value; and processing the plurality of chroma samples associated with the first block using the chroma scaling factor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This disclosure claims priority to U.S. Provisional Application No. 62 / 865,815, filed on June 24, 2019, which is incorporated herein by reference in its entirety. Technical Field

[0003] The present disclosure relates generally to video processing and, more particularly, to methods and systems for performing luma mapping with chroma scaling. Background Art

[0004] A video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, the video can be compressed before storage or transmission and decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are many video coding formats that use standardized video coding techniques. The most common ones are based on prediction, transform, quantization, entropy coding, and loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard and the Versatile Video Coding (VVC / H.266) standard AVS standard, specify specific video coding formats and are developed by standardization organizations. As more and more video standards adopt advanced video coding technologies, the coding efficiency of new video coding standards is also getting higher and higher. Summary of the Invention

[0005] Embodiments of the present disclosure provide methods and systems for performing in-loop luminance mapping and cross-component linear models with chroma scaling.

[0006] In one exemplary embodiment, the method includes receiving data representing a first block and a second block in an image, the data including a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block, determining an average of the plurality of luma samples associated with the second block; determining a chroma scaling factor for the first block based on the average; and processing the plurality of chroma samples associated with the first block using the chroma scaling factor.

[0007] In some embodiments, the system includes: a memory for storing a set of instructions; and at least one processor configured to execute the set of instructions to cause the system to: receive data representing a first block and a second block in an image, the data including a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determine an average of the plurality of luma samples associated with the second block; determine a chroma scaling factor for the first block based on the average; and process the plurality of chroma samples associated with the first block using the chroma scaling factor. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and accompanying drawings. The various features shown in the drawings are not drawn to scale.

[0009] Figure 1 The structure of an example video sequence according to some embodiments of the present disclosure is shown.

[0010] Figure 2A A schematic diagram illustrating an example encoding process according to some embodiments of the present disclosure is shown.

[0011] Figure 2B A schematic diagram illustrating another example encoding process according to some embodiments of the present disclosure is shown.

[0012] Figure 3A A schematic diagram illustrating an example decoding process according to some embodiments of the present disclosure is shown.

[0013] Figure 3B A schematic diagram illustrating another example decoding process according to some embodiments of the present disclosure is shown.

[0014] Figure 4 A block diagram illustrating an example apparatus for encoding or decoding video according to some embodiments of the present disclosure is shown.

[0015] Figure 5 A schematic diagram illustrating an exemplary Luma Mapping with Chroma Scaling (LMCS) process according to some embodiments of the present disclosure is shown.

[0016] Figure 6 is a tile group level syntax table for LMCS according to some embodiments of the present disclosure.

[0017] Figure 7 is a syntax table for LMCS according to some embodiments of the present disclosure.

[0018] Figure 8 is a slice-level syntax table for LMCS according to some embodiments of the present disclosure.

[0019] Figure 9 is a syntax table of the LMCS piecewise linear model according to some embodiments of the present disclosure.

[0020] Figure 10 An example of sample positions for deriving a and b according to some embodiments of the present disclosure is shown.

[0021] Figure 11 is a table for deriving chroma prediction mode from luma mode when CCLM is enabled according to some embodiments of the present disclosure.

[0022] Figure 12is a syntax structure of an exemplary coding tree unit according to some embodiments of the present disclosure.

[0023] Figure 13 is an exemplary dual-tree partitioned syntax structure according to some embodiments of the present disclosure.

[0024] Figure 14 is a syntax structure of an exemplary coding tree unit according to some embodiments of the present disclosure.

[0025] Figure 15 is an exemplary dual-tree partitioned syntax structure according to some embodiments of the present disclosure.

[0026] Figure 16A An exemplary chroma tree partitioning according to some embodiments of the present disclosure is shown.

[0027] Figure 16B An exemplary luma tree partitioning according to some embodiments of the present disclosure is shown.

[0028] Figure 17 A simplification of an exemplary averaging operation according to some embodiments of the present disclosure is shown.

[0029] Figure 18 An example of samples used in an average calculation to derive a chroma scaling factor according to some embodiments of the present disclosure is shown.

[0030] Figure 19 An example of chroma scaling factor derivation for blocks at the right or bottom border of an image according to some embodiments of the present disclosure is shown.

[0031] Figure 20 is an exemplary coding tree unit syntax structure according to some embodiments of the present disclosure.

[0032] Figure 21 is another exemplary coding tree unit syntax structure according to some embodiments of the present disclosure.

[0033] Figure 22 is an exemplary modified signal of the LMCS piecewise linear model at the slice level according to some embodiments of the present disclosure.

[0034] Figure 23 is a flowchart of an exemplary method for processing video content according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0035] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings and, unless otherwise indicated, the same numbers in different figures represent the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with the aspects related to the present invention as described in the appended claims. Unless otherwise specifically indicated, the term "or" includes all possible combinations unless the combination is not feasible. For example, if a component is stated to include A or B, then, unless otherwise explicitly indicated or not feasible, the component may include A, or B, or A and B. As a second example, if a component is stated to include A, B, or C, then, unless otherwise explicitly indicated or not feasible, the component may include A, or B, or C, or A and B, or A and C, or B and, or A and B and C.

[0036] A video is a set of static images (or "frames") that store visual information in a temporal sequence. Video capture devices (such as cameras) can be used to capture and store these images in a temporal sequence. Video playback devices (such as televisions, computers, smartphones, tablets, video players, or any end-user device with a display) can be used to display these images in a temporal sequence. Furthermore, in some applications, video capture devices can transmit captured video in real time to video playback devices (such as computers with monitors), for purposes such as surveillance, conferencing, and live broadcasts.

[0037] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a general-purpose computer processor) or dedicated hardware. The compression module is generally referred to as an "encoder," and the decompression module is generally referred to as a "decoder." Encoders and decoders can be collectively referred to as "codecs." Encoders and decoders can be implemented in any of a variety of suitable hardware, software, or combinations of the two. For example, hardware implementations of encoders and decoders can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, a codec can decompress video using a first coding standard and recompress the decompressed video using a second coding standard. In this case, the codec can be referred to as a "transcoder."

[0038] Video encoding processes identify and retain useful information that can be used to reconstruct an image, while ignoring less important information for reconstruction. If the ignored, less important information cannot be fully reconstructed, the encoding process is called "lossy." Otherwise, it is called "lossless." Most encoding processes are lossy, a trade-off made to reduce required storage space and transmission bandwidth.

[0039] The useful information of the image being coded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously coded and reconstructed image). Such changes can include changes in pixel position, brightness, or color, with position changes being of greatest interest. A change in the position of a group of pixels representing an object can reflect the object's motion between the reference image and the current image.

[0040] A picture that is encoded without reference to another picture (i.e., it is its own reference picture) is called an "I picture." A picture that is encoded using a previous picture as a reference picture is called a "P picture." A picture that is encoded using both a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B picture."

[0041] As mentioned above, one of the goals of developing new video coding technologies is to improve coding efficiency, that is, to use less coded data to represent the same picture quality. The present disclosure provides a method and system for performing luminance mapping with chroma scaling. Luma mapping is the process of mapping luminance samples for use in the loop filter, while chroma scaling is the process of scaling chroma residual values that depends on luminance. To obtain the same subjective quality as HEVC / H.265 using half the bandwidth, LMCS mainly includes two parts: 1) the process of mapping input luminance code values to a set of new code values for use within the encoding loop; and 2) the process of scaling chroma residual values that depends on luminance. The luminance mapping process improves the coding efficiency of standard and high dynamic range video signals by better utilizing the range of luminance code values allowed by the specified bit depth.

[0042] Figure 1 The structure of an example video sequence 100 used in video encoding according to some embodiments of the present disclosure is illustrated. Video sequence 100 can be live video or video that has been captured and archived. Video 100 can be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). Video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored on a storage device), or a video providing interface (e.g., a video broadcast transceiver) that receives video from a video content provider.

[0043] like Figure 1 As shown, video sequence 100 may include a series of images arranged in time along a time axis, including images 102, 104, 106, and 108. Images 102-106 are consecutive, and there may be more images between images 106 and 108. Figure 1 In FIG, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108, as indicated by the arrow. In some embodiments, the reference picture of a picture (e.g., picture 104) may be a picture that is not immediately before or after the picture. For example, the reference picture of picture 104 may be a picture that is before picture 102. It should be noted that the reference pictures of pictures 102-106 are only examples, and the present invention is not limited to such pictures. Figure 1 An embodiment of a reference image is shown in .

[0044] Typically, video codecs do not encode or decode an entire image at once due to the computational complexity of such tasks. Instead, they may divide the image into elementary segments and encode or decode the image segment by segment. Such elementary segments are referred to in this disclosure as Basic Processing Units ("BPUs"). For example, Figure 1 Structure 110 in shows an example structure of an image (e.g., any of images 102-108) of video sequence 100. In structure 110, the image is divided into 4×4 basic processing units, whose boundaries are represented by dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit in an image can have variable sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of arbitrary shape and size. The size and shape of the basic processing unit for an image can be selected based on a balance between coding efficiency and the level of detail to be retained in the basic processing unit.

[0045] A basic processing unit may be a logical unit that may include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit of a color image may include a luma component (Y) representing non-color luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components may have the same size as the basic processing unit. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated for each of its luma and chroma components.

[0046] Video encoding has multiple stages of operation, examples of which are given in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to be processed, so it can be further divided into segments, and these further divided segments are referred to as "basic processing sub-units" in the present invention. In some embodiments, the basic processing sub-unit may be referred to as a "block" in some video coding standards (e.g., MPEG series, H.261, H.263 or H.264 / AVC), or as a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-unit may have the same size as the basic processing unit or a smaller size. Similar to the basic processing unit, the basic processing sub-unit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing sub-unit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to a deeper level according to processing needs. It is also important to note that different stages can use different schemes to divide the basic processing units.

[0047] For example, in the mode decision phase (an example of which will be Figure 2B ), the encoder can decide what prediction mode (e.g., intra-image prediction or inter-image prediction) to use for a basic processing unit, but the basic processing unit may be too large to make this decision. The encoder can split the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and decide the prediction type to use for each individual basic processing sub-unit.

[0048] In another example, in the prediction phase (an example of which will be Figure 2A), the encoder can perform prediction operations at the level of basic processing sub-units (e.g., CUs). However, in some cases, the basic processing sub-units may still be too large to process. The encoder can further split the basic processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which prediction operations can be performed.

[0049] In another example, during the transformation phase (an example of which will be Figure 2A (Detailed in H.265 / HEVC or H.266 / VVC), the encoder can perform transform operations for the residual basic processing sub-unit (e.g., CU), however, in some cases, the basic processing sub-unit may still be too large to process. The encoder can further split the basic processing sub-unit into smaller parts (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which level the transform operations can be performed. It should be noted that the division schemes for the same basic processing sub-unit in the prediction stage and the transform stage can be different. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.

[0050] exist Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, whose boundaries are represented by dotted lines. Different basic processing units of the same image can be divided into basic processing sub-units in different schemes.

[0051] In some embodiments, to provide parallel processing and fault resilience for video encoding and decoding, an image can be divided into multiple regions for processing, so that the encoding or decoding process for one region of the image does not depend on information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving coding efficiency. Furthermore, when data in one region is damaged during processing or network transmission is lost, the codec can correctly encode or decode other regions of the same image without relying on the damaged or lost data, thereby providing fault resilience. In some video coding standards, an image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different images in the video sequence 100 can have different partitioning schemes for dividing the image into multiple regions.

[0052] For example, in Figure 1In FIG, the structure 110 is divided into three regions 114, 116 and 118, whose boundaries are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that Figure 1 The basic processing units, basic processing sub-units, and regions of the structure 110 are merely examples, and the present disclosure is not limited to the embodiments thereof.

[0053] Figure 2A FIG2 illustrates a schematic diagram of an example encoding process 200A according to some embodiments of the present disclosure. According to process 200A, an encoder may encode a video sequence 202 into a video bitstream 228. Similar to Figure 1 The video sequence 100 in Figure 1 As shown, the video sequence 202 may include a set of images (referred to as "original images") arranged in time sequence. Figure 1 In the structure 110 in FIG. 1 , each original image of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can perform process 200A at the basic processing unit level for each original image of the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, where the encoder can encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for each region (e.g., regions 114-118) of the original image of the video sequence 202.

[0054] exist Figure 2A200A, an encoder may provide a basic processing unit (BPU) of an original image of a video sequence 202 (referred to as an "original BPU") to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a "forward path." During process 200A, after the quantization stage 214, the encoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the next iteration of process 200A in the prediction stage 204. Components 218, 220, 222, and 224 of process 200A can be referred to as a "reconstruction path." The reconstruction path can be used to ensure that the encoder and decoder use the same reference data for prediction.

[0055] The encoder can iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original image (in the reconstruction path). After encoding all the original BPUs of the original image, the encoder can continue to encode the next image in the video sequence 202.

[0056] Referring to process 200A, an encoder may receive a video sequence generated by a video capture device (e.g., a camera) 202. The term "receive" as used herein may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or in any way inputting data.

[0057] In the prediction stage 204, in the current iteration, the encoder can receive the original BPU and the prediction reference 224 and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 can be generated from the reconstruction path in the previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting the prediction data 206 that can be used to reconstruct the original BPU into the predicted BPU 208 from the prediction data 206 and the prediction reference 224.

[0058] Ideally, the predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 typically differs slightly from the original BPU. To account for this difference, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210. For example, the encoder can subtract the value (e.g., grayscale value or RGB value) of the corresponding pixel of the predicted BPU 208 from the corresponding pixel value of the original BPU. Each pixel of the residual BPU 210 can have a residual value, which can be the result of the subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significantly reducing quality. Therefore, the original BPU is compressed.

[0059] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", each of which is associated with a "transform coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency-varying (e.g., frequency-varying) component of the residual BPU 210. No basis pattern can be copied from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the variation of the residual BPU 210 into the frequency domain. Such a decomposition is analogous to a discrete Fourier transform of a function, where the basis patterns are analogous to the basis functions of the discrete Fourier transform (e.g., trigonometric functions) and the transform coefficients are analogous to the coefficients associated with the basis functions.

[0060] Different transform algorithms can use different base patterns. Various transform algorithms can be used in the transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform in the transform stage 212 is reversible. That is, the encoder can restore the residual BPU 210 by performing the inverse operation of the transform (referred to as an "inverse transform"). For example, to restore the pixels of the residual BPU 210, the inverse transform may be to multiply the values of the corresponding pixels of the base pattern by their associated coefficients and add the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transform algorithm (and therefore the same base pattern). Therefore, the encoder can only record the transform coefficients, and the decoder can reconstruct the residual BPU 210 based on these coefficients without receiving the base pattern from the encoder. Compared to the residual BPU 210, the transform coefficients may have fewer bits, but they can be used to reconstruct the residual BPU 210 without significantly reducing quality. As a result, the residual BPU 210 is further compressed.

[0061] The encoder can further compress the transform coefficients during the quantization stage 214. During the transform process, different basis patterns can represent different frequencies of change (e.g., the frequency of changes in brightness). Because the human eye is generally better at detecting low-frequency changes, the encoder can ignore information about high-frequency changes without significantly degrading decoding quality. For example, during the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to the nearest integer. This operation can convert some transform coefficients of high-frequency basis patterns to zero, while transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, thereby further compressing the quantized transform coefficients. The quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed as transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0062] Because the encoder ignores the remainder of this division during rounding operations, the quantization stage 214 can be lossy. Generally, the quantization stage 214 can contribute the most information loss in process 200A. The greater the information loss, the fewer bits may be required to quantize the transform coefficients 216. To achieve different degrees of information loss, the encoder can use different values for the quantization parameter or any other parameter of the quantization process.

[0063] In the binary encoding stage 226, the encoder can encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder can use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 can be further packaged for network transmission.

[0064] Referring to the reconstruction path of process 200A, at the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. At the inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0065] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, the stages of process 200A can be performed by the encoder in a different order. In some embodiments, one or more stages of process 200A can be combined into a single stage. In some embodiments, a single stage of process 200A can be divided into multiple stages. For example, transform stage 212 and quantization stage 214 can be combined into a single stage. In some embodiments, process 200A can include additional stages. In some embodiments, process 200A can be omitted. Figure 2A one or more stages in a process.

[0066] Figure 2B A schematic diagram illustrating another example encoding process 200B according to some embodiments of the present disclosure is shown. Process 200B may be modified from process 200A. For example, process 200B may be used by encoders conforming to hybrid video coding standards (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and separates the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0067] In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-frame prediction") can use pixels from one or more coded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more coded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded images. Temporal prediction can reduce the temporal redundancy inherent in the image.

[0068] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction stage 2042 and a temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-frame prediction. For an original BPU of the picture being encoded, the prediction reference 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same picture. The encoder may generate a predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, and the like. In some embodiments, the encoder may perform extrapolation at the pixel level, for example, by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The neighboring BPUs used for extrapolation may be positioned relative to the original BPU in various directions, such as vertically (e.g., on top of the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below and to the left, below and to the right, above and to the left, or above and to the right of the original BPU), or in any direction defined in the video coding standard being used. For intra prediction, the prediction data 206 may include, for example, the location (eg, coordinates) of the used neighboring BPUs, the size of the used neighboring BPUs, extrapolated parameters, the direction of the used neighboring BPUs relative to the original BPU, and the like.

[0069] As another example, during the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that were encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference image may be a BPU encoded and reconstructed. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When generating all reconstructed BPUs for the same image, the encoder may generate the reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching region within a range of the reference image (referred to as a "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a location in the reference image with the same coordinates as the original BPU in the current image and may extend outward by a predetermined distance. When the encoder identifies an area in the search window that is similar to the original BPU (e.g., using a pixel recursive algorithm, a block matching algorithm, etc.), the encoder may determine such an area as a matching region. The matching region may have different dimensions than the original BPU (e.g., smaller, equal, larger, or in a different shape). Because the reference image and the current image are temporally separated on the timeline (e.g., as Figure 1As shown in the figure, the matching region can be considered to "move" to the location of the original BPU over time. The encoder can record the direction and distance of this movement as a "motion vector". When using multiple reference images (e.g., Figure 1 The encoder can search for matching areas and determine the motion vector associated with each reference image. In some embodiments, the encoder can assign weights to the pixel values of the matching areas of each matching reference image.

[0070] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.

[0071] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move a matching area of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., Figure 1 The encoder may move the matching regions of the reference image based on their respective motion vectors and the average pixel value of the matching regions. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of the respective matching reference images, the encoder may add the weighted sum of the pixel values of the moved matching regions.

[0072] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference pictures in the same temporal direction as the current picture. For example, Figure 1 104 is a unidirectional inter-frame prediction image, where the reference image (i.e., image 102) precedes image 104. Bidirectional inter-frame prediction can use one or more reference images in both temporal directions relative to the current image. For example, Figure 1 Picture 106 in is a bi-directionally inter-predicted picture, where the reference pictures (ie, pictures 104 and 108 ) are relative to picture 104 in both temporal directions.

[0073] Still referring to the forward path of process 200B, after the spatial prediction 2042 and temporal prediction stages 2044, in the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction and inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform rate-distortion optimization techniques, in which the encoder can select a prediction mode based on the bit rate of candidate prediction modes and the distortion of the reconstructed reference image under the candidate prediction modes to minimize the value of a cost function. Based on the selected prediction mode, the encoder can generate a corresponding prediction BPU 208 and prediction data 206.

[0074] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU in the current picture that has been encoded and reconstructed), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If inter-prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current picture in which all BPUs have been encoded and reconstructed), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortion introduced by inter-prediction (e.g., blocking artifacts). The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference picture can be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., as an inter-prediction reference picture for a future picture in the video sequence 202). The encoder may store one or more reference pictures in a buffer 234 for use in a temporal prediction stage 2044. In some embodiments, the encoder may encode parameters of the loop filter (e.g., loop filter strength) in a binary encoding stage 226, along with the quantized transform coefficients 216, prediction data 206, and other information.

[0075] Figure 3A A schematic diagram illustrating an example decoding process 300A according to some embodiments of the present disclosure is shown. Process 300A may correspond to Figure 2A In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder may decode the video bitstream 228 into a video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to the compression and decompression processes (e.g., Figures 2A-2BTypically, the video stream 304 is different from the video sequence 202, similar to Figures 2A-2B 200A and 200B, the decoder may perform process 300A at a basic processing unit (BPU) level for each picture encoded in the video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of decoding process 300A. In some embodiments, the decoder may perform process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded in the video bitstream 228.

[0076] exist Figure 3A In the process 300A, the decoder may provide a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to the prediction stage 204 to generate the predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 to perform a prediction operation in the next iteration of the process 300A.

[0077] The decoder may iteratively perform process 300A to decode each coded BPU of a coded picture and generate prediction reference 224 for use in encoding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to a video stream 304 for display and continue decoding the next coded picture in the video bitstream 228.

[0078] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in the binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, the transform type, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted over the network in packet form, the decoder can depacketize the video bitstream 228 before providing it to the binary decoding stage 302.

[0079] Figure 3B A schematic diagram of another example decoding process 300B according to some embodiments of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by decoders compliant with hybrid video coding standards (e.g., the H.26x series). Compared to process 300A, process 300B additionally separates the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0080] In process 300B, for a coded basic processing unit (referred to as a "current BPU") of a coded image being decoded (referred to as a "current image"), the prediction data 206 decoded by the decoder 302 from the binary decoding stage may include various types of data, depending on the prediction mode used by the encoder to encode the current BPU. For example, if the encoder encodes the current BPU using intra-frame prediction, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters for the intra-frame prediction operation, and the like. The parameters for the intra-frame prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, extrapolation parameters, the orientation of the neighboring BPUs relative to the original BPU, and the like. For another example, if the encoder encodes the current BPU using inter-frame prediction, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-frame prediction, parameters for the inter-frame prediction operation, and the like. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights respectively associated with the reference images, the positions (e.g., coordinates) of one or more matching regions in each reference image, one or more motion vectors respectively associated with the matching regions, etc.

[0081] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Figure 2B The details of performing such spatial prediction or temporal prediction are described in detail in

[15] and will not be repeated here. After performing the spatial prediction or temporal prediction, the decoder may generate a predicted BPU 208. The decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as shown in FIG. Figure 3A As described in.

[0082] In process 300B, the decoder may provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-frame prediction in the spatial prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter-frame prediction in the temporal prediction stage 2044, then after generating the prediction reference 224 (e.g., a reference picture in which all BPUs have been decoded), the encoder may provide the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may Figure 2B The loop filter is applied to the prediction reference 224 in the manner described in

[15] . The loop filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., as an inter-frame prediction reference picture for a future coded picture in the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that the current BPU is encoded using inter-frame prediction, the prediction data may further include parameters of the loop filter (e.g., loop filter strength).

[0083] Figure 4 FIG is a block diagram of an example apparatus 400 for encoding or decoding a video according to some embodiments of the present disclosure. Figure 4As shown, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 can become a special-purpose machine for video encoding or decoding. The processor 402 can be any type of circuit capable of operating or processing information. For example, the processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chip (SoCs), application-specific integrated circuits (ASICs), and the like. In some embodiments, the processor 402 may also be a group of processors grouped into a single logical component. For example, as shown in the figure. As Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0084] The apparatus 400 may also include a memory 404 configured to store data (eg, a set of instructions, computer code, intermediate data, etc.). Figure 4 As shown, the stored data may include program instructions (e.g., program instructions for implementing stages in process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any number of random access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a group of memories grouped into a single logical component ( Figure 4 not shown).

[0085] The bus 410 may be a communication device for transmitting data between components within the apparatus 400 , such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a universal serial bus port, a peripheral component interconnect port), etc.

[0086] For ease of explanation and to avoid ambiguity, in this disclosure, processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry." The data processing circuitry may be implemented entirely in hardware or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry may be a single standalone module or may be fully or partially integrated into any other component of apparatus 400.

[0087] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, any combination of wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.

[0088] In some embodiments, the apparatus 400 may optionally further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), etc.

[0089] It should be noted that a video codec (e.g., a codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules in apparatus 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of apparatus 400, such as program instructions that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of apparatus 400, such as dedicated data processing circuits (e.g., FPGAs, ASICs, NPUs, etc.).

[0090] Figure 5 A schematic diagram of an exemplary Luma Mapping with Chroma Scaling (LMCS) process 500 is shown according to some embodiments of the present disclosure. For example, process 500 may be used by a decoder compliant with a Hybrid Video Coding standard (e.g., the H.26x series). LMCS is a Figure 2B A new processing block is applied before the loop filter 232. The LMCS may be referred to as a reshaper.

[0091] The LMCS process 500 may include in-loop mapping of luma component values based on an adaptive piecewise linear model and luma-dependent chroma residual scaling of chroma components.

[0092] As shown in the figure. Figure 5 The in-loop mapping of luma component values based on the adaptive piecewise linear model may include a forward mapping stage 518 and an inverse mapping stage 508. The luma-dependent chroma residual scaling of the chroma components may include a chroma scaling 520.

[0093] The sample values before mapping or after inverse mapping can be referred to as samples in the original domain, and the sample values after mapping and before inverse mapping can be referred to as samples in the mapped domain. When LMCS is enabled, certain stages in process 500 can be performed in the mapped domain rather than the original domain. It will be appreciated that the forward mapping stage 518 and the inverse mapping stage 508 can be enabled / disabled at the sequence level using the SPS flag.

[0094] like Figure 5 As shown, Q -1 &T -1 Stage 504, reconstruction 506 and intra prediction 514 are performed in the mapped domain. -1 &T -1 Stage 504 may include inverse quantization and inverse transformation, reconstruction 506 may include luma prediction and addition of luma residual, and intra prediction 508 may include luma intra prediction.

[0095] The loop filter 510, motion compensation stages 516 and 530, intra prediction stage 528, reconstruction stage 522, and decoded picture buffers (DPBs) 512 and 526 are performed in the original (i.e., non-mapped) domain. In some embodiments, the loop filter 510 may include deblocking, adaptive loop filter (ALF), and sample adaptive offset (SAO), the reconstruction stage 522 may include the addition of chroma prediction and chroma residual, and the DPBs 512 and 526 may store decoded pictures as reference pictures.

[0096] In some embodiments, a method of processing video content using a luminance mapping with a piecewise linear model may be applied.

[0097] In-loop mapping of the luma component can adjust the signal statistics of the input video to improve compression efficiency by redistributing codewords across the dynamic range. Luma mapping uses a forward mapping function "FwdMap" and a corresponding inverse mapping function "InvMap". The "FwdMap" function is signaled using a piecewise linear model with 16 equal parts. The "InvMap" function does not need to be signaled but is derived from the "FwdMap" function.

[0098] The signal representation of the piecewise linear model is Figure 6 Table 1 and Figure 7 As shown in Table 2, and later in VVC draft 5, the signal of the piecewise linear model changes as follows Figure 8 Table 3 and Figure 9 Table 4 shows the syntax structure of the tile group header and slice header. Table 1 and Table 3 show the syntax structure of the tile group header and slice header. First, the shaper model parameter exists flag can be used to signal whether the brightness mapping model exists in the target tile group or target slice. If the brightness mapping model exists in the current tile group / slice, the shaper model parameter exists flag can be used to signal whether the brightness mapping model exists in the target tile group or target slice. Figure 7 Table 2 and Figure 9 The syntax elements shown in Table 4 of represent the piecewise linear model parameters corresponding to the target tile group or target slice in tile_group_reshaper_model() / lmcs_data(). The piecewise linear model divides the dynamic range of the input signal into 16 equal segments. For each segment, its linear mapping parameters can be represented using multiple codewords assigned to the segment. Taking 10-bit input as an example, by default, for each of the 16 segments of the input, 64 codewords can be assigned to the segment. Multiple signal codewords can be used to calculate the scaling factor and adjust the mapping function for the segment accordingly. Figure 7 Table 2 and Figure 9 Table 4 also defines the minimum and maximum indices for signaling the number of codewords, such as "reshaper_model_min_bin_idx" and "reshaper_model_delta_max_bin_idx" in Table 2, and "lmcs_min_bin_idx" and "lmcs_delta_max_bin_idx" in Table 4. If a segment index is less than "reshaper_model_min_bin_idx" or "lmcs_min_bin_idx" or greater than "15-reshaper_model_max_bin_idx" or "15-lmcs_delta_max_bin_idx", the number of codewords for that segment is not signaled and is inferred to be zero. In other words, no codeword is assigned and no mapping / scaling is applied to the segment.

[0099] At the tile group header level or the slice header level, another reshaper enable flag (e.g., “tile_group_reshaper_enable_flag” or “slice_lmcs_enabled_flag”) may be signaled to indicate Figure 5Whether the LMCS process described in

[0066] is applied to the target tile group or target slice. If the reshaper is enabled for the target tile group or target slice, and if the target tile group or target slice does not use dual-tree partitioning, a further chroma scaling enable flag may be signaled to indicate whether chroma scaling is enabled for the target slice group or target slice. It will be appreciated that dual-tree partitioning may also be referred to as a chroma separation tree. Below, the present disclosure will explain dual-tree partitioning in more detail.

[0100] The piecewise linear model can be constructed as follows based on the signaled syntax elements in Table 2 or Table 4. The i-th segment (i=0...15) of the "FwdMap" piecewise linear model can be defined by two input pivot points InputPivot[] and two mapped pivot points MappedPivot[]. The mapped pivot points MappedPivot[] can be the output of the "FwdMap" piecewise linear model. Assuming that the bit depth of the exemplary input video is 10 bits, InputPivot[] and MappedPivot[] can be calculated based on the following signal syntax. It will be understood that the bit depth can be different from 10 bits.

[0101] a) Use the syntax elements in Table 2:

[0102] 1) OrgCW=64

[0103] 2)For i=0:16,InputPivot[i]=i*OrgCW

[0104] 3)For i=reshaper_model_min_bin_idx:reshaper_model_max_bin_idx,

[0105] 4) For i = 0:16, MappedPivot[i] is calculated as follows:

[0106] MappedPivot[0] = 0;

[0107] for(i=0;i<16;i++)

[0108] MappedPivot[i+1]=MappedPivot[i]+SignaledCW[i]

[0109] b) Using the syntax elements in Table 4:

[0110] 1) OrgCW=64

[0111] 2)For i=0:16,InputPivot[i]=i*OrgCW

[0112] 3)For i=lmcs_min_bin_idx:lmcsl_max_bin_idx,

[0113] 4) For i = 0:16, MappedPivot[i] is calculated as follows:

[0114] MappedPivot[0] = 0;

[0115] for(i=0;i<16;i++)

[0116] MappedPivot[i+1]=MappedPivot[i]+SignaledCW[i]

[0117] The inverse mapping function "InvMap" is also defined by InputPivot[] and MappedPivot[]. Unlike "FwdMap", for the "InvMap" piecewise linear model, the two input pivot points of each segment are defined by MappedPivot[], and the two output pivot points are defined by InputPivot[]. In this way, the input of "FwdMap" is divided into equal segments, but it is not guaranteed that the input of "InvMap" is divided into equal segments.

[0118] like Figure 5 As shown, for inter-frame coded blocks, motion compensated prediction can be performed in the mapping domain. In other words, after motion compensated prediction 516, Y is calculated based on the reference signal in the DPB. pred , the “FwdMap” function 518 may be applied to map the luma prediction block in the original domain to the mapped domain Y′ pred =FwdMap(Y pred ). For intra-coded blocks, the "FwdMap" function is not applied because the reference samples used in intra prediction are already in the mapped domain. After reconstructing block 506, Y r The "InvMap" function 508 may be applied to convert the reconstructed luminance values in the mapped domain back to the original domain. The "InvMap" function 508 can be applied to both intra- and inter-coded luminance blocks.

[0119] The brightness mapping process (forward or inverse mapping) can be implemented using a lookup table (LUT) or using on-the-fly calculations. If LUTs are used, the tables "FwdMapLUT[]" and "InvMapLUT[]" can be pre-calculated and stored for use at the tile group level or slice level, and the forward and inverse mapping can be simply implemented as FwdMap(Y pred )=FwdMapLUT[Y pred ] and InvMap(Y r )=InvMapLUT[Y r ].

[0120] Alternatively, on-the-fly calculations can be used. Take the forward mapping function "FwdMap" as an example. To determine the segment to which a luma sample belongs, the sample value can be right-shifted by 6 bits (assuming 10-bit video, corresponding to 16 equal segments) to obtain the segment index. Then, the linear model parameters of the segment are retrieved and applied on-the-fly to calculate the mapped luma value. The calculation of the "FwdMap" function is as follows:

[0121] Y′ pred =FwdMap(Y pred )=((b2-b1) / (a2-a1))*(Y pred -a1)+b1

[0122] Where "i" is the fragment index, al is InputPivot[i], a2 is InputPivot[i+1], b1 is MappedPivot[i], b2 is MappedPivot[i+1],

[0123] The "InvMap" function can be calculated on the fly in a similar manner, except that a conditional check needs to be applied instead of a simple right shift when determining which segment a sample value belongs to, since there is no guarantee that the segments in the mapped domain are of equal size.

[0124] In some embodiments, a method for processing video content using luma-dependent chroma residual scaling may be provided.

[0125] Chroma residual scaling can be used to compensate for the interaction between the luma signal and the chroma signal corresponding to the luma signal. Whether chroma residual scaling is enabled can be signaled at the tile group level or the slice level. Figure 6 Table 1 and Figure 8As shown in Table 3 in

[0045] , if luma mapping is enabled and if dual-tree partitioning is not applied to the current tile group, an additional flag (e.g., "tile_group_reshaper_chroma_residual_scale_flag" or "slice_chroma_residual_scale_flag") is signaled to indicate whether luma-dependent chroma residual scaling is enabled. When luma mapping is not used or dual-tree partitioning is used in the target tile group (or target slice), luma-dependent chroma residual scaling may be disabled accordingly. In addition, luma-dependent chroma residual scaling may be disabled for chroma blocks with an area less than or equal to 4.

[0126] The chroma residual scaling depends on the average value of the luma prediction blocks (for intra and inter coded blocks) corresponding to the chroma signal. The average value of the luma prediction block "avgY'" can be determined using the following equation

[0127]

[0128] Chroma Residual Scaling "C scaleInv The value of the chroma scaling factor can be determined using the following steps:

[0129] 1) Find the index Y of the piecewise linear model to which avgY' belongs according to the InvMap function Idx ,

[0130] 2) C scaleInv =cScaleInv[Y Idx ], where cScaleInv[] is a precomputed LUT with, for example, 16 fragments.

[0131] In some embodiments, in the LMCS method, a pre-calculated LUT "cScaleInv[i]" can be derived based on the 64-entry static LUT "ChromaResidualScaleLut" and the signal codeword "SignaledCW[i]" value, where i ranges from 0 to 15, as shown below.

[0132] ChromaResidualScaleLut

[64] ={16384,16384,16384,16384,16384,16384,16384,8192,8192,8192,81 92,5461,5461,5461,5461,4096,4096,4096,4096,3277,3277,3277,3277,2731,2731,2731,2731,2341, 2341,2341,2048,2048,2048,1820,1820,1820,1638,1638,1638,1489,1489,1489,1365,1365,1365,1260,1260,1260,1260,1170,1170,1170,1092,1092,1092,1024,1024,1024};

[0133] shiftC=11

[0134] -if(SignaledCW[i]=0)

[0135] cScaleInv[i]=(1< <shiftC)

[0136] -otherwise

[0137] cScaleInv[i]=ChromaResidualScaleLut[(SignaledCW[i]>>1)-1]

[0138] For example, assuming the input is 10 bits, the static LUT "ChromaResidualScaleLut[]" contains 64 entries, and the signal codeword "SignaledCW[]" is in the range of [0,128], so the chroma scaling factor LUT "cScalelnv[]" is constructed using division by 2 (or right shift by 1). The LUT "cScalelnv[]" can be constructed at the tile group (or slice level).

[0139] If the current block can be coded using intra, CIIP, or intra block copy (IBC) mode, avgY' can be determined as the average of the intra, CIIP, or IBC predicted luminance values. Otherwise, avgY' is calculated as the average luminance value of the forward-mapped inter prediction (i.e., Figure 5 Y' pred ). IBC can also be called Current Picture Reference (CPR) mode. Unlike brightness mapping performed based on samples, “C ScaleInv" is a constant value for the entire chroma block. Use "C ScaleInv ", chroma residual scaling can be applied at the decoder side as follows:

[0140]

[0141] in is the reconstructed chroma residual of the current block. At the encoder side, forward chroma residual scaling (before being transformed and quantized) is performed as follows: Encoder side: C ResScale =C Res *C Scale =C Res / C ScaleInv .

[0142] In some embodiments, a method for processing video content using cross-component linear model prediction may be provided.

[0143] To reduce cross-component redundancy, the Cross-Component Linear Model (CCLM) prediction mode can be used. In CCLM, chroma samples are predicted based on the reconstructed luma samples of the same coding unit (CU) using a linear model, as shown below:

[0144] pred C (i,j)=α·rec L ′(i,j)+β

[0145] where pred C (i, j) represents the predicted chroma sample in CU, rec L (i, j) represents the downsampled reconstructed luma sample of the same CU.

[0146] The linear model parameters a and β are derived based on the relationship between luma and chroma values from two sample positions. The two sample positions may be comprised of a first luma sample position having a maximum luma sample value and a second luma sample position having a minimum luma sample value, in a set of downsampled adjacent luma samples, and their corresponding chroma samples. The linear model parameters a and β are obtained according to the following equations.

[0147]

[0148] β=Y b -α·X b

[0149] where Y a and X a Represents the luminance value and chrominance value of the first luminance sample position respectively. b and Y b Represent the luminance value and chrominance value of the second luminance sampling position respectively.

[0150] Figure 10 illustratively shows examples of sample positions involved in CCLM mode according to some embodiments of the present disclosure.

[0151] The calculation of parameter a can be implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and parameter a are expressed in exponential notation. For example, diff is approximated as a 4-bit significand and an exponent. Therefore, for 16 values of the significand, the table of 1 / diff is reduced to 16 elements, as shown below:

[0152] DivTable[]={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}

[0153] The table "DivTable[]" can reduce the complexity of the calculation and also reduce the memory size required to store the required table.

[0154] In addition to the top position and the left position being used together to calculate the linear model coefficients, they can also be used alternately in the other two LM modes, called LM_A and LM_L modes.

[0155] In LM_A mode, only the samples in the top position are used to calculate the linear model coefficients. To obtain more samples, the top position can be extended to cover (W+H) samples. In LM_L mode, only the samples in the left position are used to calculate the linear model coefficients. To obtain more samples, the left position can be extended to cover (H+W) samples.

[0156] For a non-square block, the upper template is expanded to W+W, and the left template is expanded to H+H.

[0157] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters can be applied to the luma samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter can be specified by the SPS level flag. The two downsampling filters are as follows, corresponding to the contents of "type-0" and "type-2" respectively.

[0158]

[0159] It can be understood that when the upper reference line is located at a CTU boundary, only one luma line (usually in the line buffer for intra prediction) is used to calculate the downsampled luma samples.

[0160] This parameter calculation can be performed as part of the decoding process, not just as an encoder search operation. Therefore, no syntax is used to transmit the α and β values to the decoder. The α and β parameters are calculated separately for each chroma component.

[0161] For chroma intra mode coding, a total of eight intra modes are allowed. These modes include five traditional intra modes and three cross-component linear model modes (e.g., CCLM, LM_A, and LM_L). Figure 9 Table 5 shows the process for representing and deriving chroma modes when CCLM is enabled. The chroma mode encoding of a chroma block can depend on the intra-frame prediction mode of the luma block corresponding to the chroma block. Because the separate block partitioning structure for luma and chroma components is enabled in Islice (described below), a chroma block can correspond to multiple luma blocks. Therefore, for the chroma derivation mode (DM), the intra-frame prediction mode of the corresponding luma block covering the center position of the current chroma block is inherited.

[0162] In some embodiments, a method for processing video content using dual-tree partitioning may be provided.

[0163] In the VVC draft, the coding tree scheme supports the ability for luminance and chrominance to have separate block tree partitions. This is also called dual tree partitioning. In the VVC draft, the signal representation of the dual tree partition is Figure 12 Table 6 and Figure 13 As shown in Table 7 of VVC Draft 5, the dual tree partitioning is Figure 14 Table 8 and Figure 15 9 of . When the sequence level control flag for signaling in the SPS (e.g., "qtbtt_dual_tree_intra_flag") is turned on and when the target tile group (or target slice) is intra-coded, the block partitioning information can be signaled first for luma and then for chroma separately. For inter-coded tile groups / slices (e.g., P and B tile groups / slices), dual tree partitioning is not allowed. When separate block tree mode is applied, luma coding tree blocks (CTBs) are partitioned into CUs by a first coding tree structure and chroma CTBs are partitioned into chroma CUs by a second coding tree structure, as shown in FIG. Figure 13 As shown in Table 7.

[0164] While allowing different partitioning for luma and chroma blocks, this can cause issues for coding tools with dependencies between different color components. For example, when applying LMCS, the average of the luma blocks corresponding to the target chroma block can be used to determine the scaling factor to be applied to the target chroma block. When using dual-tree partitioning, determining the average for the luma block incurs latency across the entire CTU. For example, if a CTU's luma block is split vertically once and its chroma blocks are split horizontally once, both luma blocks of the CTU must be decoded and the average calculated before the first chroma block of the CTU can be decoded. In VVC, CTUs can be up to 128×128 in terms of luma samples, significantly increasing the decoding latency of chroma blocks. Therefore, VVC drafts 4 and 5 prohibit the combination of dual-tree partitioning and luma-dependent chroma scaling. Chroma scaling can be forced off when dual-tree partitioning is enabled for the target tile group (or target slice). Note that the luma mapping portion of LMCS is still allowed with dual-tree partitioning, as it operates only on the luma component and does not have cross-color component dependencies.

[0165] Another example of a coding tool that relies on correlation between color components to achieve better coding efficiency is called the Cross Component Linear Model (CCLM), which was discussed above. In CCLM, adjacent reconstructed luma and chroma samples can be used to derive cross-component parameters. And the cross-component parameters can be applied to the corresponding reconstructed luma samples of the target chroma block to derive predictors for the chroma components. When using dual-tree partitioning, the luma and chroma partitions are not guaranteed to be aligned. Therefore, CCLM cannot be started for a chroma block until all corresponding luma blocks containing samples used in CCLM have been reconstructed.

[0166] Figures 16A-16B An exemplary chroma tree partitioning and an exemplary luma tree partitioning according to some embodiments of the present disclosure are illustrated. Figure 16A An exemplary partition structure of a chroma block 1600 is illustrated. Figure 16B The diagram corresponds to Figure 16A An exemplary division structure of the chrominance block 1600 and the luminance block 1610. Figure 16A In , the chroma block 1600 is divided into four sub-blocks, the sub-block in the lower left corner is further divided into four sub-blocks, and the block with a grid pattern is the current block to be predicted. Figure 16B In Figure 1, luma block 1610 is horizontally bisected into two sub-blocks. The area with the grid pattern corresponds to the target chroma block to be predicted. To derive CCLM parameters, the adjacent reconstructed sample values, represented by the empty circles, are required. Therefore, prediction of the target chroma block cannot begin until reconstruction of the bottom luma block is complete, introducing a significant delay.

[0167] In some embodiments, a method for processing video content using a virtual pipeline data unit is provided.

[0168] In the VVC standardization, the concept of Virtual Pipeline Data Unit (VPDU) was introduced to enable more user-friendly hardware implementation. A VPDU is defined as a non-overlapping M×M-luminance (L) / N×N-chrominance (C) unit in the image. In the hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. Different stages process different VPDUs simultaneously. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is important to keep the VPDU size small. In VVC, the VPDU size is set to 64×64 samples. Therefore, all coding tools used in VVC cannot violate the VDPU limit. For example, the maximum transform size can only be 64×64 because the entire transform block needs to be operated on in the same pipeline stage. Due to the VPDU limit, intra-frame prediction blocks should also be no larger than 64x64. Therefore, in an intra-coded tile group / slice (e.g., Itile group / slice), the CTU is forced to be split into four 64x64 blocks (if the CTU is larger than 64×64), and each 64x64 block can be further split into two using a dual-tree structure. Therefore, when dual-tree is enabled, the common root of the luma partition tree and the chroma partition tree is at a 64x64 block size.

[0169] There are several problems with current LMCS and CCLM designs.

[0170] First, for example, the derivation of the tile group-level chroma scaling factor LUT "cScaleInv[]" does not scale easily. The derivation currently depends on a constant chroma LUT "ChromaResidualScaleLut" with 64 entries. For 10-bit video with 16 fragments, an additional step of division by 2 must be applied. When the number of fragments changes (for example, if 8 fragments are used instead of 16 fragments), the derivation must be changed to apply an additional step of division by 4 instead of 2. This extra step may cause a loss of precision.

[0171] Second, for example, to calculate Y for obtaining the chroma scaling factor Idx , the average value of the entire luma block is used. Considering the maximum CTU size is 128×128, the average luma value can be calculated based on 16384 (128×128) luma samples, which is relatively complex. Furthermore, if the encoder chooses a 128×128 luma block partition, the block is more likely to contain homogeneous content. Therefore, a subset of the luma samples in the block is sufficient to calculate the luma average.

[0172] Third, during the dual-tree partitioning, chroma scaling is set to off to avoid potential pipeline issues in the hardware decoder. However, if an explicit signal is used to indicate the chroma scaling factor to be applied (instead of using the corresponding luma samples to derive it), this dependency can be avoided. Enabling chroma scaling in intra-coded tile groups / slices can further improve the coding efficiency.

[0173] Fourth, traditionally, for each of the 16 segments, the delta codeword value is signaled. It has been observed that only a limited number of different codewords are typically used for the 16 segments. Therefore, the overhead of signaling can be further reduced.

[0174] Fifth, the parameters of the CCLM are derived from blocks that are causally adjacent to the block that is the target chroma block using luma and chroma reconstruction samples. In the dual-tree partitioning, the luma and chroma block partitions are not necessarily aligned. Therefore, the target chroma block can correspond to multiple luma blocks or a luma block that is larger in area than the target chroma block. To derive the CCLM parameters for the target chroma block, all corresponding luma blocks must first be reconstructed, as Figures 16A-16B shown. This causes latency in the pipeline implementation and reduces the throughput of the hardware decoder.

[0175] To address the above problems, the present disclosure provides the following embodiments.

[0176] Embodiments of the present disclosure provide a video content processing method by removing the chroma scaling LUT

[0177] As described above, the 64-entry chroma LUT is not easily scalable and may cause problems when using other piecewise linear models (e.g., 8, 4, 64 segments, etc.). This is also unnecessary because the chroma scaling factor can be set to the same as the luma scaling factor for the corresponding segment to achieve the same coding efficiency. In some embodiments of the present invention, denoting Y Idx as the segment index of the current chroma block, the steps to determine the chroma scaling factor are as follows:

[0178] · If Y Idx > reshaper_model_max_bin_idx or Y Idx < reshaper_model_min_bin_idx, or if SignaledCW[Y Idx = 0, then set chroma_scaling to a default value, e.g., chroma_scaling = 1.0, i.e., no scaling is applied.

[0179] · Otherwise, set chroma_scaling to SignaledCW[Y Idx / OrgCW.

[0180] The chroma scaling factors derived above have fractional precision. A fixed-point approximation can be applied to avoid hardware / software platform dependencies. Furthermore, at the decoder side, inverse chroma scaling needs to be performed. This division can be implemented using fixed-point arithmetic with multiplication followed by a right shift. The number of bits in the fixed-point approximation is denoted as CSCALE_FP_PREC. The following can be used to determine the inverse chroma scaling factor in fixed-point precision:

[0181] inverse_chroma_scaling[Y Idx ]=((1<<(luma_bit_depth-

[0182] log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(SignaledCW[Y Idx ]

[0183] >>1)) / SignaledCW[Y Idx ];

[0184] Where luma_bit_depth is the luma bit depth, TOTAL_NUMBER_PIECES is the total number of pieces in the piecewise linear model, which is set to 16 in VVC draft 4. Note that the inverse_chroma_scaling value may only need to be calculated once for each tile group / slice, and the above division is an integer division operation.

[0185] Further quantization can be applied to derive the chroma scaling and inverse scaling factors. For example, the inverse chroma scaling factor can be calculated for all even (2×m) values of SignaledCW, and the odd (2×m+1) values of SignaledCW can use the chroma scaling factor of the scaling factor of the adjacent even value. In other words, the following steps can be used:

[0186]

[0187] The above embodiment with quantized chroma scaling factors can be further generalized. For example, an inverse chroma scaling factor can be calculated for every n-th value of SignaledCW, with all other adjacent values sharing the same chroma scaling factor. For example, "n" can be set to 4, meaning that every 4 adjacent codeword values share the same inverse chroma scaling factor value. The value of "n" is preferably a power of 2, which allows the use of shifts to calculate divisions. Denoting the value of log2(n) as LOG2_n, the above can be modified to: tempCW = SignaledCW[i]>>LOG2_n)< <LOG2_n。

[0188] Finally, the value of LOG2_n may be a function of the number of pieces used in the piecewise linear model. If fewer pieces are used, it may be beneficial to use a larger LOG2_n. For example, if the value of TOTAL_NUMBER_PIECES is less than or equal to 16, then LOG2_n may be set to 1 + (4 - log2(TOTAL_NUMBER_PIECES)). If TOTAL_NUMBER_PIECES is greater than 16, then LOG2_n may be set to 0.

[0189] Embodiments of the present disclosure provide a method for processing video content by simplifying average calculation of luma prediction blocks.

[0190] As mentioned above, in order to determine the current chrominance block "Y Idx ", the average value of the corresponding luminance blocks can be used. However, for large blocks, the averaging process may involve a large number of luminance samples. In the worst case, the averaging process may involve 128×128 luminance samples,

[0191] The embodiments of the present disclosure provide a simplified averaging process to reduce the worst case to using only N×N luma samples (N is a power of 2).

[0192] In some embodiments, if both dimensions of a non-two-dimensional luma block are less than or equal to a preset threshold N (in other words, at least one of the two dimensions is greater than N), then "downsampling" can be applied using only N locations in that dimension. Without loss of generality, let's take the horizontal dimension as an example. If the width is greater than N, then only the sample at position x is used for averaging, where x = i × (width >> log2(N)), i = 0, ... N-1.

[0193] Figure 17 The figure illustrates an exemplary simplification of an averaging operation according to some embodiments of the present disclosure. In this example, N is set to 4, and only 16 luma samples (shaded samples) in the block are used for averaging. It will be appreciated that the value of N is not limited to 4. For example, N can be set to any value that is a power of 2. In other words, N can be 1, 2, 4, 8, and so on.

[0194] In some embodiments, different values of N can be applied in the horizontal and vertical dimensions. In other words, the worst-case scenario for the averaging operation may be to use N×M samples. In some embodiments, the number of samples used in the averaging process can be limited regardless of the dimensions. For example, a maximum of 16 samples can be used. These 16 samples can be distributed in the horizontal or vertical dimensions in the form of 1×16, 16×1, 2×8, 8×2, 4×4, or a form suitable for the shape of the target block. For example, if the block is narrow and tall, use 2×8, if the block is wide and short, use 8×2, and if the block is square, use 4×4.

[0195] While this simplification can cause the average to differ from the true average across the luminance block, any such difference is likely to be small. This is because the content within the block tends to be more uniform when larger block sizes are chosen.

[0196] In addition, the decoder-side motion vector refinement (DMVR) mode is a complex process in the VVC standard, especially for the decoder. This is because DMVR requires the decoder to perform a motion search to derive motion vectors before applying motion compensation. The bidirectional optical flow (BDOF) mode in the VVC standard further complicates the situation because BDOF is an additional sequential process that needs to be applied after DMVR to obtain the luma prediction block. Since chroma scaling requires the average value of the corresponding luma prediction blocks, DMVR and BDOF can be applied before the average value is calculated, which will cause delay issues.

[0197] To address this latency issue, in some embodiments of the present disclosure, a luma prediction block is used to calculate the average luma value before DMVR and BDOF, and the average luma value is used to obtain the chroma scaling factor. This allows chroma scaling to be applied in parallel with the DMVR and BDOF processes, thus significantly reducing latency.

[0198] Various variations of latency reduction consistent with the present disclosure can be envisioned. In some embodiments, this latency reduction can also be combined with the simplified averaging process discussed above, which uses only a portion of the luma prediction block to calculate the average luma value. In some embodiments, the luma prediction block can be used after the DMVR process and before the BDOF process to calculate the average luma value. The average luma value is then used to obtain the chroma scaling factor. This design allows chroma scaling to be applied in parallel with the BDOF process while maintaining accuracy in determining the chroma scaling factor. The DMVR process can refine the motion vectors, so that using prediction samples with refined motion vectors after the DMVR process can be more accurate than using prediction samples with motion vectors before the DMVR process.

[0199] In addition, in the VVC standard, the CU syntax structure (e.g., coding_unit()) may include a syntax element "cu_cbf" to indicate whether there are any non-zero residual coefficients in the target CU. At the TU level, the TU syntax structure transform_unit() includes syntax elements tu_cbf_cb and tu_cbf_cr to indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the target TU. In VVC draft 4, if chroma scaling is enabled at the tile group level or slice level, the averaging of the corresponding luminance block can be called. The present disclosure also provides a method for bypassing the luminance averaging process. Consistent with the disclosed embodiments, since the chroma scaling process is applied to the residual chroma coefficients, if there are no non-zero chroma coefficients, the luminance averaging process can be bypassed. This can be determined according to the following conditions:

[0200] Condition 1: cu_cbf is equal to 0

[0201] Condition 2: tu_cbf_cr and tu_cbf_cb are both equal to 0

[0202] When condition 1 or condition 2 is met, the brightness averaging process can be bypassed.

[0203] In the above embodiments, only N×N samples of the prediction block are used to derive the average value, which simplifies the averaging process. For example, when N is equal to 1, only the top left corner sample of the prediction block can be used. However, even this simplified case requires the prediction block to be generated first, which incurs a delay. Therefore, in some embodiments, it is expected that the reference luma samples can be used directly to derive the chroma scaling factors. This allows the decoder to derive the chroma scaling factors in parallel with the luma prediction process, thereby reducing delays. In other words, intra-frame prediction and inter-frame prediction are handled separately.

[0204] In the case of intra prediction, adjacent samples that have been decoded in the same image can be used as reference samples to generate the predicted block. These reference samples include samples at the top of the target block, the left side of the target block, and the upper left corner of the target block. The average of all these reference samples can be used to derive the chroma scale factor. Alternatively, the average of only a portion of these reference samples can be used. For example, only the M reference samples (e.g., M=3) closest to the upper left corner of the target block can be averaged.

[0205] As another example, the M reference samples averaged to derive the chroma scaling factor are not closest to the upper left position but are distributed along the upper and left boundaries of the target block, as shown in Figure 18 shown. Figure 18 An exemplary sample is shown for use in the average calculation to derive the chroma scaling factor. Figure 18 As shown. Figure 18 In , the exemplary samples are shown in solid dashed boxes. Figure 18 In each of the example blocks 1801-1807 shown, one, two, three, four, five, six, and eight samples are averaged, respectively. The average calculation in the present disclosure can be replaced by a weighted average, where different samples can have different weights in the average calculation. For example, the sum of the weights can be a power of 2 to avoid division operations in the average calculation.

[0206] In the case of inter-frame prediction, reference samples from a temporal reference picture can be used to generate the prediction block. These reference samples are identified by a reference picture index and a motion vector. If the motion vector has fractional precision, interpolation can be applied. To calculate the average of the reference samples, the interpolated reference samples can be used, or the reference samples before interpolation can also be interpolated (i.e., the motion vectors are clipped to integer precision). Consistent with the disclosed embodiments, all reference samples can be used to calculate the average. Alternatively, only a portion of the reference samples (e.g., the reference samples corresponding to the upper left corner of the target block) can be used to calculate the average.

[0207] As shown in the figure. Figure 5 In

[15] , intra prediction is performed in the reconstructed domain, while inter prediction is performed in the original domain. Therefore, for inter prediction, the prediction block is forward mapped, and the average value of the luma block is calculated using the luma prediction block after forward mapping. To reduce latency, the average value can be calculated using the luma prediction block before forward mapping. For example, the entire luma block before forward mapping, an N×N portion of the luma block before forward mapping, or the top-left sample of the luma block before forward mapping can be used.

[0208] An embodiment of the present disclosure also provides a method for processing dual-tree partitioned video content with chroma scaling.

[0209] Because reliance on luma blocks complicates hardware design, chroma scaling can be disabled for intra-coded tile groups / slices with dual-tree partitioning enabled. However, this restriction results in a loss in coding efficiency.

[0210] Because CTU is the common root of the luma coding tree and the chroma coding tree, deriving the chroma scaling factor at the CTU level can eliminate the dependency between chroma and luma in the dual-tree partitioning. For example, the reconstructed luma samples or chroma samples adjacent to the CTU are used to derive the chroma scaling factor. The chroma scaling factor can then be used for all chroma samples within the CTU. In this example, the above-mentioned method of averaging reference samples can be applied to average the adjacent reconstructed samples of the CTU. The average value of all these reference samples can be used to derive the chroma scaling factor. Alternatively, only the average value of a portion of these reference samples can be used. For example, only the M reference samples (e.g., M = 4, 8, 16, 32, or 64) closest to the upper left corner of the target block are averaged.

[0211] Because CTU is the common root of the luma coding tree and the chroma coding tree, deriving the chroma scaling factor at the CTU level can eliminate the dependency between chroma and luma in the dual-tree partitioning. For example, the reconstructed luma samples or chroma samples adjacent to the CTU are used to derive the chroma scaling factor. The chroma scaling factor can then be used for all chroma samples within the CTU. In this example, the above-mentioned method of averaging reference samples can be applied to average the adjacent reconstructed samples of the CTU. The average value of all these reference samples can be used to derive the chroma scaling factor. Alternatively, only the average value of a portion of these reference samples can be used. For example, only the M reference samples (e.g., M = 4, 8, 16, 32, or 64) closest to the upper left corner of the target block can be averaged.

[0212] However, for a CTU on the bottom or right boundary of the image, all samples of the CTU may not be within the image boundary, such as Figure 19 In this case, only adjacent reconstructed samples on the CTU boundary within the image boundary ( Figure 19 Gray samples in the image can be used to derive the chroma scaling factor. However, a variable number of samples in the average calculation requires a division operation, which is not desirable in hardware implementation. Therefore, an embodiment of the present invention provides a method for padding the image boundary samples with a fixed number of powers of 2, thereby avoiding the division operation in the average calculation. For example, as shown in the figure. Figure 19 In the example, the padding samples outside the bottom boundary of the image ( Figure 19 ) is generated from sample 1905 which is the closest padding sample among all the samples on the bottom border of the image. In addition to the CTU level chroma scaling factor derivation, the chroma scaling factors can be derived on a fixed grid. Considering that a virtual pipeline data unit (VPDU) is defined as a data unit processed by a pipeline stage, the chroma scaling factors can be derived at the VPDU level. In VVC draft 5, VPDU is defined as a 64×64 block on a luma sample grid. Therefore, an embodiment of the present invention provides a chroma scaling factor derivation with a 64×64 block granularity. In VVC draft 6, VPDU is defined as M×M blocks on a luma sample grid, where M is the smaller of the CTU size and 64. The CTU level derivation method explained previously can also be used for the VPDU level.

[0213] In some embodiments, in addition to deriving the chroma scaling factor on a fixed grid with a grid size smaller than the CTU, each CTU derives the factor only once and uses it for all grid cells (e.g., VPDUs) within that CTU. For example, the chroma scaling factor is derived on the first VPDU of a CTU, and this factor is used for all VPDUs within that CTU. It can be understood that the method at the VPDU level is equivalent to the CTU-level derivation using a limited number of adjacent samples (e.g., only using the adjacent samples corresponding to the first VPDU in the CTU).

[0214] Calculate avgY' at the CTU level, VPDU level, or any other fixed-size block unit level by averaging the sample values of the corresponding luma block, and determine the segment index Y Idx , and as an alternative to obtaining the chroma scaling factor inverse_chroma_scaling[Y Idx , the chroma scaling factor can also be explicitly signaled in the bitstream to avoid dependence on luma in the case of binary tree partitioning.

[0215] The chroma scaling index can be signaled at multiple levels. For example, the chroma scaling index can be signaled together with the chroma prediction mode at the coding unit (CU) level, as shown in Figure 20 Table 10 of Figure 21 and Figure 20 Element 2002 in Figure 21 and element 2102 in

[0216] According to the possible values of lmcs_chroma_scaling_idx, the cost of signaling it may be too high, especially for small blocks. Therefore, in some embodiments of the present disclosure, Figure 20 the signaling conditions in Table 10 of Figure 202002 in lmcs_chroma_scaling_idx) is only signaled if the target block contains more than N chroma samples or if the width of the target block is greater than a given width W and / or the height is greater than a given height H. For smaller blocks, if lmcs_chroma_scaling_idx is not signaled, its chroma scaling factor can be determined at the decoder end. For example, the chroma scaling factor can be set to 1.0 in floating point precision. In some embodiments, the value of lmcs_chroma_scaling_idx can be set by default at the tile group header level or the slice header level ( Figure 10 1), blocks (e.g., tiles) without signaled lmcs_chroma_scaling_idx can use the default index at the tile group / slice level to derive the chroma scaling factor for the block. In some embodiments, the chroma scaling factor for a tile can be inherited from its neighbors (e.g., top or left neighbors) that have explicitly expressed scaling factors.

[0217] In addition to signaling this syntax element of "lmcs_chroma_scaling_idx" at the CU level, it can also be signaled at the CTU level. However, when the maximum CTU size in VVC is 128×128, chroma scaling at the CTU level according to the syntax element signaled by "lmcs_chroma_scaling_idx" may be too coarse. Therefore, in some embodiments of the present disclosure, this syntax element of "lmcs_chroma_scaling_idx" may be signaled using a fixed granularity. For example, for an area with 16×16 samples (or an area with 64×64 samples for a VPDU), a lmcs_chroma_scaling_idx may be signaled and applied to samples in an area with 16×16 samples (or an area with 64×64 samples).

[0218] The range of lmcs_chroma_scaling_idx for the target tile group / slice depends on how many chroma scaling factor values are allowed in the target tile group / slice. The range of lmcs_chroma_scaling_idx can be determined using the existing method in VVC, which relies on the 64-entry chroma LUT discussed above. Alternatively, it can be determined using the chroma scaling factor calculation described above.

[0219] As an example, the value of LOG2_n is set to 2 (i.e., "n" is set to 4) for the "quantization" method described above, and the codeword assignment for each tile in the piecewise linear model of the target tile group / slice is set as follows: {0, 65, 66, 64, 67, 62, 62, 64, 64, 64, 67, 64, 64, 62, 61, 0}. Then, there are only 2 possible scaling factor values for the entire tile group, because any codeword value from 64 to 67 can have the same scaling factor value (e.g., 1.0 in decimal precision), and any codeword value from 60 to 63 can have the same scaling factor value (e.g., 60 / 64=0.9375 in fractional precision). For the two ends without assigned codewords, the chroma scaling factor can be set to 1.0 by default. Therefore, in this example, one bit is sufficient to represent the lmcs_chroma_scaling_idx for the blocks in the target slice. Depending on the chroma scaling factor signal indication level, the block may consist of a CU, a CTU or a fixed region.

[0220] In addition to deriving the number of possible chroma scaling factor values using a piecewise linear model, in some embodiments, the encoder can signal a set of chroma scaling factor values in the tile group / slice header. Then, at the block level, the chroma scaling factor value can be determined using this setting and the lmcs_chroma_scaling_idx value for that block.

[0221] Alternatively, to reduce signal representation costs, the chroma scaling factor can be predicted from neighboring blocks. For example, a flag can be used to indicate that the chroma scaling factor of the target block is equal to the chroma scaling factor of the neighboring block of the target block. The neighboring block can be the top or left neighboring block. Therefore, for the target block, a maximum of 2 bits can be used for signal representation. For example, of the 2 bits, the first bit can indicate whether the chroma scaling factor of the target block is equal to the chroma scaling factor of the left neighboring block of the target block, and the second bit can indicate whether the chroma scaling factor of the target block is equal to the chroma scaling factor of the top neighboring block of the target block. If the values of both bits indicate that the chroma scaling factor of the target block is equal to the chroma scaling factor of the top neighboring block or the left neighboring block, the lmcs_chroma_scaling_idx syntax can be signaled.

[0222] Depending on the possibility of different values of “lmcs_chroma_scaling_idx”, variable length codewords may be used for “code_lmcs_chroma_scaling_idx” to reduce the average code length.

[0223] Context-based adaptive binary arithmetic coding (CABAC) may be applied to encode the lmcs_chroma_scaling_idx of the target block. The CABAC context associated with the target block may depend on the lmcs_chroma_scaling_idx of the target block's neighboring blocks. For example, the left neighboring block or the top neighboring block may be used to form the CABAC context. Regarding the binarization of lmcs_chroma_scaling_idx, truncated Rice binarization may be used to binarize lmcs_chroma_scaling_idx.

[0224] By signaling "lmcs_chroma_scaling_idx", the encoder can select adaptive lmcs_chroma_scaling_idx based on the rate-distortion cost. Therefore, using rate-distortion optimization to select lmcs_chroma_scaling_idx improves coding efficiency, which helps offset the increase in signaling cost.

[0225] Embodiments of the present disclosure also provide a method for processing video content by signaling a LMCS piecewise linear model.

[0226] Although the LMCS method in VVC draft 4 uses a piecewise linear model with 16 segments, many unique values of SignaledCW[i] in a tile group / slice are often much smaller than 16. For example, some of the 16 segments may use the default number of codewords "OrgCW", and some of the 16 segments may have the same number of codewords. Therefore, when signaling the LMCS piecewise linear model, multiple unique codewords may be signaled in the form of "listUniqueCW[]", and then, for each segment of the LMCS piecewise linear model, the index of listUniqueCW[] may be sent for target segment codeword selection.

[0227] exist Figure 12 A modified syntax table is provided in Table 12 of , where syntax elements 2202 and 2204 shown in italics are modified according to this embodiment.

[0228] The semantics of the disclosed signaling method are as follows, with underscores varying:

[0229] reshaper_model_min_bin_idx specifies the minimum bin (or segment) index to be used during the reshaper construction process. The value of reshape_model_min_bin_idx should be in the range of 0 to MaxBinIdx, inclusive. The value of MaxBinIdx should be equal to 15.

[0230] reshaper_model_delta_max_bin_idx specifies the maximum allowed bin (or fragment) index MaxBinldx minus the maximum bin index to be used during the reshaper construction process. The value of reshape_model_max_bin_idx is set equal to MaxBinIdx - reshape_model_min_bin_idx.

[0231] reshaper_model_bin_delta_abs_cw_prec_minus1 plus 1 specifies the number of bits used to represent the syntax reshape_model_bin_delta_abs_CW[i].

[0232] reshaper_model_bin_num_unique_cw_minus1 plus 1 specifies the codeword array listUniqueCW size .

[0233] reshaper_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value for the i-th bin.

[0234] reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshape_model_bin_delta_abs_CW[i] as follows:

[0235] -If reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is positive.

[0236] -Otherwise (reshape_model_bin_delta_sign_CW_flag[i] is equal to 0), the corresponding variable RspDeltaCW[i] is negative.

[0237] When reshape_model_bin_delta_sign_CW_flag[i] is not present, it is inferred to be equal to 0.

[0238] The variable RspDeltaCW[i] is exported as RspDeltaCW[i]=(1-2*reshape_model_bin_delta_sign_CW[i])*reshape_model_bin_delta_abs_CW[i]

[0239] The variable listUniqueCW[0] is set equal to OrgCW .variable listUniqueCW[i] with i=1… reshaper_model_bin_num_unique_cw minus 1, including The endpoint is derived as follows:

[0240] - The variable OrgCW is set to be equal to (1 < <BitDepth Y ) / (MaxBinIdx+1).

[0241] -listUniqueCW[i] =OrgCW+RspDeltaCW[i-1]

[0242] reshaper_model_bin_cw_idx[i] specifies the array listUniqueCW[] used to derive RspCW[i] The value of reshaper_model_bin_cw_idx[i] should be between 0 and (reshaper_model_bin_num_unique_ cw_minus1+1), including the endpoints .

[0243] RspCW[i] is derived as follows:

[0244] -If reshaper_model_min_bin_idx<=i<=reshaper_model_max_bin_idx

[0245] RspCW[i]= listUniqueCW[reshaper_model_bin_cw_idx[i]] .

[0246] -Otherwise, RspCW[i]=0.

[0247] If BitDepth Y If the value of is equal to 10, the value of RspCW[i] can be in the range of 32 to 2*OrgCW-1.

[0248] Embodiments of the present disclosure provide a method for processing video content with conditional chroma scaling at the block level.

[0249] like Figure 6As shown in Table 1 of , whether chroma scaling is applied can be determined by the tile_group_reshaper_chroma_residual_scale_flag signaled at the tile group / slice level. However, it may be beneficial to determine whether chroma scaling is applied at the block level. For example, in some embodiments, a CU level flag may be signaled to indicate whether chroma scaling is applied to the target block. The presence of the CU level flag may be conditional on the tile group level flag "tile_group_reshaper_chroma_residual_scale_flag." In other words, the CU level flag may be signaled only when chroma scaling is allowed at the tile group / slice level. The CU level flag may allow the encoder to choose whether to use chroma scaling based on whether it would be beneficial for the target block to use chroma scaling, but this may also result in signaling overhead.

[0250] Consistent with the disclosed embodiments, to avoid the overhead of the signaling described above, whether chroma scaling is applied to a block can depend on the prediction mode of the target block. For example, if the target block is inter-predicted, the prediction signal tends to be good, especially if its reference picture is closer in temporal distance. In this case, chroma scaling can be bypassed because the residual is expected to be very small. For example, pictures in higher temporal levels tend to have reference pictures that are close in temporal distance, and chroma scaling can be disabled for blocks in these pictures that use nearby reference pictures. The difference in picture order count (POC) between the target picture and the reference picture of the target block can be used to determine whether this condition is met.

[0251] In some embodiments, chroma scaling can be disabled for all inter-coded blocks. In some embodiments, chroma scaling can be disabled for all intra-coded blocks. In some embodiments, chroma scaling can be disabled for the combined intra / inter prediction (CIIP) mode defined in the VVC standard.

[0252] In the VVC standard, the CU syntax structure "coding_unit()" may contain the syntax element "cu_cbf" to indicate whether there are any non-zero residual coefficients in the target CU. At the TU level, the TU syntax structure "transform_unit()" may include the syntax elements "tu_cbf_cb" and "tu_cbf_cr" to indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the target TU. The chroma scaling process may be conditional on these flags. As described above, if there are no non-zero residual coefficients, the averaging of the corresponding luminance chroma scaling process may be called, and then the chroma scaling process may be bypassed, and the present disclosure provides a method for bypassing the luminance averaging process.

[0253] Embodiments of the present disclosure provide a method for processing video content with CCLM parameter derivation,

[0254] As mentioned earlier, in VVC5, the CCLM parameters used to predict the target chroma block are derived from the luminance and chroma reconstruction samples of the adjacent blocks. In the case of a dual tree, the luminance block partitioning and the chroma block partitioning may not be aligned. In other words, in order to derive the CCLM parameters of an N×M chroma block, multiple adjacent luminance blocks or luminance blocks larger than 2N×2M (in the case of the color format 4:2:0) may be reconstructed, resulting in delays.

[0255] For example, to reduce latency, CCLM parameters are derived at the CTU / VPDU level. Reconstructed luminance and chrominance samples from neighboring CTU / VPDUs can be used to derive CCLM parameters. And the derived parameters can be applied to all blocks within a CTU / VPDU. For example, the formula described in Cross Component Linear Model Prediction can be used to derive the CCLM parameters with X a and Y a Parameters, where X a and Y a are the luminance and chrominance values of the luminance sample position with the maximum luminance sample value among the adjacent luminance samples of the CTU / VPDU. b and Y b The luminance value and chrominance value of the luminance sample position with the minimum luminance sample among the adjacent luminance samples of the CTU / VPDU are respectively represented. For those skilled in the art, any other derivation process can be used in combination with the CTU / VPDU level parameter derivation concept proposed in this article.

[0256] In addition to the CTU / VPDU level CCLM parameter derivation, such a derivation process can be performed on a fixed luma grid. In VVC draft 5, when dual-tree partitioning is used, separate luma and chroma partitioning can start from a 64×64 luma grid. In other words, the split from a 128×128 CTU to a 64×64 CU can be performed jointly, rather than separately for luma and chroma. Therefore, as another example, CCLM parameters can be derived on a 64×64 luma grid. The adjacent reconstructed luma and chroma samples of the 64×64 grid cell can be used to derive CCLM parameters for all chroma blocks within the 64×64 grid cell. Compared to the CTU level derivation where luma samples can reach 128×128, the 64×64 unit level derivation can be more accurate and also does not have the pipeline delay issues in the current VVC draft 5. In addition to this embodiment, the CCLM parameter derivation can be further simplified by skipping the derivation of certain grids. For example, CCLM parameters are derived only on the first 64x64 block within a CTU, and derivation for subsequent 64x64 blocks in the same CTU is skipped. Parameters derived based on the first 64x64 block are available to all blocks within the CTU.

[0257] Figure 23 A flow chart of an exemplary method 2300 for processing video content according to some embodiments of the present disclosure is illustrated. In some embodiments, the method 2300 may be performed by a codec (e.g., Figures 2A-2B The encoder or Figures 3A-3B For example, the codec may be implemented as one or more software or hardware components of an apparatus (e.g., apparatus 400) for encoding or transcoding a video sequence. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a decoded compressed video sequence (e.g., video stream 304). In some embodiments, the video sequence may be a monitoring video sequence, which may be executed by a monitoring device (e.g., processor 402) associated with a processor of the apparatus. Figure 4 The video sequence may include multiple images. The apparatus may perform method 2300 at the image level. For example, the apparatus may process one image at a time during method 2300. For another example, the apparatus may process multiple images at a time during method 2300. Method 2300 may include the following steps.

[0258] In step 2302, data representing a first block and a second block in an image is received. The plurality of blocks may include the first block and the second block. In some embodiments, the first block may be a target chrominance block (e.g., Figure 16A The second block may be a coding tree block (CTB), a transform unit (TU), or a virtual pipeline data unit (VPDU). A virtual pipeline data unit is a non-overlapping unit in an image, and its size is less than or equal to the size of the coding tree unit of the image. For example, when the size of the CTU is 128×128 pixels, the size of the VPDU may be less than the size of the CTU, and the size of the VPDU (e.g., 64x64 pixels) may be proportional to the buffer size in most pipeline stages of the hardware (e.g., a hardware decoder).

[0259] In some embodiments, the coding tree block may be a luma block corresponding to a target chroma block (e.g., Figure 16B Thus, the data may include a plurality of chroma samples associated with the first block and a plurality of luminance samples associated with the second block. The plurality of chroma samples associated with the first block may include: a plurality of chroma residual samples within the first block.

[0260] At step 2304, an average value of a plurality of luma samples associated with the second block may be determined. The plurality of luma samples may include a reference Figure 18-19 As an example, Figure 19As shown, the plurality of luma samples may include a plurality of reconstructed luma samples (e.g., shadow samples 1905 and padding samples 1903) on the left boundary 1901 of the second block (e.g., 1902) or on the upper boundary of the second block. It should be understood that the plurality of reconstructed luma samples may belong to an adjacent reconstructed luma block (e.g., 1904).

[0261] Method 2300 may also include determining whether a first luma sample among a plurality of luma samples associated with the second block is outside the boundaries of the image; and in response to determining that the first luma sample is outside the boundaries of the image, setting the value of the first luma sample to the value of a second luma sample among the plurality of luma samples that is within the boundaries of the image. The boundary of the image may include one of a right boundary of the image and a bottom boundary of the image. For example, it may be determined that padding sample 1903 is outside the bottom boundary of the image, and therefore the value of padding sample 1903 is set to the value of shadow sample 1905, which is the sample closest to padding sample 1903 among all samples at the bottom boundary of the image.

[0262] It should be understood that when the second block (e.g., 1902) crosses the boundary of the image, padding samples (e.g., padding sample 1903) can be created so that the number of multiple brightness samples can be a fixed number, which is usually a power of 2 to avoid division operations.

[0263] At step 2306, a chroma scaling factor for the first block may be determined based on the average value. Figure 18 As discussed, in intra-frame prediction, decoded samples in adjacent blocks of the same image can be used as reference samples to generate a predicted block. For example, the average value of the samples in the adjacent blocks can be used as the luma average value to determine the chroma scaling factor of the target block (e.g., the first block in this example), and the luma average value of the second block can be used to determine the chroma scaling factor of the first block.

[0264] At step 2308, a plurality of chroma samples associated with the first block may be processed using the chroma scaling factor. Figure 5 As discussed, multiple chroma scaling factors can be used to construct a chroma scaling factor LUT at the tile group level and applied to the reconstructed chroma residual of the target block at the decoder side. Similarly, chroma scaling factors can also be applied at the encoder side.

[0265] In some embodiments, a non-transitory computer-readable storage medium comprising instructions is also provided, and the instructions can be executed by a device for performing the above method (e.g., the disclosed encoder and decoder). Common forms of non-transitory media include, for example, floppy disks, disks, hard disks, solid-state drives, tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with hole patterns, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, buffers, registers, any other memory chips or cassettes, and network versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.

[0266] The embodiments may be further described using the following terms:

[0267] 1. A computer-implemented method for processing video content, the method comprising:

[0268] receiving data representing a first block and a second block in an image, the data comprising a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block;

[0269] determining an average of the plurality of luma samples associated with the second block;

[0270] determining a chroma scaling factor for the first block based on the average value; and

[0271] A plurality of chroma samples associated with the first block are processed using the chroma scaling factor.

[0272] 2. A method according to clause 1, wherein the plurality of luma samples associated with the second block comprises:

[0273] A plurality of reconstructed luma samples on a left boundary of the second block or on an upper boundary of the second block.

[0274] 3. The method according to clause 2, further comprising:

[0275] determining whether a first luma sample among the plurality of luma samples associated with the second block is outside a boundary of the image; and

[0276] In response to determining that the first luma sample is outside the image boundary, the value of the first luma sample is set to the value of a second luma sample of the plurality of luma samples that is within the image boundary.

[0277] 4. The method according to clause 3, further comprising:

[0278] determining whether a first luma sample among the plurality of luma samples associated with the second block is outside a boundary of the image; and

[0279] In response to determining that the first luma sample is outside the image boundary, the value of the first luma sample is set to a value of a second luma sample of the plurality of luma samples that is located on the image boundary.

[0280] 5. The method of clause 4, wherein the border of the image is one of a right border of the image and a bottom border of the image.

[0281] 6. A method according to any of clauses 1-5, wherein the second block is a coding tree block, a transform unit or a virtual pipeline data unit, wherein the size of the virtual pipeline data unit is equal to or smaller than the size of a coding tree unit of the picture.

[0282] 7. The method of clause 6, wherein the virtual pipeline data units are non-overlapping units in the image.

[0283] 8. A method according to any of clauses 1-7, wherein the plurality of chroma samples associated with the first block comprises: a plurality of chroma residual samples within the first block.

[0284] 9. The method according to any one of clauses 1 to 8, wherein the first block is a target chrominance block, and the second block is a luminance block corresponding to the target chrominance block.

[0285] 10. A video content processing system, comprising:

[0286] a memory for storing a set of instructions; and

[0287] at least one processor configured to execute the set of instructions to cause the system to:

[0288] receiving data representing a first block and a second block in an image, the data comprising a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block;

[0289] determining an average of the plurality of luma samples associated with the second block;

[0290] determining a chroma scaling factor for the first block based on the average value; and

[0291] The plurality of chroma samples associated with the first block are processed using the chroma scaling factor.

[0292] 11. The system of clause 10, wherein the plurality of luma samples associated with the second block comprises:

[0293] A plurality of reconstructed luma samples on a left boundary of the second block or on an upper boundary of the second block.

[0294] 12. The system of clause 11, wherein the at least one processor is configured to execute the set of instructions so that the system further performs:

[0295] determining whether a first luma sample among the plurality of luma samples associated with the second block is outside a boundary of the image; and

[0296] In response to determining that the first luma sample is outside the image boundary, the value of the first luma sample is set to the value of a second luma sample of the plurality of luma samples that is within the image boundary.

[0297] 13. The system of clause 12, wherein the second luminance sample is on a boundary of the image.

[0298] 14. The system of clause 13, wherein the border of the image is one of a right border of the image and a bottom border of the image.

[0299] 15. The system of any of clauses 10-14, wherein the second block is a coding tree block, a transform unit or a virtual pipeline data unit, wherein a size of the pipe virtual pipeline data unit is equal to or smaller than a size of a coding tree unit of a pipe picture.

[0300] 16. The system of clause 15, wherein the virtual pipeline data units are non-overlapping units in the pipe image.

[0301] 17. The system of any of clauses 10-16, wherein the plurality of chroma samples associated with the first block comprises: a plurality of chroma residual samples within the first block.

[0302] 18. A system according to any of clauses 10-17, wherein the first block is a target chroma block and the second block is a luma block corresponding to the target chroma block.

[0303] 19. A non-transitory computer-readable medium having stored thereon a set of instructions executable by at least one processor of a computer system to cause the computer system to perform a method for processing video content, the method comprising:

[0304] receiving data representing a first block and a second block in an image, the data comprising a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block;

[0305] determining an average of the plurality of luma samples associated with the second block;

[0306] determining a chroma scaling factor for the first block based on the average value; and

[0307] A plurality of chroma samples associated with the first block are processed using the chroma scaling factor.

[0308] 20. The non-transitory computer-readable medium of clause 19, wherein the plurality of luma samples associated with the second block comprises:

[0309] A plurality of reconstructed luma samples on a left boundary of the second block or on an upper boundary of the second block.

[0310] It should be noted that the relational terms such as "first" and "second" in this document are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "include", "have", "include" and "including" and other similar forms are equivalent in meaning and are open-ended, in that the one or more items following any of these words are not intended to be an exhaustive list of such items or items, or to be limited to the listed items.

[0311] As used herein, unless specifically stated otherwise, the term "or" encompasses all possible combinations unless not feasible. For example, if a database is stated to contain either A or B, then, unless explicitly stated otherwise or not feasible, the database may contain either A, or B, or A and B. As a second example, if a database is stated to contain either A, B, or C, then, unless explicitly stated otherwise or not feasible, the database may contain either A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.

[0312] It will be appreciated that the above embodiments may be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it may be stored in the above-mentioned computer-readable medium. The software may execute the disclosed method when executed by a processor. The computing units and other functional units described in the present disclosure may be implemented by hardware, or software, or a combination of hardware and software. It will also be appreciated by those skilled in the art that multiple of the above-mentioned modules / units may be combined into one module / unit, and each of the above-mentioned modules / units may be further divided into multiple submodules / subunits.

[0313] In the foregoing description, embodiments have been described with reference to many specific details, which may vary depending on the implementation. Certain modifications and variations may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art from consideration of the description and practice of the invention disclosed herein. The description and embodiments are to be considered exemplary only, with the true scope and spirit of the invention being indicated by the claims. The order of steps shown in the figures is also intended to be for illustrative purposes only and is not intended to be limited to any particular order of steps. Therefore, it will be understood by those skilled in the art that these steps may be performed in different orders while implementing the same method.

[0314] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications may be made to these embodiments. Therefore, although specific terms are used, they are used in a general and descriptive sense only and not for the purpose of limitation.

Claims

1. A computer-implemented method for processing video content, the method comprising: receiving data representing a first block and a second block in an image, the data comprising a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determining whether a first luma sample among the plurality of luma samples associated with the second block is outside a boundary of the image; and In response to determining that the first luma sample is outside the right border of the image, setting the value of the first luma sample to the value of a second luma sample of the plurality of luma samples that is within the right border of the image; determining an average value of the plurality of luma samples associated with the second block; determining a chroma scaling factor for the first block based on the average value; and The plurality of chroma samples associated with the first block are processed using the chroma scaling factor.

2. The method of claim 1 , wherein the plurality of luma samples associated with the second block comprises: A plurality of reconstructed luma samples on a left boundary of the second block or on an upper boundary of the second block.

3. The method according to claim 1, wherein The second block is a coding tree block, a transform unit, or a virtual pipeline data unit, wherein a size of the virtual pipeline data unit is equal to or smaller than a size of a coding tree unit of the picture.

4. The method according to claim 3, wherein The virtual pipeline data units are non-overlapping units in the image.

5. The method of claim 1 , wherein the plurality of chroma samples associated with the first block comprises: A plurality of chroma residual samples within the first block.

6. The method according to claim 1, wherein The first block is a target chrominance block, and the second block is a luminance block corresponding to the target chrominance block.

7. A video content processing system comprising: A memory, the memory being configured to store a set of instructions; and at least one processor configured to execute the set of instructions to cause the system to perform: receiving data representing a first block and a second block in an image, the data comprising a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determining whether a first luma sample among the plurality of luma samples associated with the second block is outside a boundary of the image; and In response to determining that the first luma sample is outside the right border of the image, setting the value of the first luma sample to the value of a second luma sample of the plurality of luma samples that is within the right border of the image; determining an average of the plurality of luma samples associated with the second block; determining a chroma scaling factor for the first block based on the average value; and The plurality of chroma samples associated with the first block are processed using the chroma scaling factor.

8. The system of claim 7, wherein the plurality of luma samples associated with the second block comprises: A plurality of reconstructed luma samples on a left boundary of the second block or on an upper boundary of the second block.

9. The system of claim 7, wherein the second luminance sample is on a boundary of the image. 10 . The system of claim 7 , wherein the second block is a coding tree block, a transform unit, or a virtual pipeline data unit, wherein a size of the virtual pipeline data unit is equal to or smaller than a size of a coding tree unit of a picture.

11. The system according to claim 10, wherein: The virtual pipeline data units are non-overlapping units in the image.

12. The system of claim 7, wherein the plurality of chroma samples associated with the first block comprises: A plurality of chroma residual samples within the first block.

13. The system according to claim 7, wherein: The first block is a target chrominance block, and the second block is a luminance block corresponding to the target chrominance block.

14. A non-transitory computer-readable medium having stored thereon a bitstream of a video, the bitstream being generated by a method executed by a video processing device, the method comprising: receiving data representing a first block and a second block in an image, the data comprising a plurality of chroma samples associated with the first block and a plurality of luma samples associated with the second block; determining whether a first luma sample among the plurality of luma samples associated with the second block is outside a boundary of the image; and In response to determining that the first luma sample is outside the right border of the image, setting the value of the first luma sample to the value of a second luma sample of the plurality of luma samples that is within the right border of the image; determining an average value of a plurality of luma samples associated with the second block; determining a chroma scaling factor for the first block based on the average value; and A plurality of chroma samples associated with the first block are processed using the chroma scaling factor.

15. The non-transitory computer-readable medium of claim 14, wherein the plurality of luma samples associated with the second block comprises: A plurality of reconstructed luma samples on a left boundary of the second block or on an upper boundary of the second block.