Method and system for processing video content

Feature-based video processing with chroma scaling addresses high bandwidth and storage issues in HD video surveillance by classifying content features and applying scale factors, achieving reduced bit rates and costs.

JP7825087B2Active Publication Date: 2026-03-05HFI INNOVATION INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025027351
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-12
Filing Date
2025-02-21
Publication Date
2026-03-05
Estimated Expiration
2040-02-28

AI Technical Summary

Technical Problem

High bandwidth and storage requirements for high-definition video surveillance due to high bitrates and continuous monitoring, limiting large-scale deployment.

Method used

Implementing feature-based video processing with chroma scaling, using a feature classifier to detect and classify video content features, associating different priority levels with bit rates and parameter sets for encoding, and applying luma and chroma scale factors to reduce bit rate without significant information loss.

Benefits of technology

Significantly reduces bandwidth and storage costs while maintaining video quality by customizing encoding for various surveillance scenarios, improving coding efficiency and reducing bit rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007825087000005
    Figure 0007825087000005
  • Figure 0007825087000006
    Figure 0007825087000006
  • Figure 0007825087000007
    Figure 0007825087000007
Patent Text Reader

Abstract

To provide a method and system for processing video content.SOLUTION: The method includes: receiving a chrome block and a luma block that are associated with a picture; determining luma scaling information associated with the luma block; determining a chroma scaling factor based on the luma scaling information; and processing the chroma block using the chroma scaling factor.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 813,728, filed March 4, 2019, and U.S. Provisional Patent Application No. 62 / 817,546, filed March 12, 2019, both of which are incorporated by reference in their entireties.

[0002] Technical Field FIELD OF THE DISCLOSURE

[0002] The present disclosure relates generally to video processing, and more particularly to methods and systems for in-loop mapping with chroma scaling. [Background technology]

[0003] background

[0003] Video coding systems are often used to compress digital video signals, for example, to reduce the storage space consumed or to reduce the amount of transmission bandwidth consumed associated with such signals. With the increasing popularity of high-definition (HD) video (e.g., having a resolution of 1920x1080 pixels) in various applications of video compression, such as online video streaming, video conferencing, or video surveillance, there is a constant demand to develop video coding tools that can increase the efficiency of compressing video data.

[0004] For example, video surveillance applications have become more widely used in many application scenarios (e.g., security, traffic, and environmental monitoring), resulting in a rapid increase in the number and resolution of surveillance devices. Many video surveillance application scenarios choose to provide users with HD video to capture more information, and HD video has more pixels per frame to capture such information. However, HD video bitstreams can have high bitrates, requiring high bandwidth for transmission and large storage space. For example, a surveillance video stream with an average resolution of 1920x1080 may require a bandwidth of as much as 4Mbps for real-time transmission. Furthermore, video surveillance typically involves continuous monitoring 24 hours a day, 7 days a week, which can significantly test the capacity of storage systems when storing video data. Therefore, the high bandwidth and large storage space requirements of HD video are major limitations to the large-scale deployment of HD video in video surveillance. Summary of the Invention [Means for solving the problem]

[0005] Disclosure Overview

[0005] Embodiments of the present disclosure provide a method for processing video content, which may include receiving chroma blocks and luma blocks associated with a picture, determining luma scale information associated with the luma blocks, determining a chroma scale factor based on the luma scale information, and processing the chroma blocks using the chroma scale factor.

[0006]

[0006] Embodiments of the present disclosure provide an apparatus for processing video content. The apparatus may include a memory that stores a set of instructions; and a processor, coupled to the memory, configured to execute the set of instructions to cause the apparatus to receive chroma blocks and luma blocks associated with a picture, determine luma scale information associated with the luma blocks, determine a chroma scale factor based on the luma scale information, and process the chroma blocks using the chroma scale factor.

[0007]

[0007] Embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of an apparatus to cause the apparatus to perform a method for processing video content, the method including receiving chroma blocks and luma blocks associated with a picture, determining luma scale information associated with the luma blocks, determining chroma scale factors based on the luma scale information, and processing the chroma blocks using the chroma scale factors.

[0008] BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Embodiments and various aspects of the present disclosure are set forth in the following detailed description and accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]

[0009] [Figure 1] 9 illustrates an example structure of a video sequence according to some embodiments of the present disclosure. [Figure 2A]

[0010] 1 illustrates a schematic diagram of an example encoding process according to some embodiments of the present disclosure. [Figure 2B]

[0011] 10 shows a schematic diagram of another example of an encoding process according to some embodiments of the present disclosure. [Figure 3A]

[0012] 1 illustrates a schematic diagram of an example of a decoding process according to some embodiments of the present disclosure. [Figure 3B]

[0013] 10 shows a schematic diagram of another example of a decoding process according to some embodiments of the present disclosure. [Figure 4]

[0014] 1 illustrates a block diagram of an example of a device for encoding or decoding video, according to some embodiments of the present disclosure. [Figure 5]

[0015] 1 shows a schematic diagram of an exemplary luma mapping with chroma scaling (LMCS) process, according to some embodiments of the present disclosure. [Figure 6]

[0016] 1 illustrates a tile group level syntax table for an LMCS piecewise linear model, according to some embodiments of the present disclosure. [Figure 7]

[0017] 10 illustrates another tile group level syntax table for an LMCS piecewise linear model, in accordance with some embodiments of the present disclosure. [Figure 8]

[0018] 1 is a table of a syntax structure for a coding tree unit, according to some embodiments of the present disclosure. [Figure 9]

[0019] 1 is a table of a syntax structure for dual tree splitting, according to some embodiments of the present disclosure. [Figure 10]

[0020] 10 illustrates an example of simplifying averaging of luma prediction blocks according to some embodiments of the present disclosure. [Figure 11]

[0021] 1 is a table of a syntax structure for a coding tree unit, according to some embodiments of the present disclosure. [Figure 12]

[0022] 10 is a table of syntax elements for modified signaling of LMCS piecewise linear models at the tile group level, in accordance with some embodiments of the present disclosure. [Figure 13]

[0023] 1 is a flow diagram of a method for processing video content according to some embodiments of the present disclosure. [Figure 14]

[0024] 1 is a flow diagram of a method for processing video content according to some embodiments of the present disclosure. [Figure 15]

[0025] 10 is a flow diagram of another method for processing video content according to some embodiments of the present disclosure. [Figure 16]

[0026] 10 is a flow diagram of another method for processing video content according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010] Detailed Description

[0027] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numerals in different figures, unless otherwise indicated, represent the same or similar elements. The implementations described in the following description of exemplary embodiments do not represent all implementations consistent with the present invention. Rather, they are merely examples of devices and methods consistent with aspects related to the present invention as recited in the appended claims. Unless otherwise specified, the term "or" encompasses all possible combinations unless impracticable. For example, if a component is stated to include A or B, that component can include A, or B, or A and B, unless otherwise specified or impracticable. As a second example, if a component is stated to include A, B, or C, that component can include A, or B, or C, or A and B, or A and C, or A, B, and C, unless otherwise specified or impracticable.

[0011]

[0028] A video is a set of still pictures (or "frames") arranged in chronological order to store visual information. A video capture device (e.g., a camera) can be used to capture and store the pictures in chronological order, and a video playback device (e.g., a television, computer, smartphone, tablet computer, video player, or any end-user terminal with display capabilities) can be used to display the pictures in chronological order. Furthermore, in some applications, a video capture device can transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as surveillance, conferencing, or live broadcasting.

[0012]

[0029] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission and decompressed before display. This compression and decompression can be implemented by software executed by a processor (e.g., a general-purpose computer processor) or dedicated hardware. The compression module is generally referred to as an "encoder," and the decompression module is generally referred to as a "decoder." The encoder and decoder can be collectively referred to as a "codec." The encoder and decoder can be implemented as various suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoder and decoder may include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoder and decoder may include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, and the H.26x series. In some applications, a codec may decompress video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec may be called a "transcoder."

[0013]

[0030] A video coding process can identify and retain useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, the coding process can be called "lossy." Otherwise, the coding process can be called "lossless." Most coding processes are lossy; this is a tradeoff to reduce the required storage space and transmission bandwidth.

[0014]

[0031] Useful information about the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include pixel position changes, luminance changes, or color changes, of which position changes are the most relevant. Position changes of pixels representing an object can reflect the object's movement between the reference picture and the current picture.

[0015]

[0032] A picture that is coded without reference to another picture (i.e., such a picture is its own reference picture) is called an "I-picture." A picture that is coded using a past picture as a reference picture is called a "P-picture." A picture that is coded using both past and future pictures as reference pictures (i.e., the referencing is "bidirectional") is called a "B-picture."

[0016]

[0033] As mentioned above, video surveillance using HD video faces the challenges of high bandwidth and large storage requirements. To address this challenge, the bit rate of the encoded video can be reduced. Among I-, P-, and B-pictures, I-pictures have the highest bit rate. Because the background of most surveillance video is nearly static, one way to reduce the overall bit rate of the encoded video can be to use fewer I-pictures for video encoding.

[0017]

[0034] However, since I-pictures are generally undominant in coded video, the improvement of using fewer I-pictures may be trivial. For example, in a typical video bitstream, the ratio of I-pictures, B-pictures, and P-pictures may be 1:20:9, with I-pictures accounting for less than 10% of the total bitrate. In other words, removing all I-pictures in such an example may only result in a 10% bitrate reduction.

[0018]

[0035] The present disclosure provides methods, devices, and systems for feature-based video processing for video surveillance. "Features" herein refer to content features related to video content within a picture, motion features related to motion estimation for encoding or decoding a picture, or both. For example, the content features can be pixels within one or more consecutive pictures of a video sequence, where the pixels relate to at least one of an object, a scene, or an environmental event within the picture. In another example, the motion features can include information related to the video coding process, an example of which is described in more detail below.

[0019]

[0036] In the present disclosure, a feature classifier can be used to detect and classify one or more features of pictures of a video sequence when encoding the pictures. Different classes of features can be associated with different priority levels, which are further associated with different bit rates for encoding. Different priority levels can be associated with different parameter sets for encoding, which can result in different levels of encoding quality. The higher the priority level, the higher the video quality that the associated parameter set can result in. Such feature-based video processing can significantly reduce the bit rate for surveillance video without causing significant information loss. In addition, embodiments of the present disclosure customize the corresponding relationship between priority levels and parameter sets for various application scenarios (e.g., security, traffic, environmental monitoring, etc.), thereby significantly improving video coding quality and significantly reducing bandwidth and storage costs.

[0020]

[0037] 1 illustrates an example structure of a video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be live video or captured and archived video. The video 100 can be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.

[0021]

[0038] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with more pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture for a picture (e.g., picture 104) need not be immediately before or immediately after that picture. For example, the reference picture for picture 104 may be a picture that precedes picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the example shown in FIG.

[0022]

[0039] Typically, video codecs do not encode or decode an entire picture at once because such a task is computationally complex. Rather, video codecs may divide a picture into elementary segments and encode or decode the picture segment by segment. This disclosure refers to such elementary segments as basic processing units ("BPUs"). For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) in video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) and as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). A basic processing unit can have variable sizes within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of a basic processing unit can be selected for a picture based on a balance between coding efficiency and the level of detail one wishes to preserve within the basic processing unit.

[0023]

[0040] A basic processing unit may be a logical unit that may include various types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements of the basic processing unit, where the luma and chroma components may have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit can be repeated for each of its luma and chroma components.

[0024]

[0041] Video coding involves multiple operational stages, examples of which are detailed in Figures 2A-2B and 3A-3B. For each stage, the size of the basic processing unit may still be too large to process and therefore may be further divided into segments referred to in this disclosure as "basic processing sub-units." In some embodiments, the basic processing sub-units may be referred to as "blocks" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-units may have the same or smaller size as the basic processing units. Like the basic processing units, the basic processing sub-units are also logical units that may contain various types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit can be repeated on each of its luma and chroma components. It should be noted that such division can be performed on further levels depending on the processing needs. It should also be noted that different stages can divide the basic processing unit using different schemes.

[0025]

[0042] For example, in a mode decision stage (one example of which is detailed in FIG. 2B ), the encoder may decide which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, which may be too large for such a decision to be made. The encoder may divide the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing sub-unit.

[0026]

[0043] In another example, in the prediction stage (one example of which is detailed in FIG. 2A), the encoder can perform prediction operations at the level of elementary processing sub-units (e.g., CUs). However, in some cases, elementary processing sub-units may still be too large to process. The encoder can further divide the elementary processing sub-units into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform prediction operations at that level.

[0027]

[0044] In another example, in the transform stage (one example of which is detailed in FIG. 2A ), the encoder can perform a transform operation on a residual elementary processing sub-unit (e.g., a CU). However, in some cases, the elementary processing sub-unit may still be too large to process. The encoder can further divide the elementary processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transform operation at that level. It should be noted that the division scheme of the same elementary processing sub-unit may be different between the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0028]

[0045] In structure 110 of Figure 1, basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown by dotted lines. Different basic processing units of the same picture can be divided into basic processing sub-units in different ways.

[0029]

[0046] In some implementations, to provide parallel processing and error resilience for video encoding and decoding, a picture can be divided into regions for processing, allowing the encoding or decoding process to be independent of information from any other region of the picture. In other words, each region of a picture can be processed independently. This allows a codec to process different regions of a picture in parallel, thereby increasing coding efficiency. Furthermore, if data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thereby providing error resilience. Some video coding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different pictures in video sequence 100 may have different partitioning schemes for dividing the picture into regions.

[0030]

[0047] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.

[0031]

[0048] FIG. 2A shows a schematic diagram of an example encoding process 200A according to some embodiments of the present disclosure. An encoder may follow process 200A to encode a video sequence 202 into a video bitstream 228. Similar to video sequence 100 of FIG. 1, video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to structure 110 of FIG. 1, each original picture of video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of the basic processing units for each original picture of video sequence 202. For example, the encoder may perform process 200A in an iterative manner, where the encoder may encode a basic processing unit within one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions (e.g., regions 114-118) of each original picture of video sequence 202.

[0032]

[0049] In FIG. 2A , an encoder may feed a basic processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder may feed the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may feed the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the “forward path.” During process 200A, the encoder may feed quantized transform coefficients 216 after quantization stage 214 to inverse quantization stage 218 and inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to a predicted BPU 208 to generate a prediction reference 224 that is used in prediction stage 204 of the next iteration of process 200A. The components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0033]

[0050] The encoder may perform process 200A iteratively to encode each original BPU of the original picture (in the forward path) and generate a predicted reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.

[0034]

[0051] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or any action in any manner to input data.

[0035]

[0052] In the prediction stage 204 of the current iteration, the encoder receives the original BPU and a prediction reference 224 and may perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of a previous iteration of the process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting prediction data 206 from the prediction data 206 and the prediction reference 224 that can be used to reconstruct the original BPU as a predicted BPU 208.

[0036]

[0053] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, the encoder may subtract predicted BPU 208 from the original BPU after generating it to generate residual BPU 210. For example, the encoder may subtract the values ​​(e.g., grayscale or RGB values) of pixels of predicted BPU 208 from the values ​​of corresponding pixels of the original BPU. As a result of such subtraction between corresponding pixels of the original BPU and predicted BPU 208, each pixel of residual BPU 210 may have a residual value. Compared to the original BPU, predicted data 206 and residual BPU 210 may have fewer bits, which can be used to reconstruct the original BPU without significant loss of quality.

[0037]

[0054] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variation frequency (e.g., luminance variation frequency) component of the residual BPU 210. None of the basis patterns can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, such decomposition allows the variation of the residual BPU 210 to be decomposed in the frequency domain. Such decomposition is analogous to the discrete Fourier transform of a function, with the basis patterns being analogous to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients being analogous to the coefficients associated with the basis functions.

[0038]

[0055] Different transform algorithms can use different basis patterns. Different transform algorithms can be used in transform stage 212, such as a discrete cosine transform, a discrete sine transform, etc. The transform in transform stage 212 is reversible. That is, the encoder can reconstruct residual BPU 210 by inverting the transform (referred to as an "inverse transform"). For example, to reconstruct a pixel of residual BPU 210, the inverse transform can multiply the value of the corresponding pixel in the basis pattern by the associated coefficient and add the products to obtain a weighted sum. In video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients from which the decoder can reconstruct residual BPU 210 without receiving the basis pattern from the encoder. Although the transform coefficients may have fewer bits compared to residual BPU 210, they can still be used to reconstruct residual BPU 210 without significant loss of quality. Thus, residual BPU 210 is further compressed.

[0039]

[0056] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different fluctuation frequencies (e.g., luminance fluctuation frequencies). Because the human eye is generally good at recognizing low-frequency fluctuations, the encoder can ignore high-frequency fluctuation information without causing significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest neighbor. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process is also reversible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients by the inverse operation of quantization (referred to as "dequantization").

[0040]

[0057] Quantization stage 214 may be lossy because the encoder ignores the remainder of such a division in a rounding operation. Typically, quantization stage 214 may contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0041]

[0058] In binary coding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary coding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in binary coding stage 226, such as a prediction mode used in prediction stage 204, parameters of the prediction operation, the type of transform in transform stage 212, parameters of the quantization process (e.g., quantization parameters), and encoder control parameters (e.g., bitrate control parameters). The encoder may generate a video bitstream 228 using the output data of binary coding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0042]

[0059] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to a predicted BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0043]

[0060] It should be noted that other variations of process 200A can be used to encode video sequence 202. In some embodiments, an encoder can perform the stages of process 200A in a different order. In some embodiments, one or more stages of process 200A can be combined into a single stage. In some embodiments, a single stage of process 200A can be separated into multiple stages. For example, transform stage 212 and quantization stage 214 can be combined into a single stage. In some embodiments, process 200A can include additional stages. In some embodiments, process 200A can omit one or more stages in FIG. 2A .

[0044]

[0061] 2B shows a schematic diagram of another example encoding process 200B according to some embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder conforming to a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B further includes a mode decision stage 230 and separates prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0045]

[0062] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels of one or more neighboring BPUs already coded in the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions of one or more neighboring pictures already coded to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0046]

[0063] Referring to process 200B, in the forward path, the encoder performs prediction operations in a spatial prediction step 2042 and a temporal prediction step 2044. For example, in the spatial prediction step 2042, the encoder may perform intra prediction. With respect to an original BPU of a picture being coded, the prediction reference 224 may include one or more neighboring BPUs coded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating, for each pixel of the predicted BPU 208, the value of the corresponding pixel. The neighboring BPUs used for extrapolation may be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., bottom-left, bottom-right, top-left, or top-right of the original BPU), or any direction specified within the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the positions (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientation of the neighboring BPUs used relative to the original BPU, etc.

[0047]

[0064] In another example, the encoder may perform inter-prediction in the temporal prediction stage 2044. With respect to the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, a reference picture may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. Once all the reconstructed BPUs of the same picture are generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference picture. The position of the search window in the reference picture may be determined based on the position of the original BPU in the current picture. For example, the search window may be centered in the reference picture at a location having the same coordinates as the original BPU in the current picture and may extend over a predetermined distance. When the encoder identifies a region within the search window that is similar to the original BPU (e.g., by using a pel recursion algorithm, a block matching algorithm, etc.), the encoder can determine that region as a matching region. The matching region may have different dimensions (e.g., smaller, equal, larger, or different shape) than the original BPU. Because the reference picture and the current picture are separated in time in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the position of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." If multiple reference pictures are used (e.g., picture 106 in FIG. 1), the encoder can find the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching region in each matching reference picture.

[0048]

[0065] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. In inter-prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, etc.

[0049]

[0066] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift the matching region of the reference picture according to the motion vector, within which the encoder may predict the original BPU of the current picture. When multiple reference pictures are used (e.g., picture 106 of FIG. 1), the encoder may shift the matching region of the reference picture according to the respective motion vectors and average the pixel values ​​of the matching region. In some embodiments, if the encoder assigns weights to the pixel values ​​of the matching region of each matching reference picture, the encoder may perform a weighted sum of the pixel values ​​of the shifted matching region.

[0050]

[0067] In some embodiments, inter-prediction may be unidirectional or bidirectional. Unidirectional inter-prediction may use one or more reference pictures that are in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter-prediction may use one or more reference pictures that are in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture in which the reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0051]

[0068] Continuing with the forward path of process 200B, after spatial prediction step 2042 and temporal prediction step 2044, in mode decision step 230, the encoder may select a prediction mode (e.g., one of intra-prediction or inter-prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique, in which the encoder may select a prediction mode to minimize the value of a cost function depending on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.

[0052]

[0069] In the reconstruction path of process 200B, if an intra-prediction mode is selected in the forward path, after generating a prediction reference 224 (e.g., a current BPU that has been coded and reconstructed in a current picture), the encoder can directly feed the prediction reference 224 to a spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If an inter-prediction mode is selected in the forward path, after generating a prediction reference 224 (e.g., a current picture in which all BPUs have been coded and reconstructed), the encoder can feed the prediction reference 224 to a loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) caused by inter-prediction. The encoder can apply various loop filter techniques in the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop filtered reference pictures may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., for use as inter-predicted reference pictures for future pictures in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in a binary coding stage 226.

[0053]

[0070] FIG. 3A illustrates a schematic diagram of an example decoding process 300A according to some embodiments of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A of FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 of FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B of FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, a decoder may perform process 300A in an iterative manner, and the decoder may decode a basic processing unit within one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for a region of each picture (eg, regions 114-118) encoded in video bitstream 228.

[0054]

[0071] 3A , a decoder may feed a portion of a video bitstream 228 associated with a basic processing unit of a coded picture (referred to as a “coded BPU”) to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may feed the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may feed the prediction data 206 to a prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may feed the predicted reference 224 to the prediction stage 204 for performing the prediction operation in the next iteration of the process 300A.

[0055]

[0072] The decoder may iteratively perform process 300A to decode each coded BPU of the coded picture and generate a predicted reference 224 for coding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder may output the picture to the video stream 304 for display and proceed to decode the next coded picture in the video bitstream 228.

[0056]

[0073] In binary decoding stage 302, the decoder may perform an inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), etc. In some embodiments, if video bitstream 228 is transmitted in packets over a network, the decoder may depacketize video bitstream 228 before feeding it to binary decoding stage 302.

[0057]

[0074] 3B shows a schematic diagram of another example decoding process 300B according to some embodiments of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B further divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.

[0058]

[0075] In process 300B, for a coded basic processing unit (referred to as a "current BPU") of a coded picture being decoded (referred to as a "current picture"), prediction data 206 decoded by the decoder from binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to code the current BPU. For example, if intra prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, sizes of the neighboring BPUs, parameters of extrapolation, directions of the neighboring BPUs relative to the original BPU, etc. In another example, if inter prediction was used by the encoder to code the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. Parameters for inter-prediction operations may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, the locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, etc.

[0059]

[0076] Based on the prediction mode indicator, the decoder can determine whether to perform spatial prediction (e.g., intra-prediction) in spatial prediction step 2042 or temporal prediction (e.g., inter-prediction) in temporal prediction step 2044. Details of performing such spatial or temporal prediction are shown in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder can generate a predicted BPU 208. As described in FIG. 3A, the decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.

[0060]

[0077] In process 300B, the decoder may feed the predicted reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation within the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., to extrapolate the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B . The loop-filtered reference pictures may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as inter-predicted reference pictures for future coded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator in the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include loop filter parameters (e.g., loop filter strength).

[0061]

[0078] FIG. 4 is a block diagram of an example device 400 for encoding or decoding video, in accordance with some embodiments of the present disclosure. As shown in FIG. 4, device 400 may include a processor 402. When processor 402 executes the instructions described herein, device 400 may be a dedicated machine for encoding or decoding video. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any combination of any number of central processing units (“CPUs”), graphics processing units (“GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), etc. In some embodiments, processor 402 may be a set of processors grouped together as a single logical entity. For example, as shown in Figure 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0062]

[0079] Device 400 may also include memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to operate on or process the data for processing. Memory 404 may include high-speed random access storage or non-volatile storage. In some embodiments, memory 404 may include any combination of any number of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a collection of memories (not shown in FIG. 4) grouped together as a single logical entity.

[0063]

[0080] Bus 410, such as an internal bus (e.g., a CPU memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like, may be a communication device that transfers data between components within device 400.

[0064]

[0081] For ease of explanation and without ambiguity, this disclosure will collectively refer to the processor 402 and other data processing circuitry as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuitry may be a single, independent module, or may be fully or partially combined within any other component of the device 400.

[0065]

[0082] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communications network, etc.) In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.

[0066]

[0083] In some embodiments, device 400 may optionally further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, peripherals may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or input interface coupled to a video archive), etc.

[0067]

[0084] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) can be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the stages of process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).

[0068]

[0085] 5 shows a schematic diagram of an exemplary luma mapping with chroma scaling (LMCS) process 500 according to some embodiments of the present disclosure. For example, the process 500 may be used by a decoder that complies with a hybrid video coding standard (e.g., the H.26x series). The LMCS is a new processing block that is applied before the loop filter 232 in FIG. 2B. The LMCS may also be referred to as a reshaper.

[0069]

[0086] The LMCS process 500 may include an in-loop mapping of luma component values ​​based on an adaptive piecewise linear model and luma-dependent chroma residual scaling of chroma components.

[0070]

[0087] 5, in-loop mapping of luma component values ​​based on an adaptive piecewise linear model may include a forward mapping stage 518 and an inverse mapping stage 508. Luma-dependent chroma residual scaling of chroma components may include chroma scaling 520.

[0071]

[0088] The sample values ​​before mapping or after de-mapping can be referred to as samples in the original domain, and the sample values ​​after mapping and before de-mapping can be referred to as samples in the mapped domain. When LMCS is enabled, some stages in the process 500 can be performed in the mapped domain rather than the original domain. It will be appreciated that the forward mapping stage 518 and the de-mapping stage 508 can be enabled / disabled at the sequence level using the SPS flag.

[0072]

[0089] As shown in Figure 5, Q -1 &T -1 Steps 504, reconstruction 506, and intra prediction 508 are performed within the map domain. -1 &T -1 Stage 504 may include inverse quantization and inverse transform, reconstruction 506 may include summing luma prediction and luma residual, and intra prediction 508 may include luma intra prediction.

[0073]

[0090] The loop filter 510, motion compensation stages 516 and 530, intra prediction stage 528, reconstruction stage 522, and decoded picture buffers (DPBs) 512 and 526 are performed in the original (i.e., unmapped) domain. In some embodiments, the loop filter 510 may include deblocking, an adaptive loop filter (ALF), and a sample adaptive offset (SAO), the reconstruction stage 522 may include chroma prediction and addition of chroma residual, and the DPBs 512 and 526 may store decoded pictures as reference pictures.

[0074]

[0091] In some embodiments, luma mapping according to a piecewise linear model may be applied.

[0075]

[0092] In-loop mapping of the luma component can improve compression efficiency by adjusting the signal statistics of the input video through a redistribution of codewords across the dynamic range. Luma mapping can be performed by a forward mapping function "FwdMap" and a corresponding inverse mapping function "InvMap." The "FwdMap" function is signaled using a piecewise linear model with 16 equal sections. The "InvMap" function does not need to be signaled, but is instead derived from the "FwdMap" function.

[0076]

[0093] The signaling of the piecewise linear model is shown in Table 1 of Figure 6 and Table 2 of Figure 7. Table 1 of Figure 6 shows the tile group header syntax structure. As shown in Figure 6, a reshaper model parameters present flag is signaled to indicate whether there is a luma mapping model in the current tile group. If there is a luma mapping model in the current tile group, the corresponding piecewise linear model parameters can be signaled in tile_group_reshaper_model() using the syntax elements shown in Table 2 of Figure 7. The piecewise linear model divides the dynamic range of the input signal into 16 equal partitions. For each of the 16 equal partitions, the linear mapping parameters of the partition are represented using the number of codewords assigned to the partition. Take a 10-bit input as an example. Each of the 16 partitions may have 64 codewords assigned to that partition by default. The signaled number of codewords can be used to calculate a scale factor and adjust the mapping function for that partition accordingly. Table 2 in Figure 7 also inclusively defines the minimum index "reshaper_model_min_bin_idx" and maximum index "reshaper_model_max_bin_idx" for which the number of codewords is signaled. If the partition index is smaller than reshaper_model_min_bin_idx or larger than reshaper_model_max_bin_idx, the number of codewords for that partition is not signaled and is inferred to be zero (i.e., no codewords are assigned to that partition and no mapping / scaling is applied).

[0077]

[0094] After tile_group_reshaper_model() is signaled, another reshaper enable flag "tile_group_reshaper_enable_flag" is signaled at the tile group header level to indicate whether the LMCS process shown in Figure 8 is applied to the current tile group. If the reshaper is enabled for the current tile group and the current tile group does not use dual tree partitioning, an additional chroma scaling enable flag is signaled to indicate whether chroma scaling is enabled for the current tile group. Dual tree partitioning can also be called chroma separate tree.

[0078]

[0095] The piecewise linear model can be constructed based on the signaled syntax elements in Table 2 of Figure 7 as follows: Each ith piece, i=0,1,...,15, of the "FwdMap" piecewise linear model is defined by two input pivot points InputPivot[] and two output (mapped) pivot points MappedPivot[]. InputPivot[] and MappedPivot[] are calculated based on the signaled syntax as follows (without loss of generality, we assume the bit depth of the input video is 10 bits): 1)OrgCW=64 2) For i=0:16, InputPivot[i]=i*OrgCW 3)i=reshaper_model_min_bin_idx: in reshaper_model_max_bin_idx SignaledCW[i]=OrgCW+(1¬2*reshape_model_bin_delta_sign_CW[i])*reshape_model_bin_delta_abs_CW[i]; 4) For i = 0:16, calculate MappedPivot[i] as follows: MappedPivot[0]=0; (i=0; i<16; i++) MappedPivot[i+1]=MappedPivot[i]+SignaledCW[i]

[0079]

[0096] The inverse mapping function "InvMap" can also be defined by InputPivot[] and MappedPivot[]. Unlike "FwdMap", in "InvMap" piecewise linear models, the two input pivot points for each piece can be defined by MappedPivot[] and the two output pivot points can be defined by InputPivot[], which is the opposite of "FwdMap". In this way, the input of "FwdMap" is divided into equal pieces, but the input of "InvMap" is not guaranteed to be divided into equal pieces.

[0080]

[0097] As shown in Figure 5, for inter-coded blocks, motion compensation prediction can be performed in the map domain. In other words, after motion compensation prediction 516, Y is calculated based on the reference signal in the DPB. pred and apply a “FwdMap” function 518 to map the luma prediction blocks in the original region to the map region (Y′ pred =FwdMap(Y pred )) For intra-coded blocks, the "FwdMap" function is not applied since the reference samples used in intra prediction are already in the map region. After the reconstructed block 506, Y r An "InvMap" function 508 can be applied to convert the reconstructed luma values ​​in the map domain back to reconstructed luma values ​​in the original domain (

number

[0081]

[0098] The luma mapping process (forward or inverse mapping) can be implemented using a look-up table (LUT) or using on-the-fly calculations. If a LUT is used, the "FwdMapLUT[]" and "InvMapLUT[]" tables can be pre-calculated and pre-stored for use at the tile group level, and the forward and inverse mappings can be simply calculated as FwdMap(Y pred )=FwdMapLUT[Y pred ], and InvMap(Y r )=InvMapLUT[Y r ], or in-place calculations can be used. Take the forward mapping function "FwdMap" as an example. To locate the partition to which a luma sample belongs, the sample value can be right-shifted by 6 bits (corresponding to 16 equal partitions, assuming 10-bit video) to get the partition index. The linear model parameters for that partition are then taken and applied in-place to calculate the mapped luma value. The FwdMap function is evaluated as follows: Y'pred=FwdMap(Y pred )=((b2-b1) / (a2-a1))*(Y pred -a1)+b1 where "i" is the partition index, a1 is InputPivot[i], a2 is InputPivot[i+1], b1 is MappedPivot[i], and b2 is MappedPivot[i+1].

[0082]

[0099] The "InvMap" function can be computed on the fly in a similar manner, except that the partitions in the map region are not guaranteed to be of equal size, so a conditional check must be applied instead of a simple right bit shift when finding the partition to which a sample value belongs.

[0083]

[0100] In some embodiments, luma-dependent chroma residual scaling can be performed.

[0084]

[0101] Chroma residual scaling is designed to compensate for the interaction between a luma signal and its corresponding chroma signal. Whether chroma residual scaling is enabled is also signaled at the tile group level. As shown in Table 1 of Figure 6, when luma mapping is enabled and dual-tree partitioning is not applied to the current tile group, an additional flag (e.g., tile_group_reshaper_chroma_residual_scale_flag) is signaled to indicate whether luma-dependent chroma residual scaling is enabled. When luma mapping is not used or dual-tree partitioning is used within the current tile group, luma-dependent chroma residual scaling is automatically disabled. Furthermore, luma-dependent chroma residual scaling can be disabled for chroma blocks whose area is 4 or less.

[0085]

[0102] Chroma residual scaling depends on the average value of the corresponding luma prediction block (for both intra-coded and inter-coded blocks). The average of the luma prediction block, avgY', can be calculated as follows:

number

[0086]

[0103] C ScaleInv The value of is calculated using the following steps: 1) Index Y of the piecewise linear model to which avgY' belongs Idx based on the InvMap function. 2) C ScaleInv =cScaleInv[Y Idx ] holds, where cScaleInv[] is a pre-computed LUT with 16 segments.

[0087]

[0104] In the current LMCS method in VTM4, the pre-computed LUT cScaleInv[i] (where i is in the range 0 to 15) is derived based on the values ​​of the 64-entry static LUT ChromaResidualScaleLut and SignaledCW[i] as follows: ChromaResidualScaleLut

[64] ={16384, 16384, 16384, 16384, 16384, 16384, 16384, 8192, 8192, 8192, 8192, 5461, 5461, 5461, 5461, 4096, 4096, 4096, 4096, 3277, 3277, 3277, 3277, 2731, 2731, 2731, 2731, 2341, 2341, 2341, 2048, 2048, 2048, 1820, 1820, 1820, 1638, 1638, 1638, 1638, 1489, 1489, 1489, 1489, 1365, 1365, 1365, 1260, 1260, 1260, 1260, 1170, 1170, 1170, 1092, 1092, 1092, 1024, 1024, 1024, 1024}; shiftC=11 -If (SignaledCW[i]==0) holds, cScaleInv[i]=(1< <shiftC) - Otherwise cScaleInv[i]=ChromaResidualScaleLut[(SignaledCW[i]>>1)-1]

[0088]

[0105] The static table ChromaResidualScaleLut[] contains 64 entries, and SignaledCW[] is in the range [0,128] (assuming 10-bit input). Therefore, we use a division by 2 (e.g., a right shift by 1) to construct the chroma scale factor LUT cScaleInv[]. The chroma scale factor LUT cScaleInv[] can contain multiple chroma scale factors. The LUT cScaleInv[] is constructed at the tile group level.

[0089]

[0106] If the current block is coded using intra, CIIP, or intra block copy (IBC, also known as current picture reference or CPR) mode, avgY' is calculated as the average of the intra, CIIP, or IBC predicted luma values. Otherwise, avgY' is calculated as the average of the forward-mapped inter-predicted luma values ​​(i.e., Y' in Figure 5). pred ) is calculated as the average of the C ScaleInv is a constant value for all chroma blocks. C ScaleInv chroma residual scaling is applied at the decoder side as follows:

number

[0090]

[0107] however

number

[0091]

[0108] In some embodiments, a dual tree split may be performed.

[0092]

[0109] In Draft 4 of VVC, the coding tree method supports the ability for luma and chroma to have separate block tree partitions. This is also called dual-tree partitioning. The signaling of dual-tree partitioning is shown in Table 3 of Figure 8 and Table 4 of Figure 9. When the sequence-level control flag "qtbtt_dual_tree_intra_flag" signaled in the SPS is turned on and the current tile group is intra-coded, block partition information can be signaled separately, first for luma and then for chroma. Dual-tree partitioning is not allowed for inter-coded tile groups (P and B tile groups). When the separate block tree mode is applied, the luma coding tree block (CTB) is partitioned into CUs by one coding tree structure, and the chroma CTB is partitioned into chroma CUs by another coding tree structure, as shown in Table 4 of Figure 9.

[0093]

[0110] If luma and chroma are allowed to have different partitions, problems can arise for coding tools with dependencies between various color components. For example, in the case of LMCS, the average value of the corresponding luma blocks is used to find the scale factor to apply to the current block. If dual trees are used, this can incur latency for the entire CTU. For example, if the luma blocks of a CTU are partitioned once vertically and the chroma blocks of the CTU are partitioned once horizontally, both luma blocks of the CTU are decoded before the first chroma block of the CTU can be decoded (so that the average value needed to calculate the chroma scale factor can be calculated). In VVC, a CTU can be as large as 128x128 in units of luma samples. Such large latency can be very problematic for hardware decoder pipeline designs. Therefore, VVC Draft 4 can prohibit the combination of dual-tree partitioning and luma-dependent chroma scaling. If dual-tree partitioning is enabled for the current tile group, chroma scaling can be forced off. Note that the luma mapping part of the LMCS is still allowed in the dual-tree case, since it only affects the luma component and does not have the issue of dependencies across color components. Another example of a coding tool that relies on dependencies between color components to achieve better coding efficiency is called a cross-component linear model (CCLM).

[0094]

[0111] Therefore, the derivation of the tile group level chroma scale factor LUT, cScaleInv[], cannot be easily extended. The derivation process currently relies on a constant chroma LUT, ChromaResidualScaleLut, with 64 entries. For 10-bit video with 16 divisions, an additional step of divide by 2 must be applied. If the number of divisions changes, for example, if 8 divisions are used instead of 16 divisions, the derivation process must be modified to apply a divide by 4 instead of a divide by 2. This additional step not only causes a loss of precision, but is also inelegant and unnecessary.

[0095]

[0112] Additionally, the partition index Y of the current chroma block used to obtain the chroma scale factor Idx To calculate the luma mean, the average value of all luma blocks can be used. This is undesirable and often unnecessary. Consider a maximum CTU size of 128x128. In this case, the average luma value is calculated based on 16,384 (128x128) luma samples, and such calculations are complex. Furthermore, if a 128x128 luma block partition is selected by the encoder, the block is likely to contain homogeneous content. Therefore, a subset of the luma samples in the block may be sufficient to calculate the luma mean.

[0096]

[0113] In dual-tree partitioning, chroma scaling may be turned off to avoid potential pipeline issues in the hardware decoder. However, this dependency can be avoided if explicit signaling is used to indicate the chroma scale factor to be applied, instead of using the corresponding luma samples to derive the chroma scale factor to be applied. Enabling chroma scaling within intra-coded tile groups may further improve coding efficiency.

[0097]

[0114] The signaling of piecewise linear parameters can be further improved. Currently, delta codeword values ​​are signaled for each of the 16 segments. It is recognized that often only a limited number of different codewords are used for the 16 segments. Thus, the signaling overhead can be further reduced.

[0098]

[0115] SUMMARY OF THE INVENTION An embodiment of the present disclosure provides a method for processing video content by removing a chroma scaling LUT.

[0099]

[0116] As described above, it may be difficult to expand the 64-entry chroma LUT, which can be a problem when other partitioning linear models are used (e.g., 8-partition, 4-partition, 64-partition, etc.). To achieve the same coding efficiency, the chroma scale factor can be set to the same as the corresponding luma scale factor of the partition, so such expansion is not necessary. In some embodiments of the present disclosure, the chroma scale factor (chroma_scaling) can be determined based on the partition index "Y Idx " of the current chroma block as follows. · When Y Idx >reshaper_model_max_bin_idx, Y Idx <reshaper_model_min_bin_idx or SignaledCW[Y Idx =0 holds, set chroma_scaling to the default and chroma_scaling = 1.0. · Otherwise, set chroma_scaling to SignaledCW[Y Idx / OrgCW.

[0100]

[0117] When chroma_scaling = 1.0, scaling is not applied.

[0101]

[0118] The chroma scale factor determined above may have fractional precision. It will be understood that fixed-point approximation can be applied to avoid dependencies on the hardware / software platform. Furthermore, inverse chroma scaling can be performed on the decoder side. Therefore, division can be implemented by fixed-point arithmetic using multiplication followed by a right shift. The inverse chroma scale factor "inverse_chroma_scaling[]" at fixed-point precision can be determined as follows based on the number of bits within the fixed-point approximation "CSCALE_FP_PREC". inverse_chroma_scaling[Y Idx]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(SignaledCW[Y Idx ]>>1)) / SignaledCW[Y Idx ] where luma_bit_depth is the luma bit depth, and TOTAL_NUMBER_PIECES is the total number of pieces in the piecewise linear model, which is set to 16 in VVC Draft 4. It will be understood that the value of "inverse_chroma_scaling[]" may only need to be calculated once per tile group, and the above division is an integer division operation.

[0102]

[0119] Further quantization can be applied to determine the chroma and inverse scale factors. For example, for every even (2*m) value of "SignaledCW", an inverse chroma scale factor can be calculated, and odd (2*m+1) values ​​of "SignaledCW" reuse the chroma scale factor of the adjacent even-valued scale factor. In other words, the following can be used: for(i=reshaper_model_min_bin_idx; i<=reshaper_model_max_bin_idx; i++) { tempCW=SignaledCW[i]>>1)<<1; inverse_chroma_scaling[i]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(tempCW>>1)) / tempCW; }

[0103]

[0120] The quantization of the chroma scale factors can be further generalized. For example, an inverse chroma scale factor "inverse_chroma_scaling[]" can be calculated for every nth value of "SignaledCW," with all other adjacent values ​​sharing the same chroma scale factor. For example, "n" can be set to 4. Thus, every four adjacent codeword values ​​can share the same inverse chroma scale factor value. In some embodiments, the value of "n" can be a power of 2, which allows for the use of shifts to calculate divisions. Expressing the value of log2(n) as LOG2_n, the above equation "tempCW=SignaledCW[i]>>1)<<1" can be modified as follows: tempCW=SignaledCW[i]>>LOG2_n)< <LOG2_n

[0104]

[0121] In some embodiments, the value of LOG2_n may be a function of the number of pieces used in the piecewise linear model. If fewer pieces are used, it may be beneficial to use a larger LOG2_n. For example, if TOTAL_NUMBER_PIECES is 16 or less, LOG2_n may be set to 1 + (4 - log2(TOTAL_NUMBER_PIECES)). If TOTAL_NUMBER_PIECES is greater than 16, LOG2_n may be set to 0.

[0105]

[0122] SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a method for processing video content by simplifying the averaging of luma prediction blocks.

[0106]

[0123] As discussed above, the current chroma block's division index "Y Idx To determine ", the average value of the corresponding luma block can be used. However, for large block sizes, the averaging process may involve a large number of luma samples. In the worst case, 128x128 luma samples may be involved in the averaging process.

[0107]

[0124] Embodiments of the present disclosure provide a simplified averaging process to reduce the worst case to using only NxN luma samples (N is a power of 2).

[0108]

[0125] In some embodiments, if both dimensions of a 2D luma block are less than or equal to a preset threshold M (in other words, at least one of the two dimensions is greater than M), then "downsampling" can be applied to use only M positions within that dimension. Without loss of generality, take the horizontal dimension as an example. If the width is greater than M, then only samples at position x, where x=i×(width>>log2(M)), i=0,...M-1, are used for averaging.

[0109]

[0126] 10 shows an example of applying the proposed simplification to calculate the average of a 16x8 luma block. In this example, M is set to 4, and only 16 luma samples in the block (shaded samples) are used for averaging. It will be understood that the preset threshold M is not limited to 4, and M can be set to any value that is a power of 2. For example, the preset threshold M can be 1, 2, 4, 8, etc.

[0110]

[0127] In some embodiments, the horizontal and vertical dimensions of the luma blocks may have different preset thresholds M. In other words, the worst case for the averaging operation may use M1xM2 samples.

[0111]

[0128] In some embodiments, the number of samples in the averaging process can be limited without considering the dimensions. For example, a maximum of 16 samples can be used, and the samples can be distributed in the horizontal or vertical dimensions in a 1x16, 16x1, 2x8, 8x2, or 4x4 format, and any format that fits the shape of the current block can be chosen. For example, if the block is vertical, a matrix of 2x8 samples can be used, if the block is horizontal, a matrix of 8x2 samples can be used, and if the block is square, a matrix of 4x4 samples can be used.

[0112]

[0129] It will be appreciated that if a large block size is chosen, the content within the block will tend to be more homogeneous, so although the above simplifications may cause differences between the average value and the true average of all luma blocks, such differences may be small.

[0113]

[0130] Furthermore, decoder-side motion vector refinement (DMVR) requires the decoder to perform motion estimation to derive motion vectors before motion compensation can be applied. Therefore, DMVR mode can be complicated within the VVC standard, especially for the decoder. Bidirectional optical flow (BDOF) mode within the VVC standard can further complicate this situation, as BDOF is an additional sequential process that must be applied after DMVR to obtain the luma prediction block. Because chroma scaling requires the average value of the corresponding luma prediction block, DMVR and BDOF can be applied before the average value can be calculated.

[0114]

[0131] To solve this latency issue, some embodiments of the present disclosure use luma prediction blocks before DMVR and BDOF to calculate an average luma value, and then use the average luma value to obtain a chroma scale factor, which allows chroma scaling to be applied in parallel with the DMVR and BDOF process, thus significantly reducing latency.

[0115]

[0132] Consistent with this disclosure, variations in latency reduction are contemplated. In some embodiments, this latency reduction can be combined with the simplified averaging process described above, which uses only a portion of the luma prediction block to calculate the average luma value. In some embodiments, the luma prediction block can be used after the DMVR process and before the BDOF process to calculate the average luma value. The average luma value is then used to obtain the chroma scale factor. This design allows chroma scaling to be applied in parallel with the BDOF process while maintaining accuracy in determining the chroma scale factor. Because the DMVR process may refine motion vectors, using prediction samples with refined motion vectors after the DMVR process may be more accurate than using prediction samples with motion vectors before the DMVR process.

[0116]

[0133] Furthermore, in the VVC standard, the CU syntax structure "coding_unit()" includes a syntax element "cu_cbf" to indicate whether there are any non-zero residual coefficients in the current CU. At the TU level, the TU syntax structure "transform_unit()" includes syntax elements "tu_cbf_cb" and "tu_cbf_cr" to indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the current TU. Conventionally, in VVC Draft 4, when chroma scaling is enabled at the tile group level, averaging of the corresponding luma block is always invoked.

[0117]

[0134] Embodiments of the present disclosure further provide a method for processing video content by bypassing the luma averaging process. Consistent with the disclosed embodiments, a chroma scaling process is applied to residual chroma coefficients, so that the luma averaging process can be bypassed if there are no non-zero chroma coefficients. This can be determined based on the following conditions: Condition 1: cu_cbf is equal to 0 Condition 2: tu_cbf_cr and tu_cbf_cb are both equal to 0

[0118]

[0135] As discussed above, "cu_cbf" can indicate whether there are any non-zero residual coefficients in the current CU, and "tu_cbf_cb" and "tu_cbf_cr" can indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the current TU. If condition 1 or condition 2 is met, the luma averaging process can be bypassed.

[0119]

[0136] In some embodiments, only NxN samples of the predicted block are used to derive the average value, which simplifies the averaging process. For example, if N is equal to 1, only the top-left sample of the predicted block is used. However, this simplified averaging process using the predicted block still requires the generation of the predicted block, thereby incurring latency.

[0120]

[0137] In some embodiments, reference luma samples can be used directly to generate chroma scale factors. This allows the decoder to derive scale factors in parallel with the luma prediction process, thereby reducing latency. In the following, intra-prediction and inter-prediction using reference luma samples are described separately.

[0121]

[0138] In exemplary intra prediction, decoded neighboring samples within the same picture can be used as reference samples for generating the predicted block. These reference samples may include, for example, a sample at the top of the current block, a sample to the left of the current block, or a sample at the top left of the current block. An average of these reference samples can be used to derive the chroma scale factor. In some embodiments, an average of a subset of these reference samples can be used. For example, only the K reference samples (e.g., K=3) closest to the top left position of the current block are averaged.

[0122]

[0139] In exemplary inter-prediction, reference samples from temporal reference pictures can be used to generate the prediction block. These reference samples are identified by a reference picture index and a motion vector. Interpolation can be applied if the motion vector has fractional precision. The reference samples used to determine the average of the reference samples can include pre-interpolation or post-interpolation reference samples. Pre-interpolation reference samples can include motion vectors clipped to integer precision. Consistent with disclosed embodiments, all of the reference samples can be used to calculate the average. Alternatively, only a portion of the reference samples (e.g., the reference sample corresponding to the top-left position of the current block) can be used to calculate the average.

[0123]

[0140] As shown in Figure 5, intra prediction (e.g., intra prediction 514 or 528) can be performed in the reshaped region while inter prediction is performed in the original region. Thus, in inter prediction, forward mapping can be applied to the predicted block, and the luma predicted block after forward mapping is used to calculate the average. To reduce latency, the average can be calculated using the predicted block before forward mapping. For example, the block before forward mapping, an NxN portion of the block before forward mapping, or the top-left sample of the block before forward mapping can be used.

[0124]

[0141] Embodiments of the present disclosure further provide a method for processing video content with chroma scaling for dual-tree partitioning.

[0125]

[0142] Since the dependency on luma can cause hardware design complications, chroma scaling can be turned off for intra-coded tile groups, which allows for dual-tree partitioning. However, this restriction can cause a loss of coding efficiency. Sample values ​​of the corresponding luma blocks are averaged to calculate avgY', and the partition index Y Idx Determine the chroma scale factor inverse_chroma_scaling[YIdx Instead of obtaining [], the chroma scale factor can be explicitly signaled in the bitstream to avoid the dependency on luma in the case of dual-tree splitting.

[0126]

[0143] The chroma scale index can be signaled at various levels. For example, as shown in Table 5 of FIG. 11, the chroma scale index can be signaled at the coding unit (CU) level together with the chroma prediction mode. To determine the chroma scale factor of the current chroma block, the syntax element "lmcs_scaling_factor_idx" can be used. If there is no "lmcs_scaling_factor_idx", it can be inferred that the chroma scale factor of the current chroma block is equal to 1.0 with floating-point precision or equivalently (1<<CSCALE_FP_PREC) with fixed-point precision. The allowable value range of "lmcs_chroma_scaling_idx" is determined at the tile group level and will be described later.

[0127]

[0144] Depending on the possible values ​​of "lmcs_chroma_scaling_idx," signaling costs may be high, especially for small blocks. Therefore, in some embodiments of the present disclosure, the signaling conditions in Table 5 of FIG. 11 may additionally include a condition for block size. For example, this syntax element "lmcs_chroma_scaling_idx" (highlighted in italics and gray shading) may be signaled only if the current block contains more than a given number of chroma samples, or if the current block has a width greater than a given width W or a height greater than a given height H. For smaller blocks, if "lmcs_chroma_scaling_idx" is not signaled, the decoder may determine its chroma scale factor. In some embodiments, the chroma scale factor may be set to 1.0 with floating-point precision. In some embodiments, a default "lmcs_chroma_scaling_idx" value may be added at the tile group header level (see 1 in FIG. 6). Small blocks that do not have a signaled "lmcs_chroma_scaling_idx" can use this tile group level default index to derive the corresponding chroma scale factor. In some embodiments, the chroma scale factor of a small block can be inherited from its neighbors (e.g., top or left neighbors) that explicitly signaled a scale factor.

[0128]

[0145] In addition to signaling the syntax element "lmcs_chroma_scaling_idx" at the CU level, this syntax element can also be signaled at the CTU level. However, given that the maximum CTU size in VVC is 128x128, performing the same scaling at the CTU level may be too coarse. Therefore, in some embodiments of the present disclosure, this syntax element "lmcs_chroma_scaling_idx" can be signaled using a fixed granularity. For example, one "lmcs_chroma_scaling_idx" is signaled for each 16x16 region within a CTU, and it applies to all samples within that 16x16 region.

[0129]

[0146] The range of "lmcs_chroma_scaling_idx" for the current tile group depends on the number of chroma scale factor values ​​allowed in the current tile group. The number of chroma scale factor values ​​allowed in the current tile group can be determined based on a 64-entry chroma LUT as discussed above. Alternatively, the number of chroma scale factor values ​​allowed in the current tile group can be determined using the chroma scale factor calculations discussed above.

[0130]

[0147] For example, in the "quantize" method, the value of LOG2_n can be set to 2 (i.e., "n" is set to 4), and the codeword assignment for each segment in the piecewise linear model for the current tile group can be set as follows: {0, 65, 66, 64, 67, 62, 62, 64, 64, 64, 67, 64, 64, 62, 61, 0}. Since any codeword value between 64 and 67 can have the same scale factor value (1.0 in fractional precision), and any codeword value between 60 and 63 can have the same scale factor value (60 / 64 = 0.9375 in fractional precision), there are only two possible scale factor values ​​for the entire tile group. The two extreme segments, which do not have any codewords assigned, have the chroma scale factor set to 1.0 by default. Therefore, in this example, one bit is sufficient to signal "lmcs_chroma_scaling_idx" for the blocks in the current tile group.

[0131]

[0148] In addition to using a piecewise linear model to determine possible chroma scale factor values, the encoder can signal a set of chroma scale factor values ​​in the tile group header, and then at the block level, use that set of chroma scale factor values ​​and the value of the block's "lmcs_chroma_scaling_idx" to determine the chroma scale factor value for the block.

[0132]

[0149] CABAC coding can be applied to code "lmcs_chroma_scaling_idx". The CABAC context of a block can depend on the "lmcs_chroma_scaling_idx" of its neighboring blocks. For example, the left block or the top block can be used to form the CABAC context. Regarding the binarization of this "lmcs_chroma_scaling_idx" syntax element, the same truncated rice binarization applied to the ref_idx_10 and ref_idx_11 syntax elements in VVC Draft 4 can be used to binarize "lmcs_chroma_scaling_idx".

[0133]

[0150] The advantage of signaling 'chroma_scaling_idx' is that the encoder can select the best 'lmcs_chroma_scaling_idx' in terms of rate-distortion cost. Selecting 'lmcs_chroma_scaling_idx' using rate-distortion optimization can improve coding efficiency, which may help offset the increased signaling cost.

[0134]

[0151] Embodiments of the present disclosure further provide a method for processing video content with LMCS piecewise linear model signaling.

[0135]

[0152] Although the LMCS method uses a piecewise linear model with 16 partitions, the number of unique values ​​of "SignaledCW[i]" within a tile group is likely to be much less than 16. For example, some of the 16 partitions may use the default number of codewords "OrgCW", and some of the 16 partitions may have the same number of codewords as each other. Thus, an alternative way of signaling the LMCS piecewise linear model may include signaling the number of unique codewords "listUniqueCW[]" and sending an index for each of the partitions to indicate the element of "listUniqueCW[]" for the current partition.

[0136]

[0153] The revised syntax table is shown in Figure 12. Table 6 in Figure 12 highlights new or revised syntax in italics and gray shading.

[0137]

[0154] The semantics of the disclosed signaling method are as follows, with changes underlined: reshaper_model_min_bin_idx specifies the minimum bin (or partition) index used in the reshaper construction process. The value of reshaper_model_min_bin_idx shall be in the range 0 to MaxBinIdx. The value of MaxBinIdx shall be equal to 15. reshaper_model_delta_max_bin_idx specifies the maximum allowed bin (or partition) index MaxBinIdx minus the maximum bin index used in the reshaper construction process. The value of reshaper_model_max_bin_idx is set equal to MaxBinIdx-reshape_model_delta_max_bin_idx. reshaper_model_bin_delta_abs_cw_prec_minus1 plus 1 specifies the number of bits used to represent the syntax reshape_model_bin_delta_abs_CW[i]. reshaper_model_bin_num_unique_cw_minus1 plus 1 specifies the size of the codeword array listUniqueCW. reshaper_model_bin_delta_abs_CW[i] specifies the absolute delta codeword value for the ith bin. reshaper_model_bin_delta_sign_CW_flag[i] specifies the sign of reshaper_model_bin_delta_abs_CW[i] as follows: - If reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is positive. Otherwise (reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0) the corresponding variable RspDeltaCW[i] is negative. If reshape_model_bin_delta_sign_CW_flag[i] is missing, the corresponding variable RspDeltaCW[i] is inferred to be equal to 0. The variable RspDeltaCW[i] is derived as RspDeltaCW[i]=(1-2*reshape_model_bin_delta_sign_CW[i])*reshape_model_bin_delta_abs_CW[i]. The variable listUniqueCW[0] is set equal to OrgCW. The variables listUniqueCW[i] for i=1... reshaper_model_bin_num_unique_cw_minus1 are as follows: It is derived as follows: - Set the variable OrgCW to (1< <BitDepth Y ) / (MaxBinIdx+1). - listUniqueCW[i] =OrgCW+RspDeltaCW[i-1] reshaper_model_bin_cw_idx[i] specifies the index into the array listUniqueCW[] used to derive RspCW[i]. The value of reshaper_model_bin_cw_idx[i] shall be in the range 0 to (reshaper_model_bin_num_unique_cw_minus1+1). RspCW[i] is derived as follows: - If reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx holds, RspCW[i]= listUniqueCW[reshaper_model_bin_cw_idx[i]]. - Otherwise RspCW[i]=0. BitDepth Y If the value of is equal to 10, the value of RspCW[i] can be in the range of 32 to 2*OrgCW-1.

[0138]

[0155] Embodiments of the present disclosure further provide a method for processing video content with conditional chroma scaling at the block level.

[0139]

[0156] As shown in Table 1 of Figure 6, whether chroma scaling is applied can be determined by the "tile_group_reshaper_chroma_residual_scale_flag" signaled at the tile group level.

[0140]

[0157] However, it may be beneficial to determine whether to apply chroma scaling at the block level. For example, in some disclosed embodiments, a CU-level flag may be signaled to indicate whether chroma scaling is applied to the current block. The presence of the CU-level flag may be conditioned based on the tile group-level flag "tile_group_reshaper_chroma_residual_scale_flag." That is, the CU-level flag may be signaled only if chroma scaling is allowed at the tile group level. Although the encoder is allowed to choose whether to use chroma scaling based on whether chroma scaling is beneficial for the current block, it may also incur significant signaling overhead.

[0141]

[0158] Consistent with the disclosed embodiments, to avoid the above signaling overhead, whether chroma scaling is applied to a block can be conditioned based on the prediction mode of the block. For example, if a block is inter-predicted, the prediction signal tends to be good, especially if its reference picture is close in terms of temporal distance. Therefore, chroma scaling can be bypassed because the residual is expected to be very small. For example, pictures in higher temporal levels tend to have reference pictures that are close in terms of temporal distance. For a block, chroma scaling can be disabled in pictures that use nearby reference pictures. To determine whether this condition is met, the difference in Picture Order Count (POC) between the current picture and the block's reference picture can be used.

[0142]

[0159] In some embodiments, chroma scaling can be disabled for all inter-coded blocks. In some embodiments, chroma scaling can be disabled for combined intra / inter prediction (CIIP) modes defined in the VVC standard.

[0143]

[0160] In the VVC standard, the CU syntax structure "coding_unit()" includes a syntax element "cu_cbf" to indicate whether there are any non-zero residual coefficients in the current CU. At the TU level, the TU syntax structure "transform_unit()" includes syntax elements "tu_cbf_cb" and "tu_cbf_cr" to indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the current TU. The chroma scaling process can be conditioned based on these flags. As described above, if there are no non-zero residual coefficients, the averaging of the corresponding luma-chroma scaling process can be invoked. By invoking averaging, the chroma scaling process can be bypassed.

[0144]

[0161] FIG. 13 shows a flow diagram of a computer-implemented method 1300 for processing video content. In some embodiments, method 1300 may be performed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec may be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding or converting a video sequence into another code. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence may be a surveillance video sequence that may be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). A video sequence may include multiple pictures. The device may perform method 1300 at the picture level. For example, the device may process pictures one at a time in method 1300. In another example, the device may process multiple pictures at a time in method 1300. Method 1300 may include steps such as:

[0145]

[0162] At step 1302, chroma blocks and luma blocks associated with a picture may be received. It will be appreciated that a picture may be associated with chroma and luma components. Thus, a picture may be associated with chroma blocks that include chroma samples and luma blocks that include luma samples.

[0146]

[0163] In step 1304, luma scale information associated with the luma block may be determined. In some embodiments, the luma scale information may be a syntax element signaled in the data stream for the picture or a variable derived based on a syntax element signaled in the data stream for the picture. For example, the luma scale information may include "reshape_model_bin_delta_sign_CW[i] and reshape_model_bin_delta_abs_CW[i]" as described in the above equations, and / or "SignaledCW[i]" as described in the above equations, etc. In some embodiments, the luma scale information may include a variable determined based on the luma block. For example, an average luma value may be determined by calculating the average value of luma samples adjacent to the luma block (such as luma samples in the top row of the luma block and luma samples in the left column of the luma block).

[0147]

[0164] In step 1306, a chroma scale factor can be determined based on the luma scale information.

[0148]

[0165] In some embodiments, a luma scale factor for a luma block may be determined based on luma scale information. For example, according to the above equation of "inverse_chroma_scaling[i]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(tempCW>>1)) / tempCW", a luma scale factor may be determined based on luma scale information (e.g., "tempCW"). Then, a chroma scale factor may be further determined based on the value of the luma scale factor. For example, the chroma scale factor may be set equal to the value of the luma scale factor. It will be appreciated that further calculations may be applied to the value of the luma scale factor before being set as the chroma scale factor. As another example, the chroma scale factor may be set equal to "SignaledCW[Y Idx] / OrgCW” and sets the partition index of the current chroma block “Y” based on the average luma value associated with the luma block. Idx " can be determined.

[0149]

[0166] In step 1308, a chroma block may be processed using a chroma scale factor. For example, a residual of the chroma block may be processed using the chroma scale factor to generate a scaled residual of the chroma block. The chroma block may be a Cb chroma component or a Cr chroma component.

[0150]

[0167] In some embodiments, a chroma block may be processed if a condition is met. For example, the condition may include that a target coded unit associated with a picture does not have a non-zero residual or that a target transform unit associated with a picture does not have a non-zero chroma residual. The fact that a target coded unit does not have a non-zero residual may be determined based on the value of a first coded block flag of the target coded unit. The fact that a target transform unit does not have a non-zero chroma residual may be determined based on the value of a second coded block flag of a first component of the target transform unit and the value of a third coded block flag of the second component. For example, the first component may be a Cb component, and the second component may be a Cr component.

[0151]

[0168] It will be appreciated that each step of method 1300 may be performed as a separate method, for example, the method for determining a chroma scale factor described in step 1308 may be performed as a separate method.

[0152]

[0169] FIG. 14 shows a flow diagram of a computer-implemented method 1400 for processing video content. In some embodiments, method 1300 may be performed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec may be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding or converting a video sequence into another code. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence may be a surveillance video sequence that may be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). A video sequence may include multiple pictures. The device may perform method 1400 at the picture level. For example, the device may process pictures one at a time in method 1400. In another example, the device may process multiple pictures at a time in method 1400. Method 1400 may include steps such as:

[0153]

[0170] At step 1402, chroma blocks and luma blocks associated with a picture may be received. It will be understood that a picture may be associated with chroma components and luma components. Thus, a picture may be associated with chroma blocks including chroma samples and luma blocks including luma samples. In some embodiments, a luma block may include NxM luma samples, where N may be the width of the luma block and M may be the height of the luma block. As discussed above, the luma samples of the luma block may be used to determine the partition index of the target chroma block. Thus, a luma block associated with a picture of a video sequence may be received. It will be understood that N and M may have the same value.

[0154]

[0171] In step 1404, a subset of NxM luma samples may be selected depending on at least one of N and M exceeding a threshold. To accelerate the determination of the partition index, the luma block may be "downsampled" if certain conditions are met. In other words, a subset of luma samples within the luma block may be used to determine the partition index. In some embodiments, the certain condition is that at least one of N and M exceeds a threshold. In some embodiments, the threshold may be based on at least one of N and M. The threshold may be a power of two. For example, the threshold may be 4, 8, 16, etc. Taking 4 as an example, if N or M is greater than 4, a subset of luma samples may be selected. In the example of Figure 10, both the width and height of the luma block exceed the threshold of 4, so a subset of 4x4 samples is selected. It will be appreciated that subsets such as 2x8, 1x16, etc. may also be selected for processing.

[0155]

[0172] In step 1406, the mean value of the subset of NxM luma samples may be determined.

[0156]

[0173] In some embodiments, determining the average value may further include determining whether a second condition is met, and determining the average value of the subset of NxM luma samples in response to determining that the second condition is met. For example, the second condition may include that a target coded unit associated with the picture has no non-zero residual coefficients or that the target coded unit has no non-zero chroma residual coefficients.

[0157]

[0174] In step 1408, a chroma scale factor based on the average value may be determined. In some embodiments, to determine the chroma scale factor, a partition index of the chroma block may be determined based on the average value; whether the partition index of the chroma block satisfies a first condition may be determined; and the chroma scale factor may be set to a default value in response to the partition index of the chroma block satisfying the first condition. The default value may indicate that chroma scaling is not applied. For example, the default value may be 1.0 with decimal precision. It will be understood that a fixed-point approximation may be applied to the default value. In response to the partition index of the chroma block not satisfying the first condition, the chroma scale factor may be determined based on the average value. More specifically, the chroma scale factor may be set to SignaledCW[Y Idx ] / OrgCW and the target chroma block division index 'Y Idx " can be determined based on the average value of the corresponding luma block.

[0158]

[0175] In some embodiments, the first condition may include the partition index of the chroma block being greater than the maximum index of the signaled codeword or less than the minimum index of the signaled codeword. The maximum and minimum indices of the signaled codeword may be determined as follows:

[0159]

[0176] The codewords can be generated using a piecewise linear model (e.g., LMCS) based on the input signal (e.g., luma samples). As discussed above, the dynamic range of the input signal can be divided into several partitions (e.g., 16 partitions), and each partition of the input signal can be used to generate a bin of the codeword as an output. Thus, each bin of the codeword can have a bin index corresponding to a partition of the input signal. In this example, the bin index can range from 0 to 15. In some embodiments, the output (i.e., codeword) value can be between a minimum value (e.g., 0) and a maximum value (e.g., 255), and multiple codewords having values ​​between the minimum and maximum values ​​can be signaled. Bin indices for the signaled multiple codewords can be determined. Of the bin indices for the signaled multiple codewords, a maximum bin index and a minimum bin index for the bins of the signaled multiple codewords can further be determined.

[0160]

[0177] In addition to the chroma scale factors, method 1400 can further determine luma scale factors based on the bins of the signaled codewords. The luma scale factors can be used as inverse chroma scale factors. The equations for determining the luma scale factors are described above and will not be described again here. In some embodiments, multiple signaled adjacent codewords share a luma scale factor. For example, two or four signaled adjacent codewords can share the same luma scale factor, which can reduce the burden of determining the luma scale factor.

[0161]

[0178] In step 1410, chroma blocks can be processed using chroma scale factors. As discussed above with respect to Figure 5, multiple chroma scale factors can be constructed in a chroma scale factor LUT at the tile group level and applied to the reconstructed chroma residual of the target block at the decoder side. Similarly, chroma scale factors can also be applied at the encoder side.

[0162]

[0179] It will be appreciated that each step of method 1400 may be performed as a separate method, for example, the method for determining the chroma scale factor described in step 1308 may be performed as a separate method.

[0163]

[0180] FIG. 15 shows a flow diagram of a computer-implemented method 1500 for processing video content. In some embodiments, method 1500 may be performed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec may be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding or converting a video sequence into another code. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence may be a surveillance video sequence that may be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). A video sequence may include multiple pictures. The device may perform method 1500 at the picture level. For example, the device may process pictures one at a time in method 1500. In another example, the device may process multiple pictures at a time in method 1500. Method 1500 may include steps such as:

[0164]

[0181] In step 1502, it may be determined whether there is a chroma scale index in the received video data.

[0165]

[0182] In step 1504, in response to determining that there is no chroma scale index in the received video data, it may be determined that no chroma scaling is applied to the received video data.

[0166]

[0183] In step 1506, in response to determining that there is chroma scaling in the received video data, a chroma scale factor may be determined based on the chroma scale index.

[0167]

[0184] FIG. 16 shows a flow diagram of a computer-implemented method 1600 for processing video content. In some embodiments, method 1600 may be performed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec may be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding or converting a video sequence into another code. In some embodiments, the video sequence may be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence may be a surveillance video sequence that may be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). A video sequence may include multiple pictures. The device may perform method 1600 at the picture level. For example, the device may process pictures one at a time in method 1600. In another example, the device may process multiple pictures at a time in method 1600. Method 1600 may include steps such as:

[0168]

[0185] At step 1602, a number of unique codewords to be used for the dynamic range of the input video signal may be received.

[0169]

[0186] In step 1604, an index may be received.

[0170]

[0187] In step 1606, at least one of the plurality of unique code words may be selected based on the index.

[0171]

[0188] In step 1608, a chroma scale factor may be determined based on the selected at least one codeword.

[0172]

[0189] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which can be executed by an apparatus (such as the disclosed encoders and decoders) to perform the above-described methods. Common non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM, EPROM, Flash EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. An apparatus may include one or more processors (CPUs), input / output interfaces, a network interface, and / or memory.

[0173]

[0190] It will be understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, the software can be stored in the above computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computational units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above modules / units can be combined into one module / unit, and that each of the above modules / units can be further divided into multiple sub-modules / sub-units.

[0174]

[0191] Embodiments may be further described using the following clauses: 1. A computer-implemented method for processing video content, comprising: receiving chroma blocks and luma blocks associated with a picture; determining luma scale information associated with a luma block; determining a chroma scale factor based on the luma scale information; and Processing chroma blocks using chroma scale factors A method comprising: 2. Determining a chroma scale factor based on luma scale information; determining a luma scale factor for the luma block based on the luma scale information; Determining a chroma scale factor based on the value of the luma scale factor 2. The method of clause 1, further comprising: 3. Determining a chroma scale factor based on the value of the luma scale factor; Setting the chroma scale factor equal to the value of the luma scale factor 3. The method of clause 2, further comprising: 4. Processing chroma blocks using chroma scale factors determining whether a first condition is met; and processing the chroma block using the chroma scale factor in response to determining that the first condition is met; or Bypassing processing of the chroma block with the chroma scale factor in response to determining that the first condition is not satisfied. Do one of the following: 4. The method of any one of clauses 1 to 3, further comprising: 5. The first condition is The target coding unit associated with the picture does not have a non-zero residual, or The target transform unit associated with the picture does not have a non-zero chroma residual. 5. The method according to clause 4, comprising: 6. The target coding unit does not have a non-zero residual, determined based on the value of a first coded block flag of the target coding unit; determining that the target transform unit does not have a non-zero chroma residual based on values ​​of second coded block flags of a first chroma component and values ​​of third coded block flags of a second luma chroma component of the target transform unit; The method described in clause 5. 7. The value of the first coded block flag is 0; the value of the second coded block flag and the value of the third coded block flag are 0; The method described in clause 6. 8. Processing chroma blocks using chroma scale factors Processing the residual of the chroma block using the chroma scale factor 8. The method according to any one of clauses 1 to 7, comprising: 9. A device for processing video content, comprising: a memory for storing a set of instructions; coupled to the memory, receiving chroma blocks and luma blocks associated with a picture; determining luma scale information associated with a luma block; determining a chroma scale factor based on the luma scale information; and Processing chroma blocks using chroma scale factors a processor configured to execute a set of instructions to cause the device to Including, equipment. 10. When determining a chroma scale factor based on luma scale information, determining a luma scale factor for the luma block based on the luma scale information; Determining a chroma scale factor based on the value of the luma scale factor 10. The apparatus of clause 9, wherein the processor is configured to execute a set of instructions to cause the apparatus to further: 11. When determining a chroma scale factor based on the value of a luma scale factor: Setting the chroma scale factor equal to the value of the luma scale factor 11. The apparatus of clause 10, wherein the processor is configured to execute a set of instructions to cause the apparatus to further: 12. When processing chroma blocks using chroma scale factors, determining whether a first condition is met; and processing the chroma block using the chroma scale factor in response to determining that the second condition is met; or Bypassing processing of the chroma block with the chroma scale factor in response to determining that the second condition is not satisfied. Do one of the following: 12. The apparatus of any one of clauses 9 to 11, wherein the processor is configured to execute a set of instructions to cause the apparatus to further: 13. The first condition is: The target coding unit associated with the picture does not have a non-zero residual, or The target transform unit associated with the picture does not have a non-zero chroma residual. The equipment referred to in clause 12, including: 14. The target coding unit does not have a non-zero residual, determined based on the value of a first coded block flag of the target coding unit; determining that the target transform unit does not have a non-zero chroma residual based on values ​​of a second coded block flag of a first chroma component and values ​​of a third coded block flag of the second chroma component of the target transform unit; Equipment as described in clause 13. 15. The value of the first coded block flag is 0; the value of the second coded block flag and the value of the third coded block flag are 0; Equipment as described in clause 14. 16. When processing chroma blocks using chroma scale factors, Processing the residual of the chroma block using the chroma scale factor 16. The apparatus of any one of clauses 9 to 15, wherein the processor is configured to execute a set of instructions to cause the apparatus to further: 17. A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of a device to cause the device to perform a method for processing video content, the method comprising: receiving chroma blocks and luma blocks associated with a picture; determining luma scale information associated with a luma block; determining a chroma scale factor based on the luma scale information; and Processing chroma blocks using chroma scale factors 1. A non-transitory computer-readable storage medium comprising: 18. A computer-implemented method for processing video content, comprising: receiving chroma blocks and luma blocks associated with a picture, the luma blocks including NxM luma samples; selecting a subset of NxM luma samples in response to at least one of N and M being above a threshold; determining a mean value of a subset of NxM luma samples; determining a chroma scale factor based on the average value; and Processing chroma blocks using chroma scale factors A method comprising: 19. A computer-implemented method for processing video content, comprising: determining whether a chroma scale index is present in the received video data; determining, in response to determining that there is no chroma scale index in the received video data, that chroma scaling is not applied to the received video data; and In response to determining that there is chroma scaling in the received video data, determining a chroma scale factor based on the chroma scale index. A method comprising: 20. A computer-implemented method for processing video content, comprising: receiving a plurality of unique codewords for use in a dynamic range of the input video signal; receiving an index; selecting at least one of the plurality of unique code words based on the index; and determining a chroma scale factor based on the selected at least one codeword; A method comprising:

[0175]

[0192] In addition to implementing the above methods using computer-readable program code, the above methods can also be implemented in the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Thus, such controllers can be considered hardware components, and devices contained within the controllers and configured to implement various functions can also be considered structures within the hardware components. Or, devices configured to implement various functions can even be considered both software modules configured to implement the methods and structures within the hardware components.

[0176]

[0193] The present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, assemblies, data structures, classes, etc., used to perform particular tasks or implement particular abstract data types. Embodiments of the present disclosure may also be implemented in distributed computing environments, where tasks are performed using remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may reside in both local and remote computer storage media, including storage devices.

[0177]

[0194] It should be noted that relative terms such as "first" and "second" are used herein merely to distinguish one entity or operation from another and do not require or imply any actual relationship or order between those entities or operations. Furthermore, the terms "comprising," "having," "containing," and "including," and other similar forms of words, are intended to be equivalent in meaning and are open-ended in that the items following any one of these terms are not intended to be an exhaustive list of such items or to be limited to only the items listed.

[0178]

[0195] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the described embodiments may be made. Other embodiments may become apparent to those skilled in the art from consideration of the specification and practice of the disclosure set forth herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims. The order of steps depicted in the figures is for illustrative purposes only and is not intended to be limited to any particular order of steps. As such, one skilled in the art will recognize that steps may be performed in different orders while implementing the same method.

Claims

1. reconstructing a chroma block based on a plurality of luma samples associated with the chroma block; reconstructing the chroma blocks determining whether the chroma block has a non-zero chroma residual; and and bypassing a process of averaging the plurality of luma samples in response to determining that the chroma block does not have a non-zero chroma residual, wherein the averaging process is used to reconstruct the chroma block. Including, Methods for computer-implemented image processing.

2. reconstructing the chroma blocks in response to determining that the chroma block has one or more non-zero chroma residuals, performing the averaging process to determine an average value of the plurality of luma samples, scaling the one or more non-zero chroma residuals based on the average value, and reconstructing the chroma block based on the scaled one or more non-zero chroma residuals. Further comprising: The method of claim 1.

3. scaling the one or more non-zero chroma residuals based on the average values ​​of the plurality of luma samples; determining a chroma scale factor based on the average value; and applying the chroma scale factor to the one or more non-zero chroma residuals; The method of claim 2 , comprising:

4. The method of claim 1 , wherein determining whether the chroma block has no non-zero chroma residual is based on a value of a coded block flag associated with the chroma block.

5. determining that the chroma block has one or more non-zero chroma residuals in response to the value of the coded block flag being equal to one; The method of claim 4 further comprising:

6. determining that the chroma block does not have a non-zero chroma residual in response to the value of the coded block flag being equal to 0; The method of claim 4 further comprising:

7. Whether the chroma block has no non-zero chroma residual is determined by: a value of a first coded block flag associated with a first chroma component of the chroma block; and a value of a second coded block flag associated with a second chroma component of the chroma block; The method of claim 1 , wherein the determination is based on:

8. determining that the chroma block has one or more non-zero chroma residuals in response to at least one of the value of the first coded block flag or the value of the second coded block flag being equal to one; The method of claim 7 further comprising:

9. determining that the chroma block does not have a non-zero chroma residual in response to the value of the first coded block flag and the value of the second coded block flag both being equal to 0; The method of claim 7 further comprising:

10. one or more memories that store a set of instructions; one or more processors; An apparatus comprising: the one or more processors: configured to execute the set of instructions to cause the device to reconstruct a chroma block based on a plurality of luma samples associated with the chroma block; When reconstructing the chroma blocks, determining whether the chroma block has a non-zero chroma residual; and and bypassing a process of averaging the plurality of luma samples in response to determining that the chroma block does not have a non-zero chroma residual, wherein the averaging process is used to reconstruct the chroma block. the one or more processors are configured to execute the set of instructions to further cause the device to: device.

11. the one or more processors: in response to determining that the chroma block has one or more non-zero chroma residuals, performing the averaging process to determine an average value of the plurality of luma samples, scaling the one or more non-zero chroma residuals based on the average value, and reconstructing the chroma block based on the scaled one or more non-zero chroma residuals.

11. The device of claim 10, configured to execute the set of instructions to further cause the device to:

12. the one or more processors: determining a chroma scale factor based on the average value of the plurality of luma samples; and applying the chroma scale factor to the one or more non-zero chroma residuals; 12. The device of claim 11, configured to execute the set of instructions to further cause the device to:

13. the one or more processors: determining whether the chroma block does not have a non-zero chroma residual based on a value of a coded block flag associated with the chroma block; 11. The device of claim 10, configured to execute the set of instructions to further cause the device to:

14. the one or more processors: determining that the chroma block has one or more non-zero chroma residuals in response to the value of the coded block flag being equal to one; 14. The device of claim 13, configured to execute the set of instructions to further cause the device to:

15. the one or more processors: determining that the chroma block does not have a non-zero chroma residual in response to the value of the coded block flag being equal to 0; 14. The device of claim 13, configured to execute the set of instructions to further cause the device to:

16. 1. A method for storing a video bitstream, the method comprising: reconstructing a chroma block based on a plurality of luma samples associated with the chroma block; generating a bitstream containing coded information for reconstructing the chroma blocks; storing the bitstream on a non-transitory computer-readable medium; Including, reconstructing the chroma blocks determining whether the chroma block has a non-zero chroma residual; and and bypassing a process of averaging the plurality of luma samples in response to determining that the chroma block does not have a non-zero chroma residual, wherein the averaging process is used to reconstruct the chroma block. A method comprising:

17. The method comprises: in response to determining that the chroma block has one or more non-zero chroma residuals, performing the averaging process to determine an average value of the plurality of luma samples, scaling the one or more non-zero chroma residuals based on the average value, and reconstructing the chroma block based on the scaled one or more non-zero chroma residuals. Further comprising:

17. The method of claim 16.

18. scaling the one or more non-zero chroma residuals based on the average values ​​of the plurality of luma samples; determining a chroma scale factor based on the average value; and applying the chroma scale factor to the one or more non-zero chroma residuals; 18. The method of claim 17, comprising:

19. 17. The method of claim 16, wherein the coded information includes a coded block flag that indicates whether the chroma block has no non-zero chroma residual.

20. The coded information is a first coded block flag associated with a first chroma component of the chroma block; and a second coded block flag associated with a second chroma component of the chroma block; Including, a value of at least one of the first coded block flag or the second coded block flag equal to 1 indicates that the chroma block has one or more non-zero chroma residuals; 17. The method of claim 16, indicating that the chroma block does not have non-zero chroma residual in response to the value of the first coded block flag and the value of the second coded block flag both being equal to 0.

Citation Information

Patent Citations

  • Interactions between in-loop reshaping and inter coding tools

    WO2020156526A1

  • Signaling of in-loop reshaping information using parameter sets

    WO2020156529A1