Method and system for processing video content
By processing video content using chroma and luma blocks with calculated scale factors, the method addresses the high bandwidth and storage issues in HD video surveillance, achieving efficient bitrate reduction and improved encoding quality.
Patent Information
- Application Number
- JP2025027351
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-03-12
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-02-28
AI Technical Summary
High bandwidth and storage requirements for high-definition (HD) video in video surveillance applications, due to the high bitrate of HD video bitstreams.
A method for processing video content that involves receiving chroma blocks and luma blocks, determining luma scale information, calculating a chroma scale factor based on the luma scale information, and processing the chroma blocks using this scale factor.
This approach reduces the bitrate of video data without significant information loss, thereby decreasing bandwidth and storage demands, and improving encoding quality for various surveillance applications.
Smart Images

Figure 2025090618000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications
[0001] This disclosure claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 813,728, filed Mar. 4, 2019, and U.S. Provisional Patent Application No. 62 / 817,546, filed Mar. 12, 2019, both of which are hereby incorporated by reference in their entireties.
[0002] Technical Field
[0002] This disclosure generally relates to video processing, and more particularly to methods and systems for performing in - loop mapping by chroma scaling.
Background Art
[0003] Background
[0003] Video coding systems are often used to compress digital video signals, for example, to reduce the memory space consumed or to reduce the consumption of transmission bandwidth associated with such signals. As the popularity of high - definition (HD) video (e.g., having a resolution of 1920×1080 pixels) grows in various applications of video compression such as online video streaming, video conferencing, or video surveillance, there is an ongoing need to develop video coding tools that can improve the compression efficiency of video data.
[0004]
[0004] For example, video surveillance applications are being used more extensively and widely in many application scenarios (such as security, traffic, environmental monitoring, etc.), and the number and resolution of surveillance devices are increasing rapidly. Many video surveillance application scenarios choose to provide users with HD video in order to capture more information, and HD video has more pixels per frame to capture such information. However, an HD video bitstream can have a high bitrate that requires high bandwidth for transmission and large space for storage. For example, a surveillance video stream with an average resolution of 1920x1080 may require a bandwidth of 4 Mbps for real-time transmission. Furthermore, video surveillance generally performs round-the-clock monitoring for 7 days x 24 hours, which can significantly test the capabilities of the storage system when storing video data. Therefore, the demand for high bandwidth and large storage space for HD video has become the main limitation for the large-scale deployment of HD video in video surveillance.
Summary of the Invention
Means for Solving the Problems
[0005] Summary of the Disclosure
[0005] Embodiments of the present disclosure provide a method for processing video content. The method may include receiving chroma blocks and luma blocks related to a picture, determining luma scale information related to the luma blocks, determining a chroma scale factor based on the luma scale information, and processing the chroma blocks using the chroma scale factor.
[0006]
[0006] Embodiments of the present disclosure provide a device for processing video content. The device may include a memory storing a set of instructions, and a processor coupled to the memory and configured to execute the set of instructions to cause the device to receive chroma blocks and luma blocks related to a picture, determine luma scale information related to the luma blocks, determine a chroma scale factor based on the luma scale information, and process the chroma blocks using the chroma scale factor.
[0007]
[0007] Embodiments of the present disclosure provide a non - transitory computer - readable storage medium storing a set of instructions executable by one or more processors of a device to cause the device to perform a method for processing video content. The method includes receiving chroma blocks and luma blocks related to a picture, determining luma scale information related to the luma blocks, determining a chroma scale factor based on the luma scale information, and processing the chroma blocks using the chroma scale factor.
[0008] Brief Description of the Drawings
[0008] Embodiments and various aspects of the present disclosure are shown in the following detailed description and the accompanying drawings. The various features shown in the figures are not drawn to scale.
Brief Description of the Drawings
[0009]
Figure 1
[0009] An example structure of a video sequence according to some embodiments of the present disclosure is shown.
Figure 2A
[0010] A schematic diagram of an example of an encoding process according to some embodiments of the present disclosure is shown.
Figure 2B
[0011] A schematic diagram of another example of an encoding process according to some embodiments of the present disclosure is shown.
Figure 3A
[0012] A schematic diagram of an example of a decoding process according to some embodiments of the present disclosure is shown.
Figure 3B
[0013] A schematic diagram of another example of a decoding process according to some embodiments of the present disclosure is shown.
Figure 4
[0014] A block diagram of an example of a device for encoding or decoding video according to some embodiments of the present disclosure is shown.
Figure 5
[0015] Schematic diagrams of an exemplary luma mapping with chroma scaling (LMCS) process according to some embodiments of the present disclosure are shown.
Figure 6
[0016] A syntax table at the tile group level for an LMCS partitioned linear model according to some embodiments of the present disclosure is shown.
Figure 7
[0017] Another syntax table at the tile group level for an LMCS partitioned linear model according to some embodiments of the present disclosure is shown.
Figure 8
[0018] A table of the syntax structure of a coded tree unit according to some embodiments of the present disclosure.
Figure 9
[0019] A table of the syntax structure of a dual tree split according to some embodiments of the present disclosure.
Figure 10
[0020] An example of simplifying the averaging of luma prediction blocks according to some embodiments of the present disclosure is shown.
Figure 11
[0021] A table of the syntax structure of a coded tree unit according to some embodiments of the present disclosure.
Figure 12
[0022] A table of syntax elements related to the modified signaling of the LMCS partitioned linear model at the tile group level according to some embodiments of the present disclosure.
Figure 13
[0023] A flowchart of a method for processing video content according to some embodiments of the present disclosure.
Figure 14
[0024] A flowchart of a method for processing video content according to some embodiments of the present disclosure.
Figure 15
[0025] A flowchart of another method for processing video content according to some embodiments of the present disclosure.
Figure 16
[0026] A flowchart of another method for processing video content according to some embodiments of the present disclosure.
Best Mode for Carrying Out the Invention
[0010] Detailed Description
[0027] Next, a detailed reference is made to exemplary embodiments shown in the accompanying drawings with examples. The following description refers to the accompanying drawings, in which the same numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementation forms described in the following description of the exemplary embodiments do not represent all implementation forms that conform to the present invention. Rather, they are merely examples of devices and methods that conform to aspects related to the present invention recited in the appended claims. Unless otherwise specified, the term "or" includes all possible combinations except when not executable. For example, when it is stated that a certain component may include A or B, unless otherwise specified or not executable, the component may include A, or B, or A and B. As a second example, when it is stated that a certain component may include A, B, or C, unless otherwise specified or not executable, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0011]
[0028] Video is a set of still pictures (or "frames") arranged in chronological order for storing visual information. A video capture device (such as a camera) can be used to capture and store those pictures in chronological order, and a video playback device (such as a TV, computer, smartphone, tablet computer, video player, or any end-user terminal having a display function) can be used to display such pictures in chronological order. Further, in some applications, the video capture device can transmit the captured video in real time to a video playback device (such as a computer having a monitor) for monitoring, conferencing, or live broadcasting, etc.
[0012]
[0029] In order to reduce the memory space and transmission bandwidth required by such applications, the video can be compressed before being stored and transmitted, and decompressed before being displayed. This compression and decompression can be implemented by software executed by a processor (such as the processor of a general-purpose computer) or dedicated hardware. The module for compression is generally called an "encoder", and the module for decompression is generally called a "decoder". The encoder and decoder can be collectively called a "codec". The encoder and decoder can be implemented as various suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder may include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder may include program code, computer-executable instructions, firmware, or algorithms or processes implemented by any suitable computer fixed in a computer-readable medium. The compression and decompression of video can be implemented by various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be called a "transcoder".
[0013]
[0030] The video encoding process can identify and retain the useful information that can be used to reconstruct the picture, and can ignore the information that is not important for reconstruction. If the ignored unimportant information cannot be completely reconstructed, such an encoding process can be called "irreversible". Otherwise, such an encoding process can be called "reversible". Most encoding processes are irreversible, which is a trade-off for reducing the required memory space and transmission bandwidth.
[0014]
[0031] The useful information of the symbolized picture (referred to as the "current picture") includes the changes with respect to the reference picture (for example, a picture that was symbolized and reconstructed in the past). Such changes can include changes in pixel position, luminance, or color, among which the position change is the most relevant. The position change of the group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.
[0015]
[0032] A picture that is coded without referring to another picture (i.e., such a picture is its own reference picture) is called an "I picture". A picture that is coded using a past picture as a reference picture is called a "P picture". A picture that is coded using both a past picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B picture".
[0016]
[0033] As described above, video surveillance using HD video faces the issues of high bandwidth and large storage requirements. To address this issue, the bit rate of the coded video can be reduced. Among I pictures, P pictures, and B pictures, I pictures have the highest bit rate. Since the background of most surveillance videos is almost static, one way to reduce the overall bit rate of the coded video may be to use fewer I pictures for video coding.
[0017]
[0034] However, since I pictures are generally not dominant in the coded video, the improvement measure of using fewer I pictures may be minor. For example, in a typical video bitstream, the ratio of I pictures, B pictures, and P pictures may be 1:20:9, and the I pictures may account for less than 10% of the total bit rate. In other words, in such an example, even if all I pictures are removed, the reduced bit rate may be only 10%.
[0018]
[0035] The present disclosure provides a method, apparatus, and system for characteristic-based video processing for video surveillance. As used herein, "characteristics" refers to content characteristics related to video content within a picture, motion characteristics related to motion estimation for encoding or decoding a picture, or both. For example, the content characteristics can be pixels within one or more consecutive pictures of a video sequence, and the pixels are related to at least one of an object, a scene, or an environmental event within the picture. In another example, the motion characteristics can include information related to the encoding process of the video, examples of which will be described in detail later.
[0019]
[0036] In the present disclosure, when encoding a picture of a video sequence, a classifier can be used to detect and classify one or more characteristics of the picture of the video sequence. Different classes of characteristics can be associated with different priority levels, and such priority levels are further associated with different bitrates for encoding. The different priority levels can be associated with different parameter sets for encoding, thereby resulting in different encoding quality levels. The higher the priority level, the higher the quality of the video that the associated parameter set can result in. Such characteristic-based video processing can significantly reduce the bitrate for surveillance videos without causing significant information loss. In addition, embodiments of the present disclosure can customize the corresponding relationship between the priority level and the parameter set for various application scenarios (such as security, traffic, environmental monitoring, etc.), thereby significantly improving the encoding quality of the video and significantly reducing the bandwidth and storage cost.
[0020]
[0037] FIG. 1 shows the structure of an example of a video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live relay video or a captured and archived video. The video 100 can be a real video, a video generated by a computer (such as a computer game video), or a combination thereof (such as a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (such as a camera), a video archive including videos captured in the past (such as video files stored in a storage device), or a video feed interface (such as a video broadcast transceiver) for receiving videos from a video content provider.
[0021]
[0038] As shown in FIG. 1, the video sequence 100 can include a series of pictures temporally arranged along a timeline including pictures 102, 104, 106, and 108. Pictures 102 to 106 are consecutive, and there are more pictures between picture 106 and picture 108. In FIG. 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102 as indicated by the arrow. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108 as indicated by the arrows. In some embodiments, the reference picture of a picture (such as picture 104) may not be the picture immediately before or after that picture. For example, the reference picture of picture 104 can be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102 to 106 are only examples, and the present disclosure is not limited to the examples shown in FIG. 1 for the embodiments of the reference pictures.
[0022]
[0039] Typically, a video codec does not encode or decode an entire picture at once because such a task is computationally complex. Instead, a video codec divides a picture into basic segments and can encode or decode the picture segment by segment. In the present disclosure, such a basic segment is referred to as a basic processing unit ("BPU"). For example, structure 110 in FIG. 1 shows an example of the structure of a picture (e.g., any one of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, and its boundaries are indicated by dashed lines. In some embodiments, the basic processing unit can be referred to as a "macroblock" within some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), and can be referred to as a "coding tree unit" ("CTU") within some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have a variable size within a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for a picture based on the balance between coding efficiency and the level of detail to be maintained within the basic processing unit.
[0023]
[0040] The basic processing unit can be a logical unit that can include various types of video data groups stored in a computer memory (e.g., within a video frame buffer). For example, the basic processing unit of a color picture can include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements of the basic processing unit where the luma component and the chroma components can have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma component and the chroma components can be referred to as "coding tree blocks" ("CTBs"). Any operation performed on the basic processing unit can be repeated for each of its luma component and chroma components.
[0024]
[0041] The encoding of an image has multiple operation stages, examples of which are detailed in FIGS. 2A-2B and FIGS. 3A-3B. For each stage, the size of the basic processing unit may still be too large to process, and thus can be further divided into segments called "basic processing sub-units" in the present disclosure. In some embodiments, the basic processing sub-unit can be called a "block" within some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or can be called a "coding unit" ("CU") within some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing sub-unit can have a size equal to or smaller than that of the basic processing unit. Similar to the basic processing unit, the basic processing sub-unit is a logical unit that can include various types of video data groups (e.g., Y, Cb, Cr, and related syntax elements) stored in a computer memory (e.g., within a video frame buffer). Any operation performed on the basic processing sub-unit can be repeated for each of its luma and chroma components. It should be noted that such division can be performed at a further level according to the need for processing. It should also be noted that various stages can divide the basic processing unit in various ways.
[0025]
[0042] For example (one example of which is detailed in FIG. 2B) in the mode decision stage, the encoder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, and the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into a plurality of basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing sub-unit.
[0026]
[0043] In another example, (one example of which is detailed in FIG. 2A) in the prediction stage, the coder can perform a prediction operation at the level of a basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC) and perform the prediction operation at that level.
[0027]
[0044] In another example, (one example of which is detailed in FIG. 2A) in the transformation stage, the coder can perform a transformation operation on a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC) and perform the transformation operation at that level. It should be noted that the same basic processing subunit division method can be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU can have different sizes and numbers.
[0028]
[0045] In the structure 110 of FIG. 1, the basic processing unit 112 is further divided into 3x3 basic processing subunits, and its boundaries are indicated by dotted lines. Different basic processing units of the same picture can be divided into basic processing subunits in different ways.
[0029]
[0046] In some embodiments, to provide parallel processing and error resilience functions for video encoding and decoding, a picture can be divided into regions for processing, so that for the regions of the picture, the encoding or decoding process can be made independent of the information of any other region of the picture. In other words, each region of the picture can be processed independently. By doing so, the codec can process different regions of the picture in parallel, and thus the encoding efficiency can be improved. Further, if the data of a region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the damaged or lost data, thus providing an error resilience function. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions, namely "slices" and "tiles". It should also be noted that various pictures of video sequence 100 may have various splitting methods for dividing the picture into regions.
[0030]
[0047] For example, in FIG. 1, structure 110 is divided into three regions 114, 116, and 118, and its boundaries are shown as solid lines within structure 110. Region 114 contains four basic processing units. Each of regions 116 and 118 contains six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in FIG. 1 are only examples, and the present disclosure does not limit its embodiments.
[0031]
[0048] FIG. 2A shows a schematic diagram of an example of an encoding process 200A according to some embodiments of the present disclosure. The encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 may include a set of pictures (referred to as “original pictures”) arranged in chronological order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or areas for processing. In some embodiments, the encoder can execute process 200A at the level of the basic processing unit for each original picture of the video sequence 202. For example, the encoder can execute process 200A in an iterative manner, and the encoder can encode a basic processing unit within one iteration of process 200A. In some embodiments, the encoder can execute process 200A in parallel for the areas (e.g., areas 114-118) of each original picture of the video sequence 202.
[0032]
[0049] In FIG. 2A, the coder can feed the original basic processing unit of the video sequence 202 (referred to as the “original BPU”) to the prediction stage 204 to generate prediction data 206 and the predicted BPU 208. The coder can subtract the predicted BPU 208 from the original BPU to generate the residual BPU 210. The coder can feed the residual BPU 210 to the transform stage 212 and the quantization stage 214 to generate the quantized transform coefficients 216. The coder can feed the prediction data 206 and the quantized transform coefficients 216 to the binary coding stage 226 to generate the video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the “forward path”. During process 200A, the coder can feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 after the quantization stage 214 to generate the reconstructed residual BPU 222. The coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the prediction reference 224 used in the prediction stage 204 of the next iteration of process 200A. The components 218, 220, 222, and 224 of process 200A can be referred to as the “reconstruction path”. The reconstruction path can be used to ensure that both the coder and the decoder use the same reference data for prediction.
[0033]
[0050] The coder can repeatedly execute process 200A to encode each original BPU of the original picture (within the forward path) and generate the predicted reference 224 for encoding the next original BPU of the original picture (within the reconstruction path). After encoding all the original BPUs of the original picture, the coder can proceed to encode the next picture in the video sequence 202.
[0034]
[0051] Referring to process 200A, the coder can receive video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive (s)" can refer to receiving for inputting data, inputting, obtaining, retrieving, getting, reading out, accessing, or any action in any way.
[0035]
[0052] In the prediction stage 204 of the current iteration, the coder can receive the original BPU and prediction criterion 224 and perform a prediction operation to generate prediction data 206 and predicted BPU 208. Prediction criterion 224 can be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction stage 204 is to reduce the redundancy of information by extracting prediction data 206 that can be used to reconstruct the original BPU as predicted BPU 208 from prediction data 206 and prediction criterion 224.
[0036]
[0053] Ideally, predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such a difference, the coder can subtract predicted BPU 208 from the original BPU after generating it to generate residual BPU 210. For example, the coder can subtract the pixel value (e.g., grayscale value or RGB value) of predicted BPU 208 from the corresponding pixel value of the original BPU. As a result of such subtraction between the corresponding pixel of the original BPU and predicted BPU 208, each pixel of residual BPU 210 can have a residual value. Compared with the original BPU, prediction data 206 and residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly degrading the quality.
[0037]
[0054] To further compress the residual BPU 210, at the transformation stage 212, the coder can reduce the spatial redundancy of the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", where each basis pattern is associated with a "transformation coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency component (e.g., a luminance frequency component) of the residual BPU 210. None of the basis patterns can be reproduced from any combination (e.g., a linear combination) of any other basis patterns. In other words, such decomposition can decompose the variations of the residual BPU 210 into the frequency domain. Such decomposition is similar to the discrete Fourier transform of a function, the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transformation coefficients are similar to the coefficients associated with the basis functions.
[0038]
[0055] Various transformation algorithms can use various basis patterns. For example, at the transformation stage 212, various transformation algorithms such as the discrete cosine transform, the discrete sine transform, etc. can be used. The transformation at the transformation stage 212 is reversible. That is, the coder can restore the residual BPU 210 by the inverse operation of the transformation (referred to as "inverse transformation"). For example, to restore the pixels of the residual BPU 210, the inverse transformation can be to multiply the values of the corresponding pixels of the basis patterns by their respective associated coefficients and add the products to yield a weighted sum. In a video coding standard, both the coder and the decoder can use the same transformation algorithm (and thus the same basis patterns). Therefore, the coder can record only the transformation coefficients that can reconstruct the residual BPU 210 from there without the decoder receiving the basis patterns from the coder. The transformation coefficients may have fewer bits compared to the residual BPU 210, but those transformation coefficients can be used to reconstruct the residual BPU 210 without significantly degrading the quality. Therefore, the residual BPU 210 is further compressed.
[0039]
[0056] The coder can further compress the transform coefficients in the quantization stage 214. In the transform process, various basis patterns can represent various fluctuation frequencies (e.g., luminance fluctuation frequencies). Since the human eye is generally good at recognizing low-frequency fluctuations, the coder can ignore the information of high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the coder can generate the quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as the "quantization parameter") and rounding the quotient to its nearest neighbor. After such an operation, some of the transform coefficients of the high-frequency basis pattern can be converted to zero, and the transform coefficients of the low-frequency basis pattern can be converted to smaller integers. The coder can ignore the quantized transform coefficients 216 with zero values, thereby further compressing the transform coefficients. The quantization process is also reversible, and the quantized transform coefficients 216 can be reconstructed into the transform coefficients by an inverse operation of quantization (referred to as "inverse quantization").
[0040]
[0057] Since the coder ignores the remainder of such division in the rounding operation, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can contribute to the largest information loss within the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain various levels of information loss, the coder can use various values of the quantization parameter or any other parameter of the quantization process.
[0041]
[0058] In the binary encoding stage 226, the coder can encode the prediction data 206 and the quantized transform coefficients 216 using binary encoding techniques such as, for example, entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other reversible or irreversible compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the coder can encode other information such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), coder control parameters (e.g., bitrate control parameters), etc. in the binary encoding stage 226. The coder can generate a video bitstream 228 using the output data of the binary encoding stage 226. In some embodiments, the video bitstream 228 can be further packetized for network transmission.
[0042]
[0059] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the coder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transform stage 220, the coder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction criterion 224 to be used in the next iteration of process 200A.
[0043]
[0060] It should be noted that other variations of process 200A can be used to encode the video sequence 202. In some embodiments, the encoder can execute the steps of process 200A in a different order. In some embodiments, one or more steps of process 200A can be combined into a single step. In some embodiments, a single step of process 200A can be divided into multiple steps. For example, the transform step 212 and the quantization step 214 can be combined into a single step. In some embodiments, process 200A can include additional steps. In some embodiments, process 200A can omit one or more steps in FIG. 2A.
[0044]
[0061] FIG. 2B shows a schematic diagram of another example 200B of an encoding process according to some embodiments of the present disclosure. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode decision step 230, and divides the prediction step 204 into a spatial prediction step 2042 and a temporal prediction step 2044. The reconstruction path of process 200B additionally includes a loop filter step 232 and a buffer 234.
[0045]
[0062] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use the pixels of one or more adjacent BPUs that have already been encoded within the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use the regions of one or more pictures that have already been encoded to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include the encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0046]
[0063] Referring to process 200B, in the forward path, the coder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the coder can perform intra prediction. With respect to the original BPU of the coded picture, the prediction reference 224 can include one or more adjacent BPUs that are coded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The coder can generate the predicted BPU 208 by extrapolating the adjacent BPUs. The extrapolation technique can include, for example, linear extrapolation or linear interpolation, polynomial extrapolation or polynomial interpolation, etc. In some embodiments, the coder can perform extrapolation at the pixel level, such as by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The adjacent BPUs used for extrapolation can be located with respect to the original BPU in various directions, such as the vertical direction (e.g., above the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., bottom left, bottom right, top left, or top right of the original BPU), or any direction defined within the video coding standard being used. In intra prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the parameters of the extrapolation, the direction of the adjacent BPUs used with respect to the original BPU, etc.
[0047]
[0064] In another example, the coder can perform inter prediction in the temporal prediction stage 2044. For the original BPU of the current picture, the prediction reference 224 can include one or more pictures (referred to as "reference pictures") that are encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures can be encoded and reconstructed for each BPU. For example, the coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same picture are generated, the coder can generate a reconstructed picture as a reference picture. The coder can perform an operation of "motion estimation" to find a matching region within the range of the reference picture (referred to as the "search window"). The position of the search window within the reference picture can be determined based on the position of the original BPU within the current picture. For example, the search window can be centered at a position having the same coordinates as the original BPU within the current picture within the reference picture and can be expanded over a predetermined distance. When the coder identifies a region similar to the original BPU within the search window (for example, by using a pel recursive algorithm, a block matching algorithm, etc.), the coder can determine that region as the matching region. The matching region can have dimensions different from (for example, smaller than, equal to, larger than, or of a different shape than) the original BPU. Since the reference picture and the current picture are temporally separated within the timeline (as shown in FIG. 1, for example), it can be considered that the matching region "moves" to the position of the original BPU as time passes. The coder can record the direction and distance of such motion as a "motion vector". When multiple reference pictures (such as picture 106 in FIG. 1) are used, the coder can search for a matching region for each reference picture and obtain its associated motion vector. In some embodiments, the coder can assign weights to the pixel values of the matching regions of the individual matching reference pictures.
[0048]
[0065] Motion estimation can be used to identify various types of motion such as translational motion, rotational motion, and scaling. In inter prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, and the like.
[0049]
[0066] To generate the predicted BPU 208, the coder can perform an operation of "motion compensation". Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vector) and the prediction reference 224. For example, the coder can move the matching region of the reference picture according to the motion vector, in which the coder can predict the original BPU of the current picture. When multiple reference pictures (such as picture 106 in FIG. 1) are used, the coder can move the matching regions of the reference pictures according to the individual motion vectors and average the pixel values of the matching regions. In some embodiments, when the coder assigns weights to the pixel values of the matching regions of the individual matching reference pictures, the coder can obtain the weighted sum of the pixel values of the moved matching regions.
[0050]
[0067] In some embodiments, the inter prediction can be either unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction with respect to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter prediction picture in which the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions with respect to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter prediction picture in which the reference pictures (i.e., pictures 104 and 108) are in both temporal directions with respect to picture 104.
[0051]
[0068] Continuing to refer to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the coder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the coder can perform rate distortion optimization techniques, in which the coder can select a prediction mode to minimize the value of a cost function according to the bit rate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the coder can generate the corresponding predicted BPU 208 and predicted data 206.
[0052]
[0069] In the reconstruction path of process 200B, when the intra prediction mode is selected within the forward path, after generating a prediction reference 224 (e.g., the current BPU that is encoded and reconstructed within the current picture), the encoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). When the inter prediction mode is selected within the forward path, after generating a prediction reference 224 (e.g., the current picture in which all BPUs are encoded and reconstructed), the encoder can feed the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) caused by inter prediction. The encoder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference picture can be stored in a buffer 234 (or "reconstructed picture buffer") for later use (e.g., for use as an inter prediction reference picture for future pictures of video sequence 202). The encoder can store one or more reference pictures in buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode the parameters of the loop filter (e.g., the strength of the loop filter) at the binary coding stage 226 along with the quantized transform coefficients 216, prediction data 206, and other information.
[0053]
[0070] FIG. 3A shows a schematic diagram of an example of a decoding process 300A according to some embodiments of the present disclosure. The process 300A can be a decompression process corresponding to the compression process 200A of FIG. 2A. In some embodiments, the process 300A can be similar to the reconstruction path of the process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 of FIGS. 2A-2B), generally the video stream 304 is not identical to the video sequence 202. Similar to the processes 200A and 200B of FIGS. 2A-2B, the decoder can execute the process 300A at the level of the basic processing unit (BPU) for each picture encoded in the video bitstream 228. For example, the decoder can execute the process 300A in an iterative manner, and the decoder can decode the basic processing unit within one iteration of the process 300A. In some embodiments, the decoder can execute the process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded in the video bitstream 228.
[0054]
[0071] In FIG. 3A, the decoder can feed a part of the video bitstream 228 related to the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode that part into predicted data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate the reconstructed residual BPU 222. The decoder can feed the predicted data 206 to the prediction stage 204 to generate the predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate the predicted reference 224. In some embodiments, the predicted reference 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can feed the predicted reference 224 for performing a prediction operation in the next iteration of process 300A to the prediction stage 204.
[0055]
[0072] The decoder can repeatedly execute process 300A to decode each encoded BPU of the encoded picture and generate a predicted reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all the encoded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.
[0056]
[0073] In the binary decoding stage 302, the decoder can perform the reverse operation of the binary coding technique (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other reversible compression algorithm) used by the encoder. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the decoder can decode other information such as, for example, the prediction mode, the parameters of the prediction operation, the type of transformation, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), etc. in the binary decoding stage 302. In some embodiments, when the video bitstream 228 is transmitted in packet units over the network, the decoder can depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.
[0057]
[0074] FIG. 3B shows a schematic diagram of another example 300B of the decoding process according to some embodiments of the present disclosure. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0058]
[0075] In Process 300B, for an encoded basic processing unit of the decoded encoded picture (referred to as the "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 may include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, if intra prediction was used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the position (e.g., coordinates) of one or more adjacent BPUs used as a reference, the size of the adjacent BPUs, extrapolation parameters, the direction of the adjacent BPUs with respect to the original BPU, etc. In another example, if inter prediction was used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. The parameters of the inter prediction operation may include, for example, the number of reference pictures related to the current BPU, the weights respectively related to the reference pictures, the position (e.g., coordinates) of one or more matching regions in each reference picture, one or more motion vectors respectively related to the matching regions, etc.
[0059]
[0076] Based on the prediction mode indicator, the decoder can determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of the execution of such spatial or temporal prediction are shown in Figure 2B and will not be repeated here. After performing such spatial or temporal prediction, the decoder can generate the predicted BPU 208. As described in Figure 3A, the decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate the prediction reference 224.
[0060]
[0077] In process 300B, the decoder can feed the predicted reference 224 for performing a prediction operation within the next iteration of process 300B to the spatial prediction stage 2042 or the temporal prediction stage 2044. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop-filtered reference picture can be stored in the buffer 234 (e.g., the decoded picture buffer in computer memory) for later use (e.g., for use as an inter prediction reference picture for future encoded pictures of the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter prediction is used to encode the current BPU, the prediction data can further include loop filter parameters (e.g., the strength of the loop filter).
[0061]
[0078] FIG. 4 is a block diagram of an example of a device 400 for encoding or decoding video according to some embodiments of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 can become a dedicated machine for encoding or decoding video. The processor 402 can be any kind of circuit that can manipulate or process information. For example, the processor 402 can include any combination of any number of central processing units ("CPUs"), graphics processing units ("GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, IP (intellectual property) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), system on chips (SoCs), application specific integrated circuits (ASICs), etc. In some embodiments, the processor 402 can also be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, the processor 402 can include a plurality of processors including processor 402a, processor 402b, and processor 402n.
[0062]
[0079] The machine 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for implementing steps within processes 200A, 200B, 300A, or 300B) and processing data (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and the processing data (e.g., via bus 410) and execute the program instructions to perform operations or processing on the processing data. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 can also be a memory bank (not shown in FIG. 4) grouped as a single logical component.
[0063]
[0080] A bus 410, such as an internal bus (e.g., a CPU memory bus) or an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), can be a communication device that transfers data between components within the device 400.
[0064]
[0081] For the sake of simplicity of explanation without causing ambiguity, in the present disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuits can be implemented entirely as hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuits can be a single independent module or can be fully or partially combined within any other component of the device 400.
[0065]
[0082] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (such as the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, and the like.
[0066]
[0083] In some embodiments, device 400 may optionally further include a peripheral device interface 408 for providing connection to one or more peripheral devices. As shown in FIG. 4, the peripheral devices may include, but are not limited to, a cursor control device (such as a mouse, touch pad, or touch screen), a keyboard, a display (such as a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (such as a camera or input interface coupled to a video archive), and the like.
[0067]
[0084] It should be noted that the video codec (such as the codec that executes processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within device 400. For example, some or all of the steps of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. In another example, some or all of the steps of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (such as FPGAs, ASICs, NPUs, etc.).
[0068]
[0085] Figure 5 shows a schematic diagram of an exemplary luma mapping (LMCS) process 500 by chroma scaling according to some embodiments of the present disclosure. For example, process 500 may be used by a decoder compliant with a hybrid video coding standard (e.g., H.26x series). LMCS is a new processing block applied before the loop filter 232 of FIG. 2B. LMCS can also be called a reshaper.
[0069]
[0086] The LMCS process 500 may include in-loop mapping of luma component values based on an adaptive segmented linear model and luma-dependent chroma residual scaling of chroma components.
[0070]
[0087] As shown in FIG. 5, the in-loop mapping of luma component values based on an adaptive segmented linear model may include a forward mapping stage 518 and an inverse mapping stage 508. The luma-dependent chroma residual scaling of chroma components may include chroma scaling 520.
[0071]
[0088] The sample values before mapping or after inverse mapping can be called samples within the original region, and the sample values after mapping and before inverse mapping can be called samples within the map region. When LMCS is enabled, some stages within process 500 can be performed within the map region instead of the original region. It will be understood that the forward mapping stage 518 and the inverse mapping stage 508 can be enabled / disabled at the sequence level using an SPS flag.
[0072]
[0089] As shown in FIG. 5, Q -1 &T -1 stage 504, reconstruction 506, and intra prediction 508 are executed within the map region. For example, Q -1 &T -1 stage 504 may include inverse quantization and inverse transformation, reconstruction 506 may include addition of luma prediction and luma residual, and intra prediction 508 may include luma intra prediction.
[0073]
[0090] The loop filter 510, motion compensation stages 516 and 530, intra prediction stage 528, reconstruction stage 522, and decoded picture buffers (DPBs) 512 and 526 are executed within the original (i.e., non-map) region. In some embodiments, the loop filter 510 can include deblocking, adaptive loop filter (ALF), and sample adaptive offset (SAO), the reconstruction stage 522 can include addition of chroma prediction and chroma residual, and the DPBs 512 and 526 can store decoded pictures as reference pictures.
[0074]
[0091] In some embodiments, luma mapping using a segmented linear model can be applied.
[0075]
[0092] In-loop mapping of the luma component can adjust the signal statistics of the input video by redistribution of the codewords over the dynamic range to improve the compression efficiency. The luma mapping can be executed by a forward mapping function “FwdMap” and a corresponding inverse mapping function “InvMap”. The “FwdMap” function is signaled using a segmented linear model with 16 equal segments. The “InvMap” function need not be signaled and is instead derived from the “FwdMap” function.
[0076]
[0093] The signaling of the differential linear model is shown in Table 1 of FIG. 6 and Table 2 of FIG. 7. Table 1 of FIG. 6 shows the tile group header syntax structure. As shown in FIG. 6, a reshaper model parameter presence flag is signaled to indicate whether there is a luma mapping model within the current tile group. If there is a luma mapping model within the current tile group, the syntax elements shown in Table 2 of FIG. 7 can be used to signal the corresponding differential linear model parameters within tile_group_reshaper_model(). The differential linear model divides the dynamic range of the input signal into 16 equal segments. For each of the 16 equal segments, the linear mapping parameters of the segment are represented using the number of codewords assigned to the segment. Taking a 10-bit input as an example. Each of the 16 segments can have 64 codewords assigned to that segment by default. The number of codewords signaled can be used to calculate the scale factor and adjust the mapping function for that segment as appropriate. Table 2 of FIG. 7 also comprehensively defines the minimum index "reshaper_model_min_bin_idx" and the maximum index "reshaper_model_max_bin_idx" for which the number of codewords is signaled. If the segment index is less than reshaper_model_min_bin_idx or greater than reshaper_model_max_bin_idx, the number of codewords for that segment is not signaled and is inferred to be zero (i.e., no codewords are assigned to that segment and no mapping / scaling is applied).
[0077]
[0094] After the `tile_group_reshaper_model()` is signaled, another reshaper enable flag, `tile_group_reshaper_enable_flag`, is signaled at the tile group header level to indicate whether the LMCS process shown in FIG. 8 is applied to the current tile group. If the reshaper is enabled for the current tile group and the current tile group does not use dual-tree splitting, a further chroma scaling enable flag is signaled to indicate whether chroma scaling is enabled for the current tile group. Dual-tree splitting can also be referred to as a chroma separate tree.
[0078]
[0095] The piecewise linear model can be constructed as follows based on the signaled syntax elements in Table 2 of FIG. 7. Each i-th segment, i = 0, 1, ..., 15, of the "FwdMap" piecewise linear model is defined by two input pivot points InputPivot[] and two output (mapped) pivot points MappedPivot[]. InputPivot[] and MappedPivot[] are calculated as follows based on the signaled syntax (assuming without loss of generality that the bit depth of the input video is 10 bits): 1) OrgCW = 64 2) For i = 0:16, InputPivot[i] = i * OrgCW 3) For i = reshaper_model_min_bin_idx: reshaper_model_max_bin_idx, SignaledCW[i] = OrgCW + (1 ¬ 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i]; 4) For i = 0:16, MappedPivot[i] is calculated as follows: MappedPivot[0] = 0; (For i = 0; i < 16; i++) MappedPivot[i + 1] = MappedPivot[i] + SignaledCW[i]
[0079]
[0096] The inverse mapping function "InvMap" can also be defined by InputPivot[] and MappedPivot[]. Different from "FwdMap", in the piecewise linear model of "InvMap", two input pivot points for each segment can be defined by MappedPivot[], and two output pivot points can be defined by InputPivot[], which is the opposite of "FwdMap". In this way, the input of "FwdMap" is divided into equal segments, but it is not guaranteed that the input of "InvMap" is divided into equal segments.
[0080]
[0097] As shown in FIG. 5, in the inter-coded block, motion compensation prediction can be executed within the mapped area. In other words, after the motion compensation prediction 516, Y can be calculated based on the reference signal in the DPB, and the "FwdMap" function 518 can be applied to map the luma prediction block in the original area to the mapped area (Y' pred = FwdMap(Y pred )). In the intra-coded block, since the reference samples used in the intra prediction already exist in the mapped area, the "FwdMap" function is not applied. After the reconstructed block 506, Y pred can be calculated. The "InvMap" function 508 can be applied to convert and return the reconstructed luma value in the mapped area to the reconstructed luma value in the original area ( r ). The "InvMap" function 508 can be applied to both intra-coded luma blocks and inter-coded luma blocks.
Number
[0081]
[0098] The luma mapping process (forward or inverse mapping) can be implemented using a look-up table (LUT) or by on-the-fly calculation. When an LUT is used, tables "FwdMapLUT[]" and "InvMapLUT[]" can be pre-calculated and pre-stored for use at the tile group level, and forward mapping and inverse mapping can be implemented simply as FwdMap(Y pred ) = FwdMapLUT[Y pred , and InvMap(Y r ) = InvMapLUT[Y r respectively. Alternatively, on-the-fly calculation can be used. Taking the forward mapping function "FwdMap" as an example. To determine the segment to which a luma sample belongs, the sample value can be right-shifted by 6 bits (corresponding to 16 equal segments assuming a 10-bit video) to obtain a segment index. Then the linear model parameters for that segment can be retrieved and applied on-the-fly to calculate the mapped luma value. The FwdMap function is evaluated as follows: Y’pred = FwdMap(Y pred ) = ((b2 - b1) / (a2 - a1))*(Y pred - a1) + b1 where "i" is the segment index, a1 is InputPivot[i], a2 is InputPivot[i + 1], b1 is MappedPivot[i], and b2 is MappedPivot[i + 1].
[0082]
[0099] The "InvMap" function can be calculated on-the-fly in a similar way, except that a conditional check must be applied when finding the segment to which a sample value belongs, because it is not guaranteed that the segments within the mapped area are of equal size.
[0083]
[0100] In some embodiments, luma-dependent chroma residual scaling can be performed.
[0084]
[0101] Chroma residual scaling is designed to compensate for the interaction between the luma signal and its corresponding chroma signal. Whether chroma residual scaling is enabled is also signaled at the tile group level. As shown in Table 1 of FIG. 6, when luma mapping is enabled and dual-tree splitting is not applied to the current tile group, an additional flag (e.g., tile_group_reshaper_chroma_residual_scale_flag) is signaled to indicate whether luma-dependent chroma residual scaling is enabled. When luma mapping is not used or dual-tree splitting is used within the current tile group, luma-dependent chroma residual scaling is automatically disabled. Further, luma-dependent chroma residual scaling can be disabled for chroma blocks whose area is 4 or less.
[0085]
[0102] Chroma residual scaling depends on the average value of the corresponding luma prediction block (for both intra-coded blocks and inter-coded blocks). The avgY’ as the average of the luma prediction block can be calculated as follows:
Number
[0086]
[0103] C ScaleInv The value of is calculated in the following steps: 1) Find the index Y of the piecewise linear model to which avgY’ belongs based on the InvMap function. Idx using the InvMap function. 2) C ScaleInv = cScaleInv[Y Idx ], where cScaleInv[] is a pre-computed 16-piece LUT.
[0087]
[0104] In the current LMCS method in VTM4, the pre-calculated LUT cScaleInv[i] (where i is in the range of 0 to 15) is derived as follows based on the 64-entry static LUT ChromaResidualScaleLut and the value of SignaledCW[i]: ChromaResidualScaleLut
[64] ={16384, 16384, 16384, 16384, 16384, 16384, 16384, 8192, 8192, 8192, 8192, 5461, 5461, 5461, 5461, 4096, 4096, 4096, 4096, 3277, 3277, 3277, 3277, 2731, 2731, 2731, 2731, 2341, 2341, 2341, 2048, 2048, 2048, 1820, 1820, 1820, 1638, 1638, 1638, 1638, 1489, 1489, 1489, 1489, 1365, 1365, 1365, 1365, 1260, 1260, 1260, 1260, 1170, 1170, 1170, 1170, 1092, 1092, 1092, 1092, 1024, 1024, 1024, 1024}; shiftC = 11 - If -(SignaledCW[i] == 0) holds, cScaleInv[i]=(1 << shiftC) - Otherwise cScaleInv[i]=ChromaResidualScaleLut[(SignaledCW[i] >> 1)-1]
[0088]
[0105] The static table ChromaResidualScaleLut[] contains 64 entries and SignaledCW[] is in the range of [0, 128] (assuming the input is 10 bits). Therefore, division by 2 (e.g., right shift by 1) is used to construct cScaleInv[] of the chroma scale factor LUT. The cScaleInv[] of the chroma scale factor LUT can contain multiple chroma scale factors. The cScaleInv[] of the LUT is constructed at the tile group level.
[0089]
[0106] When the current block is encoded using the intra, CIIP, or intra block copy (IBC, also known as current picture reference or CPR) mode, avgY’ is calculated as the average of the intra, CIIP, or IBC predicted luma values. Otherwise, avgY’ is calculated as the average of the inter predicted luma values mapped forward (i.e., Y’ in Figure 5 pred ). Different from the luma mapping performed based on samples, C ScaleInv is a constant value for all chroma blocks. C ScaleInv is used to apply chroma residual scaling on the decoder side as follows:
Number
[0090]
[0107] However
Number
[0091]
[0108] In some embodiments, dual tree splitting can be performed.
[0092]
[0109] In Draft 4 of VVC, the coding tree method supports the ability of luma and chroma to have separate block tree partitions. This is also called dual tree partitioning. The signaling of dual tree partitioning is shown in Table 3 of FIG. 8 and Table 4 of FIG. 9. When the sequence level control flag "qtbtt_dual_tree_intra_flag", which is signaled in the SPS, is turned on and the current tile group is intra-coded, the block partition information can be signaled separately for luma first and then for chroma. Dual tree partitioning is not permitted for tile groups that are inter-coded (P and B tile groups). When a separate block tree mode is applied, as shown in Table 4 of FIG. 9, the luma coding tree block (CTB) is divided into coding units (CUs) by a certain coding tree structure, and the chroma CTB is divided into chroma CUs by another coding tree structure.
[0093]
[0110] When different splits of luma and chroma are allowed, problems can arise with coding tools that have dependencies between various color components. For example, in the case of LMCS, to find the scale factor applied to the current block, the average value of the corresponding luma block is used. When a dual tree is used, this can lead to latency for the entire CTU. For example, if the luma blocks of a CTU are split vertically once and the chroma blocks of the CTU are split horizontally once, both luma blocks of the CTU are decoded before the first chroma block of the CTU can be decoded (thereby allowing the calculation of the average value necessary to calculate the chroma scale factor). In VVC, a CTU can be as large as 128x128 in units of luma samples. Such a large latency can be a significant problem for the pipeline design of a hardware decoder. Therefore, Draft 4 of VVC can prohibit the combination of dual tree splitting and luma-dependent chroma scaling. If dual tree splitting is enabled for the current tile group, chroma scaling can be forced off. Note that the luma mapping part of LMCS only affects the luma component and does not have dependency problems across color components, so it continues to be allowed even in the case of a dual tree. Another example of a coding tool that relies on dependencies between color components to achieve better coding efficiency is called the cross-component linear model (CCLM).
[0094]
[0111] Therefore, the derivation of cScaleInv[] of the chroma scale factor LUT at the tile group level cannot be easily extended. The derivation process currently depends on a fixed chroma LUT ChromaResidualScaleLut with 64 entries. For a 10-bit video with 16 segments, an additional step of division by 2 must be applied. If the number of segments changes, for example, if 8 segments are used instead of 16 segments, the derivation process must be changed to apply division by 4 instead of division by 2. This additional step not only causes a loss of precision but is also unrefined and unnecessary.
[0095]
[0112] Further, for calculating the partition index Y of the current chroma block used to obtain the chroma scale factor, the average value of all luma blocks can be used. This is undesirable and generally unnecessary. Consider the maximum CTU size of 128x128. In this case, the average luma value is calculated based on 16384 (128x128) luma samples, and such calculation is complex. Further, when a 128x128 luma block partition is selected by the coder, the block is likely to contain homogeneous content. Therefore, a subset of the luma samples within the block may be sufficient to calculate the luma average. Idx To calculate the partition index Y of the current chroma block used to obtain the chroma scale factor, the average value of all luma blocks can be used. This is undesirable and generally unnecessary. Consider the maximum CTU size of 128x128. In this case, the average luma value is calculated based on 16384 (128x128) luma samples, and such calculation is complex. Further, when a 128x128 luma block partition is selected by the coder, the block is likely to contain homogeneous content. Therefore, a subset of the luma samples within the block may be sufficient to calculate the luma average.
[0096]
[0113] In dual-tree partitioning, chroma scaling can be off to avoid potential pipeline problems of the hardware decoder. However, this dependency can be avoided if, instead of using the corresponding luma samples to derive the applied chroma scale factor, explicit signaling for indicating the applied chroma scale factor is used. Enabling chroma scaling within the tile group to be intra-coded can further improve the coding efficiency.
[0097]
[0114] The signaling of the partition linear parameters can be further improved. Currently, delta codeword values are signaled for each of the 16 partitions. It has been observed that often only a limited number of different codewords are used for these 16 partitions. Therefore, the signaling overhead can be further reduced.
[0098]
[0115] Embodiments of the present disclosure provide a method for processing video content by removing the chroma scaling LUT.
[0099]
[0116] As described above, the expansion of a 64 - entry chroma LUT can be difficult and can be a problem when other partitioning linear models are used (e.g., 8 - partition, 4 - partition, 64 - partition, etc.). To achieve the same coding efficiency, such an expansion is not necessary because the chroma scale factor can be set to the same as the corresponding luma scale factor of the partition. In some embodiments of the present disclosure, the chroma scale factor (chroma_scaling) can be determined based on the partition index "Y" of the current chroma block as follows. Idx ". · When Y Idx >reshaper_model_max_bin_idx, Y Idx <reshaper_model_min_bin_idx or SignaledCW[Y Idx =0 holds, set chroma_scaling to the default and chroma_scaling = 1.0. · Otherwise, set chroma_scaling to SignaledCW[Y Idx / OrgCW.
[0100]
[0117] When chroma_scaling = 1.0, no scaling is applied.
[0101]
[0118] The chroma scale factor determined above can have fractional precision. It will be understood that a fixed - point approximation can be applied to avoid dependencies on the hardware / software platform. Further, inverse chroma scaling can be performed on the decoder side. Therefore, division can be implemented by a fixed - point operation using multiplication followed by a right shift. The inverse chroma scale factor "inverse_chroma_scaling[]" in fixed - point precision can be determined as follows based on the number of bits within the fixed - point approximation "CSCALE_FP_PREC". inverse_chroma_scaling[Y Idx=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(SignaledCW[Y Idx >>1)) / SignaledCW[Y Idx Here, luma_bit_depth is the luma bit depth, and TOTAL_NUMBER_PIECES is the total number of segments in the partitioned linear model set to 16 in VVC Draft 4. The value of "inverse_chroma_scaling[]" may only need to be calculated once per tile group, and it should be understood that the above division is an integer division operation.
[0102]
[0119] Further quantization can be applied to determine the chroma scale factor and the inverse scale factor. For example, the inverse chroma scale factor can be calculated for all even (2×m) values of "SignaledCW", and the odd (2×m + 1) values of "SignaledCW" reuse the chroma scale factor of the adjacent even value's scale factor. In other words, the following can be used: for(i=reshaper_model_min_bin_idx; i<=reshaper_model_max_bin_idx; i++) { tempCW=SignaledCW[i]>>1)<<1; inverse_chroma_scaling[i]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(tempCW>>1)) / tempCW; }
[0103]
[0120] The quantization of the chroma scale factor can be further generalized. For example, the inverse chroma scale factor "inverse_chroma_scaling[]" can be calculated for each n-th value of "SignaledCW" while all other adjacent values share the same chroma scale factor. For example, "n" can be set to 4. Thus, the same inverse chroma scale factor value can be shared for every four adjacent codeword values. In some embodiments, the value of "n" can be a power of 2, and such a setting enables the use of shifts to calculate divisions. Representing the value of log2(n) as LOG2_n, the above equation "tempCW = SignaledCW[i] >> 1) << 1" can be modified as follows: tempCW = SignaledCW[i] >> LOG2_n) << LOG2_n
[0104]
[0121] In some embodiments, the value of LOG2_n can be a function of the number of segments used within the piecewise linear model. When fewer segments are used, it may be beneficial to use a larger LOG2_n. For example, when TOTAL_NUMBER_PIECES is 16 or less, LOG2_n can be set to 1 + (4 - log2(TOTAL_NUMBER_PIECES)). When TOTAL_NUMBER_PIECES exceeds 16, LOG2_n can be set to 0.
[0105]
[0122] Embodiments of the present disclosure provide a method for processing video content by simplifying the averaging of luma prediction blocks.
[0106]
[0123] As discussed above, the average value of the corresponding luma block can be used to determine the segmentation index "Y Idx " of the current chroma block. However, for large block sizes, the averaging process can include a large number of luma samples. In the worst case, 128x128 luma samples can be involved in the averaging process.
[0107]
[0124] Embodiments of the present disclosure provide a simplified averaging process to reduce using only a few NxN luma samples (where N is a power of 2) in the worst case.
[0108]
[0125] In some embodiments, if both dimensions of the two-dimensional luma block are less than or equal to a preset threshold M (in other words, at least one of the two dimensions exceeds M), "downsampling" can be applied to use only the positions of M within that dimension. Taking the horizontal dimension as an example without loss of generality, if the width exceeds M, only the samples at positions x (x = i×(width>>log2(M)), i = 0, … M-1) are used for averaging.
[0109]
[0126] FIG. 10 shows an example of applying the proposed simplification to calculate the average of a 16x8 luma block. In this example, M is set to 4, and only 16 luma samples (shaded samples) within the block are used for averaging. It will be understood that the preset threshold M is not limited to 4, and M can be set to any value that is a power of 2. For example, the preset threshold M can be 1, 2, 4, 8, etc.
[0110]
[0127] In some embodiments, the horizontal and vertical dimensions of the luma block can have different preset thresholds M. In other words, in the worst case of the averaging operation, M1xM2 samples can be used.
[0111]
[0128] In some embodiments, the number of samples in the averaging process can be limited without considering the dimensions. For example, up to 16 samples can be used, and those samples can be distributed within the horizontal or vertical dimension in a 1x16, 16x1, 2x8, 8x2, or 4x4 format, and any format that fits the shape of the current block can be selected. For example, if the block is vertically long, a matrix of 2x8 samples can be used, if the block is horizontally long, a matrix of 8x2 samples can be used, and if the block is square, a matrix of 4x4 samples can be used.
[0112]
[0129] When the size of the large block is selected, it will be understood that the content within the block tends to be more homogeneous. Thus, the above simplification can cause a difference between the average value and the true average of all luma blocks, but such a difference can be small.
[0113]
[0130] Further, decoder-side motion vector refinement (DMVR) requires the decoder to perform motion detection to derive motion vectors before motion compensation can be applied. Thus, the DMVR mode can be particularly complex within the VVC standard for the decoder. Bidirectional optical flow (BDOF) is an additional sequential process that must be applied after DMVR to obtain the luma prediction block, so the BDOF mode within the VVC standard can further complicate this situation. Chroma scaling requires the average value of the corresponding luma prediction block, so DMVR and BDOF can be applied before the average value can be calculated.
[0114]
[0131] To solve this latency problem, in some embodiments of the present disclosure, the average luma value is calculated using the luma prediction block before DMVR and BDOF, and the average luma value is used to obtain the chroma scale factor. This enables chroma scaling to be applied in parallel with the DMVR and BDOF processes, thus significantly reducing latency.
[0115]
[0132] In accordance with the present disclosure, modified forms of latency reduction can be considered. In some embodiments, this latency reduction can also be combined with the above-described simplified averaging process that uses only a part of the luma prediction block to calculate the average luma value. In some embodiments, the luma prediction block can be used after the DMVR process and before the BDOF process to calculate the average luma value. Then, the average luma value is used to obtain the chroma scale factor. This design enables the application of chroma scaling in parallel with the BDOF process while maintaining the accuracy of determining the chroma scale factor. Since the DMVR process may refine the motion vectors, it may be more accurate to use the predicted samples with the refined motion vectors after the DMVR process rather than with the motion vectors before the DMVR process.
[0116]
[0133] Furthermore, in the VVC standard, the CU syntax structure “coding_unit()” includes a syntax element “cu_cbf” for indicating whether there are any non-zero residual coefficients in the current CU. At the TU level, the TU syntax structure “transform_unit()” includes syntax elements “tu_cbf_cb” and “tu_cbf_cr” for indicating whether there are any non-zero chroma (Cb or Cr) residual coefficients in the current TU. Conventionally, in draft 4 of VVC, when chroma scaling is enabled at the tile group level, the averaging of the corresponding luma block is always called.
[0117]
[0134] Embodiments of the present disclosure further provide a method for processing video content by bypassing the luma averaging process. In accordance with the disclosed embodiments, since the chroma scaling process is applied to the residual chroma coefficients, the luma averaging process can be bypassed when there are no non-zero chroma coefficients. This can be determined based on the following conditions: Condition 1: cu_cbf is equal to 0 Condition 2: both tu_cbf_cr and tu_cbf_cb are equal to 0
[0118]
[0135] As discussed above, "cu_cbf" can indicate whether there are any non-zero residual coefficients in the current CU, and "tu_cbf_cb" and "tu_cbf_cr" can indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the current TU. When Condition 1 or Condition 2 is satisfied, the luma averaging process can be bypassed.
[0119]
[0136] In some embodiments, only the NxN samples of the prediction block are used to derive the average value, which simplifies the averaging process. For example, when N is equal to 1, only the top-left sample of the prediction block is used. However, this simplified averaging process using the prediction block still requires the generation of the prediction block, thereby causing latency.
[0120]
[0137] In some embodiments, the reference luma samples can be directly used to generate the chroma scale factor. This enables the decoder to derive the scale factor in parallel with the luma prediction process, thereby reducing latency. Intra prediction and inter prediction using the reference luma samples are described separately below.
[0121]
[0138] In an exemplary intra prediction, the decoded adjacent samples within the same picture can be used as reference samples for generating the prediction block. These reference samples can include, for example, the samples above the current block, the samples to the left of the current block, or the top-left samples of the current block. The average of these reference samples can be used to derive the chroma scale factor. In some embodiments, the average of some of these reference samples can be used. For example, only the K reference samples (e.g., K = 3) closest to the top-left position of the current block are averaged.
[0122]
[0139] In an exemplary inter prediction, reference samples from a temporal reference picture can be used to generate a prediction block. These reference samples are identified by a reference picture index and a motion vector. When the motion vector has fractional precision, interpolation can be applied. The reference samples used to determine the average of the reference samples can include the reference samples before or after interpolation. The reference samples before interpolation can include motion vectors that are clipped to integer precision. Consistent with the disclosed embodiments, the average can be calculated using all of the reference samples. Alternatively, the average can be calculated using only a portion of the reference samples (e.g., the reference samples corresponding to the upper left position of the current block).
[0123]
[0140] As shown in FIG. 5, while performing inter prediction within the original region, intra prediction (e.g., intra prediction 514 or 528) can be performed within the reshaped region. Accordingly, forward mapping can be applied to the prediction block in inter prediction, and the average is calculated using the luma prediction block after forward mapping. To reduce latency, the average can be calculated using the prediction block before forward mapping. For example, the block before forward mapping, the NxN portion of the block before forward mapping, or the upper left sample of the block before forward mapping can be used.
[0124]
[0141] Embodiments of the present disclosure further provide a method for processing video content with chroma scaling for dual tree splitting.
[0125]
[0142] Since the dependence on luma can cause the complexity of the hardware design to increase, chroma scaling can be turned off for the intra-coded tile groups that enable dual tree splitting. However, this limitation can cause a loss of coding efficiency. The sample values of the corresponding luma blocks are averaged to calculate avgY’, and the segmentation index Y Idx is determined, and the chroma scale factor inverse_chroma_scaling[YIdx Instead of obtaining [], the chroma scale factor can be explicitly signaled within the bitstream to avoid the dependency on luma in the case of dual-tree splitting.
[0126]
[0143] The chroma scale index can be signaled at various levels. For example, as shown in Table 5 of FIG. 11, the chroma scale index can be signaled at the coding unit (CU) level together with the chroma prediction mode. To determine the chroma scale factor of the current chroma block, the syntax element "lmcs_scaling_factor_idx" can be used. If there is no "lmcs_scaling_factor_idx", it can be inferred that the chroma scale factor of the current chroma block is equal to 1.0 with floating-point precision or equivalently (1<<CSCALE_FP_PREC) with fixed-point precision. The range of allowable values of "lmcs_chroma_scaling_idx" is determined at the tile group level and will be described later.
[0127]
[0144] Depending on the possible values of 「lmcs_chroma_scaling_idx」, the signaling cost can be high, especially for small blocks. Thus, in some embodiments of the present disclosure, the signaling conditions in Table 5 of FIG. 11 may additionally include block size conditions. For example (emphasized in italics and gray shading), this syntax element 「lmcs_chroma_scaling_idx」 can be signaled only if the current block contains more than a given number of chroma samples, or if the current block has a width greater than a given width W or a height greater than a given height H. For smaller blocks, if 「lmcs_chroma_scaling_idx」 is not signaled, the decoder side can determine the chroma scale factor. In some embodiments, the chroma scale factor can be set to 1.0 with floating-point precision. In some embodiments, a default 「lmcs_chroma_scaling_idx」 value can be added at the tile group header level (see 1 in FIG. 6). Small blocks without a signaled 「lmcs_chroma_scaling_idx」 can use this tile group level default index to derive the corresponding chroma scale factor. In some embodiments, the chroma scale factor of a small block can be inherited from its vicinity (e.g., the vicinity of the top or left side) where the scale factor is explicitly signaled.
[0128]
[0145] In addition to signaling the syntax element "lmcs_chroma_scaling_idx" at the CU level, this syntax element can also be signaled at the CTU level. However, given that the maximum CTU size in VVC is 128x128, performing the same scaling at the CTU level may be too coarse. Therefore, in some embodiments of the present disclosure, this syntax element "lmcs_chroma_scaling_idx" can be signaled using a fixed granularity. For example, one "lmcs_chroma_scaling_idx" is signaled for each 16x16 region within the CTU and applied to all samples within that 16x16 region.
[0129]
[0146] The range of "lmcs_chroma_scaling_idx" for the current tile group depends on the number of values of the chroma scale factor permitted within the current tile group. The number of values of the chroma scale factor permitted within the current tile group can be determined based on the 64-entry chroma LUT as discussed above. Alternatively, the number of values of the chroma scale factor permitted within the current tile group can be determined using the chroma scale factor calculation discussed above.
[0130]
[0147] For example, in the "quantization" method, the value of LOG2_n can be set to 2 (i.e., "n" is set to 4), and the codeword assignment for each segment in the piecewise linear model of the current tile group can be set as follows: {0, 65, 66, 64, 67, 62, 62, 64, 64, 64, 67, 64, 64, 62, 61, 0}. Then any codeword value from 64 to 67 can have the same scale factor value (1.0 with fractional precision), and any codeword value from 60 to 63 can have the same scale factor value (60 / 64 = 0.9375 with fractional precision). So there are only two possible scale factor values for all tile groups. In the two end segments where no codewords are assigned, the chroma scale factor is set to 1.0 by default. Therefore, in this example, 1 bit is sufficient to signal "lmcs_chroma_scaling_idx" for the blocks within the current tile group.
[0131]
[0148] In addition to determining the possible chroma scale factor values using the piecewise linear model, the encoder can signal a set of chroma scale factor values in the tile group header. Then at the block level, the value of that set of chroma scale factor values and the value of "lmcs_chroma_scaling_idx" of the block can be used to determine the chroma scale factor value of the block.
[0132]
[0149] CABAC coding can be applied to code 「lmcs_chroma_scaling_idx」. The CABAC context of a block may depend on the 「lmcs_chroma_scaling_idx」 of the adjacent blocks of the block. For example, the left block or the upper block can be used to form the CABAC context. Regarding the binarization of this syntax element of 「lmcs_chroma_scaling_idx」, the same truncated Rice binarization applied to the ref_idx_10 and ref_idx_11 syntax elements in Draft 4 of VVC can be used to binarize 「lmcs_chroma_scaling_idx」.
[0133]
[0150] The advantage of signaling 「chroma_scaling_idx」 is that the coder can select the best 「lmcs_chroma_scaling_idx」 regarding the rate distortion cost. Using rate distortion optimization to select 「lmcs_chroma_scaling_idx」 can improve the coding efficiency, which can help offset the increase in signaling cost.
[0134]
[0151] Embodiments of the present disclosure further provide a method for processing video content involving the signaling of an LMCS partitioned linear model.
[0135]
[0152] The LMCS method uses a partitioned linear model with 16 partitions, but the number of unique values of 「SignaledCW[i]」 within a tile group tends to be much less than 16. For example, some of the 16 partitions can use the default number of codewords 「OrgCW」, and some of the 16 partitions can have the same number of codewords as each other. Therefore, an alternative method of signaling the LMCS partitioned linear model may include signaling the number of unique codewords 「listUniqueCW[]」 and transmitting an index for each of the partitions to indicate the elements of 「listUniqueCW[]」 of the current partition.
[0136]
[0153] Shows the corrected syntax table in FIG. 12. In Table 6 of FIG. 12, new or corrected syntax is emphasized in italics and gray shading.
[0137]
[0154] The semantic rules of the disclosed signaling method are as follows, with the changed parts underlined: reshaper_model_min_bin_idx defines the minimum bin (or section) index used within the reshaper construction process. The value of reshape_model_min_bin_idx shall be within the range from 0 to MaxBinIdx. The value of MaxBinIdx shall be equal to 15. reshaper_model_delta_max_bin_idx defines the maximum bin index used within the reshaper construction process minus the maximum allowable bin (or section) index MaxBinIdx. The value of reshape_model_max_bin_idx is set to be equal to MaxBinIdx - reshape_model_delta_max_bin_idx. reshaper_model_bin_delta_abs_cw_prec_minus1 plus 1 defines the number of bits used for the representation of the syntax reshape_model_bin_delta_abs_CW[i]. reshaper_model_bin_num_unique_cw_minus1 plus 1 defines the size of the codeword array listUniqueCW. reshaper_model_bin_delta_abs_CW[i] defines the absolute delta codeword value of the i-th bin. reshaper_model_bin_delta_sign_CW_flag[i] defines the sign of reshape_model_bin_delta_abs_CW[i] as follows: - When reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is a positive value. - Otherwise (if reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is negative. If reshape_model_bin_delta_sign_CW_flag[i] does not exist, it is inferred that the corresponding variable RspDeltaCW[i] is equal to 0. The variable RspDeltaCW[i] is derived as RspDeltaCW[i] = (1 - 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i]. The variable listUniqueCW[0] is set equal to OrgCW. The variable listUniqueCW[i] for i = 1... reshaper_model_bin_num_unique_cw_minus1 is as follows It is derived as follows: - Set the variable OrgCW equal to (1 << BitDepth Y ) / (MaxBinIdx + 1). - listUniqueCW[i] = OrgCW + RspDeltaCW[i - 1] reshaper_model_bin_cw_idx[i] defines the index of the array listUniqueCW[] used to derive RspCW[i]. The value of reshaper_model_bin_cw_idx[i] shall be within the range from 0 to (reshaper_model_bin_num_unique_cw_minus1 + 1). RspCW[i] is derived as follows: - If reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx holds, RspCW[i] = listUniqueCW[reshaper_model_bin_cw_idx[i]]. - Otherwise RspCW[i] = 0. BitDepth Y When the value of is equal to 10, the value of RspCW[i] can be in the range of 32 to 2 * OrgCW - 1.
[0138]
[0155] Embodiments of the present disclosure further provide a method for processing video content with conditional chroma scaling at the block level.
[0139]
[0156] As shown in Table 1 of FIG. 6, whether chroma scaling is applied can be determined by the "tile_group_reshaper_chroma_residual_scale_flag" signaled at the tile group level.
[0140]
[0157] However, it may be beneficial to determine whether to apply chroma scaling at the block level. For example, in some of the disclosed embodiments, a CU level flag can be signaled to indicate whether chroma scaling is applied to the current block. The presence of the CU level flag can be conditioned based on the tile group level flag "tile_group_reshaper_chroma_residual_scale_flag". That is, the CU level flag can only be signaled when chroma scaling is permitted at the tile group level. The coder is permitted to choose whether to use chroma scaling based on whether it is beneficial for the current block, but this can also incur significant signaling overhead.
[0141]
[0158] Consistent with the disclosed embodiments, to avoid the above-mentioned signaling overhead, whether chroma scaling is applied to a block can be conditioned based on the prediction mode of the block. For example, when a block is inter-predicted, especially when its reference picture is close in terms of temporal distance, the predicted signal tends to be good. Therefore, since the residual is expected to be very small, chroma scaling can be bypassed. For example, pictures at a higher temporal level tend to have reference pictures that are close in terms of temporal distance. For a block, chroma scaling can be disabled within the picture that uses a nearby reference picture. To determine whether this condition is met, the difference in the picture order count (POC) between the current picture and the reference picture of the block can be used.
[0142]
[0159] In some embodiments, chroma scaling can be disabled for all blocks to be inter-coded. In some embodiments, chroma scaling can be disabled for the combined intra / inter prediction (CIIP) mode defined within the VVC standard.
[0143]
[0160] In the VVC standard, the CU syntax structure “coding_unit()” includes the syntax element “cu_cbf” to indicate whether there are any non-zero residual coefficients within the current CU. At the TU level, the TU syntax structure “transform_unit()” includes the syntax elements “tu_cbf_cb” and “tu_cbf_cr” to indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients within the current TU. Based on these flags, the chroma scaling process can be conditioned. As described above, when there are no non-zero residual coefficients, the averaging of the corresponding luma-chroma scaling process can be invoked. By invoking the averaging, the chroma scaling process can be bypassed.
[0144]
[0161] FIG. 13 shows a flowchart of a method 1300 implemented by a computer for processing video content. In some embodiments, method 1300 may be executed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 1300 at the picture level. For example, the device can process pictures one by one within method 1300. In another example, the device can process a plurality of pictures at once within method 1300. Method 1300 can include steps as follows.
[0145]
[0162] In step 1302, chroma blocks and luma blocks associated with a picture can be received. It will be understood that a picture can be associated with a chroma component and a luma component. Thus, a picture can be associated with chroma blocks including chroma samples and luma blocks including luma samples.
[0146]
[0163] In step 1304, the luma scale information related to the luma block can be determined. In some embodiments, the luma scale information can be a syntax element signaled within the data stream of the picture, or a variable derived based on the syntax element signaled within the data stream of the picture. For example, the luma scale information can include "reshape_model_bin_delta_sign_CW[i] and reshape_model_bin_delta_abs_CW[i]" described in the above equations, and / or "SignaledCW[i]" described in the above equations, etc. In some embodiments, the luma scale information can include variables determined based on the luma block. For example, the average luma value can be determined by calculating the average value of the luma samples adjacent to the luma block (such as the luma samples in the row above the luma block and the luma samples in the column to the left of the luma block, etc.).
[0147]
[0164] In step 1306, the chroma scale factor can be determined based on the luma scale information.
[0148]
[0165] In some embodiments, the luma scale factor of the luma block can be determined based on the luma scale information. For example, according to the above equation "inverse_chroma_scaling[i]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(tempCW>>1)) / tempCW", the luma scale factor can be determined based on the luma scale information (such as "tempCW"). Then, the chroma scale factor can be further determined based on the value of the luma scale factor. For example, the chroma scale factor can be set equal to the value of the luma scale factor. It will be understood that further calculations can be applied to the value of the luma scale factor before setting it as the chroma scale factor. As another example, the chroma scale factor is "SignaledCW[Y IdxIt can be set equal to " / OrgCW", and the classification index "Y" of the current chroma block can be determined based on the average luma value related to the luma block. Idx can be determined.
[0149]
[0166] In step 1308, the chroma block can be processed using a chroma scale factor. For example, to generate a scaled residual of the chroma block, the residual of the chroma block can be processed using the chroma scale factor. The chroma block can be a Cb chroma component or a Cr chroma component.
[0150]
[0167] In some embodiments, the chroma block can be processed when the condition is satisfied. For example, the condition can include that the target coding unit related to the picture has no non-zero residual, or the target transform unit related to the picture has no non-zero chroma residual. Whether the target coding unit has no non-zero residual can be determined based on the value of the first coding block flag of the target coding unit. Whether the target transform unit has no non-zero chroma residual can be determined based on the value of the second coding block flag of the first component and the value of the third coding block flag of the second component of the target transform unit. For example, the first component can be the Cb component, and the second component can be the Cr component.
[0151]
[0168] It will be understood that each step of method 1300 can be executed as an independent method. For example, the method for determining the chroma scale factor described in step 1308 can be executed as an independent method.
[0152]
[0169] FIG. 14 shows a flowchart of a method 1400 implemented by a computer for processing video content. In some embodiments, method 1300 may be executed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 1400 at the picture level. For example, the device can process pictures one by one within method 1400. In another example, the device can process a plurality of pictures at a time within method 1400. Method 1400 can include the following steps.
[0153]
[0170] In step 1402, chroma blocks and luma blocks associated with a picture can be received. It will be appreciated that a picture can be associated with a chroma component and a luma component. Thus, a picture can be associated with chroma blocks including chroma samples and luma blocks including luma samples. In some embodiments, a luma block can include NxM luma samples. N can be the width of the luma block and M can be the height of the luma block. As discussed above, the luma samples of the luma block can be used to determine the segmentation index of the target chroma block. Thus, luma blocks associated with the pictures of the video sequence can be received. It will be appreciated that N and M can have the same value.
[0154]
[0171] In step 1404, a subset of NxM luma samples can be selected in response to at least one of N and M exceeding a threshold value. To accelerate the determination of the segmentation index, if certain conditions are met, the luma block can be "downsampled". In other words, the segmentation index can be determined using a subset of the luma samples within the luma block. In some embodiments, the certain condition is that at least one of N and M exceeds a threshold value. In some embodiments, the threshold value can be based on at least one of N and M. The threshold value can be a power of 2. For example, the threshold value can be 4, 8, 16, etc. Taking 4 as an example, if N or M exceeds 4, a subset of the luma samples can be selected. In the example of FIG. 10, both the width and height of the luma block exceed the threshold value of 4, and thus a subset of 4x4 samples is selected. It will be understood that subsets such as 2x8, 1x16, etc. can also be selected for processing.
[0155]
[0172] In step 1406, the average value of a subset of NxM luma samples can be determined.
[0156]
[0173] In some embodiments, determining the average value can further include determining whether a second condition is met and, in response to determining that the second condition is met, determining the average value of a subset of NxM luma samples. For example, the second condition can include that the target coding unit related to the picture has no non-zero residual coefficients or the target coding unit has no non-zero chroma residual coefficients.
[0157]
[0174] In step 1408, a chroma scale factor based on the average value can be determined. In some embodiments, to determine the chroma scale factor, a segmentation index of a chroma block can be determined based on the average value, it can be determined whether the segmentation index of the chroma block satisfies a first condition, and the chroma scale factor can be set to a default value according to the segmentation index of the chroma block satisfying the first condition. The default value may indicate that chroma scaling is not applied. For example, the default value can be 1.0 with decimal precision. It will be understood that a fixed-point approximation can be applied to the default value. According to the segmentation index of the chroma block not satisfying the first condition, the chroma scale factor can be determined based on the average value. More specifically, the chroma scale factor can be set to SignaledCW[Y Idx / OrgCW, and the segmentation index "Y Idx " of the target chroma block can be determined based on the average value of the corresponding luma block.
[0158]
[0175] In some embodiments, the first condition may include that the segmentation index of the chroma block is greater than the maximum index of the coded word to be signaled or less than the minimum index of the coded word to be signaled. The maximum index and the minimum index of the coded word to be signaled can be determined as follows.
[0159]
[0176] The codeword can be generated using a piecewise linear model (e.g., LMCS) based on an input signal (e.g., luma samples). As discussed above, the dynamic range of the input signal can be divided into several segments (e.g., 16 segments), and each segment of the input signal can be used to generate a bin of the codeword as an output. Thus, each bin of the codeword can have a bin index corresponding to the segment of the input signal. In this example, the range of the bin index can be from 0 to 15. In some embodiments, the value of the output (i.e., the codeword) is between a minimum value (e.g., 0) and a maximum value (e.g., 255), and a plurality of codewords having values between the minimum value and the maximum value can be signaled. The bin indices of the plurality of signaled codewords can be determined. Among the bin indices of the plurality of signaled codewords, the maximum bin index and the minimum bin index of the bins of the plurality of signaled codewords can be further determined.
[0160]
[0177] In addition to the chroma scale factor, method 1400 can further determine the luma scale factor based on the bins of the plurality of signaled codewords. The luma scale factor can be used as an inverse chroma scale factor. The equation for determining the luma scale factor has been described above, and its description is omitted herein. In some embodiments, a plurality of adjacent signaled codewords share the luma scale factor. For example, two or four adjacent signaled codewords can share the same luma scale factor, which can reduce the load of determining the luma scale factor.
[0161]
[0178] In step 1410, the chroma scale factor can be used to process the chroma block. As discussed above with respect to FIG. 5, a plurality of chroma scale factors can be used to construct a chroma scale factor LUT at the tile group level and can be applied on the decoder side to the reconstructed chroma residual of the target block. Similarly, the chroma scale factor can also be applied on the encoder side.
[0162]
[0179] It will be appreciated that each step of method 1400 can be executed as an independent method. For example, the method for determining the chroma scale factor described in step 1308 can be executed as an independent method.
[0163]
[0180] FIG. 15 shows a flowchart of a method 1500 implemented by a computer for processing video content. In some embodiments, method 1500 may be executed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 1500 at the picture level. For example, the device can process pictures one by one within method 1500. In another example, the device can process a plurality of pictures at a time within method 1500. Method 1500 can include the following steps.
[0164]
[0181] In step 1502, it can be determined whether there is a chroma scale index in the received video data.
[0165]
[0182] In step 1504, in response to determining that there is no chroma scale index in the received video data, it can be determined that no chroma scaling is applied to the received video data.
[0166]
[0183] In step 1506, in response to determining that there is chroma scaling in the received video data, a chroma scale factor can be determined based on the chroma scale index.
[0167]
[0184] FIG. 16 shows a flowchart of a method 1600 implemented by a computer for processing video content. In some embodiments, method 1600 may be executed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 1600 at the picture level. For example, the device can process pictures one by one within method 1600. In another example, the device can process a plurality of pictures at a time within method 1600. Method 1600 can include the following steps.
[0168]
[0185] In step 1602, a plurality of unique codewords used for the dynamic range of the input video signal can be received.
[0169]
[0186] In step 1604, an index can be received.
[0170]
[0187] In step 1606, at least one of the plurality of unique codewords can be selected based on the index.
[0171]
[0188] In step 1608, a chroma scale factor can be determined based on at least one selected codeword.
[0172]
[0189] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by an apparatus (such as the disclosed encoder and decoder) to perform the above method. Common non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROM, EPROM, flash EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and networked versions of those. The apparatus can include one or more processors (CPUs), an input / output interface, a network interface, and / or a memory.
[0173]
[0190] It will be understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. When implemented by software, the software can be stored in the above computer-readable media. When executed by a processor, the software can perform the disclosed method. The computing units and other functional units described in this disclosure can be implemented by hardware, software, or a combination of hardware and software. It will be understood by those skilled in the art that a plurality of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into a plurality of sub-modules / sub-units.
[0174]
[0191] Embodiments can be further described using the following clauses: 1. A computer-implemented method for processing video content, comprising: Receiving chroma blocks and luma blocks related to a picture, Determining luma scale information related to the luma blocks, Determining a chroma scale factor based on the luma scale information, and Processing the chroma blocks using the chroma scale factor A method comprising. 2. Determining the chroma scale factor based on the luma scale information is Determining a luma scale factor of the luma blocks based on the luma scale information, Determining the chroma scale factor based on the value of the luma scale factor The method according to item 1, further comprising. 3. Determining the chroma scale factor based on the value of the luma scale factor is Setting the chroma scale factor equal to the value of the luma scale factor The method according to item 2, further comprising. 4. Processing the chroma blocks using the chroma scale factor is Determining whether a first condition is satisfied, and Processing the chroma blocks using the chroma scale factor in response to a determination that the first condition is satisfied, or Bypassing the processing of the chroma blocks using the chroma scale factor in response to a determination that the first condition is not satisfied Performing one of The method according to any one of items 1 to 3, further comprising. 5. The first condition is That the target coding unit related to the picture has no non-zero residual, or That the target transform unit related to the picture has no non-zero chroma residual The method according to item 4, comprising. 6. That the target coding unit has no non-zero residual is determined based on the value of the first coding block flag of the target coding unit, The fact that the target transformation unit has no non-zero chroma residual is determined based on the value of the second coding block flag of the first chroma component of the target transformation unit and the value of the third coding block flag of the second luma-chroma component. The method according to clause 5. 7. The value of the first coding block flag is 0, The value of the second coding block flag and the value of the third coding block flag are 0, The method according to clause 6. 8. Processing a chroma block using a chroma scale factor is Processing the residual of the chroma block using the chroma scale factor The method according to any one of clauses 1 to 7, including this. 9. A device for processing video content, A memory for storing a set of instructions, Connected to the memory, Receiving a chroma block and a luma block related to a picture, Determining luma scale information related to the luma block, Determining a chroma scale factor based on the luma scale information, and Processing the chroma block using the chroma scale factor A processor configured to execute a set of instructions for causing the device to perform The device including this. 10. When determining a chroma scale factor based on luma scale information, Determining a luma scale factor of the luma block based on the luma scale information, Determining a chroma scale factor based on the value of the luma scale factor The device according to clause 9, wherein the processor is configured to execute a set of instructions for further causing the device to perform this. 11. When determining a chroma scale factor based on the value of the luma scale factor, Setting the chroma scale factor equal to the value of the luma scale factor The apparatus according to clause 10, wherein the processor is configured to execute a set of instructions for causing the apparatus to further perform 12. When processing a chroma block using a chroma scale factor, determining whether a first condition is satisfied, and processing the chroma block using the chroma scale factor in response to a determination that a second condition is satisfied, or bypassing the processing of the chroma block using the chroma scale factor in response to a determination that the second condition is not satisfied performing one of The apparatus according to any one of clauses 9 to 11, wherein the processor is configured to execute a set of instructions for causing the apparatus to further perform 13. The first condition is that a target coding unit related to a picture has no non-zero residual, or that a target transform unit related to a picture has no non-zero chroma residual The apparatus according to clause 12 14. That a target coding unit has no non-zero residual is determined based on the value of a first coding block flag of the target coding unit, and that a target transform unit has no non-zero chroma residual is determined based on the value of a second coding block flag of a first chroma component of the target transform unit and the value of a third coding block flag of a second chroma component The apparatus according to clause 13 15. The value of the first coding block flag is 0, and the values of the second coding block flag and the third coding block flag are 0 The apparatus according to clause 14 16. When processing a chroma block using a chroma scale factor, processing the residual of the chroma block using the chroma scale factor The apparatus according to any one of clauses 9 to 15, wherein the processor is configured to execute a set of instructions for causing the apparatus to further perform A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of a device for causing the device to perform a method for processing video content, the method comprising: Receiving chroma blocks and luma blocks associated with a picture; Determining luma scale information associated with the luma blocks; Determining a chroma scale factor based on the luma scale information; and Processing the chroma blocks using the chroma scale factor A non-transitory computer-readable storage medium comprising the steps of: A computer-implemented method for processing video content, the method comprising: Receiving chroma blocks and luma blocks associated with a picture, the luma blocks including NxM luma samples; Selecting a subset of the NxM luma samples in response to at least one of N and M exceeding a threshold; Determining an average value of the subset of the NxM luma samples; Determining a chroma scale factor based on the average value; and Processing the chroma blocks using the chroma scale factor A method comprising the steps of: A computer-implemented method for processing video content, the method comprising: Determining whether there is a chroma scale index in the received video data; Determining that chroma scaling is not applied to the received video data in response to determining that there is no chroma scale index in the received video data; and Determining a chroma scale factor based on the chroma scale index in response to determining that there is chroma scaling in the received video data A method comprising the steps of: A computer-implemented method for processing video content, the method comprising: Receiving a plurality of unique codewords used for the dynamic range of an input video signal, Receiving an index, Selecting at least one of the plurality of unique codewords based on the index, and Determining a chroma scale factor based on the selected at least one codeword A method comprising.
[0175]
[0192] In addition to implementing the above method by using computer-readable program code, the above method can also be implemented in the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be regarded as a hardware component, and a device configured to implement various functions included in the controller can also be regarded as a structure within the hardware component. Or a device configured to implement various functions can even be regarded as both a software module configured to implement the method and a structure within the hardware component.
[0176]
[0193] The present disclosure can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, assemblies, data structures, classes, etc. used to execute a specific task or implement a specific abstract data type. Embodiments of the present disclosure can also be implemented in a distributed computing environment. In a distributed computing environment, tasks are executed by using remote processing devices connected by a communication network. In a distributed computing environment, program modules can be in local and remote computer storage media including storage devices.
[0177]
[0194] Relational terms such as "first" and "second" in this specification are used solely to distinguish one entity or operation from another entity or operation, and it should be noted that no actual relationship or order between those entities or operations is required or implied. Further, words such as "comprising", "having", "containing", and "including" and other similar forms of words are intended to be equivalent in meaning, and it is not intended that an item following any one of these words be an exhaustive listing of such item, nor is it intended to be limited to only the items listed, being intended to be non-limiting in that respect.
[0178]
[0195] In the above specification, embodiments have been described with respect to numerous specific details that may vary for each implementation form. Certain adaptations and modifications of the described embodiments can be made. By considering this specification and practicing the disclosure made clear herein, other embodiments may become apparent to those skilled in the art. This specification and the examples are to be considered solely as examples, and it is intended that the true scope and spirit of the disclosure be indicated by the appended claims. The order of the steps shown in the figures is for illustrative purposes only and is not intended to be limited to a particular order of steps. Therefore, those skilled in the art can understand that those steps can be executed in different orders while implementing the same method.
Claims
1. 1. A computer-implemented method for processing video content, comprising: receiving chroma and luma blocks associated with a picture; determining luma scale information associated with the luma block; determining a chroma scale factor based on the luma scale information; and processing said chroma block using said chroma scale factor; A method comprising:
2. determining the chroma scale factor based on the luma scale information; determining a luma scale factor for the luma block based on the luma scale information; determining the chroma scale factor based on the value of the luma scale factor; The method of claim 1 further comprising:
3. determining the chroma scale factor based on the value of the luma scale factor; setting the chroma scale factor equal to the value of the luma scale factor. The method of claim 2 further comprising:
4. processing the chroma block using the chroma scale factor, Determining whether a first condition is satisfied; and processing the chroma block using the chroma scale factor in response to the determination that the first condition is satisfied; or bypassing the processing of the chroma block with the chroma scale factor in response to the determination that the first condition is not satisfied. Do one of the following: The method of claim 1 further comprising:
5. The first condition is the target coding unit associated with the picture does not have a non-zero residual; or the target transform unit associated with the picture does not have non-zero chroma residual; The method of claim 4 , comprising:
6. The target coding unit is determined to have no non-zero residual based on a value of a first coded block flag of the target coding unit; The target transform unit is determined to have no non-zero chroma residual based on a value of a second coded block flag of a first chroma component and a value of a third coded block flag of a second luma chroma component of the target transform unit. The method according to claim 5.
7. the value of the first coded block flag is 0; the value of the second coded block flag and the value of the third coded block flag are 0; The method according to claim 6.
8. processing the chroma block using the chroma scale factor, processing the residual of the chroma block using the chroma scale factor; The method of claim 1 , comprising:
9. A device for processing video content, comprising: a memory for storing a set of instructions; coupled to the memory; receiving chroma and luma blocks associated with a picture; determining luma scale information associated with the luma block; determining a chroma scale factor based on the luma scale information; and processing said chroma block using said chroma scale factor; a processor configured to execute the set of instructions to cause the device to Including, equipment.
10. determining the chroma scale factor based on the luma scale information, determining a luma scale factor for the luma block based on the luma scale information; determining the chroma scale factor based on the value of the luma scale factor; 10. The apparatus of claim 9, wherein the processor is configured to execute the set of instructions to further cause the apparatus to:
11. determining the chroma scale factor based on the value of the luma scale factor; setting the chroma scale factor equal to the value of the luma scale factor. The apparatus of claim 10 , wherein the processor is configured to execute the set of instructions to further cause the apparatus to:
12. When processing the chroma block using the chroma scale factor, Determining whether a first condition is satisfied; and processing the chroma block using the chroma scale factor in response to the determination that a second condition is satisfied; or bypassing the processing of the chroma block with the chroma scale factor in response to the determination that the second condition is not satisfied. Do one of the following:
10. The apparatus of claim 9, wherein the processor is configured to execute the set of instructions to further cause the apparatus to:
13. The first condition is the target coding unit associated with the picture does not have a non-zero residual; or the target transform unit associated with the picture does not have non-zero chroma residual; 13. The device of claim 12, comprising:
14. The target coding unit is determined to have no non-zero residual based on a value of a first coded block flag of the target coding unit; determining that the target transform unit does not have a non-zero chroma residual based on a value of a second coded block flag of a first chroma component and a value of a third coded block flag of a second chroma component of the target transform unit; 14. The device of claim 13.
15. the value of the first coded block flag is 0; the value of the second coded block flag and the value of the third coded block flag are 0; 15. The device of claim 14.
16. When processing the chroma block using the chroma scale factor, processing the residual of the chroma block using the chroma scale factor; 10. The apparatus of claim 9, wherein the processor is configured to execute the set of instructions to further cause the apparatus to:
17. 1. A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of a device to cause the device to perform a method for processing video content, the method comprising: receiving chroma blocks and luma blocks associated with a picture; determining luma scale information associated with the luma block; determining a chroma scale factor based on the luma scale information; and processing said chroma block using said chroma scale factor; A non-transitory computer readable storage medium comprising:
Citation Information
Patent Citations
Interactions between in-loop reshaping and inter coding tools
WO2020156526A1
Signaling of in-loop reshaping information using parameter sets
WO2020156529A1