Method and system for processing video content - Patents.com
By processing video content using chroma and luma blocks with a calculated chroma scale factor, the method addresses the high bandwidth and storage issues associated with HD video surveillance, achieving reduced bitrate and efficient data handling.
Patent Information
- Application Number
- JP2021547465
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-12
- Filing Date
- 2020-02-28
- Publication Date
- 2025-05-19
- Estimated Expiration
- 2040-02-28
AI Technical Summary
High bandwidth and storage requirements for high-definition (HD) video in video surveillance applications, due to the high bitrate of HD video bitstreams.
A method for processing video content that involves receiving chroma blocks and luma blocks, determining luma scale information, calculating a chroma scale factor based on the luma scale information, and processing the chroma blocks using this chroma scale factor.
This approach reduces the bitrate of video data without significant information loss, thereby decreasing bandwidth and storage requirements for HD video surveillance.
Smart Images

Figure 0007679301000005 
Figure 0007679301000006 
Figure 0007679301000007
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications
[0001] This disclosure claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 813,728, filed Mar. 4, 2019, and U.S. Provisional Patent Application No. 62 / 817,546, filed Mar. 12, 2019, both of which are hereby incorporated by reference in their entireties.
[0002] Technical Field
[0002] This disclosure generally relates to video processing, and more particularly, to methods and systems for performing in - loop remapping by chroma scaling.
Background Art
[0003] Background
[0003] Video coding systems are often used to compress digital video signals, for example, to reduce the memory space consumed or to reduce the amount of transmission bandwidth consumed associated with such signals. As the popularity of high - definition (HD) video (e.g., having a resolution of 1920×1080 pixels) grows in various applications of video compression such as online video streaming, video conferencing, or video surveillance, there is an ongoing need to develop video coding tools that can improve the compression efficiency of video data.
[0004]
[0004] For example, video surveillance applications are being used more and more extensively in many application scenarios (such as security, traffic, environmental monitoring, etc.), and the number and resolution of surveillance devices are increasing rapidly. Many video surveillance application scenarios choose to provide users with HD video in order to capture more information, and HD video has more pixels per frame to capture such information. However, an HD video bitstream can have a high bitrate that requires high bandwidth for transmission and large space for storage. For example, a surveillance video stream with an average resolution of 1920x1080 may require a bandwidth of 4 Mbps for real-time transmission. Furthermore, video surveillance generally performs round-the-clock monitoring for 7 days x 24 hours, which can significantly test the capabilities of the storage system when storing video data. Therefore, the demand for high bandwidth and large storage space for HD video has become the main limitation for the large-scale deployment of HD video in video surveillance.
Summary of the Invention
Means for Solving the Problems
[0005] Summary of the Disclosure
[0005] Embodiments of the present disclosure provide a method for processing video content. The method may include receiving chroma blocks and luma blocks related to a picture, determining luma scale information related to the luma blocks, determining a chroma scale factor based on the luma scale information, and processing the chroma blocks using the chroma scale factor.
[0006]
[0006] Embodiments of the present disclosure provide a device for processing video content. The device may include a memory storing a set of instructions, and a processor coupled to the memory and configured to execute the set of instructions to cause the device to receive chroma blocks and luma blocks related to a picture, determine luma scale information related to the luma blocks, determine a chroma scale factor based on the luma scale information, and process the chroma blocks using the chroma scale factor.
[0007]
[0007] Embodiments of the present disclosure provide a non - transitory computer - readable storage medium storing a set of instructions executable by one or more processors of a device to cause the device to perform a method for processing video content. The method includes receiving chroma blocks and luma blocks related to a picture, determining luma scale information related to the luma blocks, determining a chroma scale factor based on the luma scale information, and processing the chroma blocks using the chroma scale factor.
[0008] Brief Description of the Drawings
[0008] Embodiments and various aspects of the present disclosure are shown in the following detailed description and the accompanying drawings. The various features shown in the figures are not drawn to scale.
Brief Description of the Drawings
[0009]
Figure 1
[0009] An example structure of a video sequence according to some embodiments of the present disclosure is shown.
Figure 2A
[0010] A schematic diagram of an example of an encoding process according to some embodiments of the present disclosure is shown.
Figure 2B
[0011] A schematic diagram of another example of an encoding process according to some embodiments of the present disclosure is shown.
Figure 3A
[0012] A schematic diagram of an example of a decoding process according to some embodiments of the present disclosure is shown.
Figure 3B
[0013] A schematic diagram of another example of a decoding process according to some embodiments of the present disclosure is shown.
Figure 4
[0014] A block diagram of an example of a device for encoding or decoding video according to some embodiments of the present disclosure is shown.
Figure 5
[0015] Schematic diagram of an exemplary luma mapping with chroma scaling (LMCS) process according to some embodiments of the present disclosure is shown.
Figure 6
[0016] Table of syntax at tile group level for the LMCS partition linear model according to some embodiments of the present disclosure is shown.
Figure 7
[0017] Table of syntax at another tile group level for the LMCS partition linear model according to some embodiments of the present disclosure is shown.
Figure 8
[0018] Table of syntax structure of coded tree units according to some embodiments of the present disclosure.
Figure 9
[0019] Table of syntax structure of dual tree splitting according to some embodiments of the present disclosure.
Figure 10
[0020] An example of simplifying the averaging of luma prediction blocks according to some embodiments of the present disclosure is shown.
Figure 11
[0021] Table of syntax structure of coded tree units according to some embodiments of the present disclosure.
Figure 12
[0022] Table of syntax elements related to the modified signaling of the LMCS partition linear model at the tile group level according to some embodiments of the present disclosure.
Figure 13
[0023] Flow diagram of a method for processing video content according to some embodiments of the present disclosure.
Figure 14
[0024] Flow diagram of a method for processing video content according to some embodiments of the present disclosure.
Figure 15
[0025] Flow diagram of another method for processing video content according to some embodiments of the present disclosure.
Figure 16
[0026] Flow diagram of another method for processing video content according to some embodiments of the present disclosure.
Best Mode for Carrying Out the Invention
[0010] Detailed Description
[0027] Next, reference will be made in detail to exemplary embodiments shown in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numerals in different figures represent the same or similar elements unless otherwise indicated. The implementation forms described in the following description of the exemplary embodiments do not represent all implementation forms that conform to the present invention. Rather, they are merely examples of devices and methods that conform to aspects related to the present invention listed in the appended claims. Unless otherwise specified, the term "or" includes all possible combinations except when not executable. For example, when it is stated that a certain component may include A or B, unless otherwise specified or not executable, the component may include A, or B, or A and B. As a second example, when it is stated that a certain component may include A, B, or C, unless otherwise specified or not executable, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0011]
[0028] Video is a set of still pictures (or "frames") arranged in chronological order for storing visual information. A video capture device (such as a camera) can be used to capture and store those pictures in chronological order, and a video playback device (such as a TV, computer, smartphone, tablet computer, video player, or any end-user terminal having a display function) can be used to display such pictures in chronological order. Further, in some applications, the video capture device can transmit the captured video in real time to a video playback device (such as a computer having a monitor) for monitoring, conferencing, or live broadcasting, etc.
[0012]
[0029] In order to reduce the memory space and transmission bandwidth required by such applications, the video can be compressed before being stored and transmitted, and decompressed before being displayed. This compression and decompression can be implemented by software executed by a processor (such as the processor of a general-purpose computer) or dedicated hardware. The module for compression is generally called an "encoder", and the module for decompression is generally called a "decoder". The encoder and decoder can be collectively called a "codec". The encoder and decoder can be implemented as various suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder may include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder may include program code, computer-executable instructions, firmware, or algorithms or processes implemented by any suitable computer fixed in a computer-readable medium. The compression and decompression of video can be implemented by various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be called a "transcoder".
[0013]
[0030] The video encoding process can identify and retain the useful information that can be used to reconstruct the picture, and ignore the information that is not important for reconstruction. If the unimportant information that is ignored cannot be completely reconstructed, such an encoding process can be called "irreversible". Otherwise, such an encoding process can be called "reversible". Most encoding processes are irreversible, which is a trade-off for reducing the required memory space and transmission bandwidth.
[0014]
[0031] The useful information of the symbolized picture (referred to as the "current picture") includes changes with respect to the reference picture (e.g., a picture that was symbolized and reconstructed in the past). Such changes can include changes in pixel position, luminance, or color, among which the position change is the most relevant. The position change of the group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.
[0015]
[0032] A picture that is coded without referring to another picture (i.e., such a picture is its own reference picture) is called an "I picture". A picture that is coded using a past picture as a reference picture is called a "P picture". A picture that is coded using both a past picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B picture".
[0016]
[0033] As described above, video surveillance using HD video faces the problem of high bandwidth and large storage requirements. To address this problem, the bit rate of the coded video can be reduced. Among the I picture, P picture, and B picture, the I picture has the highest bit rate. Since the background of most surveillance videos is almost static, one way to reduce the overall bit rate of the coded video may be to use fewer I pictures for video coding.
[0017]
[0034] However, since the I picture is generally not the main one in the coded video, the improvement measure of using fewer I pictures may be a minor matter. For example, in a typical video bitstream, the ratio of I pictures, B pictures, and P pictures may be 1:20:9, and the I pictures may account for less than 10% of the total bit rate. In other words, in such an example, even if all the I pictures are removed, the reduced bit rate may be only 10%.
[0018]
[0035] The present disclosure provides a method, apparatus, and system for characteristic-based video processing for video surveillance. As used herein, "characteristics" refers to content characteristics related to video content within a picture, motion characteristics related to motion estimation for encoding or decoding a picture, or both. For example, the content characteristics can be pixels within one or more consecutive pictures of a video sequence, and the pixels are related to at least one of an object, a scene, or an environmental event within the picture. In another example, the motion characteristics can include information related to the encoding process of the video, examples of which will be described in detail later.
[0019]
[0036] In the present disclosure, when encoding a picture of a video sequence, a classifier can be used to detect and classify one or more characteristics of the picture of the video sequence. Various classes of characteristics can be associated with various priority levels, and such priority levels are further associated with various bitrates for encoding. The various priority levels can be associated with various parameter sets for encoding, whereby different encoding quality levels can be achieved. The higher the priority level, the higher the quality of the video that the associated parameter set can provide. Such characteristic-based video processing can significantly reduce the bitrate for surveillance video without causing significant information loss. In addition, embodiments of the present disclosure can customize the corresponding relationship between the priority level and the parameter set for various application scenarios (such as security, traffic, environmental monitoring, etc.), thereby significantly improving the encoding quality of the video and significantly reducing the bandwidth and storage cost.
[0020]
[0037] FIG. 1 shows the structure of an example of video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live relay video or a captured and archived video. The video 100 can be a real video, a video generated by a computer (e.g., a computer game video), or a combination thereof (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing videos captured in the past (e.g., a video file stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) for receiving videos from a video content provider.
[0021]
[0038] As shown in FIG. 1, the video sequence 100 can include a series of pictures temporally arranged along a timeline including pictures 102, 104, 106, and 108. Pictures 102 to 106 are consecutive, and there are more pictures between picture 106 and picture 108. In FIG. 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102 as indicated by the arrow. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108 as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be the picture immediately before or after that picture. For example, the reference picture of picture 104 can be a picture preceding picture 102. It should be noted that the reference pictures of pictures 102 to 106 are merely examples, and the present disclosure does not limit the embodiments of the reference pictures to the examples shown in FIG. 1.
[0022]
[0039] Typically, a video codec does not encode or decode an entire picture at once because such a task is computationally complex. Instead, a video codec divides a picture into basic segments and can encode or decode the picture segment by segment. In the present disclosure, such a basic segment is referred to as a basic processing unit ("BPU"). For example, structure 110 of FIG. 1 shows an example of the structure of a picture (e.g., any one of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are indicated by dashed lines. In some embodiments, the basic processing unit can be referred to as a "macroblock" within some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), and can be referred to as a "coding tree unit" ("CTU") within some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have a variable size within the picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size of pixels. The size and shape of the basic processing unit can be selected for a picture based on the balance between coding efficiency and the level of detail to be maintained within the basic processing unit.
[0023]
[0040] The basic processing unit can be a logical unit that can include various types of video data groups stored in a computer memory (e.g., within a video frame buffer). For example, the basic processing unit of a color picture can include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements of the basic processing unit where the luma component and the chroma components can have the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luma component and the chroma components can be referred to as "coding tree blocks" ("CTBs"). Any operation performed on the basic processing unit can be repeated for each of its luma component and chroma components.
[0024]
[0041] The encoding of an image has multiple operation stages, and examples thereof will be detailed in FIGS. 2A to 2B and FIGS. 3A to 3B. For each stage, the size of the basic processing unit may still be too large to process, and thus it can be further divided into segments called "basic processing sub-units" in the present disclosure. In some embodiments, the basic processing sub-unit can be called a "block" within some video encoding standards (such as the MPEG family, H.261, H.263, or H.264 / AVC), or can be called a "coding unit" ("CU") within some other video encoding standards (such as H.265 / HEVC or H.266 / VVC). The basic processing sub-unit can have a size equal to or smaller than that of the basic processing unit. Similar to the basic processing unit, the basic processing sub-unit is a logical unit that can include various types of video data groups (such as Y, Cb, Cr, and related syntax elements) stored in a computer memory (such as within a video frame buffer). Any operation performed on the basic processing sub-unit can be repeatedly performed on each of its luma component and chroma component. It should be noted that such division can be performed at a further level according to the need for processing. It should also be noted that various stages can divide the basic processing unit using various methods.
[0025]
[0042] For example (one example is detailed in FIG. 2B), in the mode decision stage, the encoder can determine which prediction mode (such as intra-picture prediction or inter-picture prediction) to use for the basic processing unit, and the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into a plurality of basic processing sub-units (such as the CU in H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing sub-unit.
[0026]
[0043] In another example, (one example of which is described in detail in FIG. 2A) in the prediction stage, the coder can perform a prediction operation at the level of a basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and perform a prediction operation at that level.
[0027]
[0044] In another example, (one example of which is described in detail in FIG. 2A) in the conversion stage, the coder can perform a conversion operation on a residual basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "conversion blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and perform a conversion operation at that level. It should be noted that the splitting method of the same basic processing subunit may be different in the prediction stage and the conversion stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and conversion blocks of the same CU may have different sizes and numbers.
[0028]
[0045] In the structure 110 of FIG. 1, the basic processing unit 112 is further divided into 3x3 basic processing subunits, and its boundary is shown by a dotted line. Different basic processing units of the same picture can be divided into basic processing subunits in different ways.
[0029]
[0046] In some implementation forms, in order to provide parallel processing and error tolerance functions for video encoding and decoding, a picture can be divided into regions for processing, so that for the regions of the picture, the encoding or decoding process can be made independent of the information in any other region of the picture. In other words, each region of the picture can be processed independently. By doing so, the codec can process different regions of the picture in parallel, thus improving the encoding efficiency. Further, when the data of a region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the damaged or lost data, thus providing an error tolerance function. In some video coding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions, namely "slices" and "tiles". It should also be noted that various pictures of video sequence 100 can have various splitting methods for dividing the picture into regions.
[0030]
[0047] For example, in FIG. 1, structure 110 is divided into three regions 114, 116, and 118, and its boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in FIG. 1 are only examples, and the present disclosure does not limit its embodiments.
[0031]
[0048] Figure 2A shows a schematic diagram of an example of an encoding process 200A according to some embodiments of the present disclosure. The encoder can encode the video sequence 202 into a video bitstream 228 according to process 200A. Similar to the video sequence 100 of FIG. 1, the video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in chronological order. Similar to the structure 110 of FIG. 1, each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder can execute process 200A at the level of the basic processing unit for each original picture of the video sequence 202. For example, the encoder can execute process 200A in an iterative manner, and the encoder can encode a basic processing unit within one iteration of process 200A. In some embodiments, the encoder can execute process 200A in parallel for the regions (e.g., regions 114-118) of each original picture of the video sequence 202.
[0032]
[0049] In FIG. 2A, the coder can feed the basic processing unit of the original picture of video sequence 202 (referred to as the "original BPU") to prediction stage 204 to generate prediction data 206 and predicted BPU 208. The coder can subtract the predicted BPU 208 from the original BPU to generate residual BPU 210. The coder can feed the residual BPU 210 to transformation stage 212 and quantization stage 214 to generate quantized transform coefficients 216. The coder can feed the prediction data 206 and the quantized transform coefficients 216 to binary coding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the "forward path". During process 200A, the coder can feed the quantized transform coefficients 216 to inverse quantization stage 218 and inverse transformation stage 220 after quantization stage 214 to generate reconstructed residual BPU 222. The coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate prediction reference 224 for use in the prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referred to as the "reconstruction path". The reconstruction path can be used to ensure that both the coder and the decoder use the same reference data for prediction.
[0033]
[0050] The coder can repeatedly execute process 200A to encode each original BPU of the original picture (within the forward path) and generate prediction reference 224 for encoding the next original BPU of the original picture (within the reconstruction path). After encoding all the original BPUs of the original picture, the coder can proceed to encode the next picture in video sequence 202.
[0034]
[0051] Referring to process 200A, the coder can receive video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive (s)" can refer to receiving for inputting data, inputting, obtaining, retrieving, getting, reading, accessing, or any action in any way.
[0035]
[0052] At the prediction stage 204 of the current iteration, the coder can receive the original BPU and prediction criterion 224 and perform a prediction operation to generate prediction data 206 and predicted BPU 208. Prediction criterion 224 can be generated from the reconstruction path of the previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206 that can be used to reconstruct the original BPU as predicted BPU 208 from prediction data 206 and prediction criterion 224.
[0036]
[0053] Ideally, predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such a difference, the coder can subtract predicted BPU 208 from the original BPU after generating it to generate residual BPU 210. For example, the coder can subtract the pixel value (e.g., grayscale value or RGB value) of predicted BPU 208 from the corresponding pixel value of the original BPU. As a result of such subtraction between the corresponding pixel of the original BPU and predicted BPU 208, each pixel of residual BPU 210 can have a residual value. Compared with the original BPU, prediction data 206 and residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly degrading the quality.
[0037]
[0054] To further compress the residual BPU 210, at the transformation stage 212, the coder can reduce the spatial redundancy of the residual BPU 210 by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", where each basis pattern is associated with a "transformation coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a variable frequency (e.g., luminance variable frequency) component of the residual BPU 210. None of the basis patterns can be reproduced from any combination (e.g., linear combination) of any other basis patterns. In other words, such decomposition can decompose the variations of the residual BPU 210 into the frequency domain. Such decomposition is similar to the discrete Fourier transform of a function, the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basis functions.
[0038]
[0055] Various transformation algorithms can use various basis patterns. For example, at the transformation stage 212, various transformation algorithms such as the discrete cosine transform, discrete sine transform, etc. can be used. The transformation at the transformation stage 212 is reversible. That is, the coder can restore the residual BPU 210 by the inverse operation of the transformation (referred to as "inverse transformation"). For example, to restore the pixels of the residual BPU 210, the inverse transformation can be to multiply the values of the corresponding pixels of the basis patterns by their respective associated coefficients and add the products to obtain a weighted sum. In the video coding standard, both the coder and the decoder can use the same transformation algorithm (and thus the same basis patterns). Therefore, the coder can record only the transformation coefficients that can reconstruct the residual BPU 210 from there without the decoder receiving the basis patterns from the coder. The transformation coefficients may have fewer bits compared to the residual BPU 210, but those transformation coefficients can be used to reconstruct the residual BPU 210 without significantly degrading the quality. Therefore, the residual BPU 210 is further compressed.
[0039]
[0056] The coder can further compress the transform coefficients in the quantization stage 214. In the transform process, various basis patterns can represent various fluctuation frequencies (e.g., luminance fluctuation frequencies). Since the human eye is generally good at recognizing low-frequency fluctuations, the coder can ignore the information of high-frequency fluctuations without causing significant quality degradation during decoding. For example, in the quantization stage 214, the coder can divide each transform coefficient by an integer value (referred to as the "quantization parameter") and round the quotient to its nearest neighbor to generate the quantized transform coefficient 216. After such an operation, some of the transform coefficients of the high-frequency basis pattern can be converted to zero, and the transform coefficients of the low-frequency basis pattern can be converted to smaller integers. The coder can ignore the quantized transform coefficients 216 with zero values, thereby further compressing the transform coefficients. The quantization process is also reversible, and the quantized transform coefficients 216 can be reconstructed into transform coefficients by the inverse operation of quantization (referred to as "inverse quantization").
[0040]
[0057] Since the coder ignores the remainder of such division in the rounding operation, the quantization stage 214 can be irreversible. Typically, the quantization stage 214 can contribute to the largest information loss within the process 200A. The greater the information loss, the fewer bits the quantized transform coefficient 216 may require. To obtain various levels of information loss, the coder can use various values of the quantization parameter or any other parameter of the quantization process.
[0041]
[0058] In the binary encoding stage 226, the coder can encode the prediction data 206 and the quantized transform coefficients 216 using binary encoding techniques such as, for example, entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other reversible or irreversible compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the coder can encode other information such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transformation in the transformation stage 212, the parameters of the quantization process (e.g., quantization parameters), coder control parameters (e.g., bit rate control parameters), etc. in the binary encoding stage 226. The coder can generate a video bitstream 228 using the output data of the binary encoding stage 226. In some embodiments, the video bitstream 228 can be further packetized for network transmission.
[0042]
[0059] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the coder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transformation stage 220, the coder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction criterion 224 used in the next iteration of process 200A.
[0043]
[0060] It should be noted that other variations of process 200A can be used to encode the video sequence 202. In some embodiments, the encoder can execute the steps of process 200A in a different order. In some embodiments, one or more steps of process 200A can be combined into a single step. In some embodiments, a single step of process 200A can be divided into multiple steps. For example, the transform step 212 and the quantization step 214 can be combined into a single step. In some embodiments, process 200A can include additional steps. In some embodiments, process 200A can omit one or more steps in FIG. 2A.
[0044]
[0061] FIG. 2B shows a schematic diagram of another example 200B of an encoding process according to some embodiments of the present disclosure. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 200A, the forward path of process 200B further includes a mode decision step 230 and divides the prediction step 204 into a spatial prediction step 2042 and a temporal prediction step 2044. The reconstruction path of process 200B additionally includes a loop filter step 232 and a buffer 234.
[0045]
[0062] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use the pixels of one or more adjacent already-coded BPUs within the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use the areas of one or more already-coded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include the coded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0046]
[0063] Referring to process 200B, in the forward path, the coder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the coder can perform intra prediction. With respect to the original BPU of the encoded picture, the prediction reference 224 can include one or more adjacent BPUs that are encoded (in the forward path) and reconstructed (in the reconstruction path) within the same picture. The coder can generate the predicted BPU 208 by extrapolating the adjacent BPUs. The extrapolation technique can include, for example, linear extrapolation or linear interpolation, polynomial extrapolation or polynomial interpolation, etc. In some embodiments, the coder can perform extrapolation at the pixel level, such as by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The adjacent BPUs used for extrapolation can be located relative to the original BPU in various directions, such as in the vertical direction (e.g., above the original BPU), horizontal direction (e.g., to the left of the original BPU), diagonal direction (e.g., bottom left, bottom right, top left, or top right of the original BPU), or any direction defined within the video coding standard being used. In intra prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the parameters of the extrapolation, the direction of the adjacent BPUs used relative to the original BPU, etc.
[0047]
[0064] In another example, the coder can perform inter prediction in the temporal prediction stage 2044. For the original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that are encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures can be encoded and reconstructed for each BPU. For example, the coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same picture are generated, the coder can generate a reconstructed picture as a reference picture. The coder can perform an operation of "motion estimation" to search for a matching region within the range of the reference pictures (referred to as the "search window"). The position of the search window within the reference picture can be determined based on the position of the original BPU within the current picture. For example, the search window can be centered at a position having the same coordinates as the original BPU within the current picture within the reference picture and can be expanded over a predetermined distance. When the coder identifies a region similar to the original BPU within the search window (e.g., by using a pel recursive algorithm, a block matching algorithm, etc.), the coder can determine that region as the matching region. The matching region can have dimensions different from (e.g., smaller than, equal to, larger than, or of a different shape than) the original BPU. Since the reference picture and the current picture are temporally separated within the timeline (as shown in FIG. 1, for example), it can be considered that the matching region "moves" to the position of the original BPU as time passes. The coder can record the direction and distance of such motion as a "motion vector". When multiple reference pictures (such as picture 106 in FIG. 1) are used, the coder can search for a matching region for each reference picture and obtain its associated motion vector. In some embodiments, the coder can assign weights to the pixel values of the matching regions of the individual matching reference pictures.
[0048]
[0065] Motion estimation can be used to identify various types of motion such as translation, rotation, scaling, etc. In inter prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the matching region, the motion vector related to the matching region, the number of reference pictures, the weights related to the reference pictures, etc.
[0049]
[0066] To generate the predicted BPU 208, the coder can perform an operation of "motion compensation". Motion compensation can be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vector) and the prediction reference 224. For example, the coder can move the matching region of the reference picture according to the motion vector, in which the coder can predict the original BPU of the current picture. When multiple reference pictures (such as picture 106 in FIG. 1) are used, the coder can move the matching regions of the reference pictures according to the individual motion vectors and average the pixel values of the matching regions. In some embodiments, when the coder assigns weights to the pixel values of the matching regions of the individual matching reference pictures, the coder can obtain the weighted sum of the pixel values of the moved matching regions.
[0050]
[0067] In some embodiments, the inter prediction can be either unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction as the current picture. For example, picture 104 in FIG. 1 is an unidirectional inter prediction picture where the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions as the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter prediction picture where the reference pictures (i.e., pictures 104 and 108) are in both temporal directions with respect to picture 104.
[0051]
[0068] Continuing to refer to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the coder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the coder can perform rate-distortion optimization techniques, in which the coder can select a prediction mode to minimize the value of a cost function according to the bit rate of the candidate prediction modes and the distortion of the reconstructed reference pictures under the candidate prediction modes. Depending on the selected prediction mode, the coder can generate the corresponding predicted BPU 208 and predicted data 206.
[0052]
[0069] In the reconstruction path of process 200B, when an intra prediction mode is selected within the forward path, after generating a prediction reference 224 (e.g., the current BPU that is encoded and reconstructed within the current picture), the coder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). When an inter prediction mode is selected within the forward path, after generating a prediction reference 224 (e.g., the current picture in which all BPUs are encoded and reconstructed), the coder can feed the prediction reference 224 to the loop filter stage 232, where the coder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) caused by inter prediction. The coder can apply various loop filter techniques at the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, etc. The loop-filtered reference picture can be stored in a buffer 234 (or "reconstructed picture buffer") for later use (e.g., for use as an inter prediction reference picture for future pictures of video sequence 202). The coder can store one or more reference pictures in buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the coder can encode loop filter parameters (e.g., the strength of the loop filter) at the binary coding stage 226 along with the quantized transform coefficients 216, prediction data 206, and other information.
[0053]
[0070] FIG. 3A shows a schematic diagram of an example of a decoding process 300A according to some embodiments of the present disclosure. The process 300A may be a decompression process corresponding to the compression process 200A of FIG. 2A. In some embodiments, the process 300A may be similar to the reconstruction path of the process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss in the compression and decompression processes (e.g., the quantization stage 214 of FIGS. 2A-2B), generally the video stream 304 is not identical to the video sequence 202. Similar to the processes 200A and 200B of FIGS. 2A-2B, the decoder can execute the process 300A at the level of a basic processing unit (BPU) for each picture encoded in the video bitstream 228. For example, the decoder can execute the process 300A in an iterative manner, and the decoder can decode the basic processing unit within one iteration of the process 300A. In some embodiments, the decoder can execute the process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in the video bitstream 228.
[0054]
[0071] In FIG. 3A, the decoder can feed a portion of the video bitstream 228 associated with a basic processing unit of the encoded picture (referred to as an “encoded BPU”) to the binary decoding stage 302. At the binary decoding stage 302, the decoder can decode that portion into prediction data 206 and quantized transform coefficients 216. The decoder can feed the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder can feed the prediction data 206 to the prediction stage 204 to generate a predicted BPU 208. The decoder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a predicted reference 224. In some embodiments, the predicted reference 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can feed the predicted reference 224 for performing a prediction operation within the next iteration of process 300A to the prediction stage 204.
[0055]
[0072] The decoder can repeatedly execute process 300A to decode each encoded BPU of the encoded picture and generate a predicted reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all the encoded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.
[0056]
[0073] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary coding technique (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other reversible compression algorithm) used by the encoder. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the decoder can decode other information such as, for example, the prediction mode, the parameters of the prediction operation, the type of transformation, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), etc. in the binary decoding stage 302. In some embodiments, when the video bitstream 228 is transmitted in packet units over a network, the decoder can depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.
[0057]
[0074] FIG. 3B shows a schematic diagram of another example 300B of the decoding process according to some embodiments of the present disclosure. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0058]
[0075] In process 300B, regarding the encoded basic processing unit of the decoded encoded picture (referred to as the "current picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data depending on which prediction mode was used by the encoder to encode the current BPU. For example, when intra prediction is used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, etc. The parameters of the intra prediction operation may include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as references, the sizes of the adjacent BPUs, extrapolation parameters, the directions of the adjacent BPUs with respect to the original BPU, etc. In another example, when inter prediction is used by the encoder to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, etc. The parameters of the inter prediction operation may include, for example, the number of reference pictures related to the current BPU, the weights respectively related to the reference pictures, the positions (e.g., coordinates) of one or more matching regions in each reference picture, one or more motion vectors respectively related to the matching regions, etc.
[0059]
[0076] Based on the prediction mode indicator, the decoder can determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of the execution of such spatial or temporal prediction are shown in Figure 2B and will not be repeated here. After performing such spatial or temporal prediction, the decoder can generate the predicted BPU 208. As described in Figure 3A, the decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate the prediction reference 224.
[0060]
[0077] In process 300B, the decoder can feed the predicted reference 224 for performing a prediction operation within the next iteration of process 300B to the spatial prediction stage 2042 or the temporal prediction stage 2044. For example, when the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly feed the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolating the next BPU of the current picture). When the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture in which all BPUs are decoded), the encoder can feed the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop-filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., for use as an inter prediction reference picture for future encoded pictures of the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, if the prediction mode indicator of the prediction data 206 indicates that inter prediction is used to encode the current BPU, the prediction data can further include loop filter parameters (e.g., the strength of the loop filter).
[0061]
[0078] FIG. 4 is a block diagram of an example of a device 400 for encoding or decoding video according to some embodiments of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may become a dedicated machine for encoding or decoding video. The processor 402 may be any kind of circuit capable of manipulating or processing information. For example, the processor 402 may include any combination of any number of central processing units ("CPUs"), graphics processing units ("GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, IP (intellectual property) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), system-on-chips (SoCs), application specific integrated circuits (ASICs), etc. In some embodiments, the processor 402 may also be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, the processor 402 may include a plurality of processors including processor 402a, processor 402b, and processor 402n.
[0062]
[0079] The machine 400 may also include a memory 404 configured to store data (such as a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. 4, the stored data may include program instructions (such as program instructions for implementing steps within processes 200A, 200B, 300A, or 300B) and processing data (such as video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and processing data (e.g., via bus 410) and execute the program instructions to perform operations or processing on the processing data. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any combination of any number of random access memories (RAMs), read-only memories (ROMs), optical disks, magnetic disks, hard drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (registered trademark) (CF) cards, etc. The memory 404 can also be a memory bank (not shown in FIG. 4) grouped as a single logical component.
[0063]
[0080] A bus 410, such as an internal bus (e.g., a CPU memory bus) or an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), can be a communication device that transfers data between components within the device 400.
[0064]
[0081] For the sake of simplicity of explanation without causing ambiguity, in the present disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuit can be implemented entirely as hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuit can be a single independent module or can be fully or partially combined within any other component of the device 400.
[0065]
[0082] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (such as the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, and the like.
[0066]
[0083] In some embodiments, device 400 may optionally further include a peripheral device interface 408 for providing connection to one or more peripheral devices. As shown in FIG. 4, the peripheral devices may include, but are not limited to, a cursor control device (such as a mouse, touch pad, or touch screen), a keyboard, a display (such as a cathode ray tube display, liquid crystal display, or light emitting diode display), a video input device (such as a camera or input interface coupled to a video archive), and the like.
[0067]
[0084] It should be noted that the video codec (such as the codec that executes processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within device 400. For example, some or all of the steps of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions loadable into memory 404. In another example, some or all of the steps of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuits (such as FPGAs, ASICs, NPUs, etc.).
[0068]
[0085] Figure 5 shows a schematic diagram of an exemplary luma mapping (LMCS) process 500 by chroma scaling according to some embodiments of the present disclosure. For example, process 500 may be used by a decoder compliant with a hybrid video coding standard (e.g., H.26x series). LMCS is a new processing block applied before the loop filter 232 of FIG. 2B. LMCS can also be called a reshaper.
[0069]
[0086] The LMCS process 500 may include in-loop mapping of luma component values and luma-dependent chroma residual scaling of chroma components based on an adaptive segmented linear model.
[0070]
[0087] As shown in FIG. 5, the in-loop mapping of luma component values based on an adaptive segmented linear model may include a forward mapping stage 518 and an inverse mapping stage 508. The luma-dependent chroma residual scaling of chroma components may include chroma scaling 520.
[0071]
[0088] Sample values before mapping or after inverse mapping can be called samples in the original region, and sample values after mapping and before inverse mapping can be called samples in the map region. When LMCS is enabled, some stages within process 500 can be performed in the map region instead of the original region. It will be understood that the forward mapping stage 518 and the inverse mapping stage 508 can be enabled / disabled at the sequence level using an SPS flag.
[0072]
[0089] As shown in FIG. 5, Q -1 &T -1 stage 504, reconstruction 506, and intra prediction 508 are performed in the map region. For example, Q -1 &T -1 stage 504 may include inverse quantization and inverse transformation, reconstruction 506 may include addition of luma prediction and luma residual, and intra prediction 508 may include luma intra prediction.
[0073]
[0090] The loop filter 510, motion compensation stages 516 and 530, intra prediction stage 528, reconstruction stage 522, and decoded picture buffers (DPBs) 512 and 526 are executed within the original (i.e., non-map) region. In some embodiments, the loop filter 510 can include deblocking, an adaptive loop filter (ALF), and sample adaptive offset (SAO), the reconstruction stage 522 can include adding a chroma prediction and a chroma residual, and the DPBs 512 and 526 can store decoded pictures as reference pictures.
[0074]
[0091] In some embodiments, luma mapping using a piecewise linear model can be applied.
[0075]
[0092] In-loop mapping of the luma component can adjust the signal statistics of the input video by redistributing the codewords over the dynamic range to improve the compression efficiency. The luma mapping can be executed by a forward mapping function “FwdMap” and a corresponding inverse mapping function “InvMap”. The “FwdMap” function is signaled using a piecewise linear model with 16 equal segments. The “InvMap” function need not be signaled and is instead derived from the “FwdMap” function.
[0076]
[0093] The signaling of the differential linear model is shown in Table 1 of FIG. 6 and Table 2 of FIG. 7. Table 1 of FIG. 6 shows the tile group header syntax structure. As shown in FIG. 6, a reshaper model parameter presence flag is signaled to indicate whether there is a luma mapping model within the current tile group. If there is a luma mapping model within the current tile group, the syntax elements shown in Table 2 of FIG. 7 can be used to signal the corresponding differential linear model parameters within tile_group_reshaper_model(). The differential linear model divides the dynamic range of the input signal into 16 equal segments. For each of the 16 equal segments, the linear mapping parameters of the segment are represented using the number of codewords assigned to the segment. Taking a 10-bit input as an example. Each of the 16 segments can have 64 codewords assigned to that segment by default. The number of codewords signaled can be used to calculate the scale factor and appropriately adjust the mapping function for that segment. Table 2 of FIG. 7 also comprehensively defines the minimum index "reshaper_model_min_bin_idx" and the maximum index "reshaper_model_max_bin_idx" at which the number of codewords is signaled. If the segment index is less than reshaper_model_min_bin_idx or greater than reshaper_model_max_bin_idx, the number of codewords for that segment is not signaled and is inferred to be zero (i.e., no codewords are assigned to that segment and no mapping / scaling is applied).
[0077]
[0094] After the `tile_group_reshaper_model()` is signaled, another reshapable flag, `tile_group_reshaper_enable_flag`, is signaled at the tile group header level to indicate whether the LMCS process shown in Figure 8 is applied to the current tile group. If the reshaver is enabled for the current tile group and the current tile group does not use dual-tree splitting, a further chroma scaling enable flag is signaled to indicate whether chroma scaling is enabled for the current tile group. Dual-tree splitting can also be referred to as a chroma separate tree.
[0078]
[0095] The segmented linear model can be constructed as follows based on the signaled syntax elements in Table 2 of Figure 7. Each i-th segment of the "FwdMap" segmented linear model, where i = 0, 1,..., 15, is defined by two input pivot points InputPivot[] and two output (mapped) pivot points MappedPivot[]. InputPivot[] and MappedPivot[] are calculated as follows based on the signaled syntax (assuming without loss of generality that the bit depth of the input video is 10 bits): 1) OrgCW = 64 2) For i = 0:16, InputPivot[i] = i * OrgCW 3) For i = reshaper_model_min_bin_idx: reshaper_model_max_bin_idx, SignaledCW[i] = OrgCW + (1 ⊕ 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i]; 4) For i = 0:16, calculate MappedPivot[i] as follows: MappedPivot[0] = 0; (For i = 0; i < 16; i++) MappedPivot[i + 1] = MappedPivot[i] + SignaledCW[i]
[0079]
[0096] The inverse mapping function "InvMap" can also be defined by InputPivot[] and MappedPivot[]. Different from "FwdMap", in the piecewise linear model of "InvMap", two input pivot points of each piece can be defined by MappedPivot[], and two output pivot points can be defined by InputPivot[], which is the opposite of "FwdMap". In this way, the input of "FwdMap" is divided into equal pieces, but it is not guaranteed that the input of "InvMap" is divided into equal pieces.
[0080]
[0097] As shown in FIG. 5, in the inter-coded block, motion compensation prediction can be executed within the mapped area. In other words, after the motion compensation prediction 516, Y can be calculated based on the reference signals in the DPB, and the "FwdMap" function 518 can be applied to map the luma prediction block in the original area to the mapped area (Y' pred = FwdMap(Y pred )). In the intra-coded block, since the reference samples used in the intra prediction are already in the mapped area, the "FwdMap" function is not applied. After the reconstructed block 506, Y pred can be calculated. The "InvMap" function 508 can be applied to convert and revert the reconstructed luma value in the mapped area to the reconstructed luma value in the original area ( r ). The "InvMap" function 508 can be applied to both the intra-coded luma block and the inter-coded luma block.
Number
[0081]
[0098] The luma mapping process (forward or inverse mapping) can be implemented using a Look-Up Table (LUT) or by on-the-fly calculation. When using an LUT, the tables "FwdMapLUT[]" and "InvMapLUT[]" can be pre-computed and stored in advance for use at the tile group level, and the forward mapping and inverse mapping can be simply implemented as FwdMap(Y pred ) = FwdMapLUT[Y pred , and InvMap(Y r ) = InvMapLUT[Y r respectively. Alternatively, on-the-fly calculation can be used. Taking the forward mapping function "FwdMap" as an example. To determine the bin to which a luma sample belongs, the sample value can be right-shifted by 6 bits (corresponding to 16 equal bins assuming a 10-bit video) to obtain the bin index. Then the linear model parameters for that bin can be retrieved and applied on-the-fly to calculate the mapped luma value. The FwdMap function is evaluated as follows: Y’pred = FwdMap(Y pred ) = ((b2 - b1) / (a2 - a1))*(Y pred - a1) + b1 where "i" is the bin index, a1 is InputPivot[i], a2 is InputPivot[i + 1], b1 is MappedPivot[i], and b2 is MappedPivot[i + 1].
[0082]
[0099] The "InvMap" function can be calculated on-the-fly in a similar way, except that a conditional check must be applied when finding the bin to which the sample value belongs because the bins within the mapped region are not guaranteed to be of equal size.
[0083]
[0100] In some embodiments, luma-dependent chroma residual scaling can be performed.
[0084]
[0101] Chroma residual scaling is designed to compensate for the interaction between the luma signal and its corresponding chroma signal. Whether chroma residual scaling is enabled is also signaled at the tile group level. As shown in Table 1 of Figure 6, when luma mapping is enabled and dual-tree splitting is not applied to the current tile group, an additional flag (e.g., tile_group_reshaper_chroma_residual_scale_flag) is signaled to indicate whether luma-dependent chroma residual scaling is enabled. When luma mapping is not used or dual-tree splitting is used within the current tile group, luma-dependent chroma residual scaling is automatically disabled. Furthermore, luma-dependent chroma residual scaling can be disabled for chroma blocks whose area is 4 or less.
[0085]
[0102] Chroma residual scaling depends on the average value of the corresponding luma prediction block (for both intra-coded blocks and inter-coded blocks). The avgY’ as the average of the luma prediction block can be calculated as follows:
Equation
[0086]
[0103] C ScaleInv is calculated in the following steps: 1) Find the index Y of the piecewise-linear model to which avgY’ belongs using the InvMap function. Idx 2) C ScaleInv = cScaleInv[Y Idx ] holds, where cScaleInv[] is a pre-computed 16-piece LUT.
[0087]
[0104] In the current LMCS method in VTM4, the pre-computed LUT cScaleInv[i] (where i is in the range of 0 to 15) is derived as follows based on the 64-entry static LUT ChromaResidualScaleLut and the value of SignaledCW[i]: ChromaResidualScaleLut
[64] ={16384, 16384, 16384, 16384, 16384, 16384, 16384, 8192, 8192, 8192, 8192, 5461, 5461, 5461, 5461, 4096, 4096, 4096, 4096, 3277, 3277, 3277, 3277, 2731, 2731, 2731, 2731, 2341, 2341, 2341, 2048, 2048, 2048, 1820, 1820, 1820, 1638, 1638, 1638, 1638, 1489, 1489, 1489, 1489, 1365, 1365, 1365, 1365, 1260, 1260, 1260, 1260, 1170, 1170, 1170, 1170, 1092, 1092, 1092, 1092, 1024, 1024, 1024, 1024}; shiftC = 11 - If -(SignaledCW[i] == 0) holds, cScaleInv[i]=(1 << shiftC) - Otherwise cScaleInv[i]=ChromaResidualScaleLut[(SignaledCW[i] >> 1)-1]
[0088]
[0105] The static table ChromaResidualScaleLut[] contains 64 entries, and SignaledCW[] is in the range of [0, 128] (assuming the input is 10 bits). Therefore, division by 2 (e.g., right shift by 1) is used to construct cScaleInv[] of the chroma scale factor LUT. cScaleInv[] of the chroma scale factor LUT can contain multiple chroma scale factors. cScaleInv[] of the LUT is constructed at the tile group level.
[0089]
[0106] When the current block is encoded using the intra, CIIP, or intra block copy (IBC, also known as current picture reference or CPR) mode, avgY’ is calculated as the average of the intra, CIIP, or IBC predicted luma values. Otherwise, avgY’ is calculated as the average of the inter predicted luma values mapped forward (i.e., Y’ in Figure 5 pred ). Different from the luma mapping performed based on samples, C ScaleInv is a constant value for all chroma blocks. C ScaleInv is used to apply chroma residual scaling on the decoder side as follows:
Number
[0090]
[0107] However
Number
[0091]
[0108] In some embodiments, dual tree splitting can be performed.
[0092]
[0109] In VVC Draft 4, the coding tree method supports the ability for luma and chroma to have separate block tree partitions. This is also called dual tree partitioning. The signaling of dual tree partitioning is shown in Table 3 of FIG. 8 and Table 4 of FIG. 9. When the sequence level control flag "qtbtt_dual_tree_intra_flag", which is signaled within the SPS, is turned on and the current tile group is intra-coded, the block partition information can be signaled separately for luma first and then for chroma. Dual tree partitioning is not permitted for inter-coded tile groups (P and B tile groups). When a separate block tree mode is applied, as shown in Table 4 of FIG. 9, the luma coding tree block (CTB) is divided into coding units (CUs) by a certain coding tree structure, and the chroma CTB is divided into chroma CUs by another coding tree structure.
[0093]
[0110] When different splits for luma and chroma are allowed, problems can arise with coding tools that have dependencies between various color components. For example, in the case of LMCS, the average value of the corresponding luma block is used to find the scale factor applied to the current block. When a dual tree is used, this can lead to latency for the entire CTU. For example, if the luma blocks of a CTU are split vertically once and the chroma blocks of the CTU are split horizontally once, both luma blocks of the CTU are decoded before the first chroma block of the CTU can be decoded (thereby allowing the calculation of the average value necessary to calculate the chroma scale factor). In VVC, a CTU can be as large as 128x128 in units of luma samples. Such a large latency can be a significant problem for the pipeline design of a hardware decoder. Therefore, Draft 4 of VVC can prohibit the combination of dual tree splitting and luma-dependent chroma scaling. When dual tree splitting is enabled for the current tile group, chroma scaling can be forced off. Note that the luma mapping part of LMCS only affects the luma component and does not have problems with dependencies across color components, so it continues to be allowed even in the case of a dual tree. Another example of a coding tool that relies on dependencies between color components to achieve better coding efficiency is called the cross-component linear model (CCLM).
[0094]
[0111] Therefore, the derivation of cScaleInv[] of the chroma scale factor LUT at the tile group level cannot be easily extended. The derivation process currently depends on a fixed chroma LUT ChromaResidualScaleLut with 64 entries. For a 10-bit video with 16 segments, an additional step of division by 2 must be applied. If the number of segments changes, for example, if 8 segments are used instead of 16 segments, the derivation process must be changed to apply division by 4 instead of division by 2. This additional step not only causes a loss of precision but is also unrefined and unnecessary.
[0095]
[0112] Furthermore, the average value of all luma blocks can be used to calculate the partition index Y Idx of the current chroma block, which is used to obtain the chroma scale factor. This is undesirable and often unnecessary. Consider the maximum CTU size of 128x128. In this case, the average luma value is calculated based on 16384 (128x128) luma samples, and such calculation is complex. Furthermore, if a luma block partition of 128x128 is selected by the encoder, the block is likely to contain homogeneous content. Therefore, a subset of luma samples in the block may be sufficient to calculate the luma average.
[0096]
[0113] In dual tree partitioning, chroma scaling can be off to avoid potential pipeline problems in the hardware decoder. However, this dependency can be avoided if explicit signaling is used to indicate the chroma scale factor to be applied, instead of using the corresponding luma samples to derive the chroma scale factor to be applied. Enabling chroma scaling in intra-coded tile groups can further improve coding efficiency.
[0097]
[0114] The signaling of piecewise linear parameters can be further improved. Currently, delta codeword values are signaled for each of the 16 partitions. It is recognized that often only a limited number of different codewords are used for the 16 partitions. Thus, the signaling overhead can be further reduced.
[0098]
[0115] Embodiments of the present disclosure provide a method for processing video content by removing chroma scaling LUTs.
[0099]
[0116] As described above, the expansion of a 64-entry chroma LUT may be difficult and may be problematic when other partitioning linear models are used (e.g., 8-partition, 4-partition, 64-partition, etc.). To achieve the same coding efficiency, such an expansion is not necessary because the chroma scale factor can be set to the same as the corresponding luma scale factor of the partition. In some embodiments of the present disclosure, the chroma scale factor (chroma_scaling) can be determined based on the partition index "Y" of the current chroma block as follows. Idx ". · If Y Idx >reshaper_model_max_bin_idx, Y Idx <reshaper_model_min_bin_idx or SignaledCW[Y Idx =0 holds, set chroma_scaling to the default and chroma_scaling = 1.0. · Otherwise, set chroma_scaling to SignaledCW[Y Idx / OrgCW.
[0100]
[0117] When chroma_scaling = 1.0, scaling is not applied.
[0101]
[0118] The chroma scale factor determined above may have fractional precision. It will be appreciated that a fixed-point approximation can be applied to avoid dependencies on the hardware / software platform. Further, inverse chroma scaling can be performed on the decoder side. Therefore, division can be implemented by a fixed-point operation using multiplication followed by a right shift. The inverse chroma scale factor "inverse_chroma_scaling[]" at fixed-point precision can be determined as follows based on the number of bits within the fixed-point approximation "CSCALE_FP_PREC". inverse_chroma_scaling[Y Idx=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(SignaledCW[Y Idx >>1)) / SignaledCW[Y Idx Here, luma_bit_depth is the luma bit depth, and TOTAL_NUMBER_PIECES is the total number of segments in the segment linear model set to 16 in VVC draft 4. The value of "inverse_chroma_scaling[]" may only need to be calculated once per tile group, and it should be understood that the above division is an integer division operation.
[0102]
[0119] Further quantization can be applied to determine the chroma scale factor and the inverse scale factor. For example, the inverse chroma scale factor can be calculated for all even (2×m) values of "SignaledCW", and the odd (2×m + 1) values of "SignaledCW" reuse the chroma scale factor of the adjacent even value's scale factor. In other words, the following can be used: for(i=reshaper_model_min_bin_idx; i<=reshaper_model_max_bin_idx; i++) { tempCW=SignaledCW[i]>>1)<<1; inverse_chroma_scaling[i]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(tempCW>>1)) / tempCW; }
[0103]
[0120] The quantization of the chroma scale factor can be further generalized. For example, the inverse chroma scale factor "inverse_chroma_scaling[]" can be calculated for the n-th value of "SignaledCW" while all other adjacent values share the same chroma scale factor. For example, "n" can be set to 4. Therefore, the same inverse chroma scale factor value can be shared for every four adjacent codeword values. In some embodiments, the value of "n" can be a power of 2, and such a setting enables the use of shifts to calculate divisions. Representing the value of log2(n) as LOG2_n, the above equation "tempCW = SignaledCW[i] >> 1) << 1" can be modified as follows: tempCW = SignaledCW[i] >> LOG2_n) << LOG2_n
[0104]
[0121] In some embodiments, the value of LOG2_n can be a function of the number of partitions used within the piecewise linear model. When fewer partitions are used, it may be beneficial to use a larger LOG2_n. For example, when TOTAL_NUMBER_PIECES is 16 or less, LOG2_n can be set to 1 + (4 - log2(TOTAL_NUMBER_PIECES)). When TOTAL_NUMBER_PIECES exceeds 16, LOG2_n can be set to 0.
[0105]
[0122] Embodiments of the present disclosure provide a method for processing video content by simplifying the averaging of luma prediction blocks.
[0106]
[0123] As discussed above, to determine the partition index "Y Idx " of the current chroma block, the average value of the corresponding luma block can be used. However, for large block sizes, the averaging process can include a large number of luma samples. In the worst case, 128x128 luma samples can be involved in the averaging process.
[0107]
[0124] Embodiments of the present disclosure provide a simplified averaging process for reducing the use of only NxN luma samples (where N is a power of 2) in the worst case scenario.
[0108]
[0125] In some embodiments, if both dimensions of a 2D luma block are less than or equal to a preset threshold M (in other words, at least one of the two dimensions exceeds M), "downsampling" can be applied to use only the positions of M within that dimension. Without loss of generality, taking the horizontal dimension as an example. If the width exceeds M, only the samples at positions x (x = i×(width>>log2(M)), i = 0, … M-1) are used for averaging.
[0109]
[0126] FIG. 10 shows an example of applying the proposed simplification to calculate the average of a 16x8 luma block. In this example, M is set to 4, and only 16 luma samples (shaded samples) within the block are used for averaging. It should be understood that the preset threshold M is not limited to 4, and M can be set to any value that is a power of 2. For example, the preset threshold M can be 1, 2, 4, 8, etc.
[0110]
[0127] In some embodiments, the horizontal and vertical dimensions of a luma block can have different preset thresholds M. In other words, in the worst case of the averaging operation, M1xM2 samples can be used.
[0111]
[0128] In some embodiments, the number of samples within the averaging process can be limited without considering the dimensions. For example, a maximum of 16 samples can be used, and those samples can be distributed within the horizontal or vertical dimension in a 1x16, 16x1, 2x8, 8x2, or 4x4 format, and any format that conforms to the shape of the current block can be selected. For example, if the block is vertically long, a matrix of 2x8 samples can be used, if the block is horizontally long, a matrix of 8x2 samples can be used, and if the block is square, a matrix of 4x4 samples can be used.
[0112]
[0129] When the size of the large block is selected, it will be understood that the content within the block tends to be more homogeneous. Thus, the above simplification can cause a difference between the average value and the true average of all luma blocks, but such a difference can be small.
[0113]
[0130] Further, decoder-side motion vector refinement (DMVR) requires the decoder to perform motion detection to derive a motion vector before motion compensation can be applied. Thus, the DMVR mode can be particularly complex for the decoder within the VVC standard. Bidirectional optical flow (BDOF) is an additional sequential process that must be applied after DMVR to obtain a luma prediction block, so the BDOF mode within the VVC standard can further complicate this situation. Chroma scaling requires the average value of the corresponding luma prediction block, so DMVR and BDOF can be applied before the average value can be calculated.
[0114]
[0131] To solve this latency problem, in some embodiments of the present disclosure, a luma prediction block is used to calculate an average luma value before DMVR and BDOF, and the average luma value is used to obtain a chroma scale factor. This enables chroma scaling to be applied in parallel with the DMVR and BDOF processes, thus significantly reducing latency.
[0115]
[0132] In accordance with the present disclosure, modified forms of latency reduction can be considered. In some embodiments, this latency reduction can also be combined with the above-described simplified averaging process that uses only a part of the luma prediction block to calculate the average luma value. In some embodiments, the luma prediction block can be used after the DMVR process and before the BDOF process to calculate the average luma value. Then, the average luma value is used to obtain the chroma scale factor. This design enables the application of chroma scaling in parallel with the BDOF process while maintaining the accuracy of determining the chroma scale factor. Since the DMVR process may refine the motion vectors, it may be more accurate to use the predicted samples with the refined motion vectors after the DMVR process than with the motion vectors before the DMVR process.
[0116]
[0133] Furthermore, in the VVC standard, the CU syntax structure “coding_unit()” includes the syntax element “cu_cbf” for indicating whether there are any non-zero residual coefficients in the current CU. At the TU level, the TU syntax structure “transform_unit()” includes the syntax elements “tu_cbf_cb” and “tu_cbf_cr” for indicating whether there are any non-zero chroma (Cb or Cr) residual coefficients in the current TU. Conventionally, in draft 4 of VVC, when chroma scaling is enabled at the tile group level, the averaging of the corresponding luma block is always called.
[0117]
[0134] Embodiments of the present disclosure further provide a method for processing video content by bypassing the luma averaging process. In accordance with the disclosed embodiments, since the chroma scaling process is applied to the residual chroma coefficients, the luma averaging process can be bypassed when there are no non-zero chroma coefficients. This can be determined based on the following conditions: Condition 1: cu_cbf is equal to 0 Condition 2: both tu_cbf_cr and tu_cbf_cb are equal to 0
[0118]
[0135] As discussed above, "cu_cbf" can indicate whether there are any non-zero residual coefficients in the current CU, and "tu_cbf_cb" and "tu_cbf_cr" can indicate whether there are any non-zero chroma (Cb or Cr) residual coefficients in the current TU. When Condition 1 or Condition 2 is satisfied, the luma averaging process can be bypassed.
[0119]
[0136] In some embodiments, only the NxN samples of the prediction block are used to derive the average value, which simplifies the averaging process. For example, when N is equal to 1, only the top-left sample of the prediction block is used. However, this simplified averaging process using the prediction block still requires the generation of the prediction block, thereby causing latency.
[0120]
[0137] In some embodiments, the reference luma samples can be directly used to generate the chroma scale factor. This enables the decoder to derive the scale factor in parallel with the luma prediction process, thereby reducing latency. Hereinafter, intra prediction and inter prediction using the reference luma samples will be described separately.
[0121]
[0138] In an exemplary intra prediction, the decoded adjacent samples within the same picture can be used as reference samples for generating the prediction block. These reference samples can include, for example, the samples above the current block, the samples to the left of the current block, or the top-left sample of the current block. The average of these reference samples can be used to derive the chroma scale factor. In some embodiments, the average of some of these reference samples can be used. For example, only the K reference samples (e.g., K = 3) closest to the top-left position of the current block are averaged.
[0122]
[0139] In an exemplary inter prediction, reference samples from a temporal reference picture can be used to generate a prediction block. These reference samples are identified by a reference picture index and a motion vector. When the motion vector has fractional precision, interpolation can be applied. The reference samples used to determine the average of the reference samples can include the reference samples before or after interpolation. The reference samples before interpolation can include motion vectors that are clipped to integer precision. Consistent with the disclosed embodiments, the average can be calculated using all of the reference samples. Alternatively, the average can be calculated using only a portion of the reference samples (e.g., the reference samples corresponding to the upper left position of the current block).
[0123]
[0140] As shown in FIG. 5, while performing inter prediction within the original region, intra prediction (e.g., intra prediction 514 or 528) can be performed within the reshaped region. Thus, in inter prediction, forward mapping can be applied to the prediction block, and the average is calculated using the luma prediction block after forward mapping. To reduce latency, the average can be calculated using the prediction block before forward mapping. For example, the block before forward mapping, the NxN portion of the block before forward mapping, or the upper left sample of the block before forward mapping can be used.
[0124]
[0141] Embodiments of the present disclosure further provide a method for processing video content with chroma scaling for dual tree splitting.
[0125]
[0142] Since the dependency on luma can cause the hardware design to be complicated, chroma scaling can be turned off for the intra-coded tile groups that enable dual tree splitting. However, this limitation can cause a loss of coding efficiency. The sample values of the corresponding luma blocks are averaged to calculate avgY’, and the segmentation index Y Idx is determined, and the chroma scale factor inverse_chroma_scaling[YIdx Instead of obtaining [], the chroma scale factor can be explicitly signaled in the bitstream to avoid the dependency on luma in the case of dual-tree splitting.
[0126]
[0143] The chroma scale index can be signaled at various levels. For example, as shown in Table 5 of FIG. 11, the chroma scale index can be signaled at the coding unit (CU) level together with the chroma prediction mode. To determine the chroma scale factor of the current chroma block, the syntax element "lmcs_scaling_factor_idx" can be used. If there is no "lmcs_scaling_factor_idx", it can be inferred that the chroma scale factor of the current chroma block is equal to 1.0 with floating-point precision or equivalently (1<<CSCALE_FP_PREC) with fixed-point precision. The range of allowable values of "lmcs_chroma_scaling_idx" is determined at the tile group level and will be described later.
[0127]
[0144] Depending on the possible values of 「lmcs_chroma_scaling_idx」, the signaling cost can be high, especially for very small blocks. Thus, in some embodiments of the present disclosure, the signaling conditions in Table 5 of FIG. 11 may additionally include block size conditions. For example (emphasized in italics and shaded in gray), this syntax element 「lmcs_chroma_scaling_idx」 can be signaled only if the current block contains more than a given number of chroma samples, or if the current block has a width greater than a given width W or a height greater than a given height H. For smaller blocks, if 「lmcs_chroma_scaling_idx」 is not signaled, the decoder side can determine the chroma scale factor. In some embodiments, the chroma scale factor can be set to 1.0 with floating point precision. In some embodiments, a default 「lmcs_chroma_scaling_idx」 value can be added at the tile group header level (see 1 in FIG. 6). Small blocks without a signaled 「lmcs_chroma_scaling_idx」 can use this tile group level default index to derive the corresponding chroma scale factor. In some embodiments, the chroma scale factor of a small block can be inherited from its vicinity (e.g., the vicinity of the top or left side) where the scale factor is explicitly signaled.
[0128]
[0145] In addition to signaling the syntax element "lmcs_chroma_scaling_idx" at the CU level, this syntax element can also be signaled at the CTU level. However, given that the maximum CTU size in VVC is 128x128, performing the same scaling at the CTU level may be too coarse. Therefore, in some embodiments of the present disclosure, this syntax element "lmcs_chroma_scaling_idx" can be signaled using a fixed granularity. For example, one "lmcs_chroma_scaling_idx" is signaled for each 16x16 region within the CTU and applied to all samples within that 16x16 region.
[0129]
[0146] The range of "lmcs_chroma_scaling_idx" for the current tile group depends on the number of values of the chroma scale factor permitted within the current tile group. The number of values of the chroma scale factor permitted within the current tile group can be determined based on the 64-entry chroma LUT as discussed above. Alternatively, the number of values of the chroma scale factor permitted within the current tile group can be determined using the chroma scale factor calculation discussed above.
[0130]
[0147] For example, in the "quantization" method, the value of LOG2_n can be set to 2 (i.e., "n" is set to 4), and the codeword assignment for each segment in the piecewise linear model of the current tile group can be set as follows: {0, 65, 66, 64, 67, 62, 62, 64, 64, 64, 67, 64, 64, 62, 61, 0}. Then any codeword value from 64 to 67 can have the same scale factor value (1.0 with decimal precision), and any codeword value from 60 to 63 can have the same scale factor value (60 / 64 = 0.9375 with decimal precision). Therefore, there are only two possible scale factor values for all tile groups. In the two end segments where no codewords are assigned, the chroma scale factor is set to 1.0 by default. Thus, in this example, 1 bit is sufficient to signal "lmcs_chroma_scaling_idx" for the blocks within the current tile group.
[0131]
[0148] In addition to determining the possible chroma scale factor values using the piecewise linear model, the encoder can signal a set of chroma scale factor values in the tile group header. Then at the block level, the value of the chroma scale factor for the block can be determined using that set of chroma scale factor values and the value of "lmcs_chroma_scaling_idx" of the block.
[0132]
[0149] CABAC coding can be applied to code 「lmcs_chroma_scaling_idx」. The CABAC context of a block may depend on the 「lmcs_chroma_scaling_idx」 of the adjacent blocks of the block. For example, the left block or the upper block can be used to form the CABAC context. Regarding the binarization of this syntax element of 「lmcs_chroma_scaling_idx」, the same truncated Rice binarization applied to the ref_idx_10 and ref_idx_11 syntax elements in Draft 4 of VVC can be used to binarize 「lmcs_chroma_scaling_idx」.
[0133]
[0150] The advantage of signaling 「chroma_scaling_idx」 is that the coder can select the best 「lmcs_chroma_scaling_idx」 regarding the rate-distortion cost. Using rate-distortion optimization to select 「lmcs_chroma_scaling_idx」 can improve the coding efficiency, which can help offset the increase in signaling cost.
[0134]
[0151] Embodiments of the present disclosure further provide a method for processing video content involving signaling of an LMCS partitioned linear model.
[0135]
[0152] The LMCS method uses a partitioned linear model with 16 partitions, but the number of unique values of 「SignaledCW[i]」 within a tile group tends to be much less than 16. For example, some of the 16 partitions can use the default number of codewords 「OrgCW」, and some of the 16 partitions can have the same number of codewords as each other. Therefore, an alternative method of signaling the LMCS partitioned linear model may include signaling the number of unique codewords 「listUniqueCW[]」 and transmitting an index for each of the partitions to indicate the elements of 「listUniqueCW[]」 of the current partition.
[0136]
[0153] The corrected syntax table is shown in FIG. 12. In Table 6 of FIG. 12, new or corrected syntax is emphasized in italics and with a gray shade.
[0137]
[0154] The semantic rules of the disclosed signaling method are as follows, with the changed parts underlined: reshaper_model_min_bin_idx defines the minimum bin (or section) index used within the reshaper construction process. The value of reshape_model_min_bin_idx shall be within the range from 0 to MaxBinIdx. The value of MaxBinIdx shall be equal to 15. reshaper_model_delta_max_bin_idx defines the maximum bin index used within the reshaper construction process, which is the maximum allowable bin (or section) index MaxBinIdx minus. The value of reshape_model_max_bin_idx is set equal to MaxBinIdx - reshape_model_delta_max_bin_idx. reshaper_model_bin_delta_abs_cw_prec_minus1 plus 1 defines the number of bits used for the representation of the syntax reshape_model_bin_delta_abs_CW[i]. reshaper_model_bin_num_unique_cw_minus1 plus 1 defines the size of the codeword array listUniqueCW. reshaper_model_bin_delta_abs_CW[i] defines the absolute delta codeword value of the i-th bin. reshaper_model_bin_delta_sign_CW_flag[i] defines the sign of reshape_model_bin_delta_abs_CW[i] as follows: - When reshape_model_bin_delta_sign_CW_flag[i] is equal to 0, the corresponding variable RspDeltaCW[i] is a positive value. - Otherwise (if reshape_model_bin_delta_sign_CW_flag[i] is not equal to 0), the corresponding variable RspDeltaCW[i] is negative. If reshape_model_bin_delta_sign_CW_flag[i] does not exist, it is inferred that the corresponding variable RspDeltaCW[i] is equal to 0. The variable RspDeltaCW[i] is derived as RspDeltaCW[i] = (1 - 2 * reshape_model_bin_delta_sign_CW[i]) * reshape_model_bin_delta_abs_CW[i]. The variable listUniqueCW[0] is set equal to OrgCW. The variable listUniqueCW[i] for i = 1... reshaper_model_bin_num_unique_cw_minus1 is It is derived as follows: - Set the variable OrgCW equal to (1 << BitDepth Y ) / (MaxBinIdx + 1). - listUniqueCW[i] = OrgCW + RspDeltaCW[i - 1] reshaper_model_bin_cw_idx[i] defines the index of the array listUniqueCW[] used to derive RspCW[i]. The value of reshaper_model_bin_cw_idx[i] shall be within the range from 0 to (reshaper_model_bin_num_unique_cw_minus1 + 1). RspCW[i] is derived as follows: - When reshaper_model_min_bin_idx <= i <= reshaper_model_max_bin_idx holds, RspCW[i] = listUniqueCW[reshaper_model_bin_cw_idx[i]]. - Otherwise RspCW[i] = 0. BitDepth Y When the value of is equal to 10, the value of RspCW[i] can be within the range of 32 to 2 * OrgCW - 1.
[0138]
[0155] Embodiments of the present disclosure further provide a method for processing video content with conditional chroma scaling at the block level.
[0139]
[0156] As shown in Table 1 of FIG. 6, whether chroma scaling is applied can be determined by the "tile_group_reshaper_chroma_residual_scale_flag" signaled at the tile group level.
[0140]
[0157] However, it may be beneficial to determine whether to apply chroma scaling at the block level. For example, in some of the disclosed embodiments, a CU level flag can be signaled to indicate whether chroma scaling is applied to the current block. The presence of the CU level flag can be conditioned based on the tile group level flag "tile_group_reshaper_chroma_residual_scale_flag". That is, the CU level flag can only be signaled when chroma scaling is permitted at the tile group level. The coder is permitted to choose whether to use chroma scaling based on whether it is beneficial for the current block, but this can also incur significant signaling overhead.
[0141]
[0158] Consistent with the disclosed embodiments, to avoid the above-mentioned signaling overhead, whether chroma scaling is applied to a block can be conditioned based on the prediction mode of the block. For example, when a block is inter-predicted, especially when its reference picture is close in terms of temporal distance, the predicted signal tends to be good. Therefore, since the residual is expected to be very small, chroma scaling can be bypassed. For example, pictures at a higher temporal level tend to have reference pictures that are close in terms of temporal distance. For a block, chroma scaling can be disabled within the picture that uses a nearby reference picture. To determine whether this condition is met, the difference in the picture order count (POC) between the current picture and the reference picture of the block can be used.
[0142]
[0159] In some embodiments, chroma scaling can be disabled for all blocks to be inter-coded. In some embodiments, chroma scaling can be disabled for the combined intra / inter prediction (CIIP) mode defined within the VVC standard.
[0143]
[0160] In the VVC standard, the CU syntax structure “coding_unit()” includes a syntax element “cu_cbf” for indicating whether there are any non-zero residual coefficients within the current CU. At the TU level, the TU syntax structure “transform_unit()” includes syntax elements “tu_cbf_cb” and “tu_cbf_cr” for indicating whether there are any non-zero chroma (Cb or Cr) residual coefficients within the current TU. The chroma scaling process can be conditioned based on these flags. As described above, when there are no non-zero residual coefficients, the averaging of the corresponding luma-chroma scaling process can be invoked. By invoking the averaging, the chroma scaling process can be bypassed.
[0144]
[0161] FIG. 13 shows a flowchart of a method 1300 implemented by a computer for processing video content. In some embodiments, method 1300 may be executed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 1300 at the picture level. For example, the device can process pictures one by one within method 1300. In another example, the device can process a plurality of pictures at a time within method 1300. Method 1300 can include steps as follows.
[0145]
[0162] At step 1302, chroma blocks and luma blocks associated with a picture can be received. It will be understood that a picture can be related to a chroma component and a luma component. Thus, a picture can be related to chroma blocks including chroma samples and luma blocks including luma samples.
[0146]
[0163] In step 1304, luma scale information related to the luma block can be determined. In some embodiments, the luma scale information can be a syntax element signaled within the data stream of the picture, or a variable derived based on the syntax element signaled within the data stream of the picture. For example, the luma scale information can include "reshape_model_bin_delta_sign_CW[i] and reshape_model_bin_delta_abs_CW[i]" described in the above equations, and / or "SignaledCW[i]" described in the above equations, etc. In some embodiments, the luma scale information can include variables determined based on the luma block. For example, the average luma value can be determined by calculating the average value of luma samples adjacent to the luma block (such as luma samples within the row above the luma block and luma samples within the column to the left of the luma block, etc.).
[0147]
[0164] In step 1306, a chroma scale factor can be determined based on the luma scale information.
[0148]
[0165] In some embodiments, the luma scale factor of the luma block can be determined based on the luma scale information. For example, according to the above equation of "inverse_chroma_scaling[i]=((1<<(luma_bit_depth-log2(TOTAL_NUMBER_PIECES)+CSCALE_FP_PREC))+(tempCW>>1)) / tempCW", the luma scale factor can be determined based on the luma scale information (such as "tempCW"). Then, the chroma scale factor can be further determined based on the value of the luma scale factor. For example, the chroma scale factor can be set equal to the value of the luma scale factor. It will be understood that further calculations can be applied to the value of the luma scale factor before setting it as the chroma scale factor. As another example, the chroma scale factor is "SignaledCW[Y Idxcan be set equal to " / OrgCW", and determine the segmentation index "Y" of the current chroma block based on the average luma value related to the luma block Idx ".
[0149]
[0166] In step 1308, the chroma block can be processed using a chroma scale factor. For example, to generate a scaled residual of the chroma block, the residual of the chroma block can be processed using the chroma scale factor. The chroma block can be a Cb chroma component or a Cr chroma component.
[0150]
[0167] In some embodiments, the chroma block can be processed when a condition is satisfied. For example, the condition can include that the target coding unit related to the picture has no non-zero residual, or the target transform unit related to the picture has no non-zero chroma residual. Whether the target coding unit has no non-zero residual can be determined based on the value of the first coding block flag of the target coding unit. Whether the target transform unit has no non-zero chroma residual can be determined based on the value of the second coding block flag of the first component and the value of the third coding block flag of the second component of the target transform unit. For example, the first component can be the Cb component, and the second component can be the Cr component.
[0151]
[0168] It will be understood that each step of method 1300 can be executed as an independent method. For example, the method for determining the chroma scale factor described in step 1308 can be executed as an independent method.
[0152]
[0169] FIG. 14 shows a flowchart of a method 1400 implemented by a computer for processing video content. In some embodiments, method 1300 may be executed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 1400 at the picture level. For example, the device can process pictures one by one within method 1400. In another example, the device can process a plurality of pictures at a time within method 1400. Method 1400 can include the following steps.
[0153]
[0170] In step 1402, chroma blocks and luma blocks related to a picture can be received. It will be understood that a picture can be related to a chroma component and a luma component. Thus, a picture can be related to chroma blocks including chroma samples and luma blocks including luma samples. In some embodiments, a luma block can include NxM luma samples. N can be the width of the luma block, and M can be the height of the luma block. As discussed above, the luma samples of the luma block can be used to determine the partition index of the target chroma block. Thus, luma blocks related to the pictures of the video sequence can be received. It will be understood that N and M can have the same value.
[0154]
[0171] In step 1404, a subset of NxM luma samples can be selected in response to at least one of N and M exceeding a threshold value. To accelerate the determination of the segmentation index, if certain conditions are met, the luma block can be "downsampled". In other words, the segmentation index can be determined using a subset of the luma samples within the luma block. In some embodiments, the certain condition is that at least one of N and M exceeds the threshold value. In some embodiments, the threshold value can be based on at least one of N and M. The threshold value can be a power of 2. For example, the threshold value can be 4, 8, 16, etc. Taking 4 as an example, when N or M exceeds 4, a subset of the luma samples can be selected. In the example of FIG. 10, both the width and height of the luma block exceed the threshold value of 4, and thus a subset of 4x4 samples is selected. It will be understood that subsets such as 2x8, 1x16, etc. can also be selected for processing.
[0155]
[0172] In step 1406, the average value of a subset of NxM luma samples can be determined.
[0156]
[0173] In some embodiments, determining the average value can further include determining whether a second condition is met and, in response to determining that the second condition is met, determining the average value of a subset of NxM luma samples. For example, the second condition can include that the target coding unit related to the picture has no non-zero residual coefficients or that the target coding unit has no non-zero chroma residual coefficients.
[0157]
[0174] In step 1408, a chroma scale factor based on the average value can be determined. In some embodiments, to determine the chroma scale factor, the segmentation index of the chroma block can be determined based on the average value, it can be determined whether the segmentation index of the chroma block satisfies a first condition, and the chroma scale factor can be set to a default value according to the segmentation index of the chroma block satisfying the first condition. The default value may indicate that chroma scaling is not applied. For example, the default value can be 1.0 with decimal precision. It will be understood that a fixed-point approximation can be applied to the default value. According to the segmentation index of the chroma block not satisfying the first condition, the chroma scale factor can be determined based on the average value. More specifically, the chroma scale factor can be set to SignaledCW[Y Idx / OrgCW, and the segmentation index "Y Idx " of the target chroma block can be determined based on the average value of the corresponding luma block.
[0158]
[0175] In some embodiments, the first condition may include that the segmentation index of the chroma block is greater than the maximum index of the coded word to be signaled or less than the minimum index of the coded word to be signaled. The maximum index and the minimum index of the coded word to be signaled can be determined as follows.
[0159]
[0176] The codeword can be generated using a piecewise linear model (e.g., LMCS) based on an input signal (e.g., a luma sample). As discussed above, the dynamic range of the input signal can be divided into several segments (e.g., 16 segments), and each segment of the input signal can be used to generate a bin of the codeword as an output. Therefore, each bin of the codeword can have a bin index corresponding to the segment of the input signal. In this example, the range of the bin index can be from 0 to 15. In some embodiments, the value of the output (i.e., the codeword) is between a minimum value (e.g., 0) and a maximum value (e.g., 255), and a plurality of codewords having values between the minimum value and the maximum value can be signaled. The bin indices of the plurality of signaled codewords can be determined. Among the bin indices of the plurality of signaled codewords, the maximum bin index and the minimum bin index of the bins of the plurality of signaled codewords can be further determined.
[0160]
[0177] In addition to the chroma scale factor, method 1400 can further determine the luma scale factor based on the bins of the plurality of signaled codewords. The luma scale factor can be used as an inverse chroma scale factor. The equation for determining the luma scale factor has been described above, and its description is omitted herein. In some embodiments, a plurality of adjacent signaled codewords share the luma scale factor. For example, two or four adjacent signaled codewords can share the same luma scale factor, which can reduce the load of determining the luma scale factor.
[0161]
[0178] In step 1410, the chroma block can be processed using the chroma scale factor. As discussed above with respect to FIG. 5, a plurality of chroma scale factors can be used to construct a chroma scale factor LUT at the tile group level and can be applied on the decoder side to the reconstructed chroma residual of the target block. Similarly, the chroma scale factor can also be applied on the encoder side.
[0162]
[0179] It will be appreciated that each step of method 1400 can be executed as an independent method. For example, the method for determining the chroma scale factor described in step 1308 can be executed as an independent method.
[0163]
[0180] FIG. 15 shows a flowchart of a method 1500 implemented by a computer for processing video content. In some embodiments, method 1500 may be executed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of the device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 1500 at the picture level. For example, the device can process the pictures one by one within method 1500. In another example, the device can process a plurality of pictures at a time within method 1500. Method 1500 can include steps as follows.
[0164]
[0181] In step 1502, it can be determined whether there is a chroma scale index in the received video data.
[0165]
[0182] In step 1504, in response to determining that there is no chroma scale index in the received video data, it can be determined that no chroma scaling is applied to the received video data.
[0166]
[0183] In step 1506, in response to determining that there is chroma scaling in the received video data, a chroma scale factor can be determined based on the chroma scale index.
[0167]
[0184] FIG. 16 shows a flowchart of a method 1600 implemented by a computer for processing video content. In some embodiments, method 1600 can be executed by a codec (e.g., the encoder of FIGS. 2A-2B or the decoder of FIGS. 3A-3B). For example, the codec can be implemented as one or more software or hardware components of a device (e.g., device 400) for encoding a video sequence or converting it to another code. In some embodiments, the video sequence can be an uncompressed video sequence (e.g., video sequence 202) or a compressed video sequence to be decoded (e.g., video stream 304). In some embodiments, the video sequence can be a surveillance video sequence that can be captured by a surveillance device (e.g., the video input device of FIG. 4) associated with a processor of a device (e.g., processor 402). The video sequence can include a plurality of pictures. The device can execute method 1600 at the picture level. For example, the device can process pictures one by one within method 1600. In another example, the device can process a plurality of pictures at a time within method 1600. Method 1600 can include steps as follows.
[0168]
[0185] In step 1602, a plurality of unique codewords used for the dynamic range of the input video signal can be received.
[0169]
[0186] In step 1604, an index can be received.
[0170]
[0187] In step 1606, at least one of the plurality of unique codewords can be selected based on the index.
[0171]
[0188] In step 1608, a chroma scale factor can be determined based on at least one selected codeword.
[0172]
[0189] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) to perform the above method. Common non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROM, EPROM, flash EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and networked versions of those. The device can include one or more processors (CPUs), an input / output interface, a network interface, and / or a memory.
[0173]
[0190] It will be understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. When implemented by software, the software can be stored in the above computer-readable media. When executed by a processor, the software can perform the disclosed method. The computing units and other functional units described in this disclosure can be implemented by hardware, or software, or a combination of hardware and software. It will be understood by those skilled in the art that a plurality of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into a plurality of sub-modules / sub-units.
[0174]
[0191] Embodiments can be further described using the following clauses: 1. A computer-implemented method for processing video content, comprising: Receiving chroma blocks and luma blocks related to a picture, Determining luma scale information related to the luma blocks, Determining a chroma scale factor based on the luma scale information, and Processing the chroma blocks using the chroma scale factor A method comprising. 2. Determining the chroma scale factor based on the luma scale information is Determining a luma scale factor of the luma blocks based on the luma scale information, Determining the chroma scale factor based on the value of the luma scale factor The method according to clause 1, further comprising. 3. Determining the chroma scale factor based on the value of the luma scale factor is Setting the chroma scale factor equal to the value of the luma scale factor The method according to clause 2, further comprising. 4. Processing the chroma blocks using the chroma scale factor is Determining whether a first condition is satisfied, and Processing the chroma blocks using the chroma scale factor in response to a determination that the first condition is satisfied, or Bypassing the processing of the chroma blocks using the chroma scale factor in response to a determination that the first condition is not satisfied Performing one of The method according to any one of clauses 1 to 3, further comprising. 5. The first condition is That the target coding unit related to the picture has no non-zero residual, or That the target transform unit related to the picture has no non-zero chroma residual The method according to clause 4, comprising. 6. That the target coding unit has no non-zero residual is determined based on the value of the first coding block flag of the target coding unit, The fact that the target transformation unit has no non-zero chroma residual is determined based on the value of the second coding block flag of the first chroma component of the target transformation unit and the value of the third coding block flag of the second luma-chroma component. The method according to clause 5. 7. The value of the first coding block flag is 0, and the values of the second coding block flag and the third coding block flag are 0. The method according to clause 6. 8. Processing a chroma block using a chroma scale factor includes processing the residual of the chroma block using the chroma scale factor and is the method according to any one of clauses 1 to 7. 9. A device for processing video content, comprising a memory storing a set of instructions, coupled to the memory, receiving chroma blocks and luma blocks related to a picture, determining luma scale information related to the luma block, determining a chroma scale factor based on the luma scale information, and processing the chroma block using the chroma scale factor and a processor configured to execute a set of instructions for causing the device to perform the above. The device. 10. When determining the chroma scale factor based on the luma scale information, the device is further configured to execute a set of instructions for causing the device to determine the luma scale factor of the luma block based on the luma scale information, and determine the chroma scale factor based on the value of the luma scale factor. The device according to clause 9. 11. When determining the chroma scale factor based on the value of the luma scale factor, the device is configured to set the chroma scale factor equal to the value of the luma scale factor. The apparatus according to clause 10, wherein the processor is configured to execute a set of instructions for causing the apparatus to further perform. 12. When processing a chroma block using a chroma scale factor, determining whether a first condition is satisfied, and processing the chroma block using the chroma scale factor in response to a determination that a second condition is satisfied, or bypassing the processing of the chroma block using the chroma scale factor in response to a determination that the second condition is not satisfied performing one of The apparatus according to any one of clauses 9 to 11, wherein the processor is configured to execute a set of instructions for causing the apparatus to further perform. 13. The first condition is that a target coding unit related to a picture has no non-zero residual, or that a target transform unit related to a picture has no non-zero chroma residual The apparatus according to clause 12. 14. That a target coding unit has no non-zero residual is determined based on the value of a first coding block flag of the target coding unit, and that a target transform unit has no non-zero chroma residual is determined based on the value of a second coding block flag of a first chroma component of the target transform unit and the value of a third coding block flag of a second chroma component of the target transform unit, The apparatus according to clause 13. 15. The value of the first coding block flag is 0, and the values of the second coding block flag and the third coding block flag are 0, The apparatus according to clause 14. 16. When processing a chroma block using a chroma scale factor, processing the residual of the chroma block using the chroma scale factor The apparatus according to any one of clauses 9 to 15, wherein the processor is configured to execute a set of instructions for causing the apparatus to further perform. A non-transitory computer-readable storage medium storing a set of instructions executable by one or more processors of a device for causing the device to perform a method for processing video content, the method comprising: Receiving chroma blocks and luma blocks associated with a picture; Determining luma scale information associated with the luma blocks; Determining a chroma scale factor based on the luma scale information; and Processing the chroma blocks using the chroma scale factor A non-transitory computer-readable storage medium comprising the above. 18. A computer-implemented method for processing video content, comprising: Receiving chroma blocks and luma blocks associated with a picture, the luma blocks including NxM luma samples; Selecting a subset of the NxM luma samples in response to at least one of N and M exceeding a threshold; Determining an average value of the subset of the NxM luma samples; Determining a chroma scale factor based on the average value; and Processing the chroma blocks using the chroma scale factor A method comprising the above. 19. A computer-implemented method for processing video content, comprising: Determining whether there is a chroma scale index in the received video data; Determining that no chroma scaling is applied to the received video data in response to determining that there is no chroma scale index in the received video data; and Determining a chroma scale factor based on the chroma scale index in response to determining that there is chroma scaling in the received video data A method comprising the above. 20. A computer-implemented method for processing video content, comprising: Receiving a plurality of unique code words used for the dynamic range of an input video signal, Receiving an index, Selecting at least one of the plurality of unique code words based on the index, and Determining a chroma scale factor based on the selected at least one code word A method comprising.
[0175]
[0192] In addition to implementing the above method by using computer-readable program code, the above method can also be implemented in the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be regarded as a hardware component, and a device configured to implement various functions included in the controller can also be regarded as a structure within the hardware component. Or a device configured to implement various functions can even be regarded as both a software module configured to implement the method and a structure within the hardware component.
[0176]
[0193] The present disclosure can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, assemblies, data structures, classes, etc. used to execute a specific task or implement a specific abstract data type. Embodiments of the present disclosure can also be implemented in a distributed computing environment. In a distributed computing environment, tasks are executed by using remote processing devices connected by a communication network. In a distributed computing environment, program modules can be in local and remote computer storage media including storage devices.
[0177]
[0194] Relational terms such as "first" and "second" in this specification are used solely to distinguish one entity or operation from another entity or operation, and it should be noted that no actual relationship or order between those entities or operations is required or implied. Further, words such as "comprising", "having", "containing", and "including" and other similar forms of words are intended to be equivalent in meaning, and it is not intended that the items following any one of these words be an exhaustive listing of such items, nor is it intended to be limited to only the items listed, and it is intended to be non-limiting.
[0178]
[0195] In the above specification, embodiments have been described with respect to numerous specific details that may vary for each implementation form. Certain adaptations and modifications of the described embodiments can be made. By studying this specification and practicing the disclosure revealed herein, other embodiments may become apparent to those skilled in the art. This specification and the examples are to be considered solely as examples, and it is intended that the true scope and spirit of the disclosure be indicated by the appended claims. The order of the steps shown in the figures is for illustrative purposes only and is not intended to be limited to the order of any particular steps. Therefore, those skilled in the art can understand that those steps can be executed in different orders while implementing the same method.
Claims
1. 1. A method for decoding a bitstream to output one or more pictures of a video stream, comprising the steps of: receiving a bitstream; and Decoding one or more pictures using the coded information of the bitstream. Including, The decoding step comprises: Reconstructing a number of luma samples associated with a picture; and Reconstructing chroma blocks associated with said picture. Including, reconstructing the chroma blocks determining whether the chroma block has a non-zero residual; in response to determining that the chroma block has one or more non-zero chroma residuals, determining an average value of the reconstructed luma samples and scaling the residuals of the chroma block based on the average value prior to reconstructing the chroma block; and and bypassing the determination of the average value of the reconstructed luma samples in response to determining that the chroma block does not have a non-zero chroma residual. Including, method.
2. scaling the residual of the chroma block based on the average value of the reconstructed luma samples; determining a chroma scale factor based on the average value of the reconstructed luma samples; and applying the chroma scale factor to the residual of the chroma block. The method of claim 1 , comprising:
3. The method of claim 1 , wherein whether the chroma block has a non-zero residual is determined based on a value of a coded block flag associated with the chroma block.
4. The decoding step comprises: determining that the chroma block has one or more non-zero residuals in response to the value of the coded block flag being equal to one; The method of claim 3 further comprising:
5. The decoding step comprises: determining that the chroma block does not have a non-zero residual in response to the value of the coded block flag being equal to zero; The method of claim 3 further comprising:
6. Whether the chroma block has a non-zero residual is determined by a value of a first coded block flag associated with a first chroma component of the chroma block; and a value of a second coded block flag associated with a second chroma component of the chroma block; The method of claim 1 , wherein the determination is based on:
7. The decoding step comprises: determining that the chroma block has one or more non-zero residuals in response to at least one of the value of the first coded block flag or the value of the second coded block flag being equal to one; The method of claim 6 further comprising:
8. The decoding step comprises: determining that the chroma block does not have a non-zero residual in response to the value of the first coded block flag and the value of the second coded block flag both being equal to zero; The method of claim 6 further comprising:
9. one or more memories storing a set of instructions; one or more processors; An apparatus comprising: the one or more processors: Reconstructing a number of luma samples associated with a picture; and Reconstructing chroma blocks associated with said picture. configured to execute the set of instructions to cause the device to When reconstructing the chroma blocks, determining whether the chroma block has a non-zero residual; in response to determining that the chroma block has one or more non-zero chroma residuals, determining an average value of the reconstructed luma samples and scaling the residuals of the chroma block based on the average value prior to reconstructing the chroma block; and and bypassing the determination of the average value of the reconstructed luma samples in response to determining that the chroma block does not have a non-zero chroma residual. the one or more processors are configured to execute the set of instructions to further cause the device to: device.
10. the one or more processors: determining a chroma scale factor based on the average value of the reconstructed luma samples; and applying the chroma scale factor to the residual of the chroma block.
10. The apparatus of claim 9, further configured to execute the set of instructions to cause the apparatus to further:
11. the one or more processors: determining whether the chroma block has a non-zero residual based on a value of a coded block flag associated with the chroma block; 10. The apparatus of claim 9, further configured to execute the set of instructions to cause the apparatus to further:
12. the one or more processors: determining that the chroma block has one or more non-zero residuals in response to the value of the coded block flag being equal to one; 12. The apparatus of claim 11, configured to execute the set of instructions to cause the apparatus to further:
13. the one or more processors: determining that the chroma block does not have a non-zero residual in response to the value of the coded block flag being equal to zero; 12. The apparatus of claim 11, configured to execute the set of instructions to cause the apparatus to further:
Citation Information
Patent Citations
Color Residual Prediction for Video Coding
JP2016541164A
Signaling in-loop reshaping information using parameter sets
JP2022517683A
Systems and methods for coding video data using adaptive component scaling
WO2018016381A1
Signaling of in-loop reshaping information using parameter sets
WO2020156529A1