Method for using adaptive loop filter and system therefor
By introducing an adaptive loop filter (ALF) into the video encoding process, the problem of insufficient encoding efficiency in existing technologies is solved, achieving more efficient video encoding and lower storage and transmission requirements, and improving the ability to reconstruct video quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2026-04-10
AI Technical Summary
Existing video coding technologies are insufficient in terms of coding efficiency and compression performance. In particular, during the development of high-efficiency video coding standards such as AVS3, it is difficult to further improve coding efficiency and reduce the need for storage space and transmission bandwidth.
An adaptive loop filter (ALF) is used to process video data. By receiving the bit stream, decoding the index to determine the maximum number of image components, and using the ALF to process the pixels in the image, the encoding efficiency and compression performance are improved.
It improves the coding efficiency of video encoding, enhances the parallel processing capabilities of encoders and decoders, reduces the requirements for storage space and transmission bandwidth, and improves the ability to reconstruct video quality.
Smart Images

Figure CN116601960B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of video processing, and in particular, to systems and methods using adaptive loop filter (ALF) in video encoding and decoding. BACKGROUND
[0002] A video is a set of still pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission, and decompressed before display. The compression process is commonly referred to as encoding, and the decompression process is commonly referred to as decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Video coding standards that specify particular video coding formats are developed by standardization organizations, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the Audio Video Coding Standard (AVS) standard. As more and more advanced video coding techniques are adopted by video standards, the coding efficiency of new video coding standards is also increasing. SUMMARY
[0003] Embodiments of the present disclosure provide a method for processing video data, the method comprising: receiving a bitstream; decoding a first index from the bitstream; determining a maximum number of adaptive loop filters (ALFs) for a component of a picture based on the first index; and processing pixels in the picture using the ALFs.
[0004] Embodiments of the present disclosure provide an apparatus for processing video data, the apparatus comprising: a memory configured to store instructions; and one or more processors configured to execute the instructions to cause the apparatus to perform: receiving a bitstream; decoding a first index from the bitstream; determining a maximum number of adaptive loop filters (ALFs) for a component of a picture based on the first index; and processing pixels in the picture using the ALFs.
[0005] Embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing a set of instructions that is executable by one or more processors of an apparatus to cause the apparatus to initiate a method for processing video data. The method comprises: receiving a bitstream; decoding a first index from the bitstream; determining a maximum number of adaptive loop filters (ALFs) for a component of a picture based on the first index; and processing pixels in the picture using the ALFs.
[0006] Embodiments of the present disclosure provide a non-transitory computer readable medium storing a bitstream, the bitstream comprising a first index associated with video data, the first index being encoded based on one or more of a plurality of contexts used in binary entropy coding and indicating a maximum number of adaptive loop filters (ALF) for a component of a picture. BRIEF DESCRIPTION OF DRAWINGS
[0007] Embodiments of the present disclosure and corresponding aspects are illustrated in the following detailed description and in the accompanying drawings. Various features shown in the figures are not to scale.
[0008] Figure 1 is a schematic diagram of a structure of an example video sequence shown in accordance with Embodiment 1 of the present disclosure;
[0009] Figure 2A is a schematic diagram of an example encoding process of a hybrid video coding system shown in accordance with Embodiment 1 of the present disclosure;
[0010] Figure 2B is a schematic diagram of another example encoding process of a hybrid video coding system shown in accordance with Embodiment 1 of the present disclosure;
[0011] Figure 3A is a schematic diagram of an example decoding process of a hybrid video coding system shown in accordance with Embodiment 1 of the present disclosure;
[0012] Figure 3B is a schematic diagram of another example decoding process of a hybrid video coding system shown in accordance with Embodiment 1 of the present disclosure;
[0013] Figure 4 is a block diagram of an example apparatus for encoding or decoding video shown in accordance with Embodiment 1 of the present disclosure;
[0014] Figure 5A is a schematic diagram of an example partitioning of a picture into 16 adaptive loop filter (ALF) regions shown in accordance with Embodiment 1 of the present disclosure;
[0015] Figure 5B is a schematic diagram of a region order and a region sequence number for each of the 16 ALF regions shown in accordance with Embodiment 1 of the present disclosure;
[0016] Figure 5C is a schematic diagram of an example merging region shown in accordance with Embodiment 1 of the present disclosure;
[0017] Figure 6A is a schematic diagram of a partitioning of a picture into more than 16 ALF regions shown in accordance with Embodiment 1 of the present disclosure;
[0018] Figure 6Bis a diagram illustrating partitioning of a picture into more than 16 ALF regions according to Embodiment 1 of the present disclosure;
[0019] Figure 6C is a diagram illustrating partitioning of a picture into more than 16 ALF regions according to Embodiment 1 of the present disclosure;
[0020] Figure 7A is a diagram illustrating an exemplary ALF region partitioning mode according to Embodiment 1 of the present disclosure;
[0021] Figure 7B is a diagram illustrating another exemplary ALF region partitioning mode according to Embodiment 1 of the present disclosure;
[0022] Figure 7C is a diagram illustrating another exemplary ALF region partitioning mode according to Embodiment 1 of the present disclosure;
[0023] Figure 7D is a diagram illustrating another exemplary ALF region partitioning mode according to Embodiment 1 of the present disclosure;
[0024] Figure 8 is a diagram illustrating a flowchart for determining a region index for each sample according to Embodiment 1 of the present disclosure;
[0025] Figure 9A is a diagram illustrating an exemplary ALF region partitioning mode according to Embodiment 1 of the present disclosure;
[0026] Figure 9B is a diagram illustrating another exemplary ALF region partitioning mode according to Embodiment 1 of the present disclosure;
[0027] Figure 9C is a diagram illustrating another exemplary ALF region partitioning mode according to Embodiment 1 of the present disclosure;
[0028] Figure 9D is a diagram illustrating another exemplary ALF region partitioning mode according to Embodiment 1 of the present disclosure;
[0029] Figure 10A is a diagram illustrating an exemplary ALF region order for 64 regions according to Embodiment 1 of the present disclosure;
[0030] Figure 10B is a diagram illustrating another exemplary ALF region order for 64 regions according to Embodiment 1 of the present disclosure;
[0031] Figure 10C is a diagram illustrating another exemplary ALF region order for 64 regions according to Embodiment 1 of the present disclosure;
[0032] Figure 10D is a diagram illustrating another exemplary ALF region order for 64 regions according to embodiment 1 of the disclosure;
[0033] Figure 11A is a diagram illustrating an exemplary ALF region order for 256 regions according to embodiment 1 of the disclosure;
[0034] Figure 11B is a diagram illustrating another exemplary ALF region order for 256 regions according to embodiment 1 of the disclosure;
[0035] Figure 11C is a diagram illustrating another exemplary ALF region order for 256 regions according to embodiment 1 of the disclosure;
[0036] Figure 11D is a diagram illustrating another exemplary ALF region order for 256 regions according to embodiment 1 of the disclosure;
[0037] Figure 12A is a diagram illustrating an exemplary ALF region order for 128 regions according to embodiment 1 of the disclosure;
[0038] Figure 12B is a diagram illustrating another exemplary ALF region order for 128 regions according to embodiment 1 of the disclosure;
[0039] Figure 12C is a diagram illustrating another exemplary ALF region order for 128 regions according to embodiment 1 of the disclosure;
[0040] Figure 12D is a diagram illustrating another exemplary ALF region order for 128 regions according to embodiment 1 of the disclosure. DETAILED DESCRIPTION
[0041] Reference will now be made in detail embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which the same numbers represent the same or similar elements unless the context clearly dictates otherwise. The implementations set forth in the following description of exemplary embodiments do not represent all implementations consistent with the disclosure. Instead, they are merely
[0042] The Audio Video Coding Standard (AVS) Working Group was established in China in 2002 and is currently developing the AVS video standard, namely the third generation AVS3 video standard. The predecessors of the AVS3 standard, AVS1 and AVS2, were released in China in 2006 and 2016, respectively. In December 2017, the AVS Working Group released a Call For Proposals (CFP) to officially start the development work of the third generation AVS standard, AVS3. In December 2018, the High Performance Model (HPM) was selected by the Working Group as the new reference software platform for the development of the AVS3 standard. The initial technology of the HPM inherited from the AVS2 standard, and on this basis, more and more advanced video coding techniques were used to improve the compression performance. For example, in 2019, the first stage of the AVS3 standard has been completed, and compared with its predecessor AVS2, the coding performance has been improved by more than 20%, and the second stage of the AVS3 standard is still being developed on the basis of the first stage of the AVS3 to further improve the coding efficiency.
[0043] The AVS3 standard is recently developed and continues to include more coding techniques that provide better compression performance. The AVS3 is based on the same hybrid video coding system used in modern video compression standards (e.g., AVS1, AVS2, H.265 / HEVC, H.264 / AVC (Advanced Video Coding), MPEG2, H.263, etc.).
[0044] A video is a set of still pictures (or “frames”) arranged in a time sequence for storing visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in a time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet, a video player, or any end-user terminal with display functionality) can be used to display these pictures in a time sequence. Furthermore, in some applications, a video capture device can transmit the captured video to a video playback device (e.g., a computer with a display) in real time, for example, for surveillance, conferencing, or live streaming.
[0045] To reduce the storage space and transmission bandwidth needed for such applications, the video can be compressed before storage and transmission, and decompressed before display. The compression and decompression can be achieved by software executed by a processor (e.g., a processor of a general purpose computer) or by specialized hardware. A module for compression is generally referred to as an “encoder”, and a module for decompression is generally referred to as a “decoder”. An encoder and a decoder can be collectively referred to as a “CODEC”. An encoder and a decoder can be implemented in any of various suitable hardware, software, or combination thereof. For example, a hardware implementation of an encoder and a decoder can include circuitry, such as one or more microprocessors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field Programmable Gate Arrays (FPGAs), discrete logic or any combination thereof. A software implementation of an encoder and a decoder can include program code, computer executable instructions, firmware or any suitable computer-implemented algorithm or process fixed in a computer readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x, AVS series, etc. In some applications, a CODEC can decompress a video from a first encoding standard, and recompress the decompressed video using a second encoding standard, in which case the CODEC can be referred to as a “transcoder”.
[0046] A video encoding process can identify and retain useful information that can be used to reconstruct a picture, and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be completely reconstructed, such an encoding process can be referred to as “lossy”, otherwise can be referred to as “lossless”. Most encoding processes are “lossy”, which is a tradeoff for reducing the required storage space and transmission bandwidth.
[0047] Useful information in a picture being encoded (referred to as the “current picture”) can include changes relative to a reference picture (e.g., a previously encoded and reconstructed picture), which can include changes in position, brightness, or color of pixels, with changes in position being the most important. Changes in position of a group of pixels representing an object can reflect motion of the object between the reference picture and the current picture.
[0048] Generally, a picture that is coded without reference to another picture (i.e., the picture is its own reference picture) is referred to as an "I picture"; a picture is referred to as a "P picture" if some or all blocks (e.g., blocks referring to portions of a video picture) in the picture are obtained using intra prediction or inter prediction (e.g., uni-prediction) with one reference picture; a picture is referred to as a "B picture" if at least one block in the picture is predicted using two reference pictures (e.g., bi-prediction).
[0049] Figure 1 is a schematic diagram of an example structure of a video sequence according to Embodiment 1 of the present disclosure, where the video sequence 100 can be a video live stream or a video that has been captured and archived, and can be a real-life video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real-life video with augmented reality effects), and can be input from a video capturing device (e.g., a camera), a video archive containing previously captured videos (e.g., video files stored in a storage device), or a video feed interface (e.g., a video broadcast transceiver) to receive the video from a video content provider.
[0050] As shown in Figure 1 , the video sequence 100 can include a series of pictures arranged in time along a time line, including pictures 102, 104, 106, and 108. The pictures 102-106 are consecutive, and there can be more pictures between the pictures 106 and 108. In Figure 1 , the picture 102 is an I picture, whose reference picture is the picture 102 itself; the picture 104 is a P picture, whose reference picture is the picture 102 as indicated by the arrow; the picture 106 is a B picture, whose reference pictures are the pictures 104 and 108 as indicated by the arrows. In some embodiments, a reference picture of a picture (e.g., the picture 104) can not be adjacent to the picture, for example, the reference picture of the picture 104 can be a picture before the picture 102. It should be noted that the reference pictures of the pictures 102-106 are merely examples, and the present disclosure does not limit the embodiments of the reference pictures to the examples shown in Figure 1 .
[0051] Generally, due to the high computational complexity of the encoding and decoding processes of a video, a video codec does not encode or decode an entire picture at once, but instead, the video codec can split the picture into basic segments and encode or decode the picture segment by segment. In the present disclosure, such a basic segment is referred to as a Basic Processing Unit (BPU), for example, Figure 1The structure 110 in FIG. 1 shows an example structure of a picture of the video sequence 100 (e.g., any of the pictures 102-108). In the structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are shown with dashed lines. In some embodiments, a basic processing unit can be referred to as a “macroblock” in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as a “coding tree unit” (CTU) in some other video coding standards (e.g., H.265 / HEVC, H.266 / VVC, AVS). A basic processing unit can have a variable size in pixels in a picture, e.g., 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, or any arbitrary shape and size. The size and shape of a basic processing unit can be selected for a picture based on a balance of coding efficiency and level of detail to be preserved in the basic processing unit.
[0052] A basic processing unit can be a logical unit that can include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer), e.g., a basic processing unit of a color picture can include a luma component (Y) representing non-color luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the size of the luma and chroma components can be the same as the size of the basic processing unit. In some video coding standards (e.g., H.265 / HEVC, H.266 / VVC, AVS), the luma and chroma components can be referred to as “coding tree blocks” (CTBs), and any operation performed on a basic processing unit can be repeated on each of its luma and chroma components.
[0053] A video coding process has multiple stages of operations, examples of which are shown in Figure 2A-2B and Figure 3A-3BAs shown. For each stage, the size of the basic processing unit can still be too large to be processed, and thus the basic processing unit can be further divided into segments, which are referred to as “basic processing sub-units” in this disclosure. In some embodiments, the basic processing sub-units can be referred to as “blocks” in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as “Coding Units” (CU) in some other video coding standards (e.g., H.265 / HEVC, H.266 / VVC, AVS). The size of the basic processing sub-units can be smaller than or equal to the size of the basic processing unit. Similar to the basic processing unit, the basic processing sub-unit is also a logical unit, and can include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in a computer memory (e.g., in a video frame buffer), and any operation performed on the basic processing sub-unit can be repeated for each of its luma and chroma components. It should be noted that the basic processing sub-units can be further divided into smaller units according to processing needs. It should also be noted that different division stages can use different schemes to divide the basic processing unit.
[0054] For example, at the mode decision stage (an example of which is shown in Figure 2B ), the encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, but if the basic processing unit is too large, the encoder can not be able to make the decision, in which case the encoder can split the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC, H.266 / VVC, AVS), and then decide the prediction mode for each basic processing sub-unit individually.
[0055] For another example, at the prediction stage (an example of which is shown in Figure 2A-2B ), the encoder can perform prediction operations on the basic processing sub-units (e.g., CUs). However, in some cases, the basic processing sub-units can still be too large for the encoder to perform the prediction operations, in which case the encoder can further split the basic processing sub-units into smaller segments (e.g., referred to as “prediction blocks” or “PBs” in H.265 / HEVC, H.266 / VVC, AVS), and then perform the prediction operations on the smaller segments.
[0056] For yet another example, at the transform stage (an example of which is shown in Figure 2A-2BAs shown in the diagram, the encoder can perform transform operations on the remaining basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large for the encoder to perform transform operations. In this case, the encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC, H.266 / VVC, and AVS), and then perform transform operations on these smaller segments. It should be noted that the partitioning scheme of the same basic processing subunits can be different in the prediction and transform phases. For example, in H.265 / HEVC, H.266 / VVC, or AVS, different sizes and numbers of prediction blocks and transform blocks can be configured for the same CU.
[0057] exist Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units with boundaries shown by dashed lines. Different partitioning schemes can be adopted to divide different basic processing units into multiple basic processing sub-units.
[0058] In some implementations, to provide parallel processing and error recovery capabilities for video encoding and decoding, an image can be divided into one or more regions for processing. This allows the encoder or decoder to process these regions independently, without relying on information from any other region of the image. In other words, each region of the image can be processed independently, meaning the codec can process different regions of the image in parallel, thereby improving encoding efficiency. Furthermore, when data in a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the corrupted or lost data, thus ensuring the codec's error recovery capability. In some video coding standards, an image can be divided into different types of regions. For example, H.265 / HEVC, H.266 / VVC, and AVS provide two types of regions: "slices" and "tiles." It should also be noted that different segmentation schemes can be used to segment different images, such as different images in video sequence 100, to obtain multiple regions.
[0059] For example, in Figure 1 In the diagram, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units, while regions 116 and 118 each include six basic processing units. It should be noted that... Figure 1 The basic processing unit, basic processing subunit, and area of structure 110 are merely examples, and this disclosure does not limit its embodiments.
[0060] Figure 2AThis is a schematic diagram of an exemplary encoding process of a hybrid video encoding system according to Embodiment 1 of this disclosure. For example, encoding process 200A can be performed by an encoder. Figure 2A As shown, the encoder can encode the video sequence 202 into a video bitstream 228 according to process 200A. Similar to... Figure 1 Video sequence 100 and video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to... Figure 1 In structure 110 of the video sequence 202, each original image can be divided by the encoder into basic processing units, basic processing sub-units, or processing regions. In some embodiments, the encoder can perform process 200A for each basic processing unit of the original image of the video sequence 202. For example, the encoder can perform process 200A iteratively, wherein the encoder can encode the basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for a region (e.g., region 114-118) of each original image of the video sequence 202.
[0061] exist Figure 2A In this process, the encoder can feed the basic processing unit (referred to as "raw BPU") of the original image of the video sequence 202 to the prediction stage 204 to generate prediction data 206 and prediction BPU 208. Then, the prediction BPU 208 is subtracted from the raw BPU to generate residual BPU 210. The residual BPU 210 is fed to the transform stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. Finally, the prediction data 206 and the quantized transform coefficients 216 are fed to the binary encoding stage 226 to generate the video bitstream 228. Figure 2A Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 in process 200A can be referred to as the "forward path." During process 200A, after quantization stage 214, the encoder can feed the quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate a reconstructed residual BPU 222. The reconstructed residual BPU 222 is then added to prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referred to as the "reconstruction path" to ensure that the encoder and decoder use the same reference data for prediction.
[0062] The encoder can iteratively execute process 200A to encode each original BPU of the original image (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original image (in the reconstruction path). After encoding all the original BPUs of the original image, the encoder can continue encoding the next image in the video sequence 202.
[0063] Referring to process 200A, the encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to any action in any manner of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or using data for input.
[0064] In prediction phase 204, during the current iteration, the encoder can receive the original BPU and prediction reference 224 and perform prediction operations to generate prediction data 206 and prediction BPU 208. Prediction reference 224 can be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction phase 204 is to reduce information redundancy by extracting prediction data 206, making the prediction data usable for reconstructing the original BPU into prediction BPU 208 based on prediction data 206 and prediction reference 224.
[0065] Ideally, the predicted BPU 208 should be identical to the original BPU. However, since the prediction and reconstruction operations do not always achieve the desired results, the predicted BPU 208 is typically slightly different from the original BPU. To record this difference, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210 to record the difference. For example, the encoder can subtract the pixel values (e.g., grayscale or RGB values) of the predicted BPU 208 from the corresponding pixel values of the original BPU. Each pixel of the residual BPU 210 can have a residual value as the result of this subtraction between the corresponding pixels of the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the space footprint can be reduced by compressing the original BPU.
[0066] To further compress the residual BPU 210, at a transform stage 212, the encoder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional “basis patterns” associated with “transform coefficients.” The basis patterns can have the same size (e.g., the size of the residual BPU 210), can represent the frequency of variation of the components of the residual BPU 210 (e.g., the frequency of luminance variation), and none of the basis patterns can be copied from any combination (e.g., linear combination) of any other basis patterns. Since the basis patterns are similar to the basis functions of a discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are similar to the coefficients associated with the basis functions, the variation of the residual BPU 210 can be decomposed into the frequency domain in a manner similar to the decomposition of a function by a discrete Fourier transform, to reduce its spatial redundancy.
[0067] Different transform algorithms can use different basis patterns. Various transform algorithms can be used at the transform stage 212, such as a discrete cosine transform, a discrete sine transform, and the like. Since the transform at the transform stage 212 is invertible, the encoder can recover the residual BPU 210 by an inverse operation of the transform, called “inverse transform.” For example, to recover the pixels of the residual BPU 210, the inverse transform can be multiplying the values of the corresponding pixels of the basis patterns by the corresponding associated coefficients and adding the products to produce a weighted sum. For a video coding standard, the encoder and the decoder can use the same transform algorithm (and thus the same basis patterns). Thus, the encoder can record only the transform coefficients, and the decoder can reconstruct the residual BPU 210 from these transform coefficients without receiving the basis patterns from the encoder. The transform coefficients can have fewer bits than the residual BPU 210 but can be used to reconstruct the residual BPU 210 without significant quality degradation. Thus, the residual BPU 210 is further compressed.
[0068] The encoder can further compress the transform coefficients at a quantization stage 214. During the transform, different basis patterns can represent different frequencies of variation (e.g., frequencies of luminance variation). Since the human eye is generally better at recognizing low-frequency variations, the encoder can ignore the information of high-frequency variations without causing significant quality degradation in the decoding. For example, at the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value called “quantization scale factor” and rounding the quotient to its nearest integer. After such an operation, some of the transform coefficients of high-frequency basis patterns can be converted to zero, and the transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore the zero-valued quantized transform coefficients 216, further compressing the transform coefficients by the encoder. The quantization process is also invertible, in which the quantized transform coefficients 216 can be reconstructed to the transform coefficients in an inverse operation of quantization called “inverse quantization.”
[0069] Because the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 can be lossy. Generally, the quantization stage 214 can cause the most information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 require. To obtain different levels of information loss, the encoder can use different quantization parameter values or any other parameter of the quantization process.
[0070] In the binarization stage 226, the encoder can encode the prediction data 206 and the quantized transform coefficients 216 using a binarization technique, such as entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can also encode other information in the binarization stage 226, such as the prediction modes used in the prediction stage 204, parameters of the prediction operations, the type of transform used in the transform stage 212, parameters of the quantization process (e.g., the quantization parameter), encoder control parameters (e.g., bit rate control parameters), and the like. The encoder can use the output data of the binarization stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 can be further packetized for network transmission.
[0071] Referring to the reconstruction path of the process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transform stage 220, the encoder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of the process 200A.
[0072] It should be noted that other variations of the process 200A can be used to encode the video sequence 202. In some embodiments, the stages of the process 200A can be performed by the encoder in a different order. In some embodiments, one or more stages of the process 200A can be combined into a single stage. In some embodiments, a single stage of the process 200A can be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 can be merged into a single stage. In some embodiments, the process 200A can include additional stages. In some embodiments, the process 200A can omit one or more stages of the process 200A. Figure 2A
[0073] Figure 2B is a schematic diagram of another exemplary encoding process of a hybrid video encoding system according to Embodiment 1 of the present disclosure, process 200B can be modified from process 200A, and can be used by an encoder conforming to a hybrid video encoding standard (e.g., H.26x series). Compared with process 200A, process 200B additionally includes a mode decision stage 230 in the forward path, and separates the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. In the reconstruction path of process 200B, a loop filter stage 232 and a buffer 234 are also included.
[0074] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use pixels from one or more already encoded neighboring BPU(s) in the same picture to predict a current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPU(s) to reduce the spatial redundancy inherent in pictures. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions from one or more already encoded pictures to predict a current BPU. That is, the prediction reference 224 in temporal prediction can include encoded pictures to reduce the temporal redundancy inherent in pictures.
[0075] Referring to process 200B, in the forward path, the encoder performs prediction operations at the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, at the spatial prediction stage 2042, the encoder can perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 224 can include one or more neighboring BPU(s) in the same picture that have been encoded (in the forward path) and reconstructed (in the reconstruction path). The encoder can generate a predicted BPU 208 by extrapolating the neighboring BPU(s). Extrapolation techniques can include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder can perform extrapolation at a pixel level, e.g., by extrapolating the value of a corresponding pixel of each pixel of the predicted BPU 208. The neighboring BPU(s) used for extrapolation can be positioned from various directions relative to the original BPU, e.g., in a vertical direction (e.g., at the top of the original BPU), in a horizontal direction (e.g., at the left of the original BPU), in a diagonal direction (e.g., at the lower left, lower right, upper left, or upper right of the original BPU), or in any direction defined in the video encoding standard being used. For intra prediction, the prediction data 206 can include, for example, the position (e.g., coordinates) of the neighboring BPU(s) used, the size of the neighboring BPU(s) used, the parameters of the extrapolation, the direction of the neighboring BPU(s) used relative to the original BPU, etc.
[0076] As another example, at the temporal prediction stage 2044, the encoder can perform inter prediction. For a current picture’s original BPU, the prediction reference 224 can include one or more pictures (referred to as “reference pictures”) that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures can be BPU-encoded and reconstructed by the BPU, e.g., the encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same picture are generated, the encoder can generate a reconstructed picture as a reference picture. The encoder can perform a “motion estimation” operation to search for a matching region in a range of the reference picture (referred to as a “search window”), where the search window’s location can be determined based on the location of the original BPU in the current picture, e.g., the search window can be centered at a location in the reference picture that has the same coordinates as the original BPU in the current picture, and can extend outward by a predetermined distance. When the encoder identifies (e.g., by using a pixel recursive algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder can determine such a region as a matching region, which can have a different size (e.g., smaller than, equal to, larger than, or have a different shape) than the original BPU. Because the reference picture and the current picture are temporally separated in the timeline (e.g., as shown in Figure 1 FIG. 1), the matching region can be considered to “move” to the location of the original BPU over time, and the encoder can record the direction and distance of such “movement” as a “motion vector.” When multiple reference pictures are used (e.g., as in Figure 1 FIG. 1), the encoder can search for matching regions and determine their associated motion vectors for each reference picture. In some embodiments, the encoder can assign weights to the pixel values of the matching regions of the respective matching reference pictures.
[0077] Motion estimation can be used to identify various types of motion, e.g., translation, rotation, scaling, etc. For inter prediction, the prediction data 206 can include, e.g., the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, the weights associated with the reference pictures, etc.
[0078] To generate the prediction BPU 208, the encoder can perform a “motion compensation” operation. Motion compensation can be used to reconstruct the prediction BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder can move the matching region of the reference picture according to the motion vector, where the encoder can predict the original BPU of the current picture. When multiple reference pictures are used (e.g., as in Figure 1In some embodiments, the encoder can move the matching region of the reference picture according to the corresponding motion vector and average pixel value of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching region of the corresponding matching reference picture, the encoder can add the weighted sum of the pixel values of the moved matching region.
[0079] In some embodiments, inter prediction can be uni-directional or bi-directional. Uni-directional inter prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, if the picture 104 in FIG. 1 is a uni-directional inter prediction picture, a reference picture (e.g., the picture 102) can be used before the picture 104. Figure 1 In some embodiments, inter prediction can be uni-directional or bi-directional. Uni-directional inter prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, if the picture 104 in FIG. 1 is a uni-directional inter prediction picture, a reference picture (e.g., the picture 102) can be used before the picture 104. Figure 1 In some embodiments, inter prediction can be uni-directional or bi-directional. Uni-directional inter prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, if the picture 104 in FIG. 1 is a uni-directional inter prediction picture, a reference picture (e.g., the picture 102) can be used before the picture 104.
[0080] Still referring to the forward path of the process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, at a mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of the process 200B. For example, the encoder can perform a rate-distortion optimization technique to select the prediction mode according to the bit rate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode to minimize a cost function value. According to the selected prediction mode, the encoder can generate the corresponding predicted BPU 208 and the prediction data 206.
[0081] In the reconstruction path of process 200B, if an intra prediction mode has been selected in the forward path, after generating a prediction reference 224 (e.g., a current BPU that has been encoded and reconstructed in the current picture), the encoder can feed the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of a next BPU of the current picture). The encoder can also feed the prediction reference 224 to the in-loop filter stage 232, where an in-loop filter is applied to the prediction reference 224 to reduce or eliminate distortion (e.g., blocking artifacts) introduced during encoding of the prediction reference 224. The encoder can also apply various in-loop filter techniques at the in-loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter (ALF), etc. The in-loop filtered reference picture can be stored in a buffer 234 (or “decoded picture buffer”) for later use (e.g., as an inter prediction reference picture for future pictures of the video sequence 202). The encoder can also store one or more reference pictures in the buffer 234 for use at the temporal prediction stage 2044. In some embodiments, the encoder can encode parameters of the in-loop filter (e.g., in-loop filter strength) along with the quantized transform coefficients 216, prediction data 206, and other information at the binary encoding stage 226.
[0082] Figure 3A is a schematic diagram of an exemplary decoding process of a hybrid video coding system according to Embodiment 1 of the present disclosure, where the process 300A can be a counterpart to the compression process 200A in Figure 2A decompression process of the compression process 200A in Figure 2A-2B some embodiments, the process 300A can also be similar to the reconstruction path of the process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to the process 300A that is very similar to the video sequence 202d. But the video stream 304 is not identical to the video sequence 202 again due to information loss in the compression and decompression processes (e.g., the quantization stage 214 in Figure 2A-2B
[0083] In Figure 3A In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can perform, at the binarization decoding stage 302, the inverse operations of the binarization techniques used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other lossless compression algorithm) to decode other information, such as prediction modes, parameters of the prediction operations, transform types, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), and the like. In some embodiments, if the video bitstream 228 is transmitted over a network in packets, the decoder can depacketize the video bitstream 228 before feeding it to the binarization decoding stage 302.
[0084] The decoder can iteratively perform the process 300A to decode each coded BPU of the encoded picture and generate the prediction reference 224 for the next coded BPU of the encoded picture. After decoding all coded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and continue decoding the next encoded picture in the video bitstream 228.
[0085] In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can perform, at the binarization decoding stage 302, the inverse operations of the binarization techniques used by the encoder (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other lossless compression algorithm) to decode other information, such as prediction modes, parameters of the prediction operations, transform types, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), and the like. In some embodiments, if the video bitstream 228 is transmitted over a network in packets, the decoder can depacketize the video bitstream 228 before feeding it to the binarization decoding stage 302.
[0086] Figure 3B is a schematic diagram of another exemplary decoding process of a hybrid video coding system according to Embodiment 1 of the present disclosure, where the process 300B can be modified from the process 300A and can be used by decoders conforming to hybrid video coding standards (e.g., H.26x series). Compared to the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.
[0087] In process 300B, the decoder decodes the encoded base processing unit (referred to as "current BPU") of the encoded picture being decoded (referred to as "current picture") from binary decoding stage 302 to obtain prediction data 206, which can include various types of data depending on what prediction mode the encoder used to encode the current BPU. For example, if the encoder used intra prediction to encode the current BPU, prediction data 206 can include an intra prediction prediction mode indicator (e.g., a flag value), parameters of the intra prediction operation, etc., where the parameters obtained by the intra prediction operation can include, for example, locations (e.g., coordinates) of one or more neighboring BPUs used as reference, sizes of the neighboring BPUs, extrapolation parameters, directions of the neighboring BPUs relative to the original BPU, etc.; if the encoder used inter prediction to encode the current BPU, prediction data 206 can include an inter prediction prediction mode indicator (e.g., a flag value), parameters of the inter prediction operation, etc., where the parameters obtained by the inter prediction operation can include, for example, a number of reference pictures associated with the current BPU, weights associated with the reference pictures, respectively, locations (e.g., coordinates) of one or more matching regions in the respective reference pictures, one or more motion vectors associated with the matching regions, respectively, etc.
[0088] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) at spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) at temporal prediction stage 2044. Details of performing such spatial prediction or temporal prediction can be found in the description of FIG. 2B, which is not repeated below. After performing such spatial prediction or temporal prediction, as shown in FIG. 2B, the decoder can generate a predicted BPU 208 and add predicted BPU 208 and reconstructed residual BPU 222 to generate a prediction reference 224. Figure 2B Figure 3A
[0089] In process 300B, the decoder can feed prediction reference 224 to spatial prediction stage 2042 or temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if intra prediction is used to decode the current BPU at spatial prediction stage 2042, after generating prediction reference 224 (e.g., the decoded current BPU), the decoder can feed prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture); if inter prediction is used to decode the current BPU at temporal prediction stage 2044, after generating prediction reference 224 (e.g., the reference picture with all BPUs decoded), the decoder can feed prediction reference 224 to loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can feed prediction reference 224 to spatial prediction stage 2042 or temporal prediction stage 2044 in any suitable manner, such as by using a buffer to store prediction reference 224 and then feeding prediction reference 224 to spatial prediction stage 2042 or temporal prediction stage 2044. Figure 2B The loop filter is applied to the prediction reference 224 in the manner described in the middle, the loop-filtered reference picture can be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., as an inter prediction reference picture for future encoded pictures of the video bitstream 228). The decoder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data can further include parameters of the loop filter (e.g., loop filter strength). In some embodiments, the prediction data can further include parameters of the loop filter when the prediction mode indicator of the prediction data 206 indicates that inter prediction is used for encoding the current BPU.
[0090] Figure 4 is a block diagram of an exemplary apparatus for encoding or decoding video according to Embodiment 1 of the present disclosure, as Figure 4 As shown in FIG. 4, the apparatus 400 can include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a special-purpose machine for video encoding or decoding. The processor 402 can be any type of circuitry capable of manipulating or processing information. For example, the processor 402 can include any combination of central processing units (“CPUs”), graphics processing units (“GPUs”), neural processing units (“NPUs”), microcontroller units (“MCUs”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, Intellectual Property (“IP”) cores, programmable logic arrays (“PLAs”), programmable array logic (“PAL”), generic array logic (“GAL”), complex programmable logic devices (“CPLDs”), field-programmable gate arrays (“FPGAs”), systems on a chip (“SoCs”), application-specific integrated circuits (“ASICs”), and the like. In some embodiments, the processor 402 can also be a group of processors grouped as a single logical component. For example, as shown in FIG. 4, the processor 402 can include multiple processors, including a processor 402a, a processor 402b, and a processor 402n. Figure 4 As shown in FIG. 4, the processor 402 can include multiple processors, including a processor 402a, a processor 402b, and a processor 402n.
[0091] The device 400 may also include a memory 404 configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, such as Figure 4 As shown, the stored data may include program instructions (e.g., program instructions for implementing stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any combination of any number of random-access memory (RAM), read-only memory (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, Security Digital (SD) cards, Memory Sticks, Compact Flash (CF) cards, etc. Memory 404 may also be a group of memories grouped into single logical components. Figure 4 (Not shown in the image).
[0092] Bus 410 may be a communication device for transmitting data between components within device 400, such as an internal bus (e.g., CPU memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Fast Port), etc.
[0093] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry can be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry can be a single, independent module, or it can be wholly or partially integrated into any other component of the device 400.
[0094] Device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near-field communication (NFC) adapters, cellular network chips, etc.
[0095] In some embodiments, optionally, the apparatus 400 can further include a peripheral interface 408 to provide a connection to one or more peripheral devices. As shown, the peripheral devices can include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light-emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), and the like. Figure 4
[0096] It should be noted that a video codec (e.g., a codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any number and arrangement of software or hardware modules in apparatus 400, e.g., some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of apparatus 400, e.g., program instructions that can be loaded into memory 404. For another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of apparatus 400, e.g., a special-purpose data processing circuit (e.g., an FPGA, an ASIC, an NPU, or the like).
[0097] An adaptive loop filter (ALF) is an in-loop filter (e.g., applied by Figure 2B and 3B loop filter 232 of H.266 / VVC) that is applied to reconstructed samples to further refine the reconstructed samples and obtain the final decoded samples output by the decoder. The purpose of ALF is to improve the quality of the reconstructed samples and reduce coding errors.
[0098] ALF introduces a Wiener filter in the coding loop to minimize the mean square error between the original and decoded samples. The value of a filtered sample is generated by a weighted average of the value of the current unfiltered sample and the values of unfiltered spatial neighboring samples. The weights, referred to as the coefficients of the filter, are determined by the encoder and explicitly signaled to the decoder.
[0099] Figure 5A is a schematic diagram illustrating exemplary partitioning of a picture into 16 adaptive loop filter (ALF) regions according to Embodiment 1 of the present disclosure, as Figure 5A As shown, for example, in AVS3, in order to adapt the filter to the content, the whole picture is divided into 16 regions (4 columns x 4 rows), and each region has an index that can identify the region. Among the 16 regions, different regions can have different filters (e.g., different sets of coefficients). Therefore, for one picture, at most 16 filters (e.g., 16 sets of coefficients) can be signaled, i.e., the sets of coefficients can be signaled to the filters one set after another according to the order of the regions. Since one filter is generally applied to one region, the ALF index of each region can be easily determined.
[0100] Figure 5B is a diagram illustrating the region order and region sequence number of each region of the 16 ALF regions according to Embodiment 1 of the present disclosure, in which the curve shows the order of the regions as Figure 5B As shown, the 16 regions are numbered from 0 to 15 with region sequence numbers and are traversed along the curve.
[0101] By dividing the picture into multiple regions and deriving the coefficients of each region respectively, the effect of filter positioning can be achieved, but the cost of doing so increases with the number of coefficients that need to be signaled. Considering that for some smooth content, different regions in one picture can share similar content features, it is not necessary to derive different filters for each different region, therefore, in order to save the signaling cost of these video sequences, i.e., the cost of signaling, a region merging mechanism can be used to merge any two connected regions (e.g., regions with consecutive sequence numbers) in the order of the regions into one region with a set of filter coefficients, and the merged region can be further merged with its connected regions. For example, region 4 and region 5 can be merged into region 45, and region 45 can be continuously merged with region 3 or region 6 until the whole picture is one region. The encoder can decide how to merge the 16 regions and signal the merged regions to the decoder. Therefore, as shown above, the number of filters (e.g., the number of sets of coefficients) for each picture can change from picture to picture according to the merging result that can be decided by the encoder based on the content of the current picture.
[0102] Figure 5C is a diagram illustrating an exemplary merged region according to Embodiment 1 of the present disclosure, in which Figure 5CAs shown, region 0 can be merged region A; regions 1 to 3 can be merged into merged region B; regions 4 to 8 can be merged into merged region C, regions 9 to 12 can be merged into merged region D, and regions 13 to 15 can be merged into merged region E. Thus, there are 5 merged regions in total in one picture, and 5 sets of signals corresponding to the coefficients of the merged regions can be sent for the picture, instead of 16 sets, thereby saving signaling overhead. At the decoder side, by decoding the merged region information, the ALF index of the region can be determined, and since one filter can be applied to one merged region, the decoder can utilize the ALF to process the current region based on the ALF index. In some embodiments, during the ALF processing, one region can be divided into one or more ALF units, and the ALF can process the pixels one unit by one unit. Since the ALF units are subsets of the region, all the pixels in one ALF unit are processed by one ALF (i.e., share one set of ALF coefficients), and thus the pixels in one ALF unit have the same ALF index.
[0103] However, in the current ALF design, there are several problems.
[0104] On the one hand, in the current ALF design, one picture can be divided into at most 16 regions. But for high-resolution pictures, such as 2K and 4K, there are a lot of contents included, and when a picture is divided into 16 regions, the regions can include various contents that can require different filters, and one region can only have one filter, and applying one filter to all these samples can reduce the coding performance of the ALF, and thus dividing one picture into 16 regions can not be enough.
[0105] On the other hand, in the current ALF design, the division mode when the picture is divided into 16 regions is fixed, but since the video content of each picture is different, in some cases, different contents must be included in the same region that affects the coding performance of the ALF, and thus the fixed division mode cannot be suitable for all pictures.
[0106] On the other hand, in the current design, only one fixed order is defined for the 16 regions, and only the connected regions in the sequence can be merged, but since some spaces are not connected in the defined order, even if these spaces are adjacent, they cannot be merged. For example, as shown in Figure 5C For example, as shown, taking two spatially adjacent regions (e.g., region 2 and region 11) as an example, since they are not in the connected sequence, even if the two regions share similar contents, they cannot be merged. So this leads to the fact that the encoder has to signal two similar filters for the two regions, which increases the signaling overhead.
[0107] The present disclosure provides methods and systems for addressing some or all of the above problems.
[0108] According to some example embodiments, since a picture can be divided into more than 16 regions, in some embodiments, the picture can be divided with the same number of region columns and region rows. For example, Figure 6A is a diagram illustrating division of a picture into more than 16 ALF regions according to Embodiment 1 of the present disclosure, Figure 6B is a diagram illustrating division of a picture into more than 16 ALF regions according to Embodiment 1 of the present disclosure. As Figure 6A illustrated, the picture is divided into 8 region columns and 8 region rows to obtain 64 regions; as Figure 6B illustrated, the picture is divided into 16 region columns and 16 region rows to obtain 256 regions.
[0109] In some embodiments, the picture can also be divided into different number of region columns and region rows. Figure 6C is a diagram illustrating division of a picture into more than 16 ALF regions according to Embodiment 1 of the present disclosure, as Figure 6C illustrated, the picture is divided into 16 region columns and 8 region rows to obtain 128 regions, where the number of region columns (i.e., 16) is different from the number of region rows (i.e., 8) because generally the picture width is greater than the picture height.
[0110] While increasing the number of regions is beneficial for larger pictures, it can not be beneficial for smaller pictures because it increases the signaling burden. Therefore, in some embodiments, the number of regions can be made variable and the number can depend on the resolution of the picture, e.g., a large picture can have more ALF regions than a small picture.
[0111] In some embodiments, the number of regions can be determined by different scales or one or more thresholds. For example, the width of the picture is denoted as W and the height of the picture is denoted as H, if WxH is less than or equal to a first threshold TH1 (e.g., 1920x1080), the picture is divided into a first number of regions (e.g., 64 regions as Figure 6A illustrated); if WxH is greater than TH1 but less than a second threshold TH2 (e.g., 3840x2160), the picture is divided into a second number of regions (e.g., 128 regions as Figure 6C illustrated); if WxH is greater than or equal to TH2, the picture is divided into a third number of regions (e.g., 256 regions as Figure 6B illustrated). In some embodiments, the number of regions is proportional to the threshold. It can be appreciated that there is no limit to the number of thresholds and different number of regions can correspond to different threshold intervals.
[0112] In some embodiments, the number of regions can be determined based on width and height, respectively. If W is less than a width threshold TW, the picture is divided into 8 region columns; if W is greater than or equal to TW, the picture is divided into 16 region columns; if H is less than a height threshold TH, the picture is divided into 8 region rows; if H is greater than or equal to TH, the picture is divided into 16 region rows. It can be appreciated that there can be more than one width threshold and more than one height threshold. The number of region columns and the number of region rows correspond to different width threshold intervals and height threshold intervals, respectively.
[0113] In some embodiments, the number of regions can be selected from a set of preset region numbers according to the size of the picture. For example, a basic region (e.g., 16x16 = 256 pixels) is defined as Ba, and the number of basic regions in a picture N is calculated as N = WxH / Ba, where N is rounded to a set of preset region numbers. As shown in Figure 6A-6C the set of preset region numbers includes 64, 128, and 256, so N can be rounded to 64, 128, and 256.
[0114] In some embodiments, the number of regions can be selected from a set of preset column numbers and a set of preset row numbers based on width and height, respectively. The width and height define two basic lengths Lw and Lh. The number of basic lengths in the picture width is calculated as Nw = W / Lw, and the number of basic lengths in the picture height is calculated as Nh = H / Lh, where Nw is rounded to a set of preset column numbers and Nh is rounded to a set of preset row numbers. As shown in Figure 6A-6C the set of preset column numbers includes 8 and 16, so Nw is rounded to 8 or 16. The set of preset row numbers is 8 and 16, so Nh is rounded to 8 or 16. The encoder / decoder supports the preset column numbers and preset row numbers of the picture.
[0115] In some embodiments, to simplify the filtering operation, the regions can be aligned with the largest coding unit (LCU). That is, the boundary of each region should be a LCU boundary or a picture boundary, so that all pixels within one LCU have the same filter.
[0116] In some embodiments, variable number of regions is supported in a video sequence. That is, in a video sequence, to make the number of regions adaptive to the pictures so as to improve the efficiency of video processing, the number of regions can be different for different pictures. In some embodiments, one region has at most one ALF, but different regions can share one ALF, thus, the maximum number of ALFs can be equal to the number of regions in a picture. In some embodiments, ALF can be performed on different components (e.g., chroma component and luma component) contained in a picture, thus, the encoder can decide how many regions there are in one picture at picture level or sequence level, and signal the index of the maximum number of ALFs for the components of the picture in the bitstream. After receiving the bitstream, the decoder can determine the maximum number of ALFs for the current picture or the current video sequence according to the index parsed from the bitstream.
[0117] According to some exemplary embodiments, the present disclosure also provides different partition modes, Figure 7A-7D Examples of different partition modes according to some embodiments consistent with the present disclosure are shown.
[0118] Figure 7A is a schematic diagram of an exemplary ALF region partition mode shown according to Embodiment 1 of the present disclosure, as shown in FIG. 7, a picture can be partitioned into multiple region columns and region rows, if the picture width is not a multiple of the number of region columns, all region columns except the last column 701A can have the same width; similarly, if the picture height is not a multiple of the number of region rows, all region rows except the last row 702A have the same height. Figure 7B is a schematic diagram of another exemplary ALF region partition mode shown according to Embodiment 1 of the present disclosure; Figure 7C is a schematic diagram of another exemplary ALF region partition mode shown according to Embodiment 1 of the present disclosure; Figure 7D is a schematic diagram of another exemplary ALF region partition mode shown according to Embodiment 1 of the present disclosure Figure 7B-7D Examples of exemplary ALF region partition modes according to some embodiments of the present disclosure are shown, as Figure 7B-7D shown, if the picture width is not a multiple of the number of region columns, the first column or the last column can have a different width from other columns; if the picture height is not a multiple of the number of region rows, the height of the first row or the last row can be different from other rows. If the picture width and the picture height are not multiples of the number of region columns and the number of region rows respectively, there can be four different ALF region partition modes, for example, as shown in Figure 7A-7D For example, in Figure 7B , the first column 701B has a different width from other columns, and the last row 702B has a different height from other rows; in Figure 7CIn an embodiment, the first column 701C has a width different from other columns, and the first row 702C has a height different from other rows; in Figure 7D In an embodiment, the last column 701D has a width different from other columns, and the first row 702D has a height different from other rows.
[0119] If the picture width in units of LCU width is 30, the picture height in units of LCU height is 17, and if the picture is divided into 8 area columns and 8 area rows, for the division mode in Figure 7A each area column has a width in units of LCU width of 4, 4, 4, 4, 4, 4, 4, 2, and each area row has a height in units of LCU height of 2, 2, 2, 2, 2, 2, 2, 3; for the division mode in Figure 7B each area column has a width in units of LCU width of 2, 4, 4, 4, 4, 4, 4, 4, and each area row has a height in units of LCU height of 2, 2, 2, 2, 2, 2, 2, 3; for the division mode in Figure 7C each area column has a width in units of LCU width of 2, 4, 4, 4, 4, 4, 4, 4, and each area row has a height in units of LCU height of 3, 2, 2, 2, 2, 2, 2, 2; for the division mode in Figure 7D each area column has a width in units of LCU width of 4, 4, 4, 4, 4, 4, 4, 2, and each area row has a height in units of LCU height of 3, 2, 2, 2, 2, 2, 2, 2.
[0120] Figure 8 is a schematic diagram of a flowchart for determining an area index of each sample according to Embodiment 1 of the present disclosure, as shown in Figure 8 Method 800 can be performed by an encoder (for example, by process 200A of apparatus 400 or process 200B of apparatus 500), a decoder (for example, by process 300A of apparatus 400 or process 300B of apparatus 600), or one or more software or hardware components of an apparatus (for example, apparatus 400), as shown in Figure 2A Figure 2B Figure 3A Figure 3B Method 800 can be performed by an encoder (for example, by process 200A of apparatus 400 or process 200B of apparatus 500), a decoder (for example, by process 300A of apparatus 400 or process 300B of apparatus 600), or one or more software or hardware components of an apparatus (for example, apparatus 400), as shown in Figure 4 Figure 4 Method 800 can be performed by an encoder (for example, by process 200A of apparatus 400 or process 200B of apparatus 500), a decoder (for example, by process 300A of apparatus 400 or process 300B of apparatus 600), or one or more software or hardware components of an apparatus (for example, apparatus 400), as shown in Figure 4 Figure 8 Method 800 can include the following steps 801-805.
[0121] At step 801, a first column width and a first row height are determined. Determining the first column width further comprises: calculating a number of LCUs in a picture width by rounding up with one LCU; obtaining a column width in LCU width by dividing the number of LCUs in the picture by the number of region rows; and obtaining the first column width in pixels by multiplying the column in LCU width by the LCU width. The first row height can be determined in a similar way. In some embodiments, the first column width x_interval and the first row height y_interval can be determined as follows:
[0122] x_interval = ((( (img_width + lcu_width - 1) / lcu_width) + RE_OFFSET_X) / INTERVAL_X * lcu_width);
[0123] y_interval = ((( (img_height + lcu_height - 1) / lcu_height) + RE_OFFSET_Y) / INTERVAL_Y * lcu_height),
[0124] where img_width is the width of the picture, img_height is the height of the picture, lcu_width and lcu_height are the width and height of an LCU. INTERVAL_X is the number of region columns, and INTERVAL_Y is the number of region rows. RE_OFFSET_X and RE_OFFSET_Y are two predefined offsets.
[0125] At step 803, a second column width and a second row height are determined. Determining the second column width further comprises: calculating a first column number based on the picture and the first column width, and clipping the first column number to a region column number; calculating a second column width based on the picture width, the first column width, and the first column number; and aligning the second column width with the LCU width. The second row height can be determined in a similar way. In some embodiments, the second column width x_st_offset and the second row height y_st_offset can be determined as follows:
[0126] if (y_interval == 0)
[0127] {
[0128] y_st_offset = 0;
[0129] }
[0130] else
[0131] {
[0132] y_cnt = Clip3(0, INTERVAL_Y, (img_height + y_interval - 1) / y_interval);
[0133] y_st_offset = img_height - y_interval * (y_cnt - 1);
[0134] y_st_offset = (y_st_offset + lcu_height / 2) / lcu_height * lcu_height;
[0135] }
[0136] if (x_interval == 0)
[0137] {
[0138] x_st_offset = 0;
[0139] }
[0140] else
[0141] {
[0142] x_cnt = Clip3(0, INTERVAL_X, (img_width + x_interval - 1) / x_interval);
[0143] x_st_offset = img_width - x_interval * (x_cnt - 1);
[0144] x_st_offset = (x_st_offset + lcu_width / 2) / lcu_width * lcu_width;
[0145] }
[0146] wherein Clip3(x, y, t) is a function that clips t to y if t is greater than y, and clips t to x if t is less than x.
[0147] At step 805, the region index of the sample is determined for each sample (e.g., with the coordinator (g, i)). The coordinate (g, i) indicates the position of the sample in the picture, and the region index of the sample is the index of the region where the sample is located. For example, the coordinate (0, 0) refers to the top-left sample of the whole picture. In this embodiment, the region index is assigned in a raster scan order, e.g., from left to right, from top to bottom, starting from 0. The region index of the sample can be obtained by calculating the relationship between the coordinate (g, i) of the sample and the position of each region. Because there can be 4 different partition modes (e.g., as shown in Figure 7A-7D
[0148] y_idx = (y_interval == 0)? (INTERVAL_Y - 1) : (Clip3(0, INTERVAL_Y - 1, i / y_interval));
[0149] y_idx_offset = y_idx * INTERVAL_X;
[0150] y_index2 = (y_interval == 0 || i < y_st_offset)? (0) : (Clip3(-1, INTERVAL_Y - 2, (i - y_st_offset) / y_interval) + 1);
[0151] y_index_offset2 = y_index2 * INTERVAL_X;
[0152] x_index = (x_interval == 0)? (INTERVAL_X - 1) : (Clip3(0, INTERVAL_X - 1, g / x_interval));
[0153] x_index2 = (x_interval == 0 || g < x_st_offset)? (0) : (Clip3(-1, INTERVAL_X - 2, (g - x_st_offset) / x_interval) + 1);
[0154] 1) for the partition mode in (g, i): the region index of the sample (g, i) is y_index_offset + x_index; Figure 7A 2) for the partition mode in (g, i): the region index of the sample (g, i) is y_index_offset2 + x_index2;
[0155] Figure 7B the partition mode in the picture: the region index of sample (g,i) is y_index_offset+x_index2;
[0156] 3) for the partition mode in the picture: Figure 7C the partition mode in the picture: the region index of sample (g,i) is y_index_offset2+x_index2;
[0157] 4) for the partition mode in the picture: Figure 7D the partition mode in the picture: the region index of sample (g,i) is y_index_offset2+x_index.
[0158] Figure 9A-9D Another example of different partition modes according to some embodiments consistent with the present disclosure is shown, wherein, Figure 9A is a schematic diagram of an exemplary ALF region partition mode shown according to embodiment 1 of the present disclosure; Figure 9B is a schematic diagram of another exemplary ALF region partition mode shown according to embodiment 1 of the present disclosure; Figure 9C is a schematic diagram of another exemplary ALF region partition mode shown according to embodiment 1 of the present disclosure; Figure 9D is a schematic diagram of another exemplary ALF region partition mode shown according to embodiment 1 of the present disclosure. In this example, the picture can be partitioned into multiple region columns and region rows. If the picture width W is not a multiple of the number of region columns CN, then the width of a region column is W / CN or W / CN+1; if the picture height H is not a multiple of the number of region rows RN, then the height of a region row is H / RN or H / RN+1, where " / " denotes integer division; if the picture width in LCU width units is 19 and the picture height in LCU height units is 11, and if the picture is partitioned into 8 region columns and 8 region rows, then for the partition mode in the picture: Figure 9A the width of each region column in LCU width units is 2, 3, 2, 3, 2, 3, 2, 2, and the height of each region row in LCU height units is 1, 2, 1, 2, 1, 2, 1, 1; for the partition mode in the picture: Figure 9B the width of each region column in LCU width units is 3, 2, 3, 2, 3, 2, 2, 2, and the height of each region row in LCU height units is 1, 2, 1, 2, 1, 2, 1, 1; for the partition mode in the picture: Figure 9C the width of each region column in LCU width units is 3, 2, 3, 2, 3, 2, 2, 2, and the height of each region row in LCU height units is 2, 1, 2, 1, 2, 1, 1, 1; for the partition mode in the picture: Figure 9DIn the partition mode shown in FIG. 1, the width of each region column in LCU width is 2, 3, 2, 3, 2, 3, 2, 2, and the height of each region row in LCU height is 2, 1, 2, 1, 2, 1, 1, 1. It can be understood that columns and rows with different widths and heights can be other possible arrangements.
[0159] In some embodiments, the video sequence supports variable partition modes (e.g., the partition modes shown in FIG. 1). The encoder decides which partition mode to use at picture level or sequence level, and signals an index of the partition mode used in the bitstream. The decoder determines the partition mode of the current picture or the current sequence according to the partition mode index parsed from the bitstream. For example, an index with a value between 0 and 3 corresponds to one of the partition modes shown in FIG. 1. Figure 9A-9D Figure 9A-9D
[0160] According to some example embodiments, different region orders are proposed. In some embodiments, the Hilbert-like curve can be extended or reversed to be used as the region order, for example, Figure 10A is a schematic diagram of an example ALF region order of 64 regions according to Embodiment 1 of the present disclosure; Figure 10B is a schematic diagram of another example ALF region order of 64 regions according to Embodiment 1 of the present disclosure; Figure 10C is a schematic diagram of another example ALF region order of 64 regions according to Embodiment 1 of the present disclosure; Figure 10D is a schematic diagram of another example ALF region order of 64 regions according to Embodiment 1 of the present disclosure. As shown in FIG. 1, the picture is partitioned into 8 region columns and 8 region rows, and the curves in the figure indicate the region order. Figure 10A-10D Figure 10A-10D Four different orders are shown in FIG. 1 respectively, and the numbers in the figure represent the region serial numbers. These orders can be reversed, flipped or rotated. If the serial number i is replaced by 64-i, the order is reversed.
[0161] If the picture is scanned in raster order, the region serial number of each region in FIG. 1 is: Figure 10A
[0162] {0, 1, 14, 15, 16, 19, 20, 21, 3, 2, 13, 12, 17, 18, 23, 22, 4, 7, 8, 11, 30, 29, 24, 25, 5, 6, 9, 10, 31, 28, 27, 26, 58, 57, 54, 53, 32, 35, 36, 37, 59, 56, 55, 52, 33, 34, 39, 38, 60, 61, 50, 51, 46, 45, 40, 41, 63, 62, 49, 48, 47, 44, 43, 42},
[0163] or
[0164] {63, 62, 49, 48, 47, 44, 43, 42, 60, 61, 50, 51, 46, 45, 40, 41, 59, 56, 55, 52, 53, 33, 34, 49, 38, 58, 57, 54, 53, 32, 35, 36, 37, 5, 6, 9, 10, 31, 28, 27, 26, 4, 7, 8, 11, 30, 29, 24, 25, 3, 2, 13, 12, 17, 18, 23, 22, 0, 1, 14, 15, 16, 19, 20, 21}.
[0165] If the picture is scanned in the raster order, the region sequence number of each region in Figure 10B is:
[0166] {63, 60, 59, 58, 5, 4, 3, 0, 62, 61, 56, 57, 6, 7, 2, 1, 49, 50, 55, 54, 9, 8, 13, 14, 48, 51, 52, 53, 10, 11, 12, 15, 47, 46, 33, 32, 31, 30, 17, 16, 44, 45, 34, 35, 28, 29, 18, 19, 43, 40, 39, 36, 27, 24, 23, 20, 42, 41, 38, 37, 26, 25, 22, 21},
[0167] or
[0168] {0, 3, 4, 5, 58, 59, 60, 63, 1, 2, 7, 6, 57, 56, 61, 62, 14, 13, 8, 9, 54, 55, 50, 49, 15, 12, 11, 10, 53, 52, 51, 48, 16, 17, 30, 31, 32, 33, 46, 47, 19, 18, 29, 28, 35, 34, 45, 44, 20, 23, 24, 27, 36, 39, 40, 43, 21, 22, 25, 26, 27, 38, 41, 42}.
[0169] If the picture is scanned in the raster order, the region sequence number of each region in Figure 10C is:
[0170] {42,43,44,47,48,49,62,63,41,40,45,46,51,50,61,60,38,39,34,33,52,55,56,59,37,36,35,32,53,54,57,58,26,27,28,31,10,9,6,5,25,24,29,30,11,8,7,4,22,23,18,17,12,13,2,3,21,20,19,16,15,14,1,0},
[0171] or
[0172] {21,20,19,16,15,14,1,0,22,23,18,17,12,13,2,3,25,24,29,30,11,8,7,4,26,27,28,31,10,9,6,5,37,36,35,32,53,54,57,58,38,39,34,33,52,55,56,59,41,40,45,46,51,50,61,60,42,43,41,47,48,49,62,63}.
[0173] If the picture is scanned in raster order, then Figure 10D the region sequence number of each region in
[0174] {21,22,25,26,37,38,41,42,20,23,24,27,36,39,40,43,19,18,29,28,35,34,45,44,16,17,30,31,32,33,46,47,15,12,11,10,53,52,51,48,14,13,8,9,54,55,50,49,1,2,7,6,57,56,61,62,0,3,4,5,58,59,60,63},
[0175] or
[0176] {42,41,38,37,26,25,22,21,43,40,39,36,27,24,23,20,44,45,34,35,28,29,18,19,47,46,33,32,31,30,17,16,48,51,52,53,10,11,12,15,49,50,55,54,9,8,13,14,62,61,56,57,6,7,2,1,63,60,59,58,5,4,3,0}.
[0177] In some embodiments, a picture can be divided into 16 columns and 16 rows of regions, and the curves in the figure indicate the region order. Figure 11A-11D Four different region orders are shown, where, Figure 11A is a diagram illustrating exemplary ALF region orders for 256 regions according to embodiment 1 of the disclosure; Figure 11B is a diagram illustrating another exemplary ALF region orders for 256 regions according to embodiment 1 of the disclosure; Figure 11C is a diagram illustrating another exemplary ALF region orders for 256 regions according to embodiment 1 of the disclosure; Figure 11D is a diagram illustrating another exemplary ALF region orders for 256 regions according to embodiment 1 of the disclosure. These orders can be reversed, flipped, or rotated.
[0178] If the picture is scanned in raster order, then Figure 11A The region sequence number for each region in
[0179] {0, 3, 4, 5, 58, 59, 60, 63, 64, 65, 78, 79, 80, 83, 84, 85,
[0180] 1, 2, 7, 6, 57, 56, 61, 62, 67, 66, 77, 76, 81, 82, 87, 86,
[0181] 14, 13, 8, 9, 54, 55, 50, 49, 68, 71, 72, 75, 94, 93, 88, 89,
[0182] 15, 12, 11, 10, 53, 52, 51, 48, 69, 70, 73, 74, 95, 92, 91, 90,
[0183] 16, 17, 30, 31, 32, 33, 46, 47, 122, 121, 118, 117, 96, 99, 100, 91, 19, 18, 29, 28, 35, 34, 45, 44, 123, 120, 119, 116, 97, 98, 103, 102, 20, 23, 24, 27, 36, 39, 40, 43, 124, 125, 114, 115, 110, 109, 104, 105, 21, 22, 25, 26, 27, 38, 41, 42, 127, 126, 113, 112, 111, 108, 107, 106, 234, 233, 230, 219, 218, 217, 214, 213, 128, 129, 142, 143, 144, 147, 148, 149, 235, 232, 231, 228, 219, 216, 215, 212, 131, 130, 141, 140, 145, 146, 151, 150, 236, 237, 226, 227, 220, 221, 210, 211, 132, 135, 136, 139, 158, 157, 152, 153, 239, 238, 225, 224, 223, 222, 209, 208, 133, 134, 137, 138, 159, 156, 155, 154, 240, 243, 244, 245, 202, 203, 204, 207, 186, 185, 182, 181, 160, 163, 164, 155, 241, 242, 247, 246, 201, 200, 205, 206, 187, 184, 183, 180, 161, 162, 167, 166, 254, 253, 248, 249, 198, 199, 194, 193, 188, 189, 178, 179, 174, 173, 168, 169, 255, 252, 251, 250, 197, 196, 195, 192, 191, 190, 177, 176, 175, 172, 171, 170.
[0184] If the picture is scanned in raster order, then Figure 11B The region sequence number for each region in
[0185] {0, 1, 14, 15, 16, 19, 20, 21, 234, 235, 236, 239, 240, 241, 254, 255, 3, 2, 13, 12, 17, 18, 23, 22, 233, 232, 237, 238, 243, 242, 253, 252, 4, 7, 8, 11, 30, 29, 24, 25, 230, 231, 226, 225, 244, 247, 248, 251, 5, 6, 9, 10, 31, 28, 27, 26, 219, 228, 227, 224, 245, 246, 249, 250, 58, 57, 54, 53, 32, 35, 36, 27, 218, 219, 220, 223, 202, 201, 198, 197, 59, 56, 55, 52, 33, 34, 39, 38, 217, 216, 221, 222, 203, 200, 199, 196, 60, 61, 50, 51, 46, 45, 40, 41, 214, 215, 210, 209, 204, 205, 194, 195, 63, 62, 49, 48, 47, 44, 43, 42, 213, 212, 211, 208, 207, 206, 193, 192, 64, 67, 68, 69, 122, 123, 124, 127, 128, 131, 132, 133, 186, 187, 188, 191, 65, 66, 71, 70, 121, 120, 125, 126, 129, 130, 135, 134, 185, 184, 189, 190, 78, 77, 72, 73, 118, 119, 114, 113, 142, 141, 136, 137, 182, 183, 178, 177, 79, 76, 75, 74, 117, 116, 115, 112, 143, 140, 139, 138, 181, 180, 179, 176, 80, 81, 94, 95, 96, 97, 110, 111, 144, 145, 158, 159, 160, 161, 174, 175, 83, 82, 93, 92, 99, 98, 109, 108, 147, 146, 157, 156, 163, 162, 173, 172, 84, 87, 88, 91, 100, 103, 104, 107, 148, 151, 152, 155, 164, 167, 168, 171, 85, 86, 89, 90, 91, 102, 105, 106, 149, 150, 153, 154, 155, 166, 169, 170}.
[0186] Figure 11C The region sequence number for each region in the above is as follows:
[0187] {170,171,172,175,176,177,190,191,192,195,196,197,250,251,252,255,169,168,173,174,179,178,189,188,193,194,199,198,249,248,253,254,166,167,162,161,180,183,184,187,206,205,200,201,246,247,242,241,
[0188] 155,164,163,160,181,182,185,186,207,204,203,202,245,244,243,240,
[0189] 154,155,156,159,138,137,134,133,208,209,222,223,224,225,238,239,
[0190] 153,152,157,158,139,136,135,132,211,210,221,220,227,226,237,236,
[0191] 150,151,146,145,140,141,130,131,212,215,216,219,228,231,232,235,
[0192] 149,148,147,144,143,142,129,128,213,214,217,218,219,230,233,234,
[0193] 106,107,108,111,112,113,126,127,42,41,38,27,26,25,22,21,
[0194] 105,104,109,110,115,114,125,124,43,40,39,36,27,24,23,20,
[0195] 102,103,98,97,116,119,120,123,44,45,34,35,28,29,18,19,
[0196] 91,100,99,96,117,118,121,122,47,46,33,32,31,30,17,16,
[0197] 90, 91, 92, 95, 74, 73, 70, 69, 48, 51, 52, 53, 10, 11, 12, 15,
[0198] 89, 88, 93, 94, 75, 72, 71, 68, 49, 50, 55, 54, 9, 8, 13, 14,
[0199] 86, 87, 82, 81, 76, 77, 66, 67, 62, 61, 56, 57, 6, 7, 2, 1,
[0200] 85, 84, 83, 80, 79, 78, 65, 64, 63, 60, 59, 58, 5, 4, 3, 0}.
[0201] If the picture is scanned in raster order, then Figure 11D The region sequence number for each region in
[0202] { 170, 169, 166, 155, 154, 153, 150, 149, 106, 105, 102, 91, 90, 89, 86, 85,
[0203] 171, 168, 167, 164, 155, 152, 151, 148, 107, 104, 103, 100, 91, 88, 87, 84,
[0204] 172, 173, 162, 163, 156, 157, 146, 147, 108, 109, 98, 99, 92, 93, 82, 83,
[0205] 175, 174, 161, 160, 159, 158, 145, 144, 111, 110, 97, 96, 95, 94, 81, 80,
[0206] 176, 179, 180, 181, 138, 139, 140, 143, 112, 115, 116, 117, 74, 75, 76, 79,
[0207] 177, 178, 183, 182, 137, 136, 141, 142, 113, 114, 119, 118, 73, 72, 77, 78,
[0208] 190, 189, 184, 185, 134, 135, 130, 129, 126, 125, 120, 121, 70, 71, 66, 65,
[0209] 191, 188, 187, 186, 133, 132, 131, 128, 127, 124, 123, 122, 69, 68, 67, 64,
[0210] 192, 193, 206, 207, 208, 211, 212, 213, 42, 43, 44, 47, 48, 49, 62, 63,
[0211] 195, 194, 205, 204, 209, 210, 215, 214, 41, 40, 45, 46, 51, 50, 61, 60,
[0212] 196, 199, 200, 203, 222, 221, 216, 217, 38, 39, 34, 33, 52, 55, 56, 59,
[0213] 197, 198, 201, 202, 223, 220, 219, 218, 27, 36, 35, 32, 53, 54, 57, 58,
[0214] 250, 249, 246, 245, 224, 227, 228, 219, 26, 27, 28, 31, 10, 9, 6, 5,
[0215] 251, 248, 247, 244, 225, 226, 231, 230, 25, 24, 29, 30, 11, 8, 7, 4,
[0216] 252, 253, 242, 243, 238, 237, 232, 233, 22, 23, 18, 17, 12, 13, 2, 3,
[0217] 255, 254, 241, 240, 239, 236, 235, 234, 21, 20, 19, 16, 15, 14, 1, 0}.
[0218] In the above four region sequences, the region sequence number i can be replaced by 255-i, which means the region sequence is reversed.
[0219] In some embodiments, since the number of region columns is not equal to the number of region rows, the quasi-Hilbert curve cannot be directly applied, but a part of the quasi-Hilbert type curve can be used. For example Figure 12A is a schematic diagram of an exemplary ALF region sequence of 128 regions according to Embodiment 1 of the present disclosure; Figure 12B is a schematic diagram of another exemplary ALF region sequence of 128 regions according to Embodiment 1 of the present disclosure; Figure 12Cis a diagram illustrating another exemplary ALF region order of 128 regions according to Embodiment 1 of the present disclosure; Figure 12D is a diagram illustrating another exemplary ALF region order of 128 regions according to Embodiment 1 of the present disclosure, as Figure 12A-12D shown, the picture is divided into 16 regions columns and 8 regions rows, in this case only half of the Hilbert curve is used. Figure 12A-12D 4 region orders are given in Table 1, which can be reversed, flipped or rotated.
[0220] If the picture is scanned in raster order, then Figure 12A The region sequence number of each region in Table 1 is as follows:
[0221] {0, 3, 4, 5, 58, 59, 60, 63, 64, 65, 78, 79, 80, 83, 84, 85,
[0222] 1, 2, 7, 6, 57, 56, 61, 62, 67, 66, 77, 76, 81, 82, 87, 86,
[0223] 14, 13, 8, 9, 54, 55, 50, 49, 68, 71, 72, 75, 94, 93, 88, 89,
[0224] 15, 12, 11, 10, 53, 52, 51, 48, 69, 70, 73, 74, 95, 92, 91, 90,
[0225] 16, 17, 30, 31, 32, 33, 46, 47, 122, 121, 118, 117, 96, 99, 100, 91,
[0226] 19, 18, 29, 28, 35, 34, 45, 44, 123, 120, 119, 116, 97, 98, 103, 102,
[0227] 20, 23, 24, 27, 36, 39, 40, 43, 124, 125, 114, 115, 110, 109, 104, 105,
[0228] 21, 22, 25, 26, 27, 38, 41, 42, 127, 126, 113, 112, 111, 108, 107, 106}.
[0229] If the picture is scanned in raster order, then Figure 12B The region sequence number of each region in Table 1 is as follows:
[0230] {0, 3, 4, 5, 58, 59, 60, 63, 64, 67, 68, 69, 122, 123, 124, 127,
[0231] 1, 2, 7, 6, 57, 56, 61, 62, 65, 66, 71, 70, 121, 120, 125, 126,
[0232] 14, 13, 8, 9, 54, 55, 50, 49, 78, 77, 72, 73, 118, 119, 114, 113,
[0233] 15, 12, 11, 10, 53, 52, 51, 48, 79, 76, 75, 74, 117, 116, 115, 112,
[0234] 16, 17, 30, 31, 32, 33, 46, 47, 80, 81, 94, 95, 96, 97, 110, 111,
[0235] 19, 18, 29, 28, 35, 34, 45, 44, 83, 82, 93, 92, 99, 98, 109, 108,
[0236] 20, 23, 24, 27, 36, 39, 40, 43, 84, 87, 88, 91, 100, 103, 104, 107,
[0237] 21, 22, 25, 26, 27, 38, 41, 42, 85, 86, 89, 90, 91, 102, 105, 106}.
[0238] If the picture is scanned in raster order, then Figure 12C the region sequence number of each region in
[0239] { 106, 107, 108, 111, 112, 113, 126, 127, 42, 41, 38, 27, 26, 25, 22, 21,
[0240] 105, 104, 109, 110, 115, 114, 125, 124, 43, 40, 39, 36, 27, 24, 23, 20,
[0241] 102, 103, 98, 97, 116, 119, 120, 123, 44, 45, 34, 35, 28, 29, 18, 19,
[0242] 91, 100, 99, 96, 117, 118, 121, 122, 47, 46, 33, 32, 31, 30, 17, 16,
[0243] 90, 91, 92, 95, 74, 73, 70, 69, 48, 51, 52, 53, 10, 11, 12, 15,
[0244] 89, 88, 93, 94, 75, 72, 71, 68, 49, 50, 55, 54, 9, 8, 13, 14,
[0245] 86, 87, 82, 81, 76, 77, 66, 67, 62, 61, 56, 57, 6, 7, 2, 1,
[0246] 85, 84, 83, 80, 79, 78, 65, 64, 63, 60, 59, 58, 5, 4, 3, 0}.
[0247] If the picture is scanned in raster order, then Figure 12D The region sequence number of each region in
[0248] { 106, 105, 102, 101, 90, 89, 86, 85, 42, 41, 38, 27, 26, 25, 22, 21,
[0249] 107, 104, 103, 100, 91, 88, 87, 84, 43, 40, 39, 36, 27, 24, 23, 20,
[0250] 108, 109, 98, 99, 92, 93, 82, 83, 44, 45, 34, 35, 28, 29, 18, 19,
[0251] 111, 110, 97, 96, 95, 94, 81, 80, 47, 46, 33, 32, 31, 30, 17, 16,
[0252] 112, 115, 116, 117, 74, 75, 76, 79, 48, 51, 52, 53, 10, 11, 12, 15,
[0253] 113, 114, 119, 118, 73, 72, 77, 78, 49, 50, 55, 54, 9, 8, 13, 14,
[0254] 126, 125, 120, 121, 70, 71, 66, 65, 62, 61, 56, 57, 6, 7, 2, 1,
[0255] 127, 124, 123, 122, 69, 68, 67, 64, 63, 60, 59, 58, 5, 4, 3, 0}.
[0256] In the above four region sequences, the region sequence number i can be replaced by 127-i so that the region sequence is reversed.
[0257] In some embodiments, the region order is variable, the encoder decides which region order to use in a picture level or sequence level, and signals an index of the region order used in the bitstream. The decoder determines the region order of the current picture or the current sequence according to the region order index parsed from the bitstream. As Figure 12A-12D shown in the example, the picture supports four different orders, thus uses 2 bits to encode the index with values from 0 to 3.
[0258] Consistent with the disclosed embodiments, the partition mode and the region order can be combined. According to some exemplary embodiments, variable partition mode and variable region order are supported. The combination of the partition mode and the region order is decided by the encoder and signaled in the bitstream. The decoder determines the partition mode and the region order of the current picture or the current sequence according to the index signaled in the bitstream.
[0259] For example, the picture is partitioned into 8 region columns and 8 region rows, two partition modes as shown in Figure 7A and 7C are supported, and two region orders as shown in Figure 10A and 10B are supported, so totally 4 combinations are supported. The encoder selects one combination from the four options in the picture level or sequence level, and signals an index of the selected combination of the partition mode and the region order in the bitstream. Since there are four combinations, the index value is set from 0 to 3. The decoder decides the combination of the current picture or sequence according to the index signaled in the bitstream. For example, if the index is equal to 0, the partition mode in Figure 7A and the region order in Figure 10A are used. If the index is equal to 1, the partition mode in Figure 7B and the region order in Figure 10A are used. For example, if the index is equal to 2, the partition mode in Figure 7A and the region order in Figure 10B are used. For example, if the index is equal to 3, the partition mode in Figure 7B and the region order in Figure 10B are used.
[0260] Consistent with the disclosed embodiments, partition modes and region orders can be combined, but the combination is restricted. That is, a partition mode can only be applied with a corresponding region order for a picture or video sequence. According to some example embodiments, variable partition modes and variable region orders are supported. However, a partition mode can not be allowed to be combined with every supported region order. If a certain partition mode is selected, only a specified scan order or certain specified scans can be used. In some example embodiments, each partition mode can only be applied with a corresponding region order, and each region order can only be applied with a corresponding partition mode. Thus, an index can be used to indicate the region order and the partition mode. And a decoder determines the partition mode and the region order according to the index signaled in the bitstream. After determining the partition mode and the region order, the decoder can determine the ALF index for the current region, and use the ALF index to process the region with the correct ALF.
[0261] For example, a picture is partitioned into 8 region columns and 8 region rows, four partition modes as shown in Figure 7A-7D are supported, and four region orders as shown in Figure 10A-10D are supported. Thus, there are 16 different combinations of partition modes and scan orders in total. However, not all of the 16 combinations are allowed. Figure 7A The partition mode in Figure 10A can only be combined with the scan order in Figure 7B The partition mode in Figure 10B can only be combined with the scan order in Figure 7C The partition mode in Figure 10C can only be combined with the scan order in Figure 7D The partition mode in Figure 10D can only be combined with the scan order in That is, only four combinations are allowed. Thus, an encoder selects one from the four options for each picture or the entire sequence, and signals an index with a value from 0 to 3 in the bitstream. A decoder decides the partition mode and the region order according to the index in the bitstream. The present disclosure does not restrict the combination of partition modes and region orders. Thus, in another example, Figure 7A The partition mode in Figure 10B can be combined with the scan order in Figure 7B The partition mode in Figure 10C can be combined with the scan order in Figure 7C The partition mode in Figure 10D can be combined with the scan order in Figure 7D The partition mode in Figure 10A can be combined with the scan order in The index signaled in these embodiments indicates the partition mode and the region order.
[0262] Embodiment 2
[0263] According to an embodiment of the present disclosure, a video data processing method is also provided, including: determining a maximum number of an adaptive loop filter (ALF) of a component of a picture; processing pixels in the picture with the ALF; signaling a first index of the maximum number of the ALF of the component of the picture.
[0264] In the above embodiment of the present disclosure, the first index equal to the first value indicates that the maximum number of the ALF is 64; and the first index equal to the second value or not equal to the first value indicates that the maximum number of the ALF is 16.
[0265] In the above embodiment of the present disclosure, further including: determining an order of an ALF region of the picture; signaling a second index of the order of the ALF region of the picture.
[0266] In the above embodiment of the present disclosure, the second index uses a 2-bit code.
[0267] In the above embodiment of the present disclosure, further including: determining an ALF index based on the second index; processing the pixels in the picture according to the ALF index.
[0268] In the above embodiment of the present disclosure, determining the ALF index based on the second index includes: determining a first column width based on a picture width and a maximum coding unit (LCU) width; determining a first row height based on a picture height and an LCU height; determining a region index based on the first column width and the first row height; and determining the ALF index based on the region index and the second index.
[0269] In the above embodiment of the present disclosure, determining the ALF index based on the region index and the second index includes: determining a second column width based on the first column width, the picture width and the LCU width; determining a second row height based on the first row height, the picture height and the LCU height; determining the region index based on the first column width, the second column width, the first row height and the second row height; and determining the ALF index based on the region index and the second index.
[0270] In the above embodiment of the present disclosure, determining the region index based on the first column width and the first row height includes: determining a first horizontal region index and a second horizontal region index based on the first column width; determining a first vertical region index and a second vertical region index based on the first row height; and determining the region index based on the first horizontal region index, the second horizontal region index, the first vertical region index, the second vertical region index and the second index.
[0271] In the above embodiment of the present disclosure, determining the ALF index based on the region index and the second index includes: determining a sequence number based on the region index and the second index; and determining the ALF index based on the sequence number.
[0272] In the above embodiments of this disclosure, the sequence number = regionTable[second index][region index], where regionTable is a two-dimensional circular table, and the two-dimensional circular table is defined as regionTable[4]
[64] = {
[0273] {63,60,59,58,5,4,3,0,62,61,56,57,6,7,2,1,49,50,55,54,9,8,13,14,48,51,52,53,10,11,12,15,47,46,33,32,31,30,17,16,44,45,34,35,28,29,18,19,43,40,39,36,27,24,23,20,42,41,38,37,26,25,22,21}
[0274] {42,43,44,47,48,49,62,63,41,40,45,46,51,50,61,60,38,39,34,33,52,55,56,59,37,36,35,32,53,54,57,58,26,27,28,31,10,9,6,5,25,24,29,30,11,8,7,4,22,23,18,17,12,13,2,3,21,20,19,16,15,1
[0275] 4,1,0},
[0276] {21,22,25,26,37,38,41,42,20,23,24,27,36,39,40,43,19,18,29,28,35,34,45,44,16,17,30,31,32,33,46,47,15,12,11,10,53,52,51,48,14,13,8,9,54,55,50,49,1,2,7,6,57,56,61,62,0,3,4,5,58,59,60,63}
[0277] {0,1,14,15,16,19,20,21,3,2,13,12,17,18,23,22,4,7,8,11,30,29,24,25,5,6,9,10,31,28,27,26,58,57,54,53,32,35,36,37,59,56,55,52,33,34,39,38,60,61,50,51,46,45,40,41,63,62,49,48,47,44,43,42}
[0278] Example 3
[0279] According to an embodiment of the present disclosure, a video data processing apparatus is also provided. The apparatus includes a memory configured to store instructions, and one or more processors configured to execute the instructions to cause the apparatus to perform receiving a bitstream, decoding a first index from the bitstream, determining a maximum number of adaptive loop filters (ALFs) for a component of a picture based on the first index, and processing pixels in the picture with the ALFs.
[0280] In the embodiments of the present disclosure, the processor is further configured to execute the instructions to cause the apparatus to perform determining the maximum number of ALFs as 64 in response to the first index being equal to a first value, or determining the maximum number of ALFs as 16 in response to the first index being equal to a second value or not equal to the first value.
[0281] In the embodiments of the present disclosure, before processing the pixels in the picture with the ALFs, the processor is further configured to execute the instructions to cause the apparatus to perform decoding a second index from the bitstream, wherein the second index indicates an order of ALF regions of the picture.
[0282] In the embodiments of the present disclosure, the second index is encoded by 2 bits.
[0283] In the embodiments of the present disclosure, when processing the pixels in the picture with the ALFs, the processor is further configured to execute the instructions to cause the apparatus to perform determining an ALF index based on the second index, and processing the pixels in the picture with the ALFs according to the ALF index.
[0284] In the embodiments of the present disclosure, when determining the ALF index based on the second index, the processor is further configured to execute the instructions to cause the apparatus to perform determining a first column width based on a picture width and a maximum coding unit (LCU) width, determining a first row height based on a picture height and a LCU height, determining a region index based on the first column width and the first row height, and determining the ALF index based on the region index and the second index.
[0285] In the embodiments of the present disclosure, when determining the ALF index based on the region index and the second index, the processor is further configured to execute the instructions to cause the apparatus to perform determining a second column width based on the first column width, the picture width and the LCU width, determining a second row height based on the first row height, the picture height and the LCU height, determining the region index based on the first column width, the second column width, the first row height and the second row height, and determining the ALF index based on the region index and the second index.
[0286] In the above embodiments of the present disclosure, in determining the region index based on the first column width and the first row height, the processor is further configured to execute instructions to cause the device to: determine a first horizontal region index and a second horizontal region index based on the first column width; determine a first vertical region index and a second vertical region index based on the first row height; and determine the region index based on the first horizontal region index, the second horizontal region index, the first vertical region index, the second vertical region index, and the second index.
[0287] In the above embodiments of the present disclosure, in determining the ALF index based on the region index and the second index, the processor is further configured to execute instructions to cause the device to: determine a sequence number based on the region index and the second index; and determine the ALF index based on the sequence number.
[0288] In the above embodiments of the present disclosure, the processor is further configured to execute instructions to cause the device to determine the sequence number based on the region index and the second index as follows: sequence number = regionTable[second index][region index], where regionTable is a two-dimensional circular table, and the two-dimensional circular table is defined as regionTable[4]
[64] = {
[0289] {63, 60, 59, 58, 5, 4, 3, 0, 62, 61, 56, 57, 6, 7, 2, 1, 49, 50, 55, 54, 9, 8, 13, 14, 48, 51, 52, 53, 10, 11, 12, 15, 47, 46, 33, 32, 31, 30, 17, 16, 44, 45, 34, 35, 28, 29, 18, 19, 43, 40, 39, 36, 27, 24, 23, 20, 42, 41, 38, 37, 26, 25, 22, 21},
[0290] {42, 43, 44, 47, 48, 49, 62, 63, 41, 40, 45, 46, 51, 50, 61, 60, 38, 39, 34, 33, 52, 55, 56, 59, 37, 36, 35, 32, 53, 54, 57, 58, 26, 27, 28, 31, 10, 9, 6, 5, 25, 24, 29, 30, 11, 8, 7, 4, 22, 23, 18, 17, 12, 13, 2, 3, 21, 20, 19, 16, 15, 1
[0291] 4, 1, 0},
[0292] {21, 22, 25, 26, 37, 38, 41, 42, 20, 23, 24, 27, 36, 39, 40, 43, 19, 18, 29, 28, 35, 34, 45, 44, 16, 17, 30, 31, 32, 33, 46, 47, 15, 12, 11, 10, 53, 52, 51, 48, 14, 13, 8, 9, 54, 55, 50, 49, 1, 2, 7, 6, 57, 56, 61, 62, 0, 3, 4, 5, 58, 59, 60, 63},
[0293] {0, 1, 14, 15, 16, 19, 20, 21, 3, 2, 13, 12, 17, 18, 23, 22, 4, 7, 8, 11, 30, 29, 24, 25, 5, 6, 9, 10, 31, 28, 27, 26, 58, 57, 54, 53, 32, 35, 36, 37, 59, 56, 55, 52, 33, 34, 39, 38, 60, 61, 50, 51, 46, 45, 40, 41, 63, 62, 49, 48, 47, 44, 43, 42}
[0294] }.
[0295] Embodiment 4
[0296] According to the embodiments of the present disclosure, a device for performing video data processing is also provided, the device includes a memory configured to store instructions, and one or more processors configured to execute the instructions to cause the device to perform: determining a maximum number of adaptive loop filters (ALF) of a component of a picture; processing pixels in the picture with the ALF; signaling a first index indicating the maximum number of the ALF of the component of the picture.
[0297] In the above embodiments of the present disclosure, the first index equal to the first value indicates that the maximum number of the ALF is 64; and the first index equal to the second value or not equal to the first value indicates that the maximum number of the ALF is 16.
[0298] In the above embodiments of the present disclosure, the processor is further configured to execute the instructions to cause the device to perform: determining an order of ALF regions of the picture; signaling a second index indicating the order of the ALF regions of the picture.
[0299] In the above embodiments of the present disclosure, the second index uses a 2-bit coding.
[0300] In the above embodiments of the present disclosure, the processor is further configured to execute the instructions to cause the device to perform: determining an ALF index based on the second index; processing the pixels in the picture according to the ALF with the ALF index.
[0301] In the above embodiments of the present disclosure, when the ALF index is determined based on the second index, the processor is further configured to execute instructions to cause the apparatus to: determine a first column width based on a picture width and a maximum coding unit (LCU) width; determine a first row height based on a picture height and an LCU height; determine a region index based on the first column width and the first row height; and determine the ALF index based on the region index and the second index.
[0302] In the above embodiments of the present disclosure, when the ALF index is determined based on the region index and the second index, the processor is further configured to execute instructions to cause the apparatus to: determine a second column width based on the first column width, the picture width and the LCU width; determine a second row height based on the first row height, the picture height and the LCU height; determine the region index based on the first column width, the second column width, the first row height and the second row height; and determine the ALF index based on the region index and the second index.
[0303] In the above embodiments of the present disclosure, when the region index is determined based on the first column width and the first row height, the processor is further configured to execute instructions to cause the apparatus to: determine a first horizontal region index and a second horizontal region index based on the first column width; determine a first vertical region index and a second vertical region index based on the first row height; and determine the region index based on the first horizontal region index, the second horizontal region index, the first vertical region index, the second vertical region index and the second index.
[0304] In the above embodiments of the present disclosure, when the ALF index is determined based on the region index and the second index, the processor is further configured to execute instructions to cause the apparatus to: determine a sequence number based on the region index and the second index; and determine the ALF index based on the sequence number.
[0305] In the above embodiments of the present disclosure, the processor is further configured to execute instructions to cause the apparatus to determine the sequence number based on the region index and the second index as follows: sequence number = regionTable[second index][region index], wherein the regionTable is a two-dimensional cyclic table, and the two-dimensional cyclic table is defined as regionTable[4]
[64] = {
[0306] {63, 60, 59, 58, 5, 4, 3, 0, 62, 61, 56, 57, 6, 7, 2, 1, 49, 50, 55, 54, 9, 8, 13, 14, 48, 51, 52, 53, 10, 11, 12, 15, 47, 46, 33, 32, 31, 30, 17, 16, 44, 45, 34, 35, 28, 29, 18, 19, 43, 40, 39, 36, 27, 24, 23, 20, 42, 41, 38, 37, 26, 25, 22, 21},
[0307] {42, 43, 44, 47, 48, 49, 62, 63, 41, 40, 45, 46, 51, 50, 61, 60, 38, 39, 34, 33, 52, 55, 56, 59, 37, 36, 35, 32, 53, 54, 57, 58, 26, 27, 28, 31, 10, 9, 6, 5, 25, 24, 29, 30, 11, 8, 7, 4, 22, 23, 18, 17, 12, 13, 2, 3, 21, 20, 19, 16, 15, 1
[0308] 4, 1, 0},
[0309] {21, 22, 25, 26, 37, 38, 41, 42, 20, 23, 24, 27, 36, 39, 40, 43, 19, 18, 29, 28, 35, 34, 45, 44, 16, 17, 30, 31, 32, 33, 46, 47, 15, 12, 11, 10, 53, 52, 51, 48, 14, 13, 8, 9, 54, 55, 50, 49, 1, 2, 7, 6, 57, 56, 61, 62, 0, 3, 4, 5, 58, 59, 60, 63},
[0310] {0, 1, 14, 15, 16, 19, 20, 21, 3, 2, 13, 12, 17, 18, 23, 22, 4, 7, 8, 11, 30, 29, 24, 25, 5, 6, 9, 10, 31, 28, 27, 26, 58, 57, 54, 53, 32, 35, 36, 37, 59, 56, 55, 52, 33, 34, 39, 38, 60, 61, 50, 51, 46, 45, 40, 41, 63, 62, 49, 48, 47, 44, 43, 42}
[0311] }.
[0312] Embodiment 5
[0313] According to the embodiments of the present disclosure, a non-transitory computer readable medium storing a set of instructions is also provided, the set of instructions executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising: receiving a bitstream; decoding a first index from the bitstream; determining a maximum number of adaptive loop filters (ALFs) for a component of a picture based on the first index; and processing pixels in the picture with the ALFs.
[0314] In the above embodiments of the present disclosure, the set of instructions executable by one or more processors of an apparatus to cause the apparatus to further perform: in response to the first index being equal to a first value, determining that the maximum number of ALFs is 64; or in response to the first index being equal to a second value or not equal to the first value, determining that the maximum number of ALFs is 16.
[0315] In the embodiments of the present disclosure, before processing the pixels in the picture with the ALF, the instruction set can be executed by the one or more processors of the device to make the device further perform: decoding a second index from the bitstream, wherein the second index indicates an order of the ALF regions of the picture.
[0316] In the embodiments of the present disclosure, the second index is encoded by 2 bits.
[0317] In the embodiments of the present disclosure, when processing the pixels in the picture with the ALF, the instruction set can be executed by the one or more processors of the device to make the device further perform: determining the ALF index based on the second index; and processing the pixels in the picture with the ALF according to the ALF index.
[0318] In the embodiments of the present disclosure, when determining the ALF index based on the second index, the instruction set can be executed by the one or more processors of the device to make the device further perform: determining a first column width based on a picture width and a maximum coding unit (LCU) width; determining a first row height based on a picture height and a LCU height; determining a region index based on the first column width and the first row height; and determining the ALF index based on the region index and the second index.
[0319] In the embodiments of the present disclosure, when determining the ALF index based on the region index and the second index, the instruction set can be executed by the one or more processors of the device to make the device further perform: determining a second column width based on the first column width, the picture width and the LCU width; determining a second row height based on the first row height, the picture height and the LCU height; determining the region index based on the first column width, the second column width, the first row height and the second row height; and determining the ALF index based on the region index and the second index.
[0320] In the embodiments of the present disclosure, when determining the region index based on the first column width and the first row height, the instruction set can be executed by the one or more processors of the device to make the device further perform: determining a first horizontal region index and a second horizontal region index based on the first column width; determining a first vertical region index and a second vertical region index based on the first row height; and determining the region index based on the first horizontal region index, the second horizontal region index, the first vertical region index, the second vertical region index and the second index.
[0321] In the embodiments of the present disclosure, when determining the ALF index based on the region index and the second index, the instruction set can be executed by the one or more processors of the device to make the device further perform: determining a sequence number based on the region index and the second index; and determining the ALF index based on the sequence number.
[0322] In the above embodiments of this disclosure, the instruction set can be executed by one or more processors of the device to enable the device to further determine the sequence number based on the region index and the second index as follows: sequence number = regionTable[second index][region index], where regionTable is a two-dimensional circular table, and the two-dimensional circular table is defined as regionTable[4]
[64] = {
[0323] {63,60,59,58,5,4,3,0,62,61,56,57,6,7,2,1,49,50,55,54,9,8,13,14,48,51,52,53,10,11,12,15,47,46,33,32,31,30,17,16,44,45,34,35,28,29,18,19,43,40,39,36,27,24,23,20,42,41,38,37,26,25,22,21}
[0324] {42,43,44,47,48,49,62,63,41,40,45,46,51,50,61,60,38,39,34,33,52,55,56,59,37,36,35,32,53,54,57,58,26,27,28,31,10,9,6,5,25,24,29,30,11,8,7,4,22,23,18,17,12,13,2,3,21,20,19,16,15,1
[0325] 4,1,0},
[0326] {21,22,25,26,37,38,41,42,20,23,24,27,36,39,40,43,19,18,29,28,35,34,45,44,16,17,30,31,32,33,46,47,15,12,11,10,53,52,51,48,14,13,8,9,54,55,50,49,1,2,7,6,57,56,61,62,0,3,4,5,58,59,60,63}
[0327] {0,1,14,15,16,19,20,21,3,2,13,12,17,18,23,22,4,7,8,11,30,29,24,25,5,6,9,10,31,28,27,26,58,57,54,53,32,35,36,37,59,56,55,52,33,34,39,38,60,61,50,51,46,45,40,41,63,62,49,48,47,44,43,42}
[0328] }。
[0329] Embodiment 6
[0330] According to the embodiments of the present disclosure, a non-transitory computer readable medium storing a set of instructions is also provided, the set of instructions executable by one or more processors of an apparatus to cause the apparatus to initiate a method for performing video data processing, the method comprising: determining a maximum number of adaptive loop filter (ALF) for a component of a picture; processing pixels in the picture with the ALF; signaling a first index indicating the maximum number of ALF for the component of the picture.
[0331] In the embodiments of the present disclosure, the first index equal to the first value indicates that the maximum number of ALF is 64, and the first index equal to the second value or not equal to the first value indicates that the maximum number of ALF is 16.
[0332] In the embodiments of the present disclosure, the set of instructions executable by one or more processors of the apparatus to cause the apparatus to further perform: determining an order of ALF regions of the picture; signaling a second index indicating the order of ALF regions of the picture.
[0333] In the embodiments of the present disclosure, the second index is coded with 2 bits.
[0334] In the embodiments of the present disclosure, the set of instructions executable by one or more processors of the apparatus to cause the apparatus to further perform: determining an ALF index based on the second index; and processing the pixels in the picture with the ALF according to the ALF index.
[0335] In the embodiments of the present disclosure, when the ALF index is determined based on the second index, the set of instructions executable by one or more processors of the apparatus to cause the apparatus to further perform: determining a first column width based on a picture width and a maximum coding unit (LCU) width; determining a first row height based on a picture height and a LCU height; determining a region index based on the first column width and the first row height; determining the ALF index based on the region index and the second index.
[0336] In the embodiments of the present disclosure, when the ALF index is determined based on the region index and the second index, the set of instructions executable by one or more processors of the apparatus to cause the apparatus to further perform: determining a second column width based on the first column width, the picture width and the LCU width; determining a second row height based on the first row height, the picture height and the LCU height; determining the region index based on the first column width, the second column width, the first row height and the second row height; determining the ALF index based on the region index and the second index.
[0337] In the above embodiments of the present disclosure, when the region index is determined based on the first column width and the first row height, the instruction set can be executed by the one or more processors of the device to cause the device to further perform: determining a first horizontal region index and a second horizontal region index based on the first column width; determining a first vertical region index and a second vertical region index based on the first row height; determining the region index based on the first horizontal region index, the second horizontal region index, the first vertical region index, the second vertical region index, and the second index.
[0338] In the above embodiments of the present disclosure, when the ALF index is determined based on the region index and the second index, the instruction set can be executed by the one or more processors of the device to cause the device to further perform: determining a sequence number based on the region index and the second index; determining the ALF index based on the sequence number.
[0339] In the above embodiments of the present disclosure, the instruction set can be executed by the one or more processors of the device to cause the device to further determine the sequence number based on the region index and the second index as follows: sequence number = regionTable[second index][region index], wherein the regionTable is a two-dimensional circular table, and the two-dimensional circular table is defined as regionTable[4]
[64] = {
[0340] {63, 60, 59, 58, 5, 4, 3, 0, 62, 61, 56, 57, 6, 7, 2, 1, 49, 50, 55, 54, 9, 8, 13, 14, 48, 51, 52, 53, 10, 11, 12, 15, 47, 46, 33, 32, 31, 30, 17, 16, 44, 45, 34, 35, 28, 29, 18, 19, 43, 40, 39, 36, 27, 24, 23, 20, 42, 41, 38, 37, 26, 25, 22, 21},
[0341] {42, 43, 44, 47, 48, 49, 62, 63, 41, 40, 45, 46, 51, 50, 61, 60, 38, 39, 34, 33, 52, 55, 56, 59, 37, 36, 35, 32, 53, 54, 57, 58, 26, 27, 28, 31, 10, 9, 6, 5, 25, 24, 29, 30, 11, 8, 7, 4, 22, 23, 18, 17, 12, 13, 2, 3, 21, 20, 19, 16, 15, 1
[0342] 4, 1, 0},
[0343] {21,22,25,26,37,38,41,42,20,23,24,27,36,39,40,43,19,18,29,28,35,34,45,44,16,17,30,31,32,33,46,47,15,12,11,10,53,52,51,48,14,13,8,9,54,55,50,49,1,2,7,6,57,56,61,62,0,3,4,5,58,59,60,63},
[0344] {0,1,14,15,16,19,20,21,3,2,13,12,17,18,23,22,4,7,8,11,30,29,24,25,5,6,9,10,31,28,27,26,58,57,54,53,32,35,36,37,59,56,55,52,33,34,39,38,60,61,50,51,46,45,40,41,63,62,49,48,47,44,43,42}
[0345] }.
[0346] Example 7
[0347] According to an embodiment of the present disclosure, a non-transitory computer readable medium storing a bitstream is also provided, wherein the bitstream comprises a first index associated with video data, the first index indicating a maximum number of adaptive loop filters (ALFs) for a component of a picture.
[0348] In the above embodiments of the present disclosure, the bitstream comprises a second index associated with the video data, the second index indicating an order of ALF regions for the picture.
[0349] In some embodiments, a non-transitory computer readable storage medium comprising instructions executable by a device (e.g., the disclosed encoders and decoders) to perform the above methods is also provided. Common forms of non-transitory media include, for example, a floppy disk, flexible disk, hard disk, solid-state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, PROM, and EPROM, a FLASH-EPROM or any other flash memory, NVRAM, cache, register, any other memory chip or cartridge, and a network version thereof. The device can include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0350] It should be noted that relational terms herein, such as“first” and“second”, are used solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual relationship or order between such entities or actions. Moreover, the words“comprise”,“have”,“contain”, and“include”, and other similar forms, are intended to be open-ended and to mean that the item or items listed thereafter can be present, but are not limited to, the items listed. They are not intended to preclude the presence or addition of one or more other items.
[0351] As used herein, unless expressly stated otherwise, the term“or” includes all possible combinations of the referenced items, unless infeasible. For example, if it is stated that a database can include A or B, then unless specifically stated otherwise or infeasible, the database can include A, B, or A and B. As a second example, if it is stated that a database can include A, B, or C, then unless specifically stated otherwise or infeasible, the database can include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0352] It should be understood that the above-described embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. When executed by a processor, the software can perform the disclosed method. The computing units and other functional units described in the present disclosure can be implemented by hardware, software, or a combination of hardware and software. Those of ordinary skill in the art will also understand that multiple aforementioned modules / units can be combined into one module / unit, and each aforementioned module / unit can be further divided into multiple sub-modules / sub-units.
[0353] In the foregoing specification, embodiments have been described with reference to numerous specific details that can vary from implementation to implementation. Certain embodiments were described with reference to specific apparatuses and methods. It is intended that various embodiments can be practiced in other ways, including methods which are different from those explicitly described, and using apparatuses which are different from those explicitly shown. Certain embodiments were described above as possible combinations of enumerated items. It is intended that one of ordinary skill in the art can practice those embodiments in other ways, including combinations of items not explicitly enumerated. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims. The sequence of steps shown in the figures is also intended to be illustrative only, and is not intended to be limiting to any particular sequence of steps. Thus, those skilled in the art will appreciate that the steps can be performed in a different order than shown.
[0354] In the drawings and specification, there have been disclosed exemplary embodiments. However, many variations and modifications can be made to these embodiments. Thus, it is the intention that only such limitations be placed on the application that are imposed by the appended claims, wherefore this application should be understood to cover any and all adaptations or variations of the various embodiments and implementations disclosed above.
Claims
1. A video data processing method, characterized in that, include: Receive bit stream; The first index is obtained by decoding the bitstream; The maximum number of adaptive loop filter ALFs for determining the components of an image based on the first index includes: when the first index is equal to a first value, the maximum number of adaptive loop filter ALFs is 64; or, when the first index is equal to a second value or not equal to the first value, the maximum number of adaptive loop filter ALFs is 16. The pixels in the image are processed using the adaptive loop filter (ALF).
2. The method according to claim 1, characterized in that, Before processing the pixels in the image using the adaptive loop filter (ALF), the method further includes: A second index is obtained by decoding from the bitstream, wherein the second index is used to indicate the order of the adaptive loop filter (ALF) regions of the image.
3. The method according to claim 2, characterized in that, The second index is a two-bit code.
4. The method according to claim 2, characterized in that, Processing the pixels in the image using the adaptive loop filter (ALF) further includes: The adaptive loop filter ALF index is determined based on the second index; and The pixels in the image are processed using the adaptive loop filter ALF based on the ALF index.
5. The method according to claim 4, characterized in that, The ALF index of the adaptive loop filter is determined based on the second index, including: The width of the first column is determined based on the image width and the width of the largest coding unit (LCU). The height of the first row is determined based on the image height and LCU height. The region index is determined based on the first column width and the first row height; The adaptive loop filter ALF index is determined based on the region index and the second index.
6. The method according to claim 5, characterized in that, Determining the adaptive loop filter ALF index based on the region index and the second index includes: The second column width is determined based on the first column width, the image width, and the LCU width; The height of the second row is determined based on the height of the first row, the height of the image, and the height of the LCU. The region index is determined based on the first column width, the second column width, the first row height, and the second row height; and The adaptive loop filter ALF index is determined based on the region index and the second index.
7. The method according to claim 5, characterized in that, The region index is determined based on the first column width and the first row height, including: The first horizontal region index and the second horizontal region index are determined based on the first column width; The first vertical region index and the second vertical region index are determined based on the first row height; and The region index is determined based on the first horizontal region index, the second horizontal region index, the first vertical region index, the second vertical region index, and the second index.
8. The method according to claim 5, characterized in that, Determining the adaptive loop filter ALF index based on the region index and the second index includes: The serial number is determined based on the region index and the second index; The ALF index of the adaptive loop filter is determined based on the sequence number.
9. The method according to claim 8, characterized in that, The serial number is determined based on the region index and the second index as follows: The sequence number = region table [second index][region index], where the region table is a two-dimensional circular table, and the two-dimensional circular table is defined as region table [4][64] = { {63,60,59,58,5,4,3,0,62,61,56,57,6,7,2,1,49,50,55,54,9,8,13,14,48,51,52,53,10,11,12,15,47,46,33,32,31,30,17,16,44,45,34,35,28,29,18,19,43,40,39,36,27,24,23,20,42,41,38,37,26,25,22,21}, {42,43,44,47,48,49,62,63,41,40,45,46,51,50,61,60,38,39,34,33,52,55,56,59,37,36,35,32,53,54,57,58,26,27,28,31,10,9,6,5,25,24,29,30,11,8,7,4,22,23,18,17,12,13,2,3,21,20,19,16,15,1 4,1,0}, {21,22,25,26,37,38,41,42,20,23,24,27,36,39,40,43,19,18,29,28,35,34,45,44,16,17,30,31,32,33,46,47,15,12,11,10,53,52,51,48,14,13,8,9,54,55,50,49,1,2,7,6,57,56,61,62,0,3,4,5,58,59,60,63}, {0,1,14,15,16,19,20,21,3,2,13,12,17,18,23,22,4,7,8,11,30,29,24,25,5,6,9,10,31,28,27,26,58,57,54,53,32,35,36,37,59,56,55,52,33,34,39,38,60,61,50,51,46,45,40,41,63,62,49,48,47,44,43,42} }。 10. An apparatus for performing video data processing, comprising: Memory, configured to store instructions; One or more processors are configured to execute the instructions to cause the device to perform: Receive video sequences; One or more images in the video sequence are divided into basic processing units, basic processing sub-units, or processing regions. Encoding the basic processing unit, basic processing subunit, or processing region into a video bitstream includes: the encoder, based on the number of the basic processing unit, the basic processing subunit, or the processing region, signaling a first index indicating the maximum number of adaptive loop filters (ALFs) for the components of the image in the video bitstream, and indicating that when the first index is equal to a first value, the maximum number of adaptive loop filters (ALFs) is 64; or, indicating that when the first index is equal to a second value or not equal to the first value, the maximum number of adaptive loop filters (ALFs) is 16. The pixels in the image are processed using the adaptive loop filter (ALF).
11. The apparatus according to claim 10, characterized in that, Before processing the pixels in the image using the adaptive loop filter (ALF), the processor is also configured to execute the instructions to cause the device to perform: The signaling indicates the second index of the order of the adaptive loop filter (ALF) regions of the image.
12. The apparatus according to claim 11, characterized in that, When processing pixels in the image using the Adaptive Loop Filter (ALF), the processor is also configured to execute the instructions to cause the device to perform: The adaptive loop filter ALF index is determined based on the second index.
13. The apparatus according to claim 12, characterized in that, When determining the adaptive loop filter ALF index based on the second index, the processor is also configured to execute the instructions to cause the device to perform: The width of the first column is determined based on the image width and the width of the largest coding unit (LCU). The height of the first row is determined based on the image height and LCU height. The region index is determined based on the first column width and the first row height; The adaptive loop filter ALF index is determined based on the region index and the second index.
14. The apparatus according to claim 13, characterized in that, Determining the adaptive loop filter ALF index based on the region index and the second index includes: The second column width is determined based on the first column width, the image width, and the LCU width; The height of the second row is determined based on the height of the first row, the height of the image, and the height of the LCU. The region index is determined based on the first column width, the second column width, the first row height, and the second row height; and The adaptive loop filter ALF index is determined based on the region index and the second index.
15. A computer-readable storage medium storing an instruction set and a video bitstream, the instruction set being executable by one or more processors to generate the video bitstream, the method comprising: Receive video sequences; One or more images in the video sequence are divided into basic processing units, basic processing sub-units, or processing regions. Encoding the basic processing unit, basic processing subunit, or processing region into a video bitstream includes: the encoder, based on the number of the basic processing unit, the basic processing subunit, or the processing region, signaling a first index indicating the maximum number of adaptive loop filters (ALFs) for the components of the image in the video bitstream, and indicating that when the first index is equal to a first value, the maximum number of adaptive loop filters (ALFs) is 64; or, indicating that when the first index is equal to a second value or not equal to the first value, the maximum number of adaptive loop filters (ALFs) is 16. The pixels in the image are processed using the adaptive loop filter (ALF).
16. The computer-readable storage medium according to claim 15, characterized in that, Before processing the pixels in the image using the Adaptive Loop Filter (ALF), the instruction set may be executed by one or more processors to further perform the following: The signaling indicates the second index of the order of the adaptive loop filter (ALF) regions of the image.
17. The computer-readable storage medium according to claim 16, characterized in that, When processing pixels in the image using the Adaptive Loop Filter (ALF), the instruction set can be executed by one or more processors to further perform: The adaptive loop filter ALF index is determined based on the second index.
18. The computer-readable storage medium according to claim 17, characterized in that, When determining the adaptive loop filter ALF index based on the second index, the instruction set can be executed by one or more processors to further perform: The width of the first column is determined based on the image width and the width of the largest coding unit (LCU). The height of the first row is determined based on the image height and LCU height. The region index is determined based on the first column width and the first row height; The adaptive loop filter ALF index is determined based on the region index and the second index.
19. The computer-readable storage medium according to claim 18, characterized in that, The adaptive loop filter (ALF) index is determined based on the region index and the second index. The instruction set can be executed by one or more processors to further perform the following: The second column width is determined based on the first column width, the image width, and the LCU width; The height of the second row is determined based on the height of the first row, the height of the image, and the height of the LCU. The region index is determined based on the first column width, the second column width, the first row height, and the second row height. as well as The adaptive loop filter ALF index is determined based on the region index and the second index.
20. The computer-readable storage medium according to claim 18, characterized in that, The region index is determined based on the first column width and the first row height. The instruction set can be executed by one or more processors to further perform: The first horizontal region index and the second horizontal region index are determined based on the first column width; The first vertical region index and the second vertical region index are determined based on the height of the first row. as well as The region index is determined based on the first horizontal region index, the second horizontal region index, the first vertical region index, the second vertical region index, and the second index.
Citation Information
Patent Citations
Image filter apparatus, decoder apparatus, encoder apparatus, and data structure
WO2012137890A1