Adaptive loop filter for video coding

By applying non-zero filter coefficients to adjust the filtering range of ALF and CCALF during video encoding and decoding, the problem of low encoding efficiency in existing technologies is solved, achieving more efficient video data processing and quality improvement.

CN121773609APending Publication Date: 2026-03-31BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing video coding techniques have room for improvement in terms of coding efficiency for Adaptive Loop Filter (ALF) and Cross-Component Adaptive Loop Filter (CCALF), especially in the process of video data compression and decoding, where it is difficult to effectively reduce redundancy and maintain video quality.

Method used

By applying non-zero filter coefficients in the decoder and encoder, the filtering range of the adaptive loop filter (ALF) and the cross-component adaptive loop filter (CCALF) is dynamically adjusted to filter the current block, generating or decoding the bitstream.

Benefits of technology

It improves the efficiency of video encoding and decoding, reduces video data redundancy, enhances video quality, and optimizes the generation process of encoded bitstreams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121773609A_ABST
    Figure CN121773609A_ABST
Patent Text Reader

Abstract

Methods, apparatus, and non-transitory computer-readable storage media for video decoding and encoding are provided. A method for video decoding includes receiving, by a decoder, at least one non-zero filter coefficient; applying, by the decoder, at least one non-zero filter coefficient to at least one range of a dynamic range of adaptive loop filter (ALF) input values; and filtering, by the decoder, a reconstructed video block corresponding to the current block based on a result of applying the at least one non-zero filter coefficient to the at least one range.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is based on and claims priority to U.S. Provisional Application No. 63 / 536,926, filed September 6, 2023, entitled “Adaptive Loop Filter for Video Coding,” the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0003] This disclosure relates to video coding and compression. More specifically, this application relates to methods and apparatus for improving the coding efficiency of adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF). Background Technology

[0004] Various electronic devices (such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc.) support digital video. Electronic devices send and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited storage resources of storage devices, video data can be compressed using one or more video codec standards before it is transmitted or stored. For example, video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data to a form using a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention

[0005] Embodiments of this disclosure provide methods and apparatus for improving the coding efficiency of adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF).

[0006] According to a first aspect of this disclosure, a method for video decoding is provided, comprising: receiving at least one non-zero filter coefficient by a decoder; applying the at least one non-zero filter coefficient to at least one range in the dynamic range of an adaptive loop filter (ALF) input value by the decoder; and filtering a reconstructed video block corresponding to a current block by the decoder based on the result of applying the at least one non-zero filter coefficient to the at least one range.

[0007] According to a second aspect of this disclosure, a video encoding method is provided, comprising: transmitting at least one non-zero filter coefficient by an encoder; applying the at least one non-zero filter coefficient to at least one range of the dynamic range of an adaptive loop filter (ALF) input value by the encoder; filtering a reconstructed video block corresponding to a current block by the encoder based on the result of applying the at least one non-zero filter coefficient to the at least one range; and generating a bitstream by the encoder based on the result of filtering the reconstructed video block.

[0008] According to a third aspect of this disclosure, an apparatus for video decoding is provided. The apparatus may include: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to perform the method according to the first aspect when executing the instructions.

[0009] According to a fourth aspect of this disclosure, an apparatus for video encoding is provided. The apparatus may include: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to perform the method according to the second aspect when executing the instructions.

[0010] According to a fifth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the first aspect.

[0011] According to a sixth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the second aspect.

[0012] According to a seventh aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream to be decoded by the method according to the first aspect.

[0013] According to the eighth aspect of this disclosure, a non-transitory computer-readable storage medium is provided for storing a bit stream generated by the method according to the second aspect. Attached Figure Description

[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate examples according to this disclosure and, together with this description, serve to explain the principles of this disclosure.

[0015] Figure 1 This is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some embodiments of the present disclosure.

[0016] Figure 2 This is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0017] Figure 3 This is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0018] Figures 4A to 4E This is a block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes according to some embodiments of this disclosure.

[0019] Figure 5 This is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0020] Figure 6 This is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0021] Figure 7 This is a diagram illustrating the shape of an ALF filter according to some examples of this disclosure.

[0022] Figure 8 It is a depiction of the gradient of subsampled points based on some examples of this disclosure.

[0023] Figure 9 This is a diagram illustrating the geometric transformation of the shape of a rhombus filter according to some examples of this disclosure.

[0024] Figure 10 This is a diagram illustrating the shape of an inline filter used in an ECM according to some examples of this disclosure.

[0025] Figure 11 This is a diagram of the CCALF architecture based on some examples of this disclosure.

[0026] Figure 12This is a diagram illustrating the relative positions of filtered chroma samples in a 4:2:0 chroma format with chroma location type 0 and their support in the luminance plane. The diagram is based on some examples of this disclosure.

[0027] Figure 13 This is a diagram of a 25-tap long filter based on some examples of this disclosure.

[0028] Figure 14 This is an illustration of the shape of a filter used to predict a signal or an SAO signal, according to an example of this disclosure.

[0029] Figure 15 This is an illustration of an adjusted ALF filter shape based on some examples of this disclosure.

[0030] Figure 16 This is a diagram illustrating a computing environment coupled to a user interface according to some embodiments of the present disclosure.

[0031] Figures 17A to 17C This is a diagram illustrating symmetrical sample filling of a luminance ALF filter according to an example of this disclosure.

[0032] Figures 18A to 18B This is a diagram illustrating an existing ALF clipping operation according to an example of this disclosure and a proposed clipping operation with scaling if the sample difference is too large.

[0033] Figure 19 This is a flowchart illustrating a method for video decoding according to some examples of this disclosure.

[0034] Figure 20 This illustrates some examples corresponding to, for example, this disclosure. Figure 19 The flowchart shown is for a method used for video decoding and a method used for video encoding. Detailed Implementation

[0035] Referring now to the detailed description, examples of which are illustrated in the accompanying drawings. Numerous non-limiting details are set forth in the following detailed description to aid in understanding the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.

[0036] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this disclosure are used to distinguish objects and are not used to describe any specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in sequences other than those shown in the drawings or described in this disclosure.

[0037] Figure 1 This is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, according to some embodiments of the present disclosure. Figure 1 As shown, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a target device 14. The source device 12 and target device 14 can include any electronic device from a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and target device 14 are equipped with wireless communication capabilities.

[0038] In some implementations, target device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to target device 14. In one example, link 16 may include a communication medium enabling source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated according to a communication standard (e.g., a wireless communication protocol) and transmitted to target device 14. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, switch, base station, or any other means that may facilitate communication from source device 12 to target device 14.

[0039] In some other implementations, encoded video data can be sent from output interface 22 to storage device 32. The target device 14 can then access the encoded video data in storage device 32 via input interface 28. Storage device 32 can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, digital universal discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Target device 14 can access the stored video data from storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending encoded video data to target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Target device 14 can access the encoded video data via any standard data connection suitable for accessing encoded video data stored on the file server. Standard data connections include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or a combination of both. Transfer of encoded video data from storage device 32 can be streaming, downloading, or a combination of both.

[0040] like Figure 1 As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources or combinations of such sources, such as: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if video source 18 is a camera in a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application are generally applicable to video encoding and decoding and can be applied to wireless and / or wired applications.

[0041] Video captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be sent directly to target device 14 via output interface 22 of source device 12. Alternatively, the encoded video data can be stored on storage device 32 for later access by target device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or transmitter. The encoded video data may include a sequence of images, each of which may include one or more sample arrays, for example, luminance (Y) for monochromatic purposes only; luminance and two chrominances in the YCbCr or YCgCo domain; or green, blue, and red in the GBR (also known as RGB) domain. For ease of reference and terminology in this application, in some embodiments, the variables and terms associated with each set of three sample arrays may be referred to as luminance and chrominance, where the two chrominance arrays may be referred to as Cb and Cr, regardless of the actual color representation used. The video data may be in 4:0:0, 4:2:0, 4:2:2 or 4:4:4 chroma format, but this application is not limited to this.

[0042] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem, and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements may be included within encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.

[0043] In some embodiments, the target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with the target device 14. The display device 34 displays decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0044] Video encoder 20 and video decoder 30 can operate according to proprietary or industry standards (e.g., VVC, HEVC, MPEG-4 Part 10, AVC) or extensions of such standards. It should be understood that this application is not limited to any particular video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally understood that the video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally understood that the video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.

[0045] The video encoder 20 and video decoder 30 can be implemented as any circuit of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, and either encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.

[0046] In some implementations, components of source device 12 (e.g., video source 18, video encoder 20, or the following references) Figure 2 The components included in the video encoder 20, and at least a portion of the components in the output interface 22, and / or the components of the target device 14 (e.g., the input interface 28, the video decoder 30, or the following references) Figure 3At least a portion of the components included in the video decoder 30 and the display device 34 can operate in a cloud computing service network such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), where the cloud computing service network can provide software, platform, and / or infrastructure. In some embodiments, one or more components of the source device 12 and / or the target device 14 that are not included in the cloud computing service network can be located in one or more client devices, and these client devices can communicate with server computers in the cloud computing service network via wireless communication networks (e.g., cellular communication networks, short-range wireless communication networks, or Global Navigation Satellite System (GNSS) communication networks) or wired communication networks (e.g., local area network (LAN) communication networks or power line communication (PLC) networks). In one embodiment, at least a portion of the operations described herein can be implemented as a cloud-based service provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein can be implemented by one or more client devices. In some implementations, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of this disclosure, terms such as “cloud,” “cloud computing,” and “cloud-based” may be used interchangeably as appropriate. It should be understood that this disclosure is not limited to implementation within the aforementioned cloud computing service networks. Instead, this disclosure may also be implemented in any other type of computing environment currently known or developed in the future.

[0047] Figure 2 This is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described in this application. The video encoder 20 can perform intra-frame predictive coding and inter-frame predictive coding on video blocks within a video frame. Intra-frame predictive coding relies on spatial prediction to reduce or remove spatial redundancy in the video data within a given video frame or picture. Inter-frame predictive coding relies on temporal prediction to reduce or remove temporal redundancy in the video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video encoding and decoding, the term "frame" can be used as a synonym for the terms "image" or "picture".

[0048] like Figure 2As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copying (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove block artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of the adder 62. It should be noted that, regarding CCSAO technology, this application is not limited to the embodiments described herein, but can also be applied to situations where an offset is selected for either the luminance component and either of the two chrominance components (which may represent Y, Cb, and Cr in the YCbCr domain; Y, Cg, and Co in the YCgCo domain; or G, B, and R in the RGB domain, for the purposes of the symbols and terminology described above) to modify the other component based on the selected offset. Furthermore, it should be noted that the first component mentioned herein can be either the luminance component or either of the two chrominance components, the second component mentioned herein can be either the luminance component or either of the two chrominance components, and the third component mentioned herein can be the remaining component among the luminance component and the two chrominance components. In some examples, the loop filter may be omitted, and the decoded video block may be directly provided to the DPB 64 by the adder 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be distributed among one or more of the described fixed or programmable hardware units.

[0049] The video data storage device 40 can store video data encoded by the components of the video encoder 20. For example, it can store data from... Figure 1 The video source 18 shown receives video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used by the video encoder 20 (e.g., in intra-frame or inter-frame predictive coding modes) when encoding the video data. The video data memory 40 and DPB 64 can be formed from any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip along with other components of the video encoder 20, or off-chip relative to those components.

[0050] like Figure 2 As shown, after receiving video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a collection of video blocks) or other larger coding units (CUs) according to a predefined splitting structure associated with the video data (e.g., a quadtree (QT) structure). A video frame is, or can be considered, a two-dimensional array or matrix of sample points with sample values. Sample points in the array may also be referred to as pixels or image elements (pel). The number of sample points in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the video frame. For example, a video frame can be divided into multiple video blocks using QT segmentation. A video block is again, or can be considered, a two-dimensional array or matrix of sample points with sample values, but its dimension is smaller than that of the video frame. The number of sample points in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. A video block can be further divided into one or more block partitions or sub-blocks (which can then re-form blocks) by iteratively using, for example, QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof. It should be noted that the term “block” or “video block” as used herein can be a portion of a frame or image, particularly a rectangular (square or non-square) portion. Referring, for example, to HEVC and VVC, a block or video block can be or corresponds to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or corresponds to a corresponding block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB)) and / or sub-block.

[0051] The prediction processing unit 41 can select one of several feasible predictive coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one or more inter-frame predictive coding modes among multiple intra-frame predictive coding modes. The prediction processing unit 41 can provide the resulting intra-frame or inter-frame predictive coding block to adder 50 to generate a residual block, and to adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, segmentation information, and other such syntax information) to entropy coding unit 56.

[0052] To select a suitable intra-predictive coding mode for the current video block, the intra-predictive processing unit 46 within the prediction processing unit 41 can perform intra-predictive coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple coding passes, for example, to select a suitable coding mode for each block of video data.

[0053] In some implementations, motion estimation unit 42 determines an inter-frame prediction mode for the current video frame by generating motion vectors based on a predetermined pattern within the video frame sequence. The motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame associated with the current block being encoded in the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra-frame BC unit 48 may determine vectors (e.g., block vectors) for intra-frame BC coding in a similar manner to how motion estimation unit 42 determines motion vectors for inter-frame prediction, or the block vectors may be determined using motion estimation unit 42.

[0054] Regarding pixel differences, the predicted block for a video block can be, or can correspond to, a block or reference block of a reference frame considered to closely match the video block to be encoded. Pixel differences can be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some implementations, the video encoder 20 can compute values ​​for sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 can interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0055] The motion estimation unit 42 calculates the motion vector for a video block in an inter-frame predictive coding frame by comparing the position of the video block with the position of the predicted block in a reference frame selected from either a first reference frame list (list 0) or a second reference frame list (list 1), each of the first and second reference frame lists identifying one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44, and then to the entropy coding unit 56.

[0056] Motion compensation performed by motion compensation unit 44 may involve acquiring or generating prediction blocks based on motion vectors determined by motion estimation unit 42. Upon receiving motion vectors for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in a reference frame list within a reference frame list, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel differences by subtracting the pixel values ​​of the prediction block provided by motion compensation unit 44 from the pixel values ​​of the currently encoded video block. The pixel differences forming the residual video block may include a luminance component difference or a chrominance component difference, or both. Motion compensation unit 44 may also generate syntax elements associated with video blocks of a video frame for use by video decoder 30 when decoding video blocks of a video frame. Syntax elements may include, for example, syntax elements defining motion vectors for identifying prediction blocks, any flags indicating prediction modes, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are described separately for conceptual purposes.

[0057] In some implementations, the intra-BC unit 48 can generate vectors and acquire prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44; however, these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra-BC unit 48 can determine the intra-prediction mode to be used for encoding the current block. In some examples, the intra-BC unit 48 can, for example, use various intra-prediction modes to encode the current block during individual encoding passes and test their performance through rate-distortion analysis. Next, the intra-BC unit 48 can select a suitable intra-prediction mode from the various tested intra-prediction modes to use and generate an intra-mode indicator accordingly. For example, the intra-BC unit 48 can use rate-distortion analysis to calculate rate-distortion values ​​for the various tested intra-prediction modes and select the intra-prediction mode with the best rate-distortion characteristics from the tested modes as the suitable intra-prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra-frame BC unit 48 can calculate the ratio based on the distortion and rate for various coded blocks to determine which intra-frame prediction mode exhibits the optimal rate-distortion value for the block.

[0058] In other examples, the intra-frame BC unit 48 may use, in whole or in part, the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra-frame BC prediction according to the embodiments described herein. In any case, for intra-frame block copying, in terms of pixel differences, the predicted block may be a block considered to closely match the block to be encoded, the pixel differences may be determined by SAD, SSD, or other difference metrics, and identifying the predicted block may include calculating values ​​for sub-integer pixel positions.

[0059] Regardless of whether the predicted block comes from the same frame predicted intra-frame or from different frames predicted inter-frame, the video encoder 20 can form a residual video block by subtracting the pixel values ​​of the predicted block from the pixel values ​​of the current video block being encoded. The pixel difference forming the residual video block can include both luma component difference and chroma component difference.

[0060] As an alternative to the inter-frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above, or the intra-block copy prediction performed by the intra-BC unit 48, the intra-prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, the intra-prediction processing unit 46 can determine an intra-prediction mode for encoding the current block. To this end, the intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during individual encoding passes, and the intra-prediction processing unit 46 (or, in some examples, the mode selection unit) can select a suitable intra-prediction mode from the tested intra-prediction modes for use. The intra-prediction processing unit 46 can provide information indicating the intra-prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode into the bitstream.

[0061] After prediction processing unit 41 determines the prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 uses a transform (e.g., discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients.

[0062] The transform processing unit 52 can send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can subsequently perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.

[0063] After quantization, the entropy coding unit 56 entropy codes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval segmented entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream can then be sent to, for example,... Figure 1 The video decoder 30 shown, or archived in, for example Figure 1 The data is stored in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 can also entropy code the motion vectors and other syntax elements used for the current video frame being encoded.

[0064] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual video blocks in the pixel domain for generating reference blocks to predict other video blocks. As noted above, the motion compensation unit 44 can generate motion-compensated prediction blocks from one or more reference blocks of frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction blocks to compute sub-integer pixel values ​​for use in motion estimation.

[0065] Adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block can then be used as a prediction block by intra-frame BC unit 48, motion estimation unit 42, and motion compensation unit 44 for inter-frame prediction of another video block in subsequent video frames.

[0066] Figure 3 This is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of this application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 can perform operations in conjunction with the above. Figure 2 The decoding process described for the video encoder 20 is essentially the inverse of the encoding process. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra-frame prediction unit 84 can generate prediction data based on the intra-frame prediction mode indicator received from the entropy decoding unit 80.

[0067] In some examples, units of the video decoder 30 may be assigned tasks to perform embodiments of this application. Furthermore, in some examples, embodiments of this disclosure may be distributed across one or more units of the video decoder 30. For example, the intra-frame BC unit 85 may perform embodiments of this application individually or in combination with other units of the video decoder 30 (e.g., motion compensation unit 82, intra-frame prediction unit 84, and entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra-frame BC unit 85, and the functionality of the intra-frame BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., motion compensation unit 82).

[0068] Video data memory 79 can store video data, such as encoded video bitstreams, that will be decoded by other components of video decoder 30. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source (e.g., a camera), via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of video decoder 30 stores reference video data for use by video decoder 30 (e.g., in intra-frame or inter-frame predictive coding modes) when decoding video data. Video data memory 79 and DPB 92 can be formed of any memory device from a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, video data memory 79 and DPB 92 are shown in... Figure 3 The video data memory 79 and DPB 92 are depicted as two distinct components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip along with other components of the video decoder 30, or off-chip relative to those components.

[0069] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of encoded video frames and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators, and other syntax elements to the prediction processing unit 81.

[0070] When a video frame is encoded as an intra-predictive coded (I) frame or used as an intra-coded prediction block in other types of frames, the intra-predictive unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-predictive mode transmitted by the signal and reference data from the previous decoded block of the current frame.

[0071] When a video frame is encoded as an inter-frame predictive coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the current video frame based on motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame within a reference frame list. The video decoder 30 can construct the reference frame list, i.e., list 0 and list 1, based on the reference frames stored in the DPB 92 using a default construction technique.

[0072] In some examples, when a video block is encoded according to the intra-frame BC mode described herein, the intra-frame BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block can be located within a reconstructed region of the same image as the current video block, as defined by the video encoder 20.

[0073] Motion compensation unit 82 and / or intra-frame prediction (BC) unit 85 determine prediction information for video blocks in the current video frame by parsing motion vectors and other syntax elements, and then use this prediction information to generate prediction blocks for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) used to encode video blocks in the video frame, the inter-frame prediction frame type (e.g., B or P), construction information for one or more reference frames in the reference frame list for the frame, motion vectors for each inter-frame prediction encoded video block in the frame, the inter-frame prediction state for each inter-frame prediction encoded video block in the frame, and other information for decoding video blocks in the current video frame.

[0074] Similarly, the intra-BC unit 85 can use some of the received syntax elements, such as flags, to determine which video blocks in the frame are predicted using the intra-BC mode, which video blocks in the frame are in the reconstruction region and should be stored in the DPB 92, the block vector for each intra-BC predicted video block in the frame, the intra-BC prediction state for each intra-BC predicted video block in the frame, and other information for decoding video blocks in the current video frame.

[0075] The motion compensation unit 82 can also perform interpolation using interpolation filters, such as those used by the video encoder 20 during encoding of video blocks, to calculate interpolated values ​​for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filters used by the video encoder 20 based on the received syntax elements and use these interpolation filters to generate the prediction block.

[0076] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.

[0077] After the motion compensation unit 82 or the intra-frame BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra-frame BC unit 85. A loop filter 91 (e.g., a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 for further processing of the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a separate memory device may also store the decoded video for later presentation on a display device (e.g., ...). Figure 1 On the display device 34).

[0078] In a typical video coding process, a video sequence usually consists of an ordered set of frames or images. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of chrominance samples (Cb). SCr is a two-dimensional array of chrominance samples (Cr). In other instances, a frame may be monochrome and therefore consist of only a two-dimensional array of luma samples.

[0079] like Figure 4AAs shown, the video encoder 20 (or more specifically, the segmentation unit 45) generates an encoded representation of a frame by first segmenting the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set such that all CTUs in the video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not limited to a specific size. Figure 4B As shown, each CTU may include a CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units in the coded pixel block and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome image or an image with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding the samples of that coding tree block. The coding tree block may be an N×N sample block.

[0080] To achieve better performance, the video encoder 20 can recursively perform tree splitting on the coding tree blocks of the CTU, such as binary tree splitting, ternary tree splitting, quadtree splitting, or combinations thereof, and divide the CTU into smaller CUs. Figure 4C As described, the 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Of these four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs, 430 and CU 440, are further divided into four CUs with a block size of 8×8. Figure 4D Depicting as shown Figure 4C The final result of the CTU 400 partitioning process described in the figure is a quadtree data structure, where each leaf node of the quadtree corresponds to a CU of a corresponding size ranging from 32×32 to 8×8. Similar to... Figure 4B The CTU depicted in the image can include, for example, two corresponding coding blocks (CBs) of luminance and chrominance samples of the same frame size, as well as syntax elements for encoding the samples of the coding blocks. In monochrome images or images with three separate color planes, a CU can include a single coding block and a syntax structure for encoding the samples of the coding block. It should be noted that... Figure 4C and Figure 4DThe quadtree partitioning depicted is for illustrative purposes only, and a CTU can be split into multiple CUs based on quadtree / ternary / binary partitioning to adapt to varying local characteristics. In multi-type tree structures, a CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to binary and ternary tree structures. Figure 4E As shown, a coded block with width W and height H has five possible segmentation types: quadruple segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal triple segmentation, and vertical triple segmentation.

[0081] In some implementations, the video encoder 20 may further segment the coded blocks of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter-frame or intra-frame) is applied. The PU of the CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome image or an image with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PBs. The video encoder 20 may generate predicted luma blocks, predicted Cb blocks, and predicted Cr blocks for each PU of the CU, representing the luma PB, Cb PB, and Cr PB.

[0082] Video encoder 20 can use intra-frame prediction or inter-frame prediction to generate prediction blocks for the PU. If video encoder 20 uses intra-frame prediction to generate prediction blocks for the PU, then video encoder 20 can generate prediction blocks for the PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate prediction blocks for the PU, then video encoder 20 can generate prediction blocks for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0083] After the video encoder 20 generates a predicted luminance block, a predicted Cb block, and a predicted Cr block for one or more PUs of the CU, the video encoder 20 can generate a luminance residual block for the CU by subtracting the predicted luminance block of the CU from the original luminance coding block of the CU, such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample in one of the predicted luminance blocks of the CU and a corresponding sample in the original luminance coding block of the CU. Similarly, the video encoder 20 can generate Cb residual blocks and Cr residual blocks for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0084] In addition, such as Figure 4CAs shown, the video encoder 20 can use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block of the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A TU of the CU can include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the samples of the transform block. Therefore, each TU of the CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU can be a sub-block of the CU's luminance residual block. A Cb transform block can be a sub-block of the CU's Cb residual block. A Cr transform block can be a sub-block of the CU's Cr residual block. In a monochrome image or an image with three separate color planes, a TU can include a single transform block and syntax structures for transforming the samples of that transform block.

[0085] The video encoder 20 can apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block can be a two-dimensional array of transform coefficients. The transform coefficients can be scalars. The video encoder 20 can apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 can apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.

[0086] After generating coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), video encoder 20 can quantize the coefficient blocks. Quantization typically refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After quantizing the coefficient blocks, video encoder 20 can entropy encode the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream comprising a bit sequence that forms a representation of coded frames and associated data; the bitstream is stored in storage device 32 or transmitted to target device 14.

[0087] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements. The video decoder 30 can reconstruct frames of video data, at least in part, based on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coded blocks of the current CU by adding the samples of the prediction block of the PU for the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the coded blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.

[0088] As mentioned above, video coding primarily uses two modes (i.e., intra-frame prediction and inter-frame prediction) to achieve video compression. It should be noted that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on a reference video block.

[0089] However, with continuously improving video data capture technologies and finer video tile sizes used to preserve details in video data, the amount of data required to represent the motion vectors for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only do a set of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (also known as the "motion vector predictor" (MVP) of the current CU) by exploring their spatial and temporal correlations.

[0090] Instead of the above combination Figure 2 The method described involves encoding the actual motion vector of the current CU, determined by the motion estimation unit 42, into the video bitstream, and subtracting the motion vector prediction factor of the current CU from the actual motion vector of the current CU to produce the motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.

[0091] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coded block, both the video encoder 20 and the video decoder 30 need to employ a set of rules to construct a motion vector candidate list (also known as a "merging list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs. Then, a member is selected from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list for encoding and decoding the current CU.

[0092] Figure 16 A computing environment 1610 coupled to a user interface 1650 is shown. The computing environment 1610 may be part of a data processing server. The computing environment 1610 includes a processor 1620, memory 1630, and input / output (I / O) interface 1640.

[0093] Processor 1620 typically controls the overall operation of computing environment 1610, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1620 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 1620 may include one or more modules that facilitate interaction between processor 1620 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.

[0094] Memory 1630 is configured to store various types of data to support the operation of computing environment 1610. Memory 1630 may include predefined software 1632. Examples of such data include instructions for any application or method operating on computing environment 1610, video datasets, image data, etc. Memory 1630 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0095] I / O interface 1640 provides an interface between processor 1620 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1640 can be coupled to encoders and decoders.

[0096] This disclosure relates to video coding and compression. More specifically, this application relates to methods and apparatus for improving the coding efficiency of adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF).

[0097] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, some well-known video coding standards today include Universal Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which were jointly developed by ISO / IEC MPEG and ITU-T TVCEG. AOMedia Video1 (AV1) was developed by the Alliance for Open Media (AOM) as the successor to its previous standard VP9. Audio and Video Coding (AVS) is another series of video compression standards developed by the China Audio and Video Coding Standards Working Group, which refers to digital audio and digital video compression standards. Most existing video coding standards are built on the well-known hybrid video coding framework, which uses block-based prediction methods (e.g., inter-frame prediction, intra-frame prediction) to reduce redundancy in video images or sequences and uses transform coding to compress the energy of prediction errors. A key goal of video coding technology is to compress video data into a form using a lower bit rate while avoiding or minimizing video quality degradation.

[0098] The first generation of AVS standards included the Chinese national standards "Information Technology, Advanced Audio-Video Coding, Part 2: Video" (referred to as AVS1) and "Information Technology, Advanced Audio-Video Coding, Part 16: Broadcast Television Video" (referred to as AVS+). Compared to the MPEG-2 standard, it offered approximately 50% bitrate savings while maintaining the same perceptual quality. The video portion of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second generation of AVS standards included a series of Chinese national standards, "Information Technology, High-Efficiency Multimedia Coding" (referred to as AVS2), primarily targeting the transmission of ultra-high-definition television programs. AVS2's coding efficiency was twice that of AVS+. In May 2016, AVS2 was released as a Chinese national standard. Simultaneously, the video portion of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for application. The AVS3 standard is a new generation of video coding standards for UHD video applications, designed to surpass the coding efficiency of the latest international standard, HEVC. In March 2019, at the 68th AVS meeting, the AVS3-P2 baseline was completed, offering approximately 30% bit rate savings compared to the HEVC standard. Currently, there is reference software called the High Performance Model (HPM), maintained by the AVS group, to demonstrate a reference implementation of the AVS3 standard.

[0099] Like HEVC, the AVS3 standard is built on a block-based hybrid video coding framework. Figure 5 A block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (called a coding unit (CU)). Unlike HEVC, which only segments blocks based on quadtrees, in AVS3, a coding tree unit (CTU) is segmented into CUs to accommodate the local characteristics of variations based on quadtrees / binary trees / extended quadtrees. Furthermore, the concept of multiple segmentation unit types in HEVC is removed; that is, the separation of CUs, prediction units (PUs), and transform units (TUs) does not exist in AVS3. Instead, each CU is always used as the basic unit for both prediction and transform without further segmentation. In the tree segmentation structure of AVS3, a CTU is first segmented based on a quadtree structure. Then, the leaf nodes of each quadtree can be further segmented based on binary tree and extended quadtree structures. Figure 5In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is sent to identify which reference picture in the reference picture memory the temporal prediction signal originates from. After spatial and / or temporal prediction, a mode decision block in the encoder selects the optimal prediction mode, for example, based on a rate distortion optimization method. The prediction block is then subtracted from the current video block; and the prediction residual is decorrelated using a transform and then quantized. The quantization residual coefficients are inversely quantized and inversely transformed to form the reconstruction residuals, which are then added back to the prediction block to form the reconstructed signal of the CU. Further loop filtering, such as deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF), can be applied to the reconstructed CU before it is stored in the reference image and used as a reference for encoding future video blocks. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantization residual coefficients are sent to the entropy coding unit for further compression and packing.

[0100] The first version of the HEVC standard was completed in October 2013, offering approximately 50% bitrate savings or equivalent perceptual quality compared to its predecessor, H.264 / MPEG AVC. Despite the significant coding improvements offered by HEVC, evidence suggests that superior coding efficiency can be achieved using additional coding tools on top of HEVC. Based on this, both VCEG and MPEG began exploring new coding technologies for future video coding standardization. In October 2015, ITU-TVCEG and ISO / IEC MPEG formed a Joint Video Exploration Group (JVET) to begin significant research into advanced technologies that could significantly improve coding efficiency. JVET maintains a reference software called the Joint Exploration Model (JEM) by integrating several additional coding tools on top of the HEVC Test Model (HM).

[0101] In October 2017, the ITU-T and ISO / IEC issued a joint call for proposals (CfP) on video compression with capabilities exceeding HEVC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating a compression efficiency gain of approximately 40% compared to HEVC. Based on these evaluation results, JVET launched a new project to develop a next-generation video coding standard known as Universal Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0102] Like HEVC, VVC is built on a block-based hybrid video coding framework. Figure 5 A block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (called a coding unit (CU)). In VTM-1.0, the CU can be up to 128×128 pixels. However, unlike HEVC, which is based solely on quadtree-based block partitioning, in VVC, a coding tree unit (CTU) is partitioned into CUs to accommodate the local characteristics of variations based on quadtree / binary / tritree. Furthermore, the concept of multiple partitioning unit types in HEVC is removed; that is, the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as the basic unit for both prediction and transform without further partitioning. In the multi-type tree structure, a CTU is first partitioned using a quadtree structure. Then, each quadtree leaf node can be further partitioned using binary and ternary tree structures. Figure 4E As shown, there are five types of segmentation: quadruple segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal triple segmentation, and vertical triple segmentation. Figure 5In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is sent to identify which reference picture in the reference picture memory the temporal prediction signal originates from. After spatial and / or temporal prediction, a mode decision block in the encoder selects the optimal prediction mode, for example, based on a rate distortion optimization method. The prediction block is then subtracted from the current video block; and the prediction residual is decorrelated and quantized using a transform. The quantization residual coefficients are inversely quantized and inversely transformed to form the reconstruction residuals, which are then added back to the prediction block to form the reconstructed signal of the CU. Further loop filtering, such as deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF), can be applied to the reconstructed CU before it is placed into the reference image memory and used to encode future video blocks. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantization residual coefficients are sent to the entropy coding unit for further compression and packing to form the bitstream.

[0103] Figure 6 A general block diagram of a block-based video decoder is given. The video bitstream is first entropy-decoded at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (if intra-frame coded) or the temporal prediction unit (if inter-frame coded) to form prediction blocks. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual blocks. The prediction blocks and residual blocks are then added together. The reconstructed blocks may undergo further loop filtering before being stored in the reference picture memory. The reconstructed video in the reference picture memory is then sent to drive the display device, along with the video blocks used to predict future frames.

[0104] The main focus of this disclosure is on improved adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF). The relevant background information is described in detail in the following sections.

[0105] 1.1 ALF in VVC

[0106] 1.1.1 Filter Shape, Linear Filtering, and Adaptive Limiting

[0107] In VVC, ALF is applied to the output samples of SAO. For example... Figure 7 As shown, the luminance and chrominance components each support two filter shapes, 7 7. Rhombus shape and 5 5. Rhombus shape. In Figure 7 In this diagram, each square corresponds to a luminance or chrominance sample, and the central square corresponds to the current sample to be filtered. Filter coefficients are point-symmetric, and each integer filter coefficient is represented with 7-bit fractional precision. Furthermore, the sum of the coefficients of a filter equals 128, which is a fixed-point representation of 1.0 with 7-bit fractional precision. (1) The number of coefficients N is 13 and 7 for 7×7 and 5×5 filter shapes, respectively.

[0108] coordinate Filtered sample values ​​at By using the coefficients as follows Applied to reconstructing sample values And export: (2) ,in and Is with the first coefficients The coordinates of the corresponding reconstructed sample points. Due to the constraints in mathematical formula (1), mathematical formula (2) can be written as: (3) In VVC, the probability of cropping the difference between adjacent sample values ​​and the current sample to be filtered is added to mathematical equation (3), as shown below: (4) ,in (5) It is a limit index Determined coefficients The limiting parameters. Export as follows: (6) ,in BD It is the sample bit depth, and It can be 0, 1, 2 or 3.

[0109] 1.1.2 Adaptive Luminance Sub-block Level Filter

[0110] In VVC, sub-block level filters are adaptively applied only to the luminance component. Each 4 The 4-luminance blocks are classified based on their directionality and 2D Laplacian activity. First, the sample gradient values ​​are calculated in the horizontal, vertical, and two diagonal directions: . (7) Sub-block horizontal gradient based on sample point gradient Vertical gradient and the gradients of the two diagonals and Calculated as (8)

[0112] index and It refers to 4 The coordinates of the top left sample point in the 4-luminance block. From mathematical formula (8), it can be seen that the coverage target 4... 4 pieces of 10 The sum of the gradients of samples within a 10-level brightness window is used to classify the block. To reduce complexity, such as... Figure 8 As shown, only 10 are calculated. The gradient is displayed at every other sample point within a 10-point window. The gradient values ​​for other sample points are set to 0.

[0113] Secondly, in order to assign the directionality D, the ratio of the maximum and minimum values ​​of the horizontal and vertical gradients of the sub-blocks. (9) And the ratio of the maximum to the minimum of the diagonal gradients of the two sub-blocks. (10) With a set of thresholds and Compare them with each other: Step 1: If and ,but D It is set to 0.

[0114] Step 2: If Then, in step 3, the directionality is calculated. D Otherwise, calculate the directionality in step 4. D .

[0115] Step 3: If ,but D Set to 2, otherwise D It is set to 1.

[0116] Step 4: If ,but D Set to 4, otherwise D It is set to 3.

[0117] If not assigned in the previous steps D If the value is , then only the above will be executed. D Each sequential step in the calculation. Third, the activity value. A Calculated as >> (11) A is further mapped to the range of 0 to 4: ,in Finally, each of the four... The 4-luminance block is classified into one of 25 categories: (12) Each category can have its own assigned filter.

[0118] For each 4 Before filtering the 4 luminance blocks, geometric transformations such as 90-degree rotation, diagonal, or vertical flip are applied to the filter coefficients based on the sub-block gradient values ​​specified in Table 1. Figure 9 As shown.

[0119] Table 1 Geometric transformation based on sub-block gradient values

[0120] 1.1.3 Adaptive Coding Tree Block-Level Filter

[0121] Besides brightness 4 In addition to 4-block level filter adaptation, ALF also supports CTB level filter adaptation. The Luminance CTB can use one of the filter banks computed for the current slice or for already encoded slices. It can also use one of 16 offline-trained filter banks. In each Luminance CTB, which filter from the selected filter bank should be applied to each 4 × 4 block is determined by the class C computed for that block in mathematical formula (12).

[0122] Chroma is adaptively applied using only CTB-level filters. A maximum of eight filters can be used for the chroma components in a slice. Each CTB can select one of these filters.

[0123] 1.1.4 Syntax Design

[0124] Filter coefficients and limiting indexes are carried within the ALF APS. The ALF APS can include up to eight chromaticity filters and a luma filter bank with up to 25 filters. For each of the 25 luma categories, an index is also included. Having the same index Categories share the same filters. By merging different categories, the number of bits required to represent filter coefficients is reduced. The absolute values ​​of filter coefficients are represented using 0th-order exponential Golomb code, followed by sign bits for non-zero coefficients. When clipping is enabled, a 2-bit fixed-length code is also used to send the clipping index of each filter coefficient with the signal. The maximum storage required for ALF coefficients and clipping indices within an APS is 3480 bits. The decoder can use up to eight ALF APSs simultaneously.

[0125] The filter control syntax element includes two types of information. First, an ALF on / off flag is signaled at the sequence, picture, slice, and CTB levels. Chroma ALF can only be enabled at the picture and slice levels if Luminance ALF is enabled at the corresponding level. Second, if ALF is enabled at the picture, slice, and CTB levels, filter usage information is signaled at those levels. If all slices within a picture use the same APS, a reference ALF APS ID is encoded at the slice or picture level. Luminance components can reference up to 7 ALF APSs, and chroma components can reference 1 ALF APS. For Luminance CTB, an index indicating which ALF APS or offline-trained luminance filter bank to use is signaled. For Chroma CTB, the index indicates which filter in the referenced APS is used.

[0126] 1.1.5 Line buffer reduction

[0127] To reduce the storage requirements of ALF, VVC employs row buffer boundary handling. In VVC, the row buffer boundary is placed four luma samples and two chroma samples above the horizontal CTU boundary. When applying ALF to samples on one side of the row buffer boundary, samples on the other side of the row buffer boundary cannot be used.

[0128] 1.2 ALF in ECM

[0129] 1.2.1 Simplified Removal of ALF

[0130] ALF gradient subsampling and ALF virtual boundary processing have been removed. The block size used for classification has been reduced from 4×4 to 2×2. The filter size used for both luma and chroma (for sending ALF coefficients to the luma and chroma signals) has been increased to 9×9.

[0131] 1.2.2 ALF with a fixed filter

[0132] To filter the brightness samples, three different classifiers are used ( , and ) and three different filter banks ( , and ).Group and Includes fixed filters, with features for classifiers and Training coefficients. Transmitted via signal. The coefficients of the filter in the group. Which filter is used for a given sample is determined by the classifier used. The category assigned to the sample point Decide.

[0133] 1.2.3 Filtering

[0134] First, two 13×13 rhombic fixed filters are applied. and To derive two intermediate sample points and After that, Applied to , The adjacent samples and the samples before the deblocking filter (DBF) are used to derive the filtered samples. (13) ,in (in ]) is the neighboring sample points before DBF and the current sample point The difference in the amplitude limit between them (in () is an intermediate sample point Compared with the current sample points The difference in the amplitude limit between them (in () represents the neighboring samples before DBF and the current sample. The difference in amplitude limits between them. Here, The range of values ​​for "i" in the text is related to... The range of values ​​for "i" varies. This is used in the signal transmission filter coefficients. ,i =0,…,24. Figure 10 Presented in The filter shape.

[0135] 1.2.4 Classification

[0136] Based on directionality and activity , categorize Assign to each 2×2 block: (14) in Indicates directionality The total number.

[0137] In VVC, for example, a 1-D Laplacian operator is used to compute the horizontal, vertical, and two diagonal gradients for each sample. The sum of the gradients of the samples within a 4×4 window covering a 2×2 block of the target is used for the classifier. Furthermore, the sum of the gradients of the samples within the 12×12 window is used for the classifier. and The sum of the horizontal gradient, vertical gradient, and the two diagonal gradients are expressed as follows: , , and Directionality It is determined by comparing the following with a set of thresholds: , (15) For example, using thresholds 2 and 4.5 in VVC to derive directionality. .for and First, calculate the horizontal / vertical edge strength. and diagonal edge strength Using thresholds =[1.25, 1.5, 2, 3, 4.5, 8]. If ≤ [0], then the edge strength =0; otherwise, It is the largest integer such that >Th[ -1]. If ≤ [0], then the edge strength =0; otherwise, It is the largest integer such that > [ -1]. When > That is, when the horizontal / vertical edges are dominant, it is derived by using Table 2(a) Otherwise, the diagonal edges dominate, as derived using Table 2(b). .

[0138] Table 2. and arrive mapping

[0139] In order to obtain The sum of vertical and horizontal gradients Mapped to the range 0 to n, where n is for Equals 4, for and It equals 15.

[0140] In ALF_APS, up to four luminance filter groups are sent using signals, and each group can have up to 25 filters.

[0141] 1.2.5 Alternative 2×2 ALF Classifier

[0142] The classification in ALF is extended using an additional alternative classifier. For luminance filter banks transmitted via signaling, a signaling flag is used to indicate whether an alternative classifier is applied. Geometric transformations are not applied to alternative band-based classifiers. When applying a band-based classifier, the sum of sample values ​​for a 2×2 luminance block is first calculated. Then, the class index is calculated as follows: class_index = (sum 25) >> (sample bit depth + 2). (16)

[0143] 1.2.6 Residual-based classifier

[0144] The classification in ALF is expanded using a third classifier based on luminance residual sample values. For each 2 × 2 luminance block, the sum of the absolute values ​​of the residual samples in the adjacent 8 × 8 windows is calculated, and the class index is derived as follows: classIdx = sum >> (sample bit depth – 4).

[0145] The value of classIdx ranges from 0 to 24, the same as in ECM-8.0. It is used by a signal transmission classifier for each luminance filter bank in the APS.

[0146] 1.3 CCALF in VVC

[0147] 1.3.1 Filter Shape and Accuracy

[0148] CCALF uses luminance sample values ​​to refine chrominance sample values ​​within the ALF process. For example... Figure 11 As shown, the linear filtering operation takes the luminance sample values ​​as input and generates corrected values ​​for the chrominance sample values. The correction is applied to each chrominance component. , It is generated independently and can be represented by the following formula:

[0149] in It is the chromaticity component The location of the sample points. From The exported brightness sample point locations, It revolves around The filter supports offset. It is the chromaticity component The filter support region in the luminance. The luminance location is determined based on the spatial scaling factor between the luminance and chrominance planes. The sample values ​​in the luminance support region are also inputs to the ALF luminance level and correspond to the output of the SAO level. In some examples, chromaticity sample values ​​can be used to improve the luminance sample values ​​within the ALF process.

[0150] like Figure 12 As shown, the CCALF filter has a diamond shape. Figure 12 As shown, for a 4:2:0 video sequence, there is a chroma position type 0, which means that when the even-numbered columns of chroma samples and luminance samples are horizontally co-located and there is a vertical gap between the rows of luminance samples, the center of the rhombus is aligned with the position of the chroma sample.

[0151] Compared to regular ALF coefficients, CCALF coefficients offer greater flexibility because they do not enforce symmetric constraints. However, two constraints are enforced: 1) To maintain DC neutrality, the sum of the CCALF coefficient values ​​must be zero. Therefore, only seven of the eight CCALF coefficients need to be signaled in the bitstream and their positions derived at the decoder. The coefficient at that point.

[0152] 2) The absolute values ​​of CCALF coefficients are restricted to zero or integer powers of 2, specifically {0, 1, 2, 4, 8, 16, 32, 64}. This allows the implementation to use variable shift operations instead of CCALF multiplication if needed.

[0153] 1.3.2 Syntax Design

[0154] In the final VVC design, the maximum number of filters for each chroma component of the image is four. A different set of CCALF coefficients can be selected for each CTU of the chroma component. Similar to regular ALF coefficients, CCALF coefficients are signaled within an ALF APS. Each ALF APS contains up to four CCALF filters for each chroma component. While CCALF can be enabled at the sequence level, it can only be enabled if ALF is also enabled for that sequence. Similarly, CCALF can only be enabled at the image and slice levels if the luma ALF is enabled at the corresponding level.

[0155] 1.3.3 Reduced line buffer

[0156] As described in Section 3.1.5, the luma and chroma line buffer boundaries are four and two samples above the CTU boundary, respectively. For the 4:2:0 chroma format, this results in line buffer boundaries aligned for chroma and luma. However, for the 4:2:2 and 4:4:4 chroma formats, the chroma and luma line buffer boundaries are not aligned with each other. Due to this misalignment, CC-ALF is not applied to lines with three and four samples above the CTU boundary for the 4:2:2 and 4:4:4 chroma formats.

[0157] 1.4 CCALF in ECM

[0158] The CCALF process uses a linear filter to filter luminance sample values ​​and generates residual corrections for chrominance samples. Figure 13 The CCALF processing shown uses a large 25-tap filter. For a given slice, the encoder can collect and analyze the slice's statistics and send signals to up to 16 filters via the APS.

[0159] Figure 14 The illustration shows filter shapes for predicting signals or preceding SAO signals according to examples of this disclosure. In some embodiments, spatially neighboring pixels in the predicted signal are provided as inputs to additional ALF equations. Various filter shapes can be used to extract information from the predicted signal. For example, the filter shape can be as follows: Figure 14 The values ​​shown are 1×1, 3×3, or 5×5.

[0160] Figure 15 The diagram illustrates adjusted ALF filter shapes according to some examples of this disclosure. In some embodiments, it is provided to change the chroma ALF filter shape from a rhombus to a shape such as... Figure 15 The elongated cross shape shown is consistent with the shape of the brightness ALF filter.

[0161] Figures 17A to 17CSymmetrical sample filling for a luminance ALF filter according to an example of this disclosure is illustrated. In some embodiments, symmetrical sample filling is applied when the filter shape of an additional input whose center position is aligned with the sample to be filtered crosses a virtual boundary (line buffer boundary) or a picture (slice, tile) boundary. For example, assuming an in-line ALF filter uses residual samples as additional input, the filter shape of the fixed filter to be applied to the residual signal or the filter shape of the in-line filter directly applied to the residual signal is 7x7, and the filter shape of the residual signal whose center position is aligned with the sample to be filtered crosses a line buffer boundary, such as... Figures 17A-17C The diagram shows the symmetrical sample point filling, where... Mask the juxtaposed residual pixels of the samples to be filtered to These are the original residual samples. to These are the modified residual sample values; the thick line represents the row buffer boundary. Shaded samples indicate filled residual samples. In summary, through symmetrical sample filling, both the additional input samples that are not on the same boundary side as the sample to be filtered and the additional symmetrical input samples that are on the same boundary side as the sample to be filtered are modified symmetrically.

[0162] Although ALF and CCALF have been improved in ECM, there is still room for further improvement in their performance. According to mathematical formula (5), sample difference is trimmed before applying the ALF coefficients. Sample difference physically means adjacent samples. and center sample The similarity. In one or more examples illustrated in mathematical formula (13) in section 1.2.3 "Filtering", the sample difference can correspond to (in ])or (in When the sample difference is within a certain range, the pruning operation preserves information, and if the difference is outside that range, the information is discarded. However, the discarded information may still be meaningful and can be used to improve coding efficiency. Information or sample differences greater than a threshold can be scaled (reduced) to be combined with other ALF inputs. Figures 18A-18B The illustration is shown.

[0163] This disclosure provides methods for further improving the existing design of ALF to address the problems identified in the "Problem Statement" section. In general, the main features of the techniques proposed in this disclosure are summarized below.

[0164] 1) Before applying the ALF coefficients, divide the dynamic range of the input sample difference into N parts (N is an integer not less than 2). One or more parts are within the range [min, max], while another or some parts are outside this range. Apply M different ALF coefficients to the N parts.

[0165] 2) Each sample difference may have its own coefficient, of which at least one coefficient applies to the outer range portion.

[0166] 3) The N partitions can share the same coefficient. In some examples, an additional ALF coefficient is sent by signaling, and the two partitions below the lower limit and above the upper limit use the same coefficient.

[0167] 4) The division into two parts can be determined by sending an index indicating the range [min, max], and the index-to-range mapping table can be predefined. In one or more examples, the index can be a 2-bit index indicating the range [8, 32], [32, 128], or [128, 1024]. In one or more examples, the index can be a 2-bit index indicating the range [32, 128], [128, 512], or [512, 4096].

[0168] 5) The sample difference and the target filter sample can be the same color component, such as luminance ALF and chrominance ALF.

[0169] 6) The sample difference (luminance) and the target filtered sample (chrominance) can be different color components, such as CCALF.

[0170] Based on some examples, before applying the ALF coefficients, the input sample difference is divided into two parts: one part is within the range [min, max], and the other part is outside this range. The processor can then apply two different ALF coefficients to the two divided parts. On the encoder side, the processor can signal an additional ALF coefficient for the external part.

[0171] According to some examples, an additional coefficient is sent for each sample difference, or the same additional coefficient is shared for all sample differences outside the range.

[0172] Based on some examples, the division into two parts is determined by sending an index indicating the range [min, max], and the index-to-range mapping table can be predefined.

[0173] Based on some examples, before applying the ALF coefficients, the input sample difference is divided into three parts: one part is within the range [min, max], one part is below the lower limit, and another part is above the upper limit. The processor can apply three different ALF coefficients to the three divided parts. On the encoder side, the processor can signal two additional ALF coefficients for the external part.

[0174] According to some examples, two additional coefficients are signaled for each sample difference, or the two additional coefficients are shared for all sample differences below the lower limit or above the upper limit.

[0175] According to some examples, a processor can determine the division into 3 parts by sending an index indicating the range [min, max] using a signal; according to some examples, the index-to-range mapping table can be predefined.

[0176] Based on some examples, the sample difference and the target filtered sample in the above examples can be the same color component, such as luminance ALF and chrominance ALF.

[0177] Based on some examples, the sample difference (luminance) and the target filtered sample (chrominance) in the above examples can be different color components, such as CCALF.

[0178] In some examples, the disclosed methods can be applied independently or in combination.

[0179] Figure 19 This is a flowchart illustrating a method for video decoding according to an example of this disclosure.

[0180] In step 1910, processor 1620 may receive at least one non-zero filter coefficient on the decoder side. In some examples, the non-zero filter coefficient may indicate a non-zero scaling weight to be applied to the input value. In one or more examples, the non-zero filter coefficient or non-zero scaling weight may be a preset value in the range (0, 0.25], or more specifically in the range (0, 0.2], or most specifically in the range [0.05, 0.15]. Figure 18B In at least one example shown, at least one non-zero filter coefficient may include = 0.1.

[0181] In step 1920, processor 1620 may apply at least one non-zero filter coefficient to at least one range of the dynamic range of the adaptive loop filter (ALF) input values ​​on the decoder side. In some examples, the dynamic range refers to the range of input values ​​used by the ALF, and the input values ​​may be represented as sample differences. In one or more examples, for a positive integer n, at least one range of the dynamic range of the ALF input values ​​may be [ or( In one or more examples, at least one range of the dynamic range of the ALF input values ​​may include two ranges of the ALF input values. To apply at least one non-zero filter coefficient to said at least one range, processor 1620 may apply the at least one non-zero filter coefficient as a non-zero scaling weight to both ranges of the ALF input values, such that the ALF input values ​​contained in both ranges are scaled down according to the scaling weight received in step 1910 for later use. Figure 18B In at least one example shown, the two ranges of the ALF values ​​are [-64, -32) and (32, 64], and at least one non-zero filter coefficient includes 0.1.

[0182] In step 1930, processor 1620 can filter the reconstructed video block corresponding to the current block on the decoder side based on the result of applying at least one non-zero filter coefficient to the at least one range. In some examples, the reconstructed block can be obtained by applying an inverse quantization unit and an inverse transform unit to the current block. In one or more examples, to filter the reconstructed video block, processor 1620 can perform calculations similar to those in section "1.2.3 Filtering" above. In one or more examples, when calculating... or When the "cropped" difference between adjacent samples and the current sample R(x, y) is used, instead of cropping (i.e., applying weight 0) the ALF input values ​​contained in the at least one range, the ALF input values ​​scaled down by applying the scaling weights as described in step 1920 above are used. Figure 18B In at least one example shown, at least one range of ALF values ​​includes [-64, -32) and (32, 64], and at least one non-zero filter coefficient includes 0.1. In at least one example as shown in mathematical formula (13), for "i" is in the range of 0 to 19 and "j" is in the range of 0 to 1; and for "i" is in the range of 22 to 24 and "j" is in the range of 0 to 1.

[0183] In some examples, the processor may further receive on the decoder side an index indicating a lower limit and an upper limit that define an intermediate range between the lower limit and the upper limit, wherein the at least one range includes at least one of a first range that is below the lower limit of the intermediate range or a second range that is above the upper limit of the intermediate range; and wherein the first range and the second range are not adjacent.

[0184] In one or more examples, in step 1920, processor 1620 may perform at least one of the following actions: applying at least one non-zero filter coefficient to the first range in response to determining that the at least one range includes a first range; or applying at least one non-zero filter coefficient to the second range in response to determining that the at least one range includes a second range. Figure 18B In at least one example shown, the intermediate range is [-32, 32], the first range is [-64, -32), and the second range is (32, 64).

[0185] In some examples, the ALF input value is represented as a sample difference, which is calculated using at least one neighboring sample and a target filtered sample, wherein the at least one neighboring sample is obtained from at least one neighboring block of the current block, and the target filtered sample corresponds to the current block.

[0186] In one or more examples, the sample difference is calculated using the color components of at least one neighboring sample of the same kind as the color component of the target filtered sample; the processor 1620 may further perform the following actions on the decoder side: in response to calculating the sample difference using the luminance components of at least one neighboring sample, apply a luminance ALF filter to the luminance component of the target filtered sample based on the sample difference; or in response to calculating the sample difference using the chrominance components of at least one neighboring sample, apply a chrominance ALF filter to the chrominance component of the target filtered sample based on the sample difference.

[0187] In one or more examples, the sample difference is calculated using the color components of at least one neighboring sample of a different type than the color component of the target filtered sample; the processor 1620 may further perform the following actions on the decoder side: in response to calculating the sample difference using the luminance components of at least one neighboring sample, applying an ALF filter to the chrominance component of the target filtered sample based on the sample difference; or in response to calculating the sample difference using the chrominance components of at least one neighboring sample, applying an ALF filter to the luminance component of the target filtered sample based on the sample difference. In one or more examples, in response to calculating the sample difference using the luminance or chrominance components of at least one neighboring sample, the ALF filter is a cross-component ALF (CCALF) filter.

[0188] In one or more examples, the ALF input value is represented as sample difference, the at least one range includes both a first range and a second range, at least one non-zero filter coefficient includes a shared non-zero filter coefficient, and in step 1920, the processor 1620 may apply the shared non-zero filter coefficient to both the first range and the second range on the decoder side.

[0189] In some examples, the at least one range includes both a first range and a second range, and at least one non-zero filter coefficient includes a first non-zero filter coefficient and a second non-zero filter coefficient; and in step 1920, the processor 1620 may apply the first non-zero filter coefficient to the first range and apply the second non-zero filter coefficient to the second range on the decoder side.

[0190] In some examples, in step 1920, processor 1620 may apply at least one predefined constant as at least one non-zero filter coefficient to the at least one range on the decoder side.

[0191] In some examples, an intermediate range is defined between the lower and upper limits based on a predefined index-to-range mapping table, wherein the at least one range includes at least one of a first range that is below the lower limit of the intermediate range or a second range that is above the upper limit of the intermediate range; and wherein the first and second ranges are not adjacent.

[0192] In one or more examples, the processor 1620 may further apply intermediate-range non-zero filter coefficients to the intermediate range on the decoder side.

[0193] Figure 20 This illustrates some examples corresponding to, for example, this disclosure. Figure 19 The flowchart shown is for a method used for video decoding and a method used for video encoding.

[0194] In step 2010, the processor 1620 may signal at least one non-zero filter coefficient on the encoder side. In some examples, the non-zero filter coefficient may indicate a non-zero scaling weight to be applied to the input value. In one or more examples, the non-zero filter coefficient or non-zero scaling weight may be a preset value in the range (0, 0.25], or more specifically in the range (0, 0.2], or most specifically in the range [0.05, 0.15]. Figure 18B In at least one example shown, at least one non-zero filter coefficient may include = 0.1.

[0195] In step 2020, the processor 1620 may apply at least one non-zero filter coefficient on the encoder side to at least one range of the dynamic range of the adaptive loop filter (ALF) input values. In some examples, the dynamic range refers to the range of input values ​​used by the ALF, and the input values ​​may be represented as sample differences. In one or more examples, for a positive integer n, at least one range of the dynamic range of the ALF input values ​​may be [ or( In one or more examples, at least one range of the dynamic range of the ALF input values ​​may include two ranges of the ALF input values. To apply at least one non-zero filter coefficient to said at least one range, processor 1620 may apply the at least one non-zero filter coefficient as a non-zero scaling weight to both ranges of the ALF input values, such that the ALF input values ​​contained in both ranges are scaled down according to the non-zero scaling weight received in step 2010 for later use. Figure 18B In at least one example shown, the two ranges of the ALF values ​​are [-64, -32) and (32, 64], and at least one non-zero filter coefficient includes 0.1.

[0196] In step 2030, processor 1620 can filter the reconstructed video block corresponding to the current block on the encoder side based on the result of applying at least one non-zero filter coefficient to the at least one range. In some examples, the reconstructed block can be obtained by applying an inverse quantization unit and an inverse transform unit to the current block. In one or more examples, to filter the reconstructed video block, processor 1620 can perform calculations similar to those in section "1.2.3 Filtering" above. In one or more examples, when calculating... or When the "cropped" difference between adjacent samples and the current sample R(x, y) is used, instead of cropping (i.e., applying weight 0 to) the ALF input values ​​contained in the at least one range, the ALF input values ​​scaled down by applying non-zero scaling weights as in step 2020 are used. Figure 18B In at least one example shown, at least one range of ALF values ​​includes [-64, -32) and (32, 64], and at least one non-zero filter coefficient includes 0.1. In at least one example as shown in mathematical formula (13), for "i" is in the range of 0 to 19 and "j" is in the range of 0 to 1; and for "i" is in the range of 22 to 24 and "j" is in the range of 0 to 1.

[0197] In step 2040, the processor 1620 can generate a bitstream on the encoder side based on the result of filtering the reconstructed video blocks.

[0198] In some examples, the processor may further signal on the encoder side an index indicating the lower and upper limits of an intermediate range between the lower and upper limits, wherein the at least one range includes at least one of a first range below the lower limit of the intermediate range or a second range above the upper limit of the intermediate range; and wherein the first and second ranges are not adjacent.

[0199] In one or more examples, in step 2020, processor 1620 may perform at least one of the following actions: applying at least one non-zero filter coefficient to the first range in response to determining that the at least one range includes a first range; or applying at least one non-zero filter coefficient to the second range in response to determining that the at least one range includes a second range. Figure 18B In at least one example shown, the intermediate range is [-32, 32], the first range is [-64, -32), and the second range is (32, 64).

[0200] In some examples, the ALF input value is represented as a sample difference, which is calculated using at least one neighboring sample and a target filtered sample, wherein the at least one neighboring sample is obtained from at least one neighboring block of the current block, and the target filtered sample corresponds to the current block.

[0201] In one or more examples, the sample difference is calculated using the color components of at least one adjacent sample of the same kind as the color component of the target filtered sample; the processor 1620 may further perform the following actions on the encoder side: in response to calculating the sample difference using the luminance components of at least one adjacent sample, apply a luminance ALF filter to the luminance component of the target filtered sample based on the sample difference; or in response to calculating the sample difference using the chrominance components of at least one adjacent sample, apply a chrominance ALF filter to the chrominance component of the target filtered sample based on the sample difference.

[0202] In one or more examples, the sample difference is calculated using the color components of at least one neighboring sample of a different type than the color component of the target filtered sample; the processor 1620 may further perform the following actions on the encoder side: in response to calculating the sample difference using the luminance components of at least one neighboring sample, applying an ALF filter to the chrominance component of the target filtered sample based on the sample difference; or in response to calculating the sample difference using the chrominance components of at least one neighboring sample, applying an ALF filter to the luminance component of the target filtered sample based on the sample difference. In one or more examples, in response to calculating the sample difference using the luminance or chrominance components of at least one neighboring sample, the ALF filter is a cross-component ALF (CCALF) filter.

[0203] In one or more examples, the ALF input value is represented as a sample difference, the at least one range includes both a first range and a second range, at least one non-zero filter coefficient includes a shared non-zero filter coefficient, and in step 2020, the processor 1620 may apply the shared non-zero filter coefficient to both the first range and the second range on the encoder side.

[0204] In some examples, the at least one range includes both a first range and a second range, and at least one non-zero filter coefficient includes a first non-zero filter coefficient and a second non-zero filter coefficient; and in step 2020, the processor 1620 may apply the first non-zero filter coefficient to the first range and apply the second non-zero filter coefficient to the second range on the encoder side.

[0205] In some examples, in step 2020, the processor 1620 may apply at least one predefined constant as at least one non-zero filter coefficient to the at least one range on the encoder side.

[0206] In some examples, an intermediate range is defined between the lower and upper limits based on a predefined index-to-range mapping table, wherein the at least one range includes at least one of a first range that is below the lower limit of the intermediate range or a second range that is above the upper limit of the intermediate range; and wherein the first and second ranges are not adjacent.

[0207] In one or more examples, the processor 1620 may further apply intermediate range non-zero filter coefficients to the intermediate range on the encoder side.

[0208] In some examples, an apparatus for video encoding is provided. The apparatus includes: a processor 1620; and a memory 1640 configured to store instructions executable by the processor; wherein the processor, when executing the instructions, is configured to perform actions such as... Figures 19-20 Any of the methods shown.

[0209] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a plurality of programs in memory 1630 and / or a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. The plurality of programs can be executed by processor 1620 in computing environment 1610 to perform the above-described methods. In one example, the plurality of programs can be executed by processor 1620 in computing environment 1610 to (e.g., from...) Figure 2 The video encoder 20 in the computing environment 1610 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1620 in the computing environment 1610 to perform the above-described decoding method based on the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor 1620 in the computing environment 1610 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 1620 in the computing environment 1610 to (e.g., to...) Figure 3The video decoder 30 in the middle sends the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by the encoder (e.g., Figure 2 The video encoder 20 in the video encoder uses, for example, the encoding method described above to generate the video for the decoder (e.g., Figure 3 The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc.

[0210] In one embodiment, a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method is provided.

[0211] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 1620); and a non-transitory computer-readable storage medium or memory 1630 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.

[0212] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above. In another embodiment, a computer program product including, for example, a plurality of programs in a memory 1630, the plurality of programs being executable by a processor 1620 in a computing environment 1610 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0213] In an embodiment, the computing environment 1610 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.

[0214] In one embodiment, a method for storing a bitstream is also provided, comprising: storing the bitstream on a non-transitory computer-readable storage medium or a digital storage medium, wherein the bitstream includes encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method.

[0215] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream to be decoded by the decoder described above is also provided.

[0216] The description in this disclosure has been presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.

[0217] Unless otherwise specifically stated, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.

[0218] The examples chosen and described are intended to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it will be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.

Claims

1. A method of video decoding, comprising: receiving, by a decoder, at least one non-zero filter coefficient; applying, by the decoder, the at least one non-zero filter coefficient to at least one range in a dynamic range of adaptive loop filter (ALF) input values; and filtering, by the decoder, a reconstructed video block corresponding to a current block based on a result of applying the at least one non-zero filter coefficient to the at least one range.

2. The method of claim 1, further comprising: receiving an index indicating a lower bound and an upper bound defining an intermediate range therebetween; wherein the at least one range comprises at least one of a first range below the lower bound of the intermediate range or a second range above the upper bound of the intermediate range; and wherein the first range and the second range are not adjacent.

3. The method of claim 2, wherein, applying the at least one non-zero filter coefficient to the at least one range comprises at least one of: in response to determining that the at least one range comprises the first range, applying the at least one non-zero filter coefficient to the first range; or in response to determining that the at least one range comprises the second range, applying the at least one non-zero filter coefficient to the second range.

4. The method of claim 1, wherein, the ALF input values are expressed as sample differences; wherein the sample differences are computed using at least one neighboring sample and a target filter sample; and wherein the at least one neighboring sample is obtained from at least one neighboring block of the current block and the target filter sample corresponds to the current block.

5. The method of claim 4, wherein, the sample differences are computed using a color component of the at least one neighboring sample that is the same kind as a color component of the target filter sample; and wherein the method further comprises: in response to computing the sample differences using a luma component of the at least one neighboring sample, applying a luma ALF filter to a luma component of the target filter sample based on the sample differences; or in response to computing the sample differences using a chroma component of the at least one neighboring sample, applying a chroma ALF filter to a chroma component of the target filter sample based on the sample differences.

6. The method of claim 4, wherein, the sample differences are computed using a color component of the at least one neighboring sample that is a different kind than a color component of the target filter sample; and wherein the method further comprises: in response to computing the sample differences using a luma component of the at least one neighboring sample, applying an ALF filter to a chroma component of the target filter sample based on the sample differences; or in response to computing the sample differences using a chroma component of the at least one neighboring sample, applying the ALF filter to a luma component of the target filter sample based on the sample differences.

7. The method of claim 6, wherein, the ALF filter is a cross-component ALF (CCALF) filter using the luma component or the chroma component of the at least one neighboring sample.

8. The method of claim 2, wherein, the ALF input values are expressed as sample differences; wherein the at least one range comprises both the first range and the second range and the at least one non-zero filter coefficient comprises a shared non-zero filter coefficient; and wherein applying the at least one non-zero filter coefficient to the at least one range comprises: applying the shared non-zero filter coefficient to both the first range and the second range.

9. The method of claim 1, wherein, the at least one range comprises both the first range and the second range, and the at least one non-zero filter coefficient comprises a first non-zero filter coefficient and a second non-zero filter coefficient; and wherein applying the at least one non-zero filter coefficient to the at least one range comprises: applying the first non-zero filter coefficient to the first range and the second non-zero filter coefficient to the second range.

10. The method of claim 1, wherein, applying the at least one non-zero filter coefficient to the at least one range comprises: applying at least one predefined non-zero constant as the at least one non-zero filter coefficient to the at least one range.

11. The method of claim 1, wherein, defining an intermediate range between a lower limit and an upper limit based on a predefined index-to-range mapping table; wherein the at least one range comprises at least one of a first range below the lower limit of the intermediate range or a second range above the upper limit of the intermediate range; and wherein the first range and the second range are not adjacent.

12. The method of claim 2, further comprising: applying an intermediate range non-zero filter coefficient to the intermediate range.

13. A video encoding method, comprising: signaling, by an encoder, at least one non-zero filter coefficient; applying, by the encoder, the at least one non-zero filter coefficient to at least one range in a dynamic range of an adaptive loop filter (ALF) input value; and filtering, by the encoder, a reconstructed video block corresponding to a current block based on a result of applying the at least one non-zero filter coefficient to the at least one range; and generating, by the encoder, a bitstream based on a result of filtering the reconstructed video block.

14. The method of claim 13, further comprising: signaling an index indicating a lower limit and an upper limit defining an intermediate range between the lower limit and the upper limit; wherein the at least one range comprises at least one of a first range below the lower limit of the intermediate range or a second range above the upper limit of the intermediate range; and wherein the first range and the second range are not adjacent.

15. The method of claim 14, wherein, applying the at least one non-zero filter coefficient to the at least one range comprises at least one of: in response to determining that the at least one range comprises the first range, applying the at least one non-zero filter coefficient to the first range; or in response to determining that the at least one range comprises the second range, applying the at least one non-zero filter coefficient to the second range.

16. The method of claim 13, wherein, the ALF input value is represented as a sample difference; wherein the sample difference is calculated using at least one neighboring sample and a target filter sample; and wherein the at least one neighboring sample is obtained from at least one neighboring block of the current block, and the target filter sample corresponds to the current block.

17. The method of claim 16, wherein, calculating the sample difference using a color component of the at least one neighboring sample that is of a same category as a color component of the target filter sample; and wherein the method further comprises: in response to calculating the sample difference using a luma component of the at least one neighboring sample, applying a luma ALF filter to a luma component of the target filter sample based on the sample difference; or in response to calculating the sample difference using a chroma component of the at least one neighboring sample, applying a chroma ALF filter to a chroma component of the target filter sample based on the sample difference.

18. The method of claim 16, wherein, calculating the sample difference using a color component of the at least one neighboring sample that is of a different category than a color component of the target filter sample; and wherein the method further comprises: in response to calculating the sample difference using a luma component of the at least one neighboring sample, applying an ALF filter to a chroma component of the target filter sample based on the sample difference; or in response to calculating the sample difference using a chroma component of the at least one neighboring sample, applying the ALF filter to a luma component of the target filter sample based on the sample difference.

19. The method of claim 18, wherein, calculating the sample difference using the luma component or the chroma component of the at least one neighboring sample, the ALF filter being a cross-component ALF (CCALF) filter.

20. The method of claim 14, wherein, the ALF input value is expressed as a sample difference; wherein the at least one range includes both the first range and the second range, and the at least one non-zero filter coefficient includes a shared non-zero filter coefficient; and wherein applying the at least one non-zero filter coefficient to the at least one range comprises: applying the shared non-zero filter coefficient to both the first range and the second range.

21. The method of claim 13, wherein, the at least one range includes both the first range and the second range, and the at least one non-zero filter coefficient includes a first non-zero filter coefficient and a second non-zero filter coefficient; and wherein applying the at least one non-zero filter coefficient to the at least one range comprises: applying the first non-zero filter coefficient to the first range, and applying the second non-zero filter coefficient to the second range.

22. The method of claim 13, wherein, applying the at least one non-zero filter coefficient to the at least one range comprises: applying at least one predefined non-zero constant as the at least one non-zero filter coefficient to the at least one range.

23. The method of claim 13, wherein, defining an intermediate range between a lower bound and an upper bound based on a predefined index-to-range mapping table; wherein the at least one range includes at least one of a first range below the lower bound of the intermediate range or a second range above the upper bound of the intermediate range; and wherein the first range and the second range are not adjacent.

24. The method of claim 14, further comprising: applying an intermediate range non-zero filter coefficient to the intermediate range.

25. An apparatus for video decoding, comprising: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors, when executing the instructions, are configured to perform the method of any of claims 1-12.

26. An apparatus for video encoding, comprising: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors, when executing the instructions, are configured to perform the method of any of claims 13-24.

27. A non-transitory computer-readable storage medium for storing computer- executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any of claims 1-12.

28. A non-transitory computer-readable storage medium for storing computer- executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any of claims 13-24.

29. A non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method of any of claims 1-12.

30. A non-transitory computer-readable storage medium for storing a bitstream generated by the method of any of claims 13-24.

31. A method for storing a bitstream, comprising storing the bitstream on a non-transitory computer-readable storage medium, wherein, The bitstream comprises encoded video information generated by the method of any of claims 13-24.

32. A method for storing a bitstream, comprising storing the bitstream on a non-transitory computer-readable storage medium, wherein, The bitstream comprises encoded video information to be decoded by the method of any of claims 1-12.

33. A method for transmitting a bitstream, comprising transmitting a bitstream generated by the method of any of claims 13-24.

34. A method for receiving a bitstream, comprising receiving a bitstream to be decoded by the method of any of claims 1-12.