Methods and apparatus for adaptive loop filters and cross-component adaptive loop filters

By employing adaptive loop filters and cross-component adaptive loop filter techniques, the problem of video quality degradation in existing video encoding and decoding technologies has been solved, achieving efficient video quality improvement under limited resource conditions.

CN122139360APending Publication Date: 2026-06-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2024-11-05
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies struggle to effectively utilize redundant information in video data during compression and decoding, leading to a decline in video quality. This is especially true under conditions of limited bandwidth and storage resources, where existing loop filters cannot adequately improve video quality.

Method used

By employing adaptive loop filter (ALF) and cross-component adaptive loop filter (CCSAO) technologies, the filter parameters are adaptively adjusted to perform personalized filtering processing on the luminance and chrominance components, thereby improving video quality.

Benefits of technology

It improves video quality during video encoding and decoding, especially under conditions of limited bandwidth and storage resources, effectively reducing block artifacts and blurring, and improving encoding efficiency and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122139360A_ABST
    Figure CN122139360A_ABST
Patent Text Reader

Abstract

A method and apparatus for video decoding / encoding are provided. In the provided method, the decoder can obtain an adaptive loop filter (ALF) classifier for sub-blocks, determine the ALF based on the ALF classifier, and obtain filtered chroma samples based on the ALF and reconstructed samples of chroma samples in the sub-blocks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to Provisional Application No. 63 / 547,541, filed November 6, 2023, the contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field

[0003] This application relates to video coding and compression. More specifically, this application relates to methods and apparatus for improving adaptive loop filtering processes and cross-component adaptive loop filtering processes. Background Technology

[0004] Various electronic devices (such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc.) support digital video. Electronic devices send and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited storage resources of storage devices, video data can be compressed using one or more video codec standards before it is transmitted or stored. For example, video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data to a form using a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention

[0005] Embodiments of this disclosure provide techniques related to adaptive loop filtering and cross-component adaptive loop filtering.

[0006] In a first aspect, some embodiments of this disclosure provide a method for video decoding. In this method, the decoder can obtain an adaptive loop filter (ALF) classifier for sub-blocks. Furthermore, the decoder can determine the ALF based on the ALF classifier. The decoder can then obtain filtered chroma samples based on the ALF and reconstructed samples of chroma samples in the sub-blocks.

[0007] In a second aspect, some embodiments of this disclosure provide a method for video coding. In this method, an encoder can obtain an adaptive loop filter (ALF) classifier for sub-blocks. Furthermore, the encoder can determine the ALF based on the ALF classifier. The encoder can then obtain filtered chroma samples based on the ALF and reconstructed samples of chroma samples in the sub-blocks.

[0008] In a third aspect of this disclosure, some embodiments of the present disclosure provide an apparatus for video decoding. The apparatus may include: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to perform the method according to the first aspect when executing the instructions.

[0009] In a fourth aspect of this disclosure, some embodiments of the present disclosure provide an apparatus for video encoding. The apparatus may include: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors. Furthermore, the one or more processors are configured to perform the method according to the second aspect when executing the instructions.

[0010] In a fifth aspect of this disclosure, some embodiments of this disclosure provide a non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the first aspect.

[0011] In a sixth aspect of this disclosure, some embodiments of this disclosure provide a non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to the second aspect.

[0012] In a seventh aspect of this disclosure, some embodiments of this disclosure provide a non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method according to the first aspect.

[0013] In the eighth aspect of this disclosure, some embodiments of this disclosure provide a non-transitory computer-readable storage medium for storing a bit stream generated by the method according to the second aspect.

[0014] In a ninth aspect of this disclosure, some embodiments of this disclosure provide a method for receiving a bitstream, wherein the bitstream includes encoded video information to be decoded by the method according to the first aspect.

[0015] In a tenth aspect of this disclosure, some embodiments of this disclosure provide a method for transmitting a bitstream, wherein the bitstream includes encoded video information generated by the method according to the second aspect.

[0016] It should be understood that the foregoing general description and the following detailed description are merely examples and not limitations of this disclosure. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate examples according to this disclosure and, together with this description, serve to explain the principles of this disclosure.

[0018] Figure 1 This is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some embodiments of the present disclosure.

[0019] Figure 2 This is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0020] Figure 3 This is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0021] Figures 4A to 4E This is a block diagram illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes according to some embodiments of this disclosure.

[0022] Figure 5 This is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0023] Figure 6 This is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0024] Figure 7 This is a diagram illustrating the shape of an ALF filter according to some examples of this disclosure.

[0025] Figure 8 It is a depiction of the gradient of subsampled points based on some examples of this disclosure.

[0026] Figure 9 This is a diagram illustrating the geometric transformation of the shape of a rhombus filter according to some examples of this disclosure.

[0027] Figure 10 This is a diagram illustrating the shape of an inline filter used in an ECM according to some examples of this disclosure.

[0028] Figure 11 This is a diagram of the CCALF architecture based on some examples of this disclosure.

[0029] Figure 12 This is a diagram illustrating the relative positions of filtered chroma samples in a 4:2:0 chroma format with chroma position type 0 and their support in the luminance plane.

[0030] Figure 13 This is a diagram of a 25-tap long filter based on some examples of this disclosure.

[0031] Figure 14 These are illustrations of various online ALF filter inputs based on some examples of this disclosure.

[0032] Figure 15 This is an illustration of the filter shape for predicting a signal or a signal before SAO, according to an example of this disclosure.

[0033] Figure 16 This is an illustration of an adjusted ALF filter shape based on some examples of this disclosure.

[0034] Figures 17A to 17C This is a diagram illustrating symmetrical sample filling of a luminance ALF filter according to an example of this disclosure.

[0035] Figure 18 This is a diagram illustrating a computing environment coupled to a user interface according to some embodiments of the present disclosure.

[0036] Figure 19 This is a flowchart illustrating a method for video decoding according to some examples of this disclosure.

[0037] Figure 20 This is a flowchart illustrating a method for video encoding according to some examples of this disclosure. Detailed Implementation

[0038] Referring now to the detailed description, examples of which are illustrated in the accompanying drawings. Numerous non-limiting details are set forth in the following detailed description to aid in understanding the subject matter presented herein. However, various alternatives may be used without departing from the scope of the claims, and the subject matter may be practiced without these specific details. For example, the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.

[0039] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this disclosure are used to distinguish objects and are not used to describe any specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in sequences other than those shown in the drawings or described in this disclosure.

[0040] Figure 1 This is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel, according to some embodiments of the present disclosure. Figure 1 As shown, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a target device 14. The source device 12 and target device 14 can include any electronic device from a wide variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and target device 14 are equipped with wireless communication capabilities.

[0041] In some implementations, target device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to target device 14. In embodiments, link 16 may include a communication medium enabling source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated according to communication standards (e.g., wireless communication protocols) and transmitted to target device 14. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include a router, switch, base station, or any other means that may facilitate communication from source device 12 to target device 14.

[0042] In some other implementations, encoded video data can be sent from output interface 22 to storage device 32. The target device 14 can then access the encoded video data in storage device 32 via input interface 28. Storage device 32 can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, digital universal discs (DVDs), compact disc read-only memory (CD-ROMs), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Target device 14 can access the stored video data from storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing and sending encoded video data to target device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Target device 14 can access the encoded video data via any standard data connection suitable for accessing encoded video data stored on the file server. Standard data connections include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or a combination of both. Transfer of encoded video data from storage device 32 can be streaming, downloading, or a combination of both.

[0043] like Figure 1 As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 may include sources or combinations of such sources, such as: a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if video source 18 is a camera in a security monitoring system, source device 12 and target device 14 may form a camera phone or video phone. However, the embodiments described in this application are generally applicable to video encoding and decoding and can be applied to wireless and / or wired applications.

[0044] Video captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be sent directly to target device 14 via output interface 22 of source device 12. Alternatively, the encoded video data can be stored on storage device 32 for later access by target device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or transmitter. The encoded video data may include a sequence of images, each of which may include one or more sample arrays, for example, luminance (Y) for monochromatic purposes only; luminance and two chrominances in the YCbCr or YCgCo domain; or green, blue, and red in the GBR (also known as RGB) domain. For ease of reference and terminology in this application, in some embodiments, the variables and terms associated with each set of three sample arrays may be referred to as luminance and chrominance, where the two chrominance arrays may be referred to as Cb and Cr, regardless of the actual color representation used. The video data may be in 4:0:0, 4:2:0, 4:2:2 or 4:4:4 chroma format, but this application is not limited to this.

[0045] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem, and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 when decoding the video data. Such syntax elements may be included within encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.

[0046] In some embodiments, the target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with the target device 14. The display device 34 displays decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0047] Video encoder 20 and video decoder 30 can operate according to proprietary or industry standards (e.g., VVC, HEVC, MPEG-4 Part 10, AVC) or extensions of such standards. It should be understood that this application is not limited to any particular video encoding / decoding standard and can be applied to other video encoding / decoding standards. It is generally understood that the video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally understood that the video decoder 30 of target device 14 can be configured to decode video data according to any of these current or future standards.

[0048] The video encoder 20 and video decoder 30 can be implemented as any circuit of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When partially implemented in software, the electronic device may store instructions for software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the video encoding / decoding operations disclosed in this disclosure. Each of the video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, and either encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.

[0049] In some implementations, components of source device 12 (e.g., video source 18, video encoder 20, or the following references) Figure 2 The components included in the video encoder 20 and the output interface 22) at least a portion of the components and / or the components of the target device 14 (e.g., the input interface 28, the video decoder 30 or the following references) Figure 3At least a portion of the components included in the video decoder 30 and the display device 34 can operate in a cloud computing service network such as Software as a Service (SaaS), Platform as a Service (PaaS), or Infrastructure as a Service (IaaS), where the cloud computing service network can provide software, platform, and / or infrastructure. In some embodiments, one or more components of the source device 12 and / or the target device 14 that are not included in the cloud computing service network can be located in one or more client devices, and these client devices can communicate with server computers in the cloud computing service network via wireless communication networks (e.g., cellular communication networks, short-range wireless communication networks, or Global Navigation Satellite System (GNSS) communication networks) or wired communication networks (e.g., local area network (LAN) communication networks or power line communication (PLC) networks). In one embodiment, at least a portion of the operations described herein can be implemented as a cloud-based service provided by one or more server computers, wherein the one or more server computers are implemented by at least a portion of the components of the source device 12 and / or at least a portion of the components of the target device 14 in the cloud computing service network; and one or more other operations described herein can be implemented by one or more client devices. In some implementations, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of this disclosure, terms such as “cloud,” “cloud computing,” and “cloud-based” may be used interchangeably as appropriate. It should be understood that this disclosure is not limited to implementation within the aforementioned cloud computing service networks. Instead, this disclosure may also be implemented in any other type of computing environment currently known or developed in the future.

[0050] Figure 2 This is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described in this application. The video encoder 20 can perform intra-frame predictive coding and inter-frame predictive coding on video blocks within a video frame. Intra-frame predictive coding relies on spatial prediction to reduce or remove spatial redundancy in the video data within a given video frame or picture. Inter-frame predictive coding relies on temporal prediction to reduce or remove temporal redundancy in the video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video encoding and decoding, the term "frame" can be used as a synonym for the terms "image" or "picture".

[0051] like Figure 2As shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a segmentation unit 45, an intra-frame prediction processing unit 46, and an intra-frame block copying (BC) unit 48. In some embodiments, the video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between the adder 62 and the DPB 64 to filter block boundaries to remove block artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross-component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of the adder 62. It should be noted that, regarding CCSAO technology, this application is not limited to the embodiments described herein, but can also be applied to situations where an offset is selected for either the luminance component and either of the two chrominance components (which may represent Y, Cb, and Cr in the YCbCr domain; Y, Cg, and Co in the YCgCo domain; or G, B, and R in the RGB domain, for the purposes of the symbols and terminology described above) to modify the other component based on the selected offset. Furthermore, it should be noted that the first component mentioned herein can be either the luminance component or either of the two chrominance components, the second component mentioned herein can be either the luminance component or either of the two chrominance components, and the third component mentioned herein can be the remaining component among the luminance component and the two chrominance components. In some examples, the loop filter may be omitted, and the decoded video block may be directly provided to the DPB 64 by the adder 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be distributed among one or more of the described fixed or programmable hardware units.

[0052] The video data storage device 40 can store video data encoded by the components of the video encoder 20. For example, it can store data from... Figure 1 The video source 18 shown receives video data from the video data memory 40. The DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) used by the video encoder 20 (e.g., in intra-frame or inter-frame predictive coding modes) when encoding the video data. The video data memory 40 and DPB 64 can be formed from any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip along with other components of the video encoder 20, or off-chip relative to those components.

[0053] like Figure 2 As shown, after receiving video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting the video frame into strips, tiles (e.g., a collection of video blocks) or other larger coding units (CUs) according to a predefined splitting structure associated with the video data (e.g., a quadtree (QT) structure). A video frame is, or can be considered, a two-dimensional array or matrix of sample points with sample values. Sample points in the array may also be referred to as pixels or image elements (pel). The number of sample points in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the video frame. For example, a video frame can be divided into multiple video blocks using QT segmentation. A video block is again, or can be considered, a two-dimensional array or matrix of sample points with sample values, but its dimension is smaller than that of the video frame. The number of sample points in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. A video block can be further divided into one or more block partitions or sub-blocks (which can then re-form blocks) by iteratively using, for example, QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof. It should be noted that the term “block” or “video block” as used herein can be a portion of a frame or image, particularly a rectangular (square or non-square) portion. Referring, for example, to HEVC and VVC, a block or video block can be or corresponds to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or corresponds to a corresponding block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB)) and / or sub-block.

[0054] The prediction processing unit 41 can select one of several feasible predictive coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one or more inter-frame predictive coding modes among multiple intra-frame predictive coding modes. The prediction processing unit 41 can provide the resulting intra-frame or inter-frame predictive coding block to adder 50 to generate a residual block, and to adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (e.g., motion vectors, intra-frame mode indicators, segmentation information, and other such syntax information) to entropy coding unit 56.

[0055] To select a suitable intra-predictive coding mode for the current video block, the intra-predictive processing unit 46 within the prediction processing unit 41 can perform intra-predictive coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-predictive coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 can perform multiple coding passes, for example, to select a suitable coding mode for each block of video data.

[0056] In some implementations, motion estimation unit 42 determines an inter-frame prediction mode for the current video frame by generating motion vectors based on a predetermined pattern within the video frame sequence. The motion vectors indicate the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame associated with the current block being encoded in the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra-frame BC unit 48 may determine vectors (e.g., block vectors) for intra-frame BC coding in a similar manner to how motion estimation unit 42 determines motion vectors for inter-frame prediction, or the block vectors may be determined using motion estimation unit 42.

[0057] Regarding pixel differences, the predicted block for a video block can be, or can correspond to, a block or reference block of a reference frame considered to closely match the video block to be encoded. Pixel differences can be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some implementations, the video encoder 20 can compute values ​​for sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 can interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 can perform motion search relative to full-pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0058] The motion estimation unit 42 calculates the motion vector for a video block in an inter-frame predictive coding frame by comparing the position of the video block with the position of the predicted block in a reference frame selected from either a first reference frame list (list 0) or a second reference frame list (list 1), each of the first and second reference frame lists identifying one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44, and then to the entropy coding unit 56.

[0059] Motion compensation performed by motion compensation unit 44 may involve acquiring or generating prediction blocks based on motion vectors determined by motion estimation unit 42. Upon receiving motion vectors for the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in a reference frame list within a reference frame list, retrieve the prediction block from DPB 64, and forward the prediction block to adder 50. Adder 50 then forms a residual video block of pixel differences by subtracting the pixel values ​​of the prediction block provided by motion compensation unit 44 from the pixel values ​​of the currently encoded video block. The pixel differences forming the residual video block may include a luminance component difference or a chrominance component difference, or both. Motion compensation unit 44 may also generate syntax elements associated with video blocks of a video frame for use by video decoder 30 when decoding video blocks of a video frame. Syntax elements may include, for example, syntax elements defining motion vectors for identifying prediction blocks, any flags indicating prediction modes, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are described separately for conceptual purposes.

[0060] In some implementations, the intra-BC unit 48 can generate vectors and acquire prediction blocks in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44; however, these prediction blocks are in the same frame as the current block being encoded, and these vectors are referred to as block vectors rather than motion vectors. Specifically, the intra-BC unit 48 can determine the intra-prediction mode to be used for encoding the current block. In some examples, the intra-BC unit 48 can, for example, use various intra-prediction modes to encode the current block during individual encoding passes and test their performance through rate-distortion analysis. Next, the intra-BC unit 48 can select a suitable intra-prediction mode from the various tested intra-prediction modes to use and generate an intra-mode indicator accordingly. For example, the intra-BC unit 48 can use rate-distortion analysis to calculate rate-distortion values ​​for the various tested intra-prediction modes and select the intra-prediction mode with the best rate-distortion characteristics from the tested modes as the suitable intra-prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra-frame BC unit 48 can calculate the ratio based on the distortion and rate for various coded blocks to determine which intra-frame prediction mode exhibits the optimal rate-distortion value for the block.

[0061] In other examples, the intra-frame BC unit 48 may use, in whole or in part, the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra-frame BC prediction according to the embodiments described herein. In any case, for intra-frame block copying, in terms of pixel differences, the predicted block may be a block considered to closely match the block to be encoded, the pixel differences may be determined by SAD, SSD, or other difference metrics, and identifying the predicted block may include calculating values ​​for sub-integer pixel positions.

[0062] Regardless of whether the predicted block comes from the same frame predicted intra-frame or from different frames predicted inter-frame, the video encoder 20 can form a residual video block by subtracting the pixel values ​​of the predicted block from the pixel values ​​of the current video block being encoded. The pixel difference forming the residual video block can include both luma component difference and chroma component difference.

[0063] As an alternative to the inter-frame prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above, or the intra-block copy prediction performed by the intra-BC unit 48, the intra-prediction processing unit 46 can perform intra-frame prediction on the current video block. Specifically, the intra-prediction processing unit 46 can determine an intra-prediction mode for encoding the current block. To this end, the intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during individual encoding passes, and the intra-prediction processing unit 46 (or, in some examples, the mode selection unit) can select a suitable intra-prediction mode from the tested intra-prediction modes for use. The intra-prediction processing unit 46 can provide information indicating the intra-prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 can encode the information indicating the selected intra-prediction mode into the bitstream.

[0064] After prediction processing unit 41 determines the prediction block for the current video block via inter-frame prediction or intra-frame prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 uses a transform (e.g., discrete cosine transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients.

[0065] The transform processing unit 52 can send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can subsequently perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.

[0066] After quantization, the entropy coding unit 56 entropy codes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probabilistic interval segmented entropy (PIPE) coding, or another entropy coding method or technique. The encoded bitstream can then be sent to, for example,... Figure 1 The video decoder 30 shown, or archived in, for example Figure 1 The data is stored in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 can also entropy code the motion vectors and other syntax elements used for the current video frame being encoded.

[0067] The inverse quantization unit 58 and the inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual video blocks in the pixel domain for generating reference blocks to predict other video blocks. As noted above, the motion compensation unit 44 can generate motion-compensated prediction blocks from one or more reference blocks of frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction blocks to compute sub-integer pixel values ​​for use in motion estimation.

[0068] Adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block can then be used as a prediction block by intra-frame BC unit 48, motion estimation unit 42, and motion compensation unit 44 for inter-frame prediction of another video block in subsequent video frames.

[0069] Figure 3 This is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of this application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction unit 84, and an intra-frame BC unit 85. The video decoder 30 can perform operations in conjunction with the above. Figure 2 The decoding process described for the video encoder 20 is essentially the inverse of the encoding process. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra-frame prediction unit 84 can generate prediction data based on the intra-frame prediction mode indicator received from the entropy decoding unit 80.

[0070] In some examples, units of the video decoder 30 may be assigned tasks to perform embodiments of this application. Furthermore, in some examples, embodiments of this disclosure may be distributed across one or more units of the video decoder 30. For example, the intra-frame BC unit 85 may perform embodiments of this application individually or in combination with other units of the video decoder 30 (e.g., motion compensation unit 82, intra-frame prediction unit 84, and entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra-frame BC unit 85, and the functionality of the intra-frame BC unit 85 may be performed by other components of the prediction processing unit 81 (e.g., motion compensation unit 82).

[0071] Video data memory 79 can store video data, such as encoded video bitstreams, that will be decoded by other components of video decoder 30. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source (e.g., a camera), via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include an encoded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of video decoder 30 stores reference video data for use by video decoder 30 (e.g., in intra-frame or inter-frame predictive coding modes) when decoding video data. Video data memory 79 and DPB 92 can be formed of any memory device from a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, video data memory 79 and DPB 92 are shown in... Figure 3 The video data memory 79 and DPB 92 are depicted as two distinct components of the video decoder 30. However, in some embodiments, the video data memory 79 may be provided by the same memory device or a separate memory device. In some examples, the video data memory 79 may be on-chip along with other components of the video decoder 30, or off-chip relative to those components.

[0072] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of encoded video frames and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators, and other syntax elements to the prediction processing unit 81.

[0073] When a video frame is encoded as an intra-predictive coded (I) frame or used as an intra-coded prediction block in other types of frames, the intra-predictive unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-predictive mode transmitted by the signal and reference data from the previous decoded block of the current frame.

[0074] When a video frame is encoded as an inter-frame predictive coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the current video frame based on motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks can be generated from a reference frame within a reference frame list. The video decoder 30 can construct the reference frame list, i.e., list 0 and list 1, based on the reference frames stored in the DPB 92 using a default construction technique.

[0075] In some examples, when a video block is encoded according to the intra-frame BC mode described herein, the intra-frame BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block can be located within a reconstructed region of the same image as the current video block, as defined by the video encoder 20.

[0076] Motion compensation unit 82 and / or intra-frame prediction (BC) unit 85 determine prediction information for video blocks in the current video frame by parsing motion vectors and other syntax elements, and then use this prediction information to generate prediction blocks for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) used to encode video blocks in the video frame, the inter-frame prediction frame type (e.g., B or P), construction information for one or more reference frames in the reference frame list for the frame, motion vectors for each inter-frame prediction encoded video block in the frame, the inter-frame prediction state for each inter-frame prediction encoded video block in the frame, and other information for decoding video blocks in the current video frame.

[0077] Similarly, the intra-BC unit 85 can use some of the received syntax elements, such as flags, to determine which video blocks in the frame are predicted using the intra-BC mode, which video blocks in the frame are in the reconstruction region and should be stored in the DPB 92, the block vector for each intra-BC predicted video block in the frame, the intra-BC prediction state for each intra-BC predicted video block in the frame, and other information for decoding video blocks in the current video frame.

[0078] The motion compensation unit 82 can also perform interpolation using interpolation filters, such as those used by the video encoder 20 during encoding of video blocks, to calculate interpolated values ​​for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filters used by the video encoder 20 based on the received syntax elements and use these interpolation filters to generate the prediction block.

[0079] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy-decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.

[0080] After the motion compensation unit 82 or the intra-frame BC unit 85 generates a prediction block for the current video block based on vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra-frame BC unit 85. A loop filter 91 (e.g., a deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 for further processing of the decoded video block. In some examples, the loop filter 91 may be omitted, and the decoded video block may be directly provided to the DPB 92 by the adder 90. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a separate memory device may also store the decoded video for later presentation on a display device (e.g., ...). Figure 1 On the display device 34).

[0081] In a typical video coding process, a video sequence usually consists of an ordered set of frames or images. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of chrominance samples (Cb). SCr is a two-dimensional array of chrominance samples (Cr). In other instances, a frame may be monochrome and therefore consist of only a two-dimensional array of luma samples.

[0082] like Figure 4AAs shown, the video encoder 20 (or more specifically, the segmentation unit 45) generates an encoded representation of a frame by first segmenting the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set such that all CTUs in the video sequence have the same size, one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that this application is not limited to a specific size. Figure 4B As shown, each CTU may include a CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units in the coded pixel block and how the video sequence can be reconstructed at the video decoder 30, including inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vectors, and other parameters. In a monochrome image or an image with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding the samples of that coding tree block. The coding tree block may be an N×N sample block.

[0083] To achieve better performance, the video encoder 20 can recursively perform tree splitting on the coding tree blocks of the CTU, such as binary tree splitting, ternary tree splitting, quadtree splitting, or combinations thereof, and divide the CTU into smaller CUs. Figure 4C As described, the 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Of these four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs, 430 and CU 440, are further divided into four CUs with a block size of 8×8. Figure 4D Depicting as shown Figure 4C The final result of the CTU 400 partitioning process described in the figure is a quadtree data structure, where each leaf node of the quadtree corresponds to a CU of a corresponding size ranging from 32×32 to 8×8. Similar to... Figure 4B The CTU depicted in the image can include, for example, two corresponding coding blocks (CBs) of luminance and chrominance samples of the same frame size, as well as syntax elements for encoding the samples of the coding blocks. In monochrome images or images with three separate color planes, a CU can include a single coding block and a syntax structure for encoding the samples of the coding block. It should be noted that... Figure 4C and Figure 4DThe quadtree partitioning depicted is for illustrative purposes only, and a CTU can be split into multiple CUs based on quadtree / ternary / binary partitioning to adapt to varying local characteristics. In multi-type tree structures, a CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to binary and ternary tree structures. Figure 4E As shown, a coded block with width W and height H has five possible segmentation types: quad segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal triple segmentation, and vertical triple segmentation.

[0084] In some implementations, the video encoder 20 may further segment the coded blocks of the CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter-frame or intra-frame) is applied. The PU of the CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In a monochrome image or an image with three separate color planes, a PU may include a single PB and a syntax structure for predicting the PBs. The video encoder 20 may generate predicted luma blocks, predicted Cb blocks, and predicted Cr blocks for each PU of the CU, representing the luma PB, Cb PB, and Cr PB.

[0085] Video encoder 20 can use intra-frame prediction or inter-frame prediction to generate prediction blocks for the PU. If video encoder 20 uses intra-frame prediction to generate prediction blocks for the PU, then video encoder 20 can generate prediction blocks for the PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate prediction blocks for the PU, then video encoder 20 can generate prediction blocks for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0086] After the video encoder 20 generates a predicted luminance block, a predicted Cb block, and a predicted Cr block for one or more PUs of the CU, the video encoder 20 can generate a luminance residual block for the CU by subtracting the predicted luminance block of the CU from the original luminance coding block of the CU, such that each sample in the luminance residual block of the CU indicates the difference between a luminance sample in one of the predicted luminance blocks of the CU and a corresponding sample in the original luminance coding block of the CU. Similarly, the video encoder 20 can generate Cb residual blocks and Cr residual blocks for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0087] In addition, such as Figure 4CAs shown, the video encoder 20 can use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block of the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) sample block to which the same transform is applied. A TU of the CU can include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the samples of the transform block. Therefore, each TU of the CU can be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU can be a sub-block of the CU's luminance residual block. A Cb transform block can be a sub-block of the CU's Cb residual block. A Cr transform block can be a sub-block of the CU's Cr residual block. In a monochrome image or an image with three separate color planes, a TU can include a single transform block and syntax structures for transforming the samples of that transform block.

[0088] The video encoder 20 can apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block can be a two-dimensional array of transform coefficients. The transform coefficients can be scalars. The video encoder 20 can apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 can apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.

[0089] After generating coefficient blocks (e.g., luminance coefficient blocks, Cb coefficient blocks, or Cr coefficient blocks), video encoder 20 can quantize the coefficient blocks. Quantization typically refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After quantizing the coefficient blocks, video encoder 20 can entropy encode the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream comprising a bit sequence that forms a representation of coded frames and associated data; the bitstream is stored in storage device 32 or transmitted to target device 14.

[0090] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements. The video decoder 30 can reconstruct frames of video data, at least in part, based on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally the inverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the coded blocks of the current CU by adding the samples of the prediction block of the PU for the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the coded blocks for each CU of the frame, the video decoder 30 can reconstruct the frame.

[0091] As mentioned above, video coding primarily uses two modes (i.e., intra-frame prediction and inter-frame prediction) to achieve video compression. It should be noted that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes more to coding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block based on a reference video block.

[0092] However, with continuously improving video data capture technologies and finer video tile sizes used to preserve details in video data, the amount of data required to represent the motion vectors for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only do a set of neighboring CUs in both the spatial and temporal domains have similar video data for prediction purposes, but the motion vectors between these neighboring CUs are also similar. Therefore, the motion information of spatially neighboring CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (also known as the "motion vector predictor" (MVP) of the current CU) by exploring their spatial and temporal correlations.

[0093] Instead of the above combination Figure 2 The method described involves encoding the actual motion vector of the current CU, determined by the motion estimation unit 42, into the video bitstream, and subtracting the motion vector prediction factor of the current CU from the actual motion vector of the current CU to produce the motion vector difference (MVD) for the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.

[0094] Similar to the process of selecting a prediction block in a reference frame during inter-frame prediction of a coded block, both the video encoder 20 and the video decoder 30 need to employ a set of rules to construct a motion vector candidate list (also known as a "merging list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs. Then, a member is selected from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list for encoding and decoding the current CU.

[0095] This disclosure relates to video coding and compression. More specifically, this disclosure relates to methods and apparatus for improving the coding efficiency of adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF).

[0096] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, some well-known video coding standards today include Universal Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which were jointly developed by ISO / IEC MPEG and ITU-T TVCEG. AOMedia Video1 (AV1) was developed by the Alliance for Open Media (AOM) as the successor to its previous standard VP9. Audio and Video Coding (AVS) is another series of video compression standards developed by the China Audio and Video Coding Standards Working Group, which refers to digital audio and digital video compression standards. Most existing video coding standards are built on the well-known hybrid video coding framework, which uses block-based prediction methods (e.g., inter-frame prediction, intra-frame prediction) to reduce redundancy in video images or sequences and uses transform coding to compress the energy of prediction errors. A key goal of video coding technology is to compress video data into a form using a lower bit rate while avoiding or minimizing video quality degradation.

[0097] The first generation of AVS standards included the Chinese national standards "Information Technology, Advanced Audio-Video Coding, Part 2: Video" (referred to as AVS1) and "Information Technology, Advanced Audio-Video Coding, Part 16: Broadcast Television Video" (referred to as AVS+). Compared to the MPEG-2 standard, it offered approximately 50% bitrate savings while maintaining the same perceptual quality. The video portion of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second generation of AVS standards included a series of Chinese national standards, "Information Technology, High-Efficiency Multimedia Coding" (referred to as AVS2), primarily targeting the transmission of ultra-high-definition television programs. AVS2's coding efficiency was twice that of AVS+. In May 2016, AVS2 was released as a Chinese national standard. Simultaneously, the video portion of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for application. The AVS3 standard is a new generation of video coding standards for UHD video applications, designed to surpass the coding efficiency of the latest international standard, HEVC. In March 2019, at the 68th AVS meeting, the AVS3-P2 baseline was completed, offering approximately 30% bit rate savings compared to the HEVC standard. Currently, there is reference software called the High Performance Model (HPM), maintained by the AVS group, to demonstrate a reference implementation of the AVS3 standard.

[0098] Like HEVC, the AVS3 standard is built on a block-based hybrid video coding framework. Figure 5 A block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (called a coding unit (CU)). Unlike HEVC, which only segments blocks based on quadtrees, in AVS3, a coding tree unit (CTU) is segmented into CUs to accommodate the local characteristics of variations based on quadtrees / binary trees / extended quadtrees. Furthermore, the concept of multiple segmentation unit types in HEVC is removed; that is, the separation of CUs, prediction units (PUs), and transform units (TUs) does not exist in AVS3. Instead, each CU is always used as the basic unit for both prediction and transform without further segmentation. In the tree segmentation structure of AVS3, a CTU is first segmented based on a quadtree structure. Then, the leaf nodes of each quadtree can be further segmented based on binary tree and extended quadtree structures. Figure 5In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is sent to identify which reference picture in the reference picture memory the temporal prediction signal originates from. After spatial and / or temporal prediction, a mode decision block in the encoder selects the optimal prediction mode, for example, based on a rate distortion optimization method. The prediction block is then subtracted from the current video block; and the prediction residual is decorrelated using a transform and then quantized. The quantization residual coefficients are inversely quantized and inversely transformed to form the reconstruction residuals, which are then added back to the prediction block to form the reconstructed signal of the CU. Further loop filtering, such as deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF), can be applied to the reconstructed CU before it is stored in the reference image and used as a reference for encoding future video blocks. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantization residual coefficients are sent to the entropy coding unit for further compression and packing.

[0099] The first version of the HEVC standard was completed in October 2013, offering approximately 50% bitrate savings or equivalent perceptual quality compared to its predecessor, H.264 / MPEG AVC. Despite the significant coding improvements offered by HEVC, evidence suggests that superior coding efficiency can be achieved using additional coding tools on top of HEVC. Based on this, both VCEG and MPEG began exploring new coding technologies for future video coding standardization. In October 2015, ITU-TVCEG and ISO / IEC MPEG formed a Joint Video Exploration Group (JVET) to begin significant research into advanced technologies that could significantly improve coding efficiency. JVET maintains a reference software called the Joint Exploration Model (JEM) by integrating several additional coding tools on top of the HEVC Test Model (HM).

[0100] In October 2017, the ITU-T and ISO / IEC issued a joint call for proposals (CfP) on video compression with capabilities exceeding HEVC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating a compression efficiency gain of approximately 40% compared to HEVC. Based on these evaluation results, JVET launched a new project to develop a next-generation video coding standard known as Universal Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0101] Like HEVC, VVC is built on a block-based hybrid video coding framework. Figure 5 A block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (called a coding unit (CU)). In VTM-1.0, the CU can be up to 128×128 pixels. However, unlike HEVC, which is based solely on quadtree-based block partitioning, in VVC, a coding tree unit (CTU) is partitioned into CUs to accommodate the local characteristics of variations based on quadtree / binary / tritree. Furthermore, the concept of multiple partitioning unit types in HEVC is removed; that is, the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as the basic unit for both prediction and transform without further partitioning. In the multi-type tree structure, a CTU is first partitioned using a quadtree structure. Then, each quadtree leaf node can be further partitioned using binary and ternary tree structures. Figure 4E As shown, there are five types of segmentation: quadruple segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal triple segmentation, and vertical triple segmentation. Figure 5In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of already encoded neighboring blocks in the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is sent to identify which reference picture in the reference picture memory the temporal prediction signal originates from. After spatial and / or temporal prediction, a mode decision block in the encoder selects the optimal prediction mode, for example, based on a rate distortion optimization method. The prediction block is then subtracted from the current video block; and the prediction residual is decorrelated and quantized using a transform. The quantization residual coefficients are inversely quantized and inversely transformed to form the reconstruction residuals, which are then added back to the prediction block to form the reconstructed signal of the CU. Further loop filtering, such as deblocking filtering, sample adaptive offset (SAO), and adaptive loop filtering (ALF), can be applied to the reconstructed CU before it is placed into the reference image memory and used to encode future video blocks. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantization residual coefficients are sent to the entropy coding unit for further compression and packing to form the bitstream.

[0102] Figure 6 A general block diagram of a block-based video decoder is given. The video bitstream is first entropy-decoded at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (if intra-frame coded) or the temporal prediction unit (if inter-frame coded) to form prediction blocks. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual blocks. The prediction blocks and residual blocks are then added together. The reconstructed blocks may undergo further loop filtering before being stored in the reference picture memory. The reconstructed video in the reference picture memory is then sent to drive the display device, along with the video blocks used to predict future frames.

[0103] The main focus of this disclosure is on improving adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF). The relevant knowledge is described in detail in the following sections.

[0104] ALF in VVC

[0105] Filter shape, linear filtering and adaptive limiting

[0106] In VVC, ALF is applied to the output samples of SAO. For example... Figure 7 As shown, the luminance and chrominance components each support two filter shapes, 7 7. Rhombus shape and 5 5. Rhombus shape. In Figure 7 In this diagram, each square corresponds to a luminance or chrominance sample, and the central square corresponds to the current sample to be filtered. Filter coefficients are point-symmetric, and each integer filter coefficient is represented with 7-bit fractional precision. Furthermore, the sum of the coefficients of a filter equals 128, which is a fixed-point representation of 1.0 with 7-bit fractional precision. (1), The number of coefficients N is 13 for 7×7 and 7 for 5×5 filter shapes.

[0107] coordinate Filtered sample values ​​at By using the coefficients as follows Applied to reconstructing sample values And export: (2), in and Is with the first coefficients The coordinates of the corresponding reconstructed sample points. Due to the constraints in mathematical formula (1), mathematical formula (2) can be written as: (3).

[0108] In VVC, the probability of cropping the difference between adjacent sample values ​​and the current sample to be filtered is added to mathematical equation (3), as shown below: (4), in (5) It is a limit index Determined coefficients The limiting parameters. Export as follows: (6), in BD It is the sample bit depth, and It can be 0, 1, 2 or 3.

[0109] Luminance sub-block level filter adaptive

[0110] In VVC, sub-block level filters are adaptively applied only to the luminance component. Each 4 The 4-luminance blocks are classified based on their directionality and 2D Laplacian activity. First, the sample gradient values ​​are calculated in the horizontal, vertical, and two diagonal directions: (7) Based on the sample point gradient, the sub-block horizontal gradient Vertical gradient and the gradients of the two diagonals and Calculated as (8) index and It refers to 4 The coordinates of the top left sample point in the 4-luminance block. From mathematical formula (8), it can be seen that the coverage target 4... 4 pieces of 10 The sum of the gradients of samples within a 10-level brightness window is used to classify the block. To reduce complexity, such as... Figure 8 As shown, only 10 are calculated. The gradient is displayed at every other sample point within a 10-point window. The gradient values ​​for other sample points are set to 0.

[0111] Secondly, in order to assign the directionality D, the ratio of the maximum and minimum values ​​of the horizontal and vertical gradients of the sub-blocks. (9) And the ratio of the maximum to the minimum of the diagonal gradients of the two sub-blocks. (10) With a set of thresholds and Compare them with each other: Step 1: If and If so, then D is set to 0.

[0112] Step 2: If If the directionality D is positive, then calculate the directionality D in step 3; otherwise, calculate the directionality D in step 4.

[0113] Step 3: If If the condition is met, then D is set to 2; otherwise, D is set to 1.

[0114] Step 4: If If the condition is met, then D is set to 4; otherwise, D is set to 3.

[0115] If no value was assigned to D in the previous steps, then only each sequential step in the calculation of D described above is performed. Third, the active value A is calculated as... >> (11) A is further mapped to the range of 0 to 4: ,in Finally, each of the four... The 4-luminance block is classified into one of 25 categories: (12) Each category can have its own assigned filter.

[0116] For each 4 Before filtering the 4 luminance blocks, geometric transformations such as 90-degree rotation, diagonal, or vertical flip are applied to the filter coefficients based on the sub-block gradient values ​​specified in Table 1. Figure 9 As shown.

[0117] Table 1 Geometric transformation based on sub-block gradient values

[0118] Coding tree block-level filter adaptive

[0119] Besides brightness 4 In addition to 4-block level filter adaptation, ALF also supports CTB level filter adaptation. The Luminance CTB can use one of the filter banks computed for the current slice or for already encoded slices. It can also use one of 16 offline-trained filter banks. In each Luminance CTB, which filter from the selected filter bank should be applied to each 4 × 4 block is determined by the class C computed for that block in mathematical formula (12).

[0120] Chroma is adaptively applied using only CTB-level filters. A maximum of eight filters can be used for the chroma components in a slice. Each CTB can select one of these filters.

[0121] syntax design

[0122] Filter coefficients and limiting indexes are carried within the ALF APS. The ALF APS can include up to eight chromaticity filters and a luma filter bank with up to 25 filters. For each of the 25 luma categories, an index is also included. Having the same index Categories share the same filters. By merging different categories, the number of bits required to represent filter coefficients is reduced. The absolute values ​​of filter coefficients are represented using 0th-order exponential Golomb code, followed by sign bits for non-zero coefficients. When clipping is enabled, a 2-bit fixed-length code is also used to send the clipping index of each filter coefficient with the signal. The maximum storage required for ALF coefficients and clipping indices within an APS is 3480 bits. The decoder can use up to eight ALF APSs simultaneously.

[0123] The filter control syntax element includes two types of information. First, an ALF on / off flag is signaled at the sequence, picture, slice, and CTB levels. Chroma ALF can only be enabled at the picture and slice levels if Luminance ALF is enabled at the corresponding level. Second, if ALF is enabled at the picture, slice, and CTB levels, filter usage information is signaled at those levels. If all slices within a picture use the same APS, a reference ALF APS ID is encoded at the slice or picture level. Luminance components can reference up to 7 ALF APSs, and chroma components can reference 1 ALF APS. For Luminance CTB, an index indicating which ALF APS or offline-trained luminance filter bank to use is signaled. For Chroma CTB, the index indicates which filter in the referenced APS is used.

[0124] Line buffer reduction

[0125] To reduce the storage requirements of ALF, VVC employs row buffer boundary handling. In VVC, the row buffer boundary is placed four luma samples and two chroma samples above the horizontal CTU boundary. When applying ALF to samples on one side of the row buffer boundary, samples on the other side of the row buffer boundary cannot be used.

[0126] ALF in ECM

[0127] ALF simplified removal

[0128] ALF gradient subsampling and ALF virtual boundary processing have been removed. The block size used for classification has been reduced from 4×4 to 2×2. The filter size used for both luma and chroma (for sending ALF coefficients to the luma and chroma signals) has been increased to 9×9.

[0129] ALF with fixed filter

[0130] To filter the brightness samples, three different classifiers are used ( , and ) and three different filter banks ( , and ).Group and Includes fixed filters, with features for classifiers and Training coefficients. Transmitted via signal. The coefficients of the filter in the group. Which filter is used for a given sample is determined by the classifier used. The category assigned to the sample point Decide.

[0131] Filtering

[0132] First, two 13×13 rhombic fixed filters are applied. and To derive two intermediate sample points and After that, Applied to , The adjacent samples and the samples before the deblocking filter (DBF) are used to derive the filtered samples. (13), in It is the relationship between adjacent sample points and the current sample point. The difference in the amplitude limit between them yes Compared with the current sample points The difference in the amplitude limit between them It is the adjacent sample points before DBF and the current sample point. The amplitude difference between them. (Using signal transmission filter coefficients) , i=0,…,24. Figure 10 Presented in The filter shape.

[0133] Classification

[0134] Based on directionality and activity , categorize Assign to each 2×2 block: (14) in Indicates directionality The total number.

[0135] In VVC, for example, a 1-D Laplacian operator is used to compute the horizontal, vertical, and two diagonal gradients for each sample. The sum of the gradients of the samples within a 4×4 window covering a 2×2 block of the target is used for the classifier. Furthermore, the sum of the gradients of the samples within the 12×12 window is used for the classifier. and The sum of the horizontal gradient, vertical gradient, and the two diagonal gradients are expressed as follows: , , and Directionality It is determined by comparing the following with a set of thresholds: , (15) For example, using thresholds 2 and 4.5 in VVC to derive directionality. .for and First, calculate the horizontal / vertical edge strength. and diagonal edge strength Using thresholds =[1.25, 1.5, 2, 3, 4.5, 8]. If ≤ [0], then the edge strength =0; otherwise, It is the largest integer such that >Th[ -1]. If ≤ [0], then the edge strength =0; otherwise, It is the largest integer such that > [ -1]. When > That is, when the horizontal / vertical edges are dominant, it is derived by using Table 2(a) Otherwise, the diagonal edges dominate, as derived using Table 2(b). .

[0136] Table 2. and arrive mapping

[0137] In order to obtain The sum of vertical and horizontal gradients Mapped to the range 0 to n, where n is for Equals 4, for and It equals 15.

[0138] In ALF_APS, up to four luminance filter groups are sent using signals, and each group can have up to 25 filters.

[0139] Alternative 2×2 ALF classifier

[0140] The classification in ALF is extended using an additional alternative classifier. For luminance filter banks transmitted via signaling, a signaling flag is used to indicate whether an alternative classifier is applied. Geometric transformations are not applied to alternative strip classifiers. When applying a strip-based classifier, the sum of sample values ​​for a 2×2 luminance block is first calculated. Then, the class index is calculated as follows: class_index = (sum 25) >> (sample bit depth + 2) (16).

[0141] Residual-based classifiers

[0142] The classification in ALF is expanded using a third classifier based on luminance residual sample values. For each 2 × 2 luminance block, the sum of the absolute values ​​of the residual samples in the adjacent 8 × 8 windows is calculated, and the class index is derived as follows: classIdx = sum >> (sample bit depth – 4).

[0143] The value of classIdx ranges from 0 to 24, the same as in ECM-8.0. It is used by a signal transmission classifier for each luminance filter bank in the APS.

[0144] CCALF in VVC

[0145] Filter shape and accuracy

[0146] CCALF uses luminance sample values ​​to refine chrominance sample values ​​within the ALF process. For example... Figure 11 As shown, the linear filtering operation takes the luminance sample values ​​as input and generates corrected values ​​for the chrominance sample values. The correction is applied to each chrominance component. , Independently generated and can be represented by the following formula: in It is the chromaticity component The location of the sample points. From The exported brightness sample point locations, It revolves around The filter supports offset. It is the chromaticity component The filter support region in the luminance. The luminance location is determined based on the spatial scaling factor between the luminance and chrominance planes. The sample values ​​in the luminance support region are also the inputs to the ALF luminance level and correspond to the outputs of the SAO level.

[0147] like Figure 12 As shown, the CCALF filter has a diamond shape. Figure 12 As shown, for a 4:2:0 video sequence, there is a chroma position type 0, which means that when the even-numbered columns of chroma samples and luminance samples are horizontally co-located and there is a vertical gap between the rows of luminance samples, the center of the rhombus is aligned with the position of the chroma sample.

[0148] Compared to regular ALF coefficients, CCALF coefficients offer greater flexibility because they do not enforce symmetric constraints. However, two constraints are enforced: To maintain DC neutrality, the sum of the CCALF coefficient values ​​must be zero. Therefore, only seven of the eight CCALF coefficients need to be sent as signals in the bitstream and their positions derived at the decoder. The coefficient at that point.

[0149] The absolute values ​​of the CCALF coefficients are restricted to zero or integer powers of 2, specifically {0, 1, 2, 4, 8, 16, 32, 64}. This allows the implementation to use variable shift operations instead of CCALF multiplication if needed.

[0150] syntax design

[0151] In the final VVC design, the maximum number of filters for each chroma component of the image is four. A different set of CCALF coefficients can be selected for each CTU of the chroma component. Similar to regular ALF coefficients, CCALF coefficients are signaled within an ALF APS. Each ALF APS contains up to four CCALF filters for each chroma component. While CCALF can be enabled at the sequence level, it can only be enabled if ALF is also enabled for that sequence. Similarly, CCALF can only be enabled at the image and slice levels if the luma ALF is enabled at the corresponding level.

[0152] Line buffer reduction

[0153] As described in Section 3.1.5, the luma and chroma line buffer boundaries are four and two samples above the CTU boundary, respectively. For the 4:2:0 chroma format, this results in line buffer boundaries aligned for chroma and luma. However, for the 4:2:2 and 4:4:4 chroma formats, the chroma and luma line buffer boundaries are not aligned with each other. Due to this misalignment, CC-ALF is not applied to lines with three and four samples above the CTU boundary for the 4:2:2 and 4:4:4 chroma formats.

[0154] CCALF in ECM

[0155] The CCALF process uses a linear filter to filter luminance sample values ​​and generates residual corrections for chrominance samples. Figure 13 The CCALF processing shown uses a large 25-tap filter. For a given slice, the encoder can collect and analyze the slice's statistics and send signals to up to 16 filters via the APS.

[0156] Although ALF and CCALF have been improved in ECM, there is still room for further performance improvements.

[0157] First, the online ALF filter in ECM takes spatially adjacent pixels, the fixed ALF filter result, and the spatially adjacent pixels before the deblocking filter as input. However, in addition to this information, other information such as spatially adjacent pixels in the predicted signal, spatially adjacent pixels in the residual signal, or spatially adjacent pixels before SAO can also be used as input to the online ALF filter formula, which can benefit coding performance.

[0158] Secondly, edge-based and band-based classifiers are adaptively used for the in-line ALF filter in the ECM. However, these two classifiers can be further combined to provide other classifiers, which can benefit coding performance.

[0159] Third, the filter shape of the chroma ALF in the ECM is diamond-shaped, while the filter shape of the luma ALF is long cross-shaped. From a standardization point of view, this non-uniform design may not be optimal.

[0160] Fourth, edge-based and band-based classifiers in ECM only consider pixel values ​​after SAO. However, pixel values ​​from the following stages—1) immediately before the deblocking filter, 2) the predicted signal, 3) the residual signal, and 4) immediately before SAO—can also be used to design new classifiers, which can benefit coding performance.

[0161] Fifth, edge-based and band-based classifiers in ECM only consider the luminance pixel values ​​after SAO. However, chrominance pixel values ​​can also be used to design new classifiers, which can benefit coding performance.

[0162] Sixth, similar to the following stages: 1) immediately before deblocking filtering 2) prediction signal 3) residual signal 4) immediately before SAO, the luminance pixel values ​​are saved as input to the additional online luminance ALF filter formula, and the following stages: 1) immediately before deblocking filtering 2) prediction signal 3) residual signal 4) immediately before SAO, the chrominance pixel values ​​can also be saved as input to the additional online chrominance ALF filter formula, which can benefit coding performance.

[0163] Seventh, similar to the following stages: 1) immediately before deblocking filtering 2) prediction signal 3) residual signal 4) immediately before SAO, the luminance pixel values ​​are saved as input to the additional online luminance ALF filter formula. This can also be beneficial for coding performance.

[0164] Eighth, the classifier design in ECM only considers the reconstructed pixel values. However, coding mode information (such as whether the coding block uses skip mode coding, whether the coding block uses intra-frame, inter-frame P, or inter-frame B mode coding) can also be used to design the classifier, which can benefit coding performance.

[0165] Ninth, after the online ALF filter acquires samples as additional input based on the current line buffer settings in VVC from the following stages: 1) samples immediately before deblocking 2) prediction samples 3) residual samples 4) samples immediately before SAO, an additional line buffer is needed to store the 4 lines of corresponding luminance samples and 2 lines of corresponding chrominance samples above the horizontal CTU boundary, which increases the implementation complexity.

[0166] Tenth, after the CCALF filter acquires samples as additional input based on the current line buffer settings in VVC from the following stages: 1) samples immediately before deblocking 2) prediction samples 3) residual samples 4) samples immediately before SAO, an additional line buffer is needed to store the corresponding luminance samples of the 4 lines above the horizontal CTU boundary, which increases the implementation complexity.

[0167] Eleventh, after the online ALF filter acquires samples as additional input from the following stages: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO, sample padding is required when the filter shape of the additional input, whose center position is aligned with the sample to be filtered, crosses a boundary. In some embodiments, the boundary can be the boundary line of the region including the sample to be filtered. For example, the boundary can be a virtual boundary (line buffer boundary) or a picture (slice, tile) boundary.

[0168] Twelfth, two edge-based classifiers with two different window sizes are used to train two sets of ALF fixed filters. However, in addition to edge-based classifiers, other classifiers such as band-based classifiers, residual-based classifiers, etc., can also be used to train the corresponding ALF fixed filter sets, and the online ALF filter can use the outputs of all these trained ALF fixed filter sets as additional inputs, which can be beneficial to coding performance.

[0169] Thirteenth, use spatially adjacent reconstructed pixels as input to train the ALF fixed filter. However, when training the ALF fixed filter, in addition to spatially adjacent reconstructed pixels, other spatially adjacent pixels (such as spatially adjacent pixels immediately before deblocking filtering, spatially adjacent pixels in the predicted signal, spatially adjacent pixels in the residual signal, or spatially adjacent pixels immediately before SAO) can also be used as input to the ALF fixed filter, which can be beneficial to coding performance.

[0170] Fourteenth, sub-block level filter adaptation is applied only in the luma ALF. However, besides the luma ALF, sub-block level filter adaptation can also be extended to the chroma ALF, which can benefit coding performance.

[0171] Fifteenth, sub-block level filter adaptation is applied only in the luma ALF. However, besides the luma ALF, sub-block level filter adaptation can also be extended to the CCALF, which can benefit coding performance.

[0172] Embodiments of this disclosure provide methods and apparatus for improving the coding efficiency of adaptive loop filters (ALF) and cross-component adaptive loop filters (CCALF).

[0173] The online ALF filter takes spatially adjacent pixels in the predicted signal, spatially adjacent pixels in the residual signal, or spatially adjacent pixels before SAO as additional inputs.

[0174] A classifier combining features from edge-based and band-based classifiers is used as an additional classifier for the online ALF filter.

[0175] The filter shape for the chroma ALF changed from a rhombus to a long cross shape to match the filter shape for the luminance ALF.

[0176] A classifier from the following stages—1) immediately before deblocking filtering, 2) the predicted signal, 3) the residual signal, and 4) the pixel values ​​immediately before SAO—will be used as an additional classifier for the online ALF filter.

[0177] A classifier using chroma pixel values ​​is used as an additional classifier for the online ALF filter.

[0178] The online chroma ALF filter takes spatially adjacent pixels from the chroma prediction signal, spatially adjacent pixels from the chroma residual signal, spatially adjacent pixels from the stage immediately preceding the chroma SAO, or spatially adjacent pixels from the stage immediately preceding the chroma deblocking as additional inputs.

[0179] The CCALF filter takes spatially adjacent pixels from the luminance prediction signal, spatially adjacent pixels from the luminance residual signal, spatially adjacent pixels from the stage immediately preceding the luminance SAO, or spatially adjacent pixels from the stage immediately preceding the luminance deblocking as additional inputs.

[0180] A classifier that utilizes coding mode information (such as whether the coding block is coded in skip mode, whether the coding block is coded in intra-frame, inter-frame P, or inter-frame B mode) is used as an additional classifier for the online ALF filter.

[0181] When the online ALF filter acquires samples as additional input based on the current line buffer settings in VVC from the following stages: 1) samples immediately before deblocking 2) prediction samples 3) residual samples 4) samples immediately before SAO, the 4 lines corresponding to luminance samples and 2 lines corresponding to chrominance samples above the horizontal CTU boundary are assumed to be default values, which saves these line buffers.

[0182] When the CCALF filter acquires samples as additional input from the following stages based on the current line buffer settings in VVC: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO, the 4 lines above the horizontal CTU boundary corresponding to the luminance samples are assumed to be the default values, which saves these line buffers.

[0183] When the online ALF filter acquires samples as additional input from the following stages: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO, sample filling is performed when the filter shape of the additional input (whose center position is aligned with the sample to be filtered) crosses the virtual boundary (line buffer boundary) or the picture (slice, tile) boundary.

[0184] Additional sets of classifiers, such as band-based classifiers and residual-based classifiers, are used to train the ALF fixed filter. Then, the outputs of these additional sets of ALF fixed filters, along with the outputs of the original two sets of ALF fixed filters trained based on two edge-based classifiers, are used as inputs to the online ALF filter.

[0185] When training the ALF fixed filter, spatially adjacent reconstructed pixels, along with spatially adjacent pixels immediately preceding the deblocking filter, spatially adjacent pixels in the predicted signal, spatially adjacent pixels in the residual signal, or spatially adjacent pixels immediately preceding the SAO, are used as inputs to the ALF fixed filter.

[0186] Sub-block level filters are adaptively applied to chroma ALF, where edge-based, band-based, or residual-based classifiers are utilized in the chroma ALF.

[0187] Sub-block level filters are adaptively applied to CCALF, which utilizes edge-based, band-based, or residual-based classifiers.

[0188] According to an embodiment, when the online ALF filter takes samples from the pre-deblocking signal, prediction signal, residual signal, or pre-SAO signal as additional input, four luminance samples and two chrominance samples (the current line buffer setting in the VVC, which can be adjusted according to custom settings) above the horizontal CTU boundary in the pre-deblocking signal, prediction signal, residual signal, or pre-SAO signal are filled with predefined values.

[0189] According to an embodiment, when the CCALF filter takes luminance samples from the pre-deblocking signal, prediction signal, residual signal, or pre-SAO signal as additional input, four luminance samples above the horizontal CTU boundary in the pre-deblocking signal, prediction signal, residual signal, or pre-SAO signal (the current line buffer setting in the VVC, which can be adjusted according to custom settings) are filled with predefined values.

[0190] According to the embodiment, when the online ALF filter takes samples immediately before deblocking, prediction samples, residual samples, or samples immediately before SAO as additional inputs, sample filling is performed when the filter shape of the additional input (whose center position is aligned with the sample to be filtered) crosses the virtual boundary (row buffer boundary) or the picture (slice, tile) boundary.

[0191] According to the embodiment, additional sets of ALF fixed filters are trained using band-based classifiers, residual-based classifiers, etc. Furthermore, the outputs of these additional sets of ALF fixed filters, together with the outputs of the original two sets of ALF fixed filters trained based on two edge-based classifiers, are used as inputs to the online ALF filter.

[0192] According to an embodiment, when training the ALF fixed filter, spatially adjacent reconstructed pixels, together with spatially adjacent pixels immediately preceding the deblocking filter, spatially adjacent pixels in the predicted signal, spatially adjacent pixels in the residual signal, or spatially adjacent pixels immediately preceding the SAO, are used as inputs to the ALF fixed filter.

[0193] According to an embodiment, sub-block level filter adaptation is applied to the chroma ALF, wherein an edge-based classifier, a band-based classifier, or a residual-based classifier is used in the chroma ALF.

[0194] According to an embodiment, sub-block level filter adaptation is applied to CCALF, wherein edge-based classifiers, band-based classifiers, or residual-based classifiers are used in CCALF.

[0195] In some embodiments of this disclosure, the disclosed methods may be applied independently or in combination.

[0196] Information used as additional ALF inputs, including predictions, residuals, or information prior to SAO.

[0197] According to one or more embodiments of this disclosure, information prior to prediction, residuals, or SAO is used as input to the additional ALF formula. Different methods can be used to achieve this objective.

[0198] Figure 14 Presents the inputs of the online ALF filter. The online ALF filter can take all or a subset of the following as additional inputs: prediction samples, output samples obtained by feeding prediction samples into a fixed filter trained offline, residual samples, output samples obtained by feeding residual samples into a fixed filter trained offline, reconstructed samples immediately before SAO, and output samples obtained by feeding reconstructed samples immediately before SAO into a fixed filter trained offline.

[0199] In the first method, spatially neighboring pixels in the predicted signal are proposed as input to the additional ALF formula. Various filter shapes can be used to extract information from the predicted signal. For example, the filter shape could be as follows: Figure 15 The values ​​shown are 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from the predicted signal. In one embodiment, the limiting difference between surrounding pixels and the current pixel in the predicted signal is used as the ALF formula input. In another example, the limiting difference between surrounding pixels and co-pixels in the predicted signal, and the limiting difference between co-pixels and the current pixel in the predicted signal, are used as the ALF formula input.

[0200] Besides applying additional online ALF filter taps directly to the predicted signal, additional online ALF filter taps can also be applied to intermediate results obtained by feeding the predicted signal to fixed filters. Various fixed filters can be applied to filter the predicted signal to obtain intermediate results, which can collect information from the predicted signal over a large receptive field. For example, the two 13×13 diamond fixed filters used in the ALF of ECM can be used to filter the predicted signal to obtain intermediate results. When fixed filters are applied to the predicted signal, the block classification result can directly utilize the block classification result calculated for the signal immediately following the SAO, or the block classification result recalculated based on the predicted signal. When fixed filters are applied to the predicted signal, an intermediate result can be obtained using a single fixed filter trained on one block classifier, or two or more intermediate results can be obtained using two or more fixed filters trained on two or more block classifiers. In video coding standards, several sets of fixed filters are typically prepared, and a set of fixed filters can be selected from them through a Rate Distortion Optimization (RDO) process. For example, in ECM, a fixed filter (containing two 13×13 diamond fixed filters) is selected from two groups via an RDO procedure, and the group index is sent to the decoder. When a fixed filter is applied to the prediction signal, the group index of the prediction signal can be the same as the group index of the signal after SAO, or different from the group index of the signal after SAO based on a predefined criterion (in ECM, there are two groups, so if the group index of the signal after SAO is 0, the group index of the prediction signal is 1; if the group index of the signal after SAO is 1, the group index of the prediction signal is 0), or determined by the RDO procedure for the prediction signal, wherein in the first and second cases, it is not necessary to send the group index of the prediction signal to the decoder, and in the third case, it is necessary to send the group index of the prediction signal to the decoder.

[0201] When additional online filter taps are applied to the intermediate results obtained by feeding the predicted signal to a fixed filter, various filter shapes can be used to extract information from the intermediate results. For example, the filter shape can be as follows: Figure 15 The results are shown as 1×1, 3×3, or 5×5. Various formula forms can be used to extract information from the intermediate results. In one embodiment, the limiting difference between surrounding pixels and the current pixel in the intermediate results is used as the ALF formula input. In another example, the limiting difference between surrounding pixels and co-pixels in the intermediate results, and the limiting difference between co-pixels and the current pixel in the intermediate results, are used as the ALF formula input.

[0202] It should be noted that the additional online ALF filter tap can be applied only to the predicted signal, only to the intermediate result obtained by feeding the predicted signal to a fixed filter, or to both the predicted signal and the intermediate result obtained by feeding the predicted signal to a fixed filter. For example, in AI (intra-frame) testing, the additional online ALF filter tap is applied only to the predicted signal; in RA (random access) testing, the additional online ALF filter tap is applied to both the predicted signal and the intermediate result obtained by feeding the predicted signal to a fixed filter.

[0203] In the second method, spatially adjacent pixels in the residual signal are proposed as input to the additional ALF formula. Various filter shapes can be used to extract information from the residual signal. For example, the filter shape could be as follows: Figure 15 The values ​​shown are 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from the residual signal. In this embodiment, the clipping results of co-position pixels in the residual signal are used as input to the ALF formula.

[0204] Besides applying additional online ALF filter taps directly to the residual signal, additional online ALF filter taps can also be applied to intermediate results obtained by feeding the residual signal to a fixed filter. Various fixed filters can be applied to filter the residual signal to obtain intermediate results, which can collect residual signal information over a large receptive field. For example, the two 13 × 13 diamond fixed filters used in the ALF of an ECM can be used to filter the residual signal to obtain intermediate results. In one or more examples, consider that for the predicted signal and the pre-SAO signal, the range is exactly the same as the range of the post-SAO signal (i.e., (0, 1024)), which is positive, but for the residual signal, the range can be positive or negative. Therefore, when a fixed filter is applied to the residual signal, the filtering result can be clipped to different ranges, such as (-1024, 1024), (-512, 512), (-256, 256), (-128, 128), etc. When a fixed filter is applied to the residual signal, the block-level classification result can directly utilize the block-level classification result calculated for the signal after SAO, or the block-level classification result recalculated based on the residual signal. When a fixed filter is applied to the residual signal, an intermediate result can be obtained using a single fixed filter trained on one block-level classifier, or two or more intermediate results can be obtained using two or more fixed filters trained on two or more block-level classifiers. When a fixed filter is applied to the residual signal, the group index used for the residual signal can be the same as the group index of the signal after SAO, or different from the group index of the signal after SAO based on a predefined criterion (in ECM, there are two groups, so if the group index of the signal after SAO is 0, the group index used for the residual signal is 1; if the group index of the signal after SAO is 1, the group index used for the residual signal is 0), or determined by the RDO process for the residual signal. In the first and second cases, it is not necessary to send the group index used for the residual signal to the decoder, while in the third case, it is necessary to send the group index used for the residual signal to the decoder.

[0205] When additional online filter taps are applied to intermediate results obtained by feeding the residual signal into a fixed filter, various filter shapes can be used to extract information from the intermediate results. For example, the filter shape can be as follows: Figure 15 The results are shown as 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from the intermediate results. In this embodiment, the clipping results of co-pixels in the intermediate results are used as inputs to the ALF formula.

[0206] It should be noted that the additional online ALF filter tap can be applied only to the residual signal, only to the intermediate result obtained by feeding the residual signal to a fixed filter, or to both the residual signal and the intermediate result obtained by feeding the residual signal to a fixed filter. For example, in AI (intra-frame) testing, the additional online ALF filter tap is applied only to the residual signal; in RA (random access) testing, the additional online ALF filter tap is applied to both the residual signal and the intermediate result obtained by feeding the residual signal to a fixed filter.

[0207] In the third method, spatially neighboring pixels from the pre-SAO signal stage are proposed as input to the additional ALF formula. Various filter shapes can be used to extract information from the pre-SAO signal. For example, the filter shape could be as follows: Figure 15 The values ​​shown are 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from the pre-SAO signal. In one embodiment, the limiting difference between surrounding pixels and the current pixel in the pre-SAO signal is used as the ALF formula input. In another example, the limiting difference between surrounding pixels and co-pixels in the pre-SAO signal, and the limiting difference between co-pixels and the current pixel in the pre-SAO signal, are used as the ALF formula input.

[0208] Besides applying additional online ALF filter taps directly to the pre-SAO signal, additional online ALF filter taps can also be applied to intermediate results obtained by feeding the pre-SAO signal to a fixed filter. Various fixed filters can be applied to filter the pre-SAO signal to obtain intermediate results, allowing for the collection of pre-SAO signal information over a large receptive field. For example, the two 13 × 13 diamond fixed filters used in the ALF of an ECM can be used to filter the pre-SAO signal to obtain intermediate results. When a fixed filter is applied to the pre-SAO signal, the block classification result can directly utilize the block classification result calculated for the post-SAO signal, or the block classification result recalculated based on the pre-SAO signal. When applying a fixed filter to the pre-SAO signal, an intermediate result can be obtained using a single fixed filter trained on one block classifier, or two or more intermediate results can be obtained using two or more fixed filters trained on two or more block classifiers. When a fixed filter is applied to the pre-SAO signal, the group index of the pre-SAO signal can be the same as the group index of the post-SAO signal, or different from the group index of the post-SAO signal based on a predefined standard (in ECM, there are two groups, so if the group index of the post-SAO signal is 0, then the group index of the pre-SAO signal is 1; if the group index of the post-SAO signal is 1, then the group index of the pre-SAO signal is 0), or determined by the RDO process for the pre-SAO signal. In the first and second cases, it is not necessary to send the group index of the pre-SAO signal to the decoder, while in the third case, it is necessary to send the group index of the pre-SAO signal to the decoder.

[0209] When applying additional online filter taps to the intermediate results obtained by feeding the pre-SAO signal to a fixed filter, various filter shapes can be used to extract information from the intermediate results. For example, the filter shape can be as follows: Figure 15 The results are shown as 1×1, 3×3, or 5×5. Various formula forms can be used to extract information from the intermediate results. In one embodiment, the limiting difference between surrounding pixels and the current pixel in the intermediate results is used as the ALF formula input. In another example, the limiting difference between surrounding pixels and co-pixels in the intermediate results, and the limiting difference between co-pixels and the current pixel in the intermediate results, are used as the ALF formula input.

[0210] It should be noted that the additional online ALF filter tap can be applied only to the SAO pre-signal, or only to the intermediate result obtained by feeding the SAO pre-signal to a fixed filter, or both the SAO pre-signal and the intermediate result obtained by feeding the SAO pre-signal to a fixed filter. For example, in AI (Intra-Frame) testing, the additional online ALF filter tap is applied only to the SAO pre-signal; in RA (Random Access) testing, the additional online ALF filter tap is applied to both the SAO pre-signal and the intermediate result obtained by feeding the SAO pre-signal to a fixed filter.

[0211] The fourth method proposes using information from the predicted signal, residual signal, or pre-SAO signal as input to the ALF formula. The methods proposed in the first, second, and third methods can be combined to implement the fourth method.

[0212] A new classifier that combines features from edge-based and band-based classifiers.

[0213] According to one or more embodiments of this disclosure, features from an edge-based classifier and a band-based classifier are combined to derive a new classifier for an online ALF filter. Different methods can be used to achieve this goal.

[0214] The first method proposes to first calculate the directionality of sub-blocks of the brightness component. D Then, the sum of the sample values ​​of the sub-blocks is calculated and mapped to an index referencing the band-based classifier, and the class index of the sub-block is calculated as follows: (17), in B It is an index calculated based on a classifier with reference to the band. Indicates directionality D The total number. In the embodiment, for a 2×2 luminance block, the directionality D Calculated as in ECM Same, and B Calculated as B = (sum 5) >> (sample bit depth + 2) (18).

[0215] In the second method, it is proposed to first calculate the activity value of the sub-block of the luminance component. A Then, the sum of the sample values ​​of the sub-blocks is calculated and mapped to an index referencing the band-based classifier, and the class index of the sub-block is calculated as follows: (19), in B It is an index calculated based on a classifier with reference to the band. Indicates activity valueA The total number. In the embodiment, for a 2×2 luminance block, the activity value A Calculated as in ECM Same, and B Calculated as B = (sum 5) >> (sample bit depth + 2) (20).

[0216] In the third method, it is proposed to first calculate the index of the sub-blocks of the brightness component of the reference edge-based classifier, then calculate the sum of the sample values ​​of the sub-blocks and map them to the index of the reference band-based classifier, and the class index of the sub-block is calculated as follows: (twenty one), in B It is an index calculated based on a classifier with reference to the band. This represents the total number of indices computed using an edge-based classifier. E It is an index calculated with reference to an edge-based classifier. In the embodiment, for a 2×2 brightness block, the index is... E Calculated as in ECM Same, and B Calculated as B = (sum 2) >> (sample bit depth + 2) (22).

[0217] Adjust the shape of the chroma ALF filter to match the shape of the luminance ALF filter.

[0218] In a third aspect of this disclosure, it is proposed to change the shape of the chroma ALF filter from a rhombus to a long cross shape that matches the shape of the luminance ALF filter, such as... Figure 16 As shown.

[0219] A new classifier utilizing pixel values ​​from the stage before deblocking filtering

[0220] According to one or more embodiments of this disclosure, a new classifier for an online ALF filter is derived using pixel values ​​from a stage immediately preceding deblocking filtering. Different methods can be used to achieve this objective.

[0221] The first method proposes to first calculate the directionality of sub-blocks of the brightness component. D Then, the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-deblocking filtering stage of the sub-block, or within adjacent N×N windows surrounding the sub-block, is calculated and mapped to a difference index, and the class index of the sub-block is calculated as follows: (twenty three), in It is a poor index. Indicates directionality D The total number. In the embodiment, for a 2×2 luminance block, the directionality D Calculated as in ECM Same, and Calculated as (twenty four) in It is the sum of the differences between 2×2 luminance blocks, or the sum of the differences between adjacent N×N (such as 8×8) windows surrounding a 2×2 luminance block.

[0222] In the second method, it is proposed to first calculate the activity value of the sub-block of the luminance component. A Then, the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-deblocking filtering stage of the sub-block, or within adjacent N×N windows surrounding the sub-block, is calculated and mapped to a difference index, and the class index of the sub-block is calculated as follows: (25), in It is a poor index. Indicates activity value A The total number. In the embodiment, for a 2×2 luminance block, the activity value A Calculated as in ECM Same, and Calculate as shown in equation (24).

[0223] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component by referring to an edge-based classifier, then calculate the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-deblocking filtering stage of the sub-block or in adjacent N×N windows around the sub-block, and map this to the difference index. The class index of the sub-block is then calculated as follows: (26), in It is a poor index. This represents the total number of indices computed using an edge-based classifier. E It is an index calculated with reference to an edge-based classifier. In the embodiment, for a 2×2 brightness block, the index is... E Calculated as in ECM Same, and Calculate as shown in equation (24).

[0224] In the fourth method, it is proposed to first calculate the indexed sub-blocks of the luminance component. BThen, the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-deblocking filtering stage of the sub-block, or within adjacent N×N windows surrounding the sub-block, is calculated and mapped to a difference index, and the class index of the sub-block is calculated as follows: (27), in It is a poor index. This indicates the total number of values. In the embodiment, for a 2×2 luminance block, the index is... B Calculated as B = (sum 8) >> (sample bit depth + 2) (28), and Calculate as shown in equation (24).

[0225] In the fifth method, it is proposed to calculate the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-deblocking filtering stage of the sub-block or in adjacent N×N windows around the sub-block, then map the sum of the differences to the difference index, and use the difference index as the category index.

[0226] In the sixth method, an edge-based classifier or a band-based classifier is proposed to be calculated based on the sample values ​​from the pre-deblocking filtering stage, wherein the calculation method is the same as that of the original edge-based classifier or band-based classifier calculated based on the sample values ​​after SAO.

[0227] A new classifier that utilizes pixel values ​​in the predicted signal

[0228] According to one or more embodiments of this disclosure, a novel classifier for an online ALF filter is derived using pixel values ​​in the predicted signal. Different methods can be used to achieve this objective.

[0229] The first method proposes to first calculate the directionality of sub-blocks of the brightness component. D Then, the sum of the differences between the sample points after SAO and the predicted signals of the sub-block, or between the co-located samples in the adjacent N×N windows surrounding the sub-block, is calculated and mapped to the difference index. The class index of the sub-block is then calculated as follows: (29), in It is a poor index. Indicates directionality D The total number. In the embodiment, for a 2×2 luminance block, the directionality D Calculated as in ECM Same, and Calculated as (30) in It is the sum of the differences between 2×2 luminance blocks, or the sum of the differences between adjacent N×N (such as 8×8) windows surrounding a 2×2 luminance block.

[0230] In the second method, it is proposed to first calculate the activity value of the sub-block of the luminance component. A Then, the sum of the differences between the sample points after SAO and the predicted signals of the sub-block, or between the co-located samples in the adjacent N×N windows surrounding the sub-block, is calculated and mapped to the difference index. The class index of the sub-block is then calculated as follows: (31), in It is a poor index. Indicates activity value A The total number. In the embodiment, for a 2×2 luminance block, the activity value A Calculated as in ECM Same, and Calculate as in equation (30).

[0231] In the third method, it is proposed to first calculate the index of the sub-block of the brightness component by referring to an edge-based classifier, then calculate the sum of the differences between the sample after SAO and the predicted signal of the sub-block or the co-located samples in the adjacent N×N window around the sub-block, and map this to the difference index. The class index of the sub-block is then calculated as follows: (32), in It is a poor index. This represents the total number of indices computed using an edge-based classifier. E It is an index calculated with reference to an edge-based classifier. In the embodiment, for a 2×2 brightness block, the index is... E Calculated as in ECM Same, and Calculate as in equation (30).

[0232] In the fourth method, it is proposed to first calculate the indexed sub-blocks of the luminance component. B Then, the sum of the differences between the sample points after SAO and the predicted signals of the sub-block, or between the co-located samples in the adjacent N×N windows surrounding the sub-block, is calculated and mapped to the difference index. The class index of the sub-block is then calculated as follows: (33), in It is a poor index. This indicates the total number of values. In the embodiment, for a 2×2 luminance block, the index is... B Calculated as B = (sum 8) >> (sample bit depth + 2) (34), and Calculate as in equation (30).

[0233] In the fifth method, it is proposed to calculate the sum of the differences between the sample points after SAO and the predicted signals of the sub-block or between the co-located samples in the adjacent N×N windows around the sub-block, and then map the sum of the differences to the difference index, and use the difference index as the category index.

[0234] In the sixth method, an edge-based classifier or a band-based classifier is proposed to be calculated based on the sample values ​​in the predicted signal, wherein the calculation method is the same as that of the original edge-based classifier or band-based classifier calculated based on the sample values ​​after SAO.

[0235] A new classifier utilizing pixel values ​​in the residual signal

[0236] According to one or more embodiments of this disclosure, a novel classifier for an online ALF filter is derived using pixel values ​​in the residual signal. Different methods can be used to achieve this objective.

[0237] The first method proposes to first calculate the directionality of sub-blocks of the brightness component. D Then, the sum of pixel values ​​in the residual signal of the sub-block or in the adjacent N×N window surrounding the sub-block is calculated and mapped to the residual index, and the class index of the sub-block is calculated as follows: (35), in It is a residual index. Indicates directionality D The total number. In the embodiment, for a 2×2 luminance block, the directionality D Calculated as in ECM Same, and Calculated as (36) in It is the sum of pixel values ​​in the residual signal of a 2×2 luminance block, or the sum of pixel values ​​in the residual signal of an adjacent N×N (such as 8×8) window surrounding the 2×2 luminance block.

[0238] In the second method, it is proposed to first calculate the activity value of the sub-block of the luminance component. A Then, the sum of pixel values ​​in the residual signal of the sub-block or in the adjacent N×N window surrounding the sub-block is calculated and mapped to the residual index, and the class index of the sub-block is calculated as follows: (37), in It is a residual index. Indicates activity value A The total number. In the embodiment, for a 2×2 luminance block, the activity value A Calculated as in ECM Same, and Calculate as in equation (36).

[0239] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component by referring to an edge-based classifier, then calculate the sum of pixel values ​​in the residual signal of the sub-block or in the adjacent N×N window around the sub-block, and map it to the residual index. The class index of the sub-block is then calculated as follows: (38), in It is a residual index. This represents the total number of indices computed using an edge-based classifier. E It is an index calculated with reference to an edge-based classifier. In the embodiment, for a 2×2 brightness block, the index is... E Calculated as in ECM Same, and Calculate as in equation (36).

[0240] In the fourth method, it is proposed to first calculate the indexed sub-blocks of the luminance component. B Then, the sum of pixel values ​​in the residual signal of the sub-block or in the adjacent N×N window surrounding the sub-block is calculated and mapped to the residual index, and the class index of the sub-block is calculated as follows: (39), in It is a residual index. This indicates the total number of values. In the embodiment, for a 2×2 luminance block, the index is... B Calculated as B = (sum 8) >> (sample bit depth + 2) (40), and Calculate as in equation (36).

[0241] In the fifth method, it is proposed to calculate the sum of pixel values ​​in the residual signal of the sub-block or in the adjacent N×N window around the sub-block, then map the sum of residual values ​​to the residual index, and use the residual index as the category index.

[0242] The sixth method proposes to first calculate the sum of the absolute values ​​of pixel values ​​in the residual signal of the sub-block or in the adjacent N×N windows surrounding the sub-block, and then map this sum to the absolute value of the residual index. Then, the sum of pixel values ​​in the residual signal of the sub-block or in the adjacent N×N window surrounding the sub-block is calculated and mapped to the sign of the residual index. And the category index of the sub-block is calculated as (41) in This represents the total number of absolute values ​​of the residual index. In one example, for a 2×2 luminance block, the absolute value of the residual index is... Calculated as = ( 8) >> (sample bit depth + 2) (42) in It is the sum of the absolute values ​​of the pixel values ​​in the residual signal of the 2×2 luminance block, or the sum of the absolute values ​​of the pixel values ​​in the residual signal of the adjacent N×N (such as 8×8) windows surrounding the 2×2 luminance block, and Calculated as in equation (36). In this disclosure, N can be any integer based on the application or other factors.

[0243] In the seventh method, a reference edge-based classifier is proposed, which calculates the index of the sub-block based on the pixel values ​​in the residual signal, and then uses the index as the category index.

[0244] In the eighth method, a reference edge-based classifier is proposed, which calculates the index of the sub-block based on the absolute value of the pixel value in the residual signal, and then uses the index as the category index.

[0245] A new classifier utilizing pixel values ​​from the pre-SAO stage

[0246] According to one or more embodiments of this disclosure, a new classifier for an online ALF filter is derived using pixel values ​​from the pre-SAO stage. Different methods can be used to achieve this goal.

[0247] The first method proposes to first calculate the directionality of sub-blocks of the brightness component. D Then, the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-SAO stage of the sub-block or within adjacent N×N windows surrounding the sub-block is calculated and mapped to a difference index, and the class index of the sub-block is calculated as follows: (43), in It is a poor index. Indicates directionality D The total number. In the embodiment, for a 2×2 luminance block, the directionality D Calculated as in ECM Same, and Calculated as (44) in It is the sum of the differences between 2×2 luminance blocks, or the sum of the differences between adjacent N×N (such as 8×8) windows surrounding a 2×2 luminance block. In this disclosure, N can be any integer based on the application or other factors.

[0248] In the second method, it is proposed to first calculate the activity value of the sub-block of the luminance component. A Then, the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-SAO stage of the sub-block or within adjacent N×N windows surrounding the sub-block is calculated and mapped to a difference index, and the class index of the sub-block is calculated as follows: (45), in It is a poor index. Indicates activity value A The total number. In the embodiment, for a 2×2 luminance block, the activity value A Calculated as in ECM Same, and Calculate as shown in equation (44).

[0249] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component by referring to an edge-based classifier, then calculate the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-SAO stage of the sub-block or in adjacent N×N windows around the sub-block, and map this to the difference index. The class index of the sub-block is then calculated as follows: (46), in It is a poor index. This represents the total number of indices computed using an edge-based classifier. E It is an index calculated with reference to an edge-based classifier. In the embodiment, for a 2×2 brightness block, the index is... E Calculated as in ECM Same, and Calculate as shown in equation (44).

[0250] In the fourth method, it is proposed to first calculate the indexed sub-blocks of the luminance component. BThen, the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-SAO stage of the sub-block or within adjacent N×N windows surrounding the sub-block is calculated and mapped to a difference index, and the class index of the sub-block is calculated as follows: (47), in It is a poor index. This indicates the total number of values. In the embodiment, for a 2×2 luminance block, the index is... B Calculated as B = (sum 8) >> (sample bit depth + 2) (48), and Calculate as shown in equation (44).

[0251] In the fifth method, it is proposed to calculate the sum of the differences between samples from the post-SAO stage and co-located samples from the pre-SAO stage of a sub-block or within an adjacent N×N window surrounding the sub-block, then map the sum of the differences to a difference index, and use the difference index as a category index. In this disclosure, N can be any integer based on the application or other factors.

[0252] In the sixth method, an edge-based classifier or a band-based classifier is proposed to be calculated based on the sample values ​​from the pre-SAO stage, wherein the calculation method is the same as that of the original edge-based classifier or band-based classifier calculated based on the sample values ​​after SAO.

[0253] A new classifier using chroma pixel values

[0254] According to one or more embodiments of this disclosure, chroma pixel values ​​are used to derive a new classifier for an online ALF filter. Different methods can be used to achieve this goal.

[0255] In the first method, it is proposed to first calculate the indexed sub-blocks of the luminance component. Then calculate the indexed values ​​of the corresponding U and V components. and And the category index of the sub-block is calculated as (49), in, , and It refers to the Y, U, and V indices calculated based on the band-based classifier. and This represents the total number of U and V band index values. In the embodiment, for a 2×2 luminance block, , and Calculated as = (sumY 6) >> (sample bit depth + 2) (50), = (sumU 2) >> (sample bit depth + 2) (51), = (sumV 2) >> (sample bit depth + 2) (52).

[0256] Chromaticity information from stages such as deblocking, prediction, pre-residual, or pre-SAO is used as additional chroma ALF input.

[0257] According to one or more embodiments of this disclosure, chromaticity information from stages such as deblocking, prediction, pre-residual, or pre-SAO is used as input to the additional chromaticity ALF formula. This objective can be achieved using different methods.

[0258] In the first method, spatially neighboring pixels in the chroma prediction signal are proposed as input to the additional chroma ALF formula. Various filter shapes can be used to extract information from the chroma prediction signal. For example, the filter shape could be as follows: Figure 15 The diagram shows 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from the chroma prediction signal. In one embodiment, the limiting difference between surrounding pixels and the current chroma pixel in the chroma prediction signal is used as the chroma ALF formula input. In another example, the limiting difference between surrounding pixels and co-pixels in the chroma prediction signal, and the limiting difference between co-pixels and the current chroma pixel in the chroma prediction signal, are used as the chroma ALF formula input.

[0259] In the second method, spatially adjacent pixels in the chroma residual signal are proposed as input to the additional chroma ALF formula. Various filter shapes can be used to extract information from the chroma residual signal. For example, the filter shape could be as follows: Figure 15 The values ​​shown are 1×1, 3×3, or 5×5. Information from the chroma residual signal can be extracted using various formulaic forms. In this embodiment, the clipping result of the co-position pixels in the chroma residual signal is used as the input to the chroma ALF formula.

[0260] In the third method, spatially neighboring pixels from the pre-stage chroma SAO signal are proposed as input to the additional chroma ALF formula. Various filter shapes can be used to extract information from the pre-stage chroma SAO signal. For example, the filter shape could be as follows: Figure 15The diagram shows 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from the chroma SAO pre-stage signal. In one embodiment, the limiting difference between the surrounding pixels of the signal from the chroma SAO pre-stage and the current chroma pixel is used as the chroma ALF formula input. In another example, the limiting difference between the surrounding pixels of the signal from the chroma SAO pre-stage and the co-pixels of the signal from the chroma SAO pre-stage, and the limiting difference between the co-pixels of the signal from the chroma SAO pre-stage and the current chroma pixel are used as the chroma ALF formula input.

[0261] In the fourth method, spatially neighboring pixels from the pre-chroma deblocking stage signal are proposed as inputs to the additional chroma ALF formula. Various filter shapes can be used to extract information from the pre-chroma deblocking stage signal. For example, the filter shape could be as follows: Figure 15 The diagram shows 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from the pre-chroma deblocking stage signal. In one embodiment, the limiting difference between the surrounding pixels and the current chroma pixel from the pre-chroma deblocking stage signal is used as the chroma ALF formula input. In another example, the limiting difference between the surrounding pixels and the co-pixels from the pre-chroma deblocking stage signal, and the limiting difference between the co-pixels and the current chroma pixel from the pre-chroma deblocking stage signal, are used as the chroma ALF formula input.

[0262] The fifth method proposes using information from the chromaticity prediction signal, residual signal, pre-SAO signal, or pre-deblocking signal as input to the chromaticity ALF formula. This fifth method can be implemented by combining the utilization methods proposed in the first, second, third, and fourth methods.

[0263] Luminance information from the pre-deblocking, prediction, residual, or pre-SAO stages is used as additional CCALF input.

[0264] According to one or more embodiments of this disclosure, luminance information from the pre-deblocking, prediction, residual, or pre-SAO stages is used as additional CCALF formula inputs. Different methods can be used to achieve this goal.

[0265] In the first method, spatially neighboring pixels in the brightness prediction signal are proposed as input to an additional CCALF formula. Various filter shapes can be used to extract information from the brightness prediction signal. For example, the filter shape could be as follows: Figure 12 The 3x4 diagram is shown. Information from the luminance prediction signal can be extracted using various formulaic forms. In one embodiment, the difference between the surrounding pixels in the luminance prediction signal and the current corresponding luminance pixel is used as the CCALF formula input. In another example, the differences between the surrounding pixels in the luminance prediction signal and the co-pixels in the current corresponding luminance prediction signal, and the differences between the co-pixels in the current corresponding luminance prediction signal and the current corresponding luminance pixel are used as the CCALF formula input.

[0266] In the second method, spatially adjacent pixels in the luminance residual signal are proposed as input to an additional CCALF formula. Various filter shapes can be used to extract information from the luminance residual signal. For example, the filter shape could be as follows: Figure 12 The 3x4 diagram is shown. Information from the luminance residual signal can be extracted using various formulaic forms. In this embodiment, co-position pixels in the luminance residual signal are used as inputs to the CCALF formula.

[0267] In the third method, spatially neighboring pixels from the pre-SAO phase signal of lumenity are proposed as input to the additional CCALF formula. Various filter shapes can be used to extract information from the pre-SAO phase signal of lumenity. For example, the filter shape could be as follows: Figure 12 The diagram shows a 3x4 area. Information can be extracted from the pre-Luminance SAO stage signal using various formulaic forms. In one embodiment, the difference between the surrounding pixels from the pre-Luminance SAO stage signal and the current corresponding luminance pixel is used as the CCALF formula input. In another example, the differences between the surrounding pixels from the pre-Luminance SAO stage signal and the corresponding pixels in the current corresponding Luminance SAO stage signal, and the differences between the corresponding pixels in the current corresponding Luminance SAO stage signal and the current corresponding luminance pixel are used as the CCALF formula input.

[0268] In the fourth method, spatially neighboring pixels from the pre-luminosity deblocking stage signal are proposed as input to an additional CCALF formula. Various filter shapes can be used to extract information from the stage immediately preceding the luminosity deblocking signal. For example, the filter shape could be as follows: Figure 12 The diagram shows a 3x4 area. Information can be extracted from the pre-luminance deblocking stage signal using various formulaic forms. In one embodiment, the difference between the surrounding pixels from the pre-luminance deblocking stage signal and the current corresponding luminance pixel is used as the CCALF formula input. In another example, the differences between the surrounding pixels from the pre-luminance deblocking stage signal and the corresponding pixels in the current corresponding luminance deblocking signal, and the differences between the corresponding pixels in the current corresponding luminance deblocking signal and the current corresponding luminance pixel are used as the CCALF formula input.

[0269] The fifth method proposes using information from brightness prediction, residuals, and pre-SAO or pre-deblocking signals as input to the CCALF formula. This fifth method can be implemented by combining the utilization methods proposed in the first, second, third, and fourth methods.

[0270] A new classifier utilizing encoded pattern information

[0271] According to one or more embodiments of this disclosure, coding mode information (such as whether a coding block is encoded using a skip mode, or whether a coding block is encoded using intra-frame, inter-frame P, or inter-frame B modes) is used to derive a new classifier for an online ALF filter. This objective can be achieved using different methods.

[0272] In the first method, it is proposed to record whether the encoded block is encoded in a skip mode during the encoding and decoding process, and then use this information to design a new classifier. In one embodiment, a classifier with two categories corresponding to whether the skip mode is true or false is added as a new classifier. In another example, a classifier that combines the skip mode information with EO or BO is added as a new classifier.

[0273] In the second method, it is proposed to record whether the encoded block is encoded in intra-frame mode, inter-frame P mode, or inter-frame B mode during the encoding and decoding process, and then use this information to design a new classifier. In one embodiment, a classifier with three categories corresponding to intra-frame mode, inter-frame P mode, or inter-frame B mode is added as a new classifier. In another example, a classifier that combines intra-frame, inter-frame P, or inter-frame B mode information with EO or BO is added as a new classifier.

[0274] The third method proposes using coding mode information, such as whether the coding block is coded in skip mode, or whether it is coded in intra-frame, inter-frame P, or inter-frame B modes, to design a new classifier. The methods proposed in the first and second methods can be combined to implement the third method.

[0275] The line buffer used for appending ALF inputs is reduced.

[0276] According to one or more embodiments of this disclosure, when the online ALF filter acquires samples as additional input from the following stages: 1) pre-block samples, 2) prediction samples, 3) residual samples, 4) pre-SAO samples, etc., line buffers exist to store these samples. To reduce the line buffer requirements for these additional inputs, based on the current line buffer settings in the VVC, the 4 lines corresponding to luma samples and the 2 lines corresponding to chroma samples above the horizontal CTU boundary are assumed to be default values, which saves these line buffers. Different methods can be used to achieve this goal.

[0277] In the first method, based on the current row buffer settings in VVC, it is proposed that the four rows of luminance residual samples and two rows of chrominance residual samples above the horizontal CTU boundary are zero values, and that they come from the following stages: 1) pre-block samples 2) prediction samples 3) pre-SAO samples. The four rows of luminance samples and two rows of chrominance samples above the horizontal CTU boundary are co-position sample values ​​from the samples after the SAO stage.

[0278] In the second method, based on the current line buffer settings in VVC, it is proposed that the four rows of luminance samples and two rows of chrominance samples above the horizontal CTU boundary, which come from the following stages: 1) pre-block samples, 2) prediction samples, 3) residual samples, and 4) samples before SAO, are assumed to be the nearest sample values ​​in the horizontal CTU boundary in a repetitive manner.

[0279] In the third method, based on the current row buffer settings in VVC, hypotheses are proposed from the following stages: 1) Samples immediately before deblocking 2) Predicted samples 3) Residual samples 4) Samples immediately before SAO, etc., above the horizontal CTU boundary, 4 rows of luminance samples and 2 rows of chrominance samples, in a mirrored manner, wherein the first row of luminance samples and the first row of chrominance samples above the horizontal CTU boundary are assumed to be the corresponding sample values ​​in the horizontal CTU boundary, the second row of luminance samples and the second row of chrominance samples above the horizontal CTU boundary are assumed to be the corresponding sample values ​​in the first row of samples below the horizontal CTU boundary, and so on.

[0280] It should be noted that the 4 rows of luminance samples and 2 rows of chrominance samples above the horizontal CTU boundary are the current VVC line buffer settings, and specific values ​​can be adjusted according to custom settings.

[0281] The line buffer used for appending CCALF inputs is reduced.

[0282] According to one or more embodiments of this disclosure, when the CCALF filter acquires samples as additional input from the following stages: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO, a line buffer exists to store these samples. To reduce the line buffer requirements for these additional inputs, the four lines corresponding to luminance samples above the horizontal CTU boundary are assumed to be default values ​​based on the current line buffer settings in the VVC, which saves these line buffers. Different methods can be used to achieve this goal.

[0283] In the first method, based on the current line buffer settings in VVC, it is proposed that the four lines of luminance residual samples above the horizontal CTU boundary are zero values, and the four lines of luminance samples above the horizontal CTU boundary from the following stages: 1) samples immediately before deblocking 2) prediction samples 3) samples immediately before SAO are assumed to be co-position sample values ​​from the stage samples immediately after SAO.

[0284] In the second method, based on the current line buffer settings in VVC, four lines of luminance samples above the horizontal CTU boundary are proposed in a repetitive manner, which are assumed to be the nearest sample value in the horizontal CTU boundary, and are derived from the following stages: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO.

[0285] In the third method, based on the current row buffer settings in VVC, assumptions are made about the following stages: 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately before SAO, etc., four rows of luminance samples above the horizontal CTU boundary, in a mirrored manner, wherein the first row of luminance samples above the horizontal CTU boundary is assumed to be the corresponding sample value in the horizontal CTU boundary, the second row of luminance samples above the horizontal CTU boundary is assumed to be the corresponding sample value in the first row of samples below the horizontal CTU boundary, and so on.

[0286] It should be noted that the four lines of luminance samples above the horizontal CTU boundary are the current VVC line buffer settings, and specific values ​​can be adjusted according to custom settings.

[0287] Sample filling for additional ALF inputs

[0288] According to one or more embodiments of this disclosure, when an online ALF filter acquires samples as additional input from the following stages: 1) samples immediately preceding deblocking, 2) prediction samples, 3) residual samples, and 4) samples immediately preceding SAO, sample padding is performed when the filter shape of the additional input (whose center position is aligned with the sample to be filtered) crosses a virtual boundary (line buffer boundary) or an image (slice, tile) boundary. This objective can be achieved using different methods.

[0289] In the first method, symmetrical sample padding is applied when the filter shape of the additional input, whose center position is aligned with the sample to be filtered, crosses a virtual boundary (row buffer boundary) or an image (slice, tile) boundary. For example, assuming an online ALF filter uses residual samples as additional input, the filter shape of the fixed filter to be applied to the residual signal or the filter shape of the online filter directly applied to the residual signal is 7x7, and the filter shape of the residual signal whose center position is aligned with the sample to be filtered crosses a row buffer boundary, such as... Figures 17A-17C The diagram shows the symmetrical sample point filling, where... Masking the co-position residual pixels of the samples to be filtered to These are the original residual samples. to These are the modified residual sample values; the thick line represents the line buffer boundary. Shaded samples indicate filled residual samples. In summary, through symmetrical sample filling, both the additional input samples that are not on the same boundary side as the original additional input sample to be filtered, and the additional symmetrical input samples that are on the same boundary side as the original additional input sample to be filtered, are modified symmetrically.

[0290] In the second method, when the filter shape of an additional input whose center position is aligned with the sample point to be filtered crosses a virtual boundary (row buffer boundary) or an image (slice, tile) boundary, repeating sample filling is applied. Through repeating filling, additional input samples that are not on the same boundary side as the corresponding additional input sample point to be filtered are filled in the same way as symmetrical sample filling, while additional input samples that are on the same boundary side as the corresponding additional input sample point to be filtered remain unchanged.

[0291] ALF fixed filter with additional classifier

[0292] According to one or more embodiments of this disclosure, a classifier, a residual-based classifier, or the like is used to train an additional set of ALF fixed filters. The outputs of these additional sets of ALF fixed filters are then used as inputs to additional online ALF filters. Different methods can be used to achieve this goal.

[0293] In the first method, different sets of ALF fixed filters are first trained using different band classifiers. Different band classifiers can be defined based on different window sizes. For example, there are two band classifiers. For the first band classifier, the sum of sample values ​​for a 2×2 luminance block is calculated and mapped to the band classifier index, as follows: class_index = (sum 25) >> (sample bit depth + 2) For the second band classifier, the sum of sample values ​​in adjacent 8×8 windows surrounding the 2×2 brightness block is calculated and mapped to the band classifier index, as follows: class_index = (sum 25) >> (Sample bit depth + 6) Different band classifiers can also be defined based on different category numbers. For example, there are two band classifiers. For the first band classifier, the sum of the sample values ​​of the 2×2 brightness block is calculated and mapped to the band classifier index, as follows: class_index = (sum 25) >> (sample bit depth + 2) The number of categories in the band classifier is 25. For the second band classifier, the sum of the sample values ​​of the 2×2 brightness block is calculated and mapped to the band classifier index, as shown below: class_index = (sum 100) >> (sample bit depth + 2) The number of classes with the classifier is 100. ALF fixed filters with different taps can be used. For example, the fixed filter is a 13×13 rhombus shape. After training different sets of ALF fixed filters based on different classifiers, intermediate results can be obtained by feeding the reconstructed pixel values ​​into the new trained ALF fixed filters. The online ALF filter can then use these intermediate results as additional input. Various filter shapes can be used to extract information from the intermediate results. For example, the filter shape could be as follows: Figure 15 The results are shown as 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from intermediate results. For example, the clipping difference between surrounding pixels and the current pixel in the intermediate result is used as input to an additional online ALF filter.

[0294] In the second method, different sets of ALF fixed filters are first trained using different residual-based classifiers. Different residual-based classifiers can be defined based on different window sizes. For example, there are two residual-based classifiers. For the first residual-based classifier, the sum of the absolute values ​​of residual samples in adjacent 8×8 windows surrounding a 2×2 brightness block is calculated and mapped to the residual-based classifier index, as follows: classIdx = sum >> (sample bit depth – 4) The value of classIdx ranges from 0 to 24. For the second residual-based classifier, the sum of the absolute values ​​of residual samples in adjacent 12×12 windows surrounding the 2×2 brightness block is calculated and mapped to the residual-based classifier index, as follows: classIdx = sum >> (sample bit depth – 4) The value of classIdx ranges from 0 to 24. Different residual-based classifiers can also be defined based on different class numbers. For example, there are two residual-based classifiers. For the first residual-based classifier, the sum of the absolute values ​​of the residual samples in the adjacent 8×8 windows surrounding the 2×2 brightness block is calculated and mapped to the residual-based classifier index, as follows: classIdx = sum >> (sample bit depth – 4) The value of classIdx ranges from 0 to 24. For the second residual-based classifier, the sum of the absolute values ​​of the residual samples in the adjacent 8×8 windows surrounding the 2×2 brightness block is calculated and mapped to the residual-based classifier index, as follows: classIdx = sum >> (sample bit depth – 4) The value of classIdx ranges from 0 to 49. ALF fixed filters with different taps can be used. For example, an ALF fixed filter can be a 13×13 diamond shape. After training different sets of ALF fixed filters based on different residual-based classifiers, intermediate results can be obtained by feeding the reconstructed pixel values ​​into the new trained ALF fixed filters. The online ALF filter can then use these intermediate results as additional input. Various filter shapes can be used to extract information from the intermediate results. For example, the filter shape could be as follows: Figure 15 The results are shown as 1×1, 3×3, or 5×5. Various formulaic forms can be used to extract information from intermediate results. For example, the clipping difference between surrounding pixels and the current pixel in the intermediate result is used as input to an additional online ALF filter.

[0295] In the third approach, the methods presented in the first and second approaches can be combined. For example, two sets of ALF fixed filters can be trained using a classifier with a band and a residual-based classifier. The outputs of the two newly trained ALF fixed filters are then used as additional inputs to the online ALF filter. It should be noted that, in addition to the newly trained ALF fixed filters, the two ALF fixed filters trained according to the edge-based classifier are already included in the original design of the ALF in the ECM.

[0296] ALF fixed filter with additional input

[0297] According to one or more embodiments of this disclosure, when training the ALF fixed filter, spatially adjacent pixels immediately preceding the deblocking filter, spatially adjacent pixels in the predicted signal, spatially adjacent pixels in the residual signal, or spatially adjacent pixels immediately preceding the SAO are used as inputs to the additional ALF fixed filter. This objective can be achieved using different methods.

[0298] In the first method, the spatially adjacent pixels immediately preceding the deblocking filter are used as inputs to the additional ALF fixed filter during training. Various filter shapes can be used to extract information from the spatially adjacent pixels immediately preceding the deblocking filter. For example, the filter shape could be as follows: Figure 15 The shapes shown are 1×1, 3×3, 5×5, or 13×13 rhombuses. Various formulas can be used to extract information from spatially adjacent pixels immediately preceding the deblocking filter. For example, the clipping difference between the current pixel and the surrounding pixels immediately preceding the deblocking filter is used as input to an additional ALF fixed filter.

[0299] In the second method, spatially neighboring pixels in the predicted signal are used as inputs to an additional ALF fixed filter during ALF fixed filter training. Various filter shapes can be used to extract information from the predicted signal. For example, the filter shape could be as follows: Figure 15 The 1×1, 3×3, 5×5, or 13×13 rhombus shapes shown can be used to extract information from the predicted signal using various formulas. For example, the amplitude difference between the surrounding pixels and the current pixel in the predicted signal is used as the input to an additional ALF fixed filter.

[0300] In the third method, when training the ALF fixed filter, spatially neighboring pixels in the residual signal are used as inputs to an additional ALF fixed filter. Various filter shapes can be used to extract information from the residual signal. For example, the filter shape can be as follows: Figure 15 The residual signal can be presented in 1×1, 3×3, 5×5, or 13×13 rhombus shapes. Various formulaic forms can be used to extract information from the residual signal. For example, the clipping results of surrounding pixels in the residual signal are used as input to an additional ALF fixed filter.

[0301] In the fourth method, when training the ALF fixed filter, the spatially adjacent pixels immediately preceding the SAO are used as input to the additional ALF fixed filter. Various filter shapes can be used to extract information from the spatially adjacent pixels immediately preceding the SAO. For example, the filter shape could be as follows: Figure 15 The shapes shown are 1×1, 3×3, 5×5, or 13×13 rhombuses. Various formulas can be used to extract information from spatially adjacent pixels immediately preceding the SAO. For example, the clipping difference between the surrounding pixels immediately preceding the SAO and the current pixel is used as input to an additional ALF fixed filter.

[0302] In the fifth method, the methods presented in the first, second, third, and fourth methods can be combined. For example, both the spatially adjacent pixels immediately preceding the deblocking filter and the spatially adjacent pixels in the residual signal are used as inputs to the additional ALF fixed filter when training the ALF fixed filter.

[0303] Chromatic ALF with Sub-block Level Filter Adaptation

[0304] According to one or more embodiments of this disclosure, sub-block level filter adaptation is applied in a chroma ALF, wherein an edge-based classifier, a band-based classifier, or a residual-based classifier is used in the chroma ALF. Different methods can be used to achieve this objective.

[0305] In the first method, an edge-based classifier is used in the chroma ALF. The edge-based classifier can be defined at different sub-block levels, such as block sizes of 4×4, 2×2, or 1×1. The edge-based classifier can be computed in different ways. In the first example, the edge-based classifier is computed based on the reconstructed signal in the luma component. For example, an edge-based classifier used in the luma ALF and computed based on the post-SAO signal in the luma component is reused in the chroma ALF. In the second example, the edge-based classifier is computed based on the reconstructed signal in the chroma component. For example, a first edge-based classifier is computed based on the post-SAO signal in the Cb component, and then another edge-based classifier is computed based on the post-SAO signal in the Cr component. Finally, an edge-based classifier combining the first and second edge-based classifiers is used in the chroma ALF. It should be noted that when computed based on the post-SAO signal in the Cb or Cr component, the computation process can be the same as when computed based on the post-SAO signal in the luma component. In the third example, an edge-based classifier combining the edge-based classifier computed in the first example and the edge-based classifier computed in the second example is used for chroma ALF.

[0306] It should be noted that when reusing the edge-based classifier from the luma ALF for the chroma ALF, or recompiling the edge-based classifier for the chroma ALF based on the chroma components, the specific number of classes used in the chroma ALF may be the same as or different from the number of classes used in the luma ALF. For example, the number of classes used in the luma ALF is 25, which is obtained by combining activity (5) and directionality (5), while the number of classes used in the chroma ALF is 10, which is obtained by combining activity (2) and directionality (5).

[0307] In the second method, a band-based classifier is used in the chroma ALF. The band-based classifier can be defined at different sub-block levels, such as block sizes of 4×4, 2×2, or 1×1. The band-based classifier can be computed in different ways. In the first example, the band-based classifier is computed based on the reconstructed signal in the luma component. For example, a band-based classifier used in the luma ALF and computed based on the post-SAO signal in the luma component is reused in the chroma ALF. In the second example, the band-based classifier is computed based on the reconstructed signal in the chroma component. For example, a first band-based classifier is computed based on the post-SAO signal in the Cb component, and then another band-based classifier is computed based on the post-SAO signal in the Cr component. Finally, a band-based classifier combining the first and second band-based classifiers is used in the chroma ALF. It should be noted that when computed based on the post-SAO signal in the Cb or Cr component, the computation process can be the same as when computed based on the post-SAO signal in the luma component. In the third example, a band-based classifier combining the band-based classifiers calculated in the first and second examples is used in the chroma ALF. It should be noted that when reusing the band-based classifier from the luma ALF for the chroma ALF, or recalculating the band-based classifier based on the chroma components for the chroma ALF, the specific number of classes used in the chroma ALF can be the same as or different from the number of classes used in the luma ALF. For example, the luma ALF uses 25 classes, while the chroma ALF uses 10.

[0308] In the third method, a residual-based classifier is used in the chroma ALF. The residual-based classifier can be defined at different sub-block levels, such as block sizes of 4×4, 2×2, or 1×1. The residual-based classifier can be computed in different ways. In the first example, the residual-based classifier is computed based on the residual signal in the luminance component. For example, a residual-based classifier used in the luminance ALF and computed based on the residual signal in the luminance component is reused in the chroma ALF. In the second example, a residual-based classifier is computed based on the residual signal in the chroma component. For example, a first residual-based classifier is computed based on the residual signal in the Cb component, and then another residual-based classifier is computed based on the residual signal in the Cr component. Finally, a residual-based classifier combining the first and second residual-based classifiers is used in the chroma ALF. It should be noted that when computed based on the residual signal in the Cb or Cr component, the computation process can be the same as when computed based on the residual signal in the luminance component. In the third example, a residual-based classifier combining the residual-based classifiers computed in the first and second examples is used for the chroma ALF. It should be noted that when reusing the residual-based classifier from the luma ALF for the chroma ALF, or recompiling the residual-based classifier for the chroma ALF based on the chroma components, the specific number of classes used in the chroma ALF can be the same as or different from the number of classes used in the luma ALF. For example, the luma ALF uses 25 classes, while the chroma ALF uses 10.

[0309] In the fourth method, the methods presented in the first, second, and third methods can be combined. For example, in chroma ALF, both the edge-based classifier presented in the first method and the band-based classifier presented in the second method can be utilized.

[0310] CCALF with Sub-block Level Filter Adaptation

[0311] According to one or more embodiments of this disclosure, sub-block level filter adaptation is applied in CCALF, wherein an edge-based classifier, a band-based classifier, or a residual-based classifier is used in CCALF. Different methods can be used to achieve this objective.

[0312] In the first method, an edge-based classifier is used in the CCALF. The edge-based classifier can be defined at different sub-block levels, such as block sizes of 4×4, 2×2, or 1×1. The edge-based classifier can be computed in different ways. In the first example, the edge-based classifier is computed based on the reconstructed signal in the luminance component. For example, an edge-based classifier used in the luminance ALF and computed based on the post-SAO signal in the luminance component is reused in the CCALF for both the Cb and Cr components. In the second example, the edge-based classifier is computed based on the reconstructed signal in the chrominance component. For example, for the CCALF in the Cb component, the edge-based classifier is computed based on the post-SAO signal in the Cb component; for the CCALF in the Cr component, the edge-based classifier is computed based on the post-SAO signal in the Cr component. It should be noted that when computed based on the post-SAO signal in the Cb or Cr component, the computation process can be the same as when computed based on the post-SAO signal in the luminance component. In the third example, the edge-based classifier that combines the edge-based classifier calculated in the first example and the edge-based classifier calculated in the second example for the Cb component is used for the Cb component; the edge-based classifier that combines the edge-based classifier calculated in the first example and the edge-based classifier calculated in the second example for the Cr component is used for the Cr component.

[0313] It should be noted that when reusing the edge-based classifier from the luminance ALF for CCALF, or recompiling the edge-based classifier for CCALF based on the chrominance components, the specific number of classes used in CCALF may be the same as or different from the number of classes used in the luminance ALF. For example, the number of classes used in the luminance ALF is 25, which is obtained by combining activity (5) and directionality (5), while the number of classes used in CCALF is 10, which is obtained by combining activity (2) and directionality (5).

[0314] In the second method, a band-based classifier is utilized in CCALF. The band-based classifier can be defined at different sub-block levels, such as block sizes of 4×4, 2×2, or 1×1. The band-based classifier can be computed in different ways. In the first example, the band-based classifier is computed based on the reconstructed signal in the luminance component. For example, a band-based classifier used in the luminance ALF and computed based on the post-SAO signal in the luminance component is reused in both the Cb and Cr components of the CCALF. In the second example, the band-based classifier is computed based on the reconstructed signal in the chrominance component. For example, for the CCALF in the Cb component, the band-based classifier is computed based on the post-SAO signal in the Cb component; for the CCALF in the Cr component, the band-based classifier is computed based on the post-SAO signal in the Cr component. It should be noted that when computed based on the post-SAO signal in the Cb or Cr component, the computation process can be the same as when computed based on the post-SAO signal in the luminance component. In the third example, a band-based classifier combining the band-based classifier calculated in the first example and the band-based classifier calculated in the second example for the Cb component is used for the Cb component; a band-based classifier combining the band-based classifier calculated in the first example and the band-based classifier calculated in the second example for the Cr component is used for the Cr component. It should be noted that when reusing the edge-based classifier from the luminance ALF for the CCALF, or recalculating the edge-based classifier for the CCALF based on the chrominance component, the specific number of classes used in the CCALF can be the same as or different from the number of classes used in the luminance ALF. For example, the number of classes used in the luminance ALF is 25, while the number of classes used in the CCALF is 10.

[0315] In the third method, a residual-based classifier is utilized in CCALF. The residual-based classifier can be defined at different sub-block levels, such as block sizes of 4×4, 2×2, or 1×1. The residual-based classifier can be computed in different ways. In the first example, the residual-based classifier is computed based on the residual signal in the luminance component. For example, a residual-based classifier used in the luminance ALF and computed based on the residual signal in the luminance component is reused in both the Cb and Cr components of the CCALF. In the second example, the residual-based classifier is computed based on the residual signal in the chrominance component. For example, for the Cb component of the CCALF, the residual-based classifier is computed based on the residual signal in the Cb component; for the Cr component of the CCALF, the residual-based classifier is computed based on the residual signal in the Cr component. It should be noted that when computed based on the residual signal in the Cb or Cr component, the computation process can be the same as when computed based on the residual signal in the luminance component. In the third example, a residual-based classifier combining the residual-based classifier computed in the first example and the residual-based classifier computed in the second example for the Cb component is used for the Cb component; a residual-based classifier combining the residual-based classifier computed in the first example and the residual-based classifier computed in the second example for the Cr component is used for the Cr component. It should be noted that when reusing the residual-based classifier from the luminance ALF for the CCALF, or recompiling the residual-based classifier for the CCALF based on the chrominance component, the specific number of classes used in the CCALF can be the same as or different from the number of classes used in the luminance ALF. For example, the number of classes used in the luminance ALF is 25, while the number of classes used in the CCALF is 10.

[0316] In the fourth approach, the methods presented in the first, second, and third approaches can be combined. For example, CCALF utilizes both the edge-based classifier presented in the first approach and the band-based classifier presented in the second approach.

[0317] Figure 18 A computing environment 1810 coupled to a user interface 1850 is shown. The computing environment 1810 may be part of a data processing server. The computing environment 1810 includes a processor 1820, memory 1830, and input / output (I / O) interfaces 1840.

[0318] Processor 1820 typically controls the overall operation of computing environment 1810, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1820 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 1820 may include one or more modules that facilitate interaction between processor 1820 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, graphics processing unit (GPU), etc.

[0319] Memory 1830 is configured to store various types of data to support the operation of computing environment 1810. Memory 1830 may include predefined software 1832. Examples of such data include instructions for any application or method operating on computing environment 1810, video datasets, image data, etc. Memory 1830 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0320] I / O interface 1840 provides an interface between processor 1820 and peripheral interface modules (such as keyboards, click wheels, buttons, etc.). Buttons may include, but are not limited to, home buttons, start scan buttons, and stop scan buttons. I / O interface 1840 can be coupled to encoders and decoders.

[0321] In an embodiment, a non-transitory computer-readable storage medium is also provided, including, for example, a plurality of programs in a memory 1830 and / or a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. The plurality of programs can be executed by a processor 1820 in a computing environment 1810 to perform the above-described methods. In an embodiment, the plurality of programs can be executed by a processor 1820 in a computing environment 1810 to (e.g., from...) Figure 2The video encoder 20 in the computing environment 1810 receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 1820 in the computing environment 1810 to perform the above-described decoding method based on the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor 1820 in the computing environment 1810 to perform the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 1820 in the computing environment 1810 to (e.g., to...) Figure 3 The video decoder 30 in the middle sends the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store data generated by the encoder (e.g., Figure 2 The video encoder 20 in the video encoder uses, for example, the encoding method described above to generate the video for the decoder (e.g., Figure 3 The video decoder 30 in the video decoder uses a bitstream or data stream that includes encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) when decoding video data. Non-transitory computer-readable storage media can be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage devices, etc.

[0322] In one embodiment, a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method is provided. In another embodiment, a bitstream comprising encoded video information generated by the above-described encoding method or encoded video information to be decoded by the above-described decoding method is provided.

[0323] In one embodiment, a computing device is also provided, comprising: one or more processors (e.g., processor 1820); and a non-transitory computer-readable storage medium or memory 1830 therein storing a plurality of programs executable by the one or more processors, wherein the one or more processors are configured to perform the methods described above when executing the plurality of programs.

[0324] In one embodiment, a computer program product having instructions for storing or transmitting a bitstream is also provided, the bitstream including encoded video information generated by the encoding method described above or encoded video information to be decoded by the decoding method described above. In another embodiment, a computer program product including, for example, a plurality of programs in a memory 1830, the plurality of programs being executable by a processor 1820 in a computing environment 1810 to perform the methods described above. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0325] In an embodiment, the computing environment 1810 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.

[0326] Figure 19 This is a flowchart illustrating a method for video decoding according to some examples of this disclosure. At step 1901, the processor 1820 on the decoder side may obtain an adaptive loop filter (ALF) classifier for the sub-block. In some embodiments, the ALF classifier may be an edge-based, band-based, or residual-based classifier selected based on predefined criteria or characteristics of the video content. Each classifier type provides a different method for determining how the ALF will be applied to enhance chroma quality. In some embodiments, the ALF classifier for filtering luma samples may be used as an ALF for enhancing chroma quality, or the ALF classifier may be computed based on features of the chroma samples. At step 1902, the processor 1820 may determine the ALF based on the ALF classifier. The ALF classifier analyzes the image to classify different sub-blocks based on its features. For each classified sub-block, a filter is determined specifically for that category. At step 1903, the processor 1820 may obtain the filtered sample of the chroma sample based on the ALF and the reconstructed sample of the chroma sample in the sub-block. ALF-based filtering outputs filtered chromaticity samples with enhanced chromaticity representation, thereby reducing noise or artifacts, enhancing edge sharpness, or resolving residual differences.

[0327] In some embodiments, obtaining an ALF classifier for a sub-block includes: obtaining a luminance ALF classifier as an ALF classifier, wherein the luminance ALF is used to determine the luminance ALF for luminance samples in the sub-block.

[0328] In some embodiments, obtaining an ALF classifier for a sub-block includes: obtaining a first feature based on a first ALF classifier and a first signal in a first chromaticity component of a chromaticity sample; obtaining a second feature based on a second ALF classifier and a second signal in a second chromaticity component of a chromaticity sample; and deriving a combined classifier as an ALF classifier based on the first and second features.

[0329] In some embodiments, obtaining an ALF classifier for a sub-block includes: obtaining a luminance ALF classifier for determining the luminance ALF for luminance samples in the sub-block; obtaining a first feature based on a first ALF classifier and a first signal in a first chromaticity component of a chromaticity sample; obtaining a second feature based on a second ALF classifier and a second signal in a second chromaticity component of a chromaticity sample; deriving a combined classifier based on the first and second features; and deriving an ALF classifier based on the luminance ALF and the combined classifier.

[0330] In some embodiments, the number of classes in the ALF classifier is less than the number of classes in the luminance ALF classifier used to determine the luminance ALF for luminance samples in a sub-block.

[0331] In some embodiments, the number of categories in the ALF classifier is 10.

[0332] In some embodiments, the ALF includes a cross-component adaptive loop filter (CCALF).

[0333] In some embodiments, obtaining an ALF classifier for a sub-block includes: obtaining a luminance ALF classifier as an ALF classifier for both the first chromaticity component and the second chromaticity component, wherein the luminance ALF is used to determine the luminance ALF for luminance samples in the sub-block.

[0334] In some embodiments, obtaining an ALF classifier for a sub-block includes one of the following operations: obtaining a first ALF classifier based on a first signal in a first chromaticity component; or obtaining a second ALF classifier based on a second signal in a second chromaticity component.

[0335] In some embodiments, obtaining an ALF classifier for a sub-block includes one of the following operations: obtaining a luminance ALF classifier, the luminance ALF being used to determine the luminance ALF for luminance samples in the sub-block; obtaining a first ALF classifier based on a first signal in a first chromaticity component; and obtaining a combined ALF classifier based on the luminance ALF classifier and the first ALF classifier as an ALF classifier for the first chromaticity component; or obtaining a luminance ALF classifier, the luminance ALF being used to determine the luminance ALF for luminance samples in the sub-block; obtaining a second ALF classifier based on a second signal in a second chromaticity component; and obtaining a combined ALF classifier based on the luminance ALF classifier and the second ALF classifier as an ALF classifier for the second chromaticity component.

[0336] In some embodiments, the ALF classifier includes any one or any combination of an edge-based ALF classifier, a band-based ALF classifier, or a residual-based ALF classifier.

[0337] In some embodiments, the block size of the sub-block includes any one of 4×4, 2×2, or 1×1.

[0338] Figure 20This is a flowchart illustrating a method for video encoding according to some examples of this disclosure. At step 2001, the processor 1820 on the encoder side may obtain an adaptive loop filter (ALF) classifier for the sub-block. In some embodiments, the ALF classifier may be an edge-based, band-based, or residual-based classifier selected based on predefined criteria or characteristics of the video content. Each classifier type provides a different approach for determining how the ALF will be applied to enhance chroma quality. In some embodiments, the ALF classifier for filtering luma samples may be used as an ALF for enhancing chroma quality, or the ALF classifier may be computed based on features of the chroma samples. At step 2002, the processor 1820 may determine the ALF based on the ALF classifier. The ALF classifier analyzes the image to classify different sub-blocks based on its features. For each classified sub-block, a filter is determined specifically for that category. At step 2003, the processor 1820 may obtain the filtered sample of the chroma sample based on the ALF and the reconstructed sample of the chroma sample in the sub-block. ALF-based filtering outputs filtered chromaticity samples with enhanced chromaticity representation, thereby reducing noise or artifacts, enhancing edge sharpness, or resolving residual differences.

[0339] In some embodiments, obtaining an ALF classifier for a sub-block includes: obtaining a luminance ALF classifier as an ALF classifier, wherein the luminance ALF is used to determine the luminance ALF for luminance samples in the sub-block.

[0340] In some embodiments, obtaining an ALF classifier for a sub-block includes: obtaining a first feature based on a first ALF classifier and a first signal in a first chromaticity component of a chromaticity sample; obtaining a second feature based on a second ALF classifier and a second signal in a second chromaticity component of a chromaticity sample; and deriving a combined classifier as an ALF classifier based on the first and second features.

[0341] In some embodiments, obtaining an ALF classifier for a sub-block includes: obtaining a luminance ALF classifier for determining the luminance ALF for luminance samples in the sub-block; obtaining a first feature based on a first ALF classifier and a first signal in a first chromaticity component of a chromaticity sample; obtaining a second feature based on a second ALF classifier and a second signal in a second chromaticity component of a chromaticity sample; deriving a combined classifier based on the first and second features; and deriving an ALF classifier based on the luminance ALF and the combined classifier.

[0342] In some embodiments, the number of classes in the ALF classifier is less than the number of classes in the luminance ALF classifier used to determine the luminance ALF for luminance samples in a sub-block.

[0343] In some embodiments, the number of categories in the ALF classifier is 10.

[0344] In some embodiments, the ALF includes a cross-component adaptive loop filter (CCALF).

[0345] In some embodiments, obtaining an ALF classifier for a sub-block includes: obtaining a luminance ALF classifier as an ALF classifier for both the first chromaticity component and the second chromaticity component, wherein the luminance ALF is used to determine the luminance ALF for luminance samples in the sub-block.

[0346] In some embodiments, obtaining an ALF classifier for a sub-block includes one of the following operations: obtaining a first ALF classifier based on a first signal in a first chromaticity component; or obtaining a second ALF classifier based on a second signal in a second chromaticity component.

[0347] In some embodiments, obtaining an ALF classifier for a sub-block includes one of the following operations: obtaining a luminance ALF classifier, the luminance ALF being used to determine the luminance ALF for luminance samples in the sub-block; obtaining a first ALF classifier based on a first signal in a first chromaticity component; and obtaining a combined ALF classifier based on the luminance ALF classifier and the first ALF classifier as an ALF classifier for the first chromaticity component; or obtaining a luminance ALF classifier, the luminance ALF being used to determine the luminance ALF for luminance samples in the sub-block; obtaining a second ALF classifier based on a second signal in a second chromaticity component; and obtaining a combined ALF classifier based on the luminance ALF classifier and the second ALF classifier as an ALF classifier for the second chromaticity component.

[0348] In some embodiments, the ALF classifier includes any one or any combination of an edge-based ALF classifier, a band-based ALF classifier, or a residual-based ALF classifier.

[0349] In some embodiments, the block size of the sub-block includes any one of 4×4, 2×2, or 1×1.

[0350] In one embodiment, a method for storing a bitstream is also provided, comprising: storing the bitstream on a digital storage medium, wherein the bitstream includes coded video information generated by the above-described encoding method or coded video information to be encoded by the above-described encoding method.

[0351] In one embodiment, a method for transmitting a bitstream generated by the encoder described above is also provided. In another embodiment, a method for receiving a bitstream to be encoded by the encoder described above is also provided.

[0352] The description in this disclosure has been presented for illustrative purposes and is not intended to be exhaustive or limited to this disclosure. Many modifications, variations, and alternative embodiments will be apparent to those skilled in the art from the teachings presented in the foregoing description and the associated drawings.

[0353] Unless otherwise specifically stated, the order of steps in the method according to this disclosure is intended to be illustrative only, and the steps of the method according to this disclosure are not limited to the specific order described above, but may be changed according to actual circumstances. Furthermore, at least one step in the method according to this disclosure may be adjusted, combined, or omitted as needed.

[0354] The examples chosen and described are intended to explain the principles of this disclosure and to enable others skilled in the art to understand the various embodiments of this disclosure, and preferably to utilize the basic principles and various embodiments with various modifications suitable for the intended particular purpose. Therefore, it will be understood that the scope of this disclosure is not limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of this disclosure.

Claims

1. A method for video decoding, comprising: The decoder obtains an adaptive loop filter (ALF) classifier for the sub-blocks; The decoder determines the ALF based on the ALF classifier; as well as The decoder obtains the filtered sample points of the chroma sample points based on the ALF and the reconstructed sample points of the chroma sample points in the sub-block.

2. The method according to claim 1, wherein, Obtaining the ALF classifier for the sub-block includes: A luminance ALF classifier is obtained as the ALF classifier, and the luminance ALF is used to determine the luminance ALF for the luminance samples in the sub-block.

3. The method according to claim 1, wherein, Obtaining the ALF classifier for the sub-block includes: A first feature is obtained based on the first ALF classifier and the first signal in the first chromaticity component of the chromaticity sample point; A second feature is obtained based on the second ALF classifier and the second signal in the second chromaticity component of the chromaticity sample points; and A combined classifier is derived based on the first feature and the second feature as the ALF classifier.

4. The method according to claim 1, wherein, Obtaining the ALF classifier for the sub-block includes: A luminance ALF classifier is obtained, which is used to determine the luminance ALF for luminance samples in the sub-block; A first feature is obtained based on the first ALF classifier and the first signal in the first chromaticity component of the chromaticity sample point; The second feature is obtained based on the second ALF classifier and the second signal in the second chromaticity component of the chromaticity sample point; A combined classifier is derived based on the first feature and the second feature; and The ALF classifier is derived based on the brightness ALF and the combined classifier.

5. The method according to claim 1, wherein, The number of categories in the ALF classifier is less than the number of categories in the luminance ALF classifier, which is used to determine the luminance ALF for the luminance samples in the sub-block.

6. The method according to claim 5, wherein, The ALF classifier has 10 classes.

7. The method according to claim 1, wherein, The ALF includes a cross-component adaptive loop filter (CCALF).

8. The method according to claim 7, wherein, Obtaining the ALF classifier for the sub-block includes: A luminance ALF classifier is obtained as the ALF classifier for both the first chromaticity component and the second chromaticity component, and the luminance ALF is used to determine the luminance ALF for the luminance samples in the sub-block.

9. The method according to claim 7, wherein, Obtaining the ALF classifier for the sub-block includes one of the following operations: A first ALF classifier is obtained based on the first signal in the first chromaticity component; or A second ALF classifier is obtained based on the second signal in the second chromaticity component.

10. The method according to claim 7, wherein, Obtaining the ALF classifier for the sub-block includes one of the following operations: A luminance ALF classifier is obtained, the luminance ALF being used to determine the luminance ALF for luminance samples in the sub-block; a first ALF classifier is obtained based on a first signal in a first chromaticity component; and a combined ALF classifier is obtained based on the luminance ALF classifier and the first ALF classifier as the ALF classifier for the first chromaticity component. or, A luminance ALF classifier is obtained, the luminance ALF being used to determine the luminance ALF for luminance samples in the sub-block; a second ALF classifier is obtained based on a second signal in the second chromaticity component; and a combined ALF classifier is obtained based on the luminance ALF classifier and the second ALF classifier as the ALF classifier for the second chromaticity component.

11. The method according to claim 1, wherein, The ALF classifier includes any one or any combination of edge-based ALF classifiers, band-based ALF classifiers, or residual-based ALF classifiers.

12. The method according to claim 1, wherein, The size of the sub-block can be any one of 4×4, 2×2 or 1×1.

13. A method for video encoding, comprising: The encoder obtains an adaptive loop filter (ALF) classifier for the sub-blocks; The encoder determines the ALF based on the ALF classifier; as well as The encoder obtains the filtered sample points of the chroma sample points based on the ALF and the reconstructed sample points of the chroma sample points in the sub-block.

14. The method according to claim 13, wherein, Obtaining the ALF classifier for the sub-block includes: A luminance ALF classifier is obtained as the ALF classifier, and the luminance ALF is used to determine the luminance ALF for the luminance samples in the sub-block.

15. The method according to claim 13, wherein, Obtaining the ALF classifier for the sub-block includes: A first feature is obtained based on the first ALF classifier and the first signal in the first chromaticity component of the chromaticity sample point; A second feature is obtained based on the second ALF classifier and the second signal in the second chromaticity component of the chromaticity sample points; and A combined classifier is derived based on the first feature and the second feature as the ALF classifier.

16. The method according to claim 13, wherein, Obtaining the ALF classifier for the sub-block includes: A luminance ALF classifier is obtained, which is used to determine the luminance ALF for luminance samples in the sub-block; A first feature is obtained based on the first ALF classifier and the first signal in the first chromaticity component of the chromaticity sample point; The second feature is obtained based on the second ALF classifier and the second signal in the second chromaticity component of the chromaticity sample point; A combined classifier is derived based on the first feature and the second feature; and The ALF classifier is derived based on the brightness ALF and the combined classifier.

17. The method according to claim 13, wherein, The number of categories in the ALF classifier is less than the number of categories in the luminance ALF classifier, which is used to determine the luminance ALF for the luminance samples in the sub-block.

18. The method according to claim 17, wherein, The ALF classifier has 10 classes.

19. The method according to claim 13, wherein, The ALF includes a cross-component adaptive loop filter (CCALF).

20. The method according to claim 19, wherein, Obtaining the ALF classifier for the sub-block includes: A luminance ALF classifier is obtained as the ALF classifier for both the first chromaticity component and the second chromaticity component, and the luminance ALF is used to determine the luminance ALF for the luminance samples in the sub-block.

21. The method according to claim 19, wherein, Obtaining the ALF classifier for the sub-block includes one of the following operations: A first ALF classifier is obtained based on the first signal in the first chromaticity component; or A second ALF classifier is obtained based on the second signal in the second chromaticity component.

22. The method according to claim 19, wherein, Obtaining the ALF classifier for the sub-block includes one of the following operations: A luminance ALF classifier is obtained, the luminance ALF being used to determine the luminance ALF for luminance samples in the sub-block; a first ALF classifier is obtained based on a first signal in a first chromaticity component; and a combined ALF classifier is obtained based on the luminance ALF classifier and the first ALF classifier as the ALF classifier for the first chromaticity component. or, A luminance ALF classifier is obtained, the luminance ALF being used to determine the luminance ALF for luminance samples in the sub-block; a second ALF classifier is obtained based on a second signal in the second chromaticity component; and a combined ALF classifier is obtained based on the luminance ALF classifier and the second ALF classifier as the ALF classifier for the second chromaticity component.

23. The method according to claim 13, wherein, The ALF classifier includes any one or any combination of edge-based ALF classifiers, band-based ALF classifiers, or residual-based ALF classifiers.

24. The method according to claim 13, wherein, The size of the sub-block can be any one of 4×4, 2×2 or 1×1.

25. An apparatus for video decoding, comprising: One or more processors; as well as A memory, coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method according to any one of claims 1-12 when executing the instructions.

26. A non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 1-12.

27. An apparatus for video encoding, comprising: One or more processors; as well as A memory, coupled to the one or more processors and configured to store instructions executable by the one or more processors. The one or more processors are configured to perform the method according to any one of claims 13-24 when executing the instructions.

28. A non-transitory computer-readable storage medium for storing computer-executable instructions, which, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to any one of claims 13-24.

29. A non-transitory computer-readable storage medium for storing a bit stream to be decoded by the method according to any one of claims 1-12.

30. A non-transitory computer-readable storage medium for storing a bit stream to be generated by the method according to any one of claims 13-24.

31. A method for receiving a bit stream, wherein, The bitstream includes encoded video information to be decoded by the method according to any one of claims 1-12.

32. A method for transmitting a bit stream, wherein, The bitstream includes encoded video information generated by the method according to claims 13-24.