METHOD AND APPARATUS FOR ADAPTIVE LOOP FILTERING AND Cross-COMPONENT ADAPTIVE LOOP FILTER

By combining offline and online filters during video encoding and decoding, using spatial adjacent sample point information, the problem of insufficient encoding and decoding efficiency in the prior art is solved, and more efficient video encoding and decoding and storage optimization are achieved.

CN120266472APending Publication Date: 2025-07-04BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202380081768.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-03
Filing Date
2023-12-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has the problem of insufficient encoding and decoding efficiency in adaptive loop filtering and cross-component adaptive loop filter design, especially when processing spatial and temporal redundancy of video data, it is impossible to fully utilize adjacent sample information for optimization.

Method used

During the video encoding and decoding process, offline training fixed filters and online filters are used to obtain spatial adjacent sample points information, and filter processing is performed to reduce storage space and improve encoding and decoding efficiency.

Benefits of technology

It improves the encoding and decoding efficiency during the video encoding and decoding process, reduces the line buffer space for storing adjacent samples, and improves video quality and compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266472A_ABST
    Figure CN120266472A_ABST
Patent Text Reader

Abstract

A method and apparatus for video decoding / encoding are provided. In the provided method, a decoder obtains at least one secondary signal by applying at least one fixed filter to a primary signal. The at least one fixed filter is trained offline. The decoder obtains one or more spatially adjacent samples associated with the current sample. One or more spatially adjacent sample points are from at least one of the primary signal or the at least one secondary signal. The decoder obtains filtered sample points by applying at least one online filter to one or more spatially adjacent sample points associated with the current sample point.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to Provisional Application No. 63 / 430,015, filed on December 3, 2022, the entire content of which is incorporated herein by reference in its entirety. Technical field

[0003] This application relates to video encoding, decoding, and compression. More specifically, this application relates to methods and apparatuses for improving adaptive loop filtering processes and cross - component adaptive loop filtering processes. Background art

[0004] Various electronic devices (such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc.) support digital video. Electronic devices send, receive, or otherwise transmit digital video data via a communication network and / or store digital video data on a storage device. Since the bandwidth capacity of the communication network is limited and the memory resources of the storage device are limited, video data can be compressed using video encoding and decoding according to one or more video encoding and decoding standards before the video data is transmitted or stored. For example, video encoding and decoding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), Moving Picture Experts Group (MPEG) coding and decoding, etc. Video encoding and decoding typically employ prediction methods (such as inter - frame prediction, intra - frame prediction, etc.) that utilize the redundancy inherent in video data. Video encoding and decoding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing a degradation in video quality. Summary of the invention

[0005] Embodiments of the present disclosure provide techniques related to adaptive loop filtering and cross - component adaptive loop filtering.

[0006] In a first aspect, some embodiments of the present disclosure provide a method for video decoding, the method including: obtaining, by a decoder, at least one secondary signal by applying at least one fixed filter to a primary signal, wherein the at least one fixed filter is offline - trained; obtaining, by the decoder, one or more spatially - adjacent samples associated with a current sample, wherein the one or more spatially - adjacent samples are from at least one of the primary signal or the at least one secondary signal; and obtaining, by the decoder, a filtered sample by applying at least one online filter to the one or more spatially - adjacent samples associated with the current sample.

[0007] In a second aspect, some embodiments of the present disclosure provide a method for video encoding, the method comprising: obtaining, by an encoder, at least one secondary signal by applying at least one fixed filter to a primary signal, wherein the at least one fixed filter is offline trained; obtaining, by the encoder, one or more spatially neighboring samples associated with a current sample, wherein the one or more spatially neighboring samples are from at least one of the primary signal or the at least one secondary signal; and obtaining, by the encoder, a filtered sample by applying at least one online filter to the one or more spatially neighboring samples associated with the current sample.

[0008] In a third aspect, some embodiments of the present disclosure provide a method for video decoding, the method comprising: obtaining, by a decoder, a plurality of spatially neighboring samples associated with a current sample, wherein the plurality of spatially neighboring samples are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking; obtaining, by the decoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce a line buffer space for storing the plurality of spatially neighboring samples; and obtaining, by the decoder, a filtered sample based on the plurality of filtered input samples and the current sample.

[0009] In a fourth aspect, some embodiments of the present disclosure provide a method for video decoding, the method comprising: obtaining, by a decoder, a plurality of spatially neighboring samples associated with a current sample in a first channel, wherein the plurality of spatially neighboring samples are in a second channel and are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking; obtaining, by the decoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce a line buffer space for storing the plurality of spatially neighboring samples; and obtaining, by the decoder, a filtered sample in the first channel based on the plurality of filtered input samples and the current sample in the first channel.

[0010] In a fifth aspect, some embodiments of the present disclosure provide a method for video coding, the method comprising: obtaining, by an encoder, a plurality of spatially neighboring samples associated with a current sample, wherein the plurality of spatially neighboring samples are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking; obtaining, by the encoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce a line buffer space for storing the plurality of spatially neighboring samples; and obtaining, by the encoder, a filtered sample based on the plurality of filtered input samples and the current sample.

[0011] In a sixth aspect, some embodiments of the present disclosure provide a method for video coding, the method comprising: obtaining, by an encoder, a plurality of spatially neighboring samples associated with a current sample in a first channel, wherein the plurality of spatially neighboring samples are in a second channel and are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking; obtaining, by the encoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce a line buffer space for storing the plurality of spatially neighboring samples; and obtaining, by the encoder, a filtered sample in the first channel based on the plurality of filtered input samples and the current sample in the first channel.

[0012] It should be understood that the foregoing general description and the following detailed description are merely exemplary and not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings incorporated in the specification and constituting a part of this specification illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0014] Figure 1 is a block diagram showing an exemplary system for encoding and decoding video blocks according to some embodiments of the present disclosure.

[0015] Figure 2 is a block diagram showing an exemplary video encoder according to some embodiments of the present disclosure.

[0016] Figure 3 is a block diagram showing an exemplary video decoder according to some embodiments of the present disclosure.

[0017] Figures 4A to 4E is a block diagram showing how a frame is recursively divided into a plurality of video blocks of different sizes and shapes according to some embodiments of the present disclosure.

[0018] Figure 5 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0019] Figure 6 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0020] Figure 7 illustrates an ALF filter shape according to some examples of the present disclosure.

[0021] Figure 8 depicts a subsampled sample gradient according to some examples of the present disclosure.

[0022] Figure 9 illustrates a geometric transformation of a diamond filter shape according to some examples of the present disclosure.

[0023] Figure 10 illustrates an online filter shape used in ECM according to some examples of the present disclosure.

[0024] Figure 11 illustrates a CCALF architecture according to some examples of the present disclosure.

[0025] Figure 12 illustrates the relative positions of filtered chroma samples and their support in the luma plane for a 4:2:0 chroma format with chroma position type 0.

[0026] Figure 13 illustrates a 25-tap long filter according to some examples of the present disclosure.

[0027] Figure 14 illustrates various online ALF filter inputs according to some examples of the present disclosure.

[0028] Figure 15 illustrates a filter shape for a prediction signal or a signal before SAO according to an example of the present disclosure.

[0029] Figure 16 illustrates an adjusted ALF filter shape according to some examples of the present disclosure.

[0030] Figure 17 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.

[0031] Figure 18 is a flowchart illustrating a method for video encoding according to some examples of the present disclosure.

[0032] Figure 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.

[0033] Figure 20 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure.

[0034] Figure 21 is a flowchart illustrating a method for video encoding according to some examples of the present disclosure.

[0035] Figure 22 is a flowchart illustrating a method for video encoding according to some examples of the present disclosure.

[0036] Figure 23 is a diagram illustrating a computing environment coupled to a user interface according to some embodiments of the present disclosure. Detailed Description

[0037] Reference will now be made in detail to the detailed description, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, the subject matter may be practiced without these specific details, and various alternatives may be used without departing from the scope of the claims. For example, the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and in the drawings are used to distinguish objects and not to describe any particular order or sequence. It should be understood that such data may be interchanged where appropriate so that the embodiments of the present disclosure described herein can be implemented in sequences other than those shown in the drawings or described in the present disclosure.

[0039] Figure 1 is a block diagram showing an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. As Figure 1 shown, system 10 includes a source device 12 that generates and encodes video data that will later be decoded by a destination device 14. The source device 12 and the destination device 14 may include any of a variety of electronic devices, including cloud servers, server computers, desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some embodiments, the source device 12 and the destination device 14 are equipped with wireless communication capabilities.

[0040] In some embodiments, the target device 14 may receive encoded video data to be decoded via the link 16. The link 16 may include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the target device 14. In an embodiment, the link 16 may include a communication medium that enables the source device 12 to send the encoded video data directly to the target device 14 in real time. The encoded video data may be modulated according to a communication standard (e.g., a wireless communication protocol) and sent to the target device 14. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium may include routers, switches, base stations, or any other device that may facilitate communication from the source device 12 to the target device 14.

[0041] In some other embodiments, the encoded video data may be sent from the output interface 22 to the storage device 32. Subsequently, the encoded video data in the storage device 32 may be accessed by the target device 14 via the input interface 28. The storage device 32 may include any data storage medium in various distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In another example, the storage device 32 may correspond to a file server or another intermediate storage device that may hold the encoded video data generated by the source device 12. The target device 14 may access the stored video data from the storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing the encoded video data and sending the encoded video data to the target device 14. Exemplary file servers include web servers (e.g., for websites), File Transfer Protocol (FTP) servers, Network Attached Storage (NAS) devices, or local disk drives. The target device 14 may access the encoded video data through any standard data connection suitable for accessing the encoded video data stored on the file server, and the standard data connection includes a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a Digital Subscriber Line (DSL), a cable modem, etc.), or a combination of both a wireless channel and a wired connection. The transmission of the encoded video data from the storage device 32 may be a streaming transmission, a download transmission, or a combination of both a streaming transmission and a download transmission.

[0042] As Figure 1As shown in FIG. 0, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include a source such as, for example, a video capture device (e.g., a camera), a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video. As an example, if the video source 18 is a camera of a security monitoring system, the source device 12 and the destination device 14 may form a camera phone or a video phone. However, the embodiments described in this application can generally be applied to video encoding and decoding, and can be applied to wireless and / or wired applications.

[0043] The captured, pre-captured, or computer-generated video may be encoded by the video encoder 20. The encoded video data may be sent directly to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may also (or alternatively) be stored on a storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.

[0044] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem, and receives the encoded video data via a link 16. The encoded video data transmitted via the link 16 or provided on the storage device 32 may include various syntax elements generated by the video encoder 20 for use by the video decoder 30 when decoding the video data. Such syntax elements may be included in the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.

[0045] In some embodiments, the destination device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data to the user, and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0046] Video encoder 20 and video decoder 30 may operate according to proprietary standards or industry standards (e.g., VVC, HEVC, Part 10 of MPEG-4, AVC) or extensions of such standards. It should be understood that the present application is not limited to a particular video coding / decoding standard and may be applicable to other video coding / decoding standards. It is generally considered that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is also generally considered that the video decoder 30 of the destination device 14 may be configured to decode video data according to any of these current or future standards.

[0047] Video encoder 20 and video decoder 30 may be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic devices, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video coding / decoding operations disclosed in the present disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, and any of the encoders or decoders may be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0048] In some embodiments, at least a portion of the components of the source device 12 (e.g., video source 18, video encoder 20 or the components included in video encoder 20 described below with reference to Figure 2 and output interface 22) and / or the components of the destination device 14 (e.g., input interface 28, video decoder 30 or the following reference Figure 3At least some of the components described as being included in video decoder 30, as well as in display device 34, may operate in a cloud computing service network such as software as a service (SaaS), platform as a service (PaaS), or infrastructure as a service (IaaS), where the cloud computing service network may provide software, platform, and / or infrastructure. In some embodiments, one or more components not included in the cloud computing service network in source device 12 and / or target device 14 may be provided in one or more client devices, and the one or more client devices may communicate with server computers in the cloud computing service network via a wireless communication network (e.g., a cellular communication network, a short-range wireless communication network, or a global navigation satellite system (GNSS) communication network) or a wired communication network (e.g., a local area network (LAN) communication network or a power line communication (PLC) network). In an embodiment, at least some of the operations described herein may be implemented as cloud-based services provided by one or more server computers, where the one or more server computers are implemented by at least some of the components in source device 12 and / or at least some of the components in target device 14 in the cloud computing service network; and one or more other operations described herein may be implemented by one or more client devices. In some embodiments, the cloud computing service network may be a private cloud, a public cloud, or a hybrid cloud. Without departing from the scope of the present disclosure, terms such as "cloud", "cloud computing", "cloud-based", etc. may be used interchangeably as appropriate herein. It should be understood that the present disclosure is not limited to implementation in the above-described cloud computing service network. Instead, the present disclosure may also be implemented in any other type of computing environment known currently or developed in the future.

[0049] Figure 2 FIG. 4 is a block diagram showing an exemplary video encoder 20 according to some embodiments described in the present application. Video encoder 20 may perform intra prediction coding and inter prediction coding on video blocks within a video frame. Intra prediction coding relies on spatial prediction to reduce or remove spatial redundancy in the video data within a given video frame or picture. Inter prediction coding relies on temporal prediction to reduce or remove temporal redundancy in the video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video coding and decoding, the term "frame" may be used as a synonym for the term "image" or "picture".

[0050] As Figure 2As shown in FIG. 0, video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, adder 50, transform processing unit 52, quantization unit 54, and entropy encoding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, segmentation unit 45, intra prediction processing unit 46, and intra block copy (BC) unit 48. In some embodiments, video encoder 20 further includes inverse quantization unit 58, inverse transform processing unit 60, and adder 62 for video block reconstruction. A loop filter such as a deblocking filter 63 may be located between adder 62 and DPB 64 to filter block boundaries to remove blocking artifacts from the reconstructed video. In addition to the deblocking filter, another loop filter (e.g., sample adaptive offset (SAO) filter, cross-component sample adaptive offset (CCSAO) filter, and / or adaptive loop filter (ALF)) may be used to filter the output of adder 62. It should be noted that for the CCSAO technique, the present application is not limited to the embodiments described herein, but may also be applied to cases where an offset is selected for any one of the luminance component, Cb chrominance component, and Cr chrominance component based on any one of the luminance component, Cb chrominance component, and Cr chrominance component to modify the any one of the luminance component, Cb chrominance component, and Cr chrominance component based on the selected offset. In addition, it should be noted that the first component mentioned herein may be any one of the luminance component, Cb chrominance component, and Cr chrominance component, the second component mentioned herein may be any one of the luminance component, Cb chrominance component, and Cr chrominance component, and the third component mentioned herein may be the remaining one of the luminance component, Cb chrominance component, and Cr chrominance component. In some examples, the loop filter may be omitted, and the decoded video block may be directly provided to DPB 64 by adder 62. Video encoder 20 may take the form of fixed or programmable hardware units, or may be distributed among one or more of the illustrated fixed or programmable hardware units.

[0051] Video data memory 40 may store video data to be encoded by components of video encoder 20. The video data in video data memory 40 may be obtained, for example, from Figure 1 video source 18 shown in FIG. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use by video encoder 20 (e.g., in intra or inter prediction coding modes) when encoding video data. Video data memory 40 and DPB 64 may be formed of any of a variety of memory devices. In various examples, video data memory 40 may be on-chip with other components of video encoder 20, or off-chip relative to those components.

[0052] As Figure 2As shown, after receiving video data, the segmentation unit 45 within the prediction processing unit 41 segments the video data into video blocks. This segmentation may also include segmenting a video frame into strips, tiles (e.g., a set of video blocks), or other larger coding units (CUs) according to a predefined splitting structure associated with the video data, such as a quadtree (QT) structure. A video frame is or can be regarded as a two-dimensional sample array or matrix having sample values. The samples in the array can also be referred to as pixels or picture elements (pels). The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. For example, a video frame can be divided into multiple video blocks by using QT segmentation. A video block is again or can be regarded as a two-dimensional sample array or matrix having sample values, but its dimensions are smaller than those of the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. By, for example, iteratively using QT segmentation, binary tree (BT) segmentation, or ternary tree (TT) segmentation, or any combination thereof, the video block can be further segmented into one or more block partitions or sub-blocks (which can again form blocks). It should be noted that the term "block" or "video block" as used herein can be a part of a frame or picture, especially a rectangular (square or non-square) part. Referring to, for example, HEVC and VVC, a block or video block can be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU) and / or can be or correspond to the corresponding block (e.g., coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB)) and / or sub-block.

[0053] The prediction processing unit 41 can select one of multiple feasible prediction coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one of multiple intra prediction coding modes or one of multiple inter prediction coding modes. The prediction processing unit 41 can provide the resulting intra prediction coding block or inter prediction coding block to the adder 50 to generate a residual block, and provide it to the adder 62 to reconstruct the coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements (such as motion vectors, intra mode indicators, segmentation information, and other such syntax information) to the entropy coding unit 56.

[0054] To select a suitable intra prediction coding mode for the current video block, the intra prediction processing unit 46 within the prediction processing unit 41 may perform intra prediction coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block to be encoded to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter prediction coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 may execute multiple coding channels, for example, to select a suitable coding mode for each block of video data.

[0055] In some embodiments, the motion estimation unit 42 determines an inter prediction mode for the current video frame by generating a motion vector according to a predetermined pattern within the video frame sequence, the motion vector indicating the displacement of a video block within the current video frame relative to a prediction block within a reference video frame. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector that estimates the motion for the video block. For example, the motion vector may indicate the displacement of a video block within the current video frame or picture relative to a prediction block within a reference frame related to the current block being encoded within the current frame. The predetermined pattern may designate video frames in the sequence as P frames or B frames. The intra BC unit 48 may determine a vector (e.g., a block vector) for intra BC coding in a manner similar to the motion vector determined by the motion estimation unit 42 for inter prediction, or may utilize the motion estimation unit 42 to determine the block vector.

[0056] In terms of pixel difference, the prediction block for a video block may be or may correspond to a block or reference block of a reference frame considered to closely match the video block to be encoded, and the pixel difference may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some embodiments, the video encoder 20 may calculate values for sub - integer pixel positions of the reference frames stored in the DPB 64. For example, the video encoder 20 may interpolate values for quarter - pixel positions, eighth - pixel positions, or other fractional pixel positions of the reference frames. Thus, the motion estimation unit 42 may perform a motion search with respect to full - pixel positions and fractional pixel positions and output a motion vector with fractional - pixel accuracy.

[0057] The motion estimation unit 42 calculates a motion vector for a video block in an inter - prediction - coded frame by comparing the position of the video block with the position of a prediction block of a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1), each of the first reference frame list and the second reference frame list identifying one or more reference frames stored in the DPB 64. The motion estimation unit 42 sends the calculated motion vector to the motion compensation unit 44 and then to the entropy coding unit 56.

[0058] The motion compensation performed by the motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by the motion estimation unit 42. After receiving the motion vector for the current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists in the reference frame list, retrieve the prediction block from the DPB 64, and forward the prediction block to the adder 50. Then, the adder 50 forms a residual video block of pixel differences by subtracting the pixel values of the prediction block provided by the motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel differences forming the residual video block may include luminance component differences or chrominance component differences or both. The motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use by the video decoder 30 when decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vectors for identifying the prediction blocks, any flags indicating the prediction mode, or any other syntax information described herein. It should be noted that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes.

[0059] In some embodiments, the intra BC unit 48 may generate vectors and obtain prediction blocks in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, but the prediction blocks are in the same frame as the current block being encoded, and the vectors are referred to as block vectors rather than motion vectors. Specifically, the intra BC unit 48 may determine the intra prediction mode to be used for encoding the current block. In some examples, the intra BC unit 48 may, for example, use various intra prediction modes to encode the current block during a separate encoding pass and test their performance through rate-distortion analysis. Next, the intra BC unit 48 may select a suitable intra prediction mode to use from among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values for the various tested intra prediction modes using rate-distortion analysis and select the intra prediction mode having the best rate-distortion characteristics among the tested modes as the suitable intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original unencoded block that was encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. The intra BC unit 48 may calculate a ratio based on the distortion and rate for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0060] In other examples, the intra BC unit 48 may perform such functions for intra BC prediction according to the embodiments described herein, in whole or in part, using the motion estimation unit 42 and the motion compensation unit 44. In either case, for intra block copy, in terms of pixel differences, the predicted block may be a block that is considered to closely match the block to be encoded, the pixel differences may be determined by SAD, SSD, or other difference metrics, and identifying the predicted block may include calculating values for sub-integer pixel positions.

[0061] Regardless of whether the predicted block is from the same frame according to intra prediction or from a different frame according to inter prediction, the video encoder 20 may form a pixel difference value by subtracting the pixel values of the predicted block from the pixel values of the current video block being encoded, thereby forming a residual video block. The pixel difference values for forming the residual video block may include both a luminance component difference and a chrominance component difference.

[0062] As an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44 or the intra block copy prediction performed by the intra BC unit 48 as described above, the intra prediction processing unit 46 may perform intra prediction on the current video block. Specifically, the intra prediction processing unit 46 may determine an intra prediction mode for encoding the current block. To this end, the intra prediction processing unit 46 may, for example, encode the current block using various intra prediction modes during a separate encoding pass, and the intra prediction processing unit 46 (or in some examples, the mode selection unit) may select a suitable intra prediction mode from the tested intra prediction modes to use. The intra prediction processing unit 46 may provide information indicating the intra prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode into the bitstream.

[0063] After the prediction processing unit 41 determines a predicted block for the current video block via inter prediction or intra prediction, the adder 50 forms a residual video block by subtracting the predicted block from the current video block. The residual video data in the residual block may be included in one or more TUs and provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform (such as a discrete cosine transform (DCT) or a conceptually similar transform).

[0064] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may subsequently perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.

[0065] After quantization, an entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy encoding method or technique. Then, the encoded bitstream can be sent to a video decoder 30 as shown in Figure 1 or archived in a storage device 32 as shown in Figure 1 for later sending to or retrieval by the video decoder 30. The entropy encoding unit 56 can also entropy encode motion vectors and other syntax elements for the current video frame being encoded.

[0066] An inverse quantization unit 58 and an inverse transform processing unit 60 respectively apply inverse quantization and inverse transform to reconstruct a residual video block in the pixel domain for generating a reference block for predicting other video blocks. As pointed out above, a motion compensation unit 44 can generate a motion compensation prediction block from one or more reference blocks of frames stored in a DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.

[0067] An adder 62 adds the reconstructed residual block to the motion compensation prediction block generated by the motion compensation unit 44 to generate a reference block for storage in the DPB 64. Then, the reference block can be used as a prediction block by an intra BC unit 48, a motion estimation unit 42, and the motion compensation unit 44 to perform inter prediction on another video block in a subsequent video frame.

[0068] Figure 3 is a block diagram showing an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 can perform a decoding process that is substantially inverse to the encoding process described above in connection with Figure 2 the video encoder 20. For example, the motion compensation unit 82 can generate prediction data based on motion vectors received from the entropy decoding unit 80, while the intra prediction unit 84 can generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit 80.

[0069] In some examples, units of the video decoder 30 may be assigned tasks to perform embodiments of the present application. Additionally, in some examples, embodiments of the present disclosure may be distributed among one or more of the units of the video decoder 30. For example, the intra BC unit 85 may perform embodiments of the present application either alone or in combination with other units of the video decoder 30 (such as the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80). In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 (such as the motion compensation unit 82).

[0070] The video data memory 79 may store video data to be decoded by other components of the video decoder 30, such as an encoded video bitstream. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, from a local video source (such as a camera), via wired or wireless network communication of video data, or by accessing a physical data storage medium (such as a flash drive or a hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use by the video decoder 30 (such as in an intra or inter prediction coding mode) when decoding video data. The video data memory 79 and the DPB 92 may be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are depicted in Figure 3 as two different components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be on-chip with other components of the video decoder 30 or off-chip relative to those components.

[0071] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantized coefficients, motion vectors, or intra prediction mode indicators, as well as other syntax elements. The entropy decoding unit 80 then forwards the motion vectors or intra prediction mode indicators, as well as other syntax elements, to the prediction processing unit 81.

[0072] When a video frame is encoded as an intra-predicted coded (I) frame or for intra-coded prediction blocks in other types of frames, the intra prediction unit 84 of the prediction processing unit 81 may generate prediction data for a video block of the current video frame based on the intra prediction mode signaled and reference data from previously decoded blocks of the current frame.

[0073] When a video frame is encoded as an inter-predicted coded (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for a video block of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each of the prediction blocks may be generated from a reference frame within one of the reference frame lists. The video decoder 30 may construct the reference frame lists, i.e., list 0 and list 1, using a default construction technique based on the reference frames stored in the DPB 92.

[0074] In some examples, when a video block is encoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within the reconstructed region of the same picture as the current video block defined by the video encoder 20.

[0075] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for a video block of the current video frame by parsing the motion vectors and other syntax elements, and then use the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) for encoding a video block of the video frame, the inter prediction frame type (e.g., B or P), the construction information for one or more of the reference frame lists for the frame, the motion vector for each inter-predicted coded video block of the frame, the inter prediction status for each inter-predicted coded video block of the frame, and other information for decoding a video block in the current video frame.

[0076] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as flags, to determine that the current video block is predicted using the intra BC mode, the construction information for which video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, the block vector for each intra BC predicted video block of the frame, the intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding a video block in the current video frame.

[0077] The motion compensation unit 82 may also perform interpolation using an interpolation filter such as the one used by the video encoder 20 during encoding of a video block to calculate the interpolated values for sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 based on the received syntax elements and use these interpolation filters to generate the prediction block.

[0078] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.

[0079] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vectors and other syntax elements, the adder 90 reconstructs the decoded video block for the current video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter 91 (e.g., deblocking filter, SAO filter, CCSAO filter, and / or ALF) may be located between the adder 90 and the DPB 92 to further process the decoded video block. In some examples, the loop filter 91 may be omitted and the decoded video block may be provided directly by the adder 90 to the DPB 92. Then, the decoded video blocks in a given frame are stored in the DPB 92, which stores the reference frames for subsequent motion compensation of the next video blocks. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later presentation on a display device (e.g., Figure 1 display device 34).

[0080] In a typical video coding and decoding process, a video sequence generally includes an ordered collection of frames or pictures. Each frame may include three arrays of samples, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other instances, a frame may be monochrome and thus include only one two-dimensional array of luminance samples.

[0081] As Figure 4AAs shown, video encoder 20 (or more specifically, splitting unit 45) generates an encoded representation of a frame by first splitting the frame into a set of CTUs. A video frame may include an integer number of CTUs that are sequentially ordered from left to right and top to bottom in raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in the sequence parameter set such that all CTUs in the video sequence have the same size of one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application need not be limited to a specific size. As Figure 4B As shown, each CTU may include one CTB of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the nature of different types of units that encode pixel blocks and how the video sequence can be reconstructed at video decoder 30, including inter-prediction or intra-prediction, intra-prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding the samples of the coding tree block. The coding tree block may be a block of N×N samples.

[0082] To achieve better performance, video encoder 20 may recursively perform tree splitting, such as binary tree splitting, ternary tree splitting, quadtree splitting, or a combination thereof, on the coding tree blocks of a CTU and divide the CTU into smaller CUs. As Figure 4C As depicted, a 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs with a block size of 16×16. The two 16×16 CUs 430 and 440 are each further divided into four CUs with a block size of 8×8. Figure 4D depicts a quadtree data structure showing the final result of the splitting process of CTU 400 as depicted in Figure 4C Each leaf node of the quadtree corresponds to a CU with a corresponding size ranging from 32×32 to 8×8. Similar to the CTU depicted in Figure 4B Each CU may include a CB of luminance samples of a frame of the same size and two corresponding coding blocks of chrominance samples, and syntax elements for encoding the samples of the coding blocks. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and a syntax structure for encoding the samples of the coding block. It should be noted that Figure 4C and Figure 4DThe quadtree partitioning depicted is for illustrative purposes only, and a CTU can be split into multiple CUs based on quadtree / tritree / binary tree partitioning to adapt to varying local characteristics. In a multi-type tree structure, a CTU is partitioned according to a quadtree structure, and each quadtree leaf CU can be further partitioned according to binary and tritree structures. As Figure 4E shown, there are five possible partitioning types for a coding block with width W and height H, namely, quaternary partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.

[0083] In some embodiments, the video encoder 20 may further partition the coding block of a CU into one or more (M×N) PBs. A PB is a rectangular (square or non-square) sample block to which the same prediction (inter-frame or intra-frame) is applied. The PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PB. In a monochrome picture or a picture with three separate color planes, the PU may include a single PB and a syntax structure for predicting the PB. The video encoder 20 may generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma PB, Cb PB, and Cr PB of each PU of the CU.

[0084] The video encoder 20 may use intra-frame prediction or inter-frame prediction to generate a predicted block for the PU. If the video encoder 20 uses intra-frame prediction to generate a predicted block for the PU, the video encoder 20 may generate the predicted block for the PU based on the decoded samples of the frame associated with the PU. If the video encoder 20 uses inter-frame prediction to generate a predicted block for the PU, the video encoder 20 may generate the predicted block for the PU based on the decoded samples of one or more frames other than the frame associated with the PU.

[0085] After the video encoder 20 generates a predicted luma block, a predicted Cb block, and a predicted Cr block for one or more PUs of the CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predicted luma block of the CU from the original luma coding block of the CU, such that each sample in the luma residual block of the CU indicates the difference between the luma sample in one of the predicted luma blocks of the CU and the corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, such that each sample in the Cb residual block of the CU indicates the difference between the Cb sample in one of the predicted Cb blocks of the CU and the corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between the Cr sample in one of the predicted Cr blocks of the CU and the corresponding sample in the original Cr coding block of the CU.

[0086] In addition, as Figure 4CAs shown in, video encoder 20 may use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block of a CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. The TU of a CU may include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the samples of the transform block. Thus, each TU of a CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.

[0087] Video encoder 20 may apply one or more transforms to the luminance transform block of a TU to generate a luminance coefficient block for the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalars. Video encoder 20 may apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0088] After generating the coefficient block (e.g., luminance coefficient block, Cb coefficient block, or Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to the process by which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 may entropy code the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream including a bit sequence that forms a representation of the encoded frame and associated data, and the bitstream is stored in storage device 32 or sent to target device 14.

[0089] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data at least in part based on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally inverse to the encoding process performed by the video encoder 20. For example, the video decoder 30 may perform an inverse transform on a coefficient block associated with a TU of the current CU to reconstruct a residual block associated with the TU of the current CU. The video decoder 30 also reconstructs the encoded block of the current CU by adding the samples of the prediction block of the PU of the current CU to the corresponding samples of the transform block of the TU of the current CU. After reconstructing the encoded block of each CU for a frame, the video decoder 30 may reconstruct the frame.

[0090] As described above, video coding mainly uses two modes (i.e., intra-frame prediction (or intra prediction) and inter-frame prediction (or inter prediction)) to achieve video compression. It should be noted that IBC can be regarded as intra prediction or a third mode. Between the two modes, since motion vectors are used to predict the current video block based on reference video blocks, inter prediction contributes more to the coding efficiency than intra prediction.

[0091] However, with continuously improving video data capture technologies and finer video block sizes for preserving details in video data, the amount of data required to represent the motion vectors for the current frame has also increased significantly. One way to overcome this challenge benefits from the fact that not only a set of neighboring CUs in both the spatial domain and the temporal domain have similar video data for prediction purposes, but also the motion vectors between these neighboring CUs are similar. Therefore, the motion information of spatial neighboring CUs and / or temporal co-located CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (which is also referred to as the "motion vector prediction value" (MVP) of the current CU) by exploring their spatial and temporal correlations.

[0092] Instead of encoding the actual motion vector of the current CU determined by the motion estimation unit 42 into the video bitstream as described above in conjunction with Figure 2 the motion vector difference (MVD) for the current CU is generated by subtracting the motion vector prediction value of the current CU from the actual motion vector of the current CU. By doing so, it is not necessary to encode the motion vectors determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent the motion information in the video bitstream can be significantly reduced.

[0093] Similar to the process of selecting a prediction block in a reference frame during inter prediction of a coding block, both the video encoder 20 and the video decoder 30 need to adopt a set of rules for constructing a motion vector candidate list (also referred to as a "merge list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally collocated CUs of the current CU, and then selecting a member from the motion vector candidate list as the motion vector prediction value for the current CU. By doing so, it is not necessary to send the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector prediction value within the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector prediction value within the motion vector candidate list to encode and decode the current CU.

[0094] This disclosure relates to video coding, decoding, and compression. More specifically, this disclosure relates to methods and apparatuses for improving the coding and decoding efficiency of an adaptive loop filter (ALF) and a cross-component adaptive loop filter (CCALF).

[0095] Various video coding and decoding techniques can be used to compress video data. Video coding and decoding is performed according to one or more video coding and decoding standards. For example, nowadays, some well-known video coding and decoding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which are jointly developed by ISO / IEC MPEG and ITU-T VCEG. AV1 (Alliance for Open Media Video 1) is developed by the Alliance for Open Media (AOM) as a successor to its previous standard VP9. Audio Video Coding Standard (AVS) (which refers to digital audio and digital video compression standards) is another series of video compression standards developed by the Audio and Video Coding Standard Workgroup of China. Most existing video coding and decoding standards are based on the well-known hybrid video coding framework, that is, using block-based prediction methods (e.g., inter prediction, intra prediction) to reduce the redundancy present in a video image or sequence, and using transform coding to compress the energy of the prediction error. An important goal of video coding and decoding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing video quality degradation.

[0096] The first-generation AVS standards include the Chinese national standard "Information Technology - Advanced Audio Video Coding - Part 2: Video" (referred to as AVS1) and "Information Technology - Advanced Audio Video Coding - Part 16: Broadcast Television Video" (referred to as AVS+). Compared with the MPEG-2 standard, the first-generation AVS standards can provide approximately 50% bitrate savings at the same perceived quality. The video part of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second-generation AVS standards include the Chinese national standard "Information Technology - Efficient Multimedia Coding" (referred to as AVS2) series, which is mainly targeted at the transmission of additional HD TV programs. The coding and decoding efficiency of AVS2 is twice that of AVS+. In May 2016, AVS2 was released as a Chinese national standard. At the same time, the video part of the AVS2 standard was submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international application standard. The AVS3 standard is a new-generation video coding and decoding standard for UHD video applications, aiming to exceed the coding and decoding efficiency of the latest international standard HEVC. In March 2019, at the 68th AVS conference, the AVS3-P2 baseline was completed, which provides approximately 30% bitrate savings over the HEVC standard. Currently, there is a reference software called the High-Performance Model (HPM), maintained by the AVS working group to demonstrate the reference implementation of the AVS3 standard.

[0097] Like HEVC, the AVS3 standard is built on a block-based hybrid video coding and decoding framework. Figure 5 The block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (referred to as coding units (CUs)). Different from HEVC which partitions blocks only based on quadtrees, in AVS3, a coding tree unit (CTU) is split into multiple CUs to adapt to different local characteristics based on quadtrees / binary trees / extended quadtrees. Additionally, the concept of multiple partition unit types in HEVC is removed, that is, the splitting of CUs, prediction units (PUs), and transform units (TUs) does not exist in AVS3; instead, each CU always serves as the basic unit for both prediction and transformation without further partitioning. In the tree partition structure of AVS3, a CTU is first partitioned based on a quadtree structure. Then, each quadtree leaf node can be further partitioned based on binary tree and extended quadtree structures. In Figure 5In [the encoder], spatial prediction and / or temporal prediction can be performed. Spatial prediction (or “intra prediction”) uses pixels of samples (referred to as reference samples) from already decoded neighboring blocks in the same video picture / strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also referred to as “inter prediction” or “motion-compensated prediction”) uses the reconstructed pixels from already decoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. Also, if multiple reference pictures are supported, a reference picture index is additionally sent, which is used to identify which reference picture in the reference picture buffer the temporal prediction signal is from. After performing spatial and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode, for example, based on the rate-distortion optimization method. Then, the predicted block is subtracted from the current video block; and the prediction residual is decorrelated and then quantized using a transform. The quantized residual coefficients are dequantized and inverse-transformed to form a reconstructed residual, which is then added back to the predicted block to form the reconstructed signal of the CU. Before placing the reconstructed CU in the reference picture buffer and using it as a reference for decoding future video blocks, further loop filters, such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF), can be applied to the reconstructed CU. To form the output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit for further compression and packaging.

[0098] The first version of the HEVC standard was completed in October 2013, which provides approximately 50% bitrate savings or equivalent perceptual quality compared to the previous generation video coding standard H.264 / MPEG-4 AVC. Although the HEVC standard provides significant coding improvements over its predecessor, there is evidence that coding efficiency better than HEVC can be achieved using additional coding tools. Based on this, both VCEG and MPEG have started exploring new coding techniques for future video coding standardization. The ITU-T VCEG and ISO / IEC MPEG established a Joint Video Exploration Team (JVET) in October 2015 to start major research on advanced techniques that can significantly improve coding efficiency. JVET maintains a reference software called the Joint Exploration Model (JEM) by integrating multiple additional coding tools on top of the HEVC Test Model (HM).

[0099] In October 2017, ITU-T and ISO / IEC issued a Call for Proposals (CfP) on video compression with capabilities beyond HEVC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating a compression efficiency improvement of approximately 40% over HEVC. Based on such evaluation results, JVET launched a new project to develop a new generation of video coding standard named Versatile Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate the reference implementation of the VVC standard.

[0100] Similar to HEVC, VVC is built on a block-based hybrid video coding framework. Figure 5 The block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (referred to as coding units (CUs)). In VTM-1.0, a CU can be up to 128×128 pixels. However, different from HEVC which partitions blocks only based on quad-trees, in VVC, a coding tree unit (CTU) is split into multiple CUs to adapt to different local characteristics based on quad- / bi- / tri-trees. Additionally, the concept of multiple partition unit types in HEVC is removed, i.e., there is no longer a division of CUs, prediction units (PUs), and transform units (TUs) in VVC; instead, each CU always serves as the basic unit for both prediction and transform without further partitioning. In the multi-type tree structure, a CTU is first partitioned according to the quad-tree structure. Then, each quad-tree leaf node can be further partitioned according to the binary tree structure and the ternary tree structure. As Figure 4E shown, there are five splitting types, quaternary split, horizontal binary split, vertical binary split, horizontal ternary split, and vertical ternary split. In Figure 5In it, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra prediction") uses pixels of samples (referred to as reference samples) from already decoded neighboring blocks in the same video picture / strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also referred to as "inter prediction" or "motion-compensated prediction") uses the reconstructed pixels from already decoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal of a given CU is typically represented by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. Similarly, if multiple reference pictures are supported, a reference picture index is additionally sent, which is used to identify which reference picture in the reference picture storage the temporal prediction signal comes from. After performing spatial and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode, for example, based on the rate-distortion optimization method. Then, the prediction block is subtracted from the current video block; and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are dequantized and inverse-transformed to form a reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Before placing the reconstructed CU in the reference picture storage area and using it to decode future video blocks, further loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF), can be applied to the reconstructed CU. To form the output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit for further compression and packing to form the bitstream.

[0101] Figure 6 A general block diagram of a block-based video decoder is given. First, the video bitstream is entropy decoded at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (in the case of intra coding / decoding) or the temporal prediction unit (in the case of inter coding / decoding) to form a prediction block. The residual transform coefficients are sent to the dequantization unit and the inverse transform unit to reconstruct the residual block. Then, the prediction block and the residual block are added together. The reconstructed block can be further loop filtered and then stored in the reference picture storage area. Then, the reconstructed video in the reference picture storage area is sent out to drive a display device and is used to predict future video blocks.

[0102] The focus of this disclosure is on improving the adaptive loop filter (ALF) and the cross-component adaptive loop filter (CCALF). The relevant knowledge is elaborated in detail in the following sections.

[0103] ALF in VVC

[0104] Filter Shape, Linear Filtering, and Adaptive Clipping

[0105] In VVC, ALF is applied to the output samples of SAO. As Figure 7 shown, two filter shapes are supported for the luminance and chrominance components respectively: 7×7 rhombus and 5×5 rhombus. In Figure 7 , each square corresponds to a luminance or chrominance sample, and the central square corresponds to the current sample to be filtered. The filter coefficients are point-symmetric, and each integer filter coefficient is represented with 7-bit fractional precision. Additionally, the sum of the coefficients of a filter is equal to 128, which is the 7-bit fractional precision fixed-point representation of 1.0:

[0106]

[0107] where, for the 7×7 and 5×5 filter shapes, the number of coefficients N is equal to 13 and 7 respectively.

[0108] The value of the filtered sample at coordinates (x,y) is obtained by applying the coefficient c i to the value of the reconstructed sample R(x,y) as follows:

[0109]

[0110] where, (x + x i , y + y i ) and (x - x i , y - y i ) are the coordinates of the reconstructed samples corresponding to the i-th coefficient c i . Due to the constraints in Equation (1), Equation (2) can be written as:

[0111]

[0112] In VVC, the possibility of clipping the difference between the values of neighboring samples and the current sample to be filtered is added in Equation (3) as follows:

[0113]

[0114] where,

[0115]

[0116] b i is the clipping parameter of the coefficient c i , which is determined by the clipping index d i . b i is obtained as follows:

[0117]

[0118] Among them, BD is the depth of the sample point position, and d i can be 0, 1, 2, or 3.

[0119] Luma sub-block level filter adaptation

[0120] In VVC, sub-block level filter adaptation is only applied to the luma component. Each 4×4 luma block is classified based on its directionality and 2D Laplacian activity. First, the sample point gradient values in the horizontal, vertical, and two diagonal directions are calculated:

[0121] H k,l = |2R(k, l) - R(k - 1, l) - R(k + 1, l)|,

[0122] V k,l = |2R(k, l) - R(k, l - 1) - R(k, l + 1)|,

[0123] D0 k,l = |2R(k, l) - R(k - 1, l - 1) - R(k + 1, l + 1)|,

[0124] D1 k,l = |2R(k, l) - R(k - 1, l + 1) - R(k + 1, l - 1)|. (7)

[0125] Based on the sample point gradient, the sub-block horizontal gradient g h , vertical gradient g v , and the two diagonal gradients g d0 and g d1 are calculated as

[0126]

[0127] The indices i and j refer to the coordinates of the top-left sample point in the 4×4 luma block. It can be seen from Equation (8) that the sum of the sample point gradients within the 10×10 luma window covering the target 4×4 block is used to classify the block. To reduce the complexity, only the gradients of every other sample point in the 10×10 window are calculated, as Figure 8 shown. The other sample point gradient values are set to 0.

[0128] Secondly, to assign the directionality D, the ratio of the maximum to the minimum of the horizontal gradient and the vertical gradient of the sub-block

[0129]

[0130] , and the ratio of the maximum to the minimum of the two sub-block diagonal gradients

[0131]

[0132] Compared with a set of thresholds t1 and t2:

[0133] Step 1: If and then D is set to 0.

[0134] Step 2: If then the directivity D is calculated in Step 3, otherwise it is calculated in Step 4.

[0135] Step 3: If then D is set to 2, otherwise D is set to 1.

[0136] Step 4: If then D is set to 4, otherwise D is set to 3.

[0137] Each subsequent step in the above D calculation is only executed when no value has been assigned to D in the previous steps. Third, the activity value A is calculated as

[0138]

[0139] A is further mapped to the range from 0 to 4: where, {Q n} = {0, 1, 2, 2, 2, 2, 2, 3, 3, 3, 3, 3, 3, 3, 3, 4}. Finally, each 4×4 luminance block is classified into one of the following 25 classes:

[0140]

[0141] Each class can be assigned its own filter.

[0142] Before filtering each 4×4 luminance block, a geometric transformation, such as a 90-degree rotation, diagonal or vertical flip, is applied to the filter coefficients, as Figure 9 shown, depending on the sub-block gradient value, as specified in Table 1.

[0143] Table 1 Geometric Transformations Based on Sub-block Gradient Values

[0144] Sub-block gradient value Transformation <![CDATA[g d1 <g d0 and g h <g v > No transformation <![CDATA[g d1 <g d0 and g v ≤g h > Diagonal flip <![CDATA[g do ≤g d1 and g h <g v > Vertical flip <![CDATA[g do ≤g d1 and g v ≤g h > 90-degree rotation

[0145] Coding Tree Block-level Filter Adaptation

[0146] In addition to the luminance 4×4 block-level filter adaptation, ALF also supports CTB-level filter adaptation. The luminance CTB can use the filter bank calculated for the current strip, or one of the filter banks calculated for the coded strips. It can also use one of the 16 offline-trained filter banks. In each luminance CTB, which filter in the selected filter bank should be applied to each 4×4 block is determined by the class C calculated for that block according to Equation (12).

[0147] The chroma is adaptively filtered only using CTB-level filters. In a slice, the chroma component can use up to 8 filters. Each CTB can select one of these filters.

[0148] Syntax design

[0149] The filter coefficients and truncation indices are carried in the ALF APS. The ALF APS can include up to 8 chroma filters and one luma filter bank (which can contain up to 25 filters). Each of the 25 luma classes also includes an index i C . Classes with the same index i C share the same filter. By combining different classes, the number of bits required to represent the filter coefficients can be reduced. The absolute value of the filter coefficients is represented using a Golomb code of order 0, followed by the sign bit of the non-zero coefficients. When truncation is enabled, a two-bit fixed-length code is also used to signal the truncation index for each filter coefficient. The storage required for the ALF coefficients and truncation indices within the APS is at most 3480 bits. The decoder can use up to 8 ALF APSs simultaneously.

[0150] The filter control syntax elements include two types of information. First, an ALF on / off flag is signaled at the sequence, picture, slice, and CTB levels. The chroma ALF can only be enabled at the corresponding level if the luma ALF is enabled at the picture and slice levels. Second, if the ALF is enabled at the picture, slice, and CTB levels, the filter usage information is signaled at that level. The referenced ALF APS ID is encoded and decoded at the slice level, or at the picture level if all slices within the picture use the same APS. The luma component can reference up to 7 ALF APSs, while the chroma component can reference up to 1 ALF APS. For a luma CTB, an index is signaled to indicate which ALF APS or offline-trained luma filter bank is used. For a chroma CTB, the index indicates which filter within the referenced APS is used.

[0151] Row buffer reduction

[0152] To reduce the storage requirements of the ALF, VVC employs row buffer boundary handling. In VVC, the row buffer boundary is located 4 luma samples and 2 chroma samples above the horizontal CTU boundary. When the ALF is applied to samples on one side of the row buffer boundary, the samples on the other side of the row buffer boundary cannot be used.

[0153] ALF in ECM

[0154] ALF deletion simplification

[0155] The ALF gradient subsampling and ALF virtual boundary processing are removed. The block size for classification is reduced from 4×4 to 2×2. The filter sizes for luminance and chrominance, for which the ALF coefficients are signaled, are increased to 9×9.

[0156] ALF with fixed filters

[0157] To filter the luminance samples, three different classifiers (C0, C1, and C2) and three different filter banks (F0, F1, and F2) are used. The groups F0 and F1 contain fixed filters whose coefficients are trained for the classifiers C0 and C1. The filter coefficients in F2 are signaled. Which filter in group F i is used for a given sample is determined by the class C i assigned to that sample by the classifier C i decided.

[0158] Filtering

[0159] First, two 13×13 diamond-shaped fixed filters F0 and F1 are applied to obtain two intermediate samples R0(x,y) and R1(x,y). Then, F2 is applied to R0(x,y), R1(x,y), neighboring samples, and the samples before the deblocking filter (DBF) to obtain the filtered sample

[0160]

[0161] where, f i,j is the clipped difference between the neighboring sample and the current sample R(x,y), g i is the clipped difference between R i-20 (x,y) and the current sample R(x,y), h i,j is the clipped difference between the neighboring sample before the DBF and the current sample R(x,y). The filter coefficients c i , i = 0, … 24 are signaled. The filter shape of F2 is as Figure 10 shown.

[0162] Classification

[0163] Based on the directionality D i and activity a class C i is assigned to each 2×2 block:

[0164]

[0165] where, M D,i represents the total number of directionality D i s.

[0166] Similar to VVC, the horizontal, vertical, and two diagonal gradient values of each sample point are calculated using a 1-D Laplacian operator. The sum of the sample point gradients within a 4×4 window covering the target 2×2 block is used for classifier C0, and the sum of the sample point gradients within a 12×12 window is used for classifiers C1 and C2. The sums of the horizontal, vertical, and two diagonal gradients are denoted as and the directionality D i is determined by comparing

[0167]

[0168] with a set of thresholds. The directionality D2 in VVC is obtained using thresholds 2 and 4.5. For D0 and D1, the horizontal / vertical edge strength and the diagonal strength are calculated using the thresholds Th = [1.25, 1.5, 2, 3, 4.5, 8]. If then the edge strength is 0; otherwise, is the largest integer that satisfies If then the edge strength is 0; otherwise, is the largest integer that satisfies When , i.e., when the horizontal / vertical edges are dominant, D i is obtained using Table 2(a); otherwise, when the diagonal edges are dominant, D i is obtained using Table 2(b).

[0169] Table 2. and to D i mapping

[0170]

[0171] To obtain the sum A of the vertical and horizontal gradients i is mapped to the range from 0 to n, where, for n is equal to 4, and for and is equal to 15.

[0172] In ALF_APS, at most 4 luminance filter banks are transmitted by signal, and each bank can have at most 25 filters.

[0173] Alternative 2×2 ALF classifier

[0174] The classification in ALF is extended by additional alternative classifiers. For the luminance filter bank transmitted with signals, a flag will be transmitted with signals to indicate whether the alternative classifier is applied. Geometric transformation is not applied to the alternative band classifier. When applying the band-based classifier, first the sum of the sample values of a 2×2 luminance block is calculated. Then the class index is calculated as follows,

[0175] class_index = (sum * 25) >> (sample bit depth + 2) (16).

[0176] CCALF in VVC

[0177] Filter shape and precision

[0178] CCALF uses luminance sample values to refine the chrominance sample values in the ALF process. As Figure 11 shown, the linear filtering operation takes the luminance sample values as input and generates correction values for the chrominance sample values. The correction is generated independently for each of the following

[0179] chrominance component i, i ∈ {Cb, Cr} and can be expressed as: where (x, y) is the sample position of chrominance component i, (x C , y C ) is the luminance sample position derived from (x, y), (x0, y0) is the filter support offset around (x C , y C ), and S i is the filter support area of chrominance component i in luminance. The luminance position (x C , y C ) is determined based on the spatial scaling factor between the luminance plane and the chrominance plane. The sample values in the luminance support area are also the input of the ALF luminance stage and correspond to the output of the SAO stage.

[0180] As Figure 12 shown, the CCALF filter has a diamond shape. As Figure 12 shown, for a 4:2:0 video sequence with chrominance position type 0 (i.e., when the chrominance samples are co-located with the even columns of the luminance samples in the horizontal direction and are between the rows of the luminance samples in the vertical direction), the center of the diamond is aligned with the chrominance sample position.

[0181] Since the symmetry constraint is not enforced, the CCALF coefficients have greater flexibility compared to the conventional ALF coefficients. However, there are also two limitations:

[0182] To maintain DC neutrality, the sum of the CCALF coefficient values is required to be zero. Therefore, only seven out of the eight CCALF coefficients need to be signaled in the bitstream, and the coefficient at position (x C , y C ) is derived at the decoder.

[0183] The absolute value of the CCALF coefficients is limited to zero or an integer power of two, specifically {0, 1, 2, 4, 8, 16, 32, 64}. This enables the implementation to use variable shift operations instead of multiplication for CCALF as needed.

[0184] Syntax Design

[0185] In the final design of VVC, the maximum number of filters for each chrominance component of a picture is four. A different set of CCALF coefficients can be selected for each CTU with a chrominance component. Similar to the conventional ALF coefficients, the CCALF coefficients are signaled within the ALF APS. Each ALF APS can contain up to four CCALF filters for each chrominance component. Although CCALF can be enabled at the sequence level, it can only be enabled when ALF is also enabled for the sequence. Similarly, CCALF can only be enabled at the corresponding level when the luminance ALF is enabled at the picture and slice levels.

[0186] Line Buffer Reduction

[0187] As described in Section 3.1.5, the line buffer boundaries for luminance and chrominance are located four samples and two samples above the CTU boundary, respectively. For the 4:2:0 chrominance format, this results in the alignment of the line buffer boundaries for chrominance and luminance. However, for the 4:2:2 and 4:4:4 chrominance formats, the line buffer boundaries for chrominance and luminance are not aligned with each other. Due to this misalignment, CC-ALF is not applied to the third and fourth rows of samples above the CTU boundary for the 4:2:2 and 4:4:4 chrominance formats.

[0188] CCALF in ECM

[0189] The CCALF process uses a linear filter to filter the luminance sample values and generate a residual correction for the chrominance samples. A 25-tap large filter is used in the CCALF process, as Figure 13 shown. For a given slice, the encoder can collect the statistics of the slice, analyze them, and can signal up to 16 filters via the APS.

[0190] Although ALF and CCALF have been improved in the ECM, there is still room for further performance improvement.

[0191] First, the online ALF filter in the ECM takes as input the spatially neighboring pixels, the fixed ALF filter result, and the spatially neighboring pixels before the deblocking filter. However, in addition to this information, other information (such as spatially neighboring pixels in the prediction signal, spatially neighboring pixels in the residual signal, or spatially neighboring pixels before SAO) can also be used as input to the online ALF filter equation, which may be beneficial to the coding and decoding performance.

[0192] Second, in the ECM, the online ALF filter adaptively uses an edge-based classifier and a band-based classifier. However, these two classifiers can be further combined to provide other classifiers, which may be beneficial to the coding and decoding performance.

[0193] Third, in the ECM, the filter shape of the chroma ALF is diamond-shaped, while the filter shape of the luma ALF is long cross-shaped. From the perspective of normalization, this non-uniform design may not be optimal.

[0194] Fourth, the edge-based classifier and the band-based classifier in the ECM only consider the pixel values after SAO. However, after saving the pixel values from these stages: 1) immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before SAO as input to the online ALF filter equation, these pixel values can also be used to design a new classifier, which may be beneficial to the coding and decoding performance.

[0195] Fifth, the edge-based classifier and the band-based classifier in the ECM only consider the luma pixel values after SAO. However, the chroma pixel values can also be used to design a new classifier, which may be beneficial to the coding and decoding performance.

[0196] Sixth, similar to saving the luma pixel values from these stages: 1) immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before SAO as input to an additional online luma ALF filter equation, the chroma pixel values from these stages: 1) before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before SAO can also be saved as input to an additional online chroma ALF filter equation, which may be beneficial to the coding and decoding performance.

[0197] Seventh, similar to saving the luma pixel values from these stages: 1) immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before SAO as input to an additional online luma ALF filter equation, the luma pixel values from these stages: 1) immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before SAO can also be saved as input to an additional CCALF filter equation, which may be beneficial to the coding and decoding performance.

[0198] Eighth, the classifier design in the ECM only considers the reconstructed pixel values. However, the coding mode information (such as whether the coding block is coded / decoded using the skip mode, or whether the coding block is coded / decoded using the intra, inter-P, or inter-B mode) can also be used to design the classifier, which may be beneficial to the coding / decoding performance.

[0199] Ninth, after the online ALF filter takes the samples from stages such as 1) the samples immediately before deblocking, 2) the predicted samples, 3) the residual samples, and 4) the samples immediately before SAO as additional inputs, according to the current line buffer settings in VVC, additional line buffers are required to save 4 corresponding luminance samples and 2 corresponding chrominance samples above the horizontal CTU boundary, which increases the implementation complexity.

[0200] Tenth, after the CCALF filter takes the samples from stages such as 1) the samples immediately before deblocking, 2) the predicted samples, 3) the residual samples, and 4) the samples immediately before SAO as additional inputs, according to the current line buffer settings in VVC, additional line buffers are required to save 4 corresponding luminance samples above the horizontal CTU boundary, which increases the implementation complexity.

[0201] In the present disclosure, to solve the problems pointed out in the "Problem Description" section, a method for further improving the existing design of ALF is provided. Generally, the main features of the technology proposed in the present disclosure are summarized as follows.

[0202] The online ALF filter takes the spatially adjacent pixels in the prediction signal, the spatially adjacent pixels in the residual signal, or the spatially adjacent pixels before SAO as additional inputs.

[0203] A classifier that combines the features of an edge-based classifier and a band-based classifier is used as an additional classifier for the online ALF filter.

[0204] The filter shape of the chrominance ALF is changed from a rhombus to a long cross to unify it with the filter shape of the luminance ALF.

[0205] A classifier that uses the pixel values from stages such as 1) immediately before the deblocking filter, 2) the prediction signal, 3) the residual signal, and 4) immediately before SAO is used as an additional classifier for the online ALF filter.

[0206] A classifier that uses chrominance pixel values is used as an additional classifier for the online ALF filter.

[0207] The online chrominance ALF filter takes the spatially adjacent pixels in the chrominance prediction signal, the spatially adjacent pixels in the chrominance residual signal, the spatially adjacent pixels from the stage immediately before chrominance SAO, or the spatially adjacent pixels from the stage immediately before chrominance deblocking as additional inputs.

[0208] The CCALF filter takes as additional inputs spatially neighboring pixels in the luma prediction signal, spatially neighboring pixels in the luma residual signal, spatially neighboring pixels from the stage immediately before luma SAO, or spatially neighboring pixels from the stage immediately before luma deblocking.

[0209] A classifier that utilizes coding mode information (such as whether a coding block is coded / decoded in skip mode, or whether a coding block is coded / decoded in intra, inter P, or inter B mode) is used as an additional classifier for the online ALF filter.

[0210] When the online ALF filter takes as additional inputs samples from stages such as 1) samples immediately before deblocking, 2) predicted samples, 3) residual samples, 4) samples immediately before SAO, etc., according to the current row buffer settings in VVC, 4 corresponding luma samples and 2 corresponding chroma samples above the horizontal CTU boundary are assumed to be default values, which can save these row buffers.

[0211] When the CCALF filter takes as additional inputs samples from stages such as 1) samples immediately before deblocking, 2) predicted samples, 3) residual samples, 4) samples immediately before SAO, etc., according to the current row buffer settings in VVC, 4 corresponding luma samples above the horizontal CTU boundary are assumed to be default values, which can save these row buffers.

[0212] In some embodiments of the present disclosure, the disclosed methods can be applied independently or jointly.

[0213] Prediction, residual, or information before SAO is used as additional ALF input

[0214] According to one or more embodiments of the present disclosure, prediction, residual, or information before SAO is used as additional ALF equation inputs. Different methods can be used to achieve this goal.

[0215] Figure 14 The online ALF filter inputs are presented. The online ALF filter can take all or a subset of the following as additional inputs: predicted samples, output samples obtained by feeding the predicted samples to a fixed filter trained offline, residual samples, output samples obtained by feeding the residual samples to a fixed filter trained offline, reconstructed samples immediately before SAO, and output samples obtained by feeding the reconstructed samples immediately before SAO to a fixed filter trained offline.

[0216] In the first method, it is proposed to use spatially adjacent pixels in the prediction signal as additional inputs to the ALF equation. Various filter shapes can be used to extract information from the prediction signal. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract information from the prediction signal. In an embodiment, the clipped difference between the surrounding pixels and the current pixel in the prediction signal is used as an input to the ALF equation. In another example, the clipped difference between the surrounding pixels and the co-located pixels in the prediction signal, and the clipped difference between the co-located pixels and the current pixel in the prediction signal are used as inputs to the ALF equation.

[0217] In addition to directly applying the additional on-line ALF filter taps to the prediction signal, the additional on-line ALF filter taps can also be applied to the intermediate result obtained by feeding the prediction signal into a fixed filter. Various fixed filters can be applied to filter the prediction signal to obtain an intermediate result, which can concentrate the prediction signal information in a larger receptive field. For example, the two 13×13 diamond-shaped fixed filters used in the ALF of ECM can be used to filter the prediction signal to obtain an intermediate result. When applying a fixed filter to the prediction signal, the block-level classification result can directly utilize the block-level classification result calculated for the signal immediately after SAO, or be recalculated based on the prediction signal. When applying a fixed filter to the prediction signal, one fixed filter trained based on one block-level classifier can be used to obtain one intermediate result, or two or more fixed filters trained based on two or more block-level classifiers can be used to obtain two or more intermediate results. In video coding standards, several sets of fixed filters are usually prepared, and one set of fixed filters can be selected from them through a rate-distortion optimization (RDO) process. For example, in ECM, one set of fixed filters (including two 13×13 diamond-shaped fixed filters) is selected from two sets through the RDO process, and the set index is transmitted to the decoder. When applying a fixed filter to the prediction signal, the set index of the prediction signal can be the same as the set index of the signal immediately after SAO, or different from the set index of the signal immediately after SAO based on a predefined criterion (in ECM, there are two sets, so if the set index of the signal immediately after SAO is 0, the set index of the prediction signal is 1; if the set index of the signal immediately after SAO is 1, the set index of the prediction signal is 0), or the set index for the prediction signal can be determined through the RDO process, where in the first and second cases, it is not necessary to transmit the set index of the prediction signal to the decoder, while in the third case, it is necessary to transmit the set index of the prediction signal to the decoder.

[0218] When applying additional online filter taps to the intermediate result obtained by feeding a prediction signal into a fixed filter, various filter shapes can be used to extract information from the intermediate result. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract information from the intermediate result. In an embodiment, the clipped difference between the surrounding pixels and the current pixel in the intermediate result is used as the ALF equation input. In another example, the clipped difference between the surrounding pixels and the co-located pixels in the intermediate result, and the clipped difference between the co-located pixels and the current pixel in the intermediate result are used as the ALF equation input.

[0219] It should be noted that the additional online ALF filter taps can be applied only to the prediction signal, or only to the intermediate result obtained by feeding the prediction signal into a fixed filter, or applied to both the prediction signal and the intermediate result obtained by feeding the prediction signal into a fixed filter. For example, in the AI (all intra-frame) test, the additional online ALF filter taps are applied only to the prediction signal; in the RA (random access) test, the additional online ALF filter taps are applied to the prediction signal and the intermediate result obtained by feeding the prediction signal into a fixed filter.

[0220] In the second method, it is proposed to use spatially neighboring pixels in the residual signal as an additional input to the ALF equation. Various filter shapes can be used to extract information from the residual signal. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract information from the residual signal. In an embodiment, the clipped result of the co-located pixels in the residual signal is used as the input to the ALF equation.

[0221] In addition to directly applying the additional on-line ALF filter taps to the residual signal, the additional on-line ALF filter taps can also be applied to the intermediate result obtained by feeding the residual signal to a fixed filter. Various fixed filters can be applied to filter the residual signal to obtain an intermediate result, which can concentrate the residual signal information in a larger receptive field. For example, these two 13×13 diamond-shaped fixed filters used in the ALF in ECM can be used to filter the residual signal to obtain an intermediate result. In one or more examples, considering that for the prediction signal and the signal before SAO, the range is the same as that of the signal after SAO, i.e., (0, 1024), which is positive, but for the residual signal, the range may be positive or negative. Therefore, when applying a fixed filter to the residual signal, the filtering result can be truncated to different ranges, such as (-1024, 1024), (-512, 512), (-256, 256), (-128, 128), etc. When applying a fixed filter to the residual signal, the block-level classification result can directly utilize the block-level classification result calculated for the signal immediately after SAO, or be recalculated based on the residual signal. When applying a fixed filter to the residual signal, one fixed filter trained based on one block-level classifier can be used to obtain one intermediate result, or two or more fixed filters trained based on two or more block-level classifiers can be used to obtain two or more intermediate results. When applying a fixed filter to the residual signal, the group index of the residual signal can be the same as the group index of the signal immediately after SAO, or different from the group index of the signal immediately after SAO based on a predefined criterion (in ECM, there are two groups, so if the group index of the signal immediately after SAO is 0, the group index of the residual signal is 1; if the group index of the signal immediately after SAO is 1, the group index of the residual signal is 0), or the group index for the residual signal can be determined through the RDO process, where in the first and second cases, it is not necessary to transmit the group index of the residual signal to the decoder, while in the third case, it is necessary to transmit the group index of the residual signal to the decoder.

[0222] When applying the additional on-line filter taps to the intermediate result obtained by feeding the residual signal to a fixed filter, various filter shapes can be used to extract the information in the intermediate result. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract the information in the intermediate result. In an embodiment, the truncated result of the co-located pixels in the intermediate result is used as the input to the ALF equation.

[0223] It should be noted that the additional on-line ALF filter taps can be applied only to the residual signal, or only to the intermediate result obtained by feeding the residual signal to a fixed filter, or to both the residual signal and the intermediate result obtained by feeding the residual signal to a fixed filter. For example, in the AI (all intra-frame) test, the additional on-line ALF filter taps are applied only to the residual signal; in the RA (random access) test, the additional on-line ALF filter taps are applied to both the residual signal and the intermediate result obtained by feeding the residual signal to a fixed filter.

[0224] In the third method, it is proposed to use the spatially neighboring pixels from the stage of the signal immediately before SAO as the additional ALF equation input. Various filter shapes can be used to extract the information in the signal before SAO. For example, the filter shape can be 1×1, 3×3 or 5×5, as Figure 15 shown. Various equation forms can be used to extract the information in the signal before SAO. In an embodiment, the clipped difference between the surrounding pixels and the current pixel in the signal before SAO is used as the input to the ALF equation. In another example, the clipped difference between the surrounding pixels and the co-located pixels in the signal before SAO, and the clipped difference between the co-located pixels and the current pixel in the signal before SAO are used as the ALF equation input.

[0225] In addition to directly applying the additional in-line ALF filter taps to the signal immediately before SAO, the additional in-line ALF filter taps can be applied to the intermediate result obtained by feeding the signal immediately before SAO to a fixed filter. Various fixed filters can be applied to filter the signal immediately before SAO to obtain an intermediate result, which can concentrate the signal information immediately before SAO in a larger receptive field. For example, these two 13×13 diamond-shaped fixed filters used in the ALF in ECM can be used to filter the signal immediately before SAO to obtain an intermediate result. When a fixed filter is applied to the signal immediately before SAO, the block-level classification result can directly utilize the block-level classification result calculated for the signal immediately after SAO, or be recalculated based on the signal immediately before SAO. When a fixed filter is applied to the signal immediately before SAO, one fixed filter trained based on one block-level classifier can be utilized to obtain one intermediate result, or two or more fixed filters trained based on two or more block-level classifiers can be utilized to obtain two or more intermediate results. When a fixed filter is applied to the signal immediately before SAO, the group index of the signal immediately before SAO can be the same as the group index of the signal immediately after SAO, or different from the group index of the signal immediately after SAO based on a predefined criterion (in ECM, there are two groups, so if the group index of the signal immediately after SAO is 0, the group index of the signal immediately before SAO is 1; if the group index of the signal immediately after SAO is 1, the group index of the signal immediately before SAO is 0), or the group index of the signal immediately before SAO can be determined through the RDO process, where in the first and second cases, it is not necessary to transmit the group index of the signal immediately before SAO to the decoder, while in the third case, it is necessary to transmit the group index of the signal immediately before SAO to the decoder.

[0226] When applying the additional in-line filter taps to the intermediate result obtained by feeding the signal immediately before SAO to a fixed filter, various filter shapes can be used to extract the information in the intermediate result. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract the information in the intermediate result. In an embodiment, the clipped difference between the surrounding pixels and the current pixel in the intermediate result is used as the input to the ALF equation. In another example, the clipped difference between the surrounding pixels and the co-located pixels in the intermediate result, and the clipped difference between the co-located pixels and the current pixel in the intermediate result are used as the input to the ALF equation.

[0227] It should be noted that the additional online ALF filter taps can be applied only to the signal immediately before SAO, or only to the intermediate result obtained by feeding the signal immediately before SAO into a fixed filter, or applied simultaneously to the signal immediately before SAO and the intermediate result obtained by feeding the signal immediately before SAO into a fixed filter. For example, in the AI (All Intra) test, the additional online ALF filter taps are applied only to the signal immediately before SAO; in the RA (Random Access) test, the additional online ALF filter taps are applied to the signal immediately before SAO and the intermediate result obtained by feeding the signal immediately before SAO into a fixed filter.

[0228] In the fourth method, it is proposed to use the information in the predicted signal, the residual signal, or the signal before SAO as the input to the ALF equation. The utilization methods proposed in the first, second, and third methods can be combined to implement the fourth method.

[0229] New classifier combining features of edge-based classifier and band-based classifier

[0230] According to one or more embodiments of the present disclosure, the features of the edge-based classifier and the features of the band-based classifier are combined to obtain a new classifier for the online ALF filter. Different methods can be used to achieve this goal.

[0231] In the first method, it is proposed to first calculate the directionality D of the sub-block of the luminance component, then calculate the sum of the sample values of the sub-block and map it to the index of the reference band-based classifier, and the class index of the sub-block is calculated as

[0232] C = B * M D + D (17),

[0233] where B is the index calculated by the reference band-based classifier, and M D represents the total number of directionality D. In the embodiment, for a 2×2 luminance block, the directionality D is calculated in the same way as D2 in ECM, and B is calculated as

[0234] B = (sum * 5) >> (sample bit depth + 2) (18).

[0235] In the second method, it is proposed to first calculate the activity value A of the sub-block of the luminance component, then calculate the sum of the sample values of the sub-block and map it to the index of the band-based classifier, and the class index of the sub-block is calculated as

[0236] C = B * M A + A (19),

[0237] where B is the index calculated by the reference band-based classifier, and M ARepresents the total number of active values A. In an embodiment, for a 2×2 luminance block, the active value A is calculated in the same way as in ECM and B is calculated as

[0238] B = (sum * 5) >> (sample bit depth + 2) (20).

[0239] In the third method, it is proposed to first calculate the index of the sub - block of the luminance component with reference to an edge - based classifier, then calculate the sum of the sample values of the sub - block and map it to the index of a band - based classifier, and the class index of the sub - block is calculated as

[0240] C = B * M E + E (21),

[0241] where B is the index calculated with reference to the band - based classifier, M E represents the total number of indices calculated with reference to the edge - based classifier, and E is the index calculated with reference to the edge - based classifier. In an embodiment, for a 2×2 luminance block, the index E is calculated in the same way as C2 in ECM, and B is calculated as

[0242] B = (sum * 2) >> (sample bit depth + 2) (22).

[0243] Adjust the chrominance ALF filter shape to be unified with the luminance ALF filter shape

[0244] In a third aspect of the present disclosure, it is proposed to change the chrominance ALF filter shape from a rhombus to a long cross - shape as shown in Figure 16 to unify it with the luminance ALF filter shape.

[0245] New classifier using pixel values from the stage immediately before the deblocking filter

[0246] According to one or more embodiments of the present disclosure, pixel values from the stage immediately before the de - blocking filter are used to derive a new classifier for the in - line ALF filter. Different methods can be used to achieve this goal.

[0247] In the first method, it is proposed to first calculate the directionality D of the sub - block with the luminance component, then calculate the sum of the differences between the samples from the stage immediately after SAO and the co - located samples from the sub - block from the stage immediately before the de - blocking filter and map it to a difference index, and the class index of the sub - block is calculated as

[0248] C = Dif * M D + D (23),

[0249] where Dif is the difference index, M DRepresents the total number of directions D. In an embodiment, for a 2×2 luminance block, the direction D is calculated in the same way as D2 in ECM, and Dif is calculated as

[0250] Dif = sum Dif >0? 2 : (sum Dif <0? 0 : 1) (24).

[0251] In the second method, it is proposed to first calculate the activity value A of the sub-block of the luminance component, then calculate the sum of the differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before the deblocking filter of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0252] C = Dif * M A + A (25),

[0253] where Dif is the difference index and M A represents the total number of activity values A. In an embodiment, for a 2×2 luminance block, the activity value A is calculated in the same way as in ECM and Dif is calculated according to Equation (24).

[0254] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component with reference to an edge-based classifier, then calculate the sum of the differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before the deblocking filter of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0255] C = Dif * M E + E (26),

[0256] where Dif is the difference index and M E represents the total number of indices calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In an embodiment, for a 2×2 luminance block, the index E is calculated in the same way as C2 in ECM, and Dif is calculated according to Equation (24).

[0257] In the fourth method, it is proposed to first calculate the band index B of the sub-block with the luminance component, then calculate the sum of the differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before the deblocking filter of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0258] C = Dif * M B + B (27),

[0259] where Dif is the difference index and M BRepresents the total number of values with a value. In an embodiment, for a 2×2 luminance block, the band index B is calculated as

[0260] B = (sum * 8) >> (sample bit depth + 2) (28),

[0261] and Dif is calculated according to Equation (24).

[0262] In the fifth method, it is proposed to calculate the sum of the differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before the deblocking filter in the sub-block, then map the sum of the differences to a difference index, and use the difference index as the class index.

[0263] In the sixth method, it is proposed to calculate an edge-based classifier or a band-based classifier based on the sample values from the stage immediately before the deblocking filter, and the calculation method is the same as the calculation method of the original edge-based classifier or band-based classifier calculated based on the sample values after SAO.

[0264] New classifier using pixel values in the prediction signal

[0265] According to one or more embodiments of the present disclosure, the pixel values in the prediction signal are used to obtain a new classifier for the online ALF filter. Different methods can be used to achieve this goal.

[0266] In the first method, it is proposed to first calculate the directionality D of the sub-block with a luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples in the prediction signal of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0267] C = Dif * M D + D (29),

[0268] where Dif is the difference index and M D represents the total number of directionality D. In an embodiment, for a 2×2 luminance block, the directionality D is calculated in the same way as D2 in ECM, and Dif is calculated as

[0269] Dif = sum Dif > 0? 2 : (sum Dif < 0? 0 : 1) (30).

[0270] In the second method, it is proposed to first calculate the activity value A of the sub-block with a luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples in the prediction signal of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0271] C = Dif * M A + A (31),

[0272] where Dif is the difference index, and M A represents the total number of active values A. In an embodiment, for a 2×2 luminance block, the active value A is calculated in the same way as in ECM, and Dif is calculated according to Equation (30).

[0273] In the third method, it is proposed to first calculate the index of a sub-block of the luminance component with reference to an edge-based classifier, then calculate the sum of the differences between the samples after SAO and the co-located samples in the predicted signal of the sub-block and map it to the difference index, and the class index of the sub-block is calculated as

[0274] C = Dif * M E + E (32),

[0275] where Dif is the difference index, and M E represents the total number of indices calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In an embodiment, for a 2×2 luminance block, the index E is calculated in the same way as C2 in ECM, and Dif is calculated according to Equation (30).

[0276] In the fourth method, it is proposed to first calculate the band index B of a sub-block of the luminance component, then calculate the sum of the differences between the samples after SAO and the co-located samples in the predicted signal of the sub-block and map it to the difference index, and the class index of the sub-block is calculated as

[0277] C = Dif * M B + B (33),

[0278] where Dif is the difference index, and M B represents the total number of band values. In an embodiment, for a 2×2 luminance block, the band index B is calculated as

[0279] B = (sum * 8) >> (sample bit depth + 2) (34),

[0280] and Dif is calculated according to Equation (30).

[0281] In the fifth method, it is proposed to calculate the sum of the differences between the samples after SAO and the co-located samples in the predicted signal of the sub-block, then map the sum of the differences to the difference index, and use the difference index as the class index.

[0282] In the sixth method, it is proposed to calculate an edge-based classifier or a band-based classifier based on the sample values in the predicted signal, and the calculation method is the same as the original calculation method of the edge-based classifier or the band-based classifier calculated based on the sample values after SAO.

[0283] New classifier using pixel values in the residual signal

[0284] According to one or more embodiments of the present disclosure, pixel values in a residual signal are utilized to derive a new classifier for an online ALF filter. Different methods can be used to achieve this goal.

[0285] In the first method, it is proposed to first calculate the directionality D of a sub-block of the luminance component, then calculate the sum of the pixel values in the residual signal of the sub-block and map it to a residual index, and the class index of the sub-block is calculated as

[0286] C = Resi * M D + D (35),

[0287] where Resi is the residual index and M D represents the total number of directionality D. In an embodiment, for a 2×2 luminance block, the directionality D is calculated in the same way as D2 in ECM, and Resi is calculated as

[0288] Resi = sum Resi > 0? 2 : (sum Resi < 0? 0 : 1) (36).

[0289] In the second method, it is proposed to first calculate the activity value A of a sub-block of the luminance component, then calculate the sum of the pixel values in the residual signal of the sub-block and map it to a residual index, and the class index of the sub-block is calculated as

[0290] C = Resi * M A + A (37),

[0291] where Resi is the residual index and M A represents the total number of activity value A. In an embodiment, for a 2×2 luminance block, the activity value A is calculated in the same way as in ECM and Resi is calculated according to equation (36).

[0292] In the third method, it is proposed to first calculate the index of a sub-block with a luminance component by referring to an edge-based classifier, then calculate the sum of the pixel values in the residual signal of the sub-block and map it to a residual index, and the class index of the sub-block is calculated as

[0293] C = Resi * M E + E (38),

[0294] where Resi is the residual index and M EIndicates the total number of indices calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In an embodiment, for a 2×2 luminance block, the index E is calculated in the same way as C2 in ECM, and Resi is calculated according to Equation (36).

[0295] In the fourth method, it is proposed to first calculate the index B with respect to the sub-blocks of the luminance component, then calculate the sum of the pixel values in the residual signal of the sub-block and map it to a residual index, and the class index of the sub-block is calculated as

[0296] C = Resi * M B + B (39),

[0297] where Resi is the residual index and M B indicates the total number of band values. In an embodiment, for a 2×2 luminance block, the index B is calculated as

[0298] B = (sum * 8) >> (sample bit depth + 2) (40),

[0299] and Resi is calculated according to Equation (36).

[0300] In the fifth method, it is proposed to calculate the sum of the pixel values in the residual signal of the sub-block, then map the sum of the residual values to a residual index, and use the residual index as the class index.

[0301] New classifier using pixel values from the stage immediately before SAO

[0302] According to one or more embodiments of the present disclosure, pixel values from the stage immediately before SAO are used to derive a new classifier for the online ALF filter. Different methods can be used to achieve this goal.

[0303] In the first method, it is proposed to first calculate the directionality D of the sub-blocks with the luminance component, then calculate the sum of the differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before SAO of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0304] C = Dif * M D + D (41),

[0305] where Dif is the difference index and M D indicates the total number of directionality D. In an embodiment, for a 2×2 luminance block, the directionality D is calculated in the same way as D2 in ECM, and Dif is calculated as

[0306] Dif = sum Dif > 0? 2 : (sum Dif<0? 0: 1)(42).

[0307] In the second method, it is proposed to first calculate the activity value A of the sub-block of the luminance component, then calculate the sum of the differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before SAO of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0308] C = Dif * M A + A(43),

[0309] where Dif is the difference index and M A represents the total number of the activity value A. In the embodiment, for a 2×2 luminance block, the activity value A is calculated in the same way as in ECM and Dif is calculated according to Equation (42).

[0310] In the third method, it is proposed to first calculate the index of the sub-block of the luminance component with reference to an edge-based classifier, then calculate the sum of the differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before SAO of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0311] C = Dif * M E + E(44),

[0312] where Dif is the difference index and M E represents the total number of the indexes calculated with reference to the edge-based classifier, and E is the index calculated with reference to the edge-based classifier. In the embodiment, for a 2×2 luminance block, the index E is calculated in the same way as C2 in ECM, and Dif is calculated according to Equation (42).

[0313] In the fourth method, it is proposed to first calculate the band index B of the sub-block of the luminance component, then calculate the sum of the differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before SAO of the sub-block and map it to a difference index, and the class index of the sub-block is calculated as

[0314] C = Dif * M B + B(45),

[0315] where Dif is the difference index and M B represents the total number of the band values. In the embodiment, for a 2×2 luminance block, the band index B is calculated as

[0316] B = (sum * 8)>>(sample bit depth + 2)(46),

[0317] and Dif is calculated according to Equation (42).

[0318] In the fifth method, it is proposed to calculate the sum of differences between the samples from the stage immediately after SAO and the co-located samples from the stage immediately before SAO in the sub-block, then map the sum of differences to a difference index, and use the difference index as the class index.

[0319] In the sixth method, it is proposed to calculate an edge-based classifier or a band-based classifier based on the sample values from the stage immediately before SAO, where the calculation method is the same as that of the edge-based classifier or the band-based classifier calculated based on the sample values after SAO originally.

[0320] New classifier using chrominance pixel values

[0321] According to one or more embodiments of the present disclosure, chrominance pixel values are used to obtain a new classifier for the online ALF filter. Different methods can be used to achieve this goal.

[0322] In the first method, it is proposed to first calculate the band index B of the sub-block of the luminance component Y , then calculate the band indices B U and B V of the corresponding U and V components, and the class index of the sub-block is calculated as

[0323] C = B Y * M U * M V + B U * M V + B V (47),

[0324] where B Y , B U and B V are the Y, U, and V indices calculated with reference to the band-based classifier, and M U and M V represent the total number of U and V band index values. In an embodiment, for a 2×2 luminance block, B Y , B U and B V are calculated as

[0325] B Y = (sumY * 6) >> (sample bit depth + 2) (48),

[0326] B U = (sumU * 2) >> (sample bit depth + 2) (49),

[0327] B V = (sumV * 2) >> (sample bit depth + 2) (50).

[0328] Chrominance information from the stages immediately before deblocking, in prediction, in residual, or immediately before SAO is used as additional chrominance ALF input

[0329] According to one or more embodiments of the present disclosure, chrominance information from stages immediately before deblocking, in prediction, in residual, or immediately before SAO is used as an additional input to the chrominance ALF equation. Different methods can be used to achieve this goal.

[0330] In the first method, it is proposed to use spatially adjacent pixels in the chrominance prediction signal as an additional input to the chrominance ALF equation. Various filter shapes can be used to extract information from the chrominance prediction signal. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract information from the chrominance prediction signal. In an embodiment, the clipped difference between the surrounding pixels and the current chrominance pixel in the chrominance prediction signal is used as an input to the chrominance ALF equation. In another example, the clipped difference between the surrounding pixels and the co-located pixels in the chrominance prediction signal, and the clipped difference between the co-located pixels and the current chrominance pixel in the chrominance prediction signal are used as inputs to the chrominance ALF equation.

[0331] In the second method, it is proposed to use spatially adjacent pixels in the chrominance residual signal as an additional input to the chrominance ALF equation. Various filter shapes can be used to extract information from the chrominance residual signal. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract information from the chrominance residual signal. In an embodiment, the clipped result of the co-located pixels in the chrominance residual signal is used as an input to the chrominance ALF equation.

[0332] In the third method, it is proposed to use spatially adjacent pixels from the stage of the signal immediately before chrominance SAO as an additional input to the chrominance ALF equation. Various filter shapes can be used to extract information from the stage of the signal immediately before chrominance SAO. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract information from the stage of the signal immediately before chrominance SAO. In an embodiment, the clipped difference between the surrounding pixels and the current chrominance pixel from the stage of the signal immediately before chrominance SAO is used as an input to the chrominance ALF equation. In another example, the clipped difference between the surrounding pixels and the co-located pixels from the stage of the signal immediately before chrominance SAO, and the clipped difference between the co-located pixels and the current chrominance pixel from the stage of the signal immediately before chrominance SAO are used as inputs to the chrominance ALF equation.

[0333] In the fourth method, it is proposed to use the spatial neighboring pixels from the stage of the signal immediately before chroma deblocking as additional chroma ALF equation inputs. Various filter shapes can be used to extract the information from the stage of the signal immediately before chroma deblocking. For example, the filter shape can be 1×1, 3×3, or 5×5, as Figure 15 shown. Various equation forms can be used to extract the information from the stage of the signal immediately before chroma deblocking. In an embodiment, the clipped difference between the surrounding pixels from the stage of the signal immediately before chroma deblocking and the current chroma pixel is used as the chroma ALF equation input. In another example, the clipped difference between the surrounding pixels from the stage of the signal immediately before chroma deblocking and the co-located pixels from the stage of the signal immediately before chroma deblocking, and the clipped difference between the co-located pixels from the stage of the signal immediately before chroma deblocking and the current chroma pixel are used as the inputs to the chroma ALF equation.

[0334] In the fifth method, it is proposed to use the information in the chroma prediction, residual, before SAO, or before deblocking signals as the chroma ALF equation input. The utilization methods proposed in the first, second, third, and fourth methods can be combined to implement the fifth method.

[0335] Luminance information from the stages immediately before deblocking, in prediction, in residual, or immediately before SAO is used as additional CCALF input

[0336] According to one or more embodiments of the present disclosure, the luminance information from the stages immediately before deblocking, in prediction, in residual, or immediately before SAO is used as additional CCALF equation inputs. Different methods can be used to achieve this goal.

[0337] In the first method, it is proposed to use the spatial neighboring pixels in the luminance prediction signal as additional CCALF equation inputs. Various filter shapes can be used to extract the information in the luminance prediction signal. For example, the filter shape can be 3×4, as Figure 12 shown. Various equation forms can be used to extract the information in the luminance prediction signal. In an embodiment, the difference between the surrounding pixels in the luminance prediction signal and the current corresponding luminance pixel is used as the CCALF equation input. In another example, the difference between the surrounding pixels in the luminance prediction signal and the co-located pixels in the current corresponding luminance prediction signal, and the difference between the co-located pixels in the current corresponding luminance prediction signal and the current corresponding luminance pixel are used as the CCALF equation inputs.

[0338] In the second method, it is proposed to use the spatial neighboring pixels in the luminance residual signal as additional CCALF equation inputs. Various filter shapes can be used to extract the information in the luminance residual signal. For example, the filter shape can be 3×4, asFigure 12 As shown. Various equation forms can be used to extract information from the luminance residual signal. In an embodiment, co-located pixels in the luminance residual signal are used as inputs to the CCALF equation.

[0339] In a third method, it is proposed to use spatially adjacent pixels from the stage of the signal immediately preceding the luminance SAO as additional inputs to the CCALF equation. Various filter shapes can be used to extract information from the stage of the signal immediately preceding the luminance SAO. For example, the filter shape can be 3×4, as Figure 12 shown. Various equation forms can be used to extract information from the stage of the signal immediately preceding the luminance SAO. In an embodiment, the difference between the surrounding pixels from the stage of the signal immediately preceding the luminance SAO and the current corresponding luminance pixel is used as an input to the CCALF equation. In another example, the difference between the surrounding pixels from the stage of the signal immediately preceding the luminance SAO and the co-located pixels in the signal immediately preceding the current corresponding luminance SAO, and the difference between the co-located pixels in the signal immediately preceding the current corresponding luminance SAO and the current corresponding luminance pixel are used as inputs to the CCALF equation.

[0340] In a fourth method, it is proposed to use spatially adjacent pixels from the stage of the signal immediately preceding the luminance deblocking as additional inputs to the CCALF equation. Various filter shapes can be used to extract information from the stage of the signal immediately preceding the luminance deblocking. For example, the filter shape can be 3×4, as Figure 12 shown. Various equation forms can be used to extract information from the stage of the signal immediately preceding the luminance deblocking. In an embodiment, the difference between the surrounding pixels from the stage of the signal immediately preceding the luminance deblocking and the current corresponding luminance pixel is used as an input to the CCALF equation. In another example, the difference between the surrounding pixels from the stage of the signal immediately preceding the luminance deblocking and the co-located pixels in the signal immediately preceding the current corresponding luminance deblocking, and the difference between the co-located pixels in the signal immediately preceding the current corresponding luminance deblocking and the current corresponding luminance pixel are used as inputs to the CCALF equation.

[0341] In a fifth method, it is proposed to use information in the luminance prediction, residual, before SAO, or before deblocking signals as inputs to the CCALF equation. The utilization methods proposed in the first, second, third, and fourth methods can be combined to implement the fifth method.

[0342] New classifier using coding mode information

[0343] According to one or more embodiments of the present disclosure, a new classifier for an online ALF filter is derived using coding mode information (such as whether a coding block is encoded / decoded using a skip mode, or whether a coding block is encoded / decoded using an intra, inter-P, or inter-B mode). Different methods can be used to achieve this goal.

[0344] In a first method, it is proposed to record whether a coding block is encoded / decoded using a skip mode during the encoding and decoding processes, and then use this information to design a new classifier. In an embodiment, a classifier with 2 classes (corresponding to skip mode being true or false respectively) is added as the new classifier. In another example, a classifier that combines skip mode information with EO or BO is added as the new classifier.

[0345] In a second method, it is proposed to record whether a coding block is encoded / decoded using an intra mode, an inter-P mode, or an inter-B mode during the encoding and decoding processes, and then use this information to design a new classifier. In an embodiment, a classifier with 3 classes (corresponding to the intra mode, the inter-P mode, or the inter-B mode respectively) is added as the new classifier. In another example, a classifier that combines intra, inter-P, or inter-B mode information with EO or BO is added as the new classifier.

[0346] In a third method, it is proposed to design a new classifier by considering both types of coding mode information (whether a coding block is encoded / decoded using a skip mode, and whether a coding block is encoded / decoded using an intra, inter-P, or inter-B mode). The utilization methods proposed in the first and second methods can be combined to implement the third method.

[0347] Reduction of the row buffer for additional ALF input

[0348] According to one or more embodiments of the present disclosure, when the online ALF filter takes samples from these stages such as 1) samples immediately before deblocking, 2) predicted samples, 3) residual samples, 4) samples immediately before SAO, etc. as additional inputs, a line buffer is required to save these samples. To reduce the line buffer requirements for these additional inputs, according to the current line buffer settings in VVC, 4 corresponding luminance samples and 2 corresponding chrominance samples above the horizontal CTU boundary are assumed to be default values, which can save these line buffers. Different methods can be used to achieve this goal.

[0349] In a first method, according to the current line buffer settings in VVC, it is proposed to assume that 4 lines of luminance residual samples and 2 lines of chrominance residual samples above the horizontal CTU boundary are zero values, and assume that 4 lines of luminance samples and 2 lines of chrominance samples above the horizontal CTU boundary from these stages such as 1) samples immediately before deblocking, 2) predicted samples, 3) samples immediately before SAO, etc. are the co-located sample values from the stage immediately after SAO.

[0350] In the second method, according to the current line buffer settings in VVC, it is proposed to assume that 4 rows of luma samples and 2 rows of chroma samples above the horizontal CTU boundary from 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, 4) samples immediately before SAO, etc. are set in a repeated manner with the corresponding nearest sample values ​​in the horizontal CTU boundary.

[0351] In the third method, according to the current line buffer setting in VVC, it is proposed to assume that 4 rows of luma samples and 2 rows of chroma samples above the horizontal CTU boundary from 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, 4) samples immediately before SAO, etc. are set in a mirrored manner, wherein the first row of luma samples and the first row of chroma samples above the horizontal CTU boundary are assumed to be the corresponding sample values ​​within the horizontal CTU boundary, the second row of luma samples and the second row of chroma samples above the horizontal CTU boundary are assumed to be the corresponding sample values ​​in the first row of samples below the horizontal CTU boundary, and so on.

[0352] It should be noted that 4 rows of luma samples and 2 rows of chroma samples above the horizontal CTU boundary are the current VVC line buffer settings, and the specific values ​​can be adjusted according to custom settings.

[0353] Reduction of the row buffer for additional CCALF input

[0354] According to one or more embodiments of the present disclosure, when the CCALF filter uses samples from 1) samples immediately before deblocking, 2) prediction samples, 3) residual samples, 4) samples immediately before SAO, etc. as additional inputs, a line buffer will be required to save these samples. In order to reduce the line buffer requirements of these additional inputs, according to the current line buffer settings in VVC, the 4 lines of corresponding luminance samples above the horizontal CTU boundary are assumed as default values, which can save these line buffers. Different methods can be used to achieve this goal.

[0355] In the first method, according to the current line buffer setting in VVC, it is proposed to assume that the 4 rows of luma residual samples above the horizontal CTU boundary are zero values, and the 4 rows of luma samples above the horizontal CTU boundary from 1) samples immediately before deblocking, 2) prediction samples, 3) samples immediately before SAO, etc. are assumed to be the same-position sample values ​​from the samples immediately after SAO.

[0356] In the second method, according to the current line buffer setting in VVC, it is proposed that 4 lines of luma samples above the horizontal CTU boundary from these stages such as 1) samples immediately before deblocking, 2) predicted samples, 3) residual samples, 4) samples immediately before SAO, etc. are assumed to be set in a repeating manner with the corresponding nearest sample values in the horizontal CTU boundary.

[0357] In the third method, according to the current line buffer setting in VVC, it is proposed that 4 lines of luma samples above the horizontal CTU boundary from these stages such as 1) samples immediately before deblocking, 2) predicted samples, 3) residual samples, 4) samples immediately before SAO, etc. are assumed to be set in a mirroring manner, where the first line of luma samples above the horizontal CTU boundary is assumed to be the corresponding sample value within the horizontal CTU boundary, the second line of luma samples above the horizontal CTU boundary is assumed to be the corresponding sample value in the first line of samples below the horizontal CTU boundary, and so on.

[0358] It should be noted that the 4 lines of luma samples above the horizontal CTU boundary are the current VVC line buffer setting, and the specific values can be adjusted according to custom settings.

[0359] Figure 17 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure. In step 1701, method 1700 includes: obtaining, by a decoder, at least one secondary signal by applying at least one fixed filter to a primary signal, where the at least one fixed filter is offline trained. In step 1702, method 1700 includes: obtaining, by the decoder, one or more spatially neighboring samples associated with a current sample, where the one or more spatially neighboring samples are from at least one of the primary signal or the at least one secondary signal. In step 1703, method 1700 includes: obtaining, by the decoder, filtered samples by applying at least one online filter to the one or more spatially neighboring samples associated with the current sample.

[0360] In an embodiment, the primary signal includes any one of a prediction signal, a residual signal, or a reconstructed signal, and the reconstructed signal includes at least one sample before sample adaptive offset (SAO) filtering.

[0361] In an embodiment, applying at least one fixed filter to the primary signal includes: determining, by the decoder, the at least one fixed filter using a first block-level classification result calculated for a signal including at least one sample after SAO filtering; or calculating, by the decoder, a second block-level classification result based on the primary signal.

[0362] In an embodiment, obtaining at least one secondary signal by applying at least one fixed filter to a primary signal includes: obtaining a plurality of secondary signals by applying a plurality of fixed filters to the primary signal, wherein the plurality of fixed filters are trained based on different block-level classifiers.

[0363] In an embodiment, a plurality of fixed filter banks are provided, and applying at least one fixed filter to a primary signal includes: the decoder determining at least one fixed filter by selecting the fixed filter bank indicated by a first set of indices, wherein the first set of indices is the same as a second set of indices for the signal after SAO filtering; or the decoder determining at least one fixed filter by selecting the fixed filter bank indicated by a first set of indices, wherein the first set of indices is different from the second set of indices for the signal after SAO filtering based on a predefined criterion; or in response to a set of indices received by the decoder from the encoder, determining at least one fixed filter by selecting the fixed filter bank indicated by the set of indices.

[0364] In an embodiment, two fixed filter banks are provided, and the first set of indices for the primary signal is different from the second set of indices for the signal after SAO filtering based on a predefined criterion; wherein method 1700 further includes: in response to the second set of indices for the signal after SAO filtering being 0, determining the first set of indices for the primary signal to be 1; or in response to the second set of indices for the signal after SAO filtering being 1, determining the first set of indices for the primary signal to be 0.

[0365] In an embodiment, one of the plurality of fixed filter banks includes two 13×13 diamond-shaped fixed filters.

[0366] In an embodiment, the primary signal includes a residual signal, and method 1700 further includes: the decoder truncating at least one secondary signal into at least one updated range.

[0367] In an embodiment, at least one updated range includes at least one of (-1024, 1024), (-512, 512), (-256, 256), or (-128, 128).

[0368] In an embodiment, applying at least one online filter to one or more spatially neighboring samples associated with a current sample includes: applying a plurality of online filters to one or more spatially neighboring samples, wherein the plurality of online filters are associated with multiple filter shapes.

[0369] In an embodiment, the multiple filter shapes include any one or any combination of 1×1, 3×3, or 5×5.

[0370] In an embodiment, the primary signal includes a prediction signal or a reconstructed signal before SAO filtering, and obtaining a filtered sample by applying at least one online filter to one or more spatially neighboring samples associated with a current sample includes: obtaining, by a decoder, an interception difference based on one or more spatially neighboring samples and the current sample; and obtaining, by the decoder, the filtered sample by applying at least one online filter to the interception difference.

[0371] In an embodiment, the primary signal includes a prediction signal or a reconstructed signal before SAO filtering, and obtaining a filtered sample by applying at least one online filter to one or more spatially neighboring samples associated with a current sample includes: obtaining, by a decoder, an interception difference based on one or more spatially neighboring samples and a co-located sample corresponding to the one or more spatially neighboring samples, and an interception difference based on the co-located sample and the current sample; and obtaining, by the decoder, the filtered sample by applying at least one online filter to the interception difference. In some embodiments, the one or more spatially neighboring samples may be from a primary signal or a secondary signal. If the one or more spatially neighboring samples are from the primary signal, the co-located sample is from the primary signal; if the spatially neighboring samples are from the secondary signal, the co-located sample is from the secondary signal.

[0372] In an embodiment, the primary signal includes a residual signal, and obtaining a filtered sample by applying at least one online filter to one or more spatially neighboring samples associated with a current sample includes: obtaining, by a decoder, an interception result based on one or more spatially neighboring samples; and obtaining, by the decoder, the filtered sample by applying at least one online filter to the interception result.

[0373] In an embodiment, obtaining one or more spatially neighboring samples associated with a current sample includes: determining, in response to performing a full intra test, one or more spatially neighboring samples from the primary signal; or determining, in response to performing a random access test, one or more spatially neighboring samples from the primary signal and at least one secondary signal.

[0374] Figure 18 is a flowchart illustrating a method for video coding according to some examples of the present disclosure. In step 1801, method 1800 includes: obtaining, by an encoder, at least one secondary signal by applying at least one fixed filter to a primary signal, where the at least one fixed filter is offline trained. In step 1802, method 1800 includes: obtaining, by the encoder, one or more spatially neighboring samples associated with a current sample, where the one or more spatially neighboring samples are from at least one of the primary signal or at least one secondary signal. In step 1803, method 1800 includes: obtaining, by the encoder, a filtered sample by applying at least one online filter to the one or more spatially neighboring samples associated with the current sample.

[0375] In an embodiment, the primary signal includes any one of a prediction signal, a residual signal, or a reconstruction signal, and the reconstruction signal includes at least one sample before sample adaptive offset (SAO) filtering.

[0376] In an embodiment, applying at least one fixed filter to the primary signal includes: determining, by an encoder, at least one fixed filter using a first block-level classification result calculated for a signal including at least one sample after SAO filtering; or calculating, by the encoder, a second block-level classification result based on the primary signal.

[0377] In an embodiment, obtaining at least one secondary signal by applying at least one fixed filter to the primary signal includes: obtaining a plurality of secondary signals by applying a plurality of fixed filters to the primary signal, where the plurality of fixed filters are trained based on different block-level classifiers.

[0378] In an embodiment, a plurality of fixed filter banks are provided, and applying at least one fixed filter to the primary signal includes: determining, by an encoder, at least one fixed filter by selecting a fixed filter bank indicated by a first set of indices, where the first set of indices is the same as a second set of indices for a signal after SAO filtering; or determining, by the encoder, at least one fixed filter by selecting a fixed filter bank indicated by a first set of indices, where the first set of indices is different from a second set of indices for a signal after SAO filtering based on a predefined criterion; or the encoder sending a set of indices to a decoder to indicate selection of a fixed filter bank among the plurality of fixed filter banks, where the set of indices is determined through a rate-distortion optimization process.

[0379] In an embodiment, two fixed filter banks are provided, and a first set of indices for the primary signal is different from a second set of indices for a signal after SAO filtering based on a predefined criterion; where method 1800 further includes: determining, in response to the second set of indices for a signal after SAO filtering being 0, that the first set of indices for the primary signal is 1; or determining, in response to the second set of indices for a signal after SAO filtering being 1, that the first set of indices for the primary signal is 0.

[0380] In an embodiment, one filter bank among the plurality of fixed filter banks includes two 13×13 diamond-shaped fixed filters.

[0381] In an embodiment, the primary signal includes a residual signal, and method 1800 further includes: the encoder truncating at least one secondary signal into at least one updated range.

[0382] In an embodiment, at least one updated range includes at least one of (-1024, 1024), (-512, 512), (-256, 256), or (-128, 128).

[0383] In an embodiment, applying at least one online filter to one or more spatially neighboring samples associated with a current sample includes: applying a plurality of online filters to one or more spatially neighboring samples, where the plurality of online filters are associated with a plurality of filter shapes.

[0384] In an embodiment, the plurality of filter shapes include any one or any combination of 1×1, 3×3, or 5×5.

[0385] In an embodiment, the primary signal includes a prediction signal or a reconstructed signal before SAO filtering, and obtaining a filtered sample by applying at least one online filter to one or more spatially neighboring samples associated with a current sample includes: obtaining, by an encoder, an interception difference based on one or more spatially neighboring samples and the current sample; and obtaining, by the encoder, the filtered sample by applying at least one online filter to the interception difference.

[0386] In an embodiment, the primary signal includes a prediction signal or a reconstructed signal before SAO filtering, and obtaining a filtered sample by applying at least one online filter to one or more spatially neighboring samples associated with a current sample includes: obtaining, by an encoder, an interception difference based on one or more spatially neighboring samples and corresponding co-located samples of the one or more spatially neighboring samples, and an interception difference based on the co-located samples and the current sample; and obtaining, by the encoder, the filtered sample by applying at least one online filter to the interception difference. In some embodiments, one or more spatially neighboring samples may be from the primary signal or the secondary signal. If one or more spatially neighboring samples are from the primary signal, the co-located samples are from the primary signal; if the spatially neighboring samples are from the secondary signal, the co-located samples are from the secondary signal.

[0387] In an embodiment, the primary signal includes a residual signal, and obtaining a filtered sample by applying at least one online filter to one or more spatially neighboring samples associated with a current sample includes: obtaining, by an encoder, an interception result based on one or more spatially neighboring samples; and obtaining, by the encoder, the filtered sample by applying at least one online filter to the interception result.

[0388] In an embodiment, obtaining one or more spatially neighboring samples associated with a current sample includes: determining, in response to performing a full intra-frame test, one or more spatially neighboring samples from the primary signal; or determining, in response to performing a random access test, one or more spatially neighboring samples from the primary signal and at least one secondary signal.

[0389] Figure 19 is a flowchart illustrating a method for video decoding according to some examples of the present disclosure. In step 1901, method 1900 includes: obtaining, by a decoder, a plurality of spatially neighboring samples associated with a current sample, wherein the plurality of spatially neighboring samples are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking. In step 1902, method 1900 includes: obtaining, by the decoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples. In step 1903, method 1900 includes: obtaining, by the decoder, a filtered sample based on the plurality of filtered input samples and the current sample.

[0390] In an embodiment, obtaining a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples includes: updating a plurality of luminance samples above a horizontal boundary of a coding tree unit and a plurality of chrominance samples above the horizontal boundary.

[0391] In an embodiment, the plurality of luminance samples include 4 rows of luminance samples above the horizontal boundary and 2 rows of chrominance samples above the horizontal boundary.

[0392] In an embodiment, the plurality of spatially neighboring samples are from a residual signal, and updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples to zero values.

[0393] In an embodiment, the plurality of spatially neighboring samples are from any one of a prediction signal, a signal before SAO filtering, or a signal before deblocking, and updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples to the co-located sample values of the samples after SAO.

[0394] In an embodiment, updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples by respectively copying the luminance samples of the corresponding rows and the chrominance samples of the corresponding rows, wherein the luminance samples of the corresponding rows and the chrominance samples of the corresponding rows are both located at the horizontal boundary in the coding tree unit.

[0395] In an embodiment, updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples by respectively mirroring the luminance samples of the corresponding rows and the chrominance samples of the corresponding rows, wherein the luminance samples of the corresponding rows and the chrominance samples of the corresponding rows are both located at the horizontal boundary in the coding tree unit or below the horizontal boundary.

[0396] Figure 20FIG. is a flowchart illustrating a method for video decoding according to some examples of the present disclosure. In step 2001, method 2000 includes: obtaining, by a decoder, a plurality of spatially neighboring samples associated with a current sample in a first channel, wherein the plurality of spatially neighboring samples are in a second channel and are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking. In step 2002, method 2000 includes: obtaining, by the decoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples. In step 2003, method 2000 includes: obtaining, by the decoder, a filtered sample in the first channel based on the plurality of filtered input samples and the current sample in the first channel.

[0397] In an embodiment, the first channel is a chrominance channel and the second channel is a luminance channel.

[0398] In an embodiment, obtaining a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples includes: updating a plurality of samples in the second channel that are above the horizontal boundary of the coding tree unit.

[0399] In an embodiment, the plurality of samples in the second channel includes 4 rows of samples in the second channel that are above the horizontal boundary.

[0400] In an embodiment, the plurality of spatially neighboring samples are from a residual signal, and updating the plurality of samples in the second channel includes: updating the plurality of samples in the second channel to zero values.

[0401] In an embodiment, the plurality of spatially neighboring samples are from any one of a prediction signal, a signal before SAO filtering, or a signal before deblocking, and updating the plurality of samples in the second channel includes: updating the plurality of samples in the second channel to the co-located sample values of the samples after SAO.

[0402] In an embodiment, updating the plurality of samples in the second channel includes: updating the plurality of samples in the second channel by copying the samples of the corresponding row at the horizontal boundary in the coding tree unit in the second channel.

[0403] In an embodiment, updating the plurality of samples in the second channel includes: updating the plurality of samples in the second channel by mirroring the samples of the corresponding row at or below the horizontal boundary in the coding tree unit in the second channel.

[0404] Figure 21FIG. is a flowchart of a method for video encoding according to some examples of the present disclosure. In step 2101, method 2100 includes: obtaining, by an encoder, a plurality of spatially neighboring samples associated with a current sample, wherein the plurality of spatially neighboring samples are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking. In step 2102, method 2100 includes: obtaining, by the encoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples. In step 2103, method 2100 includes: obtaining, by the encoder, a filtered sample based on the plurality of filtered input samples and the current sample.

[0405] In an embodiment, obtaining a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples includes: updating a plurality of luminance samples above a horizontal boundary of a coding tree unit and a plurality of chrominance samples above the horizontal boundary.

[0406] In an embodiment, the plurality of luminance samples include 4 rows of luminance samples above the horizontal boundary and 2 rows of chrominance samples above the horizontal boundary.

[0407] In an embodiment, the plurality of spatially neighboring samples are from a residual signal, and updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples to zero values.

[0408] In an embodiment, the plurality of spatially neighboring samples are from any one of a prediction signal, a signal before SAO filtering, or a signal before deblocking, and updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples to the co-located sample values of the samples after SAO.

[0409] In an embodiment, updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples by respectively copying the luminance samples of the corresponding rows and the chrominance samples of the corresponding rows, wherein the luminance samples of the corresponding rows and the chrominance samples of the corresponding rows are both located at the horizontal boundary in the coding tree unit.

[0410] In an embodiment, updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples by respectively mirroring the luminance samples of the corresponding rows and the chrominance samples of the corresponding rows, wherein the luminance samples of the corresponding rows and the chrominance samples of the corresponding rows are both located at the horizontal boundary in the coding tree unit or below the horizontal boundary.

[0411] Figure 22is a flowchart illustrating a method for video coding according to some examples of the present disclosure. In step 2201, method 2200 includes: obtaining, by an encoder, a plurality of spatially neighboring samples associated with a current sample in a first channel, wherein the plurality of spatially neighboring samples are in a second channel and are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking. In step 2202, method 2200 includes: obtaining, by the encoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples. In step 2203, method 2200 includes: obtaining, by the encoder, a filtered sample in the first channel based on the plurality of filtered input samples and the current sample in the first channel.

[0412] In an embodiment, the first channel is a chrominance channel, and the second channel is a luminance channel.

[0413] In an embodiment, obtaining a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples includes: updating a plurality of samples in the second channel that are above the horizontal boundary of a coding tree unit.

[0414] In an embodiment, the plurality of samples in the second channel include 4 rows of samples in the second channel that are above the horizontal boundary.

[0415] In an embodiment, the plurality of spatially neighboring samples are from a residual signal, and updating the plurality of samples in the second channel includes: updating the plurality of samples in the second channel to zero values.

[0416] In an embodiment, the plurality of spatially neighboring samples are from any one of a prediction signal, a signal before SAO filtering, or a signal before deblocking, and updating the plurality of samples in the second channel includes: updating the plurality of samples in the second channel to the co-located sample values of the samples after SAO.

[0417] In an embodiment, updating the plurality of samples in the second channel includes: updating the plurality of samples in the second channel by copying the samples of the corresponding row at the horizontal boundary in the coding tree unit in the second channel.

[0418] In an embodiment, updating the plurality of samples in the second channel includes: updating the plurality of samples in the second channel by mirroring the samples of the corresponding row at or below the horizontal boundary in the coding tree unit in the second channel.

[0419] Figure 23Shows a computing environment 2310 coupled to a user interface 2350. The computing environment 2310 can be part of a data processing server. The computing environment 2310 includes a processor 2320, a memory 2330, and an input / output (I / O) interface 2340.

[0420] The processor 2320 generally controls the overall operation of the computing environment 2310, such as operations associated with display, data acquisition, data communication, and image processing. The processor 2320 can include one or more processors for executing instructions to perform all or some of the steps in the above methods. In addition, the processor 2320 can include one or more modules that facilitate the interaction between the processor 2320 and other components. The processor can be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a graphics processing unit (GPU), etc.

[0421] The memory 2330 is configured to store various types of data to support the operation of the computing environment 2310. The memory 2330 can include a predetermined software 2332. Examples of the above data include instructions for any application or method operating on the computing environment 2310, video data sets, image data, etc. The memory 2330 can be implemented by using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.

[0422] The I / O interface 2340 provides an interface between the processor 2320 and peripheral interface modules (such as a keyboard, click wheel, buttons, etc.). The buttons can include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 2340 can be coupled to an encoder and a decoder.

[0423] In an embodiment, a non-transitory computer-readable storage medium is also provided, which includes, for example, a plurality of programs in the memory 2330 and / or stores a bitstream generated by the above encoding method or a bitstream to be decoded by the above decoding method. The plurality of programs can be executed by the processor 2320 in the computing environment 2310 to perform the above methods. In an embodiment, the plurality of programs can be executed by the processor 2320 in the computing environment 2310 to (for example, from Figure 2The video encoder 20) in receives a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.), and can also be executed by the processor 2320 in the computing environment 2310 to perform the above decoding method according to the received bitstream or data stream. In another example, multiple programs can be executed by the processor 2320 in the computing environment 2310 to perform the above encoding method to encode video information (e.g., video blocks representing video frames, and / or one or more associated syntax elements, etc.) into a bitstream or data stream, and can also be executed by the processor 2320 in the computing environment 2310 to (e.g., send Figure 3 the video decoder 30) in the bitstream or data stream. Alternatively, a non-transitory computer-readable storage medium may store a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames, and / or one or more associated syntax elements, etc.) generated by an encoder (e.g., Figure 2 the video encoder 20) in using, for example, the above encoding method for use by a decoder (e.g., Figure 3 the video decoder 30) in the bitstream or data stream when decoding video data. The non-transitory computer-readable storage medium may be, for example, ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0424] In an embodiment, there is provided a bitstream generated by the above encoding method or a bitstream to be decoded by the above decoding method. In an embodiment, there is provided a bitstream including encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method.

[0425] In an embodiment, there is further provided a computing device, which includes: one or more processors (e.g., the processor 2320); and a non-transitory computer-readable storage medium or memory 2330 in which multiple programs that can be executed by the one or more processors are stored, wherein the one or more processors are configured to execute the above method when executing the multiple programs.

[0426] In an embodiment, there is further provided a computer program product having instructions for storing or transmitting a bitstream, the bitstream including encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method. In an embodiment, there is further provided a computer program product including multiple programs in, for example, the memory 2330, the multiple programs can be executed by the processor 2320 in the computing environment 2310 to perform the above method. For example, the computer program product may include a non-transitory computer-readable storage medium.

[0427] In an embodiment, the computing environment 2310 may be implemented by one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described methods.

[0428] In an embodiment, a method for storing a bitstream is further provided, including: storing a bitstream on a digital storage medium, where the bitstream includes encoded video information generated by the above encoding method or encoded video information to be decoded by the above decoding method.

[0429] In an embodiment, a method for transmitting a bitstream generated by the above encoder is further provided. In an embodiment, a method for receiving a bitstream to be decoded by the above decoder is further provided.

[0430] The description of the present disclosure is presented for illustrative purposes and is not intended to be exhaustive or limiting to the present disclosure. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art from the above description and the associated drawings.

[0431] Unless otherwise specifically stated, the order of the steps of the method according to the present disclosure is only illustrative, and the steps of the method according to the present disclosure are not limited to the specific order described above, but may be changed according to the actual situation. In addition, at least one of the steps of the method according to the present disclosure may be adjusted, combined, or deleted according to actual needs.

[0432] The examples are selected and described to explain the principles of the present disclosure and to enable other technicians in the art to understand the various embodiments of the present disclosure and to best utilize the basic principles and various embodiments with various modifications suitable for the intended specific purposes. Therefore, it should be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed embodiments, and modifications and other embodiments are intended to be included within the scope of the present disclosure.

Claims

1. A method for video decoding, comprising: obtaining, by a decoder, at least one secondary signal by applying at least one fixed filter to a primary signal, wherein the at least one fixed filter is offline trained; obtaining, by the decoder, one or more spatially neighboring samples associated with a current sample, wherein the one or more spatially neighboring samples are from at least one of the primary signal or the at least one secondary signal; and obtaining, by the decoder, a filtered sample by applying at least one online filter to the one or more spatially neighboring samples associated with the current sample.

2. The method according to claim 1, wherein, The primary signal includes any one of a prediction signal, a residual signal, or a reconstructed signal, and the reconstructed signal includes at least one sample before sample adaptive offset (SAO) filtering.

3. The method according to claim 1, wherein Applying the at least one fixed filter to the primary signal includes: determining, by the decoder, the at least one fixed filter using a first block-level classification result calculated for a signal including at least one sample after SAO filtering; or calculating, by the decoder, a second block-level classification result based on the primary signal.

4. The method according to claim 1, wherein, Obtaining the at least one secondary signal by applying the at least one fixed filter to the primary signal includes: obtaining a plurality of secondary signals by applying a plurality of fixed filters to the primary signal, wherein the plurality of fixed filters are trained based on different block-level classifiers.

5. The method according to claim 1, wherein Providing a plurality of fixed filter banks, and applying the at least one fixed filter to the primary signal includes: determining, by the decoder, the at least one fixed filter by selecting a fixed filter bank indicated by a first set of indices, wherein the first set of indices is the same as a second set of indices for a signal after SAO filtering; or determining, by the decoder, the at least one fixed filter by selecting a fixed filter bank indicated by a first set of indices, wherein the first set of indices is different from a second set of indices for a signal after SAO filtering based on a predefined criterion; or determining, in response to a set of indices received by the decoder from an encoder, the at least one fixed filter by selecting a fixed filter bank indicated by the set of indices.

6. The method according to claim 5, wherein, Providing two fixed filter banks, and a first set of indices for the primary signal is different from a second set of indices for a signal after SAO filtering based on a predefined criterion; wherein the method further includes: determining, in response to the second set of indices for the signal after SAO filtering being 0, the first set of indices for the primary signal to be 1; or determining, in response to the second set of indices for the signal after SAO filtering being 1, the first set of indices for the primary signal to be 0.

7. The method according to claim 5, wherein, One filter bank of the plurality of fixed filter banks includes two 13×13 diamond fixed filters.

8. The method according to claim 1, wherein The primary signal includes a residual signal, and the method further includes: clipping, by the decoder, the at least one secondary signal to at least one updated range.

9. The method according to claim 8, wherein The range of the at least one update includes at least one of (-1024, 1024), (-512, 512), (-256, 256), or (-128, 128).

10. The method according to claim 1, wherein, Applying the at least one online filter to the one or more spatially neighboring samples associated with the current sample includes: Applying a plurality of online filters to the one or more spatially neighboring samples, wherein the plurality of online filters are associated with a plurality of filter shapes.

11. The method according to claim 1, wherein, The plurality of filter shapes include any one or any combination of 1×1, 3×3, or 5×5.

12. The method according to claim 1, wherein, The primary signal includes a prediction signal or a reconstructed signal before SAO filtering, and obtaining the filtered sample by applying the at least one online filter to the one or more spatially neighboring samples associated with the current sample includes: Obtaining, by the decoder, an interception difference based on the one or more spatially neighboring samples and the current sample; and Obtaining, by the decoder, the filtered sample by applying the at least one online filter to the interception difference.

13. The method according to claim 1, wherein, The primary signal includes a prediction signal or a reconstructed signal before SAO filtering, and obtaining the filtered sample by applying the at least one online filter to the one or more spatially neighboring samples associated with the current sample includes: Obtaining, by the decoder, an interception difference based on the one or more spatially neighboring samples and co-located samples corresponding to the one or more spatially neighboring samples, and an interception difference based on the co-located samples and the current sample; and Obtaining, by the decoder, the filtered sample by applying the at least one online filter to the interception difference.

14. The method according to claim 1, wherein, The primary signal includes a residual signal, and obtaining the filtered sample by applying the at least one online filter to the one or more spatially neighboring samples associated with the current sample includes: Obtaining, by the decoder, an interception result based on the one or more spatially neighboring samples; and Obtaining, by the decoder, the filtered sample by applying the at least one online filter to the interception result.

15. The method according to claim 1, wherein Obtaining the one or more spatially neighboring samples associated with the current sample includes: Determining the one or more spatially neighboring samples from the primary signal in response to performing a full intra test; or Determining the one or more spatially neighboring samples from the primary signal and the at least one secondary signal in response to performing a random access test.

16. A method for video coding, comprising: Obtaining, by an encoder, at least one secondary signal by applying at least one fixed filter to a primary signal, wherein the at least one fixed filter is offline trained; Obtaining, by the encoder, one or more spatially neighboring samples associated with a current sample, wherein the one or more spatially neighboring samples are from at least one of the primary signal or the at least one secondary signal; and Obtaining, by the encoder, a filtered sample by applying at least one online filter to the one or more spatially neighboring samples associated with the current sample.

17. The method according to claim 16, wherein, The main signal includes any one of a prediction signal, a residual signal, or a reconstruction signal, and the reconstruction signal includes at least one sample before sample adaptive offset (SAO) filtering.

18. The method according to claim 16, wherein, Applying the at least one fixed filter to the main signal includes: determining, by the encoder, the at least one fixed filter using a first block-level classification result calculated for a signal including at least one sample after SAO filtering; or calculating, by the encoder, a second block-level classification result based on the main signal.

19. The method according to claim 16, wherein, Obtaining the at least one secondary signal by applying the at least one fixed filter to the main signal includes: obtaining a plurality of secondary signals by applying a plurality of fixed filters to the main signal, where the plurality of fixed filters are trained based on different block-level classifiers.

20. The method according to claim 16, wherein, Providing a plurality of fixed filter banks, and applying the at least one fixed filter to the main signal includes: determining, by the encoder, the at least one fixed filter by selecting a fixed filter bank indicated by a first set of indices, where the first set of indices is the same as a second set of indices for a signal after SAO filtering; or determining, by the encoder, the at least one fixed filter by selecting a fixed filter bank indicated by a first set of indices, where the first set of indices is different from a second set of indices for a signal after SAO filtering based on a predefined criterion; or sending, by the encoder, a set of indices to a decoder to indicate selection of a fixed filter bank from the plurality of fixed filter banks, where the set of indices is determined by a rate-distortion optimization process.

21. The method according to claim 20, wherein, Providing two fixed filter banks, and a first set of indices for the main signal is different from a second set of indices for the signal after SAO filtering based on a predefined criterion; wherein the method further includes: in response to the second set of indices for the signal after SAO filtering being 0, determining the first set of indices for the main signal to be 1; or in response to the second set of indices for the signal after SAO filtering being 1, determining the first set of indices for the main signal to be 0.

22. The method according to claim 20, wherein, One filter bank of the plurality of fixed filter banks includes two 13×13 diamond-shaped fixed filters.

23. The method according to claim 16, wherein, The main signal includes a residual signal, and the method further includes: clipping, by the encoder, the at least one secondary signal to at least one updated range.

24. The method according to claim 16, wherein The at least one updated range includes at least one of (-1024, 1024), (-512, 512), (-256, 256), or (-128, 128).

25. The method according to claim 16, wherein, Applying the at least one online filter to the one or more spatially neighboring samples associated with the current sample includes: applying a plurality of online filters to the one or more spatially neighboring samples, where the plurality of online filters are associated with a plurality of filter shapes.

26. The method according to claim 16, wherein, The plurality of filter shapes include any one or any combination of 1×1, 3×3, or 5×5.

27. The method according to claim 16, wherein, The primary signal includes a prediction signal or a reconstructed signal before SAO filtering, and obtaining the filtered sample by applying the at least one online filter to one or more spatial neighboring samples associated with the current sample includes: obtaining, by the encoder, an interception difference based on the one or more spatial neighboring samples and the current sample; and obtaining, by the encoder, the filtered sample by applying the at least one online filter to the interception difference.

28. The method according to claim 16, wherein, The primary signal includes a prediction signal or a reconstructed signal before SAO filtering, and obtaining the filtered sample by applying the at least one online filter to one or more spatial neighboring samples associated with the current sample includes: obtaining, by the encoder, an interception difference based on the one or more spatial neighboring samples and co-located samples corresponding to the one or more spatial neighboring samples, and an interception difference based on the co-located samples and the current sample; and obtaining, by the encoder, the filtered sample by applying the at least one online filter to the interception difference.

29. The method according to claim 16, wherein, The primary signal includes a residual signal, and obtaining the filtered sample by applying the at least one online filter to one or more spatial neighboring samples associated with the current sample includes: obtaining, by the encoder, an interception result based on the one or more spatial neighboring samples; and obtaining, by the encoder, the filtered sample by applying the at least one online filter to the interception result.

30. The method according to claim 16, wherein, Obtaining the one or more spatial neighboring samples associated with the current sample includes: determining, in response to performing a full-frame intra test, the one or more spatial neighboring samples from the primary signal; or determining, in response to performing a random access test, the one or more spatial neighboring samples from the primary signal and the at least one secondary signal.

31. An apparatus for video decoding, the apparatus comprising: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors are configured to perform the method according to any one of claims 1 to 15 when executing the instructions.

32. An apparatus for video encoding, the apparatus comprising: one or more processors; and a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors are configured to perform the method according to any one of claims 16 to 30 when executing the instructions.

33. A non-transitory computer-readable storage medium for storing computer-executable instructions, the computer-executable instructions causing the one or more computer processors to perform the method according to any one of claims 1 to 15 when executed by the one or more computer processors.

34. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method according to any one of claims 16 to 30.

35. A non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method according to any one of claims 1 to 15.

36. A non-transitory computer-readable storage medium for storing a bitstream generated by the method according to any one of claims 16 to 30.

37. A method for video decoding, comprising: obtaining, by a decoder, a plurality of spatially neighboring samples associated with a current sample, wherein the plurality of spatially neighboring samples are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking; obtaining, by the decoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce a line buffer space for storing the plurality of spatially neighboring samples; and obtaining, by the decoder, a filtered sample based on the plurality of filtered input samples and the current sample.

38. The method according to claim 37, wherein, Obtaining the plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples includes: updating a plurality of luminance samples above a horizontal boundary of a coding tree unit and a plurality of chrominance samples above the horizontal boundary.

39. The method according to claim 38, wherein, The plurality of luminance samples include 4 rows of luminance samples above the horizontal boundary and 2 rows of chrominance samples above the horizontal boundary.

40. The method according to claim 38, wherein, The plurality of spatially neighboring samples are from the residual signal, and updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples to zero values.

41. The method according to claim 38, wherein, The plurality of spatially neighboring samples are from any one of the prediction signal, the signal before SAO filtering, or the signal before deblocking, and updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance to the co-located sample values of the samples after SAO.

42. The method according to claim 38, wherein Updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples by respectively copying the luminance samples of the corresponding row and the chrominance samples of the corresponding row, wherein the luminance samples of the corresponding row and the chrominance samples of the corresponding row are both at the horizontal boundary in the coding tree unit.

43. The method according to claim 38, wherein Updating the plurality of luminance samples and the plurality of chrominance samples includes: updating the plurality of luminance samples and the plurality of chrominance samples by respectively mirroring the luminance samples of the corresponding row and the chrominance samples of the corresponding row, wherein the luminance samples of the corresponding row and the chrominance samples of the corresponding row are both at or below the horizontal boundary in the coding tree unit.

44. A method for video decoding, comprising: The decoder obtains a plurality of spatially neighboring samples associated with a current sample in a first channel, where the plurality of spatially neighboring samples are in a second channel and are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking; The decoder obtains a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples; and The decoder obtains a filtered sample in the first channel based on the plurality of filtered input samples and the current sample in the first channel.

45. The method according to claim 44, wherein, The first channel is a chrominance channel, and the second channel is a luminance channel.

46. The method according to claim 44, wherein, Obtaining the plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples includes: Updating a plurality of samples in the second channel that are above the horizontal boundary of the coding tree unit.

47. The method according to claim 46, wherein, The plurality of samples in the second channel include 4 rows of samples above the horizontal boundary in the second channel.

48. The method according to claim 46, wherein, The plurality of spatially neighboring samples are from the residual signal, and Updating the plurality of samples in the second channel includes: Updating the plurality of samples in the second channel to zero values.

49. The method according to claim 46, wherein, The plurality of spatially neighboring samples are from any one of the prediction signal, the signal before SAO filtering, or the signal before deblocking, and Updating the plurality of samples in the second channel includes: Updating the plurality of samples in the second channel to the co-located sample values of the samples after SAO.

50. The method according to claim 46, wherein, Updating the plurality of samples in the second channel includes: Updating the plurality of samples in the second channel by copying the samples of the corresponding row at the horizontal boundary in the coding tree unit in the second channel.

51. The method according to claim 46, wherein, Updating the plurality of samples in the second channel includes: Updating the plurality of samples in the second channel by mirroring the samples of the corresponding row at or below the horizontal boundary in the coding tree unit in the second channel.

52. A method for video coding, including: The encoder obtains a plurality of spatially neighboring samples associated with a current sample, where the plurality of spatially neighboring samples are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking; The encoder obtains a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples; and The encoder obtains a filtered sample based on the plurality of filtered input samples and the current sample.

53. The method according to claim 52, wherein, Obtaining the plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the line buffer space for storing the plurality of spatially neighboring samples includes: Update a plurality of luminance samples above the horizontal boundary of a coding tree unit and a plurality of chrominance samples above the horizontal boundary.

54. The method according to claim 53, wherein, The plurality of luminance samples include 4 rows of luminance samples above the horizontal boundary and 2 rows of chrominance samples above the horizontal boundary.

55. The method according to claim 53, wherein, The plurality of spatially neighboring samples are from the residual signal, and Updating the plurality of luminance samples and the plurality of chrominance samples includes: Updating the plurality of luminance samples and the plurality of chrominance samples to zero values.

56. The method according to claim 53, wherein, The plurality of spatially neighboring samples are from any one of the prediction signal, the signal before SAO filtering, or the signal before deblocking, and Updating the plurality of luminance samples and the plurality of chrominance samples includes: Updating the plurality of luminance samples and the plurality of chrominance to the co-located sample values of the samples after SAO.

57. The method according to claim 53, wherein, Updating the plurality of luminance samples and the plurality of chrominance samples includes: Updating the plurality of luminance samples and the plurality of chrominance samples by respectively copying the luminance samples of the corresponding row and the chrominance samples of the corresponding row, wherein the luminance samples of the corresponding row and the chrominance samples of the corresponding row are both located at the horizontal boundary in the coding tree unit.

58. The method according to claim 53, wherein Updating the plurality of luminance samples and the plurality of chrominance samples includes: Updating the plurality of luminance samples and the plurality of chrominance samples by respectively mirroring the luminance samples of the corresponding row and the chrominance samples of the corresponding row, wherein the luminance samples of the corresponding row and the chrominance samples of the corresponding row are both located at the horizontal boundary in the coding tree unit or below the horizontal boundary.

59. A method for video coding, comprising: Obtaining, by an encoder, a plurality of spatially neighboring samples associated with a current sample in a first channel, wherein the plurality of spatially neighboring samples are in a second channel and are from any one of a prediction signal, a residual signal, a signal before sample adaptive offset (SAO) filtering including a plurality of samples before SAO filtering, or a signal before deblocking including a plurality of samples before deblocking; Obtaining, by the encoder, a plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the row buffer space for storing the plurality of spatially neighboring samples; and Obtaining, by the encoder, a filtered sample in the first channel based on the plurality of filtered input samples and the current sample in the first channel.

60. The method according to claim 59, wherein, The first channel is a chrominance channel, and the second channel is a luminance channel.

61. The method according to claim 59, wherein, Obtaining the plurality of filtered input samples by updating the plurality of spatially neighboring samples to reduce the row buffer space for storing the plurality of spatially neighboring samples includes: Updating a plurality of samples above the horizontal boundary of a coding tree unit in the second channel.

62. The method according to claim 61, wherein, The plurality of samples in the second channel include 4 rows of samples above the horizontal boundary in the second channel.

63. The method according to claim 61, wherein, The plurality of spatially neighboring samples are from the residual signal, and Updating the plurality of samples in the second channel includes: Updating the plurality of samples in the second channel to zero values.

64. The method according to claim 61, wherein, The plurality of spatially neighboring samples are from any one of the prediction signal, the signal before SAO filtering, or the signal before deblocking, and Updating the plurality of samples in the second channel includes: Updating the plurality of samples in the second channel to the co-located sample values of the samples after SAO.

65. The method according to claim 61, wherein, Updating the plurality of samples in the second channel includes: Updating the plurality of samples in the second channel by copying the samples of the corresponding row located at the horizontal boundary in the coding tree unit in the second channel.

66. The method according to claim 61, wherein, Updating the plurality of samples in the second channel includes: Updating the plurality of samples in the second channel by mirroring the samples of the corresponding row located at or below the horizontal boundary in the coding tree unit in the second channel.

67. An apparatus for video decoding, the apparatus comprising: One or more processors; And A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors are configured to execute the method according to any one of claims 37 to 43 when executing the instructions.

68. An apparatus for video encoding, the apparatus comprising: One or more processors; And A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors are configured to execute the method according to any one of claims 52 to 58 when executing the instructions.

69. An apparatus for video decoding, the apparatus comprising: One or more processors; And A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors are configured to execute the method according to any one of claims 44 to 51 when executing the instructions.

70. An apparatus for video encoding, the apparatus comprising: One or more processors; And A memory coupled to the one or more processors and configured to store instructions executable by the one or more processors, wherein the one or more processors are configured to execute the method according to any one of claims 59 to 66 when executing the instructions.

71. A non-transitory computer-readable storage medium for storing computer-executable instructions, the computer-executable instructions causing the one or more computer processors to execute the method according to any one of claims 37 to 43 when executed by the one or more computer processors.

72. A non-transitory computer-readable storage medium for storing computer-executable instructions, the computer-executable instructions causing the one or more computer processors to execute the method according to any one of claims 44 to 51 when executed by the one or more computer processors.

73. A non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method according to any one of claims 37 to 43.

74. A non-transitory computer-readable storage medium for storing a bitstream generated by the method according to any one of claims 44 to 51.

75. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to execute the method according to any one of claims 52 to 58.

76. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to execute the method according to any one of claims 59 to 66.

77. A non-transitory computer-readable storage medium for storing a bitstream to be decoded by the method according to any one of claims 52 to 58.

78. A non-transitory computer-readable storage medium for storing a bitstream generated by the method according to any one of claims 59 to 66.

Citation Information

Cited By

  • Adaptive compensation method in video image coding process and related device

    CN121126002A

  • Adaptive compensation method in video image encoding process and related device

    CN121126002B