Cross-component sample offset filtering with interpolation filter taps

By using cross-component sample offset (CCSO) filtering technology and interpolation filter in video encoding and decoding, the problem of poor cross-component sample offset processing in the prior art is solved, and higher video quality and encoding efficiency are achieved.

CN120202664APending Publication Date: 2025-06-24TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079525.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2023-11-30
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are difficult to effectively handle cross-component sample offsets, resulting in reduced video quality and low encoding efficiency.

Method used

A cross-component sample offset (CCSO) filtering technology is used to generate sample offsets at fractional pixel positions through a CCSO filter and apply it to reconstruct the second color component of the video. Interpolation samples are generated in combination with an interpolation filter to improve accuracy.

Benefits of technology

Improve the quality and efficiency of video encoding and decoding, reduce encoding artifacts and improve the detail fidelity of reconstructed videos by effectively processing sample offsets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120202664A_ABST
    Figure CN120202664A_ABST
Patent Text Reader

Abstract

The present disclosure generally describes a set of advanced video coding techniques, and in particular relates to cross-component sample offset (CCSO) filtering. For example, a CCSO filter with fractional CCSO filtering tap positions may be used to generate sample offsets from one color component of the reconstructed video. The sample offset may then be applied to a second color component of the reconstructed video. The reconstructed sample of the first color component may be first interpolated by an interpolation filter to generate an interpolated sample of the first color component at the fractional CCSO tap position. The interpolated sample may be used as an input to a CCSO filter for generating a sample offset.
Need to check novelty before this filing date? Find Prior Art

Description

INCORPORATION BY REFERENCE

[0001] This disclosure claims priority benefit of U.S. Non - Provisional Application No. 18 / 523,374, filed on November 29, 2023, entitled "CROSS COMPONENT SAMPLE OFFSET FILTERING WITH INTERPOLATED FILTER TAPS", and U.S. Provisional Application No. 63 / 536,012, filed on August 31, 2023, entitled "CCSO WITH INTERPOLATED FILTER TAPS", the entire contents of both of which are hereby incorporated by reference. TECHNICAL FIELD

[0002] This disclosure generally describes a set of advanced video decoding techniques and, in particular, relates to cross - component sample offset (CCSO) filtering. BACKGROUND ART

[0003] Uncompressed digital video can include a sequence of pictures and may have specific bit - rate requirements for storage, data - processing, and transmission bandwidth in streaming applications. One purpose of video encoding and decoding can be to reduce redundancy in the uncompressed input video signal through various compression techniques. SUMMARY OF THE DISCLOSURE

[0004] This disclosure generally describes a set of advanced video encoding and decoding techniques and, in particular, relates to cross - component sample offset (CCSO) filtering. For example, a CCSO filter with fractional CCSO filter tap positions can be used to generate a sample offset based on one color component of a reconstructed video. The sample offset can then be applied to a second color component of the reconstructed video. The reconstructed samples of the first color component can first be interpolated by an interpolation filter to generate interpolated samples of the first color component at the fractional CCSO tap positions. The interpolated samples can be used as inputs to the CCSO filter for generating the sample offset.

[0005] In some example implementations, a method is disclosed. The method may include: reconstructing a video frame from a video bitstream to generate reconstruction samples of at least a first color component and a second color component of the video frame; receiving at least one syntax element from the video bitstream to determine a cross-component sample offset (CCSO) filter, the CCSO filter being applicable to the reconstruction samples of a CCSO filtering unit in the first color component, the CCSO filter being characterized by a set of CCSO filter tap positions relative to a center tap; in response to the CCSO filter including CCSO filter tap positions at fractional pixel positions, determining an interpolation filter and applying the interpolation filter to the reconstruction samples of the CCSO filtering unit in the first color component to generate interpolation samples of the CCSO filtering unit in the first color component at the fractional pixel positions; generating a sample offset for each sample position in the CCSO filtering unit by processing the reconstruction samples and the interpolation samples of the CCSO filtering unit in the first color component using the CCSO filter; and applying the sample offset to the reconstruction samples of the CCSO filtering unit in the second color component.

[0006] In the above example implementations, the at least one syntax element includes an indication of the CCSO filter tap positions at the fractional pixel positions. The interpolation filter is used to perform multi-tap interpolation. The interpolation filter may for example include a 2-tap, 4-tap, 6-tap or 8-tap interpolation filter.

[0007] In any of the above example implementations, the number of taps in the interpolation filter for the multi-tap interpolation depends on the position of the reconstruction samples in the CCSO filtering unit relative to the boundary of the CCSO filtering unit.

[0008] In any of the above example implementations, the interpolation filter is used to: perform the multi-tap interpolation in the horizontal direction only when the CCSO filter includes only fractional tap positions in the horizontal direction; perform the multi-tap interpolation in the vertical direction only when the CCSO filter includes only fractional tap positions in the vertical direction; or, when the CCSO filter includes fractional tap positions in the horizontal direction and fractional tap positions in the vertical direction, perform the multi-tap interpolation by interpolating in one of the vertical and horizontal directions to generate intermediate interpolation samples and then interpolating the intermediate interpolation samples in the other of the vertical and horizontal directions.

[0009] In any of the above example implementations, the indication of the CCSO filter tap position at the fractional pixel position includes a first signaling syntax element and a second signaling syntax element, where the first signaling syntax element and the second signaling syntax element are respectively used to indicate the CCSO fractional pixel accuracy for determining the CCSO filter tap position at the fractional pixel position and the phase at the CCSO fractional pixel accuracy.

[0010] In any of the above example implementations, the first signaling syntax element includes an index for indicating one of a set of predetermined CCSO fractional pixel accuracies, and the set of predetermined CCSO fractional pixel accuracies includes at least one of 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, 1 / 8 pixel accuracy, 1 / 16 pixel accuracy, and 1 / 32 accuracy.

[0011] In any of the above example implementations, the second signaling syntax element includes a phase index from a set of predefined phases determined according to the CCSO fractional pixel accuracy indicated by the first signaling syntax element.

[0012] In any of the above example implementations, the at least one syntax element further includes a flag located before the first signaling syntax element and the second signaling syntax element, and the flag is used to indicate that the CCSO filter includes one or more fractional tap positions.

[0013] In any of the above example implementations, the CCSO filter indicated by the at least one syntax element is selected from a plurality of allowed CCSO filters, and the plurality of allowed CCSO filters do not overlap in the tap direction.

[0014] In any of the above example implementations, the interpolation filter is only applied to one or more selected color components. The one or more selected color components are predetermined. The one or more selected color components may include a luminance color component. The second color component includes one of a plurality of chrominance components.

[0015] In any of the above example implementations, the set of CCSO filter tap positions is implicitly derived using adjacent samples of the current block.

[0016] In any of the above example implementations, the at least one syntax element is used to: indicate a selection from a CCSO filter having fractional tap positions or a CCSO filter having only integer tap positions using a separate index space. Alternatively, the at least one syntax element is used to: indicate a selection from a CCSO filter having fractional tap positions or a CCSO filter having only integer tap positions using a single index space.

[0017] In some implementations, a video encoding or decoding device is disclosed. The device may include circuitry configured to implement any of the above methods.

[0018] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform a method for video decoding and / or encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0020] Figure 1 A schematic diagram showing a simplified block diagram of a communication system (100) according to an example embodiment is shown.

[0021] Figure 2 A schematic diagram showing a simplified block diagram of a communication system (200) according to an example embodiment is shown.

[0022] Figure 3 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment is shown.

[0023] Figure 4 A schematic diagram showing a simplified block diagram of a video encoder according to an example embodiment is shown.

[0024] Figure 5 A block diagram of a video encoder according to another example embodiment is shown.

[0025] Figure 6 A block diagram of a video decoder according to another example embodiment is shown.

[0026] Figure 7 An exemplary adaptive loop filter according to an embodiment of the present disclosure is shown.

[0027] Figures 8A to 8D Examples of downsampling positions for calculating gradients in the vertical, horizontal, and two diagonal directions according to an embodiment of the present disclosure are shown.

[0028] Figure 8EShows an example way of determining block directionality based on various gradients for use in an Adaptive Loop Filter (ALF).

[0029] Figure 9A And Figure 9B Shows a modified block classification at a virtual boundary according to an example embodiment of the present disclosure.

[0030] Figures 10A to 10F Shows an exemplary adaptive loop filter having a padding operation at a corresponding virtual boundary according to an embodiment of the present disclosure.

[0031] Figure 11 Shows an example of picture quadtree segmentation with maximum coding unit alignment according to an embodiment of the present disclosure.

[0032] Figure 12 Shows corresponding to according to an example embodiment of the present disclosure Figure 11 The quadtree segmentation pattern.

[0033] Figure 13 Shows a cross-component filter for generating a chrominance component according to an example embodiment of the present disclosure.

[0034] Figure 14 Shows an example of a cross-component ALF filter according to an embodiment of the present disclosure.

[0035] Figure 15 Shows an exemplary position of chrominance samples relative to luminance samples according to an embodiment of the present disclosure.

[0036] Figure 16 Shows an example of direction search for a block according to an embodiment of the present disclosure.

[0037] Figure 17 Shows an example of subspace projection according to an embodiment of the present disclosure.

[0038] Figure 18 Shows an example position of cross-component sample offset CCSO filtering in the loop filter process.

[0039] Figure 19 Shows an example of a filter support region in a CCSO filter according to an embodiment of the present disclosure.

[0040] Figure 20 Shows an example implementation of the shape of a 3-tap CCSO filter according to an embodiment of the present disclosure.

[0041] Figure 21 Shows an example interpolation CCSO filter tap for CCSO filtering.

[0042] Figure 22 Shows examples of allowed 1 / 4 pixel interpolation CCSO filter taps and integer pixel interpolation CCSO filter taps for CCSO filtering.

[0043] Figure 23 Shows a flowchart outlining a process (2300) according to an embodiment of the present disclosure.

[0044] Figure 24 Shows a schematic diagram of a computer system according to an example embodiment. Detailed Description

[0045] Throughout the specification and claims, terms may have nuanced meanings that are suggested or implied in context, in addition to the explicitly stated meanings. The phrases "in one embodiment / implementation" or "in some embodiments / implementations" as used herein do not necessarily refer to the same embodiment / implementation, and the phrases "in another embodiment / implementation" or "in other embodiments" as used herein do not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include all or part of the combination of exemplary embodiments / implementations.

[0046] Generally, terms can be understood, at least in part, based on their usage in context. For example, terms such as "and", "or", or "and / or" as used herein can have various context-dependent meanings. Generally, if "or" is used to relate a list (such as A, B, or C), it is intended to mean A, B, and C (used in an inclusive sense here) as well as A, B, or C (used in an exclusive sense here). Additionally, the terms "one or more", "at least one", "a", "an", or "the" as used herein can be used in the singular or plural, at least in part depending on the context. Furthermore, the terms "based on" or "determined by" can be understood to not necessarily convey a set of exclusive factors, but can allow for additional factors that are not necessarily described again explicitly, at least in part depending on the context.

[0047] Figure 1 Shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The communication system (100) includes a plurality of terminal devices (such as 110, 120, 130, and 140) that can communicate with each other via, for example, a network (150). In Figure 1In the example, the first pair of terminal devices (110) and (120) can perform unidirectional data transmission. For example, the terminal device (110) can encode video data (such as a video picture stream captured by the terminal device (110)) into one or more encoded bitstreams for transmission over the network (150). The terminal device (120) can receive the encoded video data from the network (250), decode the encoded video data to recover the video pictures, and display the video pictures based on the recovered video data. Unidirectional data transmission can be implemented in applications such as media services.

[0048] In another example, the second pair of terminal devices (130) and (140) can perform two-way transmission of encoded video data, for example, during a video conferencing application. In the example, for two-way data transmission, each of the terminal devices (130) and (140) can encode video data (such as a video picture stream captured by the terminal device) for transmission, and each of the terminal devices (130) and (140) can also receive the encoded video data from the other terminal device among the terminal devices (130) and (140) to recover and display the video pictures.

[0049] In Figure 1 the example, the terminal device can be implemented as a server, a personal computer, and a smart phone, but the applicability of the basic principles of the present disclosure is not limited thereto. Embodiments of the present disclosure can be implemented in a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, a dedicated video conferencing device, etc. The network (150) represents any number or type of network that conveys encoded video data between terminal devices, including, for example, wired (wired) and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched channels, packet-switched channels, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0050] As an example of the application of the disclosed subject matter, Figure 2 shows the placement of a video encoder and a video decoder in a video streaming environment. The disclosed subject matter is equally applicable to other video applications, including, for example, video conferencing, digital TV broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0051] As Figure 2As shown, the video streaming system may include a video capture subsystem (213), which may include a video source (201) such as a digital camera. The video source creates an uncompressed video picture stream (202) or video image stream (202). In an example, the video picture stream (202) includes samples taken by the digital camera of the video source 201. Compared with the encoded video data (204) (or encoded video bitstream), the video picture stream (202) is depicted as a thick line to emphasize the high data volume and can be processed by an electronic device (220), which includes a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the uncompressed video picture stream (202), the encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize the lower data volume and may be stored on a streaming server (205) for future use. One or more streaming client subsystems (such as Figure 2 the client subsystem (206) and the client subsystem (208) in

[0052] Figure 3 FIG. shows a block diagram of a video decoder (310) of an electronic device (330) according to any embodiment of the present disclosure below. The electronic device (330) may include a receiver (331) (such as a receiving circuit). The video decoder (310) may be used in place of Figure 2 the video decoder (210) in the example.

[0053] As Figure 3As shown, a receiver (331) may receive one or more encoded video sequences from a channel (301). To prevent network jitter and / or handle playback timing, a buffer memory (315) may be provided between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) may reconstruct symbols (321) from the encoded video sequences. The categories of these symbols include information for managing the operation of a video decoder (310), as well as potential information for controlling a rendering device such as a display (312) (e.g., a display screen). The parser (320) may parse / entropy decode the encoded video sequences. The parser (320) may extract from the encoded video sequences subgroup parameter sets for at least one subgroup of pixels in the video decoder. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (320) may also extract information from the encoded video sequences, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. The reconstruction of the symbols (321) may involve multiple different processing or functional units. The units involved and the manner in which they are involved may be controlled by subgroup control information parsed by the parser (320) from the encoded video sequences.

[0054] A first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive from the parser (320) quantized transform coefficients as symbols (321) and control information, including information indicating which type of inverse transform is to be used, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block including sample values, which may be input into an aggregator (355).

[0055] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra picture prediction unit (352). In some cases, the intra picture prediction unit (352) may use the surrounding block information that has been reconstructed and stored in the current picture buffer (358) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (358) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (355) may add the predictive information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.

[0056] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-coded and potentially motion-compensated block. In this case, the motion compensation prediction unit (353) may access the reference picture memory (357) based on the motion vectors to extract samples for inter picture prediction. After motion compensating the extracted reference samples according to the sign (321) belonging to the block, these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (the output of unit 351 may be referred to as residual samples or a residual signal), thereby generating output sample information.

[0057] The output samples of the aggregator (355) may be employed in the loop filter unit (356) by various loop filtering techniques, and the loop filter unit (356) includes various types of loop filters. The output of the loop filter unit (356) may be a sample stream, which may be output to the rendering device (312) and stored in the reference picture memory (357) for subsequent inter picture prediction.

[0058] Figure 4 A block diagram of a video encoder (403) according to an example embodiment of the present disclosure is shown. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may further include a transmitter (440) (e.g., a transmission circuit). The video encoder (403) may be used to replace Figure 4 the video encoder (403) in the example.

[0059] A video encoder (403) may receive video samples from a video source (401). According to some example embodiments, the video encoder (403) may encode and compress pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Enforcing an appropriate encoding speed is a function of a controller (450). In some embodiments, the controller (450) may be functionally coupled to and control other functional units as described below. Parameters set by the controller (450) may include rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.

[0060] In some example embodiments, the video encoder (403) may be used to operate in an encoding loop. The encoding loop may include a source encoder (430) and a (local) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs symbols to create sample data in a manner similar to how a (remote) decoder creates sample data, although the embedded decoder 433 processes an unentropy - coded encoded video stream encoded by the source encoder 430 (because in the video compression techniques contemplated by the disclosed subject matter, any compression between symbols and the encoded video bitstream during the entropy encoding process may be lossless). At this point, it can be observed that, except for parsing / entropy decoding which may only exist in the decoder, any decoder technique must also exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations, which are related to the decoding part of the encoder. Thus, the description of encoder techniques may be simplified because encoder techniques are inverse to the decoder techniques described comprehensively. Only more detailed descriptions of the encoder are provided in specific fields or aspects below.

[0061] During operation, in some example implementations, the source encoder (430) may perform motion - compensated predictive coding. This motion - compensated predictive coding performs predictive coding on an input picture by referring to one or more previously encoded pictures in the video sequence designated as "reference pictures".

[0062] The local video decoder (433) may decode the encoded video data of pictures that may be designated as reference pictures. The local video decoder (433) replicates the decoding process that may be performed by a video decoder on a reference picture and may cause the reconstructed reference picture to be stored in a reference picture cache (434). In this way, the video encoder (403) may locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture that will be obtained by a distal (remote) video decoder.

[0063] The predictor (435) may perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) may search in the reference picture memory (434) for sample data (as candidate reference pixel blocks) or certain metadata that can serve as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc.

[0064] The controller (450) may manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.

[0065] The outputs of all the above functional units may be entropy encoded in the entropy encoder (445). The transmitter (440) may buffer the encoded video sequence created by the entropy encoder (445) to prepare for transmission over a communication channel (460), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) may merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0066] The controller (450) may manage the operations of the video encoder (403). During encoding, the controller (450) may assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures may generally be assigned to any of the following picture types: intra picture (I picture), predicted picture (P picture), bi-predicted picture (B picture), multi-predicted picture. The source pictures may generally be spatially subdivided into a plurality of sample coding blocks, as described in further detail below.

[0067] Figure 5 A diagram of a video encoder (503) according to another example embodiment of the present disclosure is shown. The video encoder (503) is for receiving sample values within a current video picture in a sequence of video pictures in a processing block (e.g., a prediction block), and encoding the processing block into an encoded picture that is part of an encoded video sequence. The example video encoder (403) may be used in place of Figure 4 the video encoder (403) in the example.

[0068] For example, the video encoder (503) receives a matrix of sample values for a processing block. Then the video encoder (503) uses, for example, rate-distortion optimization (RDO) to determine whether it is best to encode the processing block using an intra mode, an inter mode, or a bi-prediction mode.

[0069] In Figure 5In the example of FIG. 5 , the video encoder ( 503 ) includes: Figure 5 The example arrangement shown is an inter-frame encoder (530), an intra-frame encoder (522), a residual calculator (523), a switch (526), ​​a residual encoder (524), a common controller (521), and an entropy encoder (525) coupled together.

[0070] The inter-frame encoder (530) is used to perform the following operations: receive samples of a current block (e.g., a processing block); compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture in a display order); generate inter-frame prediction information (e.g., redundant information description, motion vector, merge mode information according to the inter-frame coding technology); and calculate the inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technology.

[0071] The intra-frame encoder (522) is used to perform the following operations: receive samples of a current block (e.g., a processing block); compare the block with encoded blocks in the same picture; generate quantization coefficients after transformation; and in some cases also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to one or more intra-frame coding techniques).

[0072] The general controller (521) may be used to determine general control data and control other components of the video encoder (503) based on the general control data, such as determining a prediction mode for a block and providing a control signal to a switch (526) based on the prediction mode.

[0073] The residual calculator (523) may be used to calculate the difference (residual data) between the received block and the prediction result of the block selected from the intra encoder (522) or the inter encoder (530). The residual encoder (524) may be used to encode the residual data to generate a transform coefficient. The transform coefficient is then subjected to a quantization process to obtain a quantized transform coefficient. In various example embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is used to perform an inverse transform and generate decoded residual data. The entropy encoder (525) may be used to format the bitstream to include the encoded blocks and perform entropy encoding.

[0074] Figure 6 A diagram of an example video decoder (610) according to another embodiment of the present disclosure is shown. The video decoder (610) is used to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an example, the video decoder (610) may be used instead of Figure 4 A video decoder (410) is shown in an example.

[0075] exist Figure 6In the example of, the video decoder (610) includes an entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-frame decoder (672) coupled together as shown in the example arrangement of Figure 6 The entropy decoder (671) can be used to reconstruct certain symbols according to the encoded picture, and these symbols represent the syntax elements that make up the encoded picture. The inter-frame decoder (680) can be used to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information. The intra-frame decoder (672) can be used to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information. The residual decoder (673) can be used to perform inverse quantization to extract the dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The reconstruction module (674) can be used to combine the residual output by the residual decoder (673) with the prediction result (which can be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, and the reconstructed block forms a part of the reconstructed picture, and the reconstructed picture is in turn a part of the reconstructed video.

[0076] It should be noted that any suitable technology can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In some example embodiments, one or more integrated circuits can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). One or more processors executing software instructions can be used to implement (810).

[0077] In some example implementations, a loop filter can be included in the encoder and decoder to reduce coding artifacts and improve the quality of the decoded picture. For example, the loop filter 356 can be included as

[0078] a part of the decoder 330 of Figure 3 . For another example, the loop filter can be Figure 4Part of the embedded decoder unit 433 in the encoder 420. These filters are referred to as loop filters because they are included in the decoding loop for video blocks in the decoder or encoder. Each loop filter can be associated with one or more filtering parameters. These filtering parameters can be predefined or can be derived by the encoder during the encoding process. These filtering parameters (if derived by the encoder) or their indices (if predefined) can be included in the final bitstream in coded form. The decoder can then parse these filtering parameters from the bitstream and perform loop filtering during decoding based on the parsed filtering parameters.

[0079] In various aspects, various loop filters can be used to reduce encoding artifacts and improve decoded video quality. Such loop filters can include, but are not limited to, one or more deblocking filters, adaptive loop filters (ALF), cross-component adaptive loop filters (CC-ALF), constrained directional enhancement filters (CDEF), sample adaptive offset (SAO) filters, cross-component sample offset (CCSO) filters, and local sample offset (LSO) filters. These filters can be interdependent or not. These filters can be arranged in the decoding loop of the decoder or encoder in any suitable order compatible with their interdependence (if any). These various loop filters are described in more detail in the following disclosure.

[0080] The encoder / decoder can apply an adaptive loop filter (ALF) with block-based filter adaptation to reduce artifacts. The adaptability of the ALF is reflected in that the filtering coefficients / parameters or their indices are written into the bitstream and can be designed based on the image content and the distortion of the reconstructed picture. The ALF can be applied to reduce the distortion introduced by the encoding process and improve the quality of the reconstructed image.

[0081] For the luminance component, for example, one filter from a plurality of filters (e.g., 25 filters) can be selected for a luminance block (e.g., a 4×4 luminance block) based on the direction and activity of the local gradient. The filter coefficients of these filters can be derived by the encoder during the encoding process and signaled to the decoder in the bitstream.

[0082] The ALF can have any suitable shape and size. Refer to Figure 7For example, the ALF can have a diamond shape. For instance, ALF(710) is a 5×5 diamond shape and ALF(711) is a 7×7 diamond shape. In ALF(710), thirteen (13) elements can be used during the filtering process, which form a diamond shape. For these 13 elements, seven values (such as C0 - C6) can be used and arranged in the manner of the shown example. In ALF(711), twenty-five (25) elements can be used during the filtering process, which form a diamond shape. For these 25 elements, thirteen (13) values (such as C0 - C12) can be used and arranged in the manner of the shown example.

[0083] Reference Figure 7 , in some examples, one of the two diamond shapes (710) and (711) of the ALF filter can be selected for processing luminance or chrominance blocks. For example, the 5×5 diamond shape filter (710) can be applied to the chrominance component (e.g., chrominance block, chrominance CB), while the 7×7 diamond shape filter (711) can be applied to the luminance component (e.g., luminance block, luminance CB). In the ALF, other suitable shapes and sizes can be used. For example, a 9×9 diamond shape filter can be used.

[0084] The filter coefficients at the positions indicated by the values (e.g., C0 - C6 in (710) or C0 - C12 in (711)) can be non-zero. Additionally, when the ALF includes a clipping function, the clipping values at these positions can be non-zero. The clipping function can be used to limit the upper bound of the filter values in the luminance block or chrominance block.

[0085] In some implementations, a specific ALF to be applied to a specific block of the luminance component can be based on the classification of the luminance block. For the block classification of the luminance component, a 4×4 block (or luminance block, luminance CB) can be classified or categorized into one of multiple (e.g., 25) categories corresponding to, for example, 25 different ALFs (e.g., 25 7×7 ALFs with different filter coefficients). The classification index C can be derived using Equation (1) based on the quantization values of the directionality parameter D and the activity value A to derive the classification index C.

[0086] To calculate the directionality parameter D and the quantization value the gradient g in the vertical direction can be calculated using the following one-dimensional Laplacian operator respectively v , the gradient g in the horizontal direction h and the gradients g in the two diagonal directions (e.g., d1 and d2) d1 and g d2 . Among them, the indices i and j refer to the coordinates of the upper-left sample within a 4×4 block, and R(k, l) indicates the reconstructed sample at the coordinates (k, l). Directions (such as d1 and d2) refer to two diagonal directions.

[0087] To reduce the complexity of the above block classification, downsampled one-dimensional Laplacian operator calculation can be applied. Figures 8A to 8D Examples of downsampling positions for calculating gradients are respectively shown, where the gradients are the gradients g Figure 8A in the vertical direction ( v ), the gradients g Figure 8B in the horizontal direction ( h ), and the gradients g Figure 8C and g Figure 8D in the two diagonal directions d1 ( d1 ) and d2 ( d2 ). In Figure 8A , the label "V" shows the downsampling position for calculating the vertical gradient g v . In Figure 8B , the label "H" shows the downsampling position for calculating the horizontal gradient g h . In Figure 8C , the label "D1" shows the downsampling position for calculating the d1 diagonal gradient g d1 . In Figure 8D , the label "D2" shows the downsampling position for calculating the d2 diagonal gradient g d2 . Figure 8A And Figure 8B show that the same downsampling position can be used for gradient calculations in different directions. In some other implementations, different downsampling schemes can be used for all directions. In still some other implementations, different downsampling schemes can be used for different directions.

[0088] The maximum value h of the horizontal gradient g v and the vertical gradient g and the minimum value can be set as: The maximum value d1 of the two diagonal gradients g d2 and g and the minimum value can be set as: The directional parameter D can be derived based on the above values and the following two thresholds t1 and t2. Step 1: If (1) and (2) If it is true, then D is set to 0. Step 2: If Then proceed to Step 3; otherwise proceed to Step 4. Step 3: If Then D is set to 2; otherwise D is set to 1. Step 4: If Then D is set to 4; otherwise D is set to 3.

[0089] In other words, the directivity parameter D is represented by several discrete levels and is determined based on the gradient value distribution between the horizontal and vertical directions and between the two diagonal directions of the luminance block, as Figure 8E shown.

[0090] The activity value A can be calculated as: Therefore, the activity value A represents a comprehensive measure of the horizontal one-dimensional Laplacian operator and the vertical one-dimensional Laplacian operator. The activity value A of the luminance block can be further quantized to a range of, for example, 0 to 4 (including the end values), and the quantization value is represented as

[0091] For the luminance component, then the classification index C calculated as above can be used to select one of multiple rhombus-shaped AFL filter classes (e.g., 25 classes). In some implementations, for the chrominance component in the picture, block classification may not be applied, and thus a set of ALF coefficients can be applied to each chrominance component. In such implementations, although there may be multiple sets of ALF coefficients available for the chrominance component, the determination of the ALF coefficients may not depend on any classification of the chrominance block.

[0092] Geometric transformations can be applied to the filter coefficients and the corresponding filter clipping values (also known as clipping values). Before filtering a block (e.g., a 4×4 luminance block), for example, based on the gradient values calculated for the block (e.g., g v , g h , g d1 and / or g d2 ), geometric transformations such as rotation or diagonal flipping and vertical flipping can be applied to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l). Applying geometric transformations to the filter coefficients f(k,l) and the corresponding filter clipping values c(k,l) can be equivalent to applying geometric transformations to the samples in the region supported by the filter. Geometric transformations can make different blocks to which ALF is applied more similar by aligning their respective directivities.

[0093] Three geometric transformation options, including diagonal flipping, vertical flipping, and rotation, can be performed respectively as described in Equation (9) to Equation (11). f D (k, l) = f(l, k), c D (k, l) = c(l, k) Equation (9) f V (k, l) = f(k, K - l - 1), c V (k, l) = c(k, K - l - 1) Equation (10) f R (k, l) = f(K - l - 1, k), c R (k, l) = c(K - l - 1, k) Equation (11) Where K represents the size of the ALF or filter, and 0 ≤ k, l ≤ K - 1 are the coordinates of the coefficients. For example, the position (0, 0) is at the upper left corner of the filter f or the cropping value matrix (or cropping matrix) c, and the position (K - 1, K - 1) is at the lower right corner of the filter f or the cropping value matrix (or cropping matrix) c. These transformations can be applied to the filter coefficients f(k, l) and the cropping values c(k, l) according to the gradient values calculated for this block. Table 1 summarizes an example of the relationship between the transformations and the four gradients. Table 1 Mapping of Gradients Calculated for Blocks and Transformations Gradient value Transformation <![CDATA[g d2 <g d1 and g h <g v > No transformation <![CDATA[g d2 <g d1 and g v <g h > Diagonal flip <![CDATA[g d1 <g d2 and g h <g v > Vertical flip <![CDATA[g d1 <g d2 and g v <g h > Rotation

[0094] In some embodiments, the ALF filter parameters derived by the encoder can be written into the Adaptive Parameter Set (APS) for the picture. In the APS, one or more sets of luminance filter coefficients and cropping value indices (e.g., up to 25 sets) can be written. These sets can be indexed in the APS. In an example, one set in one or more sets can include luminance filter coefficients and one or more cropping value indices. One or more sets of chrominance filter coefficients and cropping value indices (e.g., up to 8 sets) can be derived by the encoder and signaled. To reduce signaling overhead, the filter coefficients for different classifications (e.g., with different classification indices) of the luminance component can be merged. In the slice header, the index of the APS for the current slice can be written. In another example, the signaling transmission of the ALF can be CTU - based.

[0095] In an embodiment, a cropping value index (also referred to as a cropping index) can be decoded from the APS. The cropping value index can be used, for example, to determine a corresponding cropping value based on the relationship between the cropping value index and the corresponding cropping value. This relationship can be predefined and stored in the decoder. In an example, this relationship is described by one or more tables, such as a table of cropping value indices and corresponding cropping values for the luminance component (e.g., for luminance CB), and a table of cropping value indices and corresponding cropping values for the chrominance component (e.g., for chrominance CB). The cropping value can depend on the bit depth B. The bit depth B can refer to the internal bit depth, the bit depth of the reconstructed samples in the CB to be filtered, etc. In some examples, Equation (12) can be used to obtain the table of cropping values (e.g., for luminance and / or for chrominance). AlfClip = {round(2 B-α*n ), n ∈ [0..N - 1]} Equation (12) where AlfClip is the cropping value, B is the bit depth (e.g., bitDepth), N (e.g., N = 4) is the number of allowed cropping values, α is a predefined constant value. In an example, α is equal to 2.35. n is the cropping value index (also referred to as the cropping index or clipIdx). Table 2 shows an example of the table obtained using Equation (12), where N = 4. In Table 2, the cropping index n can be 0, 1, 2, and 3 (up to N - 1). Table 2 can be used for luminance blocks or chrominance blocks. Table 2 AlfClip can depend on the bit depth B and clipIdx

[0096] In the slice header for the current slice, one or more APS indices (e.g., up to 7 APS indices) can be written to specify the luminance filter bank available for the current slice. The filtering process can be controlled at one or more appropriate levels, such as the picture level, the slice level, the CTB level, etc. In an example embodiment, the filtering process can be further controlled at the CTB level. A flag can be written to indicate whether ALF is applied to the luminance CTB. The luminance CTB can select a set of filters from multiple sets of fixed filters (e.g., 16 sets of fixed filters) and the filter bank written in the APS (e.g., up to 25 filters derived by the encoder as described above and also referred to as the written filter bank). A group index of the filter can be written for the luminance CTB to indicate the filter bank to be applied (e.g., one of the multiple sets of fixed filters and the written filter bank). The multiple sets of fixed filters can be predefined and hard - coded in the encoder and decoder, and can be referred to as predefined filter banks. Therefore, it is not necessary to write the predefined filter coefficients.

[0097] For the chrominance component, the APS index can be written into the slice header to indicate the chrominance filter bank to be used for the current slice. At the CTB level, if there are more than one set of chrominance filters in the APS, a filter bank index can be written for each chrominance CTB.

[0098] The filter coefficients can be quantized with a norm equal to 128. To reduce the multiplication complexity, bitstream consistency can be applied such that the coefficient values at non-central positions can be in the range of -27 to 27-1 (including the end values). In the example, the central position coefficient is not written into the bitstream, and the central position coefficient can be considered equal to 128.

[0099] In some embodiments, the syntax and semantics of the crop index and crop value are defined as follows: alf_luma_clip_idx[sfIdx][j] can be used to specify the crop index of the crop value to be used before multiplying the j-th coefficient of the written luminance filter indicated by sfIdx. The requirements for bitstream consistency can include that the value of alf_luma_clip_idx[sfIdx][j] should be in the range of, for example, 0 to 3 (including the end values), where sfIdx = 0 to alf_luma_num_filters_signalled_minus1 and j = 0 to 11.

[0100] The luminance filter crop value AlfClipL[adaptation_parameter_set_id] with elements AlfClipL[adaptation_parameter_set_id][filtIdx][j] (where filtIdx = 0 to NumAlfFilters-1 and j = 0 to 11) can be derived as specified in Table 2, depending on the bitDepth set equal to BitDepthY and the clipIdx set equal to alf_luma_clip_idx[alf_luma_coeff_delta_idx[filtIdx]][j].

[0101] Alf_chroma_clip_idx[altIdx][j] can be used to specify the crop index of the crop value to be used before multiplying the j-th coefficient of the alternative chrominance filter with index altIdx. The requirements for bitstream consistency can include that the value of alf_chroma_clip_idx[altIdx][j] should be in the range of 0 to 3 (including the end values), where altIdx = 0 to alf_chroma_num_alt_filters_minus1 and j = 0 to 5.

[0102] The chroma filter clip value AlfClipC[adaptation_parameter_set_id][altIdx] with element AlfClipC[adaptation_parameter_set_id][altIdx][j] can be derived as specified in Table 2, where altIdx = 0 to alf_chroma_num_alt_filters_minus1 and j = 0 to 5, depending on the bitDepth set equal to BitDepthC and the clipIdx set equal to alf_chroma_clip_idx[altIdx][j].

[0103] In an embodiment, the filtering process can be described as follows. On the decoder side, when ALF is enabled for a CTB, the samples R(i,j) within the CU (or CB) of the CTB can be filtered to produce the filtered sample value R'(i,j) using Equation (13) as follows. In the example, each sample in the CU is filtered. R′(i,j) = R(i,j) + ((∑ k≠0 ∑ l≠0 f(k,l) × K(R(i + k,j + l) - R(i,j), c(k,l)) + 64) >> 7) Equation (13) where f(k,l) represents the decoded filter coefficient, K(x,y) is the clipping function, and c(k,l) represents the decoded clipping parameter (or clip value). The variables k and l can vary between -L / 2 and L / 2, where L represents the filter length (e.g., for the luminance component, Figure 7 for the example diamond filter 711 in, L = 7; for the chrominance component, Figure 7 for the example diamond filter 710 in, L = 5). The clipping function K(x,y) = min(y, max(-y,x)) corresponds to the clipping function Clip3(-y,y,x). By introducing the clipping function K(x,y), the loop filtering method (such as ALF) becomes a non-linear process and can be referred to as non-linear ALF.

[0104] The selected clip value can be encoded into the "alf_data" syntax element as follows: A suitable coding scheme (e.g., the Golomb coding scheme) can be used to encode the clip index corresponding to the selected clip value as shown in Table 2, for example. This coding scheme can be the same as the coding scheme used to encode the filter bank index.

[0105] In an embodiment, a virtual boundary filtering process can be used to reduce the line buffer requirements of the ALF. Accordingly, modified block classification and filtering can be applied to samples near the CTU boundary (e.g., a horizontal CTU boundary). The virtual boundary (930) can be defined as a line by translating the horizontal CTU boundary (920) by "N samples " samples, as Figure 9A shown, where N samples can be a positive integer. In an example, for the luminance component, N samples equals 4; for the chrominance component, N samples equals 2.

[0106] Referring to Figure 9A , for the luminance component, modified block classification can be applied. In an example, for the one-dimensional Laplacian gradient calculation of a 4×4 block (910) above the virtual boundary (930), only the samples above the virtual boundary (930) are used. Similarly, referring to Figure 9B , for the one-dimensional Laplacian gradient calculation of a 4×4 block (911) below the virtual boundary (931) translated from the CTU boundary (921), only the samples below the virtual boundary (931) are used. By considering the reduced number of samples used in the one-dimensional Laplacian gradient calculation, the quantization of the activity value A can be scaled accordingly.

[0107] For the filtering process, a symmetric padding operation at the virtual boundary can be used for both the luminance component and the chrominance component. Figures 10A to 10F An example of such modified ALF filtering for the luminance component at the virtual boundary is shown. When the sample to be filtered is below the virtual boundary, the adjacent samples above the virtual boundary can be filled. When the sample to be filtered is above the virtual boundary, the adjacent samples below the virtual boundary can be filled. Referring to Figure 10A , the adjacent sample C0 can be filled with the sample C2 below the virtual boundary (1010). Referring to Figure 10B , the adjacent sample C0 can be filled with the sample C2 above the virtual boundary (1020). Referring to Figure 10C , the adjacent samples C1 to C3 can be filled with the samples C5 to C7 below the virtual boundary (1030) respectively; the sample C0 can be filled with the sample C6. Referring to Figure 10D , the adjacent samples C1 to C3 can be filled with the samples C5 to C7 above the virtual boundary (1040) respectively; the sample C0 can be filled with the sample C6. Referring to Figure 10E , the adjacent samples C4 to C8 can be filled with the samples C10, C11, C12, C11, and C10 below the virtual boundary (1050) respectively; the samples C1 to C3 can be filled with the samples C11, C12, C11; the sample C0 can be filled with the sample C12. Referring to Figure 10F, the adjacent samples C4 to C8 can be filled with the samples C10, C11, C12, C11, and C10 located above the virtual boundary (1060) respectively; the samples C1 to C3 can be filled with the samples C11, C12, C11; and the sample C0 can be filled with the sample C12.

[0108] In some examples, the above description can be appropriately adjusted when the sample and the adjacent sample are located on the left (or right) and right (or left) sides of the virtual boundary.

[0109] Picture quad-tree segmentation with largest coding unit (LCU) alignment can be used. To improve the coding efficiency, an adaptive loop filter based on coded unit synchronous picture quad-tree can be used in video coding and decoding. In the example, the luminance picture can be segmented into multiple multi-level quad-tree partitions, and the boundary of each partition is aligned with the boundary of the largest coding unit (LCU). Each partition can have a filtering process and thus can be called a filter unit or filtering unit (FU).

[0110] The process of two-pass encoding in Example 2 is described as follows. During the first pass of encoding, the quad-tree segmentation mode and the best filter (or optimal filter) of each FU can be determined. During this determination process, the filtering distortion can be estimated by fast filtering distortion estimation (FFDE). According to the determined quad-tree segmentation mode and the selected filters of the FUs (e.g., all FUs), the reconstructed picture can be filtered. During the second pass of encoding, CU synchronous ALF on / off control can be performed. According to the ALF on / off result, the picture after the first pass of filtering can be partially restored through the reconstructed picture.

[0111] A top-down segmentation strategy can be adopted to divide the picture into multi-level quad-tree partitions by using the rate-distortion criterion. Each partition can be called an FU. The segmentation process can align the quad-tree partitions with the LCU boundary, as Figure 11 shown. Figure 11 An example of LCU-aligned picture quad-tree segmentation according to an embodiment of the present disclosure is shown. In the example, the coding order of the FUs follows the z-scan order. For example, referring to Figure 11 , the picture is segmented into ten FUs (e.g., FU0 to FU9, with a segmentation depth of 2, where FU0, FU1, and FU9 are the first-level FUs, and FU s, FU7 and FU8 are the second - depth - level FUs, and FU3 to FU6 are the third - depth - level FUs), and the coding order is from FU0 to FU9. For example, FU0, FU1, FU2, FU3, FU4, FU5, FU6, FU7, FU8, and FU9.

[0112] To indicate the picture quadtree segmentation mode, a segmentation flag can be encoded ("1" indicates quadtree segmentation, "0" indicates no quadtree segmentation) and the segmentation flag is transmitted in a z - scan order. Figure 12 shows the quadtree segmentation mode corresponding to Figure 11 in accordance with an embodiment of the present disclosure. As Figure 12 shown in the example in, the quadtree segmentation flag is encoded in a z - scan order.

[0113] The filter for each FU can be selected from two sets of filters based on the rate - distortion criterion. The first set can have newly derived 1 / 2 - symmetric square and diamond filters for the current FU. The second set can come from a delay - filter buffer. The delay - filter buffer can store the filters previously derived for FUs in the previous picture. The filter with the minimum rate - distortion cost in the two sets of filters can be selected for the current FU. Similarly, if the current FU is not the smallest FU and can be further divided into four sub - FUs, the rate - distortion costs of these four sub - FUs can be calculated. By recursively comparing the rate - distortion costs of the segmented case and the non - segmented case, the picture quadtree segmentation mode (in other words, whether the quadtree segmentation of the current FU should stop) can be determined.

[0114] In some examples, the maximum quadtree segmentation level or depth can be limited to a predefined number. For example, the maximum quadtree segmentation level or depth can be 2, and thus the maximum number of FUs can be 16 (or the fourth power of the maximum depth number). During the quadtree segmentation determination, the relevant values of the Wiener coefficients for the 16 FUs at the bottom quadtree level (the smallest FUs) can be reused. The remaining FUs can derive the Wiener filters of the remaining FUs based on the correlation of the 16 FUs at the bottom quadtree level. Thus, in the example, only one frame - buffer access is required to derive the filter coefficients of all FUs.

[0115] After determining the quadtree splitting pattern, to further reduce the filtering distortion, CU synchronous ALF on / off control can be performed. By comparing the filtered distortion and the non-filtered distortion, a leaf CU can explicitly turn on / off the ALF in the corresponding local region. By re-designing the filter coefficients according to the ALF on / off result, the coding efficiency can be further improved. In an example, this re-design process requires additional frame buffer accesses. Thus, in some examples, such as the coding unit synchronous picture quadtree-based adaptive loop filter (CS-PQALF) encoder design, no re-design process is required after the CU synchronous ALF on / off determination to minimize the number of frame buffer accesses.

[0116] The cross-component filtering process can apply a cross-component filter, such as a cross-component adaptive loop filter (CC-ALF). The cross-component filter can use the luminance sample values of the luminance component (e.g., luminance CB) to refine the chrominance component (e.g., chrominance CB corresponding to the luminance CB). In an example, the luminance CB and the chrominance CB are included in one CU.

[0117] Figure 13 A cross-component filter (e.g., CC-ALF) for generating a chrominance component according to an example embodiment of the present disclosure is shown. For example, Figure 13 A filtering process for a first chrominance component (e.g., first chrominance CB), a second chrominance component (e.g., second chrominance CB), and a luminance component (e.g., luminance CB) is shown. The luminance component can be filtered by a sample adaptive offset (SAO) filter (1310) to generate a SAO-filtered luminance component (1341). This SAO-filtered luminance component (1341) can be further filtered by an ALF luminance filter (1316) to become a filtered luminance CB (1361) (e.g., “Y”).

[0118] The first chrominance component can be filtered by the SAO filter (1312) and the ALF chrominance filter (1318) to generate a first intermediate component (1352). Additionally, the SAO-filtered luminance component (1341) can be filtered for the first chrominance component by a cross-component filter (such as CC-ALF) (1321) to generate a second intermediate component (1342). Subsequently, a filtered first chrominance component (1362) (e.g., "Cb") can be generated based on at least one of the second intermediate component (1342) and the first intermediate component (1352). In an example, the filtered first chrominance component (1362) (e.g., "Cb") can be generated by combining the second intermediate component (1342) and the first intermediate component (1352) with an adder (1322). The example cross-component adaptive loop filtering process for the first chrominance component can thus include steps performed by the CC-ALF (1321) and steps performed by, for example, the adder (1322).

[0119] The above description can be applied to the second chrominance component. The second chrominance component can be filtered by the SAO filter (1314) and the ALF chrominance filter (1318) to generate a third intermediate component (1353). Additionally, the SAO-filtered luminance component (1341) can be filtered for the second chrominance component by a cross-component filter (such as CC-ALF) (1331) to generate a fourth intermediate component (1343). Subsequently, a filtered second chrominance component (1363) (e.g., "Cr") can be generated based on at least one of the fourth intermediate component (1343) and the third intermediate component (1353). In an example, the filtered second chrominance component (1363) (e.g., "Cr") can be generated by combining the fourth intermediate component (1343) and the third intermediate component (1353) with an adder (1332). In an example, the cross-component adaptive loop filtering process for the second chrominance component can thus include steps performed by the CC-ALF (1331) and steps performed by, for example, the adder (1332).

[0120] The cross-component filters (such as CC-ALF (1321), CC-ALF (1331)) can operate by applying a linear filter with any suitable filter shape to the luminance component (or luminance channel) to refine each chrominance component (e.g., the first chrominance component, the second chrominance component). The CC-ALF exploits the correlation between color components to reduce the coding distortion of one color component based on samples of another color component.

[0121] Figure 14Shows an example of a CC-ALF filter (1400) according to an embodiment of the present disclosure. The filter (1400) may include non-zero filter coefficients and zero filter coefficients. The filter (1400) has a diamond shape (1420) formed by filter coefficients (1410) (indicated by black-filled circles). In the example, the non-zero filter coefficients in the filter (1400) are included in the filter coefficients (1410), and the filter coefficients not included in the filter coefficients (1410) are zero. Thus, the non-zero filter coefficients in the filter (1400) are included in the diamond shape (1420), and the filter coefficients not included in the diamond shape (1420) are zero. In the example, the number of filter coefficients of the filter (1400) is equal to the number of filter coefficients (1410), which is Figure 14 18 in the example shown.

[0122] CC-ALF may include any suitable filter coefficients (also referred to as CC-ALF filter coefficients). Return reference Figure 13 , CC-ALF (1321) and CC-ALF (1331) may have the same filter shape (e.g., Figure 14 the diamond shape (1420) shown in ) and the same number of filter coefficients. In the example, the values of the filter coefficients in CC-ALF (1321) are different from the values of the filter coefficients in CC-ALF (1331).

[0123] Generally, the filter coefficients in CC-ALF (e.g., non-zero filter coefficients derived by an encoder) may be transmitted, for example, in APS. In the example, the filter coefficients may be scaled by a factor (e.g., 2 10 ) and may be rounded for fixed-point representation. The application of CC-ALF may be controlled on variable block sizes and is written by a context coding flag (e.g., CC-ALF enable flag) received for each block of samples. The context coding flag (e.g., the CC-ALF enable flag) may be written at any suitable level (e.g., block level). For each chrominance component, the block size and the CC-ALF enable flag may be received at the slice level. In some examples, block sizes of 16×16, 32×32, and 64×64 (in terms of chrominance samples) may be supported.

[0124] In the example, the syntax changes of CC-ALF are described in Table 3 below. Table 3 Syntax Changes of CC-ALF

[0125] The semantics of the above example CC-ALF-related syntax may be described as follows:

[0126] alf_ctb_cross_component_cb_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] being equal to 0 can indicate that the cross-component Cb filter is not applied to the block of Cb color component samples at the luminance position (xCtb, yCtb).

[0127] alf_ctb_cross_component_cb_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] not being equal to 0 can indicate that the cross-component Cb filter represented by alf_ctb_cross_component_cb_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] is applied to the block of Cb color component samples at the luminance position (xCtb, yCtb).

[0128] alf_ctb_cross_component_cr_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] being equal to 0 can indicate that the cross-component Cr filter is not applied to the block of Cr color component samples at the luminance position (xCtb, yCtb).

[0129] alf_ctb_cross_component_cr_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] not being equal to 0 can indicate that the cross-component Cr filter represented by alf_ctb_cross_component_cr_idc[xCtb >> CtbLog2SizeY][yCtb >> CtbLog2SizeY] is applied to the block of Cr color component samples at the luminance position (xCtb, yCtb).

[0130] Examples of chroma sampling formats are described below. Generally, a luma block may correspond to one or more chroma blocks, e.g., two chroma blocks. The number of samples in each chroma block may be less than the number of samples in the luma block. The chroma subsampling format (also referred to as chroma subsampling format, e.g., specified by chroma_format_idc) may indicate the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) between each chroma block and the corresponding luma block. The chroma subsampling scheme may be specified in a 4:x:y format for a nominal 4 (horizontal) by 4 (vertical) block, where x is the horizontal chroma subsampling factor (the number of chroma samples retained in the first row of the block), and y is the number of chroma samples retained in the second row of the block. In an example, the chroma subsampling format may be 4:2:0, indicating that both the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 2, as Figure 15 shown in Figure 15 A to

[0131] Figure 15 B. In another example, the chroma subsampling format may be 4:2:2, indicating that the chroma horizontal subsampling factor (e.g., SubWidthC) is 2 and the chroma vertical subsampling factor (e.g., SubHeightC) is 1. In yet another example, the chroma subsampling format may be 4:4:4, indicating that both the chroma horizontal subsampling factor (e.g., SubWidthC) and the chroma vertical subsampling factor (e.g., SubHeightC) are 1. Thus, the chroma sample format or type (also referred to as chroma sample position) may represent the relative position of chroma samples in a chroma block with respect to at least one corresponding luma sample in a luma block. Figure 15 A to Figure 15 B illustrate exemplary positions of chroma samples with respect to luma samples according to embodiments of the present disclosure. Referring to Figure 15The luminance samples (1501) shown in A may represent a part of a picture. In the example, a luminance block (such as a luminance CB) includes luminance samples (1501). The luminance block may correspond to two chrominance blocks with a chrominance downsampling format of 4:2:0. In the example, each chrominance block includes chrominance samples (1503). Each chrominance sample (such as chrominance sample (1503(1)) corresponds to four luminance samples (such as luminance samples (1501(1))-(1501(4))). In the example, the four luminance samples are the upper left sample (1501(1)), the upper right sample (1501(2)), the lower left sample (1501(3)), and the lower right sample (1501(4)). The chrominance sample (such as (1503(1))) may be located at a left center position between the upper left sample (1501(1)) and the lower left sample (1501(3)), and the chrominance sample type of the chrominance block having the chrominance sample (1503) may be referred to as chrominance sample type 0. Chrominance sample type 0 indicates a relative position 0 corresponding to the left center position in the middle of the upper left sample (1501(1)) and the lower left sample (1501(3)). The four luminance samples (such as (1501(1)) to (1501(4))) may be referred to as the adjacent luminance samples of the chrominance sample (1503)(1).

[0132] In the example, each chrominance block may include chrominance samples (1504). The above reference description of the chrominance samples (1503) may apply to the chrominance samples (1504), and thus for the sake of brevity, the detailed description may be omitted. Each chrominance sample (1504) may be located at the center position of four corresponding luminance samples, and the chrominance sample type of the chrominance block having the chrominance sample (1504) may be referred to as chrominance sample type 1. Chrominance sample type 1 indicates a relative position 1 corresponding to the center position of the four luminance samples (such as (1501(1)) to (1501(4))). For example, one of the chrominance samples (1504) may be located in the central part of the luminance samples (1501(1)) to (1501(4)).

[0133] In the example, each chrominance block includes chrominance samples (1505). Each chrominance sample (1505) may be located at an upper left position co-located with the upper left sample of four corresponding luminance samples (1501), and the chrominance sample type of the chrominance block having the chrominance sample (1505) may be referred to as chrominance sample type 2. Thus, each chrominance sample (1505) is co-located with the upper left sample of the four luminance samples (1501) corresponding to the respective chrominance sample. Chrominance sample type 2 indicates a relative position 2 corresponding to the upper left position of the four luminance samples (1501). For example, one of the chrominance samples (1505) may be located in the upper left part of the luminance samples (1501(1)) to (1501(4)).

[0134] In an example, each chrominance block includes chrominance samples (1506). Each chrominance sample (1506) may be located at a top center position between a corresponding top left sample and a corresponding top right sample, and the chrominance sample type of the chrominance block having the chrominance sample (1506) may be referred to as chrominance sample type 3. Chrominance sample type 3 indicates a relative position 3 corresponding to the top center position between the top left sample and the top right sample. For example, one of the chrominance samples (1506) may be located in the top center portion of the luminance samples (1501(1)) to (1501(4)).

[0135] In an example, each chrominance block includes chrominance samples (1507). Each chrominance sample (1507) may be located at a bottom left position co-located with the bottom left sample of four corresponding luminance samples (1501), and the chrominance sample type of the chrominance block having the chrominance sample (1507) may be referred to as chrominance sample type 4. Thus, each chrominance sample (1507) is co-located with the bottom left sample of the four luminance samples (1501) corresponding to the respective chrominance sample. Chrominance sample type 4 indicates a relative position 4 corresponding to the bottom left position of the four luminance samples (1501). For example, one of the chrominance samples (1507) may be located in the bottom left portion of the luminance samples (1501(1)) to (1501(4)).

[0136] In an example, each chrominance block includes chrominance samples (1508). Each chrominance sample (1508) is located at a bottom center position between the bottom left sample and the bottom right sample, and the chrominance sample type of the chrominance block having the chrominance sample (1508) may be referred to as chrominance sample type 5. Chrominance sample type 5 indicates a relative position 5 corresponding to the bottom center position between the bottom left sample and the bottom right sample of the four luminance samples (1501). For example, one of the chrominance samples (1508) may be located between the bottom left sample and the bottom right sample of the luminance samples (1501(1)) to (1501)(4)).

[0137] Generally, any suitable chrominance sample type can be used for a chrominance downsampling format. Chrominance sample types 0 to chrominance sample type 5 provide exemplary chrominance sample types described in a chrominance downsampling format 4:2:0. Additional chrominance sample types can be used for the chrominance downsampling format 4:2:0. In addition, other chrominance sample types and / or variants of chrominance sample types 0 to chrominance sample type 5 can be used for other chrominance downsampling formats, such as 4:2:2, 4:4:4, etc. In an example, a chrominance sample type combining chrominance samples (1505) and chrominance samples (1507) can be used for a chrominance downsampling format 4:2:2.

[0138] In another example, a luminance block is considered to have alternating rows, such as row (1511) and row (1512), which respectively include two top samples (such as (1501(1)) and (1501)(2))) of four luminance samples (such as (1501(1)) to (1501)(4))), and two bottom samples (such as (1501(3)) and (1501)(4))) of four luminance samples (such as (1501(1) to (1501(4))). Thus, row (1511), row (1513), row (1515) and row (1517) can be referred to as the current row (also known as the top field); and row (1512), row (1514), row (1516) and row (1518) can be referred to as the next row (also known as the bottom field). Four luminance samples (such as (1501(1) to (1501(4))) are located in the current row (such as (1511)) and the next row (such as (1512)). The above relative chrominance positions 2 and relative chrominance position 3 are located in the current row, the above relative chrominance positions 0 and relative chrominance position 1 are located between each current row and the corresponding next row, and the above relative chrominance positions 4 and relative chrominance position 5 are located in the next row.

[0139] Chrominance samples (1503), chrominance samples (1504), chrominance samples (1505), chrominance samples (1506), chrominance samples (1507) or chrominance samples (1508) are located in rows (1551) to (1554) in each chrominance block. The specific positions of rows (1551) to (1554) may depend on the chrominance sample type of the chrominance samples. For example, for chrominance sample (1503) with the corresponding chrominance sample type 0 and chrominance sample (1504) with the corresponding chrominance sample type 1, row (1551) is located between row (1511) and row (1512). For chrominance sample (1505) with the corresponding chrominance sample type 2 and chrominance sample (1506) with the corresponding chrominance sample type 3, row (1551) is co-located with the current row (1511). For chrominance sample (1507) with the corresponding chrominance sample type 4 and chrominance sample (1508) with the corresponding chrominance sample type 5, row (1551) is co-located with the next row (1512). The above description can be appropriately applied to rows (1552) to (1554), and for the sake of brevity, the detailed description is omitted.

[0140] Any suitable scanning method can be used to display, store, and / or transmit the luminance blocks and corresponding chrominance blocks described above in Figure 15 A. In some example implementations, progressive scanning can be used.

[0141] Alternatively, interlaced scanning can be used, as Figure 15As shown in B. As described above, the chroma downsampling format can be 4:2:0 (e.g., chroma_format_idc equals 1). In the example, the variable chroma position type (e.g., ChromaLocType) can indicate the current line (e.g., ChromaLocType is chroma_sample_loc_type_top_field) or the next line (e.g., ChromaLocType is chroma_sample_loc_type_bottom_field). The current lines (1511), (1513), (1515), and (1517) and the next lines (1512), (1514), (1516), and (1518) can be scanned separately. For example, the current lines (1511), (1513), (1515), and (1517) can be scanned first, and then the next lines (1512), (1514), (1516), and (1518). The current line can include luminance samples (1501), and the next line can include luminance samples (1502).

[0142] Similarly, the corresponding chroma blocks can be scanned in an interlaced manner. The lines (1551) and (1553) including chroma samples (1503), (1504), (1505), (1506), (1507), or (1508) without padding can be referred to as the current line (or current chroma line), and the lines (1552) and (1554) including chroma samples (1503), (1504), (1505), (1506), (1507), or (1508) with gray padding can be referred to as the next line (or next chroma line). In the example, during interlaced scanning, the lines (1551) and (1553) can be scanned first, and then the lines (1552) and (1554).

[0143] In addition to the above ALF, a Constrained Direction Enhancement Filter (CDEF) can also be used for loop filtering in video coding and decoding. The in-loop CDEF can be used to filter out coding artifacts (such as quantization ringing artifacts) while preserving the details of the image. In some coding techniques, a Sample Adaptive Offset (SAO) algorithm can be employed to achieve a similar goal by defining signal offsets for different classes of pixels. Different from SAO, CDEF is a non-linear spatial filter. In some examples, the design of the CDEF filter is constrained to be easily vectorized (e.g., it can be implemented using single instruction multiple data (SIMD) operations). This is not the case for other non-linear filters such as median filters and bilateral filters.

[0144] The CDEF design stems from the following observations. In some cases, the amount of ringing artifacts in a coded image can be approximately proportional to the quantization step size. The smallest details retained in the quantized image are also proportional to the quantization step size. Thus, maintaining image details would require a smaller quantization step size, which would produce a higher amount of undesired quantization ringing artifacts. Fortunately, for a given quantization step size, the amplitude of the ringing artifacts can be smaller than the amplitude of the details, providing an opportunity to design the CDEF to achieve a balance by filtering out the ringing artifacts while retaining sufficient details.

[0145] The CDEF can first identify the direction of each block. Then, the CDEF can perform adaptive filtering along the identified direction and, to a lesser extent, along the direction rotated 45° from the identified direction. The filter strength can be explicitly written, allowing for a high degree of control over the degree of blurring of the details. An efficient encoder search can be designed for the filter strength. The CDEF can be based on two in-loop filters, and the combined filter can be used for video coding and decoding. In some example implementations, the CDEF filter can follow the deblocking filter in in-loop filtering.

[0146] The direction search can operate on the reconstructed pixels (or samples) after, for example, the deblocking filter, as Figure 16As shown. Since the reconstructed pixels are available to the decoder, these directions may not require signaling. The direction search can operate on blocks of a suitable size (e.g., an 8×8 block), which is small enough to adequately handle non-linear edges (such that the edges appear straight enough in the filtering block) and large enough to reliably estimate the direction when applied to the quantized image. Having a constant direction over an 8×8 region makes vectorization of the filter easier. For each block, the direction that best matches the pattern in the block can be determined by minimizing a difference metric between the quantized block and each fully oriented block, such as the sum of squared differences (SSD), RMS error, etc. In an example, a fully oriented block (e.g., Figure 16 one of (1620)) refers to a block in which all pixels in a row along one direction have the same value. Figure 16 FIG. shows an example of a direction search for an 8×8 block (1610) according to an example embodiment of the present disclosure. In Figure 16 the example shown, the 45-degree direction (1623) in a set of directions (1620) is selected because the 45-degree direction (1623) minimizes the error (1640). For example, the error for the 45-degree direction is 12, and this error is the minimum error within the error range of 12 to 87 indicated by the row (1640).

[0147] An example non-linear low-pass direction filter is described in further detail below. Identifying the direction can help align the filter taps along the identified direction to reduce ringing artifacts while preserving directional edges or patterns. However, in some examples, directional filtering alone is not sufficient to reduce ringing artifacts. Additional filter taps are desired on pixels not along the main direction (e.g., the identified direction). To reduce the risk of blurring, the additional filter taps can be processed more conservatively. Thus, the CDEF can define a main tap and a secondary tap. In some example implementations, the complete two-dimensional (2-D) CDEF filter can be represented as:

[0148] In Equation (14), D represents the damping parameter, S (p) and S (s) represent the strength of the main tap and the strength of the secondary tap, respectively. The function round(·) can round away from zero, and Let \(w\) denote the filter weights, and \(f(d, S, D)\) denote a constraint function that operates on the difference \(d\) (e.g., \(d = x(m, n)-x(i, j)\)) between a filtered pixel (e.g., \(x(i, j)\)) and each neighboring pixel (e.g., \(x(m, n)\)). When the difference is small, \(f(d, S, D)\) can be equal to the difference \(d\) (e.g., \(f(d, S, D)=d\)), and thus the filter can behave as a linear filter. When the difference is large, \(f(d, S, D)\) can be equal to 0 (e.g., \(f(d, S, D)=0\)), which effectively ignores the filter tap.

[0149] As another in-loop processing component, a set of in-loop recovery schemes can be used for post-deblocking processing in video coding and decoding to generally denoise and enhance edge quality beyond the deblocking operation. The set of in-loop recovery schemes can be switched in tiles of appropriate size within a frame (or picture). Some examples of in-loop recovery schemes are described below based on a separable symmetric Wiener filter and a dual self-guided filter that combines subspace projection. Since content statistics can vary significantly within a frame, the filters can be integrated within a switchable framework where different filters can be triggered in different regions of the frame.

[0150] An example separable symmetric Wiener filter is described below. The Wiener filter can be used as one of the switchable filters. Each pixel (or sample) in a degraded frame can be reconstructed as a non-causal filtered version of the pixels within a \(w\times w\) window around that pixel, where \(w = 2r + 1\) and for integer \(r\), \(w\) is odd. The two-dimensional filter taps can be represented by a vector \(F\) in the form of a column vector quantization with \(w\times1\) elements, and direct linear minimum mean square error (LMMSE) optimization can result in filter parameters given by \(F = H^{-1}M\), where \(H\) equals \(E[XX^T]\), and \(H\) is the autocovariance of \(x\), where \(x\) is the column vector quantization version of \(w\) samples within a \(w\times w\) window around the pixel, and where \(M\) equals \(E[YX^T]\), representing the cross-correlation vector between \(x\) and the scalar source sample \(y\) to be estimated. The encoder can be configured to estimate \(H\) and \(M\) based on the deblocked frame and the implementation in the source, and send the resulting filter \(F\) to the decoder. However, in some example implementations, instead of transmitting \(w\) 2 ×1 elements of the column vector quantization of the filter taps, the filter can be -1 represented by a smaller set of parameters that can be T derived from the full set of filter taps. For example, the 2 filter can be represented by a set of coefficients T that can be used to reconstruct the filter taps. The encoder 2A significant bitrate penalty may occur during tapping. Additionally, non-separable filtering can make decoding extremely complex. Therefore, multiple additional constraints can be imposed on the properties of F. For example, F can be constrained to be separable, such that the filtering can be implemented as separable horizontal and vertical w-tap convolutions. In the example, each of the horizontal filter and the vertical filter is constrained to be symmetric. Additionally, in some example implementations, it can be assumed that the sum of the horizontal filter coefficients and the vertical filter coefficients is 1.

[0151] Dual self-guided filtering combined with subspace projection can also be used as one of the switchable filters for in-loop recovery and is described below. In some example implementations, guided filtering can be used for image filtering, where a local linear model is used to calculate the filtered output y based on the unfiltered samples x. The local linear model can be written as: y = Fx + G Equation (15) where F and G can be determined based on the statistics of the degraded image and the guidance image (also referred to as the guide image) in the neighborhood of the filtered pixels. If the guidance image is the same as the degraded image, the resulting self-guided filtering can have an edge-preserving smoothing effect. According to some aspects of the present disclosure, the specific form of self-guided filtering can depend on two parameters: radius r and noise parameter e, listed as follows:

[0152] 1. Obtain the mean μ and variance σ of the pixels in the (2r + 1)×(2r + 1) window around each pixel 2 . For example, obtain the mean μ and variance σ of the pixel 2 which can be effectively implemented by box filtering based on integral imaging.

[0153] 2. Based on Equation (16), calculate the parameters f and g for each pixel. f = σ 2 / (σ 2 + e); g = (1 - f)μ Equation (16)

[0154] 3. Calculate F and G for each pixel as the average of the values of the parameters f and g in the 3×3 window around the pixel for use.

[0155] Dual self-guided filtering can be controlled by the radius r and the noise parameter e, where a larger radius r may imply a higher spatial variance and a higher noise parameter e may imply a higher range variance.

[0156] Figure 17 Shows an example of subspace projection according to an example embodiment of the present disclosure. InFigure 17 In the example shown, subspace projection can use inexpensive recoveries X1 and X2 to produce a final recovery X that is closer to the source Y f . Even if the inexpensive recoveries X1 and X2 are not close to the source Y, as long as the inexpensive recoveries X1 and X2 are moving in the correct direction, appropriate multipliers {α, β} can bring the inexpensive recoveries X1 and X2 closer to the source Y. For example, the final recovery X can be obtained based on Equation (17) below f . X f = X + α(X1 - X) + β(X2 - X) Equation (17)

[0157] In addition to the above deblocking filter, ALF, CDEF, and loop recovery, a loop filtering method called cross-component sample offset (CCSO) filter or CCSO can also be implemented during the loop filtering process to reduce the distortion of the reconstructed samples (also called reconstructed samples). The CCSO filter can be placed anywhere in the loop filtering stage. Regarding the deblocking, CDEF, and LR filters, an example of the CCSO filter is shown in Figure 18 . During the CCSO filtering process, a non-linear mapping can be used to determine the output offset based on the processed input reconstructed samples of the first color component. During the filtering process of CCSO, the output offset can be added to the reconstructed samples of the second color component

[0158] The input reconstructed samples can come from the first color component located in the filter support region, as Figure 19 shown Figure 19 . Specifically, an example of the filter support region in the CCSO filter according to an embodiment of the present disclosure is shown. The filter support region can include four reconstructed samples: p0 and p1 Figure 19 . In the example of, the two input reconstructed samples are located on both sides of the central sample in the vertical direction. In the example, the central sample (represented by rl) in the first color component (e.g., the luminance component) is co-located with the sample to be filtered (represented by rc) in the second color component (e.g., the chrominance component). When processing the input reconstructed samples, the following steps can be applied

[0159] Step 1: Calculate the difference values (e.g., differences) between the four reconstructed samples p0 and p1 and the central sample rl, and the difference values are respectively represented as m0 and m1. For example, the difference value between p0 and rl is m0

[0160] Step 2: The difference values m0 and m1 can be further quantized into multiple (e.g., 2) discrete values. For example, for m0 and m1, the quantization values can be respectively represented as d0 and d1. In the example, based on the following quantization process, the quantization value of each of d0 and d1 can be -1, 0, or 1 di = -1 if mi < -N, Equation (18) di = 0 if -N <= mi <= N, Equation (19) di = 1 if mi > N, Equation (20) where N is the quantization step, and example values of N are 4, 8, 12, 16, etc., and di and mi refer to the quantization value and the difference value respectively, where i is 0, 1, 2, or 3.

[0161] The quantization values d0 to d3 can be used to identify combinations of non - linear mappings. In Figure 19 the example shown, the CCSO filter has two filter inputs d0 and d1, and each filter input can take one of three quantization values (e.g., -1, 0, and 1), so the total number of combinations is 9 (e.g., 3 2 , that is, the number of quantization values to the power of the number of differences). Table 4 below is an example look - up table (LUT) showing the 9 combinations and the offset values.. Table 4 Example LUT used in CCSO Combination index d0 d1 Offset 0 -1 -1 s0 1 -1 0 s1 2 -1 1 s2 3 0 -1 s3 4 0 0 s4 5 0 1 s5 6 1 -1 s6 7 1 0 s7 8 1 1 s8

[0162] The last column can represent the output offset value for each combination, which can be looked up according to the difference. The output offset value can be an integer, such as 0, 1, -1, 3, -3, 5, -5, -7, etc. The first column represents the index of these combinations assigned to the quantized d0 and d1. The middle column represents all possible combinations of the quantized d0 and d1 (there are three possible quantization levels). The offset column can include the actual offset values. Alternatively, there can be a finite number of allowed offset values, and the offset column in Table 4 can include the indices of the allowed offset values. Thus, the terms "offset value" and "index" can be used interchangeably.

[0163] The final filtering process of the CCSO filter can be applied as follows: f′ = clip(f + s), Equation (21) where f is the reconstructed sample to be filtered, and s is the output offset value retrieved from the LUT, for example. In the example shown in Equation (21), the filtered sample value f' of the reconstructed sample f to be filtered can be further clipped to a range related to the bit depth.

[0164] As Figure 19As shown, an example CCSO filtering of the reconstructed samples rc of the second color of p0 and p1 using the first color (corresponding to the samples c of the first color) can be referred to as a 3-tap CCSO filter design. Alternatively, other CCSO designs with different numbers of filter taps can be used. For example, two additional taps can be added in the horizontal direction to make it a 5-tap filter (4 difference inputs), with the same quantization levels as above, and the number of difference combinations can be 81 (3 4 ).

[0165] Figure 20 An example implementation of different 3-tap CCSO filter shapes according to an embodiment of the present disclosure is shown. The term "CCSO filter shape" is used to represent the number of taps and the positions of the taps for CCSO filtering. For a specific number of taps, the number of filter shapes can be predefined as CCSO filter options. For example, in Figure 20 , for 3-tap CCSO filtering, any one of 6 different example filter shapes can be defined. Each filter shape can define the positions of three reconstructed samples (also referred to as three taps) in the first component (also called the first color component). The three reconstructed samples can include a center sample (denoted as c) and two symmetrically distributed samples, and these two symmetrically distributed samples are represented by the same number (one of the numbers from 1 to 6) in Figure 20 . In the example, the reconstructed samples to be filtered in the second color component are co-located with the center sample c. For clarity, Figure 20 the reconstructed samples to be filtered in the second color component are not shown.

[0166] For CCSO filtering, as described above, each CCSO filter basically corresponds to a LUT. The number of entries (rows) in the LUT is determined by the number of difference and quantization level combinations, as shown in Table 4 above. A CCSO filter can be associated with one or more CCSO filters or filtering parameters, including but not limited to filter shape, quantization step (and quantization level), and number of frequency bands. The CCSO filter or filtering parameters can alternatively be referred to as CCSO parameters.

[0167] In various implementations herein, a CCSO filtering unit refers to a portion of a reconstructed video frame of any size processed by a CCSO filter.

[0168] Accordingly, the above disclosure provides a CCSO filtering process for reconstructed samples of a second color component, which at a particular CCSO filtering tap, uses reconstructed samples of a first color co-located with the reconstructed samples of the second color and reconstructed samples adjacent to the co-located reconstructed samples of the first color component. The input to the CCSO filtering process is the reconstructed samples of the first color component. The CCSO filtering process generates a corresponding offset that will be applied or added to the reconstructed samples of the second color component to output adjusted reconstructed samples of the second color component. In some implementations, the first color component and the second color component may be different. For example, the first color component may be a luminance color component and the second color component may be a chrominance color component. In some other implementations, even if the first color component and the second color component are the same color component, such a filtering process may also be referred to as cross-component offset filtering. For example, the CCSO process may involve a first color component and a second color component that are both luminance color components.

[0169] As described above, CCSO filtering is specified by its shape, which includes the number of CCSO filtering taps and the positions of the CCSO filtering taps. For example, as Figure 20 shown, a 3-tap CCSO filter may be used, and the three CCSO taps may be located at Figure 20 any of the example positions marked in Figure 20 Thus, the example CCSO filtering scheme of provides six different exemplary 3-tap CCSO filters, each 3-tap CCSO filter corresponding to a LUT for finding the offset of the cross-component offset samples. The allowed CCSO filters or filter shapes may be predefined and may be indexed, and the CCSO filter being used is written into the bitstream by using the corresponding index in the predefined CCSO filters. The LUT corresponding to the allowed CCSO filters may be predefined or may be adaptively derived. If the CCSO filter is adaptively derived, then the CCSO filter needs to be written into the bitstream.

[0170] In the above disclosure, the CCSO filter is described in the context of integer CCSO filter taps. In other words, the CCSO filtering taps are located at integer sample positions relative to the center tap of the CCSO filter, as Figure 20 shown. In some additional example implementations, the CCSO filtering taps may not need to be located at integer sample positions. In other words, the CCSO filtering tap positions may be at fractional sample positions. A fractional sample position refers to a position where the horizontal or / and vertical pixel coordinate values are fractional.

[0171] Figure 21Example 3-tap CCSO filter schemes including taps at fractional positions are shown, where solid circles refer to the integer positions used in the cross-component offset filter scheme above, and hollow (non-solid) circles refer to CCSO filter taps located at fractional positions. Each unique numeric label indicates a possible 3-tap CCSO filter. As an example, Figure 21 twelve different 3-tap CCSO filters are shown. Each of these 3-tap CCSO filters can be characterized by its horizontal or vertical direction relative to the center tap. Figure 21 The CCSO filters shown in Figure 21 are only examples. Specific filter taps can be located at other integer or fractional positions, and each CCSO filter can include more than 3 CCSO filter taps. The CCSO filter tap positions can be fractional positions only in the horizontal direction, such as Figure 21 the CCSO filter tap positions 5, 6, and 11 in Figure 21 . The CCSO filter tap positions can be fractional positions only in the vertical direction, such as Figure 21 the CCSO filter tap positions 7, 8, and 12 in

[0172] As Figure 21 shown, the CCSO tap positions can be defined by phase and precision. Precision is used to represent the fractional precision of the tap position. The phase of the tap position is used to indicate which of the possible fractional positions is selected for the tap at a particular precision. As described in further detail below, the phase and precision used to select a CCSO filter or CCSO filter shape can be written in the bitstream or predefined. The CCSO filter will be defined by the number of CCSO filter taps and their positions (including precision and phase).

[0173] To apply a CCSO filter with fractional tap positions, it will be necessary to first determine the input sample values at the corresponding fractional tap positions. In some example implementations, the sample values at these fractional positions can be generated via an interpolation process using the reconstructed samples at integer positions. This interpolation process can be performed by using an interpolation filter. Once the interpolated samples are generated, the filter taps are applied to the interpolated samples located at fractional positions. Thus, the CCSO filtering process allowing for fractional tap positions involves: interpolating the reconstructed samples in the second color component to generate interpolated samples at the fractional tap positions via interpolation filtering; and performing CCSO filtering using the interpolated samples of the second color component to generate the sample offset to be applied to the first color component.

[0174] In the disclosure herein, the terms "CCSO filter tap" or "CCSO filtering tap" are used to refer to the position or location point of the CCSO filtering tap of a CCSO filter. Each CCSO filter may be associated with a number of CCSO filtering taps (e.g., 3 taps in a 3-tap CCSO filter). As described above, each tap in the CCSO filter taps for each CCSO filter may be located at an integer pixel position or a fractional pixel position. The terms "interpolated CCSO filter tap" or "interpolated CCSO filtering tap" or "interpolated filter tap" are used to refer to the value of a CCSO filter or filtering tap generated after applying an interpolation filter further disclosed hereinabove and hereinafter. A distinction can be made between two types of filters disclosed herein: a CCSO filter that includes CCSO filter taps or CCSO filtering taps; and an interpolation filter. The interpolation filter is used to generate sample values for CCSO filter taps (especially fractional taps of a CCSO filter).

[0175] In some example implementations, the above interpolation may be performed using a 2-tap interpolation filter (e.g., a bilinear filter), a 4-tap interpolation filter, or a 6-tap interpolation filter, an 8-tap interpolation filter, or an interpolation filter having any other number of interpolation filtering taps. The interpolation filter is applied to integer samples in a second color component (the input color component of CCSO filtering). For example, for 4-tap interpolation, samples at fractional positions are interpolated from 4 adjacent reconstructed integer samples. The different filtering coefficients in the interpolation filter may depend on the fractional position of the CCSO filtering tap.

[0176] In some example implementations, the number of interpolation filtering taps for an interpolation filter may vary according to the position of the fractional samples to be generated relative to the boundaries of the corresponding CCSO filtering unit. For example, a longer filter (e.g., an 8-tap filter) may be applied to internal samples of a CCSO filtering unit, while shorter filters (e.g., 2-tap filters) are used at the upper and lower boundaries of the CCSO filtering unit. This scheme may help reduce memory as it has less buffering requirements for adjacent CCSO filter units.

[0177] In some example implementations, for a fractional CCSO filter tap position having a fractional horizontal pixel coordinate and an integer vertical pixel coordinate (e.g., Figure 21 CCSO filter tap positions 5, 6, and 11 therein), horizontal filtering interpolation with multiple interpolation filter taps may be applied to derive sample values at such fractional CCSO filter tap positions.

[0178] Similarly, in some example implementations, for a fractional CCSO filter tap position having a fractional vertical pixel coordinate and an integer horizontal pixel coordinate (e.g.,Figure 21 For the CCSO filter tap positions 7, 8, and 12 in

[0179] Similarly, in some example implementations, for CCSO filter tap positions with fractional horizontal and fractional vertical pixel coordinates (e.g., Figure 21 the CCSO filter tap positions 9 and 10 in

[0180] First, horizontal (or vertical) interpolation filtering is applied to the reconstructed samples at integer positions to derive intermediate sample values at the fractional horizontal (or vertical) filter tap positions, and then these intermediate sample values are used to interpolate the sample values at the filter tap positions (fractional vertical coordinate and fractional horizontal coordinate). Figure 21 In a specific example implementation, for CCSO filter tap positions with fractional horizontal and fractional vertical pixel coordinates (e.g., Figure 21 the CCSO filter tap positions 9 and 10 in

[0181] Regarding the CCSO filter tap positions of the CCSO filter, in some example implementations, the phase and / or precision of the CCSO filter tap positions used in the horizontal and / or vertical interpolation process to derive sample values at fractional CCSO filter tap positions can be written into the bitstream at the frame level, or at any other level where cross-component sample offset control parameters are written.

[0182] For example, multiple pixel precisions of the CCSO filter tap positions can be predefined (including but not limited to 1 pixel, 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, 1 / 16 pixel, 1 / 32 pixel), and the selection between these precisions can be written into the bitstream at one of various levels.

[0183] In addition, for each selected precision, a set of allowed phases can be predefined, and the phase selection between these allowed predefined phases can be written into the bitstream. For example, for 1 / 4 pixel precision, the allowed phases can be 1 / 4 pixel, 1 / 2 pixel, or 3 / 4 pixel, and the selection between these phases can be written into the bitstream. The signaling can be explicitly included in the bitstream or can be implicitly derived based on other information extracted from the bitstream.

[0184] In some example implementations, a flag can be first written in the bitstream to indicate whether to apply integer-position CCSO filter taps or fractional-position CCSO filter taps. If the flag (e.g., "is_CCSO_fractional") is written with a value indicating that the fractional-position CCSO filter taps are not applied, the integer CCSO filter shape (a predefined CCSO filter shape, or one of a set of predefined integer CCSO filter shapes written, or a CCSO filter shape derived based on other information in the bitstream) can be determined. If the flag is written with a value indicating that the fractional-position CCSO filter taps are applied, other syntax can be further written in the bitstream to specify the fractional precision (e.g., 1 / 2 pixel, 1 / 4 pixel, 1 / 8 pixel, etc.). Thereafter, for each precision, other syntax is written to indicate the phase (or position) of the CCSO filter taps. This dual-syntax signaling of precision and phase can help achieve better coding efficiency because for lower precision (larger fractional pixel values), the signaling may require fewer bits as the bit length used to represent the possible number of phases will be less than that for higher precision.

[0185] In some other example implementations, the phase and / or precision of the CCSO filter tap positions can be predefined and hard-coded, and thus, there is no need to include signaling in the bitstream.

[0186] In some example implementations, a set of CCSO filters can be allowed (each CCSO filter associated with multiple CCSO filter tap positions, which can be fractional tap positions), where the directions of the filter tap positions between any two allowed CCSO filters do not overlap.

[0187] Figure 22 An example is shown which shows a set of allowed three-tap CCSO filters with 1 / 2 pixel precision (such that the CCSO filters can have filter taps at integer positions or 1 / 2 pixel positions), including 8 exemplary allowed CCSO filters 1, 2, 3, 4, 5, 6, 7, and 8. The lines connecting the taps in these filters do not overlap in direction. Under the constraint that the directions of the lines connecting the filter taps in each allowed CCSO filter do not overlap, Figure 21The 3-tap CCSO filters 9 and 2 (also with 1 / 2 pixel precision) in Figure 22 will not be allowed in Figure 22 because the directions of the filter taps connecting the two 3-tap CCSO filters overlap. The reason can be that those co-directional filter taps do not provide a significant difference in CCSO filtering performance, and thus the co-directional filter taps do not need to be two different options among the other CCSO filter options selected by the encoder during the encoding process, and including only one as a CCSO filter option also helps to reduce the total number of CCSO filter options, and thus reduce the number of bits for writing the selection of the CCSO filter in the bitstream.

[0188] In some example implementations, the interpolated CCSO filter taps may be applicable only to the loop filtering of the selected color component. For example, the interpolated CCSO filter taps may be applicable only to the loop filtering of the luminance color component. In other words, the interpolated CCSO filter may be applied only to the luminance samples to generate the CCSO sample offset. Specifically, the interpolated CCSO filter taps may be applicable only to the cross-component loop filtering using the luminance component as the input, and the output sample offset is applied to the chrominance samples. For other cases, the fractional CCSO filter taps that require interpolation may not be allowed.

[0189] In some example implementations, whether the interpolated CCSO filter taps can be applied to a specific color component for generating the sample offset can be written in the following headers in high-level syntax, including but not limited to the sequence header, frame header, slice header, tile header, and maximum coded block header.

[0190] In some example implementations, regarding the input of CCSO filtering or the position where the generated sample offset is applied, for different color components, the number (or position) of the fractional CCSO filter taps can be different. In other words, the CCSO filtering processes for different input color components for generating the sample offset can select different CCSO filters, and the different CCSO filters can have different numbers of fractional CCSO filter taps. Similarly, the CCSO filtering processes for generating sample offsets for different color components can select different CCSO filters, and the different CCSO filters can have different numbers of fractional CCSO filter taps for applying the sample offset to different color components.

[0191] In some example implementations, the interpolated CCSO filter taps may be applicable only to a given fractional precision. For example, the interpolated CCSO filter taps may be usable only at 1 / 2 pixel positions. For another example, the interpolated CCSO filter taps may be usable only at 1 / 2 pixel positions and 1 / 4 pixel positions.

[0192] In some example implementations, the selection of the interpolation CCS0 filter taps can be derived from neighboring samples of the current block. As such, it may not be necessary to write this selection in the bitstream, and the decoder can determine the taps.

[0193] In some example implementations, the selection of the interpolation filter taps can be written at a predefined level, where the cross-component sample offset is applied, e.g., at the frame level or slice level.

[0194] In some example implementations, a first flag can be written to indicate whether integer filter taps or fractional filter taps are currently being used. Thereafter, if the first flag is written to indicate that integer filter taps are being applied, a second index is written to indicate which integer filter taps are being applied. Otherwise, if a second flag is written to indicate that fractional filter taps are being applied, a third index is written to indicate which fractional filter taps are being applied. In such an implementation, the possible integer CCSO filters and the possible fractional CCSO filters can be grouped and indexed separately for selection by the encoder for CCSO filtering and writing the selected (integer or fractional) CCSO filter into the bitstream. This two-set signaling scheme may be more efficient in terms of signaling bit length when the number of possible integer CCSO filters and the number of possible fractional CCSO filters differ significantly.

[0195] In some other example implementations, all candidate CCSO filters and CCSO filter tap positions (whether integer or fractional) can be pooled together, and the index of the selected CCSO filter can be written in one index space in the bitstream to indicate which CCSO filter is being applied (which CCSO taps are being applied).

[0196] Figure 23A flowchart 2300 is shown. The logical flow 2300 starts at S2301. In S2310, video frames are reconstructed from a video bitstream to generate reconstructed samples of at least a first color component and a second color component of the video frames. In S2320, at least one syntax element is received from the video bitstream to determine a cross-component sample offset (CCSO) filter that is applicable to the reconstructed samples of CCSO filter units in the first color component, and the CCSO filter can be characterized by a set of CCSO filter tap positions relative to a center tap. In S2330, in response to the CCSO filter including CCSO filter tap positions at fractional pixel positions, an interpolation filter is determined and applied to the reconstructed samples of CCSO filter units in the first color component to generate interpolated samples of CCSO filter units in the first color component at the fractional pixel positions. In S2340, sample offsets are generated for each sample position in the CCSO filter units by processing the reconstructed samples and the interpolated samples of the CCSO filter units in the first color component using the CCSO filter. In S2350, these sample offsets are applied to the reconstructed samples of CCSO filter units in the second color component. The logical flow 2300 ends at S2399.

[0197] The above operations can be combined or arranged in any number or order as needed. Two or more of these steps and / or operations can be executed in parallel. The embodiments and implementations in the present disclosure can be used alone or in any combination. Additionally, each of the methods (or embodiments), encoders, and decoders can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. The embodiments in the present disclosure can be applied to luminance blocks or chrominance blocks. The term "block" can be interpreted as a prediction block, a coding block, or a coding unit (i.e., CU). The term "block" is also used herein to refer to a transform block. In the following items, when referring to "block size", it can refer to the width or height of the block, or the maximum of the width and height, or the minimum of the width and height, or the area size (width * height), or the aspect ratio of the block (width:height, or height:width).

[0198] The above techniques can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 24 A computer system (2400) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0199] Computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be assembled, compiled, linked, or through similar mechanisms to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode, etc.

[1100] These instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[1101] Figure 24 The components of the computer system (2400) shown are exemplary in nature and are not intended to impose any limitations on the scope of use or function of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one component or combination of components shown in the exemplary embodiments of the computer system (2400).

[1102] The computer system (2400) may include certain human - machine interface input devices. The human - machine interface input devices may include one or more of the following devices (only one of each is depicted): keyboard (2401), mouse (2402), touchpad (2403), touch screen (2410), data glove (not shown), joystick (2405), microphone (2406), scanner (2407), camera (2408).

[1103] The computer system (2400) may also include certain human - machine interface output devices. Such human - machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human - machine interface output devices may include tactile output devices (such as the tactile feedback of the touch screen (2410), data glove (not shown), or joystick (2405), and can also be a tactile feedback device that is not used as an input device), audio output devices (such as: speakers (2409), headphones (not depicted)), visual output devices (such as a screen (2410) including a cathode ray tube (CRT) screen, liquid crystal display (LCD) screen, plasma screen, organic light - emitting diode (OLED) screen, each screen having or not having touch - screen input function, each screen having or not having tactile feedback function, and some of the screens being capable of outputting two - dimensional visual output or output beyond three - dimensional through means such as stereoscopic picture output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted)), and printers (not depicted).

[1104] The computer system (2400) may also include human-accessible storage devices and their associated media, such as optical media including a CD / DVD ROM / RW (2420) having a CD / DVD or similar medium (2421), a thumb drive (2422), a removable hard disk drive or solid state drive (2423), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.

[1105] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[1106] The computer system (2400) may also include an interface (2454) to one or more communication networks (2455). The network can be, for example, a wireless network, a wired network, an optical network. The network can further be a local area network, a wide area network, a metropolitan area network, vehicle and industrial networks, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including Global System for Mobile Communications (GSM), 3G, 4G, 5G, LTE, etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CAN bus, etc.

[1107] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the kernel (2440) of the computer system (2400).

[1108] The kernel (2440) may include one or more central processing units (CPUs) (2441), a graphics processing unit (GPU) (2442), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (2443), a hardware accelerator (2444) for certain tasks, a graphics adapter (2450), etc. These devices can be connected via a system bus (2448) together with a read-only memory (ROM) (2445), a random access memory (2446), and an internal mass storage such as an internal non-user-accessible hard disk drive, SSD, etc. (2447). In some computer systems, the system bus (2448) can be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus (2448) of the kernel or attached to the system bus (1448) of the kernel via a peripheral bus (2449). In an example, a screen (2410) can be connected to the graphics adapter (2450). The architecture of the peripheral bus includes PCI, USB, etc.

[1109] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be those specially designed and constructed for the purposes of this disclosure, or the medium and the computer code may be of the type well-known and available to those having skill in the art of computer software.

[1110] Although the present disclosure has described multiple exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method, comprising: Reconstructing a video frame from a video bitstream to generate reconstruction samples of at least a first color component and a second color component of the video frame; Receiving at least one syntax element from the video bitstream to determine a cross-component sample offset (CCSO) filter, the CCSO filter being applicable to the reconstruction samples of a CCSO filtering unit in the first color component, the CCSO filter being characterized by a set of CCSO filter tap positions relative to a center tap; In response to the CCSO filter including CCSO filter tap positions at fractional pixel positions, determining an interpolation filter and applying the interpolation filter to the reconstruction samples of the CCSO filtering unit in the first color component to generate interpolation samples of the CCSO filtering unit in the first color component at the fractional pixel positions; Generating a sample offset for each sample position among a plurality of sample positions in the CCSO filtering unit by processing the reconstruction samples and the interpolation samples of the CCSO filtering unit in the first color component using the CCSO filter; And Applying the sample offset to the reconstruction samples of the CCSO filtering unit in the second color component.

2. The method according to claim 1, wherein The at least one syntax element includes an indication of the CCSO filter tap positions at the fractional pixel positions.

3. The method according to claim 2, wherein The interpolation filter is used to perform multi-tap interpolation.

4. The method according to claim 3, wherein The interpolation filter includes an interpolation filter with 2 taps, 4 taps, 6 taps or 8 taps.

5. The method according to claim 3, wherein The number of taps in the interpolation filter for the multi-tap interpolation depends on the position of the reconstruction samples in the CCSO filtering unit relative to the boundary of the CCSO filtering unit.

6. The method according to claim 3, wherein, The interpolation filter is used for the following operations: Performing the multi-tap interpolation in the horizontal direction only when the CCSO filter includes only fractional tap positions in the horizontal direction; Performing the multi-tap interpolation in the vertical direction only when the CCSO filter includes only fractional tap positions in the vertical direction; or When the CCSO filter includes fractional tap positions in the horizontal direction and fractional tap positions in the vertical direction, performing the multi-tap interpolation by interpolating in one of the vertical direction and the horizontal direction to generate intermediate interpolation samples, and then interpolating the intermediate interpolation samples in the other of the vertical direction and the horizontal direction.

7. The method according to claim 2, wherein The indication of the CCSO filter tap positions at the fractional pixel positions includes a first signaling syntax element and a second signaling syntax element, the first signaling syntax element and the second signaling syntax element being respectively used to indicate the CCSO fractional pixel precision for determining the CCSO filter tap positions at the fractional pixel positions and the phase at the CCSO fractional pixel precision.

8. The method according to claim 7, wherein The first signaling syntax element includes an index for indicating one of a set of predetermined CCSO fractional pixel precisions, the set of predetermined CCSO fractional pixel precisions including at least one of the following: 1 / 2 pixel precision, 1 / 4 pixel precision, 1 / 8 pixel precision, 1 / 16 pixel precision, 1 / 32 precision.

9. The method according to claim 7, wherein The second signaling syntax element includes a phase index of one of a set of predefined phases determined according to the CCSO fractional pixel precision indicated by the first signaling syntax element.

10. The method according to claim 7, wherein, The at least one syntax element further includes a flag located before the first signaling syntax element and the second signaling syntax element, the flag being used to indicate that the CCSO filter includes one or more fractional tap positions.

11. The method according to claim 1, wherein, The CCSO filter indicated by the at least one syntax element is selected from a plurality of allowed CCSO filters, and the plurality of allowed CCSO filters do not overlap in the tap direction.

12. The method according to claim 1, wherein, The interpolation filter is only applied to one or more selected color components.

13. The method according to claim 12, wherein, The one or more selected color components are predetermined.

14. The method according to claim 13, wherein, The one or more selected color components include a luminance color component.

15. The method according to claim 14, wherein, The second color component includes one of a plurality of chrominance components.

16. The method according to claim 1, wherein, The fractional pixel position includes a specific predefined fractional pixel position.

17. The method according to claim 1, wherein, The set of CCSO filter tap positions is implicitly derived using neighboring samples of the current block.

18. The method according to claim 1, wherein, The at least one syntax element is used to: indicate a selection from a CCSO filter with fractional tap positions or a CCSO filter with only integer tap positions using an independent index space.

19. The method according to claim 1, wherein, The at least one syntax element is used to: indicate a selection from a CCSO filter with fractional tap positions or a CCSO filter with only integer tap positions using a single index space.

20. An apparatus, comprising a memory and a processor, the memory for storing instructions, the processor for executing the instructions to perform the following operations: Reconstruct a video frame from a video bitstream to generate reconstructed samples of at least a first color component and a second color component of the video frame; Receive at least one syntax element from the video bitstream to determine a cross-component sample offset (CCSO) filter, the CCSO filter being applicable to the reconstructed samples of the CCSO filtering units in the first color component, the CCSO filter being characterized by a set of CCSO filter tap positions relative to a center tap; In response to the CCSO filter including CCSO filter tap positions at fractional pixel positions, determine an interpolation filter and apply the interpolation filter to the reconstructed samples of the CCSO filtering units in the first color component to generate interpolated samples of the CCSO filtering units in the first color component at the fractional pixel positions; Generate a sample offset for each sample position among a plurality of sample positions in the CCSO filtering unit by processing the reconstructed samples and the interpolated samples of the CCSO filtering units in the first color component using the CCSO filter; and Apply the sample offset to the reconstructed sample of the CCSO filtering unit in the second color component.