Adaptive Loop Filtering for the Output from Offline Fixed Filtering

A two-stage filtering process combining fixed and adaptive filters addresses the challenges of video coding in advanced standards by enhancing image quality and reducing complexity, achieving better coding efficiency.

JP2025522665APending Publication Date: 2025-07-17TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024547670
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-06-09
Filing Date
2023-06-14
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in effectively utilizing both fixed and adaptive filters to enhance image quality while maintaining efficient compression, particularly in advanced standards like VVC, where the complexity and overhead of adaptive loop filtering can be high.

Method used

A two-stage filtering process is employed, where fixed filters with pre-defined coefficients are applied first, followed by adaptive filters with changeable coefficients, to improve image quality and reduce the number of coefficients needed for adaptive filtering, thereby enhancing coding efficiency.

Benefits of technology

This approach achieves improved image quality with reduced computational complexity and overhead, maintaining or exceeding the performance of traditional adaptive filtering methods while optimizing coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522665000001_ABST
    Figure 2025522665000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide methods and apparatus for video and / or image coding. The apparatus includes a processing circuit that receives a bitstream including pictures. The processing circuit applies one or more first fixed filters having fixed filter coefficients to samples within the picture to obtain one or more first filtered outputs for each of the samples within the picture. After applying the one or more first fixed filters, the processing circuit applies one or more second adaptive filters having changeable coefficients to the one or more first filtered outputs to obtain a second filtered output for a current sample among the samples, and decodes the picture based at least on the second filtered sample for the current sample within the picture. Each coefficient of the second adaptive filter is applicable to a corresponding one of the one or more first filtered outputs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Related Applications] This application claims the benefit of priority of U.S. Patent Application No. 18 / 208,163, "ADAPTIVE LOOP FILTERING ON OUTPUT(S) FROM OFFLINE FIXED FILTERING", filed on June 9, 2023, which claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 389,653, "Adaptive Loop Filter on Offline Fixed Filtering", filed on July 15, 2022. The disclosures of the foregoing applications are hereby incorporated herein by reference in their entirety.

[0002] [Technical Field] The present disclosure generally describes embodiments related to video coding.

Background Art

[0003] The background description provided herein is for the purpose of generally presenting the background of the present disclosure. The research of the presently named inventors, to the extent it is not considered prior art at the time of filing in the context of the research described in this background chapter, is not expressly or implicitly admitted as prior art to the present disclosure, similar to aspects of the description that may not be considered prior art.

[0004] Image / video compression helps transfer image / video files among various devices, storage, and networks while minimizing quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancies. For example, a video codec can use a technique called intra prediction that can compress images based on spatial redundancy. For example, in intra prediction, reference data from the current picture being reconstructed can be used for sample prediction. In another example, a video codec can use a technique called inter prediction that can compress images based on temporal redundancy. For example, in inter prediction, motion compensation can be used to predict samples within the current picture from previously reconstructed pictures. Motion compensation is generally indicated by a motion vector (MV).

SUMMARY OF THE INVENTION

[0005] Disclosed aspects provide methods and apparatuses for video and / or picture encoding / decoding. The method includes receiving a bitstream including a picture. One or more first filters can be applied to samples within the picture to obtain one or more first filtered outputs of each of the samples within the picture. The samples within the picture can include a current sample filtered by a second filter. Each of the one or more first filtered outputs can include a first filtered sample that is a sample within the picture filtered by each first filter within the one or more first filters. The second filter can be applied to the one or more first filtered outputs to obtain a second filtered sample of the current sample. Each coefficient of the second filter can be applied to a corresponding one of the one or more first filtered outputs.

[0006] The machine includes a processing circuit that receives a bitstream including a picture. The processing circuit applies one or more first filters to samples within the picture to obtain one or more respective first filtered outputs of the samples within the picture. The samples within the picture can include a current sample that is filtered by a second filter. Each of the one or more first filtered outputs includes a first filtered sample that is a sample within the picture filtered by each first filter within the one or more first filters. The processing circuit can apply the second filter to the one or more first filtered outputs to obtain a second filtered sample of the current sample. Each coefficient of the second filter is applied to a corresponding one of the one or more first filtered outputs.

[0007] In an embodiment, the method includes receiving a bitstream including a picture. The method includes applying one or more first fixed filters having fixed filter coefficients to samples within the picture to obtain one or more respective first filtered outputs of the samples within the picture. After applying the one or more first fixed filters, the method applying one or more second adaptive filters having changeable coefficients to the one or more first filtered outputs to obtain a second filtered output of a current sample among the samples, and decoding the picture based at least on the second filtered sample of the current sample within the picture.

[0008] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video encoding / decoding, cause the computer to execute a method for video encoding / decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Further features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

[0010]

Figure 1

[0011]

Figure 2

[0012]

Figure 3

[0013]

Figure 4

[0014]

Figure 5

[0015]

Figure 6A

Figure 6B

Figure 6C

Figure 6D

[0016]

Figure 7

[0017]

Figure 8

Figure 9

Figure 10

Figure 11

[0018]

Figure 12

[0019]

Figure 13

[0020]

Figure 14

Mode for Carrying Out the Invention

[0021] FIG. 1 shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is a video encoder and a video decoder in a streaming environment, which is an example of the application of the disclosed subject matter. The disclosed subject matter is equally applicable to, for example, storage of compressed video in digital media including video conferencing, digital TV, streaming services, CDs, DVDs, memory sticks, etc., other video-enabled applications, etc.

[0022] A video processing system (100) includes, for example, a video source (101) that generates an uncompressed video picture stream (102), and a capture subsystem (113) that can include, for example, a digital camera. In one example, the video picture stream (102) includes samples captured by a digital camera. The video picture stream (102) is shown in bold lines to emphasize its high data volume when compared to the encoded video data (104) (or coded video bitstream), and can be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) includes hardware, software, or a combination thereof, and can enable or implement aspects of the disclosed subject matter as detailed below. The encoded video data (104) (or encoded video bitstream) is shown in thin lines to emphasize its low data volume when compared to the video picture stream (102), and can be stored in a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems (106) and (108) of FIG. 1, can access the streaming server (105) to read copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include, for example, a video decoder (110) within an electronic device (130). The video decoder (110) decodes an input copy (107) of the encoded video data and generates an output video picture stream (111) that can be rendered on a display (112) (such as a display screen) or another rendering device (not shown). In some streaming systems, the encoded video data (104), (107), and (109) (such as a video bitstream) can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is known informally as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC.

[0023] Note that the electronic devices (120) and (130) may include other components (not shown). For example, the electronic device (120) can include a video decoder (not shown), and the electronic device (130) can also include a video encoder (not shown).

[0024] FIG. 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). In the example of FIG. 1, the video decoder (210) can be used in place of the video decoder (110).

[0025] The receiver (231) can receive one or more coded video sequences decoded by the video decoder (210). In an embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequence may be received from a channel (201) which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams that may be transferred to respective using entities (not shown). The receiver (231) may separate the coded video sequence from the other data. To remove network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter, “parser (220)”). In certain applications, the buffer memory (215) is part of the video decoder (210). Alternatively, it may be external to the video decoder (210) (not shown). Still alternatively, for example, to remove network jitter, in addition to another buffer memory (215) that may be external to the video decoder (210) or internal to the video decoder (210) to process playout timing, a buffer memory (not shown) may be present. When the receiver (231) is receiving data controllably from a storage / transfer device with sufficient bandwidth or from an isosynchronous network, the buffer memory (215) may not be necessary or may be made small. For use in a best-effort packet network such as the Internet, a buffer memory (215) may be required, may be relatively large, advantageously be of an adaptive size, and may be implemented at least in part outside the operating system or similar elements (not shown) external to the video decoder (210).

[0026] Video decoder (210) may include a parser (220) to reconstruct symbols (221) from the coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (210), and in some cases, information for controlling a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but can be coupled to the electronic device (230) as shown in FIG. 2. The control information for the rendering device may be in the form of an SEI (Supplemental Enhancement Information) message or a VUI (Video Usability Information) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence can follow video coding techniques or standards and can follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser (220) may extract a set of subgroup parameters from the coded video sequence based on at least one parameter corresponding to at least one subgroup of pixels in the video decoder for at least one of the subgroups. The subgroups may include GOP (Groups of Picture), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.

[0027] The parser (220) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (215) to generate symbols (221).

[0028] The reconstruction of symbol (221) may include a plurality of different units depending on the type of the coded video picture or a portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is included can be controlled by group control information parsed by parser (220) from the coded video sequence. Such a flow of subgroup control information between parser (220) and the following plurality of units is not shown for clarity.

[0029] Beyond the function blocks already mentioned, video decoder (210) can be conceptually subdivided into a number of functional units, as will be described later. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0030] The first unit is scaler / inverse transform unit 251. Scaler / inverse transform unit (251) receives, as symbol (221) from parser (220), quantized transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. Scaler / inverse transform unit (251) can output a block including sample values that can be input to aggregator (255).

[0031] In some cases, the output samples of the scaler / inverse transform unit (251) may be related to intra-coded blocks. An intra-coded block is a block that does not use prediction information from previously reconstructed pictures but can use prediction information from the previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (255) adds, in some cases, for each sample, the prediction information generated by the intra prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0032] In other cases, the output samples of the scaler / inverse transform unit (251) may be related to inter-coded, and in some cases motion-compensated, blocks. In such cases, the motion compensation prediction unit (253) can access the reference picture memory (257) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbol (221) related to the block, these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) to generate the output sample information (in this case, called the residual samples or the residual signal). The address in the reference picture memory (257) where the motion compensation prediction unit (253) fetches the prediction samples can be controlled by the available motion vectors of the motion compensation prediction unit (253) in a form of a symbol (221) that may have, for example, X, Y, and reference picture components. Motion compensation can include interpolation of the sample values fetched from the reference picture memory (257) when an exact motion vector of sub-samples is in use, a motion vector prediction mechanism, etc.

[0033] The output samples of the aggregator (255) can undergo various loop filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the coded video sequence (also called the coded video bitstream) and that enable the loop filter unit (256) to use symbols (221) from the parser (220). Video compression can respond not only to previously reconstructed and loop-filtered sample values, but also to meta information obtained during the decoding of a previous portion (in decoding order) of the coded picture or coded video sequence.

[0034] The output of the loop filter unit (256) can be a sample stream that can be output to the rendering device (212) and stored in the reference picture memory (257) for use in future inter-picture prediction.

[0035] Once a particular coded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a fresh current picture buffer can be reallocated before starting the reconstruction of subsequent coded pictures.

[0036] The video decoder (210) may perform a decoding operation according to a standard such as ITU-T Rec.H.265 or a predetermined video compression technique. In the sense that the coded video sequence conforms to both the video compression technique or standard and the profile documented in the video compression technique or standard, the coded video sequence may conform to the syntax specified by the video compression technique or standard in use. Specifically, the profile can select specific tools from all the tools available in the video compression technique or standard as tools that can be used only under the profile. Also, what is necessary for compliance can be that the complexity of the coded video sequence is within the limits determined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (e.g., measured in megasamples per second), the maximum reference picture size, etc. The limits set by the level can be further restricted in some cases through the HRD (Hypothetical Reference Decoder) specification and the metadata for HRD buffer management signaled in the coded video sequence.

[0037] In an embodiment, the receiver (231) may receive additional (redundant) data together with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0038] FIG. 3 shows an exemplary block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used instead of the video encoder (103) in the example of FIG. 1.

[0039] The video encoder (303) may receive video samples from a video source (301) (not part of the electronic device (320) in the example of FIG. 3) that can capture a video image to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0040] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media providing system, the video source (301) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may subsequently be provided as a plurality of individual pictures that give motion when viewed in succession. The pictures themselves may be organized as a spatial array of pixels. Each pixel may contain one or more samples depending on the sampling structure, color space, etc. in use. One of ordinary skill in the art can immediately understand the relationship between pixels and samples. The following description focuses on samples.

[0041] According to an embodiment, the video encoder (303) may code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other time constraints required. Implementing an appropriate coding speed is one function of the control unit (350). In some embodiments, the control unit (350) controls other functional units described below and is functionally coupled to the other functional units. The coupling is not shown for clarity. Parameters set by the control unit (350) may include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, GOP (group of pictures) layout, maximum motion vector search range, etc. The control unit (350) may be configured to have other appropriate functions related to the video encoder (303) optimized for a specific system design.

[0042] In some embodiments, the video encoder (303) is configured to operate within a coding loop. As a very simplified explanation, in one example, the coding loop may include a source coder (330) (which is responsible for generating symbols, such as a symbol stream, based on the input picture and reference pictures to be coded), and a (local) decoder (333) built into the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in the same way as a (remote) decoder does. The reconstructed sample stream (sample data) is input into the reference picture memory (334). When the decoding of the symbol stream results in a bit-exact result independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the exact same sample values as the decoder would "see" as reference picture samples when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example due to channel errors) is also used similarly in several related arts.

[0043] The operation of the "local" decoder (333) can be the same as that of a "remote" decoder such as the video decoder (210) detailed above in relation to FIG. 2. However, referring briefly to FIG. 2, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (345) and the parser (220) can be lossless, the entropy decoding part of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).

[0044] In an embodiment, decoder techniques other than parsing / entropy decoding present in a decoder exist in a corresponding encoder in the same or substantially the same functional form. Accordingly, the disclosed subject matter focuses on decoder operations. The description of encoder techniques can be omitted since they are the reverse of the decoder techniques described comprehensively. In certain areas, more detailed descriptions are provided below.

[0045] During operation, in some examples, the source coder (330) may perform motion-compensated predictive coding. This predictively codes an input picture by referring to one or more previous coded pictures from a video sequence designated as a "reference picture". In this method, the coding engine (332) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that may be selected as a predictive reference for the input picture.

[0046] The local video decoder (333) may decode the coded video data of a picture that may be designated as a reference picture based on the symbols generated by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data can be decoded in a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a reproduction of the source video sequence with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder for the reference picture, resulting in a reconstructed reference picture to be stored in the reference picture memory (334). Thus, the video encoder (303) may store a copy of the reconstructed reference picture having the same content as the reconstructed reference picture obtained by the remote video decoder (in the absence of transmission errors).

[0047] The predictor (335) may perform predictive search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for sample data (such as a candidate reference pixel block) or specific metadata such as reference picture motion vectors, block shapes, etc. that can function as an appropriate predictive reference for the new picture. The predictor (335) may operate on a sample block - pixel block basis to find an appropriate predictive reference. In some examples, the input picture may have a predictive reference drawn from a plurality of reference pictures stored in the reference picture memory (334) as determined by the search results obtained by the predictor (335).

[0048] The control unit (350) may manage the coding operations of the source coder (330), including, for example, setting parameters and subgroup parameters used for the coding of video data.

[0049] The outputs of all the aforementioned functional units may undergo entropy coding in the entropy coder (345). The entropy coder (345) converts the symbols generated by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc.

[0050] The transmitter (340) may buffer the coded video sequence generated by the entropy coder (345) for transmission via a communication channel (360) which may be a hardware / software link to a storage device capable of storing the coded video data. The transmitter (340) may merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).

[0051] The control unit (350) may manage the operation of the video encoder (303). During coding, the control unit (350) may assign to each coded picture a type of the specific coded picture that may affect the coding technique applicable to each picture. For example, a picture may often be assigned as one of the following picture types.

[0052] An intra picture (I picture) may be a picture that can be coded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, for example, IDR (Independent Decoder Refresh) pictures. Those skilled in the art recognize the variations of I pictures and their individual applications and characteristics.

[0053] A predictive picture (P picture) may be a picture that can be coded and decoded using intra prediction or inter prediction, typically using one motion vector and a reference index to predict the sample values of each block.

[0054] A bi-directionally predictive picture (B picture, Bi-directionally Predictive Picture (B Picture)) may be a picture that can be coded and decoded using intra prediction or inter prediction, using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0055] Source pictures are generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and may be coded block by block. The blocks may be coded predictively by reference to other (already coded) blocks determined by the coding assignment applied to each picture of the block. For example, blocks of an I picture may be coded non-predictively, or they may be coded predictively by reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be coded predictively via spatial prediction or via temporal prediction by reference to one previously coded reference picture. Blocks of a B picture may be coded predictively via spatial prediction or via temporal prediction by reference to one or two previously coded reference pictures.

[0056] Video encoder (303) may perform coding operations in accordance with a predetermined video coding technology or standard such as ITU-T Rec. H.265. In such operations, video encoder (303) may perform various compression operations including predictive coding operations that utilize temporal and spatial redundancies in the input video sequence. The coded video data may thus conform to a syntax specified by the video coding technology or standard being used.

[0057] In one embodiment, transmitter (340) may transmit additional data along with the coded video. Source coder (330) may include such data as part of the coded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0058] Video may be captured as a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (which may be abbreviated as intra-prediction) utilizes the spatial correlation within a given picture, and inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture during encoding / decoding is referred to as the current picture and is partitioned into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been previously coded and is still buffered in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension to identify the reference picture when multiple reference pictures are in use.

[0059] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. According to the bi-prediction technique, two reference pictures such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order respectively) are used. A block within the current picture can be coded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by the combination of the first reference block and the second reference block.

[0060] Furthermore, in order to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.

[0061] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed within a unit of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are partitioned into coding tree units (CTUs) for compression. CTUs within a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Usually, a CTU includes three coding tree blocks (CTBs), namely one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a 64×64 pixel CTU can be partitioned into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter-prediction type or an intra-prediction type. The CU is partitioned into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Usually, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed within a unit of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0062] Note that the video encoders (103) and (303), and the video decoders (110) and (210) can be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303), and the video decoders (110) and (210) can be implemented using one or more processors that execute software instructions.

[0063] Block-based video / image coding architectures, such as classical block-based hybrid video coding architectures, can be applied, for example, in VVC. Certain coding tools can be included in the basic building blocks, for example, to improve compression in VVC.

[0064] In an embodiment, a quadtree and multi-type tree (QT+MTT) scheme that uses a four-way split followed by two-way and three-way splits for the partitioning structure, as in VVC, is used to replace the quadtree with multiple partition types used, for example, in HEVC. Individual partitioning tree structures can be supported for each luma channel and chroma channel. In the case of inter-frames, the luma channel and chroma channel within one CTU can share the same coding tree structure. In the case of intra-frames, the luma channel and chroma channel can have individual trees to improve the coding efficiency of the chroma channel.

[0065] In VVC, various inter-prediction modes can be used. For an inter-predicted CU, the motion parameters can include a motion vector, one or more reference picture indices, a reference picture list use index, and additional information on specific coding functions used for generating the inter-predicted samples. The motion parameters can be signaled explicitly or implicitly. If a CU is coded in skip mode, the CU can be associated with a PU and can have no significant residual coefficients, coded motion vector delta or MV difference (e.g., MVD), or reference picture index. The merge mode can be specified when the motion parameters of the current CU are obtained from neighboring CUs that include spatial and / or temporal candidates and additional information such as those introduced optionally in VVC. The merge mode can be applied not only to skip mode but also to inter-predicted CUs. As an example, an alternative to the merge mode is the explicit transmission of motion parameters, where, for example, motion information including MV, corresponding reference picture indices for each reference picture list, and reference picture list use flags and other information are signaled explicitly for each CU.

[0066] In embodiments such as VVC, the VVC Test model (VTM) reference software includes one or more refined inter prediction coding tools, including extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode using symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), and the like.

[0067] In some embodiments, for an intra-predicted CU, the samples of the intra-predicted CU are predicted from reference samples in neighboring blocks to the left and above the current CU (e.g., the intra-predicted CU), and the current CU has been previously decoded before in-loop filtering within the same picture. In an example such as HEVC, 35 intra-picture prediction modes including planar mode, DC mode (e.g., reference sample average mode), and 33 directional angle modes can be used. Without limitation, (i) 93 intra-picture direction prediction angles including wide-angle intra prediction (WAIP) mode, (ii) two sets of 4-tap interpolation filters, position-dependent prediction combination (PDPC), multiple reference line (MRL), cross-component linear model (CCLM), and intra sub-partition (ISP), and other various tools can be used for intra prediction in, for example, VVC.

[0068] To achieve better energy compression of residual data and further reduce the quantization error of the transformed coefficients, for example, in VVC, the following tools can be used. These tools can include, but are not limited to, non-square transform, multiple transform selection (MTS) including explicit MTS and implicit MTS, low-frequency non-separable transform (LFNST), subblock transform (SBT), dependent quantization (DQ), joint coding of chroma residual (JCCR), etc.

[0069] In some examples such as VVC, a remapping operation and three in-loop filters can be sequentially applied to the reconstructed frame or picture to remove different types of artifacts. For example, sample-based processing such as Luma Mapping with Chroma Scaling (LMCS) processing is performed. Next, a deblocking filter can be used to reduce blocking artifacts. The sample-adaptive offset (SAO) filter can be applied to the deblocked picture to attenuate ringing and banding artifacts. An alternative loop-filter (ALF) can be applied to reduce other potential distortions introduced by the transform and quantization processes. In some examples such as VVC, two operations can be included in the ALF. The first operation can be a block-based ALF, such as an ALF that uses block-based filter adaptation for both luma and chroma samples, and the second operation can be a cross-component alternative loop filter (CC-ALF) for chroma samples only.

[0070] In some examples such as VVC, two filter shapes (e.g., two diamond filter shapes) can be used for the block-based ALF. FIG. 4 shows exemplary ALF filter shapes including a 5×5 diamond shape (left) and a 7×7 diamond shape (right) according to an embodiment of the present disclosure. In one example, the 7×7 diamond shape is applied to the luma component and the 5×5 diamond shape is applied to the chroma component.

[0071] Based on the direction (or directivity) and the activity of the local gradient, one of up to 25 filters can be selected for a block (e.g., a 4×4 block or each 4×4 block). According to the directionality and activity of the local gradient, a block (e.g., a 4×4 block) can be classified and categorized into one of 25 classes. Each class can be assigned each filter coefficient. Prior to filtering, geometric transformations such as 90-degree rotation, diagonal inversion, and vertical inversion can be applied to the filter shape according to the gradient value calculated for the block. The geometric transformation is equivalent to applying the geometric transformation to the samples within the filter support region. The motivation for performing the geometric transformation may include performing the ALF more similarly for each block by aligning the directionality of each block.

[0072] In addition to filter adaptation at the luma block level (e.g., 4×4 block level), filter adaptation at the CTU level can be used in the ALF. Each CTU can use one of a filter set calculated from the current slice, one of the filter sets signaled in already-coded slices, or one of pre-defined filter sets (e.g., 16 offline-trained filter sets). Within each CTU, the selected filter set can be applied to each 4×4 block. The filter coefficients and clipping indexes can be transmitted (or signaled) in an adaptive parameter set (APS) (e.g., multiple AFL APS) for the ALF. The ALF APS can include one luma filter set including up to 8 chroma filters and up to 25 filters. The index i c indicating the luma filter class can be included for each of the 25 luma classes. In an example, to reduce signaling overhead, the filter coefficients of different classifications of the luma component can be merged. By merging different classes, the number of bits indicating the filter coefficients can be reduced.

[0073] FIG. 5 shows an example of CC-ALF. CC-ALF can refine chroma sample values within ALF processing using luma sample values (e.g., SAO luma ). A linear filtering operation (e.g., CC-ALF Cb or CC-ALF Cr ) can receive a luma sample (e.g., SAO luma ) as input and generate correction values (ΔRC b , ΔRC r ) for chroma sample values. The correction values can be generated independently for each chroma component (e.g., Cb or Cr), as shown in FIG. 5. In one example, the correction values (e.g., ΔRC b or ΔRC r ) can be added to the output from ALF for the chroma component (e.g., ALF chroma ) to generate a filtered chroma output such as Cb or Cr.

[0074] For block classification of the luma component, a block (e.g., a luma block) can be categorized or classified as one of a plurality of (e.g., 25) classes. The classification index C can be derived using the following equation (1) based on the directionality parameter D and the quantized value of activity A^:

Equation

Equation

[0075] To reduce the complexity of the above block classification, subsampled 1D Laplacian calculations can be applied. FIGS. 6A to 6D each show an example of the subsampled positions used to calculate the vertical gradient g v (FIG. 6A), the horizontal gradient g h (FIG. 6B), the diagonal gradient g d1 (FIG. 6C), and the diagonal gradient g d2 (FIG. 6D). The same subsampled positions can be used for gradient calculations in different directions. In FIG. 6A, the label "V" indicates the subsample positions for calculating the vertical gradient g v . In FIG. 6B, the label "H" indicates the subsample positions for calculating the horizontal gradient g h . In FIG. 6C, the label "D1" indicates the subsample positions for calculating the diagonal gradient g d1 . In FIG. 6D, the label "D2" indicates the subsample positions for calculating the diagonal gradient g d2 .

[0076] The maximum values g v and g h of the gradients in the horizontal and vertical directions g h,v max and the minimum values g h,v min can be set as follows.

Equation

Equation

Equation

[0077] The activity value A can be calculated as follows:

Equation

[0078] For example, for the chroma components in a picture, block classification is not applied, and thus a single set of ALF coefficients can be applied to each chroma component.

[0079] Geometric transformations can be applied to the filter coefficients and the corresponding filter clipping values (also called clipping values). Before filtering a block (e.g., a 4×4 luma block), geometric transformations such as rotation, diagonal, and vertical flipping can be applied to, for example, the gradient values (e.g., g v , g h , g d1 , and / or g d2 ) calculated for the block, to the filter coefficients f(k, l) and the corresponding filter clipping values c(k, l). The geometric transformation applied to the filter coefficients f(k, l) and the corresponding filter clipping values c(k, l) is equivalent to applying a geometric transformation to the samples within the region supported by the filter. Geometric transformations can make the different blocks to which ALF is applied more similar by aligning each direction.

[0080] Three geometric transformations including diagonal flipping, vertical flipping, and rotation can each be executed and described as follows in equations (9) to (10):

Equation

[0081] In embodiments such as the explorative compression model 5 (ECM5), ALF gradient subsampling and ALF virtual boundary processing are not used. The block size for classification can be reduced, for example, from 4×4 to 2×2. The filter sizes of the luma and chroma components for which the ALF coefficients are signaled can be increased to 9×9 (e.g., 9×9 diamond shape). To filter the luma samples, three different classifiers C0, C1, and C A such as different classifiers and three different filter sets F0, F1 and F A such as different filter sets can be used. The three classifiers C0, C1, and C A can each correspond to the three filter sets F0, F1, and F A respectively. The filter sets F0 and F1 can include fixed filters (e.g., stored in the decoder and encoder) with coefficients trained for the classifiers C0 and C1. In one example, the filter sets F0 and F1 are not signaled. The filter set F A can be called an adaptive filter. The coefficients of the set of filters F A (e.g., ALF) can be signaled. The set (e.g., F0, F1, or F AWhich filter in ( ) is used for a given sample (or current sample) R(x0,y0) is determined by the class (e.g., C0, C1, or C A ) assigned to the sample using a classifier (e.g., C0, C1, or C A ). In one example, two fixed filters F0 and F1 (e.g., two 13×13 diamond-shaped fixed filters F0 and F1) are applied to derive two intermediate outputs including, for example, respective intermediate samples R0(x,y) and R1(x,y). The filter (or adaptive filter) F A is applied to R0(x,y), R1(x,y), and neighboring samples of the current sample R(x0,y0), and the filtered sample R~(x0,y0) can be determined as follows.

Equation

Equation

[0082] In embodiments such as ECM5, the class indicated by classifier C i (e.g., C0, C1, or C A ) is the directionality (or directionality parameter) D as shown in Equation (13).i and Activity A i ^(or the quantized value A of the activity i ^), it can be assigned to a block (e.g., each 2x2 block). Here, ii is 0, 1, or A. [Number] Here, M D,i can represent the total number of the direction D i . In the example, the class C0 determined by Equation (13) indicates the filter within the filter set F0. The class C1 determined by Equation (13) indicates the filter within the filter set F1. The class C A determined by Equation (13) indicates the filter within the filter set F A .

[0083] Similar to the description related to Equations (1) to (5), in VVC etc., for the horizontal gradient g h i , the vertical gradient g v i , and the two diagonal gradients g d1 i and g d2 i , the values can be calculated for each sample using the 1-D Laplacian. Here, i is 0, 1, or A corresponding to the classifier C0, C1, or C A . The sum of the sample gradients within a window (e.g., a 4x4 window) covering the target block (e.g., the target 2×2 block) can be used for the classifier C0 (e.g., i = 0). The sum of the sample gradients within a window (e.g., a 12x12 window) can be used for the classifiers C1 (e.g., i = 1) and C A (e.g., i is A). The sum of the horizontal, vertical, and two diagonal gradients are respectively represented by g h i 、 g v i g d1 i g d2 i and the direction Di is r h,v i and r d1,d2 i can be determined by comparing with a threshold value.

Number

[0084] Directionality D A can be derived using two threshold values (e.g., 2 and 4.5), such as in VVC. For directionality D0 and D1, the edge strength (e.g., the horizontal edge strength relative to the vertical edge strength) E HV i and the edge strength (e.g., the diagonal edge strength) E D i can be calculated. Threshold values Th = [1.25, 1.5, 2, 3, 4.5, 8] can be used. Here, Th[0] to Th[5] are 1.25, 1.5, 2, 3, 4.5, and 8 respectively. The edge strength E HV i is r h,v i can be set to 0 when ≤ Th[0]. Otherwise, the edge strength E HV i is r h,v i > Th[E HV i - 1] can be made the largest integer like this. The edge strength E D i is r d1,d2 i can be set to 0 when ≤ Th[0]. Otherwise, the edge strength E D i is r d1,d2 i > Th[E D i - 1] can be made the largest integer like this. r h,v i > r d1,d2 i , that is, when the horizontal / vertical edge is dominant, D i can be derived using Figure 7A. Alternatively, r h,vi ≤ r d1,d2 i In the case where the diagonal edge is dominant, D i can be derived using FIG. 7B.

[0085] A i To obtain Â, the sum of the vertical and horizontal gradients A i can be mapped (e.g., quantized) to the range from 0 to n, where n is equal to 4 for  and 15 for A0̂ and A1̂. In some examples, in an APS including ALF filter coefficients (e.g., ALF_APS), up to four luma filter sets can be signaled, and each set can have up to 25 filters. A ^ is equal to 4 for A, and equal to 15 for A0^ and A1^. In some examples, in an APS (e.g., ALF_APS) including ALF filter coefficients, up to four luma filter sets can be signaled, and each set can have up to 25 filters.

[0086] Class C corresponding to the adaptive filter set F based on Equation (13) A can be calculated based on the gradient and thus can be called a gradient-based classifier. Classification in ALF can be extended with alternative classifiers such as band-based classifiers. For the case of signaled luma filter sets, a flag can be signaled to indicate whether an alternative classifier (e.g., band-based classifier) is applied. Geometric transformations are not applied to alternative band-based classifiers. When a band-based classifier is applied, the sum of the sample values of a block (e.g., a 2×2 luma block) can be calculated. The class index (or class A ) can be calculated using Equation (15). index ) can be calculated using Equation (15). [Number] The sample bit depth indicates the number of bits per sample. In the example, a class index such as the class index determined using Equation (15) can indicate a filter within the adaptive filter set F A can indicate a filter within the adaptive filter set F.

[0087] The offline filtering taps can be extended for use with the ALF. The offline filtered taps can provide additional information for ALF rumafiltering. In one example, the extension of the offline filtered taps is used to improve the performance of the ALF. FIG. 8 shows an example of an extended filter that includes filtering taps (or filtering coefficients). The extended filter of FIG. 8 can be, for example, an adaptive filter (such as an ALF) extended from a filter within the adaptive filter set F A described above. The filtering taps in the adaptive filter shown in FIG. 8 include a first tap or coefficient (e.g., also called a spatial tap) including c0 to c 19 , and can include a second tap or coefficient including c 20 to c 27 . In the example shown in FIG. 8, the spatial taps (e.g., c0 to c 19 ) are maintained in diamonds (e.g., c0 to c 19 in FIG. 8 corresponds to c0 to c 19 in Equation (12)), and the number of second taps (e.g., c 20 to c 27 ) applied to the results (or outputs) (e.g., R0 and R1) from fixed filters (e.g., F0 and F1) is increased from, for example, 2 (e.g., c 20 to c 21 ) in Equation (12) to 8 (e.g., c 20 to c 27 ) in FIG. 8. The second taps (e.g., c 20 to c 27 ) can be applied to the output / results of the fixed filters and can thus be called taps based on the fixed filter results. In one example, coefficients c 20 to c 26 are applied to R0 and coefficient c 27 is applied to R1. In various examples, an adaptive filter including filtering taps c 20 to c 27 is signaled.

[0088] In the related art as shown in FIG. 8, a first tap (e.g., c0 to c19 such as a spatial filter tap) and a second tap (e.g., c 20 ~c 27 ) are both used to filter the samples. Additional taps or coefficients (e.g., c 22 ~c 27 ) can improve the coding performance. For example, the filtered samples from the filter of FIG. 8 can be more accurate (e.g., have a larger signal-to-noise ratio (SNR)) than the filtered samples from the filter of Equation (12). Comparing the filtering coefficients of FIG. 8 and Equation 12, in the example of FIG. 8, more coefficients (e.g., c 21 ~c 27 ) need to be signaled, and thus, even when using the filter shown in FIG. 8, it may not be possible to maximize the coding efficiency.

[0089] Aspects of the present disclosure provide techniques for ALF in offline fixed filtering. The techniques can combine an ALF (e.g., F A ) with an offline trained filter (e.g., F0 and / or F1) used in the ALF. The ALF can be performed on an image or picture filtered by an offline trained filter (e.g., F0, F1, etc.) including offline trained fixed filter taps.

[0090] Embodiments of the present disclosure describe filtering processes. A first filtering in the filtering process (e.g., filtering by a fixed filter) can be applied to samples such as samples within a picture using one or more first filters (e.g., one or more fixed filters) to determine one or more corresponding first filtered outputs. A second filter in the filtering process (or a second adaptive filter with variable coefficients) can be applied to the one or more first filtered outputs to determine a second filtered output (also referred to as the final filtered output). For example, a sample within a picture can include the current sample being filtered by the second filter, and the second filter is applied to the one or more first filtered outputs to obtain the second filtered sample of the current sample. According to one embodiment of the present disclosure, each coefficient of the second filter can be applied to a corresponding one of the one or more first filtered outputs. For example, each coefficient of the second filter is either (i) directly applied to a corresponding one of the one or more first filtered outputs or (ii) indirectly applied to a result (e.g., a difference such as a clipped difference) based on a corresponding one of the one or more first filtered outputs.

[0091] In some examples, one or more second adaptive filters are applied to the one or more corresponding first filtered outputs.

[0092] In one example, a sample within a picture includes samples related to the current sample, such as samples within a region (e.g., a surrounding region) surrounding the current sample.

[0093] In one embodiment, the filters among the one or more first filters are from a predefined set of filters. In one embodiment, the second filter is selected from an adaptive filter set signaled in a bitstream.

[0094] Samples in the picture before the first filtering (e.g., denoted by R(x, y)) can be called unfiltered samples because they are not filtered by one or more first filters or second filters. Unfiltered samples (e.g., R(x, y)) may be filtered by other filters such as the SAO filter. Each of the one or more first filtered outputs can include a corresponding first filtered sample corresponding to an unfiltered sample in the picture. The above filtering process can include: (i) a first filtering step in which one or more first filters are applied to unfiltered samples in the picture, and (ii) a second filtering step in which a second filter is applied to one or more first filtered outputs (e.g., including each first filtered sample in the picture). The filtering process including the first filtering step and the second filtering step can be called a complete two-stage filtering process, and each coefficient of the second filter is applied to a first filtered sample, or a difference (e.g., a difference or a clipped difference) between (i) a first filtered sample and (ii) a corresponding unfiltered sample at the same position in the picture.

[0095] The filtering process (e.g., a complete two-stage filtering process) can be applied to filter samples in any suitable block such as a luma block, a chroma block, an inter-predicted block, or an intra-predicted block.

[0096] In related filtering processes such as the filtering described by Equation (12), samples in the picture are filtered by fixed filters F0 and F1, and then by an adaptive filter F A However, a subset of the coefficients of the adaptive filter F A (e.g., c 20 ~c 21) is only applied to the outputs (e.g., R0 and R1) from the fixed filters F0 and F1, and the other coefficients of the adaptive filter F A (e.g., c0 to c 19 ) are applied to the unfiltered samples, for example, to the difference (e.g., the clipped difference) between each neighboring sample (i.e., the unfiltered neighboring sample) and the current sample (e.g., the unfiltered current sample).

[0097] The complete two-step filtering process can be more beneficial for video / image coding compared to related filtering processes such as the filtering described in Equation (12). In the complete two-step filtering process, the second filter can be applied to the output filtered by the first filter, and the output filtered by the first filter can have a larger SNR than the unfiltered samples. Therefore, the second filter used in the complete two-stage filtering process can include fewer coefficients than the coefficients used in the adaptive filter F A in Equation (12), and can achieve the same or better SNR as the adaptive filter F A in Equation (12). Therefore, the complete two-stage filtering process can achieve the same or better SNR as the adaptive filter F A in Equation (12), and can increase the coding efficiency (e.g., signal only fewer coefficients). For example, when the adaptive filter F A shown in FIGS. 9 to 11 is used in the complete two-step filtering process, the number of coefficients of the adaptive filter F A shown in FIGS. 9 to 11 are 21, 22, and 13 respectively, which is less than the number of coefficients (e.g., 28) of the adaptive filter F A shown in FIG. 8.

[0098] One or more first filters may be non-linear filters that non-linearly modify samples within a picture. The one or more first filters can be predefined. The one or more first filters are known to or can be stored in an image coder or a video coder such as a decoder, an encoder, and / or the like. The one or more first filters can include fixed filters (e.g., offline fixed filters) having coefficients trained offline such as the above-described F0 and F1, and the one or more first filtered outputs can include the above-described R0 and R1.

[0099] In one example, each of the one or more first filters is selected from a set of fixed filters based on each classifier. The one or more first filters can include p filters (e.g., p fixed filters) having coefficients trained for a corresponding classifier C i where p is a positive number and i can range from 0 to p-1. The p filters can include filters from a plurality of fixed filters F0, filters from a plurality of fixed filters F1, …, and filters from a plurality of fixed filters F p-1 For the sake of brevity, the p filters can be referred to as fixed filters F0, F1, …, and F p-1 The coefficients of the plurality of fixed filters F0, the plurality of fixed filters F1, …, and the plurality of fixed filters F p-1 can be predefined and, for example, not signaled. The corresponding classifier C i can be determined using Equation (13). The fixed filter (e.g., F0, F1, …, or F p-1 ) can be selected from each of the plurality of fixed filters F i based on the classifier C i For example, the p filters include two first filters. As shown in Equation (13), the first filter is selected from a plurality of fixed filters F0 based on classifier C0, and the first filter is selected from a plurality of fixed filters F1 based on classifier C1.

[0100] The second filter can be selected from a plurality of adaptive filters (e.g., a plurality of ALF filters) FA based on classifier C A . In one example, classifier C i (e.g., C0, C1, ..., C p-1 ) or the class indicated by C A is assigned to a block in the picture. The block can have any suitable size, such as 2×2, 4×4, etc. In one example, the size of the block is 1×1, and the class indicated by classifier C i or C A is assigned to a sample in the picture.

[0101] In one example, the coefficients of a plurality of adaptive filters (e.g., a plurality of ALF filters) F A are signaled in the header like N A APS when N A is a positive integer (e.g., 8). From the N A APS, APS for a plurality of blocks (e.g., slices or a plurality of blocks in a picture) can be selected (e.g., one APS is selected from 8 APS). Next, classifier C A (e.g., classifier C A for each block or each sample) is determined based on Equation (13) or Equation (15). The second filter for a block (or sample) is selected from the plurality of adaptive filters signaled in the selected APS based on classifier C A .

[0102] In one embodiment, one or more first filters (e.g., p filters) include fixed filter coefficients, are not signaled, and the second filter includes coefficients that are changeable and signaled in a bitstream such as APS.

[0103] In one embodiment, two different types of filters (e.g., cF and aF) are used in a filtering process (e.g., a complete two-stage filtering process). The first type of filter cF can include the above-described p filters. The second type of filter aF can include a plurality of adaptive filters (e.g., a plurality of ALF filters) F A such as adaptive filters. The coefficients of the second type of filter aF can be signaled. Which filter (e.g., which adaptive filter) from the second type of filter aF to use for the current sample can be determined by the classifier C A using the class assigned to the current sample.

[0104] In one embodiment, the first type of filter cF (e.g., p fixed filters) is applied to obtain intermediate filtered outputs (also referred to as first filtered outputs) such as R0, R1, ..., and R p-1 from each of the filters F0, F1, ..., and F p-1 for samples within the picture. Each first filtered output can include an intermediate filtered sample or a first filtered sample corresponding to the sample within the picture (e.g., R0(x,y), R1(x,y), ..., R p-1 (x,y)). (x,y) can represent the sample position within the picture. R(x,y) can indicate the input sample (or unfiltered sample) located at the coordinates (or position) (x,y) before applying ALF filtering (e.g., including the first type of filter cF and the second type of filter aF). The input sample (or unfiltered sample) indicated by R(x,y) can be the input sample to F0, F1, .., or F p-1 . As described above, R(x,y) may be filtered by another filter (e.g., SAO filter) different from the two different types of filters (e.g., cF and aF) used in the complete two-stage filtering process.

[0105] After deriving the intermediate filtered output, the second filter (e.g., a filter of aF) is applied to R0(x,y), R1(x,y),..., or R p-1 (x,y) to derive the filtered current sample R~(x0,y0) as follows:

Equation

[0106] g m,0 represents the difference (e.g., the clipped difference) between the first filtered current sample R m (x0,y0) and the unfiltered current sample R(x0,y0). g k,i is the unfiltered sample R(x,y) and the corresponding first filtered sample R i at the sample position (also called the filtering position) i corresponding to the coefficient c kIt can represent the sum of the differences (e.g., clipped differences) between (x, y). i can refer to one or more sample positions (x, y) at or around the current sample position (x0, y0), where i = 0,..., n - 1. Referring to Equation (16), each coefficient of the second filter is applied to the corresponding first filtered output. For example, each c from coefficient c0 to c n-1 each c i is directly applied to g k,i and thus is indirectly applied to the first filtered output R i (x, y) at the position (x, y) related to c k Each c from coefficient c0 to c n+p-2 each c i is directly applied to g m,0 and thus is indirectly applied to the first filtered output R i at the current sample position (x0, y0) related to c m (x0, y0).

[0107] In one embodiment, k is a predetermined constant parameter such as 0, 1, 2, 3, etc. Classes such as pictures, slices, CTUs, those indicated by classifiers, etc. can each have a k value. In another embodiment, k is signaled in a syntax of an appropriate level such as the CTU level or an appropriate header in a slice header, picture header, SPS, PPS, etc.

[0108] Referring to Equation (16), in one embodiment, the one or more first filters include a first filter (e.g., F0). The first filtered output (e.g., R0) in the one or more first filtered outputs can include the first filtered samples (e.g., R0(x, y)) from the first filter. The plurality of coefficients of the second filter (e.g., c0 to c in Equation (16)) n-1) can be applied to the first filtered output. In one embodiment, the one or more first filters include at least another first filter (e.g., F1), and the one or more first filtered outputs include at least another first filtered output (e.g., R1), and at least one coefficient of a second filter different from the plurality of coefficients (e.g., c in Equation (16)) n ) can be applied to each of the at least another first filtered output.

[0109] In one embodiment, k is 0, and Equation (16) can become the following Equation (17):

Number

[0110] In one example, k is 3, and Equation (16) can become the following Equation (18):

Number

[0111] In one embodiment, three different sets of filters (e.g., F0, F1, F A ) are used. The filter sets F0 and F1 can include fixed filters with coefficients trained for the classifiers C0 and C1. The coefficients of the filter F A can be signaled. In one example, in the encoder, two fixed filters F0 and F1 (e.g., two 13×13 diamond-shaped fixed filters F0 and F1) are applied to derive two intermediate filtered outputs including the samples R0(x, y) and R1(x, y), respectively. Then, the second filter or adaptive filter F A is applied to R0(x, y) and R1(x, y) to derive the current sample filtered as follows. [Number] Here, k ∈ {0, 1} may be a pre-defined parameter or may be signaled in the syntax. For example, k is 0 or 1. The second filter can include filter coefficients or coefficients c i , i = 0,..., n - 1. The filter coefficients c i (i = 0,..., n - 2) can be associated with g k,i . The filter coefficient c i, i = 0, ..., n - 1 can be signaled. The position of a sample within a picture can be represented by (x, y). R(x0, y0) can represent the current sample that has not been filtered. R~(x0, y0) can represent the final filtered current sample after applying the first type of filter (e.g., F0 and F1) and the second filter F A and can represent the current sample that has been filtered finally.

[0112] g 1-k,0 can represent the difference (e.g., the clipped difference) between the first filtered current sample R 1-k (x0, y0) and the current sample R(x0, y0) that has not been filtered. g k,i can represent the sum of the differences (e.g., the clipped differences) between the unfiltered sample R(x, y) and the corresponding first filtered sample R k (x, y) at the sample position i (also called the filtering position). i can refer to the current sample position (x0, y0) or one or more sample positions (x, y) in the vicinity (e.g., surrounding) (x0, y0).

[0113] In one embodiment, when k is 0, Equation (19) can become Equation (20):

Equation

[0114] FIG. 9 shows an example of the second filter F including the filtering taps (or filtering coefficients) c i , i = 0, ..., 20 when n is 21 in Equation (20). The filtering coefficients in F A can include a first coefficient (e.g., c0~c A including) and a second coefficient (e.g., c 19 ). In one example, the first coefficient c 20 (e.g., c0~c i ) is g 19 ), and the second coefficient c 0,iis applied. That is, the first filter coefficient c i is associated with g 0,i , where i ranges from 0 to 19, and the second coefficient c20 is applied to g 1,0 . In various examples, an adaptive filter including filtering taps c0 to c 20 is signaled.

[0115] Referring to FIG. 9, the area supported by the second filter F A can include the position (x0, y0) where the current sample is located and the surrounding area of the current sample. The surrounding area can include samples around the current sample position (x0, y0) (e.g., the periphery). The area supported by the second filter F A can refer to the area within the picture that can be used to obtain the current sample that has been second-filtered using the second filter F A for the sample (e.g., the first filtered sample). The second coefficient (e.g., c 20 ) can be associated with the position (x0, y0) of the current sample, and the first coefficients (e.g., c0 to c 19 ) can be associated with the surrounding area of the current sample.

[0116] In one example, the area supported by the second filter F A (corresponding to coefficients c0 to c 20 ) can have a symmetric diamond shape. The surrounding area of the current sample (corresponding to coefficients c0 to c 19 ) can have a symmetric diamond shape.

[0117] g 1,0 (also denoted as g1(x,y)) can be the difference (e.g., the clipped difference) between the first filtered current sample R1(x0,y0) and the unfiltered current sample R(x0,y0). g 0,ican be the sum of the differences (e.g., clipped differences) between the unfiltered sample R(x, y) and the corresponding first filtered sample R0(x, y) at the sample position (or filtering position) i where i ranges from 0 to 19. i can refer to two sample positions in the surrounding area of the current sample position (x0, y0). For example, the two sample positions correspond to the same coefficient c i corresponds to.

[0118] Referring to FIG. 9, when i ranges from 0 to 19, for example, when i = 11, the two sample positions correspond to the coefficient c i (e.g., c11), and are located at positions related to c 11 such as (x0 - 1, y0 - 1) and (x0 + 1, y0 + 1) related to c i . For example, g 0,11 = g0(x0 - 1, y0 - 1)+ g0(x0 + 1, y0 + 1). g0(x0 - 1, y0 - 1) represents the first difference (e.g., clipped difference) between R(x0 - 1, y0 - 1) and R0(x0 - 1, y0 - 1), and g0(x0 + 1, y0 + 1) can represent the second difference (e.g., clipped difference) between R(x0 + 1, y0 + 1) and R0(x0 + 1, y0 + 1). Therefore, g 0,11 is the sum of the first difference and the second difference.

[0119] The above description is also applicable when i ranges from 0, 1,..., or 19. For example, when i = 16, the two sample positions correspond to the coefficient c 16 and are located at positions (x0 - 4, y0) and (x0 + 4, y0). For example, g 0,16 = g0(x0 - 4, y0)+ g0(x0 + 4, y0). g0(x0 - 4, y0) represents the difference (e.g., clipped difference) between R(x0 - 4, y0) and R0(x0 - 4, y0), and g0(x0 + 4, y0) can represent the second difference (e.g., clipped difference) between R(x0 + 4, y0) and R0(x0 + 4, y0).

[0120] FIG. 10 shows a second filter F including filtering taps (or filtering coefficients) c i , where i = 0, ..., 21, when n is 22 in Equation (20). i The filtering coefficients in F A can include a first coefficient (e.g., including c0~c 20 ) and a second coefficient (e.g., c 21 ). A In one example, the first coefficient c i (e.g., c0~c 20 ) is applied to g 0,i . That is, the first filter coefficient c i is associated with g 0,i , where i ranges from 0 to 20, and the second coefficient c 21 is applied to g 1,0 . A In various examples, an adaptive filter including filtering taps c0~c 21 is signaled. 20 The second coefficient c 21 21 i (e.g., c0~c 20 ) 20 is applied to g 0,i . 0,i That is, the first filter coefficient c i i is associated with g 0,i , 0,i where i ranges from 0 to 20, and the second coefficient c 21 21 is applied to g 1,0 . 1,0 In various examples, an adaptive filter including filtering taps c0~c 21 is signaled. 21

[0121] The region supported by the second filter F in FIG. 10 A can be the same as the region supported by the second filter F in FIG. 9 A . Referring to FIG. 10, the second coefficient (e.g., c 21 ) can be associated with the current sample position (x0, y0), and the first coefficient (e.g., c0~c 20 ) can be associated with the region supported by the second filter F A including the current sample position (x0, y0). A The region supported by the second filter F in FIG. 10 A can be the same as the region supported by the second filter F in FIG. 9 A . Referring to FIG. 10, the second coefficient (e.g., c 21 ) can be associated with the current sample position (x0, y0), and the first coefficient (e.g., c0~c 20 ) can be associated with the region supported by the second filter F A including the current sample position (x0, y0). A In one example, the region supported by the second filter F A can have a symmetric diamond shape. 21 20 A In one example, the region supported by the second filter F A can have a symmetric diamond shape. A

[0122] g 1,0 (or g1(x, y)) can be the difference (e.g., the clipped difference) between the first filtered current sample R1(x0, y0) and the unfiltered current sample R(x0, y0). 1,0 g 1,0 (or g1(x, y)) can be the difference (e.g., the clipped difference) between the first filtered current sample R1(x0, y0) and the unfiltered current sample R(x0, y0). 0,i ​​​​​can be the sum of the differences (e.g., clipped differences) between the unfiltered sample R(x, y) and the corresponding first filtered sample R0(x, y) at the sample position (or filtering position) i where i ranges from 0 to 20. i can refer to the current sample position (x0, y0) or one or two sample positions in the surrounding region of the current sample. For example, the one or two sample positions can have the same coefficient c i corresponding to it.

[0123] Referring to FIG. 10, when i ranges from 0 to 19, for example 11, the two sample positions have the coefficient c i (e.g., c 11 ) corresponding to them, and are located at positions related to c 11 such as (x0 - 1, y0 - 1) and (x0 + 1, y0 + 1) related to c i . For example, as shown in FIG. 9, g 0,11 = g0(x0 - 1, y0)+g0(x0 + 1, y0). This description also applies when i is 0, 1,..., or 19.

[0124] Referring to FIG. 10, when i is 20, the sample position (e.g., the current sample position (x0, y0)) corresponds to the coefficient c 20 . Accordingly, g 0,20 = g0(x0, y0). g0(x0, y0) can represent the difference (e.g., clipped difference) between R(x0, y0) and R0(x0, y0).

[0125] In one embodiment, two different sets of filters (e.g., F0 and F A ) can be used. The sets of filters F0 and F A have been described above. Applying the fixed filter F0 can derive an intermediate filtered output including the sample R0(x, y). Then, applying the second filter or adaptive filter F A to R0(x, y) can derive the currently filtered sample as follows. [Number]

[0126] In one example, the coefficients c0 to c n-1 are applied to g 0,i That is, the filter coefficient c i is associated with g 0,i where i ranges from 0 to n-1. In various examples, an adaptive filter including filtering taps c0 to c n-1 is signaled. g 0,i can represent the sum of the differences (e.g., clipped differences) between the unfiltered sample R(x,y) and the corresponding first filtered sample R0(x,y) at the sample position (filtering position) i. As depicted in FIG. 9 or 10, i can refer to the surrounding region of the current sample position (x0,y0) or one or more sample positions in the current sample. For example, the one or more sample positions correspond to the coefficient c i .

[0127] In one embodiment, two different types of filters (e.g., cF and aF) are used in a filtering process (e.g., a complete two-stage filtering process). The first type of filter cF can include the p filters described above. The second type of filter aF can include adaptive filters such as the plurality of adaptive filters (e.g., a plurality of ALF filters) F A described above. The coefficients of the second type of filter aF can be signaled. Which filter (e.g., which adaptive filter) from the second type of filter aF to use for the current sample can be determined by the class assigned to the current sample using a classifier C A .

[0128] In one embodiment, as described above, the first type of filter cF (e.g., p fixed filters) is applied to obtain R0, R1,..., and R from each of the filters F0, F1,..., and F p-1 for the samples within the picture p-1Obtain an intermediate filtered output (also referred to as the first filtered output), etc. Each first filtered output can include intermediate filtered samples or first filtered samples corresponding to samples in the picture (e.g., R0(x,y), R1(x,y),..., R p-1 (x,y)). (x,y) can represent the sample position in the picture. After deriving the intermediate filtered output, apply a second filter (e.g., the filter of aF) to R0(x,y), R1(x,y),..., or R p-1 (x,y) to derive the filtered current sample R~(x0,y0) as follows:

Equation

[0129] Referring to Equation (22), in one example, one or more coefficients can be applied to the first filtered output R k For example, the n0 coefficient can be applied to the first filtered output R0 via the term g 0,i The n1 coefficient is applied to the first filtered output R1 via the term g 1,i and so on, the n p-1 coefficient is applied to the first filtered output R p-1,i via the term g p-1 and so on.

[0130] In one embodiment, three different sets of filters (e.g., F0, F1, F as described above A ) are used. For example, filter sets F0 and F1 include fixed filters with coefficients trained for classifiers C0 and C1. The coefficients of filter F A can be signaled. In one example, filters F0 and F1 (e.g., two 13×13 diamond-shaped fixed filters F0 and F1) are applied to derive two intermediate filtered outputs each including respective samples R0(x, y) and R1(x, y). Then, a second or adaptive filter F A is applied to R0(x, y) and R1(x, y) to derive the currently filtered sample as follows.

Number

[0131] Figure 11 shows an example of the filtering taps (or filtering coefficients) c i , i = 0, ..., n - 1 included in the second filter F A when m is 7 and n is 14 in Equation (23). The filtering coefficients in F A can include a first coefficient (e.g., including c0~c6) and a second coefficient (e.g., including c7~c 13 included). In one example, the first coefficient c i (e.g., including c0~c6) is applied to g 0,i . That is, the filter coefficient c i is associated with g 0,i and i ranges from 0 to 6, and the second coefficient (e.g., including c7~c 13 included) is applied to g 1,i and i ranges from 7 to 13. In various examples, an adaptive filter including the filtering taps c0~c 13 is signaled.

[0132] The region supported by the second filter F A in Figure 11 can be diamond-shaped. Referring to Figure 11, the first coefficient (e.g., c0~c6) and the second coefficient (e.g., c7~c13 ) can be associated with the area supported by the second filter F including the current sample position (x0, y0) and the surrounding area. A

[0133] g k,i (For example, g 0,i or g 1,i ) can be the sum of the differences (e.g., clipped differences) between the unfiltered sample R(x, y) and the corresponding first filtered sample Rk(x, y) (e.g., the sample position (or filtering position) of R0(x, y) or R1(x, y)) i. i can refer to the current sample position (x0, y0) or one or two sample positions in the surrounding area of the current sample. For example, the one or two sample positions correspond to the same coefficient c i .

[0134] Referring to FIG. 11, the first coefficients include c0 to c6. When i is 0 to 5, for example 1, the two sample positions correspond to the coefficient c i (e.g., c1), and are located at positions related to c such as (x0 - 1, y0 - 1) and (x0 + 1, y0 + 1) related to c1. Therefore, for example, as shown in FIGS. 9 and 10, g i = g0(x0 - 1, y0 - 1) + g0(x0 + 1, y0 + 1). This explanation also applies when i is 0, 1,..., or 5. When i is 4, the two sample positions correspond to the coefficient c 0,1 (e.g., c4), and are located at positions related to c such as (x0 - 2, y0) and (x0 + 2, y0) related to c4. Therefore, as shown in FIGS. 9 and 10, g i = g0(x0 - 2, y0) + g0(x0 + 2, y0). When i is 6, the sample position (e.g., the current sample position (x0, y0)) corresponds to the coefficient c6. Therefore, g i = g0(x0, y0). g0(x0, y0) can indicate the difference (e.g., clipped difference) between R(x0, y0) and R0(x0, y0). 0,4 0,6

[0135] Referring to FIG. 11, the second coefficient includes coefficients c7 to c 13 When i is from 7 to 12, for example 8, the two sample positions correspond to the coefficient c i (for example, c1), and are located at positions related to c i such as (x0 - 1, y0 - 1) and (x0 + 1, y0 + 1) related to c8. For example, g 1,8 = g1(x0 - 1, y0 - 1) + g1(x0 + 1, y0 + 1). This explanation also applies when i is 7, 1,..., or 12. When i is 11, the two sample positions correspond to the coefficient c i (for example, c 11 ), and are located at positions related to c 11 such as (x0 - 2, y0) and (x0 + 2, y0) related to c i . Therefore, as shown in FIGS. 9 and 10, g 1,11 = g1(x0 - 2, y0) + g1(x0 + 2, y0). When i is 13, the sample position (for example, the current sample position (x0, y0)) corresponds to the coefficient c 13 . Therefore, g 1,13 = g1(x0, y0). g1(x0, y0) can represent the difference between R(x0, y0) and R1(x0, y0) (for example, the clipped difference).

[0136] FIG. 12 shows a flowchart illustrating an overview of a process (1200) according to an embodiment of the present disclosure. The process (1200) can be used in a video encoder. In various embodiments, the process (1200) is executed by a processing circuit such as a processing circuit that executes the functions of the video encoder (103), a processing circuit that executes the functions of the video encoder (303), etc. In some embodiments, the process (1200) is implemented by software instructions, and thus when the processing circuit executes the software instructions, the processing circuit executes the process (1200). The process (1200) can perform a filtering process such as the complete two - stage filtering process described above. The process (1200) starts at (S1201) and proceeds to (S1210).

[0137] In (S1210), one or more first filters (e.g., F0, F1, etc.) are applied to samples in the picture to obtain one or more respective first filtered outputs (e.g., R0, R1, etc.) of the samples in the picture. The samples in the picture can include the current sample R0(x0, y0) filtered by a second filter (e.g., the second filter F A ). Each of the one or more first filtered outputs can include a first filtered sample that is a sample in the picture filtered by each first filter in the one or more first filters.

[0138] In one example, the samples in the picture include samples related to the current sample, such as samples within the surrounding region of the current sample. For example, the samples in the surrounding region are related to the coefficients c0 to c in FIG. 9 19 .

[0139] In (S1220), the second filter is applied to the one or more first filtered outputs to obtain a second filtered sample (e.g., R ~ (x0, y0) of equations (16) to (23)) of the current sample. Each coefficient of the second filter can be applied to a corresponding one of the one or more first filtered outputs (e.g., R0, R 1、 , etc.).

[0140] Next, process (1200) proceeds to (S1299) and ends.

[0141] Process (1200) can include a first filtering step as executed in (S1210) and a second filtering step as executed in (S1220).

[0142] The process (1200) can be appropriately adapted. The steps of the process (1200) can be changed and / or omitted. Additional steps can be added. Any appropriate implementation order can be used. In one example, prior to (S1210), samples within a picture are inter- or intra-coded.

[0143] FIG. 13 shows a flowchart illustrating an overview of a process (1300) according to an embodiment of the present disclosure. The process (1300) can be used in a video / image decoder. In various embodiments, the process (1300) is executed by a processing circuit such as a processing circuit that executes the functions of a video decoder (110), a processing circuit that executes the functions of a video decoder (210), and the like. In some embodiments, the process (1300) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1300). The process (1300) can perform a filtering process such as the complete two-stage filtering process described above. The process (1300) starts at (S1301) and proceeds to (S1310).

[0144] (S1310), a bitstream including a picture can be received.

[0145] (S1320), one or more first filters, for example, one or more first fixed filters (e.g., F0, F1, etc.) having constant filter coefficients, are applied to samples within the picture to obtain one or more respective first filtered outputs (e.g., R0, R1, etc.) of the samples within the picture. Samples within the picture can include the current sample R0(x0,y0) filtered by a second filter (e.g., second filter F A ). Each of the one or more first filtered outputs can include a first filtered sample that is a sample within the picture filtered by each first filter within the one or more first filters.

[0146] In one example, the samples in the picture include samples related to the current sample, such as samples within the surrounding area of the current sample. For example, the samples in the surrounding area are related to the coefficients c0 to c of FIG. 9 19 are related to.

[0147] In one example, each of the one or more first filters is from a respective predefined filter set. In one embodiment, the second filter is selected from an adaptive filter set signaled in the bitstream.

[0148] (S1330), the second filter can be applied to the one or more first filtered outputs to obtain the second filtered sample of the current sample (e.g., R~(x0,y0) of equations (16) to (23)). Each coefficient of the second filter can be applied to a corresponding one of the one or more first filtered outputs (e.g., R0, R 1、 etc.). In one embodiment, the second filter is an adaptive filter with changeable coefficients. In one embodiment, the second filter can be applied to the one or more first filtered outputs after step (S1320).

[0149] In one example, the one or more first filters include a first filter (e.g., F0), and the first filtered output (e.g., R0) in the one or more first filtered outputs includes the first filtered sample R0(x,y) from the first filter, and a plurality of coefficients of the second filter (e.g., the coefficients c0 to c of FIG. 9 19 , the coefficients c0 to c of FIG. 10 20 , or the coefficients c0 to c6 of FIG. 11) can be applied to the first filtered output.

[0150] In one example as shown in FIG. 9, a plurality of coefficients of the second filter applied to the first filtered output (e.g., the coefficients c0 to c of FIG. 9 19) Each of them is associated with one or more positions within the region of the picture. In one example, the region includes only the surrounding area of the current sample. In one example, the region includes the surrounding area of the current sample and the current sample. For each of a plurality of coefficients (e.g., c0, c1,..., or c in FIG. 9) 19 ) one or more differences corresponding to one or more positions within the region of the picture can be determined. Referring to FIG. 9, the difference (e.g., g0(x0 - 1, y0 - 1)) for each position (e.g., (x0 - 1, y0 - 1)) among the one or more positions can be between the unfiltered sample R(x0 - 1, y0 - 1) located at each position (e.g., (x0 - 1, y0 - 1)) and the first filtered sample (e.g., R0(x0 - 1, y0 - 1)). Each coefficient (e.g., c 11 ) can be applied to the sum of one or more differences (e.g., g0(x0 - 1, y0 - 1) and g0(x0 + 1, y0 + 1)) (e.g., g 0,11 = g0(x0 - 1, y0 - 1) + g0(x0 + 1, y0 + 1)).

[0151] In one embodiment, the one or more first filters include at least another first filter (e.g., F1), and the one or more first filtered outputs include at least another first filtered output (e.g., R1). At least one coefficient (e.g., c in FIG. 9) of a second filter different from the plurality of coefficients (e.g., coefficients c0 to c in FIG. 9) 19 ) can be applied to at least another first filtered output (e.g., R1). In one embodiment, for each first filtered output (e.g., R1) of at least another first filtered output (e.g., R1), the difference (e.g., the clipped difference) (e.g., g 20 ) between the unfiltered current sample R(x0, y0) at the current sample position and each first filtered output can be obtained. Each coefficient (e.g., c in FIG. 9) of the second filter 1,0 ) is applied to the difference (e.g., g 20 ) (e.g., g 1,0=g1(x0,y0)) can be applied to.

[0152] In one embodiment, the one or more first filters include a first filter denoted as F0 and a first filter F1, and the one or more first filtered outputs include a first filtered output denoted as R0 and another first filtered output R1 that is the output from the first filter F1. The difference between the unfiltered current sample at the current sample position and the first filtered sample of the first filtered output R1 can be obtained, and as described with reference to Equation (19) and FIGS. 9 to 10, the coefficients of a second filter different from the plurality of coefficients can be applied to the difference.

[0153] In an example as shown in FIG. 9, the region has a diamond shape symmetric about the current sample, excluding the current sample, and each of the plurality of coefficients (e.g., coefficients c0 to c 19 ) of the second filter (e.g., c 11 ) is associated with two positions (e.g., (x0 - 1, y0 - 1), (x0 + 1, y0 + 1)) within a region symmetric with respect to the current sample.

[0154] As shown in FIG. 10, in one example, the region includes the current sample, and the plurality of coefficients (e.g., coefficients c0 to c 20 ) include a first coefficient (e.g., c 20 ) and a second coefficient (e.g., c0 to c 19 ) associated with (or applied to) the current sample, and each of the second coefficients (e.g., c 11 ) is associated with two positions (e.g., (x0 - 1, y0 - 1), (x0 + 1, y0 + 1)) within a region symmetric with respect to the current sample.

[0155] In one embodiment, the one or more first filters include a plurality of first filters (e.g., F0, F1,..., and F p-1 ). The one or more first filtered outputs each include a plurality of first filtered outputs from the plurality of first filters (e.g., R0, R1,..., and Rp-1 ) and includes. Each of the second filters includes a plurality of subsets (e.g., p subsets) of coefficients applied to the plurality of first filtered outputs. Each subset of the coefficients of the second filter can be applied to each of the first filtered outputs. In one example, referring to Equation 22, each of the subsets of the coefficients of the second filter is associated with one or more positions in each region of the picture. Each region can include the surrounding region of the current sample, and the samples in the picture can include the samples associated with the current sample (e.g., within the surrounding region). Each region can include or exclude the current sample. For each coefficient of each subset of the coefficients, one or more differences corresponding to one or more positions within each region of the picture can be determined. The difference for each position of the one or more positions can be between the unfiltered sample located at each position and each of the first filtered outputs. Each coefficient can be applied to the sum of the one or more differences.

[0156] In one example, the plurality of first filters includes a first filter F0 and a first filter F1, and the plurality of first filtered outputs include a first filtered output R0 from the first filter F0 and a first filtered output R1 from the first filter F1. In one example, referring to FIG. 11, the regions associated with the first filter F0 and the first filter F1 are the same and are diamond-shaped (e.g., 5×5 diamond shape), and the region includes the current sample and the surrounding region of the current sample. Referring to FIG. 11, the regions associated with the first filter F0 and the first filter F1 are the same as the regions supported by the second filter F A is the same.

[0157] Next, the process proceeds to (S1399) and ends. The process (1300) can be appropriately adapted. The steps of the process (1300) can be changed and / or omitted. Additional steps can be added. Any appropriate implementation order can be used. In one example, the picture is decoded based at least on the second filtered sample of the current sample.

[0158] The various embodiments described in process (1300) can be applied to the process (1200) used in the encoding process.

[0159] Embodiments of the present disclosure may be used separately or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.

[0160] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 14 shows a computer system (1400) suitable for implementing a particular embodiment of the subject matter of the present disclosure.

[0161] The computer software can be processed by mechanisms such as assembly, compilation, linking, etc., to generate code including instructions executable directly or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., and can be coded using any suitable machine code or computer language.

[0162] The instructions can be executed on various computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.

[0163] The components shown in FIG. 14 of the computer system (1400) are exemplary in nature and do not imply any limitation as to the use or functionality of the computer software implementing the embodiments of the present disclosure. Further, the component configuration should not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (1400).

[0164] The computer system (1400) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, sensory input (e.g., keystrokes, swipes, data glove operations), voice input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human interface device can also be used to capture certain media that are not necessarily directly related to conscious input by humans, such as voice (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a digital camera), video (e.g., including 2D video, 3D video, stereoscopic video).

[0165] The input human interface device may include one or more of a keyboard (1401), a mouse (1402), a trackpad (1403), a touch screen (1410), a data glove (not shown), a joystick (1405), a microphone (1406), a scanner (1407), a camera (1408) (only one of which is shown).

[0166] The computer system (1400) may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through sensory output, sound, light, and smell / taste. Such human interface output devices may include sensory output devices (e.g., a touch screen (1410), a data glove (not shown), or a joystick (1405 (sensory feedback, but there may also be sensory feedback devices that do not function as input devices)), a voice output device (e.g., a speaker (1409), headphones (not shown), a visual output device (e.g., a screen (1410), a CRT screen, an LCD screen, a plasma screen, an OLED screen, each with or without touch screen input capabilities, each with or without sensory feedback capabilities, some of which may output two-dimensional visual output or output of three dimensions or more through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown))).

[0167] The computer system (1400) may also include a human-accessible memory device and related media such as an optical medium including a CD / DVD ROM / RW (1420) with a medium (1421) such as a CD / DVD, a thumb drive (1422), a removable hard drive or solid state drive (1423), legacy magnetic media such as tapes and floppy disks (not shown), and devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown).

[0168] One skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.

[0169] The computer system (1400) may also include an interface (1454) to one or more communication networks (1455). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial including CANBus, etc. A particular network generally requires an external network interface attached to a particular general-purpose data port or peripheral device bus (1449) (e.g., the USB port of the computer system (1400)). Others are generally integrated into the core of the computer system (1400) by attachment to a system bus as described later (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using these networks, the computer system (1400) can communicate with other entities. Such communication can be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANbus to a particular CANbus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Particular protocols and protocol stacks can be used for each of the above-described networks and network interfaces.

[0170] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (1440) of the computer system (1400).

[0171] The core (1440) may include one or more central processing units (CPUs) (1441), a graphics processing unit (GPU) (1442), a dedicated programmable processing unit in the form of an FPGA (1443), a hardware accelerator for specific tasks (1444), a graphics adapter (1450), etc. These devices may be connected through a system bus (1448) together with a built-in mass storage device (1447) such as a read-only memory (ROM) (1445), a random access memory (1446), an internal hard drive not accessible to users, an SSD, etc. In some computer systems, the system bus (1448) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attachable directly to the system bus (1448) of the core or through a peripheral device bus (1449). In the example, the screen (1410) can be connected to the graphics adapter (1450). Architectures of peripheral device buses include PCI, USB, etc.

[0172] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can be combined to execute specific instructions capable of generating the aforementioned computer code. The computer code can be stored in the ROM (1445) or the RAM (1446). Temporary data can also be stored in the RAM (1446), while permanent data can be stored, for example, in the built-in mass storage device (1447). Fast storage and reading to / from any of the memory devices can be enabled through the use of a cache memory that can be closely associated with one or more of the CPU (1441), GPU (1442), mass storage device (1447), ROM (1445), RAM (1446), etc.

[0173] A computer-readable medium may have computer code for performing operations implemented by various computers. The medium and the computer code may be specially designed and configured for the purposes of this disclosure or may be of the kind well known and available to those skilled in the field of computer software.

[0174] By way of example and not limitation, a computer system (1400) having an architecture, and specifically a core (1440), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied within one or more tangible computer-readable media. Such computer-readable media can be associated with a specific storage device of the core (1440) with non-transitory characteristics such as an on-core mass storage device (1447) or ROM (1445), and media associated with a user-accessible mass storage device as described above. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1440). The computer-readable media can include one or more memory devices or chips as required by a particular need. The software can cause the core (1440) and specifically the processor (including a CPU, GPU, FPGA, etc.) therein to execute a specific process or a specific portion of a specific process described herein, including the definition of a data structure stored in a RAM (1446) and the modification of the data structure according to the process defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of logic hardwired or other implementation within a circuit (e.g., an accelerator (1444)) operable with or instead of the software to execute a specific process or a specific portion of a specific process described herein. References to software include logic and, where appropriate, vice versa. References to a computer-readable media can, where appropriate, include a circuit (such as an integrated circuit (IC)) for storing software for execution, a circuit for implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0175] The use of "at least one" or "one" in this disclosure is intended to cover any one or combination of the recited elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to cover only A, only B, only C, or any combination thereof. Reference to one of A or B and one of A and B is intended to cover A or B or (A and B). The use of "one" does not exclude any combination of the recited elements where applicable, such as where the elements are not mutually exclusive.

[0176] This disclosure has described several exemplary embodiments, but alternatives, substitutions, and various equivalents exist and are within the scope of this disclosure. It will be apparent to those skilled in the art that, although not explicitly shown or described herein, many systems and methods can be devised that implement the principles of this disclosure and are thus within the spirit and scope of this disclosure.

Claims

1. A method for video decoding in a video coder, comprising: receiving a bitstream including pictures; applying one or more first fixed filters having fixed filter coefficients to samples in the picture, obtaining a first filtered output for each of one or more of the samples in the picture; after applying the one or more first fixed filters, applying one or more second adaptive filters having changeable coefficients to the one or more first filtered outputs, obtaining a second filtered sample of a current sample in the samples; decoding the picture based at least on the second filtered sample of the current sample in the picture; A method comprising the above steps.

2. Each of the one or more first filtered outputs includes a first filtered sample which is the sample in the picture filtered by a respective first fixed filter of the one or more first fixed filters, the one or more second adaptive filters include a second adaptive filter, wherein each coefficient of the second adaptive filter is applied to a corresponding one of the one or more first filtered outputs. The method according to claim 1.

3. The one or more first fixed filters include a first fixed filter, a first filtered output of the one or more first filtered outputs includes the first filtered sample from the first fixed filter, the step of applying the one or more second adaptive filters includes applying a plurality of coefficients of the second adaptive filter to the first filtered output. The method according to claim 2.

4. The samples in the picture include samples related to the current sample, each of the plurality of coefficients of the second adaptive filter applied to the first filtered output is related to one or more positions within a region of the picture, the region including a surrounding region of the current sample, for each of the plurality of coefficients, the step of applying the plurality of coefficients includes Determining one or more differences corresponding to the one or more positions within the region of the picture, wherein the difference for each of the one or more positions is between the unfiltered sample located at each position and the first filtered sample, step; Applying each of the coefficients to the sum of the one or more differences; The method according to claim 3, comprising.

5. The one or more first fixed filters include at least another first fixed filter, The one or more first filtered outputs include at least another first filtered output, The step of applying the one or more second adaptive filters further includes the step of applying at least one coefficient of the second adaptive filter different from the plurality of coefficients to the at least another first filtered output, the method according to claim 3.

6. The step of applying the at least one coefficient, For each first filtered output of the at least another first filtered output, Determining the difference between the unfiltered current sample and the first filtered sample of each of the first filtered outputs at the same position as the current sample; Applying each coefficient of the second adaptive filter to the difference; The method according to claim 5, comprising.

7. The one or more first fixed filters include the first fixed filter indicated as F0 and the first fixed filter F 1 and The one or more first filtered outputs are R 0 the first filtered output shown as R, and another first filtered output R which is the output from the first fixed filter F 1 1 and include​ The step of applying the one or more second adaptive filters further includes, Determining a difference between the unfiltered current sample and the first filtered sample R at the same position as the current sample 1 and a step of determining a difference between the first filtered sample of Applying the coefficients of the second adaptive filter different from the plurality of coefficients to the difference; The method according to claim 4, comprising.

8. The region has a symmetric diamond shape around the current sample excluding the current sample, Each of the plurality of coefficients of the second adaptive filter is associated with two positions within the region symmetric with respect to the current sample, the method according to claim 7.

9. The region includes the current sample, The plurality of coefficients include a first coefficient and a second coefficient related to the current sample, Each of the second coefficients is associated with two positions within the region symmetric with respect to the current sample, the method according to claim 7.

10. The one or more first fixed filters include a plurality of first fixed filters, Each of the one or more first fixed filtered outputs includes a plurality of first filtered outputs from the plurality of first fixed filters. Each of the second adaptive filters includes a plurality of subsets of coefficients applied to the plurality of first filtered outputs. The method of claim 2, wherein applying the one or more second adaptive filters includes applying each subset of coefficients of the second adaptive filters to respective ones of the first filtered outputs. **Claim 11** Each of the subsets of coefficients of the second adaptive filter is associated with one or more positions in each region of the picture, each of the regions including a surrounding region of the current sample, and the samples in the picture including samples associated with the current sample. For each coefficient of each subset of coefficients, applying the subset of coefficients includes: determining the one or more differences corresponding to the one or more positions in each of the regions of the picture, wherein the difference for each position of the one or more positions is between an unfiltered sample located at each position and a respective first filtered output; applying each of the coefficients to the sum of the one or more differences; The method of claim 10, comprising: **Claim 12** The plurality of first fixed filters include first fixed filter F 0 and first fixed filter F 1 and The plurality of first filtered outputs are the first filtered output R 0 from the first fixed filter F 0 and the first filtered output R 1 from the first fixed filter F 1 The method according to claim 11, comprising: **Claim 13** the first fixed filter F 0 and the region related to the first fixed filter F 1 are the same and diamond-shaped, The method of claim 12, wherein the region includes the current sample. **Claim 14** the first fixed filter F 0 and the first fixed filter F 1 has a 13×13 symmetric diamond shape, the method according to claim 8. **Claim 15** The method of claim 2, wherein the number of coefficients of the second adaptive filter is 14, 21, or 22. **Claim 16** Apparatus for video decoding, comprising a processing circuit, the processing circuit being configured to: receive a bitstream including a picture; apply one or more first fixed filters having fixed filter coefficients to samples in the picture to obtain one or more respective first filtered outputs of the samples in the picture; after applying the one or more first fixed filters, apply one or more second adaptive filters having changeable coefficients to the one or more first filtered outputs to obtain a second filtered sample of a current sample in the samples; decode the picture based at least on the second filtered sample of the current sample in the picture. Apparatus, so configured. **Claim 17** Each of the one or more first filtered outputs is a first filtered sample that is a sample within the picture filtered by a respective first fixed filter of the one or more first fixed filters, The one or more second adaptive filters include a second adaptive filter, The apparatus according to claim 16, wherein each coefficient of the second adaptive filter is applied to a corresponding one of the one or more first filtered outputs.

18. The one or more first fixed filters include a first fixed filter, The first filtered output of the one or more first filtered outputs includes the first filtered sample from the first fixed filter, The apparatus according to claim 17, wherein applying the one or more second adaptive filters includes applying a plurality of coefficients of the second adaptive filter to the first filtered output.

19. The samples in the picture include samples related to the current sample, Each of the plurality of coefficients of the second adaptive filter applied to the first filtered output is related to one or more positions within a region of the picture, the region including the surrounding region of the current sample, For each of the plurality of coefficients, the processing circuit, Determine the one or more differences corresponding to the one or more positions within the region of the picture, wherein the difference for each position of the one or more positions is between an unfiltered sample located at each position and the first filtered sample, Applying each of the coefficients to the sum of the one or more differences, The apparatus according to claim 18, configured as such.

20. A computer program, which when executed by a computer for video decoding, causes the computer to execute the method according to any one of claims 1 to 15.

21. A method for video encoding in a video encoder, Generating a bitstream including a picture, Applying one or more first fixed filters having a certain filter coefficient to samples within the picture, obtaining a first filtered output for each of one or more of the samples within the picture, wherein the samples within the picture include current samples filtered by one or more second adaptive filters having changeable coefficients, and each of the one or more first filtered outputs includes a first filtered sample which is the sample within the picture filtered by each first filter within the one or more first filters, the step; Applying the second adaptive filter to the one or more first filtered outputs, obtaining a second filtered sample of the current sample within the sample, and each coefficient of the second adaptive filter is applied to a corresponding one of the one or more first filtered outputs, the step; A method including.