Systems and methods for filtered intra block copy prediction

By applying shift-invariant filtering to generate a filtered predictive block for FIBC mode, the system addresses inefficiencies in video coding, enhancing compression and decoding efficiency in advanced standards like H.266/VVC.

JP2026506519APending Publication Date: 2026-02-25GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025544487
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-13
Filing Date
2024-01-30
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Existing video coding techniques face challenges in efficiently utilizing filtered intra block copy prediction (FIBC) for improved compression and decoding efficiency, particularly in advanced standards like H.266/VVC, due to limitations in shift-invariant filtering and predictive block generation.

Method used

Implementing a system that enables filtered intra block copy (FIBC) mode by applying a shift-invariant weighting filter to a reference block, generating a filtered predictive block, and using it for decoding the current block, enhancing encoding and decoding processes in video codecs.

Benefits of technology

Improves encoding and decoding efficiency by effectively utilizing FIBC mode, leading to enhanced compression and quality in video coding, particularly in advanced standards like H.266/VVC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026506519000001_ABST
    Figure 2026506519000001_ABST
Patent Text Reader

Abstract

According to another aspect of the present disclosure, a system is provided. The system may include a processor and a memory storing instructions. The instructions stored in the memory, when executed by the processor, cause the processor to identify whether a filtered intra block copy (FIBC) mode is enabled for a current block. The instructions, when executed by the processor, cause the processor to generate a filtered predictive block by applying a shift-invariant weighting filter to a reference block in response to the FIBC mode being enabled for the current block. The instructions, when executed by the processor, cause the processor to encode the current block based on the filtered predictive block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to U.S. Provisional Application No. 63 / 442,408, filed January 31, 2023, entitled "IMPROVEMENTS FOR INTRA BLOCK COPY PREDICTION FOR VIDEO CODING," and U.S. Provisional Application No. 63 / 459,234, filed April 13, 2023, entitled "SYSTEM AND METHOD FOR FILTERED INTRA BLOCK COPY PREDICTION FOR VIDEO CODING," the entire contents of which are incorporated herein by reference.

[0002] The present embodiment relates to a video codec. [Background technology]

[0003] Digital video has become mainstream and is widely used in applications such as digital television, video telephony, and teleconferencing. These digital video applications are made possible by advances in computing and communications technology and efficient video coding techniques. Various video coding techniques can be used to compress video data, and one or more video coding standards can be used to encode the video data. Exemplary video coding standards may include, but are not limited to, General Purpose Video Coding (H.266 / VVC), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Expert Group (MPEG) coding, enhanced video coding model (ECM), etc. Summary of the Invention

[0004] According to another aspect of the present disclosure, a system is provided. The system may include a processor and a memory storing instructions. The instructions stored in the memory, when executed by the processor, cause the processor to identify whether a filtered intra block copy (FIBC) mode is enabled for a current block. The instructions stored in the memory, when executed by the processor, cause the processor to generate a filtered predictive block by applying a shift-invariant weighting filter to a reference block in response to the FIBC mode being enabled for the current block. The instructions stored in the memory, when executed by the processor, cause the processor to decode the current block based on the filtered predictive block.

[0005] According to another aspect of the present disclosure, a system is provided. The system may include a processor and a memory storing instructions. The instructions stored in the memory, when executed by the processor, cause the processor to identify whether FIBC mode is enabled for a current block. The instructions stored in the memory, when executed by the processor, cause the processor to generate a filtered predictive block by applying a shift-invariant weighting filter to a reference block in response to FIBC mode being enabled for the current block. The instructions stored in the memory, when executed by the processor, cause the processor to decode the current block based on the filtered predictive block.

[0006] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium is provided that stores instructions that, when executed by a processor, cause the processor to identify whether FIBC mode is enabled for a current block. When executed by the processor, the instructions cause the processor to generate a filtered prediction block by applying a shift-invariant weighting filter to a reference block in response to FIBC mode being enabled for the current block. When executed by the processor, the instructions stored in a memory cause the processor to decode the current block based on the filtered prediction block.

[0007] According to yet another aspect of the present disclosure, there is provided an encoding method by an encoder. The method may include, by a processor, identifying whether FIBC mode is enabled for a current block. In response to FIBC mode being enabled for the current block, the method may include, by the processor, generating a filtered predictive block by applying a shift-invariant weighting filter to a reference block. The method may include, by the processor, encoding the current block based on the filtered predictive block.

[0008] According to yet another aspect of the present disclosure, a system is provided. The system may include a processor and a memory storing instructions. The instructions stored in the memory, when executed by the processor, cause the processor to identify whether FIBC mode is enabled for a current block. The instructions stored in the memory, when executed by the processor, cause the processor to generate a filtered predictive block by applying a shift-invariant weighting filter to a reference block in response to FIBC mode being enabled for the current block. The instructions, when executed by the processor, cause the processor to encode the current block based on the filtered predictive block.

[0009] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium is provided that stores instructions that, when executed by a processor, cause the processor to identify whether FIBC mode is enabled for a current block. When executed by the processor, the instructions cause the processor to generate a filtered prediction block by applying a shift-invariant weighting filter to a reference block in response to FIBC mode being enabled for the current block. When executed by the processor, the instructions cause the processor to encode the current block based on the filtered prediction block.

[0010] These illustrative examples are not intended to limit or define the present disclosure, but rather to provide examples to facilitate understanding of the present disclosure. Additional examples are described in specific embodiments and further explanations are provided. [Brief explanation of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram of an example encoding system according to some embodiments of the present disclosure. [Figure 2] FIG. 1 is a block diagram of an example decoding system according to some embodiments of the present disclosure. [Figure 3] 2 is a detailed block diagram of an example encoder in the encoding system of FIG. 1 according to some embodiments of the present disclosure. [Figure 4] 3 is a detailed block diagram of an example decoder in the decoding system of FIG. 2 according to some embodiments of the present disclosure. [Figure 5] FIG. 2 illustrates an example picture divided into coding tree units (CTUs) according to some embodiments of the present disclosure. [Figure 6] FIG. 2 illustrates an example CTU divided into coding units (CUs) according to some embodiments of the present disclosure. [Figure 7]10A-10C illustrate example visualizations of a current CU block and reconstructed samples that are spatially adjacent and non-adjacent to the current block, according to some embodiments of the present disclosure. [Figure 8] 1A-1C illustrate example visualizations of angular modes of VVC according to some embodiments of the present disclosure. [Figure 9A] FIG. 1 illustrates a representation of spatial geometric partitioning mode (SGPM) signaling according to some embodiments of the present disclosure. [Figure 9B] FIG. 1 illustrates an example template for generating a candidate list according to some embodiments of the present disclosure. [Figure 10] 10A and 10B show visualization examples of an inter-angle mode (parallel mode) parallel to the GPM partition boundary (see (a)), an inter-angle mode (vertical mode) perpendicular to the GPM partition boundary (see (b)), an inter-prediction plane mode (see (c)), and an SGPM with intra-prediction and inter-prediction (see (d)) according to some embodiments of the present disclosure. [Figure 11] FIG. 1 illustrates an adaptive SGPM blending scheme according to some embodiments of the present disclosure. [Figure 12] FIG. 2 illustrates an example visualization of intra block copy (IBC) prediction according to some embodiments of the present disclosure. [Figure 13] FIG. 1 illustrates an intra template matching prediction (IntraTMP) mode for ECM according to some embodiments of the present disclosure. [Figure 14] FIG. 2 illustrates a decoder-side intra mode derivation (DIMD) mode for ECM according to some embodiments of the present disclosure. [Figure 15]FIG. 1 illustrates a template-based intra mode derivation (TIMD) mode for ECM according to some embodiments of the present disclosure. [Figure 16A] FIG. 1 illustrates low-frequency non-separable transform (LFNST) kernels for 4×N and N×4 block sizes in VVC according to some embodiments of the present disclosure. [Figure 16B] FIG. 10 illustrates LFNST kernels for 8×N and N×8 block sizes in VVC according to some embodiments of the present disclosure. [Figure 17] FIG. 1 illustrates various intra-angle prediction modes according to some embodiments of the present disclosure. [Figure 18A] FIG. 10 illustrates LFNST kernels for 4×N and N×4 block sizes in ECM according to some embodiments of the present disclosure. [Figure 18B] FIG. 10 illustrates LFNST kernels for 8×N and N×8 block sizes in ECM according to some embodiments of the present disclosure. [Figure 18C] FIG. 10 illustrates an LFNST kernel for a 16x16 block size in ECM according to some embodiments of the present disclosure. [Figure 19] FIG. 1 illustrates an example visualization of the spatial support of a learned filter according to some embodiments of the present disclosure. [Figure 20] FIG. 1 illustrates an example visualization of a reference block / template for IBC, where the reference block is identified by a block vector (bvx, bvy) according to some embodiments of the present disclosure. [Figure 21] FIG. 1 illustrates a visualization example of a local illumination compensation (LIC) inter-prediction technique according to some embodiments of the present disclosure. [Figure 22] 1 is a flowchart of an exemplary method of video decoding according to some embodiments of the present disclosure. [Figure 23]1 is a flowchart of an example method of video encoding according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0012] The drawings, which are incorporated in and constitute a part of this specification, illustrate examples of the present disclosure and, together with the description, serve to further explain the principles of the present disclosure and to enable those skilled in the art to make and use the disclosure. Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.

[0013] While certain configurations and arrangements are described, it should be understood that this is for illustrative purposes only. A person skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. Clearly, those skilled in the art will appreciate that the present disclosure can be used in a variety of other applications.

[0014] In this specification, descriptions such as "one embodiment," "embodiment," "exemplary embodiment," "some embodiments," and "particular embodiments" indicate that the described embodiment may include a particular feature, configuration, or characteristic, but not all embodiments necessarily include the particular feature, configuration, or characteristic. Furthermore, such expressions do not necessarily refer to the same embodiment. Furthermore, when a particular feature, configuration, or characteristic is described in connection with an embodiment, such feature, configuration, or characteristic can be realized in combination with other embodiments within the knowledge of a person skilled in the art, whether or not explicitly described.

[0015] Generally, terms can be understood, at least in part, from their usage in context. For example, the term "one or more," as used herein, may be used in a singular sense to describe any feature, configuration, or characteristic, or in a plural sense to describe a combination of features, configurations, or characteristics, depending, at least in part, on the context. Similarly, terms such as "one," "an," or "the" are also understood to convey singular or plural usage, depending, at least in part, on the context. Additionally, the term "based on" may be understood as not necessarily intended to convey an exclusive set of factors, and may allow for the existence of additional factors not necessarily explicitly stated, but depending, at least in part, on the context.

[0016] Various aspects of the video encoding system will now be described with reference to various apparatus and methods. These apparatus and methods will be described in the detailed description that follows and illustrated by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as "elements"). These elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends on the particular application and design constraints imposed on the overall system.

[0017] The techniques described herein can be used for various video codec applications. As described herein, video codecs include video encoding and decoding. Video encoding and decoding can be performed on a block-by-block basis. For example, encoding / decoding processes such as transform, quantization, prediction, in-loop filtering, and reconstruction can be performed on a coding block, a transform block, or a prediction block. As described herein, a block to be coded / decoded is referred to as a "current block." For example, the current block can represent a coding block, a transform block, or a prediction block from a current coding / decoding process. Furthermore, it should be understood that, as used in this disclosure, the term "unit" refers to a basic unit for performing a particular coding / decoding process, and the term "block" refers to an array of samples of a predetermined size. Unless otherwise specified, "block" and "unit" can be used interchangeably.

[0018] FIG. 1 is a block diagram of an exemplary encoding system 100 according to some embodiments of the present disclosure. FIG. 2 is a block diagram of an exemplary decoding system 200 according to some embodiments of the present disclosure. Each system 100 or 200 may be applied to or integrated with various systems and devices capable of data processing, such as computers and wireless communication devices. For example, system 100 or 200 may be all or part of a mobile phone, desktop computer, laptop computer, tablet, vehicle computer, gaming console, printer, positioning device, wearable electronic device, smart sensor, virtual reality (VR) device, argument reality (AR) device, or any other suitable electronic device with data processing capabilities. As shown in FIGS. 1 and 2, system 100 or 200 may include a processor 102, memory 104, and interface 106. While these components are shown connected to each other via a bus, other connection types are also acceptable. It should be understood that system 100 or 200 may include any other suitable components for performing the functions described herein.

[0019] The processor 102 may include a microprocessor (such as a graphics processing unit (GPU), image signal processor (ISP), central processing unit (CPU), digital signal processor (DSP), tensor processing unit (TPU), vision processing unit (VPU), neural processing unit (NPU), synergistic processing unit (SPU), or physics processing unit (PPU)), microcontroller unit (MCU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), programmable logic device (PLD), state machine, gated logic, discrete hardware circuitry, and other suitable hardware configured to perform various functions described throughout this disclosure. While only one processor is shown in FIGS. 1 and 2, it should be understood that multiple processors may be included. The processor 102 may be a hardware device having one or more processing cores, and is capable of executing software.Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Software may include computer instructions written in an interpreted language, a compiled language, or machine code. Other techniques for instructing hardware are also permitted under the broad category of software.

[0020] Memory 104 may broadly include both memory (also referred to as primary / system memory) and storage (also referred to as secondary memory). For example, memory 104 may include random-access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferroelectric RAM (FRAM), electrically erasable programmable ROM (EEPROM), compact disc read-only memory (CD-ROM) or other optical disk storage, hard disk drive (HDD) (e.g., magnetic disk storage or other magnetic storage), flash drive, solid-state drive (SSD), or any other medium that can be used to carry or store desired program code in the form of instructions that can be accessed and executed by processor 102. In a broad sense, memory 104 may be embodied by any computer-readable medium (e.g., non-transitory computer-readable medium). Although only one memory is shown in Figures 1 and 2, it should be understood that multiple memories may be included.

[0021] The interface 106 may broadly include a data interface and a communication interface configured to receive and transmit signals in the process of receiving and transmitting information to and from other external network elements. For example, the interface 106 may include input / output (I / O) devices and wired or wireless transceivers. While only one interface is shown in Figures 1 and 2, it should be understood that multiple interfaces may be included.

[0022] The processor 102, memory 104, and interface 106 may be implemented in various forms in system 100 or 200 to perform encoding and decoding functions. In some embodiments, the processor 102, memory 104, and interface 106 of system 100 or 200 are implemented (e.g., integrated) on one or more system-on-chip (SoC). In one example, the processor 102, memory 104, and interface 106 may be integrated on an application processor (AP) SoC, which processes applications including video encoding and decoding applications in an operating system (OS) environment. In another example, the processor 102, memory 104, and interface 106 may be integrated on a processor chip dedicated to video codecs, such as a GPU or ISP chip dedicated to image and video processing within a real-time operating system (RTOS).

[0023] As shown in FIG. 1 , in the encoding system 100, the processor 102 may include one or more modules, such as an encoder 101. While FIG. 1 illustrates the encoder 101 within a single processor 102, it is understood that the encoder 101 may include one or more sub-modules that may be implemented on different processors, either adjacent or remote to each other. The encoder 101 (and any corresponding sub-modules or sub-units) may be a hardware unit (e.g., part of an integrated circuit) of the processor 102, which is designed for use with other components or software units that are implemented by the processor 102 executing at least a portion (e.g., instructions) of a program. The program instructions may be stored in a computer-readable medium, such as the memory 104, and when executed by the processor 102, the instructions may perform a process having one or more functions related to video encoding, such as picture splitting, inter-prediction, intra-prediction, transform, quantization, filtering, entropy coding, etc., as described in more detail below.

[0024] Similarly, as shown in FIG. 2, in decoding system 200, processor 102 may include one or more modules, such as decoder 201. While FIG. 2 illustrates decoder 201 within a single processor 102, it will be understood that decoder 201 may include one or more sub-modules that may be implemented on different processors, either adjacent or remote to each other. Decoder 201 (and any corresponding sub-modules or sub-units) may be a hardware unit (e.g., part of an integrated circuit) of processor 102, designed for use with other components or software units implemented by processor 102 executing at least a portion of a program (e.g., instructions). The program instructions may be stored in a computer-readable medium, such as memory 104, and when executed by processor 102, the instructions may perform processes having one or more functions related to video decoding, such as entropy decoding, inverse quantization, inverse transform, inter-prediction, intra-prediction, filtering, etc., as described in more detail below.

[0025] FIG. 3 is a detailed block diagram of an exemplary encoder 101 in the encoding system 100 of FIG. 1 according to some embodiments of the present disclosure. As shown in FIG. 3, the encoder 101 may include a partitioning module 302, an inter-prediction module 304, an intra-prediction module 306, a transform module 308, a quantization module 310, an inverse quantization module 312, an inverse transform module 314, a filter module 316, a buffer module 318, and an encoding module 320. It should be understood that each element shown in FIG. 3 is shown independently to represent different characteristic functions in a video encoder, and does not imply that each component is formed by a separate hardware or software unit. That is, each element is included as listed for convenience of explanation, and at least two elements may be combined to form one element, or one element may be divided into multiple elements to perform the function. It should also be understood that some elements are not required to perform the functions described in this disclosure, but may be optional elements for improving performance. It should further be understood that these elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends on the particular application and design constraints imposed on the encoder 101.

[0026] The partitioning module 302 may be configured to partition an input picture of video into at least one processing unit. A picture may be a frame of video or a field of video. In some embodiments, a picture includes a luma sample array in monochrome format, or one luma sample and two corresponding chroma sample arrays. In this case, the processing unit may be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The partitioning module 302 may encode the picture by partitioning the picture into multiple combinations of coding units, prediction units, and transform units, and selecting a combination of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function).

[0027] Similar to H.265 / HEVC, H.266 / VVC is a block-based hybrid spatial and temporal predictive coding scheme. As shown in FIG. 5, during encoding, an input picture 500 is first divided into square blocks (CTUs) 502 by a partitioning module 302. For example, the CTUs 502 may be blocks of 128×128 pixels. As shown in FIG. 6, each CTU 502 in the input picture 500 may be divided into one or more CUs 602 by the partitioning module 302 and used for prediction and transform. Unlike H.265 / HEVC, in H.266 / VVC, the CUs 602 may be rectangular or square and may be coded without further division into prediction units or transform units. For example, as shown in FIG. 6, dividing the CTUs 502 into CUs 602 may include quadtree partitioning (shown by solid lines), binary tree partitioning (shown by dashed lines), and ternary tree partitioning (shown by dot-dashed lines). According to some embodiments, each CU 602 may be the same size as its root CTU 502 or may be a subdivision of the root CTU 502 as small as a 4x4 block.

[0028] Referring to FIG. 3 , the inter prediction module 304 may be configured to perform inter prediction on the prediction unit, and the intra prediction module 306 may be configured to perform intra prediction on the prediction unit. It may determine whether to use inter prediction or intra prediction for the prediction unit, and determine specific information (e.g., intra prediction mode, motion vector, reference picture, etc.) based on each prediction method. In this case, the processing unit for performing prediction may be different from the processing unit for determining the prediction method and specific content. For example, the prediction method and prediction mode may be determined in the prediction unit, and transformation may be performed in the transform unit. Residual coefficients of a residual block between the generated prediction block and the original block may be input to the transform module 308. Furthermore, prediction mode information, motion vector information, etc. for prediction may be coded into a bitstream by the coding module 320 along with the residual coefficients or quantization level. It should be understood that in certain coding modes, the original block may be coded as is without generating a prediction block via the prediction module 304 or 306. It should also be understood that in certain coding modes, prediction, transformation, and / or quantization may be skipped.

[0029] In some embodiments, the inter prediction module 304 may predict a prediction unit based on information of at least one picture before or after the current picture, and in some cases, may predict a prediction unit based on information of a subregion coded in the current picture. The inter prediction module 304 may include sub-modules such as a reference picture interpolation module, a motion prediction module, and a motion compensation module (not shown). For example, the reference picture interpolation module may receive reference picture information from the buffer module 318 and generate integer-pixel pixel information from the reference picture. For luma pixels, a discrete cosine transform (DCT)-based 8-tap interpolation filter with varying filter coefficients may be used to generate integer-pixel pixel information in 1 / 4-pixel increments. For chroma signals, a DCT-based 4-tap interpolation filter with varying filter coefficients may generate integer-pixel pixel information in 1 / 8-pixel increments. The motion prediction module may perform motion prediction based on a reference picture interpolated by a reference picture interpolator. Various methods can be used to calculate motion vectors, such as a full search-based block matching algorithm (FBMA), a three-step search (TSS), and a new three-step search algorithm (NTS). The motion vector may have a motion vector value in units of 1 / 2, 1 / 4, 1 / 16, or integer pixels based on interpolated pixels. The motion prediction module may predict the current prediction unit by changing the motion prediction method. Various methods can be used as motion prediction methods, such as a skip method, a merge method, an Advanced Motion Vector Prediction (AMVP) method, and an intra-block copy method.

[0030] Continuing to refer to FIG. 3 , in some embodiments, the intra prediction module 306 may generate a prediction unit based on information about reference pixels surrounding the current block (pixel information within the current picture). The reference pixels may be located on a reference line that is not adjacent to the current block. If a neighboring block of the current prediction unit is a block on which inter prediction has been performed and the reference pixels are pixels on which inter prediction has been performed, the reference pixels included in the neighboring block on which inter prediction has been performed may be used instead of the reference pixel information of the neighboring block on which intra prediction has been performed. That is, if reference pixels are unavailable, at least one of the available reference pixels may be used instead of the unavailable reference pixel information. In intra prediction, prediction modes may include an angular prediction mode that uses reference pixel information depending on the prediction direction, and a non-angular prediction mode that does not use direction information when performing prediction. The mode for predicting luma information may be different from the mode for predicting chroma information, and the chroma information may be predicted using intra prediction mode information for predicting luma information or predicted luma signal information. When performing intra prediction, if the size of the prediction unit is the same as the size of the transform unit, intra prediction may be performed on the prediction unit based on the left pixel, the top left pixel, and the top pixel of the prediction unit. However, when performing intra prediction, if the size of the prediction unit is different from the size of the transform unit, intra prediction may be performed using reference pixels based on the transform unit.

[0031] The intra prediction method may generate a predicted block after applying an adaptive intra smoothing (AIS) filter to reference pixels based on the prediction mode. Various types of AIS filters may be applied to the reference pixels. To perform the intra prediction method, the intra prediction mode of a current prediction unit may be predicted from the intra prediction mode of a prediction unit existing in the vicinity of the current prediction unit. When predicting the prediction mode of the current prediction unit using mode information predicted from a neighboring prediction unit, if the intra prediction mode of the current prediction unit is the same as that of the neighboring prediction unit, predetermined flag information may be used to transmit information indicating that the prediction mode of the current prediction unit is the same as that of the neighboring prediction unit. If the prediction mode of the current prediction unit and the prediction mode of the neighboring prediction unit are different from each other, the prediction mode information of the current block may be coded using additional flag information.

[0032] 3, a residual module may be generated, which includes a prediction module that performs prediction based on the prediction unit generated by prediction module 304 or prediction module 306 and residual coefficient information (also referred to herein as "residual"), which is a difference value between the prediction unit and the original block. The generated residual block may be input to transform module 308. Additional details of the residual and transform for video coding are provided below.

[0033] Hybrid video coding systems first exploit redundancy in a video signal by applying inter- or intra-prediction tools to each CU. The difference between the original samples of a CU and its predicted block is commonly referred to as the residual. After prediction, the residual may still be highly spatially correlated. While conditional entropy coding can capture some of the spatial dependency between neighboring samples, it is computationally impractical to create an entropy coding statistical model that can fully exploit the spatial correlation of the residual. In contrast, transform coding is a practical and effective method for spatially decorrelating the residual.

[0034] For example, the transform module 308 may transform the residual using an integerized version of a two-dimensional discrete cosine transform (DCT), which may be applied separably in the horizontal and vertical directions. For an M×N block of residual samples (where M is the width of the block and N is the height of the block), the transform module 308 may apply an M×M DCT to each row to generate intermediate transform coefficients, and then apply an N×N DCT to each column of the intermediate transform coefficients to obtain transform coefficients.

[0035] For intra-coded CUs (also referred to herein as "intra CUs"), spatially neighboring reconstructed samples are used to predict the current block, and the intra prediction mode is written once for the entire CU. Each CU consists of one or more co-located coding blocks (CBs) corresponding to the color components of a video sequence. For example, consumer video commonly employs the 4:2:0 chrominance format, in which each CU consists of a luma CB and two chroma CBs with a quarter of the luma CB's samples. Intra prediction and transform coding are performed at the prediction block (PB) and transform block (TB) levels, respectively. Except for the intra subpartition (ISP) mode and implicit partitioning, each CB consists of a single TB. For luma CBs, the maximum edge length of a TB is 64 and the minimum edge length is 4. Furthermore, a luma TB is further specified as a W × H rectangular block with width W and height H, where W, H ∈ {4, 8, 16, 32, 64}. For chromaticity CB, the maximum side length of TB is 32, and chromaticity TB is a rectangular block of size W × H, with width W and height H, where W, H ∈ {2, 4, 8, 16, 32}, but blocks of shape 2 × H and 4 × 2 are excluded to address memory architecture and throughput requirements.

[0036] 7 illustrates an example visualization 700 of a current CU block 702 and reconstructed samples that are spatially adjacent and non-adjacent to the current block, according to some aspects of the present disclosure. In FIG. 7, the numbers 0, 1, 2, ... indicate pixel line indices associated with the current CU block 702.

[0037] In VVC, intra-predicted samples for a current block are generated using reference samples obtained from reconstructed samples of neighboring blocks. For a W×H block, the reference samples are spatially adjacent to the current block and consist of a vertical line of 2 H reconstructed samples located to the left of the block and extending downward, a reconstructed sample at the top left, and a horizontal line of 2 W reconstructed samples located above the current block and extending to the right. This "L"-shaped set of samples may be referred to as a "reference line" in this disclosure. The reference line directly adjacent to the current CU block 702 is shown in Figure 7 as the line with index 0.

[0038] Similar to AVC and HEVC, VVC also supports angular intra prediction modes. Angular intra prediction is a directional intra prediction method. Compared to HEVC, VVC's angular intra prediction is modified by improving prediction accuracy and adapting to a new segmentation framework. The former is achieved by increasing the number of angular prediction directions and more accurate interpolation filters, while the latter is achieved by introducing a wide-angle intra prediction mode. In VVC, the number of directional modes available for a given block increases to 65 directions from 33 directions in HEVC. Figure 8 shows VVC angular modes 800.

[0039] Directions with even indices between 2 and 66 correspond to the directions of angle modes supported by HEVC. For square blocks, the top and left sides of the block are assigned the same number of angle modes. On the other hand, rectangular intra blocks, which do not exist in HEVC, are the central part of the VVC partitioning scheme, and additional intra prediction directions are assigned to the long sides of the block. The additional modes assigned along the long sides correspond to prediction directions with angles greater than 45° relative to the horizontal or vertical modes, and are therefore called wide-angle intra prediction (WAIP) modes. As shown in Figure 8, the WAIP mode for a given mode index is defined by mapping the original directional mode to a mode with an index offset equal to 1 and an inverse direction. For a given rectangular block, the aspect ratio, i.e., the ratio of width to height, is used to determine which angle mode replaces the corresponding wide-angle mode.

[0040] For square blocks in VVC, each pair of horizontally or vertically adjacent predicted samples is predicted from a pair of adjacent reference samples. Conversely, WAIP extends the angular range of directional prediction beyond 45°, so that for coding blocks predicted using WAIP mode, adjacent predicted samples can be predicted from non-adjacent reference samples.

[0041] In addition to the directly adjacent rows of neighboring samples, either of the two non-adjacent reference lines (line 1 and line 2) shown in Figure 7 may contain input samples for intra prediction in VVC. For ECM, more non-adjacent reference lines can be used. The use of neighboring and non-adjacent reference samples is called multiple reference line (MRL) prediction.

[0042] The intra modes available for MRL are DC mode and angular prediction mode. However, for a given block, not all of these modes can be combined with MRL. MRL modes are always combined with modes in the VVC Most Probable Mode (MPM) list. This combination means that when non-adjacent reference lines are used, the intra prediction mode can be one of multiple MPMs. This design of MPM-based MRL prediction modes is motivated by the observation that non-adjacent reference lines are primarily advantageous for texture patterns with sharp and highly directional edges. In these cases, MPM is selected more frequently because there is usually a strong correlation between the texture patterns of neighboring blocks and the current block. On the other hand, the selection of a non-MPM for intra prediction indicates that edges are unevenly distributed in neighboring blocks, and therefore, MRL prediction modes are expected to be less useful in such situations. Furthermore, it has been observed that MRL does not provide additional coding gain when the intra prediction mode is a planar mode. This is because this mode is typically used for smooth regions. Therefore, MRL always excludes planar modes, which are one of multiple MPMs. The angle or DC prediction process in MRL is very similar to that for directly adjacent reference lines. However, for non-integer slope angle modes, a DCT-based interpolation filter (DCTIF) is always used. This design choice is consistent with experimental results and empirical observations: MRL is advantageous for sharp and highly directional edges, and DCTIF is more appropriate because it preserves more high frequencies than some other filters.

[0043] From a hardware design perspective, applying multiple reference lines, as proposed in the initial method, requires the additional cost of line buffers to hold the additional reference lines. In a typical hardware design, line buffers are part of the on-chip memory architecture for image and video coding, and minimizing their on-chip area is crucial. To solve this problem, the MRL is disabled and not written to the coding unit attached to the top boundary of the CTU. In this way, the additional buffers to hold non-adjacent reference lines are separated by 128, where 128 is the width of the maximum unit size.

[0044] In some known methods, intra prediction methods have been proposed to improve the accuracy of intra prediction. More specifically, if the current block is a luminance block, is coded in a non-integer tilt angle mode rather than an ISP mode, and the block size (width x height) is greater than 16, two prediction blocks generated from two different reference lines are "fused", and the prediction fusion is calculated as a weighted sum of the two prediction blocks. More specifically, the current signaling method in the bitstream is used to signal the index i(line i ), and the predicted block generated from that reference line using the selected intra prediction mode is i ), where p() denotes the operation of generating a prediction block from a reference line having a given intra prediction mode. In a known method, the reference line line i+1 is implicitly selected as the second reference line. That is, the second reference line is at an index position farther away from the current block than the first reference line. Similarly, the predicted block generated from the second reference line is p(line i+1 ) The weighted sum of the two prediction blocks is obtained by the following Equation 1 and is used as the prediction of the current block according to Equation 1:

[0045] [Formula 1] p fusion =w0*p(line i )+w1*p(line i+1 ) where p fusion represents the fusion prediction, and w0 and w1 are two weighting factors, which are set to 3 / 4 and 1 / 4, respectively, in the experiments.

[0046] Some known methods propose a spatial geometric partitioning mode (SGPM), which can partition a CU into two parts for which different intra prediction modes are available. The new mode is conceptually similar to the geometric partitioning mode (GPM) applied to inter prediction in VVC. However, due to the large number of possible combinations of partitioning and intra prediction modes, SGPM uses a different signaling mechanism. To more efficiently represent the required partitioning and prediction information in the bitstream, a candidate list is adopted, and only candidate indices are written to the bitstream. Each candidate in the list can derive a combination of one partition mode and two intra prediction modes, as shown in the representation of SGPM signaling 900 in FIG. 9A.

[0047] A template is used to generate the candidate list. An example of a template 901 is shown in FIG. 9B, where the template width is set to 4. In the current version of SGPM adopted by ECM, the template width is set to 1. For each possible combination of one partition mode and two intra-prediction modes, a prediction for the template is generated, where partition weights are extended to the template. These combinations are arranged in ascending order of the sum of absolute transformed difference (SATD) cost between the template prediction and reconstruction. The length of the candidate list is set equal to 16, and these candidates are considered the most probable SGPM combinations for the current block. Both the encoder and decoder use the template to construct the same candidate list. To reduce the complexity of candidate list construction, both the number of possible partition modes and the number of possible intra-prediction modes (IPMs) are limited. For example, in the current version of SGPM adopted by ECM, the number of possible partition modes is limited to a predetermined set of 26 partitions covering various partition directions and positions.

[0048] Each of the two IPM candidate lists corresponding to the two SGPM partitions is constructed by adding available IPM candidates and then trimming them, if necessary, to a predetermined limit of three candidates. Some IPM candidates are inherited from the intra-inter GPM mode already employed in ECM. Figure 10 illustrates a visualization example 1000 of an inter-angle mode parallel to the GPM partition boundary (parallel mode) (see (a)), an inter-angle mode perpendicular to the GPM partition boundary (vertical mode) (see (b)), an inter-prediction planar mode (see (c)), and an SGPM with intra- and inter-prediction (see (d)) according to some embodiments of the present disclosure. These modes may be added as available IPM candidates for the SGPM.

[0049] Furthermore, template-based intra-mode derivation (TIMD) can be used to derive intra-prediction modes as available IPM candidates for SGPM, for example, using only limited horizontal and vertical neighbors (using a top or left template) to derive TIMD intra-prediction modes for IPM to form SGPM.

[0050] For some CU block sizes, SGPM may be implicitly disabled. The range of block sizes for which SGPM can be used (e.g., a CU-level flag can be written to indicate whether SGPM is used) was originally inherited from the intra-inter GPM mode. In the current version of SGPM adopted by ECM, the range of block sizes is further extended to blocks smaller than 4x8, 8x4, 4x16, and 16x4. In short, the block sizes for which SGPM can be used can be described by the following rules: 4 ≤ width ≤ 64, 4 ≤ height ≤ 64, width < height x 8, height < width x 8, width x height ≥ 32.

[0051] To better predict pixels at the boundary between two prediction portions, an adaptive SGPM blending scheme 1100 can be used, as shown in Figure 11, in which a weighted average of the two prediction portions is used in the transition region surrounding the SGPM partition. The width of the transition region is called the blend width, and the blend width is adaptively determined based on the block size. No writes are required for adaptive blending.

[0052] 11, the blend width specified for the GPM tools of VVC and ECM is τ. Then, the adaptive SGPM blend width can be determined based on the CU block width and height as follows:

[0053] If min(width, height)==4, choose 1 / 2τ.

[0054] Otherwise, choose τ if min(width, height)==8.

[0055] Otherwise, if min(width, height)==16, choose 2τ.

[0056] Otherwise, if min(width, height)==32, then choose 4τ.

[0057] Otherwise, choose 8τ.

[0058] FIG. 12 illustrates a visualization example of IBC prediction 1200 according to some embodiments of the present disclosure. When a CU predicts using intra block copy mode, it writes a block vector (BV) to indicate which block in the same picture is to be copied for use as a predictor for the current block. Writing the block vector can be performed by writing a block vector difference (BVD) to the bitstream, and the block vector can be determined by adding the BVD to the block vector predictor. Alternatively, if a block vector from a previous CU exactly matches the current block vector, it can be written using a merge flag. Regardless of the writing mechanism, the block vector points to a location in the same picture to indicate a sample block equal to the current CU, which is then used as the prediction block for the current CU. Some restrictions may apply to the block vector, since it must point to a location in the current picture that has already been decoded before the current CU. The illustration in FIG. 12 summarizes the IBC concept in HEVC and VVC, where each tiled square represents a coding tree unit (CTU). Gray-shaded regions indicate regions that have already been coded, while white-shaded regions indicate regions to be coded. HEVC's IBC generally allows BVs to point to any block included in the use of the gray-shaded regions. This freedom is partially restricted when the sps_entropy_coding_sync_enabled_flag is written to the bitstream, in order to support the wavefront parallel processing (WPP) feature. In this case, an IBC block vector cannot point to a region within two or more CTUs to the right of the current CTU immediately above the CTU row, as shown by the scribed CTUs in Figure 12. VVC's IBC is obviously even more restricted, only allowing block vectors to use CTUs to the left of the current CTU, indicated by the dashed box, as reference regions.Current IBC tools in ECM extend the search scope to VVC.

[0059] FIG. 13 illustrates an intra-template mapping prediction (IntraTMP) mode 1300 for ECM according to some embodiments of the present disclosure. Referring to FIG. 13, intra-template mapping prediction (IntraTMP) is similar to IBC because the current CU is also predicted by a sample block from the current picture. However, unlike IBC, no block vectors are written to the bitstream. Instead, the decoder 201 compares L-shaped or other shaped templates of reconstructed samples adjacent to the current CU with L-shaped templates of candidate predictors within a predetermined search area. The IntraTMP predictor block is determined by finding the best candidate template that matches the current CU template. The best match is determined by finding the template that minimizes the sum of absolute difference (SAD) or sum of absolute transformed difference (SATD), or by comparing hash values ​​between the templates. The search algorithm over the search space can be exhaustive (e.g., by scanning the template in the search space using a shift in sample resolution) or fast (e.g., by first performing a coarse search and then performing a local refinement search around the best match from the coarse search). Regardless, the search algorithm can be performed similarly in both the encoder 101 and the decoder 201, and the encoder 101 and the decoder 201 can implicitly know the IntraTMP predictor without writing it to the bitstream. In Figure 13, the current CU template 1302 and the best matching template 1304 are shown with diagonal shading, and the other templates 1306 in the search space are shown with dashed shading.

[0060] FIG. 14 is a diagram illustrating a decoder-side intra-mode derivation (DIMD) mode 1400 for ECM according to some embodiments of the present disclosure.

[0061] Referring to FIG. 14, in DIMD, an intra-prediction mode (or multiple intra-prediction modes) is implicitly derived from an L-shaped template 1404 (hereinafter referred to as "template 1404") of reconstructed samples neighboring the current CU 1402. The template 1404 has a size of three samples wide. The decoder 201 moves a 3x3 gradient analysis window 1406 over the template 1404. At each position, a Sobel filter is applied to calculate the local gradient. The set of 3x3 samples at a position in the template 1404 is called T k Assuming that , the Sobel filter can be written as Equation 2.

[0062]

number

[0063] Next, the horizontal gradient G k,x and vertical gradient G k,y are estimated by taking the dot products shown in Equation 3 and Equation 4 below, respectively.

[0064] [Formula 3] G k,x =T k M x [Formula 4] G k,y =T k M y

[0065] Local gradient magnitude G k and the local gradient angle θ k can be estimated according to Equations 5 and 6, respectively.

[0066]

number

[0067] Local gradient angle θ k is the intra-angle prediction direction IPM kFor example, an angle of 0 degrees corresponds to horizontal intra prediction mode 18. In practice, IPM k is generated by the decoder 201 using a fast implementation method such as a lookup table. k,x and G k,y At the start of the DIMD method, we initialize an empty histogram H (with zeros in each entry) with a size equal to the number of intra-prediction modes. When the DIMD method performs gradient analysis on each local window, the histogram H is updated according to Equation 7.

[0068] [Formula 7] H[IPM k ]+=G k

[0069] Thus, for an intra-prediction mode, each local gradient analyzes a "vote." At the end of the DIMD method, the intra-prediction mode with the highest count in H can be selected as the single representative intra-prediction mode for the current CU. Alternatively, multiple intra-prediction modes can be obtained in order of highest count from H.

[0070] 15 is a diagram illustrating a template-based intra-mode derivation (TIMD) mode 1500 for ECM according to some embodiments of the present disclosure. Referring to FIG. 15, in TIMD, an intra-prediction mode (or multiple intra-prediction modes) is implicitly derived from an upper and left template 1504 of reconstructed samples neighboring a current CU 1502.

[0071] A set of candidate intra-prediction modes is searched from a most probable mode (MPM) list, which is constructed from the intra-prediction modes used by neighboring CUs. Then, for each candidate intra-prediction mode, a prediction of the template 1504 is generated from the template reference sample 1506 using an intra-angle prediction method. The candidate intra-prediction mode that generates a template predictor that best matches the template 1504 is selected as the TIMD intra-prediction mode. The best match can be determined by finding the predictor that minimizes the sum of absolute differences (SAD) or the sum of absolute transform differences (SATD), or by comparing hash values ​​between the predictor and the template. Alternatively, multiple intra-prediction modes can be obtained in order of increasing SAD / SATD.

[0072] Matrix weighted intra prediction (MIP), IBC, and IntraTMP methods can be effective intra prediction modes in ECM. Greater gains can be achieved by combining MIP with a non-separable primary transform (NSPT), IBC with LFNST and NSPT, and IntraTMP with LFNST and NSPT, where the intra prediction modes are derived in the manner described in connection with the solution in FIG. 4 below.

[0073] The transform module 308 may transform the video signal in the residual block from the pixel domain to a transform domain (e.g., frequency domain, depending on the transform method). It should be understood that in some examples, the transform module 308 may be skipped and the video signal may not be transformed to a transform domain.

[0074] The quantization module 310 is configured to quantize coefficients at each position in the coding block to generate a quantization level for that position. The current block may be a residual block. That is, the quantization module 310 may perform a quantization process on each residual block. The residual block may include N×M positions (samples), each associated with a transformed or untransformed video signal / data (such as luma and / or chroma information), where N and M are positive integers. In this disclosure, before quantization, the transformed or untransformed video signal at a particular position is referred to herein as a "coefficient." After quantization, the quantized value of the coefficient is referred to herein as a "quantization level" or "level."

[0075] Quantization can be used to reduce the dynamic range of a transformed or untransformed video signal, thereby representing the video signal with fewer bits. Quantization typically involves division by a quantization step length followed by rounding, while dequantization (called inverse quantization) involves multiplication by the quantization step length. The quantization step length can be indicated by a quantization parameter (QP). This type of quantization process is called scalar quantization. Quantization of all coefficients within a coding block can be completed independently, and some existing video compression standards, such as H.264 / AVC and H.265 / HEVC, use this type of quantization method. The QP in quantization can affect the picture bitrate for encoding / decoding video. For example, a higher QP results in a lower bitrate, while a lower QP results in a higher bitrate.

[0076] For an N×M coding block, a particular coding scan order can convert the two-dimensional (2D) coefficients of the block into a one-dimensional (1D) order that can be used for quantizing and encoding the coefficients. Generally, the coding scan starts from the upper-left corner of the coding block and stops at the lower-right corner of the coding block or the last non-zero coefficient / level in the lower-right direction. It should be understood that the coding scan order can include any suitable order, such as a zigzag scan order, a vertical (column) scan order, a horizontal (row) scan order, a diagonal scan order, or any combination thereof. The quantization of coefficients in the coding block can utilize coding scan order information. For example, it may depend on the state of the previous quantization level along the coding scan order. To further improve coding efficiency, the quantization module 310 can use multiple quantizers (e.g., two scalar quantizers). Which quantizer is used to quantize a current coefficient can depend on previous information about the current coefficient in the coding scan order. This quantization process is called dependent quantization.

[0077] Referring to FIG. 3, the encoding module 320 is configured to encode the quantization level for each position within the coding block into a bitstream. In some embodiments, the encoding module 320 may perform entropy encoding on the coding block. Entropy encoding converts each quantization level into a corresponding binary representation (e.g., a binary sequence) using various binarization methods (e.g., Golomb-Rice binarization). The binary representation can then be further compressed using an entropy encoding algorithm. The compressed data may be added to the bitstream. In addition to the quantization levels, the encoding module 320 may encode various other information, such as block type information of the coding unit, prediction mode information, partition unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information input from the prediction modules 304 and 306. In some embodiments, the encoding module 320 may perform residual symbolization on the coding block to convert the quantization levels into a bitstream. For example, after quantization, there may be N×M quantization levels for an N×M block. These N×M levels may be zero or non-zero values. If the non-zero levels are non-binary, we can further binarize the non-zero levels into a binary sequence using a combination of TR (Truncated Rice) and finite EGk binarization.

[0078] Non-binary syntax elements can be mapped to binary code words. The bijective mapping between symbols and code words (usually using simple structured codes) is called binarization. Binary arithmetic coding can be used to encode binary symbols (also called sequences) for both binary syntax elements and code words for non-binary data. The core coding engine of context-adaptive binary arithmetic coding (CABAC) can support two modes of operation: a context coding mode, which encodes sequences using an adaptive probability model, and a less complex bypass mode, which uses a fixed probability of 1 / 2. The adaptive probability model is also called a context, and the assignment of a probability model to each sequence is called context modeling.

[0079] 3, the inverse quantization module 312 is configured to inverse quantize the quantization levels, and the inverse transform module 314 is configured to inverse transform the coefficients transformed by the transform module 308. The reconstructed residual blocks generated by the inverse quantization module 312 and the inverse transform module 314 may be combined with the prediction units predicted by the prediction modules 304 or 306 to generate reconstructed blocks.

[0080] The filter module 316 may include at least one of a deblocking filter, a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF). The deblocking filter can remove block artifacts caused by boundaries between blocks in the reconstructed picture. The SAO module can calibrate the offset of the deblocked video relative to the original video on a pixel-by-pixel basis. The ALF can be performed based on values ​​obtained by comparing the reconstructed and filtered video with the original video. The buffer module 318 can be configured to store the reconstructed blocks or pictures calculated by the filter module 316 and can provide the reconstructed and stored blocks or pictures to the inter prediction module 304 when inter prediction is performed.

[0081] FIG. 4 is a detailed block diagram of an exemplary decoder 201 in the decoding system 200 of FIG. 2 according to some embodiments of the present disclosure. As shown in FIG. 4, the decoder 201 may include a decoding module 402, an inverse quantization module 404, an inverse transform module 406, an inter-prediction module 408, an intra-prediction module 410, a filter module 412, and a buffer module 414. It should be understood that each element shown in FIG. 4 is shown independently to represent different characteristic functions in a video decoder, and does not imply that each component is formed by a separate hardware or software unit. That is, each element is listed as an element for convenience of explanation, and at least two elements among multiple elements may be combined to form one element, or one element may be divided into multiple elements to perform the function. It should also be understood that some elements are not required to perform the functions described in this disclosure, but may be optional elements for improving performance. It should also be understood that these elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends on the particular application and design constraints imposed on the decoder 201 .

[0082] When a video bitstream is input from a video encoder (e.g., encoder 101), the input bitstream may be decoded by decoder 201 in a process that is the reverse of that of the video encoder. Therefore, for ease of explanation, some of the decoding details described above with respect to encoding may be omitted. The decoding module 402 may be configured to decode the bitstream to obtain various information about the quantization level of each position within the coding block encoded in the bitstream. In some embodiments, the decoding module 402 may perform entropy decoding (decompression) corresponding to the entropy encoding (compression) performed by the encoder, such as variable-length coding (VLC), context-adaptive variable-length coding (CAVLC), CABAC, syntax-based binary arithmetic coding (SBAC), or PIPE coding, to obtain a binary representation (e.g., a binary sequence). The decoding module 402 may convert the binary representation into quantization levels using Golomb-Rice binarization (e.g., including EGk binarization and binarization combining TR and finite EGk). In addition to the quantization levels of positions within a transform unit, the decoding module 402 can decode various other information, such as parameters used for Golomb-Rice binarization (e.g., Rice parameters), block type information of the coding unit, prediction mode information, partition unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. During the decoding process, the decoding module 402 can perform reordering on the bitstream to reconstruct and reorder data from 1D order into 2D reordered blocks according to a reverse scan method based on the coding scan order used by the encoder.

[0083] The inverse quantization module 404 may be configured to inverse quantize the quantization level of each position of the coding block (e.g., a 2D reconstruction block) to obtain a coefficient for each position. In some embodiments, the inverse quantization module 404 may also perform dependent inverse quantization based on quantization parameters provided by the encoder, which include information related to the quantizers used in the dependent quantization, such as the quantization step length used by each quantizer.

[0084] The inverse transform module 406 may be configured to perform an inverse transform (e.g., an inverse DCT, an inverse discrete sine transform (DST), and an inverse KLT) for each of the DCT, DCT and KLT, LFNST, and / or NSPT performed by the encoder to convert data back from the transform domain (e.g., coefficients) to the pixel domain (e.g., luma and / or chroma information). In some embodiments, the inverse transform module 406 may selectively perform the transform operations (e.g., DCT, DST, KLT, LFNST, NSPT) according to multiple pieces of information, such as a prediction method, a size of the current block, a prediction direction, etc.

[0085] For example, a separable transform applies a one-dimensional transformation in the horizontal and vertical directions, respectively, while a two-dimensional non-separable transform applies it directly to the input sample block. One desirable property of a transform is that the transform vector spans the space of the input samples. This means that any input vector (e.g., any combination of input sample values) can be represented by a weighted sum of the transform vectors. For a spanning transform, one necessary condition is that there must be at least as many transform vectors as the number of dimensions of the input space; in other words, the number of output transform coefficients must be at least equal to the number of input samples. For example, the one-dimensional DCT of VVC is a spanning transform. And in a spanning non-separable transform, if the input sample block is an M × N residual, the transform also outputs an M × N block of transform coefficients, which can be realized by a matrix implementation of (M × N) × (M × N) multiplications.

[0086] To derive a non-separable transform that produces coding gain for a particular directional feature, the transform can be trained. For example, a representative set of residual blocks corresponding to the directional feature of interest can be grouped together, and then a Karhunen-Loeve transform (KLT) can be calculated based on the covariance matrix of the set of residual blocks. This process can be repeated for K different sets of residual blocks. Then, in this example, an overall transform kernel of dimensionality (MxN)x(MxN)xK is derived.

[0087] As described in this section, spanning non-separable transforms have two challenges. First, they have high computational complexity. Non-separable transforms are usually learned and therefore generally not factorizable. The matrix implementation of the spanning non-separable transform in the above example results in a complexity of M × N multiplications per sample. Second, one or more transform kernels occupy a large amount of memory in the encoder 101 and decoder 201. In the above example, a single kernel adaptable to K different directional features has (M × N) × (M × N) × K weights. The kernel can only be applied to residual blocks of size M × N. To make a non-separable transform applicable to multiple block sizes, a transform kernel must be learned for each discrete block size.

[0088] In VVC, the LFNST tool was introduced and many modifications were made to solve the above-mentioned problems of spanning non-separable transformation.

[0089] First, according to the first amendment, the LFNST tool applies to a wide range of block sizes, but only two LFNST kernels are defined. For example, a small LFNST kernel is applied to blocks of size 4xN or Nx4 (N≧4). A large LFNST kernel is applied to all large block sizes (e.g., 8x8 and above).

[0090] Figure 16A is a diagram illustrating an LFNST kernel 1600 for 4xN and Nx4 block sizes in VVC according to some embodiments of the present disclosure. Figure 16B is a diagram illustrating an LFNST kernel 1601 for 8xN and Nx8 block sizes in VVC according to some embodiments of the present disclosure.

[0091] Figures 16A and 16B show the sample positions that LFNST operates on. For example, from the encoder's perspective, for a 4xN or Nx4 block size, the top-left 4x4 sample positions (shown in the shaded area in Figure 16A) are transformed by a small LFNST. The remaining sample positions (shown in the white area in Figure 16A) are ignored, or "zeroed out." From the decoder's perspective, an inverse LFNST is applied to generate the top-left 4x4 sample, and the remaining samples are filled with zeros. A similar strategy is applied to larger block sizes, where LFNST operates on the three top-left 4x4 blocks of sample positions (shown in the shaded area in Figure 16B). The remaining sample positions are zeroed out.

[0092] As a result of the "zero-out" strategy, the size of the LFNST is significantly reduced compared to a full-size transform applied to all sample positions. However, it is inherently lossy and cannot recover values ​​at sample positions ignored by the LFNST. If the LFNST tool were applied directly to the residual samples, such losses would be too great for the LFNST tool to be useful. However, the LFNST is called a secondary transform because it is applied after the separable DCT has already been performed in the encoder and operates on the primary transform coefficients to generate secondary transform coefficients. In other words, the DCT can be considered a linear transform. According to embodiments of the present disclosure, the leftmost sample position in a block of primary transform coefficients corresponds to the horizontal low frequencies of the DCT, and the topmost sample position corresponds to the vertical low frequencies of the DCT. By preferentially transforming and reconstructing the top-left sample position in the decoder 201, the LFNST can reconstruct low-frequency information from the original residual. As mentioned above, transforms generate coding gain due to their energy compaction properties, and it is well established that the variance (energy) of camera images and video signals is primarily concentrated in the low-frequency DCT coefficients. Thus, although "zeroing out" prevents LFNST from losslessly reconstructing any residual block, in practice the loss may be minimal for most categories of image and video signals.

[0093] The second modification is that in both the small and large LFNST kernels, the applied transform is not a spanning transform. From the encoder 101's perspective, the number of output (secondary transform) coefficients is less than the number of input (primary transform) coefficients. For example, the small LFNST kernel takes 4 × 4 = 16 primary transform coefficients as input but produces only eight output secondary transform coefficients. The large LFNST kernel takes 3 × 4 × 4 = 48 input primary transform coefficients and outputs eight secondary transform coefficients. The use of a non-spanning transform results in additional reconstruction loss. However, this loss can be traded off in a controlled manner for reduced implementation complexity. First, a spanning non-separable transform can be designed using the KLT method described above. Following this method, the basis vectors of the transform correspond to feature vectors of the covariance matrix calculated from a representative set of residual blocks. These feature vectors can be ordered by importance according to their corresponding feature values, and the most important feature vectors can be selected to construct a non-spanning non-separable transform. For example, the eight feature vectors with the largest feature values ​​can be selected and used to form a non-spanning transform for a small LFNST kernel.

[0094] In summary, compared to spanning non-separable transforms, the above two modifications significantly reduce the complexity of the LFNST kernel. For small blocks, using a small LFNST kernel reduces the potential complexity from (4 × N) × (4 × N) (N ≥ 4) multiplications per transform block to 16 × 8 multiplications. For large blocks, using a large LFNST kernel reduces the potential complexity from (8 × N) × (8 × N) (N ≥ 8) multiplications per transform block to 48 × 8 multiplications.

[0095] The LFNST kernel does not contain only one transformation matrix. Multiple transformation matrices are learned to achieve better coding gain for various image and video signals. The number of different transformation matrices is the product of the third and fourth dimensions of the LFNST kernel. That is, the dimensionality of a small LFNST kernel is 16x8x2x4, and the dimensionality of a large LFNST kernel is 48x8x2x4. The specific transformation matrix of a transform block is selected through a mixture of explicit signaling and implicit selection, so the LFNST kernel is expressed with two additional dimensions.

[0096] Explicit writing is performed by writing an LFNST index into the bitstream, which can take on values ​​of 0, 1, or 2, where 0 indicates that LFNST is not used for the transform block, and values ​​of 1 or 2 indicate selection of the third dimension of the LFNST kernel. The drawback of potential reconstruction loss caused by the simplifications due to zeroing out and non-spanning is ameliorated by the explicit writing mechanism. While using LFNST results in excessive reconstruction loss for the transform block, the LFNST tool can be disabled by writing the LFNST index to 0.

[0097] LFNST enables implicit selection by restricting it to coding units that use intra prediction. Intra prediction generates a prediction block for a coding unit from neighboring reference samples adjacent to the top and left of the current block. The intra prediction mode specifies a specific method for constructing the prediction block in the bitstream. Simple methods of intra prediction include averaging reference samples ("DC" mode) or constructing an affine interpolation between several reference samples ("Planar" mode). However, most intra prediction modes are reserved for specifying an intra angle direction, where the prediction block is constructed by assuming that the reference sample values ​​are copied along a specific direction. When an intra angle direction is used, it can be a strong hint to the directional characteristics of the residual block. An implicit selection of the LFNST transform is performed by mapping the intra prediction mode to one of four possible values ​​of the "transform set index," which is used to index the fourth dimension of the LFNST kernel. The mapping used in VVC is shown in Table 1.

[0098] [Table 1]

[0099] Intra-prediction modes 0 and 1 correspond to intra-prediction plane mode and intra-DC prediction mode, respectively. These intra-prediction modes 0 and 1 are treated as special cases by mapping to transform set index 0. Otherwise, the remaining intra-prediction modes correspond to the intra-angle directions partially shown in Figure 17. Intra-prediction mode 2 corresponds to diagonal intra-angle prediction from the bottom left. Increasing intra-prediction mode numbers correspond to clockwise rotation of the intra-prediction direction, where intra-prediction mode 34 corresponds to diagonal intra-angle prediction from the top left and intra-prediction mode 66 corresponds to diagonal intra-angle prediction from the top right.

[0100] For intra-prediction modes greater than 34 (which corresponds to an intra-angular prediction direction clockwise from the diagonal starting from the top left), the selected LFNST transform matrix is ​​applied to the primary transform coefficients in a transposed manner. In one implementation, this can be performed by scanning the primary transform coefficients in a transposed direction before applying the LFNST transform. For example, from the perspective of encoder 101, if the current block is predicted using intra-prediction mode 2, the primary transform coefficients can be reordered from a two-dimensional mode within the block to a one-dimensional vector using a row-major scan before applying the selected LFNST transform matrix T. And, in this example, if the current block is instead predicted using intra-prediction mode 66 and the same write LFNST indexes are used, the primary transform coefficients are alternatively reordered to a one-dimensional vector using a column-major scan before applying the same LFNST transform matrix T. In another embodiment, the same current block with intra-prediction mode 66 can be equivalently transformed by alternatively reordering the rows of transform matrix T while still performing a row-major scan on the primary transform coefficients.

[0101] More generally, the application of the LFNST transformation matrix can be written as follows: Let p be the linear transformation coefficient located at row y and column x. x,y where A×B is the number of dimensions of the LFNST transform matrix T, where A is the number of secondary transform coefficients and B is the number of non-zeroed primary transform coefficients. Next, in the case of an intra prediction mode of 34 or less, the primary transform coefficient p x,y An arbitrary scan order for constructing a one-dimensional vector P using may be defined in Equation 8 below:

[0102]

number

[0103] For intra prediction modes greater than 34, P is constructed alternatively by adopting the transposed scan order defined in Equation 9 below.

[0104]

number

[0105] The forward LFNST transform can be understood as a matrix multiplication S=TP, where S is a one-dimensional vector of secondary transform coefficients. In practice, since all multiplications are implemented in integer arithmetic, the transform is implemented as S=n(TP), where n() represents the normalization operation required for the integerized LFNST to approximate the ideal transform represented in floating point. The secondary transform coefficients are written back into the transform block in a forward diagonal scan order. Using the same notation as described above and the convention that the (0,0) location in a conventional DCT corresponds to "low frequency" or "DC," the scan order s can be defined according to Equation 10 below.

[0106]

number

[0107] From the decoder 201's perspective, the inverse LFNST transform is a matrix multiplication P=n(T T S), in other words, the inverse transformation is performed by transposing the matrix T.

[0108] The transposition of the primary transform coefficients for intra prediction modes greater than 34 allows the same LFNST transform matrix to be shared for symmetric intra angular prediction directions.

[0109] In post-VVC exploration efforts, an extension of LFNST is proposed and integrated into ECM. The LFNST tool in ECM relaxes some of the complexity reduction imposed by the original LFNST tool employed in VVC, achieving enhanced coding gains.

[0110] ECM has three LFNST kernels. Similar to VVC's LFNST tool, in most cases, significant portions of transform blocks are zeroed out, as shown in FIGS. 18A-18C. For example, FIG. 18A illustrates an LFNST kernel 1800 for 4×N and N×4 block sizes in ECM according to some embodiments of the present disclosure. FIG. 18B illustrates an LFNST kernel 1801 for 8×N and N×8 block sizes in ECM according to some embodiments of the present disclosure. FIG. 18C illustrates an LFNST kernel 1803 for 16×16 block size in ECM according to some embodiments of the present disclosure.

[0111] In Figures 18A-18C, the shaded areas indicate the locations of primary transform coefficients in the ECM that are affected by LFNST, and the white areas indicate which transform coefficient locations are zeroed out. For blocks of size 4xN or Nx4 (where N >= 4), a small LFNST kernel is used on the top-left 4x4 block of primary transform coefficients. For blocks of size 8xN or Nx8 (where N >= 8), a medium LFNST kernel is used on the four top-left 4x4 blocks of primary transform coefficients. For blocks of size 16x16 or larger, a large LFNST kernel is used on the six top-left 4x4 blocks of primary transform coefficients.

[0112] For the small LFNST kernel, the LFNST kernel size in ECM is 16x16x3x35. For the medium LFNST kernel, the LFNST kernel size in ECM is 64x32x3x35. For the large LFNST kernel, the LFNST kernel size in ECM is 96x32x3x35. Compared to the LFNST tool in VVC, the range of written LFNST indices increases from 2 to 3, and the number of LFNST transform sets increases from 4 to 35. This means that for each of the three indices, there are 35 LFNST transform matrices. Table 2 shows the mapping from intra-prediction modes to LFNST transform set indices. Similar to the LFNST tool in VVC, when the intra-prediction mode is greater than 34, the primary transform coefficients are transposed.

[0113] [Table 2]

[0114] The complexity burden of the LFNST tool can be evaluated in three ways. First, there is the additional memory burden imposed on the decoder 201 because the decoder must store the LFNST kernel. Second, there is the worst-case multiplication of each sample that the decoder 201 must perform if the LFNST tool is used. Third, there are the additional multiplications per sample that the encoder 101 would use if the full search of the LFNST tool were performed. By these three metrics, the enhanced LFNST proposed in ECM is more complex than the LFNST in VVC. However, in terms of total multiplications per sample, the worst-case decoder complexity may still be lower than the worst-case decoder complexity of other transform options.

[0115] For a matrix multiplication realization of the DCT applied separably to a transform of size MxN, there are (M + N) multiplications per sample. Thus, the worst-case complexity appears to be a maximum of (M + N). In practice, the complexity can be reduced by alternative realizations of the DCT, such as butterfly factorization, but it is still useful to evaluate the complexity of matrix multiplication realizations. In ECM, the separable DCT is extended, and the largest transform becomes a 128-point DCT. Then, the worst-case complexity of the separable DCT can be 128 + 128 = 256 multiplications per sample.

[0116] The worst-case decoder complexity of LFNST for ECM can be evaluated by considering several different block sizes. For a fair comparison, the evaluation includes the cost of performing the linear transform. For a 4x4 block, the linear transform involves 4 + 4 = 8 multiplications per sample. LFNST involves 16x16 matrix multiplications, i.e., 16 multiplications per sample. Therefore, the total cost of LFNST for a 4x4 block is 24 multiplications per sample.

[0117] For a 4x8 block, a naive implementation of the DCT linear transform typically requires eight 4x4 transforms along the short dimension and four 8x8 transforms along the long dimension, resulting in a total of 4 + 8 = 12 multiplications per sample. However, because the LFNST reconstructs only nonzero coefficient values ​​within the top-left 4x4 block of the primary transform coefficient locations, an optimized decoder performs only four 4x4 transforms along the short dimension and then four 4x8 transforms along the long dimension, resulting in a total of 2 + 4 = 6 multiplications per sample. Because the order of separable transforms is typically fixed, in the worst-case scenario, the decoder 201 can first perform four 4x8 transforms along the long dimension. Then, the decoder 201 can perform eight 4x4 transforms along the short dimension, resulting in 4 + 4 = 8 multiplications per sample. The LFNST is still a 16x16 matrix multiplication, and its cost is amortized over larger blocks, resulting in 8 multiplications per sample. Then the worst-case cost of LFNST for a 4x8 block is 6 multiplications per sample. The same principle typically applies to 4xN or Nx4 block sizes. Thus, the multiplications per sample for a 4xN or Nx4 block are always less than or equal to the multiplications per sample for a 4x4 block.

[0118] For an 8x8 block, the linear transform consists of 8 + 8 = 16 multiplications per sample. The LFNST consists of 64 x 32 matrix multiplications, i.e., 32 multiplications per sample. Then, the total cost of the LFNST for an 8x8 block is 48 multiplications per sample.

[0119] For an 8x16 block, assume again that decoder 201 takes advantage of the zero-out property of the LFNST reconstruction. Only the top-left 8x8 block of primary transform coefficient positions is nonzero. Decoder 201 may take advantage of this by performing only eight 8x8 transforms along the short dimension. Then, decoder 201 may perform eight 8x16 transforms along the long dimension. This results in a total of 4 + 8 = 12 multiplications per sample. Alternatively, decoder 201 may first perform eight 8x16 transforms along the long dimension. Then, decoder 201 may perform sixteen 8x8 transforms along the short dimension. This results in 8 + 8 = 16 multiplications per sample. The LFNST adds another (64 x 32) / (8 x 16) = 16 multiplications per sample, resulting in an overall worst-case complexity of 32 multiplications per sample. As noted above, the multiplications per sample of an 8xN or Nx8 block are always less than or equal to the multiplications per sample of an 8x8 block.

[0120] For a 16x16 block, the zero-out property of the LFNST reconstruction means that only six 4x4 blocks of primary transform coefficients in the pattern (shown in Figure 18C) have nonzero values. For simplicity, we assume a more relaxed mode, where the top-left 12x12 block of the primary transform position may have nonzero values. First, the decoder can take advantage of this by performing only twelve 12x16 transforms in one dimension. Next, the decoder 201 can perform sixteen 12x16 transforms in the second dimension, which involves 9 + 12 = 21 multiplications per sample. The LFNST involves (96x32) / (16x16) = 12 multiplications per sample, resulting in an overall complexity of 33 multiplications per sample.

[0121] For an M×N block, where M, N≧16, decoder 201 can first perform twelve 12×M transforms in one dimension. Next, decoder 201 can perform M 12×N transforms in the second dimension, resulting in (12×12) / N+12 multiplications per sample to perform a separable DCT. The worst-case complexity then occurs for the smallest value of N=16, which is 21 multiplications per sample, equal to the complexity of a 16×16 block. The LFNST adds another (96×32) / (M×N) multiplications per sample, which is always less than or equal to the multiplications per sample of a 16×16 block. Therefore, for large M×N block sizes, the overall complexity of the LFNST in ECM is always less than or equal to the multiplications per sample of a 16×16 block.

[0122] A comprehensive evaluation of the decoder complexity of the LFNST for ECM across different block sizes shows that the worst-case complexity is 48 multiplications per sample (occurs for 8 × 8 blocks). This worst-case complexity includes the cost of implementing a separable DCT using a matrix multiplication implementation, but due to possible optimizations using LFNST zeroing, this is significantly less than the worst-case complexity of implementing a separable DCT alone (which is estimated at 256 multiplications per sample). Even assuming a more realistic implementation of a separable DCT with butterfly factorization, the worst-case complexity of the LFNST still occurs for 8 × 8 blocks, which includes the cost of 32 multiplications per sample from the LFNST and the cost of the butterfly DCT. In such cases, the LFNST may be worst-case compared to the cost of a butterfly DCT applied separably to 256 × 256-sized blocks.

[0123] As mentioned above, significant complexity reductions can be achieved by using non-separable quadratic transforms due to the use of zero-out in selected linear transform coefficient regions. However, further coding can be achieved using NSPT. Initial work on non-separable linear transforms showed that significant gains (average rate reduction of 3.43% according to the Bjontegaard metric) could be achieved, despite the complexity of the implemented transforms and the kernel weights being obtained by overfitting a test dataset.

[0124] A practical implementation of NSPT is proposed. For example, NSPT is applicable only to a few block sizes: 4x4, 4x8, 8x4, and 8x8. For these block sizes, NSPT replaces both the linear transform and LFNST. Similar to LFNST, an NSPT kernel is trained, where the selection of the appropriate matrix for a particular block is guided by both the interpolated index and the implicit selection by the intra-prediction mode. Four types of NSPT kernels are proposed. For 4x4 blocks, a small NSPT kernel with a size of 16x16x3x35 is used. For 4x8 and 8x4 blocks, a medium NSPT kernel with a size of 32x20x3x35 is used. For 8x8 blocks, a large NSPT kernel with a size of 64x32x3x35 is used.

[0125] According to the present disclosure, the zeroing method may be defined as a reduction in the input dimension of a transform kernel, where the input dimension corresponds to the first dimension in the transform kernel dimension notation of the present disclosure. Reducing the input dimension of a forward transform is equivalent to reducing the support of the transform. For example, LFNST zeroing corresponds to a reduction in the number of DCT primary transform coefficients on which the forward LFNST operates to generate secondary transform coefficients. In the proposed NSPT, the transform operates directly on the residual coefficients, so reducing the first dimension of the NSPT kernel involves reducing the number of residual coefficients on which the forward NSPT operates to generate primary transform coefficients. In the above-mentioned known method, the size of the first dimension of each NSPT kernel is always equal to the number of samples in a block, so zeroing as defined by the present disclosure is not used. However, in this known method, zeroing is instead defined as a reduction in the output dimension of a transform kernel, where the output dimension is the second dimension in the transform kernel dimension notation of the present disclosure. This definition is not ambiguous in this method. This is because no reduction is performed on the input side of the NSPT kernel, and an example of the more commonly used term "zeroing" is given. However, for the sake of consistency and clarity in this disclosure, "zero-out" is defined to describe a reduction in the input dimension of a transform kernel, and a reduction in the output dimension is marked as a non-spanning or lossy transform. For the medium and large NSPT kernels, the second dimension is smaller than the first dimension, which means that the NSPT in these cases is a lossy transform.

[0126] Similar to LFNST, an NSPT index is written into the bitstream, and the NSPT index can take values ​​of 0, 1, 2, or 3, where 0 indicates that NSPT is not used for the transform block, and values ​​1-3 indicate the selection within the corresponding NSPT kernel along the third dimension. In the same manner as the extended LFNST in ECM, the selection of the NSPT kernel along the fourth dimension is determined by a mapping from the intra prediction mode shown in Table 3. Similar to LFNST, if the intra prediction mode is greater than 34 (which means the intra angle direction is diagonally up-left and clockwise), the input of the transform is transposed. However, for NSPT, the input consists of residual coefficients rather than primary transform coefficients.

[0127] [Table 3]

[0128] For an M×N residual block, if the intra prediction mode is 34 or less, the residual sample located at row y, column x is defined as r x,y and the dimension of the LFNST transform matrix T selected from the NSPT kernel for an M×N shaped block is A×B, where A is the number of linear transform coefficients and B=M×N is the number of residual samples in the block. Then, to construct a one-dimensional vector R by the residual samples, any scan order defined according to Equation 11 below can be adopted.

[0129]

number

[0130] For intra prediction modes greater than 34, R is alternatively constructed with a transposed scan order according to Equation 12.

[0131]

number

[0132] Additionally, if the intra prediction mode is greater than 34, the LFNST transform matrix T is alternatively selected from the NSPT kernel for blocks of NxM shape. For square block shapes, this is the same kernel, so T will be the same transform matrix. However, for 4x8 or 8x4 block sizes, a transform matrix from a different NSPT kernel is selected.

[0133] The forward NSPT transform can be realized as P = n(TR), where P is a one-dimensional vector of NSPT transform coefficients and n() represents the normalization operation required to approximate the ideal transform where the integerized NSPT is represented in floating point. We write the transform coefficients back into the transform block in a forward diagonal scan order. Using the same notation as above and the convention that the (0,0) location in a conventional DCT corresponds to "low frequency" or "DC," the scan order can be written according to Equation 13:

[0134]

number

[0135] From the decoder 201's perspective, the inverse NSPT transform is a matrix multiplication R=n(T T P), in other words, the inverse transformation is performed by transposing the matrix T.

[0136] The kernel size and block sizes for which NSPT is effective are designed to make NSPT practical. This can be confirmed by comparing the complexity of NSPT for each block size with that of the corresponding LFNST it replaces. For 4x4 blocks, LFNST has a complexity of 24 multiplications per sample, while NSPT has a complexity of 16 multiplications per sample. For 4x8 and 8x4 blocks, LFNST has a complexity of 16 multiplications per sample, while NSPT has a complexity of 20 multiplications per sample. For 8x8 blocks, LFNST has a complexity of 48 multiplications per sample, while NSPT has a complexity of 32 multiplications per sample. In some cases, NSPT has lower complexity than the LFNST it replaces, but in other cases, NSPT is more complex. Overall, the NSPT complexity is designed to avoid increasing the burden on the encoder. From the decoder 201's perspective, the worst-case complexity of NSPT, which is 32 multiplications per sample, is still lower than the worst-case complexity of all transform options that the decoder must support. Therefore, there is no increase in worst-case decoder complexity.

[0137] To improve the coding performance of the NSPT tool, an extension to larger block sizes is envisioned. In this proposal, NSPT can be extended to block sizes beyond the original set of 4x4, 4x8, 8x4, and 8x8. For block sizes of 4x16, 16x4, 8x16, and 16x8, NSPT is additionally applied, replacing LFNST. For block sizes of 4x16 and 16x4, an NSPT kernel with dimensions of 64x24x3x35 is used. For block sizes of 8x16 and 16x8, an NSPT kernel with dimensions of 128x40x3x35 is used.

[0138] In some methods, when IBC mode is selected, a block is directly copied from a previously reconstructed region in the same frame to be used as the prediction block for the current block. This does not consider spatial information between neighboring pixels, which may result in undesirably low prediction accuracy. Furthermore, a mode derived from DIMD or TIMD can be used to determine which LFNST or NSPT is used for intraTMP and IBC. DIMD and TIMD still use neighboring reconstructed pixels around the current block, which may not be optimal. This disclosure provides several solutions to further improve IBC and intraTMP.

[0139] For example, in some implementations, this disclosure proposes a filtered IBC (FIBC) prediction mode, where the IBC block vectors (bv x ,bv y ) are further filtered by the online learned filter. The resulting filtered block, rather than the reference block, is used as a predictor for the current block.

[0140] 19 is a diagram illustrating an example visualization of the spatial support of a trained filter 1900 (hereinafter referred to as "filter 1900") according to some embodiments of the present disclosure. FIG. 20 is a diagram illustrating an example visualization of a reference block / template for an IBC 2000 according to some embodiments of the present disclosure, where a reference block 2006 is a block vector (bv x ,bv y )

[0141] 19, filter 1900 may include one bias weight and five spatial weights at the central "C" position and the four adjacent "W," "N," "E," and "S" positions. In this example, the region of support S={(0,0),(-1,0),(0,1),(1,0),(0,-1)}, where (0,0), (-1,0), (0,1), (1,0), and (0,-1) represent the C, W, N, E, and S positions, respectively.

[0142] Assuming that IBC is one of the modes in the intra prediction module 410, a filter 1900 can be applied to the moving support region on the reconstructed block based on shift-invariant weighting. Let R and F be the reference block and the filtered prediction block, respectively, and R x,y and F x,y Let denote the sample and filtered sample at the xth column and yth row of R and F, respectively. Then, a general filter can be applied according to Equation 14:

[0143]

number

[0144] where a i,j where b and b are the learned filter weights that are adaptively adjusted using the pixels of the reconstructed block, B is a bias term that can be set as the mean luminance value of the input video depth (e.g., for a 10-bit video, the input video depth is 512), and S is the finite domain of support of the filter.

[0145] The intra-prediction module 410 can learn the filter weights by minimizing the error between the template of the reference block and the template of the current block. An example of an L-shaped template (e.g., current template 2004 and / or reference template 2008) is shown in Figure 20. However, the template can also take other shapes, such as a top-right rectangle or a left-right rectangle. Let RT and CT be the reference template 2008 of the reference block 2006 and the current template 2004 of the current block 2002, respectively, and RT x,y and CT x,y Let be the sample in the xth column and yth row of RT and CT, respectively. Then, the filter weights are selected to minimize the mean-square error (MSE) between the filtered RT and the template of the current block according to Equation 15.

[0146]

number

[0147] It is well known that this is the minimization of a quadratic cost function, and its solution can be found by solving a matrix equation. For regular structures generated by shift-invariant filters, the solution can be computed efficiently. For example, the matrix can be decomposed first by LDL decomposition, and then its subcomponents can be efficiently inverted.

[0148] When solving for filter weights and applying a trained filter, the support of the filter may exceed the range of available samples. In such cases, boundary padding can be used to generate values ​​for these ranges.

[0149] For example, if FIBC is enabled in the sequence parameter set (SPS), picture header (PH), picture parameter set (PPS), or slice header, one high-level flag can be written. If the IBC tool is enabled for the current video sequence, a CU-level flag indicating whether IBC is selected or not is written. If the current CU is predicted by IBC, an additional flag indicating whether to use the original IBC or FIBC is input, assuming that FIBC is enabled by the high-level flag. Thus, in some implementations, the intra prediction module 410 can identify whether to perform FIBC based on the high-level flag and the CU-level flag.

[0150] In some implementations, when a CU-level IBC flag is written, FIBC can be implicitly selected (e.g., FIBC replaces the original IBC mode if FIBC is enabled by a high-level flag), where the intra-prediction module 410 can identify whether FIBC is enabled for the current block 2002 based solely on the CU-level IBC flag.

[0151] In some implementations, FIBC can completely replace the original IBC mode, i.e., the existing IBC high-level flag instead indicates whether the FIBC tool is enabled, and the existing CU-level IBC flag instead indicates the FIBC mode. Here, the intra-prediction module 410 can identify whether FIBC has been selected for the current block 2002 based on the existing IBC high-level flag and identify which FIBC mode has been selected based on the CU-level IBC flag.

[0152] In some implementations, the intra prediction module 410 determines multiple reference blocks via IBC block vector signaling to form the final predictor for the current CU. 1 ,R 2 ,…R nLet,Σ k w k For some fusion weights that satisfy =1, the fusion combination of the reference blocks is P=wR 1 +w2R 2 +...w n R n Then, in this configuration, the final predictor is generated by the intra prediction module by further filtering the fused combination with the trained filter according to Equation 16.

[0153]

number

[0154] The filter weights are calculated by the intra prediction module 410 using the template RT 1 ,RT 2 ,…RT n It can be learned by minimizing the weighted error between the current block's template and the current block's template.

[0155]

number

[0156] In some implementations, the intra prediction module 410 determines multiple reference blocks by an IBC search to form the final predictor for the current CU. As described above, the reference blocks are 1 ,R 2 ,…R n Let,Σ k w k For some fusion weights that satisfy =1, the fusion combination of the reference blocks is P=wR 1 +w2R 2 +...w n R n Then, the configuration generates the final IBC predictor by further filtering the fused combination with the trained filter, as described above.

[0157] When intraTMP or IBC is selected as the prediction mode for the current CU, the intra prediction module 410 can apply a non-separable transform using the LFNST or NSPT tools. To apply a non-separable transform, the intra prediction module 410 first identifies an intra prediction mode for selecting a transform matrix. As described above in FIG. 14, the intra prediction module 410 can derive the intra prediction mode by performing a DIMD process on templates neighboring the current CU. However, for both intraTMP and IBC, a reference block for the current block is available, and when the intra prediction module 410 estimates the intra prediction mode for the current block, the reference block can better represent the current block than the neighboring reconstructed pixels surrounding the current block.

[0158] In some implementations, to determine which LFNST or NSPT transform matrix to apply to the prediction residual of the current block, it is proposed that the reference block 2006 for intraTMP and IBC shown in Figure 20 be used by the intra prediction module 410 to derive the intra mode in the same scheme as DIMD and TIMD. In some implementations, the intra prediction module 410 can perform a DIMD process on the reference block 2006 pointed to by the intraTMP or IBC block vector. The intra prediction module 410 can then use the intra prediction mode derived by the DIMD process to select an LFNST or NSPT matrix (if any one of these non-separable transform tools is selected).

[0159] A common tool employed in hybrid video coding systems in VVC, HEVC, and many other practical video coding standards predicts video pixels or samples in a current frame awaiting decoding using pixels or samples from other frames that have already been reconstructed. Coding tools that follow this general architecture are commonly referred to as "inter-prediction" tools, and the reconstructed frame may be referred to as a "reference frame." In still video scenes, inter-prediction can be achieved simply by decoding pixels or samples in a current frame based on co-located pixels or samples from a reference frame. However, in video scenes involving motion, an inter-prediction tool with motion compensation must be used. For example, a "current block" of samples in a current frame can be predicted based on a "prediction block" or "reference block" of samples in a reference frame, where the "prediction block" or "reference block" is determined by first decoding a "motion vector," which indicates the location of the prediction block in the reference frame relative to the current block in the current frame. More complex inter-prediction tools are used to take advantage of video scenes with complex motion (e.g., occlusion or affine motion).

[0160] FIG. 21 is a diagram illustrating an example visualization of a local illumination compensation (LIC) inter-prediction technique 2100 according to some embodiments of the present disclosure.

[0161] Referring to FIG. 21, illumination changes (e.g., blinking) may occur between a current picture and its reference picture. Conventional motion compensation, including affine-based motion compensation, cannot capture such illumination changes. LIC is an inter-prediction technique that models the local illumination change between a current block 2102 and its predicted block as a linear function of the local illumination change between a current block template and a reference block template. For an M×N current block, a reference block of the same size, M×N, is determined based on a reference frame using transmitted or derived side information, such as a reference index, motion model mode selection, and motion vector. The current block template is determined based on reconstructed pixels 2104 in a region adjacent to the current block 2102. Similarly, the reference block template is determined based on reconstructed pixels 2108 in a region adjacent to the reference block 2106. In addition to the example shown in FIG. 21, templates of other shapes can also be used. A reference block 2106 for the current block 2102 is determined based on known motion information, and the reference block 2106 for the current block 2102 is represented in brown in the reference frame. Therefore, in the reference frame shown in green, a template of the same shape in the reference frame is determined.

[0162] The reconstructed template in the reference frame and the reconstructed template in the current frame are used to model a function that reflects the local illumination change of the current block 2102. Typically, a polynomial function can be used to model such a relationship. However, modeling a high-order polynomial function may be too complex. A simple linear function can provide a good tradeoff between model complexity and accuracy. The linear function can be expressed with a slope a and an offset b. More specifically, the luminance samples of the predicted block p[x] are calculated according to Equation 18.

[0163] [Formula 18] p[x]=a*r[x]+b

[0164] where r[x] is the luma sample at location x in the reference picture pointed to by the MV, and p[x] is the luma prediction sample generated by applying the local illumination compensation model. The model parameters a and b can be derived by fitting the model to the luma samples in the current block template and the reference block template. For example, the model can be fitted using the least mean square (LMS) method. Because both the current block template and the reference block template are reconstructed and available to the decoder when decoding the current block, the model parameters can be determined implicitly, and no signaling overhead of a and b is required. However, in LIC mode, the LIC flag can be written to indicate the use of LIC.

[0165] Using the techniques described above, if IBC mode is selected, a block is directly copied from a previously reconstructed region in the same frame, or a filtered block is directly copied, and used as the prediction block for the current block 2102. Unfortunately, if the current block is coded using IBC merge mode, FIBC mode is not allowed, which may not be optimal.

[0166] To overcome these challenges of existing IBCs, the present disclosure proposes several improvements to the filtered IBC (FIBC) mode, as described below.

[0167] In some implementations, this disclosure proposes that when a current block / CU is coded using IBC merge mode, the FIBC prediction mode may be implicitly written. More specifically, when the prediction mode of a current block is written as IBC merge mode, the prediction may be a reference block pointed to by a block vector inherited from a previous CU. Additionally and / or alternatively, the prediction may be a filtered block generated by filtering the reference block using an online trained filter.

[0168] In some implementations, this disclosure proposes that the inherited block vector can be inherited from a previous CU predicted by IBC, FIBC, intraTMP, or filtered intraTMP. If the inherited block vector is inherited from a previous CU predicted by IBC or intraTMP mode, the current block can be predicted by the reference block pointed to by the inherited block vector. If the inherited block vector is inherited from a previous CU predicted by filtered IBC or filtered intraTMP, and the template regions of the reference block and the current block are available for filter decision, the current block can be predicted by the filtered block. If the inherited block is inherited from a previous CU predicted by filtered IBC or filtered intraTMP, but the template region is unavailable, the current block can be predicted by the reference block. Whether the current block is predicted by the reference block or the filtered block can be determined without decoding an additional FIBC flag from the bitstream.

[0169] This disclosure further proposes that if a block is coded using IBC-local illuminance compensation (IBC-LIC), combined IBC and intra prediction (CIIP), or combined IBC and geometry partition mode (GPM), FIBC may not be allowed for these modes. Furthermore, if reconstructed-reordered IBC is allowed, FIBC may only be allowed for regular IBC and not for horizontally and vertically flipped IBC. For example, if a current block is coded according to IBC-LIC, IBC-CIIP, IBC-GPM, or horizontally and vertically flipped IBC, the reference block pointed to by the IBC BV may not be filtered. Furthermore, if a current block is coded according to IBC-LIC, IBC-CIIP, IBC-GPM, or horizontally and vertically flipped IBC, the FIBC flag may not be coded in or transmitted in the bitstream.

[0170] Additionally and / or alternatively, the inter prediction module 408 and the intra prediction module 410 may be configured to generate a prediction block based on information related to prediction block generation provided by the decoding module 402 and information on a previously decoded block or picture provided by the buffer module 414. As described above, when performing intra prediction in the same manner as the operation of an encoder, if the size of the prediction unit and the size of the transform unit are the same, intra prediction may be performed in the prediction unit based on pixels located to the left, upper left, and top of the prediction unit. However, when performing intra prediction, if the size of the prediction unit and the size of the transform unit are different, intra prediction may be performed using reference pixels based on the transform unit.

[0171] For example, the inter prediction module 408 may be configured to receive from the encoder a bitstream including a reference frame, a current frame, and an indication of a weighting factor associated with a multimedia home platform (MHP) process. The inter prediction module 408 may be configured to perform an MHP process on a CU located in the current frame based on a search block in the reference frame (e.g., a reference frame and / or a reference template). In some embodiments, to perform the MHP process, the inter prediction module 408 may be configured to perform template matching on a CU located in the current frame based on the search block in the reference frame and the weighting factor to obtain motion information. In some embodiments, to perform the MHP process, the inter prediction module 408 may be configured to identify a weighting factor index associated with the weighting factor based on the template matching. The inter prediction module 408 may be configured to identify a weighting factor sign of the weighting factor based on an indication included in the bitstream. The inter prediction module 408 may be configured to perform the inter prediction process and decode the bitstream based on the current frame, the reference frame, the weighting factor index, and the weighting factor sign of the weighting factor.

[0172] A combined reconstructed block or picture from the output of the inverse transform module 406 and the prediction module 408 or the prediction module 410 may be provided to a filter module 412. The filter module 412 may include a deblocking filter, an offset correction module, and an ALF. A buffer module 414 may store the reconstructed picture or block and use it as a reference picture or block for the inter prediction module 408, which may output the reconstructed picture.

[0173] According to the scope of the present disclosure, the encoding module 320 and the decoding module 402 may be configured to encode video pictures using a quantization level binarization scheme with a Rice parameter adapted to the bit depth and / or bit rate to improve encoding efficiency.

[0174] 22 is a flowchart of an example method 2200 of video decoding according to some embodiments of the present disclosure. Method 2200 may be performed by a system such as decoding system 200, decoder 201, or intra-prediction module 410. Method 2200 may include operations 2202-2214, which are described below. It should be understood that some of these steps may be optional, and some steps may be performed simultaneously or in a different order than shown in FIG. 22.

[0175] Referring to Figure 22, the system can parse the bitstream to decode at least one flag at 2202. For example, referring to Figures 2 and 4, the decoder 201 can parse the bitstream encoded by the encoder 101 based on at least one flag. By parsing the bitstream, the decoder 201 can obtain at least one flag that can indicate FIBC.

[0176] At 2204, the system may identify whether FIBC mode is enabled for the current block. For example, referring to FIG. 4, the intra prediction module 410 may identify whether FIBC mode is enabled for the current block based on at least one flag parsed from the bitstream. In some implementations, the intra prediction module 410 may identify that IBC mode has been selected for the current block in response to a first flag indicating that IBC mode has been selected for the current block. In some implementations, the first flag may be a CU-level flag. In some implementations, the intra prediction module 410 may identify that FIBC mode is enabled for the current block in response to a second flag indicating that FIBC mode is enabled for the current block. In some implementations, the second flag may be associated with an SPS, PH, PPS, or slice header. In some implementations, the intra prediction module 410 may identify that FIBC mode is enabled for the current block in response to a CU level indicating that FIBC mode is enabled for the current block. In some implementations, the intra prediction module 410 can identify that FIBC tools are enabled in response to an IBC high-level flag indicating that IBC tools are enabled. In some implementations, the intra prediction module 410 can identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that IBC mode is selected for the current block.

[0177] At 2206, the system can select a filter weight set for a shift-invariant weighting filter, where the filter weight set minimizes the MSE between the filtered reference template associated with the reference block and the current template associated with the current block. For example, referring to Figures 4, 19, and 20, the filter 1900 may include one bias weight and five spatial weights at a central "C" position and four adjacent "W," "N," "E," and "S" positions. In this example, the region of support S = {(0,0), (-1,0), (0,1), (1,0), (0,-1)}, where (-1,0), (0,1), (1,0), and (0,-1) represent the W, N, E, and S positions, respectively. The intra-prediction module 410 can apply the filter 1900 based on shift-invariant weighting to the region of support moving over the reconstructed block. Let R and F be the reference block and the filtered prediction block, respectively, and R x,y and F x,y Let RT and CT denote the samples and filtered samples in the xth column and yth row of R and F. A general filter can then be applied according to Equation 14. The intra prediction module 410 can learn the filter weights by minimizing the error between the template of the reference block and the template of the current block. FIG. 20 shows an example of an L-shaped template (e.g., current template 2004 and / or reference template 2008). However, the templates can also take other shapes, such as a top-rectangle or a left-rectangle. Let RT and CT be the reference template 2008 of the reference block 2006 and the current template 2004 of the current block 2002, respectively, and RT x,y and CT x,yLet be the samples in the xth column and yth row of RT and CT, respectively. Then, according to Equation 15, the filter weights are selected to minimize the MSE between the filtered RT and the template of the current block. This is the minimization of a quadratic cost function, the solution of which is well known to be obtained by solving a matrix equation. For regular structures generated by shift-invariant filters, the solution can be computed efficiently. For example, we can first decompose the matrix by LDL decomposition, and then efficiently invert its subcomponents.

[0178] At 2208, the system can generate boundary padding for unavailable pixels in the reference block and the template associated with the reference block. For example, with reference to Figures 4, 19, and 20, when solving for filter weights and applying a learned filter, the support of the filter may exceed the range of available samples. In such cases, boundary padding can be used to generate values ​​for these regions.

[0179] At 2210, the system can generate a shift-invariant weighting filter based on the set of filter weights. For example, referring to Figures 4, 19, and 20, filter 1900 may include one bias weight and five spatial weights at the central "C" position and the four adjacent "W," "N," "E," and "S" positions. In this example, the region of support S = {(0,0), (-1,0), (0,1), (1,0), (0,-1)}, where (-1,0), (0,1), (1,0), and (0,-1) represent the W, N, E, and S positions, respectively. The intra-prediction module 410 can apply filter 1900 based on shift-invariant weighting to the region of support moving over the reconstructed block. Let R and F be the reference block and the filtered prediction block, respectively, and R x,y and F x,yLet RT and CT denote the samples and filtered samples in the xth column and yth row of R and F. A general filter can then be applied according to Equation 14. The intra prediction module 410 can learn the filter weights by minimizing the error between the template of the reference block and the template of the current block. FIG. 20 shows an example of an L-shaped template (e.g., current template 2004 and / or reference template 2008). However, the templates can also take other shapes, such as a top-rectangle or a left-rectangle. Let RT and CT be the reference template 2008 of the reference block 2006 and the current template 2004 of the current block 2002, respectively, and RT x,y and CT x,y Let be the samples in the xth column and yth row of RT and CT, respectively. Then, according to Equation 15, the filter weights are selected to minimize the MSE between the filtered RT and the template of the current block. This is the minimization of a quadratic cost function, the solution of which is well known to be obtained by solving a matrix equation. For regular structures generated by shift-invariant filters, the solution can be computed efficiently. For example, we can first decompose the matrix by LDL decomposition, and then efficiently invert its subcomponents.

[0180] At 2212, in response to FIBC mode being enabled for the current block, the system can generate a filtered predicted block by applying a shift-invariant weighting filter to the reference block. For example, with reference to Figures 4 and 20, the intra prediction module 410 can generate a filtered predicted block by applying the filter 1900 to the reference template 2008 and / or the reference block 2006.

[0181] At 2214, the system can decode the current block based on the filtered prediction block. For example, referring to Figures 4 and 20, the intra prediction module 410 can decode the current block 2002 based on the filtered prediction block.

[0182] 23 is a flowchart of an example method 2300 of video encoding according to some embodiments of the present disclosure. The method 2300 may be performed by a system such as the encoding system 100, the encoder 101, or the intra-prediction module 306. The method 2300 may include operations 2302-2312, which are described below. It should be understood that some of these steps may be optional, and some steps may be performed simultaneously or in a different order than shown in FIG. 23.

[0183] At 2302, the system may identify whether FIBC mode is enabled for the current block. For example, referring to FIG. 3 , the intra prediction module 306 may identify whether FIBC mode is enabled for the current block based on at least one flag parsed from the bitstream. In some implementations, the intra prediction module 306 may identify that IBC mode has been selected for the current block in response to a first flag indicating that IBC mode has been selected for the current block. In some implementations, the first flag may be a CU-level flag. In some implementations, the intra prediction module 306 may identify that FIBC mode is enabled for the current block in response to a second flag indicating that FIBC mode is enabled for the current block. In some implementations, the second flag may be associated with an SPS, PH, PPS, or slice header. In some implementations, the intra prediction module 306 may identify that FIBC mode is enabled for the current block in response to a CU level indicating that FIBC mode is enabled for the current block. In some implementations, the intra prediction module 306 can identify that FIBC tools are enabled in response to an IBC high-level flag indicating that IBC tools are enabled. In some implementations, the intra prediction module can identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that IBC mode is selected for the current block.

[0184] At 2304, the system can select a filter weight set for a shift-invariant weighting filter, where the filter weight set minimizes the MSE between the filtered reference template associated with the reference block and the current template associated with the current block. For example, referring to Figures 3, 19, and 20, the filter 1900 may include one bias weight and five spatial weights at a central "C" position and four adjacent "W," "N," "E," and "S" positions. In this example, the region of support S = {(0,0), (-1,0), (0,1), (1,0), (0,-1)}, where (-1,0), (0,1), (1,0), and (0,-1) represent the W, N, E, and S positions, respectively. The intra-prediction module 306 can apply the filter 1900 based on shift-invariant weighting to the region of support moving over the reconstructed block. Let R and F be the reference block and the filtered prediction block, respectively, and R x,y and F x,y Let RT and CT denote the samples and filtered samples in the xth column and yth row of R and F. A general filter can then be applied according to Equation 14. The intra prediction module 306 can learn the filter weights by minimizing the error between the template of the reference block and the template of the current block. FIG. 20 shows an example of an L-shaped template (e.g., current template 2004 and / or reference template 2008). However, the templates can also take other shapes, such as a top-rectangle or a left-rectangle. Let RT and CT be the reference template 2008 of the reference block 2006 and the current template 2004 of the current block 2002, respectively, and RT x,y and CT x,yLet be the samples in the xth column and yth row of RT and CT, respectively. Then, according to Equation 15, the filter weights are selected to minimize the MSE between the filtered RT and the template of the current block. This is the minimization of a quadratic cost function, the solution of which is well known to be obtained by solving a matrix equation. For regular structures generated by shift-invariant filters, the solution can be computed efficiently. For example, we can first decompose the matrix by LDL decomposition, and then efficiently invert its subcomponents.

[0185] At 2306, the system can generate boundary padding for unavailable pixels in the reference block and the template associated with the reference block. For example, with reference to Figures 3, 19, and 20, when solving for filter weights and applying a learned filter, the support of the filter may exceed the range of available samples. In such cases, boundary padding can be used to generate values ​​for these regions.

[0186] At 2308, the system can generate a shift-invariant weighting filter based on the set of filter weights. For example, referring to Figures 3, 19, and 20, the filter 1900 may include one bias weight and five spatial weights at the central "C" position and the four adjacent "W," "N," "E," and "S" positions. In this example, the region of support S = {(0,0), (-1,0), (0,1), (1,0), (0,-1)}, where (-1,0), (0,1), (1,0), and (0,-1) represent the W, N, E, and S positions, respectively. The intra-prediction module 410 can apply the filter 1900 based on shift-invariant weighting to the region of support moving over the reconstructed block. Let R and F be the reference block and the filtered prediction block, respectively, and R x,y and F x,yLet RT and CT denote the samples and filtered samples in the xth column and yth row of R and F. A general filter can then be applied according to Equation 14. The intra prediction module 306 can learn the filter weights by minimizing the error between the template of the reference block and the template of the current block. FIG. 20 shows an example of an L-shaped template (e.g., current template 2004 and / or reference template 2008). However, the templates can also take other shapes, such as a top-rectangle or a left-rectangle. Let RT and CT be the reference template 2008 of the reference block 2006 and the current template 2004 of the current block 2002, respectively, and RT x,y and CT x,y Let be the samples in the xth column and yth row of RT and CT, respectively. Then, according to Equation 15, the filter weights are selected to minimize the MSE between the filtered RT and the template of the current block. This is the minimization of a quadratic cost function, the solution of which is well known to be obtained by solving a matrix equation. For regular structures generated by shift-invariant filters, the solution can be computed efficiently. For example, we can first decompose the matrix by LDL decomposition, and then efficiently invert its subcomponents.

[0187] At 2310, in response to FIBC mode being enabled for the current block, the system can generate a filtered predicted block by applying a shift-invariant weighting filter to the reference block. For example, with reference to Figures 3 and 20, the intra prediction module 306 can generate a filtered predicted block by applying the filter 1900 to the reference template 2008 and / or the reference block 2006.

[0188] At 2312, the system can encode the current block based on the filtered predicted block. For example, referring to Figures 3 and 20, the intra prediction module 306 can encode the current block 2002 based on the filtered predicted block.

[0189] In various aspects of the present disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. A storage medium may be any available medium that can be accessed by a processor (such as processor 102 in FIGS. 1 and 2 ). By way of non-limiting example, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, HDD (such as magnetic disk storage or other magnetic storage), flash drive, SSD, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system (such as a mobile device or computer). As used herein, disk and disk include CDs, laser disks, optical disks, digital video disks (DVDs), and floppy disks, where disks typically reproduce data magnetically and disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0190] According to one aspect of the present disclosure, a decoding method by a decoder is provided. The method may include identifying, by a processor, whether FIBC mode is enabled for a current block. In response to FIBC mode being enabled for the current block, the method may include generating, by the processor, a filtered prediction block by applying a shift-invariant weighting filter to a reference block. The method may include decoding, by the processor, the current block based on the filtered prediction block.

[0191] In some implementations, the method may include selecting, by a processor, a set of filter weights for a shift-invariant weighting filter, the set of filter weights minimizing an MSE between a filtered reference template associated with a reference block and a current template associated with a current block. In some implementations, the method may include generating, by a processor, a shift-invariant weighting filter based on the set of filter weights.

[0192] In some implementations, the method may include generating, by a processor, border padding for unavailable pixels in the reference block and the template associated with the reference block. In some implementations, a shift-invariant weighting filter can be generated based on the border padding to generate the filtered prediction block.

[0193] In some implementations, the method may include parsing, by a processor, the bitstream to identify at least one flag.

[0194] In some implementations, the processor identifying whether the FIBC mode is enabled for the current block may include the processor identifying that the IBC mode is selected for the current block in response to a first flag indicating that the IBC mode is selected for the current block. In some implementations, the first flag may be a CU-level flag. In some implementations, the processor identifying whether the FIBC mode is enabled for the current block may include the processor identifying that the FIBC mode is enabled for the current block in response to a second flag indicating that the FIBC mode is enabled for the current block. In some implementations, the second flag may be associated with an SPS, a PH, a PPS, or a slice header.

[0195] In some implementations, the processor identifying whether FIBC mode is enabled for the current block may include the processor identifying that FIBC mode is enabled for the current block in response to the CU level indicating that FIBC mode is enabled for the current block.

[0196] In some implementations, the processor identifying whether FIBC mode is enabled for the current block may include the processor identifying that the FIBC tools are enabled in response to an IBC high-level flag indicating that the IBC tools are enabled. In some implementations, the processor identifying whether FIBC mode is enabled for the current block includes the processor identifying that FIBC mode is enabled for the current block in response to a CU-level flag indicating that IBC mode has been selected for the current block.

[0197] In some implementations, if a CU is coded using IBC merge mode, FIBC mode may be identified as valid. In some implementations, the filtered prediction block may be identified based on a block vector inherited from the previous current block.

[0198] In some implementations, in response to the previous current block being predicted in IBC mode or intraTMP mode, the current block can be decoded based on the reference block pointed to by the inherited block vector. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being available for filter decision, the filtered predicted block is determined. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being unavailable for filter decision, the current block can be decoded based on the reference block.

[0199] In some implementations, if the current block is coded using IBC-LIC or CIIP or GPM, FIBC mode may be identified as valid when the CU is coded.

[0200] According to another aspect of the present disclosure, a system is provided. The system may include a processor and a memory storing instructions. The instructions stored in the memory, when executed by the processor, cause the processor to identify whether FIBC mode is enabled for a current block. The instructions, when executed by the processor, cause the processor to generate a filtered predictive block by applying a shift-invariant weighting filter to a reference block in response to FIBC mode being enabled for the current block. The instructions, when executed by the processor, cause the processor to decode the current block based on the filtered predictive block.

[0201] In some implementations, the instructions stored in the memory, when executed by the processor, can also cause the processor to select a set of filter weights for a shift-invariant weighting filter, the set of filter weights minimizing an MSE between a filtered reference template associated with the reference block and a current template associated with the current block. In some implementations, the instructions stored in the memory, when executed by the processor, can also cause the processor to generate a shift-invariant weighting filter based on the set of filter weights.

[0202] In some implementations, the instructions stored in the memory, when executed by the processor, can also cause the processor to generate border padding for unavailable pixels in the reference block and the template associated with the reference block. In some implementations, a shift-invariant weighting filter can be generated based on the border padding to generate the filtered prediction block.

[0203] In some implementations, the instructions stored in the memory, when executed by a processor, to identify whether the FIBC mode is enabled for the current block cause the processor to identify that the IBC mode has been selected for the current block in response to a first flag indicating that the IBC mode has been selected for the current block. In some implementations, the first flag may be a CU-level flag. In some implementations, the instructions stored in the memory, when executed by a processor, to identify whether the FIBC mode is enabled for the current block cause the processor to identify that the FIBC mode is enabled for the current block in response to a second flag indicating that the FIBC mode is enabled for the current block. In some implementations, the second flag may be associated with an SPS, a PH, a PPS, or a slice header.

[0204] In some implementations, instructions stored in memory, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that FIBC mode is enabled for the current block.

[0205] In some implementations, instructions stored in memory, when executed by a processor, to identify whether FIBC mode is enabled for the current block cause the processor to identify that the FIBC tool is enabled in response to an IBC high-level flag indicating that the IBC tool is enabled. In some implementations, instructions stored in memory, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that IBC mode has been selected for the current block.

[0206] In some implementations, if a CU is coded using IBC merge mode, FIBC mode may be identified as valid. In some implementations, the filtered prediction block may be identified based on a block vector inherited from the previous current block.

[0207] In some implementations, in response to the previous current block being predicted in IBC mode or intraTMP mode, the current block can be decoded based on the reference block pointed to by the inherited block vector. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being available for filter decision, the filtered predicted block is determined. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being unavailable for filter decision, the current block can be decoded based on the reference block.

[0208] In some implementations, if the current block is coded using IBC-LIC or CIIP or GPM, FIBC mode may be identified as valid when the CU is coded.

[0209] According to another aspect of the present disclosure, a non-transitory computer-readable medium is provided that stores instructions that, when executed by a processor, cause the processor to identify whether FIBC mode is enabled for a current block. When executed by the processor, the instructions cause the processor, in response to FIBC mode being enabled for the current block, to generate a filtered prediction block by applying a shift-invariant weighting filter to a reference block. When executed by the processor, the instructions cause the processor to decode the current block based on the filtered prediction block.

[0210] In some implementations, the instructions, when executed by a processor, can also cause the processor to select a set of filter weights for a shift-invariant weighting filter, the set of filter weights minimizing an MSE between a filtered reference template associated with a reference block and a current template associated with a current block. In some implementations, the instructions, when executed by a processor, can also cause the processor to generate a shift-invariant weighting filter based on the set of filter weights.

[0211] In some implementations, the instructions, when executed by a processor, may also cause the processor to generate border padding for unavailable pixels in the reference block and the template associated with the reference block. In some implementations, a shift-invariant weighting filter may be generated based on the border padding to generate the filtered prediction block.

[0212] In some implementations, the instructions, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that IBC mode has been selected for the current block in response to a first flag indicating that IBC mode has been selected for the current block. In some implementations, the first flag may be a CU-level flag. In some implementations, the instructions, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that FIBC mode is enabled for the current block in response to a second flag indicating that FIBC mode is enabled for the current block. In some implementations, the second flag may be associated with an SPS, a PH, a PPS, or a slice header.

[0213] In some implementations, to identify whether FIBC mode is enabled for the current block, the instruction, when executed by the processor, causes the processor to identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that FIBC mode is enabled for the current block.

[0214] In some implementations, the instructions, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that the FIBC tool is enabled in response to an IBC high-level flag indicating that the IBC tool is enabled. In some implementations, the instructions, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that the FIBC mode is enabled for the current block in response to a CU-level flag indicating that IBC mode is selected for the current block.

[0215] In some implementations, if a CU is coded using IBC merge mode, FIBC mode may be identified as valid. In some implementations, the filtered prediction block may be identified based on a block vector inherited from the previous current block.

[0216] In some implementations, in response to the previous current block being predicted in IBC mode or intraTMP mode, the current block can be decoded based on the reference block pointed to by the inherited block vector. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being available for filter decision, the filtered predicted block is determined. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being unavailable for filter decision, the current block can be decoded based on the reference block.

[0217] In some implementations, if the current block is coded using IBC-LIC or CIIP or GPM, FIBC mode may be identified as valid when the CU is coded.

[0218] According to one aspect of the present disclosure, there is provided an encoding method by an encoder. The method may include a processor identifying whether FIBC mode is enabled for a current block. In response to FIBC mode being enabled for the current block, the method may include the processor generating a filtered prediction block by applying a shift-invariant weighting filter to a reference block. The method may include the processor encoding the current block based on the filtered prediction block.

[0219] In some implementations, the method may include a processor selecting a set of filter weights for a shift-invariant weighting filter, the set of filter weights minimizing an MSE between a filtered reference template associated with a reference block and a current template associated with a current block. In some implementations, the method may include a processor generating a shift-invariant weighting filter based on the set of filter weights.

[0220] In some implementations, the method may include generating, by a processor, border padding for unavailable pixels in the reference block and a template associated with the reference block. In some implementations, a shift-invariant weighting filter can be generated based on the border padding to generate the filtered prediction block.

[0221] In some implementations, the method may include a processor parsing the bitstream to identify at least one flag.

[0222] In some implementations, the processor identifying whether the FIBC mode is enabled for the current block may include the processor identifying that the IBC mode is selected for the current block in response to a first flag indicating that the IBC mode is selected for the current block. In some implementations, the first flag may be a CU-level flag. In some implementations, the processor identifying whether the FIBC mode is enabled for the current block may include the processor identifying that the FIBC mode is enabled for the current block in response to a second flag indicating that the FIBC mode is enabled for the current block. In some implementations, the second flag may be associated with an SPS, a PH, a PPS, or a slice header.

[0223] In some implementations, the processor identifying whether FIBC mode is enabled for the current block may include the processor identifying that FIBC mode is enabled for the current block in response to the CU level indicating that FIBC mode is enabled for the current block.

[0224] In some implementations, the processor identifying whether the FIBC mode is enabled for the current block may include the processor identifying that the FIBC tool is enabled in response to the IBC high-level flag indicating that the IBC tool is enabled. In some implementations, the processor identifying whether the FIBC mode is enabled for the current block may include the processor identifying that the FIBC mode is enabled for the current block in response to the CU-level flag indicating that the IBC mode has been selected for the current block.

[0225] In some implementations, if a CU is coded using IBC merge mode, FIBC mode may be identified as valid. In some implementations, the filtered prediction block may be identified based on a block vector inherited from the previous current block.

[0226] In some implementations, in response to the previous current block being predicted in IBC mode or intraTMP mode, the current block can be coded based on the reference block pointed to by the inherited block vector. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being available for filter decision, the filtered predicted block is determined. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being unavailable for filter decision, the current block can be coded based on the reference block.

[0227] In some implementations, if the current block is coded using IBC-LIC or CIIP or GPM, FIBC mode may be identified as valid when the CU is coded.

[0228] According to another aspect of the present disclosure, a system is provided. The system may include a processor and a memory storing instructions. The instructions stored in the memory, when executed by the processor, cause the processor to identify whether FIBC mode is enabled for a current block. The instructions stored in the memory, when executed by the processor, cause the processor to generate a filtered predictive block by applying a shift-invariant weighting filter to a reference block in response to FIBC mode being enabled for the current block. The instructions, when executed by the processor, cause the processor to encode the current block based on the filtered predictive block.

[0229] In some implementations, the instructions stored in the memory, when executed by the processor, can also cause the processor to select a set of filter weights for a shift-invariant weighting filter, the set of filter weights minimizing an MSE between a filtered reference template associated with the reference block and a current template associated with the current block. In some implementations, the instructions stored in the memory, when executed by the processor, can also cause the processor to generate a shift-invariant weighting filter based on the set of filter weights.

[0230] In some implementations, the instructions stored in the memory, when executed by the processor, can also cause the processor to generate border padding for unavailable pixels in the reference block and the template associated with the reference block. In some implementations, a shift-invariant weighting filter can be generated based on the border padding to generate the filtered prediction block.

[0231] In some implementations, the instructions stored in the memory, when executed by a processor, to identify whether the FIBC mode is enabled for the current block cause the processor to identify that the IBC mode has been selected for the current block in response to a first flag indicating that the IBC mode has been selected for the current block. In some implementations, the first flag may be a CU-level flag. In some implementations, the instructions stored in the memory, when executed by a processor, to identify whether the FIBC mode is enabled for the current block cause the processor to identify that the FIBC mode is enabled for the current block in response to a second flag indicating that the FIBC mode is enabled for the current block. In some implementations, the second flag may be associated with an SPS, a PH, a PPS, or a slice header.

[0232] In some implementations, instructions stored in memory, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that FIBC mode is enabled for the current block.

[0233] In some implementations, instructions stored in memory, when executed by a processor, to identify whether FIBC mode is enabled for the current block cause the processor to identify that the FIBC tool is enabled in response to an IBC high-level flag indicating that the IBC tool is enabled. In some implementations, instructions stored in memory, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that IBC mode has been selected for the current block.

[0234] In some implementations, if a CU is coded using IBC merge mode, FIBC mode may be identified as valid. In some implementations, the filtered prediction block may be identified based on a block vector inherited from the previous current block.

[0235] In some implementations, in response to the previous current block being predicted in IBC mode or intraTMP mode, the current block can be coded based on the reference block pointed to by the inherited block vector. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being available for filter determination, the filtered predicted block is determined. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being unavailable for filter determination, the current block can be coded based on the reference block.

[0236] In some implementations, if the current block is coded using IBC-LIC or CIIP or GPM, FIBC mode may be identified as valid when the CU is coded.

[0237] According to another aspect of the present disclosure, a non-transitory computer-readable medium is provided that stores instructions that, when executed by a processor, cause the processor to identify whether FIBC mode is enabled for a current block. When executed by the processor, the instructions cause the processor to generate a filtered prediction block by applying a shift-invariant weighting filter to a reference block in response to FIBC mode being enabled for the current block. When executed by the processor, the instructions cause the processor to encode the current block based on the filtered prediction block.

[0238] In some implementations, the instructions, when executed by a processor, can also cause the processor to select a set of filter weights for a shift-invariant weighting filter, the set of filter weights minimizing an MSE between a filtered reference template associated with a reference block and a current template associated with a current block. In some implementations, the instructions, when executed by a processor, can also cause the processor to generate a shift-invariant weighting filter based on the set of filter weights.

[0239] In some implementations, the instructions, when executed by a processor, may also cause the processor to generate border padding for unavailable pixels in the reference block and the template associated with the reference block. In some implementations, a shift-invariant weighting filter may be generated based on the border padding to generate the filtered prediction block.

[0240] In some implementations, the instructions, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that IBC mode has been selected for the current block in response to a first flag indicating that IBC mode has been selected for the current block. In some implementations, the first flag may be a CU-level flag. In some implementations, the instructions, when executed by a processor to identify whether FIBC mode is enabled for the current block, cause the processor to identify that FIBC mode is enabled for the current block in response to a second flag indicating that FIBC mode is enabled for the current block. In some implementations, the second flag may be associated with an SPS, a PH, a PPS, or a slice header.

[0241] In some implementations, to identify whether FIBC mode is enabled for the current block, the instruction, when executed by the processor, causes the processor to identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that FIBC mode is enabled for the current block.

[0242] In some implementations, to identify whether FIBC mode is enabled for the current block, the instructions, when executed by a processor, can cause the processor to identify that the FIBC tool is enabled in response to an IBC high-level flag indicating that the IBC tool is enabled. In some implementations, to identify whether FIBC mode is enabled for the current block, the instructions, when executed by a processor, can also cause the processor to identify that FIBC mode is enabled for the current block in response to a CU-level flag indicating that IBC mode is selected for the current block.

[0243] In some implementations, if a CU is coded using IBC merge mode, FIBC mode may be identified as valid. In some implementations, the filtered prediction block may be identified based on a block vector inherited from the previous current block.

[0244] In some implementations, in response to the previous current block being predicted in IBC mode or intraTMP mode, the current block can be coded based on the reference block pointed to by the inherited block vector. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being available for filter determination, the filtered predicted block is determined. In some implementations, in response to the previous current block being predicted based on filtered IBC mode or filtered intraTMP mode and the reference block and template region for the current block being unavailable for filter determination, the current block can be coded based on the reference block.

[0245] In some implementations, if the current block is coded using IBC-LIC or CIIP or GPM, FIBC mode may be identified as valid when the CU is coded.

[0246] The foregoing description of the examples demonstrates the general nature of the present disclosure, such that those skilled in the art can readily modify and / or adapt various applications of these examples by applying their knowledge in the art without undue experimentation and without departing from the general concepts of the disclosure. Accordingly, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed examples based on the teaching and guidance presented herein. It is to be understood that the phraseology or terminology used herein is for purposes of description, not limitation, as would be interpreted by one of ordinary skill in the art based on the teaching and guidance.

[0247] The embodiments of the present disclosure are described above by functional building blocks illustrating the realization of specified functions and relationships thereof. For convenience of explanation, the boundaries of these functional building blocks are arbitrarily defined herein. Alternative boundaries can be defined as long as the specified functions and relationships thereof are appropriately performed.

[0248] The Summary and Abstract sections may describe one or more exemplary embodiments of the present disclosure contemplated by the inventors, but are not all, and are therefore not intended to limit the scope of the present disclosure and the appended claims in any way.

[0249] Various functional blocks, modules, and steps are disclosed above. The arrangements provided are exemplary and not limiting. Thus, the functional blocks, modules, and steps may be rearranged or combined in ways different from the examples provided above. Similarly, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permissible.

[0250] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. A decoding method by a decoder, comprising: The processor identifying whether a filtered intra block copy (FIBC) mode is enabled for the current block; In response to the FIBC mode being enabled for the current block, the processor generates a filtered prediction block by applying a shift-invariant weighting filter to a reference block; the processor decoding the current block based on the filtered prediction block.

2. The decoding method comprises: the processor selecting a set of filter weights for the shift-invariant weighting filter, the set of filter weights minimizing a mean square error (MSE) between a filtered reference template associated with the reference block and a current template associated with the current block; generating the shift-invariant weighting filter based on the set of filter weights. The decoding method of claim 1 .

3. The decoding method comprises: the processor further generating border padding for unavailable pixels in the reference block and a template associated with the reference block; the shift-invariant weighting filter is generated based on the boundary padding. The decoding method of claim 1 .

4. The decoding method comprises: the processor further comprising parsing the bitstream to identify at least one flag. The decoding method of claim 1 .

5. The processor identifying whether the FIBC mode is enabled for the current block includes: In response to a first flag indicating that an intra block copy (IBC) mode has been selected for the current block, the processor identifies that the IBC mode has been selected for the current block, the first flag being a coding unit (CU) level flag; and and in response to a second flag indicating that the FIBC mode is enabled for the current block, the processor identifying that the FIBC mode is enabled for the current block, wherein the second flag is associated with a sequence parameter set (SPS), a picture header (PH), a picture parameter set (PPS), or a slice header.

5. The decoding method according to claim 4.

6. The processor identifying whether the FIBC mode is enabled for the current block includes: In response to a coding unit (CU) level indicating that the FIBC mode is enabled for the current block, the processor identifies that the FIBC mode is enabled for the current block.

5. The decoding method according to claim 4.

7. The processor identifying whether the FIBC mode is enabled for the current block includes: In response to the IBC high level flag indicating that the IBC tool is valid, the processor identifies that the IBC tool is valid; and in response to a coding unit (CU) level flag indicating that an IBC mode has been selected for the current block, the processor identifying that the FIBC mode is enabled for the current block.

5. The decoding method according to claim 4.

8. If the coding unit (CU) is coded using an IBC merge mode, the FIBC mode is identified as valid; the filtered prediction block is identified based on a block vector inherited from a previous current block; The decoding method of claim 1 .

9. In response to the previous current block being predicted by an IBC mode or an intra-template matching (intraTMP) mode, decoding the current block based on a reference block pointed to by the inherited block vector; determining the filtered predicted block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and a reference block and a template region for the current block being available for filter determination; and decoding the current block based on the reference block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and the reference block and a template region for the current block being unavailable for filter determination. The decoding method according to claim 8.

10. If the current block is coded using IBC Linear Illumination Compensation (IBC-LIC), or a combination of IBC and Intra Prediction (CIIP), or a combination of IBC and Geometry Partition Mode (GPM), the FIBC mode is identified as valid when a coding unit (CU) is coded. The decoding method of claim 1 .

11. 1. A system comprising: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: Identifying whether a filtered intra block copy (FIBC) mode is enabled for the current block; generating a filtered prediction block by applying a shift-invariant weighting filter to a reference block in response to the FIBC mode being enabled for the current block; decoding the current block based on the filtered prediction block.

12. The instructions stored in the memory, when executed by the processor, cause the processor to: selecting a set of filter weights for the shift-invariant weighting filter, the set of filter weights minimizing a mean squared error (MSE) between a filtered reference template associated with the reference block and a current template associated with the current block; generating the shift-invariant weighting filter based on the set of filter weights. The system of claim 11.

13. The instructions stored in the memory, when executed by the processor, cause the processor to: generating border padding for unavailable pixels in the reference block and a template associated with the reference block; the shift-invariant weighting filter is generated based on the boundary padding. The system of claim 11.

14. To identify whether the FIBC mode is enabled for the current block, the instructions stored in the memory, when executed by the processor, cause the processor to: identifying that the IBC mode has been selected for the current block in response to a first flag indicating that the IBC mode has been selected for the current block, the first flag being a coding unit (CU) level flag; identifying that the FIBC mode is enabled for the current block in response to a second flag indicating that the FIBC mode is enabled for the current block, the second flag being associated with a sequence parameter set (SPS), a picture header (PH), a picture parameter set (PPS), or a slice header. The system of claim 13.

15. To identify whether the FIBC mode is enabled for the current block, the instructions stored in the memory, when executed by the processor, cause the processor to: and identifying that the FIBC mode is enabled for the current block in response to a coding unit (CU) level indicating that the FIBC mode is enabled for the current block. The system of claim 13.

16. To identify whether the FIBC mode is enabled for the current block, the instructions stored in the memory, when executed by the processor, cause the processor to: identifying the FIBC tool as valid in response to the IBC high-level flag indicating that the IBC tool is valid; and identifying that the FIBC mode is enabled for the current block in response to a coding unit (CU) level flag indicating that the IBC mode is selected for the current block. The system of claim 13.

17. If the coding unit (CU) is coded using an IBC merge mode, the FIBC mode is identified as valid; the filtered prediction block is identified based on a block vector inherited from a previous current block; The system of claim 11.

18. In response to the previous current block being predicted by an IBC mode or an intra-template matching (intraTMP) mode, decoding the current block based on a reference block pointed to by the inherited block vector; determining the filtered predicted block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and a reference block and a template region for the current block being available for filter determination; and decoding the current block based on the reference block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and the reference block and a template region for the current block being unavailable for filter determination.

20. The system of claim 17.

19. If the current block is coded using IBC Linear Illumination Compensation (IBC-LIC), or a combination of IBC and Intra Prediction (CIIP), or a combination of IBC and Geometry Partition Mode (GPM), the FIBC mode is identified as valid when a coding unit (CU) is coded. The system of claim 11.

20. A non-transitory computer-readable medium storing instructions that, when executed by the processor, cause the processor to: Identifying whether a filtered intra block copy (FIBC) mode is enabled for the current block; generating a filtered prediction block by applying a shift-invariant weighting filter to a reference block in response to the FIBC mode being enabled for the current block; and decoding the current block based on the filtered prediction block.

21. The instructions, when executed by the processor, cause the processor to: selecting a set of filter weights for the shift-invariant weighting filter, the set of filter weights minimizing a mean squared error (MSE) between a filtered reference template associated with the reference block and a current template associated with the current block; generating the shift-invariant weighting filter based on the set of filter weights.

21. The non-transitory computer-readable medium of claim 20.

22. The instructions, when executed by the processor, cause the processor to: generating border padding for unavailable pixels in the reference block and a template associated with the reference block; the shift-invariant weighting filter is generated based on the boundary padding.

21. The non-transitory computer-readable medium of claim 20.

23. To identify whether the FIBC mode is enabled for the current block, the instructions, when executed by the processor, cause the processor to: identifying that the IBC mode has been selected for the current block in response to a first flag indicating that the IBC mode has been selected for the current block; and identifying that the FIBC mode is enabled for the current block in response to a second flag indicating that the FIBC mode is enabled for the current block.

23. The non-transitory computer-readable medium of claim 22.

24. the first flag is a coding unit (CU) level flag, the second flag is associated with a sequence parameter set (SPS), a picture header (PH), a picture parameter set (PPS), or a slice header; 24. The non-transitory computer-readable medium of claim 23.

25. To identify whether the FIBC mode is enabled for the current block, the instructions, when executed by the processor, cause the processor to: and identifying that the FIBC mode is enabled for the current block in response to a coding unit (CU) level indicating that the FIBC mode is enabled for the current block.

23. The non-transitory computer-readable medium of claim 22.

26. To identify whether the FIBC mode is enabled for the current block, the instructions, when executed by the processor, cause the processor to: identifying the FIBC tool as valid in response to the IBC high-level flag indicating that the IBC tool is valid; and identifying that the FIBC mode is enabled for the current block in response to a coding unit (CU) level flag indicating that the IBC mode is selected for the current block.

23. The non-transitory computer-readable medium of claim 22.

27. If the coding unit (CU) is coded using an IBC merge mode, the FIBC mode is identified as valid; the filtered prediction block is identified based on a block vector inherited from a previous current block; 21. The non-transitory computer-readable medium of claim 20.

28. In response to the previous current block being predicted by an IBC mode or an intra-template matching (intraTMP) mode, decoding the current block based on a reference block pointed to by the inherited block vector; determining the filtered predicted block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and a reference block and a template region for the current block being available for filter determination; and decoding the current block based on the reference block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and the reference block and a template region for the current block being unavailable for filter determination.

28. The non-transitory computer-readable medium of claim 27.

29. If the current block is coded using IBC Linear Illumination Compensation (IBC-LIC), or a combination of IBC and Intra Prediction (CIIP), or a combination of IBC and Geometry Partition Mode (GPM), the FIBC mode is identified as valid when a coding unit (CU) is coded.

21. The non-transitory computer-readable medium of claim 20.

30. An encoding method by an encoder, comprising: The processor identifying whether a filtered intra block copy (FIBC) mode is enabled for the current block; In response to the FIBC mode being enabled for the current block, the processor generates a filtered prediction block by applying a shift-invariant weighting filter to a reference block; the processor encoding the current block based on the filtered prediction block.

31. The encoding method comprises: the processor selecting a set of filter weights for the shift-invariant weighting filter, the set of filter weights minimizing a mean square error (MSE) between a filtered reference template associated with the reference block and a current template associated with the current block; generating the shift-invariant weighting filter based on the set of filter weights.

31. The encoding method of claim 30.

32. The encoding method comprises: The processor generates border padding for unavailable pixels in the reference block and a template associated with the reference block, the shift-invariant weighting filter is generated based on the boundary padding.

31. The encoding method of claim 30.

33. The encoding method comprises: the processor further comprising parsing the bitstream to identify at least one flag.

31. The encoding method of claim 30.

34. The processor identifying whether the FIBC mode is enabled for the current block includes: In response to a first flag indicating that an intra block copy (IBC) mode has been selected for the current block, the processor identifies that the IBC mode has been selected for the current block, the first flag being a coding unit (CU) level flag; and and in response to a second flag indicating that the FIBC mode is enabled for the current block, the processor identifying that the FIBC mode is enabled for the current block, wherein the second flag is associated with a sequence parameter set (SPS), a picture header (PH), a picture parameter set (PPS), or a slice header.

34. The encoding method of claim 33.

35. The processor identifying whether the FIBC mode is enabled for the current block includes: In response to a coding unit (CU) level indicating that the FIBC mode is enabled for the current block, the processor identifies that the FIBC mode is enabled for the current block.

34. The encoding method of claim 33.

36. The processor identifying whether the FIBC mode is enabled for the current block includes: In response to the IBC high level flag indicating that the IBC tool is valid, the processor identifies that the IBC tool is valid; and in response to a coding unit (CU) level flag indicating that an IBC mode has been selected for the current block, the processor identifying that the FIBC mode is enabled for the current block.

34. The encoding method of claim 33.

37. If the coding unit (CU) is coded using an IBC merge mode, the FIBC mode is identified as valid; the filtered prediction block is identified based on a block vector inherited from a previous current block; 31. The encoding method of claim 30.

38. In response to the previous current block being predicted by an IBC mode or an intra-template matching (intraTMP) mode, encoding the current block based on a reference block pointed to by the inherited block vector; determining the filtered predicted block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and a reference block and a template region for the current block being available for filter determination; In response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and the reference block and a template region for the current block being unavailable for filter determination, encoding the current block based on the reference block.

38. The encoding method of claim 37.

39. If the current block is coded using IBC Linear Illumination Compensation (IBC-LIC), or a combination of IBC and Intra Prediction (CIIP), or a combination of IBC and Geometry Partition Mode (GPM), the FIBC mode is identified as valid when a coding unit (CU) is coded.

31. The encoding method of claim 30.

40. 1. A system comprising: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: Identifying whether a filtered intra block copy (FIBC) mode is enabled for the current block; generating a filtered prediction block by applying a shift-invariant weighting filter to a reference block in response to the FIBC mode being enabled for the current block; encoding the current block based on the filtered prediction block.

41. The instructions stored in the memory, when executed by the processor, cause the processor to: selecting a set of filter weights for the shift-invariant weighting filter, the set of filter weights minimizing a mean squared error (MSE) between a filtered reference template associated with the reference block and a current template associated with the current block; generating the shift-invariant weighting filter based on the set of filter weights.

41. The system of claim 40.

42. The instructions stored in the memory, when executed by the processor, cause the processor to: generating border padding for unavailable pixels in the reference block and a template associated with the reference block; the shift-invariant weighting filter is generated based on the boundary padding.

41. The system of claim 40.

43. To identify whether the FIBC mode is enabled for the current block, the instructions stored in the memory, when executed by the processor, cause the processor to: identifying that the IBC mode has been selected for the current block in response to a first flag indicating that the IBC mode has been selected for the current block, the first flag being a coding unit (CU) level flag; identifying that the FIBC mode is enabled for the current block in response to a second flag indicating that the FIBC mode is enabled for the current block, the second flag being associated with a sequence parameter set (SPS), a picture header (PH), a picture parameter set (PPS), or a slice header.

43. The system of claim 42.

44. To identify whether the FIBC mode is enabled for the current block, the instructions stored in the memory, when executed by the processor, cause the processor to: and identifying that the FIBC mode is enabled for the current block in response to a coding unit (CU) level indicating that the FIBC mode is enabled for the current block.

43. The system of claim 42.

45. To identify whether the FIBC mode is enabled for the current block, the instructions stored in the memory, when executed by the processor, cause the processor to: identifying the FIBC tool as valid in response to the IBC high-level flag indicating that the IBC tool is valid; and identifying that the FIBC mode is enabled for the current block in response to a coding unit (CU) level flag indicating that the IBC mode is selected for the current block.

43. The system of claim 42.

46. If the coding unit (CU) is coded using an IBC merge mode, the FIBC mode is identified as valid; the filtered prediction block is identified based on a block vector inherited from a previous current block; 40. The system of claim 39.

47. In response to the previous current block being predicted by an IBC mode or an intra-template matching (intraTMP) mode, encoding the current block based on a reference block pointed to by the inherited block vector; determining the filtered predicted block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and a reference block and a template region for the current block being available for filter determination; In response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and the reference block and a template region for the current block being unavailable for filter determination, encoding the current block based on the reference block.

47. The system of claim 46.

48. If the current block is coded using IBC Linear Illumination Compensation (IBC-LIC), or a combination of IBC and Intra Prediction (CIIP), or a combination of IBC and Geometry Partition Mode (GPM), the FIBC mode is identified as valid when a coding unit (CU) is coded.

40. The system of claim 39.

49. A non-transitory computer-readable medium storing instructions that, when executed by the processor, cause the processor to: Identifying whether a filtered intra block copy (FIBC) mode is enabled for the current block; generating a filtered prediction block by applying a shift-invariant weighting filter to a reference block in response to the FIBC mode being enabled for the current block; and encoding the current block based on the filtered prediction block.

50. The instructions, when executed by the processor, cause the processor to: selecting a set of filter weights for the shift-invariant weighting filter, the set of filter weights minimizing a mean squared error (MSE) between a filtered reference template associated with the reference block and a current template associated with the current block; generating the shift-invariant weighting filter based on the set of filter weights.

50. The non-transitory computer-readable medium of claim 49.

51. The instructions, when executed by the processor, cause the processor to: generating border padding for unavailable pixels in the reference block and a template associated with the reference block; the shift-invariant weighting filter is generated based on the boundary padding.

50. The non-transitory computer-readable medium of claim 49.

52. To identify whether the FIBC mode is enabled for the current block, the instructions, when executed by the processor, cause the processor to: identifying that the IBC mode has been selected for the current block in response to a first flag indicating that the IBC mode has been selected for the current block; and identifying that the FIBC mode is enabled for the current block in response to a second flag indicating that the FIBC mode is enabled for the current block.

52. The non-transitory computer-readable medium of claim 51.

53. the first flag is a coding unit (CU) level flag, the second flag is associated with a sequence parameter set (SPS), a picture header (PH), a picture parameter set (PPS), or a slice header; 53. The non-transitory computer-readable medium of claim 52.

54. To identify whether the FIBC mode is enabled for the current block, the instructions, when executed by the processor, cause the processor to: and identifying that the FIBC mode is enabled for the current block in response to a coding unit (CU) level indicating that the FIBC mode is enabled for the current block.

52. The non-transitory computer-readable medium of claim 51.

55. To identify whether the FIBC mode is enabled for the current block, the instructions, when executed by the processor, cause the processor to: identifying the FIBC tool as valid in response to the IBC high-level flag indicating that the IBC tool is valid; and identifying that the FIBC mode is enabled for the current block in response to a coding unit (CU) level flag indicating that the IBC mode is selected for the current block.

52. The non-transitory computer-readable medium of claim 51.

56. If the coding unit (CU) is coded using an IBC merge mode, the FIBC mode is identified as valid; the filtered prediction block is identified based on a block vector inherited from a previous current block; 50. The non-transitory computer-readable medium of claim 49.

57. In response to the previous current block being predicted by an IBC mode or an intra-template matching (intraTMP) mode, encoding the current block based on a reference block pointed to by the inherited block vector; determining the filtered predicted block in response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and a reference block and a template region for the current block being available for filter determination; In response to the previous current block being predicted by a filtered IBC mode or a filtered intraTMP mode and the reference block and a template region for the current block being unavailable for filter determination, encoding the current block based on the reference block.

57. The non-transitory computer-readable medium of claim 56.

58. If the current block is coded using IBC Linear Illumination Compensation (IBC-LIC), or a combination of IBC and Intra Prediction (CIIP), or a combination of IBC and Geometry Partition Mode (GPM), the FIBC mode is identified as valid when a coding unit (CU) is coded.

50. The non-transitory computer-readable medium of claim 49.