Method and apparatus for filter interpretation prediction

JP2026532658APending Publication Date: 2026-09-30GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026518685
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-29
Filing Date
2024-09-27
Publication Date
2026-09-30

Smart Images

  • Figure 2026532658000001_ABST
    Figure 2026532658000001_ABST
Patent Text Reader

Abstract

According to one aspect of the present disclosure, a decoding method is provided. The method may include a processor obtaining a plurality of reference blocks. The method may include the processor generating an FInter prediction for the current block based on the plurality of reference blocks in response to a determination that a Filter Inter (FInter) prediction mode has been selected for the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross-reference to Related Applications) This application claims priority to U.S. Provisional Application No. 63 / 541,716, entitled "FILTERED INTER PREDICITON FOR VIDEO CODING", filed on September 29, 2023, the entire content of which is incorporated herein by reference.

[0002] Embodiments of the present disclosure relate to video coding.

Background Art

[0003] Digital video has become mainstream and is used in a wide range of applications including digital television, video telephony, and teleconferencing. These digital video applications are enabled by advances in computing and communication technologies, as well as efficient video coding techniques. Various video coding techniques can be used to compress video data, whereby coding of video data can be performed using one or more video coding standards. Exemplary video coding standards include, but are not limited to, Versatile Video Coding (H.266 / VVC), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Expert Group (MPEG) coding, Enhanced Compression Model (ECM), and the like.

Summary of the Invention

[0004] According to one aspect of the present disclosure, a decoding method is provided. The method may include, by a processor, obtaining a plurality of reference blocks. The method may include, in response to determining that a filtered inter (FInter) prediction mode is selected for a current block, generating, by the processor, FInter prediction for the current block based on the plurality of reference blocks.

[0005] According to another aspect of the present disclosure, a decoder is provided. The decoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to: obtain a plurality of reference blocks; and in response to determining that the FInter prediction mode is selected for the current block, generate an FInter prediction of the current block based on the plurality of reference blocks.

[0006] According to another aspect of the present disclosure, a decoding apparatus is provided. The decoding apparatus may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to: obtain a plurality of reference blocks; and in response to determining that the FInter prediction mode is selected for the current block, generate an FInter prediction of the current block based on the plurality of reference blocks.

[0007] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions is provided. The instructions, when executed by a processor of a decoder, cause the processor of the decoder to: obtain a plurality of reference blocks; and in response to determining that the FInter prediction mode is selected for the current block, generate an FInter prediction of the current block based on the plurality of reference blocks.

[0008] According to an aspect of the present disclosure, an encoding method is provided. The method may include obtaining, by a processor, a plurality of reference blocks. The method may include generating, by the processor, an FInter prediction of the current block based on the plurality of reference blocks in response to determining that the FInter prediction mode is selected for the current block.

[0009] According to another aspect of the present disclosure, an encoder is provided. The encoder may include a processor and a memory for storing instructions. The memory stores instructions, which, when executed by the processor, cause the processor to: acquire a plurality of reference blocks and, in response to determining that a FInter prediction mode has been selected for the current block, generate a FInter prediction for the current block based on the plurality of reference blocks.

[0010] According to another aspect of the present disclosure, an encoding device is provided. The encoding device may include a processor and a memory for storing instructions. The memory stores instructions, which, when executed by the processor, can cause the processor to: acquire a plurality of reference blocks and, in response to determining that a FInter prediction mode has been selected for the current block, generate a FInter prediction for the current block based on the plurality of reference blocks.

[0011] In yet another aspect of the present disclosure, a non-temporary computer-readable medium for storing instructions is provided. When the instructions are executed by the encoder's processor, the processor of the encoder can be instructed to acquire a plurality of reference blocks and, in response to determining that a FInter prediction mode has been selected for the current block, to generate a FInter prediction for the current block based on the plurality of reference blocks.

[0012] In yet another aspect of this disclosure, a non-temporary computer-readable medium for storing a bitstream is provided. The bitstream may be generated according to one or more operations disclosed herein.

[0013] These exemplary embodiments are mentioned not to limit or restrict the disclosure, but to provide examples to aid its understanding. Additional embodiments are described in “Modes for Carrying Out the Invention,” where further explanation is provided. [Brief explanation of the drawing]

[0014] [Figure 1] A block diagram of an exemplary coding system according to some embodiments of this disclosure is shown. [Figure 2] An exemplary block diagram of a decoding system according to some embodiments of this disclosure is shown. [Figure 3] A detailed block diagram of an exemplary encoder in the encoding system of Figure 1, according to some embodiments of this disclosure, is shown. [Figure 4] A detailed block diagram of an exemplary decoder in the decoding system shown in Figure 2, according to some embodiments of the present disclosure, is shown. [Figure 5] The following are exemplary pictures divided into coding tree units (CTUs) according to some embodiments of the present disclosure. [Figure 6] An exemplary CTU divided into coding units (CUs) according to some embodiments of this disclosure is shown. [Figure 7] The following are schematic visualizations of the current CU block and reconstructed samples spatially adjacent and non-adjacent to the current block, according to some embodiments of the present disclosure. [Figure 8] A schematic visualization of the angular modes of VVC according to some embodiments of this disclosure is shown. [Figure 9A] Representations of spatial geometric partitioning mode (SGPM) signaling according to some embodiments of this disclosure are shown. [Figure 9B] An exemplary template used to generate a candidate list according to some embodiments of this disclosure is shown. [Figure 10]The following are schematic visualizations of inter-angle modes parallel to the GPM division boundary (parallel mode) (see (a)), inter-angle modes perpendicular to the GPM division boundary (perpendicular mode) (see (b)), and inter-predictive plane modes (see (c)), as well as SGPM with intra and inter-predictive modes (see (d)), according to some embodiments of the present disclosure. [Figure 11] This disclosure describes an adaptive SGPM blending scheme according to some embodiments of this disclosure. [Figure 12] The following are illustrative visualizations of intrablock copy (IBC) predictions according to some embodiments of this disclosure. [Figure 13] The diagrams shown here illustrate intra-template matching prediction (intraTMP) according to some embodiments of this disclosure. [Figure 14] The following are some embodiments of the ECM decoder-side intra-mode derivation (DIMD) modes. [Figure 15] The following are some embodiments of the present disclosure of template-based intra-mode derivation (TIMD) modes for ECM. [Figure 16A] The present disclosure shows low-frequency non-separated transform (LFNST) kernels for 4×N and N×4 block sizes in VVC, according to some embodiments of this disclosure. [Figure 16B] The LFNST kernels for 8×N and N×8 block sizes in VVC, according to some embodiments of this disclosure, are shown. [Figure 17] Various intra-angle prediction modes according to some embodiments of this disclosure are shown. [Figure 18A] The LFNST kernels for 4×N and N×4 block sizes in the ECM, according to some embodiments of this disclosure, are shown. [Figure 18B] The LFNST kernels for 8×N and N×8 block sizes in the ECM, according to some embodiments of this disclosure, are shown. [Figure 18C] The following are some embodiments of the LFNST kernel for a 16x16 block size in an ECM. [Figure 19] The following are illustrative visualizations of interprediction modes relating to several aspects of this disclosure. [Figure 20] The following are illustrative visualizations of the spatial support of learned filters according to some embodiments of this disclosure. [Figure 21] This disclosure shows illustrative visualizations of reference blocks / templates for FInter prediction according to some embodiments of this disclosure. [Figure 22A] A flowchart illustrating an exemplary decoding method according to some embodiments of this disclosure is shown. [Figure 22B] A flowchart illustrating an exemplary decoding method according to some embodiments of this disclosure is shown. [Figure 23A] A flowchart of an exemplary encoding method according to some embodiments of this disclosure is shown. [Figure 23B] A flowchart of an exemplary encoding method according to some embodiments of this disclosure is shown. [Modes for carrying out the invention]

[0015] The drawings are incorporated into the specification and constitute part thereof, illustrating embodiments of the disclosure, and together with the specification, further illustrate the principles of the disclosure and enable those skilled in the art to implement the disclosure.

[0016] Embodiments of this disclosure will be described with reference to the drawings.

[0017] While several configurations and arrangements are discussed, it should be understood that these are for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of this disclosure. It will also be apparent to those skilled in the art that this disclosure can be applied to a variety of other uses.

[0018] References in the specification such as “one embodiment,” “embodiment,” “exemplary embodiment,” “several embodiments,” and “specific embodiment” indicate that the described embodiment may include certain features, structures, or characteristics, but it should be noted that not all embodiments necessarily include those specific features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same embodiment. Moreover, if certain features, structures, or characteristics are described in relation to an embodiment, whether explicitly described or not, applying such features, structures, or characteristics in relation to other embodiments would be within the knowledge of those skilled in the art.

[0019] In general, terms can be understood at least partially from their usage in context. For example, the term “one or more” as used herein may, at least partially depending on the context, be used in a singular sense to describe any feature, structure, or characteristic, or in a plural sense to describe a combination of features, structures, or characteristics. Similarly, terms such as “one,” “one kind,” or “the said” can also be understood, at least partially depending on the context, to convey either singular or plural use. In addition, the term “based on” is not necessarily intended to convey an exclusive set of factors, and can also be understood, at least partially depending on the context, to allow for the presence of additional factors that are not necessarily explicitly described.

[0020] Various embodiments of video coding systems are described with reference to various devices and methods. These devices and methods are described in the following “Modes for Carrying Out the Invention” and are illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as “Elements”). These Elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such Elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the overall system.

[0021] The techniques described herein can be used in a variety of video coding applications. As described herein, video coding includes both encoding and decoding of video. Video encoding and decoding can be performed in blocks. For example, encoding / decoding processes such as transformation, quantization, prediction, in-loop filtering, and reconstruction can be performed on an encoding block, a transformation block, or a prediction block. As described herein, the block to be encoded / decoded is referred to as the “current block.” For example, the current block may represent an encoding block, a transformation block, or a prediction block according to the current encoding / decoding process. Furthermore, as used in this disclosure, the term “unit” is understood to refer to a basic unit for performing a particular encoding / decoding process, and the term “block” is understood to refer to a sample array of a given size. Unless otherwise specified, “block” and “unit” may be used interchangeably.

[0022] Figure 1 shows a block diagram of an exemplary encoding system 100 according to some embodiments of the present disclosure. Figure 2 shows a block diagram of an exemplary decoding system 200 according to some embodiments of the present disclosure. Each system 100 or 200 may be applied to or integrated into a variety of data-processing systems and devices, such as computers and wireless communication devices. For example, system 100 or 200 may be all or part of a mobile phone, desktop computer, laptop computer, tablet, in-vehicle computer, game console, printer, positioning device, wearable electronic device, smart sensor, virtual reality (VR) device, augmented reality (AR) device, or any other suitable electronic device with data processing capabilities. As shown in Figures 1 and 2, system 100 or 200 may include a processor 102, memory 104, and interface 106. These components are shown connected to each other by a bus, but other types of connections are also permitted. It is understood that system 100 or 200 may include any other suitable components for performing the functions described herein.

[0023] Processor 102 may include microprocessors such as graphics processing units (GPUs), image signal processors (ISPs), central processing units (CPUs), digital signal processors (DSPs), tensor processing units (TPUs), vision processing units (VPUs), neural processing units (NPUs), synergistic processing units (SPUs), or physical processing units (PPUs), microcontroller units (MCUs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gate logic, discrete hardware circuits, and other suitable hardware configured to perform various functions described throughout this disclosure. Although only one processor is shown in Figures 1 and 2, it will be understood that multiple processors may be included. Processor 102 may be a hardware device having one or more processing cores. Processor 102 may execute software. Software should be broadly interpreted to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or other names. Software may include computer instructions written in interpreted languages, compiled languages, or machine code. Other techniques for instructing hardware are also permitted under the broad category of software.

[0024] Memory 104 can broadly include both memory (also known as primary / system memory) and storage (also known as secondary memory). For example, memory 104 can include random access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferroelectric RAM (FRAM®), electrically erasable programmable ROM (EEPROM), hard disk drives (HDDs) such as compact disk read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash drives, solid state drives (SSDs), or any other medium that can be used to transmit or store desired program code in the form of instructions accessible and executable by processor 102. Broadly speaking, memory 104 can be embodied by any computer-readable medium such as non-temporary computer-readable medium. Although only one memory is shown in Figures 1 and 2, it is understood that multiple memories may be included.

[0025] Interface 106 may broadly include data and communication interfaces configured to send and receive signals in the process of sending and receiving information with other external network elements. For example, interface 106 may include input / output (I / O) devices and wired or wireless transceivers. Although only one interface is shown in Figures 1 and 2, it is understood that multiple interfaces may be included.

[0026] The processor 102, memory 104, and interface 106 may be implemented in various forms in system 100 or 200 to perform video coding functions. In some embodiments, the processor 102, memory 104, and interface 106 of system 100 or 200 are implemented (e.g., integrated) on one or more system-on-chip (SoCs). In one example, the processor 102, memory 104, and interface 106 may be integrated into an application processor (AP) SoC that handles application processing in an operating system (OS) environment, including the execution of video coding and decoding applications. In another example, the processor 102, memory 104, and interface 106 may be integrated into a dedicated processor chip for video coding, such as a GPU or ISP chip dedicated to image and video processing in a real-time operating system (RTOS).

[0027] As shown in Figure 1, in the encoding system 100, the processor 102 may include one or more modules, such as an encoder 101. Although Figure 1 shows that the encoder 101 resides within a single processor 102, it is understood that the encoder 101 may include one or more submodules that can be implemented on different processors located in close proximity to each other or at a distance from each other. The encoder 101 (and any corresponding submodule or subunit) may be a hardware unit of the processor 102 designed for use with other components (e.g., part of an integrated circuit), or a software unit implemented by the processor 102 by executing at least a portion of a program (e.g., instructions). The program instructions may be stored in a computer-readable medium such as memory 104, and when executed by the processor 102, they may perform processing having one or more functions related to video encoding, such as picture splitting, inter-prediction, intra-prediction, transformation, quantization, filtering, and entropy coding, as will be described in detail below.

[0028] Similarly, as shown in Figure 2, in the decoding system 200, the processor 102 may include one or more modules, such as the decoder 201. Although Figure 2 shows that the decoder 201 resides within a single processor 102, it is understood that the decoder 201 may include one or more submodules that can be implemented on different processors located in close proximity to each other or at a distance from each other. The decoder 201 (and any corresponding submodule or subunit) may be a hardware unit of the processor 102 designed for use with other components (e.g., part of an integrated circuit), or a software unit implemented by the processor 102 by executing at least a portion of a program (e.g., instructions). The program instructions may be stored in a computer-readable medium such as memory 104, and when executed by the processor 102, they may perform processing having one or more functions related to video decoding, such as entropy decoding, inverse quantization, inverse transform, inter-prediction, intra-prediction, filtering, etc., as will be described in detail below.

[0029] Figure 3 shows a detailed block diagram of an exemplary encoder 101 in the encoding system 100 of Figure 1, according to some embodiments of the present disclosure. As shown in Figure 3, the encoder 101 may include a splitting module 302, an inter-prediction module 304, an intra-prediction module 306, a transform module 308, a quantization module 310, an inverse quantization module 312, an inverse transform module 314, a filter module 316, a buffer module 318, and an encoding module 320. Each element shown in Figure 3 is shown independently to represent a different characteristic function in the video encoder, and it is understood that this does not mean that each component is formed by separate hardware or a single software configuration unit. That is, for convenience of explanation, each element is listed and included as an element, and at least two elements may be combined to form a single element, or one element may be divided into multiple elements to function. It is also understood that some of the elements may not be essential elements for performing the functions described in the present disclosure, but may be optional elements for improving performance. Furthermore, it is understood that these elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the encoder 101.

[0030] The splitting module 302 may be configured to split a video input picture into at least one processing unit. The picture may be a video frame or a video field. In some embodiments, the picture includes an array of lumar samples in monochrome format, or an array of lumar samples and two corresponding arrays of chroma samples. The processing unit may be a prediction unit (PU), a transformation unit (TU), or an encoding unit (CU). The splitting module 302 may encode the picture by splitting the picture into a plurality of combinations of encoding units, prediction units, and transformation units, and selecting a combination of encoding units, prediction units, and transformation units based on a predetermined criterion (e.g., a cost function).

[0031] Like H.265 / HEVC, H.266 / VVC is a block-based hybrid spatial and temporal prediction coding scheme. As shown in Figure 5, during coding, the input picture 500 is first divided into square blocks—CTU502—by the partitioning module 302. For example, a CTU502 may be a 128×128 pixel block. As shown in Figure 6, each CTU502 in the input picture 500 may be divided into one or more CU602 by the partitioning module 302, which can be used for prediction and transformation. Unlike H.265 / HEVC, in H.266 / VVC, a CU602 can be rectangular or square and can be coded without being further divided into prediction or transformation units. For example, as shown in Figure 6, dividing a CTU502 into CU602 may include quadtree partitioning (shown by solid lines), binary tree partitioning (shown by dashed lines), and ternary tree partitioning (shown by dotted lines). According to some embodiments, each CU602 may be the same size as its root CTU502, or it may be a subdivision of the root CTU502 that is at least a 4x4 block.

[0032] Referring to Figure 3, the inter-prediction module 304 may be configured to perform inter-prediction on prediction units, and the intra-prediction module 306 may be configured to perform intra-prediction on prediction units. A decision can be made to use inter-prediction or perform intra-prediction on a prediction unit, and specific information (e.g., intra-prediction mode, motion vector, reference picture, etc.) can be determined according to each prediction method. In this case, the processing unit for performing the prediction may differ from the processing unit for determining the prediction method and specific content. For example, the prediction method and prediction mode may be determined in the prediction unit, and the transformation may be performed in the transformation unit. The residual coefficients in the residual block between the generated prediction block and the original block may be input to the transformation module 308. Furthermore, prediction mode information, motion vector information, etc., used for prediction, along with residual coefficients or quantization levels, can be encoded into a bitstream by the encoding module 320. It is understood that in certain encoding modes, the original block can be encoded directly without generating a prediction block via the prediction modules 304 or 306. It is also understood that in certain encoding modes, prediction, transformation, and / or quantization may be skipped.

[0033] In some embodiments, the interpretation module 304 may predict prediction units based on information about at least one picture that precedes or follows the current picture, and possibly based on information about a subregion already encoded within the current picture. The interpretation module 304 may include submodules such as a reference picture interpolation module, a motion prediction module, and a motion compensation module (not shown). For example, the reference picture interpolation module may receive reference picture information from the buffer module 318 and generate pixel information for an integer number of pixels from the reference picture. For luminance pixels, pixel information for an integer number of pixels in 1 / 4 pixel units may be generated using a discrete cosine transform (DCT) based 8-tap interpolation filter with varying filter coefficients. For chrominance signals, pixel information for an integer number of pixels in 1 / 8 pixel units may be generated using a DCT-based 4-tap interpolation filter with varying filter coefficients. The motion prediction module may perform motion prediction based on the reference picture interpolated by the reference picture interpolation unit. Various methods can be used to compute motion vectors, including the Full Search-Based Block Matching Algorithm (FBMA), the Three-Step Search (TSS), and the Novel Three-Step Search Algorithm (NTS). Motion vectors may have motion vector values ​​in units of 1 / 2, 1 / 4, or 1 / 16 pixels, or integer pixels (integer pel), based on interpolated pixels. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods can be used as motion prediction methods, including the skip method, merge method, advanced motion vector prediction (AMVP) method, and intra-block copy method.

[0034] Continuing to refer to Figure 3, in some embodiments, the intra-prediction module 306 may generate prediction units based on information of reference pixels around the current block (which is pixel information within the current picture). Reference pixels may be located on reference lines not adjacent to the current block. If a neighboring block of the current prediction unit is a block on which inter-prediction has been performed, and therefore the reference pixel is a pixel on which inter-prediction has been performed, then the reference pixels included in the inter-predicted block may be used instead of the reference pixel information of the neighboring block on which intra-prediction has been performed. That is, if a reference pixel is unavailable, at least one of the available reference pixels may be used instead of the unavailable reference pixel information. In intra-prediction, the prediction modes may include an angle prediction mode that uses reference pixel information depending on the prediction direction, and a non-angle prediction mode that does not use direction information when performing prediction. The mode for predicting luminance information may differ from the mode for predicting chrominance information, and the intra-prediction mode information or predicted luminance signal information used to predict luminance information may be used to predict chrominance information. If the size of the prediction unit is the same as the size of the transformation unit when an intra-prediction is performed, the intra-prediction for the prediction unit may be performed based on the left pixel, the top-left pixel, and the top pixel of the prediction unit. However, if the size of the prediction unit is different from the size of the transformation unit when an intra-prediction is performed, the intra-prediction may be performed using reference pixels based on the transformation unit.

[0035] An intra-prediction method may generate a prediction block after applying an adaptive intra-smoothing (AIS) filter to a reference pixel according to the prediction mode. The type of AIS filter applied to the reference pixel can vary. To perform the intra-prediction method, the intra-prediction mode of the current prediction unit may be predicted from the intra-prediction modes of prediction units in the vicinity of the current prediction unit. When the prediction mode of the current prediction unit is predicted using mode information predicted from neighboring prediction units, if the intra-prediction mode of the current prediction unit is the same as that of the neighboring prediction units, information indicating that the prediction mode of the current prediction unit is the same as that of the neighboring prediction units may be transmitted using predetermined flag information. If the prediction mode of the current prediction unit and the prediction modes of the neighboring prediction units are different, the prediction mode information of the current block may be encoded with additional flag information.

[0036] As shown in Figure 3, a prediction module may be generated that performs a prediction based on the prediction module generated by prediction module 304 or 306, and a residual module may be generated that includes residual coefficient information (also referred to herein as “residuals”), which is the difference between the predicted unit and the original block. The generated residual block may be input to transformation module 308. Additional details of residuals and transformations for video coding are provided below.

[0037] In hybrid video coding systems, redundancy within the video signal is first utilized by applying an inter- or intra-prediction tool to each CU. The difference between the original sample of a CU and the predicted block of that CU is generally called the residual. Even after prediction, residuals may still be highly spatially correlated. While conditional entropy coding can capture some spatial dependence between adjacent samples, it is computationally impractical to form an entropy coding statistical model that fully utilizes the spatial correlation within residuals. In contrast, transform coding is a practical and effective method for spatially decorrelating residuals.

[0038] For example, the transformation module 308 can transform residuals using an integerized version of the two-dimensional discrete cosine transform (DCT), which can be applied separately in the horizontal and vertical directions. For an M × N residual sample block (where M is the width of the block and N is the height of the block), the transformation module 308 may obtain the transformation coefficients by applying an M × M DCT to each row to generate intermediate transformation coefficients, and then applying an N × N DCT to each column of the intermediate transformation coefficients.

[0039] For intra-encoded CUs (also referred to herein as “intraCUs”), spatially adjacent reconstructed samples are used to predict the current block, and the intra-prediction mode is signaled once for the entire CU. Each CU consists of one or more collocated coded blocks (CBs) corresponding to the color components of the video sequence. For example, consumer video typically takes a 4:2:0 chroma format, in which case each CU consists of one luma CB and two chroma CBs with a quarter of the samples of the luma CB. Intra-prediction and transform coding are performed at the predictive block (PB) level and the transform block (TB) level, respectively. Each CB consists of a single TB, except in intra-subpartition (ISP) mode and in implicit splitting. For luma CBs, the maximum side length of the TB is 64 and the minimum side length is 4. Furthermore, a luma TB is further specified as a W×H rectangular block with width W and height H, where W, H ∈ {4, 8, 16, 32, 64}. In the case of chroma CB, the maximum TB side length is 32, and chroma TB is a rectangular W×H block with width W and height H. Here, W, H ∈ {2, 4, 8, 16, 32}, but to address memory architecture and throughput requirements, blocks of the shapes 2×H and 4×2 are excluded.

[0040] Figure 7 shows schematic visualizations 700 of the current CU block 702, as well as reconstructed samples that are spatially adjacent to and non-adjacent to the current block, according to several aspects of the present disclosure. In Figure 7, the numbers 0, 1, 2, ... indicate the pixel line indices related to the current CU block 702.

[0041] In VVC, the intra-predicted sample of the current block is generated using a reference sample obtained from a reconfigured sample of an adjacent block. In the case of a W×H block, the reference sample is spatially adjacent to the current block and consists of a vertical line of 2·H reconfigured samples extending downward to the left of the block, a reconfigured sample in the upper left, and a horizontal line of 2·W reconfigured samples extending to the right above the current block. This "L"-shaped set of samples may be referred to as the "reference line" in this disclosure. The reference line directly adjacent to the current CU block 702 is shown as the line at index 0 in Figure 7.

[0042] Like AVC and HEVC, VVC also supports an intra-angle prediction mode. Intra-angle prediction is a directional intra-prediction method. Compared to HEVC, VVC's intra-angle prediction has been modified by improving prediction accuracy and adapting to a new partitioning framework. The former was achieved by increasing the number of angle prediction directions and using more accurate interpolation filters, while the latter was achieved by introducing a wide-angle intra-prediction mode. In VVC, the number of directional modes available in a given block has increased from 33 HEVC directions to 65 directions. VVC's 800 angle modes are shown in Figure 8.

[0043] Directions with even indices between 2 and 66 are equivalent to the angular mode directions supported in HEVC. For square-shaped blocks, the same number of angular modes are assigned to the top and left sides of the block. On the other hand, rectangular intrablocks, which do not exist in HEVC, are a central part of the VVC partitioning scheme, and additional intraprediction directions are assigned to the longer sides of the block. The additional modes assigned along the longer sides are called wide-angle intraprediction (WAIP) modes because they correspond to prediction directions with angles greater than 45° to the horizontal or vertical modes. A WAIP mode for a given mode index is defined by mapping the original direction mode to a mode with the opposite direction and an index offset equal to 1, as shown in Figure 8. For a given rectangular block, the aspect ratio, i.e., the ratio of height to width, is used to determine which angular modes are replaced by their corresponding wide-angle modes.

[0044] In the case of square blocks in VVC, each pair of horizontally or vertically adjacent predicted samples is predicted from a pair of adjacent reference samples. In contrast, WAIP extends the angular range of directional prediction beyond 45°, and therefore, for coded blocks predicted in WAIP mode, adjacent predicted samples may be predicted from non-adjacent reference samples.

[0045] In addition to the direct adjacent line of neighboring samples, one of the two non-adjacent reference lines (line 1 and line 2) shown in Figure 7 may contain an input sample for intra-prediction in VVC. In the case of ECM, more non-adjacent reference lines can be used. The use of adjacent and non-adjacent reference samples is called multiple reference line (MRL) prediction.

[0046] The intra-modes available for MRL are DC mode and angle prediction mode. However, not all of these modes can be combined with MRL for a given block. MRL modes are always coupled with modes in the most probable mode (MPM) list in VVC. This coupling means that when non-adjacent reference lines are used, the intra-prediction mode is one of the MPMs. Such a design of MPM-based MRL prediction modes is based on the observation that non-adjacent reference lines are primarily beneficial for texture patterns with sharp, strongly directional edges. In these cases, MPMs are chosen much more frequently because there is usually a strong correlation between the texture patterns of adjacent blocks and the current block. On the other hand, choosing a non-MPM for intra-prediction is an indicator that the edges are not consistently distributed across adjacent blocks, and therefore, in this case, the MRL prediction mode is expected to be of little use. Furthermore, since this mode is typically used for smooth regions, it has been observed that MRL does not provide additional coding gains when the intra-prediction mode is Planar mode. Therefore, MRL always excludes the Planar mode, which is one of the MPMs. The angle or DC prediction processing in MRL is very similar to that for directly adjacent reference lines. However, for angle modes with non-integer gradients, a DCT-based interpolation filter (DCTIF) is always used. This design choice is justified by experimental results, which are consistent with the empirical observation that MRL is primarily beneficial for edges with sharp and strong directionality, where DCTIF is better suited to retaining more high frequencies than some other filters.

[0047] From a hardware design perspective, applying multiple reference lines as proposed in earlier methods incurs the extra cost of line buffers used to hold additional reference lines. In typical hardware designs, line buffers are part of the on-chip memory architecture for image and video coding, and minimizing their on-chip area is crucial. To address this issue, MRLs are disabled and signaling is not performed for coding units associated with the upper boundary of the CTU. Thus, the extra buffer for holding non-adjacent reference lines is limited to 128, which is the width of the maximum unit size.

[0048] In some known methods, intra-prediction fusion methods have been proposed to improve the accuracy of intra-prediction. More specifically, if the current block is a rumor block, not in ISP mode, coded in non-integer gradient angle mode, and the block size (width × height) is greater than 16, then two prediction blocks generated from two different reference lines are "fused," and this prediction fusion is calculated as a weighted sum of the two prediction blocks. More specifically, the first reference line (linei) at index i is specified using the current signaling method in the bitstream, and the prediction block generated from that reference line using the selected intra-prediction mode is p(line i This is expressed as p(·), where p(·) represents the process of generating a prediction block from a reference line in a given intra-prediction mode. In known ways, the reference line i+1 The second reference line is implicitly selected. That is, the second reference line is at an index position one position further from the current block relative to the first reference line. Similarly, the predicted block generated from the second reference line is p(line i+1 It is expressed as ). The weighted sum of the two prediction blocks is obtained as follows and functions as the predictor of the current block according to equation (1).

[0049]

number

[0050] In several known methods, a spatial geometric partitioning mode (SGPM) has been proposed, which allows a CU to be divided into two parts where different intra-prediction modes are available. The new mode is conceptually similar to the geometric partitioning mode (GPM) applied to interprediction in VVCs. However, because numerous combinations of partitions and intra-prediction modes are possible, SGPM employs a different signaling mechanism. To more efficiently represent the required partition and prediction information in a bitstream, a candidate list is employed, and only the candidate index is signaled within the bitstream. As shown in the representation of SGPM signaling 900 in Figure 9A, each candidate in the list can derive a combination of one partitioning mode and two intra-prediction modes.

[0051] The template is used to generate the candidate list. An example of template 901 is shown in Figure 9B, where the template width is set to 4. In the current version of SGPM used in ECM, the template width is set to 1. For each possible combination of one partition mode and two intra-prediction modes, a prediction is generated for the template, and the partition weights are extended to the template. These combinations are ranked in ascending order of absolute transform difference sum (SATD) cost between the template prediction and reconstruction. The length of the candidate list is set to 16, and these candidates are considered the most likely SGPM combination for the current block. Both the encoder and decoder use the template to construct the same candidate list. To reduce the complexity of constructing the candidate list, both the number of possible partition modes and the number of possible intra-prediction modes (IPMs) are limited. For example, in the current version of SGPM used in ECM, the number of possible partition modes is limited to a predefined set of 26 partitions covering various partition directions and positions.

[0052] Each of the two IPM candidate lists corresponding to the two SGPM partitions is constructed by adding available IPM candidates and then pruning them as needed down to a predefined limit of three candidates. Some IPM candidates are inherited from intra-inter-GPM modes already employed in the ECM. Figure 10 is a schematic visualization of inter-angle mode parallel to the GPM partition boundary (parallel mode) (see (a)), inter-angle mode perpendicular to the GPM partition boundary (perpendicular mode) (see (b)), and inter-predictive plane mode (see (c)), as well as SGPM with intra and inter-prediction (see (d)), according to some embodiments of the present disclosure. These modes can be added as available IPM candidates for SGPM.

[0053] Furthermore, template-based intra-mode derivation (TIMD) can be used to derive intra-predictive modes that are available IPM candidates for SGPM. For example, to derive TIMD intra-predictive modes for forming the IPM of SGPM, only horizontal and vertical adjacent blocks (using the top or left template) are used.

[0054] For some CU block sizes, SGPM can be implicitly disabled. The range of block sizes in which SGPM can be used (for example, a CU level flag can be signaled to indicate whether SGPM is used or not) was originally inherited from intra-inter-GPM mode. In the current version of SGPM used in ECM, the range of block sizes has been further extended to smaller blocks of 4x8, 8x4, 4x16, and 16x4 dimensions. In summary, the block sizes in which SGPM can be used can be described by the rule 4 <= width <= 64, 4 <= height <= 64, width < height * 8, height < width * 8, width * height >= 32.

[0055] To better predict pixels at the boundary between two prediction segments, an adaptive SGPM blending scheme 1100 can be used, where a weighted average of the two prediction segments is used in the transition region around the SGPM partition, as shown in Figure 11. The width of the transition region is called the blending width, which is adaptively determined according to the block size. Signaling is not required for adaptive blending.

[0056] Referring to Figure 11, let τ be the blending width specified in the GPM tool for VVC and ECM. Then, depending on the width and height of the CU block, the adaptive SGPM blending width can be determined as follows.

[0057] If min(width, height) == 4, then 1 / 2τ is selected.

[0058] Instead, if min(width, height) == 8, then τ is selected.

[0059] Instead, if min(width, height) == 16, 2τ is selected.

[0060] Instead, if min(width, height) == 32, 4τ is selected.

[0061] Otherwise, 8τ is selected.

[0062] Figure 12 shows an exemplary visualization of an IBC prediction 1200 according to some embodiments of the present disclosure. When a CU is predicted by intrablock copy mode, a block vector (BV) is signaled to indicate which block in the same picture is copied to act as a predictor for the current block. This signaling of the block vector may be performed by signaling a block vector difference (BVD) in the bitstream, so that the block vector can be determined by adding the BVD to the block vector predictor. Alternatively, if the block vector from the previous CU exactly matches the current block vector, it may be signaled by a merge flag. Regardless of the signaling mechanism, the block vector points to a location in the same picture and indicates a sample block of the same size as the current CU that is used as the predictor block for the current CU. Since the block vector must point to a location in the current picture that has already been decoded before the current CU, some constraints may apply to the block vector. The illustration in Figure 12 summarizes the concept of IBC in HEVC and VVC, where each tiled square in the figure represents a coding tree unit (CTU). The gray shaded areas indicate already coded regions, while the white areas indicate regions to be coded. In HEVC, IBC generally allows a BV to point to any block contained within the gray shaded areas. This degree of freedom is partially restricted when the sps_entropy_coding_sync_enabled_flag is signaled within the bitstream, the purpose of which is to support wavefront parallel processing (WPP) functionality. In such cases, as indicated by the crossed-out CTUs in Figure 12, an IBC block vector cannot point to any region in two or more CTUs to the right of the current CTU in the CTU row immediately above it. IBC in VVC is significantly restricted, allowing block vectors to use only the CTU to the left of the current CTU as the reference region, which is indicated by the dotted box. Current IBC tools in ECM have an extended search range compared to VVC.

[0063] Figure 13 shows Figure 1300 of intra-template matching prediction (intraTMP) according to some embodiments of this disclosure.

[0064] Referring to Figure 13, intraTMP is an intra-prediction mode similar to IBC in that the current CU1322 is also predicted by sample blocks from the current picture. intraTMP can only be selected as a prediction mode for CUs of size 64x64 or less. However, unlike IBC, the block vector 1324 is not signaled in the bitstream in intraTMP. Instead, the decoder 201 compares a predefined L-shaped or other shaped template of reconstructed samples adjacent to the current CU1322 with a template of the same shape of a candidate predictor in a given search region. If the template is L-shaped, both adjacent samples to the left and above the current CU1322 or the intraTMP predictor 1326 are used. Let TmpW be the width of the left template region and TmpH be the width of the upper template region. Other template shapes include a left-side template containing only the left template region and an upper-side template containing only the upper template region.

[0065] The intraTMP predictor block is determined by finding the best candidate template that matches the current CU template. The best match can be determined by finding the template that minimizes the absolute difference sum (SAD) or absolute transformation difference sum (SATD), or by comparing hash values ​​between templates. The search algorithm may be exhaustive within the search domain (e.g., by scanning templates across the entire search domain with a shift in sample resolution) or fast (e.g., by performing a coarse search first, and then performing a local refinement search around the best match from the coarse search). In any case, the search algorithm is executed in exactly the same way by both the encoder and decoder, thereby the intraTMP predictor is implicitly recognized by both encoder 101 and decoder 201 without requiring signaling in the bitstream. Figure 13 shows an example of intraTMP where the current CU template and the best match template are indicated by shaded lines.

[0066] Continuing to refer to Figure 13, for the intraTMP predictor 1326 to be selected, the sample block corresponding to the intraTMP predictor 1326 must be completely contained within the search area. The search area is shown in Figure 13 by a dashed shadow. Within the current CTU 1320, the search area is limited to a rectangular sample block, with one corner bounded by the upper left corner of the current CTU 1320 and the other corner bounded by the upper left corner of the current CU 1322.

[0067] Outside the current CTU1320, the search region is limited by imposing a maximum length (searchRangeWidth 1328, searchRangeHeight 1330) on the intraTMP block vector, where searchRangeWidth 1328 and searchRangeHeight 1330 are set to be proportional to the dimensions of the current CU1322. That is, searchRangeWidth = a * BlkW and searchRangeHeight = a * BlkH, where "a" is a constant controlling the gain / complexity tradeoff, and BlkW and BlkH are the width and height of the current CU1322, respectively. Here, in the ECM-7.0 test software, "a" is set to 5. searchRangeHeight 1330 limits the length of the block vector only in the negative vertical direction (i.e., towards the top of the picture). For block vectors with a positive vertical component, the search region is limited by the bottom boundary of the current CTU row. For example, in Figure 13, the search area extends to the lower boundary of the left CTU1332, regardless of the value of searchRangeHeight 1330. Furthermore, these limitations on the search range do not apply to the current CTU1320. For example, in the case of a small CU where searchRangeWidth 1328 and searchRangeHeight 1330 may be smaller than the dimensions of the current CTU1320, the search area still extends to the upper left corner of the current CTU1320.

[0068] In addition to the constraints imposed by the search region, the intraTMP predictor 1326 and its template must consist of samples available for intra-prediction. For example, the boundary of the search region is still overridden by the boundaries of pictures, slices, or tiles. Let (currCuX,currCuY) be the coordinates of the top-left corner of the current CU relative to the current picture. Then, the left boundary of the intraTMP search region is initially intraTmpLeftBound = currCuX - searchRangeWidth. To account for picture boundaries, the left boundary is clipped to allow a sample width of TmpW in the predictor template: intraTmpLeftBound = max(intraTmpLeftBound, TmpW).

[0069] To speed up the template matching process, the search area is initially scanned horizontally or vertically in increments of 2 pixels at a time. This is also known as the search subsampling coefficient of 2. This reduces the complexity of the template matching search by a quarter. After the best match is found in the initial search, a refinement process is performed. Refinement is performed by narrowing the range around the best match and conducting a second template matching search. In ECM-7.0, the narrowed range is set to BlkH / 2.

[0070] Figure 14 shows a decoder-side intra-mode derivation (DIMD) mode for an ECM according to some embodiments of the present disclosure.

[0071] Referring to FIG. 14, in DIMD, an intra prediction mode (or a plurality of intra prediction modes) is implicitly derived from an L-shaped template 1404 of reconstructed samples adjacent to a current CU 1402 (hereinafter referred to as "template 1404"). The template 1404 has a size of 3 samples in width. A decoder 201 moves a 3×3 gradient analysis window 1406 over the template 1404. At each position, a local gradient is calculated by applying a Sobel filter. Let T be the set of 3x3 samples at one position in the template 1404 k , then the Sobel filter can be described according to formula (2).

[0072]

Numerical Formula

[0073]

Numerical Formula

[0074] The magnitude G of the local gradient k and the angle θ of the local gradient k can be estimated according to formulas (5) and (6), respectively.

[0075]

Numerical Formula

[0076] The angle θ of the local gradient k can be associated with an intra angular prediction direction IPM k . For example, an angle of 0 degrees corresponds to horizontal intra prediction mode 18. In practice, IPM k is obtained by using a high-speed implementation such as a lookup table based on G k,x and G k,yIt can be directly estimated by decoder 201. At the start of the DIMD method, an empty histogram H (with zeros entered in each entry) is initialized to a size equal to the number of intra-prediction modes. As the DIMD method performs gradient analysis on each local window, the histogram H is updated according to equation (7).

[0077]

number

[0078] Figure 15 shows a template-based intra-mode derivation (TIMD) mode for ECM according to some embodiments of the present disclosure. Referring to Figure 15, in TIMD, the intra-prediction mode (or multiple intra-prediction modes) is implicitly derived from the template 1504 above and to the left of the reconstructed sample adjacent to the current CU 1502.

[0079] A set of candidate intra-prediction modes is retrieved from the Most Probable Mode (MPM) list, which is constructed from intra-prediction modes used by adjacent CUs. Next, for each candidate intra-prediction mode, an intra-angle prediction method is used to generate a prediction for template 1504 from template reference sample 1506. The candidate intra-prediction mode that generates the template predictor that best matches template 1504 is selected as the TIMD intra-prediction mode. The best match can be determined by finding a predictor that minimizes the absolute difference sum (SAD) or absolute transformation difference sum (SATD), or by comparing the hash values ​​between the predictor and the template. Alternatively, multiple intra-prediction modes can be obtained in an order in which the SAD / SATD increases.

[0080] Matrix-weighted intra-prediction (MIP), IBC, and IntraTMP methods can be effective intra-prediction modes in ECM. Greater gains can be achieved by combining MIP with unseparated linear transform (NSPT), IBC with LFNST and NSPT, and IntraTMP with LFNST and NSPT, where the intra-prediction modes are derived as illustrated in the solution shown in Figure 4 below.

[0081] The conversion module 308 can convert the video signal within the residual block from the pixel domain to the conversion domain (e.g., the frequency domain depending on the conversion method). In some examples, the conversion module 308 can be skipped, and it is understood that the video signal does not need to be converted to the conversion domain.

[0082] The quantization module 310 may be configured to quantize the coefficients at each position within an encoding block to generate a quantization level for each position. The current block may be a residual block. That is, the quantization module 310 can perform a quantization process on each residual block. A residual block may contain N × M positions (samples), each position associated with a transformed or untransformed video signal / data, such as lumer and / or chroma information, where N and M are positive integers. In this disclosure, before quantization, the transformed or untransformed video signal at a particular position is referred to herein as a “coefficient”. After quantization, the quantized value of a coefficient is referred herein as a “quantization level” or “level”.

[0083] Quantization can be used to reduce the dynamic range of a converted or unconverted video signal, thereby allowing fewer bits to be used to represent the video signal. Quantization typically involves division by the quantization step size followed by rounding, while inverse quantization involves multiplication by the quantization step size. The quantization step size can be specified by a quantization parameter (QP). Such a quantization process is called scalar quantization. Quantization of all coefficients within a coding block can be performed independently, and this type of quantization is used in several existing video compression standards, such as H.264 / AVC and H.265 / HEVC. The QP in quantization can affect the bitrate used to encode / decode the video picture. For example, a higher QP may result in a lower bitrate, and a lower QP may result in a higher bitrate.

[0084] For an N×M coding block, a specific coding scan order can be used to convert the block's two-dimensional (2D) coefficients into a one-dimensional (1D) order for coefficient quantization and coding. Typically, the coding scan begins at the top-left corner and stops at the bottom-right corner of the coding block or the last non-zero coefficient / level in the bottom-right direction. It is understood that the coding scan order can include any appropriate order, such as a zigzag scan order, a vertical (column) scan order, a horizontal (row) scan order, a diagonal scan order, or any combination thereof. The quantization of coefficients within a coding block can utilize coding scan order information. For example, it may depend on the state of the previous quantization level along the coding scan order. To further improve coding efficiency, multiple quantizers (e.g., two scalar quantizers) can be used in the quantization module 310. Which quantizer is used to quantize the current coefficient may depend on the information preceding the current coefficient in the coding scan order. Such a quantization process is called dependent quantization.

[0085] Referring to Figure 3, the coding module 320 may be configured to encode the quantization level at each position within the coding block into a bitstream. In some embodiments, the coding module 320 may perform entropy coding on the coding block. Entropy coding may convert each quantization level into a corresponding binary representation, such as a binary bin, using various binarization methods, such as Golomb-Rice binarization. The binary representation can then be further compressed using an entropy coding algorithm. The compressed data can be added to the bitstream. In addition to quantization levels, the coding module 320 may encode various other information, such as block type information, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information input from, for example, prediction modules 304 and 306. In some embodiments, the coding module 320 may perform residual coding on the coding block to convert the quantization levels into a bitstream. For example, after quantization, there may be N×M quantization levels for an N×M block. These N × M levels can be zero or non-zero values. If the non-zero levels are not binary, they can be further binarized to binary bins, for example, using combination truncated rice (TR) and restricted EGk binarization.

[0086] Non-binary syntax elements can be mapped to binary codewords. The bijective mapping between a symbol and a codeword (usually using a simple structured code) is called binarization. Both binary symbols (also called bins) of binary syntax elements and codewords of non-binary data can be encoded using binary arithmetic coding. The core coding engine for context-adaptive binary arithmetic coding (CABAC) can support two modes of operation: a context coding mode where bins are encoded with an adaptive probabilistic model, and a less complex bypass mode using a fixed probability of 1 / 2. The adaptive probabilistic model is also called a context, and the assignment of a probabilistic model to individual bins is called context modeling.

[0087] As shown in Figure 3, the inverse quantization module 312 may be configured to inverse quantize the quantization level, and the inverse transform module 314 may be configured to inverse transform the coefficients transformed by the transform module 308. The reconstructed residual blocks generated by the inverse quantization module 312 and the inverse transform module 314 can be combined with predicted units predicted through the prediction module 304 or 306 to generate a reconstructed block.

[0088] The filter module 316 may include at least one of a deblocking filter, a sample-adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can remove block distortion caused by boundaries between blocks in the reconstructed picture. The SAO module can correct the offset relative to the original video on a pixel-by-pixel basis for the video on which deblocking has been performed. The ALF may be performed based on a value obtained by comparing the reconstructed and filtered video with the original video. The buffer module 318 may be configured to store the reconstructed blocks or pictures calculated via the filter module 316, and the reconstructed and stored blocks or pictures may be provided to the inter-prediction module 304 when inter-prediction is performed.

[0089] Figure 4 shows a detailed block diagram of an exemplary decoder 201 in the decoding system 200 of Figure 2, according to some embodiments of the present disclosure. As shown in Figure 4, the decoder 201 may include a decoding module 402, an inverse quantization module 404, an inverse transform module 406, an interpretation module 408, an intrapretation module 410, a filter module 412, and a buffer module 414. Each element shown in Figure 4 is shown independently to represent a distinct characteristic function in the video decoder, and it is understood that this does not mean that each component is formed by separate hardware or a single software configuration unit. That is, for convenience of explanation, each element is listed and included as an element, and at least two elements may be combined to form a single element, or one element may be divided into multiple elements to function. It is also understood that some of the elements may not be essential elements for performing the functions described in the present disclosure, but may be optional elements for improving performance. Furthermore, it is understood that these elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on decoder 201.

[0090] When a video bitstream is input from a video encoder (e.g., encoder 101), the input bitstream can be decoded by decoder 201 in the reverse process of the video encoder. Therefore, for the sake of simplicity, some of the decoding details described above with respect to encoding can be omitted. Decoding module 402 may be configured to decode the bitstream to obtain various information encoded in the bitstream, such as the quantization level at each position in the encoded block. In some embodiments, decoding module 402 may perform entropy decoding (decompression) (e.g., variable-length coding (VLC), context-adaptive variable-length coding (CAVLC), CABAC, syntax-based binary arithmetic coding (SBAC), PIPE coding, etc.) corresponding to the entropy coding (compression) performed by the encoder to obtain a binary representation (e.g., binary bin). Decoding module 402 may further convert the binary representation to quantization levels using Golomb-Rice binarization, including, for example, EGk binarization and combined TR, as well as restricted EGk binarization. In addition to the quantization level of the position within the conversion unit, the decoding module 402 can decode various other information such as parameters used for Golomb-Rice binarization (e.g., Rice parameters), block type information of the coding unit, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. During the decoding process, the decoding module 402 may perform rearrangement on the bitstream to reconstruct and rearrange the data from a 1D sequence into 2D reconstructed blocks via a reverse scan method based on the coding scan order used by the encoder.

[0091] The inverse quantization module 404 may be configured to inverse quantize the quantization level of each position in an encoded block (e.g., a 2D reconstructed block) to obtain the coefficients for each position. In some embodiments, the inverse quantization module 404 may also perform dependent inverse quantization based on quantization parameters provided by the encoder, such as information related to the quantizers used in dependent quantization, e.g., the quantization step size used by each quantizer.

[0092] The inverse transform module 406 may be configured to perform inverse transforms such as inverse DCT, inverse discrete sine transform (DST), and inverse KLT on the DCT, DST, and KLT, LFNST, and / or NSPT performed by the encoder, respectively, and to convert the data back from the transform region (e.g., coefficients) to the pixel region (e.g., lumens and / or chroma information). In some embodiments, the inverse transform module 406 may selectively perform transform operations (e.g., DCT, DST, KLT, LFNST, NSPT) according to a plurality of pieces of information such as the prediction method, the current block size, and the prediction direction.

[0093] For example, a separable transformation applies a one-dimensional transformation separately in the horizontal and vertical directions, while a two-dimensional inseparable transformation applies it directly to a block of input samples. One desirable property of a transformation is that the transformation vector spans the space of input samples. This means that any input vector (e.g., any combination of input sample values) can be represented by a weighted sum of transformation vectors. One requirement for a transformation to be spanning is that there are at least as many transformation vectors as there are dimensions in the input space; in other words, the number of output transformation coefficients is at least equal to the number of input samples. For example, a one-dimensional DCT in VVC is a spanning transform. And in the case of a spanned inseparable transform, if the block of input samples is M×N residuals, then the transformation also outputs an M×N block of transformation coefficients, which can be achieved by implementing (M×N)×(M×N) matrix multiplication.

[0094] To derive an inseparable transform that can produce a coding gain for a specific directional feature, a transform can be learned. For example, a representative set of residual blocks corresponding to a directional feature of interest can be grouped, and the Carunen-Lobe transform (KLT) can be computed from the covariance matrix of that set of residual blocks. This process can be repeated over K different sets of residual blocks. In this example, the overall transform kernel is derived in dimensions of (M×N)×(M×N)×K.

[0095] As described in this section, span inseparable transforms have two problems. First, they are computationally complex. Inseparable transforms are usually learned and therefore generally cannot be factored. The implementation of the span inseparable transform matrix in the example above results in a complexity of M × N multiplications per sample. The second problem is that the transform kernel occupies a large amount of storage in encoder 101 and decoder 201. In the example above, a single kernel adaptable to K different directional features has (M × N) × (M × N) × K weights. This kernel can only be applied to residual blocks of size M × N. In order to make the inseparable transform applicable to multiple block sizes, a transform kernel must be learned for each discrete block size.

[0096] In VVC, the LFNST tool was introduced with numerous modifications to address the aforementioned issues regarding span-inseparable transformations.

[0097] Firstly, according to the initial changes, the LFNST tool applies to a wide range of block sizes, but only two LFNST kernels are defined. For example, a smaller LFNST kernel applies to blocks of size 4×N or N×4 (if N≧4). A larger LFNST kernel applies to all larger block sizes (e.g., 8×8 and above).

[0098] Figure 16A shows Figure 1600 of LFNST kernels for 4×N and N×4 block sizes in VVC according to some embodiments of the present disclosure. Figure 16B shows Figure 1601 of LFNST kernels for 8×N and N×8 block sizes in VVC according to some embodiments of the present disclosure.

[0099] Figures 16A and 16B show the sample locations on which LFNST acts. For example, from the encoder's perspective, for a 4×N or N×4 block size, the top-left 4×4 sample location (shown as the shaded area in Figure 16A) is transformed by a small LFNST. The remaining sample locations (shown as the white area in Figure 16A) are ignored or "zeroed out". From the decoder's perspective, an inverse LFNST is applied to generate the top-left 4×4 sample, and the remaining samples are filled with zeros. A similar policy applies to larger block sizes, where LFNST acts on three top-left 4×4 blocks of the sample location (shown as the shaded area in Figure 16B). The remaining sample locations are zeroed out.

[0100] By using the "zero-out" policy, the size of the LFNST is significantly reduced compared to the full-size transformation applied to all sample positions. However, this is inherently lossy, and the values ​​of sample positions ignored by the LFNST cannot be recovered. If the LFNST tool were applied directly to residual samples, such a loss would be too large to be useful. However, since the LFNST is applied after a separable DCT has already been performed in the encoder and acts on the linear transformation coefficients to generate the quadratic transformation coefficients, the LFNST is called a quadratic transformation. In other words, the DCT can be considered a linear transformation. According to embodiments of this disclosure, the leftmost sample position in a block of linear transformation coefficients corresponds to the horizontal low frequencies of the DCT, and the uppermost sample position corresponds to the vertical low frequencies of the DCT. By preferentially transforming and reconstructing the upper-left sample position in decoder 201, the LFNST can reconstruct the low-frequency information from the original residuals. As mentioned above, the transformation generates coding gain due to its energy compression characteristics, and it is well established that the dispersion (energy) of the image and video signals captured by the camera is concentrated mainly in the low-frequency DCT coefficients. Therefore, while "zero-out" prevents LFNST from reconstructing any residual block losslessly, in practice, the loss can be minimized for most classes of image and video signals.

[0101] The second change is that, for both small and large LFNST kernels, the transformation applied is not a span transformation. From the perspective of encoder 101, the number of output (quadratic) coefficients is less than the number of input (linear) coefficients. For example, a smaller LFNST kernel takes 4 × 4 = 16 linear transformation coefficients as input but produces only 8 output quadratic transformation coefficients. A larger LFNST kernel takes 3 × 4 × 4 = 48 input linear transformation coefficients and outputs 8 quadratic transformation coefficients. The use of a non-spanning transform results in further reconstruction loss. However, this loss can be traded off in a controlled manner for the reduction in complexity achieved. A non-spanning transform can first be designed by the KLT method described above. Following this method, the basis vectors of the transformation correspond to eigenvectors of a covariance matrix computed from a representative set of residual blocks. These eigenvectors can be ranked in order of importance by their corresponding eigenvalues, and the most important eigenvectors are selected to construct the non-spanning transform. For example, to form a nonspan transform of a smaller LFNST kernel, we can select the eight eigenvectors that have the largest eigenvalues.

[0102] In summary, the two changes described above significantly reduce the complexity of the LFNST kernel compared to spanned non-separable transforms. For smaller blocks, using a smaller LFNST kernel reduces the potential complexity from (4 × N) × (4 × N) multiplications per transform block (for N ≥ 4) to 16 × 8 multiplications. For larger blocks, using a larger LFNST kernel reduces the potential complexity from (8 × N) × (8 × N) multiplications per transform block (for N ≥ 8) to 48 × 8 multiplications.

[0103] An LFNST kernel does not contain only one transformation matrix. Multiple transformation matrices are learned to achieve better coding gains across various image and video signals. The number of different transformation matrices is the product of the third and fourth dimensions of the LFNST kernel. A small LFNST kernel has dimensions of 16 × 8 × 2 × 4, and a large LFNST kernel has dimensions of 48 × 8 × 2 × 4. Because the specific transformation matrices of the transformation blocks are selected by a mixture of explicit signaling and implicit selection, the LFNST kernel is represented by two additional dimensions.

[0104] Explicit signaling is performed by an LFNST index signaled within the bitstream, which can take the value 0, 1, or 2. Here, 0 indicates that LFNST is not used for the transformation block, and values ​​of 1 or 2 indicate a selection in the third dimension of the LFNST kernel. The drawbacks of potential reconstruction loss due to zero-out and non-span simplification are mitigated by the explicit signaling mechanism. If the use of LFNST results in excessive reconstruction loss for the transformation block, the LFNST tool can be disabled by signaling an LFNST index of 0.

[0105] An implicit selection is achieved by restricting LFNST to coding units that use intra-prediction only. Intra-prediction generates a prediction block for the coding unit from adjacent reference samples above and to the left of the current block. The specific method of constructing the prediction block is signaled in the bitstream by the intra-prediction mode. Simple methods of intra-prediction include taking the average of the reference samples ("DC" mode) or constructing affine interpolation between several reference samples ("Planar" mode). However, most intra-prediction modes are reserved to signal the intra-angular direction in which the prediction block is constructed by assuming that the values ​​of the reference samples are replicated along a particular direction. When the intra-angular direction is used, it can be a strong hint of the directional characteristics of the residual block. The implicit selection of the LFNST transform is performed by mapping the intra-prediction mode to one of four possible values ​​of a "transformation set index," which is used to index to the fourth dimension of the LFNST kernel. The mappings used in VVC are shown in Table 1.

[0106] [Table 1]

[0107] Intra-prediction modes 0 and 1 correspond to the intra-prediction Planar mode and intra-DC prediction mode, respectively. These modes are treated as special cases by mapping them to transformation set index 0. Otherwise, the remaining intra-prediction modes correspond to the intra-angle direction 1700, partially shown in Figure 17. Intra-prediction mode 2 corresponds to diagonal intra-angle prediction from the lower left. Increasing the intra-prediction mode number corresponds to a clockwise rotation of the intra-prediction direction, with intra-prediction mode 34 corresponding to diagonal intra-angle prediction from the upper left, and intra-prediction mode 66 corresponding to diagonal intra-angle prediction from the upper right.

[0108] For intra-prediction modes greater than 34 (corresponding to intra-angle prediction directions clockwise from the diagonal direction from the top left), the selected LFNST transformation matrix is ​​applied in a transposed manner to the linear transformation coefficients. In one implementation, this can be done by scanning the linear transformation coefficients in the transpose direction before applying the LFNST transformation. For example, from the perspective of encoder 101, if the current block is predicted by intra-prediction mode 2, the linear transformation coefficients may be rearranged from a 2D pattern in the block to a 1D vector by a row-major scan before applying the selected LFNST transformation matrix T. Then, for this example, if the current block is instead predicted by intra-prediction mode 66 and the same signaled LFNST index is used, the linear transformation coefficients are rearranged to a 1D vector by a column-major scan before applying the same LFNST transformation matrix T. In another implementation, the same current block with intra-prediction mode 66 can be equivalently transformed by rearranging the rows of the transformation matrix T, while still performing a row-major scan on the linear transformation coefficients.

[0109] More generally, the application of the LFNST transformation matrix can be described as follows: the linear transformation coefficients located in the yth row and xth column are p x,y Let the LFNST transformation matrix T be expressed as follows, with dimension A × B. Here, A is the number of quadratic transformation coefficients, and B is the number of linear transformation coefficients that are not zeroed out. Then, for intra prediction modes of 34 or less, the linear transformation coefficients p are used to construct a one-dimensional vector P. x,y Any traversal order via can be defined by the following equation (8).

[0110]

number

[0111]

number

[0112]

number

[0113] By transposing the linear transformation coefficients for intra-prediction modes greater than 34, it becomes possible to share the same LFNST transformation matrix for symmetric intra-angle prediction directions.

[0114] In exploratory activities following VVC, extensions to LFNST were proposed and integrated into ECM. The LFNST tool in ECM relaxes some of the complexity reduction requirements imposed on the original LFNST tool in VVC in order to improve coding gains.

[0115] The ECM has three LFNST kernels. Similar to the LFNST tool in VVC, in most cases, the majority of the translated blocks are zeroed out, as shown in Figures 18A to 18C. For example, Figure 18A shows Figure 1800 of the LFNST kernel for 4×N and N×4 block sizes in the ECM according to some embodiments of the Disclosure. Figure 18B shows Figure 1801 of the LFNST kernel for 8×N and N×8 block sizes in the ECM according to some embodiments of the Disclosure. Figure 18C shows Figure 1803 of the LFNST kernel for 16×16 block size in the ECM according to some embodiments of the Disclosure.

[0116] In Figures 18A to 18C, the shaded areas indicate the locations of the linear transformation coefficients on which LFNST acts in the ECM, and the white areas indicate which transformation coefficient locations are zeroed out. For blocks of size 4×N or N×4 (if N≧4), a small LFNST kernel is used for the top-left 4×4 linear transformation coefficient. For blocks of size 8×N or N×8 (if N≧8), an intermediate LFNST kernel is used for the four top-left 4×4 blocks of the linear transformation coefficient. For blocks of 16×16 or larger, a large LFNST kernel is used for the six top-left 4×4 blocks of the linear transformation coefficient.

[0117] In ECM, the sizes of the LFNST kernels are 16×16×3×35 for the small LFNST kernel, 64×32×3×35 for the medium LFNST kernel, and 96×32×3×35 for the large LFNST kernel. Compared to the LFNST tool in VVC, the range of LFNST indices being signaled has increased from 2 to 3, and the number of LFNST transformation sets has increased from 4 to 35. This means that there are 35 LFNST transformation matrices for each of the three indices. The mapping from intra-prediction modes to LFNST transformation set indices is shown in Table 2. Similar to the LFNST tool in VVC, the linear transformation coefficients are transposed when the intra-prediction mode is greater than 34.

[0118] [Table 2]

[0119] The complexity burden of the LFNST tool can be evaluated in three ways. First, there is the additional storage burden imposed on decoder 201 because it must store the LFNST kernel. Second, there is the worst-case number of sample-by-sample multiplications that decoder 201 must perform if the LFNST tool is used. Third, there is the additional number of sample-by-sample multiplications that encoder 101 uses if a full search is performed against the LFNST tool. Looking at all three of these metrics, the extended LFNST proposed in ECM is more complex than the LFNST in VVC. However, in terms of the total number of sample-by-sample multiplications, the worst-case decoder complexity can still be smaller than the worst-case decoder complexity of other conversion options.

[0120] For a separable and applicable DCT matrix multiplication implementation for M×N size transformations, the number of multiplications per sample is M+N. Therefore, the worst-case complexity occurs when the value of (M+N) is maximum. In practice, complexity can be reduced by alternative implementations of DCT, such as butterfly factorization, but it is still useful to evaluate the complexity of the matrix multiplication implementation. ECM extends separable DCTs, with the largest transformation being a 128-point DCT. In this case, the worst-case complexity of the separable DCT can be 128+128=256 multiplication operations per sample.

[0121] The worst-case decoder complexity of LFNST in ECM can be evaluated by considering several different block sizes. For a fair comparison, the evaluation includes the cost of performing a linear transformation. For a 4x4 block, the linear transformation involves 4+4=8 multiplication operations per sample. LFNST involves 16x16 matrix multiplication, i.e., 16 multiplication operations per sample. Therefore, the total cost of LFNST for a 4x4 block is 24 multiplication operations per sample.

[0122] For a 4x8 block, a simple implementation of the DCT's linear transformation typically requires eight 4x4 transformations along the shorter dimension and four 8x8 transformations along the longer dimension, resulting in a total of 4+8=12 multiplication operations per sample. However, because LFNST reconstructs only the non-zero coefficient values ​​in the top-left 4x4 block of the linear transformation coefficient position, an optimized decoder can take advantage of this by performing only four 4x4 transformations along the shorter dimension and then four 4x8 transformations along the longer dimension, resulting in a total of 2+4=6 multiplication operations per sample. Since the order of separable transformations is usually fixed, in the worst-case scenario, decoder 201 can first perform four 4x8 transformations along the longer dimension. Then, decoder 201 can perform eight 4x4 transformations along the shorter dimension. This results in 4+4=8 multiplication operations per sample. LFNST is still a 16x16 matrix multiplication, but its cost is amortized over larger blocks, resulting in 8 multiplication operations per sample. Therefore, the worst-case cost of LFNST in a 4x8 block is 16 multiplications per sample. The same principle usually applies to 4xN or Nx4 block sizes as well. Thus, the number of multiplication operations required per sample in a 4xN or Nx4 block is always less than or equal to the number of multiplication operations required per sample in a 4x4 block.

[0123] For an 8x8 block, the linear transformation involves 8+8=16 multiplications per sample. LFNST consists of 64x32 matrix multiplications, i.e., 32 multiplications per sample. Therefore, the total cost of LFNST for an 8x8 block is 48 multiplications per sample.

[0124] For the 8x16 block, again assume that decoder 201 takes advantage of the zero-out characteristic of LFNST reconstruction. Only the top-left 8x8 block of the linear transformation coefficient position is non-zero. Decoder 201 can take advantage of this by performing only 8 8x8 transformations along the short dimension. Next, decoder 201 can perform 8 8x16 transformations along the long dimension. This requires a total of 4+8=12 multiplications per sample. Alternatively, decoder 201 can first perform 8 8x16 transformations along the long dimension. Next, decoder 201 can perform 16 8x8 transformations along the short dimension. This may require 8+8=16 multiplications per sample. LFNST adds an additional (64×32) / (8×16)=16 multiplications per sample, resulting in an overall worst-case complexity of 32 multiplication operations per sample. As mentioned earlier, the number of multiplications per sample for an 8xN or Nx8 block is always less than or equal to the number of multiplications per sample for an 8x8 block.

[0125] For a 16x16 block, the zero-out characteristic of the LFNST reconstruction means that only six 4x4 blocks (as shown in Figure 18C) of the linear transformation coefficients in the pattern have non-zero values. For clarity, let's assume a looser pattern where the top-left 12x12 block of the linear transformation position may have non-zero values. Firstly, the decoder can take advantage of this by performing only 12 12x16 transformations in one dimension. Secondly, decoder 201 can perform 16 12x16 transformations in the second dimension, which involves 9 + 12 = 21 multiplications per sample. The LFNST involves (96 × 32) / (16 × 16) = 12 multiplications per sample, resulting in an overall complexity of 33 multiplication operations per sample.

[0126] For an M×N block (where M, N≧16), decoder 201 can first perform 12 12×M transformations along one dimension. Next, decoder 201 can perform M 12×N transformations in the second dimension, resulting in (12×12) / N+12 multiplications per sample to perform a separable DCT. Then, at the minimum value of N=16, the worst-case complexity arises, i.e., 21 multiplications per sample, which is the same as the complexity of a 16×16 block. LFNST adds an additional (96×32) / (M×N) multiplications per sample, which is always less than or equal to the number of multiplications per sample for a 16×16 block. Therefore, for larger M×N block sizes, the overall complexity of LFNST in ECM is always less than or equal to the number of multiplications per sample for a 16×16 block.

[0127] A detailed evaluation of the decoder complexity of LFNST in ECM across different block sizes revealed that the worst-case complexity is 48 multiplications per sample (occurring in the case of an 8x8 block). This worst-case complexity includes the cost of performing a separable DCT using a matrix multiplication implementation, but due to optimizations possible by zeroing out from LFNST, it is significantly smaller than the worst-case complexity of performing the separable DCT alone (estimated to be 256 multiplications per sample). Assuming a more realistic implementation of the separable DCT with butterfly factorization, the worst-case complexity of LFNST still occurs in an 8x8 block, where the cost of 32 multiplications per sample from LFNST is added to the cost of the butterfly DCT. In this case, LFNST can be the worst-case compared to the cost of the butterfly DCT applied separately to a 256x256 block.

[0128] As seen above, the use of non-separable quadratic transformations can significantly reduce complexity by using zero-out on the selected linear transformation coefficient region. However, further coding can be done using NSPT. Initial investigations into non-separable linear transformations revealed that the implemented transformations were very complex, and significant gains (a 3.43% decrease in average rate according to the Bjontegaard metric) could be achieved despite the kernel weights being overfitted to the test dataset.

[0129] This specification proposes practical implementations of NSPT. For example, NSPT can be applied only to certain small block sizes such as 4x4, 4x8, 8x4, and 8x8. For these block sizes, NSPT is an alternative to both linear transformations and LFNST. Similar to LFNST, the NSPT kernel is also trained, and the selection of the appropriate matrix for a given block is guided by both a signaling index and implicit selection via an intra-predictive mode. Four types of NSPT kernels are proposed. For 4x4 blocks, a small NSPT kernel with a size of 16x16x3x35 is used. For 4x8 and 8x4 blocks, an intermediate NSPT kernel with a size of 32x20x3x35 is used. For 8x8 blocks, a large NSPT kernel with a size of 64x32x3x35 is used.

[0130] According to this disclosure, a zero-out method can be defined as a reduction in the input dimension of the transform kernel, which corresponds to the first dimension in the representation of the transform kernel dimensions in this disclosure. A reduction in the input dimension of a forward transform is equivalent to a reduction in the range of support of the transform. For example, zero-out of an LFNST corresponds to a reduction in the number of linear transform coefficients of a DCT on which the forward LFNST acts to generate quadratic transform coefficients. In the proposed NSPT, since the transform acts directly on the residual coefficients, a reduction in the first dimension of the NSPT kernel includes a reduction in the number of residual coefficients on which the forward NSPT acts to generate linear transform coefficients. In the known scheme described above, the size of the first dimension of each NSPT kernel is always equal to the number of samples in the block, so zero-out as defined in this disclosure is not used. However, in the known proposal, zero-out is instead defined as a reduction in the output dimension of the transform kernel, which corresponds to the second dimension in the representation of the transform kernel dimensions in this disclosure. Such a definition is not ambiguous in the proposal, because no reduction is performed on the input side of the NSPT kernel, and it illustrates the more commonly used term "zero-out". However, for consistency and clarity in this disclosure, “zero-out” is defined as describing a reduction in the input dimension of the transform kernel, while a reduction in the output dimension is labeled as a nonspan or lossy transform. For intermediate and large NSPT kernels, the second dimension is smaller than the first dimension, which means that the NSPT in these cases is a lossy transform.

[0131] Similar to LFNST, the NSPT index is signaled in the bitstream, and the NSPT index can take values ​​of 0, 1, 2, or 3. Here, 0 indicates that NSPT is not used in the transformation block, and values ​​1-3 indicate selection along the third dimension in the corresponding NSPT kernel. The selection along the fourth dimension of the NSPT kernel is determined by mapping from the intra-prediction mode as shown in Table 3, in the same way as in the extended LFNST in ECM. Similar to LFNST, if the intra-prediction mode is greater than 34 (meaning the intra-angle direction is clockwise with respect to the diagonal direction from the top left), the transformation input is transposed. However, in the case of NSPT, the input consists of residual coefficients rather than linear transformation coefficients.

[0132] [Table 3]

[0133] For M×N shaped residual blocks, and when the intra-prediction mode is 34 or less, the residual sample located in the yth row and xth column is r x,y Let the LFNST transformation matrix T selected from the NSPT kernel for an M×N shaped block have dimension A×B, where A is the number of linear transformation coefficients and B=M×N is the number of residual samples in the block. Then, any scan order via residual samples to construct a one-dimensional vector R can be defined by equation (11) below.

[0134]

number

[0135]

number

[0136] The forward NSPT transformation can be implemented as P=n(TR), where P is a one-dimensional vector of NSPT transformation coefficients, and n(·) represents the normalization operation required to approximate the ideal transformation represented by the integerized NSPT as a floating-point number. In the forward diagonal scan order, the transformation coefficients are written back to the transformation block. Following the same notation introduced above, and the rule that the (0,0) position corresponds to "low frequency" or "DC" in conventional DCTs, the scan order is described according to equation (13).

[0137]

number

[0138] Figure 19 shows an exemplary visualization of the interprediction mode 1900 according to several aspects of the present disclosure.

[0139] For inter-coded CUs (also referred to as interCUs in this disclosure), reconstructed samples within a temporal reference frame are used to predict the current block, and the inter-prediction mode is signaled once for the entire CU. Inter-picture prediction utilizes the temporal correlation between pictures to derive motion-compensated predictions (MCPs) for image sample blocks.

[0140] In such a block-based MCP, the video picture is divided into rectangular blocks. Assuming uniform motion within each block, for each block, a corresponding block can be found in the previously decoded picture that acts as a predictor. Figure 19 shows a general concept of an MCP based on a translational motion model. Using the translational motion model, the position of a block in the previously decoded picture is determined by a motion vector (mv x ,mv y This is shown by mv. x and mv y These specify the horizontal and vertical displacement relative to the current block position, respectively. Motion vector (mv x ,mv y The motion vector may have fractional sample precision to more accurately capture the motion of the underlying object. If the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture to derive the predicted signal. The previously decoded picture is called the reference picture and is indicated by the reference index Δt to the reference picture list. These translational motion model parameters (e.g., motion vector and reference index) are further called motion data. Modern video coding standards allow two types of picture-to-picture prediction: unidirectional prediction and bidirectional prediction. In the case of bidirectional prediction, two sets of motion data (mv x0 ,mv y0 ,Δt0) and (mv x1 ,mv y1 Two MCPs (which may be from different pictures) are generated using Δt1), and then they are combined to obtain the final MCP. Furthermore, multiple hypotheses (two or more) can be employed to form the final inter-prediction. The same or different weights can be applied to each MCP. The reference pictures that can be used in the bidirectional prediction are stored in two separate lists, List 0 and List 1. Motion data is derived in encoder 101 using motion estimation processing. Since motion estimation is not specified in the video standard, different encoders can take advantage of different complexity and quality trade-offs in their implementation.

[0141] Block motion data can correlate with adjacent blocks. To leverage this correlation, motion data is not coded directly into the bitstream but rather predictively based on adjacent motion data. Predictive coding of motion vectors can be done with Advanced Motion Vector Prediction (AMVP), where the best predictor for each motion block signals to the decoder. Furthermore, interpredictive block merging derives all motion data for a block from adjacent blocks, including both spatial and temporal blocks. Similar to intraprediction, the residuals between the original pixels and the interprediction can be further transformed and then coded into the bitstream.

[0142] Using existing techniques, reference blocks are copied directly from the reconstruction region of the previous reference frame, and if the inter-mode is selected, the weighted sum of these reference blocks acts as the predicted block for the current block. Since no spatial information between adjacent pixels is considered, prediction accuracy may be excessively limited. To overcome these and other challenges, this disclosure provides a filtered inter (FInter) prediction mode. In some implementations, the inter-motion vector (mv x ,mv y The reconstructed pixels within the reference block indicated by ) are further filtered through an online-learned filter. Instead of the reference block, the resulting filtered block (unidirectional prediction) or the sum of multiple weighted and filtered blocks (bidirectional or multiple hypotheses) is used as the predictor for the current block. In some implementations, multiple reference blocks can be determined through multiple motion data signaling to form the final predictor for the current CU. Here, the final predictor is generated by further filtering the fused combination of reference blocks with the learned filter. The following provides additional details of exemplary FInter prediction modes in relation to Figures 20 and 21.

[0143] Figure 20 shows an exemplary visualization of the spatial support of a learned filter 2000 (hereinafter referred to as "filter 2000") according to some embodiments of the present disclosure, where reference block 2106 is the motion vector (mv x ,mv y It is identified by ). Figures 20 and 21 are explained in combination.

[0144] Referring to Figure 20, the Inter prediction module 408 may apply filter 2000 for FInter prediction as shift-invariant weighting across support regions moving on the reconstructed block. Let R and F be the reference block and the filtered prediction block, respectively. x,y and F x,y Let R and F be the sample and filtered sample in the x-th column and y-th row, respectively. A filter can be applied according to equation (14).

[0145]

number

[0146] The filter 2000 may include one bias weight and five spatial weights at the central "C" position and the four adjacent "W", "N", "E", and "S" positions. In this example, the support region S = {(0,0),(-1,0),(0,1),(1,0),(0,-1)}, where (0,0), (-1,0), (0,1), (1,0), and (0,-1) represent the positions C, W, N, E, and S, respectively.

[0147] In some implementations, the support region may have different shapes and sizes. Additionally and / or alternatively, the filter may be extended by a nonlinear term. For example, the squares of the pixel values ​​at several positions within the support region S2 can be used. In one example, the support region for the square term is S2 = {(0,0)}, i.e., the value at the central position is squared. A filter including the square nonlinear term may be applied according to equation (15).

[0148]

number

[0149]

number

[0150] When calculating filter coefficients and applying the learned filter, the filter's support may extend beyond the available sample range, as shown in the blue area in Figure 21. In such cases, values ​​in these regions are generated by using boundary padding.

[0151] For bidirectional and multiple-hypothesis prediction, the inter-prediction module 408 can independently learn multiple filters by using templates around multiple reference blocks indicated by corresponding motion data and current blocks. The inter-prediction module 408 can first individually filter the reference blocks using the corresponding learned filters. Then, the inter-prediction module 408 can merge the multiple filtered reference blocks to form the final prediction. The entire bidirectional and multiple-hypothesis processing can be similar to inter-prediction, except that it is proposed to generate the final prediction using filtered reference blocks instead of the reference blocks themselves.

[0152] For example, if FInter is allowed in a Sequence Parameter Set (SPS), Picture Header (PH), Picture Parameter Set (PPS), or Slice Header, one high-level flag can be signaled. If FInter prediction mode is enabled for the current video sequence, and the current CU is predicted by FInter prediction, an additional flag is signaled to indicate whether the original FInter prediction or FInter is used.

[0153] In some implementations, FInter can completely replace the original interpretation mode. That is, the FInter high-level flag instead signals whether or not the FInter prediction mode is used for all interpretation blocks.

[0154] In some other implementations, the interpretation module 408 can determine multiple reference blocks through multiple motion data signaling to form the final predictor of the current CU. 1 ,R 2 ,…R n And, Σ k w k P=w1R is the fused combination of reference blocks for some fusion weights that satisfy =1. 1 +w2R 2 At the same time... lol n R n In this implementation, the interpretation module 408 can generate a final predictor by further filtering the fused combination using the learned filter according to equation (17).

[0155]

number

[0156]

number

[0157] For example, the interprediction module 408 may be configured to receive a bitstream from the encoder containing a reference frame, a current frame, and instructions for weighting coefficients associated with a multimedia home platform (MHP) process. The interprediction module 408 may be configured to perform the MHP process on CUs located in the current frame based on search blocks in the reference frame (e.g., the reference frame and / or reference template). In some embodiments, to perform the MHP process, the interprediction module 408 may be configured to perform template matching on CUs located in the current frame based on search blocks in the reference frame and weighting coefficients to obtain motion information. In some embodiments, to perform the MHP process, the interprediction module 408 may be configured to identify the weighting coefficient index associated with the weighting coefficients based on template matching. The interprediction module 408 may be configured to identify the weighting coefficient code of the weighting coefficients based on instructions contained in the bitstream. The interprediction module 408 may be configured to perform the interprediction process based on the current frame, reference frame, weighting coefficient index, and weighting coefficient code of the weighting coefficients and decode the bitstream.

[0158] A reconstructed block or picture formed by combining the outputs of the inverse transform module 406 and the prediction module 408 or 410 may be provided to the filter module 412. The filter module 412 may include a deblocking filter, an offset correction module, and an ALF. The buffer module 414 can store the reconstructed picture or block and use it as a reference picture or reference block for the interprediction module 408, and can also output the reconstructed picture.

[0159] Within the scope of this disclosure, the encoding module 320 and the decoding module 402 may be configured to improve coding efficiency by employing a quantization level binarization scheme having a Rice parameter suitable for the bit depth and / or bitrate when encoding the picture of the video.

[0160] Figures 22A and 22B show flowcharts of exemplary decoding methods 2200 according to some embodiments of the present disclosure. Method 2200 may be performed by a system such as a decoding system 200, a decoder 201, or an interpretation module 408 (these are just a few examples). Method 2200 may include operations 2202–2220 described below. It will be understood that some steps may be optional, some steps may be performed simultaneously, or may be performed in an order different from the order shown in Figures 22A and 22B.

[0161] Referring to Figure 22A, in 2202, the system can acquire multiple reference blocks. In some implementations, acquiring multiple reference blocks by the processor may include analyzing the bitstream to acquire multiple motion data. In some implementations, acquiring multiple reference blocks by the processor may include generating multiple reference blocks based on multiple motion data. For example, referring to Figure 4, the interpretation module 408 can acquire multiple reference blocks. In some examples, the interpretation module 408 can acquire multiple reference blocks based on motion data.

[0162] In 2204, the system may analyze the bitstream to obtain a first flag. In some implementations, the first flag may be signaled at the SPS level, PH level, PPS level, or SH level. For example, referring to Figure 4, the inter-prediction module 408 may analyze the bitstream to obtain a first flag. The first flag may indicate whether or not FInter prediction mode is currently enabled for the block.

[0163] In 2206, in response to the first flag indicating that FInter prediction mode is currently enabled for the block, the system may analyze the bitstream to obtain a second flag. For example, referring to Figure 4, if the first flag indicates that FInter prediction mode is enabled, the inter-prediction module 408 may analyze the bitstream to obtain a second flag. The second flag may indicate whether normal inter-prediction or FInter prediction is currently selected for the block.

[0164] In 2208, the system may determine whether FInter prediction mode is selected for the current block based on a second flag. For example, referring to Figure 4, the inter-prediction module 408 may determine whether FInter prediction mode is selected for the current block based on a second flag. For example, a second flag having a first value may indicate that FInter prediction is selected, and a second flag having a second value may indicate that normal inter-prediction is selected.

[0165] In 2210, the system can analyze the bitstream and obtain a flag. For example, referring to Figure 4, the interprediction module 408 can analyze the bitstream and obtain a single flag when the FInter prediction mode replaces the normal interprediction mode.

[0166] In 2212, the system may determine, based on the flag, whether or not the FInter prediction mode is selected for all inter-prediction blocks. For example, referring to Figure 4, if the single flag has a first value, the inter-prediction module 408 may determine that the FInter prediction mode is selected for all inter-prediction blocks, and if the single flag has a second value, the inter-prediction module 408 may determine that the FInter prediction mode is not selected for all inter-prediction blocks.

[0167] In 2214, the system may select a set of filter weights for the FInter prediction filter that minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. For example, referring to Figure 4, the inter-prediction module 408 predicts the template RT of the reference block according to equation (18). 1 RT 2 ,…RT n The filter weights can be learned / selected by minimizing the weighted error between each of the filters and the template of the current block. In some implementations, the interprediction module 408 may select the filter weights to minimize the mean squared error (MSE) between the filtered RT and the template of the current block according to equation (16).

[0168] Referring to Figure 22B, in 2216, the system may generate an FInter prediction filter based on the set of filter weights. In some implementations, the FInter prediction for the current block may be generated based on the FInter prediction filter. For example, referring to Figure 4, the inter-prediction module 408 may generate an FInter prediction filter based on selected filter weights.

[0169] In 2218, in response to the determination that the FInter prediction mode has been selected for the current block, the system generates an FInter prediction for the current block based on a plurality of reference blocks. In some implementations, the generation of an FInter prediction for the current block based on a plurality of reference blocks by the processor in response to the determination that the FInter prediction mode has been selected for the current block may include filtering each of the plurality of reference blocks individually using corresponding filters to obtain a plurality of filtered reference blocks. In some implementations, the generation of an FInter prediction for the current block based on a plurality of reference blocks by the processor in response to the determination that the FInter prediction mode has been selected for the current block may include merging the plurality of filtered reference blocks to obtain a plurality of merged and filtered reference blocks. In some implementations, the generation of an FInter prediction for the current block based on a plurality of reference blocks by the processor in response to the determination that the FInter prediction mode has been selected for the current block may include generating an FInter prediction for the current block based on the merged and filtered plurality of reference blocks. For example, referring to Figure 4, for bidirectional and multiple-hypothesis prediction, the inter-prediction module 408 may individually learn multiple filters by using templates around multiple reference blocks indicated by corresponding multiple motion data and current blocks. The inter-prediction module 408 may first individually filter the reference blocks using the corresponding learned filters. Then, the inter-prediction module 408 may merge the multiple filtered reference blocks to form the final prediction. The entire bidirectional and multiple-hypothesis process may be similar to inter-prediction, except that it is proposed to generate the final prediction using filtered reference blocks instead of the reference blocks themselves.In some other implementations, in response to the decision that an FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include fusing the multiple reference blocks to obtain a fusing set of reference blocks. In some other implementations, in response to the decision that an FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include filtering the fusing set of reference blocks to obtain a fusing and filtered set of reference blocks. In some other implementations, in response to the decision that an FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include generating an FInter prediction for the current block based on the fusing and filtered set of reference blocks. For example, referring to Figure 4, the interpretation module 408 can determine multiple reference blocks through multiple motion data signaling to form the final predictor of the current CU. The reference blocks are R. 1 ,R 2 ,…R n And, Σ k w k P=w1R is the fused combination of reference blocks for some fusion weights that satisfy =1. 1 +w2R 2 At the same time... lol n R n In this implementation, the interpretation module 408 can generate a final predictor by further filtering the fused combination using the learned filters according to equation (17).

[0170] In 2220, in response to the determination that the FInter prediction mode is not selected for the current block, the system generates a normal inter prediction for the current block based on multiple reference blocks. For example, referring to Figure 4, if the FInter prediction mode is not selected for the current block, the inter prediction module 408 may generate a normal inter prediction for the current block.

[0171] Figures 23A and 23B show flowcharts of exemplary coding methods 2300 according to some embodiments of the present disclosure. Method 2300 may be performed by a system such as, for example, a coding system 100, an encoder 101, or an interpretation module 304 (these are just a few examples). Method 2300 may include operations 2302–2320 described below. It will be understood that some steps may be optional, some steps may be performed simultaneously, or may be performed in an order different from the order shown in Figures 23A and 23B.

[0172] Referring to Figure 23A, in 2302, the system can acquire multiple reference blocks. In some implementations, acquiring multiple reference blocks by the processor may include encoding multiple motion data. In some implementations, acquiring multiple reference blocks by the processor may include generating multiple reference blocks based on multiple motion data. For example, referring to Figure 3, the interpretation module 304 can acquire multiple reference blocks. In some examples, the interpretation module 304 can acquire multiple reference blocks based on motion data.

[0173] In 2304, the system may encode a first flag. In some implementations, the first flag may be signaled at the SPS level, PH level, PPS level, or SH level. For example, referring to Figure 3, the inter-prediction module 304 may encode a first flag. The first flag may indicate whether or not the FInter prediction mode is enabled for the current block.

[0174] In 2306, in response to the first flag indicating that FInter prediction mode is enabled for the current block, the system may encode a second flag. For example, referring to Figure 3, if the first flag indicates that FInter prediction mode is enabled, the inter-prediction module 304 may encode a second flag. The second flag may indicate whether normal inter-prediction or FInter prediction is selected for the current block.

[0175] In 2308, the system may determine whether FInter prediction mode is selected for the current block based on a second flag. For example, referring to Figure 3, the inter-prediction module 304 may determine whether FInter prediction mode is selected for the current block based on a second flag. For example, a second flag having a first value may indicate that FInter prediction is selected, and a second flag having a second value may indicate that normal inter-prediction is selected.

[0176] In 2310, the system can encode a single flag. For example, referring to Figure 3, the interprediction module 304 can encode a single flag when the FInter prediction mode replaces the normal interprediction mode.

[0177] In 2312, the system may determine, based on the flag, whether or not the FInter prediction mode is selected for all inter-prediction blocks. For example, referring to Figure 3, if the single flag has a first value, the inter-prediction module 304 may determine that the FInter prediction mode is selected for all inter-prediction blocks, and if the single flag has a second value, the inter-prediction module 304 may determine that the FInter prediction mode is not selected for all inter-prediction blocks.

[0178] In 2314, the system may select a set of filter weights for the FInter prediction filter that minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. For example, referring to Figure 3, the inter-prediction module 304 predicts the template RT of the reference block according to equation (18). 1 RT 2 ,…RT n The filter weights can be learned / selected by minimizing the weighted error between each of the filters and the template of the current block. In some implementations, the interprediction module 304 may select the filter weights to minimize the mean squared error (MSE) between the filtered RT and the template of the current block according to equation (16).

[0179] Referring to Figure 23B, in 2316, the system may generate an FInter prediction filter based on the set of filter weights. In some implementations, the FInter prediction for the current block may be generated based on the FInter prediction filter. For example, referring to Figure 3, the inter-prediction module 304 may generate an FInter prediction filter based on selected filter weights.

[0180] In 2318, in response to the determination that the FInter prediction mode has been selected for the current block, the system generates an FInter prediction for the current block based on a plurality of reference blocks. In some implementations, the generation of an FInter prediction for the current block based on a plurality of reference blocks by the processor in response to the determination that the FInter prediction mode has been selected for the current block may include filtering each of the plurality of reference blocks individually using corresponding filters to obtain a plurality of filtered reference blocks. In some implementations, the generation of an FInter prediction for the current block based on a plurality of reference blocks by the processor in response to the determination that the FInter prediction mode has been selected for the current block may include merging the plurality of filtered reference blocks to obtain a plurality of merged and filtered reference blocks. In some implementations, the generation of an FInter prediction for the current block based on a plurality of reference blocks by the processor in response to the determination that the FInter prediction mode has been selected for the current block may include generating an FInter prediction for the current block based on the merged and filtered plurality of reference blocks. For example, referring to Figure 3, for bidirectional and multiple-hypothesis prediction, the inter-prediction module 304 may individually learn multiple filters by using templates around multiple reference blocks indicated by corresponding multiple motion data and current blocks. The inter-prediction module 304 may first individually filter the reference blocks using the corresponding learned filters. Then, the inter-prediction module 304 may merge the multiple filtered reference blocks to form the final prediction. The entire bidirectional and multiple-hypothesis process may be similar to inter-prediction, except that it is proposed to generate the final prediction using filtered reference blocks instead of the reference blocks themselves.In some other implementations, in response to the decision that an FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include fusing the multiple reference blocks to obtain a fusing set of reference blocks. In some other implementations, in response to the decision that an FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include filtering the fusing set of reference blocks to obtain a fusing and filtered set of reference blocks. In some other implementations, in response to the decision that an FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include generating an FInter prediction for the current block based on the fusing and filtered set of reference blocks. For example, referring to Figure 3, the interpretation module 304 can determine multiple reference blocks through multiple motion data signaling to form the final predictor of the current CU. The reference blocks are R. 1 ,R 2 ,…R n And, Σ k w k P=w1R is the fused combination of reference blocks for some fusion weights that satisfy =1. 1 +w2R 2 At the same time... lol n R n In this implementation, the interpretation module 304 can generate a final predictor by further filtering the fused combination using the learned filter according to equation (17).

[0181] In 2320, in response to the determination that the FInter prediction mode is not selected for the current block, the system generates a normal inter prediction for the current block based on multiple reference blocks. For example, referring to Figure 3, if the FInter prediction mode is not selected for the current block, the inter prediction module 304 may generate a normal inter prediction for the current block.

[0182] In various embodiments of this disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the functions may be stored as instructions on a non-temporary computer-readable medium. The computer-readable medium includes computer storage media. The storage medium may be any available medium accessible by a processor (e.g., processor 102 in Figures 1 and 2). Such computer-readable media may include, but are not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, HDD (such as magnetic disk storage or other magnetic storage device), flash drive, SSD, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible by a processing system (such as a mobile device or computer). As used herein, disk and disc include CD, laserdisc, optical disc, digital video disc (DVD), and floppy disk, where disk typically reproduces data magnetically, while disc uses a laser to reproduce data optically. The above combinations should also be included within the scope of computer-readable media.

[0183] According to one aspect of the present disclosure, a decoding method is provided. The method may include a processor obtaining a plurality of reference blocks. The method may include the processor generating a FInter prediction for the current block based on the plurality of reference blocks in response to a determination that a FInter prediction mode has been selected for the current block.

[0184] In some implementations, in response to a decision that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor individually filtering each of the multiple reference blocks using corresponding filters to obtain multiple filtered reference blocks. In some implementations, in response to a decision that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor fusing multiple filtered reference blocks to obtain multiple fusing and filtered reference blocks. In some implementations, in response to a decision that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor generating an FInter prediction for the current block based on the fusing and filtered multiple reference blocks.

[0185] In some implementations, the method may include the processor analyzing the bitstream to obtain a first flag. In some implementations, the method may include the processor analyzing the bitstream to obtain a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, the method may include the processor determining, based on the second flag, whether FInter prediction mode is selected for the current block.

[0186] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.

[0187] In some implementations, the method may include the processor generating a normal inter prediction for the current block based on a number of reference blocks in response to the determination that the FInter prediction mode is not selected for the current block.

[0188] In some implementations, the method may include the processor analyzing the bitstream to obtain a flag. In some implementations, the method may include the processor determining, based on the flag, whether the FInter prediction mode is selected for all inter-prediction blocks.

[0189] In some implementations, obtaining multiple reference blocks by the processor may include the processor analyzing a bitstream to obtain multiple motion data. In some implementations, obtaining multiple reference blocks by the processor may also include the processor generating multiple reference blocks based on multiple motion data.

[0190] In some implementations, when it is determined that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor fusing the multiple reference blocks to obtain a fusing set of reference blocks. In some implementations, when it is determined that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor filtering the fusing set of reference blocks to obtain a fusing and filtered set of reference blocks. In some implementations, when it is determined that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor generating an FInter prediction for the current block based on the fusing and filtered set of reference blocks.

[0191] In some implementations, the method may include the processor selecting a set of filter weights for an FInter prediction filter that minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, the method may include the processor generating an FInter prediction filter based on the set of filter weights. In some implementations, an FInter prediction for the current block may be generated based on the FInter prediction filter.

[0192] A decoder is provided according to another aspect of the present disclosure. The decoder may include a processor and a memory for storing instructions. The memory stores instructions which, when executed by the processor, cause the processor to acquire a plurality of reference blocks. The memory stores instructions which, when executed by the processor, cause the processor to generate a FInter prediction for the current block based on the plurality of reference blocks in response to a determination that a FInter prediction mode has been selected for the current block.

[0193] In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that a FInter prediction mode has been selected for the current block. When these instructions are executed by the processor, they cause the processor to individually filter each of the multiple reference blocks using the corresponding filters to obtain multiple filtered reference blocks. In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that a FInter prediction mode has been selected for the current block. When these instructions are executed by the processor, they cause the processor to merge the multiple filtered reference blocks to obtain multiple merged and filtered reference blocks. In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that a FInter prediction mode has been selected for the current block. When these instructions are executed by the processor, they cause the processor to generate a FInter prediction for the current block based on the merged and filtered multiple reference blocks.

[0194] In some implementations, memory stores instructions that, when executed by the processor, cause the processor to analyze the bitstream and obtain a first flag. In some implementations, memory stores instructions that, when executed by the processor, cause the processor to analyze the bitstream and obtain a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, memory stores instructions that, when executed by the processor, cause the processor to determine, based on the second flag, whether FInter prediction mode is selected for the current block.

[0195] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.

[0196] In some implementations, memory stores instructions that, when executed by the processor, cause the processor to generate a normal inter prediction for the current block based on multiple reference blocks in response to determining that the FInter prediction mode is not selected for the current block.

[0197] In some implementations, memory stores instructions that, when executed by the processor, cause the processor to analyze the bitstream and obtain a flag. In some implementations, memory stores instructions that, when executed by the processor, cause the processor to determine, based on the flag, whether or not the FInter prediction mode is selected for all inter prediction blocks.

[0198] In some implementations, memory stores instructions to obtain multiple reference blocks, and when these instructions are executed by the processor, they cause the processor to parse the bitstream and obtain multiple motion data. In some implementations, memory stores instructions to obtain multiple reference blocks, and when these instructions are executed by the processor, they cause the processor to generate multiple reference blocks based on multiple motion data.

[0199] In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, and when these instructions are executed by the processor, they cause the processor to merge the multiple reference blocks and obtain a merged set of reference blocks. In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, and when these instructions are executed by the processor, they cause the processor to filter the merged set of reference blocks and obtain a merged and filtered set of reference blocks. In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, and when these instructions are executed by the processor, they cause the processor to generate a FInter prediction for the current block based on the merged and filtered set of reference blocks.

[0200] In some implementations, memory stores instructions that, when executed by the processor, cause the processor to select a set of filter weights for a FInter prediction filter that minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, memory stores instructions that, when executed by the processor, cause the processor to generate a FInter prediction filter based on this set of filter weights. In some implementations, the FInter prediction for the current block can be generated based on the FInter prediction filter.

[0201] According to another aspect of the present disclosure, a decoding device is provided. The decoding device may include a processor and a memory for storing instructions. The memory stores instructions which, when executed by the processor, cause the processor to acquire a plurality of reference blocks. The memory stores instructions which, when executed by the processor, cause the processor to generate a FInter prediction for the current block based on the plurality of reference blocks in response to a determination that a FInter prediction mode has been selected for the current block.

[0202] In yet another aspect of the present disclosure, a non-temporary computer-readable medium for storing instructions is provided. When the instructions are executed by the decoder's processor, the instructions cause the decoder's processor to acquire a plurality of reference blocks. When the instructions are executed by the decoder's processor, the instructions cause the decoder's processor to generate a FInter prediction for the current block based on the plurality of reference blocks in response to determining that a FInter prediction mode has been selected for the current block.

[0203] In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the decoder's processor, can cause the decoder's processor to individually filter each of the multiple reference blocks using the corresponding filter to obtain multiple filtered reference blocks. In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the decoder's processor, can cause the decoder's processor to merge the multiple filtered reference blocks to obtain multiple merged and filtered reference blocks. In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the decoder's processor, can cause the decoder's processor to generate a FInter prediction for the current block based on the merged and filtered multiple reference blocks.

[0204] In some implementations, when the instruction is executed by the decoder's processor, it can cause the decoder's processor to analyze the bitstream and obtain a first flag. In some implementations, when the instruction is executed by the decoder's processor, it can cause the decoder's processor to analyze the bitstream and obtain a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, when the instruction is executed by the decoder's processor, it can cause the decoder's processor to determine, based on the second flag, whether or not FInter prediction mode is selected for the current block.

[0205] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.

[0206] In some implementations, when the instruction is executed by the decoder's processor, it can cause the decoder's processor to generate a normal inter prediction for the current block based on multiple reference blocks in response to determining that the FInter prediction mode is not selected for the current block.

[0207] In some implementations, when the instruction is executed by the decoder's processor, it can cause the decoder's processor to analyze the bitstream and obtain a flag. In some implementations, when the instruction is executed by the decoder's processor, it can cause the decoder's processor to determine, based on the flag, whether or not the FInter prediction mode is selected for all inter prediction blocks.

[0208] In some implementations, to obtain multiple reference blocks, the instruction, when executed by the decoder's processor, can cause the decoder's processor to analyze the bitstream and obtain multiple motion data. In some implementations, to obtain multiple reference blocks, the instruction, when executed by the decoder's processor, can cause the decoder's processor to generate multiple reference blocks based on multiple motion data.

[0209] In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the decoder's processor, can cause the decoder's processor to merge the multiple reference blocks and obtain a merged set of reference blocks. In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the decoder's processor, can cause the decoder's processor to filter the merged set of reference blocks and obtain a merged and filtered set of reference blocks. In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the decoder's processor, can cause the decoder's processor to generate a FInter prediction for the current block based on the merged and filtered set of reference blocks.

[0210] In some implementations, the instruction, when executed by the decoder's processor, can cause the decoder's processor to select a set of filter weights for a FInter prediction filter that minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, the instruction, when executed by the decoder's processor, can cause the decoder's processor to generate a FInter prediction filter based on this set of filter weights. In some implementations, the FInter prediction for the current block may be generated based on the FInter prediction filter.

[0211] According to one aspect of the present disclosure, an encoding method is provided. The method may include a processor obtaining a plurality of reference blocks. The method may include the processor generating a FInter prediction for the current block based on the plurality of reference blocks in response to a determination that a FInter prediction mode has been selected for the current block.

[0212] In some implementations, in response to a decision that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor individually filtering each of the multiple reference blocks using corresponding filters to obtain multiple filtered reference blocks. In some implementations, in response to a decision that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor fusing multiple filtered reference blocks to obtain multiple fusing and filtered reference blocks. In some implementations, in response to a decision that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor generating an FInter prediction for the current block based on the fusing and filtered multiple reference blocks.

[0213] In some implementations, the method may include the processor encoding a first flag. In some implementations, the method may include the processor encoding a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, the method may include the processor determining, based on the second flag, whether FInter prediction mode is selected for the current block.

[0214] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.

[0215] In some implementations, the method may include the processor generating a normal inter prediction for the current block based on a number of reference blocks in response to the determination that the FInter prediction mode is not selected for the current block.

[0216] In some implementations, the method may include the processor encoding a flag. In some implementations, the method may include the processor determining, based on the flag, whether the FInter prediction mode is selected for all inter-prediction blocks.

[0217] In some implementations, obtaining multiple reference blocks by the processor may involve the processor encoding multiple motion data. In some implementations, obtaining multiple reference blocks by the processor may involve the processor generating multiple reference blocks based on multiple motion data.

[0218] In some implementations, when it is determined that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor fusing the multiple reference blocks to obtain a fusing set of reference blocks. In some implementations, when it is determined that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor filtering the fusing set of reference blocks to obtain a fusing and filtered set of reference blocks. In some implementations, when it is determined that the FInter prediction mode has been selected for the current block, the processor generating an FInter prediction for the current block based on multiple reference blocks may include the processor generating an FInter prediction for the current block based on the fusing and filtered set of reference blocks.

[0219] In some implementations, the method may include the processor selecting a set of filter weights for an FInter prediction filter that minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, the method may include the processor generating an FInter prediction filter based on the set of filter weights. In some implementations, an FInter prediction for the current block may be generated based on the FInter prediction filter.

[0220] An encoder is provided according to another aspect of the present disclosure. The encoder may include a processor and a memory for storing instructions. The memory stores instructions which, when executed by the processor, cause the processor to acquire a plurality of reference blocks. The memory stores instructions which, when executed by the processor, cause the processor to generate a FInter prediction for the current block based on the plurality of reference blocks in response to a determination that a FInter prediction mode has been selected for the current block.

[0221] In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that a FInter prediction mode has been selected for the current block. When these instructions are executed by the processor, they cause the processor to individually filter each of the multiple reference blocks using the corresponding filters to obtain multiple filtered reference blocks. In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that a FInter prediction mode has been selected for the current block. When these instructions are executed by the processor, they cause the processor to merge the multiple filtered reference blocks to obtain multiple merged and filtered reference blocks. In some implementations, memory stores instructions to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that a FInter prediction mode has been selected for the current block. When these instructions are executed by the processor, they cause the processor to generate a FInter prediction for the current block based on the merged and filtered multiple reference blocks.

[0222] In some implementations, memory stores an instruction that, when executed by the processor, causes the processor to encode a first flag. In some implementations, memory stores an instruction that, when executed by the processor, causes the processor to encode a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, memory stores an instruction that, when executed by the processor, causes the processor to determine, based on the second flag, whether FInter prediction mode is selected for the current block.

[0223] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.

[0224] In some implementations, memory stores instructions that, when executed by the processor, cause the processor to generate a normal inter prediction for the current block based on multiple reference blocks in response to determining that the FInter prediction mode is not selected for the current block.

[0225] In some implementations, memory stores instructions that, when executed by the processor, allow the processor to encode a flag. In some implementations, memory stores instructions that, when executed by the processor, allow the processor to determine, based on the flag, whether or not the FInter prediction mode is selected for all inter-prediction blocks.

[0226] In some implementations, to obtain a plurality of reference blocks, a memory stores instructions that, when executed by a processor, enable the processor to encode a plurality of motion data. In some implementations, to obtain a plurality of reference blocks, a memory stores instructions that, when executed by a processor, enable the processor to generate the plurality of reference blocks based on the plurality of motion data.

[0227] In some implementations, to generate FInter prediction for a current block based on a plurality of reference blocks in response to determining that the FInter prediction mode is selected for the current block, a memory stores instructions that, when executed by a processor, enable the processor to fuse the plurality of reference blocks to obtain a plurality of fused reference blocks. In some implementations, to generate FInter prediction for a current block based on a plurality of reference blocks in response to determining that the FInter prediction mode is selected for the current block, a memory stores instructions that, when executed by a processor, enable the processor to filter the plurality of fused reference blocks to obtain a plurality of fused and filtered reference blocks. In some implementations, to generate FInter prediction for a current block based on a plurality of reference blocks in response to determining that the FInter prediction mode is selected for the current block, a memory stores instructions that, when executed by a processor, enable the processor to generate the FInter prediction of the current block based on the plurality of fused and filtered reference blocks.

[0228] In some implementations, a memory stores instructions that, when executed by a processor, cause the processor to select a set of filter weights for an FInter prediction filter that minimizes the MSE between a reference template associated with a reference block and a current template associated with a current block. In some implementations, a memory stores instructions that, when executed by a processor, cause the processor to generate an FInter prediction filter based on the set of filter weights. In some implementations, FInter prediction of a current block can be generated further based on the FInter prediction filter.

[0229] According to another aspect of the present disclosure, an encoding apparatus is provided. The encoding apparatus may include a processor and a memory that stores instructions. The memory stores instructions that, when executed by the processor, cause the processor to obtain a plurality of reference blocks. The memory stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction of a current block based on the plurality of reference blocks in response to determining that an FInter prediction mode is selected for the current block.

[0230] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions is provided. The instructions, when executed by a processor of an encoder, cause the processor of the encoder to obtain a plurality of reference blocks. The instructions, when executed by the processor of the encoder, cause the processor of the encoder to generate an FInter prediction of a current block based on the plurality of reference blocks in response to determining that an FInter prediction mode is selected for the current block.

[0231] In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the encoder's processor, can cause the encoder's processor to individually filter each of the multiple reference blocks using the corresponding filter to obtain multiple filtered reference blocks. In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the encoder's processor, can cause the encoder's processor to merge the multiple filtered reference blocks to obtain multiple merged and filtered reference blocks. In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the encoder's processor, can cause the encoder's processor to generate a FInter prediction for the current block based on the merged and filtered multiple reference blocks.

[0232] In some implementations, the instruction, when executed by the encoder's processor, can cause the encoder's processor to encode a first flag. In some implementations, the instruction, when executed by the encoder's processor, can cause the encoder's processor to encode a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, the instruction, when executed by the encoder's processor, can cause the encoder's processor to determine, based on the second flag, whether FInter prediction mode is selected for the current block.

[0233] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.

[0234] In some implementations, the instruction, when executed by the encoder's processor, can cause the encoder's processor to generate a normal inter prediction for the current block based on multiple reference blocks in response to determining that the FInter prediction mode is not selected for the current block.

[0235] In some implementations, the instruction can cause the encoder's processor to encode a flag when executed by the encoder's processor. In some implementations, the instruction can cause the encoder's processor to determine, based on the flag, whether or not the FInter prediction mode is selected for all inter prediction blocks when executed by the encoder's processor.

[0236] In some implementations, to obtain multiple reference blocks, the instruction can cause the encoder's processor to encode multiple motion data when executed by the encoder. In some implementations, to obtain multiple reference blocks, the instruction can cause the encoder's processor to generate multiple reference blocks based on multiple motion data when executed by the encoder.

[0237] In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the encoder's processor, can cause the encoder's processor to merge the multiple reference blocks and obtain a merged set of reference blocks. In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the encoder's processor, can cause the encoder's processor to filter the merged set of reference blocks and obtain a merged and filtered set of reference blocks. In some implementations, in order to generate a FInter prediction for the current block based on multiple reference blocks in response to a decision that the FInter prediction mode has been selected for the current block, the instruction, when executed by the encoder's processor, can cause the encoder's processor to generate a FInter prediction for the current block based on the merged and filtered set of reference blocks.

[0238] In some implementations, the instruction, when executed by the encoder's processor, can cause the encoder's processor to select a set of filter weights for a FInter prediction filter that minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, the instruction, when executed by the encoder's processor, can cause the encoder's processor to generate a FInter prediction filter based on this set of filter weights. In some implementations, the FInter prediction for the current block may be generated based on the FInter prediction filter.

[0239] In yet another aspect of this disclosure, a non-temporary computer-readable medium for storing a bitstream is provided. The bitstream may be generated according to one or more of the operations disclosed herein.

[0240] The foregoing description of embodiments clarifies the general nature of the disclosure, and those skilled in the art can readily modify and / or adapt such embodiments for various uses without excessive experimentation and without departing from the general concepts of the disclosure by applying their knowledge of the art. Accordingly, based on the teachings and guidance presented herein, such adaptations and modifications are intended to fall within the meaning and scope of the equivalent embodiments disclosed herein. Expressions or terms herein are for illustrative purposes only and not limiting, and it is understood that such terms or expressions herein should be interpreted by those skilled in the art in light of the teachings and guidance.

[0241] Embodiments of this disclosure are described above with the help of function building blocks that show implementations of specified functions and their relationships. The boundaries of these function building blocks are arbitrarily defined herein for the sake of explanation. Alternative boundaries can be defined as long as the specified functions and their relationships are properly performed.

[0242] The sections describing the summary and abstract of the invention may include one or more, but not all, exemplary embodiments of the present disclosure as envisioned by the inventors, and are therefore not intended to limit the scope of the present disclosure and the accompanying claims in any way.

[0243] Various functional blocks, modules, and steps are disclosed above. The provided arrangements are illustrative and not limiting. Thus, functional blocks, modules, and steps may be rearranged or combined in ways different from the examples provided above. Similarly, some embodiments may include only a subset of functional blocks, modules, and steps, and any such subset is acceptable.

[0244] The breadth and scope of this disclosure should not be limited by any of the exemplary embodiments described above, but should be defined solely in accordance with the following claims and their equivalents.

Claims

1. A decoding method using a decoder, The processor can obtain multiple reference blocks, A decoding method comprising: the processor generating a Finter prediction for the current block based on the plurality of reference blocks in response to a determination that a Finter prediction mode has been selected for the current block.

2. In response to the determination that the Finter prediction mode has been selected for the current block, the processor generates the Finter prediction for the current block based on the plurality of reference blocks. The processor filters each of the multiple reference blocks individually using the corresponding filter to obtain the filtered multiple reference blocks. The aforementioned processor merges the filtered plurality of reference blocks to obtain a plurality of merged and filtered reference blocks, The processor includes generating the Finter prediction for the current block based on the fused and filtered plurality of reference blocks, The decoding method according to claim 1.

3. The aforementioned decoding method is The aforementioned processor analyzes the bitstream to obtain a first flag, In response to the first flag indicating that the Finter prediction mode is enabled for the current block, the processor analyzes the bitstream to obtain a second flag, The processor further includes determining whether a Finter prediction mode is selected for the current block based on the second flag, The decoding method according to claim 1.

4. The first flag is signaled at the Sequence Parameter Set (SPS) level, Picture Header (PH) level, Picture Parameter Set (PPS) level, or Slice Header (SH) level. The decoding method according to claim 3.

5. The aforementioned decoding method is In response to the determination that the Finter prediction mode is not selected for the current block, the processor further includes generating a normal inter prediction for the current block based on the plurality of reference blocks. The decoding method according to claim 3.

6. The aforementioned decoding method is The aforementioned processor analyzes the bitstream and obtains a flag, The processor further includes determining whether the Finter prediction mode is selected for all interpretation blocks based on the flag, The decoding method according to claim 1.

7. The processor obtaining the aforementioned multiple reference blocks means that The aforementioned processor analyzes the bitstream to acquire multiple motion data, The processor includes generating the plurality of reference blocks based on the plurality of motion data, The decoding method according to claim 1.

8. In response to the determination that the Finter prediction mode has been selected for the current block, the processor generates the Finter prediction for the current block based on the plurality of reference blocks. The aforementioned processor merges the plurality of reference blocks to obtain the merged plurality of reference blocks, The aforementioned processor filters the merged reference blocks to obtain the merged and filtered reference blocks, The processor includes generating the Finter prediction for the current block based on the fused and filtered plurality of reference blocks, The decoding method according to claim 1.

9. The processor selects a set of filter weights for a Finter prediction filter that minimizes the mean squared error (MSE) between the reference template associated with the reference block and the current template associated with the current block, The processor further includes generating the Finter prediction filter based on the set of filter weights, The Finter prediction for the current block is further generated based on the Finter prediction filter, The decoding method according to claim 1.

10. It is a decoder, Processor and The system includes a memory for storing instructions, and when an instruction is executed by the processor, the processor receives the instructions. Obtaining multiple reference blocks, A decoder that, in response to a determination that a filter prediction mode has been selected for the current block, generates a filter prediction for the current block based on the plurality of reference blocks.

11. In response to the determination that a Finter prediction mode has been selected for the current block, the instruction, when executed by the processor, causes the processor to generate the Finter prediction for the current block based on the plurality of reference blocks. The process involves individually filtering each of the multiple reference blocks using the corresponding filter to obtain the filtered multiple reference blocks, The process involves merging the filtered reference blocks to obtain a merged and filtered set of reference blocks, To generate the Finter prediction for the current block based on the merged and filtered plurality of reference blocks, The decoder according to claim 10.

12. When the aforementioned instruction is executed by the processor, the processor will be instructed to: The bitstream is analyzed to obtain the first flag, In response to the first flag indicating that the Finter prediction mode is enabled for the current block, the bitstream is analyzed to obtain a second flag, Based on the second flag, determine whether or not the Finter prediction mode is selected for the current block, and perform the following: The decoder according to claim 10.

13. The first flag is signaled at the Sequence Parameter Set (SPS) level, Picture Header (PH) level, Picture Parameter Set (PPS) level, or Slice Header (SH) level. The decoder according to claim 12.

14. When the aforementioned instruction is executed by the processor, the processor will be instructed to: In response to the determination that the Finter prediction mode is not selected for the current block, the system is made to generate a normal inter prediction for the current block based on the plurality of reference blocks. The decoder according to claim 12.

15. When the aforementioned instruction is executed by the processor, the processor will be instructed to: The process involves parsing a bitstream to obtain a flag, Based on the aforementioned flag, determine whether or not the Finter prediction mode is selected for all interpretation blocks, and perform the following: The decoder according to claim 10.

16. In order to obtain the aforementioned multiple reference blocks, the instruction, when executed by the processor, causes the processor to: This involves analyzing a bitstream to obtain multiple motion data points, To generate the multiple reference blocks based on the multiple motion data, and to perform the following: The decoder according to claim 10.

17. In response to the determination that a Finter prediction mode has been selected for the current block, the instruction, when executed by the processor, causes the processor to generate the Finter prediction for the current block based on the plurality of reference blocks. The process involves merging the aforementioned multiple reference blocks to obtain the merged multiple reference blocks, The process involves filtering the merged reference blocks to obtain the merged and filtered reference blocks, To generate the Finter prediction for the current block based on the merged and filtered plurality of reference blocks, The decoder according to claim 10.

18. When the aforementioned instruction is executed by the processor, the processor will be instructed to: Selecting a set of filter weights for a Finter prediction filter that minimizes the mean squared error (MSE) between the reference template associated with the reference block and the current template associated with the current block, The process involves generating the Finter prediction filter based on the set of filter weights, and then performing the following: The Finter prediction for the current block is further generated based on the Finter prediction filter, The decoder according to claim 10.

19. A decoding device, Processor and The system includes a memory for storing instructions, and when an instruction is executed by the processor, the processor receives the instructions. Obtaining multiple reference blocks, A decoding device that, in response to a determination that a filter prediction mode has been selected for the current block, generates a filter prediction for the current block based on the plurality of reference blocks.

20. A non-temporary computer-readable medium for storing instructions, wherein, when the instructions are executed by the decoder's processor, the decoder's processor, Obtaining multiple reference blocks, A non-temporary computer-readable medium that, in response to a determination that a filter prediction mode has been selected for the current block, causes the system to generate a filter prediction for the current block based on the plurality of reference blocks.

21. In response to the determination that a Finter prediction mode has been selected for the current block, the instruction is executed by the processor of the decoder to generate the Finter prediction for the current block based on the plurality of reference blocks, and the processor of the decoder is instructed to do so. The process involves individually filtering each of the multiple reference blocks using the corresponding filter to obtain the filtered multiple reference blocks, The process involves merging the filtered reference blocks to obtain a merged and filtered set of reference blocks, To generate the Finter prediction for the current block based on the merged and filtered plurality of reference blocks, The non-temporary computer-readable medium according to claim 20.

22. When the aforementioned instruction is executed by the processor of the decoder, the processor of the decoder will be instructed to: The bitstream is analyzed to obtain the first flag, In response to the first flag indicating that the Finter prediction mode is enabled for the current block, the bitstream is analyzed to obtain a second flag, Based on the second flag, determine whether or not the Finter prediction mode is selected for the current block, and perform the following: The non-temporary computer-readable medium according to claim 20.

23. The first flag is signaled at the Sequence Parameter Set (SPS) level, Picture Header (PH) level, Picture Parameter Set (PPS) level, or Slice Header (SH) level. The non-temporary computer-readable medium according to claim 22.

24. When the aforementioned instruction is executed by the processor of the decoder, the processor of the decoder will be instructed to: In response to the determination that the Finter prediction mode is not selected for the current block, the system is made to generate a normal inter prediction for the current block based on the plurality of reference blocks. The non-temporary computer-readable medium according to claim 22.

25. When the aforementioned instruction is executed by the processor of the decoder, the processor of the decoder will be instructed to: The process involves parsing a bitstream to obtain a flag, Based on the aforementioned flag, determine whether or not the Finter prediction mode is selected for all interpretation blocks, and perform the following: The non-temporary computer-readable medium according to claim 20.

26. In order to obtain the aforementioned multiple reference blocks, the instruction, when executed by the decoder's processor, causes the decoder's processor to: This involves analyzing a bitstream to obtain multiple motion data points, To generate the multiple reference blocks based on the multiple motion data, and to perform the following: The non-temporary computer-readable medium according to claim 20.

27. In response to the determination that a Finter prediction mode has been selected for the current block, the instruction is executed by the processor of the decoder to generate the Finter prediction for the current block based on the plurality of reference blocks, and the processor of the decoder is instructed to do so. The process involves merging the aforementioned multiple reference blocks to obtain the merged multiple reference blocks, The process involves filtering the merged reference blocks to obtain the merged and filtered reference blocks, To generate the Finter prediction for the current block based on the merged and filtered plurality of reference blocks, The non-temporary computer-readable medium according to claim 20.

28. When the aforementioned instruction is executed by the processor of the decoder, the processor of the decoder will be instructed to: Selecting a set of filter weights for a Finter prediction filter that minimizes the mean squared error (MSE) between the reference template associated with the reference block and the current template associated with the current block, The process involves generating the Finter prediction filter based on the set of filter weights, and then performing the following: The Finter prediction for the current block is further generated based on the Finter prediction filter, The non-temporary computer-readable medium according to claim 20.

29. An encoding method using an encoder, The processor can obtain multiple reference blocks, An encoding method comprising: the processor generating a Finter prediction for the current block based on the plurality of reference blocks in response to a determination that a Finter prediction mode has been selected for the current block.

30. In response to the determination that the Finter prediction mode has been selected for the current block, the processor generates the Finter prediction for the current block based on the plurality of reference blocks. The processor filters each of the multiple reference blocks individually using the corresponding filter to obtain the filtered multiple reference blocks. The aforementioned processor merges the filtered plurality of reference blocks to obtain a plurality of merged and filtered reference blocks, The processor includes generating the Finter prediction for the current block based on the fused and filtered plurality of reference blocks, The encoding method according to claim 29.

31. The aforementioned encoding method is The processor encodes the first flag, In response to the first flag indicating that the Finter prediction mode is enabled for the current block, the processor encodes a second flag. The processor further includes determining whether a Finter prediction mode is selected for the current block based on the second flag, The encoding method according to claim 29.

32. The first flag is signaled at the Sequence Parameter Set (SPS) level, Picture Header (PH) level, Picture Parameter Set (PPS) level, or Slice Header (SH) level. The encoding method according to claim 31.

33. The aforementioned encoding method is In response to the determination that the Finter prediction mode is not selected for the current block, the processor further includes generating a normal inter prediction for the current block based on the plurality of reference blocks. The encoding method according to claim 31.

34. The aforementioned encoding method is The aforementioned processor encodes the flag, The processor further includes determining whether the Finter prediction mode is selected for all interpretation blocks based on the flag, The encoding method according to claim 29.

35. The processor obtaining the aforementioned multiple reference blocks means that The aforementioned processor encodes multiple motion data, The processor includes generating the plurality of reference blocks based on the plurality of motion data, The encoding method according to claim 29.

36. In response to the determination that the Finter prediction mode has been selected for the current block, the processor generates the Finter prediction for the current block based on the plurality of reference blocks. The aforementioned processor merges the plurality of reference blocks to obtain the merged plurality of reference blocks, The aforementioned processor filters the merged reference blocks to obtain the merged and filtered reference blocks, The processor includes generating the Finter prediction for the current block based on the fused and filtered plurality of reference blocks, The encoding method according to claim 29.

37. The aforementioned encoding method is The processor selects a set of filter weights for a Finter prediction filter that minimizes the mean squared error (MSE) between the reference template associated with the reference block and the current template associated with the current block, The processor further includes generating the Finter prediction filter based on the set of filter weights, The Finter prediction for the current block is further generated based on the Finter prediction filter, The encoding method according to claim 29.

38. It is an encoder, Processor and The system includes a memory for storing instructions, and when an instruction is executed by the processor, the processor receives the instructions. Obtaining multiple reference blocks, An encoder that, in response to a determination that a filter interface (Finter) prediction mode has been selected for the current block, causes to generate a Finter prediction for the current block based on the plurality of reference blocks.

39. In response to the determination that a Finter prediction mode has been selected for the current block, the instruction, when executed by the processor, causes the processor to generate the Finter prediction for the current block based on the plurality of reference blocks. The process involves individually filtering each of the multiple reference blocks using the corresponding filter to obtain the filtered multiple reference blocks, The process involves merging the filtered reference blocks to obtain a merged and filtered set of reference blocks, To generate the Finter prediction for the current block based on the merged and filtered plurality of reference blocks, The encoder according to claim 38.

40. When the aforementioned instruction is executed by the processor, the processor will be instructed to: Encoding the first flag, The first flag indicates that the Finter prediction mode is enabled for the current block, and the second flag is encoded in response to this. Based on the second flag, determine whether or not the Finter prediction mode is selected for the current block, and perform the following: The encoder according to claim 38.

41. The first flag is signaled at the Sequence Parameter Set (SPS) level, Picture Header (PH) level, Picture Parameter Set (PPS) level, or Slice Header (SH) level. The encoder according to claim 40.

42. When the aforementioned instruction is executed by the processor, the processor will be instructed to: In response to the determination that the Finter prediction mode is not selected for the current block, the system is made to generate a normal inter prediction for the current block based on the plurality of reference blocks. The encoder according to claim 40.

43. When the aforementioned instruction is executed by the processor, the processor will be instructed to: Encoding the flag, Based on the aforementioned flag, determine whether or not the Finter prediction mode is selected for all interpretation blocks, and perform the following: The encoder according to claim 38.

44. In order to obtain the aforementioned multiple reference blocks, the instruction, when executed by the processor, causes the processor to: Encoding multiple motion data, To generate the multiple reference blocks based on the multiple motion data, and to perform the following: The encoder according to claim 38.

45. In response to the determination that a Finter prediction mode has been selected for the current block, the instruction, when executed by the processor, causes the processor to generate the Finter prediction for the current block based on the plurality of reference blocks. The process involves merging the aforementioned multiple reference blocks to obtain the merged multiple reference blocks, The process involves filtering the merged reference blocks to obtain the merged and filtered reference blocks, To generate the Finter prediction for the current block based on the merged and filtered plurality of reference blocks, The encoder according to claim 38.

46. When the aforementioned instruction is executed by the processor, the processor will be instructed to: Selecting a set of filter weights for a Finter prediction filter that minimizes the mean squared error (MSE) between the reference template associated with the reference block and the current template associated with the current block, The process involves generating the Finter prediction filter based on the set of filter weights, and then performing the following: The Finter prediction for the current block is further generated based on the Finter prediction filter, The encoder according to claim 38.

47. An encoding device, Processor and The system includes a memory for storing instructions, and when an instruction is executed by the processor, the processor receives the instructions. Obtaining multiple reference blocks, An encoding device that, in response to a determination that a filter prediction mode has been selected for the current block, generates a filter prediction for the current block based on the plurality of reference blocks.

48. A non-temporary computer-readable medium for storing instructions, wherein, when the instructions are executed by the encoder's processor, the encoder's processor, Obtaining multiple reference blocks, A non-temporary computer-readable medium that, in response to a determination that a filter prediction mode has been selected for the current block, causes the system to generate a filter prediction for the current block based on the plurality of reference blocks.

49. In response to the determination that a Finter prediction mode has been selected for the current block, the instruction is executed by the processor of the encoder to generate the Finter prediction for the current block based on the plurality of reference blocks, and the processor of the encoder is instructed to do so. The process involves individually filtering each of the multiple reference blocks using the corresponding filter to obtain the filtered multiple reference blocks, The process involves merging the filtered reference blocks to obtain a merged and filtered set of reference blocks, To generate the Finter prediction for the current block based on the merged and filtered plurality of reference blocks, The non-temporary computer-readable medium according to claim 48.

50. When the aforementioned instruction is executed by the processor of the encoder, the processor of the encoder will be instructed to: Encoding the first flag, The first flag indicates that the Finter prediction mode is enabled for the current block, and the second flag is encoded in response to this. Based on the second flag, determine whether or not the Finter prediction mode is selected for the current block, and perform the following: The non-temporary computer-readable medium according to claim 48.

51. The first flag is signaled at the Sequence Parameter Set (SPS) level, Picture Header (PH) level, Picture Parameter Set (PPS) level, or Slice Header (SH) level. The non-temporary computer-readable medium according to claim 50.

52. When the aforementioned instruction is executed by the processor of the encoder, the processor of the encoder will be instructed to: In response to the determination that the Finter prediction mode is not selected for the current block, the system is made to generate a normal inter prediction for the current block based on the plurality of reference blocks. The non-temporary computer-readable medium according to claim 50.

53. When the aforementioned instruction is executed by the processor of the encoder, the processor of the encoder will be instructed to: Encoding the flag, Based on the aforementioned flag, determine whether or not the Finter prediction mode is selected for all interpretation blocks, and perform the following: The non-temporary computer-readable medium according to claim 48.

54. In order to obtain the plurality of reference blocks, the instruction, when executed by the processor of the encoder, instructs the processor of the encoder to Encoding multiple motion data, To generate the multiple reference blocks based on the multiple motion data, and to perform the following: The non-temporary computer-readable medium according to claim 48.

55. In response to the determination that a Finter prediction mode has been selected for the current block, the instruction is executed by the processor of the encoder to generate the Finter prediction for the current block based on the plurality of reference blocks, and the processor of the encoder is instructed to do so. The process involves merging the aforementioned multiple reference blocks to obtain the merged multiple reference blocks, The process involves filtering the merged reference blocks to obtain the merged and filtered reference blocks, To generate the Finter prediction for the current block based on the merged and filtered plurality of reference blocks, The non-temporary computer-readable medium according to claim 48.

56. When the aforementioned instruction is executed by the processor of the encoder, the processor of the encoder will be instructed to: Selecting a set of filter weights for a Finter prediction filter that minimizes the mean squared error (MSE) between the reference template associated with the reference block and the current template associated with the current block, The process involves generating the Finter prediction filter based on the set of filter weights, and then performing the following: The Finter prediction for the current block is further generated based on the Finter prediction filter, The non-temporary computer-readable medium according to claim 48.

57. A non-temporary computer-readable medium for storing a bitstream, wherein the bitstream is generated by an encoding method according to any one of claims 29 to 37.