Method and apparatus for filtered inter prediction
By introducing a filtered inter-frame prediction mode into video coding and utilizing the processor to acquire and process multiple reference blocks, the problem of low inter-frame prediction efficiency in existing technologies is solved, achieving more efficient video coding and decoding effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2024-09-27
- Publication Date
- 2026-05-01
AI Technical Summary
Existing video coding techniques suffer from inefficiency in inter-frame prediction, especially when processing complex video data, where it is difficult to effectively utilize reference blocks for efficient prediction.
The filter inter-frame prediction (FInter) mode is adopted. The processor obtains multiple reference blocks, generates the filter inter-frame prediction of the current block, and uses the filter module to perform filtering processing to improve prediction accuracy and efficiency.
It improves the efficiency and quality of video coding, especially in the processing of complex video data, enabling more accurate prediction using reference blocks and enhancing the performance of encoding and decoding.
Smart Images

Figure CN121970326A_ABST
Abstract
Description
Cross-references to related applications on methods and apparatus for inter-frame prediction using filtering
[0001] This application claims priority to U.S. Provisional Application No. 63 / 541,716, filed September 29, 2023, entitled “Filtered Inter Prediction for Video Coding,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The embodiments disclosed herein relate to video coding. Background Technology
[0003] Digital video has become mainstream and is used in a wide range of applications, including digital television, video telephony, and teleconferencing. These digital video applications are feasible due to advancements in computing and communication technologies, as well as efficient video coding techniques. Various video coding techniques can be used to compress video data, allowing the use of one or more video coding standards to encode the video data. Exemplary video coding standards may include, but are not limited to, Universal Video Coding (H.266 / VVC), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, Enhanced Video Coding Model (ECM), etc. Summary of the Invention
[0004] According to one aspect of this disclosure, a decoding method is provided. The method may include acquiring a plurality of reference blocks by a processor. The method may include: in response to determining that a filtered inter (FInter) prediction mode has been selected for the current block, generating an FInter prediction for the current block based on the plurality of reference blocks by the processor.
[0005] According to another aspect of this disclosure, a decoder is provided. The decoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to fetch a plurality of reference blocks. The memory stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0006] According to another aspect of this disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to acquire a plurality of reference blocks. The memory stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0007] According to another aspect of this disclosure, a non-transitory computer-readable medium storing instructions is provided. When executed by a decoder's processor, the instructions cause the decoder's processor to acquire a plurality of reference blocks. When executed by a decoder's processor, the instructions cause the decoder's processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0008] According to one aspect of this disclosure, an encoding method is provided. The method may include acquiring a plurality of reference blocks by a processor. The method may also include: in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on the plurality of reference blocks.
[0009] According to another aspect of this disclosure, an encoder is provided. The encoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to acquire a plurality of reference blocks. The memory stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0010] According to another aspect of this disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to fetch a plurality of reference blocks. The memory stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0011] According to another aspect of this disclosure, a non-transitory computer-readable medium storing instructions is provided. When executed by a processor of an encoder, the instructions cause the encoder's processor to acquire a plurality of reference blocks. When executed by a processor of the encoder, the instructions cause the encoder's processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0012] According to another aspect of this disclosure, a non-transitory computer-readable medium for storing bit streams is provided. Bit streams can be generated according to one or more of the operations disclosed herein.
[0013] These illustrative embodiments are mentioned not to limit or restrict this disclosure, but to provide examples to aid in understanding it. Further embodiments are described in the detailed description, and further description is provided therein. Attached Figure Description
[0014] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments of the present disclosure and, together with the specification, further serve to explain the principles of the present disclosure and enable those skilled in the art to make and use the present disclosure.
[0015] Figure 1 shows a block diagram of an exemplary encoding system according to some embodiments of the present disclosure.
[0016] Figure 2 shows a block diagram of an exemplary decoding system according to some embodiments of the present disclosure.
[0017] Figure 3 shows a detailed block diagram of an exemplary encoder in the encoding system of Figure 1 according to some embodiments of the present disclosure.
[0018] Figure 4 shows a detailed block diagram of an exemplary decoder in the decoding system of Figure 2 according to some embodiments of the present disclosure.
[0019] Figure 5 shows an exemplary image of a code tree unit (CTU) divided according to some embodiments of the present disclosure.
[0020] Figure 6 illustrates an exemplary CTU divided into coding units (CUs) according to some embodiments of the present disclosure.
[0021] Figure 7 shows a schematic visualization of the current CU block and reconstructed samples that are spatially adjacent and non-adjacent to the current block according to some embodiments of the present disclosure.
[0022] Figure 8 shows a schematic visualization of the angle modes of VVC according to some embodiments of the present disclosure.
[0023] Figure 9A illustrates a representation of Spatial Geometry Partitioning Mode (SGPM) signaling according to some embodiments of the present disclosure.
[0024] Figure 9B shows an example template for generating a candidate list according to some embodiments of the present disclosure.
[0025] Figure 10 illustrates schematic visualizations of inter-frame angle mode (parallel mode) parallel to the GPM partition boundary (see (a)), inter-frame angle mode (vertical mode) perpendicular to the GPM partition boundary (see (b)), and inter-frame prediction plane mode (see (c)) according to some embodiments of the present disclosure, as well as SGPM with intra-frame and inter-frame prediction (see (d)).
[0026] Figure 11 illustrates an adaptive SGPM hybrid scheme according to some embodiments of the present disclosure.
[0027] Figure 12 shows an example visualization of intra-block copy (IBC) prediction according to some embodiments of the present disclosure.
[0028] Figure 13 illustrates an intra-template matching prediction (intraTMP) diagram according to some embodiments of the present disclosure.
[0029] Figure 14 illustrates a decoder-side intra-frame mode derivation (DIMD) mode for ECM according to some embodiments of the present disclosure.
[0030] Figure 15 illustrates a template-based intra-frame mode derivation (TIMD) mode for ECM according to some embodiments of the present disclosure.
[0031] Figure 16A illustrates a 4 in VVC according to some embodiments of the present disclosure. N and N Four-block low-frequency non-separable transform (LFNST) cores.
[0032] Figure 16B illustrates 8 in VVC according to some embodiments of the present disclosure. N and N Eight LFNST cores.
[0033] Figure 17 illustrates various intra-frame angle prediction modes according to some embodiments of the present disclosure.
[0034] Figure 18A illustrates an ECM for 4 according to some embodiments of the present disclosure. N and N Four LFNST cores of different sizes.
[0035] Figure 18B illustrates an ECM for 8 according to some embodiments of the present disclosure. N and N Eight LFNST cores.
[0036] Figure 18C illustrates an LFNST core for a 16×16 block size in an ECM according to some embodiments of the present disclosure.
[0037] Figure 19 shows an example visualization of inter-frame prediction modes according to some aspects of this disclosure.
[0038] Figure 20 shows an example visualization of spatial support for learning filters according to some embodiments of the present disclosure.
[0039] Figure 21 shows an example visualization of a reference block / template for FInter prediction according to some embodiments of this disclosure.
[0040] Figures 22A and 22B illustrate flowcharts of exemplary decoding methods according to some embodiments of the present disclosure.
[0041] Figures 23A and 23B illustrate flowcharts of exemplary encoding methods according to some embodiments of the present disclosure.
[0042] Embodiments of this disclosure will be described with reference to the accompanying drawings. Detailed Implementation
[0043] While some configurations and arrangements have been discussed, it should be understood that this is for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of this disclosure. It will be apparent to those skilled in the art that this disclosure can also be used in a variety of other applications.
[0044] It should be noted that references to "one embodiment," "embodiment," "example embodiment," "some embodiments," "certain embodiments," etc., in the specification indicate that the described embodiments may include specific features, structures, or characteristics, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in connection with an embodiment, whether explicitly described or not, implementing such a feature, structure, or characteristic in conjunction with other embodiments will be within the knowledge of those skilled in the art.
[0045] Generally, terms can be understood, at least in part, from their usage in context. For example, the term "one or more," as used herein, can be used, at least in part, to describe any feature, structure, or characteristic in a singular sense, or in a plural sense, to describe a combination of features, structures, or characteristics. Similarly, terms such as "a," "an," or "the" can again be understood to convey either a singular or a plural usage, at least in part, depending on the context. Furthermore, the term "based on" can be understood not necessarily to convey an exclusive set of factors, but rather to allow for the presence of additional factors that are not necessarily explicitly described again, at least in part, depending on the context.
[0046] Various aspects of a video coding system will now be described with reference to various apparatuses and methods. These apparatuses and methods will be described in the following specific embodiments and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively, “elements”). These elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the system as a whole.
[0047] The techniques described herein can be used in a variety of video coding applications. As described herein, video coding includes both encoding and decoding video. Video encoding and decoding can be performed on a block-by-block basis. For example, encoding / decoding processes such as transform, quantization, prediction, loop filtering, reconstruction, etc., can be performed on encoded blocks, transform blocks, or prediction blocks. As described herein, the block to be encoded / decoded will be referred to as the "current block". For example, depending on the current encoding / decoding process, the current block can represent an encoded block, a transform block, or a prediction block. Furthermore, it should be understood that the term "unit" as used in this disclosure refers to a basic unit for performing a particular encoding / decoding process, and the term "block" refers to a sample array of a predetermined size. Unless otherwise stated, "block" and "unit" are used interchangeably.
[0048] Figure 1 shows a block diagram of an exemplary encoding system 100 according to some embodiments of the present disclosure. Figure 2 shows a block diagram of an exemplary decoding system 200 according to some embodiments of the present disclosure. Each system 100 or 200 can be applied to or integrated into a variety of systems and devices capable of data processing, such as computers and wireless communication devices. For example, system 100 or 200 can be all or part of a mobile phone, desktop computer, laptop computer, tablet computer, in-vehicle computer, game console, printer, positioning device, wearable electronic device, smart sensor, virtual reality (VR) device, augmented reality (AR) device, or any other suitable electronic device with data processing capabilities. As shown in Figures 1 and 2, system 100 or 200 may include a processor 102, a memory 104, and an interface 106. These components are shown as being connected to each other via a bus, but other connection types are also permitted. It should be understood that system 100 or 200 may include any other suitable components for performing the functions described herein.
[0049] Processor 102 may include microprocessors such as graphics processing units (GPUs), image signal processors (ISPs), central processing units (CPUs), digital signal processors (DSPs), tensor processing units (TPUs), vision processing units (VPUs), neural processing units (NPUs), coprocessor units (SPUs), or physical processing units (PPUs), microcontroller units (MCUs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), state machines, gated logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described throughout this disclosure. Although only one processor is shown in Figures 1 and 2, it should be understood that multiple processors may be included. Processor 102 may be a hardware device having one or more processing cores. Processor 102 may execute software. Software should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, application programs, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description languages, or others. Software can include computer instructions written in interpreted languages, compiled languages, or machine code. Other technologies used to direct hardware also fall under the broad category of software.
[0050] Memory 104 can broadly include both memory (also known as main / system memory) and storage devices (also known as auxiliary memory). For example, memory 104 may include random access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferroelectric RAM (FRAM), electrically erasable programmable ROM (EEPROM), optical disc read-only memory (CD-ROM) or other optical disc storage devices, hard disk drive (HDD) (such as disk storage devices or other magnetic storage devices), flash memory drive, solid-state drive (SSD), or any other medium that can be used to carry or store desired program code in the form of instructions that can be accessed and executed by processor 102. More broadly, memory 104 can be embodied in any computer-readable medium (such as non-transitory computer-readable media). Although only one memory is shown in Figures 1 and 2, it should be understood that multiple memories may be included.
[0051] Interface 106 can broadly include data interfaces and communication interfaces, the latter being configured to receive and transmit signals in the process of receiving and transmitting information with other external network elements. For example, interface 106 may include input / output (I / O) devices and wired or wireless transceivers. Although only one interface is shown in Figures 1 and 2, it should be understood that multiple interfaces may be included.
[0052] Processor 102, memory 104, and interface 106 may be implemented in various forms within system 100 or 200 for performing video encoding functions. In some embodiments, processor 102, memory 104, and interface 106 of system 100 or 200 are implemented (e.g., integrated) on one or more system-on-a-chip (SoC). In one example, processor 102, memory 104, and interface 106 may be integrated into an application processor (AP) SoC that handles application processing within an operating system (OS) environment, including running video encoding and decoding applications. In another example, processor 102, memory 104, and interface 106 may be integrated into a dedicated processor chip for video encoding, such as a GPU or ISP chip dedicated to image and video processing in a real-time operating system (RTOS).
[0053] As shown in Figure 1, in the encoding system 100, the processor 102 may include one or more modules, such as encoder 101. Although Figure 1 shows encoder 101 within a single processor 102, it should be understood that encoder 101 may include one or more sub-modules that may be implemented on different processors, either close to or far from each other. Encoder 101 (and any corresponding sub-modules or sub-units) may be a hardware unit (e.g., a portion of an integrated circuit) of processor 102, designed for use with other components or software units implemented by processor 102 by executing at least a portion of a program (e.g., instructions). The instructions of the program may be stored on a computer-readable medium such as memory 104, and when executed by processor 102, it may perform processes having one or more functions related to video encoding, such as image partitioning, inter-frame prediction, intra-frame prediction, transform, quantization, filtering, entropy coding, etc., as described in detail below.
[0054] Similarly, as shown in Figure 2, in the decoding system 200, the processor 102 may include one or more modules, such as the decoder 201. Although Figure 2 shows the decoder 201 within a single processor 102, it should be understood that the decoder 201 may include one or more sub-modules that may be implemented on different processors, either close to or far from each other. The decoder 201 (and any corresponding sub-modules or sub-units) may be a hardware unit (e.g., a portion of an integrated circuit) of the processor 102, designed for use with other components or software units implemented by the processor 102 by executing at least a portion of a program (e.g., instructions). The instructions of the program may be stored on a computer-readable medium such as memory 104, and when executed by the processor 102, it may perform processes having one or more functions related to video decoding, such as entropy decoding, inverse quantization, inverse transform, inter-frame prediction, intra-frame prediction, filtering, etc., as described in detail below.
[0055] Figure 3 shows a detailed block diagram of an exemplary encoder 101 in the encoding system 100 of Figure 1 according to some embodiments of the present disclosure. As shown in Figure 3, the encoder 101 may include a partitioning module 302, an inter-frame prediction module 304, an intra-frame prediction module 306, a transform module 308, a quantization module 310, a dequantization module 312, an inverse transform module 314, a filter module 316, a buffer module 318, and an encoding module 320. It should be understood that each element shown in Figure 3 is shown independently to represent different features in the video encoder, and does not mean that each component is formed by a separate hardware configuration unit or a single software. That is, for ease of illustration, each element is listed as a unit, and at least two elements may be combined to form a single element, or a single element may be divided into multiple elements to perform functions. It should also be understood that some elements are not essential elements for performing the functions described in the present disclosure, but may be optional elements for improving performance. It should also be understood that these elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these components are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on encoder 101.
[0056] The partitioning module 302 can be configured to partition the input image of a video into at least one processing unit. The image can be a frame or a field of a video. In some embodiments, the image includes a monochrome luminance sample array, or a luminance sample array and two corresponding chrominance sample arrays. In this case, the processing unit can be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The partitioning module 302 can partition the image into a combination of multiple coding units, prediction units, and transform units, and encode the image by selecting the combination of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function).
[0057] Similar to H.265 / HEVC, H.266 / VVC is a block-based hybrid spatial and temporal predictive coding scheme. As shown in Figure 5, during encoding, the input image 500 is first divided into multiple square blocks – CTUs 502 – by a partitioning module 302. For example, a CTU 502 can be a 128×128 pixel block. As shown in Figure 6, each CTU 502 in the input image 500 can be further divided by the partitioning module 302 into one or more CUs 602, which can be used for prediction and transformation. Unlike H.265 / HEVC, in H.266 / VVC, CUs 602 can be rectangular or square and can be encoded without further partitioning into prediction units or transformation units. For example, as shown in Figure 6, partitioning a CTU 502 into CUs 602 can include quadtree partitioning (indicated by solid lines), binary tree partitioning (indicated by dashed lines), and ternary tree partitioning (indicated by dotted lines). According to some embodiments, each CU 602 can be as large as its root CTU 502, or as small as a 4×4 block subdivision of the root CTU 502.
[0058] Referring to Figure 3, the inter-frame prediction module 304 can be configured to perform inter-frame prediction on the prediction unit, and the intra-frame prediction module 306 can be configured to perform intra-frame prediction on the prediction unit. It can be determined whether to use inter-frame prediction or perform intra-frame prediction for the prediction unit, and specific information (e.g., intra-frame prediction mode, motion vector, reference image, etc.) is determined based on each prediction method. In this case, the processing unit used to perform the prediction can be different from the processing unit used to determine the prediction method and specific content. For example, the prediction method and prediction mode can be determined in the prediction unit, and the transformation can be performed in the transform unit. The residual coefficients in the residual block between the generated prediction block and the original block can be input to the transform module 308. Additionally, the prediction mode information, motion vector information, etc., used for prediction can be encoded into the bitstream by the encoding module 320 along with the residual coefficients or quantization level. It should be understood that in some encoding modes, the original block can be encoded as is without generating a prediction block through the prediction modules 304 or 306. It should also be understood that in some encoding modes, prediction, transformation, and / or quantization can be skipped.
[0059] In some embodiments, the inter-frame prediction module 304 may predict prediction units based on information about at least one image preceding or following the current image, and in some cases, the inter-frame prediction module 304 may predict prediction units based on information about a portion of the currently encoded image that has already been encoded. The inter-frame prediction module 304 may include sub-modules such as a reference image interpolation module, a motion prediction module, and a motion compensation module (not shown). For example, the reference image interpolation module may receive reference image information from the buffer module 318 and generate pixel information for an integer number of pixels from the reference image. In the case of luminance pixels, an 8-tap interpolation filter based on Discrete Cosine Transform (DCT) with varying filter coefficients may be used to generate pixel information for an integer number of pixels in units of 1 / 4 pixels. In the case of chrominance signals, a 4-tap interpolation filter based on DCT with varying filter coefficients may be used to generate pixel information for an integer number of pixels in units of 1 / 8 pixels. The motion prediction module may perform motion prediction based on a reference image interpolated by the reference image interpolation portion. Various methods, such as Full-Search-Based Block Matching (FBMA), Three-Step Search (TSS), and the novel Three-Step Search (NTS), can be used to compute motion vectors. Based on interpolated pixels, motion vectors can have values in units of 1 / 2, 1 / 4, or 1 / 16 pixels, or integer pixels. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods, such as skipping methods, merging methods, Advanced Motion Vector Prediction (AMVP) methods, and intra-block copying methods, can be used as motion prediction methods.
[0060] Referring again to FIG3, in some embodiments, the intra-prediction module 306 may generate prediction units based on information of reference pixels surrounding the current block (which are pixel information in the current image). The reference pixels may be located in reference lines that are not adjacent to the current block. When a block in the neighborhood of the current prediction unit is a block that has already undergone inter-frame prediction and therefore the reference pixel is a pixel that has already undergone inter-frame prediction, the reference pixel information of the block in the neighborhood that has undergone intra-frame prediction can be replaced by a reference pixel included in the block that has already undergone inter-frame prediction. That is, when a reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information. In intra-frame prediction, the prediction mode may have an angular prediction mode that uses reference pixel information according to the prediction direction and a non-angular prediction mode that does not use direction information when performing prediction. The mode used to predict luminance information may be different from the mode used to predict chromatic difference information, and chromatic difference information can be predicted using intra-frame prediction mode information used to predict luminance information or predicted luminance signal information. If the size of the prediction unit is the same as the size of the transform unit when performing intra-prediction, intra-prediction can be performed based on the left, top-left, and top pixels of the prediction unit. However, if the size of the prediction unit is different from the size of the transform unit when performing intra-prediction, intra-prediction can be performed using reference pixels based on the transform unit.
[0061] Intra-prediction methods can generate prediction blocks after applying an adaptive intra-smoothing (AIS) filter to a reference pixel according to the prediction mode. The type of AIS filter applied to the reference pixel can vary. To perform the intra-prediction method, the intra-prediction mode of the current prediction unit can be predicted based on the intra-prediction modes of prediction units existing in the neighborhood of the current prediction unit. When using mode information predicted from neighboring prediction units to predict the prediction mode of the current prediction unit, if the intra-prediction mode of the current prediction unit is the same as the intra-prediction modes of the neighboring prediction units, predetermined flag information can be used to send information indicating that the prediction mode of the current prediction unit is the same as the prediction modes of the neighboring prediction units; and if the prediction mode of the current prediction unit and the prediction modes of the neighboring prediction units are different from each other, additional flag information can be used to encode the prediction mode information of the current block.
[0062] As shown in Figure 3, a residual module can be generated. This residual module includes a prediction module that performs prediction based on the prediction module generated by prediction module 304 or 306 and residual coefficient information (also referred to herein as "residual"), which is the difference between the prediction unit and the original block. The generated residual block can be input into the transform module 308. More details regarding residuals and transforms used for video coding will now be provided.
[0063] In hybrid video coding systems, redundancy in the video signal is first exploited by applying inter-frame or intra-frame prediction tools to each control unit (CU). The difference between the original samples of a CU and its predicted block is often referred to as the residual. Even after prediction, the residual can still be highly spatially correlated. While conditional entropy coding can capture some spatial dependencies between adjacent samples, it is computationally impractical to formulate an entropy coding statistical model that fully utilizes the spatial correlation in the residual. In contrast, transform coding is a practical and efficient method for spatial decorrelation of the residual.
[0064] For example, the transformation module 308 can use an integer version of the two-dimensional discrete cosine transform (DCT) to transform the residuals, which can be applied independently in the horizontal and vertical directions. For an M×N residual sample block (where M is the width of the block and N is the height of the block), the transformation module 308 can obtain the transformation coefficients by applying the M×M DCT to each row to generate intermediate transformation coefficients, and then applying the N×N DCT to each column of the intermediate transformation coefficients.
[0065] For intra-coded CUs (also referred to as “intra-CUs” in this document), spatially adjacent reconstructed samples are used to predict the current block, and the intra-prediction mode is signaled once for the entire CU. Each CU consists of one or more coded blocks (CBs) corresponding to the color components of the video sequence. For example, consumer video typically uses a 4:2:0 chroma format, in which case each CU consists of a luma CB and two chroma CBs with a quarter sample of the luma CB. Intra-prediction and transform coding are performed at the prediction block (PB) and transform block (TB) levels, respectively. Each CB consists of a single TB, except in the case of intra-fractional subdivision (ISP) mode and implicit splitting. For luma CBs, the maximum side length of a TB is 64, and the minimum side length is 4. Furthermore, luma TBs are further specified as W×H rectangular blocks of width W and height H, where W, H ∈ {4, 8, 16, 32, 64}. For chroma CBs, the maximum TB side length is 32, and the chroma TB is a W×H rectangular block of width W and height H. Here, W, H ∈ {2, 4, 8, 16, 32}, but blocks with shapes of 2×H and 4×2 are excluded in order to address memory architecture and throughput requirements.
[0066] Figure 7 shows a schematic visualization 700 of the current CU block 702 and reconstructed samples that are spatially adjacent and non-adjacent to the current block according to some aspects of this disclosure. In Figure 7, the numbers 0, 1, 2, ... indicate the pixel line indices associated with the current CU block 702.
[0067] In VVC, reference samples obtained from reconstructed samples of neighboring blocks are used to generate intra-prediction samples for the current block. For a W×H block, the reference samples are spatially adjacent to the current block and consist of a vertical line of 2.H reconstructed samples extending downwards to the left of the block, a reconstructed sample at the top left corner, and a horizontal line of 2.W reconstructed samples extending to the right above the current block. This set of "L"-shaped samples may be referred to as "reference lines" in this disclosure. The reference line directly adjacent to the current CU block 702 is shown as the line with index 0 in Figure 7.
[0068] Similar to AVC and HEVC, VVC also supports intra-angle prediction modes. Intra-angle prediction is a directional intra-prediction method. Compared to HEVC, VVC's intra-angle prediction is modified by increasing prediction accuracy and adapting to the new partitioning frame. The former is achieved by increasing the number of angle prediction directions and using more accurate interpolation filters, while the latter is achieved by introducing a wide-angle intra-prediction mode. In VVC, the number of directional modes available for a given block increases from 33 HEVC directions to 65 directions. Figure 8 depicts VVC's angle mode 800.
[0069] Orientations with even indices between 2 and 66 are equivalent to the orientations of the angle modes supported in HEVC. For square-shaped blocks, an equal number of angle modes are assigned to the top and left sides of the block. On the other hand, rectangular intra-blocks, which do not exist in HEVC, are the central part of the VVC partitioning scheme, where additional intra-prediction orientations are assigned to the longer sides of the block. The additional modes assigned along the longer sides are called Wide-Angle Intra-Prediction (WAIP) modes because they correspond to prediction orientations with an angle greater than 45° relative to the horizontal or vertical modes. As shown in Figure 8, a WAIP mode for a given mode index is defined by mapping the original orientation mode to a mode with the opposite orientation and an index offset of 1. For a given rectangular block, the aspect ratio, i.e., the ratio of width to height, is used to determine which angle modes will be replaced by the corresponding wide-angle modes.
[0070] For a square block in VVC, each pair of horizontally or vertically adjacent predicted samples is predicted from a pair of adjacent reference samples. In contrast, WAIP extends the angular range of directional prediction to more than 45°, so for a coded block predicted using the WAIP pattern, adjacent predicted samples can be predicted from non-adjacent reference samples.
[0071] In addition to the directly adjacent lines of adjacent samples, one of the two non-adjacent reference lines (line 1 and line 2) depicted in Figure 7 can include input samples for intra-frame prediction in VVC. For ECM, more non-adjacent reference lines can be used. The use of adjacent and non-adjacent reference samples is called multi-reference line (MRL) prediction.
[0072] The intra-frame modes available for MRL are DC mode and angle prediction mode. However, not all of these modes can be combined with MRL for a given block. The MRL mode is always coupled with a mode from the most probable mode (MPM) list in VVC. This coupling means that if non-adjacent reference lines are used, the intra-frame prediction mode will be one of the MPMs. The design motivation for this MPM-based MRL prediction mode stems from the observation that non-adjacent reference lines primarily favor texture patterns with sharp / clear and strongly oriented edges. In these cases, MPMs are chosen more frequently because there is often a strong correlation between the texture patterns of neighboring blocks and the current block. On the other hand, choosing a non-MPM for intra-frame prediction is an indication that edges are inconsistently distributed in neighboring blocks, and therefore, the MRL prediction mode is not expected to be very useful in this case. Furthermore, it has been observed that when the intra-frame prediction mode is a planar mode, MRL does not provide additional coding gain because this mode is typically used for smooth regions. Therefore, MRL excludes planar modes, which are always one of the MPMs. The angle or DC prediction process in MRL is very similar to the case of directly adjacent reference lines. However, for angular patterns with non-integer slopes, DCT-based interpolation filters (DCTIF) are always used. This design choice is supported by both experimental results and empirical observations that MRL is best suited for sharp and strongly directional edges, where DCTIF is more appropriate because it preserves more high frequencies than some other filters.
[0073] From a hardware design perspective, applying multiple reference lines, as proposed in the initial approach, requires additional cost for line buffers used to hold / save the additional reference lines. In typical hardware designs, line buffers are part of the on-chip memory architecture used for image and video encoding, and minimizing their on-chip area is crucial. To address this, the MRL is disabled, and no signal is given for encoding units attached to the top boundary of the CTU. In this way, the additional buffers used to hold / save non-adjacent reference lines are bounded by 128, which is the width of the maximum unit size.
[0074] In some known methods, intra-prediction fusion methods have been proposed to improve the accuracy of intra-prediction. More specifically, if the current block is a luma block, and it is encoded using a non-integer slope angle mode instead of the ISP mode, and the block size (width * height) is greater than 16, then two prediction blocks generated from two different reference lines will be "fused," where the prediction fusion is calculated as a weighted sum of the two prediction blocks. More specifically, the first reference line at index i is specified using the current signaling method in the bitstream (…). The prediction block generated from the reference line using the selected intra-frame prediction mode is represented as... ,in This refers to the operation of generating prediction blocks from a reference line with a given intra-frame prediction mode. In known methods, the reference line... It is implicitly chosen as the second reference line. That is, the second reference line is an index position further away from the current block relative to the first reference line. Similarly, the predicted block generated from the second reference line is represented as... The weighted sum of the two prediction blocks is obtained as follows and used as the predictor for the current block according to equation (1).
[0075] (1) Among them, Indicates fusion prediction, and These are two weighting factors, which were set to 3 / 4 and 1 / 4 respectively in the experiment.
[0076] Among known methods, Spatial Geometric Partitioning Mode (SGPM) has been proposed, which allows the CU to be divided into two parts that can use different intra-prediction modes. This new mode is conceptually similar to Geometric Partitioning Mode (GPM), which is applied to inter-frame prediction in VVC. However, since a large number of combinations of partitioning and intra-prediction modes are possible, SGPM uses a different signaling mechanism. To more efficiently express the necessary partitioning and prediction information in the bitstream, a candidate list is employed, and candidate indices are signaled only in the bitstream. As shown in the representation of SGPM signaling 900 in Figure 9A, each candidate in the list can derive a combination of one partitioning mode and two intra-prediction modes.
[0077] A template is used to generate the candidate list. An example of template 901 is shown in Figure 9B, where the template width is set to 4. In the current version of SGPM used by ECM, the template width is set to 1. For each possible combination of one partitioning mode and two intra-prediction modes, a prediction is generated for the template, and the partitioning weights are extended to the template. These combinations are ordered in ascending order of the sum of the absolute transform difference (SATD) costs between the template's prediction and reconstruction. The length of the candidate list is set to 16, and these candidates are considered the most probable SGPM combinations for the current block. Both the encoder and decoder use the template to build the same candidate list. To reduce the complexity of building the candidate list, the number of possible partitioning modes and the number of possible intra-prediction modes (IPMs) are limited. For example, in the current version of SGPM used by ECM, the number of possible partitioning modes is limited to a predefined set of 26 partitions, covering a variety of partitioning orientations and locations.
[0078] Each of the two IPM candidate lists corresponding to the two SGPM partitions is constructed by adding available IPM candidates and then, if necessary, pruning to a predefined limit of three candidates. Some IPM candidates are inherited from intra-inter-frame GPM modes already adopted in the ECM. Figure 10 is a schematic visualization 1000 of inter-frame angle modes parallel to the GPM partition boundary (parallel mode) (see (a)), inter-frame angle modes perpendicular to the GPM partition boundary (vertical mode) (see (b)), and inter-frame prediction plane modes (see (c)) according to some embodiments of this disclosure, as well as SGPM with intra-frame and inter-frame prediction (see (d)). These modes can be added as available IPM candidates for SGPM.
[0079] Additionally, template-based intra-mode derivation (TIMD) can be used to derive intra-prediction modes as available IPM candidates for SGPM. For example, TIMD intra-prediction modes for IPMs forming SGPM can be derived using only horizontally and vertically oriented adjacent blocks (using top or left templates).
[0080] For some CU block sizes, SGPM can be implicitly disabled. The range of block sizes for which SGPM can be used (e.g., a CU level flag can be signaled to indicate whether SGPM is used) was originally inherited from the intra-inter-frame GPM mode. In the current version of SGPM adopted by ECM, the range of block sizes has been further extended to 4. 8, 8 4, 4 16 and 16 Smaller blocks of size 4. In summary, the block size that can be used with SGPM can be described by the following rules: 4 <= width <= 64, 4 <= height <= 64, width < height * 8, height < width * 8, width * height >= 32.
[0081] To better predict pixels located at the boundary between two prediction portions, an adaptive SGPM blending scheme 1100 can be used, as shown in Figure 11, where a weighted average of the two prediction portions is used in the transition region surrounding the SGPM partition. The width of the transition region, called the blending width, is adaptively determined based on the block size. Adaptive blending requires no signaling.
[0082] Referring to Figure 11, let the blend width specified for the GPM tool in VVC and ECM be... Then, the adaptive SGPM blend width can be determined based on the CU block width and height, as follows: • If min(width, height) == 4, then select... Otherwise, if min(width, height) == 8, then select... Otherwise, if min(width, height) == 16, then select... Otherwise, if min(width, height) == 32, then select... ; and otherwise, choose .
[0083] Figure 12 illustrates an example visualization 1200 of IBC prediction according to some embodiments of this disclosure. When predicting a CU via intra-block copy mode, a block vector (BV) is signaled to indicate which block within the same image will be copied to be used as the predictor for the current block. Signaling the block vector can be performed by signaling the block vector difference (BVD) in the bitstream, making it possible to determine the block vector by adding the BVD to the block vector predictor. Alternatively, if the block vector from a previous CU is an exact match for the current block vector, it can be signaled via a merge flag. Regardless of the signaling mechanism used, the block vector points to a location within the same image to indicate a sample block of the same size as the current CU, which is used as the predictor block for the current CU. Some limitations can be applied to the block vector, as it must point to a location in the current image that has been decoded before the current CU. The illustration in Figure 12 summarizes the IBC concept in HEVC and VVC, where each tiled square shape in the figure represents a coding tree unit (CTU). Gray shaded areas represent regions that have already been encoded, while white shaded areas represent regions to be encoded. In HEVC, IBC generally allows the BV to point to any block contained within the gray shaded area. This freedom is partially restricted when the `sps_entropy_coding_sync_enabled_flag` is signaled in the bitstream to support wavefront parallel processing (WPP). In this case, the IBC block vector cannot point to any region within two or more CTUs to the right of the current CTU in the CTU row directly above, as shown by the crossed-out CTUs in Figure 12. IBC in VVC is significantly more restricted, only allowing the block vector to use the CTU to the left of the current CTU as a reference region, indicated by the dashed box. Compared to VVC, the current IBC tool in ECM expands the search range.
[0084] Figure 13 illustrates Figure 1300 of intra-template matching prediction (intraTMP) according to some embodiments of the present disclosure.
[0085] Referring to Figure 13, intraTMP is an intra-frame prediction mode similar to IBC, because the current CU 1322 is also predicted from sample blocks from the current image. IntraTMP can be selected only as the prediction mode for CUs of size 64x64 or smaller. However, unlike IBC, in intraTMP, the block vector 1324 is not signaled in the bitstream. Instead, the decoder 201 compares a predefined L-shaped or other shaped template of the reconstructed samples adjacent to the current CU 1322 with templates of the same shape from candidate predictors within a predetermined search area. For L-shaped templates, the left and top adjacent samples of the current CU 1322 or the intraTMP predictor 1326 are used. Let the width of the left template region be TmpW, and the width of the top template region be TmpH. Other template shapes include a left template and a top template, where the left template includes only the template region on the left, and the top template includes only the template region on the top.
[0086] The intraTMP predictor block is determined by finding the best candidate template that matches the current CU template. This can be done by finding a template that minimizes the sum of absolute differences (SAD) or the sum of absolute transformed differences (SATD), or by comparing the hash values between templates. The search algorithm within the search region can be exhaustive (e.g., by scanning templates across the search region using a sample resolution offset) or fast (e.g., by performing a coarse search first, followed by a local fine search around the best match location obtained from the coarse search). In any case, the search algorithm performed by the encoder and decoder is identical, so that both encoder 101 and decoder 201 implicitly know the intraTMP predictor without needing to be notified via signaling in the bitstream. Figure 13 shows an example of intraTMP, where the current CU template and the best matching template are shaded with diagonal lines.
[0087] Referring again to Figure 13, for the intraTMP predictor 1326 to be selected, the sample block corresponding to the intraTMP predictor 1326 must be completely contained within the search region. The search region is shown in Figure 13 by dashed shading. Within the current CTU 1320, the search region is restricted to a rectangular sample block bounded at one corner by the upper left corner of the current CTU 1320 and at the other corner by the upper left corner of the current CU 1322.
[0088] Outside the current CTU 1320, the search region is limited by imposing a maximum length on the intraTMP block vector of (searchRangeWidth 1328, searchRangeHeight 1330), where searchRangeWidth 1328 and searchRangeHeight 1330 are set to be proportional to the size of the current CU 1322. That is, searchRangeWidth = a * BlkW and searchRangeHeight = a * BlkH, where "a" is a constant controlling the gain / complexity tradeoff, and BlkW and BlkH are the width and height of the current CU 1322, respectively. Here, "a" is set to 5 in the ECM-7.0 test software. searchRangeHeight 1330 limits the length of the block vector only in the negative vertical direction (i.e., in the direction to the top of the image). For block vectors with a positive vertical component, the search region is limited by the bottom boundary of the current CTU row. For example, in Figure 13, the search area extends to the bottom boundary of the left CTU 1332, regardless of the value of searchRangeHeight 1330. Furthermore, these limitations on the search range do not apply to the current CTU 1320. For instance, for a small CU where searchRangeWidth 1328 and searchRangeHeight 1330 may be smaller compared to the size of the current CTU 1320, the search area still extends to the top left corner of the current CTU 1320.
[0089] In addition to the constraints imposed by the search region, the intraTMP predictor 1326 and its template must consist of samples available for intra-frame prediction. For example, the boundaries of the search region are still covered by image, tile, or patch boundaries. Let the coordinates of the top-left corner of the current CU relative to the current image be (currCuX, currCuY). Then, the left boundary of the intraTMP search region is initially intraTmpLeftBound = currCUX - searchRangeWidth. To account for image boundaries, the left boundary is cropped to allow a TmpW sample width for the predictor template. intraTmpLeftBound = max(intraTmpLeftBound, TmpW). To accelerate the template matching process, the search region is initially traversed horizontally or vertically in increments of 2 pixels each time. This is also referred to as a search subsampling factor of 2. This results in a 4x reduction in template matching search complexity. After finding the best match from the initial search, a correction process is performed. The correction is accomplished by performing a second template matching search around the best match with a reduced range. In ECM-7.0, the reduced range is set to BlkH / 2.
[0090] Figure 1400 illustrates a decoder-side intra-frame mode derivation (DIMD) mode for ECM according to some embodiments of the present disclosure.
[0091] Referring to Figure 14, in DIMD, intra-prediction modes (or multiple intra-prediction modes) are implicitly derived from an L-shaped template 1404 (hereinafter referred to as "template 1404") of reconstructed samples adjacent to the current CU 1402. The template 1404 is 3 sample widths in size. The decoder 201 moves a 3×3 gradient analysis window 1406 over the template 1404. At each location, a local gradient is computed by applying a Sobel filter. Assume the 3x3 sample set at a location in the template 1404 is... Then the Sobel filter can be described by equation (2).
[0092] (2).
[0093] Then, the horizontal gradient is estimated by taking the dot product shown in equations (3) and (4) below, respectively. and vertical gradient .
[0094] (3); and (4).
[0095] The magnitude of the local gradient and the angle of local gradient It can be estimated based on equations (5) and (6) respectively.
[0096] (5); and (6).
[0097] Angle of local gradient It can predict direction with intra-frame angle. Related. For example, a 0-degree angle corresponds to horizontal intra-frame prediction mode 18. In fact, Decoder 201 can utilize fast implementations such as table lookups from... and Direct estimation. At the start of the DIMD method, an empty histogram H is initialized (filled with zeros in each entry), the size of which is equal to the number of intra-predictive modes. When the DIMD method performs gradient analysis on each local window, the histogram H is updated according to equation (7).
[0098] (7).
[0099] Therefore, each local gradient analysis "votes" for an intra-prediction mode. At the end of the DIMD method, the intra-prediction mode with the highest count in H can be selected as the single representative intra-prediction mode for the current CU. Alternatively, multiple intra-prediction modes can be obtained from H in descending order of their counts.
[0100] Figure 15 illustrates a template-based intra-frame mode derivation (TIMD) pattern for ECM according to some embodiments of the present disclosure. Referring to Figure 15, in TIMD, an intra-frame prediction pattern (or multiple intra-frame prediction patterns) is implicitly derived from templates 1504 above and to the left of the reconstructed sample adjacent to the current CU 1502.
[0101] A set of candidate intra-prediction modes is searched from the Most Probable Mode (MPM) list, which is constructed from the intra-prediction modes used by neighboring CUs. Then, for each candidate intra-prediction mode, the prediction result of template 1504 is generated based on template reference sample 1506 using the intra-angle prediction method. The candidate intra-prediction mode that produces the template predictor that best matches template 1504 is selected as the TIMD intra-prediction mode. The best match can be determined by finding the predictor that minimizes the sum of absolute differences (SAD) or the sum of absolute transform differences (SATD), or by comparing the hash values between the predictor and the template. Alternatively, multiple intra-prediction modes can be obtained in ascending order of SAD / SATD.
[0102] Matrix-weighted intra-prediction (MIP), IBC, and IntraTMP methods can be effective intra-prediction modes in ECM. Greater gains can be achieved by combining MIP with a non-separable primary transform (NSPT), IBC with LFNST and NSPT, and IntraTMP with LFNST and NSPT. The derivation of the intra-prediction modes is described below with reference to the solution in Figure 4.
[0103] Transform module 308 can transform the video signal in the residual block from the pixel domain to the transform domain (e.g., to the frequency domain, depending on the transform method). It should be understood that in some examples, transform module 308 can be skipped, and the video signal may not be transformed to the transform domain.
[0104] Quantization module 310 is configurable to quantize the coefficients at each location in the coded block to generate a quantization level for each location. The current block can be a residual block. That is, quantization module 310 can perform quantization processing on each residual block. A residual block can include N×M locations (samples), each location being associated with a transformed or untransformed video signal / data (such as luminance and / or chrominance information), where N and M are positive integers. In this disclosure, the transformed or untransformed video signal at a particular location before quantization is referred to herein as a “coefficient”. After quantization, the quantized value of the coefficient is referred to herein as a “quantization level” or “level”.
[0105] Quantization can be used to reduce the dynamic range of transformed or untransformed video signals, allowing fewer bits to be used to represent the signal. Quantization typically involves dividing by the quantization step size and subsequent rounding, while inverse quantization (also known as dequantization) involves multiplying by the quantization step size. The quantization step size can be indicated by the quantization parameter (QP). This quantization process is called scalar quantization. Quantization of all coefficients within a coded block can be performed independently, and this method has been adopted by several existing video compression standards, such as H.264 / AVC and H.265 / HEVC. The QP in quantization can affect the bitrate used to encode / decode the image in the video. For example, a higher QP results in a lower bitrate, while a lower QP results in a higher bitrate.
[0106] For an N×M coded block, the two-dimensional (2D) coefficients of the block can be converted into a one-dimensional (1D) sequence for coefficient quantization and encoding using a specific coding scan order. Typically, the coding scan begins at the top left corner and stops at the bottom right corner or the last non-zero coefficient / level in the bottom-right direction of the coded block. It should be understood that the coding scan order can include any suitable order, such as a zigzag scan order, a vertical (column) scan order, a horizontal (row) scan order, a diagonal scan order, or any combination thereof. The quantization of coefficients within the coded block can utilize coding scan order information. For example, it can depend on the state of the previous quantization level along the coding scan order. To further improve coding efficiency, the quantization module 310 can use more than one quantizer (e.g., two scalar quantizers). Which quantizer will be used to quantize the current coefficient can depend on information preceding the current coefficient in the coding scan order. Such a quantization process is called dependent quantization.
[0107] Referring to Figure 3, the encoding module 320 can be configured to encode the quantization level at each position in the coded block into a bitstream. In some embodiments, the encoding module 320 can perform entropy coding on the coded block. Entropy coding can use various binarization methods (such as Golomb-Rice binarization) to convert each quantization level into a corresponding binary representation (such as binary bits). The binary representation can then be further compressed using an entropy coding algorithm. The compressed data can be added to the bitstream. In addition to quantization levels, the encoding module 320 can encode various other information, such as block type information of the coding unit, prediction mode information, partitioning unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information input from, for example, prediction modules 304 and 306. In some embodiments, the encoding module 320 can perform residual coding on the coded block to convert the quantization levels into a bitstream. For example, after quantization, for an N×M block, there can be N×M quantization levels. These N×M levels can be zero or non-zero values. If the non-zero levels are not binary, they can be further binarized into binary bins, for example, using a binarization method that combines truncated Rice (TR) codes and finite K-order Exp-Golomb codes (EGk).
[0108] Non-binary syntax elements can be mapped to binary codewords. The bijective mapping between symbols and codewords (usually using simple structured coding) is called binarization. Binary arithmetic coding can be used to encode both binary syntax elements and the binary symbols (also called bins) used for non-binary data. The core coding engine of Context Adaptive Binary Arithmetic Coding (CABAC) supports two operating modes: context coding mode (where bins are encoded using an adaptive probability model) and a less complex bypass mode using a fixed probability of 1 / 2. The adaptive probability model is also called the context, and assigning the probability model to each bin is called context modeling.
[0109] As shown in Figure 3, the dequantization module 312 can be configured to inverse quantization of the quantization level, and the inverse transform module 314 can be configured to inverse transform the coefficients transformed by the transform module 308. The reconstructed residual block generated by the dequantization module 312 and the inverse transform module 314 can be combined with the prediction unit predicted by the prediction module 304 or 306 to generate a reconstructed block.
[0110] Filter module 316 may include at least one of a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can eliminate block distortion caused by boundaries between blocks in the reconstructed image. The SAO module can correct the offset relative to the original video on a pixel-by-pixel basis for the video that has been deblocked. ALF can be performed based on values obtained by comparing the reconstructed and filtered video with the original video. Buffer module 318 can be configured to store the reconstructed blocks or images calculated by filter module 316, and can provide the reconstructed and stored blocks or images to inter-frame prediction module 304 when inter-frame prediction is performed.
[0111] Figure 4 shows a detailed block diagram of an exemplary decoder 201 in the decoding system 200 of Figure 2 according to some embodiments of the present disclosure. As shown in Figure 4, the decoder 201 may include a decoding module 402, a dequantization module 404, an inverse transform module 406, an inter-frame prediction module 408, an intra-frame prediction module 410, a filter module 412, and a buffer module 414. It should be understood that each element shown in Figure 4 is shown independently to represent different features in the video decoder, and does not mean that each component is formed by a separate hardware or single software configuration unit. That is, for ease of illustration, each element is listed as an element, and at least two elements may be combined to form a single element, or an element may be divided into multiple elements to perform functions. It should also be understood that some elements are not essential elements for performing the functions described in the present disclosure, but may be optional elements for improving performance. It should also be understood that these elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the decoder 201.
[0112] When a video bitstream / bitstream is input from a video encoder (e.g., encoder 101), the input bitstream can be decoded by decoder 201 in the reverse process of the video encoder. Therefore, for ease of description, some decoding details described above regarding encoding can be omitted. Decoding module 402 can be configured to decode the bitstream to obtain various information encoded into the bitstream, such as the quantization level at each position in the coded block. In some embodiments, decoding module 402 can perform entropy decoding (decompression) corresponding to entropy encoding (compression) performed by the encoder, such as, for example, variable-length coding (VLC), context-adaptive variable-length coding (CAVLC), CABAC, syntax-based binary arithmetic coding (SBAC), PIPE coding, etc., to obtain a binary representation (e.g., binary bins). Decoding module 402 can also convert the binary representation to a quantization level using Columbus-Rice binarization (including, for example, EGk binarization and binarization combining TR and finite EGk). In addition to the quantization level of the position in the transform unit, the decoding module 402 can also decode various other information, such as parameters used for Columbus-Rice binarization (e.g., Rice parameters), block type information of the coding unit, prediction mode information, partitioning unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. During the decoding process, the decoding module 402 can perform rearrangement on the bitstream to reconstruct and rearrange the data from 1D order into 2D rearranged blocks by a reverse scan method based on the coding scan order used by the encoder.
[0113] The dequantization module 404 can be configured to dequantize the quantization level at each location of the encoded block (e.g., a 2D reconstructed block) to obtain coefficients at each location. In some embodiments, the dequantization module 404 can also perform dependent dequantization based on quantization parameters provided by the encoder, which include information related to the quantizers used in dependent quantization, such as the quantization step size used by each quantizer.
[0114] The inverse transform module 406 can be configured to perform inverse transforms, such as inverse DCT, inverse Discrete Sine Transform (DST), and inverse KLT, LFNST, and / or NSPT, respectively, performed by the encoder, to transform data from the transform domain (e.g., coefficients) back to the pixel domain (e.g., luminance and / or chrominance information). In some embodiments, the inverse transform module 406 can selectively perform transform operations (e.g., DCT, DST, KLT, LFNST, NSPT) based on multiple pieces of information such as the prediction method, the size of the current block, and the prediction direction.
[0115] For example, when a separable transformation applies a one-dimensional transformation in the horizontal and vertical directions respectively, a two-dimensional inseparable transformation is applied directly to the input sample block. An desirable property of the transformation is that the transformation vector spans the space of the input samples. This means that any input vector (e.g., any combination of input sample values) can be represented by a weighted sum of transformation vectors. For a transformation to span, a necessary condition is that the number of variable vectors must be at least as many as the dimension of the input space, or in other words, the number of output transformation coefficients must be at least equal to the number of input samples. For example, a one-dimensional DCT in VVC is a spanning transform. Then, for a spanning inseparable transformation, if the input sample block is a... If there is a residual block, the transformation will also output a... Transform coefficient block, this process can be achieved through ( )×( This is achieved through matrix multiplication.
[0116] To derive a non-separable transform that produces encoding gain for a specific directional feature, this transform can be learned. For example, a group of representative residual blocks corresponding to the directional feature of interest can be grouped, and the Karhunen-Loeve transform (KLT) can be computed based on the covariance matrix of this group of residual blocks. This process can be repeated for K different groups of residual blocks. Then, in this example, the final derived dimension is... The overall transformation kernel.
[0117] As discussed in this section, there are two problems with crossing inseparable transformations. First, the computational complexity is high. Because inseparable transformations are often learned, they are usually not factorable. The matrix implementation of crossing inseparable transformations in the example above results in a computational complexity of O(n log n) per sample. This involves multiple multiplication operations. The second problem is that the transform kernel occupies a significant amount of storage space in both encoder 101 and decoder 201. In the example above, a single kernel adaptable to K different directional features has... Each weight. This kernel can only be applied to kernels of size [size missing]. The residual blocks. In order to allow the non-separable transformation to be applied to multiple block sizes, a transformation kernel must be learned for each discrete block size.
[0118] In VVC, the LFNST tool was introduced and modified in many ways to address the problems that exist across the aforementioned inseparable transformations.
[0119] First, based on the initial revision, although the LFNST tool is applicable to various block sizes, only two LFNST cores are defined. For example, for a block size of... or For blocks with N≥4, apply a smaller LFNST kernel. For all larger block sizes (e.g., 8×8 or larger), apply a larger LFNST kernel.
[0120] Figure 16A illustrates some embodiments of VVC according to the present disclosure. and Figure 1600 shows a block-sized LFNST core. Figure 16B illustrates some embodiments of the present disclosure for use in VVC. and Figure 1601 shows a block-sized LFNST core.
[0121] Figures 16A and 16B show the sample locations on which the LFNST acts. For example, from the encoder's perspective, for or For a block size of 16A, the top 4x4 sample locations (indicated by the shaded area in Figure 16A) are transformed using a small LFNST. The remaining sample locations (indicated by the white area in Figure 16A) are ignored or "zeroed out." From the decoder's perspective, an inverse LFNST is applied to generate the top 4x4 samples, while the remaining samples are filled with zeros. A similar strategy is applied to larger block sizes, where the LFNST is applied to the three top 4x4 blocks of sample locations (indicated by the shaded area in Figure 16B). The remaining sample locations are then zeroed out.
[0122] Due to the use of a "zeroing" strategy, the size of LFNST is significantly reduced compared to a full-size transform applied to all sample locations. However, this method is inherently lossy and cannot recover values at sample locations ignored by LFNST. If the LFNST tool is applied directly to the residual samples, this loss would be too severe, rendering the LFNST tool ineffective. However, LFNST is called a secondary transform because it is applied after the separable DCT at the encoder has been performed, acting on the primary transform coefficients to generate secondary transform coefficients. In other words, the DCT can be considered the primary transform. According to embodiments of this disclosure, the leftmost sample location in the primary transform coefficient block corresponds to the horizontal low-frequency portion of the DCT, while the topmost sample location corresponds to the vertical low-frequency portion of the DCT. By prioritizing the transformation and reconstruction of the upper left sample location at decoder 201, LFNST is able to reconstruct low-frequency information from the original residuals. As previously mentioned, the transform produces coding gain due to its energy compression characteristics, and it has been well established that the variance (energy) of the image and video signal captured by the camera is mainly concentrated in the low-frequency DCT coefficients. Therefore, although “zeroing” prevents LFNST from reconstructing arbitrary residual blocks losslessly, in practice, this loss is minimal for most types of image and video signals.
[0123] The second modification is that, for both small and large LFNST kernels, the applied transform is not a cross-transform. From the encoder 101's perspective, the number of output (secondary transform) coefficients is less than the number of input (primary transform) coefficients. For example, the smaller LFNST kernel will... The main transform coefficients are used as input, but only 8 output second-order transform coefficients are generated. A larger LFNST kernel will... The master transform coefficients are taken as input, and eight quadratic transform coefficients are output. The use of a non-spanning transform introduces a further reconstruction loss. However, this loss can be balanced in a controlled manner with a reduction in implementation complexity. The spanning inseparable transform can be designed first using the KLT method described above. By following this method, the basis vectors of the transform correspond to eigenvectors of the covariance matrix computed from a set of representative residual blocks. These eigenvectors can be sorted by importance based on their corresponding eigenvalues, with the most important eigenvectors selected to construct the non-spanning inseparable transform. For example, the eight eigenvectors with the largest eigenvalues can be selected to form the non-spanning transform for a smaller LFNST kernel.
[0124] In summary, compared to crossing inseparable transforms, the two modifications described above significantly reduce the complexity of the LFNST kernel. For smaller blocks, using a smaller LFNST kernel reduces the potential complexity from that of each transform block. The number of multiplications (for N≥4) is reduced to This involves multiple multiplications. For larger blocks, using a larger LFNST kernel reduces the potential complexity from one-third of the transformation block's complexity. The number of multiplications (for N≥8) is reduced to Multiplication.
[0125] The LFNST kernel contains more than one transformation matrix. To achieve better coding gain on various image and video signals, multiple transformation matrices need to be learned. The number of different transformation matrices is the product of the 3rd and 4th dimensions of the LFNST kernel: the smaller LFNST kernel has a dimension of... The dimension of a larger LFNST kernel is Because the specific transformation matrix of the transform block is selected through a hybrid approach of explicit signaling and implicit selection, the LFNST kernel is represented using two additional dimensions.
[0126] Explicit signaling is performed by using a signaled LFNST index in the bitstream, which can take values of 0, 1, or 2. A value of 0 indicates that LFNST is not used for the transform block, while values of 1 or 2 indicate the selection in the third dimension of the LFNST core. The drawback of potential reconstruction loss due to zeroing and non-crossing simplification is mitigated by the explicit signaling mechanism. Although using LFNST can lead to excessive reconstruction loss in the transform block, the LFNST tool can be disabled by signaling that the LFNST index is 0.
[0127] Implicit selection is achieved by limiting LFNST to coding units that use intra-prediction. Intra-prediction generates a prediction block for a coding unit based on the top and left neighboring reference samples of the current block. The specific method for constructing the prediction block is signaled in the bitstream via the intra-prediction mode. Simple methods of intra-prediction include averaging the reference samples (“DC” mode) or constructing affine interpolation between some reference samples (“Plane” mode). However, most intra-prediction modes are reserved for signaling the intra-angular direction, where the prediction block is constructed by assuming that the reference sample values are copied along a specific direction. When using the intra-angular direction, it can be a strong cue for the directional characteristics of the residual block. The LFNST transform is implicitly selected by mapping the intra-prediction mode to one of four possible values of a “transform set index” used to index into the fourth dimension of the LFNST kernel. The mapping used in VVC is shown in Table 1.
[0128]
[0129] Table 1: Mapping from Intra-Prediction Modes to LFNST Transform Set Indices. Intra-prediction modes 0 and 1 correspond to the intra-prediction plane mode and the intra-DC prediction mode, respectively. These modes are treated as special cases by mapping to transform set index 0. Otherwise, the remaining intra-prediction modes correspond to the intra-angular direction 170°, partially shown in Figure 17. Intra-prediction mode 2 corresponds to intra-angular prediction in the diagonal direction starting from the lower left corner. Increasing the intra-prediction mode number corresponds to a clockwise rotation of the intra-prediction direction, where intra-prediction mode 34 corresponds to intra-angular prediction in the diagonal direction starting from the upper left corner, and intra-prediction mode 66 corresponds to intra-angular prediction in the diagonal direction starting from the upper right corner.
[0130] For intra-prediction modes greater than 34 (corresponding to intra-angle prediction directions that rotate clockwise from the top-left diagonal direction), the selected LFNST transform matrix is applied to the major transform coefficients in a transposed manner. In one implementation, this can be performed by scanning the major transform coefficients in the transposed direction before applying the LFNST transform. For example, from the encoder 101's perspective, if the current block is predicted using intra-prediction mode 2, the major transform coefficients can be rearranged from their two-dimensional pattern in the block into a one-dimensional vector by a row-major scan before applying the selected LFNST transform matrix T. Then, for this example, if the current block is instead predicted using intra-prediction mode 66, and the same signaled LFNST index is used, the major transform coefficients will instead be rearranged into a one-dimensional vector by a column-major scan before applying the same LFNST transform matrix T. In another implementation, the same current block with intra-prediction mode 66 can be equivalently transformed by still performing a row-major scan on the major transform coefficients, but instead rearranging the rows of the transform matrix T.
[0131] More generally, the application of the LFNST transform matrix can be described as follows. Let the principal transform coefficients located in the y-th row and x-th column be denoted as... And let the dimension of the LFNST transformation matrix T be... Where A is the number of second-order transform coefficients and B is the number of principal transform coefficients not set to zero. Then, for intra-frame prediction modes <= 34, the principal transform coefficients p are used... x,y The arbitrary scanning order for constructing a one-dimensional vector P can be defined by the following equation (8).
[0132] (8).
[0133] For intra-prediction modes greater than 34, P is instead constructed by the transpose scan order defined by equation (9) as shown below.
[0134] (9).
[0135] The forward LFNST transformation can be understood as matrix multiplication. ,in It is a one-dimensional vector of the coefficients of the quadratic transformation. In practice, the transformation is implemented as... Because all multiplication is implemented using integer operations, and The normalization operation required to represent the integerized LFNST approximation of the ideal transform in floating-point form. The quadratic transform coefficients are written back to the transform block in a forward diagonal scan order. Following the same notation as described above, and the convention that the (0,0) position in conventional DCT corresponds to "low frequency" or "DC", the scan order s can be defined according to equation (10) shown below.
[0136] (10).
[0137] From the perspective of decoder 201, the inverse LFNST transform can be represented as matrix multiplication. Or, in other words, the inverse transformation is performed by transposing matrix T.
[0138] For intra-prediction modes with a value greater than 34, transposing the master transform coefficients allows the sharing of the same LFNST transform matrix for symmetric intra-angle prediction directions.
[0139] In exploratory activities following VVC, an extension of LFNST was proposed and integrated into ECM. The LFNST tool in ECM relaxes some of the complexity reduction requirements imposed on the original LFNST tool in VVC to improve coding gain.
[0140] The ECM has three LFNST cores. Similar to the LFNST tool in VVC, in most cases, the majority of the transform block is set to zero, as shown in Figures 18A-18C. For example, Figure 18A illustrates an ECM used in some embodiments of this disclosure for... and Figure 1800 shows a block-sized LFNST core. Figure 18B illustrates an ECM used in some embodiments of this disclosure. and Figure 1801 shows a block-sized LFNST core. Figure 18C shows Figure 1803 of an ECM for a 16x16 block-sized LFNST core according to some embodiments of this disclosure.
[0141] In Figures 18A-18C, the shaded areas indicate the positions of the principal transform coefficients on which the LFNST in the ECM acts, while the white areas indicate which transform coefficient positions are set to zero. For a size of or (Where N≥4) blocks are in the top left corner. A small LFNST kernel is used on the main transform coefficients. For a size of or (Where N≥8) blocks, located in the top left of the four main transform coefficients. Use a medium-sized LFNST core on the block. For Or a larger block, in the upper left of the 6th of the main transform coefficients. A larger LFNST core is used on the block.
[0142] The sizes of LFNST cores in ECM are as follows: The size of a small LFNST core is... The size of the medium-sized LFNST core is The size of a large LFNST core is Compared to the LFNST tool in VVC, the range of LFNST indices notified by the signal increases from 2 to 3, and the number of LFNST transform sets increases from 4 to 35. This means that each of the three indices has 35 LFNST transform matrices. The mapping from intra-prediction modes to LFNST transform set indices is shown in Table 2. Similar to the LFNST tool in VVC, when the intra-prediction mode is greater than 34, the master transform coefficients are transposed.
[0143]
[0144] Table 2: The mapping from intra-prediction modes to LFNST transform set indices in ECM can be evaluated in three ways to assess the complexity burden of the LFNST tool. First, there is an additional storage burden imposed on decoder 201, as it must store the LFNST kernel. Second, the worst-case number of multiplications that decoder 201 must perform per sample if the LFNST tool is used. Third, the additional number of multiplications per sample used by encoder 101 if a full search is performed on the LFNST tool. By all three metrics, the extended LFNST proposed in ECM is more complex than the LFNST in VVC. However, in terms of the total number of multiplications per sample, the worst-case decoder complexity can still be less than the worst-case decoder complexity of other transform options.
[0145] For can be applied independently The matrix multiplication implementation of the size-transformed DCT, with the number of multiplications per sample being... Therefore, the worst-case complexity occurs when the value of (M+N) is maximized. In practice, the complexity can be reduced using alternative implementations of the DCT, such as butterfly factorization, but it is still convenient to evaluate the complexity of the matrix multiplication implementation. The ECM extends separable DCT so that the largest transformation is a 128-point DCT. Then, the worst-case complexity of separable DCT might be per sample... This is the first multiplication operation.
[0146] The worst-case decoder complexity of LFNST in ECM can be evaluated by considering various different block sizes. For a fair comparison, the evaluation includes the cost of performing the master transform. The block, the main transform, includes each sample's... This is a multiplication operation. LFNST includes... Matrix multiplication involves performing 16 multiplications for each sample. Therefore, The total cost of the LFNST for a block is 24 multiplications per sample.
[0147] for A naive implementation of the DCT principal transform typically requires eight blocks along the short dimension. Transformation and four along the long dimension The transformation results in a total of [number] transformations required for each sample. This is the second multiplication operation. However, because LFNST only reconstructs the upper left of the principal transform coefficient positions... The non-zero coefficient values in the block, so the optimized decoder will be able to execute only 4 along the short dimension. Transform, then perform 4 operations along the long dimension. Transformation to take advantage of this advantage results in a total of [number] tests required per sample. This involves four multiplication operations. Since the order of separable transformations is usually fixed, in the worst-case scenario, decoder 201 can first perform four multiplication operations along the long dimension. Transformation. Then, decoder 201 can perform 8 transformations along the short dimension. Transformation. This requires each sample to undergo... This is the second multiplication operation. LFNST is still... Matrix multiplication, whose cost is amortized over larger blocks, results in 8 multiplication operations per sample. Then, The worst-case cost of LFNST for a block is 16 multiplications per sample. The same principle generally applies to or Block size. Therefore, or The number of multiplication operations required for each sample in a block is always less than or equal to 1. The number of multiplication operations required for each sample in the block.
[0148] for The block, the main transform, includes each sample's... This is the second multiplication. LFNST is... The matrix is composed of 32 multiplications per sample. Then, The total cost of the LFNST for a block is 48 multiplications per sample.
[0149] for This block again assumes that decoder 201 utilizes the zeroing property of LFNST reconstruction. Only the upper left of the main transform coefficient position... The block is non-zero. Decoder 201 can execute 8 steps only along the short dimension. This is utilized through transformation. Then, decoder 201 can perform eight transformations along the long dimension. Transformation. This results in a total of [number] samples required. This is the second multiplication. Alternatively, decoder 201 can first perform 8 multiplications along the long dimension. Transformation. Then, decoder 201 can perform 16 transformations along the short dimension. Transformation. This may require each sample to be transformed. LFNST adds an extra multiplication step for each sample. This results in a total of 32 multiplications per sample in the worst-case scenario. As mentioned earlier, or The number of multiplications for each sample in the block is always less than or equal to The number of multiplications for each sample in the block.
[0150] for The zero-set characteristic of LFNST reconstruction means that only six 4x4 blocks of the principal transform coefficients in the pattern (as shown in Figure 18C) have non-zero values. For simplicity, let's assume a more relaxed pattern where the top-left 12x12 block of the principal transform location can have non-zero values. First, the decoder can perform only 12 operations in one dimension. This is utilized through transformation. Then, decoder 201 can perform 16 transformations in the second dimension. Transformation, which includes each sample Multiplication. LFNST includes each sample. This results in a total complexity of 33 multiplications for each sample.
[0151] for Blocks (where M, N≥16), decoder 201 can first execute 12 in one dimension. Transformation. Then, decoder 201 can perform M transformations in the second dimension. The transformation requires each sample to be processed. This involves multiple multiplications to perform a separable DCT. Then, for the minimum N=16, the worst-case complexity occurs, which is 21 multiplications per sample. The block complexity is equal. LFNST adds an extra layer of complexity for each sample. This multiplication is always less than or equal to... The number of multiplications for each sample in the block. Therefore, for larger blocks... Block size; the overall complexity of LFNST in ECM is always equal to or less than [the block size]. The number of multiplications for each sample in the block.
[0152] After a detailed evaluation of the decoder complexity of LFNST in ECM at different block sizes, the worst-case complexity is shown to be 48 multiplications per sample (occurring in...). (In the case of blocks). This worst-case complexity includes the cost of implementing the separable DCT using matrix multiplication, but it is significantly smaller than the worst-case complexity of executing the separable DCT alone (which is evaluated as 256 multiplications per sample) due to possible optimizations from zeroing out LFNST. If a more practical implementation of the separable DCT with butterfly factorization is assumed, the worst-case complexity of LFNST still appears in In the block, the cost of 32 multiplications for each sample from LFNST is added to the cost of the butterfly DCT. In this case, it is applied separably / separately. Compared to the cost of large and small butterfly DCTs, LFNST is probably the worst-case scenario.
[0153] As seen above, the complexity can be significantly reduced by using an inseparable quadratic transform with zeroing over the selected master transform coefficient region. However, further encoding can be performed using NSPT. Preliminary studies of the inseparable master transform show that significant gains can be achieved (a 3.43% reduction in average rate according to the Bjontegaard metric), despite the complexity of the implemented transform and the fact that the kernel weights were obtained through overfitting to the test dataset.
[0154] This paper presents a practical implementation of NSPT. For example, NSPT can be applied to only a small portion of the block size: , , and For these block sizes, NSPT replaces both the main transform and LFNST. Like LFNST, the NSPT kernel also needs to be trained, where the selection of the appropriate matrix for a specific block is guided by both signaling indexing and implicit selection via intra-frame prediction modes. Four NSPT kernels are proposed. Block, using size The small NSPT core. For and Block, using size Medium NSPT cores. For Block, using size Large NSPT core.
[0155] According to this disclosure, the zeroing method can be defined as a reduction in the input dimension of the transform kernel, which corresponds to a first dimension in the representation of transform kernel dimensions in this disclosure. A reduction in the input dimension of the forward transform is equivalent to a reduction in the support range of the transform. For example, zeroing the LFNST corresponds to a reduction in the number of DCT principal transform coefficients on which the forward LFNST acts to produce the quadratic transform coefficients. In the proposed NSPT, the transform acts directly on the residual coefficients, so a reduction in the first dimension of the NSPT kernel includes a reduction in the number of residual coefficients on which the forward NSPT acts to produce the principal transform coefficients. In the known schemes described above, the size of the first dimension of each NSPT kernel is always equal to the number of samples in the block, therefore zeroing as defined in this disclosure is not used. However, in this known proposal, zeroing is alternatively defined as a reduction in the output dimension of the transform kernel, which is a second dimension in the representation of transform kernel dimensions in this disclosure. This definition is not ambiguous in this proposal because reduction is never performed on the input side of the NSPT kernel, and an example of the more generally used term "zeroing" is given. However, for consistency and clarity in this disclosure, "zeroing" is defined as describing a reduction in the input dimension of the transform kernel, while a reduction in the output dimension is labeled as a non-crossing or lossy transform. For medium and large NSPT kernels, the second dimension is smaller than the first dimension, which means that the NSPT in these cases is a lossy transform.
[0156] Similar to LFNST, the NSPT index is signaled in the bitstream. This index can take values of 0, 1, 2, or 3, where 0 indicates that the NSPT is not used for the transform block, and values 1-3 indicate selection along the third dimension within the corresponding NSPT kernel. The selection along the fourth dimension of the NSPT kernel is determined by mapping from the intra-prediction modes as shown in Table 3, in the same manner as extended LFNST in ECM. As with LFNST, the transform input is transposed when the intra-prediction mode is greater than 34 (meaning the intra-angle direction is clockwise from the diagonal upper-left direction). However, for NSPT, the input consists of residual coefficients instead of the master transform coefficients.
[0157]
[0158] Table 3: Mapping from Intra-Prediction Mode to NSPT Transform Set Index in ECM The shape of the residual block, and when the intra-frame prediction mode is less than or equal to 34, let the residual sample located at the y-th row and x-th column be represented as And let it be used for The LFNST transformation matrix T selected in the NSPT kernel of the shape block has dimension. , where A is the number of principal transform coefficients, and It is the number of residual samples in the block. Then, the arbitrary scanning order of constructing a one-dimensional vector R through the residual samples can be defined according to equation (11) shown below.
[0159] (11).
[0160] For intra-prediction modes greater than 34, R is alternatively constructed from the transposed scan order according to equation (12).
[0161] (12).
[0162] Additionally, when the intra-prediction mode is greater than 34, from the used The LFNST transformation matrix T is chosen from the NSPT kernel of the shape block. For square block shapes, this is the same kernel, therefore T is the same transformation matrix. However, in or For block sizes, the transformation matrix is selected from different NSPT kernels.
[0163] The forward NSPT transform can be implemented as , where P is a one-dimensional vector of NSPT transform coefficients, and This represents the normalization operation necessary to approximate the integerized NSPT to an ideal transform expressed in floating-point form. The transform coefficients are written back to the transform block in forward diagonal scan order. Following the same notation as described above, and the convention that the (0,0) position in conventional DCT corresponds to "low frequency" or "DC", the scan order is described according to equation (13): (13).
[0164] From the perspective of decoder 201, the inverse NSPT transform can be represented as matrix multiplication. Or, in other words, the inverse transformation is achieved by transposing the matrix T.
[0165] Figure 19 shows an example visualization of inter-frame prediction mode 1900 according to some aspects of this disclosure.
[0166] For inter-frame coded CUs (also referred to as inter-frame CUs in this disclosure), reconstructed samples from a time reference frame are used to predict the current block, and the inter-frame prediction mode is signaled only once for the entire CU. Inter-image prediction utilizes the temporal correlation between images to derive motion-compensated prediction (MCP) for image sample blocks.
[0167] For this block-based MCP, the video image is divided into rectangular blocks. Assuming uniform motion within a block, for each block, a corresponding block in the previously decoded image can be found as a predictor. Figure 19 illustrates the general concept of MCP based on a translational motion model. Using the translational motion model, the position of a block in the previously decoded image is determined by the motion vector (…). ) indicates that among them and These represent the horizontal and vertical displacements relative to the current block's position in the horizontal and vertical directions, respectively. Motion vector ( It can have subpixel sampling precision to capture the motion of the underlying object more accurately. When the corresponding motion vector has fractional sampling precision, interpolation is applied to the reference image to derive the prediction signal. The previously decoded image is called the reference image and is indexed by a list of reference images. Indicators. These translational motion model parameters (e.g., motion vectors and reference indices) are further referred to as motion data. Modern video coding standards allow two types of image / frame prediction: one-way prediction and two-way prediction. In the case of two-way prediction, two sets of motion data ( )and( This is used to generate two MCPs (potentially from different images), which are then combined to obtain the final MCP. Furthermore, multiple hypotheses (more than two) can be employed to form the final inter-frame prediction. The same or different weights can be applied to each MCP. Reference images that can be used in bidirectional prediction are stored in two separate lists, namely list 0 and list 1. Motion data is derived at encoder 101 using a motion estimation process. Motion estimation is not specified in the video standard, therefore different encoders can utilize different complexity-quality tradeoffs in their implementations.
[0168] Motion data for a block can be correlated with neighboring blocks. To utilize this correlation, motion data is not directly encoded into the bitstream, but rather predictively encoded based on neighboring motion data. This predictive encoding of motion vectors can be performed using Advanced Motion Vector Prediction (AMVP), where the best predictor for each motion block is signaled to the decoder. Furthermore, inter-frame prediction blocks are merged to derive all motion data for a block based on neighboring blocks, including both spatial and temporal blocks. Similar to intra-frame prediction, the residual between the original pixels and the inter-frame prediction can be further transformed before being encoded into the bitstream.
[0169] Using existing techniques, reference blocks are directly copied from the reconstructed region of a previous reference frame. Then, when an inter-frame mode is selected, the weighted sum of these reference blocks is used as the prediction block for the current block. Because it does not consider any spatial information between adjacent pixels, the prediction accuracy may be excessively limited.
[0170] To overcome these and other challenges, this disclosure provides a filtered inter-frame prediction mode. In some implementations, the inter-frame motion vector (Finter) is used to predict the prediction mode. The reconstructed pixels within the reference block pointed to by the reference block are further filtered by an online learning filter. The resulting filtered block (single prediction) or the sum of multiple weighted filtered blocks (two or multiple hypotheses), instead of the reference block, is used as the predictor for the current block. In some implementations, multiple reference blocks can be determined from multiple motion data signaling to form the final predictor for the current CU. Here, the final predictor is generated by a fusion combination of reference blocks further filtered by the learning filter. More details of an exemplary Finter prediction pattern are provided below with reference to Figures 20 and 21.
[0171] Figure 20 illustrates a visualization example of the spatial support domain of a learned filter 2000 (hereinafter referred to as "filter 2000") according to some embodiments of the present disclosure, wherein the reference block 2106 is transmitted via motion vectors ( (Identification mark). See Figures 20 and 21 for further explanation.
[0172] Referring to Figure 20, the inter-frame prediction module 408 can apply filter 2000 to FInter prediction as a shift-invariant weighted sum on the support region moved on the reconstructed block. Let R and F be the reference block and the filtered prediction block, respectively, and and Let x and y represent the sample and filtered sample in the x-th column and y-th row of R and F, respectively. The filter can be applied according to equation (14).
[0173]
[0174] in, and is the learned filter weights adaptively adjusted using the pixels of the reconstructed block, B is a bias term that can be set to an intermediate brightness value of the input video depth (e.g., 512 for 10-bit video), and S is a finite region of the filter support domain.
[0175] Filter 2000 may include a bias weight and five spatial weights located at the central "C" position and four adjacent "W", "N", "E", and "S" positions. In this example, the support region... , where (0,0), (-1,0), (0,1), (1,0) and (0,-1) represent the positions of C, W, N, E and S, respectively.
[0176] In some implementations, the support region can have different shapes and sizes. Additionally and / or alternatively, the filter can be enhanced by nonlinear terms. For example, the support region can be used. The squared pixel values at some locations. In one example, the support region for the squared terms is... - For example, the value at the center position is squared. A filter including a squared nonlinear term can be applied according to equation (15).
[0177] (15).
[0178] Inter-frame prediction module 408 can learn the filter weights by minimizing the error between the template of the reference block and the template of the current block. Figure 21 shows an example with an L-shaped template (e.g., the current template 2104 and / or the reference template 2108). However, the template can take other shapes, such as an upper rectangle or a left rectangle. Let RT and CT be the reference template 2108 of reference block 2106 and the current template 2104 of current block 2102, respectively; and let... and Let x and y represent the samples at the x-th column and y-th row of RT and CT, respectively. Then, according to equation (16), filter weights are selected to minimize the mean square error (MSE) between the filtered RT and the template of the current block.
[0179]
[0180] This is the minimization of a quadratic cost function, which can be represented as solving a set of linear equations, and equivalently, by finding the inverse of the matrix representing these linear equations. For regular structures generated by shift-invariant filters, the solution can be computed efficiently. For example, the matrix can first be decomposed using LDL decomposition, and then the inverses of its subcomponents can be efficiently obtained.
[0181] When solving for the filter coefficients, and when applying the learned filter, the filter's support domain can extend beyond the region of available samples, as shown in the blue area in Figure 21. In this case, boundary padding is used to generate values for these regions.
[0182] For bidirectional and multi-hypothesis prediction, the inter-frame prediction module 408 can learn multiple filters separately using templates derived from multiple corresponding motion data and multiple reference blocks indicated by the current block. The inter-frame prediction module 408 can first filter the reference blocks individually using the corresponding learned filters. Then, the inter-frame prediction module 408 can fuse the multiple filtered reference blocks together to form the final prediction. The entire bidirectional and multi-hypothesis prediction process can be similar to inter-frame prediction, except that it suggests using filtered reference blocks instead of the reference blocks themselves to generate the final prediction.
[0183] If FInter is enabled, for example, at the Sequence Parameter Set (SPS), Picture Header (PH), Picture Parameter Set (PPS), or Slice Header, a high-level flag can be signaled. If FInter prediction mode is enabled for the current video sequence and the current CU is predicted via inter-frame prediction, an additional flag is signaled to indicate whether raw inter-frame prediction or FInter is used.
[0184] In some implementations, FInter can completely replace the original inter-frame prediction mode. That is, the FInter advanced flag will instead signal whether the prediction mode is used for all inter-frame prediction blocks.
[0185] In some other implementations, the inter-frame prediction module 408 can determine multiple reference blocks using multiple motion data signaling to form the final predictor for the current CU. Let the reference block be... For satisfying For certain fusion weights, the fusion combination of the reference block is as follows: Then, in this embodiment, the inter-frame prediction module 408 can generate the final predictor by further filtering the fused combination using a learning filter according to equation (17).
[0186]
[0187] According to equation (18), the inter-frame prediction module 408 can minimize the template of the reference block. The filter weights are learned by weighting the error between each element and the template of the current block.
[0188] (18).
[0189] Referring again to Figure 4, additionally and / or alternatively, the inter-frame prediction module 408 and the intra-frame prediction module 410 can be configured to generate prediction blocks based on information related to the generation of prediction blocks provided by the decoding module 402 and information about previously decoded blocks or images provided by the buffer module 414. As described above, when intra-frame prediction is performed in the same manner as the encoder operation, if the size of the prediction unit and the size of the transform unit are the same, intra-frame prediction can be performed on the prediction unit based on the pixels to the left, the upper left, and the top of the prediction unit. However, when performing intra-frame prediction, if the size of the prediction unit and the size of the transform unit are different, intra-frame prediction can be performed based on the transform unit using reference pixels.
[0190] For example, inter-frame prediction module 408 may be configured to receive from the encoder a bitstream containing an indication of a reference frame, the current frame, and weighting factors associated with a Multimedia Home Platform (MHP) process. Inter-frame prediction module 408 may be configured to perform an MHP process on a CU located in the current frame based on a search block (e.g., the reference frame and / or a reference template) in the reference frame. In some embodiments, to perform the MHP process, inter-frame prediction module 408 may be configured to perform template matching on the CU located in the current frame based on the search block and weighting factors in the reference frame to obtain motion information. In some embodiments, to perform the MHP process, inter-frame prediction module 408 may be configured to identify a weighting factor index associated with the weighting factors based on template matching. Inter-frame prediction module 408 may be configured to identify the weighting factor symbol of the weighting factors based on indications included in the bitstream. Inter-frame prediction module 408 may be configured to perform an inter-frame prediction process to decode the bitstream based on the current frame, the reference frame, the weighting factor index, and the weighting factor symbol of the weighting factors.
[0191] The reconstructed block or reconstructed image, composed of the outputs of the inverse transform module 406 and the prediction modules 408 or 410, can be provided to the filter module 412. The filter module 412 may include a deblocking filter, an offset correction module, and an ALF. The buffer module 414 can store the reconstructed image or block and use it as a reference image or reference block for the inter-frame prediction module 408, and can output the reconstructed image.
[0192] Within the scope of this disclosure, the encoding module 320 and the decoding module 402 can be configured to employ a quantization-level binarization scheme with Rice parameters suitable for encoding video images to improve encoding efficiency.
[0193] Figures 22A and 22B illustrate flowcharts of an exemplary decoding method 2200 according to some embodiments of the present disclosure. Method 2200 can be performed by a system, such as decoding system 200, decoder 201, or inter-frame prediction module 408, to name just a few. Method 2200 may include operations 2202-2220 as described below. It should be understood that some steps may be optional, and some steps may be performed simultaneously, or in a different order than that shown in Figures 22A and 22B.
[0194] Referring to Figure 22A, at 2202, the system can acquire multiple reference blocks. In some embodiments, acquiring multiple reference blocks by the processor may include parsing a bitstream to obtain multiple motion data. In some embodiments, acquiring multiple reference blocks by the processor may include generating multiple reference blocks based on multiple motion data. For example, referring to Figure 4, the inter-frame prediction module 408 can acquire multiple reference blocks. In some examples, the inter-frame prediction module 408 may acquire multiple reference blocks based on motion data.
[0195] At position 2204, the system can parse the bitstream to obtain a first flag. In some implementations, the first flag can be signaled at the SPS, PH, PPS, or SH level. For example, referring to Figure 4, the inter-frame prediction module 408 can parse the bitstream to obtain the first flag. The first flag indicates whether FInter prediction mode is enabled for the current block.
[0196] At 2206, in response to the first flag indicating that FInter prediction mode is enabled for the current block, the system can parse the bitstream to obtain the second flag. For example, referring to Figure 4, when the first flag indicates that FInter prediction mode is enabled, the inter-frame prediction module 408 can parse the bitstream to obtain the second flag. The second flag can indicate whether regular inter-frame prediction or FInter prediction is selected for the current block.
[0197] At 2208, the system can determine whether FInter prediction mode has been selected for the current block based on a second flag. For example, referring to FIG4, the inter-frame prediction module 408 can determine whether FInter prediction mode has been selected for the current block based on the second flag. For example, a second flag with a first value can indicate that FInter prediction has been selected, while a second flag with a second value can indicate that regular inter-frame prediction has been selected.
[0198] At 2210, the system can parse the bitstream to obtain a flag. For example, referring to Figure 4, the inter-frame prediction module 408 can parse the bitstream to obtain a single flag when the FInter prediction mode replaces the regular inter-frame prediction mode.
[0199] At 2212, the system can determine whether to select the FInter prediction mode for all inter-frame prediction blocks based on this flag. For example, referring to FIG4, when the single flag has a first value, the inter-frame prediction module 408 can determine to select the FInter prediction mode for all inter-frame prediction blocks, and when the single flag has a second value, the inter-frame prediction module 408 determines not to select the FInter prediction mode for all inter-frame prediction blocks.
[0200] At 2214, the system can select a set of filter weights for the FInter prediction filter, which minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. For example, referring to FIG4, the inter-frame prediction module 408 can, according to equation (18), minimize the MSE between the reference template and the current template associated with the current block. The filter weights are learned / selected by weighting the error between each frame prediction module and the template of the current block. In some implementations, the inter-frame prediction module 408 may select the filter weights according to equation (16) to minimize the mean square error (MSE) between the filtered RT and the template of the current block.
[0201] Referring to Figure 22B, at 2216, the system can generate an FInter prediction filter based on this set of filter weights. In some implementations, the FInter prediction for the current block can be further generated based on the FInter prediction filter. For example, referring to Figure 4, the inter-frame prediction module 408 can generate an FInter prediction filter based on the selected filter weights.
[0202] At 2218, in response to determining that an FInter prediction mode has been selected for the current block, the system generates an FInter prediction for the current block based on multiple reference blocks. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include filtering each of the multiple reference blocks separately using corresponding filters to obtain multiple filtered reference blocks. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include fusing the multiple filtered reference blocks to obtain a filtered and fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include generating an FInter prediction for the current block based on the filtered and fused reference block. For example, referring to FIG4, for bidirectional prediction and multi-hypothesis prediction, the inter-frame prediction module 408 can learn multiple filters individually using templates around multiple reference blocks indicated by the corresponding multiple motion data and multiple reference blocks. The inter-frame prediction module 408 may first filter the reference blocks individually using the corresponding learned filters. Then, the inter-frame prediction module 408 can fuse multiple filtered reference blocks together to form a final prediction. The entire bidirectional prediction and multiple hypothesis process can be similar to inter-frame prediction, except that it is suggested to use filtered reference blocks instead of the reference blocks themselves to generate the final prediction. In some other embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include fusing multiple reference blocks to obtain a fused reference block. In some other embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include filtering the fused reference blocks to obtain a fused and filtered reference block. In some other embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include generating an FInter prediction for the current block based on the fused and filtered reference block. For example, referring to FIG4, the inter-frame prediction module 408 can determine multiple reference blocks through multiple motion data signaling to form a final predictor for the current CU. Let the reference block be... For satisfying For certain fusion weights, the fusion combination of the reference block is as follows: Then, in this embodiment, the inter-frame prediction module 408 can generate the final predictor by further filtering the fused combination with a learning filter according to equation (17).
[0203] At 2220, in response to determining that the FInter prediction mode is not selected for the current block, the system generates a regular inter-frame prediction for the current block based on multiple reference blocks. For example, referring to Figure 4, when the FInter prediction mode is not selected for the current block, the inter-frame prediction module 408 can generate a regular inter-frame prediction for the current block.
[0204] Figures 23A and 23B illustrate flowcharts of an exemplary encoding method 2300 according to some embodiments of the present disclosure. Method 2300 may be performed by a system, such as encoding system 100, encoder 101, or inter-frame prediction module 304, to name a few. Method 2300 may include operations 2302-2320, as described below. It should be understood that some steps may be optional, and some steps may be performed simultaneously, or in a different order than shown in Figures 23A and 23B.
[0205] Referring to Figure 23A, at 2302, the system can acquire multiple reference blocks. In some embodiments, acquiring multiple reference blocks by the processor may include encoding multiple motion data. In some embodiments, acquiring multiple reference blocks by the processor may include generating multiple reference blocks based on multiple motion data. For example, referring to Figure 3, the inter-frame prediction module 304 can acquire multiple reference blocks. In some examples, the inter-frame prediction module 304 may obtain reference blocks based on motion data.
[0206] At 2304, the system can encode the first flag. In some implementations, the first flag can be signaled at the SPS, PH, PPS, or SH level. For example, referring to FIG3, the inter-frame prediction module 304 can encode the first flag. The first flag can indicate whether FInter prediction mode is enabled for the current block.
[0207] At 2306, in response to the first flag indicating that FInter prediction mode is enabled for the current block, the system can encode the second flag. For example, referring to FIG3, the inter-frame prediction module 304 can encode the second flag when the first flag indicates that FInter prediction mode is enabled. The second flag can indicate whether regular inter-frame prediction or FInter prediction is selected for the current block.
[0208] At 2308, the system can determine whether FInter prediction mode has been selected for the current block based on a second flag. For example, referring to FIG3, the inter-frame prediction module 304 can determine whether FInter prediction mode has been selected for the current block based on the second flag. For example, a second flag with a first value can indicate that FInter prediction has been selected, while a second flag with a second value can indicate that regular inter-frame prediction has been selected.
[0209] At position 2310, the system can encode a flag. For example, referring to Figure 3, the inter-frame prediction module 304 can encode a single flag when the FInter prediction mode replaces the regular inter-frame prediction mode.
[0210] At 2312, the system can determine whether the FInter prediction mode has been selected for all inter-frame prediction blocks based on this flag. For example, referring to FIG3, the inter-frame prediction module 304 can determine that the FInter prediction mode is selected for all inter-frame prediction blocks when a single flag has a first value, and determine that the FInter prediction mode is not selected for all inter-frame prediction blocks when a single flag has a second value.
[0211] At 2314, the system can select a set of filter weights for the FInter prediction filter that minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. For example, referring to FIG3, the inter-frame prediction module 304 can, according to equation (18), minimize the MSE between the reference template and the current template associated with the current block. The filter weights are learned / selected by weighting the error between each frame prediction module and the template of the current block. In some implementations, the inter-frame prediction module 304 may select the filter weights according to equation (16) to minimize the mean square error (MSE) between the filtered RT and the template of the current block.
[0212] Referring to Figure 23B, at 2316, the system can generate an FInter prediction filter based on this set of filter weights. In some implementations, the FInter prediction for the current block can be further generated based on the FInter prediction filter. For example, referring to Figure 3, the inter-frame prediction module 304 can generate an FInter prediction filter based on the selected filter weights.
[0213] At 2318, in response to determining that an FInter prediction mode has been selected for the current block, the system generates an FInter prediction for the current block based on multiple reference blocks. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include filtering each of the multiple reference blocks separately using corresponding filters to obtain multiple filtered reference blocks. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include fusing the multiple filtered reference blocks to obtain a filtered and fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include generating an FInter prediction for the current block based on the filtered and fused reference block. For example, referring to FIG3, for bidirectional prediction and multi-hypothesis prediction, the inter-frame prediction module 304 can learn multiple filters individually using templates around multiple reference blocks indicated by the corresponding multiple motion data and the multiple reference blocks. The inter-frame prediction module 304 may first filter the reference blocks individually using the corresponding learned filters. Then, the inter-frame prediction module 304 can fuse multiple filtered reference blocks together to form a final prediction. The entire bidirectional prediction and multiple hypothesis process can be similar to inter-frame prediction, except that it is suggested to use filtered reference blocks instead of the reference blocks themselves to generate the final prediction. In some other embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include fusing multiple reference blocks to obtain a fused reference block. In some other embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include filtering the fused reference blocks to obtain a fused and filtered reference block. In some other embodiments, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include generating an FInter prediction for the current block based on a fused and filtered reference block. For example, referring to FIG3, the inter-frame prediction module 304 can determine multiple reference blocks through multiple motion data signaling to form a final predictor for the current CU. Let the reference block be... For satisfying For certain fusion weights, the fusion combination of the reference block is as follows: Then, in this embodiment, the inter-frame prediction module 304 can generate a final predictor by further filtering the fused combination with a learning filter according to equation (17).
[0214] At 2320, in response to determining that the FInter prediction mode is not selected for the current block, the system generates a regular inter-frame prediction for the current block based on multiple reference blocks. For example, referring to Figure 3, when the FInter prediction mode is not selected for the current block, the inter-frame prediction module 304 can generate a regular inter-frame prediction for the current block.
[0215] In all aspects of this disclosure, the functions described herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. Storage media can be any available medium that can be accessed by a processor (such as processor 102 in Figures 1 and 2). By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, HDD (such as disk storage devices or other magnetic storage devices), flash drives, SSDs, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system (such as a mobile device or computer). As used herein, disks and optical discs include CDs, laser discs, optical discs, digital video discs (DVDs), and floppy disks, wherein disks typically reproduce data magnetically, while optical discs use lasers to reproduce data optically. Combinations of the above-described storage media should also be included within the scope of computer-readable media.
[0216] According to one aspect of this disclosure, a decoding method is provided. The method may include acquiring a plurality of reference blocks by a processor. The method may also include: in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on the plurality of reference blocks.
[0217] In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include filtering each of the multiple reference blocks using corresponding filters to obtain multiple filtered reference blocks. In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include fusing the multiple filtered reference blocks to obtain a filtered and fused reference block. In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on the filtered and fused reference block may include generating an FInter prediction for the current block by the processor based on the filtered and fused reference block.
[0218] In some implementations, the method may include parsing a bitstream by a processor to obtain a first flag. In some implementations, the method may include: in response to the first flag indicating that FInter prediction mode is enabled for the current block, parsing a bitstream by a processor to obtain a second flag. In some implementations, the method may include the processor determining, based on the second flag, whether FInter prediction mode has been selected for the current block.
[0219] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.
[0220] In some implementations, the method may include: in response to determining that an FInter prediction mode has not been selected for the current block, the processor generates a regular inter-frame prediction for the current block based on multiple reference blocks.
[0221] In some implementations, the method may include having a processor parse the bitstream to obtain a flag. In some implementations, the method may include having the processor determine, based on the flag, whether the FInter prediction mode has been selected for all inter-frame prediction blocks.
[0222] In some implementations, acquiring multiple reference blocks by the processor may include parsing a bitstream to obtain multiple motion data. In some implementations, acquiring multiple reference blocks by the processor may include generating multiple reference blocks based on the multiple motion data.
[0223] In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include fusing the multiple reference blocks to obtain a fused reference block. In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include filtering the fused reference block to obtain a fused and filtered reference block. In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include generating an FInter prediction for the current block by the processor based on the fused and filtered reference block.
[0224] In some implementations, the method may include a set of filter weights selected by a processor for an FInter prediction filter, the set of filter weights minimizing the MSE between a reference template associated with a reference block and a current template associated with the current block. In some implementations, the method may include the processor generating the FInter prediction filter based on the set of filter weights. In some implementations, an FInter prediction for the current block may be further generated based on the FInter prediction filter.
[0225] According to another aspect of this disclosure, a decoder is provided. The decoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to fetch a plurality of reference blocks. The memory stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0226] In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to individually filter each of the plurality of reference blocks using a corresponding filter to obtain a plurality of filtered reference blocks. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to fuse the plurality of filtered reference blocks to obtain a filtered and fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to generate an FInter prediction for the current block based on the filtered and fused reference block.
[0227] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to parse the bitstream to obtain a first flag. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to parse the bitstream to obtain a second flag in response to a first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to determine, based on the second flag, whether FInter prediction mode has been selected for the current block.
[0228] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.
[0229] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to generate a regular inter-frame prediction for the current block based on multiple reference blocks in response to determining that an FInter prediction mode has not been selected for the current block.
[0230] In some implementations, the memory stores instructions that, when executed by the processor, enable the processor to parse the bitstream to obtain a flag. In some implementations, the memory stores instructions that, when executed by the processor, enable the processor to determine, based on the flag, whether the FInter prediction mode has been selected for all inter-frame prediction blocks.
[0231] In some implementations, to acquire multiple reference blocks, memory stores instructions that, when executed by the processor, cause the processor to parse the bitstream to obtain multiple motion data. In some implementations, to acquire multiple reference blocks, memory stores instructions that, when executed by the processor, cause the processor to generate multiple reference blocks based on the multiple motion data.
[0232] In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to fuse the multiple reference blocks to obtain a fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to filter the fused reference blocks to obtain a fused and filtered reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to generate an FInter prediction for the current block based on the fused and filtered reference blocks.
[0233] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to select a set of filter weights for the FInter prediction filter, the set of filter weights minimizing the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction filter based on the set of filter weights. In some implementations, the FInter prediction for the current block can be further generated based on the FInter prediction filter.
[0234] According to another aspect of this disclosure, an apparatus for decoding is provided. The apparatus for decoding may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to acquire a plurality of reference blocks. The memory also stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0235] According to another aspect of this disclosure, a non-transitory computer-readable medium storing instructions is provided. When executed by a decoder's processor, the instructions cause the decoder's processor to acquire a plurality of reference blocks. When executed by a decoder's processor, the instructions cause the decoder's processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0236] In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, the instruction, when executed by the decoder's processor, can cause the decoder's processor to individually filter each of the plurality of reference blocks using corresponding filters to obtain a plurality of filtered reference blocks. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, the instruction, when executed by the decoder's processor, can cause the decoder's processor to fuse the plurality of filtered reference blocks to obtain a filtered and fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, the instruction, when executed by the decoder's processor, can cause the decoder's processor to generate an FInter prediction for the current block based on the filtered and fused reference block.
[0237] In some implementations, when executed by the decoder's processor, the instruction may cause the decoder's processor to parse the bitstream to obtain a first flag. In some implementations, when executed by the decoder's processor, the instruction may cause the decoder's processor to parse the bitstream to obtain a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, when executed by the decoder's processor, the instruction may cause the decoder's processor to determine, based on the second flag, whether FInter prediction mode has been selected for the current block.
[0238] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.
[0239] In some implementations, when executed by the decoder's processor, the instruction may cause the decoder's processor to generate a regular inter-frame prediction for the current block based on multiple reference blocks in response to determining that the FInter prediction mode has not been selected for the current block.
[0240] In some implementations, when executed by the decoder's processor, the instructions may cause the decoder's processor to parse the bitstream to obtain a flag. In some implementations, when executed by the decoder's processor, the instructions may cause the decoder's processor to determine, based on the flag, whether the FInter prediction mode has been selected for all inter-frame prediction blocks.
[0241] In some implementations, to acquire multiple reference blocks, instructions executed by the decoder's processor may cause the decoder's processor to parse the bitstream to obtain multiple motion data. In some implementations, to acquire multiple reference blocks, instructions executed by the decoder's processor may cause the decoder's processor to generate multiple reference blocks based on the multiple motion data.
[0242] In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, the instruction, when executed by the decoder's processor, can cause the decoder's processor to fuse the multiple reference blocks to obtain a fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, the instruction, when executed by the decoder's processor, can cause the decoder's processor to filter the fused reference block to obtain a fused and filtered reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, the instruction, when executed by the decoder's processor, can cause the decoder's processor to generate an FInter prediction for the current block based on the fused and filtered reference block.
[0243] In some implementations, when executed by the decoder's processor, the instruction causes the decoder's processor to select a set of filter weights for the FInter prediction filter, which minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, when executed by the decoder's processor, the instruction causes the decoder's processor to generate the FInter prediction filter based on the set of filter weights. In some implementations, the FInter prediction for the current block can also be generated based on the FInter prediction filter.
[0244] According to one aspect of this disclosure, an encoding method is provided. The method may include acquiring a plurality of reference blocks by a processor. The method may also include: in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on the plurality of reference blocks.
[0245] In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include filtering each of the multiple reference blocks using corresponding filters to obtain multiple filtered reference blocks. In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include fusing the multiple filtered reference blocks to obtain a filtered and fused reference block. In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on the filtered and fused reference block may include generating an FInter prediction for the current block by the processor based on the filtered and fused reference block.
[0246] In some implementations, the method may include encoding a first flag by a processor. In some implementations, the method may include encoding a second flag by a processor in response to a first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, the method may include the processor determining, based on the second flag, whether FInter prediction mode has been selected for the current block.
[0247] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.
[0248] In some implementations, the method may include: in response to determining that an FInter prediction mode has not been selected for the current block, the processor generates a regular inter-frame prediction for the current block based on multiple reference blocks.
[0249] In some implementations, the method may include encoding a flag by a processor. In some implementations, the method may include determining, based on the flag, whether an inter-prediction mode has been selected for all inter-prediction blocks.
[0250] In some implementations, acquiring multiple reference blocks by a processor may include encoding multiple motion data by the processor. In some implementations, acquiring multiple reference blocks by a processor may include generating multiple reference blocks by the processor based on the multiple motion data.
[0251] In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include fusing the multiple reference blocks to obtain a fused reference block. In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include filtering the fused reference block to obtain a fused and filtered reference block. In some implementations, in response to determining that an FInter prediction mode has been selected for the current block, generating an FInter prediction for the current block by the processor based on multiple reference blocks may include generating an FInter prediction for the current block by the processor based on the fused and filtered reference block.
[0252] In some implementations, the method may include a set of filter weights selected by the processor for the FInter prediction filter, the set of filter weights minimizing the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, the method may include the processor generating an FInter prediction filter based on the set of filter weights. In some implementations, an FInter prediction for the current block may be further generated based on the FInter prediction filter.
[0253] According to another aspect of this disclosure, an encoder is provided. The encoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to acquire a plurality of reference blocks. The memory also stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0254] In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to individually filter each of the plurality of reference blocks using a corresponding filter to obtain a plurality of filtered reference blocks. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to fuse the plurality of filtered reference blocks to obtain a filtered and fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to generate an FInter prediction for the current block based on the filtered and fused reference block.
[0255] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to encode a first flag. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to encode a second flag in response to a first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to determine, based on the second flag, whether FInter prediction mode has been selected for the current block.
[0256] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.
[0257] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to generate a regular inter-frame prediction for the current block based on multiple reference blocks in response to determining that an FInter prediction mode has not been selected for the current block.
[0258] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to encode a flag. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to determine, based on the flag, whether the FInter prediction mode has been selected for all inter-frame prediction blocks.
[0259] In some implementations, to acquire multiple reference blocks, the memory stores instructions that, when executed by the processor, cause the processor to encode multiple motion data. In some implementations, to acquire multiple reference blocks, the memory stores instructions that, when executed by the processor, cause the processor to generate multiple reference blocks based on the multiple motion data.
[0260] In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to fuse the multiple reference blocks to obtain a fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to filter the fused reference block to obtain a fused and filtered reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, a memory instruction is stored that, when executed by the processor, causes the processor to generate an FInter prediction for the current block based on the fused and filtered reference block.
[0261] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to select a set of filter weights for the FInter prediction filter, the set of filter weights minimizing the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction filter based on the set of filter weights. In some implementations, the FInter prediction for the current block can be further generated based on the FInter prediction filter.
[0262] According to another aspect of this disclosure, an apparatus for encoding is provided. The apparatus for encoding may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to fetch a plurality of reference blocks. The memory also stores instructions that, when executed by the processor, cause the processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0263] According to another aspect of this disclosure, a non-transitory computer-readable medium storing instructions is provided. When executed by a processor of an encoder, the instructions cause the encoder's processor to acquire a plurality of reference blocks. When executed by a processor of the encoder, the instructions cause the encoder's processor to generate an FInter prediction for the current block based on the plurality of reference blocks in response to determining that an FInter prediction mode has been selected for the current block.
[0264] In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, the instruction, when executed by the encoder's processor, can cause the encoder's processor to individually filter each of the plurality of reference blocks using corresponding filters to obtain a plurality of filtered reference blocks. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, the instruction, when executed by the encoder's processor, can cause the encoder's processor to fuse the plurality of filtered reference blocks to obtain a filtered and fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on a plurality of reference blocks, the instruction, when executed by the encoder's processor, can cause the encoder's processor to generate an FInter prediction for the current block based on the filtered and fused reference block.
[0265] In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to encode a first flag. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to encode a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to determine, based on the second flag, whether FInter prediction mode has been selected for the current block.
[0266] In some implementations, the first flag can be signaled at the SPS level, PH level, PPS level, or SH level.
[0267] In some implementations, when executed by the encoder's processor, the instruction may cause the encoder's processor to generate a regular inter-frame prediction for the current block based on multiple reference blocks in response to determining that the FInter prediction mode has not been selected for the current block.
[0268] In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to encode a flag. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to determine, based on the flag, whether an inter-prediction mode has been selected for all inter-prediction blocks.
[0269] In some implementations, to acquire multiple reference blocks, instructions executed by the encoder's processor may cause the encoder's processor to encode multiple motion data. In some implementations, to acquire reference blocks, instructions executed by the encoder's processor may cause the encoder's processor to generate multiple reference blocks based on the multiple motion data.
[0270] In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, the instruction, when executed by the encoder's processor, can cause the encoder's processor to fuse the multiple reference blocks to obtain a fused reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, the instruction, when executed by the encoder's processor, can cause the encoder's processor to filter the fused reference block to obtain a fused and filtered reference block. In some embodiments, in response to determining that an FInter prediction mode has been selected for the current block to generate an FInter prediction for the current block based on multiple reference blocks, the instruction, when executed by the encoder's processor, can cause the encoder's processor to generate an FInter prediction for the current block based on the fused and filtered reference block.
[0271] In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to select a set of filter weights for the FInter prediction filter, which minimizes the MSE between the reference template associated with the reference block and the current template associated with the current block. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to generate the FInter prediction filter based on the set of filter weights. In some implementations, the FInter prediction for the current block may be further generated based on the FInter prediction filter.
[0272] According to another aspect of this disclosure, a non-transitory computer-readable medium for storing bit streams is provided. Bit streams can be generated according to one or more of the operations disclosed herein.
[0273] The foregoing description of the embodiments will reveal the general nature of this disclosure, enabling others to readily modify and / or adapt such embodiments for various applications by applying their technical knowledge in the art without departing from the general concept of this disclosure, without excessive experimentation. Therefore, based on the teachings and guidance presented herein, such modifications and adaptations should be considered to fall within the meaning and scope of equivalents of the disclosed embodiments. It should be understood that the wording or terminology herein is for descriptive and not limiting purposes, and that the terminology or terminology of this specification will be interpreted by those skilled in the art based on the teachings and guidance.
[0274] Embodiments of this disclosure have been described above using functional building blocks that illustrate implementations of specified functions and their relationships. For ease of description, the boundaries of these functional building blocks have been arbitrarily defined herein. Alternative boundaries may be defined, provided that the specified functions and their relationships are properly performed.
[0275] The summary and abstract may set forth one or more, but not all, exemplary embodiments of this disclosure as contemplated by the inventors, and are therefore not intended to limit this disclosure and the appended claims in any way.
[0276] Various functional blocks, modules, and steps have been disclosed above. The arrangements provided are illustrative and not limiting. Therefore, functional blocks, modules, and steps can be rearranged or combined in ways different from the examples provided above. Similarly, some embodiments include only a subset of functional blocks, modules, and steps, and any such subset is permitted.
[0277] The breadth and scope of this disclosure should not be limited by any of the foregoing exemplary embodiments, but should be defined solely by the appended claims and their equivalents.
Claims
1. A method for decoding using a decoder, characterized in that, The method includes: acquiring a plurality of reference blocks by a processor; and in response to determining that a filter inter-frame prediction (FInter) mode has been selected for the current block, generating an FInter prediction for the current block based on the plurality of reference blocks by the processor.
2. The method according to claim 1, wherein, The step of generating the FInter prediction for the current block by the processor based on the plurality of reference blocks in response to determining that the FInter prediction mode has been selected for the current block includes: the processor filtering each of the plurality of reference blocks using corresponding filters to obtain a plurality of filtered reference blocks; the processor fusing the plurality of filtered reference blocks to obtain a filtered and fused reference block; and the processor generating the FInter prediction for the current block based on the filtered and fused reference block.
3. The method according to claim 1, wherein, The method further includes: parsing a bitstream by the processor to obtain a first flag; parsing the bitstream by the processor to obtain a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block; and determining by the processor, based on the second flag, whether FInter prediction mode has been selected for the current block.
4. The method according to claim 3, wherein, The first flag is signaled at the sequence parameter set SPS level, image header PH level, image parameter set PPS level, or slice header SH level.
5. The method according to claim 3, wherein, The method further includes: in response to determining that the FInter prediction mode is not selected for the current block, the processor generates a regular inter-frame prediction for the current block based on the plurality of reference blocks.
6. The method according to claim 1, wherein, The method further includes: parsing the bitstream by the processor to obtain a flag; and determining by the processor, based on the flag, whether to select the FInter prediction mode for all inter-frame prediction blocks.
7. The method according to claim 1, wherein, The step of obtaining the plurality of reference blocks by the processor includes: parsing the bitstream by the processor to obtain a plurality of motion data; and generating the plurality of reference blocks by the processor based on the plurality of motion data.
8. The method according to claim 1, wherein, The step of generating the FInter prediction for the current block by the processor based on the plurality of reference blocks in response to determining that the FInter prediction mode has been selected for the current block includes: fusing the plurality of reference blocks by the processor to obtain a fused reference block; filtering the fused reference block by the processor to obtain a fused and filtered reference block; and generating the FInter prediction for the current block by the processor based on the fused and filtered reference block.
9. The method according to claim 1, wherein, The method further includes: the processor selecting a set of filter weights for the FInter prediction filter, the filter weights minimizing the mean square error (MSE) between a reference template associated with a reference block and a current template associated with the current block; and the processor generating the FInter prediction filter based on the set of filter weights, wherein the FInter prediction of the current block is further generated based on the FInter prediction filter.
10. A decoder, characterized in that, The decoder includes: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: acquire a plurality of reference blocks; and, in response to determining that a filter inter-frame prediction mode (FInter) has been selected for the current block, generate an FInter prediction for the current block based on the plurality of reference blocks.
11. The decoder according to claim 10, wherein, In response to determining that the FInter prediction mode has been selected for the current block, in order to generate the FInter prediction for the current block based on the plurality of reference blocks, the memory stores instructions that, when executed by the processor, cause the processor to: filter each of the plurality of reference blocks using corresponding filters respectively to obtain a plurality of filtered reference blocks; The multiple filtered reference blocks are fused to obtain a filtered and fused reference block; And the FInter prediction of the current block is generated based on the filtered and fused reference block.
12. The decoder according to claim 10, wherein, The memory stores instructions, which, when executed by the processor, cause the processor to: parse a bitstream to obtain a first flag; parse the bitstream to obtain a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block; and determine, based on the second flag, whether FInter prediction mode has been selected for the current block.
13. The decoder according to claim 12, wherein, The first flag is signaled at the sequence parameter set SPS level, image header PH level, image parameter set PPS level, or slice header SH level.
14. The decoder according to claim 12, wherein, The memory stores instructions that, when executed by the processor, cause the processor to: generate a regular inter-frame prediction for the current block based on the plurality of reference blocks in response to determining that the FInter prediction mode has not been selected for the current block.
15. The decoder according to claim 10, wherein, The memory stores instructions that, when executed by the processor, cause the processor to: parse the bitstream to obtain a flag; and determine, based on the flag, whether to select the FInter prediction mode for all inter-frame prediction blocks.
16. The decoder according to claim 10, wherein, In order to obtain the plurality of reference blocks, the memory stores instructions that, when executed by the processor, cause the processor to: parse a bitstream to obtain a plurality of motion data; and generate the plurality of reference blocks based on the plurality of motion data.
17. The decoder according to claim 10, wherein, In response to determining that the FInter prediction mode has been selected for the current block, in order to generate the FInter prediction for the current block based on the plurality of reference blocks, the memory stores instructions that, when executed by the processor, cause the processor to: fuse the plurality of reference blocks to obtain a fused reference block; The fused reference block is filtered to obtain a fused and filtered reference block; And the FInter prediction for the current block is generated based on the fused and filtered reference block.
18. The decoder according to claim 10, wherein, The memory stores instructions that, when executed by the processor, cause the processor to: select a set of filter weights for the FInter prediction filter, the filter weights minimizing the mean square error (MSE) between a reference template associated with a reference block and a current template associated with the current block; and generate the FInter prediction filter based on the set of filter weights, wherein the FInter prediction of the current block is further generated based on the FInter prediction filter.
19. An apparatus for decoding, characterized in that, The apparatus includes: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: acquire a plurality of reference blocks; and, in response to determining that a filter inter-frame prediction mode (FInter) has been selected for the current block, generate an FInter prediction for the current block based on the plurality of reference blocks.
20. A non-transitory computer-readable medium storing instructions, characterized in that, When executed by the decoder's processor, the instruction causes the decoder's processor to: acquire a plurality of reference blocks; and, in response to determining that a filter inter-frame FInter prediction mode has been selected for the current block, generate an FInter prediction for the current block based on the plurality of reference blocks.
21. The non-transitory computer-readable medium according to claim 20, wherein, In response to determining that the FInter prediction mode has been selected for the current block, in order to generate the FInter prediction of the current block based on the plurality of reference blocks, the instruction, when executed by the processor of the decoder, causes the processor of the decoder to: filter each of the plurality of reference blocks using corresponding filters respectively to obtain a plurality of filtered reference blocks; The multiple filtered reference blocks are fused to obtain a filtered and fused reference block; And the FInter prediction of the current block is generated based on the filtered and fused reference block.
22. The non-transitory computer-readable medium of claim 20, wherein, When the instruction is executed by the processor of the decoder, the processor of the decoder causes the processor to: parse the bitstream to obtain a first flag; In response to the first flag indicating that FInter prediction mode is enabled for the current block, the bit stream is parsed to obtain a second flag; and based on the second flag, it is determined whether FInter prediction mode has been selected for the current block.
23. The non-transitory computer-readable medium according to claim 22, wherein, The first flag is signaled at the sequence parameter set SPS level, image header PH level, image parameter set PPS level, or slice header SH level.
24. The non-transitory computer-readable medium according to claim 22, wherein, When the instruction is executed by the processor of the decoder, the processor of the decoder: in response to determining that the FInter prediction mode is not selected for the current block, generates a regular inter-frame prediction for the current block based on the plurality of reference blocks.
25. The non-transitory computer-readable medium according to claim 20, wherein, When the instruction is executed by the processor of the decoder, the processor of the decoder causes the processor to: parse the bitstream to obtain a flag; and determine, based on the flag, whether the FInter prediction mode has been selected for all inter-frame prediction blocks.
26. The non-transitory computer-readable medium of claim 20, wherein, In order to obtain the plurality of reference blocks, the instructions, when executed by the processor of the decoder, cause the processor of the decoder to: parse the bitstream to obtain a plurality of motion data; and generate the plurality of reference blocks based on the plurality of motion data.
27. The non-transitory computer-readable medium according to claim 20, wherein, In response to determining that the FInter prediction mode has been selected for the current block, in order to generate the FInter prediction for the current block based on the plurality of reference blocks, the instruction, when executed by the processor of the decoder, causes the processor of the decoder to: fuse the plurality of reference blocks to obtain a fused reference block; The fused reference block is filtered to obtain a fused and filtered reference block; And the FInter prediction for the current block is generated based on the fused and filtered reference block.
28. The non-transitory computer-readable medium according to claim 20, wherein, When executed by the processor of the decoder, the instructions cause the processor of the decoder to: select a set of filter weights for the FInter prediction filter, the filter weights minimizing the mean square error (MSE) between the reference template associated with the reference block and the current template associated with the current block; and generate the FInter prediction filter based on the set of filter weights, wherein the FInter prediction of the current block is further generated based on the FInter prediction filter.
29. A method for encoding by an encoder, characterized in that, The method includes: acquiring a plurality of reference blocks by a processor; and in response to determining that a filter inter-frame prediction (FInter) mode has been selected for the current block, generating an FInter prediction for the current block based on the plurality of reference blocks by the processor.
30. The method according to claim 29, wherein, In response to determining that the FInter prediction mode has been selected for the current block, the processor generates the FInter prediction for the current block based on the plurality of reference blocks, which includes: the processor filtering each of the plurality of reference blocks using corresponding filters to obtain a plurality of filtered reference blocks; the processor fusing the plurality of filtered reference blocks to obtain a filtered and fused reference block; and the processor generating the FInter prediction for the current block based on the filtered and fused reference block.
31. The method according to claim 29, wherein, The method further includes: encoding a first flag by the processor; encoding a second flag by the processor in response to the first flag indicating that FInter prediction mode is enabled for the current block; and determining by the processor, based on the second flag, whether FInter prediction mode has been selected for the current block.
32. The method according to claim 31, wherein, The first flag is signaled at the sequence parameter set SPS level, image header PH level, image parameter set PPS level, or slice header SH level.
33. The method according to claim 31, wherein, The method further includes: in response to determining that the FInter prediction mode is not selected for the current block, the processor generates a regular inter-frame prediction for the current block based on the plurality of reference blocks.
34. The method according to claim 29, wherein, The method further includes: encoding a flag by the processor; and determining by the processor, based on the flag, whether the FInter prediction mode has been selected for all inter-frame prediction blocks.
35. The method according to claim 29, wherein, The step of obtaining the plurality of reference blocks by the processor includes: encoding the plurality of motion data by the processor; and generating the plurality of reference blocks by the processor based on the plurality of motion data.
36. The method according to claim 29, wherein, In response to determining that the FInter prediction mode has been selected for the current block, the processor generates the FInter prediction for the current block based on the plurality of reference blocks, which includes: the processor fusing the plurality of reference blocks to obtain a fused reference block; the processor filtering the fused reference block to obtain a fused and filtered reference block; and the processor generating the FInter prediction for the current block based on the fused and filtered reference block.
37. The method according to claim 29, wherein, The method further includes: the processor selecting a set of filter weights for the FInter prediction filter, the filter weights minimizing the mean square error (MSE) between a reference template associated with a reference block and a current template associated with the current block; and the processor generating the FInter prediction filter based on the set of filter weights, wherein the FInter prediction of the current block is further generated based on the FInter prediction filter.
38. An encoder, characterized in that, The encoder includes: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: acquire a plurality of reference blocks; and, in response to determining that a filter inter-frame prediction mode (FInter) has been selected for the current block, generate an FInter prediction for the current block based on the plurality of reference blocks.
39. The encoder according to claim 38, wherein, In response to determining that the FInter prediction mode has been selected for the current block, in order to generate the FInter prediction for the current block based on the plurality of reference blocks, the memory stores instructions that, when executed by the processor, cause the processor to: filter each of the plurality of reference blocks using corresponding filters respectively to obtain a plurality of filtered reference blocks; The multiple filtered reference blocks are fused to obtain a filtered and fused reference block; And the FInter prediction of the current block is generated based on the filtered and fused reference block.
40. The encoder according to claim 38, wherein, The memory stores instructions, which, when executed by the processor, cause the processor to: encode a first flag; encode a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block; and determine, based on the second flag, whether FInter prediction mode has been selected for the current block.
41. The encoder according to claim 40, wherein, The first flag is signaled at the sequence parameter set SPS level, image header PH level, image parameter set PPS level, or slice header SH level.
42. The encoder according to claim 40, wherein, The memory stores instructions that, when executed by the processor, cause the processor to: generate a regular inter-frame prediction for the current block based on the plurality of reference blocks in response to determining that the FInter prediction mode has not been selected for the current block.
43. The encoder according to claim 38, wherein, The memory stores instructions that, when executed by the processor, cause the processor to: encode a flag; and determine, based on the flag, whether to select the FInter prediction mode for all inter-frame prediction blocks.
44. The encoder according to claim 38, wherein, In order to acquire the plurality of reference blocks, the memory stores instructions that, when executed by the processor, cause the processor to: encode the plurality of motion data; And generate the multiple reference blocks based on the multiple motion data.
45. The encoder according to claim 38, wherein, In response to determining that the FInter prediction mode has been selected for the current block, in order to generate the FInter prediction for the current block based on the plurality of reference blocks, the memory stores instructions that, when executed by the processor, cause the processor to: fuse the plurality of reference blocks to obtain a fused reference block; The fused reference block is filtered to obtain a fused and filtered reference block; And the FInter prediction for the current block is generated based on the fused and filtered reference block.
46. The encoder according to claim 38, wherein, The memory stores instructions that, when executed by the processor, cause the processor to: select a set of filter weights for the FInter prediction filter, the filter weights minimizing the mean square error (MSE) between a reference template associated with a reference block and a current template associated with the current block; and generate the FInter prediction filter based on the set of filter weights, wherein the FInter prediction of the current block is further generated based on the FInter prediction filter.
47. An apparatus for encoding, characterized in that, The apparatus includes: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: acquire a plurality of reference blocks; and, in response to determining that a filter inter-frame prediction mode (FInter) has been selected for the current block, generate an FInter prediction for the current block based on the plurality of reference blocks.
48. A non-transitory computer-readable medium storing instructions, characterized in that, When executed by the encoder's processor, the instructions cause the encoder's processor to: acquire a plurality of reference blocks; and, in response to determining that a filter inter-frame FInter prediction mode has been selected for the current block, generate an FInter prediction for the current block based on the plurality of reference blocks.
49. The non-transitory computer-readable medium according to claim 48, wherein, In response to determining that the FInter prediction mode has been selected for the current block, in order to generate the FInter prediction of the current block based on the plurality of reference blocks, the instruction, when executed by the processor of the encoder, causes the processor of the encoder to: filter each of the plurality of reference blocks using corresponding filters respectively to obtain a plurality of filtered reference blocks; The multiple filtered reference blocks are fused to obtain a filtered and fused reference block; And the FInter prediction of the current block is generated based on the filtered and fused reference block.
50. The non-transitory computer-readable medium according to claim 48, wherein, When executed by the processor of the encoder, the instructions cause the processor of the encoder to: encode a first flag; encode a second flag in response to the first flag indicating that FInter prediction mode is enabled for the current block; and determine, based on the second flag, whether FInter prediction mode has been selected for the current block.
51. The non-transitory computer-readable medium according to claim 50, wherein, The first flag is signaled at the sequence parameter set SPS level, image header PH level, image parameter set PPS level, or slice header SH level.
52. The non-transitory computer-readable medium according to claim 50, wherein, When the instruction is executed by the processor of the encoder, the processor of the encoder: in response to determining that the FInter prediction mode is not selected for the current block, generates a regular inter-frame prediction for the current block based on the plurality of reference blocks.
53. The non-transitory computer-readable medium according to claim 48, wherein, When the instruction is executed by the processor of the encoder, the processor of the encoder causes the processor to: encode a flag; and determine, based on the flag, whether to select the FInter prediction mode for all inter-frame prediction blocks.
54. The non-transitory computer-readable medium according to claim 48, wherein, In order to obtain the plurality of reference blocks, the instructions, when executed by the processor of the encoder, cause the processor of the encoder to: encode the plurality of motion data; and generate the plurality of reference blocks based on the plurality of motion data.
55. The non-transitory computer-readable medium according to claim 48, wherein, In response to determining that the FInter prediction mode has been selected for the current block, in order to generate the FInter prediction for the current block based on the plurality of reference blocks, the instruction, when executed by the processor of the encoder, causes the processor of the encoder to: fuse the plurality of reference blocks to obtain a fused reference block; The fused reference block is filtered to obtain a fused and filtered reference block; And the FInter prediction for the current block is generated based on the fused and filtered reference block.
56. The non-transitory computer-readable medium according to claim 48, wherein, When executed by the processor of the encoder, the instructions cause the processor of the encoder to: select a set of filter weights for the FInter prediction filter, the filter weights minimizing the mean square error (MSE) between the reference template associated with the reference block and the current template associated with the current block; and generate the FInter prediction filter based on the set of filter weights, wherein the FInter prediction of the current block is further generated based on the FInter prediction filter.
57. A non-transitory computer-readable medium for storing bit streams, characterized in that, The bitstream is generated according to one or more of claims 29-37.