System and method for intra template matching
By parsing the bitstream through the decoder and encoder processor, decoding the syntax elements associated with intra-template matching prediction, determining whether to enable intraTMP mode, and using the fused weight set and reference block set to decode and encode video blocks, the problem of low efficiency of intra-template matching prediction in the prior art is solved, and more efficient video coding is achieved.
Patent Information
- Application Number
- CN202480025194.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-14
- Filing Date
- 2024-04-12
- Publication Date
- 2025-11-21
AI Technical Summary
Existing video coding technologies suffer from low efficiency and high coding complexity in intra-frame template matching prediction, making it difficult to effectively utilize intra-frame template matching prediction modes for video coding.
The decoder and encoder processor parse the bitstream, decode multiple syntax elements associated with intra-template matching prediction, determine whether intraTMP mode is enabled, determine intra-template matching prediction values through fusion mode, and decode and encode video blocks using fusion weight set and reference block set.
It improves the efficiency and complexity of video coding, enhances the effectiveness of intra-frame template matching prediction, and optimizes the compression and decoding process of video data.
Smart Images

Figure CN121002864A_ABST
Abstract
Description
[0001] Cross-referencing related applications This application claims priority to U.S. Provisional Application No. 63 / 459,236, filed April 13, 2023, entitled “SYSTEMS AND METHODS FOR TNTRA TEMPLATE MATCHING”, and U.S. Provisional Application No. 63 / 459,550, filed April 14, 2023, both of which are incorporated herein by reference in their entirety. Technical Field
[0002] The embodiments disclosed herein relate to video encoding and decoding. Background Technology
[0003] Digital video has become mainstream and is used in a wide variety of applications, including digital television, video telephony, and teleconferencing. These digital video applications are made possible by advancements in computing and communication technologies, as well as efficient video coding techniques. Video data can be compressed using various video coding techniques, allowing it to be encoded using one or more video coding standards. Exemplary video coding standards may include, but are not limited to, Universal Video Coding (H.266 / VVC), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Expert Group (MPEG) coding, and Enhanced Video Coding Model (ECM). Summary of the Invention
[0004] According to one aspect of this disclosure, a decoding method performed by a decoder is provided. The method may include: a processor parsing a bitstream to decode a plurality of syntax elements associated with intra-template matching prediction (intraTMP). The method may include: the processor decoding a first syntax element from the bitstream. The method may include: the processor determining, based on the first syntax element, whether an intraTMP mode is enabled for a current block. The method may include: in response to enabling the intraTMP mode for the current block, the processor decoding a second syntax element from the bitstream. The method may include: the processor determining, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The method may include: in response to determining that the intraTMP prediction value is determined using an intraTMP fusion mode, the processor decoding a third and a fourth syntax element from the bitstream. The method may include: the processor determining a fusion weight set based on the fourth syntax element. The method may include: the processor determining the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. The method may include: the processor decoding the current block based on the intraTMP prediction value.
[0005] According to another aspect of this disclosure, a decoder is provided. The decoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to: parse a bitstream to decode a plurality of syntax elements associated with intraTMP. The memory stores instructions that, when executed by the processor, cause the processor to: decode a first syntax element from the bitstream. The memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the first syntax element, whether to enable intraTMP mode for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: decode a second syntax element from the bitstream in response to enabling intraTMP mode for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Decode a third syntax element and a fourth syntax element from the bitstream in response to determining the intraTMP prediction value through an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Determine a fusion weight set based on the fourth syntax element. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Determine the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Decode the current block based on the intraTMP prediction value.
[0006] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided that stores instructions for a decoder. The memory stores instructions that, when executed by a processor of the decoder, cause the processor to: parse a bitstream to decode a plurality of syntax elements associated with intraTMP. The memory stores instructions that, when executed by a processor of the decoder, cause the processor to: decode a first syntax element from the bitstream. The memory stores instructions that, when executed by a processor of the decoder, cause the processor to: determine, based on the first syntax element, whether to enable intraTMP mode for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: decode a second syntax element from the bitstream in response to enabling intraTMP mode for the current block. The memory stores instructions that, when executed by a processor of the decoder, cause the processor to: determine, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: in response to determining the intraTMP prediction value through an intraTMP fusion mode, decode a third syntax element and a fourth syntax element from the bitstream. The memory storage instruction, when executed by the decoder's processor, causes the decoder's processor to perform the following operations: determine the fusion weight set based on the fourth syntax element. The memory storage instruction, when executed by the decoder's processor, causes the decoder's processor to perform the following operations: determine the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. The memory storage instruction, when executed by the decoder's processor, causes the decoder's processor to perform the following operations: decode the current block based on the intraTMP prediction value.
[0007] According to another aspect of this disclosure, an encoding method performed by an encoder is provided. The method may include: a processor enabling intraTMP to encode a current block. The method may include: the processor encoding a first syntax element into a bitstream. The method may include: the processor determining, based on the first syntax element, whether to enable an intraTMP mode for the current block. The method may include: in response to enabling the intraTMP mode for the current block, the processor encoding a second syntax element into the bitstream. The method may include: the processor determining, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The method may include: in response to determining that the intraTMP prediction value is determined using an intraTMP fusion mode, the processor encoding a third and a fourth syntax element into the bitstream. The method may include: the processor determining a fusion weight set based on the fourth syntax element. The method may include: the processor determining the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. The method may include: the processor encoding the current block based on the intraTMP prediction value.
[0008] According to another aspect of this disclosure, an encoder is provided. The encoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to: enable intraTMP for encoding a current block. The memory stores instructions that, when executed by the processor, cause the processor to: encode a first syntax element into a bitstream. The memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the first syntax element, whether to enable intraTMP mode for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: encode a second syntax element into the bitstream in response to enabling intraTMP mode for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Encoding a third syntax element and a fourth syntax element into the bitstream in response to determining the intraTMP prediction value through an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Determining the fusion weight set based on the fourth syntax element. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Determining the intraTMP prediction value based on the fusion weight set and the set of reference blocks indicated by the third syntax element. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Encoding the current block based on the intraTMP prediction value.
[0009] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided that stores instructions for an encoder. When executed by a processor of the encoder, the instructions can cause the encoder processor to: enable intraTMP to encode a current block. When executed by a processor of the encoder, the instructions can cause the encoder processor to: encode a first syntax element into a bitstream. When executed by a processor of the encoder, the instructions can cause the encoder processor to: determine, based on the first syntax element, whether to enable intraTMP mode for the current block. A memory stores instructions that, when executed by the processor, can cause the processor to: encode a second syntax element into the bitstream in response to enabling intraTMP mode for the current block. When executed by a processor of the encoder, the instructions can cause the encoder processor to: determine, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: in response to determining the intraTMP prediction value through an intraTMP fusion mode, encode a third syntax element and a fourth syntax element into the bitstream. When executed by the encoder's processor, the instructions cause the encoder's processor to perform the following operations: determine the fusion weight set based on the fourth syntax element. When executed by the encoder's processor, the instructions cause the encoder's processor to perform the following operations: determine the intraTMP prediction value based on the fusion weight set and the reference block set indicated by the third syntax element. When executed by the encoder's processor, the instructions cause the encoder's processor to perform the following operations: encode the current block based on the intraTMP prediction value.
[0010] These illustrative embodiments are described not to limit or restrict this disclosure, but to provide examples that aid in understanding it. Further embodiments are described in the detailed description, and further description is provided below. Attached Figure Description
[0011] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments of the present disclosure and, together with the specification, further serve to illustrate the principles of the present disclosure and enable those skilled in the art to implement and use the present disclosure.
[0012] Figure 1 A block diagram of an exemplary encoding system according to some embodiments of the present disclosure is shown.
[0013] Figure 2 A block diagram of an exemplary decoding system according to some embodiments of the present disclosure is shown.
[0014] Figure 3 Some embodiments according to this disclosure are shown. Figure 1 A detailed block diagram of an exemplary encoder in an encoding system.
[0015] Figure 4 Some embodiments according to this disclosure are shown. Figure 2 A detailed block diagram of an exemplary decoder in a decoding system.
[0016] Figure 5 Exemplary images showing partitions into coding tree units (CTUs) according to some embodiments of the present disclosure are shown.
[0017] Figure 6 An exemplary CTU, divided into coding units (CUs) according to some embodiments of the present disclosure, is shown.
[0018] Figure 7 A schematic diagram is shown of the current CU block and spatially adjacent and non-adjacent reconstructed samples of the current block according to some embodiments of the present disclosure.
[0019] Figure 8 A schematic diagram of the angle mode of VVC according to some embodiments of the present disclosure is shown.
[0020] Figure 9A A schematic diagram of an intra-template matching prediction (intraTMP) search region according to some embodiments of the present disclosure is shown.
[0021] Figure 9B A schematic diagram of the intraTMP extended search region according to some embodiments of the present disclosure is shown.
[0022] Figure 10 The spatial components of an intraTMP filter according to some embodiments of this disclosure are shown.
[0023] Figure 11 A reference region for deriving filter coefficients for intraTMP is shown according to some embodiments of the present disclosure.
[0024] Figure 12 A diagram showing the adjacent half-pixel positions for intraTMP according to some embodiments of the present disclosure is illustrated.
[0025] Figure 13A diagram showing various template shapes for intraTMP according to some embodiments of the present disclosure is provided.
[0026] Figure 14 A diagram showing fractional block vector positions for intraTMP according to some embodiments of the present disclosure is illustrated.
[0027] Figures 15A to 15D A flowchart illustrating an exemplary method for video decoding according to some embodiments of the present disclosure is shown.
[0028] Figures 16A to 16D A flowchart illustrating an exemplary method of video encoding according to some embodiments of the present disclosure is shown.
[0029] Embodiments of this disclosure will be described with reference to the accompanying drawings. Detailed Implementation
[0030] While some configurations and arrangements have been discussed, it should be understood that this is for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of this disclosure. It will be apparent to those skilled in the art that this disclosure can also be used in a variety of other applications.
[0031] It should be noted that the use of terms such as "one embodiment," "embodiment," "exemplary embodiment," "some embodiments," and "certain embodiments" in the specification indicates that the described embodiments may include specific features, structures, or characteristics, but not every embodiment must include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an embodiment, whether explicitly described or not, implementing such a feature, structure, or characteristic in conjunction with other embodiments will fall within the knowledge scope of those skilled in the art.
[0032] Generally, terms can be understood, at least in part, based on their usage in the context. For example, the term "one or more," as used herein, can be used, at least in part, to describe any feature, structure, or characteristic in a singular sense, or a combination of features, structures, or characteristics in a plural sense, depending at least in part on the context. Similarly, terms such as "a," "an," or "the" can also be understood to convey either a singular or a plural usage, depending at least in part on the context. Furthermore, the term "based on" can be understood to not necessarily convey an exclusive set of factors, but rather to allow for the presence of additional factors that may not be explicitly described, again depending at least in part on the context.
[0033] Various aspects of a video coding system will now be described with reference to various apparatuses and methods. These apparatuses and methods will be described in the detailed description below and will be shown in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively, “elements”). These elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the overall system.
[0034] The techniques described herein can be used in a variety of video encoding and decoding applications. As described herein, video encoding and decoding include both encoding and decoding of video. Video encoding and decoding can be performed by block units. For example, encoding / decoding processes such as transform, quantization, prediction, in-loop filtering, reconstruction, etc., can be performed on coded blocks, transform blocks, or prediction blocks. As described herein, the block to be encoded / decoded will be referred to as the "current block." For example, the current block can represent a coded block, transform block, or prediction block according to the current encoding / decoding process. Furthermore, it will be understood that the term "unit" as used in this disclosure refers to a basic unit used to perform a particular encoding / decoding process, and the term "block" refers to a sample array of a predetermined size. Unless otherwise stated, "block" and "unit" are used interchangeably.
[0035] Figure 1 A block diagram of an exemplary encoding system 100 according to some embodiments of the present disclosure is shown. Figure 2 A block diagram of an exemplary decoding system 200 according to some embodiments of the present disclosure is shown. Each system 100 or 200 can be applied to or integrated into various systems and devices capable of data processing, such as computers and wireless communication devices. For example, system 100 or 200 can be all or part of a mobile phone, desktop computer, laptop computer, tablet computer, in-vehicle computer, game console, printer, positioning device, wearable electronic device, smart sensor, virtual reality (VR) device, augmented reality (AR) device, or any other suitable electronic device with data processing capabilities. Figure 7 and Figure 8 As shown, system 100 or 200 may include processor 102, memory 104, and interface 106. These components are shown interconnected via a bus, but other connection types are also permitted. It should be understood that system 100 or 200 may include any other suitable components for performing the functions described herein.
[0036] Processor 102 may include a microprocessor, such as a graphics processing unit (GPU), image signal processor (ISP), central processing unit (CPU), digital signal processor (DSP), tensor processing unit (TPU), vision processing unit (VPU), neural processing unit (NPU), synergistic processing unit (SPU), or physics processing unit (PPU), microcontroller unit (MCU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), programmable logic device (PLD), state machine, gated logic, discrete hardware circuitry, and other suitable hardware configured to perform the various functions described throughout this disclosure. Although Figure 7 and Figure 8 Only one processor is shown, but it is understood that multiple processors may be included. Processor 102 may be a hardware device having one or more processing cores. Processor 102 can execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise, software should be broadly understood to refer to instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc. Software may include computer instructions written in interpreted languages, compiled languages, or machine code. Other technologies used to indicate hardware are also permitted within the broad category of software.
[0037] Memory 104 can broadly include main memory (also known as primary / system memory) and secondary storage (also known as secondary storage). For example, memory 104 may include random-access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferro-electric RAM (FRAM), electrically erasable programmable ROM (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, hard disk drive (HDD) such as disk storage or other magnetic storage devices, flash memory drive, solid-state drive (SSD), or any other medium that can be used to carry or store desired program code in the form of instructions accessible and executable by processor 102. More broadly, memory 104 can be implemented from any computer-readable medium, such as non-transitory computer-readable media. Although Figure 7 and Figure 8 Only one memory is shown, but it is understood that multiple memories may be included.
[0038] Interface 106 can broadly include data interfaces and communication interfaces, used for receiving and transmitting signals in the process of exchanging information with other external network elements. For example, interface 106 may include input / output (I / O) devices and wired or wireless transceivers. Although Figure 7 and Figure 8 Only one memory is shown, but it should be understood that multiple interfaces may be included.
[0039] Processor 102, memory 104, and interface 106 can be implemented in various forms within system 100 or 200 for performing video encoding and decoding functions. In some embodiments, processor 102, memory 104, and interface 106 of system 100 or 200 are implemented (e.g., integrated) on one or more system-on-chip (SoC) devices. In one example, processor 102, memory 104, and interface 106 can be integrated on an application processor (AP) SoC that controls application processing within an operating system (OS) environment, including running video encoding and decoding applications. In another example, processor 102, memory 104, and interface 106 can be integrated on a dedicated processor chip for video encoding and decoding, such as a GPU or ISP chip dedicated to image and video processing within a real-time operating system (RTOS).
[0040] like Figure 1 As shown, in the encoding system 100, the processor 102 may include one or more modules, such as the encoder 101. Although Figure 1 Encoder 101 is shown within a processor 102, but it should be understood that encoder 101 may include one or more submodules that may be implemented on different processors located close to or far from each other. Encoder 101 (and any corresponding submodules or subunits) may be a hardware unit of processor 102 (e.g., part of an integrated circuit) designed for use with other components or software units implemented by processor 102 by executing at least a portion of a program (e.g., instructions). The program instructions may be stored on a computer-readable medium such as memory 104, and when executed by processor 102, may perform processes having one or more functions related to video coding, such as image segmentation, inter-frame prediction, intra-frame prediction, conversion, quantization, filtering, entropy coding, etc., as described below.
[0041] Similarly, such as Figure 2 As shown, in the decoding system 200, the processor 102 may include one or more modules, such as the decoder 201. Although Figure 2Decoder 201 is shown within a processor 102, but it should be understood that decoder 201 may include one or more submodules that may be implemented on different processors located close to or far from each other. Decoder 201 (and any corresponding submodules or subunits) may be a hardware unit of processor 102 (e.g., part of an integrated circuit) designed for use with other components or software units implemented by processor 102 by executing at least a portion of a program (e.g., instructions). The program instructions may be stored on a computer-readable medium such as memory 104, and when executed by processor 102, may perform processes having one or more functions related to video decoding, such as entropy decoding, inverse quantization, inverse transform, inter-frame prediction, intra-frame prediction, filtering, as described below.
[0042] Figure 3 Some embodiments according to this disclosure are shown. Figure 1 A detailed block diagram of an exemplary encoder 101 in the encoding system 100. (See attached diagram.) Figure 3 As shown, encoder 101 may include segmentation module 302, inter-frame prediction module 304, intra-frame prediction module 306, transform module 308, quantization module 310, inverse quantization module 312, inverse transform module 314, filtering module 316, buffer module 318, and encoding module 320. It should be understood that... Figure 3 Each element shown is presented independently in the video encoder to represent a unique function distinct from each other, and this does not mean that each component consists of a separate hardware or software configuration unit. That is, for ease of illustration, each element is included as listed as an element, but at least two elements may be combined to form a single element, or an element may be divided into multiple elements for performing a function. It is also understood that some elements are not essential for performing the functions described in this disclosure, but may be optional elements used to improve performance. It should also be understood that these elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on encoder 101.
[0043] The segmentation module 302 can be configured to divide the input video image into at least one processing unit. The image can be a frame or a field of the video. In some embodiments, the image includes a monochrome luminance sample array, or a luminance sample array and two corresponding chrominance sample arrays. In this case, the processing unit can be a prediction unit (PU), a transform unit (TU), or a coding unit (CU). The segmentation module 302 can divide the image into combinations of multiple coding units, prediction units, and transform units, and encode the image by selecting combinations of coding units, prediction units, and transform units based on a predetermined criterion (e.g., a cost function).
[0044] Similar to H.265 / HEVC, H.266 / VVC is a block-based hybrid spatial and temporal predictive coding scheme. Figure 5 As shown, during encoding, the input image 500 is first divided into square blocks CTU 502 by the segmentation module 302. For example, CTU 502 can be a 128×128 pixel block. Figure 6 As shown, each CTU 502 in the input image 500 can be segmented into one or more CUs 602 by the segmentation module 302. These CUs 602 can be used for prediction and transformation. Unlike H.265 / HEVC, in H.266 / VVC, CUs 602 can be rectangular or square and can be encoded without further segmentation into prediction units or transformation units. For example, as... Figure 6 As shown, the partitioning from CTU 502 to CU 602 may include quadtree partitioning (indicated by solid lines), binary tree partitioning (indicated by dashed lines), and ternary tree partitioning (indicated by dotted lines). According to some embodiments, each CU 602 may be as large as its root CTU 502, or it may be a subdivision as small as 4×4 blocks of the root CTU 502.
[0045] Reference Figure 4Inter-frame prediction module 304 can be configured to perform inter-frame prediction on prediction units, and intra-frame prediction module 306 can be configured to perform intra-frame prediction on prediction units. It can be determined whether inter-frame prediction or intra-frame prediction is used on prediction units, and specific information (e.g., intra-frame prediction mode, motion vector, reference image, etc.) is determined based on each prediction method. In this case, the processing unit used to perform prediction can be different from the processing unit used to determine the prediction method and specific content. For example, the prediction method and prediction mode can be determined in the prediction unit, and prediction can be performed in the transform unit. The residual coefficients in the residual block between the generated prediction block and the original block can be input to transform module 308. Furthermore, encoding module 320 can encode the prediction mode information, motion vector information, etc., used for prediction along with the residual coefficients or quantization level into the bitstream. It should be understood that in some encoding modes, the original block can be encoded as is without generating prediction blocks through prediction modules 304 or 306. It should also be understood that in some encoding modes, prediction, transform, and / or quantization can be skipped.
[0046] In some embodiments, the inter-frame prediction module 304 may predict prediction units based on information relating to at least one image preceding or following the current image, and in some cases, it may predict prediction units based on information relating to a portion of the encoded region in the current image. The inter-frame prediction module 304 may include sub-modules such as a reference image interpolation module, a motion prediction module, and a motion compensation module (not shown). For example, the reference image interpolation module may receive reference image information from the buffer module 318 and generate pixel information of an integer number of pixels or sub-pixels based on the reference image. In the case of luminance pixels, an 8-tap interpolation filter with variable filter coefficients based on discrete cosine transform (DCT) can be used to generate pixel information of an integer number of pixels or sub-pixels in units of 1 / 4 pixels. In the case of chrominance signals, a 4-tap interpolation filter with variable filter coefficients based on DCT can be used to generate pixel information of an integer number of pixels or sub-pixels in units of 1 / 8 pixels. The motion prediction module may perform motion prediction based on a reference image interpolated by the reference image interpolation module. Various methods, such as the full search-based block matching algorithm (FBMA), three-step search (TSS), and new three-step search (NTS), can be used to compute motion vectors. Motion vectors can be based on interpolated pixels having motion vector values in units of 1 / 2, 1 / 4, or 1 / 16 pixels or integer pixels. The motion prediction module can predict the current prediction unit by changing the motion prediction method. Various methods, such as skipping methods, merging methods, advanced motion vector prediction (AMVP) methods, and intra-block copying methods, can be used as motion prediction methods.
[0047] Still refer to Figure 3In some embodiments, the intra-prediction module 306 can generate prediction units based on information related to reference pixels surrounding the current block, which is pixel information in the current image. The reference pixels can be located in reference rows that are not adjacent to the current block. When a block in the neighborhood of the current prediction unit is a block that has undergone inter-frame prediction and therefore the reference pixel is a pixel that has undergone inter-frame prediction, the reference pixels included in the block that underwent inter-frame prediction can be used instead of the reference pixel information of the block in the neighborhood where intra-frame prediction was performed. That is, when a reference pixel is unavailable, at least one of the available reference pixels can be used to replace the unavailable reference pixel information. In intra-frame prediction, the prediction mode can have an angular prediction mode and a non-angular prediction mode. The angular prediction mode uses reference pixel information according to the prediction direction, while the non-angular prediction mode does not use direction information when performing prediction. The mode used to predict luminance information can be different from the mode used to predict chromatic difference information, and the intra-frame prediction mode information or the predicted luminance signal information used to predict luminance information can be used to predict chromatic difference information. If the size of the prediction cell is the same as the size of the transform cell when performing intra-prediction, intra-prediction can be performed based on the left, top-left, and top pixels of the prediction cell. However, if the size of the prediction cell is different from the size of the transform cell when performing intra-prediction, intra-prediction can be performed using reference pixels based on the transform cell.
[0048] Intra-prediction methods can generate prediction blocks after applying an adaptive intra-smoothing (AIS) filter to a reference pixel based on the prediction mode. The type of AIS filter applied to the reference pixel can vary. To perform intra-prediction, the intra-prediction mode of the current prediction unit can be predicted based on the intra-prediction modes of prediction units existing in the neighborhood of the current prediction unit. When using mode information predicted by neighboring prediction units to predict the prediction mode of the current prediction unit, if the intra-prediction mode of the current prediction unit is the same as that of the prediction units in the neighborhood, predetermined flag information can be used to send information indicating that the prediction mode of the current prediction unit is the same as that of the prediction units in the neighborhood; and if the prediction mode of the current prediction unit and the prediction modes of the prediction units in the neighborhood are different from each other, the prediction mode information of the current block can be encoded through additional flag information.
[0049] like Figure 3 As shown, the following residual block can be generated: This residual block includes the prediction unit that performed the prediction based on the prediction unit generated by prediction module 304 or 306, and residual coefficient information (also referred to herein as "residual") as the difference between the prediction unit and the original block. The generated residual block can be input to transform module 308. Further details regarding the residuals and transforms used for video encoding will now be provided.
[0050] In hybrid video coding systems, redundancy in the video signal is first utilized by applying inter-frame or intra-frame prediction tools to each control unit (CU). The difference between the original sample of a CU and its predicted block is often called the residual. Even after prediction, the residual remains highly spatially correlated. Although conditional entropy coding can capture some spatial dependencies between adjacent samples, it is computationally impractical to formulate an entropy coding statistical model that fully utilizes the spatial correlation in the residual. In contrast, transform coding is a practical and efficient method for spatial decorrelation of the residual.
[0051] For example, the transformation module 308 can transform the residuals using an integer version of the two-dimensional discrete cosine transform (DCT), which can be applied in both the horizontal and vertical directions. For an MxN residual sample block (where M is the width of the block and N is the height of the block), the transformation module 308 can obtain transformation coefficients by applying an MxM DCT to each row, thereby obtaining intermediate transformation coefficients, and then applying an NxN DCT to each column of the intermediate transformation coefficients.
[0052] For intra-coded CUs (also referred to as "intra-CUs" in this paper), spatially adjacent reconstructed samples are used to predict the current block, and the intra-prediction mode is signaled once for the entire CU. Each CU consists of one or more coded blocks (CBs) corresponding to the color components of the video sequence. For example, consumer videos are typically in a 4:2:0 chroma format, in which case each CU consists of one luma CB and two chroma CBs, with the chroma CBs having one-quarter the number of samples as luma CBs. Intra-prediction and transform coding are performed at the prediction block (PB) level and the transform block (TB) level, respectively. Each CB consists of a single TB, except in the case of intra subpartition (ISP) mode and implicit partitioning. For luma CBs, the maximum side length of the TB is 64, and the minimum side length is 4. Furthermore, the luma TB is designated as a W×H rectangular block with a width of W and a height of H, where W, H ∈ {4, 8, 16, 32, 64}. For chroma CB, the maximum TB side length is 32, and the chroma TB is a rectangular W×H block with width W and height H. Here, W, H∈{2, 4, 8, 16, 32}, but to meet memory architecture and throughput requirements, blocks with shapes of 2×H and 4×2 are excluded.
[0053] Figure 7A schematic diagram 700 shows the current CU block 702 and spatially adjacent and non-adjacent reconstructed samples of the current block according to some aspects of this disclosure. Figure 7 In the text, the numbers 0, 1, 2... indicate the pixel line index associated with the current CU block 702.
[0054] In VVC, reference samples obtained from reconstructed samples of adjacent blocks are used to generate intra-predicted samples for the current block. For a W×H block, the reference sample is spatially adjacent to the current block and consists of: a vertical line containing 2.H reconstructed samples extending downwards to the left of the current block; an upper-left reconstructed sample; and a horizontal line containing 2.W reconstructed samples extending to the right above the current block. This "L"-shaped sample set may be referred to as a "reference line" in this disclosure. The reference line directly adjacent to the current CU block 702 is in... Figure 7 The row shown is index 0.
[0055] Similar to AVC and HEVC, VVC also supports angular intra-prediction modes. Angular intra-prediction is a directional intra-prediction method. Compared to HEVC, VVC's angular intra-prediction is improved by increasing prediction accuracy and adapting to new segmentation frameworks. The former achieves this by increasing the number of angular prediction directions and using more accurate interpolation filters, while the latter achieves it by introducing wide-angle intra-prediction modes. In VVC, the number of directional modes available for a given block increases from 33 HEVC directions to 65 directions. Figure 8 The VVC angle mode 800 is described in the text.
[0056] Orientations with even indices between 2 and 66 are equivalent to the orientations of the angle modes supported in HEVC. For rectangular blocks, an equal number of angle modes are assigned to the top and left sides of the block. On the other hand, rectangular intra blocks, which do not exist in HEVC, are the central part of VVC's segmentation scheme, where additional intra-prediction orientations are assigned to the long side of the block. These additional orientations assigned along the long side are called wide-angle intra-prediction (WAIP) modes because they correspond to prediction orientations with angles greater than 45° relative to the horizontal or vertical modes. Figure 8 As shown, the WAIP mode for a given mode index is constrained by mapping the original directional mode to a mode with an index offset of 1 in the opposite direction. For a given rectangular block, the aspect ratio (i.e., width-to-height ratio) is used to determine which angular modes should be replaced with the corresponding wide-angle modes.
[0057] For square blocks in VVC, each pair of horizontally or vertically adjacent predicted samples is predicted based on pairs of adjacent reference samples. In contrast, WAIP extends the angular range of directional prediction to more than 45°, and therefore, for a coded block predicted using the WAIP pattern, adjacent predicted samples can be predicted based on non-adjacent reference samples.
[0058] Apart from directly adjacent sample lines. Figure 7 One of the two non-adjacent reference lines (line 1 and line 2) depicted in the image can also include input samples for intra-frame prediction in VVC. For ECM, more non-adjacent reference lines can be used. The use of adjacent and non-adjacent reference samples is called multiple reference line (MRL) prediction.
[0059] Intra-frame modes that can be used for MRL are DC mode and angle prediction mode. However, not all of these modes can be combined with MRL for a given block. MRL modes are always coupled with modes from the most probable mode (MPM) list in VVC. This coupling means that if non-adjacent reference lines are used, the intra-frame prediction mode is one of the MPMs. The design of such MPM-based MRL prediction modes was inspired by the observation that non-adjacent reference lines primarily favor texture patterns with sharp and highly directional edges. In these cases, MPMs are chosen more frequently because there is usually a strong correlation between the texture patterns of adjacent blocks and the texture pattern of the current block. On the other hand, choosing a non-MPM for intra-frame prediction can indicate that edges are not uniformly distributed in adjacent blocks, so the expected MRL prediction mode is less useful in this case. Furthermore, it has been observed that MRL does not provide additional coding gain when the intra-frame prediction mode is a planar mode, because this mode is typically used for smooth regions. Therefore, MRL does not include planar modes, which are always one of the MPMs. The angle or DC prediction process in MRL is very similar to the case of directly adjacent reference lines. However, for angular patterns with non-integer slopes, DCT-based interpolation filters (DCTIF) are always used. This design choice is supported by both experimental results and empirical observations that MRLs are primarily advantageous for sharp and directional edges, where DCTIF is more suitable than some other filters because it preserves more high frequencies.
[0060] From a hardware design perspective, applying multiple reference lines as proposed in the initial approach requires additional line buffer overhead, where the line buffer is used to store the additional reference lines. In typical hardware designs, the line buffer is part of the on-chip memory architecture used for image and video encoding, and minimizing the on-chip area of the line buffer is crucial. To address this issue, for encoding units attached to the top boundary of the CTU, the MRL is disabled and no signaling is used to notify the MRL. In this way, the additional buffer used to store non-adjacent reference lines is constrained to 128, where 128 is the width of the maximum unit size.
[0061] In some known methods, intra-frame prediction fusion methods have been proposed to improve the accuracy of intra-frame prediction. More specifically, if the current block is a luma block, encoded using a non-integer slope angle mode and not in ISP mode, and the block size (width * height) is greater than 16, then two prediction blocks generated based on two different reference lines will be "fused." The prediction fusion is calculated by a weighted sum of the two prediction blocks. More specifically, the first reference line with index i is specified in the bitstream using the current signaling transmission method. ), and the predicted block generated based on the reference line using the selected intra-frame prediction mode will be represented as ,in, This describes the operation of generating a prediction block based on a reference line using a given intra-frame prediction mode. In known methods, the reference line is implicitly selected. The second reference line is used as an index position that is farther from the current block relative to the first reference line. Similarly, the predicted block generated based on the second reference line is represented as follows: The weighted sum of the two prediction blocks is obtained according to equation (1), and this weighted sum is used as the prediction value of the current block.
[0062] (1), in, Indicates fusion prediction, and These are two weighting factors, and they were set to 3 / 4 and 1 / 4 respectively in the experiment.
[0063] Figure 9A A schematic diagram of an intraTMP search region 900 according to some embodiments of the present disclosure is shown.
[0064] Intra-template matching prediction (intraTMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame. The L-shaped or other shaped template of this best prediction block matches the current template. Unlike inter-block copy (IBC), no signaling is given to the block vector in the bitstream. For a predefined search range, encoder 101 searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. Encoder 101 then signals the use of this mode, and the same prediction operation is performed by decoder 201.
[0065] like Figure 9A As shown, a prediction signal is generated by matching the predefined causal neighborhood of the current block with another block in a predefined search area consisting of the current CTU, the top-left CTU, the top CTU, and the left-side CTU.
[0066] The cost function is the sum of absolute differences (SAD), or the sum of absolute transformed differences (SATD), or by comparing the hash values between templates. Within each region, the decoder 201 searches for the template with the minimum cost relative to the current template and uses its corresponding block as the prediction block.
[0067] The dimensions of the entire region (searchRangeWidth, searchRangeHeight) are set to be proportional to the block size (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. For example, searchRangeWidth = a * BlkW and searchRangeHeight = a * BlkH, where 'a' is a constant that controls the gain / complexity tradeoff. In practice, 'a' is equal to 5 in the ECM-7.0 testing software.
[0068] To speed up the template matching process, the search region is initially traversed in increments of 2 pixels. This is also known as the search subsampling factor of 2. This reduces the template matching search complexity to one-quarter. After finding the best match from the initial search, a refinement process is performed. Refinement is accomplished through a second template matching search around the best match with a reduced range. The reduced range is defined as min(BlkW, BlkH) / 2, where BlkW and BlkH are the width and height of the current block, respectively.
[0069] For CUs with a width and height of 64 or less, the intra-template matching prediction tool is enabled. This maximum CU size for intraTMP is configurable. When the decoder-side intra-mode derivation DIMD is not used for the current CU, signaling notification for the intraTMP mode is provided at the CU level via a dedicated flag.
[0070] Figure 9B A schematic diagram of an intraTMP extended search region 901 according to some embodiments of the present disclosure is shown.
[0071] The original intraTMP method implicitly selects only one block vector by searching for the minimum template matching cost. However, for content captured by a camera, template matching alone cannot find a good prediction. There are usually several blocks similar to the current block, and their template matching costs are comparable. The BV with the minimum template matching cost may not be the best prediction for the current block. To further improve coding performance, a multi-candidate approach can be used with intraTMP. In the multi-candidate approach, a prediction candidate list is constructed using candidate BVs, which are arranged in ascending order of their corresponding template matching costs. This candidate list is constructed by both encoder 101 and decoder 201. Signaling is used in the bitstream to indicate which candidate BV has been selected for the current block. This method uses template matching to select a final candidate list of promising candidates from a large number of possible BVs, and then allows encoder 101 (which can check the true coding cost of the current block for each of these candidates) to make rate-distortion optimized (RDO) decisions from the final candidate list. Compared to encoder RDO, the complexity of candidate list construction is relatively low.
[0072] The proposed syntax changes are shown in Table 1 below.
[0073] Table 1 In Table 1, intra_tmp_flag equals 1, indicating that the current block uses intraTMP, and intra_tmp_idx also indicates which BV in the candidate BV list is used to identify the predicted block.
[0074] To construct the candidate list, sparse search and refined search can be used. In the sparse search, the subsampling factor is set to 3, and the 30 BVs with the minimum SAD cost are retained after the sparse search. In the refined search, a 3x3 local search is examined around each of the 30 BVs. The 15 BVs with the minimum SAD cost after the refined search are selected to form the candidate list.
[0075] In the intraTMP fusion method, multiple intraTMP prediction blocks are first derived and then fused to produce a better overall prediction. These prediction blocks can also be called "matching blocks," "fusion blocks," or simply "predictions." The intraTMP fusion method is considered beneficial for the content captured by the camera. In summary, the intraTMP fusion method can include the following four operations: In the first operation, the intraTMP fusion method can generate multiple candidate predictions during the intraTMP search. For example, a sparse intraTMP search process is first performed with a search subsampling factor of 3. After the sparse search, a candidate list is generated using the 30 candidate BVs with the smallest template SAD. A full-pixel thinning search is performed around each candidate BV within a small region. The thinning region is a 3x3 region surrounding each of the 30 candidate BVs. Finally, the three best candidate BVs determined by the template SAD across all thinning regions are selected. The block pointed to by each of these candidate BVs is selected as the candidate prediction block for intraTMP fusion.
[0076] In the second operation, the intraTMP fusion method can select candidate predicted values for fusion. For example, for each of the three candidate predicted value blocks, a threshold is used to determine whether it should be used for fusion, as shown in equation (2) below.
[0077] Threshold= << 1 (2), in, The template SAD is the smallest among the three candidate predicted value blocks. Choose SAD<= Threshold The candidate predicted value blocks are then fused. This also determines the number of candidate predicted value blocks selected for fusion.
[0078] In the third operation, the intraTMP fusion method can calculate a weighting factor for each predicted value block selected for fusion. For example, once the predicted value blocks to be fused are determined, they are fused using weights. According to this disclosure, two methods can be used to determine these fusion weights.
[0079] In the first method of the third operation, the weights are fused. w i The calculation is performed based on their SAD. The fusion weights are calculated according to equations (3) and (4). w i .
[0080] (3), and (4).
[0081] To simplify implementation, the division operation is replaced with an integer lookup table (LUT). In the second approach for the third operation, fusion weights with fixed values are used to further reduce complexity. The weights are set to... .
[0082] In the fourth operation, the intraTMP fusion method can determine the final fusion prediction value according to equation (5).
[0083] (5), in, Let be the i-th prediction block, and n be the number of prediction blocks selected for fusion.
[0084] Still referring to the fourth operation, in the special case where only one block of predicted values is retained after the second operation, the final predicted value is calculated according to equation (6).
[0085] (6), in, For a single predicted value block, and These are the intra-frame prediction values obtained through planar mode. In this special case, the weights are set to... and .
[0086] A CU-level flag can be added to the bitstream to signal whether the intraTMP CU is predicted using the proposed intraTMP fusion method or the original intraTMP method.
[0087] For small blocks, the search range may be too small, making it unlikely to find a good match. To address this, an expanded search range is proposed, for example, `searchRangeWidth = max(a*BlkW, minSearchRange)` and `searchRangeHight = max(a*BlkH, minSearchRange)`, where `minSearchRange` is set to 128.
[0088] Additionally, in the original intraTMP, regions to the left and top of the current block are not searched when they are close to the current block. For example... Figure 9B As shown, a search is proposed for these regions.
[0089] Figure 10 Spatial components of an intraTMP filter 1000 according to some embodiments of the present disclosure are shown.
[0090] Reference Figure 10 The blocks selected via intraTMP can be further filtered to provide better predictions. To achieve this, a 6-tap filtering process is proposed, which includes a 5-tap plus sign spatial component and a bias term. The input to the spatial 5-tap component of the filter consists of a center (C) sample in the reference block and its above / north (N), below / south (S), left / west (W), and right / east (E) neighborhoods, where the center sample is located at the corresponding position of the sample in the current block to be predicted.
[0091] Figure 11 Reference region 1100 for deriving filter coefficients for intraTMP according to some embodiments of the present disclosure is shown.
[0092] Reference Figure 11 The bias term B represents the scalar offset between the input and output and is set to a medium brightness value (512 for 10-bit content). The output of the filter is calculated according to equation (7).
[0093] predLumaVal = c0C+c1N+c2S+c3E+c4W+c5B (7).
[0094] like Figure 11 As shown, the filter coefficients ci can be calculated by minimizing the mean-square error (MSE) between the filtered reference template and the current template 1. The extended region shown in dark gray supports "side samples" of the plus-shaped spatial filter. When unavailable, pixels in the extended region can be obtained by boundary padding with or without utilizing adjacent BV information.
[0095] MSE minimization is performed by calculating the autocorrelation matrix between the reference template input and the current template output. The autocorrelation matrix is decomposed using LDL (where L is the unit lower triangular matrix and D is the diagonal matrix), and the final filter coefficients are calculated using back-substitution.
[0096] The use of the filtered intraTMP mode is signaled via an encoded CU-level flag. Filtered intraTMP is considered a sub-mode of intraTMP. That is, the intraTMP filtering flag is signaled only when the intraTMP flag is true.
[0097] Furthermore, the filtering intraTMP can be selected from a list of candidate reference blocks to be filtered. The candidate list is constructed based on the lowest SAD cost of the unfiltered template. Different filter parameters are computed for each candidate in the list, and the candidate that performs best in terms of template matching cost with filtering is selected and used as the final candidate. The final prediction of the block is then generated by applying the derived filter to the best candidate.
[0098] Figure 12 A diagram showing adjacent half-pixel positions 1200 in eight directions according to some embodiments of the present disclosure is provided.
[0099] Reference Figure 12 A method with half-pixel precision, the intraTMP method, is proposed to provide better predictions. More specifically, such as... Figure 12 As shown, eight adjacent half-pixel locations are added in eight directions around the integer pixel location, and the proposed method selects one of nine locations (one integer pixel location + eight half-pixel locations) through encoder rate distortion optimization (RDO).
[0100] If the intraTMP mode is selected for the current block, a further signaling flag is provided to indicate whether integer or half-pixel precision is used. When half-pixel precision is used, a further signaling index is provided to indicate the direction of the half-pixel position. A 4-tap DCT-IF interpolation filter, [-5, 37, 37, -5], is used for half-pixel interpolation in fractional intraTMP.
[0101] Inspired by the "Combined Inter- and Intra-Prediction (CIIP)" mode, a spatially combined inter- and intra-prediction (CIIP) mode is proposed as a novel intra-prediction mode. When the spatial CIIP mode is selected, a prediction block is generated by combining intraTMP predictions with intra-predictions generated using template-based intra-mode derivation (TIMD). This combination is weighted using predefined weights. Spatial CIIP can be considered a special case of intraTMP fusion.
[0102] Figure 13 A diagram showing various template shapes for intraTMP according to some embodiments of the present disclosure is provided.
[0103] Reference Figure 13Two additional template shapes (e.g., left and top templates) are proposed for intraTMP. The left and top templates are considered as two additional intraTMP patterns. In addition, according to Equation (8), the two best candidates found using the L-shaped template are saved and linearly fused together as follows to generate a fused prediction.
[0104] (8), in, For fusion prediction, and The best and second-best L-shaped candidate predicted values are given, and and The template cost is the cost of the two L-shaped candidates obtained during the template matching process.
[0105] To signal the new mode, if intraTMP is used in the current CU, two flags are further signaled to indicate whether the L-shaped template with fusion, the left template, or the top template is applied, as detailed in Table 2 below.
[0106] Table 2: IntraTMP Mode Signaling While each of the aforementioned intraTMP methods provides coding gain individually, their combination is not straightforward. Unrestricted simultaneous enabling and indication of these intraTMP methods may result in excessive signaling overhead without providing sufficient predictive improvement to justify the bit consumption in signaling. To overcome these and other challenges, this disclosure describes a combined scheme that integrates efficient signaling techniques, providing cumulative coding gain from the combined intraTMP methods.
[0107] For example, in some implementations, this disclosure proposes that signaling notifications to intraTMP encoding tools can be delivered using the following summarized syntax elements: intra_tmp_flag if (intra_tmp_flag) { intra_tmp_fusion_flag if (intra_tmp_fusion_flag) { intra_tmp_fusion_idx intra_tmp_fusion_weight_type } else{ intra_tmp_idx intra_tmp_filter_flag if (!intra_tmp_filter_flag) { intra_tmp_sub_pel_flag if (intra_tmp_sub_pel_flag) { intra_tmp_sub_pel_direction_idx intra_tmp_sub_pel_phase_idx } } } } First, the intraTMP flag (intra_tmp_flag) can be signaled to indicate whether the current block is predicted via intraTMP. This syntax element can be decoded from the bitstream, or it can have an inferred value. For example, if the intraTMP tool is disabled at a higher syntax level (e.g., in the sequence parameter set (SPS)), the intraTMP flag can have an inferred value of 0.
[0108] If the current block is predicted by intraTMP, a signaling message can be sent to the intraTMP fusion flag (intra_tmp_fusion_flag) to indicate whether the intraTMP prediction value is determined by fusing multiple reference blocks. Regardless of the value of the intraTMP fusion flag, sparse search rounds and refined search rounds can be performed to construct a candidate list containing N intraTMP block vectors, where the N intraTMP block vectors are selected by choosing block vectors based on the SAD cost computed on the template region. Additional details of the candidate list construction process are provided below. In this scheme, N is greater than or equal to 15. Block vectors from the intraTMP candidate list are marked as... The reference block corresponding to the block vector from the intraTMP candidate list is marked as... .
[0109] If the current block is predicted using the intraTMP fusion mode (intra_tmp_fusion_flag is 1), then the signaling in the bitstream will notify both the intraTMP fusion index (intra_tmp_fusion_idx) and the intraTMP fusion weight type (intra_tmp_fusion_weight_type). The fusion prediction value is determined according to equation (9).
[0110] (9), The reference blocks selected for fusion are determined using the intraTMP fusion index. The intraTMP fusion index can take the values 0, 1, or 2. When intra_tmp_fusion_idx is 0, then A = 0 and B = 4. That is, the five reference blocks corresponding to the top five intraTMP block vectors in the candidate list are selected for fusion. When intra_tmp_fusion_idx is 1, then A = 5 and B = 9. When intra_tmp_fusion_idx is 2, then A = 10 and B = 14. midVal Set to the median value of the video sample. For example, if the bit depth is B, then midVal = 2 B-1 Example binarizations of intra_tmp_fusion_idx are shown in Table 3 below.
[0111] Table 3: Binarization of intra_tmp_fusion_idx The method for calculating the fusion weights is selected by the intraTMP fusion weight type (intra_tmp_fusion_weight_type), which is a flag that can take the value 0 or 1. For example, when intra_tmp_fusion_weight_type is 0, the weights can be determined by an algorithm based on SAD (see equations (10) and (11) below), where, Let represent the SAD cost of the i-th intraTMP candidate.
[0112] (10); and (11) When intra_tmp_fusion_weight_type is 1, the set of 6 weights can be determined by minimizing the MSE between the fusion of the adjacent template regions of the 5 reference blocks and the adjacent template regions of the current block. That is, if each reference block The adjacent template is And the adjacent template of the current block is Then the weights can be determined according to expression (12).
[0113] (12) As mentioned above, one way to solve this minimization problem is through LDL decomposition.
[0114] If the current block is predicted via intraTMP but not via intraTMP fusion (intra_tmp_fusion_flag is 0), signaling is used to instruct the intraTMP index to identify the individual block vector from the candidate list. The same candidate list is constructed regardless of whether intraTMP fusion is used. However, the candidate list construction is modified compared to the known methods described above to combine two aspects: searching multiple candidates through sparse and refined rounds, and searching for candidates using different template shapes.
[0115] In the sparse search, candidate block vectors are searched in parallel, which minimizes the SAD cost computed using L-shaped templates, left-only templates, and top-only templates. While the same block vector can be chosen, the SAD costs corresponding to templates of different shapes cannot be compared. If M candidates are maintained for each sparse search, a total of 3xM candidate BVs are maintained after the sparse search, where M candidate BVs exist for each template shape type. In the refinement search, a 3x3 local search is performed around each of the sparse candidate BVs, using the corresponding template shape. In some implementations, the search algorithm can be significantly optimized because the left-only template cost and the top-only template cost can be computed concurrently with the L-shaped template cost.
[0116] In one arrangement, the candidate list length can be increased to 19 to accommodate more types of block vectors. Up to 19 block vectors with the lowest SAD cost on the L-shaped template are selected, ordered in ascending order of SAD cost: first, the block vector with the lowest SAD cost, and the remaining block vectors are ordered in ascending order of their associated SAD cost. Similarly, up to 3 block vectors with the lowest SAD cost on only the left template are selected and ordered in ascending order, and up to 3 block vectors with the lowest SAD cost on only the left template are selected and ordered in ascending order of SAD cost. With all block vectors unique, the candidate list is constructed using 13 L-shaped template candidates, 3 left-only template candidates, and 3 top-only candidates. The L-shaped template candidates are filled with the lowest index. This is attributed to the binarization scheme used for signaling notification of the intraTMP index. This means that less overhead is used to signal these candidates. Using left-only or top-only templates to search for selected block vectors generally has a lower SAD cost because they have smaller template areas; however, L-shaped template candidate block vectors are still preferred because L-shaped templates are more likely to find good predictions for the current block. When any BV found through searches of different template shapes is identical, redundant BVs are removed from the top-only or left-only candidates. For example, if only one left-only template candidate is unique, and only two top-only template candidates are unique, the final candidate list will be constructed from 16 L-shaped template candidates, one left-only template candidate, and two top-only template candidates.
[0117] In another arrangement, the candidate list is 19 in length. At most 19 BVs with the lowest SAD cost on the L-shaped template are selected, along with at most 2 BVs with the lowest SAD cost on the left-only template and at most 2 BVs with the lowest SAD cost on the left-only template. For example, if all BVs are unique, the candidate list is constructed using 15 L-shaped template candidates, 2 left-only template candidates, and 2 top-only template candidates.
[0118] In one arrangement, the candidate list is constructed using the following candidate BVs: first, candidate BVs from an L-shaped template search, followed by candidates from other template shapes in a fixed order, such as candidates from a top-only template search, and then candidates from a left-only template search.
[0119] In another arrangement, the candidate list is constructed using the following candidate block vectors: first, candidate block vectors from the L-shaped template search; then, candidates from the template shape with the next largest area; and finally, candidates from the template shape with the smallest area. For example, if the area of only the top template is greater than the area of only the left template, the candidate list is constructed using the following candidate block vectors: first, candidate block vectors from the L-shaped template search; then, candidates from the top template search; and finally, candidates from the left template search. The template area depends on the height (h) and width (w) of the current block. For example, if h > w, then the area of only the left template is greater than the area of only the top template.
[0120] When the candidate list length is 19, the value range of the intra_tmp_idx syntax element can be from 0 to 18, and its binarization is shown in Table 4, where x represents 0 or 1.
[0121] Table 4: Binarization of intra_tmp_idx when the candidate list length is 19 In other words, when the value of intra_tmp_idx is in the range of 3 to 18, it is represented in the bitstream as follows: first 0, followed by a 4-bit fixed-length code (intra_tmp_idx - 2).
[0122] If the current block is predicted via intraTMP but not via intraTMP fusion (intra_tmp_fusion_flag is 0), then in addition to the intraTMP index, the intraTMP filter flag (intra_tmp_filter_flag) is signaled to indicate whether the selected reference block is fused. Filtering is performed to determine the prediction block. When intra_tmp_filter_flag is 1, according to equation (13), the learned filter is used... Filtering is performed to determine the prediction block. The learned filter in equation (13) It can be determined by equation (14).
[0123] (13); and (14) in, This represents the convolution operator. As shown in expression (15), it is a set of 6 coefficients. This is determined by minimizing the MSE between the filtered adjacent template regions of the selected reference block and the adjacent template regions of the current block.
[0124] (15) If the current block is predicted using intraTMP but not through intraTMP fusion and without intraTMP filtering (intra_tmp_filter_flag is 0), signaling can be used to notify the intraTMP sub-pixel flag (intra_tmp_sub_pel_flag) to indicate whether to further refine the intraTMP BV using fractional precision. If intra_tmp_sub_pel_flag is 0, the intraTMP block vector is not further modified. The predicted block is then the selected reference block. .
[0125] If intra_tmp_sub_pel_flag is 1, then the sub-pixel thinning direction (intra_tmp_sub_pel_direction_idx) and sub-pixel thinning phase (intra_tmp_sub_pel_phase_idx) are notified via signaling to further refine the intraTMP BV.
[0126] Figure 14 A diagram showing fractional block vector positions 1400 for intraTMP according to some embodiments of the present disclosure is illustrated.
[0127] Reference Figure 14 The intraTMP block vector selected by indexing to the candidate list has integer precision. Figure 14 In the diagram, the black circles represent the integer-interval coordinate space of the block vectors, with the middle circle representing the selected intraTMP block vector. The white circle represents 24 sub-pixel thinning block vectors, which can be signaled through the signaling mechanism of this scheme.
[0128] Then, as Figure 14 As indicated by the middle arrow, subpixel thinning can be performed along one of eight directions. `intra_tmp_sub_pel_direction_idx` signals the subpixel thinning direction by taking a value between 0 and 7 (represented by a 3-bit fixed-length code). The subpixel thinning block vector distance is determined by `intra_tmp_sub_pel_phase_idx`. The distance is used for signaling notification. `intra_tmp_sub_pel_phase_idx` indicates a value of 0 to indicate 1 / 4 phase, a value of 1 to indicate 1 / 2 phase, and a value of 2 to indicate 3 / 4 phase. The binarization of `intra_tmp_sub_pel_phase_idx` can be indicated according to Table 5.
[0129] Table 5: Binarization of intra_tmp_sub_pel_phase_idx The subpixel-thinned prediction block is determined by applying a 1D interpolation filter in a separable manner, where the desired interpolation is achieved using 1 / 4, 1 / 2, or 3 / 4 phase interpolation filters. The interpolation filter can reuse existing interpolation filters used for motion compensation in the ECM, or existing ECM filters used for intra-frame reference sample interpolation, or it can be an interpolation filter specifically designed for intraTMP.
[0130] The following syntax elements can be signaled in different orders, and the signaling notifications of these syntax elements are independent of each other. For example, the signaling notification order of intra_tmp_fusion_idx and intra_tmp_fusion_weight_type can be reversed, or the signaling notification order of intra_tmp_idx and intra_tmp_filter_flag can be reversed, or the signaling notification order of intra_tmp_sub_pel_phase_idx and intra_tmp_sub_pel_direction_idx can be reversed, without changing the encoding efficiency of the intraTMP encoding tool. For example, in some implementations, the syntax signaling notification order can be summarized as follows: intra_tmp_flag if (intra_tmp_flag) { intra_tmp_fusion_flag if (intra_tmp_fusion_flag) { intra_tmp_fusion_idx intra_tmp_fusion_weight_type } else { intra_tmp_idx intra_tmp_filter_flag if (!intra_tmp_filter_flag) { intra_tmp_sub_pel_flag if (intra_tmp_sub_pel_flag) { intra_tmp_sub_pel_phase_idx intra_tmp_sub_pel_direction_idx } } } } In some implementations, the intra_tmp_sub_pel_flag and intra_tmp_sub_pel_phase_idx syntax elements can be merged into a single syntax element (intra_tmp_sub_pel_precision_idx), which signals the subpixel to refine the BV distance. The distance. If intra_tmp_sub_pel_precision_idx is 0, then the phase is 0, indicating that the original block vector is used. In this case, subpixel block vector thinning is not used, therefore no signaling is given regarding the subpixel thinning direction. If intra_tmp_sub_pel_precision_idx is 1, the phase is 1 / 4. If intra_tmp_sub_pel_precision_idx is 2, the phase is 1 / 2. If intra_tmp_sub_pel_precision_idx is 3, the phase is 3 / 4. The binarization of intra_tmp_sub_pel_precision_idx can be the same as concatenating the syntax elements intra_tmp_filter_flag and intra_tmp_sub_pel_phase_idx. This binarization can be shown in Table 6.
[0131] Table 6: Binarization of intra_tmp_sub_pel_precision_idx Accordingly, the syntactic element dependencies and signaling notification order can be summarized as follows: intra_tmp_flag if (intra_tmp_flag) { intra_tmp_fusion_flag if (intra_tmp_fusion_flag) { intra_tmp_fusion_idx intra_tmp_fusion_weight_type } else { intra_tmp_idx intra_tmp_filter_flag if (!intra_tmp_filter_flag) { intra_tmp_sub_pel_precision_idx if (intra_tmp_sub_pel_precision_idx != 0) { intra_tmp_sub_pel_direction_idx } } } } Reference Figure 3 Transform module 308 can transform the video signal in the residual block from the pixel domain to the transform domain (e.g., to the frequency domain depending on the transform method). It should be understood that in some examples, transform module 308 can be skipped, and the video signal may not be transformed to the transform domain.
[0132] The quantization module 310 can be configured to quantize the coefficients at each position in the coded block to generate quantization levels for those positions. The current block can be a residual block. That is, the quantization module 310 can perform the quantization process on each residual block. The residual block can include... N × M Each location (sample) is associated with a transformed or untransformed video signal / data, such as luminance and / or chrominance information. N and M The value is a positive integer. In this disclosure, the transformed or untransformed video signal at a specific location before quantization is referred to herein as a "coefficient". After quantization, the quantized value of the coefficient is referred to herein as a "quantization level" or "level".
[0133] Quantization can be used to reduce the dynamic range of a transformed or untransformed video signal, allowing fewer bits to be used to represent the signal. Quantization typically involves dividing by the quantization step size and then rounding, while inverse quantization (also known as dequantization) involves multiplying by the quantization step size. The quantization step size can be indicated by the quantization parameter (QP). This quantization process is called scalar quantization. Quantization of all coefficients within a coded block can be performed independently, and this method is used in some existing video compression standards such as H.264 / AVC and H.265 / HEVC. The QP in quantization can affect the bitrate used to encode / decode the video image. For example, a higher QP results in a lower bitrate, and vice versa.
[0134] for N × MA coding block can be converted from two-dimensional (2D) coefficients to a one-dimensional (1D) sequence using a specific coding scan order for coefficient quantization and encoding. Typically, the coding scan begins at the top left corner and ends at the bottom right corner or the last non-zero coefficient / level in the bottom-right direction of the coding block. It should be understood that the coding scan order can include any suitable order, such as a zigzag scan order, a vertical (column) scan order, a horizontal (row) scan order, a diagonal scan order, or any combination thereof. The quantization of coefficients within a coding block can utilize coding scan order information. For example, it can depend on the state of previous quantization levels along the coding scan order. To further improve coding efficiency, the quantization module 310 can use more than one quantizer, such as two scalar quantizers. Which quantizer is used to quantize the current coefficient can depend on information preceding the current coefficient along the coding scan order. Such a quantization process is called dependent quantization.
[0135] Reference Figure 3 The encoding module 320 can be configured to encode the quantization level at each position in the coding block into a bitstream. In some embodiments, the encoding module 320 can perform entropy coding on the coding block. Entropy coding can use various binarization methods, such as Golomb-Rice binarization, to convert each quantization level into a corresponding binary representation, such as binary bits. The binary representation can then be further compressed using an entropy coding algorithm. The compressed data can be added to the bitstream. In addition to quantization levels, the encoding module 320 can also encode various other information, such as filter information input from prediction modules 304 and 306, block type information of coding units, prediction mode information, segmentation unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, and block interpolation information. In some embodiments, the encoding module 320 can perform residual coding on the coding block to convert the quantization level into a bitstream. For example, after quantization, for N × M Blocks can exist. N × M Each quantification level. These N × M Each level can be zero or non-zero. If a non-zero level is not binary, it can be further binaryized to binary bits, for example, using combined truncated Rice (TR) and restricted EGk binarization.
[0136] Non-binary syntax elements can be mapped to binary codewords. The bijective mapping between codewords and symbols, typically used in simple structured code, is also known as binarization. The binary symbols (also called bits) of both binary syntax elements and non-binary data codewords can be encoded using binary arithmetic coding. The core coding engine of Context-Adaptive Binary Arithmetic Coding (CABAC) supports two operating modes: context coding mode, where bits are encoded using an adaptive probability model; and a low-complexity bypass mode, which uses a fixed probability of 1 / 2. The adaptive probability model is also called the context, and assigning the probability model to each bit is also called context modeling.
[0137] like Figure 3 As shown, the dequantization module 312 can be configured to dequantize the quantization level, and the inverse transform module 314 can be configured to perform an inverse transform on the coefficients transformed by the transform module 308. The reconstructed residual block generated by the dequantization module 312 and the inverse transform module 314 can be combined with the prediction unit predicted by the prediction module 304 or 306 to generate a reconstructed block.
[0138] The filtering module 316 may include at least one of a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF). The deblocking filter can remove block distortion caused by boundaries between blocks in the reconstructed image. For a video that has been deblocked, the SAO module can correct the offset relative to the original video on a pixel-by-pixel basis. The ALF can be performed based on values obtained by comparing the reconstructed and filtered video with the original video. The caching module 318 can be configured to store the reconstructed blocks or images calculated by the filtering module 316, and the reconstructed and stored blocks or images can be provided to the inter-frame prediction module 304 when performing inter-frame prediction.
[0139] Figure 4 Some embodiments according to this disclosure are shown. Figure 2 A detailed block diagram of an exemplary decoder 201 in the decoding system 200. (See attached diagram.) Figure 4 As shown, decoder 201 may include decoding module 402, inverse quantization module 404, inverse transform module 406, inter-frame prediction module 408, intra-frame prediction module 410, filtering module 412, and buffer module 414. It should be understood that... Figure 4Each element shown is presented independently in the video decoder to represent a unique function distinct from each other, and this does not mean that each component consists of a separate hardware or software configuration unit. That is, for ease of illustration, each element is included as listed as an element, but at least two elements may be combined to form a single element, or an element may be divided into multiple elements for performing a function. It is also understood that some elements are not essential for performing the functions described in this disclosure, but may be optional elements used to improve performance. It should also be understood that these elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the decoder 201.
[0140] When a video stream is input from a video encoder (e.g., encoder 101), the input stream can be decoded by decoder 201 in the reverse order of the video encoder's process. Therefore, for ease of description, some decoding details described above regarding encoding can be omitted. Decoding module 402 can be configured to decode the stream to obtain various information encoded into it, such as the quantization level at each location in the encoded block. In some embodiments, decoding module 402 can perform entropy decoding (decompression) corresponding to the entropy coding (compression) performed by the encoder, such as VideoLAN coding (VLC), context-adaptive variable-length coding (CAVLC), CABAC, syntax-based binary arithmetic coding (SBAC), PIPE coding, etc., to obtain a binary representation (e.g., binary bits). Decoding module 402 can also convert the binary representation to quantization levels using Golomb-Rice binarization (including, for example, EGk binarization and combined TR and restricted EGk binarization). In addition to the quantization level of the position in the transform unit, the decoding module 402 can also decode various other information, such as parameters used for Golomb-Rice binarization (e.g., Rice parameters), block type information of the coding unit, prediction mode information, segmentation unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. During the decoding process, the decoding module 402 can perform rearrangement on the bitstream to reconstruct and rearrange the data from 1D order into 2D rearranged blocks using an inverse scanning method based on the coding scan order used by the encoder.
[0141] The dequantization module 404 can be configured to dequantize the quantization level at each location of the encoded block (e.g., a 2D reconstructed block) to obtain the coefficients at each location. In some embodiments, the dequantization module 404 can also perform dependent dequantization based on quantization parameters provided by the encoder, which include information related to the quantizers used in dependent quantization, such as the quantization step size used by each quantizer.
[0142] The inverse transform module 406 can be configured to perform inverse transforms on the DCT, DST, KLT, LFNST, and / or NSPT performed by the encoder, such as inverse DCT, inverse discrete sine transform (DST), and inverse KLT, to convert data from the transform domain (e.g., coefficients) back to the pixel domain (e.g., luminance and / or chrominance information). In some embodiments, the inverse transform module 406 can selectively perform transform operations (e.g., DCT, DST, KLT, LFNST, NSPT) based on various information such as the prediction method, the size of the current block, and the prediction direction.
[0143] Additionally and / or alternatively, the inter-frame prediction module 408 and the intra-frame prediction module 410 can be configured to generate prediction blocks based on information related to the generation of prediction blocks provided by the decoding module 402 and information about previously decoded blocks or images provided by the buffer module 414. As described above, if intra-frame prediction is performed in the same manner as the encoder operation, and the size of the prediction unit is the same as the size of the transform unit, intra-frame prediction can be performed on the prediction unit based on the pixels present to the left, the pixels present to the upper left, and the pixels present at the top of the prediction unit. However, if the size of the prediction unit and the size of the transform unit are different when performing intra-frame prediction, intra-frame prediction can be performed using reference pixels based on the transform unit.
[0144] For example, inter-frame prediction module 408 can be configured to receive a bitstream from an encoder, the bitstream including a reference frame, a current frame, and an indication of weighting factors associated with a multimedia home platform (MHP) process. Inter-frame prediction module 408 can be configured to perform an MHP process on CUs located in the current frame based on a search block (e.g., the reference frame and / or a reference template) in the reference frame. In some embodiments, to perform an MHP process, inter-frame prediction module 408 can be configured to perform template matching on CUs located in the current frame based on the search block and weighting factors in the reference frame to obtain motion information. In some embodiments, to perform an MHP process, inter-frame prediction module 408 can be configured to identify a weighting factor index associated with the weighting factors based on template matching. Inter-frame prediction module 408 can be configured to identify weighting factor symbols for the weighting factors based on indications included in the bitstream. Inter-frame prediction module 408 can be configured to perform an inter-frame prediction process based on the current frame, the reference frame, the weighting factor index, and the weighting factor symbols for the weighting factors to decode the bitstream.
[0145] The reconstructed blocks or reconstructed image, a combination of the outputs of the inverse transform module 406 and the prediction modules 408 or 410, can be provided to the filtering module 412. The filtering module 412 may include a deblocking filter, an offset correction module, and an ALF. The buffer module 414 may store the reconstructed image or reconstructed blocks and use them as reference images or reference blocks for the inter-frame prediction module 408, and may also output the reconstructed image.
[0146] Within the scope of this disclosure, the encoding module 320 and the decoding module 402 can be configured to employ a quantization level binarization scheme having a Rice parameter with a bit depth and / or bit rate suitable for encoding images of video, in order to improve encoding efficiency.
[0147] Figures 15A to 15D A flowchart of an exemplary method 1500 for video decoding according to some embodiments of the present disclosure is shown. Method 1500 can be performed by a system, such as, to name a few examples, a decoding system 200, a decoder 201, or an intra-frame prediction module 410. Method 1500 may include operations 1502 to 1552 as described below. It should be understood that some steps may be optional, some steps may be performed simultaneously, or in accordance with... Figures 15A to 15D The different sequences shown are executed in order.
[0148] Reference Figure 15A At 1502, the system can parse the bitstream to decode multiple syntax elements associated with intraTMP.
[0149] At 1504, the system can decode the first syntax element from the bitstream. For example, the intraTMP flag (intra_tmp_flag) can be signaled to indicate whether the current block is predicted via intraTMP. This syntax element can be decoded from the bitstream, or it can have an inferred value. For example, if the intraTMP tool is disabled at a higher syntax level (e.g., in the Sequence Parameter Set (SPS)), the intraTMP flag can have an inferred value of 0.
[0150] At 1506, the system can determine whether intraTMP mode is enabled for the current block based on the first syntax element. For example, the intraTMP flag (intra_tmp_flag) can be signaled to indicate whether the current block is predicted via intraTMP. This syntax element can be decoded from the bitstream, or it can have an inferred value. For example, if the intraTMP tool is disabled at a higher syntax level (e.g., in the Sequence Parameter Set (SPS)), the intraTMP flag can have an inferred value of 0.
[0151] At 1508, in response to enabling intraTMP mode for the current block, the system can decode the second syntax element from the bitstream. For example, if the current block is predicted via intraTMP, the intraTMP fusion flag (intra_tmp_fusion_flag) can be signaled to indicate whether to determine the intraTMP prediction value by fusing multiple reference blocks.
[0152] At 1510, the system can determine whether to use the intraTMP fusion mode to determine the intraTMP prediction value for the current block based on the second syntax element. For example, if the current block is predicted using the intraTMP fusion mode, the value of intra_tmp_fusion_flag can be represented as 1; otherwise, the value can be represented as 0.
[0153] At 1512, in response to determining the intraTMP prediction value via the intraTMP fusion mode, the system can decode the third and fourth syntax elements from the bitstream. For example, if the current block is predicted via the intraTMP fusion mode (intra_tmp_fusion_flag is 1), the intraTMP fusion index (intra_tmp_fusion_idx) and the intraTMP fusion weight type (intra_tmp_fusion_weight_type) are signaled in the bitstream.
[0154] At 1514, the system can determine the set of fusion weights based on the fourth syntax element. For example, the method for calculating the fusion weights is selected by the intraTMP fusion weight type (intra_tmp_fusion_weight_type), which is a flag that can take the value 0 or 1. For example, when intra_tmp_fusion_weight_type is 0, the weights can be determined by an algorithm based on SAD (see equations (10) and (11) below), where, This represents the SAD cost of the i-th intraTMP candidate. When intra_tmp_fusion_weight_type is 1, the set of 6 weights can be determined by minimizing the MSE between the fusion of the adjacent template regions of the 5 reference blocks and the adjacent template regions of the current block. That is, if each reference block The adjacent template is And the adjacent template of the current block is Then the weights can be determined according to expression (12).
[0155] Reference Figure 15B At point 1516, the system can determine the intraTMP prediction value based on the fusion weight set and the reference block set indicated by the third syntax element. For example, the intraTMP prediction value can be determined based on the fusion weights and reference blocks calculated by an algorithm based on SAD or an algorithm based on MSE.
[0156] At point 1518, in response to determining the intraTMP prediction value of the current block through the intraTMP fusion mode, the system can perform a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors. For example, if the current block is predicted via intraTMP, the intraTMP fusion flag (intra_tmp_fusion_flag) can be signaled to indicate whether the intraTMP prediction value is determined by fusing multiple reference blocks. Regardless of the value of the intraTMP fusion flag, sparse search rounds and refinement search rounds can be performed to construct a candidate list containing N intraTMP block vectors, where the N intraTMP block vectors are selected by picking block vectors based on the SAD cost computed on the template region. Further details of the candidate list construction process are provided below. In this scheme, N is greater than or equal to 15. The block vectors from the intraTMP candidate list are marked as... The reference block corresponding to the block vector from the intraTMP candidate list is marked as... .
[0157] At 1520, the system can determine the set of reference blocks selected for determining the intraTMP prediction value based on the third syntax element. For example, if the current block is predicted using the intraTMP fusion mode (intra_tmp_fusion_flag is 1), the intraTMP fusion index (intra_tmp_fusion_idx) and the intraTMP fusion weight type (intra_tmp_fusion_weight_type) are signaled in the bitstream. The fusion prediction value is determined according to Equation (9). In Equation (9), the reference blocks selected for fusion are determined by the intraTMP fusion index. The intraTMP fusion index can be 0, 1, or 2. When intra_tmp_fusion_idx is 0, then A = 0 and B = 4. That is, the five reference blocks corresponding to the top five intraTMP block vectors in the candidate list are selected for fusion ( When intra_tmp_fusion_idx is 1, then A = 5 and B = 9. When intra_tmp_fusion_idx is 2, then A = 10 and B = 14. Set to the median value of the video sample. For example, if the bit depth is B, then Example binarizations of intra_tmp_fusion_idx are shown in Table 3 below.
[0158] At position 1522, in response to the fourth syntax element including the first value, the system can determine the set of fusion weights using a SAD-based algorithm. For example, when intra_tmp_fusion_weight_type is 0, the weights can be determined by a SAD-based algorithm (see equations (10) and (11) above), where, Let represent the SAD cost of the i-th intraTMP candidate.
[0159] At 1524, in response to the inclusion of the second value in the fourth syntax element, the system can determine the set of fusion weights using an MSE-based algorithm. For example, when intra_tmp_fusion_weight_type is 1, the set of six weights can be determined instead by minimizing the MSE between the fusion of the adjacent template regions of the five reference blocks and the adjacent template regions of the current block. In other words, if each reference block The adjacent template is And the adjacent template of the current block is Then the weights can be determined according to expression (12).
[0160] At 1526, in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block, the system can decode the fifth syntax element from the bitstream. For example, if the current block is predicted via intraTMP but not via intraTMP fusion (intra_tmp_fusion_flag is 0), the signaling notifies the intraTMP index to identify the individual block vector from the candidate list.
[0161] At point 1528, the system can identify block vectors from the candidate list based on the fifth syntax element. For example, if the current block is predicted via intraTMP but not via intraTMP fusion (intra_tmp_fusion_flag is 0), signaling is used to instruct the intraTMP index to identify a single block vector from the candidate list. The same candidate list is constructed regardless of whether intraTMP fusion is used. However, the candidate list construction is modified compared to the known methods described above to combine two aspects: searching multiple candidates through sparse rounds and refined rounds, and searching for candidates using different template shapes.
[0162] Reference Figure 15C At 1530, in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block, the system can decode the sixth syntax element from the bitstream. For example, if the current block is predicted via intraTMP but not via intraTMP fusion (intra_tmp_fusion_flag is 0), then in addition to the intraTMP index, the intraTMP filter flag (intra_tmp_filter_flag) is signaled to indicate whether to pass the selected reference block. Filtering is performed to determine the prediction block.
[0163] At 1532, the system can determine whether to determine the intraTMP prediction value by filtering the selected reference block based on the sixth syntax element. For example, if the current block is predicted by intraTMP but not by intraTMP fusion (intra_tmp_fusion_flag is 0), then in addition to the intraTMP index, the intraTMP filter flag (intra_tmp_filter_flag) is signaled to indicate whether to filter the selected reference block. Filtering is performed to determine the prediction block.
[0164] At position 1534, in response to the sixth syntax element including the first value, the system can determine the intraTMP prediction value by filtering the selected reference block with a learned filter. For example, when intra_tmp_filter_flag is 1, according to equation (13), by using a learned filter... Filtering is performed to determine the prediction block. The learned filter in equation (13) It can be determined by equation (14).
[0165] At position 1536, in response to the sixth syntax element including the second value, the system can decode the seventh syntax element from the bitstream. For example, if the current block is predicted via intraTMP but not via intraTMP fusion and not via intraTMP filtering (intra_tmp_filter_flag is 0), signaling can be used to notify the intraTMP subpixel flag (intra_tmp_sub_pel_flag) to indicate whether to further refine the intraTMP BV using fractional precision.
[0166] At 1538, the system can determine whether to refine the block vector using fractional precision based on the seventh syntax element. For example, if the current block is predicted via intraTMP but not via intraTMP fusion and not via intraTMP filtering (intra_tmp_filter_flag is 0), the intraTMP subpixel flag (intra_tmp_sub_pel_flag) can be signaled to indicate whether to further refine the intraTMP BV using fractional precision.
[0167] At 1540, in response to the seventh syntax element including the first value, the system can determine through the processor whether to refine the block vector using fractional precision.
[0168] At position 1542, in response to the seventh syntax element including the second value, the system can determine whether to refine the block vector using fractional precision. For example, if `intra_tmp_sub_pel_flag` is 0, the `intraTMP` block vector is not further modified. The predicted block is then the selected reference block. .
[0169] Reference Figure 15DAt 1544, in response to the seventh syntax element including the second value, the system can decode the eighth and ninth syntax elements from the bitstream. For example, if intra_tmp_sub_pel_flag is 1, the intraTMP BV is further refined by signaling the subpixel thinning direction (intra_tmp_sub_pel_direction_idx) and the subpixel thinning phase (intra_tmp_sub_pel_phase_idx).
[0170] At position 1546, the system can determine the subpixel thinning direction based on the eighth syntax element. For example, subpixel thinning can proceed along... Figure 14 The thinning direction is selected from one of the eight directions indicated by the arrow. The intra_tmp_sub_pel_direction_idx indicates the subpixel thinning direction by taking a value in the range of 0 to 7 (represented by a 3-bit fixed-length code).
[0171] At position 1548, the system can determine the subpixel thinning phase based on the ninth syntax element. For example, the subpixel BV distance is thinned using intra_tmp_sub_pel_phase_idx. The distance is used for signaling notification. intra_tmp_sub_pel_phase_idx indicates a value of 0 to indicate 1 / 4 phase, a value of 1 to indicate 1 / 2 phase, and a value of 2 to indicate 3 / 4 phase.
[0172] At 1550, the system can refine the block vector based on the subpixel refinement direction and subpixel refinement phase. For example, the predicted subpixel refinement block can be determined by applying a one-dimensional interpolation filter in a separable manner, where the desired interpolation is achieved using 1 / 4, 1 / 2, or 3 / 4 phase interpolation filters. The interpolation filter can reuse existing interpolation filters used for motion compensation in the ECM, or existing ECM filters used for intra-frame reference sample interpolation, or it can be an interpolation filter specifically designed for intraTMP.
[0173] At position 1552, the system can decode the current block based on the intraTMP prediction value. For example, decoder 201 can decode the current block based on the intraTMP prediction value determined from the various syntax elements parsed from the bitstream as described above.
[0174] Figures 16A to 16DA flowchart of an exemplary method 1600 for video encoding according to some embodiments of the present disclosure is shown. Method 1600 may be performed by a system, such as encoding system 200, encoder 101, or intra-frame prediction module 410, among other examples. Method 1600 may include operations 1602 to 1652 as described below. It should be understood that some steps may be optional, some steps may be performed simultaneously, or the execution order may be different. Figures 16A to 16D The order shown is different.
[0175] Reference Figure 16A At position 1602, the system can determine that the current block is encoded using intraTMP. For example, encoder 101 can determine that the current block is encoded using intraTMP.
[0176] At 1604, the system can encode the first syntax element into the bitstream. For example, the intraTMP flag (intra_tmp_flag) can be signaled to indicate whether the current block is predicted via intraTMP. This syntax element can be encoded into the bitstream, or it can have an inferred value. For example, if the intraTMP tool is disabled at a higher syntax level (e.g., in the Sequence Parameter Set (SPS)), the intraTMP flag can have an inferred value of 0.
[0177] At 1606, the system can determine whether intraTMP mode is enabled for the current block based on the first syntax element. For example, the intraTMP flag (intra_tmp_flag) can be signaled to indicate whether the current block is predicted via intraTMP. This syntax element can be encoded into the bitstream, or it can have an inferred value. For example, if the intraTMP tool is disabled at a higher syntax level (e.g., in the Sequence Parameter Set (SPS)), the intraTMP flag can have an inferred value of 0.
[0178] At 1608, in response to enabling intraTMP mode for the current block, the system can encode the second syntax element into the bitstream. For example, if the current block is predicted via intraTMP, the system can signal the intraTMP fusion flag (intra_tmp_fusion_flag) to indicate whether to determine the intraTMP prediction by fusing multiple reference blocks.
[0179] At 1610, the system can determine whether to use the intraTMP fusion mode to determine the intraTMP prediction value of the current block based on the second syntax element. For example, if the current block is predicted using the intraTMP fusion mode, the value of intra_tmp_fusion_flag can be represented as 1; otherwise, it can be represented as 0.
[0180] At 1612, in response to determining that the intraTMP prediction value is determined via the intraTMP fusion mode, the system can encode the third and fourth syntax elements into the bitstream. For example, if the current block is predicted via the intraTMP fusion mode (intra_tmp_fusion_flag is 1), then the signaling in the bitstream informs both the intraTMP fusion index (intra_tmp_fusion_idx) and the intraTMP fusion weight type (intra_tmp_fusion_weight_type).
[0181] At 1614, the system can determine the set of fusion weights based on the fourth syntax element. For example, the method for calculating the fusion weights is selected by the intraTMP fusion weight type (intra_tmp_fusion_weight_type), which is a flag that can be 0 or 1. For example, when intra_tmp_fusion_weight_type is 0, the weights can be determined by an algorithm based on SAD (see equations (10) and (11) below), where, Indicates the first i The SAD cost of the intraTMP candidates. When intra_tmp_fusion_weight_type is 1, the set of 6 weights can be determined by minimizing the MSE between the fusion of the adjacent template regions of the 5 reference blocks and the adjacent template regions of the current block. In other words, if each reference block The adjacent template is And the adjacent template of the current block is Then the weights can be determined according to expression (12).
[0182] Reference Figure 16B At point 1616, the system can determine the intraTMP prediction based on the fused weight set and the reference block set indicated by the third syntax element. For example, the intraTMP prediction can be determined based on the fused weights and reference blocks computed by an algorithm based on SAD or an algorithm based on MSE.
[0183] At point 1618, in response to determining the intraTMP prediction value of the current block determined by the intraTMP fusion mode, the system can perform a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors. For example, if the current block is predicted via intraTMP, the intraTMP fusion flag (intra_tmp_fusion_flag) can be signaled to indicate whether the intraTMP prediction value is determined by fusing multiple reference blocks. Regardless of the value of the intraTMP fusion flag, sparse search rounds and refinement search rounds can be performed to construct a candidate list containing N intraTMP block vectors, where the N intraTMP block vectors are selected by picking block vectors based on the SAD cost computed on the template region. Further details of the candidate list construction process are provided below. In this scheme, N is greater than or equal to 16. The block vectors from the intraTMP candidate list are marked as... The reference block corresponding to the block vector from the intraTMP candidate list is marked as... .
[0184] At 1620, the system can determine the set of reference blocks selected for determining the intraTMP prediction value based on the third syntax element. For example, if the current block is predicted using the intraTMP fusion mode (intra_tmp_fusion_flag is 1), then the intraTMP fusion index (intra_tmp_fusion_idx) and the intraTMP fusion weight type (intra_tmp_fusion_weight_type) are signaled in the bitstream. The fusion prediction value is determined according to Equation (9). In Equation (9), the reference blocks selected for fusion are determined by the intraTMP fusion index. The intraTMP fusion index can be 0, 1, or 2. When intra_tmp_fusion_idx is 0, then A = 0 and B = 4. That is, the five reference blocks corresponding to the first five intraTMP block vectors in the candidate list are selected for fusion ( When intra_tmp_fusion_idx is 1, then A = 5 and B = 9. When intra_tmp_fusion_idx is 2, then A = 10 and B = 14. Set to the median value of the video sample. For example, if the bit depth is B, then Example binarizations of intra_tmp_fusion_idx are shown in Table 3 below.
[0185] At position 1622, in response to the fourth syntax element including the first value, the system can determine the set of fusion weights using a SAD-based algorithm. For example, when intra_tmp_fusion_weight_type is 0, the weights can be determined using a SAD-based algorithm (see equations (10) and (11) above), where, Indicates the first i The SAD cost of each intraTMP candidate.
[0186] At 1624, in response to the inclusion of a second value in the fourth syntax element, the system can determine the set of fusion weights using an MSE-based algorithm. For example, when intra_tmp_fusion_weight_type is 1, the set of six weights can be determined instead by minimizing the MSE between the fusion of the adjacent template regions of the five reference blocks and the adjacent template regions of the current block. In other words, if each reference block The adjacent template is And the adjacent template of the current block is Then the weights can be determined according to expression (12).
[0187] At 1626, in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block, the system can encode the fifth syntax element into the bitstream. For example, if the current block is predicted via intraTMP but not via intraTMP fusion (intra_tmp_fusion_flag is 0), then signaling informs the intraTMP index to identify the individual block vector from the candidate list.
[0188] At point 1628, the system can identify block vectors from the candidate list based on the fifth syntax element. For example, if the current block is predicted via intraTMP but not via intraTMP fusion (intra_tmp_fusion_flag is 0), signaling is used to instruct the intraTMP index to identify a single block vector from the candidate list. The same candidate list is constructed regardless of whether intraTMP fusion is used. However, the candidate list construction is modified compared to the known methods described above to combine two aspects: searching multiple candidates through sparse rounds and refined rounds, and searching for candidates using different template shapes.
[0189] Reference Figure 16CAt 1630, in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block, the system can encode the sixth syntax element into the bitstream. For example, if the current block is predicted via intraTMP but not via intraTMP fusion (intra_tmp_fusion_flag is 0), then in addition to the intraTMP index, the intraTMP filter flag (intra_tmp_filter_flag) is signaled to indicate whether to pass the selected reference block. Filtering is performed to determine the prediction block.
[0190] At 1632, the system can determine whether to determine the intraTMP prediction value by filtering the selected reference block based on the sixth syntax element. For example, if the current block is predicted by intraTMP but not by intraTMP fusion (intra_tmp_fusion_flag is 0), then in addition to the intraTMP index, the intraTMP filter flag (intra_tmp_filter_flag) is signaled to indicate whether to filter the selected reference block. Filtering is performed to determine the prediction block.
[0191] At position 1634, in response to the sixth syntax element including the first value, the system can determine the intraTMP prediction value by filtering the selected reference block using a learned filter. For example, when intra_tmp_filter_flag is 1, according to equation (13), by using a learned filter... Filtering is performed to determine the prediction block. The learned filter c in equation (13) can be determined by equation (14).
[0192] At position 1636, in response to the sixth syntax element including the second value, the system can encode the seventh syntax element into the bitstream. For example, if the current block is predicted via intraTMP but not via intraTMP fusion and not via intraTMP filtering (intra_tmp_filter_flag is 0), signaling can be used to notify the intraTMP subpixel flag (intra_tmp_sub_pel_flag) to indicate whether to further refine the intraTMP BV using fractional precision.
[0193] At 1638, the system can determine whether to refine the block vector using fractional precision based on the seventh syntax element. For example, if the current block is predicted by intraTMP but not by intraTMP fusion and not by intraTMP filtering (intra_tmp_filter_flag is 0), the intraTMP subpixel flag (intra_tmp_sub_pel_flag) can be signaled to indicate whether to further refine the intraTMP BV using fractional precision.
[0194] At 1640, in response to the seventh syntax element including the first value, the system can determine through the processor whether to refine the block vector using fractional precision.
[0195] At 1642, in response to the seventh syntax element including the second value, the system can determine whether to refine the block vector using fractional precision. For example, if `intra_tmp_sub_pel_flag` is 0, the `intraTMP` block vector is not further modified. The predicted block is then the selected reference block. .
[0196] Reference Figure 16D At 1644, in response to the seventh syntax element including the second value, the system can encode the eighth and ninth syntax elements into the bitstream. For example, if intra_tmp_sub_pel_flag is 1, the intraTMP BV is further refined by signaling the subpixel thinning direction (intra_tmp_sub_pel_direction_idx) and the subpixel thinning phase (intra_tmp_sub_pel_phase_idx).
[0197] At position 1646, the system can determine the subpixel thinning direction based on the eighth syntax element. For example, subpixel thinning can be along... Figure 14 The thinning direction is selected from one of the eight directions indicated by the arrow. The intra_tmp_sub_pel_direction_idx indicates the subpixel thinning direction by taking a value in the range of 0 to 7 (represented by a 3-bit fixed-length code).
[0198] At position 1648, the system can determine the subpixel thinning phase based on the ninth syntax element. For example, the subpixel BV distance is thinned using intra_tmp_sub_pel_phase_idx. The distance is used for signaling notification. intra_tmp_sub_pel_phase_idx indicates a value of 0 to indicate 1 / 4 phase, a value of 1 to indicate 1 / 2 phase, and a value of 2 to indicate 3 / 4 phase.
[0199] At 1650, the system can refine the block vector based on the subpixel refinement direction and subpixel refinement phase. For example, the predicted subpixel refinement block can be determined by applying a one-dimensional interpolation filter in a separable manner, where the desired interpolation is achieved using 1 / 4, 1 / 2, or 3 / 4 phase interpolation filters. The interpolation filter can reuse existing interpolation filters used for motion compensation in the ECM, or existing ECM filters used for intra-frame reference sample interpolation, or it can be an interpolation filter specifically designed for intraTMP.
[0200] In 1652, the system can encode the current block based on the intraTMP prediction value. For example, encoder 101 can encode the current block based on the intraTMP prediction value determined based on the various syntax elements parsed into the bitstream as described above.
[0201] In all aspects of this disclosure, the functions described herein can be implemented by hardware, software, firmware, or any combination thereof. If implemented by software, the functions can be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. The storage medium can be a processor (e.g., Figure 1 and Figure 2 The processor 102 in the document can access any available medium. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, HDD, such as disk storage or other magnetic storage devices, flash drives, SSDs, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system (e.g., a mobile device or computer). As used herein, disks and optical discs include CDs, laser discs, optical discs, digital video discs (DVDs), and floppy disks, wherein disks typically reproduce data magnetically, while optical discs reproduce data optically by laser. The combinations described above should also be included within the scope of computer-readable media.
[0202] According to one aspect of this disclosure, a decoding method performed by a decoder is provided. The method may include: a processor parsing a bitstream to decode a plurality of syntax elements associated with intraTMP. The method may include: the processor decoding a first syntax element from the bitstream. The method may include: the processor determining, based on the first syntax element, whether to enable an intraTMP mode for a current block. The method may include: in response to enabling the intraTMP mode for the current block, the processor decoding a second syntax element from the bitstream. The method may include: the processor determining, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The method may include: in response to determining the intraTMP prediction value using an intraTMP fusion mode, the processor decoding a third and a fourth syntax element from the bitstream. The method may include: the processor determining a fusion weight set based on the fourth syntax element. The method may include: the processor determining the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. The method may include: the processor decoding the current block based on the intraTMP prediction value.
[0203] In some implementations, in response to determining the intraTMP prediction value of the current block through the intraTMP fusion mode, the method may include: the processor performing a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on the SAD cost computed on the template region.
[0204] In some implementations, the processor performs sparse search and refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions. This may include: the processor identifying a set of sparse candidate block vectors in parallel during the sparse search. In some implementations, the processor performs sparse search and refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions. This may include: the processor constructing the candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using a selected template shape.
[0205] In some implementations, the method may include: the processor determining, based on the third syntax element, a set of reference blocks for determining the intraTMP prediction value.
[0206] In some implementations, in response to the fourth syntax element including a first value, the method may include: the processor determining the fusion weight set using a SAD-based algorithm. In some implementations, in response to the fourth syntax element including a second value, the method may include: the processor determining the fusion weight set using an MSE-based algorithm.
[0207] In some implementations, in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block, the method may include: the processor decoding the fifth syntax element from the bitstream. In some implementations, the method may include: the processor identifying a block vector from the candidate list based on the fifth syntax element.
[0208] In some implementations, the value range of the fifth syntax element is 0 to 18. In some implementations, when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream as follows: 0 followed by intra_tmp_idx - 2, represented by a 4-bit fixed-length code.
[0209] In some implementations, in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block, the method may include: the processor decoding a sixth syntax element from the bitstream. In some implementations, the processor determines, based on the sixth syntax element, whether to determine the intraTMP prediction value by filtering a selected reference block.
[0210] In some implementations, in response to the sixth syntax element including a first value, the method may include: the processor determining the intraTMP prediction value by filtering a selected reference block using a learned filter.
[0211] In some implementations, in response to the sixth syntax element including a second value, the method may include: the processor decoding a seventh syntax element from the bitstream. In some implementations, the method may include: the processor determining, based on the seventh syntax element, whether to refine the block vector using fractional precision.
[0212] In some implementations, in response to the seventh syntax element including a first value, the method may include: the processor determining not to refine the block vector using fractional precision. In some implementations, in response to the seventh syntax element including a second value, the method may include: the processor determining to refine the block vector using fractional precision.
[0213] In some implementations, in response to the seventh syntax element including a second value, the method may include: the processor decoding an eighth and a ninth syntax element from the bitstream. In some implementations, the method may include: the processor determining a subpixel thinning direction based on the eighth syntax element. In some implementation examples, the method may include: the processor determining a subpixel thinning phase based on the ninth syntax element. In some implementations, the method may include: the processor thinning the block vector based on the subpixel thinning direction and the subpixel thinning phase.
[0214] According to another aspect of this disclosure, a decoder is provided. The decoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to: parse a bitstream to decode a plurality of syntax elements associated with intraTMP. The memory stores instructions that, when executed by the processor, cause the processor to: decode a first syntax element from the bitstream. The memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the first syntax element, whether to enable intraTMP mode for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: decode a second syntax element from the bitstream in response to enabling intraTMP mode for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Decode a third syntax element and a fourth syntax element from the bitstream in response to determining the intraTMP prediction value through an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Determine a fusion weight set based on the fourth syntax element. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Determine the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: Decode the current block based on the intraTMP prediction value.
[0215] In some implementations, the memory stores instructions that, when executed by the processor, may also cause the processor to perform the following operations in response to determining the intraTMP prediction value of the current block through the intraTMP fusion mode: performing a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on the SAD cost computed on the template region.
[0216] In some implementations, to perform sparse search and refinement search, and to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions, the memory stores instructions that, when executed by the processor, further cause the processor to perform the following operation: concurrently identify a set of sparse candidate block vectors during the sparse search. In some implementations, to perform sparse search and refinement search, and to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions, the memory stores instructions that, when executed by the processor, further cause the processor to perform the following operation: construct a candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using a selected template shape.
[0217] In some implementations, the memory stores instructions that, when executed by the processor, also cause the processor to perform the following operations: determine, based on the third syntax element, the set of reference blocks selected for determining the intraTMP prediction value.
[0218] In some implementations, the memory stores instructions that, when executed by the processor, can also cause the processor to perform the following operation: in response to the fourth syntax element including a first value, determine the fusion weight set using a SAD-based algorithm. In some implementations, the memory stores instructions that, when executed by the processor, can also cause the processor to perform the following operation: in response to the fourth syntax element including a second value, determine the fusion weight set using an MSE-based algorithm.
[0219] In some implementations, the memory stores instructions that, when executed by the processor, may also cause the processor to: decode a fifth syntax element from the bitstream in response to a second syntax element indicating that intraTMP fusion is not enabled for the current block. In some implementations, the memory stores instructions that, when executed by the processor, may also cause the processor to: identify a block vector from the candidate list based on the fifth syntax element.
[0220] In some implementations, the value range of the fifth syntax element is 0 to 18. In some implementations, when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream as follows: 0 followed by intra_tmp_idx - 2, represented by a 4-bit fixed-length code.
[0221] In some implementations, the memory stores instructions that, when executed by the processor, may also cause the processor to: decode a sixth syntax element from the bitstream in response to a second syntax element indicating that intraTMP fusion is not enabled for the current block. In some implementations, the memory stores instructions that, when executed by the processor, may also cause the processor to: determine, based on the sixth syntax element, whether to determine the intraTMP prediction value by filtering a selected reference block.
[0222] In some implementations, the memory stores instructions that, when executed by the processor, may also cause the processor to: determine the intraTMP prediction value by filtering a selected reference block using a learned filter in response to the sixth syntax element including a first value.
[0223] In some implementations, the memory stores instructions that, when executed by the processor, can also cause the processor to: decode a seventh syntax element from the bitstream in response to the sixth syntax element including a second value. In some implementations, the memory stores instructions that, when executed by the processor, can also cause the processor to: determine, based on the seventh syntax element, whether to refine the block vector using fractional precision.
[0224] In some implementations, the memory stores instructions that, when executed by the processor, can also cause the processor to: determine, in response to the seventh syntax element including a first value, not to refine the block vector using fractional precision. In some implementations, the memory stores instructions that, when executed by the processor, can also cause the processor to: determine, in response to the seventh syntax element including a second value, to refine the block vector using fractional precision.
[0225] In some implementations, the memory stores instructions that, when executed by the processor, can further cause the processor to perform the following operations: decode an eighth and a ninth syntax element from the bitstream in response to the seventh syntax element including a second value. In some implementations, the memory stores instructions that, when executed by the processor, can further cause the processor to perform the following operations: determine a sub-pixel thinning direction based on the eighth syntax element. In some implementations, the memory stores instructions that, when executed by the processor, can further cause the processor to perform the following operations: determine a sub-pixel thinning phase based on the ninth syntax element. In some implementations, the memory stores instructions that, when executed by the processor, can further cause the processor to perform the following operations: thin the block vector based on the sub-pixel thinning direction and the sub-pixel thinning phase.
[0226] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided, which stores instructions for a decoder. The memory stores instructions that, when executed by a processor of the decoder, cause the processor of the decoder to parse a bitstream to decode a plurality of syntax elements associated with intraTMP. The memory stores instructions that, when executed by the processor of the decoder, cause the processor of the decoder to decode a first syntax element from the bitstream. The memory stores instructions that, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operation: determine, based on the first syntax element, whether intraTMP mode is enabled for the current block. The memory stores instructions that, when executed by the processor, cause the processor to perform the following operation: decode a second syntax element from the bitstream in response to the intraTMP mode being enabled for the current block. The memory stores instructions that, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operation: determine, based on the second syntax element, whether to determine an intraTMP prediction value for the current block via an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: in response to determining the intraTMP prediction value through the intraTMP fusion mode, decode the third and fourth syntax elements from the bitstream. The memory storage instruction, when executed by the decoder's processor, causes the decoder's processor to perform the following operations: determine a fusion weight set based on the fourth syntax element. The memory storage instruction, when executed by the decoder's processor, causes the decoder's processor to perform the following operations: determine the intraTMP prediction value based on the fusion weight set and the reference block set indicated by the third syntax element. The memory storage instruction, when executed by the decoder's processor, causes the decoder's processor to perform the following operations: decode the current block based on the intraTMP prediction value.
[0227] In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to perform the following operations in response to determining the intraTMP prediction value of the current block through the intraTMP fusion mode: performing sparse search and refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on the SAD cost computed on the template region.
[0228] In some implementations, to perform sparse search and refinement search, and to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions, the memory stores instructions that, when executed by the processor, further cause the processor to perform the following operation: during the sparse search, identify a set of sparse candidate block vectors in parallel. In some implementations, to perform sparse search and refinement search, and to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions, the memory stores instructions that, when executed by the processor, further cause the processor to perform the following operation: construct a candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using a selected template shape.
[0229] In some implementations, the memory stores instructions that, when executed by the processor, also cause the processor to perform the following operations: determine, based on the third syntax element, a set of reference blocks selected for determining the intraTMP prediction value.
[0230] In some implementations, the memory stores instructions that, when executed by the decoder's processor, may further cause the decoder's processor to perform the following operation: determining the fusion weight set using a SAD-based algorithm in response to the fourth syntax element including a first value. In some implementations, the memory stores instructions that, when executed by the decoder's processor, may further cause the decoder's processor to perform the following operation: determining the fusion weight set using an MSE-based algorithm in response to the fourth syntax element including a second value.
[0231] In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: decode a fifth syntax element from the bitstream in response to a second syntax element indicating that intraTMP fusion is not enabled for the current block. In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: identify a block vector from the candidate list based on the fifth syntax element.
[0232] In some implementations, the value range of the fifth syntax element is 0 to 18. In some implementations, when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream as follows: 0 followed by intra_tmp_idx - 2, represented by a 4-bit fixed-length code.
[0233] In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: decode a sixth syntax element from the bitstream in response to a second syntax element indicating that intraTMP fusion is not enabled for the current block. In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: determine, based on the sixth syntax element, whether to determine the intraTMP prediction value by filtering a selected reference block.
[0234] In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: determine the intraTMP prediction value by filtering a selected reference block using a learned filter in response to the sixth syntax element including a first value.
[0235] In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: decode a seventh syntax element from the bitstream in response to the sixth syntax element including a second value. In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: determine, based on the seventh syntax element, whether to refine the block vector using fractional precision.
[0236] In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: determine, in response to the seventh syntax element including a first value, not to refine the block vector using fractional precision. In some implementations, the memory stores instructions that, when executed by the decoder's processor, may also cause the decoder's processor to: determine, in response to the seventh syntax element including a second value, to refine the block vector using fractional precision.
[0237] In some implementations, the memory stores instructions that, when executed by the decoder's processor, can also cause the decoder's processor to perform the following operations: decoding an eighth and ninth syntax element from the bitstream in response to the seventh syntax element including a second value. In some implementations, the memory stores instructions that, when executed by the decoder's processor, can also cause the decoder's processor to perform the following operations: determining a sub-pixel thinning direction based on the eighth syntax element. In some implementations, the memory stores instructions that, when executed by the decoder's processor, can also cause the decoder's processor to perform the following operations: determining a sub-pixel thinning phase based on the ninth syntax element. In some implementations, the memory stores instructions that, when executed by the decoder's processor, can also cause the decoder's processor to perform the following operations: thinning the block vector based on the sub-pixel thinning direction and the sub-pixel thinning phase.
[0238] According to another aspect of this disclosure, an encoding method implemented by an encoder is provided. The method may include: a processor enabling intraTMP to encode a current block. The method may include: the processor encoding a first syntax element into a bitstream. The method may include: the processor determining, based on the first syntax element, whether to enable an intraTMP mode for the current block. The method may include: in response to enabling the intraTMP mode for the current block, the processor encoding a second syntax element into the bitstream. The method may include: the processor determining, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The method may include: in response to determining that the intraTMP prediction value is determined using an intraTMP fusion mode, the processor encoding a third and a fourth syntax element into the bitstream. The method may include: the processor determining a fusion weight set based on the fourth syntax element. The method may include: the processor determining the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. The method may include: the processor encoding the current block based on the intraTMP prediction value.
[0239] In some implementations, in response to determining the intraTMP prediction value of the current block through the intraTMP fusion mode, the method may include: the processor performing a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on the SAD cost computed on the template region.
[0240] In some implementations, the processor performs sparse search and refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions. This may include: the processor identifying a set of sparse candidate block vectors in parallel during the sparse search. In some implementations, the processor performs sparse search and refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions. This may include: the processor constructing the candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using a selected template shape.
[0241] In some implementations, the method may include: the processor determining, based on the third syntax element, the set of reference blocks selected for determining the intraTMP prediction value.
[0242] In some implementations, in response to the fourth syntax element including a first value, the method may include: the processor determining the fusion weight set using a SAD-based algorithm. In some implementations, in response to the fourth syntax element including a second value, the method may include: the processor determining the fusion weight set using an MSE-based algorithm.
[0243] In some implementations, in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block, the method may include: the processor encoding a fifth syntax element into the bitstream. In some implementations, the method may include: the processor identifying a block vector from the candidate list based on the fifth syntax element.
[0244] In some implementations, the value range of the fifth syntax element is 0 to 18. In some implementations, when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream as follows: 0 followed by intra_tmp_idx - 2, represented by a 4-bit fixed-length code.
[0245] In some implementations, in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block, the method may include: the processor encoding a sixth syntax element into the bitstream. In some implementations, the method may include: the processor determining, based on the sixth syntax element, whether to determine the intraTMP prediction value by filtering a selected reference block.
[0246] In some implementations, in response to the sixth syntax element including a first value, the method may include: the processor determining the intraTMP prediction value by filtering a selected reference block using a learned filter.
[0247] In some implementations, in response to the sixth syntax element including a second value, the method may include: the processor encoding a seventh syntax element into the bitstream. In some implementations, the method may include: the processor determining, based on the seventh syntax element, whether to refine the block vector using fractional precision.
[0248] In some implementations, in response to the seventh syntax element including a first value, the method may include: the processor determining not to refine the block vector using fractional precision. In some implementations, in response to the seventh syntax element including a second value, the method may include: the processor determining to refine the block vector using fractional precision.
[0249] In some implementations, in response to the seventh syntax element including a second value, the method may include: the processor encoding an eighth and a ninth syntax element into the bitstream. In some implementations, the method may include: the processor determining a subpixel thinning direction based on the eighth syntax element. In some implementations, the method may include: the processor determining a subpixel thinning phase based on the ninth syntax element. In some implementations, the method may include: the processor thinning the block vector based on the subpixel thinning direction and the subpixel thinning phase.
[0250] According to another aspect of this disclosure, an encoder is provided. The encoder may include a processor and a memory storing instructions. The memory stores instructions that, when executed by the processor, cause the processor to: enable intraTMP to encode a current block. The memory stores instructions that, when executed by the processor, cause the processor to: encode a first syntax element into a bitstream. The memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the first syntax element, whether intraTMP mode is enabled for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: encode a second syntax element into the bitstream in response to enabling intraTMP mode for the current block. The memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: in response to determining the intraTMP prediction value through an intraTMP fusion mode, encode a third syntax element and a fourth syntax element into the bitstream. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: determine a fusion weight set based on the fourth syntax element. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: determine the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. The memory storage instruction, when executed by the processor, causes the processor to perform the following operations: encode the current block based on the intraTMP prediction value.
[0251] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to perform the following operations in response to determining the intraTMP prediction value of the current block through the intraTMP fusion mode: performing a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on the SAD cost computed on the template region.
[0252] In some implementations, to perform sparse search and refinement search, and to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions, the memory stores instructions that, when executed by the processor, cause the processor to perform the following operation: concurrently identify a set of sparse candidate block vectors during the sparse search. In some implementations, to perform sparse search and refinement search, and to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions, the memory stores instructions that, when executed by the processor, cause the processor to perform the following operation: construct the candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using a selected template shape.
[0253] In some implementations, the memory stores instructions that, when executed by the processor, enable the processor to perform the following operation: determine, based on the third syntax element, the set of reference blocks selected for determining the intraTMP prediction value.
[0254] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to perform the following operation: in response to the fourth syntax element including a first value, determine the fusion weight set using a SAD-based algorithm. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to perform the following operation: in response to the fourth syntax element including a second value, determine the fusion weight set using an MSE-based algorithm.
[0255] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: encode a fifth syntax element into the bitstream in response to a second syntax element indicating that intraTMP fusion is not enabled for the current block. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: identify a block vector from the candidate list based on the fifth syntax element.
[0256] In some implementations, the value range of the fifth syntax element is 0 to 18. In some implementations, when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream as follows: 0 followed by intra_tmp_idx - 2, represented by a 4-bit fixed-length code.
[0257] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: encode a sixth syntax element into the bitstream in response to a second syntax element indicating that intraTMP fusion is not enabled for the current block. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the sixth syntax element, whether to determine the intraTMP prediction value by filtering a selected reference block.
[0258] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to perform the following operation: in response to the sixth syntax element including a first value, to determine the intraTMP prediction value by filtering a selected reference block using a learned filter.
[0259] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: encode a seventh syntax element into the bitstream in response to the sixth syntax element including a second value. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: determine, based on the seventh syntax element, whether to refine the block vector using fractional precision.
[0260] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: determine, in response to the seventh syntax element including a first value, not to refine the block vector using fractional precision. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: determine, in response to the seventh syntax element including a second value, to refine the block vector using fractional precision.
[0261] In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: encode an eighth and a ninth syntax element into the bitstream in response to the seventh syntax element including the second value. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: determine a subpixel thinning direction based on the eighth syntax element. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: determine a subpixel thinning phase based on the ninth syntax element. In some implementations, the memory stores instructions that, when executed by the processor, cause the processor to: thin the block vector based on the subpixel thinning direction and the subpixel thinning phase.
[0262] According to another aspect of this disclosure, a non-transitory computer-readable medium is provided, which stores instructions for an encoder. When executed by a processor of the encoder, the instructions can cause the encoder's processor to enable intraTMP for encoding a current block. When executed by the encoder's processor, the instructions can cause the encoder's processor to encode a first syntax element into a bitstream. When executed by the encoder's processor, the instructions can cause the encoder's processor to determine, based on the first syntax element, whether to enable intraTMP mode for the current block. Instructions stored in memory, when executed by the processor, can cause the processor to encode a second syntax element into the bitstream in response to enabling the intraTMP mode for the current block. When executed by the encoder's processor, the instructions can cause the encoder's processor to determine, based on the second syntax element, whether to determine an intraTMP prediction value for the current block using an intraTMP fusion mode. The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: in response to determining the intraTMP prediction value through an intraTMP fusion mode, encode a third syntax element and a fourth syntax element into the bitstream. When executed by the encoder's processor, the instructions cause the encoder's processor to perform the following operations: determine a fusion weight set based on the fourth syntax element. When executed by the encoder's processor, the instructions cause the encoder's processor to perform the following operations: determine the intraTMP prediction value based on the fusion weight set and a set of reference blocks indicated by the third syntax element. When executed by the encoder's processor, the instructions cause the encoder's processor to perform the following operations: encode the current block based on the intraTMP prediction value.
[0263] In some implementations, when the instructions are executed by the encoder's processor, the encoder's processor may perform the following operations in response to determining the intraTMP prediction value of the current block through the intraTMP fusion mode: performing sparse search and refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on the SAD cost calculated on the template region.
[0264] In some implementations, to perform sparse search and refinement search, and to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions, the instructions, when executed by the encoder's processor, can cause the encoder's processor to perform the following operation: concurrently identify the set of sparse candidate block vectors during the sparse search. In some implementations, to perform sparse search and refinement search, and to construct a candidate list containing N intraTMP block vectors by selecting block vectors based on SAD costs computed on template regions, the instructions, when executed by the encoder's processor, can cause the encoder's processor to perform the following operation: construct the candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using a selected template shape.
[0265] In some implementations, when the instructions are executed by the encoder's processor, the encoder's processor may perform the following operation: determine, based on the third syntax element, the set of reference blocks selected for determining the intraTMP prediction value.
[0266] In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to perform the following operations: in response to the fourth syntax element including a first value, determine the fusion weight set using a SAD-based algorithm. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to perform the following operations: in response to the fourth syntax element including a second value, determine the fusion weight set using an MSE-based algorithm.
[0267] In some implementations, when executed by the encoder's processor, the instruction may cause the encoder's processor to: encode a fifth syntax element into the bitstream in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block. In some implementations, when executed by the encoder's processor, the instruction may cause the encoder's processor to: identify a block vector from the candidate list based on the fifth syntax element.
[0268] In some implementations, the value range of the fifth syntax element is 0 to 18. In some implementations, when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream as follows: 0 followed by intra_tmp-idx - 2, represented by a 4-bit fixed-length code.
[0269] In some implementations, when executed by the encoder's processor, the instruction may cause the encoder's processor to: encode a sixth syntax element into the bitstream in response to the second syntax element indicating that intraTMP fusion is not enabled for the current block. In some implementations, when executed by the encoder's processor, the instruction may cause the encoder's processor to: determine, based on the sixth syntax element, whether to determine the intraTMP prediction value by filtering a selected reference block.
[0270] In some implementations, when the instructions are executed by the encoder's processor, the encoder's processor may perform the following operation: in response to the sixth syntax element including a first value, determine the intraTMP prediction value by filtering a selected reference block using a learned filter.
[0271] In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to perform the following operation: in response to the sixth syntax element including a second value, encode a seventh syntax element into the bitstream. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to perform the following operation: determine, based on the seventh syntax element, whether to refine the block vector using fractional precision.
[0272] In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to: determine, in response to the seventh syntax element including a first value, not to refine the block vector using fractional precision. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to: determine, in response to the seventh syntax element including a second value, to refine the block vector using fractional precision.
[0273] In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to: encode the eighth and ninth syntax elements into the bitstream in response to the seventh syntax element including the second value. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to: determine a subpixel thinning direction based on the eighth syntax element. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to: determine a subpixel thinning phase based on the ninth syntax element. In some implementations, when executed by the encoder's processor, the instructions may cause the encoder's processor to: thin the block vector based on the subpixel thinning direction and the subpixel thinning phase.
[0274] The description of the above embodiments will reveal the general nature of this disclosure, enabling others to easily modify and / or adapt these embodiments to various applications without excessive experimentation by applying technical knowledge in the art, without departing from the general conception of this disclosure. Therefore, based on the teachings and guidance set forth herein, such modifications and adaptations are intended to fall within the meaning and scope of equivalents of the disclosed embodiments. It should be understood that phrases or terms herein are for descriptive rather than limiting purposes, and that the terminology or phrases in this specification should be interpreted by those skilled in the art based on the teachings and guidance.
[0275] The embodiments of this disclosure have been described above using functional building blocks, which illustrate how specified functions and their relationships are implemented. For ease of description, the boundaries of these functional building blocks are arbitrarily defined herein. Alternative boundaries can be defined as long as the specified functions and their relationships are properly executed.
[0276] The summary and abstract sections may list one or more, but not all, exemplary embodiments of this disclosure as conceived by the inventors, and therefore they are not intended to limit this disclosure and the appended claims in any way.
[0277] Various functional blocks, modules, and steps have been disclosed above. The arrangements provided are illustrative and not limiting. Accordingly, functional blocks, modules, and steps may be rearranged or combined in ways different from the examples provided above. Similarly, some embodiments include only a subset of functional blocks, modules, and steps, and any such subset is permitted.
[0278] The breadth and scope of this disclosure should not be limited by any of the exemplary embodiments described above, but should be defined solely by the appended claims and their equivalents.
Claims
1. A decoding method performed by a decoder, the method comprising: Comprising: a processor parsing a bitstream to decode a plurality of syntax elements associated with intra template prediction (intraTMP); the processor decoding a first syntax element from the bitstream; the processor determining, based on the first syntax element, whether an intraTMP mode is enabled for a current block; in response to the intraTMP mode being enabled for the current block, the processor decoding a second syntax element from the bitstream; the processor determining, based on the second syntax element, whether an intraTMP prediction value for the current block is determined by an intraTMP merge mode; in response to determining that the intraTMP prediction value is determined by the intraTMP merge mode, the processor decoding a third syntax element and a fourth syntax element from the bitstream; the processor determining a set of merge weights based on the fourth syntax element; the processor determining the intraTMP prediction value based on the set of merge weights and a set of reference blocks indicated by the third syntax element; and the processor decoding the current block based on the intraTMP prediction value.
2. The method of claim 1, wherein, Further comprising: in response to determining that the intraTMP prediction value for the current block is determined by the intraTMP merge mode, the processor performing a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by picking block vectors according to sum of absolute difference (SAD) cost calculated on a template region.
3. The method of claim 2, wherein, the processor performing a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by picking block vectors according to SAD cost calculated on a template region, comprising: the processor identifying a set of sparse candidate block vectors in parallel during the sparse search; and the processor constructing the candidate list containing N intraTMP block vectors by searching around sparse candidate block vectors using a selected template shape.
4. The method of claim 1, wherein, Further comprising: the processor determining, based on the third syntax element, the set of reference blocks selected for determining the intraTMP prediction value.
5. The method of claim 1, wherein, Further comprising: in response to the fourth syntax element comprising a first value, the processor determining the set of merge weights using a sum of absolute difference (SAD) based algorithm; and in response to the fourth syntax element comprising a second value, the processor determining the set of merge weights using a mean square error (MSE) based algorithm.
6. The method of claim 2, wherein, Further comprising: in response to the second syntax element indicating that intraTMP merge is not enabled for the current block, the processor decoding a fifth syntax element from the bitstream; and the processor identifying a block vector from the candidate list based on the fifth syntax element.
7. The method of claim 6, wherein, the fifth syntax element has a value range of 0 to 18, and when the value of the fifth syntax element is in a range of 3 to 18, the fifth syntax element is represented in the bitstream by 0 followed by intra_tmp_idx - 2 represented by a 4-bit fixed length code.
8. The method of claim 6, wherein, Further comprising: in response to the second syntax element indicating that intra TMP blending is not enabled for the current block, the processor decodes a sixth syntax element from the bitstream; and the processor determines, based on the sixth syntax element, whether to determine the intra TMP prediction value by filtering the selected reference block.
9. The method of claim 8, wherein, Further comprising: in response to the sixth syntax element comprising a first value, the processor determines the intra TMP prediction value by filtering the selected reference block with a learned filter.
10. The method of claim 9, wherein, Further comprising: in response to the sixth syntax element comprising a second value, the processor decodes a seventh syntax element from the bitstream; and the processor determines, based on the seventh syntax element, whether to refine the block vector with fractional precision.
11. The method of claim 10, wherein, Further comprising: in response to the seventh syntax element comprising a first value, the processor determines not to refine the block vector with fractional precision; and in response to the seventh syntax element comprising a second value, the processor determines to refine the block vector with fractional precision.
12. The method of claim 11, wherein, Further comprising: in response to the seventh syntax element comprising the second value, the processor decodes an eighth syntax element and a ninth syntax element from the bitstream; the processor determines a sub-pixel refinement direction based on the eighth syntax element; the processor determines a sub-pixel refinement phase based on the ninth syntax element; and the processor refines the block vector based on the sub-pixel refinement direction and the sub-pixel refinement phase.
13. A decoder, characterized by Comprising: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: parse a bitstream to decode a plurality of syntax elements associated with intra template prediction (intra TMP); decode a first syntax element from the bitstream; determine, based on the first syntax element, whether an intra TMP mode is enabled for a current block; in response to the intra TMP mode being enabled for the current block, decode a second syntax element from the bitstream; determine, based on the second syntax element, whether to determine an intra TMP prediction value for the current block by an intra TMP blending mode; in response to determining to determine the intra TMP prediction value by an intra TMP blending mode, decode a third syntax element and a fourth syntax element from the bitstream; determine a set of blending weights based on the fourth syntax element; determine the intra TMP prediction value based on the set of blending weights and a set of reference blocks indicated by the third syntax element; and decode the current block based on the intra TMP prediction value. The memory stores instructions that, when executed by the processor, further cause the processor to:
14. The decoder of claim 13, wherein, in response to determining to determine the intra TMP prediction value for the current block by the intra TMP blending mode, perform a sparse search and a refinement search to construct a candidate list containing N intra TMP block vectors by picking a block vector according to a sum of absolute difference (SAD) cost calculated on a template region. 15. The decoder of claim 14, wherein, To perform the sparse search and the refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors according to SAD cost calculated on template regions, the memory stores instructions that, when executed by the processor, further cause the processor to perform the following operations: identify a set of sparse candidate block vectors in parallel during the sparse search; and construct the candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using the selected template shape.
16. The decoder of claim 13, wherein, The memory stores instructions that, when executed by the processor, further cause the processor to perform the following operations: determine, based on the third syntax element, that the set of reference blocks used to determine the intraTMP prediction value is selected.
17. The decoder of claim 13, wherein, The memory stores instructions that, when executed by the processor, further cause the processor to perform the following operations: in response to the fourth syntax element including a first value, determine the set of blending weights using a sum of absolute difference (SAD) based algorithm; and in response to the fourth syntax element including a second value, determine the set of blending weights using a mean square error (MSE) based algorithm.
18. The decoder of claim 14, wherein, The memory stores instructions that, when executed by the processor, further cause the processor to perform the following operations: in response to the second syntax element indicating that intraTMP blending is not enabled for the current block, decode a fifth syntax element from the bitstream; and identify a block vector from the candidate list based on the fifth syntax element.
19. The decoder of claim 18, wherein, The fifth syntax element has a value range of 0 to 18, and when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream by 0 followed by intra_tmp-idx - 2 represented by a 4-bit fixed length code.
20. The decoder of claim 18, wherein, The memory stores instructions that, when executed by the processor, further cause the processor to perform the following operations: in response to the second syntax element indicating that intraTMP blending is not enabled for the current block, decode a sixth syntax element from the bitstream; and determine, based on the sixth syntax element, whether to determine the intraTMP prediction value by filtering the selected reference blocks.
21. The decoder of claim 20, wherein, The memory stores instructions that, when executed by the processor, further cause the processor to perform the following operations: in response to the sixth syntax element including a first value, determine the intraTMP prediction value by filtering the selected reference blocks with a learned filter.
22. The decoder of claim 21, wherein, The memory stores instructions that, when executed by the processor, further cause the processor to perform the following operations: in response to the sixth syntax element including a second value, decode a seventh syntax element from the bitstream; and determine, based on the seventh syntax element, whether to refine the block vector with fractional precision.
23. The decoder of claim 22, wherein, The memory stores instructions that, when executed by the processor, further cause the processor to perform the following operations: in response to the seventh syntax element including a first value, determining not to refine the block vector with fractional precision; and in response to the seventh syntax element including a second value, determining to refine the block vector with fractional precision.
24. The decoder of claim 23, wherein, The memory stores instructions that, when executed by the processor, further cause the processor to: in response to the seventh syntax element including the second value, decode an eighth syntax element and a ninth syntax element from the bitstream; determine a sub-pixel refinement direction based on the eighth syntax element; determine a sub-pixel refinement phase based on the ninth syntax element; and refine the block vector based on the sub-pixel refinement direction and the sub-pixel refinement phase.
25. A non-transitory computer readable medium storing instructions that, when executed by a processor of a decoder, cause the processor of the decoder to: parse a bitstream to decode a plurality of syntax elements associated with intra template prediction (intraTMP); decode a first syntax element from the bitstream; based on the first syntax element, determine whether an intraTMP mode is enabled for a current block; in response to the intraTMP mode being enabled for the current block, decode a second syntax element from the bitstream; based on the second syntax element, determine whether an intraTMP prediction value for the current block is determined by an intraTMP merge mode; in response to determining that the intraTMP prediction value is determined by the intraTMP merge mode, decode a third syntax element and a fourth syntax element from the bitstream; determine a set of merge weights based on the fourth syntax element; determine the intraTMP prediction value based on the set of merge weights and a set of reference blocks indicated by the third syntax element; and decode the current block based on the intraTMP prediction value. The instructions, when executed by the processor of the decoder, cause the processor of the decoder to:
26. The non-transitory computer-readable medium of claim 25, wherein, in response to determining that the intraTMP prediction value for the current block is determined by the intraTMP merge mode, perform a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by picking block vectors according to sum of absolute difference (SAD) cost computed on a template region. To perform the sparse search and the refinement search to construct the candidate list containing N intraTMP block vectors by picking block vectors according to SAD cost computed on a template region, the instructions, when executed by the processor of the decoder, cause the processor of the decoder to:
27. The non-transitory computer-readable medium of claim 26, wherein, identify a set of sparse candidate block vectors in parallel during the sparse search; and construct the candidate list containing N intraTMP block vectors by searching around sparse candidate block vectors using a selected template shape. The instructions, when executed by the processor of the decoder, cause the processor of the decoder to:
28. The non-transitory computer-readable medium of claim 25, wherein, based on the third syntax element, determine to select the set of reference blocks for determining the intraTMP prediction value.
29. The non-transitory computer-readable medium of claim 25, wherein, The instructions, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operations: in response to the fourth syntax element comprising a first value, determine the set of blending weights using a sum of absolute difference (SAD) based algorithm; and in response to the fourth syntax element comprising a second value, determine the set of blending weights using a mean square error (MSE) based algorithm.
30. The non-transitory computer-readable medium of claim 26, wherein, The instructions, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operations: in response to the second syntax element indicating that intraTMP blending is not enabled for the current block, decode a fifth syntax element from the bitstream; and identify a block vector from the candidate list based on the fifth syntax element.
31. The non-transitory computer-readable medium of claim 30, wherein, The fifth syntax element has a value range of 0 to 18, and when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream by 0 followed by intra_tmp_idx - 2 represented by a 4-bit fixed length code.
32. The non-transitory computer-readable medium of claim 30, wherein, The instructions, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operations: in response to the second syntax element indicating that intraTMP blending is not enabled for the current block, decode a sixth syntax element from the bitstream; and based on the sixth syntax element, determine whether to determine the intraTMP prediction value by filtering the selected reference block.
33. The non-transitory computer-readable medium of claim 32, wherein, The instructions, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operations: in response to the sixth syntax element comprising a first value, determine the intraTMP prediction value by filtering the selected reference block with a learned filter.
34. The non-transitory computer-readable medium of claim 33, wherein, The instructions, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operations: in response to the sixth syntax element comprising a second value, decode a seventh syntax element from the bitstream; and based on the seventh syntax element, determine whether to refine the block vector with fractional precision.
35. The non-transitory computer-readable medium of claim 34, wherein, The instructions, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operations: in response to the seventh syntax element comprising a first value, determine not to refine the block vector with fractional precision; and in response to the seventh syntax element comprising a second value, determine to refine the block vector with fractional precision.
36. The non-transitory computer-readable medium of claim 35, wherein, The instructions, when executed by the processor of the decoder, cause the processor of the decoder to perform the following operations: in response to the seventh syntax element comprising the second value, decode an eighth syntax element and a ninth syntax element from the bitstream; determine a sub-pixel refinement direction based on the eighth syntax element; determine a sub-pixel refinement phase based on the ninth syntax element; and refine the block vector based on the sub-pixel refinement direction and the sub-pixel refinement phase. comprise:
37. A method of encoding performed by an encoder, the method comprising: a processor enabling intra template prediction (intraTMP) for encoding a current block; The processor encodes a first syntax element into a bitstream; The processor determines, based on the first syntax element, whether an intra TMP mode is enabled for a current block; In response to the intra TMP mode being enabled for the current block, the processor encodes a second syntax element into the bitstream; The processor determines, based on the second syntax element, whether an intra TMP prediction value of the current block is determined by an intra TMP merge mode; In response to determining that the intra TMP prediction value is determined by the intra TMP merge mode, the processor encodes a third syntax element and a fourth syntax element into the bitstream; The processor determines, based on the fourth syntax element, a set of fusion weights; The processor determines the intra TMP prediction value based on the set of fusion weights and a set of reference blocks indicated by the third syntax element; and The processor encodes the current block based on the intra TMP prediction value.
38. The method of claim 37, wherein, Further comprising: In response to determining that the intra TMP prediction value of the current block is determined by the intra TMP merge mode, the processor performs a sparse search and a refinement search to construct a candidate list containing N intra TMP block vectors by picking block vectors according to sum of absolute difference (SAD) cost calculated on a template region.
39. The method of claim 38, wherein, The processor performs a sparse search and a refinement search to construct a candidate list containing N intra TMP block vectors by picking block vectors according to sum of absolute difference (SAD) cost calculated on a template region, comprising: The processor identifies a set of sparse candidate block vectors in parallel during the sparse search; and The processor constructs the candidate list containing N intra TMP block vectors by searching around sparse candidate block vectors using the selected template shape.
40. The method of claim 37, wherein, Further comprising: The processor determines, based on the third syntax element, the set of reference blocks selected for determining the intra TMP prediction value.
41. The method of claim 37, wherein, Further comprising: In response to the fourth syntax element comprising a first value, the processor determines the set of fusion weights using a sum of absolute difference (SAD) based algorithm; and In response to the fourth syntax element comprising a second value, the processor determines the set of fusion weights using a mean square error (MSE) based algorithm.
42. The method of claim 38, wherein, Further comprising: In response to the second syntax element indicating that the intra TMP merge is not enabled for the current block, the processor encodes a fifth syntax element into the bitstream; and The processor identifies a block vector from the candidate list based on the fifth syntax element.
43. The method of claim 42, wherein, The fifth syntax element has a value range of 0 to 18, and when the value of the fifth syntax element is in a range of 3 to 18, the fifth syntax element is represented in the bitstream by 0 followed by intra_tmp_idx - 2 represented by a 4-bit fixed length code.
44. The method of claim 42, wherein, Further comprising: in response to the second syntax element indicating that intra TMP merging is not enabled for the current block, the processor encodes a sixth syntax element into the bitstream; and the processor determines, based on the sixth syntax element, whether to determine the intra TMP prediction value by filtering the selected reference block.
45. The method of claim 44, wherein, Further comprising: in response to the sixth syntax element comprising a first value, the processor determines the intra TMP prediction value by filtering the selected reference block with a learned filter.
46. The method of claim 45, wherein, Further comprising: in response to the sixth syntax element comprising a second value, the processor encodes a seventh syntax element into the bitstream; and the processor determines, based on the seventh syntax element, whether to refine the block vector with fractional precision.
47. The method of claim 46, wherein, Further comprising: in response to the seventh syntax element comprising a first value, the processor determines not to refine the block vector with fractional precision; and in response to the seventh syntax element comprising a second value, the processor determines to refine the block vector with fractional precision.
48. The method of claim 47, wherein, Further comprising: in response to the seventh syntax element comprising the second value, the processor encodes an eighth syntax element and a ninth syntax element into the bitstream; the processor determines, based on the eighth syntax element, a sub-pixel refinement direction; the processor determines, based on the ninth syntax element, a sub-pixel refinement phase; and the processor refines the block vector based on the sub-pixel refinement direction and the sub-pixel refinement phase.
49. An encoder comprising: Comprising: a processor; and a memory storing instructions that, when executed by the processor, cause the processor to: enable intra template prediction (intra TMP) for encoding a current block; encode a first syntax element into a bitstream; determine, based on the first syntax element, whether an intra TMP mode is enabled for the current block; in response to the intra TMP mode being enabled for the current block, encode a second syntax element into the bitstream; determine, based on the second syntax element, whether to determine an intra TMP prediction value for the current block by an intra TMP merging mode; in response to determining to determine the intra TMP prediction value by the intra TMP merging mode, encode a third syntax element and a fourth syntax element into the bitstream; determine, based on the fourth syntax element, a set of merging weights; determine the intra TMP prediction value based on the set of merging weights and a set of reference blocks indicated by the third syntax element; and encode the current block based on the intra TMP prediction value. the memory stores instructions that, when executed by the processor, cause the processor to:
50. The encoder of claim 49, wherein, in response to determining to determine the intra TMP prediction value for the current block by the intra TMP merging mode, perform a sparse search and a refinement search to construct a candidate list containing N intra TMP block vectors by picking a block vector according to a sum of absolute difference (SAD) cost calculated on a template region. 51. The encoder of claim 50, wherein, To perform a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by selecting block vectors according to SAD cost calculated on a template region, the memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: identify a set of sparse candidate block vectors in parallel during the sparse search; and construct the candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using the selected template shape.
52. The encoder of claim 49, wherein, The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: determine, based on the third syntax element, that the set of reference blocks used to determine the intraTMP prediction value is selected.
53. The encoder of claim 49, wherein, The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: in response to the fourth syntax element including a first value, determine the set of blending weights using an algorithm based on sum of absolute differences (SAD); and in response to the fourth syntax element including a second value, determine the set of blending weights using an algorithm based on mean square error (MSE).
54. The encoder of claim 50, wherein, The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: in response to the second syntax element indicating that intraTMP blending is not enabled for the current block, encode a fifth syntax element to the bitstream; and identify a block vector from the candidate list based on the fifth syntax element.
55. The encoder of claim 54, wherein, The fifth syntax element has a value range of 0 to 18, and when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream by 0 followed by intra_tmp_idx - 2 represented by a 4-bit fixed length code.
56. The encoder of claim 54, wherein, The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: in response to the second syntax element indicating that intraTMP blending is not enabled for the current block, encode a sixth syntax element to the bitstream; and determine, based on the sixth syntax element, whether to determine the intraTMP prediction value by filtering the selected reference blocks.
57. The encoder of claim 56, wherein, The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: in response to the sixth syntax element including a first value, determine the intraTMP prediction value by filtering the selected reference blocks with a learned filter.
58. The encoder of claim 57, wherein, The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: in response to the sixth syntax element including a second value, encode a seventh syntax element to the bitstream; and determine, based on the seventh syntax element, whether to refine the block vector with fractional precision.
59. The encoder of claim 58, wherein, The memory stores instructions that, when executed by the processor, cause the processor to perform the following operations: determining not to refine the block vector with fractional precision in response to the seventh syntax element including a first value; and determining to refine the block vector with fractional precision in response to the seventh syntax element including a second value.
60. The encoder of claim 59, wherein, The memory stores instructions that, when executed by the processor, cause the processor to: encode an eighth syntax element and a ninth syntax element to the bitstream in response to the seventh syntax element including the second value; determine a sub-pixel refinement direction based on the eighth syntax element; determine a sub-pixel refinement phase based on the ninth syntax element; and refine the block vector based on the sub-pixel refinement direction and the sub-pixel refinement phase.
61. A non-transitory computer readable medium storing instructions that, when executed by a processor of an encoder, cause the processor of the encoder to: enable intra template prediction (intraTMP) for encoding a current block; encode a first syntax element to a bitstream; determine whether an intraTMP mode is enabled for the current block based on the first syntax element; encode a second syntax element to the bitstream in response to the intraTMP mode being enabled for the current block; determine whether an intraTMP prediction value of the current block is determined by an intraTMP merge mode based on the second syntax element; encode a third syntax element and a fourth syntax element to the bitstream in response to determining that the intraTMP prediction value is determined by the intraTMP merge mode; determine a set of merge weights based on the fourth syntax element; determine the intraTMP prediction value based on the set of merge weights and a set of reference blocks indicated by the third syntax element; and encode the current block based on the intraTMP prediction value. The instructions, when executed by the processor of the encoder, cause the processor of the encoder to:
62. The non-transitory computer-readable medium of claim 61, wherein, perform a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by picking block vectors according to sum of absolute difference (SAD) cost calculated on a template region in response to determining that the intraTMP prediction value of the current block is determined by the intraTMP merge mode. To perform a sparse search and a refinement search to construct a candidate list containing N intraTMP block vectors by picking block vectors according to SAD cost calculated on a template region, the instructions, when executed by the processor of the encoder, cause the processor of the encoder to:
63. The non-transitory computer-readable medium of claim 62, wherein, identify a set of sparse candidate block vectors in parallel during the sparse search; and construct the candidate list containing N intraTMP block vectors by searching around the sparse candidate block vectors using the selected template shape. The instructions, when executed by the processor of the encoder, cause the processor of the encoder to: determine the set of reference blocks selected for determining the intraTMP prediction value based on the third syntax element.
64. The non-transitory computer-readable medium of claim 61, wherein, 65. The non-transitory computer-readable medium of claim 61, wherein, The instructions, when executed by the processor of the encoder, cause the processor of the encoder to perform the following operations: in response to the fourth syntax element comprising a first value, determine the set of blending weights using a sum of absolute difference (SAD) based algorithm; and in response to the fourth syntax element comprising a second value, determine the set of blending weights using a mean square error (MSE) based algorithm.
66. The non-transitory computer-readable medium of claim 62, wherein, The instructions, when executed by the processor of the encoder, cause the processor of the encoder to perform the following operations: in response to the second syntax element indicating that intraTMP blending is not enabled for the current block, encode a fifth syntax element into the bitstream; and identify a block vector from the candidate list based on the fifth syntax element.
67. The non-transitory computer-readable medium of claim 66, wherein, The fifth syntax element has a value range of 0 to 18, and when the value of the fifth syntax element is in the range of 3 to 18, the fifth syntax element is represented in the bitstream by 0 followed by intra_tmp_idx - 2 represented by a 4-bit fixed length code.
68. The non-transitory computer-readable medium of claim 66, wherein, The instructions, when executed by the processor of the encoder, cause the processor of the encoder to perform the following operations: in response to the second syntax element indicating that intraTMP blending is not enabled for the current block, encode a sixth syntax element into the bitstream; and determine whether to determine the intraTMP prediction value by filtering the selected reference block based on the sixth syntax element.
69. The non-transitory computer-readable medium of claim 68, wherein, The instructions, when executed by the processor of the encoder, cause the processor of the encoder to perform the following operations: in response to the sixth syntax element comprising a first value, determine the intraTMP prediction value by filtering the selected reference block with a learned filter.
70. The non-transitory computer-readable medium of claim 69, wherein, The instructions, when executed by the processor of the encoder, cause the processor of the encoder to perform the following operations: in response to the sixth syntax element comprising a second value, encode a seventh syntax element into the bitstream; and determine whether to refine the block vector with fractional precision based on the seventh syntax element.
71. The non-transitory computer-readable medium of claim 70, wherein, The instructions, when executed by the processor of the encoder, cause the processor of the encoder to perform the following operations: in response to the seventh syntax element comprising a first value, determine not to refine the block vector with fractional precision; and in response to the seventh syntax element comprising a second value, determine to refine the block vector with fractional precision.
72. The non-transitory computer-readable medium of claim 71, wherein, The instructions, when executed by the processor of the encoder, cause the processor of the encoder to perform the following operations: in response to the seventh syntax element comprising the second value, encode an eighth syntax element and a ninth syntax element into the bitstream; determine a sub-pixel refinement direction based on the eighth syntax element; determine a sub-pixel refinement phase based on the ninth syntax element; and refine the block vector based on the sub-pixel refinement direction and the sub-pixel refinement phase. The instructions, when executed by the processor of the encoder, cause the processor of the encoder to perform the following operations: in response to the seventh syntax element comprising the second value, encode an eighth syntax element and a ninth syntax element into the bitstream; determine a sub-pixel refinement direction based on the eighth syntax element; determine a sub-pixel refinement phase based on the ninth syntax element; and refine the block vector based on the sub-pixel refinement direction and the sub-pixel refinement phase.