Template type selection for video encoding and decoding

By allowing the encoder to select appropriate prediction modes and template types according to the encoding cost in video encoding, the problem of difficulty in effectively selecting templates in the prior art is solved, and the encoding performance and video compression efficiency are improved.

CN120019642APending Publication Date: 2025-05-16INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071619.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-10
Filing Date
2023-09-29
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When existing video encoding technologies use template-based encoding tools, it is difficult to effectively select prediction templates suitable for the current block, resulting in poor encoding performance.

Method used

A method is provided that allows the encoder to select appropriate prediction modes and template types according to encoding cost and signal the type of templates in the encoded data so that the decoder can perform appropriate reconstruction. The method involves using three types of templates: the upper template, the left template, and a combination of them to capture the statistical properties of the current block.

Benefits of technology

By dynamically selecting the appropriate template type, encoding performance is improved, video compression efficiency is enhanced, and the decoder can correctly reconstruct video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019642A_ABST
    Figure CN120019642A_ABST
Patent Text Reader

Abstract

When a template-based prediction mode, such as an intra template matching prediction mode, is used, a video coding system may use different types of templates. The "L-shaped" templates conventionally used by the template-based coding tool may be divided into a horizontal "upper" (or "upper") template and a vertical "left" template. This allows the encoder to select an appropriate portion of the template that better captures the statistical characteristics of the current block and results in an improvement in coding performance. The type of template used is signaled in the encoded data such that when a template-based prediction mode, such as an intra template matching prediction mode, is used, a decoder may use an appropriate template for reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to European application number 22306519.4 filed on October 10, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0002] The present disclosure is in the field of video compression, and at least one embodiment more particularly relates to signaling the type of template to be used for template-based encoding tools. Background Art

[0003] In order to achieve high compression efficiency, image and video coding schemes usually use prediction and transformation to fully exploit the spatial and temporal redundancy in video content. In general, intra-frame or inter-frame prediction is used to exploit the correlation within or between frames, and then the difference between the original picture block and the predicted picture block (often expressed as prediction error or prediction residual) is transformed, quantized and entropy encoded. In order to reconstruct the video, the compressed data is decoded by the inverse process corresponding to entropy coding, quantization, transformation and prediction. Summary of the invention

[0004] At least one embodiment provides the encoder with the possibility to use different types of templates for template-based tools. The "L-shaped" template conventionally used by template-based coding tools can be divided into a horizontal "above" (or "upper") template and a vertical "left" template. This allows the encoder to select an appropriate part of the template that better captures the statistical characteristics of the current block and leads to improved coding performance. The type of template used is signaled in the encoded data so that the decoder will use the appropriate template for reconstruction.

[0005] A first aspect is directed to a method, comprising: for a current block of a picture, providing information representing the type of a template to be used by a template-based prediction mode (such as an intra-frame template matching prediction mode), wherein the type of the template is selected from a first template, a second template, or a third template, wherein the first template includes a group of adjacent pixels above the current block, the second template includes a group of adjacent pixels on the left side of the current block, and the third template is a combination of the first template and the second template.

[0006] The second aspect is directed to a method, which includes: for a current block of a picture, selecting a prediction mode and a type of template based on a coding cost, predicting the block based on the selected prediction mode and the type of template, encoding the current block based on the predicted block, and providing coding information for the current block, the coding information including at least information representing the use of a template-based prediction mode and information representing the type of template according to the first aspect.

[0007] The third aspect is directed to a device comprising a processor, wherein the processor is configured to: for a current block of a picture, select a prediction mode and a type of template based on a coding cost, predict the block based on the selected prediction mode and the type of template, encode the current block based on the predicted block, and provide coding information for the current block, the coding information comprising at least information representing the use of a template-based prediction mode and information representing the type of template according to the first aspect.

[0008] The fourth aspect is directed to a method, which includes: for a current block of a picture, obtaining encoding information for the current block, the encoding information at least including information representing the use of a template-based prediction mode and information representing the type of the template according to the first aspect, predicting a block based on the prediction mode and the type of the template, and decoding the current block based on the predicted block.

[0009] The fifth aspect is directed to a device, which includes a processor, and the processor is configured to: for a current block of a picture, obtain encoding information for the current block, the encoding information at least including information representing the use of a template-based prediction mode and information representing the type of template according to the first aspect, predict the block based on the prediction mode and the type of template, and decode the current block based on the predicted block.

[0010] A sixth aspect is directed to a computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to the first aspect, the second aspect or the fourth aspect.

[0011] A seventh aspect is directed to a non-transitory computer-readable medium storing executable program instructions, wherein the executable program instructions enable a computer executing the instructions to perform the method according to the first aspect, the second aspect, or the fourth aspect.

[0012] An eighth aspect is directed to a bitstream representing a coded picture, wherein according to the first aspect, the coded picture is formatted to include a syntax element indicating the type of a template of a current block.

[0013] A ninth aspect is directed to a non-transitory computer-readable medium storing information representing a coded picture, wherein according to the method of the first aspect, the coded picture is formatted to include a syntax element indicating the type of a template of a current block.

[0014] A simplified summary of the subject matter is presented above in order to provide a basic understanding of some aspects of the present disclosure. This summary is not an extensive overview of the subject matter. It is not intended to identify the key / important elements of the embodiments or to delineate the scope of the subject matter. Its sole purpose is to present some concepts of the subject matter in a simplified form as a prelude to the more detailed description provided below. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The present disclosure may be better understood by considering the following detailed description in conjunction with the accompanying drawings, in which:

[0016] Figure 1 A block diagram of a video encoder according to one embodiment is illustrated.

[0017] Figure 2 A block diagram of a video decoder according to one embodiment is illustrated.

[0018] Figure 3 A block diagram illustrating an example of a system in which various aspects and embodiments are implemented is illustrated.

[0019] Figure 4A , 4B 4C illustrates the principle of template-based intra-frame mode derivation.

[0020] Figure 5 The principle of intra mode derivation at the decoder side is illustrated.

[0021] Figure 6 The diagram illustrates the principle of intra-frame template matching.

[0022] Figure 7 The diagram illustrates the principle of template matching.

[0023] Figure 8 Illustrated is the reference area constraint for intra block copy in template matching mode.

[0024] Fig.9A A flow diagram illustrating an example of an encoding process for template-based prediction mode in accordance with at least one embodiment is illustrated.

[0025] Fig. 9B A flow diagram illustrating an example of a decoding process for template-based prediction mode in accordance with at least one embodiment is illustrated.

[0026] Fig.10 Illustrate the types of templates.

[0027] Fig.11 An example of CTU partitioning with CU decoding order is illustrated.

[0028] Fig.12 A flow diagram illustrating an example of an encoding process using intra template matching prediction mode in accordance with at least one embodiment is illustrated.

[0029] Fig.13 A flow diagram illustrating an example of a decoding process using intra template matching prediction mode in accordance with at least one embodiment is illustrated.

[0030] It should be understood that the drawings are for purposes of illustrating examples of various aspects, features, and embodiments according to the present disclosure and are not necessarily the only possible configurations. Throughout the various figures, like reference indicators refer to the same or similar features. DETAILED DESCRIPTION

[0031] As will be described in more detail below, a video codec may involve determining a prediction block for a current block based on samples of a selected block, the block being selected in a region of decoded picture information based on a template matching process, the template matching process comprising comparing a template associated with the current block with at least one other template associated with at least one other block in the region of decoded picture information. Coding methods, decoding methods, coding devices, and decoding devices based on this principle are described.

[0032] In addition, although the principles related to a specific draft of the VVC (Versatile Video Coding) or HEVC (High Efficiency Video Coding) specification are described, the present aspects are not limited to VVC or HEVC, and the present aspects can be applied to, for example, other standards and recommendations (whether pre-existing or developed in the future), and any such standards and recommendations Extensions (including VVC and HEVC). Unless otherwise stated, or technically excluded, the aspects described in this application can be used alone or in combination.

[0033] Figure 1 A block diagram of a video encoder according to one embodiment is illustrated. Variations of this encoder 100 are contemplated, but for clarity, the encoder 100 is described below without describing all contemplated variations. Before being encoded, the data sequence may undergo a pre-encoding process (101), for example, applying a color transform to an input color picture (e.g., conversion from RGB 4:4:4 to YCbCr4:2:0), or performing a remapping of input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.

[0034] In encoder 100, as described below, pictures are encoded by encoder elements. The picture to be encoded is segmented (102) and processed in units such as CUs. Each unit is encoded using, for example, intramode or intermode. When a unit is encoded in intramode, it performs intraprediction (160). In intermode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) which of intramode or intermode is used to encode the unit and indicates the intra / inter decision by, for example, a prediction mode flag. For example, a prediction residual is calculated by subtracting (110) the predicted block from the original image block.

[0035] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy encoded (145) to output a bitstream. The encoder may skip the transform and apply quantization directly to the untransformed residual signal. The encoder may bypass both the transform and quantization, i.e., encode the residual directly without applying the transform or quantization process.

[0036] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (1050) to decode the prediction residual. The image block is reconstructed by combining (155) the decoded prediction residual and the prediction block. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset), adaptive loop filter (ALF) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (180).

[0037] Figure 2A block diagram of a video decoder according to one embodiment is illustrated. In the decoder 200, as described below, a bitstream is decoded by decoder elements. The video decoder 200 typically performs a decoding pass that is inverse to the encoding pass. The encoder 100 also typically performs video decoding as part of encoding video data. In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other encoding information. The picture segmentation information indicates how the picture is segmented. Therefore, the decoder can divide (235) the picture according to the decoded picture segmentation information. The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The image block is reconstructed by combining (255) the decoded prediction residual and the prediction block. The prediction block can be obtained (270) from an intra-frame prediction (260) or a motion compensated prediction (i.e., inter-frame prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).

[0038] The decoded picture may further undergo post-decoding processing (285), such as an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or performing an inverse remapping that is the inverse of the remapping process performed in the pre-encoding process (101). The post-decoding process may use metadata derived in the pre-encoding process and signaled in the bitstream.

[0039] Figure 3 A block diagram of an example of a system in which various aspects and embodiments are implemented is illustrated. System 1000 may be embodied as a device including various components described below, and is configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances and servers. The elements of system 1000 may be embodied in a single integrated circuit (IC), multiple ICs and / or discrete components singly or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed over multiple ICs and / or discrete components. In various embodiments, system 1000 is coupled to one or more other systems or other electronic devices via, for example, a communication bus or by a dedicated input and / or output port communication. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.

[0040] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement, for example, various aspects described in this document. The processor 1010 may be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, and the like. The processor 1010 may include embedded memory, input and output interfaces, and various other circuits as known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, magnetic disk drive, and / or optical disk drive. As non-limiting examples, storage 1040 may include internal storage, attached storage (including removable and non-removable storage), and / or network accessible storage.

[0041] The system 1000 includes an encoder / decoder module 1030, which is configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module (one or more) that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. In addition, as known to those skilled in the art, the encoder / decoder module 1030 may be implemented as a separate element of the system 1000, or may be incorporated into the processor 1010 as a combination of hardware and software.

[0042] Program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described in this document may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of the various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0043] In some embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, an operating system for the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations, such as working memory for MPEG-2 (MPEG refers to Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG H Part 2), or VVC (Versatile Video Coding, a new standard being developed by the Joint Video Experts Group JVET).

[0044] Inputs to the elements of system 1000 may be provided through various input devices, as indicated in block 1130. Such input devices include, but are not limited to: (i) an RF section that receives a radio frequency (RF) signal transmitted over the air, for example, by a broadcast device, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 3 Other examples not shown include composite video.

[0045] In various embodiments, the input devices of block 1130 have associated respective input processing elements, as known in the art. For example, the RF portion may be associated with elements adapted to: (i) select a desired frequency (also referred to as selecting a signal, or band limiting a signal to a frequency band), (ii) down-convert the selected signal, (iii) again band-limit to a narrower frequency band to select a signal frequency band that may be referred to as a channel in some embodiments, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired data packet stream. The RF portion of various embodiments includes one or more elements to perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near baseband frequency) or baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter, down-convert and filter to the desired frequency band again to perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some in these elements, and / or add other elements of similar or different functions. Adding element can include inserting element between existing elements, such as for example inserting amplifier and analog to digital converter. In various embodiments, the RF part includes antenna.

[0046] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 1000 to other electronic devices across the USB and / or HDMI connections. It is to be understood that various aspects of input processing (e.g., Reed Solomon error correction) may be implemented, for example, within a separate input processing IC or within the processor 1010 as desired. Similarly, various aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within the processor 1010 as desired. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, to operate in conjunction with memory and storage elements to process the data streams as desired for presentation on an output device.

[0047] The various elements of system 1000 may be disposed within an integrated housing. Within the integrated housing, the various elements may be interconnected and transmit data between them using a suitable connection arrangement 1140 (e.g., an internal bus as known in the art, including an Inter-IC (I2C) bus, wiring, and a printed circuit board).

[0048] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, in a wired and / or wireless medium.

[0049] In various embodiments, data is streamed or otherwise provided to the system 1000 using a wireless network such as a Wi-Fi network (e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received through a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communications. The communication channel 1060 of these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 1000 using a set-top box that delivers data through the HDMI connection of the input block 1130. Still other embodiments provide streamed data to the system 1000 using the RF connection of the input block 1130. As indicated above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use a wireless network other than WiFi, such as a cellular network or a Bluetooth network.

[0050] The system 1000 can provide output signals to various output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be used for a television, a tablet computer, a laptop, a cellular phone (mobile phone), or other devices. The display 1100 can also be integrated with other components (e.g., as in a smart phone), or separated (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 1120 include one or more of a stand-alone digital video disk (or digital versatile disk) (DVR for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functions based on the output of the system 1000. For example, a disk player performs the function of playing the output of the system 1000.

[0051] In various embodiments, control signals are communicated between the system 1000 and the display 1100, speaker 1110, or other peripheral device 1120 using signaling such as AV.link, consumer electronics control (CEC), or other communication protocols that enable inter-device control with or without user intervention. Output devices may be communicatively coupled to the system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices may be connected to the system 1000 using a communication channel 1060 via a communication interface 1050. The display 1100 and speaker 1110 may be integrated into a single unit with other components of the system 1000 in an electronic device such as, for example, a television. In various embodiments, the display interface 1070 includes a display driver such as, for example, a timing controller (T Con) chip.

[0052] For example, if the RF portion of input 1130 is part of a stand-alone set-top box, the display 1100 and speaker 1110 may alternatively be separate from one or more of the other components. In various embodiments where the display 1100 and speaker 1110 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0053] The embodiments may be implemented by computer software implemented by the processor 1010 or by hardware or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, the memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as an optical memory device, a magnetic memory device, a semiconductor-based memory device, a fixed memory, and a removable memory. As a non-limiting example, the processor 1010 may be of any type suitable for the technical environment and may cover one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0054] The technical field of the embodiment relates to the intra-frame prediction stage of a video compression scheme, and more particularly to a template matching tool. The template matching tool is based on the following assumption that if the adjacent pixels of the block to be reconstructed can be found in the reconstructed area, the reconstructed block corresponding to the matching template is likely to be similar to the block to be reconstructed. When using the template matching tool, the device first determines a group of pixels forming the neighborhood of the current brightness block for the current block, i.e., the pixels adjacent to the block. This group of pixels in the form of an L-shape is conventionally referred to as a template. Various sizes (e.g., one, two or four pixels wide) can be used as templates, but are still based on the same principle. The process on both sides is similar. In other words, it is implemented by both an encoder device and a decoder device. Only a small signaling element is required to determine which template-based tool to use. Various template matching-based tools can be used to encode and decode videos. Some of them are described below.

[0055] Figure 4A , 4B 4C illustrates the principle of template-based intra-mode derivation. Template-based intra-mode derivation (TIMD) is the first tool based on template matching. For a given luma coding block (CB) 400, the following mode derivation via TIMD is applied in the same way on the encoder side and the decoder side. For each intra-prediction mode in the most probable mode (MPM) list of the luma CB, if supplemented with a default mode is needed, this mode calculates the prediction 401, 402 of the template of the luma CB from the decoded reference sample of the template 403, and calculates the sum of the absolute transform difference (SATD) between this prediction and the template of the luma CB. The two intra-prediction modes with the smallest SATD are selected as TIMD modes. Note that for TIMD, the set of directional intra-prediction modes is extended from 65 to 129. This means that the set of possible intra-prediction modes derived via TIMD collects 131 modes. After retaining the two intra prediction modes from the first test involving the MPM list supplemented with the default mode, for each of these two modes, if this mode is neither PLANAR nor DC, TIMD also tests its two closest extended directional intra prediction modes in terms of predicted SATD. Note that in the above description, it is assumed that the template of the luma CB does not exceed the boundary of the current frame. In the case that at least a part of the template of the luma CB exceeds the boundary of the current frame, if the reconstructed left or above samples, respectively, are not available, the template area where the prediction and SATD are calculated is as follows Figure 4B and Figure 4C Modified as depicted in .

[0056] To predict the current luma CB via TIMD, the two predictions of the luma CB via the two TIMD modes resulting from the two passes are merged with weights after applying PDPC. The weights used depend on the predicted SATD of the two TIMD modes.

[0057] exist Figure 4A In the example, the current W×H brightness CB(100) is surrounded by its fully available template, which consists of w on its left side. t ×H portion (402) and the W×h portion above it t During the TIMD derivation step, the intra prediction mode tested is based on the 1+2w of the template. t +2W+2h t +2H decoded reference samples (403) are used to predict the template of the current brightness CB. In at least one embodiment, if W≤8, then w t is equal to 2, otherwise, w t is equal to 4. If H≤8, then h t is equal to 2, otherwise, h t Equals 4.

[0058] exist Figure 4B In the example, the current W×H brightness CB(400) is surrounded by its template, and only its W×h t Part (401) is available. During the TIMD derivation step, the intra prediction mode tested is based on the 1+2W+2h of the template t +2H decoded reference samples (403) are used to predict the template of the current brightness CB.

[0059] exist Figure 4C In the example, the current W×H brightness CB(400) is surrounded by its template, and only its w t The ×H part (402) is available. During the TIMD derivation step, the intra prediction mode tested is based on the 1+2w of the template. t +2W+2H decoded reference samples (403) are used to predict the template of the current brightness CB.

[0060] Figure 5 The diagram illustrates the principle of decoder-side intra mode derivation.Decoder-side intra mode derivation (DIMD) is another template-based tool.

[0061] When DIMD is applied, two intra modes are derived from the reconstructed L-shaped template. Basically, a histogram of the gradients of the template pixels is constructed and two peak points are selected as the predicted angles o, and the corresponding intra modes are selected. These two predictors are combined with the planar mode predictor, with weights derived from their histogram values. The division operation in the weight derivation is performed using the same lookup table (LUT) based integerization scheme used by the cross component linear model (CCLM) prediction tool. For example, the orientation calculation Orient = G y / G x The division operation in is computed using the following LUT-based scheme:

[0062] x=Floor(Log2(G x ))

[0063] normDiff=((G x <<4)>>x)&15

[0064] x+=(3+(normDiff!=0)?1:0)

[0065] Orient=(G y *(DivSigTable[normDiff]|8)+(1<<(x-1)))>>x

[0066] Where DivSigTable

[16] ={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0}.

[0067] The derived intra modes are included in the main list of intra most probable modes (MPMs), so the DIMD process is performed before the MPM list is constructed. The main derived intra modes of a DIMD block are stored with the block and used for MPM list construction of neighboring blocks.

[0068] The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra prediction mode for the current block based on the neighboring reconstructed Y, Cb, and Cr samples in the second neighboring rows and columns, such as Figure 5 Specifically, the horizontal gradient and the vertical gradient are calculated for each collocated reconstructed luma sample and the reconstructed Cb and Cr samples of the current chroma block to construct a histogram of oriented gradients (HoG). Then, the intra prediction mode with the largest histogram amplitude value is used to perform chroma intra prediction of the current chroma block.

[0069] When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the direct mode, the intra prediction mode with the second largest histogram magnitude value is used as the DIMD chroma mode. A coding unit level flag is signaled to indicate whether the proposed DIMD chroma mode is applied.

[0070] The cross-component linear model (MMLM) based on multiple models is another tool based on template analysis, which extends the cross-component linear model (CCLM) prediction by adding three MMLM modes. In each MMLM mode, the reconstructed template samples are divided into two classes using a threshold, which is the average of the neighboring samples of the luminance reconstruction. The least mean square (LMS) method is used to derive the linear model for each class. For the CCLM mode using a single class, the LMS method is also used to derive the linear model. Slope adjustment is applied to CCLM and MMLM predictions. The adjustment is to tilt the linear function that maps luminance values ​​to chrominance values ​​relative to the center point determined by the average luminance value of the reference samples.

[0071] Figure 6 The diagram illustrates the principle of intra template matching. Intra Template Matching (Intra TMP) is another tool based on template matching. Intra TMP is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame whose L-shaped template matches the current template. In a predefined search range, the encoder searches for the template that is most similar to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. The encoder signals the use of this mode so that the same prediction operation is performed on the decoder side.

[0072] By combining the L-shaped causal neighboring blocks of the current block with Figure 6 The prediction signal is generated by matching another block in a predefined search area in the decoder, which includes 4 regions: R1 (current CTU), R2 (top left CTU), R3 (top CTU) and R4 (left CTU). The sum of absolute differences (SAD) is used as a cost function. In each region, the decoder searches for the template with the smallest SAD relative to the current template and uses its corresponding block as the prediction block. The dimensions of the region (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is:

[0073] SearchRange_w=a*BlkW

[0074] SearchRange_h=a*BlkH

[0075] Where 'a' is a constant that controls the gain / complexity tradeoff. For example, 'a' is equal to 5.

[0076] Intra TMP prediction mode is enabled for CUs with width and height dimensions less than or equal to 64. This maximum CU size for intra template matching is configurable.

[0077] When DIMD is not used for the current CU, the intra TMP prediction mode is signaled at the CU level through a dedicated flag.

[0078] Figure 7 The diagram illustrates the principle of template matching. Template matching (TM) is another tool based on template matching. This is a decoder-side motion vector (MV) derivation method to refine the motion information of the current CU by finding the closest match between the template of neighboring samples in the current picture and the block in the reference picture (i.e., the same size as the template). A better MV is searched around the initial motion of the current CU within the [–8, +8] pixel search range. In at least one embodiment of TM, the search step size is determined based on the adaptive motion vector resolution (AMVR) mode, and TM can be cascaded with the bilateral matching process in merge mode.

[0079] In adaptive motion vector prediction (AMVP) mode, MVP candidates are determined based on template matching error to select the MVP candidate that achieves the minimum difference between the current block template and the reference block template, and then TM is performed only on this specific MVP candidate for MV refinement. TM refines this MVP candidate by using iterative diamond search, starting from full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode) in the [–8, +8] pixel search range. The AMVP candidate can be further refined by using a cross search with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode) according to the AMVR mode as specified in Table 1, followed by half-pixel accuracy and quarter-pixel accuracy in sequence. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process. During the search process, if the difference between the previous minimum cost and the current minimum cost in the iteration is less than a threshold, which is equal to the area of ​​the block, the search process terminates.

[0080] Table 1 illustrates the search method of AMVR and the merging mode with AMVR.

[0081]

[0082]

[0083] Table 1

[0084] In merge mode, a similar search method is applied to the merge candidates indicated by the merge index. As illustrated in Table 1, depending on whether an alternative interpolation filter is used according to the merged motion information (which is used when AMVR is in half-pixel mode), TM can be performed all the way down to 1 / 8 pixel MVD accuracy or skip those exceeding half-pixel MVD accuracy. In addition, when TM mode is enabled, template matching can work as an independent process or an additional MV refinement process between a block-based bilateral matching (BM) method and a sub-block-based bilateral matching (BM) method, depending on whether the BM checked according to the BM enable condition can be enabled.

[0085] Figure 8 The reference area constraints for intra block copy in template matching mode are illustrated. Intra Block Copy with Template Matching Mode (IBC-TM) is another tool based on template matching. Template matching is used in IBC for both IBC merge mode and IBC AMVP mode. Compared to the list used by the normal IBC merge mode, the IBC-TM merge list is modified so that candidates are selected according to a pruning method using the motion distance between candidates as in the normal TM merge mode. The end zero motion implementation is replaced by motion vectors to the left (-W, 0), up (0, -H), and top left (-W, -H), where W is the width of the current CU and H is the height of the current CU.

[0086] In IBC-TM merge mode, a template matching method is used to refine the selected candidates before RDO or decoding process. IBC-TM merge mode has been competed with normal IBC merge mode and the TM-Merge flag is signaled.

[0087] In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM merge list. Each of these 3 selected candidates is refined using a template matching method and ranked according to their resulting template matching cost. Then, as usual, only the first two candidates are considered in the motion estimation process.

[0088] Template matching refinement for both IBC-TM merge mode and AMVP mode is very simple since IBC motion vectors are constrained to be (i) integers and (ii) within the reference region. Figure 8 Four example reference region constraints for the current CU position are illustrated: Regions marked with an "X" symbol are not considered for refinement with respect to the current CU position.

[0089] In IBC-TM merge mode, all refinements are performed with integer precision, while in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision depending on the AMVR value. Such refinements only access samples without interpolation. In both cases, the refined motion vectors and the template used in each refinement step must respect the reference region constraints.

[0090] The embodiments described below have been designed with the foregoing in mind. Conventional template-based tools are based on the assumption that an L-shaped template captures the statistics of the current block. This assumption is used to infer the best prediction mode (TIMD, DIMD), the best copy block (intra-frame template matching), or to refine the merged motion vector (inter-frame template matching and IBC template matching). However, in some cases, the template on the left or above can be distorted, for example, by edges that result in different statistics. This usually occurs when encoding non-camera captured content (game content, screen content).

[0091] Therefore, it is suggested to rely on partial templates (above or left) instead of the conventional L-shaped template, in other words, consider an "above" template comprising pixels adjacent to and above the block, or a "left" template comprising pixels adjacent to and to the left of the block. Multiple rows or lines of pixels can also be used for the template.

[0092] At least one embodiment proposes to allow the encoder to select the type of template for a template-based coding tool among an "above" template that includes pixels adjacent to and above the block, or a "left" template that includes pixels adjacent to and to the left of the block, or a combination of an "above" template and a "left" template. In the latter case, in at least one embodiment, the template may also include an upper left element, which is in fact the case in a conventional L-shaped template. The type of template is signaled in the encoded data, and the type of template is used by the decoder to select the appropriate type of template to perform the prediction as expected using the selected template-based tool.

[0093] Fig.9A A flowchart of an example of an encoding process for a template-based prediction mode according to at least one embodiment is illustrated. This encoding process 900 is, for example, Figure 3 The device 1000 Figure 1The method is implemented by the encoder 100 of FIG. 1 . For the current block, in step 911, the device selects a template of neighboring samples according to the type of template. In step 912, the reconstructed area is analyzed to find a template that matches the selected template. In step 913, a block corresponding to the matching template is selected, and the encoding cost for the block is determined according to the template-based prediction mode. In step 910, steps 911, 912, 913 are iterated for different template-based prediction modes (with different parameter sets when available) and different types of templates. In step 920, a prediction mode (using the type of the selected template) is selected based on the encoding cost. Note that these steps 910-920 can be performed in an RDO optimization, which is conventionally part of the encoder. In addition, we describe here the case where a template-based prediction mode is selected. If another mode that does not use a template is selected as the prediction mode for the block, steps 925 to 935 are replaced by conventional encoding steps according to the selected prediction mode. In step 925, the current block is predicted based on the selected prediction and associated parameters or related data (e.g., values ​​of samples of the block selected for prediction). In step 930, the current block is then encoded based on the predicted block, and in step 935, the type of template used is signaled for the current block based on the selected template-based prediction mode.

[0094] Fig. 9B A flowchart of an example of a decoding process for a template-based prediction mode according to at least one embodiment is illustrated. This decoding process 950 is, for example, Figure 3 The device 1000 Figure 1 The method is implemented by the encoder 100 of the embodiment of the present invention. In step 960, the device obtains information that signals the use of a template-based prediction mode for the current block and the type of template selected as described above. In step 965, the device determines the template of the neighboring samples of the current block. In step 970, the device finds a matching (e.g., best matching) template in the reconstructed area of ​​the image, and in step 975, selects parameters corresponding to the matching template based on the prediction mode. In at least one embodiment, the parameter set is a sample set corresponding to the block of the matching template. Then, in step 980, these parameters are used to predict the current block. In step 985, the current block is then decoded (reconstructed) according to the convention based on the predicted block.

[0095] Fig.10The type of template is illustrated. Element 1004 represents the current coding unit. It is surrounded by the upper template 1001 and the left template 1002 adjacent to the CU. The conventional 'L-shape' can be obtained by combining the upper template 1001 and the left template 1002 with the upper left area 1003. The templates shown in this figure are multiple pixels wide (or high), however the templates described in this document can also be templates with a single pixel width (or height).

[0096] In at least one embodiment, in addition to signaling at least the first element of the use of template-based tools, additional syntax elements are also signaled to indicate that the template-based tool uses the type of template selected from the upper template, the left template, or the two templates. For example, this signaling can be done at the advanced SPS or at the slice header or CU level. It is generated by the encoder device for the block of the image of the video and provided to the decoding device through the encoded stream to achieve the correct reconstruction of the encoded block of the image of the video. The binary form of this signaling is illustrated in Table 2.

[0097]

[0098]

[0099] Table 2

[0100] In at least one embodiment, the syntax elements indicating the type of template are optimized to reduce signaling overhead. In the case of non-square blocks, for wide blocks (width is greater than height), it is unlikely to use the left template because there are more elements available in the upper template. Similarly, for thin blocks, it is unlikely to use the upper template. Therefore, in at least one embodiment, it is recommended to use a single bit signaling to indicate whether the upper template and the left template are used, or whether only one of them is used. The selection between the upper template or the left template is based on the shape of the block: the upper template is used for wide blocks, and the left template is used for tall / narrow blocks. For square blocks, the upper template is used by default. The binary form of this signaling is illustrated in Table 3.

[0101]

[0102] Table 3

[0103] In a variant embodiment, the signaling form of Table 3 is used only when the proportions of the blocks comply with certain conditions, for example, only when the ratio of the width divided by the height when the width is greater than the height or the ratio of the height divided by the width when the height is greater than the width is higher than a predetermined threshold.

[0104] exist Fig.10In another variation shown in FIG. 1 , when an upper template 1001 and a left template 1002 are used, the template also includes an upper left area 1003, so the template is 'L-shaped'.

[0105] In another variation, the template is not limited to a single row, but uses multiple rows, for example the template above uses multiple rows of pixels, while the template on the left uses multiple columns of pixels.

[0106] Fig.11 An example of CTU partitioning with CU decoding order is illustrated. It illustrates the latency issue for Template Matching (TM) prediction mode.

[0107] At least one embodiment proposes to reduce the latency caused by the TM prediction mode by signaling the template used. In fact, TM introduces some latency in the decoding of the CU, because it needs to complete the reconstruction of the surrounding environment in order to be able to calculate the current template for a specific CU. For example, the current template of the 8th CU needs to wait for the complete reconstruction of the 5th CU, the 6th CU and the 7th CU. Therefore, the process of the 8th CU cannot start until the 7th CU has not been completely reconstructed.

[0108] In general, the 'above' CU is available, but the 'left' CU may be lost. In this embodiment, it is suggested to signal at the SPS, picture header or slice header level if both (left and above) templates or only the above template can be used for TM. The latter case allows to overcome the TM latency problem.

[0109] In a variant embodiment, since some CUs do not suffer from such latency issues (e.g., the 7th CU in the figure requires reconstruction of the first CU and the second CU), it is suggested to use a signal at the CU level to use the upper template or the left template or both templates. For example, 2, 3, 4, 6, 8, and 11 use the upper template; 10 uses the left template; and 1, 5, 7, 9 use both.

[0110] In another variant, in order to keep CTU independent, it is suggested to signal whether the upper template, the left template, both templates, or no template can be used for each CU. For example, 1, 2, 3, 4, 10 do not use templates; 6, 8, 9 and 11 use the upper template; 5 and 7 use both.

[0111] Fig.12 A flowchart of an example of an encoding process using intra template matching prediction mode according to at least one embodiment is illustrated. This encoding process 1200 is, for example, Figure 3 The device 1000 Figure 1The process is implemented by the encoder 100 of the embodiment of the present invention. The process operates on the current block of the image or video. In step 1210, the encoder selects the intra-frame template matching prediction mode and the type of template for the current block. In step 1220, the encoder predicts the block using the intra-frame template matching prediction mode based on the type of template. In step 1230, the encoder encodes the predicted block. In step 1240, the encoder provides encoding information for the current block, which encoding information includes at least the use of the template intra-frame matching prediction mode and the type of template.

[0112] Fig.13 A flowchart of an example of a decoding process using intra template matching prediction mode according to at least one embodiment is illustrated. This decoding process 1300 is, for example, Figure 3 The device 1000 Figure 2 The process is implemented by the decoder 200 of the embodiment of the present invention. The process operates on the current block of the image or video. In step 1310, the decoder obtains the encoding information for the current block, and the encoding information at least includes the use of the template intra-frame matching prediction mode and the type of the template. In step 1320, the decoder predicts the block using the intra-frame template matching prediction mode based on the type of the template. In step 1330, the decoder decodes the predicted block.

[0113] At least one example of an embodiment may relate to a device comprising an apparatus as described herein and at least one of the following: (i) an antenna configured to receive a signal comprising data representing image information; (ii) a band limiter configured to limit the received signal to a frequency band comprising data representing image information; and (iii) a display configured to display an image based on the image information.

[0114] At least one example of an embodiment may involve a device as described herein, wherein the device includes one of a television, a television signal receiver, a set-top box, a gateway device, a mobile device, a cellular phone, a tablet, a computer, a laptop, or other electronic devices.

[0115] In general, another example of an embodiment may involve a bitstream or signal formatted to include syntax elements and picture information, wherein the syntax elements are generated by processing based on any one or more of the examples of embodiments of the method according to the present disclosure, and the picture information is encoded.

[0116] In general, one or more other examples of embodiments may also provide a computer-readable storage medium (e.g., a non-volatile computer-readable storage medium) on which instructions for encoding or decoding picture information such as video data are stored according to the methods or apparatus described herein. One or more embodiments may also provide a computer-readable storage medium on which a bitstream generated according to the methods or apparatus described herein is stored. One or more embodiments may also provide methods and apparatus for sending or receiving a bitstream or signal generated according to the methods or apparatus described herein.

[0117] Many examples of the embodiments described herein are described in detail and, at least to illustrate individual features, are often described in a manner that may sound restrictive. However, this is for clarity of description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, embodiments, features, etc. can also be combined and interchanged with others described in earlier submissions.

[0118] Various embodiments relate to decoding. "Decoding" as used in the present application may encompass, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such a process also includes or alternatively includes a process performed by a decoder of the various embodiments described in the present application.

[0119] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or generally to a broader decoding process will be clear based on the context of the particular description and is considered to be well understood by those skilled in the art.

[0120] Various embodiments relate to encoding. In a manner similar to the discussion above about "decoding", "encoding" as used in this application can encompass, for example, all or part of the processes performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as segmentation, differential encoding, transforms, quantization, and entropy encoding.

[0121] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or generally to a broader encoding process will be clear based on the context of the particular description and is considered to be well understood by those skilled in the art.

[0122] Note that syntax elements as used herein are descriptive terms. Therefore, they do not exclude the use of other syntax element names.

[0123] When a figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.

[0124] In general, the examples of the embodiments, implementations and aspects described herein, etc., can be implemented in, for example, methods or processes, devices, software programs, data streams or signals. Even if only discussed in the context of a single implementation form (e.g., discussed only as a method), the implementation of the features discussed can also be implemented in other forms (e.g., devices or programs). The device can be implemented in, for example, appropriate hardware, software and firmware. One or more examples of the method can be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit or a programmable logic device. The processor also includes communication equipment, such as, for example, a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate the communication of information between terminal-users. In addition, the use of the term "processor" herein is intended to broadly cover various configurations of a processor or more than one processor.

[0125] Reference to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", and other variations thereof, means that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearance of the phrase "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation", and any other variations thereof throughout this application, are not necessarily all referring to the same embodiment.

[0126] Additionally, this application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0127] Additionally, the present application may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0128] Furthermore, the present application may refer to "receiving" various pieces of information. Like "accessing," receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" is generally referred to in one way or another during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0129] It is to be understood that, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," the use of any of the following " / ," "and / or," and "at least one of" is intended to encompass selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such wording is intended to encompass selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This can be extended to as many of the listed items as are apparent to one of ordinary skill in this and related arts.

[0130] As will be apparent to one of ordinary skill in the art, embodiments may generate various signals that are formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method, or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bitstream of the described embodiment. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is well known, a signal may be transmitted over a variety of different wired or wireless links. A signal may be stored on a processor readable medium.

[0131] Various embodiments are described herein. Features of these embodiments may be provided individually or in any combination across various claim categories and types.

Claims

1. A method comprising: For a current block of a picture, information representing the type of a template to be used in an intra-frame template matching prediction mode is provided, wherein the type of the template is selected from a first template, a second template, or a third template, wherein the first template includes a group of adjacent pixels above the current block, the second template includes a group of adjacent pixels on the left side of the current block, and the third template is a combination of the first template and the second template.

2. The method according to claim 1, wherein: The information representing the type of the template is encoded as two bits of information, the first bit indicating whether a combination of the first template and the second template or a single template is used, and the second bit indicating whether the first template or the second template is used.

3. The method according to claim 1, wherein: A first type and a second type of template are represented by the same element and are selected according to the size of the current block, and wherein information representing the type of the template is encoded as a single bit.

4. A method according to any one of the preceding claims, wherein: The third template also includes a group of neighboring pixels at the upper left of the current block.

5. A method according to any one of the preceding claims, wherein: The height of the first template and the width of the second template are greater than one pixel.

6. A method comprising: For the current block of the image, Selecting the intra-frame template matching prediction mode and the type of template based on the coding cost; Based on the type of the template, predicting a block using the intra template matching prediction mode; encoding the current block based on the predicted block; as well as Provide encoding information for the current block, the information including at least information representing the use of an intra-frame template matching prediction mode and information representing the type of a template selected from a first template, a second template, or a third template, wherein the first template includes a group of adjacent pixels above the current block, the second template includes a group of adjacent pixels on the left side of the current block, and the third template is a combination of the first template and the second template.

7. A device comprising a processor, the processor being configured to: for a current block of a picture, Selecting the intra-frame template matching prediction mode and the type of template based on the coding cost; Based on the type of the template, predicting a block using the intra template matching prediction mode; encoding the current block based on the predicted block; as well as Provide encoding information for the current block, the encoding information at least including information representing the use of an intra-frame template matching prediction mode and information representing the type of a template selected from a first template, a second template, or a third template, wherein the first template includes a group of adjacent pixels above the current block, the second template includes a group of adjacent pixels on the left side of the current block, and the third template is a combination of the first template and the second template.

8. A method comprising: For the current block of the image, Obtaining encoding information for the current block, the encoding information at least including information representing the use of an intra-frame template matching prediction mode and information representing the type of a template selected from a first template, a second template, or a third template, the first template including a group of adjacent pixels above the current block, the second template including a group of adjacent pixels on the left side of the current block, and the third template being a combination of the first template and the second template; Based on the type of the template, predicting a block using the intra template matching prediction mode; as well as The current block is decoded based on the predicted block.

9. A device comprising a processor, the processor being configured to: for a current block of a picture, Obtaining encoding information for the current block, the encoding information at least including information representing the use of an intra-frame template matching prediction mode and information representing the type of a template selected from a first template, a second template, or a third template, the first template including a group of adjacent pixels above the current block, the second template including a group of adjacent pixels on the left side of the current block, and the third template being a combination of the first template and the second template; Based on the type of the template, predicting a block using the intra template matching prediction mode; as well as The current block is decoded based on the predicted block.

10. A computer program product comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 7.

11. A non-transitory computer-readable medium storing executable program instructions, the executable program instructions causing a computer executing the instructions to perform the method according to any one of claims 1 to 7.

12. A bitstream representing a coded picture, the coded picture formatted according to any one of claims 1 to 5 to include a syntax element indicating the type of template of the current block.

13. A non-transitory computer-readable medium storing information representing a coded picture, the coded picture formatted to include a syntax element indicating the type of template of a current block according to the method of any one of claims 1 to 5.