Template matching prediction with multiple template types

By optimizing video block coding through template matching prediction technology and quadtree + multi-type tree partitioning, the problem of low efficiency in reducing redundant information in existing technologies is solved, thereby improving video coding efficiency and decoding quality.

CN121176004APending Publication Date: 2025-12-19OFINNO LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480016655.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-11
Filing Date
2024-01-04
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Existing video coding technologies are not very efficient at reducing redundant information in various scenarios, especially when processing mixed videos of screen content and natural scenes, where it is difficult to effectively utilize spatial and temporal redundant information.

Method used

Template matching prediction (TMP) technology is adopted. By performing template matching on video blocks, prediction is made by taking advantage of the repetition of screen content and the motion of natural scenes. Combined with quadtree and multi-type tree partitioning, the encoding and decoding process of blocks is optimized.

Benefits of technology

It improves video encoding efficiency, especially when processing videos that mix screen content and natural scenes, reducing redundant information transmission and improving encoding efficiency and decoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121176004A_ABST
    Figure CN121176004A_ABST
Patent Text Reader

Abstract

A decoder determines a first candidate template of a first candidate reference block (RB) from a first search region and a second candidate template of a second candidate RB from a second search region. The first candidate template corresponds to a current template of a current block (CB) flipped in a direction. The second candidate template corresponds to the current template. Based on the current template, a template match (TM) cost is calculated, the TM cost including a first TM cost of the first candidate template and a second TM cost of the second candidate template. A reference template is selected from the first candidate template and the second candidate template based on the TM cost. And decoding the CB based on the RB indicated by the reference template.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing related applications

[0002] This application claims the benefits of U.S. Provisional Application No. 63 / 437,118, filed January 4, 2023, and U.S. Provisional Application No. 63 / 438,319, filed January 11, 2023, each of which is incorporated herein by reference in its entirety. Attached Figure Description

[0003] Examples of several embodiments of the various embodiments of this disclosure are described herein with reference to the accompanying drawings.

[0004] Figure 1 An exemplary video coding / decoding system in which embodiments of the present disclosure can be implemented is shown.

[0005] Figure 2 An exemplary encoder in which embodiments of the present disclosure may be implemented is shown.

[0006] Figure 3 An exemplary decoder in which embodiments of the present disclosure can be implemented is shown.

[0007] Figure 4 An example quadtree partitioning of a write code tree block (CTB) according to an embodiment of the present disclosure is shown.

[0008] Figure 5 An embodiment according to this disclosure is shown. Figure 4 The corresponding quadtree for the CTB example quadtree partitioning.

[0009] Figure 6 Example binary and ternary tree partitions are shown according to embodiments of the present disclosure.

[0010] Figure 7 An example quadtree + multi-type tree partitioning of CTB according to an embodiment of the present disclosure is shown.

[0011] Figure 8 An embodiment according to this disclosure is shown. Figure 7 The CTB example quadtree + multi-type tree partition corresponds to the quadtree + multi-type tree.

[0012] Figure 9 An example set of reference samples determined for intra-frame prediction of the current block being encoded or decoded according to embodiments of the present disclosure is shown.

[0013] Figure 10A Thirty-five intra-frame prediction modes supported by HEVC are shown according to embodiments of this disclosure.

[0014] Figure 10BThe 67 intra prediction modes supported by HEVC are shown according to embodiments of the disclosure.

[0015] Figure 11 Example angular mode predictions from a current block of Figure 9 are shown according to embodiments of the disclosure.

[0016] Figure 12 Example angular mode predictions from a current block of Figure 9 are shown according to embodiments of the disclosure.

[0017] Figure 13A An example of inter prediction performed on a current block being coded in a current picture is shown according to embodiments of the disclosure.

[0018] Figure 13B Example horizontal and vertical components of a motion vector are shown according to embodiments of the disclosure.

[0019] Figure 14 An example of bi-prediction performed on a current block is shown according to embodiments of the disclosure.

[0020] Figure 15A Example locations of five spatial candidate neighboring blocks relative to a current block being coded are shown according to embodiments of the disclosure.

[0021] Figure 15B Example locations of two temporal collocated blocks relative to a current block being coded are shown according to embodiments of the disclosure.

[0022] Figure 16 An example of IBC applied to screen content is shown according to embodiments of the disclosure.

[0023] Figure 17A An example of a template matching prediction (TMP) mode for predicting or determining a current block (CB) is shown according to some embodiments.

[0024] Figure 17B An example of a template matching prediction (TMP) mode for predicting or determining a current block (CB) is shown according to some embodiments.

[0025] Figure 18 An example of a RRIBC applied to screen content is shown according to some embodiments.

[0026] Figure 19 An example TMP mode using a candidate template flipped in the horizontal direction is shown according to some embodiments.

[0027] Figure 20AAn example TMP mode using candidate templates flipped in the vertical direction is shown, in accordance with some embodiments.

[0028] Figure 20B An example TMP mode using candidate templates flipped in the vertical direction is shown, in accordance with some embodiments.

[0029] Figure 21 An example TMP mode using multiple types of candidate templates is shown, in accordance with some embodiments.

[0030] Figure 22 An example of template matching for TMP is shown, in accordance with some embodiments.

[0031] Figure 23 A flowchart of a method of coding (e.g., encoding or decoding) a current block (CB) using template matching prediction (TMP) with multiple template types, in accordance with some embodiments, is shown.

[0032] Figure 24 A block diagram of an example computer system in which embodiments of the disclosure can be implemented is shown. DETAILED DESCRIPTION

[0033] In the following description, numerous specific details are set forth to provide a thorough understanding of the present disclosure. However, it will be apparent to one skilled in the art that the present disclosure, including the structural, system, and method embodiments, can be practiced without these specific details. The description and representation are used by those skilled in the art to most effectively convey the substance of their work to others skilled in the art. In other instances, well-known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the present disclosure.

[0034] Reference in the specification to “one embodiment,” “an embodiment,” “an example embodiment,” etc., indicates that the described embodiment can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of those skilled in the art to effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.

[0035] Moreover, various embodiments can be described as a process or method or as a system or apparatus or as an apparatus or device. While the novel concepts can be implemented in virtually any type of apparatus or device, the present disclosure can be implemented in a number of different embodiments and in many different ways. In addition, the novel concepts can be implemented using software, hardware or a combination of software and hardware. If implemented in software, the novel concepts can be realized in a number of ways, including but not limited to, as a program of instructions executed by a data processor, as a program of instructions executed by the processor of a programmable data processing apparatus, or as a program of instructions executed by a general purpose computer or a computer system. The program of instructions can be stored in a computer readable memory or storage medium of any type desired.

[0036] The term "computer readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, and various other mediums capable of storing, containing or carrying instruction(s) and / or data. A computer readable medium can include a non-transitory medium in which data can be stored and not contained in a carrier wave and / or a transitory electronic signal. Examples of a non-transitory medium can include, but are not limited to, a magnetic, optical, or other disc storage, a flash memory, a memory or memories, etc. Computer readable media can be in the form of a non-transitory medium storing code and / or machine executable instructions that can represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0037] Furthermore, embodiments can be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer program product) can be stored in a computer-readable or machine-readable medium. A processor(s) can perform the necessary tasks.

[0038] Representing a video sequence in digital form can require a large number of bits. In many applications, the data size of a video sequence in digital form can be too large for storage and / or transmission. Video encoding can be used to compress the size of a video sequence to provide more efficient storage and / or transmission. Video decoding can be used to decompress a compressed video sequence for display and / or other forms of consumption.

[0039] Figure 1An exemplary video coding / decoding system 100 in which embodiments of the disclosure can be implemented is shown. The video coding / decoding system 100 includes a source device 102, a transmission medium 104, and a destination device 106. The source device 102 encodes a video sequence 108 into a bitstream 110 for more efficient storage and / or transmission. The source device 102 can store the bitstream 110 and / or transmit the bitstream to the destination device 106 via the transmission medium 104. The destination device 106 decodes the bitstream 110 to display the video sequence 108. The destination device 106 can receive the bitstream 110 from the source device 102 via the transmission medium 104. The source device 102 and the destination device 106 can be any of a wide range of devices, including desktop computers, laptop computers, tablet computers, smartphones, wearable devices, televisions, cameras, video gaming consoles, set-top boxes, or video streaming devices.

[0040] To encode the video sequence 108 into the bitstream 110, the source device 102 can include a video source 112, an encoder 114, and an output interface 116. The video source 112 can provide or generate the video sequence 108 from a capture of a natural scene and / or a synthetically generated scene. The synthetically generated scene can be a scene that includes computer-generated graphics or screen content. The video source 112 can include a video capture device, such as a video camera, a video archive that includes previously captured natural and / or synthetically generated scenes, a video feed interface to receive captured natural and / or synthetically generated scenes from a video content provider, and / or a processor to generate synthetic scenes.

[0041] As shown in Figure 1 A video sequence, such as the video sequence 108, can include a series of pictures (also referred to as frames). The video sequence can achieve the impression of motion when the pictures of the video sequence are presented in a constant or variable temporal succession. A picture can include one or more sample arrays of intensity values. The intensity values can be taken at a series of regularly spaced locations within the picture. A color picture typically includes a luma sample array and two chroma sample arrays. The luma sample array can include intensity values that represent the luminance (or luma component Y) of the picture. The chroma sample arrays can include intensity values that represent the blue and red components (or chroma components Cb and Cr) of the picture that are separated from the luma. Other color picture sample arrays are possible based on different color schemes, such as an RGB color scheme. For color pictures, a pixel can refer to all three intensity values at a given location in the three sample arrays used to represent a color picture. A monochrome picture includes a single luma sample array. For monochrome pictures, a pixel can refer to the intensity value at a given location in the single luma sample array used to represent a monochrome picture.

[0042] The encoder 114 can encode the video sequence 108 into the bitstream 110. To encode the video sequence 108, the encoder 114 can apply one or more prediction techniques to reduce redundant information in the video sequence 108. Redundant information is information that can be predicted at the decoder and, thus, can not need to be transmitted to the decoder to accurately decode the video sequence. For example, the encoder 114 can apply spatial prediction (e.g., intra-frame or intra prediction), temporal prediction (e.g., inter-frame or inter prediction), inter-layer prediction, and / or other prediction techniques to reduce redundant information in the video sequence 108. Prior to applying one or more prediction techniques, the encoder 114 can partition pictures of the video sequence 108 into rectangular regions called blocks. The encoder 114 can then encode the blocks using one or more of the prediction techniques.

[0043] For temporal prediction, the encoder 114 can search for a block similar to the block being encoded in another picture (also referred to as a reference picture) of the video sequence 108. The block determined during the search (also referred to as a prediction block) can then be used to predict the block being encoded. For spatial prediction, the encoder 114 can form a prediction block based on data from reconstructed neighboring samples of the block to be encoded within the same picture of the video sequence 108. A reconstructed sample refers to a sample that is encoded and then decoded. The encoder 114 can determine a prediction error (also referred to as a residual) based on a difference between the block being encoded and the prediction block. The prediction error can represent non-redundant information that can be transmitted to the decoder to accurately decode the video sequence.

[0044] The encoder 114 can apply a transform (e.g., a discrete cosine transform (DCT)) to the prediction error to generate transform coefficients. The encoder 114 can form the bitstream 110 based on the transform coefficients and other information used to determine the prediction block (e.g., a prediction type, a motion vector, and a prediction mode). In some examples, the encoder 114 can perform one or more of quantization and entropy coding on the transform coefficients and / or other information used to determine the prediction block prior to forming the bitstream 110 to further reduce the number of bits needed to store and / or transmit the video sequence 108.

[0045] Output interface 116 can be configured to write and / or store bitstream 110 onto transmission medium 104 for transmission to destination device 106. Additionally or alternatively, output interface 116 can be configured to transmit, upload, and / or stream bitstream 110 to destination device 106 via transmission medium 104. Output interface 116 can include a wired and / or wireless transmitter configured to transmit, upload, and / or stream bitstream 110 according to one or more proprietary and / or standardized communication protocols, such as Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, Third Generation Partnership Project (3GPP) standards, Institute of Electrical and Electronics Engineers (IEEE) standards, Internet Protocol (IP) standards, and Wireless Application Protocol (WAP) standards.

[0046] Transmission medium 104 can include a wireless, wired, and / or computer readable medium. For example, transmission medium 104 can include one or more wires, cables, air interfaces, optical fibers, and / or flash memories. Additionally or alternatively, transmission medium 104 can include one or more networks (e.g., the Internet) or file servers for storing and / or transmitting encoded video data.

[0047] To decode bitstream 110 into video sequence 108 for display, destination device 106 can include input interface 118, decoder 120, and video display 122. Input interface 118 can be configured to read bitstream 110 stored on transmission medium 104 by source device 102. Additionally or alternatively, input interface 118 can be configured to receive, download, and / or stream bitstream 110 from source device 102 via transmission medium 104. Input interface 118 can include a wired and / or wireless receiver configured to receive, download, and / or stream bitstream 110 according to one or more proprietary and / or standardized communication protocols, such as those mentioned above.

[0048] The decoder 120 can decode the video sequence 108 from the encoded bitstream 110. To decode the video sequence 108, the decoder 120 can generate prediction blocks for the pictures of the video sequence 108 in a similar manner as the encoder 114 and determine prediction errors for the blocks. The decoder 120 can generate the prediction blocks using the prediction types, prediction modes, and / or motion vectors received in the bitstream 110 and determine the prediction errors using the transform coefficients also received in the bitstream 110. The decoder 120 can determine the prediction errors by weighting the transform basis functions using the transform coefficients. The decoder 120 can combine the prediction blocks and the prediction errors to decode the video sequence 108. In some examples, the decoder 120 can decode a video sequence that approximates the video sequence 108 due to, for example, lossy compression of the video sequence 108 by the encoder 114 and / or errors introduced into the encoded bitstream 110 during transmission to the destination device 106.

[0049] The video display 122 can display the video sequence 108 to a user. The video display 122 can include a cathode ray tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light-emitting diode (LED) display, or any other display device suitable for displaying the video sequence 108.

[0050] It is noted that the video encoding / decoding system 100 is presented by way of example and not limitation. In Figure 1 In examples, the video encoding / decoding system 100 can have other components and / or be arranged differently. For example, the video source 112 can be external to the source device 102. Similarly, the video display 122 can be external to the destination device 106 or omitted entirely in cases where the video sequence is intended for consumption by a machine and / or storage. In another example, the source device 102 can further include a video decoder and the destination device 106 can include a video encoder. In such an example, the source device 102 can be configured to further receive encoded bitstreams from the destination device 106 to support bidirectional video transmission between the devices.

[0051] In Figure 1In the example of FIG. 1, the encoder 114 and the decoder 120 can operate according to any of a number of proprietary or industry standards for video coding. For example, the encoder 114 and the decoder 120 can operate according to one or more of the following: International Telecommunication Union Telecommunication Standardization Sector (ITU-T) H.263, ITU-T H.264 and MPEG-4 Visual (also known as Advanced Video Coding (AVC)), ITU-T H.265 and MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC)), ITU-T H.265 and MPEG-I Part 3 (also known as Versatile Video Coding (VVC)), WebM VP8 and VP9 codecs, and AOMedia Video 1 (AV1).

[0052] Figure 2 An exemplary encoder 200 in which embodiments of the disclosure can be implemented is shown. The encoder 200 encodes a video sequence 202 into a bitstream 204 for more efficient storage and / or transmission. The encoder 200 can be implemented in the video coding / decoding system 100 of FIG. 1, or in any of a number of different devices, including a desktop computer, a laptop computer, a tablet computer, a smartphone, a wearable device, a television, a camera, a video game console, a set-top box, or a video streaming device. The encoder 200 includes an inter prediction unit 206, an intra prediction unit 208, combiners 210 and 212, a transform and quantization unit (TR+Q) unit 214, an inverse transform and quantization unit (iTR+iQ) 216, an entropy coding unit 218, one or more filters 220, and a buffer 222. Figure 1

[0053] The encoder 200 can partition pictures of the video sequence 202 into blocks and encode the video sequence 202 on a block-by-block basis. The encoder 200 can perform a prediction technique on a block being encoded using the inter prediction unit 206 or the intra prediction unit 208. The inter prediction unit 206 can perform inter prediction by searching for a block in another reconstructed picture (also referred to as a reference picture) of the video sequence 202 that is similar to the block being encoded. A reconstructed picture refers to a picture that is encoded and then decoded. The block determined during the search (also referred to as a prediction block) can then be used to predict the block being encoded to remove redundant information. The inter prediction unit 206 can exploit temporal redundancy or similarity in scene content from picture to picture in the video sequence 202. For example, the scene content between pictures of the video sequence 202 can be similar except for differences due to motion or affine transformation of the screen content over time.

[0054] ​Intra prediction unit 208 can perform intra prediction by forming a prediction block based on data from reconstructed neighboring samples of the block to be coded within the same picture of video sequence 202. A reconstructed sample refers to a sample that has been coded and then decoded. Intra prediction unit 208 can exploit spatial redundancy or similarity in the scene content within a picture of video sequence 202 to determine the prediction block. For example, the texture of a region of scene content in a picture can be similar to the texture in the immediately surrounding regions of the scene content in the same picture.

[0055] After prediction, combiner 210 can determine a prediction error (also referred to as a residual) based on a difference between the block being coded and the prediction block. The prediction error can represent non-redundant information that can be transmitted to a decoder to accurately decode the video sequence.

[0056] Transform and quantization unit 214 can transform and quantize the prediction error. Transform and quantization unit 214 can transform the prediction error into transform coefficients by applying, for example, a DCT, to reduce the correlated information in the prediction error. Transform and quantization unit 214 can quantize the coefficients by mapping data of the transform coefficients to a set of pre-defined representative values. Transform and quantization unit 214 can quantize the coefficients to reduce irrelevant information in bitstream 204. Irrelevant information is information that can be removed from the coefficients without producing visible and / or perceptible distortion in video sequence 202 after decoding.

[0057] Entropy coding unit 218 can apply one or more entropy coding methods to the quantized transform coefficients to further reduce the bit rate. For example, entropy coding unit 218 can apply context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and syntax-based context-adaptive binary arithmetic coding (SBAC). The entropy coded coefficients are packaged to form bitstream 204.

[0058] Inverse transform and quantization unit 216 can inverse quantize and inverse transform the quantized transform coefficients to determine a reconstructed prediction error. Combiner 212 can combine the reconstructed prediction error with the prediction block to form a reconstructed block. Filter 220 can filter the reconstructed block using, for example, a deblocking filter and / or a sample adaptive offset (SAO) filter. Buffer 222 can store the reconstructed block for use in the prediction of one or more other blocks in the same and / or different pictures of video sequence 202.

[0059] Although Figure 2 Although not shown in FIG. 2, encoder 200 also includes an encoder control unit configured to control the encoding operations of encoder 200. For example, encoder control unit can control the encoding operations of intra prediction unit 208, combiner 210, transform and quantization unit 214, entropy coding unit 218, inverse quantization and inverse transform unit 216, combiner 212, and filter 220. Figure 2One or more of the units of the encoder 200 shown in FIG. 2. The encoder control unit can control one or more units of the encoder 200 such that the bitstream 204 is generated in accordance with the requirements of any of a number of proprietary or industry video coding standards. For example, the encoder control unit can control one or more units of the encoder 200 such that the bitstream 204 is generated in accordance with one or more of the ITU-T H.263, AVC, HEVC, VVC, VP8, VP9, and AV1 video coding standards.

[0060] Within the constraints of a proprietary or industry video coding standard, the encoder control unit can attempt to minimize or reduce the bit rate of the bitstream 204 and maximize or improve the reconstructed video quality. For example, the encoder control unit can attempt to minimize or reduce the bit rate of the bitstream 204 given a level below which the reconstructed video quality cannot fall, or attempt to maximize or improve the reconstructed video quality given a level above which the bit rate of the bitstream 204 cannot exceed. The encoder control unit can determine / control one or more of the following: the partitioning of the pictures of the video sequence 202 into blocks, whether the blocks are inter-predicted by the inter-prediction unit 206 or intra-predicted by the intra-prediction unit 208, the motion vectors used for the inter-prediction of the blocks, the intra-prediction modes of the plurality of intra-prediction modes used for the intra-prediction of the blocks, the filtering performed by the filter 220, and the one or more transform types and / or quantization parameters applied by the transform and quantization unit 214. The encoder control unit can determine / control the above based on determining / control how it affects the rate-distortion metric for the block or picture being encoded. The encoder control unit can determine / control the above to reduce the rate-distortion metric for the block or picture being encoded.

[0061] After being determined, the prediction type used to encode the block (intra- or inter-prediction), the prediction information for the block (intra-prediction mode if intra-predicted, motion vectors, etc.), and the transform and quantization parameters can be sent to the entropy coding unit 218 to be further compressed, thereby reducing the bit rate. For example, the entropy coding unit 218 can apply context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and syntax-based context-adaptive binary arithmetic coding (SBAC) to compress the prediction type used to encode the block (intra- or inter-prediction), the prediction information for the block (intra-prediction mode if intra-predicted, motion vectors, etc.), and the transform and quantization parameters. The prediction type, the prediction information, and the transform and quantization parameters can be packaged with the prediction error to form the bitstream 204.

[0062] It is noted that the encoder 200 is presented by way of example and not limitation. In other examples, the encoder 200 can have other components and / or arrangements. For example, Figure 2One or more of the components shown in FIG. 1 can optionally be included in the encoder 200, such as the entropy coding unit 218 and the filter 220.

[0063] Figure 3 An exemplary decoder 300 in which embodiments of the disclosure can be implemented is shown. The decoder 300 decodes a bitstream 302 into a decoded video sequence 304 for display and / or some other form of consumption. The decoder 300 can be implemented in the video coding / decoding system 100 in FIG. 1, or in any of a variety of different devices, including a desktop computer, a laptop computer, a tablet computer, a smartphone, a wearable device, a television, a camera, a video game console, a set-top box, or a video streaming device. The decoder 300 includes an entropy decoding unit 306, an inverse transform and quantization (iTR+iQ) unit 308, a combiner 310, one or more filters 312, a buffer 314, an inter prediction unit 316, and an intra prediction unit 318. Figure 1

[0064] Although Figure 3 not shown in FIG. 1, the decoder 300 also includes a decoder control unit configured to control one or more of the units of the decoder 300 shown in FIG. 1. The decoder control unit can control one or more units of the decoder 300 such that the bitstream 302 is decoded in accordance with the requirements of any of a number of proprietary or industry video coding standards. For example, the decoder control unit can control one or more units of the decoder 300 such that the bitstream 302 is decoded in accordance with one or more of the ITU-T H.263, AVC, HEVC, VVC, VP8, VP9, and AV1 video coding standards. Figure 3

[0065] The decoder control unit can determine / control one or more of whether a block is inter predicted by the inter prediction unit 316 or intra predicted by the intra prediction unit 318, a motion vector for inter prediction of a block, an intra prediction mode of a plurality of intra prediction modes for intra prediction of a block, filtering performed by the filters 312, and one or more inverse transform types and / or inverse quantization parameters to be applied by the inverse transform and quantization unit 308. One or more of the control parameters used by the decoder control unit can be packaged in the bitstream 302.

[0066] ​​Entropy decoding unit 306 can entropy-decode the bitstream 302. For example, entropy decoding unit 306 can apply context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), and syntax-based context-adaptive binary arithmetic coding (SBAC) to decompress the prediction type (intra- or inter-prediction) used to encode the block, prediction information for the block (intra-prediction mode if intra-predicted, motion vectors, etc.), and transform and quantization parameters. Inverse transform and quantization unit 308 can inverse quantize and inverse transform the quantized transform coefficients to determine a decoded prediction error. Combiner 310 can combine the decoded prediction error with a prediction block to form a decoded block. The prediction block can be generated by intra- block estimation unit 318 or inter-block estimation unit 316, as described above with respect to encoder 200 in FIG. 2. Filter 312 can filter the decoded block using, for example, a deblocking filter and / or a sample adaptive offset (SAO) filter. Buffer 314 can store the decoded block for use in predicting one or more other blocks in the same and / or different pictures of a video sequence in bitstream 302. Decoded video sequence 304 can be output from filter 312, as shown in FIG. 3. Figure 2 It is noted that decoder 300 is presented by way of example and not limitation. In other examples, decoder 300 can have other components and / or arrangements. For example, one or more of the components shown in FIG. 3 can optionally be included in decoder 300, such as entropy decoding unit 306 and filter 312. Figure 3

[0067] It is noted that decoder 300 is presented by way of example and not limitation. In other examples, decoder 300 can have other components and / or arrangements. For example, one or more of the components shown in FIG. 3 can optionally be included in decoder 300, such as entropy decoding unit 306 and filter 312. Figure 3

[0068] It is further noted that although not shown in FIGS. 2 and 3, each of encoder 200 and decoder 300 can include an intra-block copy unit in addition to the inter- block estimation unit and the intra-block estimation unit. The intra-block copy unit can perform similarly to the inter-block estimation unit, but predict blocks within the same picture. For example, the intra-block copy unit can exploit repeating patterns that occur in screen content. Screen content can include, for example, computer-generated text, graphics, and animation. Figure 2 Figure 3 As described above, video encoding and decoding can be performed on a block-by-block basis. Based on the content of a picture, the process of partitioning a picture into blocks can be adaptive. For example, larger block partitions can be used in regions of a picture that have a high level of homogeneity to improve coding efficiency.

[0069] As described above, video encoding and decoding can be performed on a block-by-block basis. Based on the content of a picture, the process of partitioning a picture into blocks can be adaptive. For example, larger block partitions can be used in regions of a picture that have a high level of homogeneity to improve coding efficiency.

[0070] ​​​In HEVC, a picture can be partitioned into non-overlapping square blocks comprising samples in a sample array, referred to as coding tree blocks (CTBs). A CTB can have a size of 2nx2n samples, where n can be specified by a parameter of the coding system. For example, n can be 4, 5, or 6. A CTB can be further partitioned by a recursive quadtree decomposition into coding blocks (CBs) of semi-vertically and semi-horizontally sized. The CTB forms the root of the quadtree. A CB that is not further partitioned as part of the recursive quadtree decomposition can be referred to as a leaf CB of the quadtree, otherwise as a non-leaf CB of the quadtree. A CB can have a minimum size specified by a parameter of the coding system. For example, the minimum size of a CB can be 4x4, 8x8, 16x16, 32x32, or 64x64 samples. For inter and intra prediction, a CB can be further partitioned into one or more prediction blocks (PBs) for performing inter and intra prediction. A PB can be a rectangular block of samples on which the same prediction type / mode can be applied. For transform, a CB can be partitioned into one or more transform blocks (TBs). A TB can be a rectangular block of samples for which a transform size is determined.

[0071] Figure 4 An example quadtree partitioning of the CTB 400 is shown. Figure 5 An example quadtree partitioning of the CTB 400 is shown. Figure 4 Figure 4 Figure 5 As shown in FIGS. 4A and 4B, the CTB 400 is first partitioned into four CBs of semi-vertically and semi-horizontally sized. Three of the resulting CBs of the first level partitioning of the CTB 400 are leaf CBs. The three leaf CBs of the first level partitioning of the CTB 400 are labeled as 7, 8, and 9 in FIGS. 4A and 4B, respectively. Figure 4 Figure 5 The non-leaf CB of the first level partitioning of the CTB 400 is partitioned into four CBs of semi-vertically and semi-horizontally sized. Three of the resulting CBs of the second level partitioning of the CTB 400 are leaf CBs. The three leaf CBs of the second level partitioning of the CTB 400 are labeled as 0, 5, and 6 in FIGS. 4A and 4B, respectively. Figure 4 Figure 5 Finally, the non-leaf CB of the second level partitioning of the CTB 400 is partitioned into four leaf CBs of semi-vertically and semi-horizontally sized. The four leaf CBs are labeled as 1, 2, 3, and 4 in FIGS. 4A and 4B, respectively. Figure 4 Figure 5 In summary, the CTB 400 is partitioned into 10 leaf CBs labeled as 0-9, respectively. A z-scan (from left to right, from top to bottom) can be used to scan the resulting quadtree partitioning of the CTB 400 to form a sequence order for encoding / decoding the CB leaf nodes.

[0072] In summary, the CTB 400 is partitioned into 10 leaf CBs labeled as 0-9, respectively. A z-scan (from left to right, from top to bottom) can be used to scan the resulting quadtree partitioning of the CTB 400 to form a sequence order for encoding / decoding the CB leaf nodes. Figure 4 Figure 5 ​​​​​​The numerical label of each CB leaf node can correspond to the encoding / decoding sequence order, where CB leaf node 0 is encoded / decoded first, and CB leaf node 9 is encoded / decoded last. Although Figure 4 and Figure 5 Not shown, but it should be noted that each CB leaf node may include one or more PBs and TBs.

[0073] In VVC, images can be partitioned in a similar way to HEVC. The image can first be partitioned into non-overlapping square CTBs. Then, the CTBs can be further partitioned into CBs of semi-vertical and semi-horizontal sizes using a recursive quadtree partitioning method. In VVC, the leaf nodes of a quadtree can be further partitioned into CBs of unequal sizes using binary or ternary tree partitioning. Figure 6 Example binary and ternary tree partitions are shown. A binary tree partition can divide the parent block in half along the vertical direction 602 or the horizontal direction 604. The resulting partition may be half the size of the parent block. A ternary tree partition can divide the parent block into three parts along the vertical direction 606 or the horizontal direction 608. In a ternary tree partition, the middle partition may be twice the size of the other two end partitions.

[0074] Because of the addition of binary and ternary tree partitioning, the block partitioning strategy in VVC can be called quadtree + multi-type tree partitioning. Figure 7 An example quadtree + multi-type tree partitioning of CTB 700 is shown. Figure 8 It shows Figure 7 The example quadtree of CTB 700 + the corresponding quadtree of multi-type tree partitioning + multi-type tree 800. Figure 7 and Figure 8 In the diagram, quadtree partitions are shown with solid lines, and multi-type tree partitions are shown with dashed lines. For ease of explanation, the CTB 700 is shown as having the same characteristics as... Figure 4 The quadtree partitioning described is the same as that of the CTB 400. Therefore, the description of quadtree partitioning for the CTB 700 is omitted. The description of the additional multi-type tree partitioning of the CTB 700 is relative to... Figure 4 The three leaf CBs shown have been further partitioned using one or more binary and ternary tree partitions. Figure 4 In Figure 7 The three leaves CB shown in the middle that are further divided are leaves CB 5, 8 and 9.

[0075] from Figure 4 Starting with leaf CB 5, Figure 7 This shows that the leaf CB is partitioned into two CBs based on a vertical binary tree partition. The two resulting CBs are... Figure 7 and Figure 8 Leaves CB are marked as 5 and 6 respectively. Regarding...Figure 4 Leaf CB 8 in Figure 7 This leaf CB is shown to be divided into three CBs based on vertical ternary tree partitioning. Two of the three resulting CBs are leaf CBs labeled 9 and 14 in Figure 7 and Figure 8 respectively. The remaining non-leaf CB is first divided into two CBs based on horizontal binary tree partitioning, one of the two CBs is leaf CB labeled 10, and the other of the two CBs is further divided into three CBs based on vertical ternary tree partitioning. The three resulting CBs are leaf CBs labeled 11, 12, and 13 in Figure 7 and Figure 8 respectively. Finally, with respect to Leaf CB 9 in Figure 4 Figure 7 This leaf CB is shown to be divided into three CBs based on horizontal ternary tree partitioning. Two of the three CBs are leaf CBs labeled 15 and 19 in Figure 7 and Figure 8 respectively. The remaining non-leaf CB is divided into three CBs based on another horizontal ternary tree partitioning. The three resulting CBs are all leaf CBs labeled 16, 17, and 18 in Figure 7 and Figure 8 respectively.

[0076] In summary, CTB 700 is divided into 20 leaf CBs labeled 0-19 respectively. A z-scan (from left to right, from top to bottom) can be used to scan the resulting quadtree+multi-type tree partitioning of CTB 700 to form a sequence order for encoding / decoding the CB leaf nodes. Figure 7 and Figure 8 The numerical labels of each CB leaf node in Figure 7 and Figure 8 may correspond to the sequence order for encoding / decoding, where CB leaf node 0 is encoded / decoded first and CB leaf node 19 is encoded / decoded last. Although not shown in

[0077] ​In addition to specifying various blocks (e.g., CTBs, CBs, PBs, TBs), HEVC and VVC define various units. While a block can include a rectangular region of samples in a sample array, a unit can include co-located blocks of samples from different sample arrays (e.g., luma and chroma sample arrays) that form a picture as well as syntax elements and prediction data for the blocks. A coding tree unit (CTU) can include co-located CTBs of different sample arrays and can form a complete entity in a coded bitstream. A coding unit (CU) can include co-located CBs of different sample arrays and syntax structures for coding the samples of the CBs. A prediction unit (PU) can include co-located PBs of different sample arrays and syntax elements for predicting the PBs. A transform unit (TU) can include TBs of different sample arrays and syntax elements for transforming the TBs.

[0078] It should be noted that the term “block” can be used in the context of HEVC and VVC to refer to any of a CTB, CB, PB, TB, CTU, CU, PU, or TU. It should also be noted that the term “block” can be used in the context of other video coding standards to refer to similar data structures. For example, the term “block” can refer to a macroblock in AVC, a macroblock or subblock in VP8, a superblock or subblock in VP9, or a superblock or subblock in AV1.

[0079] In intra prediction, samples of a block to be encoded (also referred to as a current block) can be predicted from samples of a column immediately to the left of the current block and samples of a row immediately above the current block. The samples from the immediately neighboring column and row can be collectively referred to as reference samples. Each sample of the current block can be predicted by projecting the position of the sample in the current block onto a point along the reference samples in a given direction (also referred to as an intra prediction mode). If the projection does not fall directly on a reference sample, the sample can be predicted by interpolating between the two closest reference samples to the projected point. A prediction error (also referred to as a residual) for the current block can be determined based on the difference between the predicted sample values and the original sample values.

[0080] At an encoder, this process of predicting samples and determining prediction errors based on the difference between predicted samples and original samples can be performed for a plurality of different intra prediction modes including the non-directional intra prediction mode. The encoder can select one of the plurality of intra prediction modes and its corresponding prediction error to encode the current block. The encoder can send an indication of the selected intra prediction mode and its corresponding prediction error to a decoder to decode the current block. The decoder can decode the current block by predicting samples of the current block using the intra prediction mode indicated by the encoder and combining the predicted samples with the prediction error.

[0081] Figure 9An example set of reference samples 902 for intra prediction determination for a current block 904 being coded or decoded is shown. In Figure 9 the current block 904 corresponds to Figure 7 block 3 of the partitioned CTB 700. As described above, the numerical labels 0-19 of the blocks of the partitioned CTB 700 can correspond to a coding / decoding sequence order, and in Figure 9 examples are used as such.

[0082] Given that the current block 904 is of size w x h samples, the reference samples 902 can extend on 2w samples of the row immediately above the top row of the current block 904, 2h samples of the column immediately to the left of the current block 904, and on the top-left corner sample of the current block 904. In Figure 9 examples, the current block 904 is square, so w = h = s. To construct the set of reference samples 902, available samples from neighboring blocks of the current block 904 can be used. For example, a sample can not be available for constructing the set of reference samples 902 if the sample would be outside the picture of the current block, the sample is part of a different slice of the current block (where the concept of slice is used), and / or the sample belongs to a block that has been inter coded and indicated as constrained intra prediction. When indicated as constrained intra prediction, the intra prediction can not depend on the inter predicted block.

[0083] In addition to the above, samples that can not be available for constructing the set of reference samples 902 include samples in blocks that have not been coded at the encoder and reconstructed or decoded at the decoder based on the coding / decoding sequence order. Such a restriction can allow the same prediction result to be determined at the encoder and the decoder. In Figure 9 examples, samples from neighboring blocks 0, 1, and 2 can be available for constructing the reference samples 902, assuming that these blocks are coded and reconstructed at the encoder and decoded at the decoder before the current block 904 is coded. This assumes that no other issues (such as those described above) hinder the availability of the samples from neighboring blocks 0, 1, and 2. However, due to the coding / decoding sequence order, a portion of the reference samples 902 from neighboring block 6 can not be available.

[0084] Unavailable reference samples in the reference samples 902 can be filled with available reference samples in the reference samples 902. For example, an unavailable reference sample can be filled with the nearest available reference sample determined by moving through the reference samples 902 in a clockwise direction from the location of the unavailable reference. If no reference sample is available, the reference samples 902 can be filled with a middle value of the dynamic range of the picture being coded.

[0085] It should be noted that the reference samples 902 can be filtered based on the size of the current block 904 being coded and the applied intra prediction mode. It should also be noted that Figure 9 Only one exemplary determination of reference samples for intra prediction of a block is shown. In some proprietary and industry video coding standards, the reference samples can be determined in a different manner than described above. For example, in other cases, multiple reference lines can be used, such as in VVC.

[0086] After determining the reference samples 902 and optionally filtering them, the samples of the current block 904 can be intra predicted based on the reference samples 902. Most encoders / decoders support multiple intra prediction modes according to one or more video coding standards. For example, HEVC supports 35 intra prediction modes, including a planar mode, a DC mode, and 33 angular modes. VVC supports 67 intra prediction modes, including a planar mode, a DC mode, and 65 angular modes. The planar mode and the DC mode can be used to predict smooth and gradually changing regions of a picture. The angular modes can be used to predict regions in a picture with directional structures.

[0087] Figure 10A The 35 intra prediction modes supported by HEVC are shown. The 35 intra prediction modes are identified by indices 0 to 34. Prediction mode 0 corresponds to the planar mode. Prediction mode 1 corresponds to the DC mode. Prediction modes 2-34 correspond to the angular modes. Prediction modes 2-18 can be referred to as horizontal prediction modes because the primary source of the prediction is in the horizontal direction. Prediction modes 19-34 can be referred to as vertical prediction modes because the primary source of the prediction is in the vertical direction.

[0088] Figure 10B The 67 intra prediction modes supported by VVC are shown. The 67 intra prediction modes are identified by indices 0 to 66. Prediction mode 0 corresponds to the planar mode. Prediction mode 1 corresponds to the DC mode. Prediction modes 2-66 correspond to the angular modes. Prediction modes 2-34 can be referred to as horizontal prediction modes because the primary source of the prediction is in the horizontal direction. Prediction modes 35-66 can be referred to as vertical prediction modes because the primary source of the prediction is in the vertical direction. Because blocks in VVC can be non-square, prediction modes 35-66 can be referred to as vertical prediction modes even though the primary source of the prediction is in the horizontal direction. Figure 10B Some of the intra prediction modes shown in

[0089] To further describe the application of an intra prediction mode to determine a prediction for a current block, reference is made to Figure 11 and Figure 12 In Figure 11 , from Figure 9The current block 904 and the reference samples 902 are shown in a two-dimensional x, y plane, where the samples can be referred to as p[x][y]. To simplify the prediction process, the reference samples 902 can be placed in two one-dimensional arrays. The reference samples 902 above the current block 904 can be placed in a one-dimensional array ref1[x]:

[0090] ref1[x] = p[-1+x][-1], (x ≥ 0) (1)

[0091] The reference samples 902 to the left of the current block 904 can be placed in a one-dimensional array ref2[y]:

[0092] ref2[y] = p[-1][-1+y], (y ≥ 0) (2)

[0093] For the planar mode, a sample at position [x][y] in the current block 904 can be predicted by computing an average of two interpolated values. A first interpolated value of the two interpolated values can be based on a horizontal linear interpolation at position [x][y] in the current block 904. A second interpolated value of the two interpolated values can be based on a vertical linear interpolation at position [x][y] in the current block 904. The predicted sample p[x][y] in the current block 904 can be computed as

[0094]

[0095] where

[0096] h[x][y] = (s-x-1) · ref2[y] + (x+1) · ref1[s] (4)

[0097] may be a horizontal linear interpolation at position [x][y] in the current block 904, and

[0098] v[x][y] = (s-y-1) · ref1[x] + (y+1) · ref2[s] (5)

[0099] may be a vertical linear interpolation at position [x][y] in the current block 904.

[0100] For the DC mode, a sample at position [x][y] in the current block 904 can be predicted by an average of the reference samples 902. The predicted value sample p[x][y] in the current block 904 can be computed as

[0101]

[0102] For angled patterns, a sample at position [x][y] in the current block 904 can be predicted by projecting position [x][y] onto a point on a horizontal or vertical line that includes reference sample 902 in the direction specified by the given angled pattern. If the projection does not fall directly on a reference sample, the sample at position [x][y] can be predicted by interpolating between the two nearest reference samples of the projection point. The direction specified by the angled pattern can be defined by an angle relative to the y-axis of vertical prediction patterns (e.g., patterns 19-34 in HEVC and patterns 35-66 in VVC) and the x-axis relative to horizontal prediction patterns (e.g., patterns 2-18 in HEVC and patterns 2-34 in VVC). Provided.

[0103] Figure 12 It shows the angle The given vertical prediction pattern 906 provides the prediction of the sample at position [x][y] in the current block 904. For the vertical prediction pattern, position [x][y] in the current block 904 is projected onto a point on the horizontal line of the reference sample ref1[x] (referred to as the "projection point" in this paper). For ease of illustration, the reference sample 902 is... Figure 12 Only a portion is shown. Because... Figure 12 In the example, the projection point falls at the fractional sample position between the two reference samples, so the predicted sample p[x][y] in the current block 904 can be calculated as follows by linear interpolation between the two reference samples.

[0104] p[x][y]=(1-i f )·ref1[x+i i +1]+i f ·ref1[x+i i +2] (7)

[0105] Where i i It is the integer part of the horizontal displacement of the projection point relative to the position [x][y], and can be used as the angle of the vertical prediction pattern 906. The tangent function is calculated as follows:

[0106]

[0107] And i f It is the fractional part of the horizontal displacement of the projection point relative to the position [x][y], and can be calculated as

[0108]

[0109] in It is rounded down (integer floor).

[0110] For the horizontal prediction mode, the position [x][y] of a sample in the current block 904 can be projected onto the vertical line of the reference sample ref2[y]. The sample prediction for the horizontal prediction mode is given by

[0111] p[x][y] = (1 - i f ) · ref2[y + i i + 1] + i f · ref2[y + i i + 2] (10)

[0112] where i i is the integer part of the vertical displacement of the projection point with respect to the position [x][y], and can be computed as a function of the tangent of the angle of the horizontal prediction mode as follows

[0113]

[0114] and i f is the fractional part of the vertical displacement of the projection point with respect to the position [x][y], and can be computed as

[0115]

[0116] where is the floor function.

[0117] The interpolation functions of (7) and (10) can be implemented by an encoder or decoder, such as the encoder 200 in Figure 2 or the decoder 300 in Figure 3 , as a set of two-tap Finite Impulse Response (FIR) filters. The coefficients of the two-tap FIR filters can be given by (1 - i f ) and i f , respectively. In the above example of angular intra prediction, the predicted sample p[x][y] can be computed with a certain pre-defined level of sample accuracy, such as 1 / 32 sample accuracy. For 1 / 32 sample accuracy, the set of two-tap FIR interpolation filters can include up to 32 different two-tap FIR interpolation filters - one for each of the 32 possible values of the fractional part of the projection displacement i f . In other examples, different levels of sample accuracy can be used.

[0118] In embodiments, two-tap interpolation FIR filters can be used to predict chroma samples. For luma samples, different interpolation techniques can be used. For example, for luma samples, a four-tap FIR filter can be used to determine the predicted value of a luma sample. For example, similar to the two-tap FIR filter, the four-tap FIR filter can have coefficients that are based on if The determined coefficients. For 1 / 32 sample accuracy, a set of 32 different four-tap FIR filters can include up to 32 different four-tap FIR filters - one for each of the 32 possible values of the fractional part of the projected displacement i f In other examples, different levels of sample accuracy can be used. The set of four-tap FIR filters can be stored in a look-up table (LUT) and referenced based on i f For vertical prediction modes, the value of the predicted sample p[x][y] can be determined based on a four-tap FIR filter as follows:

[0119]

[0120] where ft[i], i = 0...3, are the filter coefficients. For horizontal prediction modes, the value of the predicted sample p[x][y] can be determined based on a four-tap FIR filter as follows:

[0121]

[0122] It should be noted that supplemental reference samples can be constructed for cases where the position [x][y] of the sample to be predicted in the current block 904 is projected to a negative x coordinate, which occurs for negative vertical prediction angles Supplemental reference samples can be constructed by projecting the reference samples in ref2[y] in the vertical line of reference samples 902 to the horizontal line of reference samples 902 using negative vertical prediction angles The supplemental reference samples can be similarly used for cases where the position [x][y] of the sample to be predicted in the current block 904 is projected to a negative y coordinate, which occurs for negative horizontal prediction angles Supplemental reference samples can be constructed by projecting the reference samples in ref1[x] on the horizontal line of reference samples 902 to the vertical line of reference samples 902 using negative horizontal prediction angles

[0123] ​An encoder can predict samples of a current block, such as current block 904, being encoded for a number of intra prediction modes as described above. For example, an encoder can predict samples of a current block for each of the 35 intra prediction modes in HEVC or the 67 intra prediction modes in VVC. For each intra prediction mode applied, an encoder can determine a prediction error for the current block based on a difference (e.g., sum of squared differences (SSD), sum of absolute differences (SAD), or sum of absolute transformed differences (SATD)) between prediction samples determined for the intra prediction mode and original samples of the current block. An encoder can select one of the intra prediction modes to encode the current block based on the determined prediction errors. For example, an encoder can select an intra prediction mode that produces a smallest prediction error for the current block. In another example, an encoder can select an intra prediction mode to encode the current block based on a rate-distortion metric (e.g., Lagrangian rate-distortion cost) determined using the prediction errors. An encoder can send an indication of the selected intra prediction mode and its corresponding prediction error to a decoder to decode the current block.

[0124] Similar to an encoder, a decoder can predict samples of a current block, such as current block 904, being decoded for intra prediction modes as explained above. For example, a decoder can receive an indication of an angular intra prediction mode from an encoder of the block. The decoder can construct a set of reference samples and perform intra prediction based on the angular intra prediction mode indicated by the encoder of the block in a similar manner as discussed above for an encoder. The decoder adds the predicted values of the samples of the block to the residuals of the block to reconstruct the block. In another embodiment, a decoder can not receive an indication of an angular intra prediction mode from an encoder of the block. Instead, the decoder can determine the intra prediction mode by other decoder-side means.

[0125] While the above description is primarily with respect to intra prediction modes in HEVC and VVC, it should be understood that the techniques of the present disclosure described above and further described below can be applied to other intra prediction modes, including intra prediction modes of other video coding standards such as VP8, VP9, AV1, and the like.

[0126] As described above, intra-frame prediction can perform video compression by leveraging the correlation between spatially adjacent samples in the same frame of a video sequence. Inter-frame prediction is another coding tool that can be used to perform video compression by leveraging the temporal correlation between sample blocks in different frames of a video sequence. Generally, objects can be seen in multiple frames of a video sequence. The objects can move (e.g., by some kind of translation and / or affine motion) or remain stationary in the multiple frames. Therefore, the current sample block in the current frame being encoded may have a corresponding sample block in a previously decoded frame that accurately predicted the current sample block. Due to the movement of the objects represented in the two blocks in the corresponding frames of those blocks, the corresponding sample block may be shifted from the current sample block. The previously decoded frame may be referred to as the reference frame, and the corresponding sample block in the reference frame may be referred to as the reference block or motion-compensated prediction. The encoder can use block matching techniques to estimate displacement (or motion) and determine the reference block in the reference frame.

[0127] Similar to intra-frame prediction, once a prediction for the current block has been determined and / or generated using inter-frame prediction, the encoder can determine the difference between the current block and the prediction. This difference may be referred to as the prediction error or residual. The encoder can then store and / or signal the prediction error and other relevant prediction information in the bitstream for use in decoding or other forms of overhead. The decoder can decode the current block by using the prediction information to predict samples of the current block and combining the predicted samples with the prediction error.

[0128] Figure 13A An example of inter-frame prediction performed for the current block 1300 in the current image 1302 being encoded is shown. The encoder (such as...) Figure 2 The encoder 200 can perform inter-frame prediction to determine and / or generate a reference block 1304 in a reference picture 1306 to predict the current block 1300. The reference picture (such as reference picture 1306) is a previously decoded picture available at the encoder and decoder. The availability of the previously decoded picture may depend on whether it is available in the decoded picture buffer when the current block 1300 is encoded or decoded. The encoder can, for example, search for a reference block similar to the current block 1300 in one or more reference pictures. The encoder can determine the “best-matching” reference block as reference block 1304 from the blocks tested during the search process. The encoder can determine that reference block 1304 is the best-matching reference block based on one or more cost criteria (such as rate distortion criteria (e.g., Lagrange rate distortion cost)). One or more cost criteria may be based on, for example, the difference between the predicted sample of reference block 1304 and the original sample of the current block 1300 (e.g., sum of squared differences (SSD), sum of absolute differences (SAD), or sum of absolute transform differences (SATD)).

[0129] The encoder can search for the reference block 1304 within a search range 1308. The search range 1308 can be located around a collocated position (or block) 1310 of the current block 1300 in the reference picture 1306. In some cases, the search range 1308 can extend at least partially outside of the reference picture 1306. When extending outside of the reference picture 1306, a constant boundary extension can be used such that the value of a sample in the row or column of the reference picture 1306 immediately adjacent to the portion of the search range 1308 that extends outside of the reference picture 1306 is used for the “sample” location outside of the reference picture 1306. All potential locations or a subset of potential locations within the search range 1308 can be searched for the reference block 1304. The encoder can utilize any of a variety of different search implementations to determine and / or generate the reference block 1304. For example, the encoder can determine a set of candidate search locations based on motion information of neighboring blocks of the current block 1300.

[0130] The encoder can search one or more reference pictures during inter prediction to determine and / or generate a best matching reference block. The reference pictures searched by the encoder can be included in one or more reference picture lists. For example, in HEVC and VVC, two reference picture lists, reference picture list 0 and reference picture list 1, can be used. A reference picture list can include one or more pictures. The reference picture 1306 of the reference block 1304 can be indicated by a reference index that points to the reference picture list that includes the reference picture 1306.

[0131] The displacement between the reference block 1304 and the current block 1300 can be interpreted as an estimate of the motion between the reference block 1304 and the current block 1300 on their respective pictures. The displacement can be represented by a motion vector 1312. For example, the motion vector 1312 can be indicated by a horizontal component (MVx) and a vertical component (MVy) relative to the position of the current block 1300. Figure 13B The horizontal and vertical components of the motion vector 1312 are shown. Motion vectors, such as the motion vector 1312, can have fractional or integer resolution. Motion vectors with fractional resolution can point between two samples in a reference picture to provide a better estimate of the motion of the current block 1300. For example, a motion vector can have a fractional sample resolution of 1 / 2, 1 / 4, 1 / 8, 1 / 16, or 1 / 32. When a motion vector points to a non-integer sample value in a reference picture, interpolation between the samples at integer positions can be used to generate the reference block and corresponding samples at fractional positions thereof. The interpolation can be performed by a filter with two or more taps.

[0132] Once the reference block 1304 for the current block 1300 is determined and / or generated using inter prediction, the encoder can determine the difference (e.g., corresponding sample-wise difference) between the reference block 1304 and the current block 1300. The difference can be referred to as prediction error or residual. The encoder can then store and / or signal the prediction error and related motion information in the bitstream for decoding or other forms of consumption. The motion information can include the motion vector 1312 and the reference index pointing to the reference picture list including the reference picture 1306. In other cases, the motion information can include an indication of the motion vector 1312 and an indication of the reference index pointing to the reference picture list including the reference picture 1306. The decoder can decode the current block 1300 by determining and / or generating the reference block 1304 that forms the prediction of the current block 1300, using the motion information, and combining the prediction with the prediction error.

[0133] In Figure 13A , inter prediction is performed using one reference picture 1306 as the prediction source for the current block 1300. Because the prediction for the current block 1300 comes from a single picture, this type of inter prediction is referred to as uni-prediction. Figure 14 Another type of inter prediction, referred to as bi-prediction, is shown being performed for the current block 1400. In bi-prediction, the prediction source for the current block 1400 comes from two pictures. Bi-prediction can be useful, for example, in cases where the video sequence includes fast motion, camera panning or zooming, or scene changes. Bi-prediction can also be used to capture a fade-out of one scene or a fade-out from one scene to another scene, where the two pictures are effectively displayed at different intensity levels at the same time.

[0134] Whether uni-prediction or both uni-prediction and bi-prediction can be used to perform inter prediction can depend on the slice type of the current block 1400. For P slices, only uni-prediction can be used to perform inter prediction. For B slices, either uni-prediction or bi-prediction can be used. When uni-prediction is performed, the encoder can determine and / or generate a reference block for predicting the current block 1400 from the reference picture list 0. When bi-prediction is performed, the encoder can determine and / or generate a first reference block for predicting the current block 1400 from the reference picture list 0 and a second reference block for predicting the current block 1400 from the reference picture list 1.

[0135] In Figure 14 , inter prediction is performed using bi-prediction, where two reference blocks 1402 and 1404 are used to predict the current block 1400. The reference block 1402 can be in the reference picture of one of the reference picture lists 0 or 1, and the reference block 1404 can be in the reference picture of the other of the reference picture lists 0 or 1. As Figure 14As shown, the reference block 1402 is in a picture that precedes the current picture of the current block 1400 in terms of picture order count (POC), and the reference block 1402 is in a picture that follows the current picture of the current block 1400 in terms of POC. In other examples, the reference pictures can precede or follow the current picture in terms of POC. POC is the order in which pictures are output from, for example, a decoded picture buffer, and is the order in which pictures are generally intended to be displayed. However, it should be noted that output pictures are not necessarily displayed, but can undergo different processing or consumption, such as transcoding. In other examples, the two reference blocks determined and / or generated using bi-prediction can be from the same reference picture. In such cases, the reference picture can be included in both reference picture list 0 and reference picture list 1.

[0136] The configurable weight and offset values can be applied to one or more inter-predicted reference blocks. The encoder can use a flag in a picture parameter set (PPS) to enable the use of weighted prediction, and signal the weight and offset parameters in the slice segment header of the current block. Different weight and offset parameters can be signaled for the luma component and the chroma components.

[0137] Once the reference blocks 1402 and 1404 for the current block 1400 are determined and / or generated using inter-prediction, the encoder can determine the differences between the current block 1400 and each of the reference blocks 1402 and 1404. These differences can be referred to as prediction errors or residuals. The encoder can then store and / or signal the prediction errors and their respective associated motion information in the bitstream for decoding or other forms of consumption. The motion information for the reference block 1402 can include a motion vector 1406 and a reference index pointing to a reference picture list that includes the reference block 1402. In other cases, the motion information for the reference block 1402 can include an indication of the motion vector 1406 and an indication of the reference index pointing to the reference picture list that includes the reference block 1402. The motion information for the reference block 1404 can include a motion vector 1408 and a reference index pointing to a reference picture list that includes the reference block 1404. In other cases, the motion information for the reference block 1404 can include an indication of the motion vector 1408 and an indication of the reference index pointing to the reference picture list that includes the reference block 1404. A decoder can decode the current block 1400 using their respective motion information by determining and / or generating the reference blocks 1402 and 1404 that together form a prediction of the current block 1400, and combining the prediction with the prediction errors.

[0138] In HEVC, VVC, and other video compression schemes, motion information can be predicted and written before being stored in the bitstream or signaled. Motion information for the current block can be predicted and written based on the motion information of its neighboring blocks. Generally, the motion information of neighboring blocks is often related to the motion information of the current block because the motion of objects represented in the current block is usually the same as or similar to the motion of objects in neighboring blocks. Two motion prediction techniques in HEVC and VVC include Advanced Motion Vector Prediction (AMVP) and Inter-Frame Prediction Block Merging.

[0139] Encoders (such as) Figure 2 The encoder (200) can use the AMVP tool to write motion vectors as the difference between the motion vector of the current block being written and the predicted motion vector (MVP). The encoder can select an MVP from a list of candidate MVPs. Candidate MVPs can be previously decoded motion vectors from neighboring blocks in the current image of the current block, or previously decoded motion vectors from blocks at or near the juxtaposition position of the current block in other reference images. Both the encoder and decoder can generate or determine the list of candidate MVPs.

[0140] After the encoder selects an MVP from the list of candidate MVPs, it can signal the selected MVP and the indication of the motion vector difference (MVD) in the bit stream. The encoder can indicate the selected MVP in the bit stream by an index pointing to the list of candidate MVPs. The MVD can be calculated based on the difference between the motion vector of the current block and the selected MVP. For example, for a motion vector represented by the horizontal component (MVx) and vertical displacement (MVy) relative to the position of the current block being written, the MVD can be represented by two components calculated as follows:

[0141] MVD x =MV x -MVP x (15)

[0142] MVD y =MV y -MVP y (16)

[0143] MVD x and MVD y These represent the horizontal and vertical components of MVD, respectively, and MVP x and MVP y These represent the horizontal and vertical components of the MVP, respectively. Decoder (such as...) Figure 3The decoder (300) in the bitstream can decode the motion vector by adding the MVD to the MVP indicated in the bitstream. The decoder can then decode the current block by determining and / or generating a reference block that forms the prediction of the current block, using the decoded motion vector, and combining the prediction with the prediction error.

[0144] In HEVC and VVC, the list of candidate MVPs for AMVP can include two candidates, referred to as Candidate A and Candidate B. Candidate A and Candidate B can contain up to two spatial candidate MVPs derived from the five spatially adjacent blocks of the current block being written. When both spatial candidate MVPs are unavailable or identical, a temporal candidate MVP derived from two temporally co-located blocks can be included. Alternatively, when spatial, temporal, or both candidates are unavailable, a zero motion vector can be included. Figure 15A The positions of five spatial candidate neighbor blocks relative to the current block 1500 being encoded are shown. The five spatial candidate neighbor blocks are denoted as A0, A1, B0, B1, and B2. Figure 15B The positions of two time-coordinated blocks relative to the current block 1500 being written are shown. The two time-coordinated blocks are represented as C0 and C1, and are contained within a reference picture that is different from the current picture of the current block 1500.

[0145] Encoders (such as) Figure 2 The encoder 200 in the code can use an inter-frame prediction block merging tool (also known as merge mode) to write motion vectors. Using merge mode, the encoder can reuse the same motion information from neighboring blocks for inter-frame prediction of the current block. Because the same motion information from neighboring blocks is used, there is no need to signal the MVD, and the signaling overhead for signaling the motion information of the current block can be very small. Similar to AMVP, both the encoder and decoder can generate a candidate list of motion information based on the neighboring blocks of the current block. The encoder can then determine the motion information for the current block being written, using (or inheriting) the motion information from one of the neighboring blocks in the candidate list. The encoder can signal the indication of the determined motion information from the candidate list in the bitstream. For example, the encoder can signal an index to the list of candidate motion information to indicate the determined motion information.

[0146] In HEVC and VVC, the list of candidate motion information used for merging modes can include, for example, Figure 15A The AMVP shown derives up to four spatial merge candidates from five spatially adjacent blocks, as follows: Figure 15B The AMVP shown contains a time merge candidate derived from two time co-location blocks, as well as additional merge candidates including bidirectional prediction candidates and zero motion vector candidates.

[0147] It should be noted that inter prediction can be performed in other ways and variations other than the above described ways. For example, motion information prediction techniques other than AMVP and merge mode are possible. Also, while the above description is mainly with respect to inter prediction modes in HEVC and VVC, it should be understood that the techniques of the present disclosure described above and further described below can be applied to other inter prediction modes, including inter prediction modes of other video coding standards such as VP8, VP9, AV1, etc. Also, history-based motion vector prediction (HMVP) as described in VVC, combined intra / inter prediction mode (CIIP), and merge mode with motion vector difference (MMVD) can also be performed and are within the scope of the present disclosure.

[0148] In inter prediction, a block matching technique can be applied to determine a reference block in a different picture to a current block being coded. Block matching techniques have also been applied to determine a reference block in the same picture to a current block being coded. However, it has been determined that for camera captured video, the reference block in the same picture to the current block determined using block matching can often not accurately predict the current block. This is often not the case for screen content video. Screen content video can include, for example, computer generated text, graphics, and animations. In screen content, there are often repeating patterns within the same picture (e.g., of text and graphics). Thus, block matching techniques applied to determine a reference block in the same picture to a current block being coded can provide efficient compression of screen content video.

[0149] Both HEVC and VVC include a prediction technique that takes advantage of the correlation between sample blocks within the same picture of screen content video. This technique is referred to as intra block copy (IBC) or current picture reference (CPR). Similar to inter prediction, an encoder can apply a block matching technique to determine a displacement vector, referred to as a block vector (BV), that indicates a relative displacement from a current block to a reference block that “best matches” the current block (or intra block compensated prediction). The encoder can determine the best matching reference block from among blocks tested during a search process similar to inter prediction. The encoder can determine that a reference block is the best matching reference block based on one or more cost criteria, such as a rate-distortion criterion (e.g., a Lagrangian rate-distortion cost). The one or more cost criteria can be based on, for example, a difference (e.g., sum of squared differences (SSD), sum of absolute differences (SAD), sum of absolute transformed differences (SATD), or a difference determined based on a hash function) between predicted samples of the reference block and original samples of the current block. The reference block can correspond to a previously decoded sample block of the current picture. The reference block can include a decoded sample block of the current picture prior to being processed by in-loop filtering operations such as deblocking or SAO filtering. Figure 16An example of IBC applied to screen content is shown. The rectangular portion with an arrow starting from its boundary is the current block being coded, and the rectangular portion pointed to by the arrow is the reference block used to predict the current block.

[0150] Once the reference block for the current block is determined and / or generated using IBC, the encoder can determine the difference (e.g., corresponding sample-wise difference) between the reference block and the current block. The difference can be referred to as prediction error or residual. The encoder can then store and / or signal the prediction error and related prediction information in the bitstream for decoding or other forms of consumption. The prediction information can include the BV. In other cases, the prediction information can include an indication of the BV. A decoder (such as decoder 300 in FIG. 3) can decode the current block using the prediction information and combining the prediction with the prediction error by determining and / or generating the reference block that forms the prediction of the current block. Figure 3

[0151] In HEVC, VVC, and other video compression schemes, the BV can be predictively coded before being stored or signaled in the bitstream. The BV of a current block can be predictively coded based on the BVs of neighboring blocks of the current block. For example, the encoder can predictively code the BV using the merge mode as explained above for inter prediction or a technique similar to AMVP also explained above for inter prediction. The technique similar to AMVP can be referred to as BV prediction and difference coding.

[0152] For BV prediction and difference coding, an encoder (such as encoder 200 in FIG. 2) can code the BV as a difference between the BV of the current block being coded and a BV prediction value (BVP). The encoder can select the BVP from a list of candidate BVPs. The candidate BVPs can be from previously decoded BVs of neighboring blocks of the current block in the current picture. Both the encoder and the decoder can generate or determine the list of candidate BVPs. Figure 2

[0153] After the encoder selects the BVP from the list of candidate BVPs, the encoder can signal an indication of the selected BVP and the BV difference (BVD) in the bitstream. The encoder can indicate the selected BVP in the bitstream by an index pointing to the list of candidate BVPs. The BVD can be calculated based on the difference between the BV of the current block and the selected BVP. For example, for a BV represented by a horizontal component (BV x ) and a vertical component (BV y ) relative to the position of the current block being coded, the BVD can be represented by two components calculated as follows:

[0154] BVD x = BV x - BVP x (17)​​

[0155] BVD y =BV y -BVP y (18)

[0156] Among them BVD x and BVD y Let BVD represent the horizontal and vertical components, respectively, and BVP... x and BVP y These represent the horizontal and vertical components of the BVP, respectively. Decoders (such as...) Figure 3 The decoder 300 in the bitstream can decode the BV by adding the BVD to the BVP indicated in the bitstream. The decoder can then decode the current block by determining and / or generating a reference block that forms the prediction of the current block, using the decoded BV, and combining the prediction with the prediction error.

[0157] In HEVC and VVC, the list of candidate BVPs can include two candidates referred to as Candidate A and Candidate B. Candidate A and Candidate B can contain up to two spatial candidate BVPs derived from the five spatially adjacent blocks of the currently encoded block, or, when spatially adjacent candidates are unavailable (e.g., because they are written in intra-frame or inter-frame mode), one or more of the last two written code BVs. The positions of the five spatial candidate adjacent blocks relative to the current block encoded using IBC are... Figure 15A The locations shown are the same for inter-frame prediction. The five spatial candidate neighbor blocks are denoted as A0, A1, B0, B1, and B2. In other embodiments, the list of candidate BVPs may contain more than two candidate BVPs.

[0158] Template matching prediction (TMP) is a prediction method that can be performed iteratively by the encoder and decoder. In TMP, the reconstruction region of the current image can be searched to find the reference template of a reference block (RB) that is the "best match" to the current template of the current block (CB). For example, multiple candidate templates can be determined / searched from the reconstruction region, and the reference template can be determined from the reconstruction region based on the template matching (TM) cost calculated for the multiple candidate templates, as will be further described below. TMP performed on the same image frame as the current block can be referred to as intra-frame TMP mode or IBC with TMP. The reference template (of the RB) indicates the position of the RB in the reconstruction region, and the RB at this position can be used to predict the CB (by the encoder) or determine the CB (by the decoder). The block vector (BV) can be determined and indicates the displacement from the current block to the reference block. For ease of reference, mentioning predicting the CB can refer to the operation of the encoder, and mentioning determining the CB can refer to the operation of the decoder.

[0159] In some examples, the encoder can encode an indication (e.g., a syntax flag or signal) that indicates a reference block for the current block is determined by applying the TMP. Based on receiving and decoding the indication, the decoder can reciprocally apply the TMP to determine the same reference block for the current block. By adding the TMP mode for coding the current block, the indication of the BV with respect to the reference block of the current block can not need to be coded by the encoder and transmitted to the decoder, thereby reducing the information to be coded and improving compression efficiency. In some examples, the BV of the current block can be stored as associated with the current block to enable the BV of the current block to be used for predictively coding a next block, e.g., in IBC merge mode or IBC AMVP mode, as will be further described below.

[0160] Figure 17A An example of a template matching prediction (TMP) mode (also referred to as an intra-TMP mode) for predicting or determining a current block (CB) 1700 is shown in accordance with some embodiments. The CB 1700 includes a rectangular block of samples in a picture or video frame of a current picture 1702 that is to be encoded by an encoder or decoded by a decoder. To perform TMP to determine a reference block (RB) 1710 for the CB 1700, a coder (e.g., an encoder or a decoder) can determine or construct a current template 1708 for the CB 1700. The coder can determine or construct the current template 1708 based on samples in a reconstructed region. In an example, the current template 1708 can include samples in the reconstructed region that are adjacent to samples of the CB 1700. For example, the current template 1708 can include samples in the reconstructed region that are to the left and / or above the CB 1700. A block vector (BV) 1730 indicates a displacement from the CB 1700 to the determined RB 1710.

[0161] After determining or constructing the current template 1708 for the CB 1700, the coder can search a TMP search region 1706 (also referred to herein as a TMP reference region) for a reference template 1712 of a reference block (RB) (e.g., the RB 1710) determined to be a “best match” to the current template 1708 of the CB 1700. For example, the coder can determine a plurality of candidate templates corresponding to a plurality of respective candidate reference blocks (RBs) 1714 from which the reference template 1712 and the reference block (RB) 1710 can be determined. As shown in the example, the candidate templates of the candidate reference blocks 1714 match the current template 1708 in shape, orientation, and size. Figure 17A

[0162] ​In some examples, the writer can search the TMP search region 1706 for a candidate template of a candidate RB that best matches the current template 1708 by determining a cost between the current template 1708 and each of the candidate templates of the candidate reference blocks 1714 in the TMP search region 1706. In examples, the cost can be based on a difference (e.g., sum of squared differences (SSD), sum of absolute differences (SAD), sum of absolute transformed differences (SATD), or a difference determined based on a hash function) between the candidate template of the candidate RB and the current template 1708. In examples, the candidate template of the candidate RB that has the lowest cost (e.g., SSD, SAD, SATD, or hash function) with the current template 1708 can be determined as the reference template of the RB that best matches the current template 1708. Figure 17A In the example shown in FIG. 17, the reference template 1712 of the RB 1710 is determined as best matching the current template 1708 (e.g., based on a cost between the reference template 1712 and the current template 1708). A block vector (BV) can indicate a displacement of the RB (e.g., the RB 1710) relative to the CB (e.g., the CB 1700).

[0163] After determining the reference template 1712 of the RB 1710, the writer can use the RB 1710 to code the CB 1700. For example, the encoder can determine a difference (e.g., corresponding sample-wise difference) between the CB 1700 and the RB 1710. The difference can be referred to as a prediction error or residual. The encoder can store and / or signal the prediction error or residual in the bitstream for decoding by the decoder.

[0164] To perform TMP for determining the CB 1700, the decoder can perform the same (or reciprocal) operations as the encoder, as described above with respect to Figure 17A For example, based on receiving an indication from the encoder that the TMP is used to predict the CB 1700 (e.g., via a flag), the decoder can similarly determine or construct the current template 1708 of the CB 1700. After determining or constructing the current template 1708, the decoder can further similarly search the TMP search region 1706 for a reference template of a RB that is determined to best match the current template 1708. For example, the decoder can determine that the reference template 1712 of the RB 1710 best matches the current template 1708 from the candidate reference blocks 1714. After determining the reference template 1712 of the RB 1710, the decoder can use the RB 1710 (corresponding to the reference template 1712) to determine the CB 1700. For example, the decoder can combine the residual received from the encoder with the RB 1710 to reconstruct the CB 1700. Thus, the BV 1730 indicating the RB 1710 can not need to be indicated by the encoder to the decoder to enable the decoder to determine the RB 1710.

[0165] In some examples, the TMP search region 1706 includes a portion of the reconstructed region of the current picture 1702. The TMP search region 1706 indicates a region that an encoder or decoder can search to find a candidate template (e.g., a candidate template of the candidate reference block 1714) to determine the reference template 1712 and the region of the corresponding RB 1710. In some examples, the TMP search region 1706 can include region 1706A, region 1706B, region 1706C, and region 1706D. With respect to the CB 1700, region 1706A (R1) can include a portion of the current CTU, region 1706C (R2) can be the top-left CTU, region 1706B (R3) can be the above CTU, and region 1706D (R4) can be the left CTU. The CTUs are a result of the picture partitioning operations described in more detail above. For example, an encoder or decoder can search within the TMP search region 1706 for a matching template. For example, the reference template 1712 of the RB 1710 can be determined to be the best match to the current template 1708 of the CB 1700 based on a SAD cost or some other cost as described above. The decoder can use the RB 1710 to predict the CB 1700 as described above.

[0166] In some examples, the size of the TMP search region 1706 (referred to as SearchRange_w, SearchRange_h) can be set to have a fixed number of SAD comparisons (or other difference comparisons) per pixel proportionally to the size of the CB 1700 (referred to as BlkW, BlkH). More specifically, the size of the TMP search region 1706 can be calculated as follows:

[0167] SearchRange_w = a * BlkW (19)

[0168] SearchRange_h = a * BlkH (20)

[0169] where "a" (or a) is a constant that controls the gain / complexity tradeoff of the encoder or decoder. For example, "a" can equal 5. In Figure 17A It should be further noted that the size of the TMP search region 1706 is illustrated by way of example and not limitation. In actual implementations, for example, the size of the regions can vary, and / or one or more of the regions can not be present. In particular, Figure 17AIn the example shown, portions of the reconstructed region directly above and directly left of CB 1700 can not be available for prediction or determination, and can be excluded from TMP search region 1720. For example, this can be because the RBs in these portions would overlap with CB 1700, which would be an invalid location for predicting or determining CB 1700. Similar limitations can also be based on unavailability of samples due to sequence order of encoding or decoding, or due to samples being outside of the TMP search region or current picture 1702.

[0170] Figure 17B An example of a template matching prediction (TMP) mode for predicting or determining a current block (CB) 1700 is shown in accordance with some embodiments. As in Figure 17A TMP search region 1720 can not necessarily encompass one or more CTUs as compared to TMP search region 1706 in (20), as described above. For example, the size of TMP search region 1720 can be based on a fixed multiple of the size of CB 1700, as shown in (19) and (20). For example, search region height 1724 can correspond to (20), and each of search region width 1722, search region width 1726, and search region width 1728 can correspond to (19). TMP search region 1720 can include region 1720A (R1), region 1720B (R3), region 1720C (R2), and region 1720D (R4). As shown, one or more regions (or portions) of TMP search region 1720 can not encompass an entire CTU (e.g., region 1720D), and can span multiple CTUs (e.g., region 1720B and region 1720C). In Figure 17B In the example of (20), the same reference template 1712 from candidate reference block 1716 of RB 1710 can have been determined to be the best match to current template 1708. Similarly, BV 1730 indicates the displacement from CB 1700 to RB 1710.

[0171] Referring back to Figure 16 In IBC mode applied to screen content, a reference block (RB) can be determined to be the reference block that “best matches” a current block. For example, the arrows correspond to block vectors (BVs) that indicate the respective displacement from a respective current block (CB) to a respective reference block that best matches the respective current block. In Figure 16 In the example shown in (20), the reference blocks match the respective current blocks, and the computed residual would be small (if not zero). However, in general, video content can be more efficiently coded by taking into account symmetry properties. For example, it has been observed that symmetry often exists in video content, particularly in text character regions and computer-generated graphics in screen content video.

[0172] In the prior art, a reconfigured rearranged intra block copy, IBC, (RRIBC) mode (e.g., also referred to as IBC-mirror mode) is introduced for screen content video coding to exploit the symmetry within the video content to further improve the coding efficiency of IBC. For example, the RRIBC mode is adopted into the enhanced compression mode (ECM) software algorithm, which is currently being coordinated by the Joint Video Exploration Team (JVET) of ITU-T Video Coding Experts Group (VCEG) and ISO / IEC MPEG as a potential enhanced video coding technology beyond VVC capability. In some examples, the RRIBC mode can be signaled with an indication (or flag) based IBC mode indicating whether flipping is applied, and if flipping is applied, further signaling an indication (or flag) indicating the flipping direction.

[0173] In some embodiments, when the RIBC mode is indicated for encoding a current block, a residual of the current block can be calculated based on samples of a reference block flipped with respect to the current block (e.g., corresponding to the original reference block encoded and decoded to form the reconstructed block) according to the flipping direction indicated for the current block. In an example, at the encoder side, the current block (to be predicted) can be flipped before matching and residual calculation, while the reference block (used to predict the current block) can be derived without flipping. Similarly, at the decoder side, the current block (flipped at the encoder) can be determined based on the reference block and the residual information, and then flipped back to recover the original orientation of the current block before it was flipped at the encoder side. In another example, instead of the current block being flipped, the reference block can be flipped such that the reference block is flipped to encode the current block (at the encoder) and flipped back (at the decoder) to recover the original orientation of the reference block at the encoder. As described in this specification, a reference to flipping the current block can alternatively refer to flipping the reference block instead of the current block, such that the reference block and the current block are flipped in orientation with respect to each other.

[0174] In an example, in the RRIBC mode, the flipping direction can include one of the horizontal direction or the vertical direction of the RRIBC coded block. In an embodiment, for a current block coded in the RRIBC mode (e.g., an IBC advanced motion vector prediction (AMVP) coded block), a first indication (e.g., a first syntax flag) can indicate / signal whether flipping (e.g., also referred to as mirroring flipping) is used to encode / decode the current block. In addition, for the current block, a second indication (e.g., a second syntax flag) can indicate / signal the direction of flipping (e.g., vertical or horizontal). For IBC merge, the flipping direction can be inherited from the neighboring block without syntax signaling. In an example, for RRIBC, flipping of the current block (or the reference block in alternative embodiments) in the horizontal direction and the vertical direction can be represented in (21) and (22), respectively:

[0175] reference(x, y) = sample(w - 1 - x, y) (21)

[0176] reference(x, y) = sample(x, h - 1 - y) (22)

[0177] where w and h are the width and height of the current block, respectively. sample(x, y) can indicate the sample value located at (x, y). reference(x, y) can indicate the corresponding reference sample value after flipping. In other words, for horizontal flipping, (21) demonstrates flipping of the current block in the horizontal direction by sampling from right to left. Similarly, for vertical flipping, (22) demonstrates flipping of the current block in the vertical direction by sampling the current block from bottom to top.

[0178] Considering the horizontal or vertical symmetry, the current block and the reference block are usually aligned horizontally or vertically, respectively. Therefore, in an example, based on the RRIBC mode and the flipping direction, the reference block can be determined from the reference region (containing the candidate reference blocks) aligned in the same flipping direction, as will be further described below. Thus, when flipping in the horizontal direction is applied / indicated, the vertical component of the BV (BV y ) (indicating the displacement from the current block to the reference block) can not need to be signaled, as it can be inferred to be equal to 0. Similarly, when flipping in the vertical direction is applied / indicated, the horizontal component of the BV (BV x ) can not need to be signaled, as it can be inferred to be equal to 0. In other words, in an example, for the current block, only one component of the BV aligned with the flipping direction can be encoded and signaled.

[0179] Figure 18 An example of the RRIBC mode applied to screen content to exploit the symmetry within the text region to improve the efficiency of coding the video content is shown. Similar to the encoder (e.g., Figure 16 described in the description of FIG. 1, the encoder 100 can include a video encoder 102 and a video preprocessor 104. The video preprocessor 104 can be configured to perform the functions of the video preprocessor 104 described in the description of FIG. 1.Figure 1 The encoder can determine that the reference block 1804 is the best matching reference block for the current block 1802, e.g., based on applying the horizontal flipping. For example, the encoder can select the reference block 1804 as the“best matching” reference block based on one or more cost criteria (e.g., rate-distortion criteria), as described above. The one or more cost criteria can be applied to the reference block 1804 flipped in the horizontal direction with respect to the current block 1802. For example, the current block 1802 can be flipped before the one or more cost criteria are applied to determine the reference block 1804. Note that for flipping in the horizontal direction, the reference block 1804 is located in a reference region that is horizontally aligned with the current block 1802, as will be further described below. Thus, a block vector (BV) 1806 indicating the displacement between the current block 1802 and the reference block 1804 can be represented as a horizontal-only component (BV x ) of the BV 1806, since the vertical component of the BV 1806 will be equal to 0 when the horizontal flipping is indicated / applied due to the constraint on the possible positions of the reference block.

[0180] For a current block coded in IBC mode, the BV of the current block can be constrained to indicate the relative displacement from the current block to a reference block within an IBC reference region. In some examples, the BVP used to predictively code the BV can be similarly constrained. This is because the BVP can be derived from the BVs of spatial neighboring blocks of the current block or previously coded BVs as explained above. Based on the BVP, the BVD can be determined as the difference between the BV and the BVP. This BVD can be encoded and transmitted with an indication of the selected BVP in the bitstream to enable the decoding of the current block, as described above. With the introduction of RIBC, a reference block that is flipped in a certain direction with respect to a current block can be constrained to (i.e., selected from) an RIBC reference region corresponding to the direction, i.e., a subset or within the IBC reference region. As in IBC mode, a BVP can be used to predictively code the BV for the current block, indicating the relative displacement from the current block to a reference block within a reference region (e.g., RIBC region). Based on the indication of the RIBC mode and based on the flipping direction of the reference block with respect to the current block, the reference region corresponding to the flipping direction (e.g., RIBC reference region) can be determined. The reference region indicates the region within the picture frame from which the reference block can be selected (e.g., after flipping the CB).

[0181] In the prior art, as described above, the TMP mode can be applied to the current block to determine a reference block in IBC. For example, a reference template corresponding to a candidate reference block that "best" matches a current template of the current block can be searched to determine a reference block used to code the current block. In this TMP mode, the reference template matches the current template in size, shape, and orientation. However, this TMP mode searches for a reference block that best predicts the current block without considering horizontal or vertical symmetry of content within the picture or video frame.

[0182] In some embodiments, to exploit horizontal or vertical symmetry of content, the TMP mode can be enhanced by considering one or more other template types when searching for a reference block. For example, the TMP mode can use candidate templates that are flipped in a certain direction relative to the current template. These candidate templates can match the current template in shape and size, but not in orientation. For example, these candidate templates can match the current template in shape, size, and orientation after being flipped in the direction. In some examples, the TMP search area of the TMP mode using flipped templates can be extended compared to the reference / search area in the RRICB mode to search for candidate reference blocks that are misaligned in the same row or column as the current block. For example, since the TMP technique can have less computational cost compared to directly searching for candidate reference blocks, extending the TMP search area can increase compression by finding reference blocks that better match the current block with a small increase in processing and / or complexity cost.

[0183] In some embodiments, the BV can indicate a displacement from the current block to the determined reference block. This BV of the current block can be stored in association with the current block to enable the BV of the current block to be used to predictively code a next block, for example in IBC merge mode or IBC AMVP mode, as will be further described below.

[0184] Figure 19 An example TMP mode using candidate templates flipped in the horizontal direction is shown in accordance with some embodiments. For ease of comparison, Figure 19 A TMP search area 1906 containing regions 1906A-D is shown. The TMP search area 1906 can be the TMP search area 1706 or the TMP search area 1720. In some embodiments, the TMP mode, such as the TMP mode described above, can be enhanced to exploit horizontal or vertical symmetry of content. For example, the TMP mode can use candidate templates that are flipped in a certain direction relative to the current template. These candidate templates can match the current template in shape and size, but not in orientation. For example, these candidate templates can match the current template in shape, size, and orientation after being flipped in the direction. In some examples, the TMP search area of the TMP mode using flipped templates can be extended compared to the reference / search area in the RRICB mode to search for candidate reference blocks that are misaligned in the same row or column as the current block. For example, since the TMP technique can have less computational cost compared to directly searching for candidate reference blocks, extending the TMP search area can increase compression by finding reference blocks that better match the current block with a small increase in processing and / or complexity cost. Figure 17AThe example TMP mode shown in FIG. 19A can be enhanced by including other types of candidate templates. For example, candidate templates of candidate RB 1908 can be determined from TMP search region 1916 to determine TM cost, similar to that described with respect to candidate templates in TMP search region 1706, to determine a best matching candidate template as a reference template. For example, reference template 1904 of reference block 1902 can be determined to best match current template 1708. As shown, block vector (BV) 1930 indicates a displacement from CB 1700 to reference block 1902 determined using TMP mode using flipped candidate templates. Similar to TMP mode, for example Figure 17A and Figure 17B As shown in FIG. 20A, an encoder can signal TMP mode using candidate templates flipped in the vertical direction to enable a decoder to reciprocally perform TMP to determine reference block 2002 also determined by the encoder. Thus, BV 2030 can not need to be signaled to the decoder and improves compression efficiency of CB 1700.

[0185] In some examples, TMP search region 1916 can include region 1916A and region 1916B, as shown in FIG. 19A and further described in Figure 23 As shown in FIG. 19A and further described, search region width 1910 and search region height 1912 can be based on a fixed multiple of the width and height of CB 1700, as described above with respect to (19) and (20). In some examples, TMP search region 1916 includes search region 1916A that is not aligned with CB 1700 in the flipping direction compared to the search region for horizontally flipped RRICB.

[0186] Figure 20A An example TMP mode using candidate templates flipped in the vertical direction is shown in accordance with some embodiments. In some embodiments, the TMP mode (such as Figure 17A The example TMP mode shown in FIG. 19A can be enhanced by including other types of candidate templates. For example, candidate templates of candidate RB 1908 can be determined from TMP search region 1916 to determine TM cost, similar to that described with respect to candidate templates in TMP search region 1706, to determine a best matching candidate template as a reference template. For example, reference template 1904 of reference block 1902 can be determined to best match current template 1708. As shown, block vector (BV) 1930 indicates a displacement from CB 1700 to reference block 1902 determined using TMP mode using flipped candidate templates. Similar to TMP mode, for example Figure 17A andFigure 17B As shown, the encoder can use a candidate template flipped in the vertical direction to signal the TMP mode, enabling the decoder to repeatedly execute the TMP to determine the reference block 2002, which is also determined by the encoder. Therefore, the BV 2030 may not need to be signaled to the decoder, thus improving the compression efficiency of the CB 1700.

[0187] In some examples, TMP search region 2006 can contain region 2006A and region 2006B, such as Figure 23 As shown and further described herein, it can be based on a search region width 2010 and a search region height 2012. In some examples, the candidate template of candidate RB 2008 can be flipped vertically relative to the current template 1708. In some examples, the search region width 2010 and search region height 2012 can be based on a fixed multiple of the width and height of CB 1700, as described above with respect to (19) and (20). In some examples, the TMP search region 2006 includes a search region 2006A that is not aligned with CB 1700 in the flipping direction, compared to the search region of the RRIBC used for vertical flipping.

[0188] Figure 20B An example TMP pattern using a candidate template that is flipped in the vertical direction is shown according to some embodiments. Figure 20B The TMP search region 2026 for candidate templates flipped vertically relative to the current template 1708 is shown. As shown, the same reference template 2004 of reference block 2002 can be identified as... Figure 20A The reference template in [the document / reference template]. Figure 20A Compared to TMP search region 2006, TMP search region 2026 shows that TMP search region 2026 includes regions 2026A and 2026B. In some examples, depending on the search region width 2020 and search region height 2022, region 2006B may not necessarily intersect with the upper and / or left boundaries of the current CTU 1704, as will be explained below. Figure 23 Further details are provided below. Similarly, depending on the search region width of 1910 and the search region height of 1912, Figure 19 The TMP search region 1916 in the code may not necessarily intersect with the upper and / or left boundaries of the current CTU 1704, as will be explained below. Figure 23 Further details are provided below.

[0189] Figure 21 Examples of TMP modes (also known as TM modes or intra-frame TMP modes) using multiple types of candidate templates according to some embodiments are shown. For example, Figure 21It is shown that candidate templates of candidate reference blocks 1716 from the TMP search region 1906 and candidate templates of candidate RBs 2008 from the TMP search region 2006 can be determined to determine a best matching candidate template as a reference template (e.g., reference template 2004). For illustration purposes, Figure 21 It is shown that types of candidate templates, including non-flipped candidate templates and flipped candidate templates in the vertical direction, are searched for as described with respect to Figure 20A and Figure 20B However, other types of candidate templates, such as flipped candidate templates in the horizontal direction, can alternatively be searched for as described in Figure 19 In some examples, multiple types of candidate templates can include flipped candidate templates in the horizontal direction and flipped candidate templates in the vertical direction.

[0190] Figure 22 An example of template matching for TMP according to some embodiments is shown. In some embodiments, a shape of a current template is defined with respect to a current block, and can abut or surround the current block, but need not be positioned immediately adjacent to the current block. The shape can include a plurality of samples in a reconstructed portion of a picture frame. For example, the plurality of samples can include a plurality of reference pixels that have been reconstructed (e.g., encoded and then decoded) and distributed along at least one of two adjacent sides (e.g., a left side and an upper side) of the current block. The plurality of reference pixels of the current block can also be referred to as first reference pixels proximate to the current block. A pixel proximate to the current block can refer to a distance between the pixel and a side of the current block closest to the pixel being less than a threshold value. The distance between a pixel and a side of the coded block can be defined by a number or count of pixels between the pixel and the side of the current block. The threshold value can equal 1, or 2, or 3, or 4, etc.

[0191] In some examples, the current template can include a first portion and a second portion, where the first portion includes a plurality of rows of (adjacently reconstructed) samples above the current block, and the second portion includes a plurality of columns of (adjacently reconstructed) samples to the left of the current block. It should be appreciated that other types of shapes that include a set of reconstructed samples defined with respect to the current block are possible. In some examples, a candidate template can be compared to a current template by comparing a pair of samples from the candidate template and the current template, respectively, where the pair of samples are iterated in a mirrored fashion depending on a flipping direction.

[0192] For example, when the direction is horizontal flipping and for a current template of template size T s , the position (x c , y c ) is the top-left corner of the current block of size W x H and the position (x ref , y ref) is the top-left corner of the reference block, a pair of samples for the second portion of the current template and the corresponding portion of the reference template is defined as {(x c -1-j, y c +i), (x ref +W+j, y ref +i)}, where j e [0, T s ), i e [0, H). For example, the size T s may be the width of the second portion. In some examples, samples in the first portion of the current template can be similarly compared to samples in the corresponding portion of the candidate template. For example, a pair of samples for the first portion of the current template and the corresponding portion of the reference template is defined as {(x c +j, y c -1-i), (x ref +W-1-j, y ref -1-i)}, where j e [0, W), i e [0, T s ). Here, the size T s may be the height of the first portion.

[0193] For example, when the direction is a vertical flip and for a current template having a template size T s , the position (x c , y c ) is the top-left corner of the current block of size W x H and the position (x ref , y ref ) is the top-left corner of the reference block, a pair of samples for the second portion of the current template and the corresponding portion of the reference template is defined as {(x c -1-j, y c +i), (x ref -1-j, y ref +H-1-i)}, where j e [0, T s ), i e [0, H). For example, the size T s may be the width of the second portion. In some examples, samples in the first portion of the current template can be similarly compared to samples in the corresponding portion of the candidate template. For example, a pair of samples for the first portion of the current template and the corresponding portion of the reference template is defined as {(x c +j, y c -1-i), (x ref +j, y ref +H+i)}, where j e [0, W), i e [0, T s ). Here, the size T s may be the height of the first portion.

[0194] Figure 22An example of template matching between a current template 2206 of a current block 2202 and a candidate template 2208 of a reference block (RB) candidate 2204A is shown in accordance with some embodiments. As shown in Figure 22 For the case of horizontal flipping, samples P idx and R idx of the current template 2206 of the current block 2202 (to be predicted) and the candidate template 2208 of the RB candidate 2204 are shown in idx For comparison of the samples of these templates to compute the matching / comparison cost, the difference between pairs of samples can be computed as:∑|P idx -R HOR |, idx = {idx VER}, where idx HOR = {D HOR ,n}, D e {"A,"B,"C,"D"}, n e [0, cbWidth - 1]; and idx VER = {D VER ,m}, D e {"E,"F,"G,"H"}, m e [0, cbHeight - 1]. In some embodiments, how portions of the current template 2206 and the candidate template 2208 will be compared is based on the distance 2210 as shown. Figure 22 The comparison shown in

[0195] Figure 23 A flowchart 2300 of a method of coding (e.g., encoding or decoding) a current block (CB) using template matching prediction (TMP) with multiple template types is shown in accordance with some embodiments. For example, a CB can be coded in a TMP mode using at least two template types. For example, a CB can be coded in a TMP mode that uses a candidate template that is flipped in a certain direction relative to a current template of the CB, as further described below. The method of flowchart 2300 can be implemented by a coder such as an encoder (e.g., encoder 200 in Figure 2 ) or a decoder (e.g., decoder 300 in Figure 3 ). In other words, Figure 23 The method shown in

[0196] Some steps of flowchart 2300 can not necessarily be in the same order as understood by those skilled in the art.

[0197] At block 2302, first candidate templates of first candidate reference blocks (RBs) are determined from a first search region (which can alternatively be referred to as a first reference region). Each of the first candidate templates corresponds to a flipped current template of the CB in a direction.

[0198] In some examples, the current template is defined with respect to the CB. The current template can include a set of reconstructed samples, such as reconstructed pixels, that are adjacent to the CB. In some examples, the current template can have an "L" shape. For example, the current template can include a first portion that includes a number of rows (e.g., 1, 2, 4, etc.) of samples above the CB and a second portion that includes a number of columns (e.g., 1, 2, 4, etc.) of samples to the left of the CB. In examples, the first portion can be adjacent to a top side of the CB and the second portion can be adjacent to a left side of the CB. In examples, the rows match the CB in width and the columns match the CB in height.

[0199] In some examples, the first candidate templates are defined with respect to respective first RB candidates. In some examples, each of the first candidate templates that corresponds to the flipped current template includes each of the first candidate templates that match the current template in shape and orientation after flipping in the direction. Each of the first candidate templates can further match the flipped current template in size. In some examples, each of the first candidate templates only differs from the flipped current template in position (or location) in the picture frame in geometry.

[0200] In some examples, based on the direction being horizontal, each of the first candidate templates includes a number of rows of samples above the respective candidate template and a number of columns of samples to the right of the respective candidate template. Similarly, based on the direction being vertical, each of the first candidate templates can include a number of rows of samples below the respective first candidate template and a number of columns of samples to the left of the respective first candidate template.

[0201] At block 2304, second candidate templates of second candidate RBs are determined from a second search region (which can alternatively be referred to as a second reference region). Each of the second candidate templates corresponds to the current template of the CB. For example, each of the second candidate templates can correspond to the current template without flipping and / or other transformations other than translation.

[0202] In some examples, the second candidate templates are defined relative to the respective second RB candidates. In some examples, each of the second candidate templates corresponding to the current template includes each of the second candidate templates that match the current template in shape and orientation. Each of the second candidate templates can further match the current template in size. In some examples, each of the second candidate templates differs from the current template only in position (or location) in the picture frame in geometry.

[0203] In some examples, each of the second candidate templates includes a plurality of rows of samples above the respective candidate template and a plurality of columns of samples to the left of the respective candidate template.

[0204] In some examples, the first search region and the second search region are each a portion of a picture frame (alternatively referred to as a video frame) of the CB. Each portion corresponds to a reconstructed portion (when performed by an encoder) or a decoded portion (when performed by a decoder) of the CB. The current template, each of the first candidate templates, and each of the second candidate templates can include a respective set of reconstructed (or decoded) samples.

[0205] In some examples, the first search region is different from the second search region. For example, the second search region includes a portion of the current coding tree unit (CTU) of the CB that is not included (or excluded) in the first search region, as shown in the examples of Figure 19 , Figure 20A and / or Figure 20B In some examples, the second search region is defined relative to a location of the CB (or relative to a current coding tree unit (CTU) of the CB). For example, as described with respect to Figure 17A , the second search region includes a first CTU above and adjacent to a current CTU in which the CB is located, a second CTU to the left and adjacent to the current CTU, a third CTU above and to the left of the current CTU, where the third CTU is adjacent to the first CTU and the second CTU, and a portion of the current CTU above and to the left of the CB. In other examples, the second search region can include a plurality of rectangles having dimensions determined based on dimensions of the CB.

[0206] In some examples, the first search region includes a first rectangular region above and to the left of the CB, and a second rectangular region above or to the left of the CB depending on the direction. The first rectangular region can be within a current CTU of the CB. In some examples, the second search region can overlap the first search region in at least the first rectangular region of the first search region. As shown in the examples of Figures 19 to 21 , the first rectangular region can have a lower right corner that intersects (or meets) a top left corner of the CB. In some examples, as shown in the examples of Figures 19 to 21As described in the detailed description, the first rectangular region includes a first width and a first height that can be based on the width and the height of the CB, respectively. For example, based on the direction being horizontal, the height of the first rectangular region can be a predefined multiple of the height of the CB. In some examples, based on the direction being horizontal, the height of the first rectangular region can be a smaller one of a predefined multiple of the height of the CB or a distance between a top side (also referred to as an upper boundary) of the CB and the current CTU of the CB. In some examples, based on the direction being horizontal, the width of the first rectangular region can be a smaller one of a predefined multiple of the width of the CB and a distance between a left side (also referred to as a left boundary) of the CB and the current CTU of the CB. In some examples, the width of the first rectangular region can be the distance between the left side of the CB and the current CTU.

[0207] Similarly, based on the direction being vertical, the width of the first rectangular region can be a predefined multiple of the width of the CB. In some examples, based on the direction being vertical, the width of the first rectangular region can be a smaller one of a predefined multiple of the width of the CB or a distance between a left side (or a left boundary) of the CB and the CTU of the CB. In some examples, based on the direction being vertical, the height of the first rectangular region can be a smaller one of a predefined multiple of the height of the CB and a distance between a top side (also referred to as an upper boundary) of the CB and the current CTU. In some examples, the height of the first rectangular region can be the distance between the top side of the CB and the current CTU.

[0208] In some examples, the second rectangular region includes a second width and a second height that are based on the width and the height of the CB, respectively. For example, as shown in Figure 19 As shown in the detailed description, based on the direction being horizontal, the second rectangular region can be adjacent to and on the left side of the CB. For example, the second rectangular region can include a second height that is the same as the height of the CB and include a second width that is based on the width of the CB.

[0209] For example, as shown in Figure 20A and / or Figure 20B As shown in the detailed description, based on the direction being vertical, the second rectangular region can be adjacent to and above the CB. For example, the second rectangular region can include a second width that is the same as the width of the CB and include a second height that is based on the height of the CB.

[0210] At block 2306, a template matching (TM) cost is calculated based on the current template, where the TM cost includes a first TM cost of the first candidate template and a second TM cost of the second candidate template. For example, the first TM cost belongs to the first candidate template, respectively, and the second TM cost belongs to the second candidate template, respectively.

[0211] In some examples, computing the TM cost includes comparing samples in each of the first candidate templates to corresponding samples in the current template flipped in the direction to compute respective first TM costs. Each of the first TM costs can be based on one or more cost criteria, such as a rate-distortion criterion (e.g., a Lagrangian rate-distortion cost), in some examples. The one or more cost criteria can be based on, for example, a difference (e.g., a sum of squared differences (SSD), a sum of absolute differences (SAD), or a sum of absolute transformed differences (SATD)) between the prediction / reference samples of each of the first candidate templates and the samples of the current template flipped in the direction.

[0212] In some examples, computing the TM cost includes comparing samples in each of the second candidate templates to corresponding samples in the current template to compute respective second TM costs. Each of the second TM costs can be based on one or more cost criteria, such as a rate-distortion criterion (e.g., a Lagrangian rate-distortion cost), similar to the costs for computing the first TM costs. The one or more cost criteria can be based on, for example, a difference (e.g., a sum of squared differences (SSD), a sum of absolute differences (SAD), or a sum of absolute transformed differences (SATD)) between the prediction / reference samples of each of the second candidate templates and the corresponding samples of the current template.

[0213] At block 2308, a reference template is selected from (e.g., at least) the first candidate template and the second candidate template based on the TM costs. In some examples, the reference template can be the candidate template selected from at least the first candidate template and the second candidate template as having a smallest TM cost among the TM costs. In other words, the reference template can be determined from the candidate templates under test as the template that "best matches" the reference template, meaning that the residual between the CB (after flipping in the direction if the reference template is from the first candidate template) and the RB (indicated by the reference template) can be reduced compared to using other blocks indicated by other candidate templates.

[0214] At block 2310, the CB is coded based on the reference block (RB) indicated by the reference template. For example, the CB can be predicted (by an encoder) or determined (by a decoder) based on the RB. The RB can be the block from which the reference template is defined.

[0215] In some embodiments, at the encoder, the RB is used to predict the CB, and coding the CB includes encoding the CB based on the RB. In some examples, encoding the CB can include determining whether to flip the CB in the direction before determining a residual for the CB based on whether the reference template is from the first candidate template or the second candidate template. The residual for the CB can then be determined based on the determining whether to flip the CB, the RB, and the CB. The determined residual (for the CB) can be transmitted in the bitstream. In some examples, based on the reference template being from the first candidate template, the residual can be determined as a difference between the RB and the CB flipped in the direction. In some examples, based on the reference template being from the second candidate template, the residual can be determined based on a difference between the RB and the CB.

[0216] In some embodiments, at the encoder, coding the CB can include transmitting in the bitstream an indication that the CB is encoded in a template matching prediction (TMP) mode that uses multiple types of candidate templates, such as using a candidate template flipped in the direction relative to a current template. The determining of the first candidate template (at block 2302) and the determining of the second candidate template (at block 2304) can be in response to the encoder being in (or operating in) such a TMP mode (also referred to herein as a reconstruction rearranged TMP mode or RR-TMP mode).

[0217] In some examples, encoding the CB can include determining a residual based on a difference between the RB and the CB flipped in the direction based on the reference template being from the first candidate template, and determining a residual based on a difference between the RB and the CB based on the reference template being from the second candidate template. The residual for the CB can then be transmitted in the bitstream.

[0218] In some embodiments, at the decoder, the RB is used to determine the CB, and coding the CB includes decoding the CB based on the RB. In some examples, decoding the CB can include receiving a residual for the CB from the bitstream. A reconstructed block can be determined based on combining the RB with the residual for the CB received from the bitstream. Whether to flip the reconstructed block in the direction can be determined based on whether the reference template is from the first candidate template or the second candidate template. For example, based on the reference template being one of (or from) the first candidate template, the decoder can determine to flip the reconstructed block, and based on the reference template being one of the second candidate template, the decoder can determine not to flip the reconstructed block. The CB can be decoded based on whether to flip the reconstructed block in the direction.

[0219] In some examples, based on the reference template being from the first candidate template, the decoder can flip the reconstructed block in the direction and decode the CB based on the flipped reconstructed block. For example, the CB can correspond to the flipped reconstructed block. In some examples, based on the reference template being from the second candidate template, the decoder can decode the CB based on the reconstructed block. For example, the CB can correspond to the reconstructed block (e.g., without further transformations such as flipping, rotation, and / or scaling, etc.).

[0220] In some examples, the decoder can receive an indication in the bitstream that the CB is encoded in a template matching prediction (TMP) mode that uses multiple types of candidate templates, e.g., using a candidate template that is flipped in the direction relative to the current template. The determination of the first candidate template (at block 2302) and the determination of the second candidate template (at block 2304) can be in response to the decoder receiving an indication of this TMP mode (also referred to herein as a reconstruction rearrangement TMP mode or RR-TMP mode).

[0221] In some examples, the decoder can receive a residual for the CB from the bitstream. Based on the selected / determined reference template being from the first candidate template, the decoder can decode the CB based on flipping the reconstructed block in the direction, where the reconstructed block is a combination of the RB and the residual for the CB. Based on the reference template being from the second candidate template, the decoder can decode the CB based on combining the RB and the residual for the CB.

[0222] In some embodiments, in addition to the first and second candidate templates, one or more types of additional candidate templates corresponding to additional candidate RBs can be further determined based on TM costs of the additional candidate templates from which the reference template can be determined. For example, Figure 23 The method of example 1 can further include determining third candidate templates of third candidate RBs from a third search region, where each of the third candidate templates corresponds to the current template flipped in a second direction. The third search region can correspond to the second direction. For example, the direction (associated with the first candidate templates) can be horizontal and the second direction can be vertical, or vice versa. In these examples, the TM costs can further include third TM costs of the third candidate templates, which can be similarly computed based on one or more cost criteria as described above with respect to computing the first TM costs of the first candidate templates, at block 2306. In these examples, the reference template can be selected from the first candidate templates, the second candidate templates, and the third candidate templates based on the TM costs, at block 2308.

[0223] It should be further noted that the above with respect to Figure 23The discussed approach can not be limited to candidate templates that are flipped in a certain direction, and can be further extended to include other types of candidate templates corresponding to other transformations performed on the current template, as will be appreciated by one of skill in the art based on the present disclosure. For example, at block 2302, instead of each of the first candidate templates corresponding to the current template (of a CB) flipped in a certain direction, each of the first candidate templates can correspond to the current template (of a CB) with a transformation applied. In some examples, each of the first candidate templates corresponding to the current template with a transformation applied can include each of the first candidate templates that match the current template in size, shape, and orientation after the transformation. In an example, the transformation can include a rigid transformation (or isometry) that does not change the size or shape after the transformation. The transformation can include one or more of a rotation and a reflection, and can exclude a translation. For example, the transformation can include rotating the current template by a predetermined amount (e.g., degrees or radians, etc.).

[0224] Embodiments of the present disclosure can be implemented in hardware that uses analog and / or digital circuits, in software that is executed by one or more general purpose or application specific processors, or in a combination of hardware and software. Thus, embodiments of the present disclosure can be implemented in an environment such as that shown in FIG. 24. Figure 24 Figure 1 Figure 2 Figure 3 The blocks depicted in the above figures can be implemented on one or more computer systems 2400. Further, each block of the flowchart illustrations, and combinations of blocks in the flowchart illustrations, can be implemented on one or more computer systems 2400.

[0225] The computer system 2400 includes one or more processors, such as processor 2404. Processor 2404 can be a specialized processor, a general purpose processor, a microprocessor, or a digital signal processor, for example. Processor 2404 can be connected to a communication infrastructure 2402 (e.g., a bus or network).

[0226] ​​​Auxiliary memory 2408 can include, for example, a hard disk drive 2410 and / or a removable storage drive 2412, representing a magnetic tape drive, an optical disk drive, etc. Removable storage drive 2412 can read from and / or write to a removable storage unit 2416 in a well-known manner. Removable storage unit 2416, represents a magnetic tape, an optical disk, etc. which is read by and written to by removable storage drive 2412. As will be appreciated, removable storage unit 2416 includes a computer usable storage medium having stored therein computer software and / or data.

[0227] In alternative implementations, auxiliary memory 2408 can include other similar means for allowing computer programs or other instructions to be loaded into computer system 2400. Such means can include, for example, a removable storage unit 2418 and an interface 2414. Examples of this can include a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a thumb drive, and USB port, and other removable storage units 2418 and interfaces 2414 which allow software and data to be transferred from the removable storage unit 2418 to computer system 2400.

[0228] Computer system 2400 can also include a communications interface 2420. Communications interface 2420 allows software and data to be transferred between computer system 2400 and external devices. Examples of communications interface 2420 can include a modem, a network interface (such as an Ethernet card), a communications port, etc. Software and data transferred via communications interface 2420 are in the form of signals which can be electronic, electromagnetic, optical or other signals capable of being received by communications interface 2420. These signals are provided to communications interface 2420 via a communications path 2422. Communications path 2422 carries signals and can be implemented using wire or cable, fiber optics, a phone line, a cellular phone link, an RF link, and / or other communications channels.

[0229] As used herein, the terms "computer program medium" and "computer-readable medium" are used to generally refer to tangible storage media such as removable storage units 2416 and 2418 or a hard disk installed in hard disk drive 2410. These computer program products are means for providing software to computer system 2400. Computer programs (also known as computer control logic) can be stored in main memory 2406 and / or secondary memory 2408. Computer programs can also be received via communications interface 2420. Such computer programs, when executed, enable computer system 2400 to implement the present disclosure as discussed herein. In particular, the computer programs, when executed, enable processor 2404 to implement the processes of the present disclosure, such as any of the methods described herein. Accordingly, such computer programs represent controllers of the computer system 2400.

[0230] In another embodiment, the features of the present disclosure can be implemented in hardware using, for example, hardware components such as application specific integrated circuits (ASICs). It is also clear for one skilled in the art that implementation of a hardware state machine for performing the functions described herein is also possible.

Claims

1. A method comprising: receiving, in a bitstream, an indication that a current block (CB) is encoded using a template matching prediction (TMP) mode that uses candidate templates that are flipped in a direction relative to a current template of the CB; based on the indication: determining first candidate templates of first candidate reference blocks (RBs) from a first search region, the first candidate templates each corresponding to the current template flipped in the direction; determining second candidate templates of second candidate RBs from a second search region, the second candidate templates each corresponding to the current template; calculating template matching (TM) costs based on the current template, the TM costs including: first TM costs of the first candidate templates; and second TM costs of the second candidate templates; selecting a reference template from the first candidate templates and the second candidate templates based on the TM costs; and decoding the CB based on a RB indicated by the reference template.

2. A method comprising: determining first candidate templates of first candidate reference blocks (RBs) from a first search region, the first candidate templates each corresponding to a current template of a current block (CB) flipped in a direction; determining second candidate templates of second candidate RBs from a second search region, the second candidate templates each corresponding to the current template of the CB; calculating template matching (TM) costs based on the current template, the TM costs including: first TM costs of the first candidate templates; second TM costs of the second candidate templates; selecting a reference template from the first candidate templates and the second candidate templates based on the TM costs; and encoding the CB based on a RB indicated by the reference template.

3. The method of claim 2, further comprising: receiving, in a bitstream, an indication that the CB is encoded in a template matching prediction (TMP) mode that uses candidate templates that are flipped in the direction relative to the current template; wherein the determining the first candidate templates and the determining the second candidate templates are based on the receiving the indication.

4. The method of any of claims 2-3, wherein the encoding the CB comprises: determining a reconstructed block based on combining the RB with a residual of the CB; determining whether to flip the reconstructed block in the direction based on whether the reference template is one of the first candidate templates or the second candidate templates; and encoding the CB based on whether the reconstructed block is flipped in the direction. receiving the residual of the CB from a bitstream.

5. The method of claim 4, further comprising:

6. The method of any of claims 4-5, wherein the CB is decoded based on the reference template being from the first candidate templates based on the reconstructed block being flipped in the direction.

7. The method of any of claims 4-6, wherein the CB is decoded based on the reference template being from the second candidate templates based on the reconstructed block. ​ 8. The method of any of claims 2-3, further comprising: receiving, from a bitstream, a residual for the CB, wherein: based on the reference template being one of the first candidate templates, decoding the CB based on flipping the RB in the direction, wherein the reconstructed block comprises combining the flipped RB with the residual; and based on the reference template being from the second candidate templates, decoding the CB based on combining the RB with the residual.

9. The method of claim 2, further comprising: transmitting, in a bitstream, an indication that the CB is encoded in a template matching prediction (TMP) mode that uses a candidate template that is flipped in the direction relative to the current template.

10. The method of claim 9, wherein the determining the first candidate templates and the determining the second candidate templates are based on being in the TMP mode.

11. The method of any of claims 2, 9, and 10, wherein the coding the CB comprises encoding the CB, the encoding the CB comprising: determining whether to flip the CB in the direction prior to determining a residual for the CB based on whether the reference template is one of the first candidate templates or the second candidate templates; determining the residual for the CB based on: the determining whether to flip the CB; the RB; and the CB; and transmitting, in a bitstream, the residual for the CB.

12. The method of claim 11, wherein based on the reference template being one of the first candidate templates, the residual is determined based on a difference between the CB flipped in the direction and the RB.

13. The method of any of claims 11-12, wherein based on the reference template being one of the second candidate templates, the residual is determined based on a difference between the CB and the RB.

14. The method of any of claims 9-10, wherein the coding the CB comprises encoding the CB, the encoding the CB comprising: determining a residual based on the CB and the RB, wherein: based on the reference template being from the first candidate templates, the residual is determined based on a difference between the CB and the RB flipped in the direction; and based on the reference template being from the second candidate templates, the residual is determined based on a difference between the CB and the RB; and transmitting, in a bitstream, the residual for the CB.

15. The method of any of claims 2-14, wherein the current template is defined relative to the CB, wherein the first candidate templates are defined relative to respective first RB candidates, and wherein the second candidate templates are defined relative to respective second RB candidates.

16. The method of any of claims 2-15, wherein each of the first candidate templates matches the current template flipped in the direction in shape and orientation. ​ ​ 17. The method of claim 17, wherein each of the first candidate templates is further matched in size to the current template flipped in the direction.

18. The method of any of claims 2-17, wherein each of the first candidate templates is matched in shape and direction to the current template but not in orientation.

19. The method of any of claims 2-18, wherein each of the second candidate templates is matched in shape and orientation to the current template.

20. The method of claim 19, wherein each of the first candidate templates is further matched in size to the current template.

21. The method of any of claims 2-20, wherein the computing the TM cost comprises: comparing samples in each of the first candidate templates to corresponding samples in the current template flipped in the direction to compute respective first TM costs; and comparing samples in each of the second candidate templates to corresponding samples in the current template to compute respective second TM costs.

22. The method of any of claims 2-21, wherein the current template comprises: a first portion comprising a plurality of rows of samples above the CB; and a second portion comprising a plurality of columns of samples to the left of the CB.

23. The method of claim 22, wherein the rows are matched in width to the CB, and wherein the columns are matched in height to the CB.

24. The method of any of claims 22-23, wherein: based on the direction being horizontal, each of the first candidate templates comprises: the plurality of rows of samples above a respective first candidate template; and the plurality of columns of samples to the right of the respective first candidate template; and based on the direction being vertical, each of the first candidate templates comprises: the plurality of rows of samples below the respective first candidate template; and the plurality of columns of samples to the left of the respective first candidate template.

25. The method of any of claims 2-24, wherein the first search region and the second search region are portions of a picture frame of the CB.

26. The method of claim 25, wherein the portions comprise reconstructed portions or decoded portions of the CB.

27. The method of any of claims 2-26, wherein the current template, each of the first candidate templates, and each of the second candidate templates comprise respective sets of reconstructed samples.

28. The method of any of claims 2-27, wherein the first search region is different from the second search region.

29. The method of any of claims 2-28, wherein the first search region is associated with a flipping direction.

30. The method of any of claims 2-29, wherein the second search region comprises: ​ ​ a portion of a first coding tree unit (CTU) above and adjacent to a current CTU in which the CB is located; a portion of a second CTU left and adjacent to the current CTU; a portion of a third CTU above and left of the current CTU, wherein the third CTU is adjacent to the first CTU and the second CTU; and a portion of the current CTU above and left of the CB.

31. The method of any of claims 2-30, wherein the first search region comprises: a first rectangular region above and left of the CB; and a second rectangular region above or left of the CB depending on the direction.

32. The method of claim 31, wherein the first rectangular region is within a current CTU of the CB.

33. The method of any of claims 31-32, wherein the first rectangular region has a lower right corner that intersects an upper left corner of the CB.

34. The method of any of claims 31-33, wherein the first rectangular region comprises a first width and a first height based on a width and a height of the CB, respectively.

35. The method of any of claims 31-34, wherein the second rectangular region comprises a second width and a second height based on the width and the height of the CB, respectively.

36. The method of any of claims 31-35, wherein based on the direction being horizontal, the second rectangular region is adjacent to the CB and left of the CB.

37. The method of claim 36, wherein the second rectangular region comprises a second height that is the same as a height of the CB and comprises a second width based on a width of the CB.

38. The method of any of claims 31-37, wherein based on the direction being vertical, the second rectangular region is adjacent to the CB and above the CB.

39. The method of claim 38, wherein the second rectangular region comprises a second width that is the same as a width of the CB and comprises a second height based on a height of the CB.

40. The method of any of claims 2-39, wherein the RB is a block from which the reference template is defined.

41. The method of any of claims 2-40, further comprising: determining third candidate templates of third candidate RBs from a third search region, wherein each of the third candidate templates corresponds to the current template flipped in a second direction, wherein the third search region is associated with the second direction, wherein: the TM cost further comprises a third TM cost of the third candidate templates; and selecting the reference template from the first candidate templates, the second candidate templates, and the third candidate templates based on the TM cost.

42. The method of claim 41, wherein the direction is horizontal and the second direction is vertical, or the direction is vertical and the second direction is horizontal. ​ 43. The method of any of claims 2-42, wherein the selecting the reference template is based on the reference template having a minimum cost in the TM cost.

44. A method comprising: based on a template matching prediction (TMP) mode using candidate templates flipped in a direction relative to a current template of a current block (CB): determining, from a first search region, first candidate templates of first candidate reference blocks (RBs), the first candidate templates each corresponding to the current template flipped in the direction; determining, from a second search region, second candidate templates of second candidate RBs, the second candidate templates each corresponding to the current template of the CB; computing, based on the current template, a template matching (TM) cost, the TM cost comprising: first TM costs of the first candidate templates; second TM costs of the second candidate templates; selecting, based on the TM cost, a reference template from the first candidate templates and the second candidate templates; encoding the CB based on a RB indicated by the reference template; and transmitting, in a bitstream, an indication that the CB is encoded using the TMP mode.

45. A decoder comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the decoder to perform the method of any of claims 1, 2-8, and 15-43.

46. An encoder comprising: one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the encoder to perform the method of any of claims 2, 9-43, and 44.

47. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any of claims 1-44. ​ ​