Search region modification for intra template matching prediction

By determining multi-directional search regions and template matching in video frames, the intra-frame TMP mode is improved, solving the problem of low video encoding and decoding efficiency in existing technologies and achieving more efficient video compression and decoding performance.

CN120982085APending Publication Date: 2025-11-18BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480023515.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-09
Filing Date
2024-04-16
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies suffer from inefficiency and poor compression in intra-frame template matching prediction, especially when processing search regions and reference block matching of video blocks, making it difficult to effectively utilize the redundancy of video data.

Method used

An improved search region and iterative search method for intra-frame TMP mode is achieved by determining search regions with different sizes and positions in video frames, performing template matching to determine reference blocks, and using a processor to generate predictive samples of video blocks.

Benefits of technology

It improves the efficiency and compression effect of video encoding and decoding, reduces bit rate requirements, and enhances video quality and compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120982085A_ABST
    Figure CN120982085A_ABST
Patent Text Reader

Abstract

A method for video encoding, a method for video decoding, and apparatuses therefor are provided. A search region for a video block is determined from a video frame of a video. The search area has a first search area size in a first direction and a second search area size in a second direction perpendicular to the first direction. The search region includes a plurality of regions, each region of the plurality of regions being within a region determined by the first search region size, the second search region size, and a position of the video block in the video frame. A reference block is determined from the search region. The template of the reference block matches the template of the video block. Prediction samples for the video block are determined based on the reference block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to U.S. Provisional Application No. 63 / 459,967, filed April 17, 2023. This application is a continuation-in-part application and further claims priority to PCT Application No. PCT / US24 / 23677, filed April 9, 2024, which in turn claims priority to U.S. Provisional Application No. 63 / 458,419, filed April 10, 2023. The contents of all the foregoing applications are incorporated herein by reference in their entirety. Technical Field

[0003] This application relates to video encoding / decoding and compression. More specifically, this application relates to video processing apparatus and methods for intra-frame template matching prediction (TMP). Background Technology

[0004] Various electronic devices (such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, video streaming devices, etc.) support digital video. Electronic devices send and receive, or otherwise transmit, digital video data via communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited storage resources of storage devices, video data can be compressed using one or more video codec standards before it is transmitted or stored. For example, video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video codecs typically employ prediction methods that utilize the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video codecs aim to compress video data to a form using a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention

[0005] Implementations of the present disclosure provide a method for video decoding. The method can include determining, by a decoder, a search region for a video block from a video frame of a video. The search region has a first search region size in a first direction and a second search region size in a second direction perpendicular to the first direction. The search region includes a plurality of regions, each of the plurality of regions being within an area determined by the first search region size, the second search region size, and a location of the video block in the video frame. The method can also include determining, by the decoder, a reference block from the search region. A template of the reference block matches a template of the video block. The method can also include determining, by the decoder, prediction samples for the video block based on the reference block.

[0006] Implementations of the present disclosure provide a method for video encoding. The method can include determining, by an encoder, a search region for a video block from a video frame of a video. The search region has a first search region size in a first direction and a second search region size in a second direction perpendicular to the first direction. The search region includes a plurality of regions, each of the plurality of regions being within an area determined by the first search region size, the second search region size, and a location of the video block in the video frame. The method can also include determining, by the encoder, a reference block from the search region. A template of the reference block matches a template of the video block. The method can also include determining, by the encoder, prediction samples for the video block based on the reference block. The method can also include generating, by the encoder, a bitstream based on the prediction samples.

[0007] Implementations of the present disclosure also provide an apparatus for video decoding. The apparatus can include a memory configured to store a bitstream and a processor coupled to the memory. The processor can be configured to perform the methods disclosed herein for video decoding to decode the bitstream.

[0008] Implementations of the present disclosure also provide an apparatus for video encoding. The apparatus can include a memory configured to store a bitstream and a processor coupled to the memory. The processor can be configured to perform the methods disclosed herein for video encoding to generate the bitstream.

[0009] Implementations of the present disclosure also provide a non-transitory computer- readable storage medium having a bitstream stored therein, the bitstream being decoded by the methods disclosed herein for video decoding.

[0010] Implementations of the present disclosure also provide a non-transitory computer- readable storage medium having a bitstream stored therein, the bitstream being generated by the methods disclosed herein for video encoding.

[0011] It will be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0012] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0013] Figure 1 FIG. 1 is a block diagram illustrating an exemplary system for encoding and decoding video blocks according to some embodiments of the present disclosure.

[0014] Figure 2 FIG. 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0015] Figure 3 FIG. 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0016] Figures 4A-4E FIG. 4 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure.

[0017] Figure 5A FIG. 5 is a diagram of search regions for intra-TMP modes in an enhanced compression model (ECM) according to some examples.

[0018] Figure 5B FIG. 6 is a block diagram illustrating reference blocks in search regions according to some embodiments of the present disclosure.

[0019] Figure 6 FIG. 7 is a flowchart of an exemplary method for intra-TMP of video frames of a video according to some embodiments of the present disclosure.

[0020] Figure 7A FIG. 8 is a diagram of modified search regions for intra-TMP modes according to some embodiments of the present disclosure.

[0021] Figure 7B FIG. 9 is another diagram of modified search regions for intra-TMP modes according to some embodiments of the present disclosure.

[0022] Figures 8A-8C FIG. 10 illustrates a method of generating sample values for unreconstructed samples according to some embodiments of the present disclosure.

[0023] Figures 9A-9B FIG. 11 illustrates performing an iterative search method on a search region with a scaling factor according to some embodiments of the present disclosure.

[0024] Figure 10 FIG. 12 is a flowchart of another exemplary method for intra-TMP of video frames of a video according to some embodiments of the present disclosure.

[0025] Figure 11 is a diagram illustrating a computing environment coupled with a user interface in accordance with some embodiments of the present disclosure. DETAILED DESCRIPTION

[0026] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description of embodiments, numerous specific details are set forth in order to provide a thorough understanding of the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that the subject matter presented can be practiced without the specific details presented. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.

[0027] It should be noted that the terms "first," "second," and the like, used in the description and the claims of the present disclosure as well as the drawings herein are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the term data herein is not limited to nor is it to be construed as a reference to such physically-embodied data only. Rather, it is also to be understood that such data is to be construed, as example, as assuming the form of a signal whether electronically, optically, or otherwise transmitted or conveyed.

[0028] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel in accordance with some embodiments of the present disclosure. As shown in Figure 1 System 10 includes a source device 12 that generates and encodes video data to be decoded at a later time by a destination device 14, as shown in FIG. 1. Source device 12 and destination device 14 can comprise any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, video streaming devices, and the like. In some embodiments, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0029] In some embodiments, destination device 14 can receive, via link 16, encoded video data to be decoded. Link 16 can comprise any type of communication medium or device capable of moving the encoded video data from source device 12 to destination device 14. In one example, link 16 can comprise a communication medium to enable source device 12 to transmit encoded video data directly to destination device 14 in real-time. The encoded video data can be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium can comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide-area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful to facilitate communication from source device 12 to destination device 14.

[0030] In other embodiments, encoded video data can be transmitted from output interface 22 to storage device 32. Subsequently, encoded video data in storage device 32 can be accessed by destination device 14 via input interface 28. Storage device 32 can include any of a variety of distributed or locally accessed data storage media such as a hard drive, Blu-ray discs, Digital Video Disk (DVD), Compact Disk-Read Only Memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data. In a further example, storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by source device 12. Destination device 14 can access stored video data from storage device 32 via streaming or download. The file server can be any type of computer

[0031] As Figure 1As shown in FIG. 1, source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 can include a source such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feed interface to receive video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of such sources. As one example, if video source 18 is a video camera of a security surveillance system, source device 12 and destination device 14 can form a camera phone or video phone. However, the implementations described in this application can be applied to video coding in general, and can have application to wireless and / or wired applications.

[0032] The captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data can also (or alternatively) be stored onto storage device 32 for later access by destination device 14 or other devices, for decoding and / or playback. Output interface 22 can further include a modem and / or a transmitter.

[0033] Destination device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 can include a receiver and / or modem and receives encoded video data over link 16. The encoded video data communicated over link 16, or provided on storage device 32, can include a variety of syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements can be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.

[0034] In some implementations, destination device 14 can include display device 34, which can be an integrated display device and an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user, and can include any of a variety of display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0035] Video encoder 20 and video decoder 30 can operate according to a proprietary standard or industry standard, such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of such standards. It should be understood that the application is not limited to a specific video coding / decoding standard and can be applicable to other video coding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is also generally contemplated that video decoder 30 of destination device 14 can be configured to decode video data according to any of these current or future standards.

[0036] Video encoder 20 and video decoder 30 can be implemented as any of a variety of suitable encoder and / or decoder circuitry, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic circuitry, software, hardware, firmware or any combinations thereof. When implemented partially in software, an electronic device can store instructions for the software in a suitable, non- transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video coding / decoding operations disclosed in the present disclosure. Each of video encoder 20 and video decoder 30 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined encoder / decoder (CODEC) in a respective device.

[0037] Figure 2 FIG. 1 is a block diagram illustrating an example video encoder 20 according to some embodiments described in the present application. Video encoder 20 can perform intra-prediction encoding and inter-prediction encoding on video blocks within a video frame. Intra-prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within neighboring video frames or pictures of a video sequence. It should be noted that in the field of video coding, the term “frame” can be used as a synonym for the term “image” or “picture.”

[0038] As Figure 2As shown in FIG. 1, video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, summer 50, transform processing unit 52, quantization unit 54, and entropy encoding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, partition unit 45, intra-prediction processing unit 46, and intra-block copy (BC) unit 48. In some implementations, video encoder 20 also includes inverse quantization unit 58, inverse transform processing unit 60, and summer 62 for video block reconstruction. A loop filter 63, such as a deblocking filter, can be located between summer 62 and DPB 64 to filter block boundaries to remove blockiness artifacts from reconstructed video. In addition to the deblocking filter, another loop filter (e.g., a sample adaptive offset (SAO) filter, a cross component sample adaptive offset (CCSAO) filter, and / or an adaptive loop filter (ALF)) can be used to filter the output of summer 62. Note that for the CCSAO technique, the present application is not limited to the embodiments described herein, but can also be applied to a case where an offset is selected for any one of a luma component, a Cb chroma component, and a Cr chroma component to modify any other one of the luma component, the Cb chroma component, and the Cr chroma component based on the selected offset. Also note that the first component referred to herein can be any one of the luma component, the Cb chroma component, and the Cr chroma component, the second component referred to herein can be any other one of the luma component, the Cb chroma component, and the Cr chroma component, and the third component referred to herein can be the remaining one of the luma component, the Cb chroma component, and the Cr chroma component. In some examples, the loop filters can be omitted, and the decoded video blocks can be provided directly by summer 62 to DPB 64. Video encoder 20 can take the form of a fixed or programmable hardware encoder, or can be dispersed in one or more of the illustrated fixed or programmable hardware encoders.

[0039] Video data memory 40 can store video data to be encoded by the components of video encoder 20. The video data in video data memory 40 can be obtained, for example, from video source 18, as shown in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use in encoding video data by video encoder 20 (e.g., in intra- or inter-coding modes). Video data memory 40 and DPB 64 can be formed by a variety of memory devices, including a Figure 1 suggested in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or pictures) for use in encoding video data by video encoder 20 (e.g., in intra- or inter-coding modes). Video data memory 40 and DPB 64 can be formed by a variety of memory devices, including a

[0040] As Figure 2As shown in FIG. 1, after receiving video data, partitioning unit 45 within prediction processing unit 41 partitions the video data into video blocks. This partitioning can also include partitioning of a video frame into slices, tiles, or other larger coding units (CUs) according to a predefined splitting structure (e.g., a quadtree (QT) structure) associated with the video data, for example. A video frame is or can be considered as a two-dimensional array or matrix of sample values. Samples in the array can also be referred to as pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. A video frame can be divided into a plurality of video blocks, for example, by using QT partitioning. A video block is or can be considered as a two-dimensional array or matrix of sample values again, but with dimensions smaller than those of the video frame. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. A video block can be further partitioned into one or more block partitions or sub-blocks (which can form blocks again) by using QT partitioning, binary tree (BT) partitioning, or ternary tree (TT) partitioning, or any combination thereof, for example, iteratively. It should be noted that the term “block” or “video block” as used herein can be a portion of a frame or picture, in particular a rectangular (square or non-square) portion. With reference to HEVC and VVC, for example, a block or video block can be or correspond to a coding tree unit (CTU), a CU, a prediction unit (PU), or a transform unit (TU) and / or can be or correspond to a respective block (e.g., a coding tree block (CTB), a coding block (CB), a prediction block (PB), or a transform block (TB)) and / or a sub-block.

[0041] Prediction processing unit 41 can select one of a plurality of possible predictive encoding modes, such as one of a plurality of intra-predictive encoding modes or one of a plurality of inter-predictive encoding modes, for the current video block based on the error results (e.g., rate and distortion levels). Prediction processing unit 41 can provide the resulting intra- or inter-predicted block to summer 50 to generate a residual block, and to summer 62 to reconstruct the encoded block for use as part of a reference frame at a later time. Prediction processing unit 41 also provides syntax elements, such as motion vectors, intra-mode indicators, partitioning information, and other such syntax information, to entropy encoding unit 56.

[0042] To select a suitable intra-prediction coding mode for the current video block, intra-prediction processing unit 46 within prediction processing unit 41 can perform intra-prediction coding of the current video block in relation to one or more neighboring blocks in the same frame as the current block being coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-prediction coding of the current video block in relation to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 can perform multiple coding passes, e.g., to select a suitable coding mode for each block of video data.

[0043] In some implementations, motion estimation unit 42 determines an inter-prediction mode for a current video frame by generating motion vectors according to a predetermined pattern within a sequence of video frames, the motion vectors indicating displacements of video blocks within the current video frame relative to prediction blocks within a reference video frame. Motion estimation performed by motion estimation unit 42 is a process of generating motion vectors that estimate the motion for video blocks. For example, a motion vector can indicate a displacement of a video block within a current video frame or picture relative to a prediction block within a reference frame that is related to a current block being coded within the current frame. The predetermined pattern can designate video frames in the sequence as P-frames or B-frames. Intra-BC unit 48 can determine vectors for intra-BC coding (e.g., block vectors) in a similar manner as the motion vectors determined by motion estimation unit 42 for inter-prediction, or can utilize motion estimation unit 42 to determine the block vectors.

[0044] In terms of pixel differences, a prediction block for a video block can be or can correspond to a block or reference block of a reference frame that is deemed to closely match the video block being coded, the pixel differences can be determined by a sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metric. In some implementations, video encoder 20 can calculate values for sub-integer pixel positions of reference frames stored in DPB 64. For example, video encoder 20 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of a reference frame. Thus, motion estimation unit 42 can perform a motion search relative to full-pixel positions and fractional-pixel positions and output motion vectors with fractional-pixel precision.

[0045] Motion estimation unit 42 calculates a motion vector for a video block in an inter-prediction coded frame by comparing a location of the video block to a location of a prediction block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of the first and second reference frame lists identifying one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44, which then sends to entropy encoding unit 56.

[0046] Motion compensation performed by motion compensation unit 44 can involve fetching or generating a prediction block based on a motion vector determined by motion estimation unit 42. After receiving a motion vector for a current video block, motion compensation unit 44 can locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to summer 50. Summer 50 then forms a residual video block of pixel difference values by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the current video block being encoded. The pixel difference values forming the residual video block can include luma component differences or chroma component differences or both. Motion compensation unit 44 can also generate syntax elements associated with the video block of the video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements can include, for example, syntax elements defining the motion vector used to identify the prediction block, any flags indicating the prediction mode, or any other syntax information described herein. It is noted that motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are illustrated separately for conceptual purposes.

[0047] In some implementations, intra BC unit 48 can generate vectors and fetch prediction blocks in a manner similar to that described above in connection with motion estimation unit 42 and motion compensation unit 44, but the prediction blocks are in the same frame as the current block being encoded, and the vectors are referred to as block vectors rather than motion vectors. Specifically, intra BC unit 48 can determine an intra prediction mode to use for encoding the current block. In some examples, intra BC unit 48 can encode the current block using various intra prediction modes, e.g., during a separate encoding pass, and test their performance through rate-distortion analysis. Next, intra BC unit 48 can select an appropriate intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, intra BC unit 48 can compute rate-distortion values for the various tested intra prediction modes using rate-distortion analysis, and select the intra prediction mode having the best rate-distortion characteristics among the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original, unencoded block that was encoded to produce the encoded block, and the bit rate (i.e., the number of bits) used to produce the encoded block. Intra BC unit 48 can compute a ratio from the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion values for the block.

[0048] In other examples, intra BC unit 48 can perform such functions for intra BC prediction according to the implementations described herein using motion estimation unit 42 and motion compensation unit 44 in whole or in part. In either case, for intra block copy, in terms of pixel difference, the prediction block can be a block that is deemed to closely match the block to be encoded, the pixel difference can be determined by SAD, SSD, or other difference metric, and identifying the prediction block can include calculating values for sub-integer pixel positions.

[0049] Regardless of whether the prediction block is from the same frame according to intra prediction or a different frame according to inter prediction, video encoder 20 can form pixel difference values by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded, thereby forming a residual video block. The pixel difference values forming the residual video block can include both luma component differences and chroma component differences.

[0050] As an alternative to inter prediction performed by motion estimation unit 42 and motion compensation unit 44 or intra block copy prediction performed by intra BC unit 48 as described above, intra prediction processing unit 46 can intra predict the current video block. In particular, intra prediction processing unit 46 can determine an intra prediction mode to use for encoding the current block. To this end, intra prediction processing unit 46 can encode the current block using various intra prediction modes, e.g., during a separate encoding pass, and intra prediction processing unit 46 (or, in some examples, a mode selection unit) can select an appropriate intra prediction mode to use from among the tested intra prediction modes. Intra prediction processing unit 46 can provide information indicative of the selected intra prediction mode for the block to entropy encoding unit 56. Entropy encoding unit 56 can encode the information indicative of the selected intra prediction mode into the bitstream.

[0051] After prediction processing unit 41 determines a prediction block for the current video block via inter prediction or intra prediction, adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block can be included in one or more TUs and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, e.g., a discrete cosine transform (DCT) or a conceptually similar transform.

[0052] Transform processing unit 52 can send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce bit rate. The quantization process can also reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting a quantization parameter. In some examples, quantization unit 54 can then perform a scan on the matrix including the quantized transform coefficients. Alternatively, entropy encoding unit 56 can perform the scan.

[0053] Following quantization, entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding, or another entropy encoding methodology or technique. The encoded bitstream can then be transmitted to video decoder 30 as shown in FIG. 3 or archived, such as in storage device 32 as shown in FIG. 3 for later transmission to or retrieval by video decoder 30. Entropy encoding unit 56 can also entropy encode motion vectors and other syntax elements for the current video frame that is being encoded. Figure 1 Figure 1

[0054] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transforms, respectively, to reconstruct the residual video blocks in the pixel domain for use in generating reference blocks for predicting other video blocks. As noted above, motion compensation unit 44 can generate motion compensated prediction blocks from one or more reference blocks of frames stored in DPB 64. Motion compensation unit 44 can also apply one or more interpolation filters to a prediction block to calculate sub-integer pixel values for use in motion estimation.

[0055] Adder 62 adds the reconstructed residual blocks to the motion compensated prediction blocks produced by motion compensation unit 44 to produce reference blocks for storage in DPB 64. The reference blocks can then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as prediction blocks to inter predict another video block in a subsequent video frame.

[0056] Figure 3 FIG. 3 is a block diagram illustrating an example video decoder 30 according to some embodiments of the present application. Video decoder 30 includes video data memory 79, entropy decoding unit 80, prediction processing unit 81, inverse quantization unit 86, inverse transform processing unit 88, adder 90, and DPB 92. Prediction processing unit 81 further includes motion compensation unit 82, intra prediction unit 84, and intra BC unit 85. Video decoder 30 can perform a decoding process generally reciprocal to the encoding process described above in connection with video encoder 20. Figure 2 The decoding process described in connection with video encoder 20 is substantially reciprocal. For example, motion compensation unit 82 can generate prediction data based on motion vectors received from entropy decoding unit 80, while intra prediction unit 84 can generate prediction data based on intra prediction mode indicators received from entropy decoding unit 80.

[0057] ​​In some examples, the components of video decoder 30 can be tasked to perform the implementations of the present application. Moreover, in some examples, the implementations of the present disclosure can be distributed among one or more of the components of video decoder 30. For example, intra BC unit 85 can perform the implementations of the present application, alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some examples, video decoder 30 can not include intra BC unit 85, and the functionality of intra BC unit 85 can be performed by other components of prediction processing unit 81, such as motion compensation unit 82.

[0058] Video data memory 79 can store video data, such as an encoded video bitstream, to be decoded by the other components of video decoder 30. The video data stored in video data memory 79 can be obtained, for example, from storage device 32, from a local video source, such as a camera, via wired or wireless network communication of video data, or by accessing physical data storage media (e.g., a flash drive or hard disk). Video data memory 79 can include an encoded picture buffer (CPB) that stores encoded video data from an encoded video bitstream. DPB 92 of video decoder 30 stores reference video data for use in decoding video data by video decoder 30 (e.g., in intra- or inter-coding modes). Video data memory 79 and DPB 92 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magneto resistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For Figure 3 illustrative purposes, video data memory 79 and DPB 92 are depicted as two distinct components of video decoder 30. But it will be readily apparent to one of ordinary skill in the art that video data memory 79 and DPB 92 can be provided by same memory device or separate memory devices. In some examples, video data memory 79 can be on-chip with other components of video decoder 30, or off-chip relative to those components.

[0059] During the decoding process, video decoder 30 receives an encoded video bitstream that represents encoded video frames and associated syntax elements. Video decoder 30 can receive the syntax elements at the video frame level and / or video block level. Entropy decoding unit 80 of video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors, or intra-prediction mode indicators, among other syntax elements. Entropy decoding unit 80 then forwards the motion vectors, or intra-prediction mode indicators, among other syntax elements, to prediction processing unit 81.

[0060] When a video frame is encoded as an intra-predicted (I) frame or an intra-coded prediction block in another type of frame, intra-prediction unit 84 of prediction processing unit 81 can generate prediction data for a video block of the current video frame based on the intra-prediction mode signaled and reference data from previously decoded blocks of the current frame.

[0061] When a video frame is encoded as an inter-predicted (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 produces one or more prediction blocks for a video block of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks can be produced from a reference frame within one of the reference frame lists. Video decoder 30 can construct the reference frame lists, i.e., List 0 and List 1, using default construction techniques based on reference frames stored in DPB 92.

[0062] In some examples, when a video block is encoded according to the intra BC modes described herein, intra BC unit 85 of prediction processing unit 81 produces a prediction block for the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. The prediction block can be within a reconstructed region of the same picture as the current video block, as defined by video encoder 20.

[0063] Motion compensation unit 82 and / or intra BC unit 85 determine the prediction information for a video block of the current video frame by parsing the motion vectors and other syntax elements, and then use the prediction information to produce the prediction block for the current video block being decoded. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode used to encode the video block of the video frame (e.g., intra- or inter-prediction), the inter-prediction frame type (e.g., B or P), the construction information for one or more of the reference frame lists for the frame, the motion vectors for each inter-predicted video block of the frame, the inter-prediction status for each inter-predicted video block of the frame, and other information used to decode the video block in the current video frame.

[0064] Similarly, intra BC unit 85 can use some of the received syntax elements, such as flags, to determine that the current video block is predicted using the intra BC mode, the construction information for which video blocks of the frame are within the reconstructed region and should be stored in DPB 92, the block vectors for each intra BC predicted video block of the frame, the intra BC prediction status for each intra BC predicted video block of the frame, and other information used to decode the video block in the current video frame.

[0065] Motion compensation unit 82 can also perform interpolation using interpolation filters as used by video encoder 20 during encoding of the video blocks to calculate interpolated values for sub-integer pixels of reference blocks. In this case, motion compensation unit 82 can determine the interpolation filters used by video encoder 20 from the syntax elements received and use these interpolation filters when generating the prediction blocks.

[0066] Dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80 using the same quantization parameter calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct the residual blocks in the pixel domain.

[0067] After motion compensation unit 82 or intra BC unit 85 generates the prediction block for the current video block based on the vectors and other syntax elements, adder 90 reconstructs the decoded video block for the current video block by adding the residual block from inverse transform processing unit 88 to the corresponding prediction block generated by motion compensation unit 82 and intra BC unit 85. In-loop filter 91, e.g., a de-blocking filter, a SAO filter, a CCSAO filter, and / or an ALF, can be located between adder 90 and DPB 92 to further process the decoded video block. In some examples, in-loop filter 91 can be omitted, and the decoded video block can be directly provided by adder 90 to DPB 92. The decoded video blocks in a given frame are then stored in DPB 92, which stores reference frames for subsequent motion compensation of video blocks. DPB 92 or a separate memory device from DPB 92 can also store decoded video for later presentation on a display device (e.g., display device 34 of FIG. 1). Figure 1

[0068] In a typical video encoding process, a video sequence generally includes an ordered set of frames or pictures. Each frame can include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other examples, a frame can be monochrome, and thus include only one two-dimensional array of luma samples.

[0069] As Figure 4A ​As shown in FIG. 1, video encoder 20 (or more specifically, partitioning unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of CTUs. A video frame can include an integer number of CTUs ordered consecutively in the raster scan order from left to right and from top to bottom. Each CTU is the largest logical coding unit and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set such that all CTUs in a video sequence have the same size of one of 128x128, 64x64, 32x32, and 16x16. It should be noted, however, that the present application is not necessarily limited to a particular size. As Figure 4B As shown in FIG. 2, each CTU can include one CTB of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements used to code the samples of the coding tree blocks. The syntax elements describe the properties of different types of units that code the pixel blocks and how the video sequence can be reconstructed at video decoder 30, including inter or intra prediction, intra prediction mode, motion vectors, and other parameters. In monochrome pictures or pictures with three separate color planes, a CTU can include a single coding tree block and syntax elements used to code the samples of the coding tree block. A coding tree block can be an NxN block of samples.

[0070] To achieve better performance, video encoder 20 can recursively perform tree partitioning, e.g., binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree blocks of a CTU and divide the CTU into smaller CUs. As Figure 4C As depicted in FIG. 3, a 64x64 CTU 400 is first divided into four smaller CUs, each having a block size of 32x32. Of the four smaller CUs, CUs 410 and 420 are each divided into four CUs having a block size of 16x16. The two 16x16 CUs 430 and 440 are each further divided into four CUs having a block size of 8x8. Figure 4D A quad tree data structure showing the final result of the partitioning process of CTU 400 as depicted in FIG. 4 is depicted in FIG. 5. Each leaf node of the quad tree corresponds to one CU of a respective size ranging from 32x32 to 8x8. Similar to the binary tree data structure depicted in FIG. 3, the quad tree data structure can be used to indicate the partitioning of a CTU into smaller CUs. Figure 4C As depicted in FIG. 6, each CU can include a CB of luma samples and two corresponding coding blocks of chroma samples of the same size of the frame, and syntax elements used to code the samples of the coding blocks. In monochrome pictures or pictures with three separate color planes, a CU can include a single coding block and syntax elements used to code the samples of the coding block. It should be noted that Figure 4B As depicted in FIG. 6, each CU can include a CB of luma samples and two corresponding coding blocks of chroma samples of the same size of the frame, and syntax elements used to code the samples of the coding blocks. In monochrome pictures or pictures with three separate color planes, a CU can include a single coding block and syntax elements used to code the samples of the coding block. It should be noted that Figure 4C and Figure 4DThe quad-tree partitioning depicted in the middle is for illustrative purposes only, and one CTU can be split into multiple CUs based on quad-tree partitioning / triple-tree partitioning / binary-tree partitioning to adapt to varying local characteristics. In the multi-type tree structure, one CTU is partitioned according to a quad-tree structure, and each quad-tree leaf CU can be further partitioned according to binary and triple tree structures. As shown in Figure 4E The coding block with width W and height H has five possible partition types, i.e., quad partition, horizontal binary partition, vertical binary partition, horizontal ternary partition, and vertical ternary partition.

[0071] In some implementations, video encoder 20 can further partition the coding block of a CU into one or more (MxN) PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. A PU of a CU can include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PBs. In a monochrome picture or a picture having three separate color planes, a PU can include a single PB and syntax structures used to predict the PB. Video encoder 20 can generate a predicted luma block, a predicted Cb block, and a predicted Cr block for the luma PB, the Cb PB, and the Cr PB of each PU of a CU.

[0072] Video encoder 20 can use intra prediction or inter prediction to generate the predicted blocks of a PU. If video encoder 20 uses intra prediction to generate the predicted blocks of a PU, video encoder 20 can generate the predicted blocks of the PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter prediction to generate the predicted blocks of a PU, video encoder 20 can generate the predicted blocks of the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0073] After video encoder 20 generates the predicted luma blocks, the predicted Cb blocks, and the predicted Cr blocks for one or more PUs of a CU, video encoder 20 can generate luma residual blocks for the CU by subtracting the predicted luma blocks of the CU from the original luma coding block of the CU, such that each sample in a luma residual block of the CU indicates a difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, video encoder 20 can generate Cb residual blocks and Cr residual blocks for the CU, respectively, such that each sample in a Cb residual block of the CU indicates a difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in a Cr residual block of the CU can indicate a difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0074] Furthermore, as Figure 4CAs shown in FIG. 1, video encoder 20 can partition a picture into one or more slices, each including one or more CUs. Video encoder 20 can partition a picture into one or more tiles, each including one or more CUs. Video encoder 20 can partition a picture into one or more warps, each including one or more CUs. Video encoder 20 can partition a picture into one or more tiles and one or more warps. In some examples, a picture can be partitioned into one or more tiles, one or more warps, and one or more slices. In other examples, a picture can be partitioned into one or more slices, one or more tiles, or one or more warps, but not into both tiles and warps. In some examples, a picture can be partitioned into one or more slices, but not into tiles or warps. In other examples, a picture can be partitioned into one or more tiles, but not into slices or warps. In other examples, a picture can be partitioned into one or more warps, but not into slices or tiles. In some examples, a picture can be partitioned into one or more slices, one or more tiles, and one or more warps. In other examples, a picture can be partitioned into one or more slices, one or more tiles, or one or more warps, but not into both tiles and warps. In some examples, a picture can be partitioned into one or more slices, but not into tiles or warps. In other examples, a picture can be partitioned into one or more tiles, but not into slices or warps. In other examples, a picture can be partitioned into one or more warps, but not into slices or tiles. In some examples, a picture can be partitioned into one or more slices, one or more tiles, and one or more warps.

[0075] Video encoder 20 can apply one or more transforms to the luma transform block of a TU to generate a luma coefficient block for the TU. A coefficient block can be a two- dimensional array of transform coefficients. A transform coefficient can be a scalar. Video encoder 20 can apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 can apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0076] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 can quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to possibly reduce the amount of data used to represent the transform coefficients, providing further compression. Following quantization of a coefficient block by video encoder 20, video encoder 20 can entropy encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 can output a bitstream that includes a bit sequence forming a representation of the encoded frame and associated data, the bitstream being saved in storage device 32 or transmitted to destination device 14.

[0077] After receiving the bitstream generated by video encoder 20, video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 can reconstruct the frames of the video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data is generally reciprocal to the encoding process performed by video encoder 20. For example, video decoder 30 can perform inverse transforms on the coefficient blocks associated with the TUs of the current CU to reconstruct the residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs the coding blocks of the current CU by adding the samples of the prediction blocks for the PUs of the current CU to corresponding samples of the transform blocks of the TUs of the current CU. After reconstructing the coding blocks for each CU of a frame, video decoder 30 can reconstruct the frame.

[0078] As noted above, video coding primarily uses two modes (i.e., intra prediction (or intra-frame prediction) and inter prediction (or inter-frame prediction)) to achieve video compression. It should be noted that IBC can be considered as a third mode of intra prediction. Between the two modes, inter prediction contributes more to coding efficiency than intra prediction due to the use of motion vectors to predict the current video block from a reference video block.

[0079] However, as video data capture technology is constantly improving and finer video block sizes are used to preserve details in the video data, the amount of data required to represent the motion vectors for the current frame also increases substantially. One way to overcome this challenge benefits from the fact that not only a group of neighboring CUs in both spatial and temporal domains have similar video data for prediction purposes, but also the motion vectors between these neighboring CUs are similar. Therefore, the motion information (e.g., motion vectors) of the spatial neighboring CUs and / or temporal collocated CUs can be used as an approximation of the motion information (e.g., motion vectors) of the current CU (which is also referred to as the “motion vector predictor” (MVP) of the current CU) by exploring their spatial and temporal correlations.

[0080] Instead of encoding the actual motion vectors of the current CUs determined by motion estimation unit 42 into the video bitstream as described above in connection with Figure 2 Instead of encoding the actual motion vectors of the current CUs determined by motion estimation unit 42 into the video bitstream as described above in connection with

[0081] Similar to the process of selecting a prediction block in a reference frame during inter prediction of a coding block, both video encoder 20 and video decoder 30 need to employ a set of rules for constructing a motion vector candidate list (also referred to as a "merge list") for a current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally collocated CUs of the current CU, and then selecting a member from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, there is no need to send the motion vector candidate list itself from video encoder 20 to video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for video encoder 20 and video decoder 30 to use the same motion vector predictor within the motion vector candidate list to encode and decode the current CU.

[0082] Intra TMP can be utilized in ECM to improve the compression efficiency of intra coding. TMP can be an intra prediction mode that generates prediction samples (or predicted samples) of a video block from a reference block in the reconstructed portion of the current frame, where the template of the reference block matches the template of the video block. According to the intra TMP design, the template of the video block can have an L shape and include causal neighboring samples of the video block in the L shape. Similarly, the template of the reference block can also have an L shape and include causal neighboring samples of the reference block in the L shape. Reference is made to Figure 5B An example template is shown. For a predefined search range, video encoder 20 can search for a reference block in the reconstructed portion of the current video frame that has a template most similar to the current template of the video block, and can use the reference block with the most similar template as a prediction block. The prediction block can include the predicted samples of the video block. Video encoder 20 can then signal the use of this intra TMP mode. Subsequently, a similar prediction operation can be performed at the decoder side to generate the prediction block at the decoder side.

[0083] Figure 5A An example search region for the intra TMP mode is shown, which bounds the region of coordinates of the top-left corner of each candidate block (e.g., each potential reference block). Figure 5A A video block 502 is shown, which is part of a CTU 504. For example, video block 504 can be a coding block. Figure 5A The search region of can include four regions (e.g., R1, R2, R3, and R4), where each region contains the coordinates of the top-left corner of a candidate block. For example, for the top region R1, a set of coordinates (x, y) of the top-left corner of a candidate block in the top region R1 can be determined as follows:

[0084] currX - SearchRange_w <= x <= currX + SearchRange_w (1)

[0085] currY - SearchRange_h <= y <= ctuY - 1 - H (2)

[0086] In the above expressions (1) and (2), currX and currY represent the horizontal and vertical coordinates of the current video block 502 (e.g., the current coding block), respectively; W and H represent the width and height of the video block 502, respectively; and ctuY represents the vertical coordinate of the CTU 504 to which the video block 502 belongs. For example, (currX, currY) can be the coordinates of the top-left corner of the video block 502.

[0087] The size of the search range (SearchRange_w, SearchRange_h) can be set to be proportional to the block size (W, H). For example, SearchRange_w and SearchRange_h can be configured as follows:

[0088] SearchRange_w = a * W (3)

[0089] SearchRange_h = a * H (4)

[0090] In the above expressions (3) and (4), "a" can be a constant that controls the trade-off between gain and complexity. In one implementation, "a" can be equal to 5.

[0091] For example, for the bottom-left region R2, a set of coordinates (x, y) of the top-left corner of a candidate block in the bottom-left region R2 can be determined as follows:

[0092] currX - SearchRange_w <= x <= currX - 1 - W (5)

[0093] currY + 1 <= y <= ctuY + ctuH - 1 - H (6)

[0094] In the above expression (6), ctuH represents the height of the CTU 504.

[0095] For example, for the left region R3, a set of coordinates (x, y) of the top-left corner of a candidate block in the left region R3 can be determined as follows:

[0096] currX - SearchRange_w <= x <= ctuX - 1 - W (7)

[0097] ctuY - 1 - H <= y <= currY (8)

[0098] In the above expressions (7) and (8), ctuX and ctuY represent the horizontal and vertical coordinates of the CTU 504, respectively. For example, (ctuX, ctuY) can be the coordinates of the top-left corner of the CTU 504.

[0099] For example, for the top-left region R4, a set of coordinates (x, y) of the top-left corner of a candidate block in the top-left region R4 can be determined as follows:

[0100] ctuX - W < = x < = currX - W (9)

[0101] ctuY - H < = y < = currY - H (10)

[0102] Referring to Figure 5A Three CTUs 506, 508, and 510 can be above the CTU 504. For example, the CTU 508 can be on top of and adjacent to the CTU 504. The CTU 506 can be on the left of and adjacent to the CTU 508. The CTU 510 can be on the right of and adjacent to the CTU 508. Another CTU 512 can be on the left of and adjacent to the CTU 504. The CTU 506, the CTU 508, the CTU 510, the CTU 512, and the CTU 504 are in the same video frame.

[0103] The top region R1 can include a first portion in the CTU 506, a second portion in the CTU 508, and a third portion in the CTU 510. A width 526 of the top region R1 in a horizontal direction (e.g., x direction) can be 2a*W. A distance 518 between a top boundary of the top region R1 and the video block 502 in a vertical direction (e.g., y direction) can be a*H. A distance 524 between the top region R1 and the CTU 504 in the vertical direction can be H.

[0104] The left region R3 can be below and adjacent to the top region R1 and can include a first portion in the CTU 506 and a second portion in the CTU 512. A distance 522 between the left region R3 and the CTU 504 in the horizontal direction can be W.

[0105] The bottom-left region R2 can be in the CTU 512. The bottom-left region R2 can be below and adjacent to the left region R3. A distance 520 between a bottom boundary of the bottom-left region R2 and a bottom boundary of the CTU 512 in the vertical direction can be H. The bottom-left region R2 can have the same width as the left region R3 in the horizontal direction.

[0106] The top-left region R4 can include a first portion in the CTU 506, a second portion in the CTU 508, a third portion in the CTU 512, and a fourth portion in the CTU 504. The top-left region R4 can be below and adjacent to the top region Rl, and can be to the right of and adjacent to the left region R3. The top-left region R4 can be a distance of W 516 from the video block 502 in a horizontal direction, and a distance of H 514 from the video block 502 in a vertical direction.

[0107] In some implementations, a difference metric such as SAD can be used as a cost function to identify a reference block in a search region. For example, within a search region, the video encoder 20 or the video decoder 30 can search for a template having a minimum SAD relative to a template of a video block, and can select a candidate block that is the template having the minimum SAD as a prediction block.

[0108] In some implementations, intra-TMP can be enabled for CUs whose size is smaller than or equal to 64 in width and height. The maximum CU size intended for intra-TMP can be configurable. In some implementations, the intra-TMP mode can be signaled at the CU level by an indication flag.

[0109] Terms such as "top," "bottom," "left," "right," "upper," "lower," "above," "below," and the like are used to describe the relative positions of regions for the purpose of describing the relative positions of regions in a particular coordinate system, as shown in the illustrated coordinate system of Figure 5A and the additional figures below. They are not intended to limit the particular positions or orientations of the regions. For example, a first region described as being in the upper left corner of a second region in the illustrated coordinate system can be considered to be at one distance relative to the second region in the vertical direction and another distance relative to the second region in the horizontal direction in that coordinate system.

[0110] Figure 5B is a block diagram illustrating a reference block 554 in a search region 560 according to some implementations of the disclosure. A template 556 of the reference block 554 can be matched to a template 552 of a video block 550 in the same video frame 561. Each of the template 556 and the template 552 has an L-shape. The reference block 554 can be determined as a prediction block for the video block 550. For example, the template 556 of the reference block 554 can have a minimum SAD relative to the template 552 of the video block 550 when compared to templates of other candidate blocks in the search region 560. Then, prediction samples in a prediction block can be generated based on the reference block 554. For example, samples of the reference block 554 can be considered as prediction samples in the prediction block.

[0111] Consistent with some embodiments of the present disclosure, an improved intra-TMP scheme is disclosed herein to significantly improve the intra coding efficiency of ECM. For example, the improved intra-TMP scheme disclosed herein can include methods and apparatuses to improve and simplify the coding efficiency of the intra-TMP tool by modifying the search region of the reconstructed samples in the same video frame for the synthesis of the coded block. As discussed with reference to Figure 5A As discussed with reference to Figure 5A and Figure 6 and Figures 7A-7B ,

[0112] Figure 6 is a flowchart of an exemplary method 600 for intra-TMP of a video frame of a video according to some embodiments of the present disclosure. The method 600 can be implemented by a processor associated with the video encoder 20 or the video decoder 30, and the method 600 can include steps 602-606 as described below. Some steps can be optional for performing the disclosure provided herein. In addition, some steps can be performed concurrently, or in a different order than as shown. Figure 6 . Figure 7A is a diagram of a modified search region for an intra-TMP mode according to some embodiments of the present disclosure. Figure 7B is another diagram of a modified search region for an intra-TMP mode according to some embodiments of the present disclosure. Figures 8A-8C shows a method of generating a sample value for an unreconstructed sample according to some embodiments of the present disclosure. The following Figure 6 , Figures 7A-7B and Figures 8A-8C are described together below.

[0113] In step 602, as shown in Figure 6 , the processor can determine a search region for a video block from a video frame. Reference is made to Figure 7AA video frame can include CTU 706, CTU 708, CTU 710, CTU 712, and CTU 704, where video block 702 is included in current CTU 704. The search region can include a first region (top-left region R4) that is a first distance 714 from video block 702 in a first direction (e.g., vertical direction) and a second distance 716 from video block 702 in a second direction (e.g., horizontal direction) that is perpendicular to the first direction. The first distance 714 in the first direction can be equal to the height of video block 702 in the first direction. The second distance 716 in the second direction can be equal to the width of video block 702 in the first direction. Figure 7A The first region R4 in can be similar to the top-left region R4 of Figure 5A , and similar descriptions will not be repeated here.

[0114] The search region can also include a second region (next-left region R2’) adjacent to the first region R4 in the first direction and a second distance 716 from the video block in the second direction. The second region R2’ can have the same width as the first region R4 in the second direction. The height of the second region R2’ can be equal to the height of video block 702 in the first direction. The search region can also include a third region (next-top-left region R4’) adjacent to the first region R4 in the second direction and a first distance 714 from the video block in the first direction. The third region R4’ can have the same height as the first region R4 in the first direction. The width of the third region R4’ is equal to the width of video block 702 in the second direction.

[0115] For example, the improved intra-TMP scheme disclosed herein can also include additional reconstructed samples (which are located above and to the left of the current video block 702) into the search region of the intra-TMP mode. These additional above-and-left reconstructed samples are located to the right of the first region R4 or below the first region R4. Specifically, as shown in Figure 7A , two new regions R2’ and R4’ can be included in the search region. For the second region R2’, a set of coordinates (x, y) of the top-left corner of a candidate block within the second region R2’ can be determined as follows:

[0116] ctuX - W <== x <== currX - W (11)

[0117] currY - H + 1 <== y <== currY (12)

[0118] For the third region R4’, a set of coordinates (x, y) of the top-left corner of a candidate block within the third region R4’ can be determined as follows:

[0119] currX - W + 1 <== x <== currX (13)

[0120] ctuY - H <= y <= currY - H (14)

[0121] Referring Figure 7A , the search area can also include a fourth region (e.g., a lower-left region R2") adjacent to the second region R2' in the first direction and a second distance 716 from the video block in the second direction. The fourth region R2" can have the same width as the first region R4 and the second region R2". The search area can also include a fifth region (e.g., an upper-right region R4") adjacent to the third region R4' in the second direction and the first distance 714 from the video block in the first direction. The fifth region R4" can have the same height as the first region R4 and the third region R4' in the first direction. A third distance 718 in the second direction between the fifth region R4" and the right boundary of the CTU 704 is equal to the width of the video block 702. A fourth distance 720 in the first direction between the fourth region R2" and the bottom boundary of the CTU 704 is equal to the height of the video block 702. The fourth region R2" is a second distance 716 from the video block 702 in the second direction. The fifth region R4" is a first distance 714 from the video block 702 in the first direction.

[0122] For example, due to the binary- or ternary-tree-based partitioning structure, some samples located at the upper-right or lower-left positions relative to the video block 702 can be available (e.g., already reconstructed) before the video block 702 is encoded / decoded, in addition to the samples located at the top or left side of the video block 702. Therefore, to further improve the intra-TMP performance, the improved intra-TMP scheme disclosed herein can also include the reconstructed samples located at the upper-right or lower-left positions relative to the video block 702 into the search area of the intra-TMP mode. Specifically, as shown in FIG. 7B, two additional regions R2" and R4" are included in the search area. For the fourth region R2", a set of coordinates (x, y) of the top-left corner of a candidate block within the fourth region R2" can be determined as follows: Figure 7A

[0123] ctuX - W <= x <= currX - W (15)

[0124] currY + 1 <= y <= ctuY + ctuH - 1 - H (16)

[0125] For the fifth region R4", a set of coordinates (x, y) of the top-left corner of a candidate block within the fifth region R4" can be determined as follows:

[0126] ​currX + 1 <= x <= ctuX + ctuW - 1 - W (17)

[0127] ctuY - H <= y <= currY - H (22a)

[0128] In the above expressions (16) and (17), ctuW and ctuH represent the width and height of the CTU 704, respectively.

[0129] In some examples, the fourth region R2" and the fifth region R4" can be determined based at least in part on a picture size of the video frame and a template size of the template of the video block 702 (or a template size of a reference block related to the video block 702). The template size can include a template width in the second direction (e.g., the horizontal direction) and a template height in the first direction (e.g., the vertical direction). The picture size can include a picture width in the second direction and a picture height in the first direction. A left-side boundary of the fourth region R2" can be determined based at least in part on the template width, and a bottom boundary of the fourth region R2" can be determined based at least in part on the picture height. A right-side boundary of the fifth region R4" can be determined based at least in part on the picture width, and a top boundary of the fifth region R4" can be determined based at least in part on the template height.

[0130] For example, for the fourth region R2", a set of coordinates (x, y) of a top-left corner of a candidate block within the fourth region R2" can be determined as follows:

[0131] max(templateWidth, ctuX - W) <= x <= currX - W (19a)

[0132] currY + 1 <= y <= min(picHeight - H, ctuY + ctuH - 1 - H) (20a)

[0133] For example, for the fifth region R4", a set of coordinates (x, y) of a top-left corner of a candidate block within the fifth region R4" can be determined as follows:

[0134] currX + 1 <= x <= min(picWidth - W, ctuX + ctuW - 1 - W) (21a)

[0135] max(templateHeight, ctuY - H) <= y <= currY - H (22a)

[0136] In the above expressions (19a), (20a), (21a), and (22a), picWidth and picHeight represent the picture width and the picture height, respectively; and templateWidth and templateHeight represent the template width and the template height, respectively.

[0137] In some examples, to control the total number of candidate blocks to be searched, the width of the fourth region R2" and the width of the fifth region R4" can be constrained. The constrained fourth region R2" can be determined as follows:

[0138] max(ctuX - W, currX - searchRange) <= x <= currX - W (19b)

[0139] currY + 1 <= y <= ctuY + ctuH - 1 - H (20b)

[0140] The constrained fifth region R4" can be determined as follows:

[0141] currX + 1 <= x <= ctuX + min(ctuX + ctuW - 1 - W, searchRange) (21b)

[0142] ctuY - H <= y <= currY - H (22b)

[0143] In the above expressions (19b) and (21b), searchRange can be the maximum horizontal region size of R2" and R4". For example, the value of searchRange can be determined to be proportional to the width of the video block 702, i.e., searchRange = m*W, where m represents a numerical coefficient.

[0144] Further referring to Figure 7A The search region can also include a sixth region Rl (a top region Rl), a seventh region R3 (e.g., a left region R3), and an eighth region R2 (a lower left region R2). The sixth region Rl can be adjacent to the first region, the third region, and the fifth region in the vertical direction. A top boundary of the sixth region Rl is at a fifth distance in the vertical direction between the video block 702 equal to a product of the height of the video block and a constant a (e.g., a*H). A sixth distance 724 in the vertical direction between the sixth region Rl and the CTU 704 is equal to the height of the video block (e.g., H).

[0145] The seventh region R3 can be adjacent to the sixth region R1 in the vertical direction and adjacent to the first region R4 and the second region R2' in the horizontal direction. The height of the seventh region R3 is equal to the sum of the height of the first region R4 and the height of the second region R2'. The distance 722 between the seventh region R3 and the CTU 704 in the horizontal direction is equal to the width (e.g., W) of the video block.

[0146] The eighth region R2 can be adjacent to the seventh region R3 in the vertical direction and adjacent to the fourth region R2" in the horizontal direction. The height of the eighth region R2 is equal to the height of the fourth region R2", and the width of the eighth region R2 is equal to the width of the seventh region R3. The distance 720 between the eighth region R2 and the bottom boundary of the CTU 712 in the vertical direction is equal to the height (e.g., H) of the video block.

[0147] The sixth region R1, the seventh region R3, and the eighth region R2 can be similar to the top region R1, the left region R3, and the lower-left region R2, respectively, of Figure 5A , which will not be repeated here. In some implementations, the first region R4, the second region R2', and the seventh region R3 can be merged into a single region R5, as shown in Figure 7B .

[0148] Consistent with some implementations of the present disclosure, a search region has a first search region size in a first direction and a second search region size in a second direction. A constraint can be applied to all regions (e.g., R1, R2, R3, R4, R2', R4, R2", and R4") included in the search region, such that each of these regions is within an area determined by the first search region size, the second search region size, and the location of the video block 702 in the video frame. In some particular implementations, the location of the video block 702 can be at the top-left sample location 730 of the video block 702. The center of the area can be located at the top-left sample location 730 of the video block 702. In some implementations, the height of the area can be twice the first search region size in the first direction, and the width of the area can be twice the second search region size in the second direction.

[0149] For example, for each region (e.g., R1, R2, R3, R4, R2', R4, R2", and R4"), a set of coordinates (x, y) of the top-left corner of a candidate block within that region can be determined as follows:

[0150] currX - searchRangeX <= x <= currX + searchRangeX (21)

[0151] currY - searchRangeY <== y <== curry + searchRangeY (24)

[0152] In the above expressions (23) and (24), searchRangeX and searchRangeY represent the second search region size (e.g., the horizontal search region size) and the first search region size (e.g., the vertical search region size), respectively. In some implementations, the first search region size and the second search region size can be proportional to the height H and the width W of the video block 702, respectively. For example, searchRangeX = m*W, and searchRangeY = n*H (e.g., m and n can be set to 5; or, m and n can be set to any other suitable values).

[0153] As for the second region R2' and the fourth region R4', all the samples included in the regions R2' and R4' are located at the top-left of the video block 702 and have been reconstructed before the video block 702 is encoded / decoded, similar to the samples of the regions Rl, R2, R3, and R4. Therefore, when a candidate block is obtained from the regions R2' and R4', no availability check needs to be performed on the candidate block. The availability check is described in more detail below.

[0154] As for the fourth region R2" and the fifth region R4", the samples in the regions R2" and R4" can be obtained from search regions below or to the right of the video block 702, which can not have been reconstructed before the video block 702 is encoded / decoded. For example, both reconstructed samples and unreconstructed samples are included in the fourth region R2" and the fifth region R4". Accordingly, when a candidate block is obtained from the regions R2" and R4", an availability check can be performed to determine whether all the samples in the respective candidate block have been reconstructed, as described in more detail below.

[0155] Referring back to Figure 6 In step 604, the processor can determine a reference block from the search regions, where a template of the reference block matches the template of the video block. Specifically, the processor can determine a plurality of candidate blocks from the search regions, and determine a plurality of templates for the plurality of candidate blocks, respectively. The processor can determine a template that matches the template of the video block from the plurality of templates, and determine the reference block as a first candidate block having the template that matches the template of the video block.

[0156] In some implementations, the processor can determine that the plurality of candidate blocks can include one or more second candidate blocks from the fourth region R2" or the fifth region R4". The one or more second candidate blocks from the fourth region R2" or the fifth region R4" can include unreconstructed samples. For each second candidate block, the processor can perform an availability check on the second candidate block.

[0157] In some examples, for each second candidate block, the processor can determine whether the second candidate block includes at least one unreconstructed sample. In response to determining that the second candidate block includes at least one unreconstructed sample, the processor can determine that the second candidate block fails the availability check. Alternatively, in response to determining that the second candidate block does not include an unreconstructed sample, the processor can determine that the second candidate block passes the availability check.

[0158] In some cases, for each second candidate block, the processor can determine whether a sample at a bottom-right corner of the second candidate block is reconstructed. In response to the sample at the bottom-right corner of the second candidate block not being reconstructed, the processor can determine that the second candidate block fails the availability check. Alternatively, in response to the sample at the bottom-right corner of the second candidate block being reconstructed, the processor can determine that the second candidate block passes the availability check.

[0159] In response to determining that a second candidate block passes the availability check, the processor can retain the second candidate block in the plurality of candidate blocks. Alternatively, in response to determining that a second candidate block fails the availability check, the processor can remove the second candidate block from the plurality of candidate blocks. For example, if all samples in a candidate block from the fourth region R2” or the fifth region R4” are reconstructed (i.e., available), the candidate block is considered as a valid intra-TMP candidate block for predicting the video block. Otherwise (i.e., at least one sample in the candidate block has not been reconstructed), the candidate block is not allowed to be referenced in the prediction of the video block.

[0160] In another example, assume that all candidate blocks have a rectangular shape and the coding order (based on the corresponding partition structure) of the video block in one picture / slice is from top to bottom and from left to right. When checking the availability of a candidate block from the fourth region R2” or the fifth region R4”, the processor only needs to check the availability of the coordinate of the bottom-right corner of the candidate block to determine whether all samples in the candidate block are reconstructed. Specifically, when the bottom-right sample of the candidate block has been reconstructed, it is guaranteed that all samples in the candidate block are available. Otherwise, when the bottom-right sample of the candidate block has not been reconstructed, there exists at least one sample in the candidate block that is not yet available to be referenced (i.e., the candidate block is not yet ready to be used as an intra-TMP candidate block).

[0161] In the above example, a candidate block in region R2” and R4” is allowed to be used as an intra-TMP candidate block only when all the samples within the candidate block have been reconstructed before the video block is encoded / decoded. This constraint can reduce the total number of valid intra-TMP candidate blocks that can be used to predict the video block, and thus limit the overall coding performance of the intra-TMP mode. To address this issue, the improved TMP scheme disclosed herein can further generate sample values for unreconstructed samples (e.g., unavailable samples) of a candidate block from neighboring available samples of the unreconstructed samples, such that the candidate block can become a valid candidate block for the intra-TMP mode.

[0162] In particular, in response to determining that the second candidate block from the fourth region R2” or the fifth region R4” fails the availability check, the processor can determine unreconstructed samples in the second candidate block. The processor can update the second candidate block by generating sample values for the unreconstructed samples, such that the updated second candidate block can be retained in the plurality of candidate blocks. In some embodiments, the sample values for the unreconstructed samples can be generated using at least one of a horizontal repetition padding method, a vertical repetition padding method, an adaptive repetition method, or a collocated copy method, as described in more detail below.

[0163] For example, the horizontal repetition padding method can be applied to generate sample values for unreconstructed samples in a candidate block from region R2” or region R4”. In particular, the sample value for each unreconstructed sample (each unavailable sample) in the candidate block can be generated by directly copying the sample value of the nearest reconstructed sample (the nearest available sample) in the horizontal direction. For example, as shown in FIG. 8, a candidate block 802 from region R2” or region R4” can include reconstructed samples (e.g., available samples) a, b, c, d, e, f, g, h, i, and j. The candidate block 802 can also include unreconstructed samples (e.g., unavailable samples) 804, 806, 808, 810, 812, and 814. For the unreconstructed samples 804, 806, and 808 in one row, the nearest reconstructed sample in the horizontal direction is the reconstructed sample i. Thus, the sample value for each unreconstructed sample 804, 806, and 808 can be generated by directly copying the sample value of the nearest reconstructed sample i in the horizontal direction. For the unreconstructed samples 810, 812, and 814 in another row, the nearest reconstructed sample in the horizontal direction is the reconstructed sample j. Thus, the sample value for each unreconstructed sample 810, 812, and 814 can be generated by directly copying the sample value of the nearest reconstructed sample j in the horizontal direction. Figure 8A

[0164] ​In another example, a vertical repetition padding method can be applied to generate sample values for unreconstructed samples in a candidate block from region R2" or region R4". Specifically, a sample value for each unreconstructed sample in the candidate block can be generated by directly copying a sample value of a nearest reconstructed sample in a vertical direction. For example, as shown in Figure 8B for unreconstructed samples 804 and 810 in a column of candidate block 802, the nearest reconstructed sample in a vertical direction is reconstructed sample f. Thus, a sample value for each unreconstructed sample 804 and 810 can be generated by directly copying a sample value of the nearest reconstructed sample f in a vertical direction. For unreconstructed samples 806 and 812 in another column of candidate block 802, the nearest reconstructed sample in a vertical direction is reconstructed sample g. Thus, a sample value for each unreconstructed sample 806 and 812 can be generated by directly copying a sample value of the nearest reconstructed sample g in a vertical direction. For unreconstructed samples 808 and 814 in yet another column of candidate block 802, the nearest reconstructed sample in a vertical direction is reconstructed sample h. Thus, a sample value for each unreconstructed sample 808 and 814 can be generated by directly copying a sample value of the nearest reconstructed sample h in a vertical direction.

[0165] In yet another example, an adaptive repetition padding method can be applied, in which an unreconstructed sample can be generated by copying a sample value of a nearest reconstructed sample in a horizontal direction (e.g., horizontal padding) or a sample value of a nearest reconstructed sample in a vertical direction (e.g., vertical padding) in the same candidate block. Whether to choose horizontal padding or vertical padding can be decided based on different methods. For example, a gradient filter (e.g., Sobel filter) can be applied to calculate gradients of already reconstructed samples inside the candidate block. If most of the gradients (e.g., more than 50%, 60%, or 70% of the gradients, etc.) are close to horizontal, then a horizontal repetition padding method can be applied. Otherwise (when most of the gradients are close to vertical), a vertical repetition padding method can be applied.

[0166] In yet another example, when a candidate block from region R2" or region R4" partially overlaps with a video block, a co-located copying method can be applied, in which unreconstructed samples in an overlapping region between the candidate block and the video block are directly copied from co-located samples in the candidate block. The co-located samples in the candidate block can have the same sample position in the candidate block as the unreconstructed samples in the video block. For example, as shown in Figure 8C candidate block 802 and video block 832 can have an overlapping region 830 Figure 8Cthe shaded regions in FIG. 8B), which can include six unreconstructed samples 804, 806, 808, 810, 812, and 814. That is, the six sample positions of the bottom-right portion of the candidate block 802 overlap with the six sample positions of the top-left portion of the video block 832. In other words, the six bottom-right samples in the candidate block 802 are located at the six top-left sample positions in the video block 832. Accordingly, the six unreconstructed bottom-right samples (804, 806, 808, 810, 812, 814) are generated by directly copying the six top-left samples (a, b, c, e, f, g) of the candidate block 802.

[0167] Consistent with some embodiments of the disclosure, different orders can be applied to scan the reconstructed samples in different regions to identify the best intra-TMP candidate block, which can result in various performance and complexity tradeoffs. The best intra-TMP candidate block can be the candidate block that has a template matching that of the video block. For example, the scanning order can be determined for the regions (e.g., as shown in FIG. 8B) such that the best intra-TMP candidate block can be searched in the regions according to the scanning order. Generally, the samples in the regions closer to the video block are more relevant to the samples in the video block, and thus can be scanned earlier. Based on such considerations, in a first example, the regions R1, R2, R3, R4, R2', R4', R2", and R4" in FIG. 8B can be scanned in the order of R4" -> R4' -> R1 -> R3 -> R4 -> R2' -> R2" -> R2. Figure 7A Figure 7B In a second example, as shown in FIG. 8C, when the regions R3, R4, and R2' are merged into a region R5, the corresponding order of scanning the regions R1, R2, R5, R2", R4', and R4" becomes R4" -> R4' -> R1 -> R5 -> R2" -> R2. Figure 7A Figure 7B

[0168] In some embodiments, the regions in FIG. 8B can be scanned in the order of R4" -> R4' -> R1 -> R3 -> R4 -> R2' -> R2" -> R2. Figure 7B ​​​The six regions R4', R4", R2", R1, R2, and R5 shown in the middle apply different scan order combinations. For example, the following scan order can be applied: R4' -> R2" -> R4" -> R1 -> R5 -> R2. In another example, the scan order R4' -> R4" -> R2" -> R1 -> R5 -> R2 can be applied. In yet another example, the scan order R4" -> R4' -> R2" -> R1 -> R5 -> R2 can be applied. In yet another example, the scan order R4" -> R2" -> R4' -> R1 -> R5 -> R2 can be applied. In yet another example, the scan order R2" -> R4" -> R4' -> R1 -> R5 -> R2 can be applied. In yet another example, the scan order R2" -> R4' -> R4" -> R1 -> R5 -> R2 can be applied. In yet another example, the scan order R4' -> R1 -> R5 -> R2 -> R2" -> R4" can be applied.

[0169] Consistent with some embodiments of the disclosure, the processor can perform an iterative search method on the search region with at least one scaling factor to identify a candidate block having a template matching the template of the video block. The candidate block having a template matching the template of the video block (also referred to as the best intra-TMP candidate block) is determined as the reference block for the video block. That is, to further reduce the complexity, the iterative search method can be applied to iteratively identify the best intra-TMP candidate block.

[0170] Specifically, in the first step of the iterative search method, the region can be sub-sampled by a scaling factor scale to find one or more initial intra-TMP candidate blocks. Then, around each initial intra-TMP candidate block, local refinement can be performed with gradually reduced search steps to find a refined intra-TMP candidate block. In each round of local refinement, the current search step is reduced to half of the previous search step. For example, in the first round of local refinement, the search step can be reduced to scale / 2, while in the second round of local refinement, the search step can be reduced to scale / 4, and so on. In another example, in the first round of local refinement, the search step can be reduced to floor(scale / 2) of the previous search step (e.g., current search step = floor(scale / 2) x previous search step), while in the second round of local refinement, the search step can be reduced to floor(scale / 4) of the previous search step (e.g., current search step = floor(scale / 4) x previous search step), and so on, where floor(·) is a floor function. Such refinement can be performed continuously until the search step is reduced to less than 1 or the refined intra-TMP candidate block does not change in the refinement process.

[0171] Additionally, during the search process, a termination criterion can be applied to terminate the entire search process across all regions or the search process in a particular region when a termination condition is satisfied. For example, the distortion value of the refined intra-TMP candidate block obtained so far during the search process can be used to determine the termination criterion. Specifically, when the distortion value is less than a predefined threshold, it is determined that the refined intra-TMP candidate block is good enough such that the search process can stop. Otherwise (i.e., the distortion value is equal to or greater than the predefined threshold), it is determined that the current refined intra-TMP candidate needs further improvement, and thus the search process continues. In some embodiments, the above termination scheme can be used to jump out (or terminate) the entire search process across all regions. Alternatively, the above termination scheme can be independently used for each region. For example, when the termination condition is satisfied (e.g., when the distortion value is less than the predefined threshold), the search process for the samples in the current region can be terminated while the search for the samples in the subsequent regions can continue to be performed.

[0172] In some examples, an adaptive refinement scheme can be applied in the search process, in which different subsampling scaling factors can be applied to different regions. For example, a region closer to the video block can be more relevant to the video block and can be searched with a smaller scaling factor, while another region farther away from the video block can be less relevant to the video block and can be searched with a larger scaling factor. In another example, scaling factor 2 (scale = 2) can be used for regions R4" and R2", and scaling factor 3 (scale = 3) can be used for other regions. In another example, scaling factor 2 (scale = 2) can be used for all regions.

[0173] Consistent with some embodiments of the disclosure, video encoder 20 can identify the best region from the search regions, generate a bitstream to include an index of the best region, and transmit the bitstream to video decoder 30. Video decoder 30 can then receive the bitstream including the index of the best region, and identify the best region from the search regions based on the index. Video decoder 30 can determine the reference block from the best region.

[0174] For example, the complexity of the intra-TMP mode is proportional to the number of candidate blocks that are examined to identify the best intra-TMP candidate block. To reduce the intra-TMP complexity, an explicit region-based intra-TMP mode can be used. For example, the total search region for the intra-TMP mode is divided into multiple regions, e.g., as Figure 7BThe regions R4", R4', R1, R5, R2", and R2" are shown such that each region can provide an independent intra-TMP candidate (i.e., the candidate block with the minimum distortion metric within the region). At the encoder side, multiple regions can be tested by rate-distortion optimization (RDO), and the index of the best region (optimal region) is identified. The index of the best region can be sent to the video decoder 30 in the bitstream. At the decoder side, after receiving the index of the best region, only the intra-TMP search process needs to be performed in the region corresponding to the index of the best region.

[0175] Referring back to FIG. 6, Figure 6 At step 606, the processor can determine the prediction samples for the video block based on the reference block. For example, the prediction samples in the prediction block can be the same as the corresponding samples in the reference block. In some embodiments, the bitstream can be generated to include encoded data associated with the difference between the video block and the prediction block, as described above with reference to FIG. 5. Figure 2 In some embodiments, the bitstream can further include an indication indicating that the intra prediction mode is intra-TMP.

[0176] Figures 9A-9B An iterative search method is shown performed on a search region with a scaling factor, according to some embodiments of the disclosure. Referring to FIG. 9, Figure 9A Taking horizontal search as an example, the region 902 can be initially subsampled by a scaling factor a to find one or more initial candidate blocks. For example, assuming a = 3, the initial search step can be equal to floor(a) * W = 3W, where W is the width of the video block, and floor(·) is the floor function. Then, around each initial candidate block, local refinement can be performed with gradually reduced search steps to find refined candidate blocks. For example, in the first round of local refinement for the first initial candidate block 904, the search step can be reduced to floor(a / 2) * W = floor(3 / 2) * W = W, so that a refined candidate block 906 can be found, as shown. Figure 9B After the first round of local refinement for the first initial candidate block 904, the search step is reduced to floor(a / 4) * W = floor(3 / 4) * W = 0. Since the search step becomes smaller than 1, the local refinement around the first initial candidate block 904 can be terminated.

[0177] Similar operations can be performed on other initial candidate blocks until the search process in the region 902 terminates. For example, if a termination condition is satisfied (e.g., the distortion value of the refined candidate block obtained so far in the search process is less than a predefined threshold), it is determined that the refined candidate block is good enough such that the search process can stop. Otherwise (i.e., the distortion value is equal to or greater than the predefined threshold), it is determined that the current refined candidate block needs further improvement, and thus the search process continues. If the entire region 902 has been searched, the search process can continue in other regions until the termination condition is satisfied.

[0178] Figure 10 is a flowchart of another exemplary method 1000 for intra-TMP of video frames of a video according to some embodiments of the present disclosure. The method 1000 can be implemented by a processor associated with the video encoder 20 or the video decoder 30, and can include steps 1002-1006 as described below. Some steps can be optional for performing the disclosure provided herein. Also, some steps can be performed concurrently, or in a different order than as shown. Figure 10

[0179] In step 1002, as shown in Figure 10 the processor can determine a search region for a video block from a video frame. The search region can have a first search region size in a first direction (e.g., a vertical search region size in a vertical direction) and a second search region size in a second direction (e.g., a horizontal search region size in a horizontal direction). The search region can include a plurality of regions, each region within an area determined by the first search region size, the second search region size, and a location of the video block in the video frame. For example, each region of the plurality of regions can be determined based on the above expressions (23) and (24).

[0180] In some embodiments, a center of the region can be located at the location of the video block. For example, the location of the video block can be at a top-left corner sample location of the video block. A height of the region can be twice the first search region size in the first direction, and a width of the region can be twice the second search region size in the second direction. The first search region size and the second search region size can be proportional to a height and a width of the video block, respectively.

[0181] ​In some embodiments, the plurality of regions can include at least one of: a first region R4 that is a first distance from the video block in the first direction and a second distance from the video block in the second direction; a second region R2' that is adjacent to the first region R4 in the first direction and the second distance from the video block in the second direction; a third region R4' that is adjacent to the first region R4 in the second direction and the first distance from the video block in the first direction; a fourth region R2" that is adjacent to the second region R2' in the first direction and the second distance from the video block in the second direction; or a fifth region R4" that is adjacent to the third region R4' in the second direction and the first distance from the video block in the first direction.

[0182] In some embodiments, the fourth region R2" and the fifth region R4" can be determined based at least in part on a template size of a template of the video block (or a template size of a reference block corresponding to the video block) and a picture size of the video frame. The template size can include a template width in the second direction and a template height in the first direction. The picture size can include a picture width in the second direction and a picture height in the first direction. A left boundary of the fourth region R2" can be determined based at least in part on the template width, and a bottom boundary of the fourth region R2" can be determined based at least in part on the picture height. A right boundary of the fifth region R4" can be determined based at least in part on the picture width, and a top boundary of the fifth region R4" can be determined based at least in part on the template height.

[0183] In some embodiments, the video frame is divided into a plurality of CTUs, and the video block is located within one of the plurality of CTUs. The search region further includes at least one of: a sixth region R1 that is adjacent to the first region R4, the third region R4', and the fifth region R4" in the first direction, wherein a distance between a top boundary of the sixth region R1 and the video block in the first direction is equal to a product of a height of the video block and a constant "a" as described above, and a distance between the sixth region R1 and the one CTU in the first direction is equal to the height of the video block; a seventh region R3 that is adjacent to the sixth region R1 in the first direction and adjacent to the first region R4 and the second region R2' in the second direction, wherein a distance between the seventh region R3 and the one CTU in the second direction is equal to a width of the video block; or an eighth region R2 that is adjacent to the seventh region R3 in the first direction and adjacent to the fourth region R2" in the second direction, wherein a distance between the eighth region R2 and the one CTU in the second direction is equal to the width of the video block.

[0184] In step 1004, the processor can determine a reference block from the search region. A template of the reference block matches a template of the video block. For example, the process in step 1004 can be performed as described above with reference to FIG. 3. Figure 6The operations described in step 604 of FIG. 6 are similar to the operations described in step 604 of FIG. 6, and similar descriptions will not be repeated herein.

[0185] In some embodiments, the processor can determine a scan order of the plurality of regions, and search the reference block in the plurality of regions based on the scan order.

[0186] In step 1006, the processor can determine the prediction samples for the video block based on the reference block. For example, the operations performed in step 1006 can be similar to the operations described above with reference to step 606 of FIG. 6, and similar descriptions will not be repeated herein. Figure 6 The operations described in step 606 of FIG. 6 are similar to the operations described in step 606 of FIG. 6, and similar descriptions will not be repeated herein.

[0187] Figure 11 A computing environment 1111 is shown coupled with a user interface 1150. The computing environment 1111 can be part of a data processing server. The computing environment 1111 includes a processor 1120, a memory 1130, and an input / output (I / O) interface 1140.

[0188] The processor 1120 generally controls the overall operation of the computing environment 1111, such as operations associated with displaying, data acquisition, data communication, and image processing. The processor 1120 can include one or more processors for executing instructions to perform all or some of the steps in the above-described methods. In addition, the processor 1120 can include one or more modules that facilitate interaction with other components of the computing environment 1111. The processor can be a central processing unit (CPU), a microprocessor, a microcontroller, a graphics processing unit (GPU), etc.

[0189] The memory 1130 is configured to store various types of data to support the operation of the computing environment 1111. The memory 1130 can include predetermined software 1132. Examples of such data include instructions for any application or method operating on the computing environment 1111, video data sets, image data, etc. The memory 1130 can be implemented by using any type of volatile or non-volatile memory devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0190] The I / O interface 1140 provides an interface between the processor 1120 and peripheral interface modules (e.g., a keyboard, a click wheel, a button, etc.). The buttons can include, but are not limited to, a home button, a start scanning button, and a stop scanning button. The I / O interface 1140 can be coupled with an encoder and a decoder.

[0191] In an embodiment, there is also provided a non-transitory computer- readable storage medium comprising a plurality of programs, e.g., in the memory 1130, which can be executed by the processor(s) 1120 in the computing environment 1111 for performing the above-described methods and / or storing a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. In one example, the plurality of programs can be executed by the processor(s) 1120 in the computing environment 1111 for receiving (e.g., from a video encoder 20 in a video encoding device 10) Figure 2 and / or associated one or more syntax elements, etc.) from a video encoder 20 in a video encoding device 10) and can also be executed by the processor(s) 1120 in the computing environment 1111 for performing the above-described decoding method in accordance with the received bitstream or data stream. In another example, the plurality of programs can be executed by the processor(s) 1120 in the computing environment 1111 for performing the above-described encoding method to encode video information (e.g., video blocks representing video frames, and / or associated one or more syntax elements, etc.) into a bitstream or data stream and can also be executed by the processor(s) 1120 in the computing environment 1111 for transmitting (e.g., to a video decoder 30 in a video decoding device 20) the bitstream or data stream. Alternatively, the non-transitory computer-readable storage medium can store a bitstream or data stream comprising encoded video information (e.g., video blocks representing encoded video frames, and / or associated one or more syntax elements, etc.) produced using, e.g., the above-described encoding method by an encoder (e.g., a video encoder 20 in a video encoding device 10) for use in decoding video data by a decoder (e.g., a video decoder 30 in a video decoding device 20). The non-transitory computer-readable storage medium can be, e.g., a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, an optical data storage device, and so on. Figure 3 Figure 2 Alternatively, the non-transitory computer-readable storage medium can store a bitstream or data stream comprising encoded video information (e.g., video blocks representing encoded video frames, and / or associated one or more syntax elements, etc.) produced using, e.g., the above-described encoding method by an encoder (e.g., a video encoder 20 in a video encoding device 10) for use in decoding video data by a decoder (e.g., a video decoder 30 in a video decoding device 20). The non-transitory computer-readable storage medium can be, e.g., a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disc, an optical data storage device, and so on. Figure 3

[0192] In an embodiment, there is provided a bitstream generated by the above-described encoding method or a bitstream to be decoded by the above-described decoding method. In an embodiment, there is provided a bitstream comprising encoded video information generated by the above-described encoding method or to be decoded by the above-described decoding method.

[0193] In an embodiment, there is also provided a computing device comprising: one or more processors (e.g., the processor(s) 1120); and a non-transitory computer-readable storage medium or memory 1130 having stored therein a plurality of programs which can be executed by the one or more processors, wherein the one or more processors, when executing the plurality of programs, are configured to perform the above-described methods.

[0194] ​​In an embodiment, there is also provided a computer program product having instructions for storing or transmitting a bitstream comprising encoded video information generated by the encoding method described above or to be decoded by the decoding method described above. In an embodiment, there is also provided a computer program product comprising a plurality of programs, for example in the memory 1130, which can be executed by the processor 1120 in the computing environment 1111 for performing the methods described above. For example, the computer program product can comprise a non-transitory computer readable storage medium.

[0195] In an embodiment, the computing environment 1111 can be implemented by one or more ASICs, DSPs, Digital Signal Processing Devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.

[0196] In an embodiment, there is also provided a method of storing a bitstream comprising storing the bitstream on a digital storage medium, wherein the bitstream comprises encoded video information generated by the encoding method described above or to be decoded by the decoding method described above.

[0197] In an embodiment, there is also provided a method for transmitting a bitstream generated by the encoder described above. In an embodiment, there is also provided a method for receiving a bitstream to be decoded by the decoder described above.

[0198] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosure as set forth in the above description and the associated drawings are presented for purposes of illustration and are not intended to be exhaustive or to limit the disclosure. Many modifications, variations, and alternative implementations will be apparent to those of ordinary skill in the art, as will further appear by the foregoing description and the associated drawings.

[0199] Unless specifically stated otherwise, the order of steps in the methods according to the disclosure is not intended to be limiting, and the steps in the methods according to the disclosure can be performed in any order according to the practical situation, and at least one of the steps in the methods according to the disclosure can be adjusted, combined, or deleted according to the practical situation.

[0200] The examples are chosen and described in order to explain the principles of the disclosure and to enable others skilled in the art to best utilize the various embodiments and implementations of the disclosure, and to best utilize the various embodiments and implementations with various modifications as are suited to the particular use contemplated. Therefore, it is to be understood that the scope of the disclosure is not to be limited to the specific examples disclosed and that modifications and other embodiments are intended to be included within the scope of this disclosure.

Claims

1. A method for video decoding, comprising: determining, by a decoder, a search region for a video block from a video frame, wherein the search region has a first search region size in a first direction and a second search region size in a second direction perpendicular to the first direction, wherein the search region comprises a plurality of regions each within an area determined by the first search region size, the second search region size, and a location of the video block in the video frame; determining, by the decoder, a reference block from the search region, wherein a template of the reference block matches a template of the video block; and determining, by the decoder, prediction samples for the video block based on the reference block.

2. The method of claim 1, wherein a center of the region is at the location of the video block, a height of the region is twice the first search region size in the first direction, and a width of the region is twice the second search region size in the second direction.

3. The method of claim 1, wherein the first search region size and the second search region size are proportional to a height and a width of the video block, respectively.

4. The method of claim 1, wherein the plurality of regions comprises at least one of: a first region, wherein the first region is a first distance from the video block in the first direction and a second distance from the video block in the second direction; a second region, wherein the second region is adjacent to the first region in the first direction and the second distance from the video block in the second direction; a third region, wherein the third region is adjacent to the first region in the second direction and the first distance from the video block in the first direction; a fourth region, wherein the fourth region is adjacent to the second region in the first direction and the second distance from the video block in the second direction; or a fifth region, wherein the fifth region is adjacent to the third region in the second direction and the first distance from the video block in the first direction.

5. The method of claim 4, wherein the fourth region and the fifth region are determined based at least in part on a template size of a template of the video block or the reference block and a picture size of the video frame.

6. The method of claim 5, wherein: the template size comprises a template width in the second direction and a template height in the first direction; the picture size comprises a picture width in the second direction and a picture height in the first direction; a left boundary of the fourth region is determined based at least in part on the template width, and a bottom boundary of the fourth region is determined based at least in part on the picture height; and a right boundary of the fifth region is determined based at least in part on the picture width, and an upper boundary of the fifth region is determined based at least in part on the template height.

7. The method of claim 4, wherein: ​ ​ The video frame includes a plurality of coding tree units (CTUs), and the video block is located within one of the CTUs; and The search region further includes at least one of: a sixth region, wherein the sixth region is adjacent to the first region, the third region, and the fifth region in the first direction, wherein a distance between an upper boundary of the sixth region and the video block in the first direction is equal to a product of a height of the video block and a constant, and a distance between the sixth region and the one CTU in the first direction is equal to the height of the video block; a seventh region, wherein the seventh region is adjacent to the sixth region in the first direction and to the first region and the second region in the second direction, wherein a distance between the seventh region and the one CTU in the second direction is equal to a width of the video block; or an eighth region, wherein the eighth region is adjacent to the seventh region in the first direction and to the fourth region in the second direction, wherein a distance between the eighth region and the one CTU in the second direction is equal to the width of the video block.

8. The method of claim 1, wherein determining the reference block from the search region comprises: determining a candidate block from one of the plurality of regions; performing an availability check on the candidate block to determine whether samples in the candidate block are reconstructed; in response to there being samples in the candidate block that have not been reconstructed, determining the candidate block as an invalid candidate block; or in response to there not being samples in the candidate block that have not been reconstructed, determining the candidate block as a valid candidate block.

9. The method of claim 8, wherein the video frame includes a plurality of coding tree units (CTUs), and the video block is located within one of the CTUs, and the plurality of regions includes at least one of: a first region, wherein the first region is a first distance from the video block in the first direction and a second distance from the video block in the second direction; a second region, wherein the second region is adjacent to the first region in the first direction and the second distance from the video block in the second direction; a third region, wherein the third region is adjacent to the first region in the second direction and the first distance from the video block in the first direction; a fourth region, wherein the fourth region is adjacent to the second region in the first direction and the second distance from the video block in the second direction; a fifth region, wherein the fifth region is adjacent to the third region in the second direction and the first distance from the video block in the first direction; a sixth region, wherein the sixth region is adjacent to the first region, the third region, and the fifth region in the first direction, wherein a distance between an upper boundary of the sixth region and the video block in the first direction is equal to a product of a height of the video block and a constant, and a distance between the sixth region and the one CTU in the first direction is equal to the height of the video block; a seventh region, wherein the seventh region is adjacent to the sixth region in the first direction and adjacent to the first region and the second region in the second direction, wherein a distance between the seventh region and the one CTU in the second direction is equal to a width of the video block; or an eighth region, wherein the eighth region is adjacent to the seventh region in the first direction and adjacent to the fourth region in the second direction, wherein a distance between the eighth region and the one CTU in the second direction is equal to the width of the video block.

10. The method of claim 9, wherein: an availability check is performed on candidate blocks in the fourth region and the fifth region; and an availability check is not performed on candidate blocks in the first region, the second region, the third region, the sixth region, the seventh region, and the eighth region, and candidate blocks in the first region, the second region, the third region, the sixth region, the seventh region, and the eighth region are determined to be valid candidate blocks.

11. The method of claim 10, wherein determining the reference block from the search region comprises: determining the reference block to be one of the valid candidate blocks having a template matching a template of the video block.

12. The method of claim 8, wherein performing the availability check on the candidate block to determine whether samples in the candidate block are reconstructed comprises: determining whether a sample at a bottom-right corner of the candidate block is reconstructed; wherein in response to the sample at the bottom-right corner being reconstructed, the candidate block is determined to be a valid candidate block, or wherein in response to the sample at the bottom-right corner not being reconstructed, the candidate block is determined to be an invalid candidate block.

13. The method of claim 1, wherein determining the reference block from the search region comprises: determining a scan order of the plurality of regions; and searching for the reference block in the plurality of regions based on the scan order.

14. A method for video coding, comprising: determining, by an encoder, a search region for a video block from a video frame, wherein the search region has a first search region size in a first direction and a second search region size in a second direction perpendicular to the first direction, wherein the search region comprises a plurality of regions each within a region determined by the first search region size, the second search region size, and a location of the video block in the video frame; determining, by the encoder, a reference block from the search region, wherein a template of the reference block matches a template of the video block; determining, by the encoder, predicted samples for the video block based on the reference block; and ​ ​ generating, by the encoder, a bitstream based on the prediction samples.

15. The method of claim 14, wherein a center of the region is located at a position of the video block, a height of the region is twice the first search region size in the first direction, and a width of the region is twice the second search region size in the second direction.

16. The method of claim 14, wherein the first search region size and the second search region size are proportional to a height and a width of the video block, respectively.

17. The method of claim 14, wherein the plurality of regions comprises at least one of: a first region, wherein the first region is a first distance from the video block in the first direction and a second distance from the video block in the second direction; a second region, wherein the second region is adjacent to the first region in the first direction and the second distance from the video block in the second direction; a third region, wherein the third region is adjacent to the first region in the second direction and the first distance from the video block in the first direction; a fourth region, wherein the fourth region is adjacent to the second region in the first direction and the second distance from the video block in the second direction; or a fifth region, wherein the fifth region is adjacent to the third region in the second direction and the first distance from the video block in the first direction.

18. The method of claim 17, wherein the fourth region and the fifth region are determined based at least in part on a template size of a template of the video block or the reference block and a picture size of the video frame.

19. The method of claim 18, wherein: the template size comprises a template width in the second direction and a template height in the first direction; the picture size comprises a picture width in the second direction and a picture height in the first direction; a left boundary of the fourth region is determined based at least in part on the template width, and a bottom boundary of the fourth region is determined based at least in part on the picture height; and a right boundary of the fifth region is determined based at least in part on the picture width, and an upper boundary of the fifth region is determined based at least in part on the template height.

20. The method of claim 17, wherein: the video frame is divided into a plurality of coding tree units (CTUs), and the video block is located within one of the CTUs; and the search region further comprises at least one of: a sixth region, wherein the sixth region is adjacent to the first region, the third region, and the fifth region in the first direction, wherein a distance between an upper boundary of the sixth region and the video block in the first direction is equal to a product of a height of the video block and a constant, and a distance between the sixth region and the one CTU in the first direction is equal to the height of the video block. ​ a seventh region, wherein the seventh region is adjacent to the sixth region in the first direction and to the first region and the second region in the second direction, wherein a distance between the seventh region and the one CTU in the second direction is equal to a width of the video block; or an eighth region, wherein the eighth region is adjacent to the seventh region in the first direction and to the fourth region in the second direction, wherein a distance between the eighth region and the one CTU in the second direction is equal to a width of the video block.

21. The method of claim 14, wherein determining the reference block from the search region comprises: determining a candidate block from one of the plurality of regions; performing an availability check on the candidate block to determine whether samples in the candidate block are reconstructed; in response to there being samples in the candidate block that have not been reconstructed, determining the candidate block as an invalid candidate block; or in response to there not being samples in the candidate block that have not been reconstructed, determining the candidate block as a valid candidate block.

22. The method of claim 21, wherein the video frame comprises a plurality of coding tree units (CTUs) and the video block is within one of the CTUs, and the plurality of regions comprises at least one of: a first region, wherein the first region is a first distance from the video block in the first direction and a second distance from the video block in the second direction; a second region, wherein the second region is adjacent to the first region in the first direction and the second distance from the video block in the second direction; a third region, wherein the third region is adjacent to the first region in the second direction and the first distance from the video block in the first direction; a fourth region, wherein the fourth region is adjacent to the second region in the first direction and the second distance from the video block in the second direction; a fifth region, wherein the fifth region is adjacent to the third region in the second direction and the first distance from the video block in the first direction; a sixth region, wherein the sixth region is adjacent to the first region, the third region, and the fifth region in the first direction, wherein a distance between a top boundary of the sixth region and the video block in the first direction is equal to a height of the video block multiplied by a constant, and a distance between the sixth region and the one CTU in the first direction is equal to the height of the video block; a seventh region, wherein the seventh region is adjacent to the sixth region in the first direction and to the first region and the second region in the second direction, wherein a distance between the seventh region and the one CTU in the second direction is equal to a width of the video block; or an eighth region, wherein the eighth region is adjacent to the seventh region in the first direction and to the fourth region in the second direction, wherein a distance between the eighth region and the one CTU in the second direction is equal to a width of the video block. an eighth region, wherein the eighth region is adjacent to the seventh region in the first direction and adjacent to the fourth region in the second direction, wherein a distance between the eighth region and the one CTU in the second direction is equal to a width of the video block.

23. The method of claim 22, wherein: performing validity check on the candidate blocks in the fourth region and the fifth region; and not performing validity check on the candidate blocks in the first region, the second region, the third region, the sixth region, the seventh region, and the eighth region, and determining the candidate blocks in the first region, the second region, the third region, the sixth region, the seventh region, and the eighth region as valid candidate blocks.

24. The method of claim 23, wherein determining the reference block from the search regions comprises: determining the reference block as one of the valid candidate blocks having a template matching a template of the video block.

25. The method of claim 21, wherein performing availability check on the candidate block to determine whether samples in the candidate block are reconstructed comprises: determining whether a sample at a bottom-right corner of the candidate block is reconstructed, wherein in response to the sample at the bottom-right corner being reconstructed, the candidate block is determined as a valid candidate block, or wherein in response to the sample at the bottom-right corner not being reconstructed, the candidate block is determined as an invalid candidate block.

26. The method of claim 14, wherein determining the reference block from the search regions comprises: determining a scan order of the plurality of regions; and searching the reference block in the plurality of regions based on the scan order.

27. An apparatus for video decoding, comprising: a memory configured to store a bitstream; and a processor coupled to the memory and configured to perform the method for video decoding of any of claims 1-13 to decode the bitstream.

28. An apparatus for video encoding, comprising: a memory configured to store a bitstream; and a processor coupled to the memory and configured to perform the method for video encoding of any of claims 14-26 to generate the bitstream.

29. A non-transitory computer-readable storage medium having a bitstream stored therein, wherein the bitstream is decoded by the method for video decoding of any of claims 1-13.

30. A non-transitory computer-readable storage medium having a bitstream stored therein, wherein the bitstream is generated by the method for video encoding of any of claims 14-26.