Method, apparatus, and computer program for video coding

Intra Template Matching (IntraTMP) mode in video coding optimizes intra block copy processes by refining candidate lists through template matching, addressing redundancy and improving compression efficiency in video coding.

JP2025523726APending Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024515595
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-10
Filing Date
2022-11-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently reducing redundancy and achieving high compression ratios through intra and inter-picture prediction, particularly in handling complex motion vectors and intra prediction directions, which can lead to increased bit usage for less likely directions.

Method used

The implementation of Intra Template Matching (IntraTMP) mode in video encoding/decoding, where a reference template is matched to a current template to determine a block vector displacement, allowing for improved intra block copy (IBC) and intraTMP-based block vector storage, refining candidate lists through template matching and scaling, to optimize prediction and reduce bit usage.

Benefits of technology

Enhances video coding efficiency by reducing bit usage for less likely prediction directions and improving compression, thereby minimizing redundancy and increasing the achievable compression ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025523726000001_ABST
    Figure 2025523726000001_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure provide a method and apparatus for video decoding. The apparatus includes a processing circuit that receives an encoded bitstream having a first block within a current picture. The processing circuit obtains prediction information indicating whether the first block is coded in an Intra Template Matching (IntraTMP) mode. When the IntraTMP mode is applied to the first block, the first block is reconstructed based on a predicted block within a reconstructed search area in the current picture. The reference template of the predicted block is matched to the current template of the first block in the IntraTMP mode. The IntraTMP-based block vector BV IntraTMP of the first block is stored. The IntraTMP-based block vector indicates a displacement in position between the current template of the first block and the reference template of the predicted block. A second block is reconstructed based on the stored IntraTMP-based block vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 388,913, filed on July 13, 2022, entitled "IBC Candidate List Construction By Using the Motion Data of Intra Template-Matching Prediction", and to U.S. Patent Application No. 17 / 984,864, filed on November 10, 2022, entitled "INTRA BLOCK COPY (IBC) CANDIDATE LIST CONSTRUCTION WITH MOTION INFORMATION OF INTRA TEMPLATE-MATCHING PREDICTION". The disclosures of these prior applications are hereby incorporated by reference in their entireties.

[0002] This disclosure describes embodiments generally related to video coding.

Background Art

[0003] The background description provided here is for the purpose of generally presenting the context of the disclosure. The work of the inventors named herein, and the descriptions, to the extent that they are not otherwise qualified as prior art at the time of filing, in the context described in this background section, are not admitted as prior art to this disclosure, either expressly or impliedly.

[0004] Uncompressed digital images and / or videos contain a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (informally also known as the frame rate), for example, it can have a picture rate of 60 pictures per second, i.e., 60Hz. Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (luminance sample resolution of 1920×1080 at a frame rate of 60Hz) requires a bandwidth close to 1.5Gbit / s. One hour of such video requires storage space exceeding 600GByte.

[0005] One purpose of encoding and decoding images and / or videos can be the reduction of redundancy in the input image and / or video signal through compression. Compression can help reduce the aforementioned bandwidth requirements and / or storage space requirements, in some cases by two or more orders of magnitude. The description here uses video encoding / decoding as an example for illustration, but without departing from the spirit of the present disclosure, the same techniques can be applied to image encoding / decoding in a similar manner. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to techniques that can reconstruct an exact replica of the original signal from the compressed original signal. When using irreversible compression, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough to be useful for the intended application of the reconstructed signal. In the case of videos, irreversible compression is widely used. The amount of allowable distortion depends on the application. For example, users of a specific consumer streaming application may tolerate higher distortion than users of a television distribution application. The achievable compression ratio reflects this, and higher allowable / tolerable distortion can result in a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially subdivided into samples of multiple blocks. If the samples of all blocks are coded in the intra mode, the picture can be an intra picture. Intra pictures and their derivatives, such as, for example, independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still picture. The samples of an intra block can be subjected to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique that minimizes sample values in the pre-transform domain. Optionally, the smaller the post-transform DC value and the smaller the AC coefficients, the fewer bits are required to represent the block with a given quantization step size after entropy coding.

[0008] For example, traditional intra coding used in MPEG-2 generation coding technology does not use intra prediction. However, some newer video compression techniques attempt to perform prediction, for example, based on surrounding sample data and / or metadata obtained during encoding and / or decoding of a block of data. Such techniques are hereinafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and not that from reference pictures.

[0009] There can be many different forms of intra prediction. In a given video coding technology, if two or more of such techniques can be used, the specific technique used can be coded as a specific intra prediction mode that uses that specific technique. In a certain case, the intra prediction mode can have sub - modes and / or parameters, and those sub - modes and / or parameters can be individually coded or included in a mode - code word that defines the prediction mode being used. Which code word to use for a given combination of mode, sub - mode, and / or parameter can affect the coding efficiency gain through intra prediction, and the same can be true for the entropy coding technology used to convert the code word into a bitstream.

[0010] A specific mode of intra prediction was introduced in H.264, improved in H.265, and further improved in more recent coding technologies such as, for example, the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). A predictor block can be formed using the adjacent sample values of already available samples. The sample values of the adjacent samples are copied into the predictor block according to a direction. The reference to the direction to use can be coded in the bitstream or can itself be predicted.

[0011] Referring to FIG. 1A, in the lower right, a subset of 9 predictor directions known from 33 possible predictor directions defined in H.265 (corresponding to 33 of the 35 intra - modes, the angular modes) is depicted. The point (101) where the arrows converge represents the sample being predicted. The arrows indicate that the sample is predicted from that direction. For example, arrow (102) indicates that sample (101) is predicted from one or more samples in the upper - right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples in the lower - left at an angle of 22.5 degrees from the horizontal.

[0012] Referring still to FIG. 1A, a square block (104) of 4×4 samples (shown by the thick dashed line) is drawn at the upper left. The square block (104) contains 16 samples, and each sample is labeled with an “S”, its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample in both the Y and X dimensions within block (104). Since this block is of size 4×4 samples, S44 is at the lower right. Further, reference samples according to a similar numbering scheme are shown. The reference samples are labeled with an “R” and the Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, the predicted samples are adjacent to the block being reconstructed, and thus there is no need to use negative values.

[0013] Intra-picture prediction functions by copying the reference sample values from adjacent samples indicated by the signaled prediction direction. For example, assume that the coded video bitstream contains signaling indicating that for this block, the prediction direction coincides with arrow (102), i.e., samples are predicted from samples at 45 degrees up and to the right from horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. And sample S44 is predicted from reference sample R08.

[0014] In certain cases, particularly when the directions cannot be evenly divided by 45 degrees, the values of multiple reference samples may be combined, for example, by interpolation, to calculate the reference samples.

[0015] As video coding technology develops, the number of possible directions is increasing. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments are being conducted to identify the most likely directions, and certain techniques in entropy coding are used to accept penalties for less likely directions and represent those likely directions with fewer bits. Furthermore, those directions themselves may be predicted from adjacent directions used in adjacent already-decoded blocks.

[0016] FIG. 1B shows a schematic diagram (110) depicting 65 intra prediction directions according to JEM to illustrate the increasingly numerous prediction directions.

[0017] The mapping of intra prediction direction bits representing directions within the encoded video bitstream can vary for each video coding technology. Such mapping can range from a simple direct mapping to complex adaptive schemes including codewords and the most probable mode, and similar techniques. However, in most cases, there may be certain directions in video content that are statistically less likely to occur than other specific directions. Since the goal of video compression is to reduce redundancy, in well-functioning video coding technology, those less likely directions will be represented with more bits than the likely directions.

[0018] Image and / or video encoding and decoding can be performed using inter-picture prediction with motion compensation. Motion compensation can be a non-reversible compression technique, and it can be related to a technique in which a block of sample data from a previously reconstructed picture or a part thereof (reference picture) is spatially shifted in a direction indicated by a motion vector (hereinafter, MV) and then used for prediction of a newly reconstructed picture or picture part. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or can have three dimensions with an indication indicating the reference picture used as the third (the latter can be indirectly considered as the temporal dimension).

[0019] In some video compression techniques, an MV applicable to a specific region of sample data can be predicted from another MV, for example, an MV preceding that MV in decoding order and related to another region of sample data spatially adjacent to the region being reconstructed. Doing so can significantly reduce the amount of data required to code that MV, thereby removing redundancy and enhancing compression. MV prediction can function effectively. This is because, for example, when coding an input video signal (known as natural video) derived from a camera, regions larger than the region to which a single MV is applicable move in a similar direction, and thus, in some cases, there is a statistical likelihood that it can be predicted using a similar motion vector derived from the MVs of adjacent regions. What this brings about is that the MV found for a given region is similar or the same as the MV predicted from surrounding MVs, and thus, after entropy coding, it can be represented with fewer bits than would be used if that MV were directly coded. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be non-lossless, for example, due to rounding errors when calculating predictors from several surrounding MVs.

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Recommendation H.265, “High Efficiency Video Coding”, December 2016). Among those numerous MV prediction mechanisms provided by H.265, a technique called “spatial merge” below will be described with reference to FIG. 2.

[0021] Referring to FIG. 2, the current block (201) has samples found by the encoder during the motion search process that it is predictable from a previous block of the same size that is spatially shifted. Instead of directly coding the MV, the MV can be derived from metadata associated with one or more reference pictures, such as from the immediately previous reference picture (in decoding order), using an MV associated with any one of five surrounding samples denoted as A0, A1, and B0, B1, B2 (202 to 206 respectively). In H.265, MV prediction can use a predictor from the same reference picture that the neighboring blocks are using. SUMMARY OF THE INVENTION

[0022] Aspects of the present disclosure provide methods and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit. The processing circuit receives an encoded video bitstream having a first block within a current picture. The processing circuit obtains prediction information indicating whether the first block is coded in an Intra Template Matching Prediction (IntraTMP) mode. In response to the IntraTMP mode being applied to the first block, the processing circuit reconstructs the first block based on a predicted block within a reconstructed search area within the current picture. In the IntraTMP mode, a reference template of the predicted block is matched to the current template of the first block. The processing circuit stores an IntraTMP-based block vector BV IntraTMP for the first block. The IntraTMP-based block vector BV IntraTMP indicates a displacement in position (also referred to as a motion vector displacement) between the current template of the first block and the reference template of the predicted block. The processing circuit reconstructs a second block based on the stored IntraTMP-based block vector BV IntraTMP . The second block can be coded in either an Intra Block Copy (IntraBC) mode (also referred to as the IBC mode) or the IntraTMP mode. In one example, the second block is within the current picture.

[0023] In one example, the first block includes one or more M×N units. The processing circuit stores an IntraTMP-based block vector BV in each M×N unit of the first block. IntraTMP In one example, the processing circuit stores the IntraTMP-based block vector BV IntraTMP with a predetermined tolerance.

[0024] In one example, the processing circuit stores the IntraTMP-based block vector BV with an accuracy indicated by syntax information in the encoded video bitstream. IntraTMP

[0025] The processing circuit can determine a reference template based on a plurality of template candidates in a reconstructed search area within the current picture. The displacement between one of the plurality of template candidates and the current template can be indicated by a vector that is (i) the block vector (BV) of a third block coded in the intra block copy (IBC) mode, or (ii) the IntraTMP-based block vector BV of a third block coded in the IntraTMP mode. IntraTMP

[0026] In one example, the processing circuit stores the IntraTMP-based block vector BV based on the template matching cost between the reference template of the prediction block and the current template of the first block being less than a threshold. IntraTMP

[0027] In one example, the processing circuit stores the template matching cost between the reference template of the prediction block and the current template of the first block.

[0028] The template matching cost can be normalized based on the number of samples in the current template.

[0029] ​​​In one embodiment, the processing circuit decodes prediction information of a first block within a current picture from an encoded video bitstream. The prediction information indicates that an intra block copy (IBC) mode is applied to the first block. The processing circuit can construct an IBC candidate list for the first block. The IBC candidate list includes a first candidate based on an IntraTMP-based block vector BV of a second block coded in an Intra Template Matching (IntraTMP) mode. The second block can be one of (i) a reconstructed block within the current picture and (ii) a temporal neighbor of the first block. The processing circuit can reconstruct the first block based on the IBC candidate list. IntraTMP

[0030] In one example, the processing circuit performs template matching on an IntraTMP-based block vector BV IntraTMP to refine the IntraTMP-based block vector BV IntraTMP and determines the first candidate as the template-matched IntraTMP-based block vector BV IntraTMP .

[0031] In one example, the IBC candidate list for the first block includes a plurality of candidates. Each of the plurality of candidates is based on an IntraTMP-based block vector BV of a block coded in (i) the IntraTMP mode IntraTMP ​、and (ii) based on one of the block vectors (BVs) of the blocks coded in IBC mode. The plurality of candidates can include a first candidate. The processing circuit can perform template matching on the plurality of candidates by determining respective template matching costs for each of the plurality of candidates based on a reference template of a reference block and a current template of the first block. The processing circuit can reorder the plurality of candidates based on the determined template matching costs. The processing circuit can reconstruct the first block based on the reordered plurality of candidates in the IBC candidate list.

[0032] In one example, for each of the plurality of candidates that is an IntraTMP-based block vector BV IntraTMP the processing circuit applies a scaling factor to the template matching cost of each candidate.

[0033] In one example, the number of one or more candidates in the IBC candidate list that are IntraTMP-based block vectors BVs IntraTMP of each block coded in IntraTMP mode is below a threshold. The one or more candidates include a first candidate.

[0034] In one example, the number of one or more candidates in the IBC candidate list is equal to the threshold. A third block is one of (i) a reconstructed block in the current picture and (ii) a temporal neighbor of the first block. For a new IntraTMP-based block vector BV IntraTMP of the third block that is not in the IBC candidate list, if the template matching cost associated with the new IntraTMP-based block vector BV IntraTMP is less than the template matching cost associated with at least one of the one or more candidates, the processing circuit replaces a certain candidate among the one or more candidates with the new IntraTMP-based block vector BV IntraTMPIt can be replaced. The candidate to be replaced may be the one with the maximum template matching cost among the one or more candidates described above.

[0035] In one example, the processing circuit may add one or more candidates from one or more blocks coded in the IBC mode to an IBC candidate list.

[0036] In one example, the IntraTMP-based block vector BV of the second block IntraTMP is obtained from a BV history table storing one or more block vectors (BVs) or one or more block displacement vectors of at least one block previously coded within the current picture.

[0037] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing a program executable by at least one processor to perform a method for video decoding.

Brief Description of the Drawings

[0038] Further features, properties, and various advantages of the matters related to the disclosure will become even more apparent from the following detailed description and the accompanying drawings.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

DETAILED DESCRIPTION OF THE INVENTION

[0039] FIG. 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional data transmission. For example, the terminal device (310) may encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to another terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission can be common in media service providing applications and the like.

[0040] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, for example, during a video conference. In bidirectional data transmission, in one example, each of the terminal devices (330) and (340) can encode video data (e.g., a stream of video pictures captured by that terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to restore the video pictures, and display the video pictures on an accessible display device according to the restored video data.

[0041] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and a smartphone, respectively, but the principles of the present disclosure may not be so limited. Embodiments of the present disclosure find use in laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. The network (350) represents any number of networks that transmit encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, a wired (wired) communication network and / or a wireless communication network. The communication network (350) may exchange data over a circuit-switched channel and / or a packet-switched channel. Representative networks include long-distance communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of this description, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described below.

[0042] FIG. 4 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application related to the matter disclosed. The matter disclosed can be equally applied to other applications where video can be used, including, for example, videoconferencing, digital TV, streaming services, and storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0043] A streaming system may include a capture subsystem (413) that can include a video source (401), such as a digital camera, that produces a stream (402) of, for example, uncompressed video pictures. In one example, the stream (402) of video pictures includes samples taken by a digital camera. The stream (402) of video pictures is drawn as a thick line to emphasize that it has a high data volume compared to the encoded video data (404) (or encoded video bitstream) and can be processed by an electronics device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosure described in more detail hereinafter. The encoded video data (404) (or encoded video bitstream) is drawn as a thin line to emphasize that it has a low data volume compared to the stream (402) of video pictures and can be stored in a streaming server (405) for later use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronics device (430). The video decoder (410) can decode an incoming copy (407) of the encoded video data and produce an outgoing stream (411) of video pictures, which can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video coding / compression standard.Examples of such standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The subject matter disclosed may be used in the context of VVC.

[0044] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).

[0045] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0046] The receiver (531) can receive one or more encoded video sequences to be decoded by the video decoder (510). In one embodiment, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences can be received from a channel (501) that can be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data together with other data, such as, for example, encoded audio data and / or auxiliary data streams, and those data can be transferred to their respective using entities (not shown). The receiver (531) can separate the encoded video sequences from the other data. To counter network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser 520 (hereinafter, “parser (520)”). In certain applications, the buffer memory (515) is part of the video decoder (510). In others, it may be external to the video decoder (510) (not shown). In still others, for example, to counter network jitter, a buffer memory (not shown) can be present external to the video decoder (510), and further, for example, to handle playback timing, another buffer memory (515) can be present inside the video decoder (510). When the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability or from a synchronous network, the buffer memory (515) may not be needed or can be made small. For use on a best-effort packet network such as the Internet, the buffer memory (515) can be needed and made relatively large and, advantageously, of an adaptable size, and can also be implemented, at least in part, by an operating system or similar element (not shown) external to the video decoder (510).

[0047] Video decoder (510) may include a parser (520) for reconstructing symbols (521) from an encoded video sequence. The categories of those symbols include information used to manage the operation of video decoder (510) and, possibly, information for controlling a rendering device, such as a renderer device (512) (e.g., a display screen), which is not an integrated part of electronics device (530) but can be coupled to electronics device (530), as shown in FIG. 5. The control information for the (one or more) rendering devices may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). Parser (520) may perform syntax analysis / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. Parser (520) can extract a set of subgroup parameters regarding at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to a group. The subgroups can include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. Parser (520) can also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the encoded video sequence information.

[0048] Parser (520) may perform entropy decoding / syntax analysis processing on the video sequence received from buffer memory (515) to produce symbols (521).

[0049] For the reconstruction of symbol (521), multiple different units may be involved depending on the type of the encoded video picture or a part thereof and other factors (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.). Which units are involved and how they are involved can be controlled by subgroup control information parsed from the encoded video sequence by parser (520). Such a flow of subgroup control information between parser (520) and the following multiple units is not illustrated for clarity.

[0050] Beyond the aforementioned functional blocks, video decoder (510) can conceptually be subdivided into a number of functional units as described later. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the matters related to the disclosure, a conceptual subdivision into the following functional units is appropriate.

[0051] The first unit is a scaler / inverse transform unit (551). Scaler / inverse transform unit (551) receives quantized transform coefficients together with control information including which transform to use, block size, quantization coefficient, quantization scaling matrix, etc. as (one or more) symbols (521) from parser (520). Scaler / inverse transform unit (551) can output a block with sample values that can be input to aggregator (555).

[0052] In some cases, the output samples of the scaler / inverse transform unit (551) may be related to intra-coded blocks. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates a block of the same size and shape as the block being reconstructed, using the surrounding already-reconstructed information fetched from the current picture buffer (558). The current picture buffer (558) buffers, for example, the partially reconstructed current picture and / or the fully reconstructed current picture. The aggregator (555) adds, in some cases, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0053] In other cases, the output samples of the scaler / inverse transform unit (551) may relate to an inter-coded, potentially motion-compensated block. In such a case, the motion-compensated prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols (521) related to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case, called the residual samples or residual signal) to generate output sample information. The address in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches the prediction samples can be controlled by the motion vector and is available to the motion-compensated prediction unit (553) in the form of, for example, symbols (521) having X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (557) when exact sub-sample motion vectors are used, a motion vector prediction mechanism, and the like.

[0054] The output samples of the aggregator (555) can be subjected to various loop filtering techniques in the loop filter unit (556). Video compression techniques can include in-loop filter techniques, which are controlled by parameters made available to the loop filter unit (556) as symbols (521) from the parser (520) included in the coded video sequence (also referred to as the coded video bitstream). Video compression can also respond to meta information obtained during decoding of a preceding portion (in decoding order) of the coded picture or coded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.

[0055] The output of the loop filter unit (556) can be a sample stream that can be output to the renderer device (512), and this can also be stored in the reference picture memory (557) for use in future inter-picture prediction.

[0056] When a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is fully reconstructed and that coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a new current picture buffer can be reallocated before starting the reconstruction of the next coded picture.

[0057] The video decoder (510) can perform decoding processing according to a predetermined video compression technology or standard such as ITU-T Recommendation H.265. The coded video sequence can conform to the syntax defined by the video compression technology or standard used, in the sense of faithfully adhering to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile can select specific tools from all the tools available in the video compression technology or standard such that only those tools available for use under that profile are selected. Also, for compliance, it is necessary that the complexity of the coded video sequence be within the range defined by the level of the video compression technology or standard. In some cases, the level restricts, for example, the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured, for example, in megasamples per second), the maximum reference picture size, etc. The restrictions set by the level can, in some cases, be further restricted through the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled in the coded video sequence.

[0058] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the (one or more) encoded video sequences. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal, spatial, or signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, and the like.

[0059] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronics device (620). For example, the electronics device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.

[0060] The video encoder (603) may receive video samples from a video source (601) (not part of the electronics device (620) in the example of FIG. 6) that may capture the (one or more) video images to be encoded by the encoder (603). In another example, the video source (601) is part of the electronics device (620).

[0061] The video source (601) can provide a source video sequence to be encoded by a video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCB, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service providing system, the video source (601) can be a storage device storing pre-prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. Video data can be provided as a plurality of individual pictures that convey motion when viewed in sequence. The pictures themselves can be organized as a spatial array of pixels, and each pixel can have one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can immediately understand the relationship between pixels and samples. The following description focuses on samples.

[0062] According to one embodiment, the video encoder (603) can code and compress pictures of the source video sequence into an encoded video sequence (643) in real time or under other time constraints required. Enforcing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described later. That coupling is not shown for clarity. Parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions related to the video encoder (603) optimized for a particular system design.

[0063] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (630) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and (one or more) reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in the same way as a (remote) decoder also does. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in a bit-exact result that is independent of the decoder position (local or remote), the content in the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is also used in some related arts.

[0064] The operation of the "local" decoder (633) can be considered the same as that of a "remote" decoder such as the video decoder (510), which has already been described in detail above in relation to FIG. 5. However, referring briefly to FIG. 5 as well, since symbols are available and the encoding / decoding of the symbols into an encoded video sequence by the entropy encoder (645) and the parser (520) can be reversible, the entropy decoding part of the video decoder (510) including the buffer memory (515) and the parser (520) does not have to be fully implemented in the local decoder (633).

[0065] In one embodiment, decoder techniques other than syntax analysis / entropy decoding that exist within the decoder exist in the corresponding encoder in the same or substantially the same functional form. Accordingly, the matters disclosed herein focus on decoder operations. The description of encoder techniques can be omitted because it is the reverse of the decoder techniques described in detail. In certain fields, more detailed descriptions are provided below.

[0066] During operation, in some examples, the source coder (630) may perform motion-compensated predictive coding in which an input picture is predictively coded against one or more previously coded pictures designated as "reference pictures" from a video sequence. Thus, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of the (one or more) reference pictures that may be selected as the (one or more) prediction reference for the input picture.

[0067] The local video decoder (633) may decode the coded video data of a picture that may be designated as a reference picture based on symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be an irreversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some error. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder on the reference picture and cause the reconstructed reference picture to be stored in the reference picture memory (634). Thus, the video encoder (603) may locally store a copy of the reconstructed reference picture having the same content as the reconstructed reference picture that would be obtained by a far-end video decoder.

[0068] Predictor (635) may perform predictive search for the coding engine (632). That is, regarding a new picture to be coded, predictor (636) may search reference picture memory (634) for sample data (as candidate reference pixel blocks) or certain metadata such as, for example, reference picture motion vectors and block shapes that can serve as appropriate prediction criteria for the new picture. Predictor (635) may operate on a per pixel block basis to find appropriate prediction criteria. Optionally, the input picture may have prediction criteria drawn from a plurality of reference pictures stored in reference picture memory (634) as determined by the search results obtained by predictor (635).

[0069] Controller (650) may manage the coding process of source coder (630), including, for example, setting parameters and subgroup parameters used to encode video data.

[0070] The outputs of all the foregoing functional units may be subjected to entropy coding in entropy coder (645). Entropy coder (645) converts the symbols generated by the various functional units into an encoded video sequence by applying reversible compression to the symbols according to techniques such as, for example, Huffman coding, variable length coding, arithmetic symbol coding, etc.

[0071] Transmitter (640) may buffer the (one or more) encoded video sequences generated by entropy coder (645) and prepare them for transmission via communication channel (660). Communication channel (660) may be a hardware / software link to a storage device that stores the encoded video data. Transmitter (640) may merge the encoded video data from video encoder (603) with other data to be transmitted, such as, for example, encoded audio data and / or auxiliary data streams (sources not shown).

[0072] The controller (650) may manage the operation of the video encoder (603). In coding, the controller (650) may assign to each coded picture a specific coded picture type that may affect the coding technique applicable to that picture. For example, a picture may often be assigned one of the following picture types.

[0073] An intra picture (I picture) may be coded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow multiple different types of intra pictures, for example including an Independent Decoder Refresh (IDR) picture. Those skilled in the art know those variants of the I picture, as well as their respective uses and characteristics.

[0074] A predicted picture (P picture) may be coded and decoded using intra prediction or inter prediction, using at most one motion vector and a reference index to predict the sample values of each block.

[0075] A bi-predicted picture (B picture) may be coded and decoded using intra prediction or inter prediction, using at most two motion vectors and a reference index to predict the sample values of each block. Similarly, a multi-predicted picture may use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0076] The source picture is generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be coded predictively by referring to other (already coded) blocks determined by the coding assignment applied to each of those blocks in their respective pictures. For example, blocks of an I picture can be coded non-predictively, or they can be coded predictively by referring to already coded blocks in the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded non-predictively or via spatial or temporal prediction by referring to a reference picture coded one ahead. Blocks of a B picture can be coded non-predictively or via spatial or temporal prediction by referring to one or two reference pictures coded ahead.

[0077] The video encoder (603) can perform coding processing according to a predetermined video coding technology or standard such as ITU-T Recommendation H.265. In its operation, the video encoder (603) can perform various compression processes including predictive coding processes that utilize the temporal and spatial redundancies in the input video sequence. The coded video data can thus conform to the syntax defined by the video coding technology or standard being used.

[0078] In one embodiment, the transmitter (640) can transmit additional data along with the encoded video. The source coder (630) can include such data as part of the encoded video sequence. The additional data can have temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.

[0079] The image can be captured as a plurality of source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation within a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded, referred to as the current picture, is divided into a plurality of blocks. When a block within the current picture is similar to a reference block within a reference picture that has been previously coded in the video and is still being buffered, that block within the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that specifies the reference picture when multiple reference pictures are being used.

[0080] In some embodiments, dual prediction techniques can be used in inter-picture prediction. According to the dual prediction technique, for example, two reference pictures are used, such as a first reference picture and a second reference picture that are both earlier than the current picture in the decoding order (however, in the display order, they can be in the past and future respectively). A block within the current picture can be coded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0081] Furthermore, merge mode techniques can be used to improve coding efficiency in inter-picture prediction.

[0082] According to some embodiments of the present disclosure, predictions, such as inter-picture prediction and intra-picture prediction, are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into a plurality of coding tree units (CTUs) for compression, and those CTUs in the picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of that CU, such as an inter prediction type or an intra prediction type. A CU is divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation during coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.

[0083] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) is configured to receive sample values of a processing block (e.g., a prediction block) in a current video picture in a sequence of video pictures and encode the processing block into an encoded picture that is part of an encoded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.

[0084] In an example of HEVC, a video encoder (703) receives a matrix of sample values for a processing block, such as an 8×8 sample of a prediction block. The video encoder (703) determines whether the processing block is best coded using an intra mode, an inter mode, or a bi-prediction mode, for example, using rate-distortion optimization. If the processing block is coded in the intra mode, the video encoder (703) can encode the processing block into the coded picture using intra prediction techniques, and if the processing block is coded in the inter mode or the bi-prediction mode, the video encoder (703) can encode the processing block into the coded picture using inter prediction techniques or bi-prediction techniques, respectively. In certain video coding techniques, the merge mode can be an inter-picture prediction sub-mode in which a motion vector is derived from one or more motion vector predictors without the benefit of the outer coded motion vector components of the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining, for example, the mode of the processing block.

[0085] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) that are coupled together as shown in FIG. 7.

[0086] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks (e.g., blocks in a previous picture and in a subsequent picture) in a reference picture, generate inter-prediction information (e.g., a description of redundant information according to an inter-coding technique, a motion vector, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using some suitable technique. In some examples, the reference picture is a reference picture decoded based on coded video information.

[0087] The intra-encoder (722) is configured to receive samples of a current block (e.g., a processing block) and, in some cases, compare the block with already-coded blocks in the same picture, generate quantized coefficients after transformation, and also generate, in some cases, intra-prediction information (e.g., intra-prediction direction information according to one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same picture.

[0088] The overall controller (721) is configured to determine overall control data and control other components of the video encoder (703) based on the overall control data. In one example, the overall controller (721) determines the mode of a block and provides a control signal to the switch (726) based on that mode. For example, when the mode is the intra mode, the overall controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), and controls the entropy encoder (725) to select the intra prediction information and include the intra prediction information in the bitstream. When the mode is the inter mode, the overall controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), and controls the entropy encoder (725) to select the inter prediction information and include the inter prediction information in the bitstream.

[0089] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) operates based on the residual data and is configured to encode the residual data to generate transform coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. Then, the transform coefficients are subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be suitably used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is suitably processed to generate a decoded picture, and the decoded picture can be buffered in a memory circuit (not shown) and, in some examples, used as a reference picture.

[0090] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to a suitable standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. Note that according to the matters related to the present disclosure, when coding a block in any merge sub-mode of the inter mode or the bi-prediction mode, there is no residual information.

[0091] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive an encoded picture that is part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.

[0092] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872), which are coupled together as shown in FIG. 8.

[0093] The entropy decoder (871) may be configured to reconstruct from the encoded picture specific symbols representing syntax elements that make up the encoded picture. Such symbols can include, for example, the mode in which a block is coded (e.g., intra mode, inter mode, bi-prediction mode, the latter two in merge sub-modes or other sub-modes), and prediction information (e.g., intra prediction information or inter prediction information, etc.) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880), respectively. The symbols can also include, for example, residual information in the form of quantized transform coefficients, and the like. In one example, when the prediction mode is inter mode or bi-prediction mode, inter prediction information is provided to the inter decoder (880), and when the prediction type is intra prediction type, intra prediction information is provided to the intra decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).

[0094] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0095] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0096] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require specific control information (including quantization parameter (QP)), and such information may be provided by the entropy decoder (871) (since this can be only low-volume control information, the data path is not shown).

[0097] The reconstruction module (874) is configured to combine the residual information output by the residual decoder (873) and the prediction result (output by the inter or intra prediction module as appropriate) in the spatial domain to form a reconstruction block. The reconstruction block can be part of a reconstructed picture, or alternatively, the reconstructed picture can be part of a reconstructed video. Note that, in order to improve visual quality, other suitable processes can be performed, such as deblocking processing and the like.

[0098] Note that the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), and (703), and the video decoders (410), (510), and (810) may be implemented using one or more processors that execute software instructions.

[0099] In VVC, various inter-prediction modes can be used. For a CU to be inter-predicted, the motion parameters can include one or more MVs, one or more reference picture indices, a reference picture list use index, and additional information about specific coding features used for sample generation by inter-prediction. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU can be associated with a PU and may not have significant residual coefficients, a coded motion vector delta or MV difference (e.g., MVD), or a reference picture index. The merge mode can be specified, and the motion parameters for the current CU can be obtained from one or more neighboring CUs including spatial and / or temporal candidates, and optionally, additional information such as introduced in VVC. The merge mode can be applied not only to skip mode but also to CUs to be inter-predicted. In one example, an alternative to the merge mode is the explicit transmission of motion parameters, where one or more MVs, the corresponding reference picture indices for each reference picture list, a reference picture list use flag, and other information are signaled explicitly for each CU.

[0100] In one embodiment, such as in VVC, the VVC Test Model (VTM) reference software includes one or more refined inter prediction coding tools, including extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode using symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8×8 motion field compression), CU-level weighted bi-prediction (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), and the like. Inter prediction and related methods are described in detail below.

[0101] Extended merge prediction may be used in some examples. For example, in one example such as in VTM4, the merge candidate list is constructed by including in order the following five types of candidates, namely, (one or more) spatial motion vector predictors (MVPs) from (one or more) spatial neighboring CUs, (one or more) temporal MVPs from (one or more) collocated CUs, (one or more) history-based MVPs from a first-in first-out (FIFO) table, (one or more) pairwise average MVPs, and (one or more) zero MVs.

[0102] The size of the merge candidate list can be signaled in the slice header. In one example, the maximum allowable size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., the merge index) can be coded using truncated unary binarization (TU). The first bin of the merge index can be coded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.

[0103] Examples of the generation process for each category of merge candidates are provided below. In one embodiment, (one or more) spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC may be the same as that in HEVC. In one example, up to four merge candidates are selected from the candidates at the positions shown in FIG. 9. FIG. 9 shows the positions of spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 9, the order of derivation is B1, A1, B0, A0, and B2. The position B2 is considered only when any of the CUs at positions A0, B0, B1, and A1 are not available (for example, because the CU belongs to another slice or another tile) or is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures that candidates with the same motion information are excluded from the candidate list so that coding efficiency is improved.

[0104] To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only the pairs connected by the arrows in FIG. 10 are considered, and a candidate is added to the candidate list only if the corresponding candidates used in the redundancy check do not have the same motion information. FIG. 10 shows the candidate pairs considered for the redundancy check of spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 10, the pairs connected by each arrow are A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, the candidates at positions B1, A0, and / or B2 can be compared with the candidate at position A1, and the candidates at positions B0 and / or B2 can be compared with the candidate at position B1.

[0105] In one embodiment, the (one or more) time candidates are derived as follows. In one example, only one time merge candidate is added to the candidate list. FIG. 11 shows exemplary motion vector scaling for a time merge candidate. To derive the time merge candidate for the current CU (1111) in the current picture (1101), a scaled MV (1121) (e.g., shown by the dotted line in FIG. 11) can be derived based on the collocated CU (1112) belonging to the collocated reference picture (1104). The reference picture list used to derive the collocated CU (1112) can be explicitly signaled in the slice header. A scaled MV (1121) regarding the time merge candidate can be obtained as shown by the dotted line in FIG. 11. The scaled MV (1121) can be scaled from the MV of the collocated CU (1112) using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (1102) of the current picture (1101) and the current picture (1101). The POC distance td can be defined as the POC difference between the collocated reference picture (1104) of the collocated picture (1103) and the collocated picture (1103). The reference picture index of the time merge candidate can be set to zero.

[0106] FIG. 12 shows exemplary candidate positions (e.g., C0 and C1) for the time merge candidate of the current CU. The position of the time merge candidate can be selected between the candidate positions C0 and C1. The candidate position C0 is located at the lower right corner of the collocated CU (1210) of the current CU. The candidate position C1 is located at the center of the collocated CU (1210) of the current CU. If the CU at the candidate position C0 is not available, intra-coded, or outside the current row of the CTU, the candidate position C1 is used to derive the time merge candidate. Otherwise, for example, if the CU at the candidate position C0 is available, intra-coded, and within the current row of the CTU, the candidate position C0 is used to derive the time merge candidate.

[0107] In video / image coding, template matching (TM) technology can be used. To further improve the compression efficiency of the VVC standard, for example, TM can be used to refine the motion vector (MV). In one example, TM is used on the decoder side. In TM mode, a template (e.g., the current template) of a block (e.g., the current block) within the current picture is constructed, and the MV can be refined by determining the closest match between the template of the block within the current picture and a plurality of possible templates (e.g., a plurality of possible reference templates) within the reference picture. In one embodiment, the template of a block within the current picture can include the adjacent reconstructed samples to the left of the block and the adjacent reconstructed samples above the block. TM can be used in video / image coding after VVC.

[0108] FIG. 13 shows an example of template matching (1300). By using TM to determine the closest match between a template (1321) of the current coding unit (CU) (1301) within the current picture (1310) and a template (e.g., one of the plurality of possible templates is template (1325)) among the plurality of possible templates within the reference picture (1311), the motion information of the current CU (1301) can be derived (e.g., deriving the final motion information from the initial motion information such as the initial MV 1302). The template (1321) of the current CU (1301) can have any suitable shape and any suitable size.

[0109] In one embodiment, the template (1321) of the current CU (1301) includes an upper template (1322) and a left template (1323). The upper template (1322) and the left template (1323) can each have any suitable shape and any suitable size.

[0110] The upper template (1322) can include samples within one or more upper adjacent blocks of the current CU (1301). In one example, the upper template (1322) includes samples of four rows within one or more upper adjacent blocks of the current CU (1301). The left template (1323) can include samples within one or more left adjacent blocks of the current CU (1301). In one example, the left template (1323) includes samples of four columns within one or more left adjacent blocks of the current CU (1301).

[0111] Each template (e.g., template (1325)) of the plurality of possible templates within the reference picture (1311) corresponds to the template (1321) within the current picture (1310). In one embodiment, the initial MV (1302) points to the reference block (1303) within the reference picture (1311) from the current CU (1301). Each template (e.g., template (1325)) of the plurality of possible templates within the reference picture (1311) and the template (1321) within the current picture (1310) can have the same shape and the same size. For example, the template (1325) of the reference block (1303) includes the upper template (1326) within the reference picture (1311) and the left template (1327) within the reference picture (1311). The upper template (1326) can include samples within one or more upper adjacent blocks of the reference block (1303). The left template (1327) can include samples within one or more left adjacent blocks of the reference block (1303).

[0112] For example, the TM cost can be determined based on pairs of templates, such as template (1321) and template (1325). The TM cost can indicate the matching between template (1321) and template (1325). An optimized MV (or final MV) can be determined based on the search around the initial MV (1302) of the current CU (1301) within the search range (1315). The search range (1315) can have any suitable shape and any suitable number of reference samples. In one example, the search range (1315) within the reference picture (1311) includes a [-L, L] pel range, where L is a positive integer such as 8 (e.g., 8 samples). For example, a difference (e.g., [0, 1]) is determined based on the search range (1315), and an intermediate MV is determined by the sum of the initial MV (1302) and the difference (e.g., [0, 1]). Based on the intermediate MV, an intermediate reference block and a corresponding template within the reference picture (1311) can be determined. The TM cost can be determined based on template (1321) and the intermediate template within the reference picture (1311). The TM cost can correspond to differences determined based on the search range (1315) (e.g., [0, 0] corresponding to the initial MV (1302), and [0, 1], etc.). In one example, the difference corresponding to the minimum TM cost is selected, and the optimized MV is the sum of the difference corresponding to the minimum TM cost and the initial MV (1302). As described above, TM can derive the final motion information (e.g., the optimized MV) from the initial motion information (e.g., the initial MV 1302).

[0113] FIG. 14 shows an example of an Intra Block Copy (IBC or IntraBC) according to an embodiment of the present disclosure. The current picture (1400) to be reconstructed includes a reconstructed area (1410) (gray area) and an area to be decoded (1420) (white area). The current block (1430) is being reconstructed by the decoder. The current block (1430) can be reconstructed from a reference block (1440) within the reconstructed area (1410). The position offset between the reference block (1440) and the current block (1430) can be referred to as a block vector (1450) (or BV (1450)). In the example of FIG. 14, the IBC reference area (1460) is within the reconstructed area (1410), the reference block (1440) is within the IBC reference area (1460), and the block vector (1450) points to the reference block (1440) within the IBC reference area (1460).

[0114] Various constraints may be applied to the BV and / or the IBC reference area.

[0115] In one embodiment, the effective memory requirement for storing reference samples used for Intra Block Copy is 1 CTU size (e.g., CTB size). In one example, the CTU size is 128×128 samples. The current CTU includes the current area being reconstructed. The current area has a size of 64×64 samples. Since the reference memory can also store the reconstructed samples within the current area, when the reference memory size is equal to the CTU size of 128×128 samples, the reference memory can store three additional 64×64 sample areas. Thus, the IBC reference area can include a specific portion of the previously reconstructed CTU while keeping the total memory requirement for storing reference samples unchanged (e.g., 128×128 samples of 1 CTU size or a total of 4 64×64 reference samples). In one example, the previously reconstructed CTU is the one adjacent to the left of the current CTU (left neighbor).

[0116] In some examples, such as in HEVC, additional memory in the DPB is used and the hardware implementation may employ external memory. Additional external memory accesses can increase the memory bandwidth and thus the implementation cost.

[0117] As described above, in order to reduce the implementation cost, the entire already reconstructed reconstructed area (1410) is not used as the IBC reference area. The IBC reference area can be restricted to be within a smaller area, such as the IBC reference area (1460) for example. The search range or the IBC reference area (1460) can be restricted for memory purposes.

[0118] In one embodiment, such as in VVC, fixed memory is used in the IBC mode (or IntraBC mode). Therefore, by implementing IBC using on-chip memory, the memory bandwidth requirements and hardware complexity can be significantly reduced. The reference sample memory (RSM) can be used to store samples of a single CTU, and the size of the RSM is a single CTU. One feature of the RSM is a continuous update mechanism that can replace the reconstructed samples of the left adjacent CTU with the reconstructed samples of the current CTU.

[0119] The BV coding in the IBC mode can adopt the method of the merge list for inter prediction. The BV may be coded either explicitly or implicitly. In the explicit mode, the BV difference (BVD) between the BV and the BV predictor (BVP) can be signaled. The BVD coding can use the MVD coding process used in the AMVP mode for inter prediction to result in the final BV (e.g., the vector sum of the BV predictor and the BVD). In one example, when the reconstructed BV points to an area outside the reference sample area, a correction is performed to remove the absolute offset in each direction using the remainder operation with the width and height of the RSM. The explicit mode may be referred to as the IBC normal mode or the IBC AMVP mode in some examples. In the implicit mode, the BV can be restored from the BVP without using the BVD in a similar manner to the coding of the MV in the merge mode used for inter prediction. The implicit mode may be referred to as the IBC merge mode in some examples.

[0120] When the IBC mode is used, an IBC candidate list including BVP candidates can be constructed. The IBC candidate list can be the IBC merge candidate list when the IBC merge mode is used. The IBC candidate list can be the IBC AMVP candidate list (or the IBC BV predictor list) when the IBC AMVP mode is used. The candidate derivation in the IBC merge candidate list or the IBC AMVP candidate list can follow the same logic as the merge candidate list used in the normal merge mode (used for inter prediction) or the AMVP candidate list used in the normal AMVP mode (used for inter prediction), respectively.

[0121] When constructing the IBC candidate list of the current block coded in the IBC mode, the BV of the spatial neighbors of the current block (e.g., two spatial neighbors) and the history-based BV predictor (HBVP) of the current block (e.g., five HBVPs) are checked and can be included in the IBC candidate list. The BV of the spatial neighbors of the current block can be referred to as spatial candidates. The HBVP of the current block can be referred to as HBVP candidates. The HBVP candidates can include the BV of the reconstructed blocks (e.g., blocks that may not be adjacent to the current block) within the current picture including the current block. In one example, only the first HBVP is compared with the spatial candidates when added to the IBC candidate list. In one example, the IBC candidate list of the current block includes temporal candidates determined based on the BV of the temporal neighbors of the current block.

[0122] In some examples, the IBC merge candidate list can include (one or more) spatial candidates, (one or more) temporal candidates, (one or more) HBVP candidates, and (one or more) BVP candidates which are pairwise candidates based on two existing BVP candidates in the IBC merge candidate list. The pairwise candidates can be determined based on two existing BVP candidates in the IBC merge candidate list. The IBC merge candidate list can include up to six BVP candidates. The IBC AMVP candidate list can include the first two BVP candidates that can be used in the IBC merge candidate list.

[0123] In one example, the IBC candidate list in the IBC mode is for both cases such as the IBC merge mode and the IBC AMVP mode. For example, the IBC merge mode can use up to six candidates of the IBC candidate list. Usually, the normal IBC mode can use only the first two candidates of the IBC candidate list.

[0124] FIG. 15 shows an example of an Intra-template matching prediction (IntraTMP) mode applied to a current block (1501) within a current picture (or current frame) (1500). The IntraTMP mode can be used in video / image coding, for example, in ECM5. In FIG. 15, the current picture (1500) includes CTUs R1-R4.

[0125] CTU R1 is the current CTU being reconstructed. The CTUs R2-R4 corresponding to the CTU above the current CTU R1, the CTU to the left of the current CTU R1, and the CTU in the upper left of the current CTU R1 have already been reconstructed. The upper left area of R1 includes regions (1521)-(1523), the current block (1501), and the current template (1511) of the current block. The current template (1511) can include adjacent samples of the current block (1501) and has already been reconstructed. In the example shown in FIG. 15, the current template (1511) is L-shaped. The current template may have another shape and can include any suitable number of samples. Region (1521) has already been reconstructed. In one example, regions (1522)-(1523) have already been reconstructed.

[0126] To encode (e.g., reconstruct) the current block (1501), the already reconstructed templates within the search area (gray area) (1510) can be checked. The search area (1510) can include R2 - R4 and the area (1521). The areas (1522) - (1523) are not included in the search area (1510). The current template (1511) of the current block (1501) can be compared with each of the respective templates within the search area (1510). The templates within the search area (1510) can have the same shape as the current template (1511). Based on the current template (1511) and each of the respective templates within the search area (1510), a template matching (TM) cost can be determined. The TM cost can be determined based on the sum of absolute differences (SAD) between the current template (1511) and each of the respective templates. Other functions such as, for example, the sum of squared errors (SSE), variance, partial SAD, or the like can also be used to determine the TM cost.

[0127] Based on the TM cost, a predicted block (1502) is determined. In one example, the predicted block (1502) corresponds to the minimum TM cost among the determined TM costs. The predicted block (1502) can be referred to as the best predicted block where the template of the predicted block (1502) (also referred to as the reference template) (1512) matches the current template (1511). The block vector BV IntraTMP (Also referred to as the IntraTMP - based block vector) (1530) can indicate the displacement in position between the reference template (1512) and the current template (1511). In one example, the displacement in position between the predicted block (1502) and the current block (1501) is the same as the displacement in position between the reference template (1512) and the current template (1511). The block vector BV IntraTMP (1530) can indicate the displacement in position between the predicted block (1502) and the current block (1501).

[0128] The IntraTMP mode is an intra prediction mode that copies a prediction block (1502) from the reconstructed part of the current picture (1500). In one example, the reference template (1512) of the prediction block (1502) matches the current template (1511). For a predetermined search range, the encoder can search for the template (e.g., the reference template (1512)) that is most similar to the current template (1511) within the reconstructed part of the current frame (1500), and use the corresponding block as the prediction block (e.g., the prediction block (1502)). The encoder can signal the use of the IntraTMP mode, for example, the same prediction operation as described above can be performed on the decoder side by the decoder.

[0129] As described above, a prediction signal can be generated by matching the current template (1511) of the current block (1501) (e.g., the L-shaped causal neighbor) with the template of another block (e.g., the prediction block (1502)) within a predetermined search range. Referring to FIG. 15, the search range can include a plurality of CTUs, such as CTU R1-R4. The search range can be predetermined. In one embodiment, only the reconstructed area (e.g., R2-R4 and areas (1521)-(1523) in FIG. 15) within the predetermined search range can be searched. Also, in some examples, a specific reconstructed area adjacent to the current block (1501) (e.g., areas (1522)-(1523) that are the upper neighbor and the left neighbor of the current block (1501)) is not searched. Therefore, in the example shown in FIG. 15, the search area (1510) within the predetermined search range includes R2-R4 and area (1521).

[0130] Within the search range (also referred to as the search area), the decoder can search for a reference template that has the minimum TM cost (e.g., the minimum SAD) for the current template (1511), and use the corresponding block of the reference template as the predicted block (e.g., (1502)). The dimensions of the search range, such as the search range width SearchRange_w and the search range height SearchRange_h, can be set to be proportional to the block dimensions, such as the block width BlkW and the block height BlkH, respectively, so as to have a fixed number of template comparisons (e.g., SAD comparisons) per pixel: SearchRange_w = a × BlkW Equation 1 SearchRange_h = a × BlkH Equation 2

[0131] The parameter "a" in Equations 1 - 2 is a constant that controls the trade - off between gain and complexity. In one example, "a" in Equations 1 - 2 is 5. The parameter "a" in Equations 1 - 2 can be constrained such that the search range is within a plurality of CTUs (e.g., the four CTUs shown in FIG. 15). In one example, the search area (1510) is within the search range.

[0132] The IntraTMP mode can be enabled for CUs having a size of 64 or less in both the horizontal and vertical directions. For example, the IntraTMP mode is enabled when the CU or block is smaller than 64×64. The maximum CU size for the IntraTMP mode can be configurable. The IntraTMP mode can be signaled at the CU level through a dedicated flag, for example, when the decoder - side intra - mode derivation (DIMD) is not used for the current CU.

[0133] For example, in one example such as in ECM5, the IntraTMP mode accesses 320 up-samples and 320 left-samples to support a 64×64 block. For example, the memory size such as 320 up-samples and 320 left-samples of a block can improve the coding efficiency of the IBC mode. The reference area or search range for the IBC mode can be extended. In one example, the reference area for the IBC mode is extended to the upper two CTU rows. FIG. 16 shows an example of a reference area for coding CTU(m,n). The integers m and n are indices representing the position of the CTU. To code CTU(m,n), the reference area can include CTUs with indices (m-2,n-2), …, (W-1,n-2), (0,n-1), …, (W-1,n-1), (0,n), …, and (m,n), where W represents the maximum horizontal index for CTUs in the current tile, slice, or picture, etc. This setting (e.g., accessing 320 up-samples and 320 left-samples to predict a block) can ensure that for a 128×128 CTU size, the IBC mode does not require additional memory in the platform of the current Essential Video Coding Test Model (ETM). The per-sample block vector search range (or referred to as the local search range) can be limited to [-(C<<1),C>>2] (or [-2C,C / 4]) in the horizontal direction and [-C,C>>2] (or [-C,C / 4]) in the vertical direction to adapt to the reference area extension. C represents the CTU size, such as 128 for example. For example, the BV of a block is limited to [-2C,C / 4]) in the horizontal direction and [-C,C / 4]) in the vertical direction.

[0134] In some embodiments as described above, the template matching search procedure is used to find the best template within the search range in the IntraTMP mode. The block vector BV between the searched reference template (e.g., (1512)) and the current template (e.g., (1511)) of the block to be IntraTMP-coded (e.g., (1501)) IntraTMP (e.g., (1530)) is not stored and thus cannot be used to code another block in either the IBC mode or the IntraTMP mode.

[0135] This disclosure describes embodiments related to the construction of an IBC candidate list (e.g., an IBC merge list or an IBC AMVP list) based on the motion data of a block coded in the IntraTMP mode. The motion data of the block coded in the IntraTMP mode can be stored in the motion storage to construct an IBC candidate list (e.g., an IBC merge list), for example, the block vector BV IntraTMP is stored in the motion buffer as a predictor for predicting another block.

[0136] According to an embodiment of this disclosure, the block vector (BV IntraTMP ) used in the IntraTMP mode between the searched reference template (e.g., (1512)) and the current template (e.g., (1511)) of the current block (e.g., (1501)) coded in the IntraTMP mode can be stored. The stored block vector (BV IntraTMP ) can be used to code (e.g., reconstruct or encode) another block, such as another block in the IntraTMP mode or the IBC mode. The another block may be within the current picture including the current block or in a picture different from the current picture.

[0137] In one embodiment, the block vector (BV IntraTMPThe memory of IntraTMP is the same as or identical to the memory of the BV used in the IBC mode, as will be described later. The current block can include one or more M×N units (e.g., in luma samples). For example, each M×N unit includes M×N samples. The integers M and N may be the same or different. The block vector (BV IntraTMP ) can be stored in M×N units such as 8×8 units (e.g., in VVC) or 4×4 units (e.g., in ECM). The block vector (BV

[0138] In one embodiment, the block vector (BV IntraTMP ) is stored using a predetermined precision such as, for example, 1 / 2 pixel (1 / 2 pel), 1 pixel (1 pel), or 2 pixels (2 pels). In one example, the predetermined precision is the only allowable precision for storing the block vector (BV IntraTMP ).

[0139] In another embodiment, the block vector (BV IntraTMP ) is stored using one of the precisions within a list of predetermined precisions. The list of predetermined precisions can include, but is not limited to, 1 / 4 pixel (1 / 4 pel), 1 / 2 pel, 1 pel, and 4 pixels (or 4 pels). An index or other information indicating which precision within the list of predetermined precisions can be selected for storing the block vector (BV IntraTMP ) can be signaled. The precision can be selected based on the index. The list of predetermined precisions may include additional precision(s) or omit one or more precisions.

[0140] In one embodiment, the current block is coded in the IntraTMP mode. (i) One or more BVs related to one or more blocks that have already been coded in the IBC mode, and / or (ii) one or more block vectors (BVs) related to one or more blocks that have already been coded in the IntraTMP modeIntraTMP ) The (one or more) reference templates indicated by IntraTMP can be used as template matching candidates in the IntraTMP mode. In one example, the starting reference template indicated by (i) BV and / or (ii) one of the block vectors (BV

[0141] In one embodiment, the current block is coded using the IntraTMP mode. Whether to store the block vector (BV IntraTMP ) of the current block may depend on the TM cost C based on the current template (e.g., (1511)) of the current block (e.g., (1501)) and the reference template (e.g., (1512)) of the predicted block (e.g., (1502)). For example, if the TM cost C is less than a given threshold, the block vector (BV IntraTMP ) of the current block is stored. If the TM cost C is not less than the given threshold, the block vector (BV IntraTMP ) of the current block is not stored.

[0142] In one embodiment, the TM cost C is normalized with respect to, for example, the size of the current template (e.g., the number of samples N CT ) in the current template. For example, the normalized TM cost C N is equal to C / N CT . The normalized TM cost C N can be stored, and the TM cost C is not stored.

[0143] In one embodiment, the partial TM cost related to the M×N unit of the current block coded by IntraTMP is stored. The partial TM cost can be determined based on a subset of the current template (e.g., (1511)) corresponding to the M×N unit and a subset of the reference template (e.g., (1512)).

[0144] As described above, the stored block vector (BV IntraTMP ) of the current block coded in the IntraTMP mode can be used to code another block in the IBC mode. In one embodiment, the first block in the current picture is to be coded in the IBC mode. An IBC candidate list for the first block can be constructed. The IBC candidate list can include a first candidate based on the block vector (BV IntraTMP )(1530) of the second block coded in the IntraTMP mode, such as the block vector (BV IntraTMP ) shown in FIG. 15 for example. The first candidate can be used as a BVP candidate in the IBC candidate list. In one example, the first candidate is the block vector (BV IntraTMP ) of the second block coded in the IntraTMP mode. The second block can be one of (i) a reconstructed block in the current picture and (ii) a temporal neighbor of the first block. In one example, the reconstructed block in the current picture is a spatial neighbor of the first block. In one example, the reconstructed block in the current picture may not be adjacent to the first block. For example, the block vector ( BVIntraTMP ) of the second block is stored in a history-based table such as a BV history table used in, for example, HBVP.

[0145] In one example, the IBC mode is the IBC merge mode, and thus, the IBC candidate list is the IBC merge candidate list (or IBC merge list). In one example, the IBC mode is the IBC AMVP mode (or IBC normal mode), and thus, the IBC candidate list is the IBC BV predictor (BVP) list.

[0146] One or more BV(s) of the block(s) coded with IBC and / or one or more block vectors (BV IntraTMP ) of the block(s) coded with IntraTMP can be used as one or more BVP candidates in the IBC candidate list, for example, when the block(s) coded with IBC and / or the block(s) coded with IntraTMP are spatial and / or temporal neighbors of the first block. Additionally, one or more BV(s) of the block(s) coded with IBC and / or one or more block vectors (BV IntraTMP ) can be used as one or more BVP candidates in the IBC candidate list when the one or more BV(s) and / or the one or more block vectors (BV IntraTMP ) are in a history-based table such as a BV history table used in HBVP.

[0147] One or more BVP candidates based on one or more block vectors (BV IntraTMP ) can be referred to as IntraTMP-based BVP candidates, and one or more BVP candidates based on one or more BV(s) can be referred to as IBC-based BVP candidates.

[0148] In one embodiment, to refine the block vector (BV IntraTMP ) of the second block, the template matching described in FIG. 13 can be applied, and the refined block vector (BV IntraTMP) is determined. The first candidate in the IBC candidate list is the refined block vector (BV IntraTMP ) of the second block.

[0149] In one embodiment, the IBC candidate list (e.g., IBC merge list or IBC BVP list) of the first block includes a plurality of IBC candidates (e.g., BVP candidates). The plurality of BVP candidates include the (one or more) BV of the (one or more) blocks coded by IBC and / or the (one or more) block vectors (BV IntraTMP ) of the (one or more) blocks coded by IntraTMP as described above. In one example, the plurality of BVP candidates includes a first candidate that is a block vector (BV IntraTMP ). The TM cost of each BVP candidate can be determined based on the current template of the first block and the reference template indicated by each BVP candidate, as described, for example, with respect to FIG. 13. The plurality of BVP candidates can be sorted based on the determined TM costs, for example, in ascending order of the determined TM costs. A sorted (e.g., rearranged) IBC candidate list is formed. An index can indicate the BVP candidates in the rearranged IBC candidate list.

[0150] In one embodiment, for each of the plurality of BVP candidates that is a block vector (BV IntraTMP ), a scaling factor is applied to the TM cost of each BVP candidate.

[0151] In one embodiment, the maximum number of IntraTMP-based BVP candidates based on the block vector (BV IntraTMP ) of the blocks coded in IntraTMP mode in the IBC candidate list is a threshold M. The blocks coded in IntraTMP can include the spatial neighbor and / or temporal neighbor of the first block. The block vector (BV IntraTMP ) can also be obtained, for example, from a history-based table as described above.

[0152] In one example, the number of IntraTMP-based BVP candidates in the IBC candidate list is equal to the threshold M. The third block (e.g., block j) is one of (i) a reconstructed block in the current picture (e.g., a spatial neighbor of the first block) and (ii) a temporal neighbor of the first block, and the third block is coded in IntraTMP mode. In one example, the third block is the adjacent block j of the first block. The block vector BV of the third block IntraTMP Whether to include it in the IBC candidate list can be determined based at least on the TM cost associated with the block vector BV of the third block IntraTMP of the third block. The TM cost associated with the block vector BV of the third block IntraTMP If the TM cost associated with the block vector BV of the third block is less than the TM cost associated with at least one of the (one or more) IntraTMP-based BVP candidates, the block vector BV of the third block IntraTMP can replace a certain candidate among the (one or more) IntraTMP-based BVP candidates. The candidate to be replaced can be the one with the maximum TM cost among the (one or more) IntraTMP-based BVP candidates.

[0153] When constructing the IBC candidate list, the (one or more) IntraTMP-based BVP candidates and the (one or more) IBC-based BVP candidates can be added in any suitable order.

[0154] In one embodiment, the (one or more) block vectors (BV IntraTMPcan be added to the IBC candidate list as (one or more) IntraTMP-based BVP candidates when constructing the IBC candidate list. In one example, the (one or more) IntraTMP-based BVP candidates are added to the IBC candidate list prior to adding the (one or more) BVs of the (one or more) blocks to be IBC-coded. In one example, only the IntraTMP-based BVP candidates with the M smallest TM costs are added to the IBC candidate list. When the number of IntraTMP-based BVP candidates reaches a threshold M (e.g., becomes equal), the (one or more) BVs of the (one or more) blocks to be IBC-coded can be added to the IBC candidate list. The (one or more) blocks to be IBC-coded can include the (one or more) spatial or temporal neighboring blocks of the first block. In one example, the (one or more) blocks to be IBC-coded include the (one or more) other reconstructed blocks in the current picture, and the (one or more) BVs of the (one or more) other reconstructed blocks in the current picture are stored in a history-based buffer such as used in HBVP.

[0155] In one example, (one or more) IBC-based BVP candidates from the (one or more) BVs of the (one or more) blocks to be IBC-coded are added to the IBC candidate list, and then, (one or more) IntraTMP-based BVP candidates from the (one or more) block vectors (BV IntraTMP ) of the (one or more) blocks to be IntraTMP-coded may be added to the IBC candidate list.

[0156] In one embodiment, the block vector (BV IntraTMP ) of the block to be IntraTMP-coded may be added to the BV history table and can be used, for example, as a history-based BV candidate / predictor derivation as described above when constructing the IBC candidate list.

[0157] Figure 17 shows a flowchart outlining an encoding process (1700) according to an embodiment of the present disclosure. The process (1700) can be used in a video encoder. The process (1700) can be executed by an apparatus for video coding that can include a processing circuit. In various embodiments, the process (1700) is executed by a processing circuit, such as, for example, the processing circuits of terminal devices (310), (320), (330), and (340), the processing circuit that executes the functions of video encoders (e.g., (403), (603), (703)), and the like. In some embodiments, the process (1700) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1700). The process starts at (S1701) and proceeds to (S1710).

[0158] (S1710), the first block within the current picture can be encoded in an Intra Template Matching Prediction (IntraTMP) mode based on a predicted block within the reconstructed search area in the current picture. The reference template of the predicted block is matched to the current template of the first block.

[0159] In one embodiment, the reference template is determined based on a plurality of template candidates within the reconstructed search area in the current picture. The displacement in position between one of the plurality of template candidates and the current template is (i) the block vector (BV) of a third block coded in an IBC mode, or (ii) the block vector BV of a third block coded in an IntraTMP mode IntraTMP and can be indicated by a vector that is

[0160] (S1720), the block vector (e.g., IntraTMP-based block vector) BV of the first block IntraTMP can be stored. The block vector BV IntraTMPindicates the displacement in position between the current template of the first block and the reference template of the predicted block.

[0161] In one embodiment, the first block includes one or more units, and each unit includes M×N samples. The block vector BV IntraTMP can be stored in each M×N unit of the first block.

[0162] In one example, the block vector BV IntraTMP is stored with a predetermined allowable accuracy. In one example, the block vector BV IntraTMP is stored in accordance with a predetermined system.

[0163] In one example, the block vector BV IntraTMP is stored when the storage condition is satisfied. The storage condition includes that the template matching cost between the reference template of the predicted block and the current template of the first block is smaller than the threshold.

[0164] In one example, the template matching cost between the reference template of the predicted block and the current template of the first block is stored. For example, the template matching cost is normalized based on the number of samples in the current template, and the normalized template matching cost is stored.

[0165] (S1730), the second block can be encoded based on the stored block vector BV IntraTMP . In one example, the second block is encoded in the IntraTMP mode or the intra-block copy (IBC) mode. For example, the IBC candidate list of the second block includes BVP candidates based on the stored block vector BV IntraTMP . The second block can be encoded based on the IBC candidate list.

[0166] Then, the process (1700) proceeds to (S1799) and ends.

[0167] The process (1700) can be suitably adapted for various scenarios, and accordingly, the steps of the process (1700) can be adjusted. One or more of the steps of the process (1700) can be adapted, omitted, repeated, and / or combined. The process (1700) can be implemented in any suitable order. Further (one or more) steps can be added.

[0168] FIG. 18 shows a flowchart outlining a decoding process (1800) according to an embodiment of the present disclosure. The process (1800) can be used in a video decoder. The process (1800) can be executed by an apparatus for video coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (1800) is executed by a processing circuit, such as, for example, the processing circuits of terminal devices (310), (320), (330), and (340), the processing circuit that executes the functions of video encoder (403), the processing circuit that executes the functions of video decoder (410), the processing circuit that executes the functions of video decoder (510), the processing circuit that executes the functions of video encoder (603), and the like. In some embodiments, the process (1800) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1800). The process starts at (S1801) and proceeds to (S1810).

[0169] At (S1810), prediction information indicating whether the first block in the current picture is coded in the Intra Template Matching (IntraTMP) mode is obtained. A coded video bitstream having the first block in the current picture can be received. In one example, the prediction information is decoded from the coded video bitstream.

[0170] When prediction information indicates that the IntraTMP mode is to be applied to the first block at (S1820), the first block can be reconstructed based on a predicted block within a reconstructed search area in the current picture. In the IntraTMP mode, a reference template of the predicted block can be matched to the current template of the first block.

[0171] In one embodiment, the reference template is determined based on a plurality of template candidates within a reconstructed search area in the current picture. A displacement in position between one of the plurality of template candidates and the current template is a vector indicated by (i) a block vector (BV) of a third block coded in the IBC mode, or (ii) a block vector BV of a third block coded in the IntraTMP mode IntraTMP , and can be indicated by a vector which is.

[0172] At (S1830), a block vector (e.g., an IntraTMP-based block vector) BV of the first block IntraTMP can be stored. The block vector BV IntraTMP indicates a displacement in position (or a motion vector displacement) between the current template of the first block and the reference template of the predicted block.

[0173] In one embodiment, the first block includes one or more units, each unit including M×N samples. The block vector BV IntraTMP can be stored in each M×N unit of the first block.

[0174] In an example, the block vector BV IntraTMP is stored with a predetermined tolerance accuracy. In an example, the block vector BV IntraTMP is stored with an accuracy indicated by syntax information (e.g., an index) in the encoded video bitstream.

[0175] In an example, the block vector BV IntraTMPIt is stored when the storage condition is satisfied. The storage condition includes that the template matching cost between the reference template of the prediction block and the current template of the first block is smaller than the threshold value.

[0176] In one example, the template matching cost between the reference template of the prediction block and the current template of the first block is stored. For example, the template matching cost is normalized based on the number of samples in the current template, and the normalized template matching cost is stored.

[0177] (In S1840), the stored block vector BV IntraTMP Based on this, the second block can be reconstructed. In one example, the second block is coded in the IntraTMP mode or the intra-block copy (IBC) mode. In one example, the second block is within the current picture. For example, the IBC candidate list of the second block includes BVP candidates based on the stored block vector BV IntraTMP Based on this IBC candidate list, the second block can be reconstructed.

[0178] The process (1800) proceeds to (S1899) and ends.

[0179] The process (1800) can be suitably adapted to various scenarios, and accordingly, the steps of the process (1800) can be adjusted. One or more of the steps of the process (1800) can be adapted, omitted, repeated, and / or combined. The process (1800) can be implemented using any suitable order. Further (one or more) steps can be added.

[0180] FIG. 19 shows a flowchart outlining an encoding process (1900) according to an embodiment of the present disclosure. The process (1900) can be used in a video encoder. The process (1900) can be executed by an apparatus for video coding that can include a processing circuit. In various embodiments, the process (1900) is executed by a processing circuit, such as, for example, the processing circuits of terminal devices (310), (319), (330), and (340), the processing circuit that executes the functions of video encoders (e.g., (403), (603), (703)), and the like. In some embodiments, the process (1900) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1900). The process starts at (S1901) and proceeds to (S1910).

[0181] (S1910), an IBC candidate list for a first block in the current picture can be constructed. The first block is to be encoded in an intra block copy (IBC) mode. The IBC candidate list for the first block includes a first candidate based on a block vector BV of a second block coded in an intra template matching prediction (IntraTMP) mode. The second block can be one of (i) a reconstructed block in the current picture and (ii) a temporal neighbor of the first block. IntraTMP In one embodiment, template matching is performed on the block vector BV

[0182] to refine the block vector BV IntraTMP The first candidate can be determined as the template-matched block vector BV IntraTMP IntraTMP IntraTMP

[0183] (S1920), the first block can be encoded based on the IBC candidate list.

[0184] ​Then, the process (1900) proceeds to (S1999) and ends.

[0185] The process (1900) can be suitably adapted for various scenarios, and accordingly, the steps of the process (1900) can be adjusted. One or more of the steps of the process (1900) can be adapted, omitted, repeated, and / or combined. The process (1900) can be implemented using any suitable order. Further (one or more) steps can be added.

[0186] In one embodiment, the IBC candidate list of the first block includes a plurality of candidates. Each of the plurality of candidates can be based on one of (i) the block vector BV of the block coded in the IntraTMP mode IntraTMP , and (ii) the block vector (BV) of the block coded in the IBC mode. The plurality of candidates includes a first candidate. As follows, template matching can be performed on the plurality of candidates. For each of the plurality of candidates, the respective template matching cost can be determined based on the reference template of the reference block and the current template of the first block. The plurality of candidates can be rearranged based on the determined template matching costs. The first block can be reconstructed based on the rearranged plurality of candidates in the IBC candidate list.

[0187] In one example, for each of the plurality of candidates that is the block vector BV IntraTMP , a scaling factor is applied to the respective template matching cost of the candidate.

[0188] In one example, the number of one or more candidates in the IBC candidate list that are the block vectors BVsIntraTMP of the respective blocks coded in the IntraTMP mode is below a threshold. The one or more candidates include the first candidate.

[0189] In one example, the number of one or more candidates in the IBC candidate list is equal to a threshold. The third block is one of (i) a reconstructed block in the current picture, and (ii) a temporal neighbor of the first block. A new block vector BV of the third block that is not in the IBC candidate list IntraTMP for the new block vector BV IntraTMP if the template matching cost associated with the new block vector BV is less than the template matching cost associated with at least one of the one or more candidates, then a certain candidate among the one or more candidates is replaced with the new block vector BV IntraTMP . The candidate to be replaced is the one with the maximum template matching cost among the one or more candidates.

[0190] In one example, one or more candidates from one or more blocks coded in IBC mode are added to the IBC candidate list.

[0191] In one example, the block vector BV of the second block IntraTMP is obtained from one or more block vectors (BVs) of at least one previously coded block in the current picture or a block vector (BV) history table storing one or more block vectors BV IntraTMP .

[0192] FIG. 20 shows a flowchart outlining a decoding process (2000) according to an embodiment of the present disclosure. The process (2000) can be used in a video decoder. The process (2000) can be executed by an apparatus for video coding that can include a receiving circuit and a processing circuit. In various embodiments, the process (2000) is executed by a processing circuit, such as, for example, the processing circuits of terminal devices (310), (320), (330), and (340), the processing circuit that executes the functions of video encoder (403), the processing circuit that executes the functions of video decoder (410), the processing circuit that executes the functions of video decoder (510), the processing circuit that executes the functions of video encoder (603), and the like. In some embodiments, the process (2000) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2000). The process starts at (S2001) and proceeds to (S2010).

[0193] (S2010), prediction information of a first block within the current picture can be decoded from the coded video bitstream. The prediction information indicates that an intra block copy (IBC) mode is applied to the first block.

[0194] (S2020), an IBC candidate list for the first block can be constructed. The IBC candidate list includes a first candidate based on a block vector BV of a second block coded in an Intra Template Matching Prediction (IntraTMP) mode. The second block can be one of (i) a reconstructed block within the current picture and (ii) a temporal neighbor of the first block. IntraTMP In one embodiment, template matching is performed on the block vector BV to refine the block vector BV. The first candidate can be determined as the template-matched block vector BV.

[0195] In one embodiment, template matching is performed on the block vector BV IntraTMP to refine the block vector BV IntraTMP . The first candidate can be determined as the template-matched block vector BV IntraTMP .

[0196] (S2030), the first block can be reconstructed based on the IBC candidate list.

[0197] The process (2000) proceeds to (S2099) and ends.

[0198] The process (2000) can be suitably adapted for various scenarios, and accordingly, the steps of the process (2000) can be adjusted. One or more of the steps of the process (2000) can be adapted, omitted, repeated, and / or combined. The process (2000) can be implemented using any suitable order. Further (one or more) steps can be added.

[0199] In one embodiment, the IBC candidate list of the first block includes a plurality of candidates. Each of the plurality of candidates can be based on one of (i) the block vector BV of the block coded in the IntraTMP mode IntraTMP , and (ii) the block vector (BV) of the block coded in the IBC mode. The plurality of candidates includes a first candidate. As follows, template matching can be performed on the plurality of candidates. For each of the plurality of candidates, the respective template matching cost can be determined based on the reference template of the reference block and the current template of the first block. The plurality of candidates can be sorted based on the determined template matching cost. The first block can be reconstructed based on the sorted plurality of candidates in the IBC candidate list.

[0200] In one example, for each of the plurality of candidates that is the block vector BV IntraTMP , a scaling factor is applied to the respective template matching cost of each candidate.

[0201] In one example, the number of one or more candidates in an IBC candidate list that are block vectors BVsIntraTMP of respective blocks coded in IntraTMP mode is below a threshold. The one or more candidates include a first candidate.

[0202] In one example, the number of one or more candidates in an IBC candidate list is equal to a threshold. A third block is one of (i) a reconstructed block in the current picture and (ii) a temporal neighbor of the first block. A new block vector BV of the third block that is not in the IBC candidate list IntraTMP for which the template matching cost associated with the new block vector BV IntraTMP is less than the template matching cost associated with at least one of the one or more candidates, then a certain candidate among the one or more candidates is replaced with the new block vector BV IntraTMP . The candidate to be replaced is the one with the maximum template matching cost among the one or more candidates.

[0203] In one example, one or more candidates from one or more blocks coded in IBC mode are added to the IBC candidate list.

[0204] In one example, the block vector BV of a second block IntraTMP is obtained from one or more block vectors (BVs) of at least one block coded previously in the current picture or a block vector (BV) history table storing one or more block vectors BV IntraTMP .

[0205] Embodiments of the present disclosure can be used separately or in combination in any order. Also, each of these methods (or embodiments), the encoder, and the decoder can be implemented by a processing circuit (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.

[0206] The above technology can be implemented as computer software using computer-readable instructions physically stored on one or more computer-readable media. For example, FIG. 21 shows a computer system (2100) suitable for implementing a particular embodiment of the disclosed matter.

[0207] The computer software can be coded using any suitable machine code or computer language such that, when subjected to assembly, compilation, linking, or similar mechanisms, it can create code having instructions that can be executed directly or via interpretation, microcode execution, and the like by one or more computer central processing units (CPUs), graphics processing units (GPUs), and the like.

[0208] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0209] The components shown in FIG. 21 with respect to the computer system (2100) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Also, the configuration of the components should not be construed as having any dependency or requirement with respect to any one or combination of the components shown in this exemplary embodiment of the computer system (2100).

[0210] The computer system (2100) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, moving a data glove, etc.), audio input (e.g., voice, clapping, etc.), visual input (e.g., gestures, etc.), olfactory input (not shown). The human interface device may also be used to capture specific media that is not necessarily directly related to conscious human input, such as, for example, audio (e.g., conversation, music, ambient sounds, etc.), images (e.g., scanned images, photographic images obtained from a still camera, etc.), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0211] The input human interface device may include one or more of a keyboard (2101), a mouse (2102), a trackpad (2103), a touch screen (2110), a data glove (not shown), a joystick (2105), a microphone (2106), a scanner (2107), a camera (2108) (each shown only once).

[0212] The computer system (2100) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (2110), a data glove (not shown), or a joystick (2105), although there may also be a tactile feedback device that does not function as an input device), audio output devices (e.g., speakers (2109), headphones (not shown), etc.), visual output devices (e.g., a touch screen (2110) including a CRT screen, an LCD screen, a plasma screen, an OLED screen (each having or not having a touch screen input function, each having or not having a tactile feedback function. Some of these can output two-dimensional visual output or output higher-dimensional output through means such as stereoscopic output), virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), etc.), and printers (not shown).

[0213] The computer system (2100) may also include human-accessible storage devices and their associated media such as, for example, an optical medium including a CD / DVD ROM / RW (2120) having a CD / DVD or similar medium (2121), a thumb drive (2122), a removable hard drive or solid state drive (2123), legacy magnetic media such as tapes and floppy disks (registered trademark, not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.

[0214] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the matters disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.

[0215] The computer system (2100) may also include an interface (2154) to one or more communication networks (2155). The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include, for example, local area networks such as Ethernet (registered trademark), wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE and the like, TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial including CANBus, and the like. A particular network generally requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus (2149) (e.g., a USB port of the computer system (2100)), and others are generally integrated into the core of the computer system (2100) by attachment to the system bus described later (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2100) can communicate with other entities. Such communication can be only one-way reception (e.g., broadcast TV), only one-way transmission (e.g., CANbus to a particular CANbus device), or two-way, for example, to other computer systems using a local or wide area digital network. Particular protocols and protocol stacks can be used on each of the networks and network interfaces as described above.

[0216] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (2140) of the computer system (2100).

[0217] The core (2140) may include one or more central processing units (CPUs) (2141), a graphics processing unit (GPU) (2142), a special programmable processing unit in the form of a field programmable gate array (FPGA) (2143), a hardware accelerator for specific tasks (2144), a graphics adapter (2150), and the like. These devices can be connected via a system bus (2148) along with a read-only memory (ROM) (2145), a random access memory (2146), an internal mass storage (2147) such as an internal hard drive, SSD, and the like that are not accessible to internal users. In some computer systems, the system bus (2148) can be made accessible in the form of one or more physical plugs to allow for expansion by additional CPUs, GPUs, and the like. Peripheral devices may be attached either directly to the system bus of the core (2148) or via a peripheral bus (2149). In one example, a touch screen (2110) can be connected to the graphics adapter (2150). The architecture of the peripheral bus includes PCI, USB, and the like.

[0218] The CPU (2141), GPU (2142), FPGA (2143), and accelerator (2144) can execute specific instructions that can be combined to form the aforementioned computer code. The computer code can be stored in the ROM (2145) or RAM (2146). Transient data can also be stored in the RAM (2146), and permanent data can be stored, for example, in the internal mass storage (2147). Fast storage and retrieval to any of the memory devices can be enabled by the use of cache memory that may be associated near one or more CPUs (2141), GPUs (2142), mass storage (2147), ROM (2145), RAM (2146), and the like.

[0219] A computer-readable medium can have computer code thereon for performing various computer-implemented processes. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well-known and available to those having skill in the computer software arts.

[0220] As an example, and not by way of limitation, a computer system having an architecture (2100), particularly a core (2140), can provide functionality as a result of software embodied on one or more tangible computer-readable media being executed by one or more processors (including CPUs, GPUs, FPGAs, accelerators, and the like). Such computer-readable media can be, for example, specific storage of the core (2140) that is non-transitory in nature, such as mass storage (2147) inside the core or ROM (2145), and media associated with user-accessible mass storage as introduced above. The software implementing various embodiments of the present disclosure can be stored on such a device and executed by the core (2140). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core (2140) and particularly the processors (including CPUs, GPUs, FPGAs, and the like) therein to define data structures stored in RAM (2146) and modify such data structures according to processes defined by the software, and execute the specific processes described herein or specific portions of the specific processes. Additionally, or alternatively, the computer system can provide functionality as a result of logic wired or otherwise embodied in a circuit (e.g., an accelerator (2144)) that operates instead of or in conjunction with software to execute the specific processes described herein or specific portions of the specific processes. References to software include logic, and vice versa where appropriate. References to computer-readable media can include circuits (e.g., integrated circuits (ICs), etc.) storing software for execution, circuits embodying logic for execution, or both where appropriate. The present disclosure includes suitable combinations of hardware and software Appendix A: Acronyms JEM: joint exploration model VVC: Versatile Video Coding (Versatile Video Coding) BMS: Benchmark Set (Benchmark Set) MV: Motion Vector (Motion Vector) HEVC: High Efficiency Video Coding (High Efficiency Video Coding) SEI: Supplementary Enhancement Information (Supplementary Enhancement Information) VUI: Video Usability Information (Video Usability Information) GOPs: Groups of Pictures (Groups of Pictures) TUs: Transform Units, (Transform Units) PU: Prediction Units (Prediction Units) CTUs: Coding Tree Units (Coding Tree Units) CTBs: Coding Tree Blocks (Coding Tree Blocks) PBs: Prediction Blocks (Prediction Blocks) HRD: Hypothetical Reference Decoder (Hypothetical Reference Decoder) SNR: Signal Noise Ratio (Signal Noise Ratio) CPUs: Central Processing Units (Central Processing Units) GPUs: Graphics Processing Units (Graphics Processing Units) CRT: Cathode Ray Tube (Cathode Ray Tube) LCD: Liquid-Crystal Display (Liquid-Crystal Display) OLED: Organic Light-Emitting Diode (Organic Light-Emitting Diode) CD: Compact Disc (Compact Disc) DVD: Digital Video Disc (Digital Video Disk) ROM: Read-Only Memory (Read-Only Memory) RAM: Random Access Memory (Random Access Memory) ASIC: Application-Specific Integrated Circuit (Application-Specific Integrated Circuit) PLD: Programmable Logic Device (Programmable Logic Device) LAN: Local Area Network (Local Area Network) GSM: Global System for Mobile communications (Global System for Mobile Communications) LTE: Long-Term Evolution (Long-Term Evolution) CANBus: Controller Area Network Bus (Controller Area Network Bus) USB: Universal Serial Bus (Universal Serial Bus) PCI: Peripheral Component Interconnect (Peripheral Component Interconnect) FPGA: Field Programmable Gate Areas (Field Programmable Gate Array) SSD: solid-state drive (Solid State Drive) IC: Integrated Circuit (Integrated Circuit) CU: Coding Unit (Coding Unit) R-D: Rate-Distortion (Rate-Distortion)

[0221] Although this disclosure describes several exemplary embodiments, there are changes, substitutions, and various equivalent alternatives that fall within the scope of the disclosure. Thus, it is understood that those skilled in the art can devise numerous systems and methods that, although not explicitly illustrated or described herein, embody the principles of the disclosure and are therefore within its spirit and scope.

Claims

1. A method for video decoding performed by a video decoder, comprising: receiving an encoded video bitstream having a first block in a current picture; obtaining prediction information indicating whether the first block is coded in an Intra Template Matching (IntraTM) mode; in response to the IntraTM mode being applied to the first block, reconstructing the first block based on a predicted block in a reconstructed search area in the current picture, wherein in the IntraTM mode, a reference template of the predicted block is matched to a current template of the first block; The step of storing the IntraTMP-based block vector BV of the first block IntraTMP wherein the IntraTMP-based block vector BV IntraTMP indicates the displacement in position between the current template of the first block and the reference template of the predicted block, and the step Reconstructing a second block coded in either an Intra Block Copy (IntraBC) mode or an IntraTMP mode based on the stored IntraTMP-based block vector BV IntraTMP and A method having the above steps.

2. The first block includes one or more M×N units. The storing step stores the IntraTM P-based block vector BV in each M×N unit of the first block IntraTMP and includes The method according to claim 1.

3. The step of storing the IntraTM P-based block vector BV IntraTMP is as follows Store the IntraTM P-based block vector BV with a specified allowable accuracy IntraTMP therein The method according to claim 1, further comprising:

4. The step of storing the IntraTM P-based block vector BV IntraTMP is as follows Store the IntraTM P-based block vector BV with the precision indicated by the syntax information in the encoded video bitstream IntraTMP therein The method according to claim 1, further comprising:

5. The step of reconstructing the first block includes: Determine the reference template based on a plurality of template candidates within the reconstructed search area in the current picture, and a displacement in position between one of the plurality of template candidates and the current template is (i) a block vector (BV) of a third block coded in the IntraBC mode, or (ii) an IntraTMP-based block vector BV of the third block coded in the IntraTMP mode IntraTMP indicated by a vector that is The method according to claim 1, further comprising:

6. The step of storing the IntraTM P-based block vector BV IntraTMP is as follows Based on the fact that the template matching cost between the reference template of the prediction block and the current template of the first block is smaller than the threshold value, the IntraTMP-based block vector BV IntraTMP is stored The method according to claim 1, further comprising:

7. The step of storing the IntraTM P-based block vector BV IntraTMP further includes storing a template matching cost between the reference template of the predicted block and the current template of the first block. The method according to claim 1, further comprising:

8. The method according to claim 7, wherein the template matching cost is normalized based on the number of samples in the current template.

9. A method for video decoding performed by a video decoder, comprising: decoding prediction information of a first block in a current picture from an encoded video bitstream, the prediction information indicating that an Intra Block Copy (IBC) mode is applied to the first block; A step of constructing an IBC candidate list of the first block, the IBC candidate list including a first candidate based on an IntraTMP-based block vector BV of a second block coded in an Intra Template Matching (IntraTMP) mode, the second block being one of (i) a reconstructed block within the current picture and (ii) a temporal neighbor of the first block IntraTMP and a step of including a first candidate based on IntraTMP , the second block being one of (i) a reconstructed block within the current picture and (ii) a temporal neighbor of the first block reconstructing the first block based on the IBC candidate list; A method having the above steps.

10. the IntraTMP-based block vector BV IntraTMP executing template matching on the IntraTMP-based block vector BV IntraTMP and refining the IntraTMP-based block vector BV Determining the first candidate as the template-matched IntraTMP-based block vector BV IntraTMP and the step of determining the first candidate as such The method according to claim 9, further comprising:

11. The IBC candidate list of the first block includes a plurality of candidates, each of the plurality of candidates being (i) an IntraTMP-based block vector BV of a block coded in the IntraTMP mode IntraTMP , and (ii) a block vector (BV) of a block coded in the IBC mode, based on one of which the plurality of candidates includes the first candidate The method further comprises, for the plurality of candidates: for each of the plurality of candidates, determining a respective template matching cost based on a reference template of a reference block and a current template of the first block; reordering the plurality of candidates based on the determined template matching costs. comprising a step of performing template matching, wherein the step of reconstructing includes reconstructing the first block based on the plurality of rearranged candidates in the IBC candidate list, The method according to claim 9.

12. Determining each of the template matching costs, For each of the plurality of candidates that is an IntraTMP-based block vector BV IntraTMP apply a scaling factor to the template matching cost of each respective candidate The method according to claim 11, comprising:

13. The method according to claim 9, wherein the number of one or more candidates in the IBC candidate list that are IntraTMP-based block vectors BVsIntraTMP of each block coded in the IntraTMP mode is less than or equal to a threshold value, and the one or more candidates include the first candidate.

14. the number of the one or more candidates in the IBC candidate list being equal to the threshold value, The step of constructing the IBC candidate list, A new IntraTMP-based block vector BV of a third block, which is one of (i) a reconstructed block in the current picture and (ii) a temporal neighbor of the first block, that is not in the IBC candidate list IntraTMP For the new IntraTMP-based block vector BV IntraTMP In response to the template matching cost associated with the new IntraTMP-based block vector BV being less than the template matching cost associated with at least one of the one or more candidates, among the one or more candidates, the candidate having the largest template matching cost among the one or more candidates is replaced with the new IntraTMP-based block vector BV IntraTMP Replace with The method according to claim 13, comprising:

15. The step of constructing the IBC candidate list, adding one or more candidates from one or more blocks coded in the IBC mode to the IBC candidate list, The method according to claim 14, comprising:

16. The IntraTMP-based block vector BV of the second block IntraTMP is obtained from a BV history table storing one or more block vectors (BVs) or one or more block displacement vectors of at least one block previously coded in the current picture, according to the method of claim 9.

17. A method for video encoding performed by a video encoder, comprising: encoding a first block in a current picture based on a predicted block in a reconstructed search area in the current picture, wherein, in an Intra Template Matching Prediction (IntraTMP) mode, a reference template of the predicted block is matched to a current template of the first block; The step of storing the IntraTMP-based block vector BV of the first block IntraTMP wherein the IntraTMP-based block vector BV IntraTMP indicates a displacement in position between the current template of the first block and the reference template of the prediction block, and a step Encoding a second block coded in either Intra Block Copy (IntraBC) mode or IntraTMP mode based on the stored IntraTMP-based block vector BV IntraTMP and A method comprising:

18. A method for video encoding performed by a video encoder, comprising: A step of constructing an IBC candidate list for a first block in a current picture, wherein the first block is encoded in an intra-block copy (IBC) mode, and the IBC candidate list includes an IntraTM-based block vector BV of a second block encoded in an intra-template matching prediction (IntraTMP) mode IntraTMP including a first candidate based on IntraTMP , wherein the second block is one of (i) a reconstructed block in the current picture and (ii) a temporal neighbor of the first block, and the step encoding the first block based on the IBC candidate list; A method comprising:

19. one or more processors; one or more memories storing a computer program; comprising: wherein the computer program causes the one or more processors to execute the method according to any one of claims 1 to 14; An apparatus.

20. one or more processors; one or more memories storing a computer program; comprising: wherein the computer program causes the one or more processors to execute the method according to claim 17 or 18; An apparatus.

21. A computer program that causes a computer to execute the method according to any one of claims 1 to 16. **Claim 22** A computer program that causes a computer to execute the method according to claim 17 or 18.

Citation Information

Patent Citations

  • Image prediction / encoding device, image prediction / encoding method, image prediction / encoding program, image prediction / decoding device, image prediction / decoding method, and image prediction / decoding program

    JP2008283662A

  • Video encoding and decoding method, device, and computer program

    JP2022505996A

  • Template matching for JVET intra prediction

    US20180241993A1

  • Method and apparatus for block vector signaling and derivation in intra picture block compensation

    US20200014934A1

  • Method and apparatus for encoding and decoding motion information

    US20200053361A1