Block vector refinement for intra-template matching prediction at the sub-block level

By refining block vectors within defined search ranges for sub-blocks using intra-template matching prediction, the method addresses inefficiencies in video coding, enhancing prediction accuracy and compression efficiency.

JP2026514634APending Publication Date: 2026-05-13TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2024-04-24
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing video coding technologies face inefficiencies in intra-template matching prediction, particularly in determining optimal block vectors for sub-blocks, leading to suboptimal compression and quality degradation in video data transmission.

Method used

The proposed method involves refining block vectors within specific search ranges for sub-blocks based on intra-template matching prediction, using predefined offsets to define search regions and determining precise block vectors through cost analysis, allowing for improved prediction accuracy and compression efficiency.

Benefits of technology

This approach enhances video coding by improving prediction accuracy and compression efficiency, resulting in better quality and reduced data transmission requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514634000001_ABST
    Figure 2026514634000001_ABST
Patent Text Reader

Abstract

The video decoder includes a processing circuit. The processing circuit is further configured to receive coded information of the current block. The coded information indicates that the current block is coded in intra-template matching prediction (intraTMP) mode. The processing circuit is configured to determine a first set of candidate block vectors (BVs) within a first search range for a first subblock of the current block. The first search range is determined based on the BV of the current block. The multiple candidate BVs represent multiple candidate predicted subblocks for the first subblock of the current block. The processing circuit is configured to determine the precise BV of the first subblock from the multiple candidate BVs determined within the first search range based on the intraTMP mode. The processing circuit is configured to reconstruct the first subblock based on the precise BV of the first subblock.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001]

[0001] Related applications This application claims priority to U.S. Provisional Application No. 63 / 462,233, “Block Vector Refinement for Intra-Template Matching Prediction at the Subblock Level,” filed on 26 April 2023, which is incorporated herein by reference in its entirety.

[0002]

[0002] Technical field This disclosure generally describes aspects related to video coding. [Background technology]

[0003]

[0003] The background explanation provided herein is for the general purpose of presenting the circumstances of this disclosure. No work currently performed under the name of the inventor is explicitly or implicitly recognized as prior art to this disclosure, to the extent that such work is described in this background section or in any other manner in which it may not be eligible to be considered prior art at the time of filing.

[0004]

[0004] Image / video compression can help transmit image / video data over various devices, storage, and networks with minimal quality degradation. In some cases, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in the current picture from previously reconstructed pictures using motion compensation. Motion compensation can be represented by motion vectors (MV). [Overview of the Initiative]

[0005]

[0005] The embodiments of the present disclosure include bitstreams, methods, and apparatus for video coding / decoding. In some examples, the apparatus for video coding / decoding includes processing circuits.

[0006]

[0006] A part of the present disclosure provides a method for processing visual media data. In this method, a bitstream of visual media data is processed according to format rules. In one example, the bitstream includes coded information of the current block in the current picture. The coded information indicates that the current block is coded in intra template matching prediction (intraTMP) mode. The format rules specify that: The predicted block of the current block is predicted from a plurality of candidate predicted blocks defined within an initial search range, based on the fact that the current block is coded in intraTMP mode. The predicted block is referenced by the block vector (BV) of the current block. A first plurality of candidate BVs is determined within a first search range for a first subblock of the current block, and a second plurality of candidate BVs is determined within a second search range for a second subblock of the current block. The formatting rules specify that the initial search range, the first search range, and the second search range contain different search regions. The formatting rules specify that the first and second search ranges are determined based on the BV of the current block. The formatting rules specify that the first multiple candidate BVs of the first subblock indicate multiple candidate predicted subblocks for the first subblock. The formatting rules specify that the second multiple candidate BVs of the second subblock indicate multiple candidate predicted subblocks for the second subblock. The formatting rules specify that the refined BV of the first subblock is determined from the first multiple candidate BVs based on the intraTMP mode. The formatting rules specify that the refined BV of the second subblock is determined from the second multiple candidate BVs based on the intraTMP mode.The formatting rules stipulate that the first subblock is processed based on the precise BV of the first subblock, and the second subblock is processed based on the precise BV of the second subblock.

[0007]

[0007] In one example, the format rule specifies that a cost value is determined between the template of the first subblock and the templates of each of the multiple candidate prediction subblocks of the first subblock. The format rule specifies that one of the first multiple candidate BVs of the first subblock is determined as the precise BV of the first subblock corresponding to the minimum cost value among the cost values ​​between the templates of the multiple candidate prediction subblocks corresponding to the first multiple candidate BVs of the first subblock and the template of the first subblock.

[0008]

[0008] In one example, the format rule specifies that the BV of the current block is defined by a first coordinate component BVx and a second coordinate component BVy. The format rule specifies that the first search range is defined by the top-left coordinate (BVx-OffsetL1, BV1y-OffsetT1) and the bottom-right coordinate (BVx+OffsetR1, BVy+OffsetB1), and that OffsetL1, OffsetT1, OffestR1, and OffsetB1 are predefined constants. The format rule specifies that the second search range is defined by the top-left coordinate (BVx-OffsetL2, BVy-OffsetT2) and the bottom-right coordinate (BVx+OffsetR2, BVy+OffsetB2), and that OffsetL2, OffsetT2, OffestR2, and OffsetB2 are predefined constants. The formatting rules stipulate that OffsetL2, OffsetT2, OffsetR2, and OffsetB2 differ from at least one corresponding offset in the first search range.

[0009]

[0009] In one example, the boundary of the first search range lies inside the boundary of the initial search range.

[0010]

[0010] In one example, the boundary of the first search range extends beyond the boundary of the initial search range.

[0011]

[0011] According to another aspect of the present disclosure, a video encoding method is provided. In one example, based on intraTMP mode, a predicted block for the current block in the current picture is determined from a plurality of candidate predicted blocks defined within an initial search range. The predicted block is referenced by the BV of the current block. A first plurality of candidate BVs are determined within a first search range for a first subblock of the current block. The initial search range and the first search range include different search regions. The first search range is determined based on the BV of the current block. The first plurality of candidate BVs for the first subblock indicate a plurality of candidate predicted subblocks for the first subblock. Based on intraTMP mode, a precise BV for the first subblock is determined from the first plurality of candidate BVs. Based on the precise BV for the first subblock, the first subblock is encoded into a bitstream.

[0012]

[0012] In one example, a cost value is determined between the template of the first subblock and the templates of each of the multiple candidate prediction subblocks. One of the first multiple candidate BVs of the first subblock is determined as the precise BV of the first subblock, which corresponds to the minimum cost value among the cost values ​​between the templates of the multiple candidate prediction subblocks corresponding to the first multiple candidate BVs of the first subblock and the template of the first subblock.

[0013]

[0013] In one example, the BV of the current block is defined by a first coordinate component BVx and a second coordinate component BVy. The first search range is defined by the top-left coordinate (BVx-OffsetL1, BV1y-OffsetT1) and the bottom-right coordinate (BVx+OffsetR1, BVy+OffsetB1). OffsetL1, OffsetT1, OffsetR1, and OffsetB1 are predefined constants.

[0014]

[0014] In one example, a second set of candidate BVs is determined within a second search range for a second subblock of the current block. The second search range is determined based on the BV of the current block. Based on the intraTMP mode, the precise BV of the second subblock is determined from the second set of candidate BVs. The second search range differs from the first search range.

[0015]

[0015] According to yet another aspect of the present disclosure, a video decoding device is provided. The device includes a processing circuit. The processing circuit is configured to receive a bitstream containing coded information of the current block in the current picture. The coded information indicates that the current block is coded based on intraTMP mode, in which case the predicted block of the current block is determined based on a cost value between the template of the current block and the template of the predicted block, and the predicted block is referenced by the BV of the current block. The processing circuit is configured to determine a first plurality of candidate BVs within a first search range for a first subblock of the current block. The first search range is determined based on the BV of the current block. The first plurality of candidate BVs indicate a plurality of candidate predicted subblocks for the first subblock of the current block. The processing circuit is configured to determine the precise BV of the first subblock from the first plurality of candidate BVs based on intraTMP mode. The processing circuit is configured to reconfigure the first subblock based on the precise BV of the first subblock.

[0016]

[0016] In one example, the processing circuit is configured to determine the cost value between the template of the first subblock and the template of each of the multiple candidate prediction subblocks. The processing circuit is configured to determine one of the first multiple candidate BVs of the first subblock as the precise BV of the first subblock that corresponds to the minimum cost value among the cost values ​​between the templates of the multiple candidate prediction subblocks corresponding to the first multiple candidate BVs of the first subblock and the template of the first subblock.

[0017]

[0017] In one example, the BV of the current block is defined by a first coordinate component BVx and a second coordinate component BVy, and the first search range is defined by the upper left coordinate (BVx-OffsetL1, BV1y-OffsetT1) and the lower right coordinate (BVx+OffsetR1, BVy+OffsetB1). OffsetL1, OffsetT1, OffsetR1, and OffsetB1 are predefined constants.

[0018]

[0018] In one example, the processing circuit is configured to determine a second set of candidate BVs within a second search range for a second subblock of the current block, the second search range being determined based on the BV of the current block. The processing circuit is configured to determine a precise BV of the second subblock from the second set of candidate BVs based on the intraTMP mode, the second search range being different from the first search range.

[0019]

[0019] In one example, the second search range is defined by the top-left coordinate (BVx-OffsetL2, BVy-OffsetT2) and the bottom-right coordinate (BVx+OffsetR2, BVy+OffsetB2). OffsetL2, OffsetT2, OffsetR2, and OffsetB2 are predefined constants and differ from at least one corresponding offset of the first search range.

[0020]

[0020] In one example, the BV of the current block is determined from a plurality of candidate BVs of the current block defined in an initial search range according to the intraTMP mode. The boundary of the first search range extends beyond the boundary of the initial search range.

[0021]

[0021] In one example, the BV of the current block is defined in an initial search range according to the intraTMP mode, and the boundary of the first search range is inside the boundary of the initial search range.

[0022]

[0022] In one example, the resolution of the BV of the current block is at either the first integer pel or the first sub-pel. The resolution of the BV of the first sub-block is at either the second integer pel or the second sub-pel. The first integer pel includes any one of 1-pel, 2-pel, 4-pel, and 8-pel, and the first sub-pel includes any one of 1 / 2-pel, 1 / 4-pel, and 1 / 8-pel. The second integer pel includes any one of 1-pel, 2-pel, 4-pel, and 8-pel, and the second sub-pel includes any one of 1 / 2-pel, 1 / 4-pel, and 1 / 8-pel.

[0023]

[0023] In one example, the processing circuit is configured to determine a predicted sub-block of the first sub-block from a plurality of candidate predicted sub-blocks. The predicted sub-block of the first sub-block is indicated by a precise BV. The processing circuit is configured to determine a reconstructed sample of the first sub-block as (i) a sample of the predicted sub-block of the first sub-block or (ii) a filtered sample of the predicted sub-block filtered based on filter coefficients.

[0024]

[0024] In one example, the processing circuit is configured to determine the BV of another block in the current picture as the precise BV of the first sub-block. The first sub-block is the sub-block of the current block that is closest to the other block. The processing circuit is configured to determine a predicted block of the other block indicated by the determined "BV of the other block".

[0025]

[0025] In one example, the processing circuit is configured to determine the BV of another block in the current picture as a weighted combination of the precise BV of the first sub-block and the precise BV of the second sub-block. The processing circuit is configured to determine a predicted block of the other block indicated by the determined "BV of the other block".

[0026]

[0026] Aspects of the present disclosure also provide an apparatus for video encoding. The apparatus for video encoding includes a processing circuit configured to perform any of the methods for video encoding described.

[0027]

[0027] Aspects of the present disclosure also provide a method for video decoding. The method includes any of the methods performed by an apparatus for video decoding.

[0028]

[0028] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods for video decoding / encoding described.

Brief Description of the Drawings

[0029]

[0029] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Figure 1]

[0030] FIG. 1 is a schematic diagram of an example block diagram of a communication system (100). [Figure 2]

[0031] Figure 2 is a schematic diagram of an example of a decoder block diagram. [Figure 3]

[0032] Figure 3 is a schematic diagram of an example block diagram of an encoder. [Figure 4]

[0033] Figure 4 is a schematic diagram of intra-template matching prediction (IntraTMP) according to some aspects of this disclosure. [Figure 5]

[0034] Figure 5 is a schematic diagram of a subblock-based IntraTMP according to certain aspects of this disclosure. [Figure 6A]

[0035] Figures 6A-6E are schematic diagrams of coding block and sub-block templates for coding blocks. [Figure 6B]

[0035] Figures 6A-6E are schematic diagrams of coding block and subblock templates of coding blocks. [Figure 6C]

[0035] Figures 6A-6E are schematic diagrams of coding block and subblock templates of coding blocks. [Figure 6D]

[0035] Figures 6A-6E are schematic diagrams of coding block and subblock templates of coding blocks. [Figure 6E]

[0035] Figures 6A-6E are schematic diagrams of coding block and subblock templates of coding blocks. [Figure 7A]

[0036] Figures 7A-7E are schematic diagrams of subblock-based IntraTMP via the intermediate results of subblock prediction. [Figure 7B]

[0036] Figures 7A-7E are schematic diagrams of subblock-based IntraTMP via the intermediate results of subblock prediction. [Figure 7C]

[0036] Figures 7A-7E are schematic diagrams of subblock-based IntraTMP via the intermediate results of subblock prediction. [Figure 7D]

[0036] Figures 7A-7E are schematic diagrams of subblock-based IntraTMP via the intermediate results of subblock prediction. [Figure 7E]

[0036] Figures 7A-7E are schematic diagrams of subblock-based IntraTMP via the intermediate results of subblock prediction. [Figure 8A]

[0037] Figures 8A-8B are schematic diagrams of the first template of a subblock in a coding block. [Figure 8B]

[0037] Figures 8A-8B are schematic diagrams of the first template of a subblock of a coding block. [Figure 9A]

[0038] Figures 9A-9B are schematic diagrams of the second template of the subblock of the coding block. [Figure 9B]

[0038] Figures 9A-9B are schematic diagrams of the second template of the subblock of the coding block. [Figure 10A]

[0039] Figures 10A-10B are schematic diagrams of the third template of the subblock of the coding block. [Figure 10B]

[0039] Figures 10A-10B are schematic diagrams of the third template of the subblock of the coding block. [Figure 11A]

[0040] Figures 11A-11B are schematic diagrams of the partition shapes of subblocks within a coding block based on the Geometric Partition Mode (GPM). [Figure 11B]

[0040] Figures 11A-11B are schematic diagrams of the partition shapes of subblocks within a coding block based on the Geometric Partition Mode (GPM). [Figure 12]

[0041] Figure 12 is a flowchart illustrating the decryption process according to certain aspects of this disclosure. [Figure 13]

[0042] Figure 13 is a flowchart illustrating an overview of the encoding process according to certain aspects of this disclosure. [Figure 14]

[0043] Figure 14 is a schematic diagram of a computer system according to one embodiment. [Modes for carrying out the invention]

[0030]

[0044] Figure 1 shows a block diagram of a video processing system (100) in some examples. The video processing system is an example application of the disclosed subject matter, a video encoder and video decoder in a streaming environment. The disclosed subject matter can be equally applied to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, and storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).

[0031]

[0045] The video processing system (100) may include a video source (101), such as a digital camera, and includes a capture subsystem (113) capable of generating a stream of, for example, uncompressed video pictures (102). In one example, the video picture stream (102) includes a sample captured by the digital camera. The video picture stream (102), which is drawn as a thick line to emphasize the large amount of data compared to encoded video data (104) (or encoded video bitstream), can be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof, and can operate or realize aspects of the disclosed subject matter as described in detail below. The encoded video data (104) (or encoded video bitstream), which is drawn as a thin line to emphasize the smaller amount of data compared to the video picture (102) stream, can be stored in the streaming server (105) for future use. One or more streaming client subsystems, such as client subsystems (106) and (108) in Figure 1, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) may include a video decoder (110) within, for example, an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an output stream (111) of a video picture that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, encoded video data (104), (107), and (109) (e.g., video bitstream) can be encoded according to a specific video coding / compression standard.Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed information may be used in the context of VVC.

[0032]

[0046] It should be noted that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may include a video decoder (not shown), and electronic device (130) may include a video encoder (not shown).

[0033]

[0047] Figure 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used in place of the video decoder (110) in the example of Figure 1.

[0034]

[0048] The receiver (231) is capable of receiving one or more coded video sequences (e.g., those contained in a bitstream) to be decoded by the video decoder (210). In one embodiment, one coded video sequence is received at a time if the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences can be received from a channel (201), which may be a hardware / software link to a storage device that stores coded video data. The receiver (231) is capable of receiving coded video data together with other data, such as coded audio data and / or auxiliary data streams, which can be transmitted using their respective entities (not shown). The receiver (231) can isolate the coded video sequences from other data. To address network jitter, the buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In certain applications, the buffer memory (215) is part of the video decoder (210). In other cases, it may be outside the video decoder (210) (not shown). In yet another example, the buffer memory (not shown) may exist outside the video decoder (210), for example to address network jitter, and furthermore, another buffer memory (212) may exist inside the video decoder (210), for example to handle playback timing. If the receiver (231) is receiving data from a store / forward device with sufficient bandwidth and controllability, or from a synchronous network, the buffer memory (215) may not be necessary, or can be made smaller.For use in best-effort packet networks such as the Internet, buffer memory (215) may be required, which may be relatively large and, advantageously, can be of an adaptive size, and may be at least partially implemented in an operating system or similar element (not shown) outside of the video decoder (210).

[0035]

[0049] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the coded video sequence. These categories of symbols include information used to manage the operation of the video decoder (210), and potentially information for controlling rendering devices, such as a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but can be coupled to the electronic device (230), as shown in Figure 2. The rendering device control information may be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not shown). The parser (220) can parse / entropy decode the received coded video sequence. The coding of the video sequence to be coded may follow video coding techniques or standards, and may follow various principles, including variable-length coding, Huffman coding, and arithmetic coding with or without context influence. The parser (220) can extract from the coded video sequence a set of subgroup parameters for at least one subgroup of pixels in the video decoder, based on at least one parameter corresponding to a group. Subgroups can include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), prediction units (PU), etc. The parser (220) can also extract from coded video sequence information such as transformation coefficients, quantization parameter values, and motion vectors (MV).

[0036]

[0050] The parser (220) can perform entropy decoding / analysis on the video sequence received from the buffer memory (215) to generate symbols (221).

[0037]

[0051] The reconstruction of symbol (221) may include multiple different units depending on the type of coded video picture or part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is included can be controlled by subgroup control information parsed from the coded video sequence by parser (220). The flow of such subgroup control information between parser (220) and subsequent units is not depicted for clarity.

[0038]

[0052] The video decoder (210) can be conceptually subdivided beyond the functional blocks already described into several functional units, as described below. In actual implementations operating under commercial constraints, many of these units can interact closely with each other and be at least partially integrated. However, for the purpose of illustrating the subject matter being disclosed, the following conceptual subdivision into functional units is appropriate.

[0039]

[0053] The first unit is the scaler / inverse unit (251). The scaler / inverse unit (251) receives not only the quantized transformation coefficients but also control information (including the transformation to be used, block size, quantization factor, quantization scaling matrix, etc.) from the parser (220) as symbols (221). The scaler / inverse unit (251) can output a block containing sample values ​​that can be input to the aggregator (255).

[0040]

[0054] In some cases, the output samples of the scaler / inverse unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use predictive information from previously reconstructed pictures, but can use predictive information from previously reconstructed portions of the current picture. Such predictive information can be provided by the intra-picture predictive unit (252). In some cases, the intra-picture predictive unit (252) uses already reconstructed surrounding information taken from the current picture buffer (258) to generate blocks of the same size and shape as the block being reconstructed. The current picture buffer (258) buffers, for example, partially reconstructed current pictures and / or fully reconstructed current pictures. The aggregator (255) may, on a sample-by-sample basis, add the predictive information generated by the intra-predictive unit (252) to the output sample information provided by the scaler / inverse unit (251).

[0041]

[0055] Otherwise, the output samples of the scaler / inverse unit (251) can be associated with blocks that are intercoded and potentially motion-compensated. In such cases, the motion-compensated prediction unit (253) can access the reference picture memory (257) to retrieve samples to be used for prediction. After motion-compensating the retrieved samples according to the symbols (221) associated with the blocks, these samples are added by the aggregator (255) to the output of the scaler / inverse unit (251) (in this case, called residual samples or residual signals) to generate output sample information. The address in the reference picture memory (257) from which the motion-compensated prediction unit (253) retrieves prediction samples can be controlled by motion vectors available to the motion-compensated prediction unit (253), for example, in the form of symbols (221) which may have X, Y, and reference picture components. Furthermore, motion compensation may include interpolation of sample values, motion vector prediction mechanisms, etc., such as those retrieved from reference picture memory (257), when accurate motion vectors of sub-samples are used.

[0042]

[0056] The output samples of the aggregator (255) may be subject to various loop filtering techniques within the loop filter unit (256). The video compression technique may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) which are made available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also depend on metadata obtained during the decoding of earlier parts (in decoding order) of the coded picture or coded video sequence, as well as on previously reconstructed and loop-filtered sample values.

[0043]

[0057] The output of the loop filter unit (256) can be a sample stream that can be output to the rendering device (212) as well as stored in reference picture memory (227) for use in future inter-picture prediction.

[0044]

[0058] A given coded picture, once fully reconfigured, can be used as a reference picture for future predictions. For example, once the coded picture corresponding to the current picture is fully reconfigured and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a fresh current picture buffer can be reallocated before starting the reconfiguration of subsequent coded pictures.

[0045]

[0059] The video decoder (210) can perform decoding operations according to a standard such as ITU-T Rec.H.265 or a specified video compression technology. The coded video sequence can conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and a profile as documented in the video compression technology or standard. Specifically, a profile allows for the selection of a particular tool from all tools available in the video compression technology or standard, with the tool being the only tool that can be used under that profile. Also, for compliance, the complexity of the coded video sequence must fall within the range defined by the level of the video compression technology or standard. In some cases, that level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted by the Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0046]

[0060] In one embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a time, space, or signal-to-noise ratio (SNR) improvement layer, redundant slices, redundant pictures, forward error correction code, etc.

[0047]

[0061] Figure 3 shows an exemplary block diagram of a video encoder (303). The video encoder (303) is contained within an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmitting circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of Figure 1.

[0048]

[0062] The video encoder (303) can receive video samples from a video source (301) (not part of the electronic device (320) in the example of Figure 3) which is capable of capturing video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0049]

[0063] The video source (301) can provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream that can be of any appropriate bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.301 YCrCB, RGB, ...), and any appropriate sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (301) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a series of individual pictures that convey motion when viewed sequentially. The pictures themselves can be organized as a spatial array of pixels, and each pixel may contain one or more samples depending on the sampling structure, color space, etc., in use. The following description focuses on samples.

[0050]

[0064] In one embodiment, the video encoder (303) can encode and compress the pictures of a source video sequence into a coded video sequence (343) in real time or under any other required time constraints. Enforcing an appropriate coding speed is one function of the controller (350). In some embodiments, the controller (350) controls and is functionally coupled to other functional units, as described below. The coupling is not depicted for clarity. Parameters set by the controller (350) may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured to have other appropriate functions related to the video encoder (303) optimized for a particular system design.

[0051]

[0065] In some embodiments, the video encoder (303) is configured to operate in a coding loop. In an extremely simplified explanation, in one example, the coding loop may include a source coder (330) (responsible for generating symbols, such as a symbol stream, based on, for example, an input picture and a reference picture to be coded) and a (local) decoder (333) built into the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data in a similar manner to how the (remote) decoder also generates. The reconstructed sample stream (sample data) is input to the reference picture memory (334). Since the decoding of the symbol stream yields a bit-exact result that is independent of the decoder's location (local or remote), the contents of the reference picture memory (334) are also bit-exact between the local encoder and the remote encoder. In other words, the encoder's prediction unit "sees" the exact same sample values ​​as the reference picture samples that the decoder would "see" if it were using predictions during decoding. This fundamental principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example, due to channel errors) is also used in several related techniques.

[0052]

[0066] It is possible to assume that the operation of the “local” decoder (333) is the same as that of a “remote” decoder, such as the video decoder (210) which has already been described in detail above in relation to Figure 2. However, as can be briefly seen with reference to Figure 2, since symbols are available and the encoding / decoding of symbols to a coded video sequence by the entropy coder (343) and parser (520) is lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).

[0053]

[0067] In one embodiment, decoder techniques other than analysis / entropy decoding present in the decoder are ideally or substantially identical in function to those present in the corresponding encoder. Therefore, the disclosed subject matter focuses on the operation of the decoder. The description of encoder techniques can be omitted, as it is the inverse of the comprehensively described decoder techniques. More detailed descriptions are provided below in specific areas.

[0054]

[0068] During operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, predictively coding the input picture by referencing one or more previously coded pictures from a video sequence designated as “reference pictures”. In this way, the coding engine (332) codes the difference between the pixel blocks of the input picture and the pixel blocks of the reference pictures that may be selected as predictive references for the input picture.

[0055]

[0069] The local video decoder (333) can decode coded video data of a picture that can be designated as a reference picture based on symbols generated by the source coder (330). The operation of the coding engine (332) can, advantageously, be a non-lossless process. If the coded video data can be decoded by a video decoder (not shown in Figure 3), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (333) can repeat the decoding process that can be performed by the video decoder on the reference picture, causing the reconstructed reference picture to be stored in the reference picture cache (334). Thus, the video encoder (303) can locally store a copy of the reconstructed reference picture with common content (assuming no transmission errors) as the reconstructed reference picture to be obtained by the video decoder at the far end.

[0056]

[0070] The predictor (335) can perform predictive searches for the coding engine (332). That is, for a new picture to be coded, the predictor (335) can search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or given metadata (reference picture motion vectors, block shapes, etc.), which may serve as appropriate predictive references for the new picture. The predictor (335) can operate on a sample block-pixel block basis to find appropriate predictive references. In some cases, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory (334), as determined by the search results obtained by the predictor (335).

[0057]

[0071] The controller (350) can manage the coding operations of the source coder (330), including, for example, setting parameters and subgroup parameters used to encode video data.

[0058]

[0072] All outputs of the aforementioned functional units can be entropy coded in the entropy coder (345). The entropy coder (345) converts the symbols generated by the various functional units into coded video sequences by applying lossless compression to the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0059]

[0073] The transmitter (340) can buffer coded video sequences, such as those created by the entropy coder (345), in preparation for transmission over the communication channel (360), which may be a hardware / software link to a storage device that stores coded video data. The transmitter (340) can merge coded video data from the video coder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (source not shown).

[0060]

[0074] The controller (350) can manage the operation of the video encoder (303). During coding, the controller (350) can assign a specific coded picture type to each coded picture, which may affect the coding techniques that can be applied to each picture. For example, a picture may often be designated as one of the following picture types:

[0075] An intra-picture (I-picture) is one that can be encoded and decoded without using any other picture in the sequence as a source for prediction. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0061]

[0076] A predictive picture (P-picture) can be encoded and decoded using intra-prediction or inter-prediction, which utilizes motion vectors and reference indices, in order to predict the sample values ​​of each block.

[0062]

[0077] A bidirectional predictive picture (B-picture) can be encoded and decoded using intra-prediction or inter-prediction, employing two motion vectors and a reference index to predict the sample values ​​for each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0063]

[0078] A source picture can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each), and each block can be coded. Blocks can be predictively coded by referencing other (already coded) blocks, as determined by the coding assignment applied to each picture in the block. For example, blocks in picture I can be unpredictably coded, or they can be predictively coded by referencing already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks in picture P can be predictively coded by spatial or temporal prediction by referencing one previously coded reference picture. Blocks in picture B can be predictively coded by spatial or temporal prediction by referencing one or two previously coded reference pictures.

[0064]

[0079] The video encoder (303) can perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec.H.265. In this operation, the video encoder (303) can perform various compression operations, including predictive coding operations that leverage temporal and spatial redundancy in the input video sequence. The coded video data can therefore conform to the syntax specified by the video coding technique or standard being used.

[0065]

[0080] In one embodiment, the transmitter (340) may transmit additional data along with the encoded video. The source coder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, and other forms of redundant data (such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.).

[0066]

[0081] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) utilizes spatial correlations within a given picture, while inter-picture prediction utilizes (temporal or other) correlations between pictures. In one example, a particular picture under encoding / decoding, called the current picture, is partitioned into blocks. If a block in the current picture is similar to a reference block in a reference picture that has been previously coded and is still buffered in the video, then the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and, if multiple reference pictures are used, can have a third dimension to identify the reference pictures.

[0067]

[0082] In some embodiments, a dual-prediction technique can be used for inter-picture prediction. According to the dual-prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture, both of which precede the current picture in the video in decoding order (although they may be past and future in display order, respectively). Blocks in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Blocks can be predicted by combinations of the first and second reference blocks.

[0068]

[0083] Furthermore, merge mode technology can be used for inter-picture prediction to improve coding efficiency.

[0069]

[0084] According to certain aspects of this disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU contains three coding tree blocks (CTBs), which are one lumen CTB and two chroma CTBs. Each CTU can be recursively quad-tree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be partitioned into one 64x64 pixel CU, four 32x32 pixel CUs, or sixteen 16x16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as inter-prediction type or intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU contains a Luma prediction block (PB) and two Chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a Luma prediction block as an example of a prediction block, the prediction block contains a matrix of values ​​(e.g., Luma values) for pixels, such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0070]

[0085] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more processors that execute software instructions.

[0071]

[0086] Aspects of this disclosure provide techniques for improving predictions built by IntraTMP.

[0072]

[0087] Intra-Template Matching Prediction (also known as IntraTMP) is a special intra-prediction mode in which the best predicted block whose L-shaped template matches the current template is copied from the reconstructed portion of the current frame (current frame). Within a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the predicted block, where the most similar template is associated with the corresponding block and the current template is associated with the current block. The encoder then signals the use of intraTMP mode, allowing the decoder to perform the same prediction operation. The matching block (or corresponding block) (402) can be shown as in Figure 4 and functions as the matching area of ​​the current CU (404).

[0073]

[0088] As shown in Figure 4, the prediction signal can be generated by matching the L-shaped causal neighbors (L-shaped template) of the current block (404) with another block within a predefined search area. An example of a predefined search area may include R1 (current CTU), R2 (upper left CTU), R3 (upper CTU), and R4 (left CTU).

[0074]

[0089] Matching is performed based on a cost function. In one embodiment, the sum of absolute differences (SAD) is used as the cost function in IntraTMP mode. Within each search area, the decoder searches for the template (406) of block (402) that has the smallest SAD for the current template (408) of the current block (404), and the block with the smallest SAD can be used as the corresponding block for the current block. The corresponding block can further function as a prediction block for the current block (404).

[0075]

[0090] The dimensions of all search areas (e.g., SearchRange_w, SearchRange_h) may be set proportionally to the block dimensions of the current block (e.g., BlkW, BlkH). Thus, a fixed number of SAD comparisons may be obtained for each pixel. For example, the dimensions of the search area (or search range) may be defined by equations 1 and 2 as follows: SearchRange_w = a * BlkW Eq. (1) SearchRange_h = a * BlkH Eq. (2) Here, "a" is a constant that controls the trade-off between the complexity of the search process and the gain. In one example, "a" is equal to 5.

[0076]

[0091] In one embodiment, to speed up the template matching process, the search range of all search domains may be subsampled by a factor of 2 (by 1 / 2). The reduced search range may result in a four-fold reduction in the template matching search. After the best match is found, a further refinement process may be performed. Refinement may be performed by a second template matching search near the best match in the reduced range. The reduced range is, min(BlkW,BlkH) / 2 It may also be defined as follows.

[0077]

[0092] The intra-template matching tool may be enabled for CUs of a predetermined size, for example, those with a width and height of 64 or less. The maximum CU size for intra-template matching may be configurable.

[0078]

[0093] In one embodiment, if decoder-side intra-mode derivation (DIMD) is not used for the current CU, an intra-template matching prediction mode may be signaled at the CU level via a dedicated flag.

[0079]

[0094] The result of intraTMP for the current block may be a reference block. A reference block may be specified by a single block vector (BV) representing the location of the reference block related to the current coding block. There may be two or more areas with different textures within the current coding block. Finding a single BV to predict / match the current coding block may be difficult. However, if the current coding block is divided into multiple subblocks, prediction becomes easier, and it may be possible to predict multiple subblocks based on two or more BVs.

[0080]

[0095] Figure 5 shows an example of intraTMP based on subblocks, where two objects (504) and (506) are contained in the current block (502). The two objects may each have a good match within the search area (508). For example, within the search area (508), object (504) may have a good match (510) and object (506) may have a good match (512). However, if the distance between the two objects is large, it may prevent the encoder from finding a single good reference block. In contrast, if the current block (502) can be divided into two subblocks by a dashed line (514) in the middle of the current block (502), each subblock may find a better match and therefore a better prediction for the entire coding block (502).

[0081]

[0096] To address the problem described in Figure 5, the present disclosure offers various solutions. In a first solution, a subblock-level search refinement step may be applied to intraTMP. In one embodiment, the subblock-level search refinement step may be an additional search step to the existing two-step search approach of intraTMP.

[0082] In one example, during the first search phase of intraTMP, a search is performed within a predefined subsampled search area (e.g., 2x2 or 3x3 subsampling, where 2x2 indicates that the search area is downsampled by half both horizontally and vertically) to derive the best (or selected) BV (denoted as BV1, for example) for a block according to the template cost of intraTMP.

[0083] In the second search phase of intraTMP, it is possible to perform a search within a predefined search range (for example, defined as min(BlkW,BlkH) / 2, where BlkW is the block width and BlkH is the block height) near the best BV1 obtained in the first phase, and to obtain an improved BV (denoted as BV1).

[0084] In the third search phase of intraTMP, the proposed subblock-level refined search may be performed on a predefined group of subblocks of a block (e.g., 8x8 or 4x4) based on BV2 (e.g., near BV2) within a predefined search range. s This can be derived for each subblock of the block. BV s It may be the same as BV2, or it may be different from BV2. Furthermore, the prediction process is BV s The process may be performed based on the following: The prediction process can be performed by directly copying the reference block, or by an interpolation method or other prediction method based on sub-pelBV, to derive the final predicted signal for the block (or subblock).

[0085]

[0097] In one embodiment, the existing second search refinement stage of intraTMP may be replaced by a subblock search refinement stage. In one example, in the first search stage of intraTMP, a search may be performed within a predefined subsampling search range (e.g., 2x2 or 3x3 subsampling) to derive the best (or selected) BV of a block according to the intraTMP template cost, indicated as BV1. In the second search stage of intraTMP, a search may be performed at the subblock level for a predefined group of subblocks (e.g., 8x8 or 4x4) of the block, within a predefined search range centered on the best BV1 obtained in the first stage. Final block vector BV at the subblock level sThis may be derived for each subblock of the block. BV s This may be the same as BV1, or it may be different from BV1. Furthermore, BV s The prediction process may be performed based on this. The prediction process may be performed to derive the final predicted signal of the block (or subblock) by directly copying the reference block, by using an interpolation method based on sub-pel BV, or by using other prediction methods.

[0086]

[0098] In the multi-stage search approach described above, different search refinements may be applied. For example, integer precision may be applied to BV in the first search stage, and fractional BV precision may be applied in other search stages.

[0087]

[0099] Problems can arise when intraTMP prediction is performed at the subblock level. Figure 6A shows an L-shaped reference template (602) for the current block (604). The L-shaped reference template (602) may be located on the top and left sides of the current block (604). Samples within the L-shaped reference template area may be reconstructed samples available from previous coding blocks. When L-shaped reference samples are applied to subblocks of the current block (604), some subblocks may not have nearby samples available as templates.

[0088] Figure 6A shows an example where the current block (604) is divided into a 2x2 grid, resulting in four subblocks. The subblocks may be designated by numbers 0-3. For subblock 0, the L-shaped template for subblock 0 may still be available, as shown in Figure 6B, because the L-shaped template for subblock 0 can apply samples from the L-shaped reference template (602). However, for subblocks 1-3, some samples may be missing from the L-shaped templates for subblocks 1-3. For example, as shown in Figure 6C, the left portion of the L-shaped template is missing for subblock 1. As shown in Figure 6D, the upper portion of the L-shaped template may be missing for subblock 2. As shown in Figure 6E, both the upper and left portions of the L-shaped template may be missing for subblock 3.

[0089]

[0100] In the second solution, intermediate reconstructed results of a subblock may be used to predict other subblocks. Figures 7A-7E show an example of sequential intraTMP prediction. As shown in Figure 7A, the current block (704) can be partitioned into four subblocks 0-3. The current block (704) can have an L-shaped template (702). In Figure 7B, subblock 0 can be reconstructed based on intraTMP using the L-shaped template (702). For subblock 1, as shown in Figure 7C, subblock 1 may have a missing left-side template. The missing left-side template may be derived from the reconstructed sample of subblock 0 after subblock 0 has been predicted. For example, an inverse transform or other approach to subblock 0 may be required to derive a residual signal. Subblock 0 can then be reconstructed using the prediction and residual signals. Other subblocks can similarly reuse intermediate reconstructed samples from previous subblocks. For subblock 2, as shown in Figure 7D, the missing upper portion of the template for subblock 2 may be derived from the reconstructed samples of subblocks 0 and 1. For subblock 3, as shown in Figure 7E, the missing left and upper portions of the template for subblock 3 may be derived from subblocks 0, 1, and 2.

[0090]

[0101] It should be noted that the search range refinement approach in the first solution may differ from the approach in the second solution, which applies the overall intraTMP at the block level. In the latter case (or the second solution), each subblock explores the entire search range using a brute-force method, and as a result, each possible candidate reference subblock is examined. In contrast, in the approach proposed in the first solution, the subblock BV may be refined based on the best (or selected) BV applied to the entire block. In other words, refinement in the first solution leads to a better starting BV.

[0091]

[0102] In the third solution, it is possible to perform prediction of subblocks of a coding block using available template information from a previous coding block. However, the inverse transform or other approach to derive the residual signal is postponed until the entire coding block is predicted. For example, as shown in Figure 7C, for subblock 1, the template for subblock 1 may contain only samples of the L-shaped template (702) located above subblock 1. For subblock 2, as shown in Figure 7D, the template for subblock 2 may contain only samples of the L-shaped template (702) located to the left of subblock 2. For subblock 3, as shown in Figure 7E, the BVs of subblock 1 and / or subblock 2 may be reused to predict subblock 3 or the best block-level BV to be used without subblock-level refinement.

[0092]

[0103] After all four subblocks 0-3 have been predicted, an inverse transform (or other approach to derive the residual signal) may be performed at the block level. Based on the residual signal at the coding block level and the combined predicted signals from the subblock levels, a reconstructed sample of the entire block (e.g., (704)) may be obtained.

[0093]

[0104] In one aspect, intraTMP may be executed at the sub-block level.

[0094]

[0105] In one aspect, after the BV for the entire coding block is derived, the search refinement may be executed.

[0095]

[0106] In one example, the BV for the entire coding block may be labeled as BV1 having component BVs BV1x and BV1y. The BV for the entire coding block may be determined based on intraTMP. For example, the BV may be selected from among a plurality of candidate BVs defined within the search range based on a template cost value according to intraTMP. The BV for the entire coding block may be regarded as an input. For the sub-blocks of the coding block, further refined search may be executed based on the input BV1. The refined search may be executed within the search range. The search range may be (BV1x - offsetL, BV1y - offsetT) such as the upper left coordinates and (BV1x + offsetR, BV1y + offsetB) such as the lower right coordinates. offsetL, offsetT, offsetR, and offsetB may be predefined constants. Based on the refined search, an improved (or refined) BV for the assumed sub-block s can be obtained, where the template cost of the improved BV s may be reduced (or corresponding) to the minimum template cost within the assumed search range. The prediction signal may be obtained using the BV s refined at the sub-block level.

[0096]

[0107] In one embodiment, for each subblock, after the subblock (e.g., subblock 0 in Figure 7B) is predicted and the prediction is obtained based on intraTMP, the corresponding inverse transform or other alternative approach can be applied to derive the residual signal. For example, the residual signal may be the difference between the subblock and the prediction. The residuals derived at the subblock level and the obtained prediction may be used to generate a reconstructed sample of the subblock. The reconstructed sample of the subblock may be used as a new template for other right and / or lower subblocks (e.g., subblock 1 and / or subblock 2 in Figure 7D). Thus, as shown in Figures 6A-6E or 7A-7E, the right subblock (e.g., subblock 1) and / or lower subblock (e.g., subblock 3) may lack all or part of the L-shaped template due to block partitioning, and the right and / or lower subblocks may still be coded by intraTMP by using the reconstructed samples of neighboring subblocks. In other words, that subblock (for example, subblock 0 in Figure 7B) could simultaneously be a coding unit, a prediction unit, and a transformation unit.

[0097]

[0108] Figures 8A and 8B show examples of subblock templates based on reconstructed samples of nearby subblocks. In one example, as shown in Figure 8A, subblock 1 may contain an L-shaped template (804). The L-shaped template (804) may contain a reconstructed sample of a previously coded subblock 0 and a portion of the L-shaped template (802) of the coding block containing subblocks 0-1. In one example, the reconstructed sample of subblock 0 may be located on a first side (e.g., the left side) of subblock 1, and the portion of the L-shaped template (802) may be located on a second side (e.g., the top side) of subblock 1. In one example, as shown in Figure 8B, the L-shaped template (806) of subblock 1 may contain a portion of the reconstructed sample of subblock 0 and a portion of the L-shaped template (802). In one example, the width W1 of the portion of the reconstructed sample of subblock 0 and the width W2 of the portion of the L-shaped template (802) may be equal. In one example, the width W1 may be half the width of subblock 0.

[0098]

[0109] For example, the size of the L-shaped template may be adaptively changed to the subblock level. For instance, an L-shaped template for a subblock might only include samples on the left, top, and upper-left sides of the subblock, and not include samples on the upper-right and lower-left sides.

[0099]

[0110] In one embodiment, the template of a subblock may be adaptively selected based on a previous coding block, but may not reuse intermediate reconstructed samples of a previously coded subblock. Thus, adjacent reconstructed samples of the previous coding block may be used, for example, some or as many as possible, to derive the template of the subblock. For example, as shown in Figure 6C, subblock 1 may lack the left-side sample to have an L-shaped template, but the upper sample in the L-shaped template (602) is available. Thus, the upper sample of subblock 1 in the L-shaped template (602) may be selected as the template.

[0100] Similar rules may apply to subblock 2, in which case only the left-hand sample of subblock 2 may be used as a template. In the case of subblock 3, subblock 3 may perform prediction using the derived BV of subblock 1 and / or subblock 2. After all subblocks have been predicted, an inverse transform or other approach to derive the residual signal may be performed at the coding block level (not the subblock level). Based on the residual signal at the block level and the collective prediction signal at the subblock level, it is possible to derive the reconstructed sample of the coding block (e.g., (604)). Thus, the coding block (e.g., (604)) may be considered as a coding unit and a transform unit, rather than a prediction unit. Alternatively, the coding block may be a collection of prediction units at the subblock level.

[0101]

[0111] In one example, for a subblock that is missing a portion of the L-shaped template, only the sample directly above / to the left may be considered as part of the template. For example, as shown in Figure 9A, in the case of subblock 1, the upper left sample of subblock 1 within the original L-shaped template (902) may not be included as part of the template for subblock 1. Only the upper sample (904) within the original L-shaped template (902) that fits within the width of subblock 1 may be defined as the template for subblock 1.

[0102] In the case of subblock 2 in Figure 9B, the upper left sample of subblock 2 within the original L-shaped template (902) does not need to be included as the template for subblock 2. Only the left-hand sample within the original L-shaped template (902) that fits within the height of subblock 2 may be defined as the template for subblock 2.

[0103]

[0112] In one example, for a subblock of a coding block that is missing a portion of an L-shaped template, the left / upper sample of the subblock within the coding block template may be considered as the template for the subblock. For example, in the case of subblock 1 in Figure 10A, the template (1004) for subblock 1 may include all the upper and left samples located within the original L-shaped template (1002) of the coding block that includes subblocks 0-1. In the case of subblock 2 in Figure 10B, the template (1006) for subblock 2 may include all the left and upper samples located within the original L-shaped template (1002) of the coding block that includes subblocks 0-2.

[0104]

[0113] In one embodiment, syntax elements or other coded information, such as a flag, may be signaled after the intraTMP flag to specify whether a particular partition scheme is enabled for the intraTMP coding block under consideration (or the coding block considered for intraTMP based on the instructions of the intraTMP flag).

[0105]

[0114] For example, only one partition scheme may be permitted.

[0106]

[0115] For example, a group of partition schemes may be allowed. A combination of disable / enable syntax elements, such as flags, or other coded information may be used to specify which partition scheme from the group of partition schemes is applied.

[0107]

[0116] In one example, a group of partition schemes may be permitted. An index can be used to specify which partition scheme within the group is to be used. The index may be specified by a syntax element or other coded information.

[0108]

[0117] In one embodiment, the search range of a coding block may be adapted to a subblock. For example, for subblock 1 in Figure 6C, the maximum allowable sample distance to the left (or the left boundary of the search range) may be subtracted by the width of the nearby subblock 0, for example, 4 (or 4 samples), because subblock 1 may be shifted 4 samples to the right (or to the right of the L-shaped template (602)). For subblock 2 in Figure 6D, the maximum allowable sample distance to the top (or the top boundary of the search range) may be subtracted by 4 (or 4 samples), because subblock 2 may be shifted 4 samples down (or to the bottom of the L-shaped template (602)).

[0109]

[0118] In one embodiment, the allowable number of subblocks within a coding block may be N × M, where N is the number of horizontal divisions (or partitions) and M is the number of vertical divisions (or partitions). In one example, as shown in Figure 6A, N=2 and M=2. However, N and M can be any appropriate number.

[0110]

[0119] In one example, N=1 or M=1. In this case, the partition can be harmonized with intra-sub-partition (ISP) coding tools, such as ISP in VVC. Sub-block partitioning in intraTMP may also be considered as harmonization with ISP coding tools.

[0111]

[0120] For example, the number of subblocks may conform to the coding block size. For instance, a fixed subblock size may be defined as 8x8 or 4x4. In this case, larger coding blocks (e.g., 32x32 or larger) may have more subblocks.

[0112]

[0121] In one embodiment, the partition shape is not limited to a rectangular or square shape. For example, the subblock may be the partition result from another existing coding tool, such as Geometric Partition Mode (GPM), where exemplary partitions of the coding block may be those shown in Figures 11A and 11B. Furthermore, the GPM is not limited to coding blocks with inter-prediction. The coding block may be coded inter-intra or intra-intra in combination.

[0113] For example, as shown in Figure 11A, block (1102) may be partitioned into partition P0 and partition P1 by partition line (1104). Partition P0 can be coded by interpretation or intrapretation. Partition P1 can be coded by interpretation or intrapretation.

[0114] As shown in Figure 11B, block (1106) may be partitioned into partition P0 and partition P1 by partition line (1108). Partition P0 can be coded by interpretation or intrapretation. Partition P1 can be coded by interpretation or intrapretation.

[0115]

[0122] In one embodiment, the subblocks do not have to be divided (or sized) equally. The number of samples in each subblock may be different.

[0116]

[0123] In one embodiment, each subblock may use multiple candidate reference blocks. The final predictor of each subblock may be derived, for example, by using a weighted average of the candidate blocks (also known as the fusion method). For example, as shown in Figure 7B, subblock 0 may have multiple candidate blocks (or candidate prediction blocks) based on intraTMP within the search range. The predictor of subblock 0 can be determined based on a weighted average of the multiple candidate prediction blocks.

[0117]

[0124] In one embodiment, the refinement search range may be set independently for each subblock of a block. For example, the values ​​of offsetL, offsetT, offsetR, and offsetB may be different for different subblocks. For instance, the values ​​of offsetL, offsetT, offsetR, and offsetB may be different for subblocks 0-3 within the current block (604).

[0118]

[0125] In one embodiment, the values ​​of offsetL, offsetT, offsetR, and offsetB may be set independently. Therefore, each of offsetL, offsetT, offsetR, and offsetB may be a value that has been set independently in advance.

[0119]

[0126] In one embodiment, the refined search range of a subblock defined by the top-left coordinate (BV1x-offsetL, BV1y-offsetT) and the bottom-right coordinate (BV1x+offsetR, BV1y+offsetB) may exceed the original (or initial) search range at the coding block level. For example, the search range for determining BV1 of a coding block based on intraTMP may be inside the refined search range for the subblocks of the coding block.

[0120]

[0127] In one embodiment, the refined search range of a subblock defined by the top-left coordinate (BV1x-offsetL, BV1y-offsetT) and the bottom-right coordinate (BV1x+offsetR, BV1y+offsetB) may be restricted to be inside the original (or initial) search range at the coding block level. For example, the search range for determining BV1 of a coding block based on intraTMP may include the refined search range for the subblocks of the coding block.

[0121]

[0128] In one embodiment, the input BV (e.g., BV1) in the coding block and / or the output BV (e.g., BV) at the subblock levels The precision of ) may be an integer pel or sub-pel (for example, half-pel or quarter-pel).

[0122]

[0129] In one embodiment, the prediction of a block or subblock may be obtained by a pixel copy operation or a filtering operation. In a filtering operation, a reference block may be set as input, and filter coefficients may be applied to the reference block. The filter coefficients may be trained or predefined based on the subblock template and the corresponding reference template. In one example, when the BV of a subblock is refined within the search range based on intraTMP, the refined BV (e.g., BV s Based on this, a reference subblock is obtained for the subblock. Predictions for the subblock may be obtained by copying samples of the reference subblock or by using filtered samples of the reference subblock. Filtered samples of the reference subblock may be obtained by applying a filter coefficient to the samples of the reference subblock.

[0123]

[0130] In one embodiment, the proposed subblock-level BV refinement described above can be added to an existing BV search process in intraTMP at the block level. Thus, the BV of a block may be determined at the block level within the search range according to intraTMP. The block vector for a subblock of a block may be refined at the subblock level based on the search range of each subblock. The search range of a subblock may be determined based on the BV of the block. For example, the BV of a block may be defined as the starting BV of a refinement search for a first subblock, and the refinement search range may be defined around the BV of the block.

[0124]

[0131] In one embodiment, the proposed subblock-level BV refinement described above may replace refinement search after the coding block BV has been derived, based on a subsampled search area in intraTMP.

[0125]

[0132] In one embodiment, various BV precisions (e.g., 1-pel, 2-pel, 4-pel, 8-pel, 1 / 2-pel, 1 / 4-pel, and 1 / 8-pel) may be applied to various BV refinement search stages. For example, the refinement search for a block's BV may use a first BV precision (e.g., integer pel), while the refinement search for a subblock's BV at the subblock level may use a second BV precision (e.g., sub-pel).

[0126]

[0133] In one embodiment, each refined BV of each subblock of the current block or each of the selected subblocks of the current block may be used as a BV predictor for a subsequent coding block to be coded after the current block, using intraTMP and / or intra-block copy (IBC) modes. In one example, the refined BV of a subblock of the current block may be inherited as the BV of a subsequent coding block. A position-based selection may be applied to define the subblock. For example, the subblock of the current block closest to the subsequent coding block may be selected. The refined BV of the selected subblock may be determined as the inherited BV of the subsequent coding block. In one example, the refined block vectors of multiple subblocks of the current block may be used to determine the BV of a subsequent coding block. The prediction for the subsequent coding block may be a weighted combination (linear combination) of multiple reference blocks indicated by the refined block vectors of multiple subblocks.

[0127]

[0134] Figure 12 shows a flowchart illustrating an overview of process (1200) according to one embodiment of the present disclosure. Process (1200) can be used in a video decoder. In various embodiments, process (1200) is performed by processing circuits such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), and so on. In some embodiments, process (1200) is performed by software instructions, and so the processing circuit performs process (1200) when it executes a software instruction. The process begins at (S1201) and proceeds to (S1210).

[0128]

[0135] In (S1210), a bitstream containing coded information for the current block in the current picture is received. The coded information indicates that the current block is coded based on intraTMP mode, in which case the predicted block of the current block is determined based on the cost value between the template of the current block and the template of the predicted block, and the predicted block is referenced by the BV of the current block.

[0129]

[0136] In (S1220), a first set of candidate BVs is determined within a first search range for the first subblock of the current block. The first search range is determined based on the BV of the current block. The first set of candidate BVs represent a set of candidate predicted subblocks for the first subblock of the current block.

[0130]

[0137] In (S1230), based on the intraTMP mode, the precise BV of the first subblock is determined from a first group of candidate BVs.

[0131]

[0138] In (S1240), the first subblock is reconstructed based on the precise BV of the first subblock.

[0132]

[0139] In one example, a cost value is determined between the template of the first subblock and the templates of each of the multiple candidate prediction subblocks. One of the first multiple candidate BVs of the first subblock is determined as the precise BV of the first subblock, corresponding to the minimum cost value among the cost values ​​between the templates of the multiple candidate prediction subblocks corresponding to the first multiple candidate BVs of the first subblock and the template of the first subblock.

[0133]

[0140] In one example, the BV of the current block is defined by a first coordinate component BVx and a second coordinate component BVy, and the first search range is defined by the top-left coordinate (BVx-OffsetL1, BV1y-OffsetT1) and the bottom-right coordinate (BVx+OffsetR1, BVy+OffsetB1). OffsetL1, OffsetT1, OffsetR1, and OffsetB1 are predefined constants.

[0134]

[0141] In one example, a second set of candidate BVs is determined within a second search range for a second subblock of the current block, and the second search range is determined based on the BV of the current block. Based on the intraTMP mode, the precise BV of the second subblock is determined from the second set of candidate BVs, and the second search range differs from the first search range.

[0135]

[0142] In one example, the second search range is defined by the top-left coordinate (BVx-OffsetL2, BVy-OffsetT2) and the bottom-right coordinate (BVx+OffsetR2, BVy+OffsetB2). OffsetL2, OffsetT2, OffsetR2, and OffsetB2 are predefined constants and differ from at least one corresponding offset of the first search range.

[0136]

[0143] In one example, the BV of the current block is determined from multiple candidate BVs of the current block defined in the initial search range according to the intraTMP mode. The boundary of the first search range extends beyond the boundary of the initial search range.

[0137]

[0144] In one example, the BV of the current block is defined in the initial search range according to the intraTMP mode, and the boundary of the first search range lies inside the boundary of the initial search range.

[0138]

[0145] In one example, the precision of the BV of the current block is in either a first integer pel or a first sub-pel. The precision of the BV of the first sub-block is in either a second integer pel or a second sub-pel. The first integer pel includes any of 1-pel, 2-pel, 4-pel, and 8-pel, and the first sub-pel includes any of 1 / 2-pel, 1 / 4-pel, and 1 / 8-pel. The second integer pel includes any of 1-pel, 2-pel, 4-pel, and 8-pel, and the second sub-pel includes any of 1 / 2-pel, 1 / 4-pel, and 1 / 8-pel.

[0139]

[0146] In one example, a predicted subblock of the first subblock is determined from multiple candidate predicted subblocks. The predicted subblock of the first subblock is represented by a precise BV. A reconstructed sample of the first subblock is determined either (i) as a sample of the predicted subblock of the first subblock, or (ii) as a filtered sample of the predicted subblock filtered based on a filter coefficient.

[0140]

[0147] In one example, the precise BV of the first subblock is determined to be the BV of another block in the current picture. The first subblock is the subblock of the current block that is closest to the other block. The predicted block of the other block, indicated by the determined "BV of the other block," is then determined.

[0141]

[0148] In one example, the BV of another block in the current picture is determined as a weighted combination of the precise BV of the first subblock and the precise BV of the second subblock. The predicted block of the other block indicated by the determined "BV of another block" is then determined.

[0142]

[0149] Next, the process proceeds to (S1299) and terminates.

[0143]

[0150] Process (1200) can be appropriately adapted. Steps in process (1200) can be modified and / or omitted. Additional steps can be added. Any appropriate execution order can be used.

[0144]

[0151] Figure 13 shows a flowchart illustrating an overview of process (1300) according to one embodiment of the present disclosure. Process (1300) can be used in a video encoder. In various embodiments, process (1300) is performed by processing circuits such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), and so on. In some embodiments, process (1300) is performed by software instructions, and so when a processing circuit executes a software instruction, the processing circuit executes process (1300). The process begins at (S1301) and proceeds to (S1310).

[0145]

[0152] In (S1310), based on the intraTMP mode, the predicted block for the current block in the current picture is determined from multiple candidate predicted blocks defined within the initial search range. The predicted block is referenced by the BV of the current block.

[0146]

[0153] In (S1320), a first set of candidate BVs is determined within a first search range for the first subblock of the current block. The initial search range and the first search range include different search regions. The first search range is determined based on the BV of the current block. The first set of candidate BVs for the first subblock represent a set of candidate predicted subblocks for the first subblock.

[0147]

[0154] In (S1330), based on the intraTMP mode, the precise BV of the first subblock is determined from a first group of candidate BVs.

[0148]

[0155] In (S1340), the first subblock is encoded into a bitstream based on the precise BV of the first subblock.

[0149]

[0156] In one example, a cost value is determined between the template of the first subblock and the templates of each of the multiple candidate prediction subblocks. One of the first multiple candidate BVs of the first subblock is determined as the precise BV of the first subblock, corresponding to the minimum cost value among the cost values ​​between the templates of the multiple candidate prediction subblocks corresponding to the first multiple candidate BVs of the first subblock and the template of the first subblock.

[0150]

[0157] In one example, the BV of the current block is defined by a first coordinate component BVx and a second coordinate component BVy. The first search range is defined by the top-left coordinate (BVx - OffsetL1, BV1y - OffsetT1) and the bottom-right coordinate (BVx + OffsetR1, BVy + OffsetB1). OffsetL1, OffsetT1, OffsetR1, and OffsetB1 are predefined constants.

[0151]

[0158] In one example, a second set of candidate BVs is determined within a second search range for a second subblock of the current block. The second search range is determined based on the BV of the current block. Based on the intraTMP mode, the precise BV of the second subblock is determined from the second set of candidate BVs. The second search range differs from the first search range.

[0152]

[0159] Next, the process proceeds to (S1399) and terminates.

[0153]

[0160] Process (1300) can be appropriately adapted. Steps in process (1200) can be modified and / or omitted. Additional steps can be added. Any appropriate execution order can be used.

[0154]

[0161] This disclosure provides a method for processing visual media data. In this method, a bitstream of visual media data is processed according to format rules. In one example, the bitstream includes coded information of the current block in the current picture. The coded information indicates that the current block is coded in intraTMP mode. The format rules specify that: the predicted block of the current block is predicted from a plurality of candidate predicted blocks defined within an initial search range, based on the current block being coded in intraTMP mode. The predicted block is referenced by the block vector (BV) of the current block. A first plurality of candidate BVs is determined within a first search range for a first subblock of the current block, and a second plurality of candidate BVs is determined within a second search range for a second subblock of the current block. The format rules specify that the initial search range, the first search range, and the second search range contain different search regions. The formatting rules specify that the first and second search ranges are determined based on the BV of the current block. The formatting rules specify that the first multiple candidate BVs of the first subblock indicate multiple candidate predicted subblocks for the first subblock. The formatting rules specify that the second multiple candidate BVs of the second subblock indicate multiple candidate predicted subblocks for the second subblock. The formatting rules specify that the precise BV of the first subblock is determined from the first multiple candidate BVs based on the intraTMP mode. The formatting rules specify that the precise BV of the second subblock is determined from the second multiple candidate BVs based on the intraTMP mode. The formatting rules specify that the first subblock is processed based on the precise BV of the first subblock, and the second subblock is processed based on the precise BV of the second subblock.

[0155]

[0162] The technologies described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 14 shows a computer system (1400) suitable for realizing a particular embodiment of the disclosed subject matter.

[0156]

[0163] Computer software can be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code containing instructions, which may be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or may be executed via interpretation, microcode execution, etc.

[0157]

[0164] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.

[0158]

[0165] The components shown in Figure 14 with respect to the computer system (1400) are illustrative and are not intended to imply any limitations on the scope or functionality of the computer software that implements the aspects of this disclosure. Furthermore, the configuration of the components should not be construed as having any dependency or condition on any one or combination of the components shown in the exemplary aspects of the computer system (1400).

[0159]

[0166] The computer system (1400) may include certain human interface input devices. Such human interface input devices may respond to input from one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), auditory input (e.g., voice, applause), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., conversations, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), or video (e.g., 2D video, 3D video including stereoscopic pictures).

[0160]

[0167] Input human interface devices may include one or more of the following (although only one of each is depicted): keyboard (1401), mouse (1402), trackpad (1403), touchscreen (1410), data glove (not shown), joystick (1405), microphone (1406), scanner (1407), and camera (1408).

[0161]

[0168] The computer system (1400) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via touch screens (1410), data gloves (not shown), joysticks (1405), although there may also be haptic feedback devices that do not function as input devices), auditory output devices (e.g., speakers (1409), headphones (not shown)), visual output devices (e.g., screens (1410) including CRT screens, LCD screens, plasma screens, OLED screens, each having or not having touch screen input functionality, each having or not having haptic feedback functionality, some of which may be capable of outputting two-dimensional visual output, three-dimensional or more output by means such as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0162]

[0169] The computer system (1400) may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (1120) using media such as CD / DVD (1421), thumb drives (1422), removable hard drives or solid-state drives (1423), legacy magnetic media such as tapes and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0163]

[0170] Those skilled in the art will also understand that the term “computer-readable medium” as used in relation to the subject matter disclosed herein does not include a transmission medium, carrier wave, or other transient signal.

[0164]

[0171] A computer system (1400) may also include an interface (1454) to one or more communication networks (1455). The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, automotive industry, real-time, latency-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide-area digital networks for television (including cable TV, satellite TV, and terrestrial TV), and the automotive industry, including CANBus. Certain networks generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (1449) (e.g., a USB port on a computer system (1400)); others are generally integrated into the core of the computer system (1400) by being attached to a system bus, as described below (e.g., an Ethernet interface is integrated into a PC computer system, and a cellular network interface is integrated into a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast television), one-way transmit-only (e.g., CANbus to a specific CANbus device), or two-way, such as to other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks can be used for each of these networks and network interfaces, as described above.

[0165]

[0172] The aforementioned human interface devices, human-accessible storage devices, and network interfaces can be mounted on the core (1440) of the computer system (1400).

[0166]

[0173] The core (1440) may include one or more central processing units (CPUs) (1441), graphics processing units (GPUs) (1442), special programmable processing units in the form of field-programmable gate areas (FPGAs) (1443), hardware accelerators for specific tasks (1444), graphics adapters (1450), etc. These devices, along with read-only memory (ROM) (1445), random-access memory (1446), and internal mass storage devices (e.g., internal non-user-accessible hard drives, SSDs, etc.) (1447), may be connected via a system bus (1448). In some computer systems, the system bus (1448) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus (1448) or via a peripheral bus (1449). In one example, a screen (1410) can be connected to a graphics adapter (1450). The peripheral bus architecture includes PCI, USB, etc.

[0167]

[0174] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can be combined to execute specific instructions that constitute the aforementioned computer code. The computer code can be stored in ROM (1445) or RAM (1446). Temporary data can be stored in RAM (1446), while persistent data can be stored, for example, in internal mass storage (1447). High-speed storage and retrieval of any memory device may be possible by utilizing cache memory, which can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage (1447), ROMs (1445), RAM (1446), etc.

[0168]

[0175] Computer-readable media can have computer code thereon for performing various computer-implemented operations. The media and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be well known and available to those skilled in the art in the field of computer software.

[0169]

[0176] As an example, and not limited to, a computer system having an architecture (1400), specifically a core (1440), can provide functionality as a result of the operation of a processor (including CPUs, GPUs, FPGAs, accelerators, etc.) that runs software embodied in one or more tangible computer-readable media. Such computer-readable media can be media related to user-accessible mass storage as described above, as well as specific storage of the core (1440) of a non-transient nature, such as mass storage (1447) or ROM (1445) within the core. Software that implements various embodiments of the present disclosure can be stored in such devices and executed by the core (1440). The computer-readable media can include one or more memory devices or chips, depending on the specific needs. The software can cause the core (1440) and specifically the processor (including CPUs, GPUs, FPGAs, etc.) within it to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (1446) and modifying such data structures according to processes defined by the software. As an addition or alternative, a computer system may provide functionality as a result of logic wired or otherwise incorporated within a circuit (e.g., an accelerator (1444)) which may perform a particular process or a particular part of a particular process as described herein, either in place of or in conjunction with software. References to software may include logic, and vice versa, as appropriate. References to computer-readable media may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. This disclosure encompasses any appropriate combination of hardware and software.

[0170]

[0177] The usage of “at least one of” or “one of” in this disclosure is intended to include any one or combination of the elements listed. For example, “at least one of A, B, or C”; “at least one of A, B, and C”; “at least one of A through C” is intended to include A only, B only, C only, or any combination thereof. The usage of “one of A or B” or “one of A and B” is intended to include A or B or (A and B). Where applicable, the usage of “one of” does not exclude any combination of the elements mentioned, for example, if the elements are not mutually exclusive.

[0171]

[0178] While this disclosure describes several exemplary embodiments, there are many variations, substitutions, and alternative equivalents that fall within the scope of this disclosure. Therefore, it will be understood that many systems and methods, not expressly illustrated or described herein, that embody the principles of this disclosure and thus fall within the spirit and scope of this disclosure, can be devised by those skilled in the art.

Claims

1. A method for processing visual media data: The process includes the step of processing the bitstream of the visual media data according to format rules, The bitstream includes coded information for the current block in the current picture, the coded information indicating that the current block is coded in intra-template matching prediction (intraTMP) mode; The aforementioned formatting rules are: The predicted block of the current block is predicted from a plurality of candidate predicted blocks defined within an initial search range based on the current block coded in intraTMP mode, and the predicted block is referenced by the block vector (BV) of the current block; The first set of candidate BVs is determined within a first search range for a first subblock of the current block, and the second set of candidate BVs is determined within a second search range for a second subblock of the current block; The initial search range, the first search range, and the second search range include different search regions; The first and second search ranges are determined based on the BV of the current block; The first multiple candidate BVs of the first subblock represent multiple candidate predicted subblocks for the first subblock; The second multiple candidate BVs of the second subblock indicate multiple candidate predicted subblocks for the second subblock; The precise BV of the first subblock is determined from the first group of candidate BVs based on the intraTMP mode; The precise BV of the second subblock is determined from the second group of candidate BVs based on the intraTMP mode; The first subblock is processed based on the precise BV of the first subblock; and The second subblock is processed based on the precise BV of the second subblock; A method that defines something.

2. In the method according to claim 1, the format rule is: The cost value is determined between the template of the first subblock and the templates of each of the multiple candidate prediction subblocks of the first subblock; and One of the first multiple candidate BVs of the first subblock is determined as the precise BV of the first subblock corresponding to the minimum cost value among the cost values ​​between the template of the multiple candidate prediction subblocks corresponding to the first multiple candidate BVs of the first subblock and the template of the first subblock; A method that defines something.

3. In the method according to claim 1, the format rule is: The BV of the current block is defined by a first component BVx and a second component BVy; The first search range is defined by the top-left coordinate (BVx-OffsetL1, BV1y-OffsetT1) and the bottom-right coordinate (BVx+OffsetR1, BVy+OffsetB1), where OffsetL1, OffsetT1, OffsetR1, and OffsetB1 are predefined constants; The second search range is defined by the top-left coordinate (BVx-OffsetL2, BVy-OffsetT2) and the bottom-right coordinate (BVx+OffsetR2, BVy+OffsetB2), where OffsetL2, OffsetT2, OffsetR2, and OffsetB2 are predefined constants; OffsetL2, OffsetT2, OffsetR2, and OffsetB2 differ from at least one corresponding offset of the first search range; A method that defines something.

4. In the method according to any one of claims 1 to 3: A method wherein the boundary of the first search range does not exceed or exceeds the boundary of the initial search range.

5. A video encoding method: A step of determining the predicted block for the current block in the current picture from a plurality of candidate predicted blocks defined within an initial search range, based on an intra-template matching prediction (intraTMP) mode, wherein the predicted block is referenced by the block vector (BV) of the current block; A step of determining a first plurality of candidate BVs within a first search range for a first subblock of the current block, wherein the initial search range and the first search range include different search regions, the first search range is determined based on the BV of the current block, and the first plurality of candidate BVs of the first subblock represent a plurality of candidate predicted subblocks for the first subblock; A step of determining the precise BV of the first subblock from the first plurality of candidate BVs based on the intraTMP mode; and A step of encoding the first subblock into a bitstream based on the precise BV of the first subblock; A method that includes this.

6. In the method of claim 5, the step of determining the precise BV further: A step of determining the cost value between the template of the first subblock and the template of each of the plurality of candidate prediction subblocks; and A step of determining one of the first multiple candidate BVs of the first subblock as the precise BV of the first subblock corresponding to the minimum cost value among the cost values ​​between the template of the candidate prediction subblock corresponding to the first multiple candidate BV of the first subblock and the template of the first subblock; A method that includes this.

7. In the method according to claim 5 or 6: The BV of the current block is defined by a first component BVx and a second component BVy; The first search range is defined by the top-left coordinate (BVx-OffsetL1, BV1y-OffsetT1) and the bottom-right coordinate (BVx+OffsetR1, BVy+OffsetB1), where OffsetL1, OffsetT1, OffsetR1, and OffsetB1 are predefined constants.

8. A video decoding device comprising a processing circuit, wherein the processing circuit is: Steps include receiving a bitstream containing coded information for the current block in the current picture, wherein the coded information indicates that the current block is coded based on intra-template matching prediction (intraTMP) mode, in which the predicted block of the current block is determined based on a cost value between the template of the current block and the template of the predicted block, and the predicted block is referenced by the block vector (BV) of the current block; A step of determining a first group of candidate BVs within a first search range for a first subblock of the current block, wherein the first search range is determined based on the BV of the current block, and the first group of candidate BVs represent a group of candidate predicted subblocks for the first subblock of the current block; A step of determining the precise BV of the first subblock from the first plurality of candidate BVs based on the intraTMP mode; and A step of reconstructing the first subblock based on the precise BV of the first subblock; A device configured to perform the following actions.

9. In the apparatus according to claim 8, the processing circuit is: A step of determining the cost value between the template of the first subblock and the template of each of the plurality of candidate prediction subblocks; and A step of determining one of the first multiple candidate BVs of the first subblock as the precise BV of the first subblock corresponding to the minimum cost value among the cost values ​​between the template of the candidate prediction subblock corresponding to the first multiple candidate BV of the first subblock and the template of the first subblock; A device configured to perform the following actions.

10. In the apparatus according to claim 8, the processing circuit is: A step of determining a second set of candidate BVs within a second search range for a second subblock of the current block, wherein the second search range is determined based on the BV of the current block; and A step of determining the precise BV of the second subblock from the second plurality of candidate BVs based on the intraTMP mode, wherein the second search range is different from the first search range; A device configured to perform the following actions.

11. In the apparatus according to claim 10: The BV of the current block is defined by a first component BVx and a second component BVy; The first search range is defined by the top-left coordinate (BVx-OffsetL1, BV1y-OffsetT1) and the bottom-right coordinate (BVx+OffsetR1, BVy+OffsetB1), where OffsetL1, OffsetT1, OffsetR1, and OffsetB1 are predefined constants; and The apparatus wherein the second search range is defined by the top-left coordinate (BVx-OffsetL2, BVy-OffsetT2) and the bottom-right coordinate (BVx+OffsetR2, BVy+OffsetB2), where OffsetL2, OffsetT2, OffsetR2, and OffsetB2 are predefined constants and differ from at least one corresponding offset of the first search range.

12. In the apparatus according to any one of claims 8 to 11: The BV of the current block is determined from a plurality of candidate BVs of the current block defined in the initial search range according to the intraTMP mode; and The device wherein the boundary of the first search range exceeds the boundary of the initial search range, or does not exceed the boundary of the initial search range.

13. In the apparatus according to any one of claims 8 to 11: The precision of the BV of the current block is in either a first integer pel or a first sub-pel; The precision of the BV of the first subblock is in either the second integer pel or the second sub-pel; The first integer pel includes any of 1-pel, 2-pel, 4-pel, and 8-pel, and the first sub-pel includes any of 1 / 2-pel, 1 / 4-pel, and 1 / 8-pel; and An apparatus in which the second integer pel includes any of 1-pel, 2-pel, 4-pel, and 8-pel, and the second sub-pel includes any of 1 / 2-pel, 1 / 4-pel, and 1 / 8-pel.

14. In the apparatus according to any one of claims 8 to 11, the processing circuit is: A step of determining a prediction subblock of the first subblock from the plurality of candidate prediction subblocks, wherein the prediction subblock of the first subblock is indicated by the precision BV; and (i) determining a reconfigured sample of the first subblock as a sample of the predicted subblock of the first subblock, or (ii) as a filtered sample of the predicted subblock filtered based on a filter coefficient; A device configured to perform the following actions.

15. In the apparatus according to any one of claims 8 to 11, the processing circuit is: (i) determining the BV of another block in the current picture as either the precise BV of the first subblock, or (ii) a weighted combination of the precise BV of the first subblock and the precise BV of the second subblock, wherein the first subblock is the subblock of the current block closest to the other block; and A step of determining the predicted block of the other block, indicated by the BV of the other block determined in the above determination step; A device configured to perform the following actions.