Template Matching Based Adaptive Block Vector Resolution (ABVR) in IBC

Template matching-based adaptive block vector resolution in IBC mode addresses the challenge of suboptimal precision in IBC, enhancing video decoding efficiency and compression performance.

JP2025539228APending Publication Date: 2025-12-04TENCENT AMERICA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025524954
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-31
Filing Date
2023-09-01
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently utilizing intra block copy (IBC) mode for improved compression efficiency, particularly in scenarios where block vector precision is not optimally determined, leading to suboptimal video decoding performance.

Method used

Implementing template matching-based adaptive block vector resolution (ABVR) in IBC mode to sort and select block vector precisions based on template matching differences, enabling selection of the most accurate precision for block vector reconstruction.

Benefits of technology

Enhances video decoding efficiency by optimizing block vector precision, thereby improving compression performance and overall video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539228000001_ABST
    Figure 2025539228000001_ABST
Patent Text Reader

Abstract

Coding information is received, indicating that an intra block copy (IBC) mode is applied to a current block. Based on the application of the IBC mode to the current block, first adaptive block vector resolution (ABVR) information included in the coding information is obtained from the received video bitstream. The first ABVR information is determined to indicate that multiple BV precisions are associated with the block vectors (BVs) of the IBC mode. The multiple BV precisions associated with the BVs of the current block are sorted based on template matching (TM) differences between a template region of the current block and each of the multiple template regions of a reference block. A specific BV precision is selected from the multiple sorted BV precisions by the TM differences. The current block is reconstructed based on at least the selected specific BV precision.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. patent application Ser. No. 63 / 438,491, entitled "TEMPLATE-MATCHING BASED ADAPTIVE BLOCK VECTOR RESOLUTION (ABVR) IN IBC," filed in 2023, which in turn claims the benefit of priority to U.S. Provisional Application No. 63 / 438,491, entitled "TEMPLATE-MATCHING BASED ADAPTIVE MOTION VECTOR RESOLUTION (AMVR) IN IBC," filed on Jan. 11, 2023. The entire disclosure of the prior application is incorporated by reference.

[0002] [Technical field] This disclosure describes embodiments that relate generally to video coding. [Background technology]

[0003] The background art discussion provided herein is intended to generally present the context for the present disclosure, and the inventors' work, to the extent described in this background art section, as well as aspects of the description that are not admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.

[0004] Image / video compression can help transmit image / video data across different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, video codecs can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, video codecs can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can generally be represented by a motion vector (MV). Summary of the Invention

[0005] Aspects of this disclosure include methods and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit.

[0006] According to one aspect of the present disclosure, a video decoding method performed in a video decoder is provided. In the method, a video bitstream including coding information for a current block is received. The coding information indicates that an intra block copy (IBC) mode is applied to the current block. Based on the application of the IBC mode to the current block, first adaptive block vector resolution (ABVR) information included in the coding information is obtained from the received video bitstream. The first ABVR information is determined to indicate that multiple BV precisions are associated with the block vectors (BVs) of the IBC mode. The multiple BV precisions associated with the BVs of the current block are sorted based on template matching (TM) differences between a template region of the current block and each of multiple template regions of a reference block. A specific BV precision is selected from the multiple sorted BV precisions according to the TM differences. The current block is reconstructed based on at least the selected specific BV precision.

[0007] In one example, a TM difference between a template region of a current block and each of the template regions of the reference blocks of the current block associated with a plurality of BV precisions is determined. The template region of the current block includes samples above and to the left of the current block. The plurality of BV precisions are sorted in ascending order based on the TM difference between the template region of the current block and the template region of the reference block of the current block.

[0008] In one example, the first ABVR information includes first precision index information indicating that a specific BV precision corresponds to a minimum TM difference among TM differences between the template region of the current block and the template region of the reference block of the current block, and the specific BV precision corresponding to the minimum TM difference is further selected from the multiple sorted BV precisions according to the first precision index information.

[0009] In one example, based on the fact that the IBC mode is applied to the current block and the first ABVR information indicating that the particular BV precision corresponds to the smallest TM difference among the TM differences, the particular BV precision is selected from the multiple sorted BV precisions corresponding to the smallest TM difference among the TM differences between the template region of the current block and the template region of the reference block of the current block.

[0010] In one example, based on the fact that the IBC mode is applied to the current block and the first ABVR information indicating that the particular precision does not correspond to the smallest TM difference, the particular BV precision is selected from the plurality of sorted BV precisions based on first precision index information included in the coding information, where the first precision index information indicates that the particular BV precision corresponds to the second smallest TM difference between the template region of the current block and the template region of the reference block of the current block.

[0011] In one example, based on the first ABVR information indicating that a predetermined BV precision is excluded from the multiple BV precisions, the subset of the multiple BV precisions excluding the predetermined BV precision are sorted in ascending order based on the TM difference between the template region of the current block and the template regions of the subset of reference blocks of the current block excluding the predetermined reference block corresponding to the predetermined BV precision.

[0012] In one example, a particular BV precision is selected from the plurality of sorted BV precisions based on first precision index information included in the coding information indicating that the particular BV precision corresponds to the smallest TM difference among the TM differences between the template region of the current block and the template regions of a subset of the reference blocks of the current block.

[0013] In one example, based on the first ABVR information indicating that one of the plurality of BV precisions does not correspond to the smallest TM difference among the TM differences, it is determined whether the second ABVR information of the coding information indicates that the plurality of BV precisions are ordered based on a predefined sequence. Based on the second ABVR information indicating that the plurality of BV precisions are ordered according to the predefined sequence, a particular BV precision is selected from the plurality of BV precisions in the predefined sequence based on second precision index information included in the coding information and indicating which of the plurality of BV precisions is to be selected.

[0014] In one example, based on the first ABVR information indicating that ABVR is enabled so that multiple BV precisions are ordered according to a predefined sequence, it is determined whether second ABVR information of the coding information indicates that the multiple BV precisions are sorted based on TM differences. Based on the second ABVR information indicating that the multiple BV precisions are sorted based on TM differences, a specific BV precision corresponding to a smallest difference among the TM differences between the template region of the current block and the template region of the reference block of the current block is selected from the multiple sorted BV precisions.

[0015] In one example, based on the first ABVR information indicating that ABVR is enabled, it is determined whether second ABVR information of the coding information indicates that the multiple BV precisions are to be reordered based on TM differences. Based on the second ABVR information being determined as indicating that the multiple BV precisions are not to be reordered based on TM differences, a specific BV precision is selected from the multiple BV precisions in a predefined sequence based on second precision index information indicating which of the multiple BV precisions is to be selected.

[0016] According to another aspect of the present disclosure, there is provided an apparatus, the apparatus including a processing circuit, the processing circuit being configured to perform any of the described methods for video decoding / encoding.

[0017] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding / encoding. [Brief explanation of the drawings]

[0018] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] FIG. 1 is a schematic diagram of a template matching process. [Figure 5] 1 shows a flowchart outlining a decoding process according to some embodiments of the present disclosure. [Figure 6] 1 shows a flowchart outlining an encoding process according to some embodiments of the present disclosure. [Figure 7] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0019] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example of an application of the disclosed subject matter, a video encoder and video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.), etc.

[0020] The video processing system (100) includes a video source (101) and a capture subsystem (113) that can include, for example, a digital camera, creating a stream of uncompressed video pictures (102). In one example, the video picture stream (102) includes samples captured by the digital camera. The video picture stream (102), shown with a thick line to emphasize its high data volume compared to the encoded video data (104) (or coded video bitstream), can be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or coded video bitstream), shown with a thin line to emphasize its low data volume compared to the video picture stream (102), can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as client subsystems 106 and 108 in FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 may include, for example, a video decoder 110 within an electronic device 130. The video decoder 110 decodes an input copy 107 of the encoded video data and creates an output stream 111 of video pictures that can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) can be encoded according to several video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, the developing video coding standard is informally known as Versatile Video Coding (VVC).The disclosed subject matter may be used in the context of VVC.

[0021] It is noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and electronic device 130 may include a video encoder (not shown).

[0022] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used in place of the video decoder (110) in the example of FIG. 1.

[0023] The receiver (231) may receive one or more coded video sequences to be decoded by the video decoder (210), for example, contained in a bitstream. In one embodiment, one coded video sequence may be received at a time, with the decoding of each coded video sequence being independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (231) may receive the coded video data along with other data (e.g., coded audio data and / or auxiliary data streams), which may be forwarded to respective using entities (not shown). The receiver (231) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the “parser” (220)). In certain applications, the buffer memory (215) is part of the video decoder (210). In other cases, it may be external to the video decoder (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder (210), for example, to prevent network jitter, and there may be yet another buffer memory (215) within the video decoder (210), for example, to handle playback timing. If the receiver (231) is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (215) may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (215) may be needed and may be relatively large, advantageously adaptively sized, and may be implemented at least in part in an operating system and similar elements (not shown) external to the video decoder (210).

[0024] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and potentially include information for controlling a rendering device, such as a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but can be coupled to the electronic device (230), as shown in FIG. 2. The rendering device control information may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0025] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to generate symbols (221).

[0026] The reconstruction of the symbols (221) can involve several different units, depending on the type of video picture or portion thereof being coded (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed by the parser (220) from the coded video sequence. The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.

[0027] In addition to the functional blocks described above, the video decoder (210) can be conceptually subdivided into multiple functional units, as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:

[0028] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbols (221) from the parser (220), along with control information (including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc.) The scalar / inverse transform unit (251) can output blocks containing sample values ​​that can be input to an aggregator (255).

[0029] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates blocks of the same size and shape as the block being reconstructed using surrounding, already reconstructed information retrieved from the current picture buffer (258). The current picture buffer (258), for example, buffers a partially reconstructed and / or fully reconstructed current picture. In some cases, the aggregator (255) adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0030] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (253) may access a reference picture memory (257) to retrieve samples used for prediction. After motion-compensating the retrieved samples according to the symbols (221) associated with the block, these samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) retrieves prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253), for example, in the form of symbols (221) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0031] The output samples of the aggregator 255 may be subjected to various loop filtering techniques in a loop filter unit 256. The video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit 256 as symbols 221 from the parser 220. The video compression may also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed loop-filtered sample values.

[0032] The output of the loop filter unit (256) can be a sample stream that can be output to a rendering device (212) and stored in a reference picture memory (257) for use in future inter-picture prediction.

[0033] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.

[0034] The video decoder (210) may perform decoding operations in accordance with a given video compression technology or standard, such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence conforms to both the syntax and profile of the video compression technology or standard as documented in the video compression technology or standard. Specifically, a profile can select certain tools from all tools available in the video compression technology or standard as the only tools for use under that profile. Compliance may also require the complexity of the coded video sequence to fall within a range defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further constrained through a Hypothetical Reference Decoder (HRD) specification and metadata about HRD buffer management signaled in the coded video sequence.

[0035] In one embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0036] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.

[0037] The video encoder (303) may receive video samples from a video source (301) (not part of the electronic device (320) in the example of FIG. 3) that may capture video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0038] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media presentation system, the video source (301) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. Video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. The following discussion focuses on samples.

[0039] According to one embodiment, the video encoder (303) may code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other required time constraints. Achieving an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) can control and is operatively coupled to other functional units, as described below. Coupling is not shown for clarity. Parameters set by the controller (350) can include rate control-related parameters (e.g., picture skip, quantization, lambda values ​​for rate-distortion optimization techniques), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other appropriate functionality associated with the video encoder (303) optimized for a particular system design.

[0040] In some embodiments, the video encoder (303) is configured to operate in a coding loop. As a very simplified description, in one example, the coding loop can include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data similar to that generated by the (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-for-bit accurate results independent of the location of the decoder (local or remote), the contents in the reference picture memory (334) are also bit-for-bit accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (including the resulting drift when synchronization cannot be maintained, for example, due to channel errors) is used in several related techniques as well.

[0041] The operation of the "local" decoder (333) can be the same as the "remote" decoder (210), such as the video decoder (210), already described in detail above in connection with Figure 2. However, briefly referring to Figure 2, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (345) and parser (220) can be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (433).

[0042] In one embodiment, decoder technology, excluding analysis / entropy decoding, present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the subject matter of the disclosure focuses on decoder operation. A description of the encoder technology can be omitted, as it is the reverse of the decoder technology, which is described generically. In certain areas, more detailed descriptions are provided below.

[0043] During operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0044] The local video decoder (333) may decode coded video data of pictures that may be designated as reference pictures based on symbols created by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a replica of the source video sequence, with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures that have common content with reconstructed reference pictures (without transmission errors) obtained by the far-end video decoder.

[0045] The predictor (335) may perform a predictive search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or specific metadata (reference picture motion vectors, block shapes, etc.), which may serve as suitable prediction references for the new picture. The predictor (335) may operate sample block-by-pixel block to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).

[0046] The controller (350) may manage the coding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0047] The output of all the above functional units may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0048] The transmitter (340) may buffer the coded video sequence produced by the entropy coder (345) and prepare it for transmission over a communication channel (360), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (340) may merge the coded video data from the video encoder (330) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (not shown).

[0049] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign a particular coding picture type to each coded picture. The coding picture type may affect the coding technique that may be applied to each picture. For example, a picture may be assigned as one of the following picture types:

[0050] Intra-pictures (I-pictures) may be coded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0051] Predictive pictures (P pictures) may be coded and decoded using intra- or inter-prediction, which uses motion vectors and reference indices to predict the sample values ​​of each block.

[0052] Bidirectionally predicted pictures (B pictures) may be coded and decoded using intra- or inter-prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0053] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0054] The video encoder (303) may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0055] In one embodiment, the transmitter (340) may transmit additional data along with the coded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other types of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0056] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0057] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. Bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, both of which precede the current picture in decoding order (but may also be past and future, respectively, in display order) in a video. A block in the current picture can be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block can be predicted by a combination of the first and second reference blocks.

[0058] Furthermore, to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.

[0059] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU can be partitioned into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the CU's prediction type, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation of coding (encoding / decoding) is performed on a prediction block basis. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0060] It is noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.

[0061] This disclosure includes aspects related to template matching-based adaptive block vector resolution (ABVR) in intra block copy (IBC).

[0062] IBC is adopted in the HEVC extension to screen content coding (SCC). IBC can significantly improve the coding efficiency of screen content material. Since IBC mode can be implemented as a block-level coding mode, block matching (BM) can be performed in the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector can be used to indicate the displacement from the current block to a reference block, which is already reconstructed in the current picture. The luma block vectors of IBC-coded CUs can be defined using integer precision. Chroma block vectors can also be rounded to integer precision. When combined with adaptive motion vector resolution (AMVR), IBC mode can switch between 1-pel and 4-pel motion vector precision. IBC-coded CUs can be treated as a third prediction mode other than intra or inter prediction modes. IBC mode can be applicable to CUs with both width and height below a certain value, such as 64 luma samples.

[0063] On the encoder side, hash-based motion estimation can be performed for IBC. The encoder can perform a rate distortion (RD) check on blocks with either width or height below a threshold, such as 16 luma samples. For non-merged motion, a block vector search can be performed first using a hash-based search. If the hash search does not return valid candidates, a block matching-based local search can be performed.

[0064] In hash-based search, hash key matching (e.g., 32-bit CRC) between the current block and the reference block can be extended to other block sizes, such as all allowed block sizes. Hash key calculation for all positions in the current picture can be based on subblocks, such as 4x4 subblocks. When the current block has a larger size, the hash key of the current block can be determined to match the hash key of the reference block when all hash keys of all 4x4 subblocks of the current block match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference can be calculated, and the one with the smallest cost can be selected.

[0065] In a block matching search, the search range can be set to cover both the previous CTU and the current CTU.

[0066] At the CU level, the IBC mode can be signaled using a flag, and the IBC mode can be signaled as IBC AMVP mode or IBC skip / merge mode as follows: (1) IBC skip / merge mode: The merge candidate index is used to indicate which block vectors in a list from neighboring candidate IBC-coded blocks can be used to predict the current block. The merge list can include, for example, spatial candidates, HMVP candidates, and pairwise candidates. (2) IBC AMVP mode: Block vector differences can be coded in the same way as motion vector differences. The block vector prediction method can use two candidates as predictors: one from the left neighbor (if IBC coded) and one from the upper neighbor. When any neighboring block is unavailable, a default block vector can be used as the predictor. A flag can also be signaled to indicate the block vector predictor index.

[0067] To further improve the compression efficiency of the VVC standard, JVET-U0100 and EE2 were planned to be carried out during the 21st and 22nd JVET meetings to evaluate advanced compression tools beyond VVC capabilities. Template matching (TM) for decoder-side motion enhancement was proposed in JVET-U010 and EE2. In TM mode, motion enhancement can be achieved by constructing a template from neighboring reconstructed samples to the left and / or above, and finding the closest match between the template in the current picture and the corresponding template in the reference frame.

[0068] As shown in Figure 4, a current block (402) may include a template (or current template) (404) that includes neighboring blocks above and to the left of the current block (402) in a current frame (406). A block (or reference template) (408) in a reference frame (410) may be indicated by an initial motion vector (412). A better motion vector (not shown) may be searched (identified) around the initial motion vector (412) of the current CU (402) within a [-8, +8] pel search range. Template matching methods such as JVET-J0021 and EE2 may be applied with the following modifications: (1) The search step size can be determined based on the AMVR mode. (2) TM can be cascaded with a bilateral matching process in merge mode.

[0069] According to the adaptive block vector resolution (ABVR) mode, for blocks coded using IBC advanced motion vector prediction (AMVP), additional ABVR signaling for multiple BV precisions (also referred to as BV resolutions), such as 1 / 2-pel and 1 / 4-pel BV precisions, can be introduced. In one example, a BV can indicate an initial reference block for a current block, and the BV precision associated with the BV can indicate candidate reference blocks near the initial reference block. Multiple BV precisions (or resolutions) can indicate a search range such that a more appropriate BV indicating a reference block with the smallest difference from the current block can be selected. When fractional BV precision (e.g., 1 / 2-pel) is enabled, the corresponding BVP and BVD can be in units of the selected precision. At the encoder side, an integer BV search method can be performed first, and then a group of N integer BVs can be selected for further fractional refinement, and 1 / 2-pel and 1 / 4-pel BV searches can be applied sequentially. The BV with the smallest RD cost in either integer or fractional precision can be signaled to the decoder side. In an exemplary design, an 8-tap DCT-IF interpolation filter can be applied to generate IBC predicted samples at fractional sample positions. Furthermore, when the interpolation process accesses reconstructed samples beyond the available reference region, a padding operation can be performed. The padding operation can copy samples from the nearest integer sample position within the valid reference region. For the IBC merge mode, fractional BVs can be inherited through spatial neighbors, which can be similar to the existing IBC inheritance process in ECMs such as ECM-7.0.

[0070] In ECM, template matching is widely adopted to refine the motion information of MVP candidates in AMVP mode and merge candidates in merge mode. Template matching is also used to sort MVP candidates in AMVP mode. For IBC, adaptive motion vector resolution (AMVR) has also been studied in ECM. However, the BV resolution indexes in adaptive BV resolution (ABVR) mode are still coded in a fixed order in IBC.

[0071] In this disclosure, a template matching (TM)-based ABVR in IBC is provided. According to IBC, a BV can be provided to indicate an initial reference block. Multiple BV precisions associated with the BV can be introduced according to ABVR. Each of the BV precisions can indicate a respective reference block around the initial reference block. The multiple BV precisions can be further sorted according to TM.

[0072] In one aspect, the TM for sorting BV precision indexes is applied in IBC AMVP mode. When sps_abvr_enabled_flag is true, the TM search procedure is performed with possible BV precision to sort BV resolution indexes in ascending order by using the TM cost between the template and the reference block.

[0073] In one embodiment, in a prediction mode such as IBC AMVP mode, TM can be applied to sort BV precision indexes. Each of the BV precision indexes can indicate a respective BV precision, such as 1 / 4 pel, 1 / 2 pel, and 1 pel. A TM search procedure can be performed on the possible BV precisions associated with IBC BVs to sort the BV resolution indexes in a predefined order (e.g., ascending order) by using the TM cost (or TM difference) between the template of the current block and the template of the reference block. In one example, when an ABVR enabled flag (e.g., sps_abvr_enabled_flag) is received and is true, the TM search procedure can be performed. When the ABVR enabled flag is true, ABVR mode can be applied to the current block, and multiple BV precisions associated with the BVs can be applied.

[0074] In one aspect, TM is applied for all BV precisions, including but not limited to ¼-pel, ½-pel, and 1-pel. The TM search procedure is performed at all BV resolutions, and then the BV precisions are sorted in ascending order by using the TM cost between the template and the reference block with the corresponding BV resolution. tm_abvr_flag being false indicates that the BV resolution with the smallest TM cost is selected. tm_abvr_precision_idx is further signaled when tm_abvr_flag is true. The selected ABVR precision is derived from the remaining BV resolutions sorted in ascending order of TM cost. A detailed description of the pseudo syntax is shown in Table 1 below.

[0075] In one embodiment, TM can be applied to multiple BV resolutions (or resolutions), including, but not limited to, ¼-pel, ½-pel, and 1-pel, for example, all BV resolutions. The TM search procedure can be performed for multiple BV resolutions (or resolutions), and then the BV resolutions are sorted in a predefined order, such as ascending order, by using the TM cost between the template of the current block and the reference block of the current block corresponding to the BV resolution. In one example, based on an initial BV indicating an initial point, multiple BV resolutions around the initial point can be determined. Each BV resolution can correspond to a respective reference block. TM can be further applied to sort the BV resolutions. For example, the TM cost (or TM difference) between the template of the current block and each template of the reference block indicated by the BV resolution can be determined. The BV resolutions can be sorted based on the TM cost in a predefined order, such as ascending order.

[0076] Table 1 shows an exemplary TM-based ABVR in IBC. As shown in Table 1, a TM-based ABVR flag (e.g., tm_abvr_flag) can be signaled. When tm_abvr_flag is false, a BV resolution corresponding to a predetermined TM cost, such as the minimum TM cost, can be selected. When tm_abvr_flag is true, a BV precision corresponding to the minimum TM cost can be excluded from the multiple BV precisions, and the remaining BV precisions can be further sorted based on their corresponding TM costs. Furthermore, TM-based precision index information (e.g., tm_abvr_precision_idx) can be signaled. The selected ABVR precision can be derived from the remaining BV resolutions sorted in a predefined order (e.g., ascending order) of TM cost based on tm_abvr_precision_idx. [Table 1]

[0077] In one aspect, TM is applied to all BV precisions, including but not limited to ¼-pel, ½-pel, and 1-pel. The TM search procedure is performed at all BV resolutions, and then the BV precisions are sorted in ascending order by using the TM cost between the template and the reference block with the corresponding BV resolution. When TM-based sorting for ABVR mode is signaled in high-level syntax, such as but not limited to SPS, PPS, slice header, etc., the selected ABVR precision is derived from the BV resolutions sorted in ascending order of TM cost. A detailed description of the pseudo-syntax is shown in Table 2 below.

[0078] In one embodiment, TM can be applied to multiple VB precisions, for example, all BV precisions, including, but not limited to, ¼-pel, ½-pel, and 1-pel. The TM search procedure can be performed for multiple BV resolutions, and then the BV precisions can be sorted in a predefined order (e.g., ascending order) by using the TM cost between the template of the current block and the template of the reference block corresponding to the BV resolution. When TM-based sorting in ABVR mode is signaled in a high-level syntax, such as, but not limited to, SPS, PPS, slice header, etc., the selected ABVR precision can be derived from the BV resolutions sorted in a predefined order (e.g., ascending order) of TM cost based on precision index information, such as TM-based precision index information (e.g., tm_abvr_precision_idx) or ABVR-based precision index information (e.g., abvr_precision_idx). In one example, the BV precision corresponding to the smallest TM cost among the TM costs can be selected from the sorted BV resolutions according to the precision index information. [Table 2]

[0079] Table 2 shows another exemplary TM-based ABVR in IBC. Assuming that coding information signaled in the video bitstream indicates that intra block copy (IBC) mode is applied to the current block, when the ABVR enabled flag (e.g., sps_abvr_enabled_flag) is true (indicating that ABVR mode is applied) and the MVD is not equal to 0, the TM cost (or TM difference) between the template of the current block and each template of the reference block of the current block can be determined, as shown in Table 2. The reference block of the current block can be indicated by BV precision (or BV resolution). The BV precisions can be sorted based on the TM cost in a predefined order, such as ascending order. Furthermore, precision index information (e.g., tm_abvr_precision_idx or abvr_precision_idx) can be signaled to indicate which of the BV precisions is selected. For example, the precision index information can indicate that the BV precision corresponding to the smallest TM cost can be selected as the ABVR precision. Table 2 shows that TM-based precision index information (e.g., tm_abvr_precision_idx) is signaled to indicate which of the BV precisions is selected. However, ABVR-based precision index information (e.g., abvr_precision_idx) can also be signaled to indicate which of the BV precisions is selected.

[0080] In one aspect, when abvr_flag is false, TM is applied to all possible BV precisions except 1 / 4 pel. The TM search procedure is performed at BV precisions including, but not limited to, 1 / 2 pel and 1 pel to sort the BV precision indexes in ascending order by using the TM cost between the template and the reference block with the corresponding BV resolution. The ABVR precision index is derived from the sorted BV precision index. A detailed description of the pseudo syntax is shown in Table 3 below.

[0081] In one embodiment, as shown in Table 3, TM can be applied to multiple possible BV precisions, e.g., all possible BV precisions, except for a predefined BV precision. The predefined BV precision can be, for example, the most frequent resolution precision, such as ¼ pel. When an ABVR flag (e.g., abvr_flag) is false, the TM search procedure can be performed for BV precisions other than the predefined BV precision. Thus, the TM search procedure can be performed for the remaining BV precisions, which can include, but are not limited to, ½ pel and 1 pel, and the BV precision indexes associated with the remaining BV precisions can be sorted in a predefined order (e.g., ascending order) by using the TM cost between the template of the current block and the template of the reference block corresponding to the BV resolution (e.g., each BV resolution corresponds to a reference block of the current block). An ABVR precision index (or TM-based index information), such as tm_abvr_precision_idx, can be further derived from the sorted BV precision index. The TM-based precision index information (e.g., tm_abvr_precision_idx) can indicate which of the BV precisions is selected. For example, the BV precision corresponding to the smallest TM cost can be selected as the ABVR precision. [Table 3]

[0082] Alternatively, when abvr_flag is false, the TM can be applied for all possible BV precisions except 1 pel.

[0083] In one aspect, the TM procedure is applied to all possible BV precisions. When the ABVR precision with the smallest TM cost is selected, tm_amvr_flag is set to true. Otherwise, tm_amvr_flag is set to false and regular ABVR signaling is still used. A detailed description of the pseudo syntax is shown in Table 4.

[0084] In one embodiment, as shown in Table 4, the TM procedure can be applied to multiple possible BV precisions, for example, all possible BV precisions. When the ABVR precision corresponding to a predetermined TM cost, such as the minimum TM cost, is selected, the TM-based ABVR flag (e.g., tm_abvr_flag) can be set as true. Otherwise, tm_abvr_flag can be set as false, and normal ABVR signaling can still be used. As shown in Table 4, when the TM-based ABVR flag is false, the ABVR flag (e.g., abvr_flag) can be signaled. When abvr_flag is true, ABVR-based precision index information (e.g., abvr_precision_idx) can further be signaled to indicate which of the BV precisions can be selected. The BV precisions can be ordered in a predefined sequence according to the ABVR mode. [Table 4]

[0085] In one aspect, the ABVR flag is signaled first to indicate whether ABVR is enabled. If ABVR is enabled, tm_abvr_flag is signaled to indicate whether the BV resolution with the lowest TM cost is selected. When tm_abvr_flag is true, the ABVR precision, excluding 1 / 4-pel luma precision, with the lowest TM cost is selected. When tm_abvr_flag is set as false, the original signaling of the MV resolution index is used. A detailed description of the pseudo syntax is shown in Table 5 below.

[0086] In one embodiment, as shown in Table 5, an ABVR flag (e.g., abvr_flag) can be first signaled to indicate whether ABVR mode is enabled. An enabled ABVR can indicate that multiple BV precisions are applied to the current block. When ABVR is enabled, a TM-based ABVR flag (e.g., tm_abvr_flag) can be further signaled to indicate whether a BV resolution corresponding to a predetermined TM cost, such as the minimum TM cost, is selected. When tm_abvr_flag is true, an ABVR precision corresponding to a predetermined TM cost, such as the minimum TM cost, is selected. In one example, the ABVR precision can exclude a predefined precision, such as ¼-pel luma precision. When tm_abvr_flag is set as false, the original signaling of the MV resolution index (or ABVR-based precision index information), such as abvr_precision_idx, can be used. The ABVR-based precision index information (e.g., abvr_precision_idx) can indicate which of the BV precisions can be selected. The BV accuracies can be ordered in a predefined sequence according to the ABVR mode, e.g., not sorted based on TM cost. [Table 5]

[0087] Alternatively, in one aspect, when tm_abvr_flag is true, the ABVR precision excluding 1-pel luma precision with the lowest TM cost is selected.

[0088] Alternatively, precisions such as 1-pel luma precision can be excluded. Thus, when tm_abvr_flag is true, ABVR precisions excluding 1-pel luma precision with a given TM cost, such as the minimum TM cost, can be selected.

[0089] 5 shows a flowchart outlining a process (500) according to one embodiment of the present disclosure. The process (500) can be used in a video decoder. In various embodiments, the process (500) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), etc. In some embodiments, the process (500) is implemented with software instructions, and thus, the processing circuit performs the process (500) when it executes the software instructions. The process begins at (S501) and proceeds to (S510).

[0090] At (S510), a video bitstream including coding information for a current block is received, where the coding information indicates that an intra block copy (IBC) mode is applied to the current block.

[0091] At (S520), based on the IBC mode being applied to the current block, first adaptive block vector resolution (ABVR) information included in the coding information is obtained from the received video bitstream. In one example, the first ABVR information includes tm_abvr_precision_idx as shown in Table 2. In one example, the first ABVR information includes abvr_flag as shown in Tables 3 and 5. In one example, the first ABVR information includes tm_abvr_flag as shown in Tables 1 and 4.

[0092] In (S530), the first ABVR information is determined to indicate that multiple BV precisions are associated with the block vector (BV) in the IBC mode. For example, the first ABVR information in Table 2 includes first precision index information (e.g., tm_abvr_precision_idx) indicating that a specific BV precision is selected from multiple BV precisions based on template matching. In another example, the first ABVR information includes abvr_flag as in Tables 3 and 5, which indicates that the ABVR mode is applied to the current block. According to the ABVR mode, multiple BV precisions associated with the BV in the IBC mode are determined.

[0093] In (S540), multiple BV resolutions associated with the BVs of the current block are sorted based on the template matching (TM) difference between the template region of the current block and each of the multiple template regions of the reference block. For example, as shown in Table 1, the TM search procedure is performed at all BV resolutions, and then the BV resolutions are sorted in ascending order by using the TM cost between the template and the reference block having the corresponding BV resolution. In another example, as shown in Table 3, the TM search procedure is performed at BV resolutions including, but not limited to, 1 / 2 pel and 1 pel to sort the BV precision indexes in ascending order by using the TM cost between the template and the reference block having the corresponding BV resolution.

[0094] At (S550), a particular BV precision is selected from the multiple sorted BV precisions by TM difference. In one example, as shown in Table 1, tm_abvr_flag being false indicates that the BV resolution with the smallest TM cost is selected. When tm_abvr_flag is true, tm_abvr_precision_idx is further signaled. The selected ABVR precision (e.g., a particular BV precision) is derived from the remaining BV resolutions sorted in ascending order of TM cost.

[0095] At (S560), the current block is reconstructed based on at least the selected particular BV precision.

[0096] In one example, a TM difference between a template region of a current block and each of the template regions of the reference blocks of the current block associated with a plurality of BV precisions is determined. The template region of the current block includes samples above and to the left of the current block. The plurality of BV precisions are sorted in ascending order based on the TM difference between the template region of the current block and the template region of the reference block of the current block.

[0097] In one example, the first ABVR information includes first precision index information indicating that a specific BV precision corresponds to a minimum TM difference among TM differences between the template region of the current block and the template region of the reference block of the current block, and the specific BV precision corresponding to the minimum TM difference is further selected from the multiple sorted BV precisions according to the first precision index information.

[0098] In one example, based on the fact that the IBC mode is applied to the current block and the first ABVR information indicating that the particular BV precision corresponds to the smallest TM difference among the TM differences, the particular BV precision is selected from the multiple sorted BV precisions corresponding to the smallest TM difference among the TM differences between the template region of the current block and the template region of the reference block of the current block.

[0099] In one example, based on the fact that the IBC mode is applied to the current block and the first ABVR information indicating that the particular precision does not correspond to the smallest TM difference, the particular BV precision is selected from the plurality of sorted BV precisions based on first precision index information included in the coding information, where the first precision index information indicates that the particular BV precision corresponds to the second smallest TM difference between the template region of the current block and the template region of the reference block of the current block.

[0100] In one example, based on the first ABVR information indicating that a predetermined BV precision is excluded from the multiple BV precisions, the subset of the multiple BV precisions excluding the predetermined BV precision are sorted in ascending order based on the TM difference between the template region of the current block and the template regions of the subset of reference blocks of the current block excluding the predetermined reference block corresponding to the predetermined BV precision.

[0101] In one example, a particular BV precision is selected from the plurality of sorted BV precisions based on first precision index information included in the coding information indicating that the particular BV precision corresponds to the smallest TM difference among the TM differences between the template region of the current block and the template regions of a subset of the reference blocks of the current block.

[0102] In one example, based on the first ABVR information indicating that one of the plurality of BV precisions does not correspond to the smallest TM difference among the TM differences, it is determined whether the second ABVR information of the coding information indicates that the plurality of BV precisions are ordered based on a predefined sequence. Based on the second ABVR information indicating that the plurality of BV precisions are ordered according to the predefined sequence, a particular BV precision is selected from the plurality of BV precisions in the predefined sequence based on second precision index information included in the coding information and indicating which of the plurality of BV precisions is to be selected.

[0103] In one example, based on the first ABVR information indicating that ABVR is enabled so that multiple BV precisions are ordered according to a predefined sequence, it is determined whether second ABVR information of the coding information indicates that the multiple BV precisions are sorted based on TM differences. Based on the second ABVR information indicating that the multiple BV precisions are sorted based on TM differences, a specific BV precision corresponding to a smallest difference among the TM differences between the template region of the current block and the template region of the reference block of the current block is selected from the multiple sorted BV precisions.

[0104] In one example, based on the first ABVR information indicating that ABVR is enabled, it is determined whether second ABVR information of the coding information indicates that the multiple BV precisions are to be reordered based on TM differences. Based on the second ABVR information being determined as indicating that the multiple BV precisions are not to be reordered based on TM differences, a specific BV precision is selected from the multiple BV precisions in a predefined sequence based on second precision index information indicating which of the multiple BV precisions is to be selected.

[0105] The process then proceeds to (S599) and ends.

[0106] The process 500 may be adapted as appropriate. Steps of the process 500 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0107] 6 shows a flowchart outlining a process (600) according to one embodiment of the present disclosure. The process (600) can be used in a video encoder. In various embodiments, the process (600) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (600) is implemented with software instructions, and thus, the processing circuit performs the process (600) when it executes the software instructions. The process begins at (S601) and proceeds to (S610).

[0108] At (S610), it is determined whether intra block copy (IBC) mode is applied to the current block in the current picture.

[0109] In (S620), a block vector (BV) associated with the current block and multiple BV precisions of the BV are determined based on whether the IBC mode is applied to the current block. For example, as shown in Table 2, when an ABVR enabled flag (e.g., sps_abvr_enabled_flag) is true, the ABVR mode can be applied to the current block and multiple BV precisions associated with the BV can be applied.

[0110] In (S630), the multiple BV resolutions of the BV associated with the current block are sorted based on the template matching (TM) difference between the template region of the current block and each of the multiple template regions of the reference block of the current block associated with the multiple BV resolutions. For example, as shown in Table 2, a TM search procedure is performed with the possible BV resolutions to sort the BV resolution indexes in ascending order by using the TM cost between the template of the current block and the template of the reference block.

[0111] At (S640), first adaptive block vector resolution (ABVR) information is encoded into the video bitstream. The first ABVR information indicates that the current block is encoded based on one of a plurality of permuted BV precisions. In one example, the first ABVR information includes tm_abvr_precision_idx, as shown in Table 2. In one example, the first ABVR information includes abvr_flag, as shown in Tables 3 and 5. In one example, the first ABVR information includes tm_abvr_flag, as shown in Tables 1 and 4.

[0112] The process then proceeds to (S699) and ends.

[0113] The process 600 may be adapted as appropriate. Steps of the process 600 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0114] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 7 illustrates a computer system (700) suitable for implementing certain embodiments of the disclosed subject matter.

[0115] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly or through interpretation, microcode execution, etc.

[0116] The instructions may be executed in various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0117] 7 for computer system (700) are exemplary in nature and are not intended to suggest any limitation regarding the scope of use or functionality of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of computer system (700).

[0118] The computer system 700 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). The human interface input devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic vision).

[0119] The human interface input devices may include one or more of a keyboard (701), a mouse (702), a trackpad (703), a touch screen (710), a data glove (not shown), a joystick (705), a microphone (706), a scanner (707), and a camera (708) (only one of each is shown).

[0120] The computer system (700) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (710), data gloves (not shown), or joystick (705), although haptic feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers (709), headphones (not shown), etc.), visual output devices (e.g., screens (710), including CRT, LCD, plasma, and OLED screens, each with or without touchscreen input capability and each with or without haptic feedback capability, some of which may provide two-dimensional visual output or greater than three-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0121] The computer system (700) may also include human-accessible storage devices and their associated media, such as optical media or similar media (721), including CD / DVD ROM / RW (720) with CDs / DVDs, thumb drives (722), removable hard drives or solid-state drives (723), legacy magnetic media such as tape and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0122] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0123] The computer system 700 may also include an interface 754 to one or more communications networks 755. Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, vehicular, industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet and WLAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial TV; and vehicular and industrial networks including CANbus. Certain networks typically require an external network interface adapter (e.g., a USB port on the computer system 700) attached to a particular general-purpose data port or peripheral bus 749; others are typically integrated into the core of the computer system 700 by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, the computer system 700 can communicate with other entities. Such communications may be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., from a particular CANbus to a particular CANbus device), or bidirectional, to other computer systems, for example, using local or wide area digital networks. As noted above, specific protocols and protocol stacks may be used in each of these networks and network interfaces.

[0124] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces can be attached to the core (740) of the computer system (700).

[0125] The core (740) may include one or more central processing units (CPUs) (741), graphics processing units (GPUs) (742), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (743), task-specific hardware accelerators (744), graphics adapters (750), etc. These devices may be connected through a system bus (748), along with read-only memory (ROM) (745), random access memory (RAM) (746), and internal mass storage (747), such as an internal non-user-accessible hard drive or SSD. In some computer systems, the system bus (748) is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (748) or through a peripheral bus (749). In one example, a screen (710) may be connected to the graphics adapter (750). Peripheral bus architectures include PCI, USB, etc.

[0126] The CPU (741), GPU (742), FPGA (743), and accelerator (744) can execute specific instructions, which, in combination, can constitute the computer code described above. The computer code can be stored in ROM (745) or RAM (746). Temporary data can be stored in RAM (746), while permanent data can be stored, for example, in internal mass storage (747). High-speed storage and retrieval from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU (741), GPU (742), mass storage (747), ROM (745), RAM (746), etc.

[0127] The computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0128] By way of example and not limitation, a computer system having the architecture (700), and in particular the core (740), can provide functionality as a result of the processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage, as introduced above, as well as media associated with the core's (740) specific storage of a non-transitory nature, such as the core's internal mass storage (747) or ROM (745). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (740). The computer-readable media can include one or more memory devices or chips according to particular needs. The software can cause the core (740), and in particular the processor therein (including a CPU, GPU, FPGA, etc.), to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM (746) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (744)), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0129] The use of "at least one" or "one" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). The use of "one" does not exclude any combination of the listed elements, where applicable, such as when elements are not mutually exclusive.

[0130] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise various systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. 1. A method of video decoding performed in a video decoder, comprising: receiving a video bitstream including coding information for a current block, the coding information indicating that an intra block copy (IBC) mode is to be applied to the current block; obtaining, from the received video bitstream, first adaptive block vector resolution (ABVR) information included in the coding information based on the IBC mode being applied to the current block; determining that the first ABVR information indicates that multiple BV precisions are associated with the IBC mode block vector (BV); sorting the plurality of BV accuracies associated with the BVs of the current block based on template matching (TM) differences between a template region of the current block and each of a plurality of template regions of a reference block; selecting a specific BV precision from the plurality of sorted BV precisions according to the TM difference; reconstructing the current block based on at least the selected particular BV precision; A method comprising:

2. The sorting step includes: determining the TM difference between the template region of the current block and each of the template regions of the reference blocks of the current block associated with the plurality of BV precisions, wherein the template region of the current block includes samples above and to the left of the current block; sorting the plurality of BV precisions in ascending order based on the TM difference between the template region of the current block and the template region of the reference block of the current block; The method of claim 1 further comprising:

3. the first ABVR information includes first precision index information indicating that the specific BV precision corresponds to a minimum TM difference among the TM differences between the template region of the current block and the template region of the reference block of the current block; 2. The method of claim 1, wherein the selecting step comprises selecting the particular BV precision corresponding to the smallest TM difference from the plurality of sorted BV precisions according to the first precision index information.

4. The selecting step includes:

2. The method of claim 1, further comprising: selecting the specific BV precision from the plurality of sorted BV precisions corresponding to the smallest TM difference among the TM differences between the template region of the current block and the template region of the reference block for the current block, based on the first ABVR information indicating that the IBC mode is applied to the current block and that the specific BV precision corresponds to the smallest TM difference among the TM differences.

5. The selecting step includes: based on the IBC mode being applied to the current block and the first ABVR information indicating that the particular precision does not correspond to the smallest TM difference, 5. The method of claim 4, further comprising: selecting the specific BV precision from the plurality of sorted BV precisions based on first precision index information included in the coding information indicating that the specific BV precision corresponds to a second smallest difference among the TM differences between the template region of the current block and the template region of the reference block of the current block.

6. The sorting step includes: based on the first ABVR information indicating that a predetermined BV precision is excluded from the plurality of BV precisions; 2. The method of claim 1, further comprising: sorting the subsets of the plurality of BV accuracies, excluding the predetermined BV accuracy, in ascending order based on a TM difference between the template region of the current block and the template regions of the subset of reference blocks of the current block, excluding the predetermined reference block corresponding to the predetermined BV accuracy.

7. 7. The method of claim 6, further comprising: selecting the particular BV precision from the plurality of sorted BV precisions based on first precision index information included in the coding information indicating that the particular BV precision corresponds to a smallest TM difference among the TM differences between the template region of the current block and the template regions of the subset of reference blocks of the current block.

8. The selecting step includes: determining whether second ABVR information of the coding information indicates that the plurality of BV precisions are ordered based on a predefined sequence, based on the first ABVR information indicating that one of the plurality of BV precisions does not correspond to the smallest TM difference among the TM differences; selecting the particular BV precision from the plurality of BV precisions in the predefined sequence based on second precision index information included in the coding information and indicating which of the plurality of BV precisions is to be selected, based on the second ABVR information indicating that the plurality of BV precisions are ordered according to the predefined sequence; The method of claim 1 further comprising:

9. The selecting step includes: determining whether second ABVR information of the coding information indicates that the plurality of BV precisions are sorted based on the TM difference, based on the first ABVR information indicating that ABVR is enabled such that the plurality of BV precisions are ordered according to a predefined sequence; selecting, from the plurality of sorted BV accuracies, the particular BV accuracies corresponding to a minimum difference among the TM differences between the template region of the current block and the template region of the reference block for the current block, based on the second ABVR information being determined as indicating that the plurality of BV accuracies are sorted based on the TM differences; The method of claim 1 further comprising:

10. The selecting step includes: determining whether the second ABVR information of the coding information indicates that the plurality of BV precisions are reordered based on the TM difference based on the first ABVR information indicating that the ABVR is enabled; selecting the particular BV precision from the plurality of BV precisions in a predefined sequence based on second precision index information indicating which of the plurality of BV precisions is to be selected based on the second ABVR information being determined as indicating that the plurality of BV precisions are not sorted based on the TM difference; The method of claim 9 further comprising:

11. 1. An apparatus including a processing circuit, Apparatus, wherein the processing circuitry is configured to perform the method of any one of claims 1 to 10.

12. A computer program product causing a computer to carry out the method according to any one of claims 1 to 10.

13. 1. A method of video encoding performed in a video encoder, comprising: determining whether an intra block copy (IBC) mode is applied to a current block in a current picture; determining a block vector (BV) associated with the current block and a plurality of BV precisions of the BV based on the IBC mode being applied to the current block; sorting the plurality of BV accuracies associated with the BVs of the current block based on template matching (TM) differences between a template region of the current block and each of a plurality of template regions of a reference block; encoding first adaptive block vector resolution (ABVR) information indicating that the current block is encoded based on one of the plurality of permuted BV precisions; A method comprising:

Citation Information

Patent Citations

  • Template matching in video coding

    US20220210438A1

  • Method and apparatus for video coding using block vector with adaptive spatial resolution

    US20240031558A1

  • Template matching in video coding

    WO2022146833A1

  • Method and device for coding video using block vector having adaptive spatial resolution

    WO2022211411A1