Partitioning Pattern Derivation Based on Template Matching
Template matching and geometric partitioning are used to derive optimal partitioning patterns for video blocks, addressing inefficiencies in existing video coding technologies and enhancing compression and decoding performance.
Patent Information
- Application Number
- JP2025514692
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-12
- Filing Date
- 2023-10-13
- Publication Date
- 2025-09-11
- Estimated Expiration
- 2043-10-13
AI Technical Summary
Existing video coding technologies face challenges in efficiently utilizing spatial and temporal redundancy for video compression, particularly in deriving optimal partitioning patterns for video blocks during encoding and decoding processes.
The method involves template matching (TM) to determine a reference block for a current block, classify its samples into classes, derive a partitioning pattern based on predetermined patterns, and reconstruct the block using the derived pattern, leveraging techniques such as geometric partitioning mode (GPM) and template matching to refine motion vectors.
This approach enhances video compression efficiency by optimizing partitioning patterns, reducing data volume, and improving decoding accuracy, aligning with emerging standards like Versatile Video Coding (VVC).
Smart Images

Figure 2025530286000001_ABST
Abstract
Description
[Technical Field]
[0001] [Related Applications] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 416,407, entitled "Template Matching Based Partitioning Pattern Derivation," filed October 14, 2022, which in turn claims the benefit of priority to U.S. Patent Application No. 18 / 379,624, entitled "TEMPLATE MATCHING BASED PARTITIONING PATTERN DERIVATION," filed October 12, 2023. The disclosures of the foregoing applications are incorporated herein by reference in their entireties.
[0002] [Technical field] This disclosure generally describes embodiments related to video coding. [Background technology]
[0003] The background description provided herein is intended to provide a general background to the present disclosure. The work of the presently named inventors is not expressly or implicitly admitted as prior art to the present disclosure, to the extent that the work described in this background section, as well as aspects of the description that may not be considered prior art at the time of filing, is not admitted as prior art to the present disclosure.
[0004] Image / video compression helps transfer image / video data between various devices, storage, and networks with minimal loss of quality. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. For example, video codecs can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, video codecs can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation is indicated by a motion vector (MV). Summary of the Invention
[0005] Aspects of the disclosure include a method and apparatus for video encoding / decoding. In some examples, an apparatus for video decoding includes a processing circuit.
[0006] According to one aspect of the present disclosure, a method for video decoding is provided. In the method, a video bitstream including a current block in a current picture is received. A reference block is determined for the current block from a plurality of candidate reference blocks based on template matching (TM) costs of the plurality of candidate reference blocks. The TM cost indicates a difference between a template of the current block and a reference template of each of the plurality of candidate reference blocks. Samples of the determined reference block are classified into a plurality of sample classes. A partitioning pattern for the current block is derived from a plurality of predetermined partitioning patterns based on the determined reference block. The derived partitioning pattern indicates a plurality of partitions of the current block. Each of the plurality of classes of samples of the determined reference block corresponds to a respective partition of the plurality of partitions of the current block. The current block is reconstructed based on the derived partitioning pattern of the current block.
[0007] In one aspect, the plurality of candidate reference blocks are from one of the current picture and the reference picture of the current block. The samples of the determined reference block are reconstructed samples.
[0008] In one example, an initial reference block is determined based on motion vector information included in a received video bitstream. A plurality of candidate reference blocks are determined within a search range of the initial reference block. A TM cost between a reference template of each of the plurality of candidate reference blocks and a template of a current block is determined. From the plurality of candidate reference blocks, a reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block is determined.
[0009] In one example, the samples of the determined reference block are classified into a first sample class and a second sample class based on a binary image segmentation algorithm.
[0010] In one example, the samples of the determined reference block are classified into a first sample class having sample values greater than a threshold value and a second sample class having sample values less than a threshold value, the threshold value being either the mean value or the median value of the samples.
[0011] In one example, the samples of the determined reference block are clustered into one or more sample classes using a clustering method.
[0012] In one example, a template region adjacent to the determined reference block is determined, and samples of the determined reference block are classified into one or more sample classes based on sample values of the template region adjacent to the determined reference block in edge detection, where the one or more sample classes include at least a first class including samples at the edge of the determined reference block and a second class including samples in an inner region of the determined reference block.
[0013] In one example, the template region includes one of: (i) multiple rows of neighboring samples above the determined reference block; (ii) multiple columns of neighboring samples to the left of the determined reference block; or (iii) a region in one or more reconstructed neighboring blocks of the determined reference block.
[0014] In one example, samples of the determined reference block are classified based on each of a plurality of partitioning boundary candidates. A plurality of partitioning boundary candidates for the current block are determined, and each of the plurality of partitioning boundary candidates for the current block corresponds to each of the plurality of partitioning boundary candidates for the determined reference block. A TM cost between each of the plurality of partitioning boundary candidates for the determined reference block samples and a corresponding one of the plurality of partitioning boundary candidates for the current block samples is determined. From the plurality of partitioning boundary candidates for the current block, a partitioning boundary corresponding to the smallest TM cost among the determined TM costs is determined. A plurality of partitions for the current block are determined based on the determined partitioning boundaries.
[0015] In one example, the binary mask is determined based on either the mean or median of the reconstructed samples of the template for the current block. A dominant sample group in the reconstructed samples of the template for the current block is determined based on the binary mask. A dominant sample group in the samples of the reference template for each of a plurality of candidate reference blocks is determined based on the binary mask. A TM cost between the dominant group in the reconstructed samples of the template for the current block and the dominant group in the samples of the reference template for each of the plurality of candidate reference blocks is determined. A reference block corresponding to the smallest TM cost among the determined TM costs between the dominant group in the reconstructed samples of the template for the current block and the dominant group in the samples of the reference template for each of the plurality of candidate reference blocks is determined from the plurality of candidate reference blocks.
[0016] In one example, multiple candidate reference blocks are determined based on multiple motion vectors associated with the current block.
[0017] In one example, a TM cost between a reference template of each of a plurality of candidate reference blocks and a template of the current block is determined. From the plurality of candidate reference blocks, a candidate reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block is determined. A motion vector (MV) is determined based on the determined candidate reference block. The motion vector indicates an offset between the current block and the determined candidate reference block. An adjusted MV is determined based on the determined MV and a motion vector difference (MVD). A reference block is determined based on the adjusted MV.
[0018] In one example, a plurality of candidate reference blocks are determined in a current picture. A TM cost between a reference template of each of the plurality of candidate reference blocks and a template of the current block is determined. From the plurality of candidate reference blocks, a candidate reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block is determined. A block vector (BV) is determined based on the determined candidate reference block, where the BV indicates an offset between the current block and the determined candidate reference block. A reference block is determined based on the determined BV.
[0019] According to another aspect of the present disclosure, an apparatus is provided. The apparatus includes a processing circuit. The processing circuit can be configured to perform any of the described methods of video decoding / encoding. For example, the processing circuit is configured to receive a video bitstream including a current block in a current picture. The processing circuit is configured to determine a reference block for the current block from a plurality of candidate reference blocks based on template matching (TM) costs of the plurality of candidate reference blocks. The TM cost indicates a difference between a template of the current block and a reference template of each of the plurality of candidate reference blocks. The processing circuit is configured to classify samples of the determined reference block into a plurality of sample classes. The processing circuit is configured to derive a partitioning pattern for the current block from a plurality of predetermined partitioning patterns based on the determined reference block. The derived partitioning pattern indicates a plurality of partitions of the current block. Each of the plurality of classes of samples of the determined reference block corresponds to a respective partition of the plurality of partitions of the current block. The processing circuit is configured to reconstruct the current block based on the derived partitioning pattern of the current block.
[0020] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding. [Brief explanation of the drawings]
[0021] Further features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings.
[0022] [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100).
[0023] [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder.
[0024] [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder.
[0025] [Figure 4] 1 shows exemplary geometric partition mode (GPM) partitions grouped by identical angles.
[0026] [Figure 5A] 1 illustrates an example GPM with intra- and inter-prediction. [Figure 5B] 1 illustrates an example GPM with intra- and inter-prediction. [Figure 5C] 1 illustrates an example GPM with intra- and inter-prediction.
[0027] [Figure 5D] 1 illustrates an example GPM with intra-prediction and intra-prediction.
[0028] [Figure 6A] 1 illustrates a first exemplary template of a block according to some embodiments of the present disclosure.
[0029] [Figure 6B] 10 illustrates a second exemplary template of a block, according to some embodiments of the present disclosure.
[0030] [Figure 7] 1 shows a flowchart outlining a decoding process according to some embodiments of the present disclosure.
[0031] [Figure 8] 1 shows a flowchart outlining an encoding process according to some embodiments of the present disclosure.
[0032] [Figure 9] FIG. 1 is a schematic diagram of a computer system, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0033] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is a video encoder and video decoder in a streaming environment, which is one example of an application of the disclosed subject matter. The disclosed subject matter is equally applicable to, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., other video-enabled applications, etc.
[0034] The video processing system 100 includes a video source 101, e.g., a capture subsystem 113, which may include a digital camera, that generates an uncompressed video picture stream 102. In one example, the video picture stream 102 includes samples captured by the digital camera. The video picture stream 102, shown in bold to emphasize its high data volume when compared to the encoded video data 104 (or coded video bitstream), may be processed by an electronic device 120 that includes a video encoder 103 coupled to the video source 101. The video encoder 103 may include hardware, software, or a combination thereof, and may enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data 104 (or coded video bitstream), shown in thin to emphasize its low data volume when compared to the video picture stream 102, may be stored on a streaming server 105 for future use. One or more streaming client subsystems, such as the client subsystems 106 and 108 of FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 may include a video decoder 110, for example, within an electronic device 130. The video decoder 110 decodes the input copy 107 of the encoded video data and generates an output video picture stream 111 that can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is known informally as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0035] It should be noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and electronic device 130 may also include a video encoder (not shown).
[0036] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used in place of the video decoder (110) in the example of FIG. 1.
[0037] The receiver (231) can receive one or more coded video sequences, e.g., included in a bitstream, to be decoded by the video decoder (210). In an embodiment, one coded video sequence is received at a time, with the decoding of each coded video sequence being independent of the decoding of other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (231) may receive the coded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which may be forwarded to respective using entities (not shown). The receiver (231) may separate the coded video sequences from other data. To eliminate network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter, "parser (220)"). In certain applications, the buffer memory (215) is part of the video decoder (210). Alternatively, it may be external to the video decoder 210 (not shown). Still alternatively, there may be a buffer memory (not shown) external to the video decoder 210, e.g., to remove network jitter, in addition to another buffer memory 215 internal to the video decoder 210, e.g., to handle playout timing. When the receiver 231 is receiving data controllably from a store / forward device of sufficient bandwidth or from an isosynchronous network, the buffer memory 215 may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory 215 may be needed, but it may be relatively large, advantageously of adaptive size, and implemented at least in part in an operating system or similar element (not shown) external to the video decoder 210.
[0038] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and information for controlling a rendering device, such as a render device (212) (e.g., a display screen), which may not be an integral part of the electronic device (230) but may be coupled to the electronic device (230) as shown in FIG. 2. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context dependency, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (220) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.
[0039] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to generate symbols (221).
[0040] The reconstruction of the symbols (221) may include several different units, depending on the type of coded video picture or portion thereof (e.g., inter- and intra-picture, inter- and intra-block) and other factors. Which units are included and how can be controlled by group control information parsed by the parser (220) from the coded video sequence. The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.
[0041] Beyond the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0042] The first unit is a scalar / inverse transform unit 251. The scalar / inverse transform unit (251) receives quantized transform coefficients and control information from the parser (220) as symbols (221), including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. The scalar / inverse transform unit (251) can output blocks containing sample values that can be input to an aggregator (255).
[0043] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers the reconstructed current picture partially and / or completely. The aggregator (255), in some cases, adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a sample-by-sample basis.
[0044] In other cases, the output samples of the scaler / inverse transform unit (251) may relate to an inter-coded, possibly motion-compensated, block. In such cases, the motion-compensated prediction unit (253) can access the reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) associated with the block, these samples may be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) to generate output sample information (in this case, referred to as residual samples or residual signals). The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches prediction samples can be controlled by the motion-compensated prediction unit (253)'s available motion vectors, e.g., in the form of symbols (221) that may have X, Y, and reference picture components. Motion compensation may include interpolation of sample values fetched from the reference picture memory (257) when sub-sample accurate motion vectors are in use, motion vector prediction mechanisms, etc.
[0045] The output samples of the aggregator (255) may undergo various loop filtering techniques in a loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression may not only be responsive to previously reconstructed, loop-filtered sample values, but may also be responsive to meta-information obtained during the decoding of previous portions (in decoding order) of the coded picture or coded video sequence.
[0046] The output of the loop filter unit (256) may be a sample stream that can be output to a render device (212) and stored in a reference picture memory (257) for use in future inter-picture prediction.
[0047] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a fresh current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.
[0048] The video decoder 210 may perform decoding operations in accordance with a standard or predetermined video compression technology, such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard in use, in the sense that the coded video sequence conforms to both the video compression technology or standard and a profile documented in the video compression technology or standard. Specifically, a profile may select certain tools from the full set of tools available in the video compression technology or standard as tools usable only under the profile. Compliance may also require that the complexity of the coded video sequence be within limits defined by the level of the video compression technology or standard. In some cases, the level may limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained through a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled within the coded video sequence.
[0049] In embodiments, the receiver (231) may receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0050] 3 shows an exemplary block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmitting circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.
[0051] The video encoder (303) may receive video samples from a video source (301) (which, in the example of FIG. 3, is not part of the electronic device (320)) that may capture video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0052] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media presentation system, the video source (301) may be a storage device that stores previously prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed sequentially, give the appearance of motion. The pictures themselves may be organized as a spatial array of pixels. Each pixel may contain one or more samples, depending on the sampling structure, color space, etc., in use. The following discussion focuses on samples.
[0053] According to an embodiment, the video encoder (303) may code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other required time constraints. Enforcing the appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is operatively coupled to other functional units, described below. The coupling is not shown for clarity. Parameters set by the controller (350) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured with other appropriate functionality associated with the video encoder (303) optimized for a particular system design.
[0054] In some embodiments, the video encoder (303) is configured to operate within a coding loop. As a highly simplified explanation, in one example, the coding loop may include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (333) built into the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to that created by a (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). When decoding the symbol stream yields bit-exact results independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" exactly the same sample values as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronism (and the resulting drift when synchronism cannot be maintained, for example due to channel errors) is similarly used in several related techniques.
[0055] The operation of the "local" decoder (333) may be the same as a "remote" decoder, such as the video decoder (210) described in detail above in connection with Figure 2. However, and referring briefly to Figure 2 as well, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333) because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (345) and parser (220) may be lossless.
[0056] In embodiments, decoder techniques, excluding analysis / entropy decoding, present in a decoder are present in the corresponding encoder in the same or substantially the same functional form. Therefore, the subject matter of the described disclosure focuses on decoder operation. A description of the encoder techniques can be omitted, as they are the reverse of the decoder techniques, which are described generically. In certain areas, more detailed descriptions are provided below.
[0057] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as "reference pictures." In this method, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as prediction references for the input picture.
[0058] The local video decoder (333) may decode coded video data of pictures that may be designated as reference pictures based on symbols generated by the source coder (330). The operation of the coding engine (332) may advantageously be lossy. When the coded video data is decoded in a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) may replicate the decoding process that may be performed by a video decoder on the reference pictures, resulting in reconstructed reference pictures to be stored in the reference picture memory (334). In this way, the video encoder (303) may store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures obtained by a far-end video decoder (absent transmission errors).
[0059] The predictor (335) may perform a predictive search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for sample data (such as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that can serve as appropriate prediction references for the new picture. The predictor (335) may operate on a sample block-pixel block basis to find an appropriate prediction reference. In some examples, an input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334), as determined by the search results obtained by the predictor (335).
[0060] The control unit (350) may manage the coding operations of the source coder (330), including, for example, setting parameters and subgroup parameters used for encoding the video data.
[0061] The output of all of the aforementioned functional units may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0062] The transmitter (340) may buffer the coded video sequence produced by the entropy coder (345) for transmission over a communication channel (360), which may be a hardware / software link to a storage device that may store the coded video data. The transmitter (340) may merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0063] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:
[0064] Intra-pictures (I-pictures) may be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of intra-pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0065] A predictive picture (P-picture) may be a picture that can be coded and decoded using intra- or inter-prediction, in most cases using motion vectors and reference indices to predict the sample values of each block.
[0066] A Bi-directionally Predictive Picture (B Picture) may be coded and decoded using intra- or inter-prediction, using two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi-predictive picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0067] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to each picture of the block. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0068] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. The coded video data may therefore conform to a syntax specified by the video coding technique or standard being used.
[0069] In one embodiment, the transmitter (340) may transmit additional data along with the coded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0070] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (sometimes abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded is called the current picture and is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are in use.
[0071] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. Bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, both of which precede the current picture in the video in decoding order (but may be past and future, respectively, in display order). A block in the current picture can be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block can be predicted by a combination of the first and second reference blocks.
[0072] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0073] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are partitioned into coding tree units (CTUs) for compression. CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Typically, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU can be partitioned into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the CU's prediction type, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels, such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0074] It is noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.
[0075] This disclosure includes aspects related to deriving partition patterns using template matching.
[0076] Geometric partitioning mode (GPM) can be applied to inter prediction, as in VVC. Geometric partitioning mode can be signaled using a flag (e.g., a CU-level flag) as a type of merge mode. Other types of merge modes include normal merge mode, merge motion vector differences (MMVD) mode, combined inter and intra prediction (CIIP) mode, and sub-block merge mode. In one example, w×h=2 m ×2 n 64 partitions are supported by the GPM for each possible CU of size m, n∈{3...6}, excluding 8x64 and 6x48.
[0077] Using GPM, a CU can be divided into two parts by a geometrically arranged line (e.g., 404). The position of the division line can be mathematically derived from the angle and offset parameters of a particular partition. FIG. 4 shows an exemplary GPM division grouped by the same angle. As shown in FIG. 4, each CU (e.g., 402) can include a respective group of partition lines. Each partition line (e.g., 404) can indicate a partitioning method and correspond to a respective offset. The group of partition lines in each CU can include up to four partition lines corresponding to a respective angle. Each part of the geometric partition within a CU can be inter-predicted using a respective motion. Uni-prediction can be allowed for each partition. For example, each part (each partition) can have one motion vector and one reference index. Uni-prediction motion constraints can be applied to ensure that two motion-compensated predictions are required for each CU, similar to traditional bi-prediction.
[0078] If the geometric partitioning mode is used for the current CU, a geometric partitioning index (e.g., angle and offset) indicating the partition mode of the geometric partition and two merge indices (one for each partition) can be further signaled. The maximum GPM candidate size can be explicitly signaled in the sequence parameter set (SPS), which specifies the syntactic binarization of the GPM merge indices. After each geometric partition is predicted, the sample values along the geometric partition edges can be adjusted using a fusion process with adaptive weights. Thus, a prediction signal for the entire CU based on the GPM can be obtained. In other prediction modes, transformation and quantization processes can be applied to the entire CU. Furthermore, the motion field of the CU predicted using the geometric partitioning mode can be stored.
[0079] The geometric partition mode can be stored in the motion field. For example, Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and Mv, the combination of Mv1 and Mv2, can be stored in the motion field of the geometric partition mode coded by the CU. The type of motion vector stored for each position in the motion field can be determined by equation (1) as follows:
number
number
[0080] Template matching can be applied to GPM. When GPM mode is enabled for a CU, a CU-level flag can be signaled to indicate whether TM applies to both geometric partitions. The motion information for each geometric partition can be refined using TM. When TM is selected, a template can be constructed using neighboring samples to the left, above, or to the left and above, according to the partition angle, as shown in Table 1. Then, using the same search pattern in merge mode with the half-pel interpolation filter disabled, the motion can be refined by minimizing the difference between the current template and the template in the reference image. Table 1 shows example templates for the first and second geometric partitions in GPM, where A represents the use of the top sample, L represents the use of the left sample, and L+A represents the use of both the left and top samples. [Table 1] Example of partition template in GPM [Table 1]
[0081] In one embodiment, the GPM candidate list can be constructed as follows: (1) Interleaved List-0MV and List-1MV candidates can be directly derived from the normal merge candidate list, and List-0MV candidates can have higher priority than List-1MV candidates. A pruning method with an adaptive threshold based on the current CU size can be applied to remove redundant MV candidates. (2) Furthermore, interleaved List-1MV and List-0MV candidates can be directly derived from the regular merge candidate list, and List-1MV candidates can have higher priority than List-0MV candidates. The same pruning method with adaptive thresholds can also be applied to remove redundant MV candidates. (3) The GPM candidate list can be padded with 0MV candidates until it is full.
[0082] GPM-MMVD and GPM-TM can be enabled exclusively for one GPM CU. For example, the GPM-MMVD syntax can be signaled first. If both GPM-MMVD control flags are equal to false, it indicates that GPM-MMVD is disabled for two GPM partitions, and a GPM-TM flag can be signaled to indicate whether template matching applies to the two GPM partitions. Otherwise, if at least one GPM-MMVD flag is equal to true, the value of the GPM-TM flag can be inferred as false.
[0083] When GPM using inter and intra prediction is applied, the final predicted sample can be generated by weighting the inter predicted sample, the predicted sample, and the intra predicted sample for each region separated by the GPM. The inter predicted sample can be derived by the inter GPM. The intra predicted sample can be derived by an intra prediction mode (IPM) candidate list and an index signaled from the encoder. The IPM candidate list size can be predefined as 3. Available IPM candidates can include a parallel angle mode (e.g., parallel mode) relative to the GPM block boundary, a perpendicular angle mode (e.g., vertical mode) relative to the GPM block boundary, and a planar mode, as shown in Figures 5A, 5B, and 5C, respectively. Furthermore, as shown in Figure 5D, GPM with intra and intra prediction can be limited to reduce the signaling overhead for IPM and avoid increasing the size of the intra prediction circuit on the hardware decoder. Furthermore, direct motion vectors and IPM storage on the GPM fusion region can be introduced to further improve coding performance.
[0084] In decoder-side intra mode derivation (DIMD) and neighboring mode-based IPM derivation, parallel modes can be registered first. Therefore, if the same IPM candidate is not in the IPM candidate list, up to two IPM candidates derived from the DIMD method and / or neighboring blocks can be registered. In neighboring mode-based IPM derivation, up to five positions can be defined to determine available neighboring blocks. However, as shown in Table 2, the five positions can be limited by the angle of the GPM block boundary when GPM with template matching (GPM-TM) is applied. As shown in Table 2, the available neighboring block positions for IPM candidate derivation can be defined based on the angle of the GPM block boundary. A and L indicate the upper and left sides of the prediction block. [Table 2] Location of available neighboring blocks for IPM candidate derivation [Table 2] In an aspect, GPM-intra can be combined with GPM with merge motion vector difference (GPM-MMVD). In an aspect, template-based intra mode derivation (TIMD) can be used to derive IPM candidates for GPM-intra to further improve coding performance. In an aspect, parallel modes can be registered first, and then TIMD, DIMD, and IPM candidates for neighboring blocks can be registered subsequently.
[0085] Current GPM designs allow only a limited set of predefined partitioning patterns using straight lines as partitioning boundaries, which may not always model the most efficient partitioning patterns, such as irregular partitioning patterns. Motion vectors can be used to identify blocks in a reference picture or a current picture, and the reconstructed values of the identified blocks can be used to derive an adaptive partitioning pattern. However, in relevant instances, signaling of additional motion vectors may be required, which may be costly to specify a partitioning pattern.
[0086] This disclosure may provide derivation of a partitioning pattern based on template matching. In this disclosure, a template may refer to samples in the neighborhood of a block, such as neighboring samples above, to the left, to the right, and / or below the block. Exemplary templates may be shown in FIGS. 6A and 6B. As shown in FIG. 6A, the template (604) may include both an upper neighboring reconstructed sample (606) and a left neighboring reconstructed sample (608) of the current block (602). The template (604) may also include a reconstructed sample (610) in the upper left corner of the current block (602). As shown in FIG. 6B, the template (611) may include an upper template (614) above the current block (612) and a left template (616) to the left of the current block (612). In one aspect, if the right or bottom sample of the current block in the current frame has already been reconstructed, the right or bottom sample can be used as the template.
[0087] In the present disclosure, template matching may be used to identify a reference block in a reference picture. The identified reference block may be used to derive a geometric partitioning pattern for a current block. Each partitioning of the derived geometric partitioning pattern for the current block may apply a different motion. In one example, multiple partitionings of the current block may be determined based on the derived geometric partitioning pattern. The derived geometric partitioning pattern may divide samples of the current block into multiple groups. Each of the multiple partitions may correspond to a respective group of samples of the current block.
[0088] In one example, a reference block can be determined for a current block from a plurality of candidate reference blocks based on template matching (TM) costs of the plurality of candidate reference blocks. The TM cost can indicate a difference between a template of the current block and a reference template of each of the plurality of candidate reference blocks. Samples (e.g., reconstructed samples) of the determined reference block can be classified into a plurality of reconstructed sample classes. A plurality of partitions of the current block corresponding to the plurality of reconstructed sample classes of the determined reference block can be determined. Each of the plurality of sample classes of the determined reference block can correspond to each of the plurality of partitions of the current block.
[0089] In one aspect, a template for a current block can be constructed using neighboring samples, such as neighboring samples to the left, above, or left and above of the current block. A template for a candidate reference block in a reference picture can be constructed using neighboring samples, such as neighboring samples to the left, above, or left and above of the reference block. Among the reference blocks, a reference block that minimizes the cost between the current template (e.g., the template of the current block) and the reference template (e.g., the template of the reference block) can be identified. The identified reference block corresponding to the minimized cost can be used to derive a partitioning pattern for the current block.
[0090] As an example, an example template for the current block can be shown in FIGS. 6A and 6B.
[0091] In one example, an initial reference block can be determined based on motion vector information included in the received video bitstream. A plurality of candidate reference blocks can be determined within a search range of the initial reference block. A TM cost between a reference template of each of the plurality of candidate reference blocks and a template of the current block can be determined. From the plurality of candidate reference blocks, a reference block corresponding to the smallest TM cost can be determined among the determined TM costs between a reference template of each of the plurality of candidate reference blocks and a template of the current block.
[0092] In one aspect, reconstructed samples of identified blocks, such as blocks identified based on template matching, can be used to derive partitioning patterns.
[0093] In one aspect, a binary image segmentation algorithm can be applied to the identified blocks, and the output binary map of the binary image segmentation algorithm can be used as the partitioning pattern.
[0094] In one example, the reconstructed samples of the determined reference block are classified into a first reconstructed sample class and a second reconstructed sample class based on a binary image segmentation algorithm.
[0095] In one aspect, a threshold value for the reconstructed samples of the identified block can be first calculated. The identified block can be divided into two partitions. The first partition can include samples having values (or sample values) greater than (or equal to or greater than) a threshold, and the second partition can include samples having values (or sample values) less than (or equal to or less than) the threshold. In one example, the threshold value can be the mean value of the reconstructed samples of the identified block. In one example, the threshold value can be the median value of the reconstructed samples of the identified block.
[0096] In one example, the reconstructed samples of the determined reference block can be classified into a first reconstructed sample class having sample values greater than a threshold and a second reconstructed sample class having sample values less than the threshold. The threshold can be based on one or more sample values of the identified block. In one example, the threshold can be one of an average value of the reconstructed samples and a median value of the reconstructed samples.
[0097] In one aspect, a clustering method can be applied to the reconstructed samples of the identified block to derive two or more classes of samples, and the samples associated with each class can constitute one of a plurality of partitions.
[0098] In one example, the clustering method may include a hierarchical model, a partitioning model, a density-based model, a model-based model, and a grid-based model, or any other suitable clustering method.
[0099] In one embodiment, an edge detection method can be applied to classify the reconstructed samples of the identified block into two or more sample classes, and the samples associated with each class can constitute one of a number of partitions.
[0100] In one example, based on edge detection, the reconstructed samples of the identified block can be classified into an edge group and an interior group. The edge group can include reconstructed samples at the edges of the identified block, and the interior group can include reconstructed samples in the interior region of the identified block. The interior region can be surrounded by the edges of the identified block.
[0101] In one example, a template region adjacent to the determined reference block can be determined. The reconstructed samples of the determined reference block can be classified into one or more reconstructed sample classes based on sample values of the template region adjacent to the determined reference block in edge detection. The one or more reconstructed sample classes can include at least a first reconstructed class including samples at the edge of the determined reference block and a second reconstructed class including samples in an interior region of the determined reference block.
[0102] In one example, the edge detection method may include Sobel, Canny, Prewitt, Roberts, fuzzy logic, or any other suitable edge detection method.
[0103] In one example, edge detection may include various mathematical methods (such as Sobel, Canny, Prewitt, Roberts, fuzzy logic, etc.) that aim to identify edges and / or curves in digital images where the image brightness changes abruptly or has discontinuities.
[0104] In one example, the template region used for edge detection can extend over multiple rows or columns of samples, or can extend from one or more reconstructed neighboring blocks.
[0105] In one example, a template region adjacent to the determined reference block can be determined, and the reconstructed samples of the determined reference block can be classified into one or more reconstructed sample classes based on the sample values of the template region adjacent to the determined reference block in edge detection.
[0106] In one aspect, if reconstruction samples are available, the width or height of the template region used for edge detection can be larger than the width or height of the current block. For example, the template region can have a size equal to twice the block width or twice the block height.
[0107] In one example, the template region may include one of: (i) multiple rows of neighboring samples above the determined reference block; (ii) multiple columns of neighboring samples to the left of the determined reference block; or (iii) a region in one or more reconstructed neighboring blocks of the determined reference block.
[0108] In one aspect, multiple partitioning boundary candidates may be applied to the reconstructed samples of the identified block. The partitioning boundary candidate that minimizes a predefined cost score may be derived as a GPM mode partitioning boundary for the current block. In one example, multiple partitions of the current block may be determined based on the derived partitioning boundaries. The derived partitioning boundaries may divide the samples of the current block into multiple groups. Each of the multiple partitions may correspond to a respective group of samples of the current block.
[0109] In one example, the reconstructed sample of the determined reference block can be classified based on each of a plurality of partitioning boundary candidates. A plurality of partitioning boundary candidates for the current block are determined, and each of the plurality of partitioning boundary candidates for the current block can correspond to each of the plurality of partitioning boundary candidates for the determined reference block. A TM cost between each of the plurality of partitioning boundary candidates for the reconstructed sample of the determined reference block and a corresponding one of the plurality of partitioning boundary candidates for the sample of the current block can be determined. From the plurality of partitioning boundary candidates for the current block, a partitioning boundary corresponding to the smallest TM cost among the determined TM costs can be determined. A plurality of partitions for the current block can be determined based on the determined partitioning boundary.
[0110] In one aspect, a binary mask can be derived based on the current template, and a template matching operation can be applied to both the current template and the reference template using the derived binary mask to identify blocks in the reference picture. The identified blocks can be used to derive a geometric partitioning pattern for the current block, and each partitioning can apply a different motion.
[0111] In one example, a binary mask can define a region of interest (ROI) of an image (or block). In the binary mask, image pixels (or samples) of the image having a first mask pixel value (e.g., 1) can belong to the ROI (or majority group). Image pixels of the image having a second mask pixel value (e.g., 0) can belong to the background (or minority group). In one example, the ROI of an image can include samples of the image that have sample values greater than a threshold sample value of the image. The threshold sample value of the image can be a mean sample value, a median sample value of the image, or a mean sample value of the image.
[0112] The binary mask may be determined based on the reconstructed sample values of the template for the current block. In one example, one of the mean and median values of the reconstructed samples of the template for the current block is used to determine the binary mask. The dominant sample group in the reconstructed samples of the template for the current block may be determined based on the frequency of the sample values or the range of the sample values. For example, the dominant sample group may be determined based on the binary mask. The dominant sample group in the sample of the reference template for each of the multiple candidate reference blocks may be determined based on the binary mask. The TM cost between the dominant group in the reconstructed sample of the template for the current block and the dominant group in the sample of the reference template for each of the multiple candidate reference blocks may be determined. From the multiple candidate reference blocks, the reference block corresponding to the smallest TM cost may be determined among the determined TM costs between the dominant group in the reconstructed sample of the template for the current block and the dominant group in the sample of the reference template for each of the multiple candidate reference blocks.
[0113] In one aspect, a threshold value for the reconstruction samples of the current template can be first calculated. At least one binary mask can be derived by labeling each sample of the current template based on whether each sample is greater than (or equal to) the threshold value. When calculating the template matching cost, samples with dominant labels can be used, and the remaining samples may not be considered for calculating the template matching cost. A dominant label indicating a majority group can indicate a label value (e.g., 0 or 1) associated with more (or fewer) samples than another label value.
[0114] In one example, the dominant label can indicate a dominant group (or majority group) of reconstructed samples of the current template. For example, if there are more reconstructed samples with sample values equal to or greater than a threshold than there are reconstructed samples with sample values less than the threshold, the dominant group can include the reconstructed samples with sample values equal to or greater than the threshold.
[0115] The threshold can be derived from one or more values of the reconstructed samples of the current template. In one embodiment, the threshold can be the mean value of the reconstructed samples of the current template. In one embodiment, the threshold can be the median value of the reconstructed samples of the current template.
[0116] In one aspect, the motion vectors between the current template and the reference template can be used for motion compensation of the associated geometric partition. The motion vectors can indicate the offset between the dominant group in the current template and the dominant group in the reference template according to a binary mask.
[0117] In one aspect, the partitioning pattern for the current block can be derived using a partitioning pattern that can be associated with an identified block coded by the GPM or that can be derived using a template matching method in a reference picture.
[0118] In one aspect, template matching can be applied using a motion vector associated with the current block to identify candidate reference blocks for template matching.
[0119] In one example, multiple candidate reference blocks can be determined based on multiple motion vectors associated with the current block.
[0120] In one aspect, a motion vector difference (MVD) can be further signaled on top of template matching to identify a reference block. The MVD can be coded using a simplified MVD coding scheme, such as MMVD. Thus, the reference block can be determined based on the MV, which is equal to the sum of the determined template and the TM determined by the MVD.
[0121] In one example, a TM cost between a reference template of each of a plurality of candidate reference blocks and a template of the current block may be determined. From the plurality of candidate reference blocks, a reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block may be determined. A motion vector (MV) may be determined based on the determined candidate reference block. The motion vector may indicate an offset between the current block and the determined candidate reference block. An adjusted MV may be determined based on the determined MV and a motion vector difference (MVD). A reference block may be determined based on the adjusted MV.
[0122] In one aspect, template matching can be applied to identify blocks in the current picture, and the identified blocks can be used to derive the partitioning pattern. Candidate reference blocks can be determined within already reconstructed regions of the current picture.
[0123] In one aspect, similar to block vectors (BVs) in intra block copy (IBC) mode, the derived BVs (and / or corresponding BV offsets) pointing to reference blocks may be at integer sample (e.g., luma sample) resolution.
[0124] In one aspect, the BV (and / or the corresponding BV offset) pointing to the reference block may be at sub-pixel resolution.
[0125] In one example, a plurality of candidate reference blocks can be determined in the current picture. A TM cost between a reference template of each of the plurality of candidate reference blocks and the template of the current block can be determined. From the plurality of candidate reference blocks, a reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block can be determined. A block vector (BV) can be determined based on the determined candidate reference blocks, and the BV can indicate an offset between the current block and the determined candidate reference block. A reference block is determined based on the determined BV.
[0126] In one aspect, an additional MV can be signaled to indicate the starting position of the identified block in the reference picture. Furthermore, template matching can be applied on the starting position to find the best matching identified block. The identified block can be further used to derive a partitioning pattern.
[0127] 7 shows a flowchart outlining a process (700) according to one embodiment of the present disclosure. The process (700) may be used in a video decoder. In various embodiments, the process (700) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (700) is implemented with software instructions, and thus the processing circuit performs the process (700) when it executes the software instructions. The process begins at (S701) and proceeds to (S710).
[0128] At (S710), a video bitstream including a current block in a current picture is received.
[0129] At (S720), a reference block is determined for the current block from a plurality of candidate reference blocks based on template matching (TM) costs of the plurality of candidate reference blocks, where the TM cost indicates a difference between the template of the current block and each reference template of the plurality of candidate reference blocks.
[0130] At (S730), the samples of the determined reference block are classified into a plurality of sample classes.
[0131] At (S740), a partitioning pattern of the current block is derived from a plurality of predetermined partitioning patterns based on the determined reference block, the derived partitioning pattern indicates a plurality of partitions of the current block, and each of the plurality of classes of samples of the determined reference block corresponds to each of the plurality of partitions of the current block.
[0132] At (S750), the current block is reconstructed based on the derived partitioning pattern of the current block.
[0133] In one aspect, the plurality of candidate reference blocks are from one of the current picture and the reference picture of the current block. The samples of the determined reference block are reconstructed samples.
[0134] In one example, an initial reference block is determined based on motion vector information included in a received video bitstream. A plurality of candidate reference blocks are determined within a search range of the initial reference block. A TM cost between a reference template of each of the plurality of candidate reference blocks and a template of a current block is determined. From the plurality of candidate reference blocks, a reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block is determined.
[0135] In one example, the samples of the determined reference block are classified into a first sample class and a second sample class based on a binary image segmentation algorithm.
[0136] In one example, the samples of the determined reference block are classified into a first sample class having sample values greater than a threshold value and a second sample class having sample values less than a threshold value, the threshold value being either the mean value or the median value of the samples.
[0137] In one example, the samples of the determined reference block are clustered into one or more sample classes using a clustering method.
[0138] In one example, a template region adjacent to the determined reference block is determined, and samples of the determined reference block are classified into one or more sample classes based on sample values of the template region adjacent to the determined reference block in edge detection, where the one or more sample classes include at least a first class including samples at the edge of the determined reference block and a second class including samples in an inner region of the determined reference block.
[0139] In one example, the template region includes one of: (i) multiple rows of neighboring samples above the determined reference block; (ii) multiple columns of neighboring samples to the left of the determined reference block; or (iii) a region in one or more reconstructed neighboring blocks of the determined reference block.
[0140] In one example, samples of the determined reference block are classified based on each of a plurality of partitioning boundary candidates. A plurality of partitioning boundary candidates for the current block are determined, and each of the plurality of partitioning boundary candidates for the current block corresponds to each of the plurality of partitioning boundary candidates for the determined reference block. A TM cost between each of the plurality of partitioning boundary candidates for the determined reference block samples and a corresponding one of the plurality of partitioning boundary candidates for the current block samples is determined. From the plurality of partitioning boundary candidates for the current block, a partitioning boundary corresponding to the smallest TM cost among the determined TM costs is determined. A plurality of partitions for the current block are determined based on the determined partitioning boundaries.
[0141] In one example, the binary mask is determined based on either the mean or median of the reconstructed samples of the template for the current block. A dominant sample group in the reconstructed samples of the template for the current block is determined based on the binary mask. A dominant sample group in the samples of the reference template for each of a plurality of candidate reference blocks is determined based on the binary mask. A TM cost between the dominant group in the reconstructed samples of the template for the current block and the dominant group in the samples of the reference template for each of the plurality of candidate reference blocks is determined. A reference block corresponding to the smallest TM cost among the determined TM costs between the dominant group in the reconstructed samples of the template for the current block and the dominant group in the samples of the reference template for each of the plurality of candidate reference blocks is determined from the plurality of candidate reference blocks.
[0142] In one example, multiple candidate reference blocks are determined based on multiple motion vectors associated with the current block.
[0143] In one example, a TM cost between a reference template of each of a plurality of candidate reference blocks and a template of the current block is determined. From the plurality of candidate reference blocks, a candidate reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block is determined. A motion vector (MV) is determined based on the determined candidate reference block. The motion vector indicates an offset between the current block and the determined candidate reference block. An adjusted MV is determined based on the determined MV and a motion vector difference (MVD). A reference block is determined based on the adjusted MV.
[0144] In one example, a plurality of candidate reference blocks are determined in a current picture. A TM cost between a reference template of each of the plurality of candidate reference blocks and a template of a current block is determined. From the plurality of candidate reference blocks, a candidate reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block is determined. A block vector (BV) is determined based on the determined candidate reference block, where the BV indicates an offset between the current block and the determined candidate reference block. A reference block is determined based on the determined BV.
[0145] Next, the process proceeds to (S799) and ends.
[0146] The process 700 may be adapted as appropriate. Steps of the process 700 may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.
[0147] 8 shows a flowchart outlining a process (800) according to one embodiment of the present disclosure. The process (800) can be used in a video encoder. In various embodiments, the process (800) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (800) is implemented by software instructions, and thus the processing circuit performs the process (800) when it executes the software instructions. The process begins at (S801) and proceeds to (S810).
[0148] In (S810), a reference block is determined from a plurality of candidate reference blocks for a current block in a current picture based on TM costs of the plurality of candidate reference blocks, where the TM cost indicates a difference between a template of the current block and each reference template of the plurality of candidate reference blocks.
[0149] In (S820), the reconstructed samples of the determined reference block are classified into a plurality of reconstructed sample classes.
[0150] At step S830, a plurality of partitions of the current block corresponding to the plurality of reconstructed sample classes of the determined reference block are determined, where each of the plurality of reconstructed sample classes of the determined reference block corresponds to each of the plurality of partitions of the current block.
[0151] At (S840), the current block is reconstructed based on the determined partitions of the current block.
[0152] Next, the process proceeds to (S899) and ends.
[0153] The process 800 may be adapted as appropriate. Steps of the process 800 may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.
[0154] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 9 illustrates a computer system (900) suitable for implementing certain embodiments of the subject matter of this disclosure.
[0155] Computer software can be coded using any suitable machine code or computer language that can be processed by mechanisms such as assembly, compilation, linking, etc. to generate code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly or through interpretation, microcode execution, etc.
[0156] The instructions may be executed by a variety of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0157] 9 of the computer system (900) are exemplary in nature and do not suggest any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Furthermore, the arrangement of components should not be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of the computer system (900).
[0158] The computer system (900) may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, through sensory input (e.g., keystrokes, swipes, data grabbing actions), audio input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a digital camera), and video (including, for example, two-dimensional video, three-dimensional video, and stereoscopic video).
[0159] The input human interface devices may include one or more of a keyboard (901), a mouse (902), a trackpad (903), a touchscreen (910), a data grab (not shown), a joystick (905), a microphone (906), a scanner (907), and a camera (908) (only one of which is shown).
[0160] The computer system (900) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses through, for example, sensory output, sound, light, and smell / taste. Such human interface output devices may include sensory output devices (e.g., sensory feedback via a touchscreen (910), a data grab (not shown), or a joystick (905; however, sensory feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers (909), headphones (not shown)), and visual output devices (e.g., a screen (910), including a CRT screen, an LCD screen, a plasma screen, and an OLED screen, each with or without touchscreen input capability and each with or without sensory feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional or higher-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown)).
[0161] The computer system (900) may also include human-accessible storage and associated media such as optical media including CD / DVD ROM / RW (920) with media such as CD / DVD (921), thumb drives (922), removable hard drives or solid state drives (923), legacy magnetic media such as tape and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.
[0162] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.
[0163] The computer system 900 may also include an interface 954 to one or more communication networks 955. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, latency-tolerant, and the like. Examples of networks include local area networks such as Ethernet; cellular networks, including WLAN, GSM, 3G, 4G, 5G, LTE, and the like; TV wired or wireless wide-area digital networks, including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks, including CAN Bus. Particular networks generally require an external network interface that is attached to a particular general-purpose data port or peripheral bus 949 (e.g., a USB port on the computer system 900). Others are generally integrated into the core of the computer system 900 by attachment to a system bus, as described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Using these networks, the computer system 900 can communicate with other entities. Such communication may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a specific CANbus device), or two-way to other computer systems, for example, using a local or wide area digital network. Specific protocols and protocol stacks may be used in each of the above networks and network interfaces.
[0164] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to the core (940) of the computer system (900).
[0165] The core (940) may include one or more central processing units (CPUs) (941), graphics processing units (GPUs) (942), dedicated programmable processing units (943) in the form of FPGAs, task-specific hardware accelerators (944), graphics adapters (950), etc. These devices may be connected through a system bus (948), along with read-only memory (ROM) (945), random access memory (946), and internal mass storage devices (947) such as internal non-user-accessible hard drives, SSDs, etc. In some computer systems, the system bus (948) is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the core's system bus (948) or through a peripheral bus (949). In an example, a screen (910) can be connected to the graphics adapter (950). Peripheral bus architectures include PCI, USB, etc.
[0166] The CPU (941), GPU (942), FPGA (943), and accelerator (944) may execute specific instructions that may be combined to generate the aforementioned computer code. The computer code may be stored in ROM (945) or RAM (946). Temporary data may also be stored in RAM (946), while permanent data may be stored, for example, in an internal mass storage device (947). Rapid storage and retrieval from any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU (941), GPU (942), mass storage device (947), ROM (945), RAM (946), etc.
[0167] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.
[0168] As an example and not by way of limitation, the computer system 900 having the architecture, and specifically the core 940, can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be specific storage of the core 940 of a non-transitory nature, such as the core's internal mass storage 947 or ROM 945, as well as media associated with user-accessible mass storage devices such as those described above. Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core 940. The computer-readable media can include one or more memory devices or chips, depending on the particular needs. The software can cause the core 940, and specifically the processor therein (including a CPU, GPU, FPGA, etc.), to perform specific processes or portions of specific processes described herein, including defining and modifying data structures stored in RAM 946 according to software-defined operations. Additionally or alternatively, the computer system may provide functionality as a result of implementation in hardwired or other circuitry (e.g., accelerator 944) that can operate in conjunction with or in place of software to perform certain processes or portions of certain processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include, where appropriate, circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that implements logic for execution, or both. The present disclosure includes any appropriate combination of hardware and software.
[0169] The use of "at least one" or "one" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). The use of "one" does not exclude any combination of the listed elements, where applicable, such as when the elements are not mutually exclusive.
[0170] While this disclosure has described several exemplary embodiments, alterations, permutations, and various substitute equivalents exist, and are encompassed within the scope of this disclosure. Those skilled in the art will appreciate that numerous systems and methods can be devised that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore are within the spirit and scope of the present disclosure.
Claims
1. 1. A method of video decoding, comprising: receiving a video bitstream including a current block in a current picture; determining a reference block from a plurality of candidate reference blocks for the current block based on template matching (TM) costs of the plurality of candidate reference blocks, the TM costs indicating a difference between a template of the current block and each reference template of the plurality of candidate reference blocks; classifying the samples of the determined reference block into a plurality of sample classes; deriving a partitioning pattern for the current block based on the determined reference block from a plurality of predetermined partitioning patterns, the derived partitioning pattern indicating a plurality of partitions of the current block, and each of the plurality of sample classes of the determined reference block corresponding to each of the plurality of partitions of the current block; reconstructing the current block based on the derived partitioning pattern of the current block; A method comprising:
2. the plurality of candidate reference blocks are from one of the current picture and a reference picture of the current block; The method of claim 1 , wherein the determined reference block samples are reconstructed samples.
3. The step of determining the reference block from the plurality of candidate reference blocks further comprises: determining an initial reference block based on motion vector information included in the received video bitstream; determining the plurality of candidate reference blocks within a search range of the initial reference block; determining the TM cost between a reference template of each of the plurality of candidate reference blocks and a template of the current block; determining, from the plurality of candidate reference blocks, the reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block; The method of claim 2 , comprising:
4. The step of classifying the samples of the determined reference block further comprises: The method of claim 1 , further comprising classifying samples of the determined reference block into a first sample class and a second sample class based on a binary image segmentation algorithm.
5. The step of classifying the samples of the determined reference block further comprises:
2. The method of claim 1, comprising classifying samples of the determined reference block into a first sample class having sample values greater than a threshold and a second sample class having sample values less than a threshold, the threshold being one of a mean value of the samples and a median value of the samples.
6. The step of classifying the samples of the determined reference block further comprises: The method of claim 1 , comprising the step of clustering the samples of the determined reference block into one or more sample classes with a clustering method.
7. The step of classifying the samples of the determined reference block further comprises: determining a template region adjacent to said determined reference block; In edge detection, classifying samples of the determined reference block into one or more sample classes based on sample values of template regions adjacent to the determined reference block, the one or more sample classes including at least a first class including samples at the edge of the determined reference block and a second class including samples in an interior region of the determined reference block; The method of claim 1 , comprising:
8. 8. The method of claim 7, wherein the template region comprises one of: (i) a plurality of rows of neighboring samples above the determined reference block; (ii) a plurality of columns of neighboring samples to the left of the determined reference block; or (iii) a region in one or more reconstructed neighboring blocks of the determined reference block.
9. The step of classifying the samples of the determined reference block further comprises classifying the samples of the determined reference block based on each of a plurality of partitioning boundary candidates; The step of deriving a partitioning pattern for the current block includes: determining a plurality of partitioning boundary candidates for the current block, each of the plurality of partitioning boundary candidates for the current block corresponding to each of the plurality of partitioning boundary candidates for the determined reference block; determining a TM cost between each of the plurality of partitioning boundary candidates of samples of the determined reference block and a corresponding one of the plurality of partitioning boundary candidates of samples of the current block; determining a partitioning boundary corresponding to a minimum TM cost among the determined TM costs from the plurality of partitioning boundary candidates of the current block; determining the plurality of partitions of the current block based on the determined partitioning boundaries; The method of claim 1 further comprising:
10. The step of determining a reference block from the plurality of candidate reference blocks includes: determining a binary mask based on one of a mean value and a median value of the reconstructed samples of the template for the current block; determining a dominant sample group in the reconstructed samples of the template of the current block based on the binary mask; determining a dominant sample group of samples of a reference template for each of the plurality of candidate reference blocks based on the binary mask; determining a TM cost between a dominant group in the reconstructed sample of the template of the current block and a dominant group in a sample of a reference template of each of the plurality of candidate reference blocks; determining, from the plurality of candidate reference blocks, a reference block corresponding to a minimum TM cost among determined TM costs between a dominant group in the reconstructed sample of the template of the current block and a dominant group in a sample of a reference template of each of the plurality of candidate reference blocks; The method of claim 1 , comprising:
11. The step of determining a reference block from the plurality of candidate reference blocks includes: The method of claim 1 , further comprising determining the plurality of candidate reference blocks based on a plurality of motion vectors associated with the current block.
12. The step of determining a reference block from the plurality of candidate reference blocks includes: determining a TM cost between a reference template of each of the plurality of candidate reference blocks and a template of the current block; determining, from the plurality of candidate reference blocks, a candidate reference block corresponding to a minimum TM cost among determined TM costs between a reference template of each of the plurality of candidate reference blocks and a template of the current block; determining a motion vector (MV) based on the determined candidate reference block, the motion vector indicating an offset between the current block and the determined candidate reference block; determining an adjusted MV based on the determined MV and a motion vector difference (MVD); determining the reference block based on the adjusted MV; The method of claim 1 further comprising:
13. the plurality of candidate reference blocks are determined in the current picture; The step of determining a reference block from the plurality of candidate reference blocks includes: determining a TM cost between a reference template of each of the plurality of candidate reference blocks and a template of the current block; determining, from the plurality of candidate reference blocks, a candidate reference block corresponding to a minimum TM cost among determined TM costs between a reference template of each of the plurality of candidate reference blocks and a template of the current block; determining a block vector (BV) based on the determined candidate reference block, the BV indicating an offset between the current block and the determined candidate reference block; determining the reference block based on the determined BV; The method of claim 2 further comprising:
14. A device, a processing circuit, the processing circuit comprising: receiving a video bitstream including a current block in a current picture; determining a reference block from the plurality of candidate reference blocks for the current block based on template matching (TM) costs of the plurality of candidate reference blocks, the TM costs indicating a difference between a template of the current block and each reference template of the plurality of candidate reference blocks; classifying the samples of the determined reference block into a plurality of sample classes; deriving a partitioning pattern for the current block based on the determined reference block from a plurality of predetermined partitioning patterns, the derived partitioning pattern indicating a plurality of partitions of the current block, and each of the plurality of sample classes of the determined reference block corresponding to each of the plurality of partitions of the current block; reconstructing the current block based on the derived partitioning pattern of the current block; The equipment is configured to:
15. the plurality of candidate reference blocks are from one of the current picture and a reference picture of the current block; The apparatus of claim 14 , wherein the samples of the determined reference block are reconstructed samples.
16. The processing circuitry determining an initial reference block based on motion vector information included in the received video bitstream; determining the plurality of candidate reference blocks within a search range of the initial reference block; determining the TM cost between a reference template of each of the plurality of candidate reference blocks and a template of the current block; determining, from the plurality of candidate reference blocks, the reference block corresponding to the smallest TM cost among the determined TM costs between the reference template of each of the plurality of candidate reference blocks and the template of the current block; 16. The device of claim 15, configured to:
17. The processing circuitry The apparatus of claim 14 , configured to classify samples of the determined reference block into a first sample class and a second sample class based on a binary image segmentation algorithm.
18. The processing circuitry 15. The apparatus of claim 14, configured to classify samples of the determined reference block into a first sample class having sample values greater than a threshold and a second sample class having sample values less than a threshold, the threshold being one of a mean value of the samples and a median value of the samples.
19. The processing circuitry 15. The apparatus of claim 14, configured to cluster samples of the determined reference block into one or more sample classes in a clustering method.
20. The processing circuitry determining a template region adjacent to the determined reference block; 15. The device of claim 14, configured to classify samples of the determined reference block into one or more sample classes based on sample values of a template region adjacent to the determined reference block in edge detection, the one or more sample classes including at least a first class including samples at an edge of the determined reference block and a second class including samples in an interior region of the determined reference block.
21. 1. A method of video encoding, comprising: determining a reference block from a plurality of candidate reference blocks for a current block in a current picture based on template matching (TM) costs of the plurality of candidate reference blocks, wherein the TM costs indicate a difference between a template of the current block and a reference template of each of the plurality of candidate reference blocks; classifying the reconstructed samples of the determined reference block into a plurality of reconstructed sample classes; determining a plurality of partitions of the current block corresponding to the plurality of reconstructed sample classes of the determined reference block, wherein each of the plurality of reconstructed sample classes of the determined reference block corresponds to a respective partition of the plurality of partitions of the current block; encoding the current block based on the determined plurality of partitions of the current block; A method comprising:
Citation Information
Patent Citations
Moving image coding system using area division
JP1989228384A
High-speed geometric mode determination method and apparatus for a video encoder
JP2010524396A
Video processing method and device
JP2021520121A
Method and device for geometric split mode with split mode rearrangement - Patents.com
JP2025503091A
Geometric partitioning modes in video coding.
JP2025504291A