Template Matching Refinement for Affine Motion

Affine motion models with template matching refinement address the inefficiencies of traditional methods by accurately compensating for complex video motions, thereby improving video coding efficiency.

JP2025534720APending Publication Date: 2025-10-17TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025521275
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-16
Filing Date
2023-10-17
Publication Date
2025-10-17

Smart Images

  • Figure 2025534720000001_ABST
    Figure 2025534720000001_ABST
Patent Text Reader

Abstract

The current block is coded in affine mode and includes a first control point located at a first corner of the current block. A current template associated with the first control point is determined. A plurality of candidate reference templates are determined in the reference picture with respect to the current template. A reference template is selected from the plurality of candidate reference templates for the current template based on template matching (TM) costs. The TM costs indicate respective differences between the current template of the first control point and each of the candidate reference templates. A first control point motion vector (CPMV) is determined based on the selected reference template, where the first CPMV indicates an offset between the selected reference template in the reference picture and the current template associated with the first control point. The current block is reconstructed based on at least the first CPMV.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001]

[0001] Incorporation by Reference This application claims the benefit of priority to U.S. Patent Application No. 18 / 380,525, entitled "Template Matching Refinement for Affine Motion," filed October 16, 2023, which in turn claims the benefit of priority to U.S. Provisional Application No. 63 / 417,279, entitled "Template Matching Refinement for Affine Motion," filed October 18, 2022. The disclosures of the prior applications are incorporated herein by reference in their entirety.

[0002]

[0002] Technical Field This disclosure describes embodiments generally related to video coding. [Background technology]

[0003] The background discussion provided herein is intended to generally present the context of the present disclosure. Work under the names of the current inventors is not admitted, expressly or impliedly, as prior art to the present disclosure to the extent that that work is described in this background section or in a descriptive manner that may not otherwise qualify as prior art as of the filing date.

[0004]

[0004] Image / video compression can help transmit image / video data across various devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can utilize a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec can utilize a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be specified by a motion vector (MV). Summary of the Invention

[0005] Aspects of the present disclosure include methods and apparatus for video encoding / decoding (i.e., video coding). In some examples, an apparatus for video decoding includes a processing circuit.

[0006] According to an aspect of the present disclosure, a video decoding method is provided. In the method, a video bitstream including a current block in a current picture is received. The current block is coded in an affine mode, and a first control point associated with the affine mode is located at a first corner of the current block. A current template associated with the first control point is determined, the current template being in a vicinity of the first control point. For the current template associated with the first control point, multiple candidate reference templates are determined in a reference picture. A reference template is selected from the multiple candidate reference templates for the current template associated with the first control point based on a template matching (TM) cost. The TM cost indicates a respective difference between the current template of the first control point and each candidate reference template. Based on the selected reference template, a first control point motion vector (CPMV) is determined, the first CPMV indicating an offset between the selected reference template in the reference picture and the current template associated with the first control point. The current block is reconstructed based on at least the first CPMV.

[0007] In one example, a first block is determined, for which a first control point is located at the center of the first block, and a current template associated with the first control point is determined as a reconstructed region located at one or a combination of (i) an upper side of the first block and (ii) a left side of the first block.

[0008]

[0008] In one example, a current template associated with a first control point is determined as a reconstructed region in the vicinity of the first control point that includes at least one of (i) a first region above the current block, or (ii) a second region to the left of the current block.

[0009] In one example, a current template associated with a first control point is determined as a reconstructed region, the first control point being at a center of the reconstructed region, the reconstructed region including at least one of (i) a first region located above the current block and extending beyond a vertical side of the current block, and (ii) a second region located to the left of the current block and extending beyond a horizontal side of the current block.

[0010]

[0010] In one example, the first region has a height equal to N samples and a width equal to the width of an affine sub-block of the current block, and the second region has a width equal to N samples and a height equal to the height of an affine sub-block of the current block, where N is a positive integer.

[0011] In one example, an initial first CPMV is determined for a first control point, and the initial first CPMV indicates an initial reference template in a reference picture. Within a search range of the initial reference template, multiple candidate reference templates are determined. The search range includes M×M pixels, where M is less than 8.

[0012] In one example, the plurality of candidate reference templates are determined within a search range of the initial reference template based on a plurality of search steps, the number of which is determined based on one of a predetermined resolution and a predetermined number.

[0013] In one example, a TM cost between a current template associated with a first control point and each of a plurality of candidate reference templates is determined, and a reference template corresponding to a minimum TM cost among the determined TM costs between the current template associated with the first control point and each of the plurality of candidate reference templates is selected from the plurality of candidate reference templates.

[0014] In one example, a template of a current block is determined, the template including a first region above the current block and a second region to the left of the current block. A first candidate CPMV is determined for a first control point based on a first candidate reference template from a plurality of candidate reference templates. A second candidate CPMV is determined for the first control point based on a second candidate reference template from a plurality of candidate reference templates. A first set of sub-block affine MVs is determined for sub-blocks of the current block that are in a vicinity of the template of the current block based on at least the first candidate CPMV. A second set of sub-block affine MVs is determined for sub-blocks of the current block that are in a vicinity of the template of the current block based on at least the second candidate CPMV. A first set of reference sub-block affine MVs is determined for sub-blocks of a collocated block of the current block that correspond to the first set of sub-block affine MVs. A second set of reference subblock affine MVs is determined for a subblock of the equivalently positioned block of the current block corresponding to the second set of subblock affine MVs. A first reference template is determined based on the first set of reference subblock affine MVs, and a second reference template is determined based on the second set of reference subblock affine MVs. A first TM cost between the template of the current block and the first reference template is determined. A second TM cost between the template of the current block and the second reference template is determined. One of the first candidate reference template and the second candidate reference template is selected as the reference template. One of the first candidate reference template and the second candidate reference template corresponds to the smaller of the first TM cost and the second TM cost.

[0015] In one example, a reference subblock is determined for each of the subblocks of the equivalently positioned block based on each of the first set of reference subblock affine MVs. A subblock template is determined for each of the reference subblocks of the subblocks of the equivalently positioned block. A first reference template is determined as a combination of the subblock templates.

[0016]

[0016] In one example, the affine mode includes one of an affine mono-prediction mode and an affine bi-prediction mode.

[0017] In one example, the current block includes a second control point at a second corner of the current block. An initial second CPMV is determined for the second control point. The second CPMV is determined for the second control point by adding a translational MV offset to the initial second CPMV, where the translational MV offset is derived based on decoder side motion vector refinement (DMVR).

[0018] According to another aspect of the disclosure, an apparatus is provided. The apparatus includes a processing circuit. The processing circuit can be configured to perform any of the described methods for video decoding / encoding. In one example, the processing circuit is configured to receive a video bitstream including a current block in a current picture, the current block being coded in affine mode, and a first control point associated with the affine mode being located at a first corner of the current block. The processing circuit is configured to determine a current template associated with the first control point, the current template being in a vicinity of the first control point. The processing circuit is configured to determine multiple candidate reference templates in a reference picture for the current template associated with the first control point. The processing circuit is configured to select a reference template from the multiple candidate reference templates for the current template associated with the first control point based on a template matching (TM) cost. The TM cost indicates a respective difference between the current template for the first control point and each candidate reference template. The processing circuit is configured to determine a first control point motion vector (CPMV) based on the selected reference template, the first CPMV indicating an offset between the selected reference template and a current template associated with the first control point in the reference picture. The processing circuit is configured to reconstruct the current block based on at least the first CPMV.

[0019]

[0019] An aspect of the present disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods described for video decoding / encoding. [Brief explanation of the drawings]

[0020]

[0020] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1]

[0021] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2]

[0022] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3]

[0023] FIG. 3 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4A]

[0024] Figure 4A is a schematic diagram of an exemplary four-parameter affine model. [Figure 4B]

[0025] Figure 4B is a schematic diagram of an exemplary six-parameter affine model. [Figure 5]

[0026] FIG. 5 is a schematic diagram of an exemplary affine motion vector field relating to sub-blocks of a block. [Figure 6]

[0027] Figure 6 is a schematic diagram of an exemplary template matching process. [Figure 7]

[0028] Figure 7 is a schematic diagram of an exemplary sub-block-based template matching process. [Figure 8]

[0029] Figure 8 shows a first exemplary template position for template-matching based refinement. [Figure 9]

[0030] Figure 9 shows a second exemplary template position for template-matching-based refinement. [Figure 10]

[0031] Figure 10 shows a third exemplary template position for template-matching-based refinement. [Figure 11]

[0032] FIG. 11 shows a flow chart outlining a decoding process according to some embodiments of the present disclosure. [Figure 12]

[0033] FIG. 12 shows a flow chart outlining an encoding process according to some embodiments of the present disclosure. [Figure 13]

[0034] FIG. 13 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0021]

[0035] 1 illustrates a block diagram of a video processing system (100) in some examples. The video processing system illustrates an example application of the disclosed subject matter: a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, and storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).

[0022]

[0036] The streaming system (100) includes a video source (101), which may include, for example, a digital camera, and a capture subsystem (113) capable of generating, for example, a stream of uncompressed video pictures (102). In one example, the video picture stream (102) includes samples captured by the digital camera. The video picture stream (102), depicted as a thick line to emphasize its large amount of data when compared to the encoded video data (104) (or coded video bitstream), may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof, and may enable or implement aspects of the disclosed subject matter, as described in more detail below. The coded video data (104) (or coded video bitstream), depicted as a thin line to emphasize its smaller amount of data when compared to the stream of video pictures (102), can be stored on a streaming server (105) for future use.

[0023] One or more streaming client subsystems, such as the client subsystems (106) and (108) of FIG. 1, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include a video decoder (110), for example, within an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an output stream (111) of video pictures that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) can be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in the context of VVC.

[0024]

[0037] It should be noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and electronic device 130 may include a video encoder (not shown).

[0025]

[0038] 2 shows an example block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used in place of the video decoder (110) in the example of FIG. 1.

[0026]

[0039] The receiver (231) can receive one or more coded video sequences, e.g., contained in a bitstream, to be decoded by the video decoder (210). In some embodiments, it can receive one coded video sequence at a time, where the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences can be received from a channel (201), which can be a hardware or software link to a storage device that stores the coded video data. The receiver (231) can receive the coded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which can be transferred using respective entities (not shown). The receiver (231) can separate the coded video sequences from other data. To address network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the "parser (220)"). In certain applications, the buffer memory (215) is part of the video decoder (210). In other cases, it may be external to the video decoder (210) (not shown). In yet another example, there may be a buffer memory (not shown) external to the video decoder (210), for example, to deal with network jitter, and there may even be another buffer memory (215) internal to the video decoder (210), for example, to handle playback timing. If the receiver (231) is receiving data from a store-and-forward device or a synchronous network with sufficient bandwidth and controllability, the buffer memory (215) may not be needed, or may be small (if present).For use in a best-effort packet network such as the Internet, a buffer memory (215) may be required, which may be relatively large and may advantageously be adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) outside the video decoder (210).

[0027]

[0040] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and, potentially, information for controlling a rendering device, such as a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but may be coupled to the electronic device (230), as shown in FIG. 2. The rendering device control information may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard, including variable length coding, Huffman coding, arithmetic coding with or without context effects, and various other principles. The parser (220) can extract from the coded video sequence a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroup can include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) can also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0028]

[0041] The parser (220) is capable of performing an entropy decoding / parsing process on the video sequence received from the buffer memory (215) to generate symbols (221).

[0029]

[0042] The reconstruction of the symbols (221) may include several different units depending on the type of coded video picture or part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are included and how can be controlled by subgroup control information parsed by the parser (220) from the coded video sequence. The flow of such subgroup control information between the parser (220) and subsequent units is not depicted for clarity.

[0030]

[0043] Beyond the functional blocks already described, the video decoder (210) may be conceptually subdivided into a number of functional units, as described below. In a practical implementation operating within commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0031]

[0044] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as well as control information (including the transform to use, block size, quantization factor, quantization scaling matrix, etc.) from the parser (220) as symbols (221). The scalar / inverse transform unit (251) can output blocks containing sample values ​​that can be input to an aggregator (255).

[0032]

[0045] In some cases, the output samples of the scalar / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates blocks of the same size and shape as the block being reconstructed using already reconstructed surrounding information retrieved from the current picture buffer (258). The current picture buffer (258), for example, buffers the partially reconstructed and / or fully reconstructed current picture. The aggregator (255) optionally adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information as provided by the scalar / inverse transform unit (251).

[0033]

[0046] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to a block that may be inter-coded and motion-compensated. In such cases, the motion-compensated prediction unit (253) may access the reference picture memory (257) to retrieve samples used for prediction. After motion-compensating the retrieved samples according to the symbols (221) associated with the block, these samples are added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) retrieves prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253), for example, in the form of symbols (221), which may have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values ​​taken from a reference picture memory (257), motion vector prediction mechanisms, etc., where sub-sample accurate motion vectors are used.

[0034]

[0047] The output samples of the aggregator (255) can be subjected to various loop filtering techniques in the loop filter unit (256). The video compression techniques can include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression can also depend on meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as on previously reconstructed loop-filtered sample values.

[0035]

[0048] The output of the loop filter unit (256) may be a stream of samples that can be output to a rendering device (212) or stored in a reference picture memory (257) for use in future inter-picture prediction.

[0036]

[0049] Once a given coded picture is fully reconstructed, it can be used as a reference picture for future predictions. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a fresh current picture buffer can be reallocated before starting the reconstruction of a subsequent coded picture.

[0037]

[0050] The video decoder (210) may perform decoding operations according to a standard or predetermined video compression technology, such as ITU-T Rec. H.265. A coded video sequence may conform to the syntax specified by the video compression technology or standard in use, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard. Specifically, a profile may select certain tools from all tools available in the video compression technology or standard as the only tools that can be used under that profile. Compliance also requires that the complexity of the coded video sequence fall within a range defined by the level of the video compression technology or standard. In some cases, the level may limit the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the levels may in some cases be further constrained by Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management signaled in the coded video sequence.

[0038]

[0051] In an embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0039]

[0052] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.

[0040]

[0053] The video encoder (303) can receive video samples from a video source (301) (which in the example of FIG. 3 is not part of the electronic device (320)) that can capture video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0041]

[0054] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (301) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed sequentially. The picture itself may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. The following discussion focuses on samples.

[0042]

[0055] According to an embodiment, the video encoder (303) is capable of coding and compressing pictures of a source video sequence into a coded video sequence (343) in real time or under any other required time constraints. Enforcing an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is functionally coupled to other functional units, as described below, whose coupling is not depicted for clarity. Parameters set by the controller (350) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured with other appropriate functions associated with the video encoder (303) optimized for a particular system design.

[0043]

[0056] In some embodiments, the video encoder (303) is configured to operate in a coding loop. As a simplified explanation, in one example, the coding loop may include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data in a manner similar to that generated by the (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-exact results independent of the location (local or remote) of the decoder, the contents of the reference picture memory (334) are also bit-exact between the local and remote encoders. In other words, the predictor in the encoder "sees" as reference picture samples exactly the same sample values ​​that the decoder would "see" if it were using prediction during decoding. This basic principle of reference picture synchronization (note that if synchronization cannot be maintained, e.g., due to channel errors, resulting in drift) is used in several related technologies as well.

[0044]

[0057] The operation of the "local" decoder (333) may be the same as that of a "remote" decoder, such as the video decoder (210) already described in detail above in connection with Figure 2. However, briefly referring also to Figure 2, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (345) and parser (220) may be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).

[0045]

[0058] In embodiments, decoder techniques other than analysis / entropy decoding present in a decoder are also present in the corresponding encoder, ideally or in substantially the same functional form. Therefore, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder techniques can be omitted, as they are the opposite of the decoder techniques described generically. In certain areas, more detailed descriptions are provided below.

[0046]

[0059] During operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that can be selected as predictive references for the input picture.

[0047]

[0060] The local video decoder (333) can decode coded video data of pictures that can be designated as reference pictures based on symbols generated by the source coder (330). The operation of the coding engine (332) can advantageously be a non-lossless process. When the coded video data can be decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence can typically be a replica of the source video sequence with some errors. The local video decoder (333) can repeat the decoding process that can be performed by the video decoder with respect to the reference pictures, causing the reconstructed reference pictures to be stored in the reference picture cache (334). In this way, the video encoder (303) can locally store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures that will be obtained by a far-end video decoder (assuming there are no transmission errors).

[0048]

[0061] The predictor (335) can perform a prediction search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) can search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or predetermined metadata (reference picture motion vectors, block shapes, etc.), which may serve as suitable prediction references for the new picture. The predictor (335) can operate on a sample-block-pixel-block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), an input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).

[0049]

[0062] The controller (350) can manage the coding operations of the source coder (330), including, for example, setting parameters and subgroup parameters used to encode the video data.

[0050]

[0063] All outputs of the aforementioned functional units can be entropy coded in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0051]

[0064] The transmitter (340) can buffer the coded video sequence, as produced by the entropy coder (345), and prepare it for transmission over a communication channel (360), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (340) can merge the coded video data from the video coder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0052]

[0065] The controller (350) can manage the operation of the video encoder (303). During coding, the controller (350) can assign a particular coded picture type to each coded picture, which can affect the coding technique that can be applied to each picture. For example, pictures can often be assigned as one of the following picture types:

[0066] An intra picture (I-picture) is one that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0053]

[0067] A predicted picture (P-picture) can be coded and decoded using intra- or inter-prediction, which uses motion vectors and reference indices to predict the sample values ​​of each block.

[0054]

[0068] Bi-directionally predicted pictures (B-pictures) can be coded and decoded using intra- or inter-prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a block.

[0055]

[0069] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and can be coded block by block. Blocks can be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the respective picture. For example, blocks of an I-picture can be non-predictively coded, or they can be predictively coded with reference to previously coded blocks of the same picture (spatial or intra prediction). Pixel blocks of a P-picture can be predictively coded with spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture can be predictively coded with spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0056]

[0070] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.266. In this operation, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. The coded video data may therefore conform to a syntax specified by the video coding technique or standard used.

[0057]

[0071] In an embodiment, the transmitter (340) can transmit additional data along with the coded video. The source coder (330) can include such data as part of the coded video sequence. The additional data can include temporal, spatial, and SNR enhancement layers, as well as other forms of redundant data (such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.).

[0058]

[0072] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture under encoding / decoding, called the current picture (or present picture), is partitioned (i.e., divided) into blocks. If a block in the current picture is similar to a reference block in a reference picture that was previously coded and is still buffered in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0059]

[0073] In some embodiments, bi-prediction techniques may be used for inter-picture prediction. Bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, that both precede the current picture in decoding order (but may be past and future, respectively, in display order) in the video. A block in the current picture may be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block may be predicted by a combination of the first and second reference blocks.

[0060]

[0074] Furthermore, to improve coding efficiency, it is possible to use merge mode techniques for inter-picture prediction.

[0061]

[0075] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be partitioned into one 64x64 pixel CU, four 32x32 pixel CUs, or sixteen 16x16 pixel CUs. In one example, each CU is analyzed to determine the CU's prediction type, such as an inter prediction type or an intra prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values ​​(e.g., luma values) for pixels, such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0062]

[0076] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In some embodiments, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.

[0063]

[0077] The present disclosure includes aspects related to template matching-based motion refinement for affine-coded blocks.

[0064]

[0078] A translational motion model can be applied to motion compensation prediction (MCP) such as that in HEVC. In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motions, and other irregular motions. Block-based affine transformation motion compensation can be applied as in the VVC test model (VTM). For example, the affine motion field of a block can be described by the motion information of two control point motion vectors (four parameters) in Figure 4A or three control point motion vectors (six parameters) in Figure 4B.

[0065]

[0079] For the four-parameter affine motion model, the motion vector at a sample location (x,y) within a block can be derived as shown in Eq. (1):

[0066]

number

[0067]

number

[0068]

number

[0069]

number

[0070]

[0080] To simplify motion-compensated prediction, block-based affine transformation prediction can be applied. As shown in Figure 5, to derive the motion vectors for each 4x4 luma subblock within the block, the motion vector (MV) of the center sample (e.g., (502)) of each subblock (e.g., (504)) of block (500) can be calculated according to one of the affine motion model formulas (1)-(4) and further rounded to 1 / 16 fractional precision. Then, a motion-compensated interpolation filter can be applied to generate a prediction for each subblock using the derived motion vector. The subblock size of the chroma components can also be set to 4x4. The motion vector for a 4x4 chroma subblock can be calculated as the average of the MVs of the four corresponding 4x4 luma subblocks.

[0071]

[0081] The basis MVs of the affine model of a coding block coded in affine merge mode (e.g., the translational part of the affine model) can, in some cases, be refined by applying only the first pass of multi-pass DMVR. For example, translational MV offsets can be added to all CPMVs of a candidate in the affine merge list if that candidate satisfies the DMVR condition. It is possible to derive motion vector offsets by minimizing the cost of bilateral matching, which can be the same as conventional DMVR. Furthermore, the DMVR condition may not be changed.

[0072]

[0082] The motion vector offset search process can be the same as the first pass of multi-pass DMVR. For example, in ECM, a 3x3 square search pattern can be used to loop (or search) through the search range [-8, +8] horizontally and [-8, +8] vertically to find the best (or selected) integer MV offset. A half-pel search can be performed around the best integer position, and an error surface estimation can be performed to find the MV offset with 1 / 16 accuracy.

[0073]

[0083] The refined CPMV can be stored as a multi-pass DMVR in the ECM for both spatial and temporal motion vector prediction.

[0074]

[0084] Template matching (TM) can be a decoder-side MV derivation method for refining the motion information of a current CU by finding the closest match between a template in the current picture (e.g., neighboring blocks above and / or to the left of the current CU) and a block in a reference picture (e.g., of the same size as the template). Figure 6 shows that a better (or selected) MV can be searched for around the initial motion (612) of the current CU (602) within a [-8, +8]-pel search range. As shown in Figure 6, the current block (602) can include a template (or current template) (604) that includes neighboring blocks above and to the left of the current block (602) in the current frame (606). A block (or reference template) (608) in the reference frame (610) can be specified by the initial MV (612). A better MV (not shown) can be searched (or identified) around the initial motion vector (612) of the current CU (602) within a [-8, +8]-pel search range.

[0075]

[0085] In one aspect, the template matching method as in JVET-J0021 can be used with the following modifications: the search step size can be determined based on the adaptive motion vector resolution (AMVR) mode, and the TM can be coupled with the bilateral matching process in merge mode.

[0076]

[0086] In advanced motion vector prediction (AMVP) mode, an MVP candidate can be determined based on the template matching error to select the one that achieves the minimum difference between the current block template and the reference block template. TM can then be performed on the determined MVP candidate for MV refinement. TM can refine the MVP candidate by using an iterative diamond search, starting with full-pel MVD precision (e.g., 4-pel for 4-pel AMVR mode) within the [-8, +8]-pel search range. The AMVP candidate can be further refined by using a cross search with full-pel MVD precision (e.g., 4-pel for 4-pel AMVR mode), followed by successive use of half-pel and quarter-pel precision depending on the AMVR mode, as specified in Table 1. The search process can ensure that the MVP candidate maintains the same MVD precision specified by the AMVR mode after the TM process. In the search process, if the difference between the previous minimum cost and the current minimum cost in the interaction is less than a threshold (eg, the area of ​​the block), the search process can terminate.

[0077] Table 1. Search patterns for AMVR mode and AMVR merge mode

[0078] [Table 1]

[0087] In merge mode, a similar search method can be applied to merge candidates indicated by the merge index. As Table 1 shows, TM can go all the way down to 1 / 8-pel MVD precision or skip precision beyond half-pel MVD precision, depending on whether an alternative interpolation filter (used when AMVR is in half-pel mode) is used according to the motion information being merged. Furthermore, when TM mode is enabled, template matching can operate as an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods or as an independent process, depending on whether BM can be enabled according to the enablement condition check.

[0079]

[0088] In adaptive reordering of merge candidates with template matching (ARMC-TM), merge candidates can be adaptively reordered in TM. This reordering method can be applied to regular merge mode, TM merge mode, and affine merge mode (except for SbTMVP candidates). In TM merge mode, merge candidates can be reordered before the refinement process.

[0080]

[0089] After the merge candidate list is constructed, the merge candidates can be divided into several subgroups. The subgroup size can be set to 5 for regular merge mode and TM merge mode. The subgroup size can be set to 3 for affine merge mode. The merge candidates within each subgroup can be sorted in ascending order according to their cost values ​​based on template matching. For simplicity, merge candidates in the latest subgroup but not the first subgroup do not need to be sorted. Zero candidates from the ARMC reordering process can be filtered out during the construction of the merge motion vector candidate list.

[0081]

[0090] The template matching cost of a merge candidate may be measured by a sum of absolute differences (SAD) between the template samples of the current block and the reference samples corresponding to the template samples of the current block. The template may include a set of reconstructed samples in the neighborhood of the current block. The template's reference samples may be located by the motion information of the merge candidate.

[0082]

[0091] If the merge candidate utilizes bi-prediction, the reference samples of the merge candidate's template may also be generated by bi-prediction.

[0083]

[0092] For a subblock-based merging candidate with a subblock size equal to W × H, the top template may include several sub-templates with a size of W × 1, and the left template may include several sub-templates with a size of 1 × H. The motion information of the subblocks in the first row and first column of the current block may be used to derive the reference samples for each sub-template.

[0084]

[0093] FIG. 7 illustrates an exemplary subblock-based template matching process (700) (or process (700)). As shown in FIG. 7, a current block (702) may be included in a current picture (704). The current block (702) may include a template (706) located above and to the left of the current block (702). The current block (702) may include multiple subblocks, such as subblocks AG, that are adjacent to the template (706). Each of the subblocks AG may include a respective motion vector (MV) that indicates a reference subblock for the respective subblock. The current block (702) may include an equivalently positioned block (708) in a reference picture (710). The equivalently positioned block (708) may include multiple equivalently positioned subblocks AG that correspond to the subblocks AG in the current block (702). Each of the equivalently positioned subblocks AG may have a respective motion vector (MV). The motion vector of the equally positioned sub-block AG may have the same direction as the motion vector of the sub-block AG. Based on the motion vector of the equally positioned sub-block AG, multiple reference sub-blocks AG for the equally positioned sub-block AG may be determined. Furthermore, a template may be determined for each of the reference sub-blocks AG. For example, a template (or sub-block template) (716) may be determined for the reference sub-block D. Based on the template of the reference sub-block AG, a reference template for the equally positioned block (708) may be determined. The reference templates for the equally positioned block (708) may include an upper reference template (721) and a left reference template (714). The reference template (712) may include a template for the reference sub-block AD, and the left reference template (714) may include templates for the reference sub-blocks A and E.Furthermore, it is possible to calculate the template matching cost between the template (706) of the current block (702) and the reference template of the equivalently positioned block (708).

[0085]

[0094] It is possible to improve coding efficiency by using methods based on affine models or DMVR methods based on CPMV. However, it is possible to further improve coding efficiency by applying MV refinement methods based on template matching.

[0086]

[0095] In the present disclosure, template matching refinement can be applied to refine CPMVs for an affine block. For example, a current template associated with a first control point of the current block can be determined, where the current template is in a vicinity of the first control point. Multiple candidate reference templates can be determined within a reference picture for the current template associated with the first control point. The reference template can be determined or selected from the multiple candidate reference templates for the current template associated with the first control point based on a TM cost. The TM cost can indicate a difference between each of the multiple candidate reference templates and the current template for the first control point. A first CPMV can be determined based on the determined (or selected) reference template, and the first CPMV can indicate an offset between the determined reference template in the reference picture and the current template associated with the first control point.

[0087]

[0096] In the present disclosure, a current template associated with a first control point of a current block may be located near the first control point. In one example, as shown in FIG. 8, the first control point (802) may be located in the upper left corner of the current block (800) and further in the center of the first block (808). The current template may include an upper region (814) located above the first block (808) and a left region (816) located to the left of the first block (808). 9, a first control point (902) may be located in the upper left corner of a current block (900). The current templates may include a top template (908) above the current block (900) and / or a left template (910) on the left side of the current block (900). Both the top template (908) and the left template (910) may be in contact with or near (or adjacent to) the first control point (902). In one example, as shown in FIG. 10, a first control point (1002) may be located in the upper left corner of a current block (1000) and may also be located within a current template, which includes a first region (1008) above the current block (1000) and a second region (1010) to the left of the current block (1000).

[0088]

[0097] In one example, to determine a reference template from the candidate reference templates for a current template associated with a first control point, a TM cost between the current template associated with the first control point and each of the plurality of candidate reference templates can be determined, and a reference template (corresponding to a minimum TM cost among the TM costs determined between the current template associated with the first control point and each of the plurality of candidate reference templates) can be determined from among the plurality of candidate reference templates.

[0089]

[0098] In one aspect, template matching refinement can be applied to one or more of the CPMVs of the affine block. The refined CPMV with the best (or selected) template matching cost for the control point can be used as the final CPMV for the control point.

[0090]

[0099] For example, the affine block may include a first control point. A current template associated with the first control point may be determined. An exemplary template for the first control point may be as shown in FIGS. 8-10. Multiple candidate reference templates may be determined within a search range of an initial candidate reference template in the reference picture. The initial candidate reference template may be specified by an initial CPMV of the first control point. The reference template may be determined or selected from the multiple candidate reference templates for the current template associated with the first control point based on a TM cost. The TM cost may indicate a difference between each of the multiple candidate reference templates and the current template for the first control point. A refined CPMV may be determined based on the determined reference template, and the refined CPMV may indicate an offset between the reference template determined in the reference picture and the current template associated with the first control point.

[0091]

[0100] In one embodiment, the template size can be set according to the affine sub-block size. Sub-block width SubW×N where N may be an integer such as 1 or 2. In one example, the left template is Sub-block height SubH×M where M can be an integer such as 1 or 2.

[0092]

[0101] In one example, as shown in FIG. 9, a current block (900) may include a first control point (902), a second control point (904), and a third control point (906). The template associated with the first control point (902) (or the current template) may include a top template (908) and a left template (910). The top template (908) may include a height equal to N samples and a width equal to the width of an affine sub-block (not shown) of the current block, and the left template (910) may include a width equal to N samples and a height equal to the height of the affine sub-block of the current block, where N is a positive integer. An exemplary affine sub-block may be shown as sub-block (504) in FIG. 5.

[0093]

[0102] In one embodiment, each control point can be refined based on a block having a size of N x M. The control point can be contained in the block. For example, the control point can be located at the center position (N / 2, M / 2) of the block. The available template area can be derived from samples in the vicinity of the block. An exemplary template area can be as shown in FIG. 8, which provides a six-parameter affine model using three control points. A similar template area can also be applied to a four-parameter affine model using two control points.

[0094]

[0103] As shown in FIG. 8 , a current block (800) may include a first control point (802), a second control point (804), and a third control point (806). Three blocks (808), (810), and (812) may be defined so that the control points are positioned at the center of the blocks. For example, a first block (808) may be determined so that the first control point (802) is positioned at the center of the first block (808). A template (or current template) associated with the first control point (802) may be determined as a reconstructed region, and the reconstructed region may include a top region (814) positioned above the first block (808) and a left region (816) positioned to the left of the first block (808).

[0095]

[0104] In one embodiment, each control point can be refined based on one or more templates in the immediate neighborhood (adjacent) of the respective control point. In one example, the template size can be set as the affine sub-block size. Figure 9 shows an example template based on a six-parameter affine with three control points. A similar template area can be applied to two control points in a four-parameter model.

[0096]

[0105] As shown in FIG. 9, the template associated with the first control point (902) can be determined as a reconstructed region, which includes an upper template (908) adjacent to the first control point (902) and located above the current block (900) or a left template (910) located to the left of the current block (900).

[0097]

[0106] In one embodiment, the template size can be set to be a multiple of the sub-block size, e.g., twice the sub-block size. An exemplary template area can be as shown in Figure 10 based on a six-parameter affine model using three control points. A similar template region can be applied to two control points of a four-parameter affine model.

[0098]

[0107] As shown in FIG. 10 , a current block (1000) may include a first control point (1002), a second control point (1004), and a third control point (1006). The template associated with the first control point (1002) may include a first region (1008) located above the current block (1000) and extending beyond a vertical edge (e.g., a left edge (1012)) of the current block (1000), and a second region (1010) located on the left side (1012) of the current block (1000) and extending beyond a horizontal edge (e.g., a bottom edge (1014)) of the current block (1000). In one example, the width of the first region (1008) may be twice the width of a sub-block (not shown) of the current block (1000). In one example, the height of the second region (1010) may be twice the height of the sub-block of the current block (1000).

[0099]

[0108] In one aspect, control points with valid reconstructed templates can be refined by template matching. For example, as shown in Figure 9, a top template (908) and a left template (910) associated with a first control point (902) can be reconstructed samples. The CPMV v0 of the first control point (902) can be refined based on the top template (908) and the left template (910) according to the TM.

[0100]

[0109] In one example, as shown in FIG. 9 , an initial first CPMV v0 may be determined for a first control point (902), where the initial first CPMV v0 may indicate an initial reference template (914) in a reference picture (916). Multiple reference template candidates (not shown) may be determined within a search range (918) that includes the initial reference template (914). A TM cost between a template associated with the first control point (902) (e.g., a top template (908) and a left template (910)) and each of the multiple candidate reference templates may be determined. A reference template corresponding to the smallest TM cost among the TM costs determined between the template associated with the first control point (902) and each of the multiple candidate reference templates may be determined or selected from the multiple candidate reference templates. A refined first CPMV may be determined based on the determined reference template, and the refined first CPMV may indicate an offset between the determined reference template in the reference picture and the template associated with the first control point (902).

[0101]

[0110] In one aspect, an inter-template matching method can be used to refine each CPMV. For example, as shown in Figure 9, the CPMV v0 of the first control point (902) can be refined based on the top template (908) and left template (910) associated with the first control point (902). The CPMV v1 of the second control point (904) can be refined based on the top template (912). The CPMV v2 of the third control point (906) can be refined based on the left template (913).

[0102]

[0111] In one embodiment, a template-to-template matching method, such as the template-to-template matching shown in FIG. 6, can be used with a reduced search range or a reduced number of search steps to reduce the complexity of the CPMV refinement.

[0103]

[0112] 9, an initial first CPMV v0 may be determined for a first control point (902), where the initial first CPMV v0 may indicate an initial reference template (914) in a reference picture (916). Multiple candidate reference templates (not shown) may be determined within a search range (918) that includes the initial reference template (914) relative to a current template (e.g., top template (908) and left template (910)) for the first control point (902). The search range (918) may include M by M pixels, where M is less than 8.

[0104]

[0113] In one example, multiple candidate reference templates can be determined within a search range 918 of the initial reference template 914 based on multiple search steps, where the number of search steps can be determined based on one of a predetermined resolution (e.g., 1 pixel or 1 / 4 pixel) and a predetermined number.

[0105]

[0114] In one aspect, for each candidate CPMV refinement, a sub-block affine MV can be generated, and a sub-block template can be generated accordingly. The candidate CPMV refinement that minimizes the sub-block template matching cost can be the CPMV refinement selected for the current affine block.

[0106]

[0115] 9, multiple candidate reference templates (not shown) may be determined for the first control point (902) within the search range (918). Candidate CPMVs (or candidate CPMV refinements) may be determined for the first control point (902) based on the candidate reference templates. Each candidate CPMV for the first control point (902) may indicate an offset between the template of the first control point (902) (e.g., top template (908) and left template (910)) and the respective candidate reference template. Similarly, candidate CPMVs may be determined for the second control point (904) and the third control point (906).

[0107] Furthermore, based on the CPMVs determined for the control points 902, 904, and 906, multiple sets of candidate CPMVs can be defined for the control points 902, 904, and 906. Each set of candidate CPMVs can construct a six-parameter affine model for the current block 900. Based on each set of candidate CPMVs, a set of sub-block affine MVs can be generated for a sub-block (e.g., sub-block 724 in FIG. 7) within the current block 900. Exemplary sub-block affine MVs can be as shown in FIG. 5.

[0108] Based on each set of subblock affine MVs (e.g., MVs (720) in FIG. 7) for the current block (900), a set of reference subblock affine MVs (e.g., MVs (722) in FIG. 7) can be determined for subblocks (e.g., subblocks (726) in FIG. 7) within the equivalently positioned block (e.g., equivalently positioned block (708) in FIG. 7) of the current block (900) that correspond to the set of subblock affine MVs for the current block (900). A reference subblock (e.g., reference subblock (718) in FIG. 7) can be determined for each subblock (e.g., subblock (726)) within the equivalently positioned block based on each of the sets of reference subblock affine MVs (e.g., MVs (722)). A subblock template (e.g., subblock template 716 in FIG. 7) can be determined for each reference subblock (e.g., reference subblock 718) of a subblock (e.g., subblock 726) in the colocated block. A set of subblock templates can be determined as a combination of subblock templates for the reference subblocks of the subblocks in the colocated block. Exemplary sets of subblock templates can be shown as top reference template 712 and left reference template 714 in FIG. 7. Based on a subblock-based template matching process such as process 700, a subblock template matching cost can be determined between the template (not shown) of the current block 900 and each set of subblock templates. The template of the current block 900 can have a configuration similar to template 706 in FIG. 7. The set of candidate CPMVs for control points (902), (904), (906) corresponding to the smallest sub-block template matching cost can be selected as refined CPMVs for block (900).

[0109]

[0116] In one embodiment, a CPMV refinement method based on template matching can be used for affine uni-prediction.

[0110]

[0117] In one embodiment, a template matching-based CPMV refinement method can be used for affine bi-prediction. In one example, each set of CPMVs in each reference list can be refined individually.

[0111]

[0118] In one embodiment, template-matching based CPMV refinement and DMVR based CPMV refinement can be used for different CPMVs within the same affine block.

[0112]

[0119] For example, as shown in FIG. 9, the current block (900) may include a first control point (902) and a second control point (904). An initial CPMV v0 of the first control point (902) may be refined based on template matching. An initial CPMV v1 of the second control point (904) may be refined by a DMVR. To refine the initial CPMV v1, a refined CPMV for the second control point (904) may be determined by adding a translational MV offset to the initial CPMV v1, where the translational MV offset may be derived based on the DMVR.

[0113]

[0120] FIG. 11 shows a flow chart outlining a process (1100) according to an embodiment of the present disclosure. The process (1100) can be used in a video decoder. In various embodiments, the process (1100) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), etc. In some embodiments, the process (1100) is implemented with software instructions, and thus, the processing circuit performs the process (1100) when it executes the software instructions. The process begins at (S1101) and proceeds to (S1110).

[0114]

[0121] At (S1110), a video bitstream including a current block in a current picture is received, the current block being coded in an affine mode, and a first control point associated with the affine mode being located at a first corner of the current block.

[0115]

[0122] At (S1120), a current template associated with the first control point is determined, where the current template is adjacent to the first control point.

[0116]

[0123] At (S1130), a plurality of candidate reference templates are determined in the reference picture for a current template associated with the first control point.

[0117]

[0124] At (S1140), a reference template is selected from the plurality of candidate reference templates for the current template associated with the first control point based on a TM cost, where the TM cost indicates a respective difference between the current template of the first control point and each of the candidate reference templates.

[0118]

[0125] At (S1150), a first CPMV is determined based on the selected reference template, where the first CPMV indicates an offset between the selected reference template and a current template associated with the first control point in the reference picture.

[0119]

[0126] In (S1160), the current block is reconstructed based on at least the first CPMV.

[0120]

[0127] In one example, a first block is determined, for which a first control point is located at the center of the first block, and a current template associated with the first control point is determined as a reconstructed region located at one or a combination of (i) the top side of the first block and (ii) the left side of the first block.

[0121]

[0128] In one example, a current template associated with a first control point is determined as a reconstructed region adjacent to the first control point that includes at least one of (i) a first region above the current block, or (ii) a second region to the left of the current block.

[0122]

[0129] In one example, the current template associated with the first control point is determined as a reconstructed region, and the first control point is at the center of the reconstructed region, the reconstructed region including at least one of (i) a first region located above the current block and extending beyond a vertical edge of the current block, and (ii) a second region located to the left of the current block and extending beyond a horizontal edge of the current block.

[0123]

[0130] In one example, the first region has a height equal to N samples and a width equal to the width of the affine sub-block of the current block, and the second region has a width equal to N samples and a height equal to the height of the affine sub-block of the current block, where N is a positive integer.

[0124]

[0131] In one example, an initial first CPMV is determined for a first control point, where the initial first CPMV indicates an initial reference template in a reference picture. Within a search range of the initial reference template, multiple candidate reference templates are determined. The search range includes M×M pixels, where M is less than 8.

[0125]

[0132] In one example, a plurality of candidate reference templates are determined within a search range of the initial reference template based on a plurality of search steps, the number of which is determined based on one of a predetermined resolution and a predetermined number.

[0126]

[0133] In one example, a TM cost between a current template associated with a first control point and each of a plurality of candidate reference templates is determined, and a reference template corresponding to a minimum TM cost among the determined TM costs between the current template associated with the first control point and each of the plurality of candidate reference templates is selected from the plurality of candidate reference templates.

[0127]

[0134] In one example, a template of a current block is determined, the template including a first region above the current block and a second region to the left of the current block. A first candidate CPMV is determined for a first control point based on a first candidate reference template from a plurality of candidate reference templates. A second candidate CPMV is determined for the first control point based on a second candidate reference template from a plurality of candidate reference templates. A first set of sub-block affine MVs is determined for sub-blocks of the current block adjacent to the template of the current block based on at least the first candidate CPMV. A second set of sub-block affine MVs is determined for sub-blocks of the current block adjacent to the template of the current block based on at least the second candidate CPMV. A first set of reference sub-block affine MVs is determined for sub-blocks of an equivalently positioned block of the current block corresponding to the first set of sub-block affine MVs. A second set of reference subblock affine MVs is determined for a subblock of the equivalently positioned block of the current block corresponding to the second set of subblock affine MVs. A first reference template is determined based on the first set of reference subblock affine MVs, and a second reference template is determined based on the second set of reference subblock affine MVs. A first TM cost between the template of the current block and the first reference template is determined. A second TM cost between the template of the current block and the second reference template is determined. One of the first candidate reference template and the second candidate reference template is selected as the reference template. One of the first candidate reference template and the second candidate reference template corresponds to the smaller of the first TM cost and the second TM cost.

[0128]

[0135] In one example, a reference subblock is determined for each of the subblocks of the equally positioned block based on each of the first set of reference subblock affine MVs, a subblock template is determined for each of the reference subblocks of the subblocks of the equally positioned block, and a first reference template is determined as a combination of the subblock templates.

[0129]

[0136] In one example, the affine mode includes one of an affine mono-prediction mode and an affine bi-prediction mode.

[0130]

[0137] In one example, the current block includes a second control point at a second corner of the current block. An initial second CPMV is determined for the second control point. The second CPMV is determined for the second control point by adding a translational MV offset to the initial second CPMV, where the translational MV offset is derived based on the DMVR.

[0131]

[0138] The process then proceeds to (S1199) and ends.

[0132]

[0139] The method 1100 may be adapted as appropriate. Steps of the process 1100 may be modified and / or omitted. Additional steps may be added. Any suitable order of execution may be used.

[0133]

[0140] FIG. 12 shows a flow chart outlining a process (1200) according to an embodiment of the present disclosure. The process (1200) can be used in a video encoder. In various embodiments, the process (1200) is performed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), etc. In some embodiments, the process (1200) is implemented by software instructions, and thus, the processing circuit performs the process (1200) when it executes the software instructions. The process begins at (S1201) and proceeds to (S1210).

[0134]

[0141] At (S1210), a current block in a current picture is determined to be coded in affine mode, and a first control point associated with the affine mode is located at a first corner of the current block.

[0135]

[0142] At (S1220), a current template associated with the first control point is determined, where the current template is adjacent to the first control point.

[0136]

[0143] At (S1230), a plurality of candidate reference templates are determined in a reference picture of the current template associated with the first control point.

[0137]

[0144] At (S1240), a reference template is determined from the plurality of candidate reference templates for the current template associated with the first control point based on a TM cost indicating a difference between the current template of the first control point and each of the plurality of candidate reference templates. In one example, determining the reference template can include selecting a reference template from the plurality of candidate reference templates for the current template associated with the first control point based on a TM cost indicating a difference between each of the plurality of candidate reference templates and the current template of the first control point.

[0138]

[0145] At (S1250), a first CPMV is determined based on the determined reference template, where the first CPMV indicates an offset between the determined reference template in the reference picture and a current template associated with the first control point.

[0139]

[0146] In (S1260), the current block is coded based on at least the first CPMV.

[0140]

[0147] The process then proceeds to (S1299) and ends.

[0141]

[0148] The method 1200 may be adapted as appropriate. Steps of the process 1200 may be modified and / or omitted. Additional steps may be added. Any suitable order of execution may be used.

[0142]

[0149] The techniques described above may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, Figure 13 illustrates a computer system (1300) suitable for implementing certain embodiments of the disclosed subject matter.

[0143]

[0150] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code that contains instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or that may be executed via interpretation, microcode execution, etc.

[0144]

[0151] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0145]

[0152] 13 for computer system 1300 are exemplary in nature and are not intended to suggest any limitation on the scope of functionality or application of the computer software implementing embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement related to any one or combination of components illustrated in the exemplary embodiment of computer system 1300.

[0146]

[0153] The computer system (1300) may include certain human interface input devices that may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), auditory input (e.g., voice, claps), visual input (e.g., gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that do not necessarily involve direct conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still-image cameras), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic pictures).

[0147]

[0154] The input human interface devices may include one or more of (only one of each is depicted) a keyboard (1301), a mouse (1302), a trackpad (1303), a touch screen (1310), a data glove (not shown), a joystick (1305), a microphone (1306), a scanner (1307), and a camera (1308).

[0148]

[0155] The computer system 1300 may also include certain human interface output devices that may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices can include haptic output devices (e.g., haptic feedback via a touch screen (1310), data gloves (not shown), and joystick (1305), although there may be haptic feedback devices that do not function as input devices), auditory output devices (e.g., speakers (1309), headphones (not shown)), visual output devices (e.g., screens (1310), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output, three-dimensional or higher output by means such as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0149]

[0156] The computer system (1300) may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (1320) using media such as CD / DVD (1321), thumb drives (1322), removable hard drives or solid state drives (1323), legacy magnetic media (not shown) such as tape and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0150]

[0157] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not encompass transmission media, carrier waves, or other transitional signals.

[0151]

[0158] The computer system (1300) may also include interfaces to one or more communications networks (1355). Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide area, metropolitan, automotive, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, and terrestrial TV), automotive networks including CANBus, etc. Particular networks typically require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (1349) (e.g., a USB port on the computer system (1300)); others are commonly integrated into the core of the computer system (1300) by attaching to a system bus, as described below (e.g., an Ethernet interface is integrated in a PC computer system, and a cellular network interface is integrated in a smartphone computer system). Using any of these networks, the computer system (1300) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional, such as with other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks, as described above, can be used with each of these networks and network interfaces.

[0152]

[0159] The aforementioned human interface devices, human accessible storage devices, and network interfaces can be attached to the core (1340) of the computer system (1300).

[0153]

[0160] The core (1340) may include one or more central processing units (CPUs) (1341), graphics processing units (GPUs) (1342), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (1343), task-specific hardware accelerators (1344), graphics adapters (1350), etc. These devices, along with read-only memory (ROM) (1345), random access memory (1346), and internal mass storage devices (e.g., internal non-user-accessible hard drives, SSDs, etc.) (1347), may be connected via a system bus (1348). In some computer systems, the system bus (1348) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1348) or via a peripheral bus (1349). In one example, a screen 1310 can be connected to a graphics adapter 1350. Peripheral bus architectures include PCI, USB, etc.

[0154]

[0161] The CPU (1341), GPU (1342), FPGA (1343), and accelerator (1344) may combine to execute specific instructions that may constitute the aforementioned computer code. The computer code may be stored in ROM (1345) or RAM (1346). Temporary data may be stored in RAM (1346), while persistent data may be stored, for example, in internal mass storage (1347). Rapid storage and retrieval from some memory device may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU (1341), GPU (1342), mass storage (1347), ROM (1345), RAM (1346), etc.

[0155]

[0162] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those having skill in the computer software arts.

[0156]

[0163] By way of example and not limitation, the architecture (1300), and in particular a computer system having a core (1340), can provide functionality as a result of operations by a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as specific storage of the core (1340) that is non-transitory in nature, such as the core's internal mass storage (1347) or ROM (1345). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1340). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (1340) and, in particular, the processor (including a CPU, GPU, FPGA, etc.) therein to perform certain processes or portions of certain processes described herein, including defining data structures stored in RAM (1346) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embedded in circuitry (e.g., accelerator (1344)), which may execute in place of or in conjunction with software to perform a particular process or portion of a particular process described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any appropriate combination of hardware and software.

[0157]

[0164] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the listed elements. For example, at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and reference to at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, or one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of listed elements, where applicable, for example, when the elements are not mutually exclusive.

[0158]

[0165] While this disclosure describes a number of exemplary embodiments, there are modifications, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be understood that those skilled in the art will be able to devise many systems and methods that, while not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.

[0159]

[0166] <<Additional Notes>> (Appendix 1) 1. A video decoding method comprising: receiving a video bitstream including a current block coded in an affine mode in a current picture, wherein a first control point associated with the affine mode is located at a first corner of the current block; determining a current template associated with the first control point, the current template being in a vicinity of the first control point; determining a plurality of candidate reference templates in a reference picture for a current template associated with the first control point; selecting a reference template from the plurality of candidate reference templates for a current template associated with the first control point based on a template matching (TM) cost indicating a respective difference between the current template for the first control point and each candidate reference template; determining a first control point motion vector (CPMV) based on a selected reference template, the first CPMV indicating an offset between the selected reference template and a current template associated with the first control point in the reference picture; and reconstructing the current block based on at least the first CPMV; A method comprising:

[0160] (Appendix 2) 10. The method of claim 1, wherein determining a current template associated with the first control point further comprises: determining the first block such that the first control point is located at the center of the first block; and determining a current template associated with the first control point as a reconstructed region located at one or a combination of (i) an upper side of the first block, and (ii) a left side of the first block; A method comprising:

[0161] (Appendix 3) 10. The method of claim 1, wherein determining a current template associated with the first control point further comprises: determining a current template associated with the first control point as a reconstructed region in the vicinity of the first control point, the current template including at least one of (i) a first region above the current block, or (ii) a second region to the left of the current block; A method comprising:

[0162] (Appendix 4) 10. The method of claim 1, wherein determining a current template associated with the first control point further comprises: determining a current template associated with the first control point as a reconstructed region, the first control point being at a center of the reconstructed region, the reconstructed region including at least one of (i) a first region located above the current block and extending beyond a vertical edge of the current block, and (ii) a second region located to the left of the current block and extending beyond a horizontal edge of the current block.

[0163] (Appendix 5) In the method described in Appendix 3: the first region has a height equal to N samples and a width equal to the width of an affine sub-block of the current block; and The method, wherein the second region comprises a width equal to N samples and a height equal to the height of an affine sub-block of the current block, where N is a positive integer.

[0164] (Appendix 6) 12. The method of claim 1, wherein determining a plurality of candidate reference templates in the reference picture further comprises: determining an initial first CPMV for the first control point, the initial first CPMV indicating an initial reference template in the reference picture; and determining the plurality of candidate reference templates within a search range of the initial reference template, the search range including M×M pixels, where M is less than 8; A method comprising:

[0165] (Appendix 7) 7. The method of claim 6, wherein the step of determining the plurality of candidate reference templates within a search range of the initial reference template further comprises: determining the plurality of candidate reference templates within a search range of the initial reference template based on a plurality of search steps, wherein the number of the plurality of search steps is determined based on one of a predetermined resolution and a predetermined number.

[0166] (Appendix 8) 10. The method of claim 1, wherein the step of selecting a reference template from the plurality of candidate reference templates further comprises: determining a TM cost between a current template associated with the first control point and each of the plurality of candidate reference templates; and selecting a reference template from the plurality of candidate reference templates that corresponds to a minimum TM cost determined between a current template associated with the first control point and each of the plurality of candidate reference templates; A method comprising:

[0167] (Appendix 9) 10. The method of claim 1, wherein the step of selecting a reference template from the plurality of candidate reference templates further comprises: determining a template of the current block including (ii) a first region above the current block, and (iii) a second region to the left of the current block; determining a first candidate CPMV for the first control point based on a first candidate reference template of the plurality of candidate reference templates and a second candidate CPMV for the first control point based on a second candidate reference template of the plurality of candidate reference templates; determining a first set of sub-block affine MVs for sub-blocks of the current block that are in a neighborhood of a template of the current block based on at least the first candidate CPMV and a second set of sub-block affine MVs for sub-blocks of the current block that are in a neighborhood of a template of the current block based on at least the second candidate CPMV; determining a first set of reference sub-block affine MVs for sub-blocks of the equivalently positioned block of the current block corresponding to the first set of sub-block affine MVs, and a second set of reference sub-block affine MVs for sub-blocks of the equivalently positioned block of the current block corresponding to the second set of sub-block affine MVs; determining a first reference template based on the first set of reference sub-block affine MVs and a second reference template based on the second set of reference sub-block affine MVs; determining a first TM cost between the template of the current block and the first reference template and a second TM cost between the template of the current block and the second reference template; and selecting one of the first candidate reference template and the second candidate reference template as the reference template, wherein the one of the first candidate reference template and the second candidate reference template corresponds to a smaller one of the first TM cost and the second TM cost; A method comprising:

[0168] (Appendix 10) 10. The method of claim 9, wherein the step of determining a first reference template based on a first set of reference sub-block affine MVs further comprises: determining a reference sub-block for each of the sub-blocks of the co-located block based on each of the first set of reference sub-block affine MVs; determining a sub-block template for each of the reference sub-blocks of the sub-blocks of the equally positioned block; and determining the first reference template as a combination of the sub-block templates; A method comprising:

[0169] (Appendix 11) 2. The method of claim 1, wherein the affine mode comprises one of an affine mono-prediction mode and an affine bi-prediction mode.

[0170] (Appendix 12) In the method described in Appendix 1: The current block includes a second control point at a second corner of the current block, and the method further comprises: determining an initial second CPMV for the second control point; and determining a second CPMV for the second control point by adding a translational MV offset to the initial second CPMV, the translational MV offset being derived based on decoder-side motion vector refinement (DMVR); A method comprising:

[0171] (Appendix 13) 1. An apparatus including a processing circuit, the processing circuit comprising: receiving a video bitstream including a current block coded in an affine mode in a current picture, wherein a first control point associated with the affine mode is located at a first corner of the current block; determining a current template associated with the first control point, the current template being in a vicinity of the first control point; determining a plurality of candidate reference templates in a reference picture for a current template associated with the first control point; selecting a reference template from the plurality of candidate reference templates for a current template associated with the first control point based on a template matching (TM) cost indicating a respective difference between the current template for the first control point and each candidate reference template; determining a first control point motion vector (CPMV) based on a selected reference template, the first CPMV indicating an offset between the selected reference template and a current template associated with the first control point in the reference picture; and reconstructing the current block based on at least the first CPMV; The apparatus is configured to:

[0172] (Appendix 14) 14. The apparatus of claim 13, wherein the processing circuitry further comprises: determining the first block such that the first control point is located at the center of the first block; and determining a current template associated with the first control point as a reconstructed region located at one or a combination of (i) an upper side of the first block, and (ii) a left side of the first block; The apparatus is configured to:

[0173] (Appendix 15) 14. The apparatus of claim 13, wherein the processing circuitry further comprises: determining a current template associated with the first control point as a reconstructed region in the vicinity of the first control point, the current template including at least one of (i) a first region above the current block, or (ii) a second region to the left of the current block; The apparatus is configured to:

[0174] (Appendix 16) 14. The apparatus of claim 13, wherein the processing circuitry: 1. An apparatus configured to perform the step of determining a current template associated with the first control point as a reconstructed region, the first control point being at a center of the reconstructed region, the reconstructed region including at least one of: (i) a first region located above the current block and extending beyond a vertical edge of the current block; and (ii) a second region located to the left of the current block and extending beyond a horizontal edge of the current block.

[0175] (Appendix 17) 16. The apparatus of claim 15, the first region has a height equal to N samples and a width equal to the width of an affine sub-block of the current block; and The second region has a width equal to N samples and a height equal to the height of an affine sub-block of the current block, where N is a positive integer.

[0176] (Appendix 18) 14. The apparatus of claim 13, wherein the processing circuitry: determining an initial first CPMV for the first control point, the initial first CPMV indicating an initial reference template in the reference picture; and determining the plurality of candidate reference templates within a search range of the initial reference template, the search range including M×M pixels, where M is less than 8; The apparatus is configured to:

[0177] (Appendix 19) 19. The apparatus of claim 18, wherein the processing circuitry: The apparatus is configured to determine the plurality of candidate reference templates within a search range of the initial reference template based on a plurality of search steps, wherein the number of the plurality of search steps is determined based on one of a predetermined resolution and a predetermined number.

[0178] (Appendix 20) 14. The apparatus of claim 13, wherein the processing circuitry: determining a TM cost between a current template associated with the first control point and each of the plurality of candidate reference templates; and selecting a reference template from the plurality of candidate reference templates that corresponds to a minimum TM cost determined between a current template associated with the first control point and each of the plurality of candidate reference templates; The apparatus is configured to:

Claims

1. 1. A video decoding method comprising: receiving a video bitstream including a current block coded in an affine mode in a current picture, wherein a first control point associated with the affine mode is located at a first corner of the current block; determining a current template associated with the first control point, the current template being in a vicinity of the first control point; determining a plurality of candidate reference templates in a reference picture for a current template associated with the first control point; selecting a reference template from the plurality of candidate reference templates for a current template associated with the first control point based on a template matching (TM) cost indicating a respective difference between the current template of the first control point and each candidate reference template; determining a first control point motion vector (CPMV) based on a selected reference template, the first CPMV indicating an offset between the selected reference template and a current template associated with the first control point in the reference picture; and reconstructing the current block based on at least the first CPMV; A method comprising:

2. 10. The method of claim 1, wherein determining a current template associated with the first control point further comprises: determining the first block such that the first control point is located at the center of the first block; and determining a current template associated with the first control point as a reconstructed region located at one or a combination of (i) an upper side of the first block, and (ii) a left side of the first block; A method comprising:

3. 10. The method of claim 1, wherein determining a current template associated with the first control point further comprises: determining a current template associated with the first control point as a reconstructed region in the vicinity of the first control point, the reconstructed region including at least one of (i) a first region above the current block, or (ii) a second region to the left of the current block; A method comprising:

4. 10. The method of claim 1, wherein determining a current template associated with the first control point further comprises: determining a current template associated with the first control point as a reconstructed region, the first control point being at a center of the reconstructed region, the reconstructed region including at least one of (i) a first region located above the current block and extending beyond a vertical edge of the current block, and (ii) a second region located to the left of the current block and extending beyond a horizontal edge of the current block.

5. 4. The method of claim 3, wherein: the first region has a height equal to N samples and a width equal to the width of an affine sub-block of the current block; and The method of claim 1, wherein the second region has a width equal to N samples and a height equal to the height of an affine sub-block of the current block, where N is a positive integer.

6. 10. The method of claim 1, wherein the step of determining a plurality of candidate reference templates in the reference picture further comprises: determining an initial first CPMV for the first control point, the initial first CPMV indicating an initial reference template in the reference picture; and determining the plurality of candidate reference templates within a search range of the initial reference template, the search range including M×M pixels, where M is less than 8; A method comprising:

7. 7. The method of claim 6, wherein the step of determining the plurality of candidate reference templates within a search range of the initial reference template further comprises: determining the plurality of candidate reference templates within a search range of the initial reference template based on a plurality of search steps, wherein the number of the plurality of search steps is determined based on one of a predetermined resolution and a predetermined number.

8. 10. The method of claim 1, wherein the step of selecting a reference template from the plurality of candidate reference templates further comprises: determining a TM cost between a current template associated with the first control point and each of the plurality of candidate reference templates; and selecting a reference template from the plurality of candidate reference templates that corresponds to a minimum TM cost determined between a current template associated with the first control point and each of the plurality of candidate reference templates; A method comprising:

9. 10. The method of claim 1, wherein the step of selecting a reference template from the plurality of candidate reference templates further comprises: determining a template for the current block including (i) a first region above the current block, and (ii) a second region to the left of the current block; determining a first candidate CPMV for the first control point based on a first candidate reference template of the plurality of candidate reference templates and a second candidate CPMV for the first control point based on a second candidate reference template of the plurality of candidate reference templates; determining a first set of sub-block affine MVs for sub-blocks of the current block that are in a neighborhood of a template of the current block based on at least the first candidate CPMV and a second set of sub-block affine MVs for sub-blocks of the current block that are in a neighborhood of a template of the current block based on at least the second candidate CPMV; determining a first set of reference sub-block affine MVs for sub-blocks of equivalently positioned blocks of the current block corresponding to the first set of sub-block affine MVs and a second set of reference sub-block affine MVs for sub-blocks of equivalently positioned blocks of the current block corresponding to the second set of sub-block affine MVs; determining a first reference template based on the first set of reference sub-block affine MVs and a second reference template based on the second set of reference sub-block affine MVs; determining a first TM cost between the template of the current block and the first reference template and a second TM cost between the template of the current block and the second reference template; and selecting one of the first candidate reference template and the second candidate reference template as the reference template, wherein the one of the first candidate reference template and the second candidate reference template corresponds to a smaller one of the first TM cost and the second TM cost; A method comprising:

10. 10. The method of claim 9, wherein the step of determining a first reference template based on the first set of reference sub-block affine MVs further comprises: determining a reference sub-block for each of the sub-blocks of the equally positioned block based on each of the first set of reference sub-block affine MVs; determining a sub-block template for each of the reference sub-blocks of the sub-blocks of the equally positioned block; and determining the first reference template as a combination of the sub-block templates; A method comprising:

11. 2. The method of claim 1, wherein the affine mode includes one of an affine mono-prediction mode and an affine bi-prediction mode.

12. 10. The method of claim 1 : The current block includes a second control point at a second corner of the current block, and the method further comprises: determining an initial second CPMV for the second control point; and determining a second CPMV for the second control point by adding a translational MV offset to the initial second CPMV, the translational MV offset being derived based on decoder-side motion vector refinement (DMVR); A method comprising:

13. A computer program product that causes a computer to carry out the method according to any one of claims 1 to 12.

14. Video processing device adapted to execute the method according to any one of claims 1 to 12.

15. 1. A video encoding method comprising: determining that a current block in the current picture is coded in affine mode, and a first control point associated with the affine mode is located at a first corner of the current block; determining a current template associated with the first control point, the current template being in a vicinity of the first control point; determining a plurality of candidate reference templates in a reference picture of the current template associated with the first control point; determining a reference template from the plurality of candidate reference templates for a current template associated with the first control point based on a template matching (TM) cost indicative of a difference between the current template of the first control point and each of the plurality of candidate reference templates; determining a first control point motion vector (CPMV) based on the reference template determined in the determining step, wherein the first CPMV indicates an offset between the determined reference template in the reference picture and a current template associated with the first control point; encoding the current block based on at least a first CPMV; A method comprising:

Citation Information

Patent Citations

  • Template matching based affine prediction for video coding

    US20220329823A1