Method and device for video processing and medium
By refining the motion vectors of control points in affine motion candidates using template matching technology, the problem of inaccurate motion estimation in existing technologies is solved, thereby improving the efficiency and effectiveness of encoding and decoding.
Patent Information
- Application Number
- CN202380090370.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-03
- Filing Date
- 2023-12-28
- Publication Date
- 2025-08-08
AI Technical Summary
In existing affine motion compensation methods, the estimation of control point motion vectors may not be consistent with the actual motion, resulting in low encoding and decoding efficiency.
Template matching technology is used to refine the motion vectors of control points in affine motion candidates, thereby improving the accuracy of motion information.
It improves the efficiency and effectiveness of video encoding and decoding.
Smart Images

Figure CN120457694A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly to affine motion candidate refinement. Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes: determining, for a conversion between a current video block of a video and a bitstream of the video, a set of control point motion vectors (CPMVs) associated with affine motion candidates for the current video block; determining a refined affine motion candidate by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs; and performing a conversion based on the refined affine motion candidate. The method according to the first aspect of the present disclosure refines the affine motion candidate and the CPMV. This improves codec efficiency and codec effectiveness.
[0005] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing. The method includes: determining a set of control point motion vectors (CPMVs) associated with affine motion candidates for a current video block of the video; determining a refined affine motion candidate by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs; and generating a bitstream based on the refined affine motion candidate.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining a set of control point motion vectors (CPMVs) associated with affine motion candidates for a current video block of the video; determining a refined affine motion candidate by applying a refinement process to at least one CPMV in the set of CPMVs based on template matching; generating a bitstream based on the refined affine motion candidate; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0012] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0014] Figure 4 Shows the locations of spatial and temporal neighboring blocks used in AMVP / Merge candidate list construction;
[0015] Figure 5 The positions of non-adjacent candidates in the ECM are shown;
[0016] Figure 6A An affine motion model based on 4-parameter control points is shown;
[0017] Figure 6B An affine motion model based on 6 parameter control points is shown;
[0018] Figure 7 The affine MVF of each sub-block is shown;
[0019] Figure 8 shows the position of the inherited affine motion prediction value;
[0020] Figure 9Shows control point motion vector inheritance;
[0021] Figure 10 The positions of candidate positions for constructing the affine Merge pattern are shown;
[0022] Figure 11 shows the spatial neighbors used to derive affine merge candidates, where Figure 11 (a) in is used to derive inherited affine Merge candidates, and Figure 11 (B) in the figure is used to derive the constructed affine Merge candidate;
[0023] Figure 12 Shows the graph from non-contiguous neighbors to constructed affine merge candidates;
[0024] Figure 13 An example of generating a HAPC is shown;
[0025] Figure 14 A diagram showing the regression-based affine merge candidate derivation;
[0026] Figure 15 shows the template matching performed on the search area around the initial MV;
[0027] Figure 16 The template and the corresponding reference template are shown;
[0028] Figure 17 A template of a block with sub-block motion and a reference template using motion information of a sub-block of a current block are shown;
[0029] Figure 18 A diagram showing derivation of a sub-CU motion field obtained by applying motion displacement based on neighboring motion information;
[0030] Figure 19 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and
[0031] Figure 20 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0032] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION
[0033] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0034] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0035] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.
[0036] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0037] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0038] Figure 1 is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0039] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0040] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0041] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0042] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0043] Figure 2 is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0044] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0045] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0046] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0047] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0048] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0049] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0050] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0051] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0052] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0053] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0054] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0055] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0056] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0057] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0058] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0059] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0060] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0061] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0062] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0063] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0064] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blockiness artifacts in the video block.
[0065] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0066] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0067] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0068] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0069] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.
[0070] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.
[0071] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.
[0072] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0073] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0074] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0075] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present disclosure relates to video coding technology. Specifically, it relates to an affine motion prediction method in video coding. This concept can be applied alone or in various combinations to any video coding standard or non-standard video codec. 2. Introduction The exponential growth of multimedia data has brought severe challenges to video coding and decoding. In order to meet the growing demand for more efficient compression technology, ITU-T and ISO / IEC have developed a series of video coding and decoding standards over the past few decades. Specifically, ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 video, and the two organizations jointly developed H.262 / MPEG-2 video, H.264 / MPEG-4 Advanced Video Codec (AVC), H.265 / HEVC and the latest VVC standards. Starting with H.262 / MPEG-2, a hybrid video coding and decoding framework was adopted, in which intra / inter prediction plus transform coding and decoding were used. 2.1. MVP in Video Codec Inter-frame prediction aims to remove temporal redundancy between adjacent frames and is an integral component of hybrid video coding frameworks. Specifically, inter-frame prediction uses the content specified by the motion vector (MV) as a predicted version of the current block to be coded, thereby transmitting only the residual signal and motion information in the bitstream. To reduce the cost of MV signaling, motion vector prediction (MVP) emerged as an effective mechanism for conveying motion information. Early strategies simply used the MV of a specified neighboring block or the median MV of neighboring blocks as the MVP. In H.265 / HEVC, a competition mechanism is introduced, in which the optimal MVP is selected from multiple candidates through rate-distortion optimization (RDO). Specifically, the Advanced MVP (AMVP) mode and Merge mode are designed with different motion information signaling strategies. In AMVP mode, the reference index, the reference AMVP candidate list, and the MVP candidate index of the motion vector difference (MVD) are signaled. In Merge mode, only the merge index of the reference merge candidate list is signaled, and all motion information associated with the merge candidate is inherited. Both the AMVP mode and the Merge mode require the construction of an MVP candidate list. The details of the construction process of these two modes are described below. AMVP mode: AMVP uses the spatiotemporal correlation between motion vectors and neighboring blocks for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by first checking the availability of the left and upper temporal neighbors, removing redundant candidates and adding zero vectors to make the candidate list length constant. Figure 4 A diagram 400 shows the locations of spatial and temporal neighboring blocks used in AMVP / Merge candidate list construction. For spatial motion vector candidate derivation, as Figure 4 As shown in Figure 1, two motion vector candidates are finally derived based on the motion vectors of blocks located at five different positions. The five neighboring blocks located at B0, B1, B2 and A0, A1 are divided into two groups, where group A includes the three spatial neighboring blocks above and group B includes the two spatial neighboring blocks on the left. In a predefined order, the two motion vector candidates are derived from the first available candidate in group A and group B respectively. For the derivation of the temporal motion vector candidate, as shown in Figure 1, Figure 4 As shown, a motion vector candidate is derived based on two different co-located positions (lower right (C0) and center (C1)) examined sequentially. To avoid MV candidate redundancy, duplicate motion vector candidates in the list are discarded. If the number of potential candidates is less than two, an additional zero motion vector candidate is added to the list. Merge mode: Similar to AMVP mode, the MVP candidate list for Merge mode also includes spatial candidates and temporal candidates. For spatial motion vector candidate derivation, after performing availability and redundancy checks, up to four candidates are selected in the order of A1, B1, B0, A0 and B2. For temporal Merge candidate (TMVP) derivation, up to one candidate is selected from two temporal neighboring blocks (C0 and C1). When there are not enough Merge candidates using spatial candidates and temporal candidates, the combined bidirectional prediction Merge candidate and zero MV candidate are added to the MVP candidate list. Once the number of available Merge candidates reaches the maximum allowed number transmitted by signal, the Merge candidate list construction process is terminated. In VVC, the construction process for Merge mode is further improved by introducing history-based MVP (HMVP), which combines the motion information of previously coded blocks that may be far away from the current block. In VVC, HMVP Merge candidates are appended to the Merge list, after spatial MVP and TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained using a first-in-first-out strategy. Whenever there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. During the VVC standardization process, non-adjacent MVP was proposed to promote better motion information derivation by adopting non-adjacent regions. In ECM software, non-adjacent MVP is inserted between TMVP and HMVP, where the distance between non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block, such as Figure 5 5 , which illustrates a diagram 500 showing the locations of non-neighboring candidates in an ECM. 2.2. Affine Motion Compensated Prediction In HEVC, only the translational motion model is applied to motion-compensated prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion-compensated prediction is applied. Figure 6A A diagram 610 showing an affine motion model based on 4-parameter control points is shown. Figure 6B FIG620 shows an affine motion model based on 6 parameter control points. Figure 11 As shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). For the 4-parameter affine motion model, the motion vector at the sample position (x,y) in the block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: Among them (mv 0x ,mv 0y ) is the motion vector of the upper left control point, (mv 1x ,mv 1y ) is the motion vector of the upper right control point, and (mv 2x ,mv 2y ) is the motion vector of the lower left control point. Figure 7 An example diagram 1200 of the affine MVF for each sub-block is shown. To simplify motion compensated prediction, block-based affine transform prediction is applied. To derive the motion vector for each 4×4 luminance sub-block, the motion vector of the center sample of each sub-block is calculated according to the above equation, as Figure 12 As shown, and rounded to 1 / 16 fractional precision. A motion compensated interpolation filter is then applied to generate a prediction for each subblock with the derived motion vector. The subblock size of the chroma component is also set to 4×4. The MV of the 4×4 chroma subblock is calculated as the average of the MVs of the upper left and lower right luma subblocks in the same 8×8 luma region. For translational motion inter prediction, there are two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.2.1. Affine Merge Prediction Affine Merge mode can be applied to CUs with width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of the spatially adjacent CUs. There can be up to five CPMVP candidates, and the index is transmitted by signal to indicate the CPMVP candidate to be used for the current CU. In VVC, the following three types of CPMV candidates are used to form the affine Merge candidate list: – Inherited affine Merge candidates, which are inferred from the CPMV of neighboring CUs, – The constructed affine Merge candidate CPMVP derived using the translation MV of the neighboring CU, –Zero MV. In VVC, there are a maximum of two inherited affine candidates derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the above neighboring CU. Figure 8 An example 800 is shown of the location of the inherited affine motion predictor. The candidate block is Figure 8As shown in . For the left prediction value, the scanning order is A0->A1, and for the above prediction value, the scanning order is B0->B1->B2. Only the first inherited candidate is selected from each side. No deduplication check is performed between two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine merge list of the current CU. Figure 9 An example diagram 900 illustrating control point motion vector inheritance is shown. Figure 9 As shown, if the neighboring lower left block A 910 is encoded and decoded in affine mode, the motion vectors v2, v3, and v4 of the upper left corner, upper right corner, and lower left corner of the CU containing block A are obtained. When block A is encoded and decoded using a 4-parameter affine model, the two CPMVs of the current CU are calculated based on and. When block A is encoded and decoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated based on v2, v3, and v4. Constructing affine candidates means constructing candidates by combining the translation motion information of the nearest neighbors of each control point. Figure 10 Figure 1000 shows the positions of candidate positions for constructing the affine Merge pattern. The motion information of the control points is obtained from Figure 10 The derivation of the specified spatial and temporal nearest neighbors is shown in CPMV. k (k=1, 2, 3, 4) represents the kth control point. For CPMV1, check B2->B3->A2 blocks and use the MV of the first available block. For CPMV2, check B1>B0 blocks, and for CPMV3, check A1>A0 blocks. For TMVP, if available, use TMVP as CPMV4. After obtaining the MVs of the four control points, construct the affine merge candidate based on that motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}. The combination of 3 CPMVs constructs a 6-parameter affine Merge candidate, and the combination of 2 CPMVs constructs a 4-parameter affine Merge candidate. To avoid the motion scaling process, the relevant combination of control point MVs is discarded if the reference indices of the control points are different. After checking the inherited affine merge candidates and the constructed affine merge candidates, if the list is still not complete, a zero MV is inserted at the end of the list. 2.2.2. Affine AMVP Prediction Affine AMVP mode can be applied to CUs with width and height both greater than or equal to 16. An affine flag in the CU level is signaled in the bitstream to indicate whether affine AMVP mode is used, and another flag is signaled to indicate 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2, and is generated in sequence by the following four types of CPMV candidates: – inherited affine AMVP candidates, which are inferred from the CPMVs of neighboring CUs, – The constructed affine AMVP candidate CPMVP is derived using the translation MV of the neighboring CU, – translation MV from neighboring CU, –Zero MV. The order in which inherited affine AMVP candidates are checked is the same as the order in which inherited affine Merge candidates are checked. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as the current block are considered. When the inherited affine motion predictor is inserted into the candidate list, no deduplication process is applied. The AMVP candidates are constructed from Figure 10 The specified spatial neighbor derivation is shown in . The same check order is used in the affine Merge candidate construction. In addition, the reference picture index of the neighboring block is also checked. The first block in the check order that uses inter-frame coding and has the same reference picture as the current CU. Only when the current CU is encoded and decoded with a 4-parameter affine mode and both mv0 and mv1 are available, they are added as a candidate in the affine AMVP list. When the current CU is encoded and decoded using a 6-parameter affine mode and all three CPMVs are available, it is added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable. If the affine AMVP list candidate is still less than 2 after inserting the valid inherited affine AMVP candidate and the constructed AMVP candidate, mv0, mv1 and mv2 will be added as translation MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, the affine AMVP list is filled with zero MVs. 2.2.3. New Affine Candidate Derivation Method in ECM-6.0 In ECM-6.0, three additional affine merge and AMVP candidate derivation methods are integrated, which are based on non-adjacent spatial domain candidates, historical parameter-based candidates, and regression-based affine candidates. 2.2.3.1. Non-adjacent airspace candidates In ECM-6.0, non-adjacent spatial neighbors are studied to provide candidates for both affine merge and affine AMVP. Figure 11 The spatial neighbors used to derive affine merge candidates are shown in Figure 2. The pattern for obtaining non-adjacent spatial candidates is shown in Figure 2. Figure 11 As shown in . Similar to the non-adjacent conventional Merge candidates, the distance between the non-adjacent spatial candidates and the current coding block is also defined based on the width and height of the current CU. use Figure 11 The motion information of the non-adjacent spatial neighbors in the is used to generate additional inherited and constructed affine merge candidates. Specifically, to generate inherited candidates, the non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from near to far). At a certain distance, only the first available neighbor that is encoded and decoded from each side of the current block (e.g., left and above) using the affine mode is included. Figure 11 As shown in (a), the inspection of the neighbors on the left and above is performed from bottom to top and from right to left respectively. For the constructed candidate, Figure 11 As shown in (b), the position of a left and upper non-adjacent spatial neighbor is first determined independently; then, the position of the upper left neighbor can be determined accordingly to form a rectangular virtual block together with the left and upper non-adjacent neighbors. Figure 12 FIG1200 shows the affine merge candidate constructed from non-adjacent neighbors. The motion information of three non-adjacent neighbors is used to form CPMVs at the top left (A), top right (B), and bottom left (C) of the virtual block, which are projected to the current CU to generate the corresponding constructed candidates, as shown in FIG1200. Figure 12 shown. 2.2.3.2. Affine Candidates Based on Historical Parameters Affine Model Inheritance Based on Historical Parameters (HAMI) allows affine models to be inherited from previously affine-encoded blocks that may not be adjacent to the current block. A History Parameter Table (HPT) is established. An entry in the HPT stores a set of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. The entries in the HPT are classified by reference list and reference index. Each reference list in the HPT supports 5 reference indices. In terms of formula, the category of the HPT (denoted as HPTCat) is calculated as HPTCat (RefList, RefIdx)= 5×RefList + min(RefIdx, 4) (3) Where RefList and RefIdx represent the reference picture list (0 or 1) and the reference index, respectively. For each category, up to seven entries can be stored, resulting in a total of 70 entries in the HPT. The number of entries for each category is initialized to zero at the beginning of each CTU row. After decoding an affine codec CU with reference list RefListCur and RefIdxcur, the affine parameters are used to update the entries in category HPTCat(RefListCur, RefIdxcur) in a manner similar to the HMVP table update. Figure 13 An example diagram 1300 of generating HAPC is shown. The candidate (HAPC) based on the historical affine parameters is obtained from Figure 13 The affine parameters of the adjacent 4x4 blocks denoted as A0, A1, B0, B1 or B2 in the HPT are derived from the corresponding entries in the HPT. The MV of the adjacent 4x4 blocks is used as the base MV. In the formal way, the MV of the current block at position (x, y) is calculated as: Where (mvhbase, mvvbase) represents the MV of the adjacent 4×4 block, (x base ,y base ) represents the center position of the neighboring 4x4 block. (x, y) can be the top-left, top-right, and bottom-left corners of the current block to obtain the angular position MV (CPMV) of the current block, or it can be the center of the current block to obtain the normal MV of the current block. Figure 13 It is shown how HAPC is derived from block A0. The affine parameters {a0, b0, c0, d0} are extracted directly from an entry of the class HPTIdx(Ref List A0, refIdx0A0) in the HPT. The affine parameters from the HPT (with the center position of A0 as the base position and the MV of block A0 as the base MV) are used together to derive the CPMV of the Affine Merge HAPC or the Affine AMVP HAPC. They can also be used to derive the MV at the center of the current block as a regular merge candidate. The HAPC can be put into a sub-block based merge candidate list, an affine AMVP candidate list, or a regular merge candidate list. In response to the introduction of the new HAPC, the size of the sub-block based merge candidate list is increased from 5 to 10 and 12 for the random access and low latency B configurations, respectively. In addition, for the random access configuration, the size of the regular merge candidate list is increased from ten to eleven to accommodate the newly added regular merge candidates. 2.2.3.3. Regression-based affine candidates In ECM-6.0, regression-based affine merge candidates are derived and added to the affine merge list. The sub-block motion fields from the previously encoded affine CU and the motion information from the neighboring sub-blocks of the current CU are used as input to the regression process to derive the proposed affine candidates. The previously encoded affine CU can be identified through non-adjacent locations and a scan of the affine HMVP table. Figure 14 1400 illustrates regression-based affine merge candidate derivation. Figure 14 As depicted in , the neighboring sub-block information of the current CU is extracted from the 4x4 sub-blocks represented by the gray area. For each sub-block, given a reference list, the corresponding motion vector and center coordinates of the sub-block can be used. For each affine CU, up to two affine candidates can be derived. One with neighboring sub-block information and one without. All candidates generated by linear regression are deduplicated and collected into a candidate subgroup. When ARMC is enabled, the TM cost-based ARMC process is applied. Afterwards, when N affine CUs are found, up to N candidates generated by linear regression are added to the affine merge list. 2.3. Template Matching Merge / AMVP Mode in ECM Template Matching (TM) Merging / AMVP mode is a decoder-side MV derivation method that refines the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the neighboring blocks above and / or to the left of the current CU) and a block in a reference picture (i.e., the same size as the template). Figure 15 An example diagram 1500 is shown showing template matching performed on a search area around an initial MV. Figure 15 As shown, a better MV is searched around the initial motion of the current CU in the [-8, +8] pixel search range. In AMVP mode, MVP candidates are determined based on template matching error, selecting the MVP candidate that achieves the minimum difference between the current block and the reference block template. The motion management (TM) then performs MV refinement only for this specific MVP candidate. The TM refines this MVP candidate using an iterative diamond search within a [–8, +8] pixel search range, starting with full-pixel MVD accuracy (or 4-pixel AMVR mode). The AMVP candidate is further refined using a cross-search using full-pixel MVD accuracy (or 4-pixel AMVR mode), followed by half-pixel and quarter-pixel searches, depending on the AMVR mode. This search process ensures that the MVP candidate, after TM processing, maintains the same MV accuracy as indicated by the adaptive motion vector resolution (AMVR) mode. In Merge mode, a similar search method is applied to the Merge candidate indicated by the Merge index. TMMerge can be performed all the way to 1 / 8 pixel MVD accuracy, or skip accuracy beyond half-pixel MVD accuracy, depending on whether the motion information is merged using an alternative interpolation filter (used when AMVR is in half-pixel mode). In addition, when TM mode is enabled, template matching can be performed as an independent process between block-based and sub-block-based bilateral matching (BM) methods or as an additional MV refinement process, depending on whether BM can be enabled according to the enable condition check. When BM and TM are enabled on a CU at the same time, the search process of TM will stop at half-pixel MVD accuracy, and the resulting MV will be further refined by using the same model-based MVD derivation method as in DMVR. 2.4. Adaptive Reordering of Merge Candidates (ARMC) Inspired by the spatial correlation between reconstructed neighboring pixels and the current codec block, we propose Adaptive Reordering of Merge Candidates (ARMC) to refine the order of candidates in a given candidate list. The basic assumption is that candidates with lower template matching costs have a higher probability of being selected by the RDO process and should therefore be placed at the front of the list to reduce signaling cost. This reordering method is applied to the normal Merge mode, the Template Matching (TM) Merge mode and the Affine Merge mode (excluding SbTMVP candidates).For the TM Merge mode, the Merge candidates are reordered before the refinement process. After building the merge candidate list, the merge candidates are divided into several subgroups. The subgroup size is set to 5. The merge candidates in each subgroup are reordered in ascending order based on the cost value based on template matching. For simplicity, merge candidates in the last subgroup but not in the first subgroup are not reordered. The template matching cost is measured by the sum of absolute differences (SAD) between the samples of the current block's template and its corresponding reference template. Figure 16 A diagram 1600 is shown of a template and a corresponding reference template. The template includes a set of reconstructed samples adjacent to the current block, while the reference template is positioned by the same motion information of the current block, such as Figure 7 When the Merge candidate uses bidirectional prediction, the reference samples of the template of the Merge candidate are also generated by bidirectional prediction. For a sub-block based Merge candidate with a sub-block size equal to Wsub*Hsub, the upper template includes several sub-templates with a size of Wsub×K, and the left template includes several sub-templates with a size of K×Hsub. Figure 17FIG1700 shows a template of a block with sub-block motion and a reference template using motion information of a sub-block of a current block. Figure 17 As shown in Figure 2, the motion information of the sub-block in the first row and first column of the current block is used to derive the reference sample of each sub-template. 2.5. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports the sub-block based temporal motion vector prediction (SbTMVP) method. Similar to TMVP, SbTMVP utilizes the motion field in the co-located picture to facilitate more accurate MVP derivation. The same co-located picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP mainly in two aspects. First, SbTMVP enables sub-CU level motion prediction, while TMVP predicts motion at the CU level; second, compared with TMVP which extracts the temporal MV from the co-located block in the co-located picture (the co-located block is the lower right or center block relative to the current CU), SbTMVP applies motion displacement before extracting the temporal motion information from the co-located picture, where the motion displacement is obtained by using the MV of one of the spatial neighboring blocks from the current CU. Figure 18 The derivation process 1800 of the sub-block-level motion field for SbTMVP is shown. Specifically, the motion information of the lower left sub-block A1 is first extracted. If any of the MVs in reference list 0 and list 1 points to the same frame, the corresponding MV is identified as the motion displacement. Otherwise, zero MV is used as the motion displacement. Once the motion displacement is determined, the sub-block level motion field is derived using the designated region in the collocated frame. Assume that A1' motion is used as the motion displacement, e.g. Figure 18 Then, for each sub-CU, the motion information of its corresponding block (the minimum motion grid covering the center sample) in the co-located picture is extracted to provide motion information, where an MV scaling operation is first performed to align the reference frame of the temporal motion vector with the reference frame of the current CU. Figure 18 Derivation of the sub-CU motion field obtained by applying motion displacement based on neighboring motion information is shown. In VVC and ECM, in addition to the CU-level MVP candidate list, a sub-CU-level MVP candidate list is constructed to provide more accurate motion prediction for the current CU, which includes the motion fields generated by both the SbTMVP and affine methods. Specifically, only one SbTMVP candidate is included and is always placed in the first entry of the constructed sub-CU-level MVP candidate list, while after performing template matching-based reordering, multiple affine candidates are included in the list, with those affine candidates with smaller costs placed in the front position. 3. Question CPMV is crucial for affine motion compensation because it provides basic motion information for all sub-blocks within a block. However, in existing CPMV derivation methods, the CPMV of the current block is estimated as the MV of the block that has already been coded, which may not guarantee consistency with the true motion. Therefore, CPMV refinement methods are highly desired to reduce the deviation between the estimated CPMV and the true motion. 4. Detailed solution In this disclosure, we propose to refine the affine CPMV using template matching. For a given affine candidate in the affine candidate list, we can further refine the CPMV using template matching, and then use the refined affine candidate to derive sub-block or pixel-level affine motion information for the current block. The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow manner. In addition, these embodiments can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB. The term "affine block" may refer to a block encoded and decoded using Affine Merge, Affine AMVP, or any other affine variant mode (i.e., Affine MMVD, etc.), which may be described by motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). The term "CPMV" may refer to motion information of an affine block at the top left corner, top right corner, and / or bottom left corner. The term "template" may refer to a reconstruction region that can be used to refine a CPMV, and may refer to a "separate template" or a "unified template." Here, a "separate template" may refer to a reconstruction region that can be used to refine an individual CPMV (i.e., a specific one or more of the top left corner, top right corner, and / or bottom left corner), while a "unified template" may refer to a reconstruction region that can be used to refine all or any (multiple) CPMVs of a block. The term "template matching cost" or "TM cost" may refer to the matching cost of a separate template or a unified template. In the present disclosure, regarding “blocks encoded and decoded using mode N”, “mode N” here can be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.), or a coding and decoding technology (e.g., DIMD, TIMD, PDPC, CCLM, CCCM, GLM, TMP, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-frame, GPM intra-frame, GPM intra-frame, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.). Note that the terms mentioned below are not limited to the specific terms defined in existing standards, and any changes in codec tools are also applicable. 1. In one example, affine motion compensation can be refined by using previously decoded samples. a) In one example, at least one CPMV may be refined. b) In one example, at least one MV of a sub-block of affine motion compensation may be refined. c) In one example, at least one affine parameter (such as a, b, c, d, e, f) may be refined. d) In one example, the previously decoded samples can be the template of the current block. e) In one example, the previously decoded sample may be a template of a reference block. f) In one example, the template representation can be used to refine the reconstructed region of the CPMV. g) In one example, for affine-coded blocks, different individual templates may be used for different control points. i. In one example, for a control point, the corresponding separate template may include sample points from adjacent and / or non-adjacent locations in the reconstructed region. 1) In one example, separate templates of all control points are collected from adjacent reconstruction regions. 2) In one example, separate template samples for some control points are collected from adjacent reconstructed areas of the current block, while for the remaining control points, template samples are collected from non-adjacent reconstructed areas. a) In one example, specifically, template samples for the upper left corner are collected from non-adjacent regions, while template samples for the upper right corner and / or lower left corner are collected from adjacent regions. 3) In one example, both adjacent and non-adjacent sample points are used for some control points. ii. In one example, the shape of a separate template may be different for different control points. 1) In one example, for certain control points, an L-shaped (eg, including both upper and left neighboring points) separate template is used. 2) In one example, for certain control points, an I-shaped or “—”-shaped template may be used (eg, including neighboring points on the left or above (but not both)). iii. In one example, which shape of the template is used for CPMV refinement can be based on the location of the control points / position. 1) In one example, the CPMV at the upper left corner of the current video unit may use an L-shaped template (eg, including both upper and left neighboring samples). 2) In one example, the CPMV at the upper right corner of the current video unit may use a '—' shape template (eg, including only upper neighboring samples). 3) In one example, the CPMV at the upper left corner of the current video unit may use an I-shaped template (eg, including only left neighboring samples). 4) In one example, for a certain CPMV, the shapes of the templates in the current picture and the reference picture are the same. a) For example, Figure 16 As depicted in , the template of the CPMV may refer to a set of neighboring samples in the current picture (eg, the template in the current picture) and a second set of neighboring samples in the reference picture (eg, the template in the reference picture). iv. In one example, the number of sample points used in the template may be different for different control points. 1) Alternatively, the number of samples for different control points is the same for the affine block. 2) For different control points, the arrangement (rows or columns) of sample points used in the template can be different. h) In one example, a unified template is used during CPMV refinement. i. In one example, the TM cost associated with the unified template is used to determine the MV shift value. ii. In one example, the TM cost associated with the unified template is used to determine the CPMV combination. iii. In one example, the unified template may include all or part of the adjacent samples of the entire block, such as Figure 16 shown. i) A template may include samples from only one component (eg luma) or from multiple components (eg luma and chroma). j) In one example, for any template, the reference template region with the same shape can be located using MV, as shown in the figure. k) In one example, the template may not necessarily include all pixels in a specific area, but may include some pixels in the specified area. 2. When constructing the affine candidate list, CPMV refinement can be performed on the potential affine candidates first, and then the refined candidates are inserted into the affine candidate list. a) In one example, alternatively, CPMV refinement is performed after building the affine candidate list. i. In one example, only the affine candidate(s) with specific index(es) need to be executed CPMV refinement. ii. In one example, all or part of the affine candidates need to perform CPMV refinement. 3. In one example, the first affine candidate list is constructed first, followed by the second affine candidate list construction process. a) For example, the input of the second affine candidate list generation can be based on the output of the first affine candidate list generation. b) For example, the first affine candidate list can be constructed without CPMV refinement. c) For example, the second affine candidate list may be generated by applying CPMV refinement to the CPMV candidates in the first affine candidate list. i. For example, at least one CPMV candidate in the first affine candidate list may be refined. ii. Alternatively, more than one CPMV candidate in the affine candidate list may be refined. iii. For example, CPMV refinement can be based on TM. d) For example, a candidate re-ranking process may be utilized to construct a first affine candidate list. i. For example, the reordering process can be based on TM. e) For example, the second affine candidate list can be constructed without any candidate re-ranking process. f) For example, different deduplication rules may be used in the first deduplication and the second deduplication. i. For example, the first affine candidate list generation may be applied in association with the first deduplication method. ii. For example, a second affine candidate list generation may be applied in association with the second deduplication method. iii. For example, the thresholds for motion similarity checking in the first deduplication method and the second deduplication method may be different. iv. For example, a threshold based on block dimensions (eg, block width and / or height) may be used in the second deduplication method. v. For example, alternatively, the second deduplication method may employ a fixed threshold. 4. For a given affine candidate, some or all of the CPMVs may be refined based on the TM, and then the refined CPMVs are used to derive affine motion information for the current block and / or sub-blocks. a) In one example, both integer precision and fractional precision can be used to refine control points. i. In one example, only integer precision is used to refine control points, and fractional precision search is skipped. 1) In one example, whether a fractional precision search is required depends on the result of an integer precision search. ii. In one example, it is proposed to use a specific interpolation filter to generate a reference template for motion vectors pointing to fractional positions. 1) In one example, a simplified interpolation filter may be applied. 2) In one example, the simplified interpolation filter may be a 2-tap bilinear, alternatively, it may be a 4-tap, 6-tap, or 8-tap filter of DCT, DST, Lanczos, or any other interpolation type. 3) In one example, a more complex interpolation filter (eg, with longer filter taps) may be applied. iii. In one example, whether to use the above method (e.g., integer precision, different interpolation filters) and / or how to use the above method can be signaled in the bitstream (e.g., in SPS, PPS, The information may be in the picture header, slice header, CTU, CU, etc.) or determined on the fly based on decoded information. 1) In one example, the method to be applied may depend on the codec tool. 2) In one example, the method to be applied may depend on the block dimension. b) In one example, different control points are refined separately, which means that the MV shift value (ie, the difference between the initial CPMV and the corresponding refined CPMV) may be different for different control points. i. In one example, all or some of the control points may first be refined separately by TM, and then one combination of control points may be refined by traversing all or some combinations of CPMVs before and after refinement (i.e., M combinations (such as M=4) for a 4-parameter model, N combinations (such as N= 8) combination) to determine and derive a set of CPMVs that minimize the TM cost of the current block. 1) In one example, in the above case, all or some of the control points may first be refined by corresponding individual templates. 2) In one example, for each combination of CPMV, sub-block level motion information is calculated for the boundary sub-blocks, and then the sub-block level motion information is calculated according to Zhang Jie 2.4 and Figure 17The unified TM cost is computed by the method described in . The best combination that produces the minimum TM cost is selected as the refined affine candidate. a) In one example, only some boundary sub-blocks need to calculate the TM cost. 3) In one example, alternatively, there is no need to loop over all combinations, and the combinations whose control points are all refined by the TM are directly used as refined affine candidates. 4) In one example, when refined affine candidates are derived, a second pass through control point refinement may be performed to further refine each control point. a) In one example, each CPMV is further iteratively refined to minimize the TM cost of the current block. In each iteration, one CPMV is refined while the others are fixed. c) In one example, alternatively, multiple control points are refined simultaneously, where the same MV shift value is shared for all or multiple control points. i. In one example, all or part of the MV displacement values in a given MV displacement set are traversed one by one. The traversed MV displacement values are assigned to all or more CPMVs, and then the motion information of the boundary sub-blocks associated with the refined CPMVs is calculated, and the TM cost is formulated accordingly. In this process, the value that produces the minimum TM cost is determined as the optimal motion displacement value, which can be finally used for refinement. CPMV. d) In one example, the refined affine candidate may replace the original affine candidate. i. In one example, the refined affine candidate will always replace the original affine candidate. ii. In one example, alternatively, the refined affine candidate will conditionally replace the original affine candidate. 1) In one example, specifically, the TM costs associated with the original CPMV (referred to as C_beforeTM) and the refined CPMV (C_afterTM) are calculated separately, and the refined affine candidate replaces the original affine candidate only when the ratio of C_afterTM to C_beforeTM is less than (or greater than) a constant or adaptively determined value TH. a) In one example, in the above case, different codec modes (eg, affine Merge / affine AMVP / affine MMVD) may have different TH value settings. iii. Alternatively, the refined affine candidate may be used as the new candidate. 1) In one example, the refined affine candidate may be placed adjacent to the original affine candidate in the affine candidate list (ie, immediately before or after the original affine candidate). 2) In one example, alternatively, the refined affine candidate can be placed at an arbitrary position in the affine candidate list. 3) In one example, the refined affine candidate can be compared with at least one candidate already in the candidate list. If they are identical or similar, it is not added to the list. 5. CPMV refinement can be used together with regression-based affine candidate inference methods. a) In one example, after refining all or some CPMVs using TM (generating Affine_model_TM), motion information of the boundary sub-blocks associated with Affine_model_TM is derived and then fed into a regression model to output a new affine model (called Affine_model_R). The TM costs of the boundary sub-blocks using Affine_model_TM and Affine_model_R are then calculated and compared, respectively. The one with the lower TM cost is determined as the final refined affine candidate. i. In one example, all or some of the CPMVs may first undergo integer precision TM refinement (producing Affine_model_TM_I), then fractional precision TM refinement is performed (producing Affine_model_TM_F). And the motion information of the boundary sub-block associated with Affine_model_TM_I is derived, which is then fed into the regression model to output a new affine model (Affine_model_R). Finally, the TM of the boundary sub-block using Affine_model_TM_F and Affine_model_R is calculated and compared. costs, and the sub-block with the smaller TM cost is determined as the final refined affine candidate. ii. In one example, only some sub-blocks may need to calculate TM costs to generate Affine_model_TM, Affine_model_TM_I and / or Affine_model_TM_F. 6. In one example, TM-based refinement can be applied to Affine Merge or Affine AMVP (Affine Inter). a) In one example, the MVP(s) of the affine AMVP may be refined based on the TM. i. Alternatively, the MVP(s) of the affine AMVP can be refined based on DMVR. 7. In one example, TM-based refinement can be applied together with DMRS-based refinement to affine-coded blocks. a) In one example, TM-based refinement can be applied before DMVR. b) In one example, TM-based refinement can be applied after DMVR. c) Alternatively, TM-based refinement can be applied to affine-coded blocks mutually exclusively with DMRS-based refinement. 8. In one example, the derivation of the TM cost may depend on whether the block is bi-directionally predicted or uni-directionally predicted. a) If the block is bi-predicted, the TM cost can be derived based on the bi-prediction on the TM. i. In one example, TM ref0 and TM ref1 References associated with List0 and List1 respectively TM, then the final reference TM (TM bi )for: TM bi =a*TM ref0 +(1-a)*TM ref1 . 1) In one example, a is equal to 0.5. 2) In one example, a is determined based on the BCW index. 3) In one example, generate a TM based on the CPMV in List 0 ref0 , and / or generate TM based on the CPMV in List 1 ref1 . b) Alternatively, if the block is bi-directionally predicted, the TM cost can be calculated for List0 and List1 separately. 9. In one example, the refinement of the CPMV can be done in an iterative manner. a) For example, in one refinement step, one CPMV is refined while the other CPMVs are fixed. b) In one example, when a subsequent CPMV is to be refined, the already refined(s) CPMV. i. In one example, alternatively, when a subsequent CPMV is to be refined, the CPMV before the refinement is used. c) In one example, the refinement of the CPMV may be done in an iterative manner for bi-directionally predicted blocks. i. In one example, the CPMV associated with list K (K=0 or 1) may be refined first, and then the CPMV associated with list (1-K) may be refined. 1) Whether and / or how to refine the CPMV in the later lists (1-K) may be determined based on the refined CPMV of the previous list K. ii. In one example, the CPMVs associated with List 0 and List 1 may be refined separately. 1) In one example, specifically, when the CPMVs in list K (K=0 or 1) are refined, for each search step, a one-way reference TM in list K is generated based on the corresponding CPMV, And the TM cost is calculated accordingly to determine the optimal MV shift value. iii. In one example, alternatively, the CPMVs associated with List 0 and List 1 may be jointly refined. 1) In one example, specifically, when the CPMV in list K (K=0 or 1) is refined, for each search step, a bidirectional reference TM is generated based on the CPMV information of the two lists. (As in Figure 8 The one that produces the minimum TM cost is determined as the optimal MV shift value. 10. CPMV can be refined in multiple rounds. a) In one example, all or part of the CPMV may be refined in each round of refinement. b) In one example, all or part of the CPMV may have been refined in a previous round of refinement, and then a later round is conducted to further refine the CPMV. 11. Whether and / or how to refine the TM-based CPMV may be determined based on the prediction direction of the current block. a) In one example, the CPMV may need to be refined by TM only when the current block is unidirectionally predicted. b) In one example, the CPMV may need to be refined by TM only when the current block is bi-directionally predicted. c) In one example, no matter whether the current block is bi-predicted or not, it may always be necessary to refine the CPMV by TM. 12. If affine prediction is used as hypothesis, the disclosed method can be applied to MHP (Multiple Hypothesis Prediction) coded blocks. 13. Whether and / or how to apply the methods disclosed above can be determined based on syntax elements. a) At least one syntax element is signaled, for example, in a bitstream. b) For example, whether and / or how the disclosed method can be applied can be signaled at the sequence level / GOP level / picture level / slice level / slice group level, e.g. in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / PPS / strip header / slice group header are transmitted through signals. c) For example, whether and / or how the disclosed method can be applied to PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of regions containing more than one sample or pixel. d) For example, whether and / or how to apply the above disclosed method may depend on coded information such as block size, color format, single / dual tree partitioning, color component, slice / picture type. e) For example, whether a syntax element is signaled (i.e., indicating whether TM refinement is applied to CPMV) It may be determined based on another syntax element.
[0076] Figure 19 FIG2 is a flowchart of a method 1900 for video processing according to an embodiment of the present disclosure. The method 1900 may be implemented for conversion between a current video block of a video and a bitstream of the video.
[0077] At block 1910, a set of control point motion vectors (CPMVs) associated with affine motion candidates for a current video block is determined. At block 1920, a refined affine motion candidate is determined by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs. At block 1930, a conversion is performed based on the refined affine motion candidate.
[0078] Method 1900 can refine affine motion candidates. For example, CPMV can be refined. In this way, codec effectiveness and codec efficiency can be improved.
[0079] In some embodiments, the affine motion candidate is replaced by the refined affine motion candidate. That is, the refined affine candidate can replace the original affine candidate. For example, in some embodiments, the refined affine candidate can always replace the original affine candidate.
[0080] In some embodiments, whether an affine motion candidate is replaced by a refined affine motion candidate is based on a condition associated with at least one CPMV, that is, the refined affine candidate will conditionally replace the original affine candidate.
[0081] In some embodiments, the condition includes a ratio of a first template matching cost of at least one refined CPMV to a second template matching cost of at least one CPMV being less than a threshold.
[0082] In some embodiments, the condition includes a ratio of a first template matching cost of at least one refined CPMV to a second template matching cost of at least one CPMV being greater than a threshold.
[0083] In some embodiments, the threshold is fixed or determined during the conversion.
[0084] In some embodiments, the threshold is determined based on a codec mode of the current video block.
[0085] In some embodiments, the first threshold for the first codec mode is different from the second threshold for the second codec mode.
[0086] In some embodiments, the encoding / decoding mode includes one of: an affine Merge mode, an affine Advanced Motion Vector Prediction (AMVP) mode, or an affine Merge with Motion Vector Difference (MMVD) mode.
[0087] In some embodiments, an affine motion candidate is replaced by a refined affine motion candidate based on a condition being met. For example, the TM costs associated with the original CPMV (referred to as C_beforeTM) and the refined CPMV (C_afterTM) are calculated separately, and the refined affine candidate replaces the original affine candidate only when the ratio of C_afterTM to C_beforeTM is less than (or greater than) a constant or adaptively determined value TH.
[0088] In some embodiments, the refined affine motion candidate is used as a new candidate that is different from the affine motion candidate. That is, the refined affine candidate can be used as a new candidate.
[0089] In some embodiments, the affine motion candidate is in an affine candidate list, and the refined affine motion candidate is placed adjacent to the affine motion candidate in the affine candidate list.
[0090] In some embodiments, the refined affine motion candidate is placed in the affine candidate list at a position immediately before or after the affine motion candidate.
[0091] In some embodiments, the affine motion candidate is in the affine candidate list, and the refined affine motion candidate is placed in the affine candidate list. For example, the refined affine candidate can be placed at any position in the affine candidate list.
[0092] In some embodiments, method 1900 further includes: determining a difference between the refined affine motion candidate and another affine motion candidate in the affine candidate list; if the difference is determined to be less than or equal to a threshold, maintaining the affine candidate list without adding the refined motion candidate to the affine candidate list; and if the difference is determined to be greater than the threshold, adding the refined affine motion candidate to the affine candidate list. In other words, the refined affine candidate can be compared with at least one candidate already in the candidate list. If they are identical or similar, it is not added to the list.
[0093] In some embodiments, method 1900 further includes: determining an affine candidate list for the current video block, the affine candidate list comprising a set of affine candidates for the current video block; and applying a CPMV refinement process to at least one affine candidate in the affine candidate list. For example, all or some of the affine candidates may require CPMV refinement.
[0094] In some embodiments, affine prediction is used as the hypothesis for a current video block that is encoded using multiple hypothesis prediction (MHP).
[0095] In some embodiments, whether and / or how the method is applied is based on syntax elements in the bitstream.
[0096] In some embodiments, the syntax element is at at least one of: sequence level, group of pictures level, picture level, slice level, or slice group level.
[0097] In some embodiments, the syntax element is included in at least one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
[0098] In some embodiments, syntax elements are indicated in regions that include more than one sample or pixel.
[0099] In some embodiments, the region includes one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.
[0100] In some embodiments, whether and / or how to apply the method is determined based on codec information of the current video block.
[0101] In some embodiments, the codec information includes at least one of the following: a block size of the current video block, a color format of the current video block, a single-tree partitioning or a dual-tree partitioning of the current video block, a color component of the current video block, a slice type of the current video block, or a picture type of the current video block.
[0102] In some embodiments, whether a first syntax element is determined based on a second syntax element, the first syntax element indicating whether a template matching based refinement process is applied to the control point motion vectors of the current video block.
[0103] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a video bitstream generated by a method performed by an apparatus for video processing. In the method, a set of affine motion candidates associated with affine motion candidates for a current video block of the video is determined. A refined affine motion candidate is determined by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs. A bitstream is generated based on the refined affine motion candidate.
[0104] According to further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, a set of affine motion candidates associated with a current video block of the video is determined. A refined affine motion candidate is determined by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs. A bitstream is generated based on the refined affine motion candidate. The bitstream is stored in a non-transitory computer-readable recording medium.
[0105] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.
[0106] Item 1. A method for video processing, comprising: determining, for a transformation between a current video block of a video and a bitstream of the video, a set of control point motion vectors (CPMVs) associated with affine motion candidates for the current video block; determining a refined affine motion candidate by applying a refinement process to at least one CPMV in the set of CPMVs based on template matching; and performing the transformation based on the refined affine motion candidate.
[0107] Item 2. The method of Item 1, wherein the affine motion candidate is replaced by the refined affine motion candidate.
[0108] Clause 3. The method of clause 1, wherein whether the affine motion candidate is replaced by the refined affine motion candidate is based on a condition associated with the at least one CPMV.
[0109] Clause 4. The method of clause 3, wherein the condition comprises a ratio of a first template matching cost of the at least one refined CPMV to a second template matching cost of the at least one CPMV being less than a threshold.
[0110] Clause 5. The method of clause 3, wherein the condition comprises a ratio of a first template matching cost of the at least one refined CPMV to a second template matching cost of the at least one CPMV being greater than a threshold.
[0111] Clause 6. The method of clause 4 or 5, wherein the threshold is fixed or determined during the conversion.
[0112] Clause 7. The method of clause 6, wherein the threshold is determined based on a codec mode of the current video block.
[0113] Item 8. The method of Item 7, wherein the first threshold for the first codec mode is different from the second threshold for the second codec mode.
[0114] Item 9. The method of Item 7, wherein the encoding / decoding mode comprises one of: an affine Merge mode, an affine Advanced Motion Vector Prediction (AMVP) mode, or an affine Merge with Motion Vector Difference (MMVD) mode.
[0115] Item 10. The method of any one of Items 4 to 9, wherein based on the condition being satisfied, the affine motion candidate is replaced by the refined affine motion candidate.
[0116] Item 11. The method of Item 1, wherein the refined affine motion candidate is used as a new candidate different from the affine motion candidate.
[0117] Item 12. The method of Item 11, wherein the affine motion candidate is in an affine candidate list, and the refined affine motion candidate is placed adjacent to the affine motion candidate in the affine candidate list.
[0118] Item 13. The method of Item 12, wherein the refined affine motion candidate is placed in the affine candidate list at a position immediately before or after the affine motion candidate.
[0119] Item 14. The method of Item 11, wherein the affine motion candidate is in an affine candidate list, and the refined affine motion candidate is placed in the affine candidate list.
[0120] Item 15. The method according to Item 11 further includes: determining a difference between the refined affine motion candidate and another affine motion candidate in the affine candidate list; if it is determined that the difference is less than or equal to a threshold, maintaining the affine candidate list without adding the refined motion candidate to the affine candidate list; and if it is determined that the difference is greater than the threshold, adding the refined affine motion candidate to the affine candidate list.
[0121] Item 16. The method according to any one of Items 1 to 15 further includes: determining an affine candidate list for the current video block, the affine candidate list comprising a set of affine candidates for the current video block; and applying a CPMV refinement process to at least one affine candidate in the affine candidate list.
[0122] Item 17. The method of any one of Items 1 to 16, wherein affine prediction is used as a hypothesis for the current video block encoded using multiple hypothesis prediction (MHP).
[0123] Clause 18. A method according to any one of clauses 1 to 17, wherein whether and / or how to apply the method is based on syntax elements in the bitstream.
[0124] Clause 19. The method of clause 18, wherein the syntax element is at at least one of: sequence level, group of pictures level, picture level, slice level, or slice group level.
[0125] Item 20. A method according to Item 18 or Item 19, wherein the syntax element is included in at least one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a temporal group header.
[0126] Clause 21. A method according to any one of clauses 18 to 20, wherein the syntax element is indicated in a region comprising more than one sample or pixel.
[0127] Item 22. The method of Item 21, wherein the region comprises one of: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.
[0128] Item 23. The method of any one of Items 1 to 22, wherein whether and / or how to apply the method is determined based on codec information of the current video block.
[0129] Item 24. A method according to Item 23, wherein the codec information includes at least one of the following: a block size of the current video block, a color format of the current video block, a single-tree partitioning or a dual-tree partitioning of the current video block, a color component of the current video block, a slice type of the current video block, or a picture type of the current video block.
[0130] Item 25. A method according to any one of items 1 to 24, wherein whether a first syntax element is determined based on a second syntax element, the first syntax element indicating whether a template matching based refinement process is applied to the control point motion vector of the current video block.
[0131] Item 26. The method of any one of Items 1 to 25, wherein the converting comprises encoding the current video block into the bitstream.
[0132] Item 27. The method of any one of Items 1 to 25, wherein the converting comprises decoding the current video block from the bitstream.
[0133] Item 28. An apparatus for video processing, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 27.
[0134] Item 29. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 27.
[0135] Item 30. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a set of control point motion vectors (CPMVs) associated with affine motion candidates for a current video block of the video; determining a refined affine motion candidate by applying a refinement process to at least one CPMV in the set of CPMVs based on template matching; and generating the bitstream based on the refined affine motion candidate.
[0136] Item 31. A method for storing a bitstream of a video, comprising: determining a set of control point motion vectors (CPMVs) associated with affine motion candidates for a current video block of the video; determining a refined affine motion candidate by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs; generating the bitstream based on the refined affine motion candidate; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0137] Figure 20A block diagram of a computing device 2000 in which various embodiments of the present disclosure may be implemented is shown. The computing device 2000 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0138] It should be understood that Figure 20 The computing device 2000 shown in FIG. 2 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.
[0139] like Figure 20 As shown, computing device 2000 comprises a general computing device 2000. Computing device 2000 may include at least one or more processors or processing units 2010, memory 2020, storage unit 2030, one or more communication units 2040, one or more input devices 2050, and one or more output devices 2060.
[0140] In some embodiments, the computing device 2000 can be implemented as any user terminal or server terminal with computing power. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 2000 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0141] The processing unit 2010 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 2020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 2000. The processing unit 2010 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0142] The computing device 2000 typically includes various computer storage media. Such media can be any media accessible by the computing device 2000, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 2020 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 2030 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk or other media that can be used to store information and / or data and can be accessed in the computing device 2000.
[0143] The computing device 2000 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 20 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0144] The communication unit 2040 communicates with another computing device via a communication medium. In addition, the functionality of the components in the computing device 2000 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 2000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0145] The input device 2050 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. The output device 2060 may be one or more of various output devices, such as a display, a speaker, a printer, and the like. With the aid of the communication unit 2040, the computing device 2000 may also communicate with one or more external devices (not shown), such as storage devices and display devices, one or more devices that enable a user to interact with the computing device 2000, or, if desired, any device that enables the computing device 2000 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).
[0146] In some embodiments, some or all components of the computing device 2000 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers in a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein may be provided by a conventional server or installed directly or otherwise on a client device.
[0147] In an embodiment of the present disclosure, the computing device 2000 may be used to implement video encoding / decoding. The memory 2020 may include one or more video encoding / decoding modules 2025 having one or more program instructions. These modules are accessible and executable by the processing unit 2010 to perform the functions of the various embodiments described herein.
[0148] In an example embodiment performing video encoding, input device 2050 may receive video data as input to be encoded 2070. The video data may be processed, for example, by video codec module 2025 to generate an encoded bitstream. The encoded bitstream may be provided as output 2080 via output device 2060.
[0149] In an example embodiment performing video decoding, input device 2050 may receive an encoded bitstream as input 2070. The encoded bitstream may be processed, for example, by video codec module 2025 to generate decoded video data. The decoded video data may be provided as output 2080 via output device 2060.
[0150] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: determining, for a conversion between a current video block of a video and a bitstream of the video, a set of control point motion vectors (CPMVs) associated with affine motion candidates for the current video block; determining a refined affine motion candidate by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs; as well as The conversion is performed based on the refined affine motion candidate. The method of claim 1 , wherein the affine motion candidate is replaced by the refined affine motion candidate. 3 . The method of claim 1 , wherein whether the affine motion candidate is replaced by the refined affine motion candidate is based on a condition associated with the at least one CPMV. 4 . The method according to claim 3 , wherein the condition includes a ratio of a first template matching cost of the at least one refined CPMV to a second template matching cost of the at least one CPMV being less than a threshold. 5 . The method according to claim 3 , wherein the condition includes a ratio of a first template matching cost of the at least one refined CPMV to a second template matching cost of the at least one CPMV being greater than a threshold. The method according to claim 4 , wherein the threshold value is fixed or determined during the conversion. The method of claim 6 , wherein the threshold is determined based on a codec mode of the current video block.
8. The method of claim 7, wherein a first threshold for a first codec mode is different from a second threshold for a second codec mode.
9. The method of claim 7, wherein the codec mode comprises one of the following: Affine Merge mode, Affine Advanced Motion Vector Prediction (AMVP) mode, or Affine Merge with Motion Vector Difference (MMVD) mode. 10 . The method according to claim 4 , wherein based on the condition being satisfied, the affine motion candidate is replaced by the refined affine motion candidate. The method of claim 1 , wherein the refined affine motion candidate is used as a new candidate different from the affine motion candidate. 12 . The method of claim 11 , wherein the affine motion candidate is in an affine candidate list, and the refined affine motion candidate is placed adjacent to the affine motion candidate in the affine candidate list. 13 . The method of claim 12 , wherein the refined affine motion candidate is placed at a position immediately before or after the affine motion candidate in the affine candidate list. The method of claim 11 , wherein the affine motion candidate is in an affine candidate list, and the refined affine motion candidate is placed in the affine candidate list.
15. The method according to claim 11, further comprising: determining a difference between the refined affine motion candidate and another affine motion candidate in an affine candidate list; if it is determined that the difference is less than or equal to a threshold, maintaining the affine candidate list without adding the refined motion candidate to the affine candidate list; as well as If it is determined that the difference is greater than the threshold, the refined affine motion candidate is added to the affine candidate list.
16. The method according to any one of claims 1 to 15, further comprising: determining an affine candidate list for the current video block, the affine candidate list comprising a set of affine candidates for the current video block; as well as A CPMV refinement process is applied to at least one affine candidate in the affine candidate list.
17. The method according to any one of claims 1 to 16, wherein affine prediction is used as a hypothesis for the current video block encoded with Multiple Hypothesis Prediction (MHP).
18. The method according to any one of claims 1 to 17, wherein whether and / or how to apply the method is based on syntax elements in the bitstream.
19. The method of claim 18, wherein the syntax element is at at least one of: sequence level, group of pictures level, picture level, slice level, or slice group level.
20. The method of claim 18 or claim 19, wherein the syntax element is included in at least one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a decoding parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a temporal group header.
21. The method according to any one of claims 18 to 20, wherein the syntax element is indicated in a region comprising more than one sample or pixel.
22. The method of claim 21, wherein the region comprises one of: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.
23. The method according to any one of claims 1 to 22, wherein whether and / or how to apply the method is determined based on codec information of the current video block.
24. The method according to claim 23, wherein the codec information includes at least one of the following: a block size of the current video block, a color format of the current video block, a single-tree partitioning or a dual-tree partitioning of the current video block, a color component of the current video block, a slice type of the current video block, or a picture type of the current video block.
25. The method of any one of claims 1 to 24, wherein whether a first syntax element is determined based on a second syntax element, the first syntax element indicating whether a template matching based refinement process is applied to the control point motion vector of the current video block.
26. The method of any one of claims 1 to 25, wherein the converting comprises encoding the current video block into the bitstream.
27. The method of any one of claims 1 to 25, wherein the converting comprises decoding the current video block from the bitstream.
28. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 27.
29. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1 to 27.
30. A non-transitory computer-readable recording medium storing a bit stream of a video, wherein the bit stream of the video is generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a set of control point motion vectors (CPMVs) associated with affine motion candidates for a current video block of the video; determining a refined affine motion candidate by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs; as well as The bitstream is generated based on the refined affine motion candidate.
31. A method for storing a bitstream of a video, comprising: determining a set of control point motion vectors (CPMVs) associated with affine motion candidates for a current video block of the video; determining a refined affine motion candidate by applying a refinement process based on template matching to at least one CPMV in the set of CPMVs; generating the bitstream based on the refined affine motion candidate; as well as The bitstream is stored in a non-transitory computer-readable recording medium.