Method and device for candidate derivation for affine merge mode in video coding

By deriving motion vector candidates from non-adjacent neighboring blocks and constructing affine models, the method addresses inefficiencies in existing video coding standards, enhancing compression efficiency and quality.

JP2026021367APending Publication Date: 2026-02-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025178002
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-16
Filing Date
2025-10-22
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in efficiently deriving motion vector candidates for affine motion prediction modes, leading to suboptimal compression efficiency and quality.

Method used

The method involves obtaining motion vector candidates from non-adjacent neighboring blocks based on scanning areas and distances, determining termination conditions, and constructing affine models using parameters from these blocks to improve motion vector prediction.

Benefits of technology

Enhances the accuracy of motion vector prediction, resulting in improved video coding efficiency and quality by optimizing the derivation of affine merge candidates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021367000001_ABST
    Figure 2026021367000001_ABST
Patent Text Reader

Abstract

To provide a method for improving affine merge candidate derivation for an affine motion prediction mode in a video encoding or decoding process.SOLUTION: A method of video decoding includes obtaining a temporal candidate list having a first list size, wherein the first list size is larger than a list size of any existing candidate list including an affine merge candidate list, an advanced motion vector prediction (AMVP) candidate list or a regular merge candidate list, and wherein the temporal candidate list includes a plurality of motion vector (MV) candidates obtained from a plurality of neighboring blocks to a current block. Furthermore, the method may include obtaining a first number of MV candidates from the temporal candidate list based on the reordered plurality of MV candidates, wherein the first number is smaller than a number of the plurality of MV candidates in the temporal candidate list.SELECTED DRAWING: Figure 29
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is filed on and claims priority to U.S. Provisional Patent Application No. 63 / 290,638, entitled "Methods and Devices for Candidate Derivation for Affine Merge Mode in Video Coding," filed December 16, 2021, which is incorporated by reference in its entirety for all purposes.

[0002] This disclosure relates to video coding and compression, and more particularly, but not exclusively, to methods and apparatus for improving affine merge candidate derivation for affine motion prediction modes in video encoding or decoding processes. [Background technology]

[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, some currently known video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), and Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMediaVideo1 (AV1) was developed by the Alliance for Open Media (AOM) as the successor to its predecessor, VP9. Audio Video Coding (AVS), which refers to digital audio and digital video compression standards, is another series of video compression standards developed by the Chinese Audio and Video Coding Standard Workgroup. Most of the existing video coding standards build on the well-known hybrid video coding framework, i.e., using block-based prediction methods (e.g., inter-prediction, intra-prediction) to reduce the redundancy present in a video image or sequence, and transform coding to reduce the energy of the prediction error. An important goal of video coding techniques is to compress video data into a format that uses a lower bit rate while avoiding or minimizing degradation of video quality.

[0004] The first-generation AVS standard includes the Chinese national standards "Information Technology, Advanced Audio Video Coding, Part 2: Video" (known as AVS1) and "Information Technology, Advanced Audio Video Coding Part 16: Radio Television Video" (known as AVS+). It can provide approximately 50% bitrate savings at the same perceptual quality compared to the MPEG-2 standard. The video part of the AVS1 standard was promulgated as a Chinese national standard in February 2006. The second-generation AVS standard includes the Chinese national standard series "Information Technology, Efficient Multimedia Coding" (known as AVS2), which is primarily targeted at the transmission of extra HD TV programs. The coding efficiency of AVS2 is twice that of AVS+. AVS2 was published as a Chinese national standard in May 2016. Meanwhile, the video part of the AVS2 standard has been submitted by the Institute of Electrical and Electronics Engineers (IEEE) as an international standard for applications. The AVS3 standard is a new generation video coding standard for UHD video applications that aims to exceed the coding efficiency of the latest international standard, HEVC. In March 2019, at the 68th AVS meeting, the AVS3-P2 baseline was finalized, which provides approximately 30% bitrate savings over the HEVC standard. Currently, there is one reference software, called the High Performance Model (HPM), maintained by the AVS group to certify reference implementations of the AVS3 standard. Summary of the Invention [Problem to be solved by the invention]

[0005] This disclosure provides examples of techniques related to improving motion vector candidate derivation for motion prediction modes in a video encoding or decoding process. [Means for solving the problem]

[0006] According to a first aspect of the present disclosure, a method for video decoding is provided. The method may include obtaining one or more motion vector (MV) candidates from a plurality of non-adjacent neighboring blocks to a current block based on at least one scanning area and at least one scanning distance, where one of the at least one scanning distances may indicate a number of blocks away from one side of the current block. Further, the method may include determining a termination condition based on the number of MV candidates obtained by scanning the at least one scanning distance within a first scanning area, where the at least one scanning area may include the first scanning area.

[0007] Additionally, the method may include, in response to determining that the completion condition is met, stopping scanning at least one scanning area, and obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more MV candidates.

[0008] According to a second aspect of the present disclosure, a method for video decoding is provided. The method may include: obtaining one or more first parameters based on one or more first neighboring blocks of a current block; obtaining one or more second parameters based on the one or more first neighboring blocks and / or one or more second neighboring blocks of the current block; constructing one or more affine models using the one or more first parameters and the one or more second parameters; and obtaining one or more CPMVs for the current block based on the one or more affine models. Furthermore, the one or more first neighboring blocks and the one or more second neighboring blocks may be obtained from multiple neighboring blocks to the current block based on at least one scanning area and at least one scanning distance. Furthermore, one of the at least one scanning distance may indicate the number of blocks away from one side of the current block, and the one or more first neighboring blocks and the one or more second neighboring blocks may be obtained by exclusively scanning at least one scanning area within the at least one scanning distance.

[0009] According to a third aspect of the present disclosure, a method of video decoding is provided. The method may include obtaining one or more motion vector prediction (MV) candidates from one or more candidate lists according to a predetermined order, where the one or more candidate lists may include an affine advanced motion vector prediction (AMVP) candidate list, a regular merge candidate list, and an affine merge candidate list, and the one or more MV candidates may be from multiple neighboring blocks to a current block. Further, the method may include obtaining one or more CPMVs for the current block based on the one or more MV candidates.

[0010] According to a fourth aspect of the present disclosure, there is provided a method for video decoding. The method may include obtaining a temporal candidate list having a first list size, where the first list size is greater than the list size of any existing candidate list, including an affine merge candidate list, an AMVP candidate list, or a regular merge candidate list, and the temporal candidate list may include multiple MV candidates obtained from multiple neighboring blocks to a current block. Further, the method may include obtaining a first number of MV candidates from the temporal candidate list based on the reordered multiple MV candidates, where the first number is less than the number of the multiple MV candidates in the temporal candidate list.

[0011] According to a fifth aspect of the present disclosure, there is provided a method for video encoding. The method may include determining one or more MV candidates from a plurality of non-adjacent neighboring blocks to a current block based on at least one scanning area and at least one scanning distance, where one of the at least one scanning distances may indicate a number of blocks away from one side of the current block. Further, the method may include determining a completion condition based on a number of MV candidates obtained by scanning at least one scanning distance within a first scanning area, where the at least one scanning area may include the first scanning area.

[0012] Additionally, the method may include, in response to determining that the completion condition is met, stopping scanning at least one scanning area and determining one or more CPMVs for the current block based on the one or more MV candidates.

[0013] According to a sixth aspect of the present disclosure, a video encoding method is provided. The method may include determining one or more first parameters based on one or more first neighboring blocks of a current block; determining one or more second parameters based on the one or more first neighboring blocks and / or one or more second neighboring blocks of the current block; constructing one or more affine models using the one or more first parameters and the one or more second parameters; and determining one or more CPMVs for the current block based on the one or more affine models. Further, the one or more first neighboring blocks and the one or more second neighboring blocks may be determined from multiple neighboring blocks to the current block based on at least one scanning area and at least one scanning distance. Moreover, one of the at least one scanning distance may indicate the number of blocks away from one side of the current block, and the one or more first neighboring blocks and the one or more second neighboring blocks may be determined by exclusively scanning at least one scanning area in at least one scanning distance.

[0014] According to a seventh aspect of the present disclosure, there is provided a method for video encoding. The method may include determining one or more MV candidates from one or more candidate lists according to a predetermined order, where the one or more candidate lists include an AMVP candidate list, a regular merge candidate list, and an affine merge candidate list, and the one or more MV candidates may be from multiple neighboring blocks to a current block. Further, the method may include determining one or more CPMVs for the current block based on the one or more MV candidates.

[0015] According to an eighth aspect of the present disclosure, there is provided a method of video encoding. The method may include determining a temporal candidate list having a first list size, where the first list size is greater than the list size of any existing candidate list, including an affine merge candidate list, an AMVP candidate list, or a regular merge candidate list, and the temporal candidate list may include multiple MV candidates obtained from multiple neighboring blocks to a current block. Further, the method may include determining a first number of MV candidates from the temporal candidate list based on the reordered multiple MV candidates, where the first number is less than the number of the multiple MV candidates in the temporal candidate list.

[0016] According to a ninth aspect of the present disclosure, there is provided an apparatus for video decoding, the apparatus including one or more processors and a memory configured to store instructions executable by the one or more processors, the one or more processors being further configured to perform a method according to the first aspect, the second aspect, the third aspect, or the fourth aspect when the instructions are executed.

[0017] According to a tenth aspect of the present disclosure, there is provided an apparatus for video encoding, the apparatus including one or more processors and a memory configured to store instructions executable by the one or more processors, the one or more processors being further configured to perform a method according to the fifth, sixth, seventh, or eighth aspects when the instructions are executed.

[0018] According to an eleventh aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium for storing computer-executable instructions, which when executed by one or more computer processors, cause the one or more computer processors to receive a bitstream and perform a method according to the first aspect, the second aspect, the third aspect, or the fourth aspect.

[0019] According to a twelfth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium for storing computer-executable instructions, which when executed by one or more computer processors, cause the one or more computer processors to perform a method according to the fifth, sixth, seventh, or eighth aspect, to encode a current block into a bitstream and transmit the bitstream.

[0020] A more particular description of embodiments of the present disclosure will be made by reference to specific embodiments illustrated in the accompanying drawings, in which the embodiments will be described and explained with additional specificity and detail, given that the drawings represent only some embodiments and therefore should not be considered limiting in scope. [Brief explanation of the drawings]

[0021] [Figure 1A] 1 is a block diagram illustrating a system for encoding and decoding video blocks according to some embodiments of the present disclosure. [Figure 1B] FIG. 2 is a block diagram of an encoder according to some embodiments of the present disclosure. [Figure 1C] 1 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 1D] 1 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 1E] 1 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 1F] 1 is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 2]FIG. 2 is a block diagram of a decoder according to some embodiments of the present disclosure. [Figure 3A] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 3B] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 3C] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 3D] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 3E] FIG. 10 illustrates an example of block partitions within a multi-type tree structure, according to some embodiments of the present disclosure. [Figure 4A] FIG. 1 illustrates an example four-parameter affine model, according to some embodiments of the present disclosure. [Figure 4B] FIG. 1 illustrates an example four-parameter affine model, according to some embodiments of the present disclosure. [Figure 4C] FIG. 10 illustrates exemplary positions of spatial merge candidates, according to some embodiments of the present disclosure. [Figure 4D] 10A-10C are diagrams illustrating candidate pairs considered for redundancy check of spatial merge candidates, according to some embodiments of the present disclosure. [Figure 4E] FIG. 10 illustrates motion vector scaling for temporal merge candidates, according to some embodiments of the present disclosure. [Figure 4F] FIG. 10 illustrates candidate positions for temporal merge candidates C0 and C1, according to some embodiments of the present disclosure. [Figure 5] FIG. 1 illustrates a six-parameter affine model, according to some embodiments of the present disclosure. [Figure 6] FIG. 10 illustrates an example of close neighboring blocks for inherited affine merge candidates, in accordance with some embodiments of the present disclosure. [Figure 7] FIG. 10 illustrates an example of close neighboring blocks for constructed affine merge candidates, in accordance with some embodiments of the present disclosure. [Figure 8] 10A-10C are diagrams illustrating non-adjacent neighboring blocks for inherited affine merge candidates, in accordance with some embodiments of the present disclosure. [Figure 9] FIG. 10 illustrates the derivation of constructed affine merge candidates using neighboring blocks, according to some embodiments of the present disclosure. [Figure 10] FIG. 10 illustrates vertical scanning of non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 11] FIG. 10 illustrates horizontal scanning of non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 12] FIG. 10 illustrates combined vertical and horizontal scanning of non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 13A] FIG. 10 illustrates a neighboring block having the same size as the current block, according to some embodiments of the present disclosure. [Figure 13B] 10A and 10B are diagrams illustrating neighboring blocks having a different size than the current block, according to some embodiments of the present disclosure. [Figure 14A] A figure illustrating an example in which the bottom-left or top-right block of the bottom-most or right-most block in the previous distance is used as the bottom-most or right-most block of the current distance, according to some embodiments of the present disclosure. [Figure 14B] 10A-10C are diagrams illustrating an example in which the left or top block of the bottom or rightmost block in the previous distance is used as the bottom or rightmost block of the current distance, according to some embodiments of the present disclosure. [Figure 15A] FIG. 10 illustrates scanning positions at bottom-left and top-right positions used for non-adjacent neighboring blocks above and to the left, according to some embodiments of the present disclosure. [Figure 15B]FIG. 10 illustrates a scanning position at the bottom right position used for both top and left non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 15C] FIG. 10 illustrates a scanning position at the bottom left position used for both top and left non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 15D] FIG. 10 illustrates a scanning position at the right-top position used for both top and left non-adjacent neighboring blocks, according to some embodiments of the present disclosure. [Figure 16] FIG. 10 illustrates a simplified scanning process for deriving pre-constructed merge candidates, according to some embodiments of the present disclosure. [Figure 17A] FIG. 10 illustrates an example of spatial neighborhoods for deriving inherited affine merge candidates, according to some embodiments of the present disclosure. [Figure 17B] FIG. 10 illustrates an example of spatial neighborhoods from which constructed affine merge candidates are derived, in accordance with some embodiments of the present disclosure. [Figure 18] 1 illustrates an example of an inheritance-based derivation method for deriving affine constructed candidates, according to some embodiments of the present disclosure. [Figure 19] FIG. 2 illustrates an example of templates and reference samples of templates in Reference List 0 and Reference List 1, according to some embodiments of the present disclosure. [Figure 20] 10A-10C are diagrams illustrating templates and reference samples of the templates for blocks with sub-block motion using sub-block motion information of a current block, according to some embodiments of the present disclosure. [Figure 21] FIG. 1 illustrates an exemplary computing environment coupled with a user interface, according to some embodiments of the present disclosure. [Figure 22] 1 is a flowchart illustrating a method for video decoding, according to some embodiments of the present disclosure. [Figure 23]23 is a flowchart illustrating a method for video encoding corresponding to the method for video decoding as shown in FIG. 22, according to some embodiments of the present disclosure. [Figure 24] 1 is a flowchart illustrating a method for video decoding, according to some embodiments of the present disclosure. [Figure 25] 25 is a flowchart illustrating a method for video encoding corresponding to the method for video decoding as shown in FIG. 24, according to some embodiments of the present disclosure. [Figure 26] 1 is a flowchart illustrating a method for video decoding, according to some embodiments of the present disclosure. [Figure 27] 27 is a flowchart illustrating a method for video encoding corresponding to the method for video decoding as shown in FIG. 26, according to some embodiments of the present disclosure. [Figure 28] 1 is a flowchart illustrating a method for video decoding, according to some embodiments of the present disclosure. [Figure 29] 29 is a flowchart illustrating a method for video encoding corresponding to the method for video decoding as shown in FIG. 28, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0022] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to aid in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0023] The terms used in the disclosure are employed only for the purpose of describing particular embodiments and are not intended to limit the disclosure. The singular forms "a / an," "said," and "the" in the disclosure and the appended claims are intended to include the plural forms as well, unless otherwise clearly indicated throughout the disclosure. Also, the term "and / or" used in the disclosure will be understood to refer to and include one or any or all possible combinations of the associated listed items.

[0024] References throughout this specification to "one embodiment," "an embodiment," "an example," "some embodiments," "some examples," or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. A feature, structure, element, or characteristic described in connection with one or some embodiments may also be applicable to other embodiments, unless expressly specified otherwise.

[0025] Throughout the disclosure, the terms "first," "second," "third," etc. all do not imply any spatial or chronological order and, unless clearly specified otherwise, are used solely as nomenclature to refer to related elements, e.g., devices, components, compositions, steps, etc. For example, a "first device" and a "second device" may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be arbitrarily named.

[0026] The terms "module," "sub-module," "circuit," "sub-circuit," "circuitry," "sub-circuitry," "unit," or "sub-unit" may include memory (shared, dedicated, or a group) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected. The components may or may not be physically attached or located near each other.

[0027] As used herein, the terms "if" or "when" may be understood to mean "upon" or "in response to," depending on the context. When these terms appear in the claims, they may not indicate that the associated limitation or feature is conditional or optional. For example, a method may include steps where: i) when or if condition X exists, a function or action X' is performed; and ii) when or if condition Y exists, a function or action Y' is performed. A method may be implemented with both the ability to perform function or action X' and the ability to perform function or action Y'. Thus, functions X' and Y' may both be performed at different times for multiple executions of the method.

[0028] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a purely software implementation, for example, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked together to perform a particular function.

[0029] 1A is a block diagram illustrating an example system 10 for encoding and decoding video blocks in parallel, according to some implementations of the present disclosure. As shown in FIG. 1A, system 10 includes a source device 12 that generates and encodes video data to be subsequently decoded by a destination device 14. Source device 12 and destination device 14 may include any of a wide variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video gaming consoles, or video streaming devices. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0030] In some implementations, destination device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to destination device 14. In one embodiment, link 16 may include a communication medium to enable source device 12 to transmit encoded video data directly to destination device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 14. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from source device 12 to destination device 14.

[0031] In some other implementations, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by destination device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessible data storage media, such as a hard drive, Blu-ray disc, digital versatile disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In further embodiments, storage device 32 may correspond to a file server or another intermediate storage device that may hold the encoded video data generated by source device 12. Destination device 14 may access the stored video data from storage device 32 via streaming or downloading. The file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to destination device 14. Exemplary file servers include web servers (e.g., for websites), file transfer protocol (FTP) servers, network-attached storage (NAS) devices, or local disk drives. Destination device 14 may access the coded video data through any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing coded video data stored on a file server. Transmission of the coded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0032] As shown in FIG. 1A , source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include sources such as a video capture device, e.g., a video camera, a video archive containing previously captured video, a video feeding interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of such sources. As one example, if video source 18 is a video camera in a security surveillance system, source device 12 and destination device 14 may form a camera phone or video phone. However, implementations described herein may be generally applicable to video coding and may be applied to wireless and / or wired applications.

[0033] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to destination device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access, decoding, and / or playback by destination device 14 or another device. Output interface 22 may further include a modem and / or a transmitter.

[0034] Destination device 14 includes an input interface 28, a video decoder 30, and a display device. Input interface 28 may include a receiver and / or a modem and may receive encoded video data over link 16. The encoded video data communicated over link 16 and provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included in encoded video data transmitted over a communications medium, stored on a storage medium, or stored on a file server.

[0035] In some implementations, destination device 14 may include a display device 34, which can be an integrated display device or an external display device configured to communicate with destination device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0036] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as VVC, HEVC, MPEG-4 Part 10, AVC, or extensions of such standards. It should be understood that the present application is not limited to any particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that video encoder 20 of source device 12 may be configured to encode video data in accordance with any of these current or later standards. Similarly, it is generally contemplated that video decoder 30 of destination device 14 may be configured to decode video data in accordance with any of these current or later standards.

[0037] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, an electronic device may store instructions for the software on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) in the respective device.

[0038] Like HEVC, VVC is built on a block-based hybrid video coding framework. FIG. 1B is a block diagram illustrating a block-based video encoder according to some implementations of the present disclosure. In encoder 100, an input video signal is processed block by block, called a coding unit (CU). Encoder 100 may be the video encoder 20 shown in FIG. 1A. In VTM-1.0, a CU can be up to 128 x 128 pixels. However, unlike HEVC, which partitions blocks only based on a quadtree, in VVC, a single coding tree unit (CTU) is divided into CUs based on a quadtree / binary tree / ternary tree to adapt to variable local characteristics. In addition, the concept of multiple partitioning unit types in HEVC has been removed; i.e., the separation of CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as the basic unit for both prediction and transform without further partitioning. In a multi-type tree structure, a CTU is first partitioned by a quadtree structure, and then each quadtree leaf node can be further partitioned by a binary tree and a ternary tree structure.

[0039] 3A-3E are schematic diagrams illustrating multi-type tree partitioning modes according to some implementations of the present disclosure, each showing five partitioning types, including quadrant partitioning (FIG. 3A), vertical bisection partitioning (FIG. 3B), horizontal bisection partitioning (FIG. 3C), vertically extended trisection partitioning (FIG. 3D), and horizontally extended trisection partitioning (FIG. 3E).

[0040] For each given video block, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra-prediction") uses pixels from samples of previously coded neighboring blocks (called reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces spatial redundancy inherent in the video signal. Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") uses reconstructed pixels from previously coded video pictures to predict the current video block. Temporal prediction reduces temporal redundancy inherent in the video signal. Temporal prediction for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal references. If multiple reference pictures are supported, a reference picture index is also transmitted, which is used to identify which reference picture in the reference picture store the temporal prediction comes from.

[0041] After spatial and / or temporal prediction, an intra / inter mode decision circuit 121 in encoder 100 chooses the best prediction mode based on, for example, a rate-distortion optimization method. Block predictor 120 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using transform circuit 102 and quantization circuit 104. The resulting quantized residual coefficients are inverse quantized by inverse quantization circuit 116 and inverse transformed by inverse transform circuit 118 to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Additionally, in-loop filtering 115, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), may be applied to the reconstructed CU before it is placed in a reference picture store in picture buffer 117 and used to code subsequent video blocks. To form the output video bitstream 114, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 106 to be further compressed and packed to form the bitstream.

[0042] For example, deblocking filters are available in the current version of VVC, along with AVC and HEVC. In HEVC, an additional in-loop filter called SAO is defined to further improve coding efficiency. In the current version of the VVC standard, an additional in-loop filter called ALF is being actively investigated and has a good chance of being included in the final standard.

[0043] These in-loop filter operations are optional. Performing these operations helps improve coding efficiency and visual quality. They can also be turned off as a decision made by the encoder 100 to save computational complexity.

[0044] It should be noted that intra prediction is typically based on unfiltered reconstructed pixels, and inter prediction is based on filtered reconstructed pixels if those filter options are turned on by the encoder 100.

[0045] FIG. 2 is a block diagram illustrating a block-based video decoder 200 that can be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related section present in encoder 100 of FIG. 1B. The block-based video decoder 200 may be the video decoder 30 shown in FIG. 1A. In the decoder 200, an incoming video bitstream 201 is first decoded through entropy decoding 202 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed through inverse quantization 204 and inverse transform 206 to obtain reconstructed prediction residuals. A block predictor mechanism, implemented in an intra / inter mode selector 212, is configured to perform either intra prediction 208 or motion compensation 210 based on the decoded prediction information. The set of unfiltered reconstructed pixels is obtained by summing, using a summer 214, the reconstructed prediction residual from the inverse transform 206 and the prediction output generated by the block predictor mechanism.

[0046] The reconstructed blocks may further pass through an in-loop filter 209 before being stored in a picture buffer 213, which serves as a reference picture store. The reconstructed video in the picture buffer 213 may be transmitted to drive a display device and may also be used to predict later video blocks. In situations where the in-loop filter 209 is turned on, a filtering operation is performed on the reconstructed pixels to derive the final reconstructed video output 222.

[0047] In the current VVC and AVS3 standards, motion information for the current coding block is either replicated from spatial or temporal neighboring blocks specified by merge candidate indexes or obtained through explicit signaling of motion estimation. The focus of this disclosure is to improve the accuracy of motion vectors for affine merge modes by improving the derivation method of affine merge candidates. To facilitate the explanation of this disclosure, the existing affine merge mode design in the VVC standard is used as an example to promote the proposed ideas. While the existing affine mode design in the VVC standard is used as an example throughout this disclosure, those skilled in the art of modern video coding technology will note that the proposed techniques may also be applied to different designs of affine motion prediction modes or other coding tools with the same or similar design spirit.

[0048] In a typical video coding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other instances, a frame may be monochrome and therefore include only one two-dimensional array of luma samples.

[0049] As shown in FIG. 1C, video encoder 20 (or, more specifically, a partitioning unit in a prediction processing unit of video encoder 20) generates a coded representation of a frame by first partitioning the frame into a set of CTUs. A video frame may include an integer number of CTUs ordered consecutively in a raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set. As a result, all CTUs in a video sequence have the same size, which may be one of 128x128, 64x64, 32x32, and 16x16. However, it should be noted that the present application is not necessarily limited to a particular size. As shown in FIG. 1D, each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements used to code the coding tree block samples. The syntax elements describe how a video sequence may be reconstructed in video decoder 30, including the nature of different types of units of coded blocks of pixels and inter or intra prediction, intra prediction mode, motion vectors, and other parameters. For monochrome pictures or pictures with three distinct color planes, a CTU may contain syntax elements used to code a single coding tree block and samples of the coding tree block. A coding tree block may be an N by N block of samples.

[0050] To achieve better performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree blocks of a CTU to divide the CTU into smaller CUs. As shown in FIG. 1E, 64×64 CTU 400 is first partitioned into four smaller CUs, each with a 32×32 block size. Among the four smaller CUs, CU 410 and CU 420 are each partitioned into four 16×16 CUs by block size. Two 16×16 CUs, 430 and 440, are each further partitioned into four 8×8 CUs by block size. FIG. 1F shows a quad tree data structure illustrating the final result of the partitioning process for CTU 400 as shown in FIG. 1E, where each leaf node of the quad tree corresponds to one CU of a respective size ranging from 32×32 to 8×8. Similar to the CTU shown in Figure 1D, each CU may contain a CB of luma samples and two corresponding coding blocks of chroma samples for identically sized frames, as well as syntax elements used to code the coding block samples. In monochrome pictures or pictures with three distinct color planes, a CU may contain a single coding block and syntax elements used to code the coding block samples. It should be noted that the quadtree partitioning shown in Figures 1E-1F is for illustrative purposes only; a CTU may be divided into CUs based on quadtree, ternary tree, or binary tree partitioning to suit variable local characteristics. In a multi-type tree structure, a CTU is partitioned using a quadtree structure, and each quadtree leaf CU may be further partitioned into binary and ternary tree structures. As shown in Figures 3A-3E, there are five possible partitioning types of coding blocks with width W and height H: quadtree partitioning, horizontal binary tree partitioning, vertical binary tree partitioning, horizontal ternary tree partitioning, and vertical ternary tree partitioning.

[0051] In some implementations, video encoder 20 may further partition a coding block of a CU into one or more M×N PBs. A PB is a rectangular (square or non-square) block of samples to which the same prediction, inter or intra, is applied. A PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements used to predict the PBs. In monochrome pictures or pictures with three separate color planes, a PU may include a single PB and syntax elements used to predict the PB. Video encoder 20 may generate predictive luma, Cb and Cr blocks for the luma, and Cb and Cr PBs for each PU of the CU.

[0052] Video encoder 20 may use intra prediction or inter prediction to generate predictive blocks for a PU. If video encoder 20 uses intra prediction to generate predictive blocks for a PU, video encoder 20 may generate predictive blocks for the PU based on decoded samples of a frame associated with the PU. If video encoder 20 uses inter prediction to generate predictive blocks for a PU, video encoder 20 may generate predictive blocks for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0053] After video encoder 20 generates predictive luma, Cb, and Cr blocks for one or more PUs of a CU, video encoder 20 may generate a luma residual block for the CU by subtracting the predictive luma block of the CU from its original luma coding block, such that each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predictive luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predictive Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates a difference between a Cr sample in one of the predictive Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0054] Further, as illustrated in FIG. 1E , video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks, respectively. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some embodiments, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In monochrome pictures or pictures with three separate color planes, a TU may include a single transform block and syntax structures used to transform the transform block samples.

[0055] Video encoder 20 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block for the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to a Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to a Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0056] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to a process in which transform coefficients are quantized to potentially reduce the amount of data used to represent the transform coefficients and provide further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 may entropy encode syntax elements that indicate the quantized transform coefficients. For example, video encoder 20 may perform CABAC on the syntax elements that indicate the quantized transform coefficients. Finally, video encoder 20 may output a bitstream that includes a sequence of bits that form a representation of the coded frame and associated data, which is either stored in storage device 32 or transmitted to destination device 14.

[0057] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally reciprocal to the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on coefficient blocks associated with TUs of the current CU to reconstruct residual blocks associated with TUs of the current CU. Video decoder 30 also reconstructs coding blocks of the current CU by adding samples of predictive blocks for PUs of the current CU to corresponding samples of transform blocks of TUs of the current CU. After reconstructing the coding blocks for each CU of the frame, video decoder 30 may reconstruct the frame.

[0058] As mentioned above, video coding achieves video compression using two main modes: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). Note that IBC can be considered either intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes to coding efficiency more than intra-frame prediction due to the use of motion vectors to predict a current video block from a reference video block.

[0059] However, ever-improving video data capture techniques and more refined video block sizes for preserving content within the video data also substantially increase the amount of data required to represent motion vectors for the current frame. One way to overcome this challenge is to take advantage of the fact that a group of neighboring CUs in both the spatial and temporal domains not only have similar video data for prediction purposes, but also have similar motion vectors between those neighboring CUs. Therefore, by utilizing their spatial and temporal correlations, also referred to as the current CU's "motion vector predictor (MVP)," it is possible to use the motion information of spatially neighboring CUs and / or temporally co-located CUs as an approximation of the current CU's motion information (e.g., motion vector).

[0060] Instead of encoding into the video bitstream the actual motion vector of the current CU determined by the motion estimation unit as described above in connection with FIG. 1B, the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to yield a motion vector difference (MVD) for the current CU. By doing so, the motion vector determined by the motion estimation unit for each CU of the frame does not need to be encoded into the video bitstream, and the amount of data used to represent motion information in the video bitstream can be significantly reduced.

[0061] Similar to the process of choosing a predictive block in a reference frame during inter-frame prediction of a code block, a set of rules is required that are adopted by both video encoder 20 and video decoder 30 to construct a motion vector candidate list (also known as a "merge list") for the current CU using those potential candidate motion vectors associated with spatially neighboring CUs and / or temporally co-located CUs of the current CU, and then select one member from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, the motion vector candidate list itself does not need to be transmitted from video encoder 20 to video decoder 30; the index of the selected motion vector predictor in the motion vector candidate list is sufficient for video encoder 20 and video decoder 30 to use the same motion vector predictor in the motion vector candidate list to encode and decode the current CU.

[0062] Affine Model In HEVC, only the translational motion model is applied for motion-compensated prediction. In the real world, many types of motion exist, such as zoom in / out, rotation, depth motion, and other irregular motion. In VVC and AVS3, affine motion-compensated prediction is applied by signaling one flag for each inter-coding block to indicate whether the translational motion model or the affine motion model is applied for inter prediction. In the current VVC and AVS3 designs, two affine modes are supported for one affine coding block, including a 4-parameter affine mode and a 6-parameter affine mode.

[0063] The four-parameter affine model has the following parameters: two parameters for translation in the horizontal and vertical directions, one parameter for zoom motion, and one parameter for rotation motion in both directions. In this model, the horizontal zoom parameter is equal to the vertical zoom parameter, and the horizontal rotation parameter is equal to the vertical rotation parameter. To achieve a better fit of the motion vectors and affine parameters, the affine parameters are derived from two MVs (also called control point motion vectors (CPMVs)) located at the top left and top right corners of the current block. As shown in Figures 4A-4B, the affine motion field of a block is described by two CPMVs (V0, V1). Based on the control point motion, the motion field (v x ,v y )teeth,

number

[0064] The 6-parameter affine model has the following parameters: two parameters for translation in each of the horizontal and vertical directions, two parameters for zoom and rotation in the horizontal direction, and two parameters for zoom and rotation in the vertical direction. The 6-parameter affine motion model is coded by three CPMVs. As shown in Figure 5, the three control points of a 6-parameter affine block are located at the top-left, top-right, and bottom-left corners of the block. The motion at the top-left control point is associated with translation, the motion at the top-right control point is associated with rotation and zoom in the horizontal direction, and the motion at the bottom-left control point is associated with rotation and zoom in the vertical direction. Compared to the 4-parameter affine motion model, the rotation and zoom motions in the horizontal direction of the 6-parameter affine model cannot be identical to those in the vertical direction. Assuming that (V0, V1, V2) are the MVs at the top left, top right, and bottom left corners of the current block in Figure 5, the motion vectors (v x ,v y )teeth,

number

[0065] Affine Merge Mode In affine merge mode, the CPMV for the current block is not explicitly signaled but is derived from neighboring blocks. In particular, in this mode, motion information of spatially neighboring blocks is used to generate the CPMV for the current block. The affine merge mode candidate list has a limited size. For example, in the current VVC design, there can be a maximum of five candidates. The encoder can evaluate and select the best candidate index based on a rate-distortion optimization algorithm. The selected candidate index is then signaled to the decoder side. Affine merge candidates can be determined in three ways. In the first method, affine merge candidates can be inherited from neighboring affine-coded blocks. In the second method, affine merge candidates can be constructed from translational MVs from neighboring blocks. In the third method, zero MVs are used as affine merge candidates.

[0066] For the inherited method, there can be at most two candidates, which are taken from the neighboring block located to the bottom left of the current block (e.g., the scanning order is from A0 to A1 as shown in Figure 6) and the neighboring block located to the top right of the current block (e.g., the scanning order is from B0 to B2 as shown in Figure 6), if available.

[0067] For the pre-constructed method, the candidates are combinations of nearby translational MVs, which can be generated by two steps.

[0068] Step 1: Obtain four translational MVs, including MV1, MV2, MV3, and MV4, from the available neighborhood. MV1: MV from one of the three neighboring blocks closest to the top left corner of the current block. As shown in Figure 7, the scanning order is B2, B3, and A2. MV2: MV from one of the two neighboring blocks closest to the top right corner of the current block. As shown in Figure 7, the scanning order is B1 and B0. MV3: MV from one of the two neighboring blocks near the bottom left corner of the current block. As shown in Figure 7, the scanning order is A1 and A0. MV4: MV from the block that is co-located in time with the neighboring block near the bottom right corner of the current block. As shown in the figure, the neighboring block is T.

[0069] Step 2: Derive combinations based on the four translational MVs from Step 1. Combination 1: MV1, MV2, MV3, Combination 2: MV1, MV2, MV4, Combination 3: MV1, MV3, MV4, Combination 4: MV2, MV3, MV4, Combination 5: MV1, MV2, Combination 6: MV1, MV3.

[0070] When the merge candidate list is not full after filling with inherited and constructed candidates, a zero MV is inserted at the end of the list.

[0071] Affine AMVP mode Affine Advanced Motion Vector Prediction (AMVP) mode can be applied to CUs with both width and height greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predictor CPMVP is signaled in the bitstream. The affine AMVP candidate list size is 2, and the affine AMVP candidate list is signaled by using the following four types of CPMV candidates in the following order: - Inherited affine AMVP candidates extrapolated from the CPMVs of neighboring CUs, - A pre-constructed affine AMVP candidate CPMVP derived using the translational MVs of nearby CUs, - Translational MVs from nearby CUs, - Temporal MV from co-located CUs, - Zero MV.

[0072] The checking order for inherited affine AMVP candidates is the same as that for inherited affine merge candidates. The only difference is that for AMVP candidates, only affine CUs with the same reference pictures as in the current block are considered. No pruning process is applied when inserting inherited affine motion predictors into the candidate list.

[0073] The pre-constructed AMVP candidate is derived from the same spatial neighborhood as in affine merge mode. The same checking order is used as in affine merge candidate construction. In addition, the reference picture indexes of neighboring blocks are also checked. The first block in the checking order that is inter-coded and has the same reference picture as the current CU is used. If the current CU is coded in 4-parameter affine mode and both mv0 and mv1 are available, mv0 and mv1 are added as one candidate to the affine AMVP candidate list. If the current CU is coded in 6-parameter affine mode and all three CPMVs are available, they are added as one candidate to the affine AMVP candidate list. Otherwise, the pre-constructed AMVP candidate is set as unavailable.

[0074] After the valid inherited and constructed affine AMVP candidates have been inserted, if there are still less than two affine AMVP list candidates, mv0, mv1, and mv2 are added in order to predict all control point MVs of the current CU as translational MVs, when available. Finally, if it is still not full, zero MVs are used to fill the affine AMVP list.

[0075] Regular Inter-Merge Mode In some embodiments, a canonical inter-merge candidate list is constructed by including the following five types of candidates in order: (1) Spatial MVP from spatially neighboring CUs, (2) Temporal MVP from the co-located CU, (3) History-based MVP from a first-in-first-out (FIFO) table, (4) Pairwise average MVP, and (5) Zero MV.

[0076] The size of the merge list is signaled in the Sequence Parameter Set header, and the maximum allowed size of the merge list is 6. For each CU code in merge mode, the index of the best merge candidate is coded using Truncated Unary Binary (TU). The first bin of the merge index is coded by context, and bypass coding is used for the other bins.

[0077] The derivation process for each category of merge candidates is provided above. In some embodiments, parallel derivation of merging candidate lists may be supported for all CUs within a certain size of an area.

[0078] Spatial candidate derivation The derivation of spatial merge candidates in VVC is identical to that in HEVC, except that the positions of the first two merge candidates are swapped. A maximum of four merge candidates are selected from the candidates located at the positions shown in Figure 4C. The derivation order is B0, A0, B1, A1, and B2. When one or more CUs in positions B0, A0, B1, and A1 are unavailable (e.g., because they belong to another slice or tile) or are intra-coded, only position B2 is considered. After the candidate in position A1 is added, the remaining candidates undergo a redundancy check, which ensures that candidates with identical motion information are removed from the list to improve coding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only pairs linked with arrows in Figure 4D are considered, and only candidates that do not have identical motion information are added to the list. FIG. 4D illustrates candidate pairs considered for spatial merge candidate redundancy check.

[0079] Temporal candidate derivation In this step, only one candidate is added to the list. In particular, in this temporal merge candidate derivation, a scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list and reference index to be used for the derivation of the co-located CU are explicitly signaled in the slice header. The scaled motion vector for the temporal merge candidate is obtained as illustrated by the dotted line in FIG. 4E and is scaled from the motion vector of the co-located CU using POC distances tb and td, where tb is defined as the POC difference between the current picture's reference picture and the current picture, and td is defined as the POC difference between the co-located picture's reference picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to zero.

[0080] The position for the temporal candidate is selected between candidates C0 and C1, as shown in Figure 4F. If the CU at position C0 is unavailable, intra-coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate.

[0081] History-based merge candidate derivation History-based MVP (HMVP) merge candidates are added to the merge list after spatial MVP and temporal motion vector prediction (TMVP). In this method, the motion information of a previously coded block is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (emptied) when a new CTU row is encountered. Whenever there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0082] The HMVP table size S can be set to 6, indicating that up to five history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is utilized, and a redundancy check is first applied to find whether an identical HMVP exists in the table. If found, the identical HMVP is removed from the table, all subsequent HMVP candidates are moved forward, and the identical HMVP is inserted into the last entry of the table.

[0083] HMVP candidates are used in the merge candidate list construction process. The most recent HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidate. Redundancy checks are applied to HMVP candidates for spatial or temporal merge candidates.

[0084] To reduce the number of operations for redundancy check, the following simplifications are introduced: First, the last two entries in the table are checked for redundancy with respect to each of the A1 and B1 spatial candidates. Second, the process of building a merge candidate list from HMVP is completed when the total number of available merge candidates reaches a maximum of the allowed merge candidates minus one.

[0085] Pairwise average merging candidate derivation A pairwise average candidate is generated by averaging a predefined pair of candidates in an existing merge candidate list using the first two merge candidates. The first merge candidate may be defined as p0Cand, and the second merge candidate may be defined as p1Cand. The averaged motion vector is calculated separately for each reference list according to the availability of motion vectors in p0Cand and p1Cand. If both motion vectors in a list are available, they are averaged even if they point to different reference pictures and the reference picture is set to one of p0Cand. If only one motion vector is available, it is used directly. If no motion vector is available, the list is kept invalid. Also, if the half-pel interpolation filter index of p0Cand and p1Cand is different, it is set to 0.

[0086] When the merge list is not full after the pairwise average merge candidates are added, zero MVPs are inserted at the end of the merge list until the maximum number of merge candidates is reached.

[0087] Adaptive Reordering of Merge Candidates by Template Matching (ARMC) The reordering method, named as ARMC, is applied to the regular merge mode, the template matching (TM) merge mode, and the affine merge mode (excluding SbTMVP candidates), where SbTMVP stands for sub-block-based temporal motion vector prediction candidates. For the TM merge mode, the merge candidates are reordered before the refinement process.

[0088] After the merge candidate list is constructed, the merge candidates are divided into several subgroups. The subgroup size is set to 5. The merge candidates in each subgroup are reordered in ascending order by cost value based on template matching. For simplicity, the merge candidates in the first subgroup but not the last subgroup are not reordered.

[0089] The template matching cost is measured by the sum of absolute differences (SAD) between the template samples of the current block and their corresponding reference samples. The template contains a set of reconstructed samples in the neighborhood of the current block. The template's reference samples are identified by the same motion information of the current block.

[0090] When a merge candidate utilizes bi-prediction, the reference samples of the merge candidate's template are also generated by bi-prediction as shown in FIG.

[0091] For a subblock-based merging candidate with a subblock size equal to W*H, the top template contains several subtemplates with a size of W×1, and the left template contains several subtemplates with a size of 1×H. W is the width of the subblock, and H is the height of the subblock. As shown in Figure 20, the motion information of the subblocks in the first row and first column of the current block is used to derive the reference sample for each subtemplate.

[0092] For the current video standards VVC and AVS, only the immediate neighboring blocks are used to derive affine merge candidates for the current block, as shown in Figures 6 and 7 for inherited and constructed candidates, respectively. To increase the diversity of merge candidates and further exploit spatial correlations, it is straightforward to extend the coverage of the neighboring blocks from the immediate area to the non-immediate area.

[0093] In the current video standards VVC and AVS, each affine inherited candidate is derived from one neighboring block with affine motion information, while each affine constructed candidate is derived from two or three neighboring blocks with translational motion information. To further exploit spatial correlation, new candidate derivation methods that combine affine and translational motion can be investigated.

[0094] The candidate derivation method proposed for the affine merge mode can be extended to other coding modes, such as the affine AMVP mode and the regular merge mode.

[0095] In this disclosure, the candidate derivation process for the affine merge mode can be extended by using not only nearby neighboring blocks but also non-neighboring neighboring blocks. The detailed methods can be summarized in the following aspects, including affine merge candidate pruning, non-neighboring neighborhood-based derivation process for affine inherited merge candidates, non-neighboring neighborhood-based derivation process for affine constructed merge candidates, inheritance-based derivation method for affine constructed merge candidates, HMVP-based derivation method for affine constructed merge candidates, and candidate derivation methods for affine AMVP mode and regular merge mode.

[0096] Affine Merge Candidate Pruning As the affine merge candidate list in typical video coding standards usually has a limited size, candidate pruning is a necessary process to remove redundant candidates. This pruning process is necessary for both inherited and constructed affine merge candidates. As explained in the introduction, the CPMV of the current block is not directly used for affine motion compensation. Instead, the CPMV needs to be transformed into translational MVs at the position of each sub-block within the current block. The transformation process is performed by adhering to a general affine model, as shown below.

number

[0097] For the 6-parameter affine model, three CPMVs are available, designated V0, V1, and V2. The six model parameters a, b, c, d, e, and f are then

number

[0098] For a four-parameter affine model, when the top-left corner CPMV and top-right corner CPMV, referred to as V0 and V1, are available, the six parameters a, b, c, d, e, and f are

number

[0099] For a four-parameter affine model, when the top-left corner CPMV and bottom-left corner CPMV, referred to as V0 and V2, are available, the six parameters a, b, c, d, e, and f are

number

[0100] In the above equations (4), (5), and (6), w and h represent the width and height of the current block, respectively.

[0101] When two merge candidate sets of CPMV are compared for redundancy check, it is proposed to check the similarity of six affine model parameters. Therefore, the candidate pruning process can be performed in two steps.

[0102] In step 1, given two candidate sets of CPMV, the corresponding affine model parameters for each candidate set are derived. More specifically, the two candidate sets of CPMV can be represented by two sets of affine model parameters, e.g., (a1, b1, c1, d1, e1, f1) and (a2, b2, c2, d2, e2, f2).

[0103] In step 2, a similarity check is performed between the two sets of affine model parameters based on one or more predefined thresholds. In one embodiment, when the absolute values ​​of (a1-a2), (b1-b2), (c1-c2), (d1-d2), (e1-e2), and (f1-f2) are all below a positive threshold, such as a value of 1, the two candidates are considered similar and one of them may be pruned / removed and not placed in the merge candidate list.

[0104] In some embodiments, the divide or right shift operation in step 1 may be removed to simplify the calculations in the CPMV pruning process.

[0105] In particular, the model parameters of c, d, e, and f can be calculated without dividing by the width w and height h of the current block. For example, taking the above equation (4) as an example, the approximate model parameters of c', d', e', and f' can be calculated as the following equation (7):

number

[0106] In the case where only two CPMVs are available, some of the model parameters are derived from other parts of the model parameters that depend on the width or height of the current block. In this case, the model parameters can be transformed to take into account the influence of the width and height. For example, in the case of Equation (5), the approximate model parameters of c', d', e', and f' can be calculated based on the following Equation (8): In the case of Equation (6), the approximate model parameters of c', d', e', and f' can be calculated based on the following Equation (9):

number

[0107] In step 2 above, a threshold is required to evaluate the similarity between two candidate sets of CPMVs. There may be multiple ways to define the threshold. In one embodiment, the threshold may be defined for each comparable parameter. Table 1 is an example of this embodiment, showing thresholds defined for each comparable model parameter. In another embodiment, the threshold may be defined by considering the size of the current coding block. Table 2 is an example of this embodiment, showing thresholds defined by the size of the current coding block.

[0108] [Table 1]

[0109] [Table 2]

[0110] In another embodiment, the thresholds can be defined by considering the weight or height of the current block. Tables 3 and 4 are examples of this embodiment. Table 3 shows thresholds defined by the width of the current coding block, and Table 4 shows thresholds defined by the height of the current coding block.

[0111] [Table 3]

[0112] [Table 4]

[0113] In another embodiment, the threshold may be defined as a group of fixed values. In another embodiment, the threshold may be defined by any combination of the above embodiments. In one example, the threshold may be defined by considering different parameters as well as the weight and height of the current block. Table 5 is an example of this embodiment, showing thresholds defined by the height of the current coding block. In any of the above proposed embodiments, the comparable parameter may represent any parameter defined in any of Equations (4) to (9), if necessary.

[0114] [Table 5]

[0115] The advantage of using transformed affine model parameters for candidate redundancy check is that it results in a unified similarity check process for candidates with different affine model types, for example, one merge candidate may use a 6-parameter affine model with three CPMVs and another candidate may use a 4-parameter affine model with two CPMVs, which takes into account the different influence of each CPMV in the merge candidate when deriving the target MV in each sub-block, and it provides the similarity significance of the two affine merge candidates relative to the width and height of the current block.

[0116] Non-proximal Neighborhood-Based Resolution Process for Affine Inherited Merge Candidates For inherited merge candidates, the non-proximal neighborhood-based derivation process can be performed in three steps: Step 1 is for candidate scanning, Step 2 is for CPMV projection, and Step 3 is for candidate pruning.

[0117] In step 1, non-adjacent neighboring blocks are scanned and selected by the following method.

[0118] Scanning Area and Distance In some embodiments, non-adjacent neighboring blocks may be scanned from the area to the left and the area above the current coding block, and the scanning distance may be defined as the number of coding blocks from the scanning position to the left or top side of the current coding block.

[0119] As shown in Figure 8, multiple lines of non-adjacent neighboring blocks can be scanned either to the left or above the current coding block. The distances shown in Figure 8 represent the number of coding blocks from each candidate position to the left or top side of the current block. For example, an area with "distance 2 (D2)" on the left side of the current block indicates that candidate neighboring blocks located within this area are two blocks away from the current block. Similar indications can be applied to other scanning areas with different distances.

[0120] In one or more embodiments, the non-neighboring neighboring blocks at each distance may have the same block size as the current coding block, as shown in FIG. 13A. As shown in FIG. 13A, non-neighboring neighboring block 1301 on the left and non-neighboring neighboring block 1302 on the top have the same size as the current block 1303. In some embodiments, the non-neighboring neighboring blocks at each distance may have a different block size than the current coding block, as shown in FIG. 13B. Neighboring block 1304 is a neighboring block of the current block 1303. As shown in FIG. 13B, non-neighboring neighboring block 1305 on the left and non-neighboring neighboring block 1306 on the top have the same size as the current block 1307. Neighboring block 1308 is a neighboring block of the current block 1307.

[0121] It should be noted that when the non-neighboring blocks at each distance have the same block size as the current coding block, the value of the block size is adaptively changed according to the partitioning granularity in each different area within the image. When the non-neighboring blocks at each distance have a block size different from that of the current coding block, the value of the block size can be predefined as a constant value, such as 4x4, 8x8, or 16x16. The 4x4 non-neighboring motion field shown in Figures 10 and 12 is an example of this case, and the motion field can be considered as a special case of, but not limited to, a sub-block.

[0122] Similarly, the non-adjacent coding blocks shown in Figure 11 may have different sizes as well. In one embodiment, the non-adjacent coding blocks may have a size as the current coding block that is adaptively changed. In another embodiment, the non-adjacent coding blocks may have a predefined size with a fixed value, such as 4x4, 8x8, or 16x16.

[0123] Based on the defined scanning distance, the total size of the scanning area on either the left or top of the current coding block may be determined by a configurable distance value. In one or more embodiments, the maximum scanning distance on the left and top sides may use the same value or different values. FIG. 13 shows an example in which the maximum distance on both the left and top sides share the same value of 2. The maximum scanning distance value(s) may be determined by the encoder side and signaled in the bitstream. Alternatively, the maximum scanning distance value(s) may be predefined as fixed value(s), such as a value of 2 or 4. When the maximum scanning distance is predefined as a value of 4, it indicates that the scanning process is completed when the candidate list is full, or that all non-adjacent neighboring blocks with a maximum distance of 4 have been scanned, whichever comes first.

[0124] In one or more embodiments, within each scanning area at a particular distance, the start and end neighborhood blocks may be position dependent.

[0125] In some embodiments, for a left scanning area, the starting neighboring block may be the adjacent lower-left block of the starting neighboring block of the neighboring scanning area with a shorter distance. For example, as shown in FIG. 8, the starting neighboring block of the "Distance 2" scanning area on the left side of the current block is the adjacent lower-left block of the starting neighboring block of the "Distance 1 (D1)" scanning area. In FIG. 8, D1, D2, and D3 represent distances 1, 2, and 3, respectively. The ending neighboring block may be the adjacent left block of the ending neighboring block of the upper scanning area with a shorter distance. For example, as shown in FIG. 8, the ending neighboring block of the "Distance 2" scanning area on the left side of the current block is the adjacent left neighboring block of the ending neighboring block of the "Distance 1" scanning area above the current block.

[0126] Similarly, for the upper scanning area, the starting neighboring block may be the right-top block adjacent to the starting neighboring block of the neighboring scanning area with a shorter distance, and the ending neighboring block may be the left-top block adjacent to the ending neighboring block of the neighboring scanning area with a shorter distance.

[0127] Scanning Order When neighboring blocks are scanned in non-contiguous areas, certain order or / and rules may be followed to determine the selection of scanned neighboring blocks.

[0128] In some embodiments, the left area may be scanned first, followed by scanning the area above. As shown in Figure 8, three lines of non-adjacent area on the left side (e.g., distance 1 (D1) to distance 3 (D3)) may be scanned first, followed by scanning three lines of non-adjacent area above the current block.

[0129] In some embodiments, the left area and the top area may be scanned alternately. For example, as shown in Figure 8, the left scanning area with "Distance 1" is scanned first, followed by scanning the top area with "Distance 1".

[0130] For scanning areas located on the same side (e.g., left or top areas), the scanning order is from the area with the shortest distance to the area with the longest distance. This order can be flexibly combined with other embodiments of the scanning order. For example, the left and top areas can be scanned alternately, and the order for the areas on the same side is scheduled to be from the shortest distance to the longest distance.

[0131] Within each scanning area at a particular distance, a scanning order can be defined. In one embodiment, for the left scanning area, scanning can start from the bottom neighboring block to the top neighboring block. For the top scanning area, scanning can start from the right block to the left block.

[0132] Scanning completed For the inherited merge candidates, neighboring blocks coded in affine mode are defined as qualified candidates. In some embodiments, the scanning process may be performed iteratively. For example, scanning performed within a particular area at a particular distance may be stopped in an instance when the first X qualified candidates are identified, where X is a predefined positive value. For example, as shown in FIG. 8, scanning within the left scanning area with a distance of 1 may be stopped when the first one or more qualified candidates are identified. Then, the next iteration of the scanning process begins by targeting another scanning area, governed by a predefined scanning order / rule.

[0133] In one or more embodiments, X may be defined for each distance. For example, at each distance, X may be set to 1, meaning that scanning is completed for each distance when the first qualified candidate is found and the scanning process is resumed from a different distance in the same area or from the same or a different distance in a different area. It should be noted that the value of X may be set as the same value or different values ​​for different distances. The scanning process for an area is fully completed when the maximum number of qualified candidates is found from all allowable distances (e.g., bounded by the maximum distance) of the area.

[0134] In another embodiment, X may be defined for an area. For example, X may be set to 3, which means that scanning is completed for the entire area (e.g., the area to the left or above the current block) when the first three qualified candidates are found and the scanning process is resumed from the same or a different distance in another area. It should be noted that the value of X may be set as the same or different values ​​for different areas. When the maximum number of qualified candidates is found from all areas, the scanning process is completed globally.

[0135] The value of X can be defined for both distance and area, for example, for each area (e.g., the area to the left or top of the current block), X is set to 3, and for each distance, X is set to 1. The value of X can be set as the same or different values ​​for different areas and distances.

[0136] In some embodiments, the scanning process may be performed continuously, for example, scanning performed within a particular area at a particular distance may be stopped in instances when all covered neighboring blocks have been scanned and no more qualified candidates have been identified, or the maximum allowed number of candidates has been reached.

[0137] During the candidate scanning process, the non-adjacent neighboring blocks of each candidate are determined and scanned by adhering to the scanning method proposed above. For easier implementation, the non-adjacent neighboring blocks of each candidate can be indicated and identified by a specific scanning position. Once the specific scanning area and distance are determined by adhering to the method proposed above, the scanning position can be determined accordingly based on the following method.

[0138] In one method, the bottom left and top right positions are used for the top and left non-adjacent neighboring blocks, as shown in Figure 15A.

[0139] Alternatively, the bottom right position is used for both the top and left non-adjacent neighboring blocks, as shown in Figure 15B.

[0140] Alternatively, the bottom left position is used for both the top and left non-adjacent neighboring blocks, as shown in Figure 15C.

[0141] Alternatively, the right-top position is used for both the top and left non-adjacent neighboring blocks, as shown in Figure 15D.

[0142] For easier illustration, in Figures 15A-15D, each non-adjacent neighboring block is assumed to have the same block size as the current block. Without loss of generality, this illustration can be easily extended to non-adjacent neighboring blocks having different block sizes.

[0143] Furthermore, step 2 can utilize the same process for CPMV projection as used in the current AVS and VVC standards, where it is assumed that the current block shares the same affine model with the selected neighboring blocks, and then the coordinates of two or three corner pixels (e.g., if the current block uses a 4-parameter model, two coordinates (the top-left pixel / sample location and the top-right pixel / sample location) are used, and if the current block uses a 6-parameter model, three coordinates (the top-left pixel / sample location, the top-right pixel / sample location, and the bottom-left pixel / sample location) are used) are plugged into equation (1) or (2), depending on whether the neighboring blocks are coded with a 4-parameter or 6-parameter affine model, to generate two or three CPMVs.

[0144] In step 3, any qualified candidate identified in step 1 and converted in step 2 may undergo a similarity check against all existing candidates already in the merge candidate list. The details of similarity check have already been explained in the "Affine Merge Candidate Pruning" section above. If the newly qualified candidate is found to be similar to any existing candidate in the candidate list, then this newly qualified candidate will be removed / pruned.

[0145] Non-proximal Neighborhood-Based Derivation Process for Affine Constructed Merge Candidates In the case of deriving inherited merge candidates, one neighboring block is identified at a time, and this single neighboring block needs to be coded in affine mode and may contain two or three CPMVs. In the case of deriving constructed merge candidates, two or three neighboring blocks are identified at a time, and each identified neighboring block does not need to be coded in affine mode, and only one translational MV is extracted from this block.

[0146] Figure 9 presents an example in which preconstructed affine merge candidates can be derived using non-neighboring neighboring blocks. In Figure 9, A, B, and C are the geographic positions of three non-neighboring neighboring blocks. A hypothetical coding block is formed using the position of A as the top-left corner, the position of B as the top-right corner, and the position of C as the bottom-left corner. Considering a hypothetical CU as an affine-coded block, the MVs at positions A', B', and C' can be derived by adhering to Equation (3), and the model parameters (a, b, c, d, e, f) can be calculated by the translational MVs at positions A, B, and C. Once derived, the MVs at positions A', B', and C' can be used as three CPMVs for the current block, and an existing process (such as the one used in the AVS and VVC standards) for generating preconstructed affine merge candidates can be used.

[0147] For the constructed merge candidates, the non-neighborhood-based derivation process can be performed in five steps. The non-neighborhood-based derivation process can be performed in a device such as an encoder or decoder in five steps: Step 1 is for candidate scanning; Step 2 is for affine model determination; Step 3 is for CPMV projection; Step 4 is for candidate generation; and Step 5 is for candidate pruning. In Step 1, non-neighborhood blocks can be scanned and selected by the following method:

[0148] Scanning Area and Distance In some embodiments, to maintain rectangular coding blocks, the scanning process is performed only for two non-neighboring neighboring blocks, and the third non-neighboring neighboring block may depend on the horizontal and vertical positions of the first and second non-neighboring neighboring blocks.

[0149] In some embodiments, as shown in Figure 9, the scanning process is performed only for the positions of B and C. The position of A can be uniquely determined by the horizontal position of C and the vertical position of B.

[0150] To form a valid hypothetical coding block, position A may need to be at least valid. The validity of position A may be defined as whether motion information is available at position A. In one embodiment, the coding block located at position A may need to be coded in inter mode so that motion information is available to form a hypothetical coding block.

[0151] In some embodiments, the scanning area and distance may be defined according to a particular scanning direction.

[0152] In some embodiments, the scanning direction may be perpendicular to the side of the current block. One example is shown in FIG. 10, where the scanning area is defined as one line of contiguous motion fields to the left or above the current block. The scanning distance is defined as the number of motion fields from the scanning position to the side of the current block. It should be noted that the size of the motion fields may depend on the maximum granularity of the applicable video coding standard. In the example shown in FIG. 10, the size of the motion fields is assumed to be set to 4x4, consistent with the current VVC standard.

[0153] In some embodiments, the scanning direction may be parallel to the sides of the current block. One example is shown in Figure 11, where the scanning area is defined as one line of contiguous coding blocks to the left or above the current block.

[0154] In some embodiments, the scanning direction can be a combination of vertical and horizontal scanning to the side of the current block. One example is shown in FIG. 12. As shown in FIG. 12, the scanning direction can also be a combination of parallel and diagonal. Scanning at position B starts from left to right, then diagonally to the block to the left and above. Scanning at position B repeats as shown in FIG. 12. Similarly, scanning at position C starts from top to bottom, then diagonally to the block to the left and above. Scanning at position C repeats as shown in FIG. 12.

[0155] Scanning Order In some embodiments, the scanning order may be defined as from a position with a smaller distance to the current coding block to a position with a larger distance, which may be applied in the case of vertical scanning.

[0156] In some embodiments, the scanning order can be defined as a fixed pattern. This fixed pattern scanning order can be used for candidate positions with similar distances. One example is the case of horizontal scanning. In one example, the scanning order can be defined as a top-to-bottom direction for the left scanning area and a left-to-right direction for the top scanning area, similar to the example shown in FIG. 11.

[0157] For the case of a combined scanning method, the scanning order can be a combination of fixed pattern and distance dependent, similar to the example shown in FIG.

[0158] Scanning completed For pre-constructed merge candidates, the qualified candidates do not need to be affine coded, since only translational MVs are required.

[0159] Depending on the number of candidates required, the scanning process may be concluded when the first X qualified candidates have been identified, where X is a positive value.

[0160] As shown in Figure 9, three corners, named A, B, and C, are needed to form a virtual coding block. For easier implementation, the scanning process in step 1 can be performed only to identify non-adjacent neighboring blocks located at corners B and C, and the coordinate of A can be accurately determined by taking the horizontal coordinate of C and the vertical coordinate of B. In this way, the formed virtual coding block is constrained to be rectangular. In cases where either point B or C is unavailable, e.g., outside the boundary, or where motion information for the non-adjacent neighboring blocks corresponding to B or C is unavailable (e.g., the block is coded in intra mode or screen-content mode), the horizontal or vertical coordinate of C can be defined as the horizontal or vertical coordinate, respectively, of the top-left point of the current block.

[0161] In another embodiment, when corner B and / or corner C are initially determined from the scanning process in step 1, non-adjacent neighboring blocks located at corners B and / or C may be identified accordingly. Second, the position(s) of corners B and / or C may be reset to a pivot point within the corresponding non-adjacent neighboring block, such as the center of mass of each non-adjacent neighboring block. For example, the center of mass may be defined as the geometric center of each neighboring block.

[0162] As shown in Figure 9, when the scanning process is performed for corner B and corner C, the process can be performed jointly or independently. In the independent scanning embodiment, the scanning method proposed previously can be applied separately to corners B and C. In the joint scanning embodiment, there can be different methods as follows:

[0163] In one embodiment, pairwise scanning may be performed. In one example of pairwise scanning, the candidate positions for corners B and C advance simultaneously. For easier illustration and without loss of generality, FIG. 17B will be taken as an example. As shown in FIG. 17B, scanning for corner B begins in a bottom-up direction from the first non-adjacent neighboring block located above the current block. Scanning for corner C begins in a right-left direction from the first non-adjacent neighboring block located to the left of the current block. Therefore, in the example shown in FIG. 17B, pairwise scanning may be defined as advancing the candidate positions for both B and C by one unit of step size, where one unit of step size is defined as the height of the current coding block for corner B and the width of the current coding block for corner C.

[0164] In another embodiment, alternative scanning may be performed. In one example of alternative scanning, the candidate positions for corners B and C are alternately advanced. In one step, only the position of B or C may be advanced, while the position of C or B remains unchanged. In one example, the position of corner B may be incrementally increased from the first non-adjacent neighboring block to a distance of the maximum number of non-adjacent neighboring blocks, while the position of corner C remains at the first non-adjacent neighboring block. In the next round, the position of corner C moves to the second non-adjacent neighboring block, and the position of corner B is again traversed from the beginning to the maximum value. The rounds continue until all combinations have been traversed.

[0165] For purposes of uniformity, the methods of defining scanning areas and distances, scanning order, and scanning completion proposed for deriving inherited merge candidates may be fully or partially reused for deriving constructed merge candidates. In one or more embodiments, the same methods defined for inherited merge candidate scanning, including but not limited to scanning areas and distances, scanning order, and scanning completion, may be fully reused for constructed merge candidate scanning.

[0166] In some embodiments, the same method defined for inherited merge candidate scanning can be partially reused for constructed merge candidate scanning. Figure 16 shows an example of this case. In Figure 16, the block size of each non-adjacent neighboring block is the same as the current block, which is defined similarly as inherited candidate scanning, but the overall process is a simplified version because scanning at each distance is limited to only one block.

[0167] 17A-17B show another example of this case, where both the non-adjacent inherited merging candidates and the non-adjacent constructed merging candidates are defined with the same block size as the current coding block, and the scanning order, scanning area, and scanning completion condition may be defined differently.

[0168] In FIG. 17A, the maximum distance for the non-close neighbor on the left is four coding blocks, and the maximum distance for the non-close neighbor on the top is five coding blocks. Also, at each distance, the scanning direction is bottom-top for the left side and right-left for the top side. In FIG. 17B, the maximum distance for non-close neighbors is four for both the left and top sides. Additionally, scanning at certain distances is unavailable because there is only one block at each distance. In FIG. 17A, if M qualified candidates are identified, the scanning operation within each distance can be completed. The value of M can be a predefined fixed value, such as a value of 1 or any other positive integer, or a signaled value determined by the encoder, or a configurable value at the encoder or decoder. In one embodiment, the value of M can be equal to the merge candidate list size.

[0169] 17A-17B, scanning operations at different distances may be completed when N qualified candidates are identified. The value of N may be a predefined fixed value, such as a value of 1 or any other positive integer, or a signaled value determined by the encoder, or a configurable value in the encoder or decoder. In one embodiment, the value of N may be the same as the merge candidate list size. In another embodiment, the value of N may be the same as the value of M.

[0170] In both Figures 17A and 17B, non-close spatial neighbors with closer distances to the current block may be prioritized, which indicates that a non-close spatial neighbor with distance i is scanned or checked before a neighbor with distance i+1, where i may be a non-negative integer representing a particular distance.

[0171] At a certain distance, at most two non-adjacent spatial neighbors are used, which means that if available, at most one neighbor from one side of the current block, e.g., from the left and from the top, is selected for inherited or constructed candidate derivation. As shown in Figure 17A, the checking order for the left and top neighbors is bottom-top and right-left, respectively. For Figure 17B, this rule can also be applied, the difference being that at any particular distance, there can be only one option per side of the current block.

[0172] For a constructed candidate, the positions of one left and one top non-adjacent spatial neighbor are first determined independently, as shown in Fig. 17B. Then, the position of the left-top neighbor, which can surround a rectangular virtual block together with the left and top non-adjacent neighbors, can be determined accordingly. Then, as shown in Fig. 9, the motion information of the three non-adjacent neighbors is used to form CPMVs at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, which are finally projected to the current CU to generate a corresponding constructed candidate.

[0173] In step 2, the translational MVs at the selected candidate positions after step 1 are evaluated and a suitable affine model can be determined. For easier illustration, and without loss of generality, Figure 9 is again used as an example.

[0174] Due to factors such as hardware constraints, implementation complexity, and different reference indexes, the scanning process may be completed before a sufficient number of candidates are identified, for example, motion information for the motion fields in one or more of the selected candidates after step 1 may not be available.

[0175] If motion information for all three candidates is available, the corresponding hypothetical coding block represents a 6-parameter affine model. If motion information for one of the three candidates is not available, the corresponding hypothetical coding block represents a 4-parameter affine model. If motion information for more than one of the three candidates is not available, the corresponding hypothetical coding block may be unusable to represent a valid affine model.

[0176] In some embodiments, if motion information is not available for the top left corner of the virtual coding block, e.g., corner A in FIG. 9, or if motion information is not available for both the top right corner, e.g., corner B in FIG. 9, and the bottom left corner, e.g., corner C in FIG. 9, the virtual block may be set as invalid and may be unusable to represent a valid model, and then steps 3 and 4 may be skipped during the current iteration.

[0177] In some embodiments, if either the top right corner, e.g., corner B in FIG. 9, or the bottom left corner, e.g., corner C in FIG. 9, is unavailable, but not both, the virtual block may represent a valid four-parameter affine model.

[0178] In step 3, if the hypothetical coding block is capable of representing a valid affine model, the same projection process used for the inherited merge candidates can be used.

[0179] In one or more embodiments, the same projection process used for the inherited merge candidate may be used, in this case the 4-parameter model represented by the hypothetical coding block from step 2 is projected onto a 4-parameter model for the current block, and the 6-parameter model represented by the hypothetical coding block from step 2 is projected onto a 6-parameter model for the current block.

[0180] In some embodiments, the affine model represented by the virtual coding block from step 2 is always projected onto a 4-parameter model or a 6-parameter model for the current block.

[0181] According to equations (5) and (6), there can be two types of four-parameter affine models: Type A is when the top-left corner CPMV and the top-right corner CPMV, designated as V0 and V1, are available; Type B is when the top-left corner CPMV and the bottom-left corner CPMV, designated as V0 and V2, are available.

[0182] In one or more embodiments, the type of the projected 4-parameter affine model is the same type of 4-parameter affine model represented by the virtual coding block, e.g., if the affine model represented by the virtual coding block from step 2 is a 4-parameter affine model of type A or B, then the projected affine model for the current block is also of type A or B, respectively.

[0183] In some embodiments, the 4-parameter affine model represented by the hypothetical coding block from step 2 is always projected onto a 4-parameter model of the same type for the current block. For example, a 4-parameter affine model of type A or B represented by a hypothetical coding block is always projected onto a 4-parameter affine model of type A.

[0184] In step 4, based on the projected CPMV after step 3, in one embodiment, the same candidate generation process used in the current VVC or AVS standard may be used. In another embodiment, the temporal motion vectors used in the candidate generation process used in the current VVC or AVS standard may not be used for the non-adjacent neighborhood block-based derivation method. When a temporal motion vector is not used, it indicates that the generated combination does not include any temporal motion vector.

[0185] In step 5, any newly generated candidate after step 4 may undergo a similarity check against all existing candidates already in the merge candidate list. The details of the similarity check have been previously explained in the "Affine Merge Candidate Pruning" section. If the newly generated candidate is found to be similar to any existing candidate in the candidate list, the newly generated candidate is removed or pruned.

[0186] Inheritance-Based Derivation Method for Affine Constructed Merge Candidates For each affine inherited candidate, all motion information is inherited from one selected spatial neighboring block coded in affine mode. The inherited information includes CPMV, reference index, prediction direction, affine model type, etc. On the other hand, for each affine constructed candidate, all motion information is constructed from two or three selected spatial or temporal neighboring blocks, and the selected neighboring blocks are not coded in affine mode, and only translational motion information is needed from the selected neighboring blocks.

[0187] In this section, a new candidate derivation method is disclosed that combines features of inherited and constructed candidates.

[0188] In some embodiments, the combination of inheritance and construction may be achieved by separating the affine model parameters into different groups, with one group of affine parameters inherited from one neighboring block and another group of affine parameters inherited from another neighboring block.

[0189] In one embodiment, the parameters of an affine model are constructed from two groups. As shown in Equation (3), an affine model may include six parameters, including a, b, c, d, e, and f. The translational parameters {a, b} may represent one group, and the non-translational parameters {c, d, e, f} may represent another group. With this grouping method, the two groups of parameters may be inherited independently from two different neighboring blocks in a first step, and then concatenated / constructed into a complete affine model in a second step. In this case, the group with non-translational parameters must be inherited from one affine-coded neighboring block, and the group with translational parameters may be from any inter-coded neighboring block that may or may not be coded in affine mode. It should be noted that the affine-coded neighboring blocks may be selected from the near affine neighboring blocks or the non-neighboring affine neighboring blocks based on the scanning method previously proposed for the affine-inherited candidate, such as the method shown in Figure 17A, which is the scanning method / rule including the scanning area and distance, scanning order, and scanning completion used in the section "Non-neighborhood-based Derivation Process for Affine-Inherited Merge Candidates," and the scanning method may be performed for both the near and non-neighboring neighboring blocks. Alternatively, the affine-coded neighboring blocks may not physically exist but may be virtually constructed from the regular inter-coded neighboring blocks, such as the method shown in Figure 17B, which is the scanning method / rule including the scanning area and distance, scanning order, and scanning completion used in the section "Non-neighborhood-based Derivation Process for Affine-Constructed Merge Candidates."

[0190] In some embodiments, the neighboring blocks associated with each group can be determined in different ways. In one method, the neighboring blocks for different groups of parameters can be all from non-proximate neighborhood areas, and the scanning method can be designed similar to previously proposed methods for non-proximate neighborhood-based derivation processes. In another method, the neighboring blocks for different groups of parameters can be all from proximate neighborhood areas, and the scanning method can be the same as the current VVC or AVS video standard. In another method, the neighboring blocks for different groups of parameters can be partial from proximate areas and partial from non-proximate neighborhood areas.

[0191] When neighboring blocks are scanned from non-adjacent neighborhood areas to construct the current type of candidate, the scanning process may be performed differently from the non-adjacent neighborhood-based derivation process for affine inherited candidates. In one or more embodiments, the scanning area, distance, and order may be defined similarly, but the scanning completion rule may be specified differently. For example, non-adjacent neighborhood blocks may be scanned exclusively within a predefined maximum distance in each area. In this case, all non-adjacent neighborhood blocks within the distance may be scanned by adhering to the scanning order. In some embodiments, the scanning area may be different. For example, in addition to the left and top areas, the adjacent and non-adjacent areas to the lower right of the current coding block may be scanned to determine neighbors for generating translational and / or non-translational parameters. Additionally, the neighbors scanned in the lower right area may be used to find co-located temporal neighbors instead of spatial neighbors. One scanning criterion may be conditional based on whether the co-located temporal neighbor to the lower right has already been used to generate the affine constructed neighborhood. If it is already in use, no scanning is performed, otherwise scanning is performed. Alternatively, if it is already in use, meaning that the temporal neighbor(s) at the same position in the bottom right are available, scanning is performed, otherwise scanning is not performed.

[0192] When several groups of affine parameters are combined to construct a new candidate, there may be several rules to be observed. The first is eligibility criteria. In one embodiment, it may be checked whether the related neighboring block or blocks for each group use the same reference picture for at least one direction or both directions. In another embodiment, it may be checked whether the related neighboring block or blocks for each group use the same precision / resolution for motion vectors.

[0193] When a criterion is checked, the first X relevant neighboring blocks per group may be used. The value of X may be defined as the same or different values ​​for different groups of parameters. For example, the first one or two neighboring blocks containing non-translational affine parameters may be used, and the first three or four neighboring blocks containing translational affine parameters may be used.

[0194] The second is a construction formula. In one embodiment, the CPMV of a new candidate can be derived in the following formula:

number

[0195] In another embodiment, the CPMV of the new candidate may be derived in the following manner:

number

[0196] FIG. 18 shows an example of an inheritance-based derivation method for deriving an affine constructed candidate. In FIG. 18, there are three steps for deriving an affine constructed candidate. In step 1, according to a specific grouping strategy, the encoder or decoder may perform scanning of nearby and non-neighboring neighboring blocks for each group. In the case of FIG. 18, two groups are defined: Neighbor 1 is coded in affine mode and provides non-translational affine parameters, and Neighbor 2 provides translational affine parameters. Neighbor 1 may be obtained according to the process in the "Non-neighborhood-based Derivation Process for Affine Inherited Merge Candidates" section shown in FIGS. 15A-15D and 17A, and Neighbor 1 may be a nearby or non-neighboring neighbor of the current block. Furthermore, Neighbor 2 may be obtained according to the process shown in FIGS. 16 and 17B.

[0197] In some embodiments, Neighbor 1, coded in affine mode, can be scanned from adjacent and / or non-adjacent areas by adhering to the scanning method proposed above. In some embodiments, Neighbor 2, coded in affine or non-affine mode, can also be scanned from adjacent or non-adjacent areas. For example, Neighbor 2 can be from one of the scanned adjacent or non-adjacent areas if the motion information has not already been used to derive some affine merge or AMVP candidates, or it can be from the bottom-right position of the current block if a co-located TMVP candidate at this position is available and / or has already been used to derive some affine merge or MVP candidates. Alternatively, a small coordinate offset (e.g., +1 or +2 or -1 or -2 for the vertical and / or horizontal coordinates) can be applied when determining the position of Neighbor 2 to provide slightly more diversified motion information for constructing new candidates.

[0198] In step 2, the parameters and positions determined in step 1 may define a specific affine model that can derive different CPMVs according to the coordinates (x, y) of the CPMVs. For example, as shown in FIG. 18, the non-translational parameters {c, d, e, f} may be obtained based on the neighborhood 1 obtained in step 1, and the translational parameters {a, b} may be obtained based on the neighborhood 2 obtained in step 1. Furthermore, the distance parameters Δw and Δh may be obtained based on the position (x1, y1) of the current block and the position (x2, y2) of neighborhood 2. The distance parameters Δw and Δh may indicate the horizontal and vertical distances between the current block and neighborhood 1 or neighborhood 2, respectively. For example, the distance parameters Δw and Δh may indicate the horizontal distance (x1-x2) between the current block and neighborhood 2, and the vertical distance (y1-y2) between the current block and neighborhood 2, respectively. In particular, Δw=x1-x2 and Δh=y1-y2.

[0199] In step 3, two or three CPMVs are derived for the current coding block, and the current coding block can be constructed to form a new affine candidate.

[0200] In some embodiments, other prediction information may be further constructed. If neighboring blocks are checked to have the same direction and / or reference picture, the prediction direction (e.g., bi-predictive or uni-predictive) and the reference picture index may be the same as that of the associated neighboring block. Alternatively, the prediction information is determined by reusing the least overlapping information among the associated neighboring blocks from different groups. For example, if only one reference index in one direction from one neighboring block is the same as the reference index in the same direction of another neighboring block, the prediction direction of the new candidate is determined as uni-predictive, and the same reference index and direction are reused.

[0201] HMVP-Based Deriving Method for Affine Constructed Merge Candidates In the case of the near-neighborhood-based derivation process, already defined in the current video standards VVC and AVS and described in the above section and in FIG. 7, a fixed order of scanning for near neighborhoods is performed to identify two or three near neighbor blocks. In the case of the non-neighborhood-based derivation process, as proposed in the previous section and in FIG. 17B, two non-neighbor neighborhoods are identified during another fixed order of scanning. In other words, for both the near-neighborhood-based derivation method and the non-neighborhood-based derivation method, a certain depth of local scanning is necessary to identify the number of neighbors. This scanning process relies on local buffering around each current block and also incurs a certain amount of computational complexity.

[0202] On the other hand, as explained in the introduction section, the HMVP merge mode is already adopted in current VVC and AVS, and the translational motion information from neighboring blocks is already stored in the history table. In this case, the scanning process can be replaced by searching the HMVP table.

[0203] Therefore, for the previously proposed non-adjacent neighborhood-based derivation process and inheritance-based derivation process, instead of the scanning method shown in Figures 17B and 18, translational motion information can be obtained from the HMVP table. However, to subsequently derive affine constructed candidates, position information, width, height, and reference information are also required, which can be accessible if the current HMVP table can be modified. Therefore, it is proposed to extend the HMVP table to store additional information in addition to the motion information of each historical neighborhood. In one embodiment, the additional information can include the position of affine or non-affine neighboring blocks, or affine motion information such as CPMV or equivalent normal motion derived from CPMV (e.g., this normal motion can be from an internal sub-block of the affine-coded neighboring block), a reference index, etc.

[0204] Candidate Derivation Methods for Affine AMVP and Regular Merge Modes As explained in the above section, for affine AMVP mode, an affine candidate list is also required to derive the CPMV predictor. As a result, all the above proposed derivation methods can be applied to the affine AMVP mode as well. The only difference is that when the above proposed derivation methods are applied in AMVP, the selected neighboring block should have the same reference picture index as the current coding block.

[0205] For the regular merge mode, the candidate list is also constructed only with translational candidate MVs, not CPMVs. In this case, all the derivation methods proposed above can still be applied by adding an additional derivation step. This additional derivation step is to derive a translational MV for the current block, which can be achieved by selecting a specific rotation position (x,y) within the current block and observing the same equation (3). In other words, to derive the CPMV of an affine block, the positions of the three corners of the block can be used as the rotation position (x,y) in equation (3), and to derive the translational MV of a regular inter-coded block, the center position of the block can be used as the rotation position (x,y) in equation (3). Once the translational MV is derived for the current block, it can be inserted into the candidate list as another candidate.

[0206] When new candidates are derived based on the methods proposed above for the affine AMVP and regular merge modes, the replacements of the new candidates may be reordered.

[0207] In one embodiment, newly derived candidates may be inserted into the affine AMVP candidate list by adhering to the following order: (1) inherited from nearby spatial neighbors; (2) constructed from close spatial neighborhoods; (3) inherited from non-proximate spatial neighbors; (4) constructed from non-proximate spatial neighborhoods; (5) translational MVs from close spatial neighbors; (6) temporal MVs from close temporal neighborhoods, and (7) Zero MV.

[0208] In another embodiment, newly derived candidates may be inserted into the affine AMVP candidate list by adhering to the following order: (1) inherited from nearby spatial neighbors; (2) constructed from close spatial neighborhoods; (3) inherited from non-proximate spatial neighbors; (4) translational MVs from close spatial neighbors; (5) constructed from non-proximate spatial neighborhoods; (6) temporal MVs from close temporal neighborhoods, and (7) Zero MV.

[0209] In another embodiment, newly derived candidates may be inserted into the affine AMVP candidate list by adhering to the following order: (1) inherited from nearby spatial neighbors; (2) constructed from close spatial neighborhoods; (3) translational MVs from close spatial neighbors; (4) inherited from non-proximate spatial neighbors; (5) constructed from non-proximate spatial neighborhoods; (6) temporal MVs from close temporal neighborhoods, and (7) Zero MV.

[0210] In another embodiment, newly derived candidates may be inserted into the affine AMVP candidate list by adhering to the following order: (1) inherited from nearby spatial neighbors; (2) constructed from close spatial neighborhoods; (3) translational MVs from close spatial neighbors; (4) Temporal MVs from close temporal neighborhoods, (5) inherited from non-proximate spatial neighbors; (6) constructed from non-proximate spatial neighborhoods, and (7) Zero MV.

[0211] In another embodiment, newly derived candidates may be inserted into the affine AMVP candidate list by adhering to the following order: (1) inherited from nearby spatial neighbors; (2) constructed from close spatial neighborhoods; (3) translational MVs from close spatial neighbors; (4) Temporal MVs from close temporal neighborhoods, (5) inherited from non-proximate spatial neighbors, and (6) Zero MV.

[0212] In another embodiment, newly derived candidates may be inserted into the regular merge candidate list by adhering to the following order: (1) Spatial MVP from close spatial neighborhoods, (2) Temporal MVP from nearby co-located neighbors, (3) spatial MVP from non-close spatial neighbors; (4) Inherited MVP from non-proximal spatially affine neighborhoods, (5) Constructed MVP from non-close spatial neighbors (6) History-based MVP from FIFO table, (7) Pairwise average MVP, and (8) Zero MV.

[0213] Reordering the affine merge candidate list In one embodiment, non-neighboring spatial merge candidates may be inserted into the affine merge candidate list by adhering to the following order: 1. subblock-based temporal motion vector prediction (SbTMVP) candidates, if available; 2. inherited from close neighbors; 3. inherited from non-neighbors; 4. constructed from close neighbors; 5. constructed from non-neighbors; 6. zero MV.

[0214] In another embodiment, non-neighboring spatial merge candidates can be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate if available, 2. inherited from close neighbors, 3. constructed from close neighbors, 4. inherited from non-neighbors, 5. constructed from non-neighbors, 6. zero MV.

[0215] In another embodiment, non-neighboring spatial merge candidates can be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate if available, 2. inherited from nearby neighbors, 3. constructed from nearby neighbors, 4. one set of zero MVs, 5. inherited from non-neighbors, 6. constructed from non-neighbors, 7. remaining zero MVs if the list is still not full.

[0216] In another embodiment, non-close spatial merge candidates can be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate, if available, inherited from close neighbors; 3. inherited from non-close neighbors by a distance less than X; 4. constructed from close neighbors; 5. constructed from non-close neighbors; 6. constructed from inherited translated and non-translated neighbors; 7. zero MV if the list is still not full.

[0217] In another embodiment, non-neighboring spatial merge candidates may be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate if available, 2. inherited from nearby neighbors, 3. inherited from non-neighbors, 4. first candidate constructed from nearby neighbors, 5. first X candidates constructed from inherited translated and non-translated neighbors, 6. constructed from non-neighbors, 7. another Y candidates constructed from inherited translated and non-translated neighbors, 8. zero MV if the list is still not full.

[0218] In some embodiments, the values ​​of X and Y may be predefined fixed values, such as a value of 2, or signaled values ​​received by the decoder (sequence / slice / block / CTU-level signaled parameters), or configurable values ​​at the encoder / decoder, or dynamically determined values ​​(e.g., X<=3, Y<=3) according to the number of available neighbors to the left and above each individual coding block, or any combination of methods for determining the values ​​of X and Y. In one embodiment, the value of X may be the same as the value of Y. In another embodiment, the value of X may be different from the value of Y.

[0219] In another embodiment, non-close spatial merge candidates can be inserted into the affine merge candidate list by adhering to the following order: 1. SbTMVP candidate, if available; 2. inherited from close neighbors, inherited from neighbors not closer than X; 4. constructed from close neighbors; 5. constructed from neighbors not closer than Y; 6. inherited from neighbors not closer than X; 7. constructed from neighbors not closer than Y; 8. zero MV. In this embodiment, the values ​​X and Y can be predefined fixed values, such as a value of 2, or signaled values ​​determined by the encoder, or configurable values ​​in the encoder or decoder. In one embodiment, the value of X can be the same as the value of Y. In another embodiment, the value of N can be different from the value of M.

[0220] In some embodiments, when a new candidate is derived by using an inheritance-based derivation method that constructs a CPMV by combining affine motion and translational MVs, the placement of this new candidate may depend on the placement of other constructed candidates.

[0221] In one embodiment, for different constructed candidates, the reordering of the affine merge candidate list may adhere to the following order: (1) Constructed from close spatial neighborhoods, (2) constructed from combining nearby spatially affine neighborhoods and translational MVs, (3) constructed from non-proximate spatial neighborhoods, and (4) It is constructed from combining non-neighboring spatially affine neighborhoods and translational MVs.

[0222] In another embodiment, for different constructed candidates, the reordering of the affine merge candidate list may adhere to the following order: (1) Constructed from close spatial neighborhoods, (2) constructed from non-proximate spatial neighborhoods; (3) constructed from combining nearby spatially affine neighborhoods and translational MVs; and (4) It is constructed from combining non-neighboring spatially affine neighborhoods and translational MVs.

[0223] Improved reordering of affine candidate lists Based on the candidate derivation method proposed above, one or more candidates may be derived for an existing affine merge candidate list, or affine AMVP candidate list, or regular merge candidate list, and the size of the corresponding list may be adjusted statistically (e.g., configurable size) or adaptively (e.g., dynamically changed according to availability at the encoder and then signaled to the decoder). It should be noted that when one or more new candidates are derived for a regular merge candidate list, the new candidates are first derived as affine candidates and then converted into translational motion vectors by using the coding block and its associated rotation position (e.g., center sample or pixel position) within the affine model before being inserted into the regular merge candidate list.

[0224] In one or more embodiments, an adaptive reordering method such as ARMC may be applied to one or more of the above candidate lists after they have been updated or constructed by adding some new candidates derived by the candidate derivation method proposed above.

[0225] In another embodiment, a temporal candidate list may be first created, and the temporal candidate list may have a size larger than an existing candidate list (e.g., an affine merge candidate list, an affine AMVP candidate list, or a regular merge candidate list). Once the temporal candidate list is constructed by adding newly derived candidates and statistically ordered by using the insertion method proposed above, an adaptive reordering method such as ARMC may be applied to reorder the temporal candidate list. After adaptive reordering, the first N candidates from the temporal candidate list are inserted into the existing candidate list, where the value of N may be a fixed value or a configurable value. In one embodiment, the value of N may be the same as the size of the existing candidate list, and the selected N candidates from the temporal candidate list are identified.

[0226] In the above application scenarios where an adaptive reordering method such as ARMC is applied, the following methods may be used to improve the performance and / or reduce the complexity of the applied reordering method.

[0227] In some embodiments, when template matching costs are used to reorder different candidates, a cost function such as the sum of absolute differences (SAD) between the template samples of the current block and their corresponding reference samples may be used. The template reference samples may be identified by the same motion information of the current block. In cases where fractional motion information is used for the current block, an interpolation filtering process may be used to generate predicted samples for the template. Because the generated predicted samples are used not for the final block reconstruction but rather to compare the motion accuracy between different candidates, the prediction accuracy of the template samples may be relaxed by using an interpolation filter with smaller taps. For example, in cases where an affine merge candidate list is adaptively reordered, a 2-tap or 4-tap interpolation filter may be used to generate predicted samples for the selected template of the current block. Alternatively, the nearest integer sample (skipping the interpolation filtering process entirely) may be used as the predicted sample for the template. When the template matching method is used to adaptively reorder candidates in other candidate lists, such as a regular merge candidate list or an affine AMVP candidate list, an interpolation filter with smaller taps may also be used.

[0228] In some embodiments, when template matching costs are used to reorder different candidates, a cost function such as the SAD between the template samples of the current block and their corresponding reference samples may be used. The corresponding reference samples may be identified at integer or fractional positions. When fractional positions are identified, a certain level of prediction accuracy may be achieved by performing interpolation filtering. Due to limited prediction accuracy, the calculated matching costs for different candidates may include noise level differences. To reduce the impact of noise level cost differences, the calculated matching costs may be adjusted by removing a few least significant bits before the candidate sorting process.

[0229] In some embodiments, if not enough candidates are derived by using different derivation methods, the candidate lists may be padded with zero MVs at the end of each list. In this case, the candidate cost may be calculated only for the first zero MV, and the remaining zero MVs may be statistically assigned arbitrarily large cost values, so that those repeated zero MVs are placed at the end of the corresponding candidate lists.

[0230] In some embodiments, all zero MVs may be statistically assigned an arbitrarily large cost value, so that all zero MVs are placed at the end of the corresponding candidate list.

[0231] In some embodiments, the previous completion method may be applied to the reordering method to reduce the complexity at the decoder side.

[0232] In one or more embodiments, when a candidate list is constructed, different types of candidates may be derived and inserted into the list. If one candidate or one type of candidate does not participate in the reordering process but is selected and signaled to the decoder, the reordering process applied to other candidates may be completed early. In one example, when applying ARMC to an affine merge candidate list, SbTMVP candidates may be excluded from the reordering process. In this case, if the signaled merge index value for an affine-coded block indicates an SbTMVP candidate at the decoder side, the ARMC process may be skipped or completed early for this affine block.

[0233] In another embodiment, if a candidate or a type of candidate does not participate in the reordering process but is selected and signaled to the decoder, both the derivation process and the reordering process for this particular candidate or this particular type of candidate may be skipped. The skipped derivation and reordering process only applies to the particular candidate or the particular type of candidate, and the remaining candidates or types of candidates are still executed. The derivation process is skipped, which indicates that the related operations for deriving the particular candidate or this particular type of candidate are skipped. Note that the predefined list position (e.g., according to the predefined insertion order) of the particular candidate or this particular type of candidate may still be maintained, and the candidate content, such as motion information, may even be invalid due to the skipped derivation process. Similarly, during the reordering process, the cost calculation for this particular candidate or this particular type of candidate may be skipped, and the list position of this particular candidate or this particular type of candidate may not be changed after reordering other candidates.

[0234] 21 illustrates a computing environment (or computing device) 2110 coupled to a user interface 2160. The computing environment 2110 can be part of a data processing server. In some embodiments, the computing device 2110 can perform any of the various methods or processes (encoding / decoding methods or processes) described below according to various embodiments of the present disclosure. The computing environment 2110 can include a processor 2120, a memory 2140, and an I / O interface 2150.

[0235] The processor 2120 typically controls the overall operation of the computing environment 2110, such as operations associated with display, data acquisition, data communication, and image processing. The processor 2120 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Additionally, the processor 2120 may include one or more modules that facilitate interaction between the processor 2120 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single chip machine, a GPU, or the like.

[0236] The memory 2140 is configured to store various types of data to support the operation of the computing environment 2110. The memory 2140 may include predetermined software 2142. Examples of such data include instructions for any applications or methods running on the computing environment 2110, video data sets, image data, etc. The memory 2140 may be implemented using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.

[0237] The I / O interface 2150 provides an interface between the processor 2120 and peripheral interface modules, such as a keyboard, click wheel, and buttons. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 2150 may be coupled to an encoder and a decoder.

[0238] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes a plurality of programs, such as contained in memory 2140, executable by processor 2120 in computing environment 2110 for performing the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.

[0239] A non-transitory computer-readable storage medium has stored thereon a plurality of programs for execution by a computing device having one or more processors, the plurality of programs, when executed by the one or more processors, causing the computing device to perform the above-described method for motion prediction.

[0240] In some embodiments, the computing environment 2110 may be implemented with one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components for performing the methods described above.

[0241] FIG. 22 is a flowchart illustrating a method for video decoding according to an embodiment of the present disclosure.

[0242] In step 2201, the processor 2120 on the decoder side may obtain one or more MV candidates from multiple non-adjacent neighboring blocks to the current block based on at least one scanning area and at least one scanning distance, where one of the at least one scanning distance indicates the number of blocks away from one side of the current block.

[0243] In step 2202, the processor 2120 may determine a completion condition based on the number of MV candidates obtained by scanning at least one scanning distance within a first scanning area, where the at least one scanning area may include the first scanning area.

[0244] In some embodiments, the processor 2120 may determine the completion condition based on the number of MV candidates obtained by scanning a first scanning distance within the first scanning area, and at least one scanning distance may include the first scanning distance.

[0245] In step 2203, the processor 2120 may stop scanning at least one scanning area in response to determining that the completion condition is met.

[0246] In some embodiments, the completion condition may include observing different scanning completion conditions, determining that the number of MV candidates acquired by scanning a first scanning distance within the first scanning area reaches a predetermined first value, determining that the number of MV candidates acquired by scanning at least one scanning distance within the first scanning area reaches a predetermined second value, or determining that the number of MV candidates acquired by scanning a first scanning distance within the first scanning area reaches a predetermined first value and that the number of MV candidates acquired by scanning at least one scanning distance within the first scanning area reaches a predetermined second value. In some embodiments, the predetermined first value is the same as the predetermined second value. In some embodiments, the predetermined first value is different from the predetermined second value.

[0247] For example, when the predetermined first value is set to 1, if the first qualified candidate is found at one distance as illustrated in FIG. 8, scanning can be completed for this distance, and the scanning process can be resumed from a different distance in the same area or from the same or a different distance in a different area.

[0248] For example, when the predetermined second value is set to 3, if the first three qualified candidates are found in one scanning area as illustrated in FIG. 8, scanning can be completed for the entire area (e.g., the area to the left or above the current block), and the scanning process can be resumed from the same or a different distance in another area.

[0249] In some embodiments, the completion condition may include determining that the number of MV candidates acquired by scanning each of the at least one scanning distance reaches a predetermined value, the predetermined value for each of the at least one scanning distance being the same or different, and the first scanning area may be scanned at the at least one scanning distance based on the completion condition until the number of MV candidates acquired from the first scanning area reaches a predetermined maximum value.

[0250] In some embodiments, the at least one scanning area may include a first scanning area and a second scanning area, where the first scanning area is scanned using a plurality of first scanning distances and the second scanning area is scanned using a plurality of second scanning distances. For example, as shown in FIGS. 8 and 14A-14B, the first scanning area may be the left scanning area, and the plurality of first scanning distances may include distance 1, distance 2, and distance 3. The first scanning distance may be one of distance 1, distance 2, or distance 3 used for scanning within the left scanning area. Additionally, the second scanning area may be the top / upper scanning area, and the plurality of scanning distances may include distance 1, distance 2, and distance 3. The second scanning distance can be one of Distance 1, Distance 2, or Distance 3 used for scanning within the top / upper scanning area.

[0251] As discussed above, in some embodiments, a scanning area to the left of the current block, i.e., an area with a "distance of 2" within the left side, indicates that a candidate neighboring block located within this area is two blocks away from the left side of the current block along the vertical direction to the left. Further, an area to the left of the current block with a "distance of 1" indicates that a candidate neighboring block located within this area is one block away from the left side of the current block along the vertical direction to the left, and an area to the left of the current block with a "distance of 3" indicates that a candidate neighboring block located within this area is three blocks away from the left side of the current block along the vertical direction to the left.

[0252] In some embodiments, an area with "distance 1" in the top / top scanning area of ​​the current block, i.e., on the left side, indicates that the candidate neighboring blocks located within this area are one block away from the top side of the current block along the vertical direction to the top. Further, an area with "distance 2" to the left of the current block indicates that the candidate neighboring blocks located within this area are two blocks away from the top side of the current block along the vertical direction to the top, and an area with "distance 3" on the top side of the current block indicates that the candidate neighboring blocks located within this area are three blocks away from the top side of the current block along the vertical direction to the top.

[0253] Further, the completion condition may include determining that the number of MV candidates acquired by scanning the first scanning area reaches a predetermined second value or determining that the number of MV candidates acquired by scanning the second scanning area reaches a predetermined third value, where the predetermined second value is the same as or different from the predetermined third value. Based on these completion conditions, the first scanning area may be scanned at a plurality of first scanning distances, and the second scanning area may be scanned at a plurality of second scanning distances, until the number of MV candidates acquired by scanning the first scanning area and the second scanning area reaches a predetermined maximum value.

[0254] For example, as shown in FIG. 8, the first scanning area may be the left scanning area and the second scanning area may be the top scanning area.

[0255] In step 2204, the processor 2120 may obtain one or more CPMVs for the current block based on the one or more candidate MVs.

[0256] FIG. 23 is a flow chart illustrating a method for video encoding that corresponds to the method for video decoding as shown in FIG.

[0257] In step 2301, the processor 2120 on the encoder side may determine one or more candidates from a plurality of non-adjacent neighboring blocks to the current block based on at least one scanning area and at least one scanning distance, one of which indicates the number of blocks away from one side of the current block.

[0258] In step 2302, the processor 2120 may determine a completion condition based on the number of MV candidates obtained by scanning at least one scanning distance within a first scanning area, where the at least one scanning area may include the first scanning area.

[0259] In some embodiments, the processor 2120 may determine the completion condition based on the number of MV candidates obtained by scanning a first scanning distance within the first scanning area, and at least one scanning distance may include the first scanning distance.

[0260] In step 2303, the processor 2120 may stop scanning at least one scanning area in response to determining that the completion condition is met.

[0261] In some embodiments, the completion condition may include observing different scanning completion conditions, determining that the number of MV candidates acquired by scanning a first scanning distance within the first scanning area reaches a predetermined first value, determining that the number of MV candidates acquired by scanning at least one scanning distance within the first scanning area reaches a predetermined second value, or determining that the number of MV candidates acquired by scanning a first scanning distance within the first scanning area reaches a predetermined first value and that the number of MV candidates acquired by scanning at least one scanning distance within the first scanning area reaches a predetermined second value. In some embodiments, the predetermined first value is the same as the predetermined second value. In some embodiments, the predetermined first value is different from the predetermined second value.

[0262] For example, when the predetermined first value is set to 1, if the first qualified candidate is found at one distance as illustrated in FIG. 8, scanning can be completed for this distance, and the scanning process can be resumed from a different distance in the same area or from the same or a different distance in a different area.

[0263] For example, when the predetermined second value is set to 3, if the first three qualified candidates are found in one scanning area as illustrated in FIG. 8, scanning can be completed for the entire area (e.g., the area to the left or above the current block), and the scanning process can be resumed from the same or a different distance in another area.

[0264] In some embodiments, the completion condition may include determining that the number of MV candidates acquired by scanning each of the at least one scanning distance reaches a predetermined value, the predetermined value for each of the at least one scanning distance being the same or different, and the first scanning area may be scanned at the at least one scanning distance based on the completion condition until the number of MV candidates acquired from the first scanning area reaches a predetermined maximum value.

[0265] In some embodiments, the at least one scanning area may include a first scanning area and a second scanning area, where the first scanning area is scanned using a plurality of first scanning distances and the second scanning area is scanned using a plurality of second scanning distances. For example, as shown in FIG. 8 , the first scanning area may be the left scanning area, and the plurality of first scanning distances may include distance 1, distance 2, and distance 3. The first scanning distance may be one of distance 1, distance 2, or distance 3 used for scanning within the left scanning area. Additionally, the second scanning area may be the top / upper scanning area, and the plurality of scanning distances may include distance 1, distance 2, and distance 3. The second scanning distance may be one of distance 1, distance 2, or distance 3 used for scanning within the top / upper scanning area.

[0266] As discussed above, in some embodiments, a scanning area to the left of the current block, i.e., an area with a "distance of 2" within the left side, indicates that a candidate neighboring block located within this area is two blocks away from the left side of the current block along the vertical direction to the left. Further, an area to the left of the current block with a "distance of 1" indicates that a candidate neighboring block located within this area is one block away from the left side of the current block along the vertical direction to the left, and an area to the left of the current block with a "distance of 3" indicates that a candidate neighboring block located within this area is three blocks away from the left side of the current block along the vertical direction to the left.

[0267] In some embodiments, an area with "distance 1" in the top / top scanning area of ​​the current block, i.e., on the left side, indicates that the candidate neighboring blocks located within this area are one block away from the top side of the current block along the vertical direction to the top. Further, an area with "distance 2" to the left of the current block indicates that the candidate neighboring blocks located within this area are two blocks away from the top side of the current block along the vertical direction to the top, and an area with "distance 3" on the top side of the current block indicates that the candidate neighboring blocks located within this area are three blocks away from the top side of the current block along the vertical direction to the top.

[0268] Further, the completion condition may include determining that the number of MV candidates acquired by scanning the first scanning area reaches a predetermined second value or determining that the number of MV candidates acquired by scanning the second scanning area reaches a predetermined third value, where the predetermined second value is the same as or different from the predetermined third value. Based on these completion conditions, the first scanning area may be scanned at a plurality of first scanning distances, and the second scanning area may be scanned at a plurality of second scanning distances, until the number of MV candidates acquired by scanning the first scanning area and the second scanning area reaches a predetermined maximum value.

[0269] For example, as shown in FIG. 8, the first scanning area may be the left scanning area and the second scanning area may be the top scanning area.

[0270] In step 2304, the processor 2120 may determine one or more CPMVs for the current block based on the one or more candidate MVs.

[0271] FIG. 24 is a flowchart illustrating a method for video decoding, according to an embodiment of the present disclosure.

[0272] In step 2401, the processor 2120 on the decoder side may obtain one or more first parameters based on one or more first neighboring blocks of the current block.

[0273] In step 2402, the processor 2120 may obtain one or more second parameters based on one or more first neighboring blocks and / or one or more second neighboring blocks of the current block.

[0274] In some embodiments, the one or more first neighboring blocks and the one or more second neighboring blocks may be obtained from a plurality of neighboring blocks to the current block based on at least one scanning area and at least one scanning distance, one of the at least one scanning distance indicating the number of blocks away from one side of the current block.

[0275] In some embodiments, the one or more first neighboring blocks and the one or more second neighboring blocks may be obtained by exclusively scanning at least one scanning area at at least one scanning distance.

[0276] For example, as shown in FIG. 18, Neighbor 1 is one first neighbor block obtained by scanning the left scanning area of ​​the current block according to the derivation process for affine inherited merge candidates, and one or more such first neighbor blocks may be obtained by exclusively scanning the left scanning area and / or the top scanning area at different scanning distances.

[0277] Further, as shown in FIG. 18, Neighbor 2 is one second neighboring block obtained by scanning the left scanning area of ​​the current block according to the derivation process for the affine constructed merge candidate, and one or more such second neighboring blocks may be obtained by exclusively scanning the left scanning area and / or the top scanning area at different scanning distances.

[0278] In some embodiments, the at least one scanning area may include a first scanning area, a second scanning area, and a third scanning area, where the first scanning area is determined according to a first maximum scanning distance indicating the maximum number of blocks away from the left side of the current block, the second scanning area is determined according to a second maximum scanning distance indicating the maximum number of blocks away from the top side of the current block, the first maximum scanning distance being the same as or different from the second maximum scanning distance, and the third scanning area is located to the lower right of the current block and includes areas adjacent to and not adjacent to the current block.

[0279] For example, as shown in FIG. 18, the first scanning area may be the left scanning area of ​​the current block, the second scanning area may be the top scanning area of ​​the current block, and the third scanning area may be the adjacent and non-adjacent areas at the bottom right of the current block.

[0280] In some embodiments, the processor 2120 may obtain one or more co-located temporal neighboring blocks from the third scanning area, and in response to determining that the one or more co-located temporal neighboring blocks are used to generate one or more affine constructed neighboring blocks, may determine not to scan the third scanning area to obtain the first neighboring block or the second neighboring block.

[0281] In some embodiments, the processor 2120 may obtain one or more co-located temporal neighboring blocks from the third scanning area and may determine to scan the third scanning area to obtain the first neighboring blocks or the second neighboring blocks in response to determining that the one or more co-located temporal neighboring blocks are used to generate one or more affine constructed neighboring blocks. Further, the processor 2120 may determine not to scan the third scanning area to obtain the first neighboring blocks or the second neighboring blocks in response to determining that the one or more co-located temporal neighboring blocks are not used to generate one or more affine constructed neighboring blocks.

[0282] In some embodiments, the one or more first neighboring blocks may include a predetermined first number of first neighboring blocks, e.g., the first one or two neighboring blocks. In some embodiments, the one or more second neighboring blocks may include a predetermined second number of second neighboring blocks, e.g., the first three or four neighboring blocks. The predetermined first number may be the same as or different from the predetermined second number.

[0283] In some embodiments, the processor 2120 may determine that a predetermined first number of first neighboring blocks are valid in response to determining that a predetermined first number of first neighboring blocks use the same reference picture for at least one motion direction, and may determine that a predetermined second number of second neighboring blocks are valid in response to determining that a predetermined second number of second neighboring blocks use the same reference picture for at least one motion direction.

[0284] In some embodiments, the processor 2120 may obtain one or more first neighboring blocks from a plurality of proximal neighboring blocks and a plurality of non-proximal neighboring blocks, where the proximal neighboring blocks are proximal to the current block and the non-proximal neighboring blocks are each located a number of blocks away from one side of the current block. Further, the processor 2120 may obtain one or more second neighboring blocks from the plurality of proximal neighboring blocks and a plurality of non-proximal neighboring blocks.

[0285] In some embodiments, in response to determining that neighboring block motion information from multiple proximal neighboring blocks and multiple non-proximal neighboring blocks is not used to derive an affine merge or AMVP candidate, processor 2120 may determine that the neighboring block is one of the one or more second neighboring blocks. Further, in response to determining that a co-located temporal motion vector prediction (TMVP) candidate is available in a third scanning area or that a co-located TMVP candidate is used to derive an affine merge or AMVP candidate, processor 2120 may determine that a co-located TMVP candidate is one of the one or more second neighboring blocks.

[0286] In some embodiments, processor 2120 may obtain adjusted positions for the one or more second neighboring blocks by applying a coordinate offset to the one or more second neighboring blocks, and may obtain one or more second parameters based on the adjusted positions of the one or more second neighboring blocks. For example, to provide slightly diversified motion information for constructing new candidates, a small coordinate offset (e.g., +1 or +2 or −1 or −2 to the vertical and / or horizontal coordinates) may be applied when determining the position of Neighbor 2 as shown in FIG. 18 .

[0287] In step 2403, the processor 2120 may construct one or more affine models using the one or more first parameters and the one or more second parameters.

[0288] In step 2404, the processor 2120 may obtain one or more CPMVs for the current block based on one or more affine models.

[0289] FIG. 25 is a flow chart illustrating a method for video encoding that corresponds to the method for video decoding as shown in FIG.

[0290] In step 2501, the processor 2120 on the encoder side may determine one or more first parameters based on one or more first neighboring blocks of the current block.

[0291] In step 2502, the processor 2120 may determine one or more second parameters based on one or more first neighboring blocks and / or one or more second neighboring blocks of the current block.

[0292] In some embodiments, one or more first neighboring blocks and one or more second neighboring blocks may be determined from a plurality of neighboring blocks to the current block based on at least one scanning area and at least one scanning distance, one of the at least one scanning distance indicating a number of blocks away from one side of the current block.

[0293] In some embodiments, the one or more first neighboring blocks and the one or more second neighboring blocks may be determined by exclusively scanning at least one scanning area at at least one scanning distance.

[0294] For example, as shown in FIG. 18, Neighbor 1 is one first neighbor block determined by scanning the left scanning area of ​​the current block according to the derivation process for affine inherited merge candidates, and one or more such first neighbor blocks may be determined by exclusively scanning the left scanning area and / or the top scanning area at different scanning distances.

[0295] Further, as shown in FIG. 18, Neighbor 2 is one second neighbor block determined by scanning the left scanning area of ​​the current block according to the derivation process for the affine constructed merge candidate, and one or more such second neighbor blocks may be determined by exclusively scanning the left scanning area and / or the top scanning area at different scanning distances.

[0296] In some embodiments, the at least one scanning area may include a first scanning area, a second scanning area, and a third scanning area, where the first scanning area is determined according to a first maximum scanning distance indicating the maximum number of blocks away from the left side of the current block, the second scanning area is determined according to a second maximum scanning distance indicating the maximum number of blocks away from the top side of the current block, the first maximum scanning distance being the same as or different from the second maximum scanning distance, and the third scanning area is located to the lower right of the current block and includes areas adjacent to and not adjacent to the current block.

[0297] For example, as shown in FIG. 18, the first scanning area may be the left scanning area of ​​the current block, the second scanning area may be the top scanning area of ​​the current block, and the third scanning area may be the adjacent and non-adjacent areas at the bottom right of the current block.

[0298] In some embodiments, the processor 2120 may determine one or more co-located temporal neighboring blocks from the third scanning area, and in response to determining that the one or more co-located temporal neighboring blocks are used to generate the one or more affine constructed neighboring blocks, may determine not to scan the third scanning area to obtain the first neighboring block or the second neighboring block. The one or more affine constructed neighboring blocks may include neighboring blocks derived based on a non-neighborhood-based derivation process for the affine constructed merge candidate, and may also include neighboring blocks derived based on a neighborhood-based derivation process.

[0299] In some embodiments, the processor 2120 may determine one or more co-located temporal neighboring blocks from the third scanning area, and may determine to scan the third scanning area to determine the first neighboring blocks or the second neighboring blocks in response to determining that the one or more co-located temporal neighboring blocks are used to generate one or more affine constructed neighboring blocks. Further, the processor 2120 may determine not to scan the third scanning area to obtain the first neighboring blocks or the second neighboring blocks in response to determining that the one or more co-located temporal neighboring blocks are not used to generate one or more affine constructed neighboring blocks.

[0300] In some embodiments, the one or more first neighboring blocks may include a predetermined first number of first neighboring blocks, e.g., the first one or two neighboring blocks. In some embodiments, the one or more second neighboring blocks may include a predetermined second number of second neighboring blocks, e.g., the first three or four neighboring blocks. The predetermined first number may be the same as or different from the predetermined second number.

[0301] In some embodiments, the processor 2120 may determine that a predetermined first number of first neighboring blocks are valid in response to determining that a predetermined first number of first neighboring blocks use the same reference picture for at least one motion direction, and may determine that a predetermined second number of second neighboring blocks are valid in response to determining that a predetermined second number of second neighboring blocks use the same reference picture for at least one motion direction.

[0302] In some embodiments, the processor 2120 may determine one or more first neighboring blocks from a plurality of proximal neighboring blocks and a plurality of non-proximal neighboring blocks, where the proximal neighboring blocks are proximal to the current block and the non-proximal neighboring blocks are each located a number of blocks away from one side of the current block. Further, the processor 2120 may determine one or more second neighboring blocks from the plurality of proximal neighboring blocks and a plurality of non-proximal neighboring blocks.

[0303] In some embodiments, processor 2120 may determine that the neighboring block is one of the one or more second neighboring blocks in response to determining that neighboring block motion information from multiple proximal neighboring blocks and multiple non-proximal neighboring blocks is not used to derive an affine merge or AMVP candidate. Further, processor 2120 may determine that a co-located temporal motion vector prediction (TMVP) candidate is available in a third scanning area or that a co-located TMVP candidate is used to derive an affine merge or AMVP candidate in response to determining that a co-located TMVP candidate is available in a third scanning area or that a co-located TMVP candidate is used to derive an affine merge or AMVP candidate.

[0304] In some embodiments, processor 2120 may determine adjusted positions for one or more second neighboring blocks by applying a coordinate offset to the one or more second neighboring blocks, and may obtain one or more second parameters based on the adjusted positions of the one or more second neighboring blocks. For example, to provide slightly diversified motion information for constructing new candidates, a small coordinate offset (e.g., +1 or +2 or −1 or −2 to the vertical and / or horizontal coordinates) may be applied when determining the position of Neighbor 2 as shown in FIG. 18.

[0305] In step 2503, the processor 2120 may construct one or more affine models using the one or more first parameters and the one or more second parameters.

[0306] In step 2504, the processor 2120 may determine one or more CPMVs for the current block based on one or more affine models.

[0307] FIG. 26 is a flowchart illustrating a method for video decoding, according to an embodiment of the present disclosure.

[0308] In step 2601, the processor 2120 on the decoder side may obtain one or more MV candidates from one or more candidate lists according to a predetermined order, where the one or more candidate lists may include an AMVP candidate list, a regular merge candidate list, and an affine merge candidate list, and the one or more MV candidates are from multiple neighboring blocks to the current block.

[0309] In some embodiments, the one or more MV candidates may include one or more MV candidates inherited from nearby spatial neighboring blocks and one or more MV candidates constructed from nearby spatial neighboring blocks, one or more translational MV candidates from nearby spatial neighboring blocks and one or more temporal MV candidates from nearby temporal neighboring blocks, and one or more MV candidates inherited from non-neighboring spatial neighboring blocks. Processor 2120 may further obtain, from the candidate list, one or more MV candidates constructed from nearby spatial neighboring blocks after the one or more MV candidates inherited from nearby spatial neighboring blocks, obtain, from the candidate list, one or more translational MV candidates from nearby spatial neighboring blocks after the one or more MV candidates constructed from nearby spatial neighboring blocks, obtain, from the candidate list, one or more temporal MV candidates from nearby temporal neighboring blocks after the one or more translational MV candidates from nearby spatial neighboring blocks, and obtain, from the candidate list, one or more MV candidates inherited from non-neighboring spatial neighboring blocks after the one or more temporal MV candidates from nearby temporal neighboring blocks. For example, a newly derived candidate can be inserted into the affine AMVP candidate list by adhering to the following order: 1. inherited from close spatial neighbors, 2. constructed from close spatial neighbors, 3. translational MVs from close spatial neighbors, 4. temporal MVs from close temporal neighbors, 5. inherited from non-close spatial neighbors, 6. zero MVs.

[0310] In some embodiments, the one or more MV candidates may include sub-block-based temporal motion vector prediction (SbTMVP) candidates, one or more MV candidates inherited from nearby neighboring blocks and one or more MV candidates inherited from non-close neighboring blocks, one or more MV candidates constructed from nearby neighboring blocks and one or more MV candidates constructed from non-close neighboring blocks, and one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks. The processor 2120 may further obtain from the candidate list one or more MV candidates inherited from nearby neighboring blocks after the SbTMVP candidate, may obtain from the candidate list one or more MV candidates inherited from non-neighboring neighboring blocks after one or more MV candidates inherited from nearby neighboring blocks, may obtain from the candidate list one or more MV candidates constructed from nearby neighboring blocks after one or more MV candidates inherited from non-neighboring neighboring blocks, may obtain from the candidate list one or more MV candidates constructed from non-neighboring neighboring blocks after one or more MV candidates constructed from nearby neighboring blocks, and may obtain from the candidate list one or more MV candidates constructed from inherited translated neighboring blocks and non-translated neighboring blocks after one or more MV candidates constructed from non-neighboring neighboring blocks. For example, newly derived candidates can be inserted into the affine AMVP candidate list by adhering to the following order: 1. SbTMVP candidate if available, 2. inherited from close neighbors, 3. inherited from non-close neighbors, 4. constructed from close neighbors, 5. constructed from non-close neighbors, 6. constructed from inherited translated and non-translated neighbors, 7. zero MV if the list is still not full.

[0311] In some embodiments, the one or more MV candidates may include one or more MV candidates including sub-block-based temporal motion vector prediction (SbTMVP) candidates, one or more MV candidates inherited from nearby neighboring blocks and one or more MV candidates inherited from non-close neighboring blocks, a first MV candidate constructed from nearby neighboring blocks and one or more MV candidates constructed from non-close neighboring blocks, and one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks, including a first number of MV candidates and a second number of MV candidates. The processor 2120 may further obtain, from the candidate list, one or more MV candidates inherited from nearby neighboring blocks after the SbTMVP candidate, may obtain, from the candidate list, one or more MV candidates inherited from non-neighboring neighboring blocks after the one or more MV candidates inherited from the nearby neighboring blocks, may obtain, from the candidate list, a first MV candidate constructed from nearby neighboring blocks after the one or more MV candidates inherited from the non-neighboring neighboring blocks, may obtain, from the candidate list, a first number of MV candidates constructed from the inherited translated and non-translated neighboring blocks after the first MV candidate constructed from the nearby neighboring blocks, may obtain, from the candidate list, one or more MV candidates constructed from non-neighboring neighboring blocks after the first number of MV candidates constructed from the inherited translated and non-translated neighboring blocks, and may obtain, from the candidate list, a second number of MV candidates after the one or more MV candidates constructed from the non-neighboring neighboring blocks.

[0312] In some embodiments, the first number and the second number may be determined in at least one of the following ways: predefining the first number as a first fixed value; predefining the second number as a second fixed value, where the first fixed value is the same as or different from the second fixed value; receiving the first number or the second number by a decoder signaled at any level; configuring the first number or the second number by a decoder or an encoder; or dynamically determining the first number and the second number according to the number of available neighboring blocks in the left and top areas of the current block.

[0313] In step 2602, the processor 2120 may obtain one or more CPMVs for the current block based on the one or more candidate MVs.

[0314] FIG. 27 is a flow chart illustrating a method for video encoding that corresponds to the method for video decoding as shown in FIG.

[0315] In step 2701, the processor 2120 on the encoder side may determine one or more MV candidates from one or more candidate lists according to a predetermined order, where the one or more candidate lists may include an AMVP candidate list, a regular merge candidate list, and an affine merge candidate list, and the one or more MV candidates are from multiple neighboring blocks to the current block.

[0316] In some embodiments, the one or more MV candidates may include one or more MV candidates inherited from nearby spatial neighboring blocks and one or more MV candidates constructed from nearby spatial neighboring blocks, one or more translational MV candidates from nearby spatial neighboring blocks and one or more temporal MV candidates from nearby temporal neighboring blocks, and one or more MV candidates inherited from non-neighboring spatial neighboring blocks. Processor 2120 may further insert into the candidate list one or more MV candidates constructed from nearby spatial neighboring blocks after the one or more MV candidates inherited from nearby spatial neighboring blocks, may insert into the candidate list one or more translational MV candidates from nearby spatial neighboring blocks after the one or more MV candidates constructed from nearby spatial neighboring blocks, may insert into the candidate list one or more temporal MV candidates from nearby temporal neighboring blocks after the one or more translational MV candidates from nearby spatial neighboring blocks, and may insert into the candidate list one or more MV candidates inherited from non-neighboring spatial neighboring blocks after the one or more temporal MV candidates from nearby temporal neighboring blocks. For example, a newly derived candidate can be inserted into the affine AMVP candidate list by adhering to the following order: 1. inherited from close spatial neighbors, 2. constructed from close spatial neighbors, 3. translational MVs from close spatial neighbors, 4. temporal MVs from close temporal neighbors, 5. inherited from non-close spatial neighbors, 6. zero MVs.

[0317] In some embodiments, the one or more MV candidates may include sub-block-based temporal motion vector prediction (SbTMVP) candidates, one or more MV candidates inherited from nearby neighboring blocks and one or more MV candidates inherited from non-close neighboring blocks, one or more MV candidates constructed from nearby neighboring blocks and one or more MV candidates constructed from non-close neighboring blocks, and one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks. The processor 2120 may further insert one or more MV candidates inherited from nearby neighboring blocks after the SbTMVP candidate, may insert one or more MV candidates inherited from non-neighboring neighboring blocks after the one or more MV candidates inherited from nearby neighboring blocks into the candidate list, may insert one or more MV candidates constructed from nearby neighboring blocks after the one or more MV candidates inherited from non-neighboring neighboring blocks into the candidate list, may insert one or more MV candidates constructed from non-neighboring neighboring blocks after the one or more MV candidates constructed from nearby neighboring blocks into the candidate list, and may insert one or more MV candidates constructed from inherited translated neighboring blocks and non-translated neighboring blocks after the one or more MV candidates constructed from non-neighboring neighboring blocks into the candidate list. For example, newly derived candidates can be inserted into the affine AMVP candidate list by adhering to the following order: 1. SbTMVP candidate if available, 2. inherited from close neighbors, 3. inherited from non-close neighbors, 4. constructed from close neighbors, 5. constructed from non-close neighbors, 6. constructed from inherited translated and non-translated neighbors, 7. zero MV if the list is still not full.

[0318] In some embodiments, the one or more MV candidates may include one or more MV candidates including sub-block-based temporal motion vector prediction (SbTMVP) candidates, one or more MV candidates inherited from nearby neighboring blocks and one or more MV candidates inherited from non-close neighboring blocks, a first MV candidate constructed from nearby neighboring blocks and one or more MV candidates constructed from non-close neighboring blocks, and one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks, including a first number of MV candidates and a second number of MV candidates. The processor 2120 may further insert into the candidate list one or more MV candidates inherited from nearby neighboring blocks after the SbTMVP candidate, may insert into the candidate list one or more MV candidates inherited from non-neighboring neighboring blocks after the one or more MV candidates inherited from nearby neighboring blocks, may insert into the candidate list a first MV candidate constructed from nearby neighboring blocks after the one or more MV candidates inherited from non-neighboring neighboring blocks, may insert into the candidate list a first number of MV candidates constructed from the inherited translated and non-translated neighboring blocks after the first MV candidate constructed from the nearby neighboring blocks, may insert into the candidate list one or more MV candidates constructed from non-neighboring neighboring blocks after the first number of MV candidates constructed from the inherited translated and non-translated neighboring blocks, and may insert into the candidate list a second number of MV candidates after the one or more MV candidates constructed from the non-neighboring neighboring blocks.

[0319] In some embodiments, the first number and the second number may be determined in at least one of the following ways: predefining the first number as a first fixed value; predefining the second number as a second fixed value, where the first fixed value is the same as or different from the second fixed value; receiving the first number or the second number by a decoder signaled at any level; configuring the first number or the second number by a decoder or an encoder; or dynamically determining the first number and the second number according to the number of available neighboring blocks in the left and top areas of the current block.

[0320] In step 2702, the processor 2120 may determine one or more CPMVs for the current block based on one or more candidate MVs.

[0321] FIG. 28 is a flowchart illustrating a method for video decoding, according to an embodiment of the present disclosure.

[0322] In step 2801, the processor 2120 on the decoder side may obtain a temporal candidate list having a first list size, where the first list size is larger than the list size of any existing candidate list, including an affine merge candidate list, an AMVP candidate list, or a regular merge candidate list, and the temporal candidate list may include multiple MV candidates obtained from multiple neighboring blocks to the current block.

[0323] In some embodiments, a temporal candidate list may be created for a corresponding existing candidate list, which may include an affine merge candidate list, an AMVP candidate list, or a regular merge candidate list. In some embodiments, one temporal candidate list may be created for each existing candidate list. For example, a first temporal candidate list may be created for the affine merge candidate list, a second temporal candidate list may be created for the AMVP candidate list, and a third temporal candidate list may be created for the regular merge candidate list.

[0324] In some embodiments, the processor 2120 may apply adaptive reordering to multiple MV candidates in a temporal candidate list and / or one or more MV candidates in one or more existing candidate lists.

[0325] In some embodiments, in response to determining that a temporal candidate list is to be created for one existing candidate list, processor 2120 may apply adaptive reordering to the MV candidates in the temporal candidate list and / or to one or more MV candidates in the corresponding existing candidate list. Further, in response to determining that a temporal candidate list is not to be created for one existing candidate list, processor 2120 may apply adaptive reordering to the one or more MV candidates in the corresponding existing candidate list.

[0326] For example, for each existing candidate list, one temporal list may or may not be created. If a temporal candidate list is created, adaptive reordering may be applied to the temporal candidate list or / and the existing list. If a temporal candidate list is not created, adaptive reordering is applied only to the existing candidate list.

[0327] In step 2802, the processor 2120 may obtain a first number of MV candidates from the temporal candidate list based on the reordered plurality of MV candidates, the first number being less than the number of the plurality of MV candidates in the temporal candidate list.

[0328] In some embodiments, the first number of MV candidates may be obtained from a plurality of MV candidates adaptively reordered based on a TM cost between the predicted sample of the template of the current block and the corresponding reference sample. Further, in response to determining that fractional motion information is used for the current block, processor 2120 may use an interpolation filter with fewer taps than a predetermined number of taps to generate the predicted sample, or may skip interpolation filtering and generating the predicted sample based on the nearest integer sample in response to determining that fractional motion information is used for the current block.

[0329] In some embodiments, processor 2120 may apply interpolation filtering to the corresponding reference sample in response to determining that the corresponding reference sample is located at a fractional position.

[0330] In some embodiments, processor 2120 may adjust the TM cost by removing one or more bits from the TM cost. For example, to reduce the impact of noise-level cost differences, the calculated matching cost may be adjusted by removing a few least significant bits before the candidate sorting process.

[0331] In some embodiments, in response to determining that the temporal candidate list includes more than one zero MV candidate, the processor 2120 may assign a predetermined fixed value to the zero MV candidates in the temporal candidate list except the first zero MV candidate, such that the zero MV candidates except the first zero MV candidate are placed at the end of the temporal candidate list and / or the corresponding existing candidate list, the predetermined fixed value being greater than a threshold value.

[0332] In some embodiments, in response to determining that the existing candidate list includes more than one zero MV candidate, the processor 2120 may assign predetermined fixed values ​​to the zero MV candidates in the existing candidate list except for the first zero MV candidate, such that the zero candidates except for the first zero MV candidate are placed at the end of the existing candidate list.

[0333] In some embodiments, in response to determining that the temporal candidate list includes at least one zero MV candidate, the processor 2120 may assign a predetermined fixed value to the at least one zero MV candidate so as to place all of the at least one zero MV candidate at the end of the temporal candidate and / or corresponding existing candidate list, the predetermined fixed value being greater than a threshold value.

[0334] In some embodiments, in response to determining that the existing candidate list includes at least one zero MV candidate, the processor 2120 may assign a predetermined fixed value to the at least one zero MV candidate so as to place all of the at least one zero MV candidate at the end of the existing candidate list.

[0335] In some embodiments, in response to determining that one MV candidate or one type of MV candidate will not participate in adaptive reordering, the processor 2120 may skip applying adaptive reordering to the remaining MV candidates in the temporal candidate list and / or corresponding existing candidate list.

[0336] In some embodiments, in response to determining that one MV candidate or one type of MV candidate will not participate in adaptive reordering, the processor 2120 may skip deriving the MV candidate or type of MV candidate and / or may skip applying adaptive reordering to the MV candidate or type of MV candidate.

[0337] In some embodiments, the processor 2120 may skip deriving the MV candidate or MV candidate type and may skip calculating the TM cost for the MV candidate or MV candidate type such that the candidate content of the MV candidate or MV candidate type is invalid and one or more positions of the MV candidate or MV candidate type remain in the temporal candidate list.

[0338] FIG. 29 is a flow chart illustrating a method for video encoding that corresponds to the method for video decoding as shown in FIG.

[0339] In step 2901, the processor 2120 on the encoder side may determine a temporal candidate list having a first list size, where the first list size is larger than the list size of any existing candidate list, including an affine merge candidate list, an AMVP candidate list, or a regular merge candidate list, and the temporal candidate list may include multiple MV candidates obtained from multiple neighboring blocks to the current block.

[0340] In some embodiments, a temporal candidate list may be created for a corresponding existing candidate list, which may include an affine merge candidate list, an AMVP candidate list, or a regular merge candidate list. In some embodiments, one temporal candidate list may be created for each existing candidate list. For example, a first temporal candidate list may be created for the affine merge candidate list, a second temporal candidate list may be created for the AMVP candidate list, and a third temporal candidate list may be created for the regular merge candidate list.

[0341] In some embodiments, the processor 2120 may apply adaptive reordering to multiple MV candidates in a temporal candidate list and / or one or more MV candidates in one or more existing candidate lists.

[0342] In some embodiments, in response to determining that a temporal candidate list is to be created for one existing candidate list, processor 2120 may apply adaptive reordering to the MV candidates in the temporal candidate list and / or to one or more MV candidates in the corresponding existing candidate list. Further, in response to determining that a temporal candidate list is not to be created for one existing candidate list, processor 2120 may apply adaptive reordering to the one or more MV candidates in the corresponding existing candidate list.

[0343] For example, for each existing candidate list, one temporal list may or may not be created. If a temporal candidate list is created, adaptive reordering may be applied to the temporal candidate list or / and the existing list. If a temporal candidate list is not created, adaptive reordering is applied only to the existing candidate list.

[0344] In step 2802, the processor 2120 may determine a first number of MV candidates from the temporal candidate list based on the reordered plurality of MV candidates, the first number being less than the number of the plurality of MV candidates in the temporal candidate list.

[0345] In some embodiments, the first number of MV candidates may be obtained from a plurality of MV candidates adaptively reordered based on a TM cost between the predicted sample of the template of the current block and the corresponding reference sample. Further, in response to determining that fractional motion information is used for the current block, processor 2120 may use an interpolation filter with fewer taps than a predetermined number of taps to generate the predicted sample, or may skip interpolation filtering and generating the predicted sample based on the nearest integer sample in response to determining that fractional motion information is used for the current block.

[0346] In some embodiments, processor 2120 may apply interpolation filtering to the corresponding reference sample in response to determining that the corresponding reference sample is located at a fractional position.

[0347] In some embodiments, processor 2120 may adjust the TM cost by removing one or more bits from the TM cost. For example, to reduce the impact of noise-level cost differences, the calculated matching cost may be adjusted by removing a few least significant bits before the candidate sorting process.

[0348] In some embodiments, in response to determining that the temporal candidate list includes more than one zero MV candidate, the processor 2120 may assign a predetermined fixed value to the zero MV candidates in the temporal candidate list except the first zero MV candidate, such that the zero MV candidates except the first zero MV candidate are placed at the end of the temporal candidate list and / or the corresponding existing candidate list, the predetermined fixed value being greater than a threshold value.

[0349] In some embodiments, in response to determining that the existing candidate list includes more than one zero MV candidate, the processor 2120 may assign predetermined fixed values ​​to the zero MV candidates in the existing candidate list except for the first zero MV candidate, such that the zero candidates except for the first zero MV candidate are placed at the end of the existing candidate list.

[0350] In some embodiments, in response to determining that the temporal candidate list includes at least one zero MV candidate, the processor 2120 may assign a predetermined fixed value to the at least one zero MV candidate so as to place all of the at least one zero MV candidate at the end of the temporal candidate and / or corresponding existing candidate list, the predetermined fixed value being greater than a threshold value.

[0351] In some embodiments, in response to determining that the existing candidate list includes at least one zero MV candidate, the processor 2120 may assign a predetermined fixed value to the at least one zero MV candidate so as to place all of the at least one zero MV candidate at the end of the existing candidate list.

[0352] In some embodiments, in response to determining that one MV candidate or one type of MV candidate will not participate in adaptive reordering, the processor 2120 may skip applying adaptive reordering to the remaining MV candidates in the temporal candidate list and / or corresponding existing candidate list.

[0353] In some embodiments, in response to determining that one MV candidate or one type of MV candidate will not participate in adaptive reordering, the processor 2120 may skip deriving the MV candidate or type of MV candidate and / or may skip applying adaptive reordering to the MV candidate or type of MV candidate.

[0354] In some embodiments, the processor 2120 may skip deriving the MV candidate or MV candidate type and may skip calculating the TM cost for the MV candidate or MV candidate type such that the candidate content of the MV candidate or MV candidate type is invalid and one or more positions of the MV candidate or MV candidate type remain in the temporal candidate list.

[0355] In some embodiments, an apparatus for video coding is provided, the apparatus including a processor 2120 and a memory 2140 configured to store instructions executable by the processor, the instructions, when executed, configured to perform any of the methods illustrated in Figures 22-29.

[0356] In some other embodiments, a non-transitory computer-readable storage medium having instructions stored thereon is provided. When the instructions are executed by a processor 2120, the instructions cause the processor to perform any of the methods illustrated in Figures 22-29. In one embodiment, a plurality of programs may be executed by the processor 2120 in the computing environment 2110 to receive (e.g., from the video encoder 20 in Figure 2) a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements), and to perform the decoding method described above in accordance with the received bitstream or data stream. In another embodiment, programs may be executed by processor 2120 in computing environment 2110 to perform the encoding method described above for encoding video information (e.g., video blocks representing video frames and / or one or more associated syntax elements) into a bitstream or data stream, and may also be executed by processor 2120 in computing environment 2110 to transmit the bitstream or data stream (e.g., to video decoder 30 in FIG. 3 ). Alternatively, a non-transitory computer-readable storage medium may store thereon a bitstream or data stream including encoded video information (e.g., video blocks representing encoded video frames and / or one or more associated syntax elements) generated by an encoder (e.g., video encoder 20 in FIG. 2 ) using, for example, the encoding method described above, for use by a decoder (e.g., video decoder 30 in FIG. 3 ) in decoding video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, or an optical data storage device.

[0357] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. This application is intended to cover any variations, uses, or adaptations of the disclosure that adhere to its general principles, including such departures from the disclosure, as come within known or customary practice in the art. It is intended that the specification and examples be considered as illustrative only.

[0358] It will be appreciated that the present disclosure is not limited to the exact embodiments described above and illustrated in the accompanying drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. 1. A method for video decoding, comprising: Obtaining one or more motion vector (MV) candidates from a plurality of non-close neighboring blocks to a current block based on at least one scanning area and at least one scanning distance, where one of the at least one scanning distance indicates a number of blocks away from one side of the current block; determining a termination condition based on a number of MV candidates obtained by scanning the at least one scanning distance within a first scanning area, wherein the at least one scanning area includes the first scanning area; responsive to determining that a completion condition is met, ceasing scanning of the at least one scanning area; and obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more MV candidates; A method comprising:

2. The completion condition is determining that the number of the MV candidates obtained by scanning a first scanning distance within the first scanning area reaches a predetermined first value; determining that the number of MV candidates obtained by scanning the at least one scanning distance within the first scanning area reaches a second predetermined value; or determining that the number of MV candidates obtained by scanning the first scanning distance in the first scanning area reaches the predetermined first value, and that the number of MV candidates obtained by scanning the at least one scanning distance in the first scanning area reaches the predetermined second value; the predetermined first value is the same as or different from the predetermined second value; The method of claim 1.

3. The completion condition is determining that the number of MV candidates obtained by scanning each of the at least one scanning distance reaches a predetermined value, wherein the predetermined value for each of the at least one scanning distance is the same or different; The method comprises: and scanning the first scanning area at the at least one scanning distance based on the completion condition until the number of the MV candidates obtained from the first scanning area reaches a predetermined maximum value. The method of claim 1.

4. the at least one scanning area includes the first scanning area and a second scanning area, the first scanning area is scanned using a plurality of first scanning distances, and the second scanning area is scanned using a plurality of second scanning distances; The completion condition is determining that the number of MV candidates obtained by scanning the first scanning area reaches a second predetermined value; or determining that the number of MV candidates obtained by scanning the second scanning area reaches a predetermined third value, wherein the predetermined second value is the same as or different from the predetermined third value; The method comprises: and scanning the first scanning area at the plurality of first scanning distances and scanning the second scanning area at the plurality of second scanning distances based on the completion condition until the number of the MV candidates obtained by scanning the first scanning area and the second area reaches a predetermined maximum value. The method of claim 1.

5. 1. A method of video decoding, comprising: Obtaining one or more first parameters based on one or more first neighboring blocks of the current block; obtaining one or more second parameters based on the one or more first neighboring blocks and / or one or more second neighboring blocks of the current block; constructing one or more affine models using the one or more first parameters and the one or more second parameters; obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more affine models; the one or more first neighboring blocks and the one or more second neighboring blocks are obtained from a plurality of neighboring blocks to the current block based on at least one scanning area and at least one scanning distance, one of the at least one scanning distance indicating a number of blocks away from one side of the current block; the one or more first neighboring blocks and the one or more second neighboring blocks are obtained by exclusively scanning the at least one scanning area in the at least one scanning distance; method.

6. the at least one scanning area includes a first scanning area, a second scanning area, and a third scanning area; the first scanning area is determined according to a first maximum scanning distance indicating a maximum number of blocks away from the left side of the current block; the second scanning area is determined according to a second maximum scanning distance indicating a maximum number of blocks away from a top side of the current block, and the first maximum scanning distance is the same as or different from the second maximum scanning distance; the third scanning area is located to the lower right of the current block and includes areas adjacent to and not adjacent to the current block; The method of claim 5.

7. obtaining one or more co-located temporal neighboring blocks from the third scanning area; determining, in response to determining that the one or more co-located temporal neighboring blocks are used to generate one or more affine constructed neighboring blocks, not to scan the third scanning area to obtain the first neighboring block or the second neighboring block; The method of claim 6 further comprising:

8. obtaining one or more co-located temporal neighboring blocks from the third scanning area; determining, in response to determining that the one or more co-located temporal neighboring blocks are used to generate one or more affine constructed neighboring blocks, to scan the third scanning area to obtain the first neighboring block or the second neighboring block; determining not to scan the third scanning area to obtain the first neighboring block or the second neighboring block in response to determining that the one or more co-located temporal neighboring blocks will not be used to generate one or more affine constructed neighboring blocks; The method of claim 6 further comprising:

9. the one or more first neighboring blocks include a predetermined first number of first neighboring blocks; the one or more second neighboring blocks include a predetermined second number of second neighboring blocks, the predetermined first number being the same as or different from a predetermined second number; The method comprises: determining that the predetermined first number of first neighboring blocks are valid in response to determining that the predetermined first number of first neighboring blocks use the same reference picture for at least one motion direction; determining that the predetermined second number of second neighboring blocks are valid in response to determining that the predetermined second number of second neighboring blocks use the same reference picture for at least one motion direction; and The method of claim 5 further comprising:

10. obtaining the one or more first neighboring blocks from a plurality of adjacent neighboring blocks and a plurality of non-adjacent neighboring blocks, wherein the plurality of adjacent neighboring blocks are adjacent to the current block and each of the plurality of non-adjacent neighboring blocks is located a number of blocks away from one side of the current block; obtaining the one or more second neighboring blocks from the plurality of adjacent neighboring blocks and the plurality of non-adjacent neighboring blocks; The method of claim 5 further comprising:

11. determining that neighboring block motion information from the plurality of proximal neighboring blocks and the plurality of non-proximal neighboring blocks is not used to derive an affine merge or advanced motion vector prediction (AMVP) candidate, the neighboring block being one of the one or more second neighboring blocks; determining that a co-located temporal motion vector prediction (TMVP) candidate within the third scanning area is available or that the co-located TMVP candidate is used to derive an affine merge or AMVP candidate, to be one of the one or more second neighboring blocks; The method of claim 10 further comprising:

12. applying a coordinate offset to the one or more second neighboring blocks to obtain an adjusted position for the one or more second neighboring blocks; obtaining the one or more second parameters based on the adjusted positions of the one or more second neighboring blocks; The method of claim 5 further comprising:

13. 1. A method for video decoding, comprising: Obtaining one or more motion vector (MV) candidates from one or more candidate lists according to a predetermined order, the one or more candidate lists including an affine advanced motion vector prediction (AMVP) candidate list, a regular merge candidate list, and an affine merge candidate list, the one or more MV candidates being from a plurality of neighboring blocks to a current block; obtaining one or more control point motion vectors (CPMVs) for the current block based on the one or more MV candidates; A method comprising:

14. The one or more MV candidates include one or more MV candidates inherited from nearby spatial neighboring blocks and one or more MV candidates constructed from nearby spatial neighboring blocks, one or more translational MV candidates from nearby spatial neighboring blocks and one or more temporal MV candidates from nearby temporal neighboring blocks, and one or more MV candidates inherited from non-neighboring spatial neighboring blocks; The method comprises: Obtaining, from a candidate list, the one or more MV candidates constructed from the nearby spatial neighboring blocks after the one or more MV candidates inherited from the nearby spatial neighboring blocks; Obtaining, from the candidate list, one or more translational MV candidates from the nearby spatial neighboring blocks after the one or more MV candidates constructed from the nearby spatial neighboring blocks; Obtaining, from the candidate list, the one or more temporal MV candidates from the nearby temporal neighboring blocks after the one or more translational MV candidates from the nearby spatial neighboring blocks; Obtaining from the candidate list the one or more MV candidates inherited from the non-neighboring spatial neighboring blocks after the one or more temporal MV candidates from the close temporal neighboring blocks; The method of claim 13 further comprising:

15. the one or more MV candidates include subblock-based temporal motion vector prediction (SbTMVP) candidates, one or more MV candidates inherited from close neighboring blocks and one or more MV candidates inherited from non-close neighboring blocks, one or more MV candidates constructed from close neighboring blocks and one or more MV candidates constructed from non-close neighboring blocks, and one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks; The method comprises: Obtaining, from a candidate list, the one or more MV candidates inherited from the adjacent neighboring blocks after the SbTMVP candidate; Obtaining, from the candidate list, one or more MV candidates inherited from non-adjacent neighboring blocks after the one or more MV candidates inherited from adjacent neighboring blocks; Obtaining from the candidate list one or more MV candidates constructed from adjacent neighboring blocks after the one or more MV candidates inherited from non-adjacent neighboring blocks; Obtaining, from the candidate list, one or more MV candidates constructed from non-adjacent neighboring blocks after the one or more MV candidates constructed from adjacent neighboring blocks; Obtaining from the candidate list one or more MV candidates constructed from inherited translational neighboring blocks and non-translated neighboring blocks after the one or more MV candidates constructed from non-neighboring neighboring blocks; The method of claim 13 further comprising:

16. the one or more MV candidates include a sub-block based temporal motion vector prediction (SbTMVP) candidate, one or more MV candidates inherited from nearby neighboring blocks and one or more MV candidates inherited from non-close neighboring blocks, a first MV candidate constructed from nearby neighboring blocks and one or more MV candidates constructed from non-close neighboring blocks, and one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks, including a first number of MV candidates and a second number of MV candidates; Obtaining, from a candidate list, the one or more MV candidates inherited from the adjacent neighboring blocks after the SbTMVP candidate; Obtaining, from the candidate list, one or more MV candidates inherited from non-adjacent neighboring blocks after the one or more MV candidates inherited from adjacent neighboring blocks; Obtaining the first MV candidate constructed from a nearby neighboring block after the one or more MV candidates inherited from non-close neighboring blocks from the candidate list; obtaining, from the candidate list, the first number of MV candidates constructed from inherited translational neighboring blocks and non-translated neighboring blocks after the first MV candidate constructed from nearby neighboring blocks; obtaining, from the candidate list, the one or more MV candidates constructed from non-adjacent neighboring blocks after the first number of MV candidates constructed from inherited translational neighboring blocks and non-translated neighboring blocks; obtaining the second number of MV candidates after the one or more MV candidates constructed from non-adjacent neighboring blocks from the candidate list; The method of claim 13 further comprising:

17. The first number and the second number are determined in the following manner: predefining the first number as a first fixed value; predefining the second number as a second fixed value, the first fixed value being the same as or different from the second fixed value; receiving, by a decoder, the first number or the second number signaled at any level; constructing the first number or the second number by the decoder or encoder; or dynamically determining the first number and the second number according to the number of available neighboring blocks in a left area and a top area of ​​the current block; The method of claim 16, wherein the at least one of the following is determined:

18. 1. A method for video decoding, comprising: obtaining a temporal candidate list having a first list size, the first list size being larger than a list size of any existing candidate list including an affine merge candidate list, an advanced motion vector prediction (AMVP) candidate list, or a regular merge candidate list, the temporal candidate list including a plurality of motion vector (MV) candidates obtained from a plurality of neighboring blocks to a current block; obtaining a first number of MV candidates from the temporal candidate list based on the reordered MV candidates, the first number being less than the number of MV candidates in the temporal candidate list; A method comprising:

19. 20. The method of claim 18, wherein the temporal candidates are created for a corresponding existing candidate list, the corresponding existing candidate list comprising the affine merge candidate list, the AMVP candidate list, or the regular merge candidate list.

20. 20. The method of claim 18, further comprising applying adaptive reordering to the plurality of MV candidates in the temporal candidate list and / or one or more MV candidates in one or more existing candidate lists.

21. responsive to determining that the temporal candidate list is to be created for an existing candidate list, applying adaptive reordering to the MV candidates in the temporal candidate list and / or the one or more MV candidates in the corresponding existing candidate list; In response to determining that a temporal candidate list is not created for an existing candidate list, applying adaptive reordering to the one or more MV candidates in the corresponding existing candidate list; 21. The method of claim 20, further comprising:

22. The first number of MV candidates are obtained from the plurality of MV candidates adaptively reordered based on a template matching (TM) cost between a predicted sample of a template of the current block and a corresponding reference sample; The method comprises: using an interpolation filter having fewer taps than a predetermined number of taps to generate the predicted samples in response to determining that fractional motion information is to be used for the current block; or responsive to determining that the fractional motion information is used for the current block, skipping interpolation filtering and generating the prediction samples based on nearest integer samples; 21. The method of claim 20, further comprising:

23. 22. The method of claim 21, further comprising applying interpolation filtering to the corresponding reference sample in response to determining that the corresponding reference sample is located at a fractional position.

24. 22. The method of claim 21, further comprising adjusting the TM cost by removing one or more bits from the TM cost.

25. in response to determining that the temporal candidate list includes more than one zero MV candidate, assigning predetermined fixed values ​​to the zero MV candidates in the temporal candidate list, except for a first zero MV candidate, so as to place the zero MV candidates, except for a first zero MV candidate, at the end of the temporal candidate list and / or a corresponding existing candidate list, the predetermined fixed value being greater than a threshold; or in response to determining that the existing candidate list includes more than one zero-MV candidate, assigning the predetermined fixed value to zero-MV candidates in the existing candidate list except for the first zero-MV candidate, such that the zero candidates except for the first zero-MV candidate are placed at the end of the existing candidate list; 20. The method of claim 18, further comprising:

26. in response to determining that the temporal candidate list includes at least one zero-MV candidate, assigning a predetermined fixed value to the at least one zero-MV candidate so as to place all of the at least one zero-MV candidate at the end of the temporal candidate and / or corresponding existing candidate list, the predetermined fixed value being greater than a threshold; or in response to determining that the existing candidate list includes at least one zero-MV candidate, assigning the predetermined fixed value to the at least one zero-MV candidate so as to place all of the at least one zero-MV candidate at the end of the existing candidate list; 20. The method of claim 18, further comprising:

27. The method of claim 20, further comprising: in response to determining that one MV candidate or one type of MV candidate will not participate in the adaptive reordering, skipping applying the adaptive reordering to remaining MV candidates in the temporal candidate list and / or corresponding existing candidate list.

28. The method of claim 20, further comprising: in response to determining that one MV candidate or one type of MV candidate will not participate in the adaptive reordering, skipping deriving the MV candidate or the type of MV candidate and / or skipping applying the adaptive reordering to the MV candidate or the type of MV candidate.

29. skipping deriving the MV candidate or the type of MV candidate such that the MV candidate or the candidate content of the type of MV candidate is invalid and one or more positions of the MV candidate or the type of MV candidate remain in the temporal candidate list; skipping calculating the TM cost for the MV candidate or the type of MV candidate; 30. The method of claim 28, further comprising:

30. 1. A method for video encoding, comprising: determining one or more motion vector (MV) candidates from a plurality of non-adjacent neighboring blocks to the current block based on at least one scanning area and at least one scanning distance, one of the at least one scanning distance indicating a number of blocks away from one side of the current block; determining a completion condition based on a number of MV candidates obtained by scanning the at least one scanning distance within a first scanning area, wherein the at least one scanning area includes the first scanning area; responsive to determining that a completion condition is met, ceasing scanning of the at least one scanning area; and determining one or more control point motion vectors (CPMVs) for the current block based on the one or more candidate MVs; A method comprising:

31. determining that the number of the MV candidates obtained by scanning a first scanning distance within the first scanning area reaches a predetermined first value; determining that the number of MV candidates obtained by scanning the at least one scanning distance within the first scanning area reaches a second predetermined value; or determining that the number of MV candidates obtained by scanning the first scanning distance in the first scanning area reaches the predetermined first value, and that the number of MV candidates obtained by scanning the at least one scanning distance in the first scanning area reaches the predetermined second value; the predetermined first value is the same as or different from the predetermined second value; 31. The method of claim 30.

32. The completion condition is determining that the number of MV candidates obtained by scanning each of the at least one scanning distance reaches a predetermined value, wherein the predetermined value for each of the at least one scanning distance is the same or different; The method comprises: and scanning the first scanning area at the at least one scanning distance based on the completion condition until the number of the MV candidates obtained from the first scanning area reaches a predetermined maximum value.

31. The method of claim 30.

33. the at least one scanning area includes the first scanning area and a second scanning area, the first scanning area is scanned using a plurality of first scanning distances, and the second scanning area is scanned using a plurality of second scanning distances; The completion condition is determining that the number of MV candidates obtained by scanning the first scanning area reaches a second predetermined value; or determining that the number of MV candidates obtained by scanning the second scanning area reaches a predetermined third value, wherein the predetermined second value is the same as or different from the predetermined third value; The method comprises: and scanning the first scanning area at the plurality of first scanning distances and scanning the second scanning area at the plurality of second scanning distances based on the completion condition until the number of the MV candidates obtained by scanning the first scanning area and the second area reaches a predetermined maximum value.

31. The method of claim 30.

34. 1. A method for video encoding, comprising: determining one or more first parameters based on one or more first neighboring blocks of the current block; determining one or more second parameters based on the one or more first neighboring blocks and / or one or more second neighboring blocks of the current block; constructing one or more affine models using the one or more first parameters and the one or more second parameters; determining one or more control point motion vectors (CPMVs) for the current block based on the one or more affine models; the one or more first neighboring blocks and the one or more second neighboring blocks are determined from a plurality of neighboring blocks to the current block based on at least one scanning area and at least one scanning distance, one of the at least one scanning distance indicating a number of blocks away from one side of the current block; the one or more first neighboring blocks and the one or more second neighboring blocks are determined by exclusively scanning the at least one scanning area at the at least one scanning distance; method.

35. the at least one scanning area includes a first scanning area, a second scanning area, and a third scanning area; the first scanning area is determined according to a first maximum scanning distance indicating a maximum number of blocks away from the left side of the current block; the second scanning area is determined according to a second maximum scanning distance indicating a maximum number of blocks away from a top side of the current block, and the first maximum scanning distance is the same as or different from the second maximum scanning distance; the third scanning area is located to the lower right of the current block and includes areas adjacent to and not adjacent to the current block; 35. The method of claim 34.

36. determining one or more co-located temporal neighboring blocks from the third scanning area; determining, in response to determining that the one or more co-located temporal neighboring blocks are used to generate one or more affine constructed neighboring blocks, not to scan the third scanning area to obtain the first neighboring block or the second neighboring block; 36. The method of claim 35, further comprising:

37. determining one or more co-located temporal neighboring blocks from the third scanning area; determining, in response to determining that the one or more co-located temporal neighboring blocks are used to generate one or more affine constructed neighboring blocks, to scan the third scanning area to obtain the first neighboring block or the second neighboring block; determining not to scan the third scanning area to obtain the first neighboring block or the second neighboring block in response to determining that the one or more co-located temporal neighboring blocks will not be used to generate one or more affine constructed neighboring blocks; 36. The method of claim 35, further comprising:

38. the one or more first neighboring blocks include a predetermined first number of first neighboring blocks; the one or more second neighboring blocks include a predetermined second number of second neighboring blocks, the predetermined first number being the same as or different from a predetermined second number; The method comprises: determining that the predetermined first number of first neighboring blocks are valid in response to determining that the predetermined first number of first neighboring blocks use the same reference picture for at least one motion direction; determining that the predetermined second number of second neighboring blocks are valid in response to determining that the predetermined second number of second neighboring blocks use the same reference picture for at least one motion direction; and 35. The method of claim 34, further comprising:

39. determining the one or more first neighboring blocks from a plurality of adjacent neighboring blocks and a plurality of non-adjacent neighboring blocks, the plurality of adjacent neighboring blocks being adjacent to the current block and the plurality of non-adjacent neighboring blocks being each located a number of blocks away from one side of the current block; determining the one or more second neighboring blocks from the plurality of close neighboring blocks and the plurality of non-close neighboring blocks; 35. The method of claim 34, further comprising:

40. determining that neighboring block motion information from the plurality of proximal neighboring blocks and the plurality of non-proximal neighboring blocks is not used to derive an affine merge or advanced motion vector prediction (AMVP) candidate, the neighboring block being one of the one or more second neighboring blocks; determining that a co-located temporal motion vector prediction (TMVP) candidate within the third scanning area is available or that the co-located TMVP candidate is used to derive an affine merge or AMVP candidate, to be one of the one or more second neighboring blocks; 40. The method of claim 39, further comprising:

41. applying a coordinate offset to the one or more second neighboring blocks to obtain an adjusted position for the one or more second neighboring blocks; determining the one or more second parameters based on the adjusted positions of the one or more second neighboring blocks; 35. The method of claim 34, further comprising:

42. 1. A method for video encoding, comprising: determining one or more motion vector (MV) candidates from one or more candidate lists according to a predetermined order, the one or more candidate lists including an affine advanced motion vector prediction (AMVP) candidate list, a regular merge candidate list, and an affine merge candidate list, the one or more MV candidates being from a plurality of neighboring blocks to a current block; determining one or more control point motion vectors (CPMVs) for the current block based on the one or more candidate MVs; A method comprising:

43. The one or more MV candidates include one or more MV candidates inherited from nearby spatial neighboring blocks and one or more MV candidates constructed from nearby spatial neighboring blocks, one or more translational MV candidates from nearby spatial neighboring blocks and one or more temporal MV candidates from nearby temporal neighboring blocks, and one or more MV candidates inherited from non-neighboring spatial neighboring blocks; The method comprises: inserting into a candidate list the one or more MV candidates constructed from the nearby spatial neighboring blocks after the one or more MV candidates inherited from the nearby spatial neighboring blocks; inserting the one or more translational MV candidates from the nearby spatial neighboring blocks into the candidate list after the one or more MV candidates constructed from the nearby spatial neighboring blocks; inserting the one or more temporal MV candidates from the nearby temporal neighboring blocks after the one or more translational MV candidates from the nearby spatial neighboring blocks into the candidate list; inserting into the candidate list the one or more MV candidates inherited from the non-neighboring spatial neighboring blocks after the one or more temporal MV candidates from the close temporal neighboring blocks; responsive to determining that the candidate list is not full, inserting into the candidate list one or more zero MV candidates after the one or more MV candidates inherited from the non-adjacent spatial neighboring blocks; 43. The method of claim 42, further comprising:

44. the one or more MV candidates include sub-block-based temporal motion vector prediction (SbTMVP) candidates, one or more MV candidates inherited from close neighboring blocks and one or more MV candidates inherited from non-close neighboring blocks, one or more MV candidates constructed from close neighboring blocks and one or more MV candidates constructed from non-close neighboring blocks, and one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks; The method comprises: inserting into a candidate list the one or more MV candidates inherited from a nearby neighbor block after the SbTMVP candidate; inserting the one or more MV candidates inherited from non-adjacent neighboring blocks into the candidate list after the one or more MV candidates inherited from adjacent neighboring blocks; inserting the one or more MV candidates constructed from adjacent neighboring blocks into the candidate list after the one or more MV candidates inherited from non-adjacent neighboring blocks; inserting the one or more MV candidates constructed from non-adjacent neighboring blocks into the candidate list after the one or more MV candidates constructed from adjacent neighboring blocks; inserting the one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks into the candidate list after the one or more MV candidates constructed from non-proximal neighboring blocks; responsive to determining that the candidate list is not full, inserting one or more zero MV candidates into the candidate list after the one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks; 43. The method of claim 42, further comprising:

45. the one or more MV candidates include a sub-block based temporal motion vector prediction (SbTMVP) candidate, one or more MV candidates inherited from nearby neighboring blocks and one or more MV candidates inherited from non-close neighboring blocks, a first MV candidate constructed from nearby neighboring blocks and one or more MV candidates constructed from non-close neighboring blocks, and one or more MV candidates constructed from inherited translational neighboring blocks and non-translational neighboring blocks, including a first number of MV candidates and a second number of MV candidates; inserting the one or more MV candidates inherited from nearby neighboring blocks after the SbTMVP candidate into a candidate list; inserting the one or more MV candidates inherited from non-adjacent neighboring blocks into the candidate list after the one or more MV candidates inherited from adjacent neighboring blocks; inserting the first MV candidate constructed from adjacent neighboring blocks into the candidate list after the one or more MV candidates inherited from non-adjacent neighboring blocks; inserting into the candidate list the first number of MV candidates constructed from inherited translational and non-translational neighboring blocks after the first MV candidate constructed from adjacent neighboring blocks; inserting into the candidate list the one or more MV candidates constructed from non-adjacent neighboring blocks after the first number of MV candidates constructed from inherited translational neighboring blocks and non-translatal neighboring blocks; inserting the second number of MV candidates into the candidate list after the one or more MV candidates constructed from non-adjacent neighboring blocks; responsive to determining that the candidate list is not full, inserting one or more zero-movement candidates into the candidate list after the second number of movement candidates; 43. The method of claim 42, further comprising:

46. The first number and the second number are determined in the following manner: predefining the first number as a first fixed value; predefining the second number as a second fixed value, the first fixed value being the same as or different from the second fixed value; signaling the first number or the second number to a decoder at any level; constructing the first number or the second number by the decoder or encoder; or dynamically determining the first number and the second number according to the number of available neighboring blocks in a left area and a top area of ​​the current block; 46. ​​The method of claim 45, wherein the determination is made in at least one of:

47. 1. A method for video encoding, comprising: determining a temporal candidate list having a first list size, the first list size being larger than a list size of any existing candidate list including an affine merge candidate list, an advanced motion vector prediction (AMVP) candidate list, or a regular merge candidate list, the temporal candidate list including a plurality of motion vector (MV) candidates obtained from a plurality of neighboring blocks to a current block; determining a first number of MV candidates from the temporal candidate list based on the reordered MV candidates, wherein the first number is less than the number of MV candidates in the temporal candidate list; A method comprising:

48. 48. The method of claim 47, wherein the temporal candidates are created for a corresponding existing candidate list, the corresponding existing candidate list comprising the affine merge candidate list, the AMVP candidate list, or the regular merge candidate list.

49. 48. The method of claim 47, further comprising applying adaptive reordering to the plurality of MV candidates in the temporal candidate list and / or one or more MV candidates in one or more existing candidate lists.

50. responsive to determining that the temporal candidate list is to be created for an existing candidate list, applying adaptive reordering to the MV candidates in the temporal candidate list and / or the one or more MV candidates in the corresponding existing candidate list; In response to determining that a temporal candidate list is not created for an existing candidate list, applying adaptive reordering to the one or more MV candidates in the corresponding existing candidate list; 50. The method of claim 49, further comprising:

51. The first number of MV candidates are obtained from the plurality of MV candidates adaptively reordered based on a template matching (TM) cost between a predicted sample of a template of the current block and a corresponding reference sample; The method comprises: using an interpolation filter having fewer taps than a predetermined number of taps to generate the predicted samples in response to determining that fractional motion information is to be used for the current block; or responsive to determining that the fractional motion information is used for the current block, skipping interpolation filtering and generating the prediction samples based on nearest integer samples; 50. The method of claim 49, further comprising:

52. 52. The method of claim 51, further comprising applying interpolation filtering to the corresponding reference sample in response to determining that the corresponding reference sample is located at a fractional position.

53. 52. The method of claim 51, further comprising adjusting the TM cost by removing one or more bits from the TM cost.

54. in response to determining that the temporal candidate list includes more than one zero-MV candidate, assigning predetermined fixed values ​​to the zero-MV candidates in the temporal candidate list, except for the first zero-MV candidate, so as to place the zero-MV candidates, except for the first zero-MV candidate, at the end of the temporal candidate list and / or a corresponding existing candidate list, the predetermined fixed value being greater than a threshold; or in response to determining that the existing candidate list includes more than one zero-MV candidate, assigning the predetermined fixed value to zero-MV candidates in the existing candidate list except for the first zero-MV candidate, such that the zero candidates except for the first zero-MV candidate are placed at the end of the existing candidate list; 48. The method of claim 47, further comprising:

55. in response to determining that the temporal candidate list includes at least one zero-MV candidate, assigning a predetermined fixed value to the at least one zero-MV candidate so as to place all of the at least one zero-MV candidate at the end of the temporal candidate and / or corresponding existing candidate list, the predetermined fixed value being greater than a threshold; or in response to determining that the existing candidate list includes at least one zero-MV candidate, assigning the predetermined fixed value to the at least one zero-MV candidate so as to place all of the at least one zero-MV candidate at the end of the existing candidate list; 48. The method of claim 47, further comprising:

56. 50. The method of claim 49, further comprising, in response to determining that one MV candidate or one type of MV candidate will not participate in the adaptive reordering, skipping applying the adaptive reordering to remaining MV candidates in the temporal candidate list and / or corresponding existing candidate list.

57. 50. The method of claim 49, further comprising: in response to determining that one MV candidate or one type of MV candidate will not participate in the adaptive reordering, skipping deriving the MV candidate or the type of MV candidate and / or skipping applying the adaptive reordering to the MV candidate or the type of MV candidate.

58. skipping deriving the MV candidate or the type of MV candidate such that the MV candidate or the candidate content of the type of MV candidate is invalid and one or more positions of the MV candidate or the type of MV candidate remain in the temporal candidate list; skipping calculating the TM cost for the MV candidate or the type of MV candidate; 58. The method of claim 57, further comprising:

59. 1. An apparatus for video decoding, comprising: one or more processors; a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors; The one or more processors, when the instructions are executed, are configured to perform the method of any one of claims 1 to 29. Device.

60. 1. An apparatus for video encoding, comprising: one or more processors; a memory coupled to the one or more processors and configured to store instructions executable by the one or more processors; The one or more processors, when the instructions are executed, are configured to perform the method of any one of claims 30 to 58. Device.

61. 30. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to receive a bitstream and perform the method of any one of claims 1 to 29 based on the bitstream.

62. 60. A non-transitory computer-readable storage medium for storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform the method of any one of claims 30 to 58 to encode a current video block into a bitstream and transmit the bitstream.